{
  "id": 93966,
  "title": "HOW TO IMPROVE THE PROBLEM SETUP?",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/93966",
  "author_name": "",
  "post_date": "2019-05-31T13:31:48.246039400Z",
  "votes": 7,
  "comment_count": 21,
  "views": 0,
  "content": "<p>I live in Chile, which is known for being the most seismic country in the world, and where the largest earthquake recorded by humanity ocurred (Valdivia, 1960). I have experienced more than once how devastating an earthquake can be for a society, so I really wish this competition contributes to improve human capacity to predict earthquakes. It is a very noble purpose and a great opportunity we have to contribute to humanity.</p>\n\n<p>However, I think this competition will not help that much in that sense because of the way the problem is configured. If this were true, maybe we are not gonna be able to contribute too much to the real purpose of this competition. But, if our community can provide suggestions about how to improve the problem configuration, that would be an important contribution.</p>\n\n<p>Why do I think that the problem configuration is not optimal? Basically because the chunks are too small and there is too few data (just 16 experiments) to train, so the optimal algorithms for this kind of problem have not the necessary input to perform their best. The first time I read about this problem, I thought that RNNs and CNNs were gonna be by far the most powerfull models, given that it is a problem based on unstructured data. But then I realized that they don't perform very well. Actually, according to what has been shared, they seem to work similar or even worse than tree-ensemble methods based on engineered features. Why? I think it is because the chunk is too small, it is just an instant, so an RNN cannot really see how the signal is comming, it cannot see what has happend before. The chunks are so small that there is too much randomness in the trends we can capture in them. Am I right?</p>\n\n<p>In my opinion,  the chunk size should be increased to have at least 3-4 seconds of information. Too much data? Well.. it doesn't need to be so densely sampled, does it? It could be sampled at a larger regular spacing to reduce the amount of data, maybe considering the mean of some size intervals.</p>\n\n<p>In summary, I think that given the problem configuration, even the best models are not gonna be the best possible solution, and will not contribute too much to the real problem.</p>\n\n<p>What do you think? Do you think the problem setup can be improved? How?</p>",
  "messages": [
    {
      "id": "540418",
      "postDate": "05/31/2019 13:31:48",
      "content": "<p>I live in Chile, which is known for being the most seismic country in the world, and where the largest earthquake recorded by humanity ocurred (Valdivia, 1960). I have experienced more than once how devastating an earthquake can be for a society, so I really wish this competition contributes to improve human capacity to predict earthquakes. It is a very noble purpose and a great opportunity we have to contribute to humanity.</p>\n\n<p>However, I think this competition will not help that much in that sense because of the way the problem is configured. If this were true, maybe we are not gonna be able to contribute too much to the real purpose of this competition. But, if our community can provide suggestions about how to improve the problem configuration, that would be an important contribution.</p>\n\n<p>Why do I think that the problem configuration is not optimal? Basically because the chunks are too small and there is too few data (just 16 experiments) to train, so the optimal algorithms for this kind of problem have not the necessary input to perform their best. The first time I read about this problem, I thought that RNNs and CNNs were gonna be by far the most powerfull models, given that it is a problem based on unstructured data. But then I realized that they don't perform very well. Actually, according to what has been shared, they seem to work similar or even worse than tree-ensemble methods based on engineered features. Why? I think it is because the chunk is too small, it is just an instant, so an RNN cannot really see how the signal is comming, it cannot see what has happend before. The chunks are so small that there is too much randomness in the trends we can capture in them. Am I right?</p>\n\n<p>In my opinion,  the chunk size should be increased to have at least 3-4 seconds of information. Too much data? Well.. it doesn't need to be so densely sampled, does it? It could be sampled at a larger regular spacing to reduce the amount of data, maybe considering the mean of some size intervals.</p>\n\n<p>In summary, I think that given the problem configuration, even the best models are not gonna be the best possible solution, and will not contribute too much to the real problem.</p>\n\n<p>What do you think? Do you think the problem setup can be improved? How?</p>",
      "rawMarkdown": "I live in Chile, which is known for being the most seismic country in the world, and where the largest earthquake recorded by humanity ocurred (Valdivia, 1960). I have experienced more than once how devastating an earthquake can be for a society, so I really wish this competition contributes to improve human capacity to predict earthquakes. It is a very noble purpose and a great opportunity we have to contribute to humanity.\n\nHowever, I think this competition will not help that much in that sense because of the way the problem is configured. If this were true, maybe we are not gonna be able to contribute too much to the real purpose of this competition. But, if our community can provide suggestions about how to improve the problem configuration, that would be an important contribution.\n\nWhy do I think that the problem configuration is not optimal? Basically because the chunks are too small and there is too few data (just 16 experiments) to train, so the optimal algorithms for this kind of problem have not the necessary input to perform their best. The first time I read about this problem, I thought that RNNs and CNNs were gonna be by far the most powerfull models, given that it is a problem based on unstructured data. But then I realized that they don't perform very well. Actually, according to what has been shared, they seem to work similar or even worse than tree-ensemble methods based on engineered features. Why? I think it is because the chunk is too small, it is just an instant, so an RNN cannot really see how the signal is comming, it cannot see what has happend before. The chunks are so small that there is too much randomness in the trends we can capture in them. Am I right?\n\nIn my opinion,  the chunk size should be increased to have at least 3-4 seconds of information. Too much data? Well.. it doesn't need to be so densely sampled, does it? It could be sampled at a larger regular spacing to reduce the amount of data, maybe considering the mean of some size intervals.\n\nIn summary, I think that given the problem configuration, even the best models are not gonna be the best possible solution, and will not contribute too much to the real problem.\n\nWhat do you think? Do you think the problem setup can be improved? How?",
      "votes": null
    },
    {
      "id": "540426",
      "postDate": "05/31/2019 13:42:15",
      "content": "<p>Organizers have done what you suggest, i.e. using chunks of several seconds, but they say they want to see how far we can go with instant information.  Larger chunks lead to way better precision of course.</p>",
      "rawMarkdown": "Organizers have done what you suggest, i.e. using chunks of several seconds, but they say they want to see how far we can go with instant information.  Larger chunks lead to way better precision of course.",
      "votes": null
    },
    {
      "id": "540451",
      "postDate": "05/31/2019 14:36:59",
      "content": "<p>On a completely separate note, I am so confused how this data can be useful. Who cares if you can predict an earthquake will happen in 10 seconds? The important thing should be if you can predict an earthquake in 1 or 2 days. If you have a prediction in 10 seconds, no one can evacuate quickly enough.</p>",
      "rawMarkdown": "On a completely separate note, I am so confused how this data can be useful. Who cares if you can predict an earthquake will happen in 10 seconds? The important thing should be if you can predict an earthquake in 1 or 2 days. If you have a prediction in 10 seconds, no one can evacuate quickly enough.",
      "votes": null
    },
    {
      "id": "540452",
      "postDate": "05/31/2019 14:40:32",
      "content": "<p>10 seconds here probably corresponds at least to 1k seconds in real world.  Remember that acoustic data is samples at 4Mhz while it is sampled at 44kHz in real world.</p>\n\n<p>Edit: another way to look at it is to look at EQ cycles.  300 years EQ cycle length is not uncommon for earthquakes.  This means that 10 seconds in this experiement correspond to 300 years in reality.</p>",
      "rawMarkdown": "10 seconds here probably corresponds at least to 1k seconds in real world.  Remember that acoustic data is samples at 4Mhz while it is sampled at 44kHz in real world.\n\n\nEdit: another way to look at it is to look at EQ cycles.  300 years EQ cycle length is not uncommon for earthquakes.  This means that 10 seconds in this experiement correspond to 300 years in reality.",
      "votes": null
    },
    {
      "id": "540458",
      "postDate": "05/31/2019 14:48:39",
      "content": "<p>Even if you can predict 1 minute before the earthquake (which I heard is currently achieved in Japan) can be very valuable. You could automate many critical processes to stop inmediately</p>",
      "rawMarkdown": "Even if you can predict 1 minute before the earthquake (which I heard is currently achieved in Japan) can be very valuable. You could automate many critical processes to stop inmediately",
      "votes": null
    },
    {
      "id": "540465",
      "postDate": "05/31/2019 14:52:59",
      "content": "<p>1k seconds in the real world, so we have ~16 minutes to get to cover. That makes more sense; you can send a national emergency through people's phones</p>",
      "rawMarkdown": "1k seconds in the real world, so we have ~16 minutes to get to cover. That makes more sense; you can send a national emergency through people's phones",
      "votes": null
    },
    {
      "id": "540466",
      "postDate": "05/31/2019 14:53:11",
      "content": "<p>I don't see the reason to do that..  Why to put a limit if you have the data and you can do better?</p>",
      "rawMarkdown": "I don't see the reason to do that..  Why to put a limit if you have the data and you can do better?",
      "votes": null
    },
    {
      "id": "540598",
      "postDate": "05/31/2019 17:48:12",
      "content": "<p><a href=\"/cpmpml\">@cpmpml</a> What you said is very weird. sampling is sampling, 4Mhz for a second is just a second. You cannot convert it like that. \nI am totally out of domain but I think it's just laboratory earthquake, not an earthquake we know from the movies or some from the experience. A different thing that can matter in f.e. nuclear energy production.\nI may be wrong, but definitely, you cannot convert 4Mhz to 44kHz and say it's just a different point of view - it has no physical meaning.</p>",
      "rawMarkdown": "cpmpml What you said is very weird. sampling is sampling, 4Mhz for a second is just a second. You cannot convert it like that. \nI am totally out of domain but I think it's just laboratory earthquake, not an earthquake we know from the movies or some from the experience. A different thing that can matter in f.e. nuclear energy production.\nI may be wrong, but definitely, you cannot convert 4Mhz to 44kHz and say it's just a different point of view - it has no physical meaning.",
      "votes": null
    },
    {
      "id": "540600",
      "postDate": "05/31/2019 17:49:41",
      "content": "<p>btw I had similar thoughts and I think I was wrong:\n<a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77708#latest-456749\">https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77708#latest-456749</a></p>",
      "rawMarkdown": "btw I had similar thoughts and I think I was wrong:\nhttps://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77708#latest-456749",
      "votes": null
    },
    {
      "id": "540604",
      "postDate": "05/31/2019 18:00:09",
      "content": "<p>For example windows larger than mean TTF could bias the model, it could see some periodic patterns in a model that could be very misleading and without much bigger dataset it would be hard to distinguish good and useless models. Comparing with that, I prefer the current size of the window. </p>",
      "rawMarkdown": "For example windows larger than mean TTF could bias the model, it could see some periodic patterns in a model that could be very misleading and without much bigger dataset it would be hard to distinguish good and useless models. Comparing with that, I prefer the current size of the window.",
      "votes": null
    },
    {
      "id": "540605",
      "postDate": "05/31/2019 18:04:29",
      "content": "<p>Ok. But here you can have a lot of data. You can generate it. So you make sure to have the amount of data and regularize correctly.</p>",
      "rawMarkdown": "Ok. But here you can have a lot of data. You can generate it. So you make sure to have the amount of data and regularize correctly.",
      "votes": null
    },
    {
      "id": "540616",
      "postDate": "05/31/2019 18:33:59",
      "content": "<blockquote>\n  <p>What you said is very weird.</p>\n</blockquote>\n\n<p>Your reaction as well.  </p>\n\n<p>I suggest you read the fourth paper cited in the introduction post.  You will see that they are looking at signal frequencies between 8 and 13 Hz, while here we deal with frequencies way way higher.  The way they process the signal is very similar to what they did with this lab experiment, except for the sampling rate.  The difference in sampling rate reflects the difference in frequency ranges.  That was my point, sorry I was not clear enough.</p>",
      "rawMarkdown": "&gt; What you said is very weird.\n\nYour reaction as well.  \n\nI suggest you read the fourth paper cited in the introduction post.  You will see that they are looking at signal frequencies between 8 and 13 Hz, while here we deal with frequencies way way higher.  The way they process the signal is very similar to what they did with this lab experiment, except for the sampling rate.  The difference in sampling rate reflects the difference in frequency ranges.  That was my point, sorry I was not clear enough.",
      "votes": null
    },
    {
      "id": "540618",
      "postDate": "05/31/2019 18:35:59",
      "content": "<blockquote>\n  <p>I had similar thoughts</p>\n</blockquote>\n\n<p>Not similar thoughts at all.  </p>",
      "rawMarkdown": "&gt; I had similar thoughts\n\nNot similar thoughts at all.",
      "votes": null
    },
    {
      "id": "540621",
      "postDate": "05/31/2019 18:42:05",
      "content": "<p>&gt; if you have the data and you can do better?</p>\n\n<p>if 10 seconds in the lab correspond to 300 years, do you really think they have the equivalent of 3 seconds of data at every location in the world where earthquakes can occur?</p>",
      "rawMarkdown": "&gt; if you have the data and you can do better?\n\nif 10 seconds in the lab correspond to 300 years, do you really think they have the equivalent of 3 seconds of data at every location in the world where earthquakes can occur?",
      "votes": null
    },
    {
      "id": "540627",
      "postDate": "05/31/2019 19:01:25",
      "content": "<p>&gt; &gt; I had similar thoughts</p>\n\n<p>&gt; Not similar thoughts at all.</p>\n\n<p>about that what I meant was that I was trying to interpret the data somehow different than suggested or said in the information from the organizer. If you are not doing this, it means I totally don't get you.</p>\n\n<p>I've read the paper. I see a lot of ambiguities in the papers and the competition. I don't trust it's well done, don't know if anyone does?</p>\n\n<p>Please explain:</p>\n\n<p>&gt; 10 seconds here probably corresponds at least to 1k seconds in real world.</p>\n\n<p>Both the sampling rate and frequencies correspond to the seconds 'we know'. What does it mean that some seconds correspond to some different seconds in some different 'real' world? </p>\n\n<p>I understand that we have some meaningless digital signal. Only the information about the sampling rate can give the signal a meaning. If it's 4MHz, it means frequencies are between 0 and 2MHz. Calculating FFT will give us 75000 real values. each value represents 2MHz/75000=frequencies so that first value represents 0-26.66 Hz and so on. And it's with losing all the temporal information, STFT will give a much worse frequency resolution.</p>\n\n<p>I don't even know how to measure precisely the singal with such high frequencies and I still don't understand how can it propagate through the ground on far distance.</p>",
      "rawMarkdown": "&gt; &gt; I had similar thoughts\n\n&gt; Not similar thoughts at all.\n\nabout that what I meant was that I was trying to interpret the data somehow different than suggested or said in the information from the organizer. If you are not doing this, it means I totally don't get you.\n\nI've read the paper. I see a lot of ambiguities in the papers and the competition. I don't trust it's well done, don't know if anyone does?\n\n Please explain:\n\n&gt; 10 seconds here probably corresponds at least to 1k seconds in real world.\n\nBoth the sampling rate and frequencies correspond to the seconds 'we know'. What does it mean that some seconds correspond to some different seconds in some different 'real' world? \n\nI understand that we have some meaningless digital signal. Only the information about the sampling rate can give the signal a meaning. If it's 4MHz, it means frequencies are between 0 and 2MHz. Calculating FFT will give us 75000 real values. each value represents 2MHz/75000=frequencies so that first value represents 0-26.66 Hz and so on. And it's with losing all the temporal information, STFT will give a much worse frequency resolution.\n\nI don't even know how to measure precisely the singal with such high frequencies and I still don't understand how can it propagate through the ground on far distance.",
      "votes": null
    },
    {
      "id": "540689",
      "postDate": "05/31/2019 21:19:34",
      "content": "<p>Good point. Still the chunk could have been double or 3 times the size it is, so you need for instance 3 to 5 years of data. You dont require the data for every part of the world, since the earthquakes are concentrated in specific places of the planet.</p>",
      "rawMarkdown": "Good point. Still the chunk could have been double or 3 times the size it is, so you need for instance 3 to 5 years of data. You dont require the data for every part of the world, since the earthquakes are concentrated in specific places of the planet.",
      "votes": null
    },
    {
      "id": "540720",
      "postDate": "06/01/2019 00:06:58",
      "content": "<p>Totally agree. Most of models in current competition fail to predict ttf after single intermediate shear stress relaxation. I would expect hundreds of such intermediate events before the real earthquake, making results obtained in the laboratory experiment kind of useless. But I'm not an any kind of specialist in a field, so can't tell for sure. </p>",
      "rawMarkdown": "Totally agree. Most of models in current competition fail to predict ttf after single intermediate shear stress relaxation. I would expect hundreds of such intermediate events before the real earthquake, making results obtained in the laboratory experiment kind of useless. But I'm not an any kind of specialist in a field, so can't tell for sure.",
      "votes": null
    },
    {
      "id": "540781",
      "postDate": "06/01/2019 04:10:58",
      "content": "<p>Agree with you, actually the first model come to my mind is LSTM, however it works so bad. I hope this competition can benefit to our real world, in saving life. And I believe everyone who participate in this competition have the same wish.</p>",
      "rawMarkdown": "Agree with you, actually the first model come to my mind is LSTM, however it works so bad. I hope this competition can benefit to our real world, in saving life. And I believe everyone who participate in this competition have the same wish.",
      "votes": null
    },
    {
      "id": "540927",
      "postDate": "06/01/2019 11:31:42",
      "content": "<p>Organizers probably have a reason to do the competition this way. We can just ask them.</p>",
      "rawMarkdown": "Organizers probably have a reason to do the competition this way. We can just ask them.",
      "votes": null
    },
    {
      "id": "540947",
      "postDate": "06/01/2019 12:16:04",
      "content": "<p>If 10 seconds correspond to 300 years, then your 3 seconds correspond to roughly one century.  Not 3 to 5 years.</p>",
      "rawMarkdown": "If 10 seconds correspond to 300 years, then your 3 seconds correspond to roughly one century.  Not 3 to 5 years.",
      "votes": null
    },
    {
      "id": "541018",
      "postDate": "06/01/2019 15:01:37",
      "content": "<p><a href=\"/cpmpml\">@cpmpml</a> Yes. That's why I agreed with your point and said that maybe chunks could be at least 2 or 3 times the size they currently are instead of 3 sec. If 10 sec correspond to 300 years, then 0.0375 (current chunk size) correspond to 1.125 years. So the current chunk size multiplied by 3 (0.1125 sec) correspond to 3.375 years</p>",
      "rawMarkdown": "cpmpml Yes. That's why I agreed with your point and said that maybe chunks could be at least 2 or 3 times the size they currently are instead of 3 sec. If 10 sec correspond to 300 years, then 0.0375 (current chunk size) correspond to 1.125 years. So the current chunk size multiplied by 3 (0.1125 sec) correspond to 3.375 years",
      "votes": null
    },
    {
      "id": "541020",
      "postDate": "06/01/2019 15:02:50",
      "content": "<p>You're right!</p>",
      "rawMarkdown": "You're right!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 540426,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "05/31/2019 13:42:15",
      "content": "<p>Organizers have done what you suggest, i.e. using chunks of several seconds, but they say they want to see how far we can go with instant information.  Larger chunks lead to way better precision of course.</p>",
      "votes": null,
      "replies": [
        {
          "id": 540466,
          "author_name": "carlospk",
          "author_url": "",
          "post_date": "05/31/2019 14:53:11",
          "content": "<p>I don't see the reason to do that..  Why to put a limit if you have the data and you can do better?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 540604,
          "author_name": "davids1992",
          "author_url": "",
          "post_date": "05/31/2019 18:00:09",
          "content": "<p>For example windows larger than mean TTF could bias the model, it could see some periodic patterns in a model that could be very misleading and without much bigger dataset it would be hard to distinguish good and useless models. Comparing with that, I prefer the current size of the window. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 540605,
          "author_name": "carlospk",
          "author_url": "",
          "post_date": "05/31/2019 18:04:29",
          "content": "<p>Ok. But here you can have a lot of data. You can generate it. So you make sure to have the amount of data and regularize correctly.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 540621,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/31/2019 18:42:05",
          "content": "<p>&gt; if you have the data and you can do better?</p>\n\n<p>if 10 seconds in the lab correspond to 300 years, do you really think they have the equivalent of 3 seconds of data at every location in the world where earthquakes can occur?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 540689,
          "author_name": "carlospk",
          "author_url": "",
          "post_date": "05/31/2019 21:19:34",
          "content": "<p>Good point. Still the chunk could have been double or 3 times the size it is, so you need for instance 3 to 5 years of data. You dont require the data for every part of the world, since the earthquakes are concentrated in specific places of the planet.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 540947,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/01/2019 12:16:04",
          "content": "<p>If 10 seconds correspond to 300 years, then your 3 seconds correspond to roughly one century.  Not 3 to 5 years.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 541018,
          "author_name": "carlospk",
          "author_url": "",
          "post_date": "06/01/2019 15:01:37",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> Yes. That's why I agreed with your point and said that maybe chunks could be at least 2 or 3 times the size they currently are instead of 3 sec. If 10 sec correspond to 300 years, then 0.0375 (current chunk size) correspond to 1.125 years. So the current chunk size multiplied by 3 (0.1125 sec) correspond to 3.375 years</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 541020,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/01/2019 15:02:50",
          "content": "<p>You're right!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 540451,
      "author_name": "returnofsputnik",
      "author_url": "",
      "post_date": "05/31/2019 14:36:59",
      "content": "<p>On a completely separate note, I am so confused how this data can be useful. Who cares if you can predict an earthquake will happen in 10 seconds? The important thing should be if you can predict an earthquake in 1 or 2 days. If you have a prediction in 10 seconds, no one can evacuate quickly enough.</p>",
      "votes": null,
      "replies": [
        {
          "id": 540452,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/31/2019 14:40:32",
          "content": "<p>10 seconds here probably corresponds at least to 1k seconds in real world.  Remember that acoustic data is samples at 4Mhz while it is sampled at 44kHz in real world.</p>\n\n<p>Edit: another way to look at it is to look at EQ cycles.  300 years EQ cycle length is not uncommon for earthquakes.  This means that 10 seconds in this experiement correspond to 300 years in reality.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 540458,
          "author_name": "carlospk",
          "author_url": "",
          "post_date": "05/31/2019 14:48:39",
          "content": "<p>Even if you can predict 1 minute before the earthquake (which I heard is currently achieved in Japan) can be very valuable. You could automate many critical processes to stop inmediately</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 540465,
          "author_name": "returnofsputnik",
          "author_url": "",
          "post_date": "05/31/2019 14:52:59",
          "content": "<p>1k seconds in the real world, so we have ~16 minutes to get to cover. That makes more sense; you can send a national emergency through people's phones</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 540598,
          "author_name": "davids1992",
          "author_url": "",
          "post_date": "05/31/2019 17:48:12",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> What you said is very weird. sampling is sampling, 4Mhz for a second is just a second. You cannot convert it like that. \nI am totally out of domain but I think it's just laboratory earthquake, not an earthquake we know from the movies or some from the experience. A different thing that can matter in f.e. nuclear energy production.\nI may be wrong, but definitely, you cannot convert 4Mhz to 44kHz and say it's just a different point of view - it has no physical meaning.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 540600,
          "author_name": "davids1992",
          "author_url": "",
          "post_date": "05/31/2019 17:49:41",
          "content": "<p>btw I had similar thoughts and I think I was wrong:\n<a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77708#latest-456749\">https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77708#latest-456749</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 540616,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/31/2019 18:33:59",
          "content": "<blockquote>\n  <p>What you said is very weird.</p>\n</blockquote>\n\n<p>Your reaction as well.  </p>\n\n<p>I suggest you read the fourth paper cited in the introduction post.  You will see that they are looking at signal frequencies between 8 and 13 Hz, while here we deal with frequencies way way higher.  The way they process the signal is very similar to what they did with this lab experiment, except for the sampling rate.  The difference in sampling rate reflects the difference in frequency ranges.  That was my point, sorry I was not clear enough.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 540618,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/31/2019 18:35:59",
          "content": "<blockquote>\n  <p>I had similar thoughts</p>\n</blockquote>\n\n<p>Not similar thoughts at all.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 540627,
          "author_name": "davids1992",
          "author_url": "",
          "post_date": "05/31/2019 19:01:25",
          "content": "<p>&gt; &gt; I had similar thoughts</p>\n\n<p>&gt; Not similar thoughts at all.</p>\n\n<p>about that what I meant was that I was trying to interpret the data somehow different than suggested or said in the information from the organizer. If you are not doing this, it means I totally don't get you.</p>\n\n<p>I've read the paper. I see a lot of ambiguities in the papers and the competition. I don't trust it's well done, don't know if anyone does?</p>\n\n<p>Please explain:</p>\n\n<p>&gt; 10 seconds here probably corresponds at least to 1k seconds in real world.</p>\n\n<p>Both the sampling rate and frequencies correspond to the seconds 'we know'. What does it mean that some seconds correspond to some different seconds in some different 'real' world? </p>\n\n<p>I understand that we have some meaningless digital signal. Only the information about the sampling rate can give the signal a meaning. If it's 4MHz, it means frequencies are between 0 and 2MHz. Calculating FFT will give us 75000 real values. each value represents 2MHz/75000=frequencies so that first value represents 0-26.66 Hz and so on. And it's with losing all the temporal information, STFT will give a much worse frequency resolution.</p>\n\n<p>I don't even know how to measure precisely the singal with such high frequencies and I still don't understand how can it propagate through the ground on far distance.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 540720,
      "author_name": "elvenmonk",
      "author_url": "",
      "post_date": "06/01/2019 00:06:58",
      "content": "<p>Totally agree. Most of models in current competition fail to predict ttf after single intermediate shear stress relaxation. I would expect hundreds of such intermediate events before the real earthquake, making results obtained in the laboratory experiment kind of useless. But I'm not an any kind of specialist in a field, so can't tell for sure. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 540781,
      "author_name": "leonshangguan",
      "author_url": "",
      "post_date": "06/01/2019 04:10:58",
      "content": "<p>Agree with you, actually the first model come to my mind is LSTM, however it works so bad. I hope this competition can benefit to our real world, in saving life. And I believe everyone who participate in this competition have the same wish.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 540927,
      "author_name": "felipefonte99",
      "author_url": "",
      "post_date": "06/01/2019 11:31:42",
      "content": "<p>Organizers probably have a reason to do the competition this way. We can just ask them.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "540418": "I live in Chile, which is known for being the most seismic country in the world, and where the largest earthquake recorded by humanity ocurred (Valdivia, 1960). I have experienced more than once how devastating an earthquake can be for a society, so I really wish this competition contributes to improve human capacity to predict earthquakes. It is a very noble purpose and a great opportunity we have to contribute to humanity.\n\nHowever, I think this competition will not help that much in that sense because of the way the problem is configured. If this were true, maybe we are not gonna be able to contribute too much to the real purpose of this competition. But, if our community can provide suggestions about how to improve the problem configuration, that would be an important contribution.\n\nWhy do I think that the problem configuration is not optimal? Basically because the chunks are too small and there is too few data (just 16 experiments) to train, so the optimal algorithms for this kind of problem have not the necessary input to perform their best. The first time I read about this problem, I thought that RNNs and CNNs were gonna be by far the most powerfull models, given that it is a problem based on unstructured data. But then I realized that they don't perform very well. Actually, according to what has been shared, they seem to work similar or even worse than tree-ensemble methods based on engineered features. Why? I think it is because the chunk is too small, it is just an instant, so an RNN cannot really see how the signal is comming, it cannot see what has happend before. The chunks are so small that there is too much randomness in the trends we can capture in them. Am I right?\n\nIn my opinion,  the chunk size should be increased to have at least 3-4 seconds of information. Too much data? Well.. it doesn't need to be so densely sampled, does it? It could be sampled at a larger regular spacing to reduce the amount of data, maybe considering the mean of some size intervals.\n\nIn summary, I think that given the problem configuration, even the best models are not gonna be the best possible solution, and will not contribute too much to the real problem.\n\nWhat do you think? Do you think the problem setup can be improved? How?",
    "540426": "Organizers have done what you suggest, i.e. using chunks of several seconds, but they say they want to see how far we can go with instant information.  Larger chunks lead to way better precision of course.",
    "540451": "On a completely separate note, I am so confused how this data can be useful. Who cares if you can predict an earthquake will happen in 10 seconds? The important thing should be if you can predict an earthquake in 1 or 2 days. If you have a prediction in 10 seconds, no one can evacuate quickly enough.",
    "540452": "10 seconds here probably corresponds at least to 1k seconds in real world.  Remember that acoustic data is samples at 4Mhz while it is sampled at 44kHz in real world.\n\n\nEdit: another way to look at it is to look at EQ cycles.  300 years EQ cycle length is not uncommon for earthquakes.  This means that 10 seconds in this experiement correspond to 300 years in reality.",
    "540458": "Even if you can predict 1 minute before the earthquake (which I heard is currently achieved in Japan) can be very valuable. You could automate many critical processes to stop inmediately",
    "540465": "1k seconds in the real world, so we have ~16 minutes to get to cover. That makes more sense; you can send a national emergency through people's phones",
    "540466": "I don't see the reason to do that..  Why to put a limit if you have the data and you can do better?",
    "540598": "cpmpml What you said is very weird. sampling is sampling, 4Mhz for a second is just a second. You cannot convert it like that. \nI am totally out of domain but I think it's just laboratory earthquake, not an earthquake we know from the movies or some from the experience. A different thing that can matter in f.e. nuclear energy production.\nI may be wrong, but definitely, you cannot convert 4Mhz to 44kHz and say it's just a different point of view - it has no physical meaning.",
    "540600": "btw I had similar thoughts and I think I was wrong:\nhttps://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77708#latest-456749",
    "540604": "For example windows larger than mean TTF could bias the model, it could see some periodic patterns in a model that could be very misleading and without much bigger dataset it would be hard to distinguish good and useless models. Comparing with that, I prefer the current size of the window.",
    "540605": "Ok. But here you can have a lot of data. You can generate it. So you make sure to have the amount of data and regularize correctly.",
    "540616": "&gt; What you said is very weird.\n\nYour reaction as well.  \n\nI suggest you read the fourth paper cited in the introduction post.  You will see that they are looking at signal frequencies between 8 and 13 Hz, while here we deal with frequencies way way higher.  The way they process the signal is very similar to what they did with this lab experiment, except for the sampling rate.  The difference in sampling rate reflects the difference in frequency ranges.  That was my point, sorry I was not clear enough.",
    "540618": "&gt; I had similar thoughts\n\nNot similar thoughts at all.",
    "540621": "&gt; if you have the data and you can do better?\n\nif 10 seconds in the lab correspond to 300 years, do you really think they have the equivalent of 3 seconds of data at every location in the world where earthquakes can occur?",
    "540627": "&gt; &gt; I had similar thoughts\n\n&gt; Not similar thoughts at all.\n\nabout that what I meant was that I was trying to interpret the data somehow different than suggested or said in the information from the organizer. If you are not doing this, it means I totally don't get you.\n\nI've read the paper. I see a lot of ambiguities in the papers and the competition. I don't trust it's well done, don't know if anyone does?\n\n Please explain:\n\n&gt; 10 seconds here probably corresponds at least to 1k seconds in real world.\n\nBoth the sampling rate and frequencies correspond to the seconds 'we know'. What does it mean that some seconds correspond to some different seconds in some different 'real' world? \n\nI understand that we have some meaningless digital signal. Only the information about the sampling rate can give the signal a meaning. If it's 4MHz, it means frequencies are between 0 and 2MHz. Calculating FFT will give us 75000 real values. each value represents 2MHz/75000=frequencies so that first value represents 0-26.66 Hz and so on. And it's with losing all the temporal information, STFT will give a much worse frequency resolution.\n\nI don't even know how to measure precisely the singal with such high frequencies and I still don't understand how can it propagate through the ground on far distance.",
    "540689": "Good point. Still the chunk could have been double or 3 times the size it is, so you need for instance 3 to 5 years of data. You dont require the data for every part of the world, since the earthquakes are concentrated in specific places of the planet.",
    "540720": "Totally agree. Most of models in current competition fail to predict ttf after single intermediate shear stress relaxation. I would expect hundreds of such intermediate events before the real earthquake, making results obtained in the laboratory experiment kind of useless. But I'm not an any kind of specialist in a field, so can't tell for sure.",
    "540781": "Agree with you, actually the first model come to my mind is LSTM, however it works so bad. I hope this competition can benefit to our real world, in saving life. And I believe everyone who participate in this competition have the same wish.",
    "540927": "Organizers probably have a reason to do the competition this way. We can just ask them.",
    "540947": "If 10 seconds correspond to 300 years, then your 3 seconds correspond to roughly one century.  Not 3 to 5 years.",
    "541018": "cpmpml Yes. That's why I agreed with your point and said that maybe chunks could be at least 2 or 3 times the size they currently are instead of 3 sec. If 10 sec correspond to 300 years, then 0.0375 (current chunk size) correspond to 1.125 years. So the current chunk size multiplied by 3 (0.1125 sec) correspond to 3.375 years",
    "541020": "You're right!"
  },
  "source": "meta"
}