{
  "id": 91125,
  "title": "Why models fail to predict TTF > 11 sec? ",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/91125",
  "author_name": "",
  "post_date": "2019-05-01T08:27:54.983953500Z",
  "votes": 24,
  "comment_count": 74,
  "views": 0,
  "content": "<p>Attached is the result of my LightGBM model (x axis) vs. true TTF values (y axis). My model fails to predict anything above approx. 10.5 ~ 11 seconds. Those data points form 14.35% of the train data and since our metric is MAE, they may be the most contributing points due to their large errors. I also observed this trend in my other models including RNN, CNN and even linear regression models. I saw this in results of public kernels too.</p>\n\n<p>I would like to hear your thought on this. What do you think is the reason? How do you think I can enhance it? </p>",
  "messages": [
    {
      "id": "525533",
      "postDate": "05/01/2019 08:27:54",
      "content": "<p>Attached is the result of my LightGBM model (x axis) vs. true TTF values (y axis). My model fails to predict anything above approx. 10.5 ~ 11 seconds. Those data points form 14.35% of the train data and since our metric is MAE, they may be the most contributing points due to their large errors. I also observed this trend in my other models including RNN, CNN and even linear regression models. I saw this in results of public kernels too.</p>\n\n<p>I would like to hear your thought on this. What do you think is the reason? How do you think I can enhance it? </p>",
      "rawMarkdown": "Attached is the result of my LightGBM model (x axis) vs. true TTF values (y axis). My model fails to predict anything above approx. 10.5 ~ 11 seconds. Those data points form 14.35% of the train data and since our metric is MAE, they may be the most contributing points due to their large errors. I also observed this trend in my other models including RNN, CNN and even linear regression models. I saw this in results of public kernels too.\n\nI would like to hear your thought on this. What do you think is the reason? How do you think I can enhance it?",
      "votes": null
    },
    {
      "id": "525553",
      "postDate": "05/01/2019 09:18:26",
      "content": "<p>Regression models predict expected value of target, and it usually not close to edge values especially when number of observation with this extreme values are small.\nYou may try something like <a href=\"https://www.kaggle.com/c/elo-merchant-category-recommendation/discussion/82036\">winner of recent ELO competition did</a> - combine regression and classification.</p>",
      "rawMarkdown": "Regression models predict expected value of target, and it usually not close to edge values especially when number of observation with this extreme values are small.\nYou may try something like [winner of recent ELO competition did](https://www.kaggle.com/c/elo-merchant-category-recommendation/discussion/82036) - combine regression and classification.",
      "votes": null
    },
    {
      "id": "525582",
      "postDate": "05/01/2019 10:36:46",
      "content": "<p>I believe this problem cannot be solved and is inherently anchored to the physics of the experiment:\nAfter each slip (quake) the crumbled parts solidify again and due to the constant speed of the piston we always expect the same TTF. But in this experiment we also have some minor slips that are kind of resetting back the state of the experiment. That adds a few seconds of TTF and it can not be predicted by a model if the coming quake is major or minor. The predictions between minor and major quakes are in line with the ground truth again. </p>\n\n<p><img src=\"https://i.imgur.com/Htg77qk.png\" alt=\"minor quake\"></p>\n\n<p>If anyone would be able to model this, expect not less then a nature paper from your results ;)</p>\n\n<p>I am very curious about the top solutions in this competition.</p>",
      "rawMarkdown": "I believe this problem cannot be solved and is inherently anchored to the physics of the experiment:\nAfter each slip (quake) the crumbled parts solidify again and due to the constant speed of the piston we always expect the same TTF. But in this experiment we also have some minor slips that are kind of resetting back the state of the experiment. That adds a few seconds of TTF and it can not be predicted by a model if the coming quake is major or minor. The predictions between minor and major quakes are in line with the ground truth again. \n\n![minor quake](https://i.imgur.com/Htg77qk.png)\n\n\nIf anyone would be able to model this, expect not less then a nature paper from your results ;)\n\nI am very curious about the top solutions in this competition.",
      "votes": null
    },
    {
      "id": "525611",
      "postDate": "05/01/2019 11:54:03",
      "content": "<p>Have a look at this</p>\n\n<p>They don't seem to have dips either they have a very good model or they have cherry picked the plot! ;)</p>\n\n<p><a href=\"https://www.researchgate.net/figure/Time-remaining-before-the-next-failure-predicted-by-the-Random-Forest-As-in-Fig-1a_fig3_313858017\">Image</a></p>\n\n<p>Edit:  It looks like they count the \"Minor Quakes\" as Quakes whereas we have to predict only the major quakes.</p>",
      "rawMarkdown": "Have a look at this\n\n\n\nThey don't seem to have dips either they have a very good model or they have cherry picked the plot! ;)\n\n[Image](https://www.researchgate.net/figure/Time-remaining-before-the-next-failure-predicted-by-the-Random-Forest-As-in-Fig-1a_fig3_313858017)\n\nEdit:  It looks like they count the \"Minor Quakes\" as Quakes whereas we have to predict only the major quakes.",
      "votes": null
    },
    {
      "id": "525621",
      "postDate": "05/01/2019 12:13:46",
      "content": "<p>My model doesn't fail this.  But there is a max value lower than 16 indeed.  This is probably because there are way less samples with very high ttf values.  Also, it may be that some ttf are very high because there are miniquakes later that relive the stress enough to postpone the real quake.  Given we only have a small snapshot, it is impossible to predict this situation.</p>\n\n<p>Edit: I see Ilu gave the same explanation.</p>",
      "rawMarkdown": "My model doesn't fail this.  But there is a max value lower than 16 indeed.  This is probably because there are way less samples with very high ttf values.  Also, it may be that some ttf are very high because there are miniquakes later that relive the stress enough to postpone the real quake.  Given we only have a small snapshot, it is impossible to predict this situation.\n\nEdit: I see Ilu gave the same explanation.",
      "votes": null
    },
    {
      "id": "525623",
      "postDate": "05/01/2019 12:18:39",
      "content": "<p>&gt; Edit: It looks like they count the \"Minor Quakes\" as Quakes whereas we have to predict only the major quakes.</p>\n\n<p>this is exactly what makes it trivial to model. Here we need to find out the difference before a minor and a major quake.</p>",
      "rawMarkdown": "&gt; Edit: It looks like they count the \"Minor Quakes\" as Quakes whereas we have to predict only the major quakes.\n\nthis is exactly what makes it trivial to model. Here we need to find out the difference before a minor and a major quake.",
      "votes": null
    },
    {
      "id": "525627",
      "postDate": "05/01/2019 12:23:00",
      "content": "<p>That implies, you have found a way to predict minor quakes?\nIncredible work and that would explain your score. Did everyone up in the top 50 find that method?</p>\n\n<p>Would you mind clarifying your statement? Or are you afraid of another \"Santander\"?</p>",
      "rawMarkdown": "That implies, you have found a way to predict minor quakes?\nIncredible work and that would explain your score. Did everyone up in the top 50 find that method?\n\nWould you mind clarifying your statement? Or are you afraid of another \"Santander\"?",
      "votes": null
    },
    {
      "id": "525630",
      "postDate": "05/01/2019 12:28:51",
      "content": "<p>Differentiation would be a lot easier if they hadn't shuffled test;)</p>",
      "rawMarkdown": "Differentiation would be a lot easier if they hadn't shuffled test;)",
      "votes": null
    },
    {
      "id": "525631",
      "postDate": "05/01/2019 12:29:20",
      "content": "<p>No, my model is able to predict ttf &gt; 11 up to some value.  Not sure why others can't actually.  </p>",
      "rawMarkdown": "No, my model is able to predict ttf &gt; 11 up to some value.  Not sure why others can't actually.",
      "votes": null
    },
    {
      "id": "525661",
      "postDate": "05/01/2019 13:37:18",
      "content": "<p><a href=\"/cpmpml\">@cpmpml</a> I've found that a couple of features make predicting TTF up to about 14 trivial, but only as a result of K-fold leakage. I'll be very impressed if you've found a way to do this without any leakage. Like <a href=\"/ilu000\">@ilu000</a> said, it's contrary to my understanding of the physics at play here.</p>",
      "rawMarkdown": "cpmpml I've found that a couple of features make predicting TTF up to about 14 trivial, but only as a result of K-fold leakage. I'll be very impressed if you've found a way to do this without any leakage. Like @ilu000 said, it's contrary to my understanding of the physics at play here.",
      "votes": null
    },
    {
      "id": "525684",
      "postDate": "05/01/2019 14:31:05",
      "content": "<p>They also use a time window of 1.8 sec (!) </p>\n\n<blockquote>\n  <p>To create a model that uncovers the physics of shear failure, we make predictions using moving time windows applied to the data. Each window is 1.8s, which is small compared to the time between fault gouge failures (8s on average).</p>\n</blockquote>\n\n<p><a href=\"https://agupubs.onlinelibrary.wiley.com/action/downloadSupplement?doi=10.1002%2F2017GL074677&amp;file=grl56367-sup-0001-supinfo.pdf\">grl56367-sup-0001-supinfo.pdf</a></p>",
      "rawMarkdown": "They also use a time window of 1.8 sec (!) \n&gt; To create a model that uncovers the physics of shear failure, we make predictions using moving time windows applied to the data. Each window is 1.8s, which is small compared to the time between fault gouge failures (8s on average).\n\n[grl56367-sup-0001-supinfo.pdf](https://agupubs.onlinelibrary.wiley.com/action/downloadSupplement?doi=10.1002%2F2017GL074677&amp;file=grl56367-sup-0001-supinfo.pdf)",
      "votes": null
    },
    {
      "id": "525702",
      "postDate": "05/01/2019 15:14:09",
      "content": "<p>I don't think I have leakage.</p>",
      "rawMarkdown": "I don't think I have leakage.",
      "votes": null
    },
    {
      "id": "525778",
      "postDate": "05/01/2019 17:34:22",
      "content": "<p>How does this all relate to a chunk size of 150000 as most kernels are using?\n1.8 seconds is much larger than 150000 datapoints.</p>",
      "rawMarkdown": "How does this all relate to a chunk size of 150000 as most kernels are using?\n1.8 seconds is much larger than 150000 datapoints.",
      "votes": null
    },
    {
      "id": "525791",
      "postDate": "05/01/2019 18:08:53",
      "content": "<p>Would you explain more? I didn't get your point.</p>",
      "rawMarkdown": "Would you explain more? I didn't get your point.",
      "votes": null
    },
    {
      "id": "525793",
      "postDate": "05/01/2019 18:14:12",
      "content": "<p><a href=\"/cpmpml\">@cpmpml</a> there are also some datapoints where true ttf is less than about 0.5 but all of my models predict them somehow randomly. look at the bottom of the figure I posted earlier; there are about 1% of points that have a very different trend than the rest of them. Do you happen to have the same problem in your models?</p>",
      "rawMarkdown": "cpmpml there are also some datapoints where true ttf is less than about 0.5 but all of my models predict them somehow randomly. look at the bottom of the figure I posted earlier; there are about 1% of points that have a very different trend than the rest of them. Do you happen to have the same problem in your models?",
      "votes": null
    },
    {
      "id": "525796",
      "postDate": "05/01/2019 18:17:28",
      "content": "<p><a href=\"/scirpus\">@scirpus</a> <a href=\"/steubk\">@steubk</a>  <a href=\"/ilu000\">@ilu000</a>  Thank you for pointing out this. I didn't know Minor quakes are a thing. I will investigate it more and share if I found a way to enhance my models by this.</p>",
      "rawMarkdown": "scirpus @steubk  @ilu000  Thank you for pointing out this. I didn't know Minor quakes are a thing. I will investigate it more and share if I found a way to enhance my models by this.",
      "votes": null
    },
    {
      "id": "525799",
      "postDate": "05/01/2019 18:22:16",
      "content": "<p>Thanks <a href=\"/ilu000\">@ilu000</a>. This graph is very informative. I am curious to see top solutions too. </p>",
      "rawMarkdown": "Thanks @ilu000. This graph is very informative. I am curious to see top solutions too.",
      "votes": null
    },
    {
      "id": "525800",
      "postDate": "05/01/2019 18:25:24",
      "content": "<p>Might be worth using the denoise kernel</p>",
      "rawMarkdown": "Might be worth using the denoise kernel",
      "votes": null
    },
    {
      "id": "525801",
      "postDate": "05/01/2019 18:25:54",
      "content": "<p>Thank you for your input. I think the count of those points (about 14% of total) is not too small for the model to completely ignore their effects. I had a glance at the link you posted; couldn't understand it :D Need to take more time to digest it.</p>",
      "rawMarkdown": "Thank you for your input. I think the count of those points (about 14% of total) is not too small for the model to completely ignore their effects. I had a glance at the link you posted; couldn't understand it :D Need to take more time to digest it.",
      "votes": null
    },
    {
      "id": "525802",
      "postDate": "05/01/2019 18:29:05",
      "content": "<p><a href=\"/mhviraf\">@mhviraf</a> The models often have a hard time distinguishing between the start and end of earthquake periods as the signal profile is very similar.</p>",
      "rawMarkdown": "mhviraf The models often have a hard time distinguishing between the start and end of earthquake periods as the signal profile is very similar.",
      "votes": null
    },
    {
      "id": "525811",
      "postDate": "05/01/2019 18:50:24",
      "content": "<p>Hein is referring to <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91125#525684\">the comment</a> by steubk.</p>",
      "rawMarkdown": "Hein is referring to [the comment](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91125#525684) by steubk.",
      "votes": null
    },
    {
      "id": "525815",
      "postDate": "05/01/2019 19:00:55",
      "content": "<p><a href=\"/scirpus\">@scirpus</a> why do you think denoised data would solve this?</p>",
      "rawMarkdown": "scirpus why do you think denoised data would solve this?",
      "votes": null
    },
    {
      "id": "525817",
      "postDate": "05/01/2019 19:04:20",
      "content": "<p><a href=\"/bigironsphere\">@bigironsphere</a> that's a good point. based on this explanation I expect to see very low (or randomly distributed) predictions for very large TTFs too but this is not the case and models could surprisingly distinguish them. </p>",
      "rawMarkdown": "bigironsphere that's a good point. based on this explanation I expect to see very low (or randomly distributed) predictions for very large TTFs too but this is not the case and models could surprisingly distinguish them.",
      "votes": null
    },
    {
      "id": "525836",
      "postDate": "05/01/2019 19:55:59",
      "content": "<p>I am with <a href=\"/bigironsphere\">@bigironsphere</a> here in terms of the physics in this experiment. Predicting TTF values above 11 sec implies an upcoming minor quake. So your model must be able to somehow predict those minor quakes <a href=\"/cpmpml\">@cpmpml</a> . \nIf this isn't some kind of leakage I am truly impressed and eagerly looking forward to your solution. \nIf you don't mind, you might play with your features a bit and find out which feature(s) is/are helping to identify the minor slips. Of course you can wait until after the competition with revealing that information ;)</p>",
      "rawMarkdown": "I am with @bigironsphere here in terms of the physics in this experiment. Predicting TTF values above 11 sec implies an upcoming minor quake. So your model must be able to somehow predict those minor quakes @cpmpml . \nIf this isn't some kind of leakage I am truly impressed and eagerly looking forward to your solution. \nIf you don't mind, you might play with your features a bit and find out which feature(s) is/are helping to identify the minor slips. Of course you can wait until after the competition with revealing that information ;)",
      "votes": null
    },
    {
      "id": "525837",
      "postDate": "05/01/2019 19:59:56",
      "content": "<p>I agree, I think it has to do with the physics of the problem.  </p>\n\n<p>The data shows 14 complete experiments.  </p>\n\n<p>Here is my assumption -  The initial conditions of the lab earthquake system at the start of each experiment are controlled to be nearly identical by controlling the quantity and type of materials and setup procedure.  With time the system itself is influenced by its own random/chaotic behavior and its internal 'state' starts to move away from initial conditions,  and this reflection of state (hopefully) shows up in the acoustic data.</p>\n\n<p>If this assumption is true the acoustic  samples early in each experimental run should look very similar, even though the experiment's final time_to_failure will be different.   </p>\n\n<p>The 'time_to_failure' data is retrospective -not causal.  At the time point the experiment is started, no one knows the time_to_failure, since it has not yet physically failed in the run. </p>\n\n<p>That is what makes this problem hard. </p>",
      "rawMarkdown": "I agree, I think it has to do with the physics of the problem.  \n\nThe data shows 14 complete experiments.  \n\nHere is my assumption -  The initial conditions of the lab earthquake system at the start of each experiment are controlled to be nearly identical by controlling the quantity and type of materials and setup procedure.  With time the system itself is influenced by its own random/chaotic behavior and its internal 'state' starts to move away from initial conditions,  and this reflection of state (hopefully) shows up in the acoustic data.\n\nIf this assumption is true the acoustic  samples early in each experimental run should look very similar, even though the experiment's final time_to_failure will be different.   \n\nThe 'time_to_failure' data is retrospective -not causal.  At the time point the experiment is started, no one knows the time_to_failure, since it has not yet physically failed in the run. \n\nThat is what makes this problem hard.",
      "votes": null
    },
    {
      "id": "525841",
      "postDate": "05/01/2019 20:06:13",
      "content": "<p>This also explains a really mysterious pattern I've been seeing in some of my cross-validation results.  On some folds, especially those which include longer ttf, I'm seeing this strange pattern of two distinct trendlines with slope of 1.  I've attached an example of my version of OP's results graph for one of my cross validation folds that exhibits this pattern.  </p>\n\n<p>You can clearly see two seperate trendlines, which after reading Ilu's hypothesis I suspect is actually evidence of the algorithm learning to correctly predict the ttf of minor earthquakes. These minor quakes occur within some but not all of our 16 quake samples, which is why this pattern only appears in some cross validation folds.</p>",
      "rawMarkdown": "This also explains a really mysterious pattern I've been seeing in some of my cross-validation results.  On some folds, especially those which include longer ttf, I'm seeing this strange pattern of two distinct trendlines with slope of 1.  I've attached an example of my version of OP's results graph for one of my cross validation folds that exhibits this pattern.  \n\nYou can clearly see two seperate trendlines, which after reading Ilu's hypothesis I suspect is actually evidence of the algorithm learning to correctly predict the ttf of minor earthquakes. These minor quakes occur within some but not all of our 16 quake samples, which is why this pattern only appears in some cross validation folds.",
      "votes": null
    },
    {
      "id": "525842",
      "postDate": "05/01/2019 20:09:44",
      "content": "<p>I wonder if the winners will be part of the nature paper the hosts are going to write after this competition</p>",
      "rawMarkdown": "I wonder if the winners will be part of the nature paper the hosts are going to write after this competition",
      "votes": null
    },
    {
      "id": "525847",
      "postDate": "05/01/2019 20:16:11",
      "content": "<p>totally agree with <a href=\"/filipmulier\">@filipmulier</a>. I wrote a similar explanation <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/89301#515975\">here</a></p>",
      "rawMarkdown": "totally agree with @filipmulier. I wrote a similar explanation [here](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/89301#515975)",
      "votes": null
    },
    {
      "id": "525854",
      "postDate": "05/01/2019 20:44:06",
      "content": "<p><a href=\"/amjad85\">@amjad85</a>  I think I read somewhere that they are willing to include winners of this competition in their next papers. Not sure if it's gonna be a Nature paper though ;)</p>",
      "rawMarkdown": "amjad85  I think I read somewhere that they are willing to include winners of this competition in their next papers. Not sure if it's gonna be a Nature paper though ;)",
      "votes": null
    },
    {
      "id": "525856",
      "postDate": "05/01/2019 20:47:45",
      "content": "<p>Aaron, In your graph I see at least 2 bands, and maybe 4.  Could this be the experiment (quake ID) showing up as contributing to the variance -  See Amjad's post on CV above. </p>",
      "rawMarkdown": "Aaron, In your graph I see at least 2 bands, and maybe 4.  Could this be the experiment (quake ID) showing up as contributing to the variance -  See Amjad's post on CV above.",
      "votes": null
    },
    {
      "id": "525871",
      "postDate": "05/01/2019 22:27:50",
      "content": "<p>I'll share my approach after competition end for sure.</p>",
      "rawMarkdown": "I'll share my approach after competition end for sure.",
      "votes": null
    },
    {
      "id": "525875",
      "postDate": "05/01/2019 22:44:33",
      "content": "<p>Filip, the image I posted was taken from one fold of a straight 5-fold split.  If it is the quake ID that is responsible for that distinctive band of low prediction for high expected ttf, then that band should disappear when I switch to 'leave one quake out' CV.  However it does not.  In fact, the band shows up clearly, but only for quakes with high maximum ttf (&gt;10).  Quakes with low ttf, around 8-9, never show this distinctive banding pattern, though there may be other patterns in the error of individual quakes. </p>",
      "rawMarkdown": "Filip, the image I posted was taken from one fold of a straight 5-fold split.  If it is the quake ID that is responsible for that distinctive band of low prediction for high expected ttf, then that band should disappear when I switch to 'leave one quake out' CV.  However it does not.  In fact, the band shows up clearly, but only for quakes with high maximum ttf (&gt;10).  Quakes with low ttf, around 8-9, never show this distinctive banding pattern, though there may be other patterns in the error of individual quakes.",
      "votes": null
    },
    {
      "id": "525889",
      "postDate": "05/02/2019 00:01:03",
      "content": "<p>Sharing one model with max predictions above 12.  The model is still fooled by the mini quake after the  7th quake.  </p>\n\n<p>As many others predicting ttf below 2 is an issue.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/525889/13136/pred.png\" alt=\"pred\"></p>",
      "rawMarkdown": "Sharing one model with max predictions above 12.  The model is still fooled by the mini quake after the  7th quake.  \n\nAs many others predicting ttf below 2 is an issue.\n\n![pred](https://storage.googleapis.com/kaggle-forum-message-attachments/525889/13136/pred.png)",
      "votes": null
    },
    {
      "id": "525906",
      "postDate": "05/02/2019 01:26:41",
      "content": "<p>I think he means that if you selected a 150000 size chuck of data - the last row of that data set becomes the target value.   </p>\n\n<p>So model is not seeing any of the times for the first 149999 rows and will therefore not predict them.  </p>\n\n<p>You could confirm this by running your model and have the y value the first row of the segment rather than the last - I think in most kernels that's a one line simple change.  </p>\n\n<p>Hmm - not sure I made it any clearer but I did say it in a different way :)</p>",
      "rawMarkdown": "I think he means that if you selected a 150000 size chuck of data - the last row of that data set becomes the target value.   \n\nSo model is not seeing any of the times for the first 149999 rows and will therefore not predict them.  \n\nYou could confirm this by running your model and have the y value the first row of the segment rather than the last - I think in most kernels that's a one line simple change.  \n\nHmm - not sure I made it any clearer but I did say it in a different way :)",
      "votes": null
    },
    {
      "id": "525907",
      "postDate": "05/02/2019 01:39:07",
      "content": "<p>When I read the 4 documents that describe the previous experiments and publication of results I see that the \"data\" is probably grabbed after the process of making quakes has been established and fairly constant.  There is a sketch in one of the pubs that shows this clearly!</p>\n\n<p>The papers indicate to me that the definition of a quake is when the stress Gage goes to zero.  We don't have the stress data - if we do it might help explain the minor quakes.  In the pubs you can see dips in the stress plots for very brief time that probably are the minor quakes we are seeing.  </p>\n\n<p>It bugs me a bit that the quake end time of zero is after the big spikes in the acoustic signal.  I assume therefore that after the big acoustic signal the stress starts dropping and that bit of calm before zero time is related to the time it takes for stress to hit the cutoff point being used.  I am not a real fan to sponsors who leave out what I think is key data from our work - in this case the stress seems very important to me.</p>\n\n<p>Suggest all read the four pubs in the Welcome - they helped me either better understand (or continue to mis-understand the experiment and data collection)</p>",
      "rawMarkdown": "When I read the 4 documents that describe the previous experiments and publication of results I see that the \"data\" is probably grabbed after the process of making quakes has been established and fairly constant.  There is a sketch in one of the pubs that shows this clearly!\n\nThe papers indicate to me that the definition of a quake is when the stress Gage goes to zero.  We don't have the stress data - if we do it might help explain the minor quakes.  In the pubs you can see dips in the stress plots for very brief time that probably are the minor quakes we are seeing.  \n\nIt bugs me a bit that the quake end time of zero is after the big spikes in the acoustic signal.  I assume therefore that after the big acoustic signal the stress starts dropping and that bit of calm before zero time is related to the time it takes for stress to hit the cutoff point being used.  I am not a real fan to sponsors who leave out what I think is key data from our work - in this case the stress seems very important to me.\n\nSuggest all read the four pubs in the Welcome - they helped me either better understand (or continue to mis-understand the experiment and data collection)",
      "votes": null
    },
    {
      "id": "525911",
      "postDate": "05/02/2019 01:47:24",
      "content": "<p>Agree, 1.8 sec makes it a lot easier.  The BIG spike is then in a lot of signals rather than 16/4194.</p>",
      "rawMarkdown": "Agree, 1.8 sec makes it a lot easier.  The BIG spike is then in a lot of signals rather than 16/4194.",
      "votes": null
    },
    {
      "id": "525912",
      "postDate": "05/02/2019 01:51:01",
      "content": "<p><a href=\"/pcjimmmy\">@pcjimmmy</a> The purpose of this competition is to see if our findings can be extrapolated to real world earthquake prediction. We cannot measure stress and strain deep underneath the Earth; we can only measure acoustic seismological waves, and in a very short timespan compared to geological scales. The data constraints are a problem, but they make sense in light of the researchers' aims.  </p>",
      "rawMarkdown": "pcjimmmy The purpose of this competition is to see if our findings can be extrapolated to real world earthquake prediction. We cannot measure stress and strain deep underneath the Earth; we can only measure acoustic seismological waves, and in a very short timespan compared to geological scales. The data constraints are a problem, but they make sense in light of the researchers' aims.",
      "votes": null
    },
    {
      "id": "525925",
      "postDate": "05/02/2019 02:47:57",
      "content": "<p>OK - makes sense that stress needs to be left out - than my only remaining issue is that the zero time for quakes was based on stress rather than big signals.  Probably would have been too hard to use real world data and quakes - the folks who participate in these challenges are pretty smart about finding leaks in the data :)</p>",
      "rawMarkdown": "OK - makes sense that stress needs to be left out - than my only remaining issue is that the zero time for quakes was based on stress rather than big signals.  Probably would have been too hard to use real world data and quakes - the folks who participate in these challenges are pretty smart about finding leaks in the data :)",
      "votes": null
    },
    {
      "id": "526033",
      "postDate": "05/02/2019 08:48:57",
      "content": "<p>Thank you for sharing <a href=\"/cpmpml\">@cpmpml</a> . Is this a LGB model?</p>",
      "rawMarkdown": "Thank you for sharing @cpmpml . Is this a LGB model?",
      "votes": null
    },
    {
      "id": "526034",
      "postDate": "05/02/2019 08:49:42",
      "content": "<p>Lgb indeed.</p>",
      "rawMarkdown": "Lgb indeed.",
      "votes": null
    },
    {
      "id": "526036",
      "postDate": "05/02/2019 08:58:48",
      "content": "<p>Very interesting. So there really seems to be a way to separate the quakes. \nYou claimed that you don't have a leak in your code. Does that include you removed the mean of the data? I saw that the model is able to cluster the quakes if the mean is used. \nIs there a way to cluster the test set (new unseen clusters) and make that information available to a LGB model?</p>\n\n<p>It seems to me that clustering/separating the quakes is the key to win this competition.</p>\n\n<p>Edit: I just saw that you have 2 negative TTF predictions. How is that possible?</p>",
      "rawMarkdown": "Very interesting. So there really seems to be a way to separate the quakes. \nYou claimed that you don't have a leak in your code. Does that include you removed the mean of the data? I saw that the model is able to cluster the quakes if the mean is used. \nIs there a way to cluster the test set (new unseen clusters) and make that information available to a LGB model?\n\nIt seems to me that clustering/separating the quakes is the key to win this competition.\n\nEdit: I just saw that you have 2 negative TTF predictions. How is that possible?",
      "votes": null
    },
    {
      "id": "526038",
      "postDate": "05/02/2019 09:02:38",
      "content": "<p>Mean of data leaks?  Interesting.  I am not using it anyway as I assumed that mean should be 0 for acoustic data and that any deviation is a measurement artifact.</p>",
      "rawMarkdown": "Mean of data leaks?  Interesting.  I am not using it anyway as I assumed that mean should be 0 for acoustic data and that any deviation is a measurement artifact.",
      "votes": null
    },
    {
      "id": "526043",
      "postDate": "05/02/2019 09:07:17",
      "content": "<p>Yes, depending on how you split your folds, the model is able to say which quake corresponds to a data chunk by evaluating the mean. The mean is slowly drifting thoughout the train set. Of course that can only give an advantage if the model is trained with a part of the same quake. \nBut I have no idea how to use that in the test set. Thus, i substracted the mean in all data chunks. </p>",
      "rawMarkdown": "Yes, depending on how you split your folds, the model is able to say which quake corresponds to a data chunk by evaluating the mean. The mean is slowly drifting thoughout the train set. Of course that can only give an advantage if the model is trained with a part of the same quake. \nBut I have no idea how to use that in the test set. Thus, i substracted the mean in all data chunks.",
      "votes": null
    },
    {
      "id": "526044",
      "postDate": "05/02/2019 09:07:48",
      "content": "<p>Negative predictions? I guess he predicted time to LAST failure and did some math with ttf? ;)</p>",
      "rawMarkdown": "Negative predictions? I guess he predicted time to LAST failure and did some math with ttf? ;)",
      "votes": null
    },
    {
      "id": "526048",
      "postDate": "05/02/2019 09:16:25",
      "content": "<p>That makes sense, <a href=\"/khahuras\">@khahuras</a> . Thanks.\nIn the papers the also tried that route and had some success. I might be giving it a try, too ;)</p>",
      "rawMarkdown": "That makes sense, @khahuras . Thanks.\nIn the papers the also tried that route and had some success. I might be giving it a try, too ;)",
      "votes": null
    },
    {
      "id": "526050",
      "postDate": "05/02/2019 09:23:40",
      "content": "<p>Well just a joke, but they are linearly the same anyway... </p>",
      "rawMarkdown": "Well just a joke, but they are linearly the same anyway...",
      "votes": null
    },
    {
      "id": "526061",
      "postDate": "05/02/2019 09:45:39",
      "content": "<blockquote>\n  <p>Negative predictions? I guess he predicted time to LAST failure and did some math with ttf? ;)</p>\n</blockquote>\n\n<p>No postprocessing or prepossessing of ttf here.  I don't know why there is some negative prediction.  </p>",
      "rawMarkdown": "&gt; Negative predictions? I guess he predicted time to LAST failure and did some math with ttf? ;)\n\nNo postprocessing or prepossessing of ttf here.  I don't know why there is some negative prediction.",
      "votes": null
    },
    {
      "id": "526084",
      "postDate": "05/02/2019 10:20:44",
      "content": "<p><a href=\"/cpmpml\">@cpmpml</a>  your predictions are up to 12 yes, but you have higher prediction for the short quakes. It's always a trade-off.</p>",
      "rawMarkdown": "cpmpml  your predictions are up to 12 yes, but you have higher prediction for the short quakes. It's always a trade-off.",
      "votes": null
    },
    {
      "id": "526089",
      "postDate": "05/02/2019 10:25:43",
      "content": "<p>Higher than what?  I don't get your point.  Predictions for short ttf peaks are lower than the one shared by Ilu elsewhere in this topic discussion for instance.</p>\n\n<p>Maybe you should show yours to make your point.</p>",
      "rawMarkdown": "Higher than what?  I don't get your point.  Predictions for short ttf peaks are lower than the one shared by Ilu elsewhere in this topic discussion for instance.\n\nMaybe you should show yours to make your point.",
      "votes": null
    },
    {
      "id": "526105",
      "postDate": "05/02/2019 11:05:22",
      "content": "<p>Please see my oof results from a LGB model using only one feature:</p>\n\n<p>for 3 folds (CV: 2.0932) we see all TTF ranges are about equal.\n<img src=\"https://i.imgur.com/2EsOJDP.png\" alt=\"3fold\"></p>\n\n<p>but for 4 folds (CV: 2.0813) the TTF ranges already differ quite a bit. Thus, the splits are very important to the model in this competition.\n<img src=\"https://i.imgur.com/oIoHv6p.png\" alt=\"4fold\"></p>\n\n<p>I am not sure if I should trust the 4fold model in this case even though it has the better CV.</p>\n\n<p>Depending on your CV setup, <a href=\"/cpmpml\">@cpmpml</a> , this might also be a reason for different TTF values right after the quakes. In my opinion standard KFold is pretty much the same way as they splitted train and test. </p>",
      "rawMarkdown": "Please see my oof results from a LGB model using only one feature:\n\nfor 3 folds (CV: 2.0932) we see all TTF ranges are about equal.\n![3fold](https://i.imgur.com/2EsOJDP.png)\n\nbut for 4 folds (CV: 2.0813) the TTF ranges already differ quite a bit. Thus, the splits are very important to the model in this competition.\n![4fold](https://i.imgur.com/oIoHv6p.png)\n\nI am not sure if I should trust the 4fold model in this case even though it has the better CV.\n\nDepending on your CV setup, @cpmpml , this might also be a reason for different TTF values right after the quakes. In my opinion standard KFold is pretty much the same way as they splitted train and test.",
      "votes": null
    },
    {
      "id": "526173",
      "postDate": "05/02/2019 13:55:25",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "526214",
      "postDate": "05/02/2019 15:31:55",
      "content": "<p><a href=\"/ilu000\">@ilu000</a> the range of values per period seems rather constant within each of your folds.  It is not the case for my model. </p>",
      "rawMarkdown": "ilu000 the range of values per period seems rather constant within each of your folds.  It is not the case for my model.",
      "votes": null
    },
    {
      "id": "526219",
      "postDate": "05/02/2019 15:47:50",
      "content": "<p>It's totally normal to see some negative predictions using gradient boosting algorithms, even when target is positive definite. It should not happen with RandomForest though. </p>",
      "rawMarkdown": "It's totally normal to see some negative predictions using gradient boosting algorithms, even when target is positive definite. It should not happen with RandomForest though.",
      "votes": null
    },
    {
      "id": "526224",
      "postDate": "05/02/2019 16:11:54",
      "content": "<p>Really? Then postprocessing of the prediction and capping the values at zero can only increase your LB score. </p>",
      "rawMarkdown": "Really? Then postprocessing of the prediction and capping the values at zero can only increase your LB score.",
      "votes": null
    },
    {
      "id": "526233",
      "postDate": "05/02/2019 16:40:01",
      "content": "<p>I expect a very small improvement but, yes, can only make it better</p>",
      "rawMarkdown": "I expect a very small improvement but, yes, can only make it better",
      "votes": null
    },
    {
      "id": "526261",
      "postDate": "05/02/2019 17:22:15",
      "content": "<blockquote>\n  <p>I expect a very small improvemen</p>\n</blockquote>\n\n<p>Indeed, let's say there is 1 prediction at -1.  If we clip it at 0, the improvement on test would be 1/2624= 0.00038</p>",
      "rawMarkdown": "&gt; I expect a very small improvemen\n\nIndeed, let's say there is 1 prediction at -1.  If we clip it at 0, the improvement on test would be 1/2624= 0.00038",
      "votes": null
    },
    {
      "id": "526291",
      "postDate": "05/02/2019 18:18:36",
      "content": "<p><a href=\"/pcjimmmy\">@pcjimmmy</a> thanks for the clarification. I got his point refering back to <a href=\"/steubk\">@steubk</a>'s comment</p>",
      "rawMarkdown": "pcjimmmy thanks for the clarification. I got his point refering back to @steubk's comment",
      "votes": null
    },
    {
      "id": "526298",
      "postDate": "05/02/2019 18:34:04",
      "content": "<p>@cpmp Oversampling? Your plot suggests over 12K points.</p>",
      "rawMarkdown": "cpmp Oversampling? Your plot suggests over 12K points.",
      "votes": null
    },
    {
      "id": "526317",
      "postDate": "05/02/2019 19:49:10",
      "content": "<p>Right, this one was with some data augmentation.</p>",
      "rawMarkdown": "Right, this one was with some data augmentation.",
      "votes": null
    },
    {
      "id": "526362",
      "postDate": "05/02/2019 21:31:17",
      "content": "<p><a href=\"/cpmpml\">@cpmpml</a>  Mine has equal heights between quakes, just like the others. If I'm not mistaken, the variable height you have in your oof is a result of quake-wise early stopping. I wonder how is that working for you in the LB. Intriguing indeed.</p>",
      "rawMarkdown": "cpmpml  Mine has equal heights between quakes, just like the others. If I'm not mistaken, the variable height you have in your oof is a result of quake-wise early stopping. I wonder how is that working for you in the LB. Intriguing indeed.",
      "votes": null
    },
    {
      "id": "526373",
      "postDate": "05/02/2019 22:33:05",
      "content": "<p>Please forgive my ignorance, but what is an oof?  I gather that it's this specific type of graph that Ilu and CPMP have been using to compare their results, but I haven't had any luck googling it to get more context.  Thanks!</p>",
      "rawMarkdown": "Please forgive my ignorance, but what is an oof?  I gather that it's this specific type of graph that Ilu and CPMP have been using to compare their results, but I haven't had any luck googling it to get more context.  Thanks!",
      "votes": null
    },
    {
      "id": "526375",
      "postDate": "05/02/2019 22:36:17",
      "content": "<p><a href=\"/aekoch95\">@aekoch95</a> <code>oof</code> stands for <code>out of fold</code> and includes the combined predictions of each validation set inside the cross validation folds. </p>",
      "rawMarkdown": "aekoch95 `oof` stands for `out of fold` and includes the combined predictions of each validation set inside the cross validation folds.",
      "votes": null
    },
    {
      "id": "526474",
      "postDate": "05/03/2019 06:16:20",
      "content": "<p><a href=\"/amjad85\">@amjad85</a> it works on my cv setting.  Lb improvement is a by product.</p>\n\n<p>Not sure why early stopping is or is not relevant.  Difference between using it or not is tiny in general.</p>",
      "rawMarkdown": "amjad85 it works on my cv setting.  Lb improvement is a by product.\n\nNot sure why early stopping is or is not relevant.  Difference between using it or not is tiny in general.",
      "votes": null
    },
    {
      "id": "526485",
      "postDate": "05/03/2019 06:34:10",
      "content": "<p>Early stopping if you split out single EQs should not work as you leak the length and it does not work for me. Really surprised that it does for you.</p>",
      "rawMarkdown": "Early stopping if you split out single EQs should not work as you leak the length and it does not work for me. Really surprised that it does for you.",
      "votes": null
    },
    {
      "id": "526497",
      "postDate": "05/03/2019 07:07:36",
      "content": "<p>I did not say I use it.</p>",
      "rawMarkdown": "I did not say I use it.",
      "votes": null
    },
    {
      "id": "526499",
      "postDate": "05/03/2019 07:13:57",
      "content": "<p>So, you did not use it ;)</p>",
      "rawMarkdown": "So, you did not use it ;)",
      "votes": null
    },
    {
      "id": "526508",
      "postDate": "05/03/2019 07:28:26",
      "content": "<p>I didn’t say that either.  What I say is that using it or not does not make much difference in general.  I am surprised to see it makes a difference for some here.  I guess it has to do with the cv setting.  Also I don’t get <a href=\"/philippsinger\">@philippsinger</a> comment on length:  which length is it?  </p>",
      "rawMarkdown": "I didn’t say that either.  What I say is that using it or not does not make much difference in general.  I am surprised to see it makes a difference for some here.  I guess it has to do with the cv setting.  Also I don’t get @philippsinger comment on length:  which length is it?",
      "votes": null
    },
    {
      "id": "526514",
      "postDate": "05/03/2019 07:37:58",
      "content": "<p>@CPMP length of EQ</p>",
      "rawMarkdown": "CPMP length of EQ",
      "votes": null
    },
    {
      "id": "526529",
      "postDate": "05/03/2019 08:26:32",
      "content": "<p>How would this leak?</p>",
      "rawMarkdown": "How would this leak?",
      "votes": null
    },
    {
      "id": "526762",
      "postDate": "05/03/2019 17:26:15",
      "content": "<p>@CPMP in the graph with your oof predictions, do you use shuffled CV or unshuffled (or CV by EQ)? Would be interested to know...</p>",
      "rawMarkdown": "CPMP in the graph with your oof predictions, do you use shuffled CV or unshuffled (or CV by EQ)? Would be interested to know...",
      "votes": null
    },
    {
      "id": "526823",
      "postDate": "05/03/2019 21:10:23",
      "content": "<blockquote>\n  <p>do you use shuffled CV or unshuffled (or CV by EQ)? </p>\n</blockquote>\n\n<p>You'll know after competition end ;)</p>",
      "rawMarkdown": "&gt; do you use shuffled CV or unshuffled (or CV by EQ)? \n\nYou'll know after competition end ;)",
      "votes": null
    },
    {
      "id": "526830",
      "postDate": "05/03/2019 21:26:10",
      "content": "<p>Hehe 👍 </p>",
      "rawMarkdown": "Hehe 👍",
      "votes": null
    },
    {
      "id": "527210",
      "postDate": "05/04/2019 20:19:53",
      "content": "<p>I've added classification of training data by type of earthquake (presence of intermediate release). For each class TTF can be predicted with much better score, than for mixed data.\nMy assumption was, that derivative of signal energy (or frequency) would be lower for long earthquakes (event escalates slowly giving chance for intermediate relieve).\nBut none of my features show more than 0.13 Pearson correlation with this classification, and models give about 34% error predicting correct class, which is not enough to use it to improve final score.\nSo I'm really curious to know what features could help reliably predict prolonged earthquake.</p>",
      "rawMarkdown": "I've added classification of training data by type of earthquake (presence of intermediate release). For each class TTF can be predicted with much better score, than for mixed data.\nMy assumption was, that derivative of signal energy (or frequency) would be lower for long earthquakes (event escalates slowly giving chance for intermediate relieve).\nBut none of my features show more than 0.13 Pearson correlation with this classification, and models give about 34% error predicting correct class, which is not enough to use it to improve final score.\nSo I'm really curious to know what features could help reliably predict prolonged earthquake.",
      "votes": null
    },
    {
      "id": "532814",
      "postDate": "05/17/2019 18:39:31",
      "content": "<p><a href=\"/ilu000\">@ilu000</a> I assume your plots show CV predictions, but you are not using it \"as is\" for the final prediction - you either average it or retrain model on all the data. So these plots are a little bit misleading. 4fold model can be good for the ensemble after averaging. Also, assuming the CPMP plot shows oof, it will not be such accurate after averaging, it will rather show constant predictions (I mean predictions will be in a given range).</p>",
      "rawMarkdown": "ilu000 I assume your plots show CV predictions, but you are not using it \"as is\" for the final prediction - you either average it or retrain model on all the data. So these plots are a little bit misleading. 4fold model can be good for the ensemble after averaging. Also, assuming the CPMP plot shows oof, it will not be such accurate after averaging, it will rather show constant predictions (I mean predictions will be in a given range).",
      "votes": null
    },
    {
      "id": "536029",
      "postDate": "05/23/2019 20:45:40",
      "content": "<p>I'm not able to predict anything over 10 seconds.</p>",
      "rawMarkdown": "I'm not able to predict anything over 10 seconds.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 525553,
      "author_name": "alexfir",
      "author_url": "",
      "post_date": "05/01/2019 09:18:26",
      "content": "<p>Regression models predict expected value of target, and it usually not close to edge values especially when number of observation with this extreme values are small.\nYou may try something like <a href=\"https://www.kaggle.com/c/elo-merchant-category-recommendation/discussion/82036\">winner of recent ELO competition did</a> - combine regression and classification.</p>",
      "votes": null,
      "replies": [
        {
          "id": 525801,
          "author_name": "mhviraf",
          "author_url": "",
          "post_date": "05/01/2019 18:25:54",
          "content": "<p>Thank you for your input. I think the count of those points (about 14% of total) is not too small for the model to completely ignore their effects. I had a glance at the link you posted; couldn't understand it :D Need to take more time to digest it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 525582,
      "author_name": "ilu000",
      "author_url": "",
      "post_date": "05/01/2019 10:36:46",
      "content": "<p>I believe this problem cannot be solved and is inherently anchored to the physics of the experiment:\nAfter each slip (quake) the crumbled parts solidify again and due to the constant speed of the piston we always expect the same TTF. But in this experiment we also have some minor slips that are kind of resetting back the state of the experiment. That adds a few seconds of TTF and it can not be predicted by a model if the coming quake is major or minor. The predictions between minor and major quakes are in line with the ground truth again. </p>\n\n<p><img src=\"https://i.imgur.com/Htg77qk.png\" alt=\"minor quake\"></p>\n\n<p>If anyone would be able to model this, expect not less then a nature paper from your results ;)</p>\n\n<p>I am very curious about the top solutions in this competition.</p>",
      "votes": null,
      "replies": [
        {
          "id": 525799,
          "author_name": "mhviraf",
          "author_url": "",
          "post_date": "05/01/2019 18:22:16",
          "content": "<p>Thanks <a href=\"/ilu000\">@ilu000</a>. This graph is very informative. I am curious to see top solutions too. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 525837,
          "author_name": "filipmulier",
          "author_url": "",
          "post_date": "05/01/2019 19:59:56",
          "content": "<p>I agree, I think it has to do with the physics of the problem.  </p>\n\n<p>The data shows 14 complete experiments.  </p>\n\n<p>Here is my assumption -  The initial conditions of the lab earthquake system at the start of each experiment are controlled to be nearly identical by controlling the quantity and type of materials and setup procedure.  With time the system itself is influenced by its own random/chaotic behavior and its internal 'state' starts to move away from initial conditions,  and this reflection of state (hopefully) shows up in the acoustic data.</p>\n\n<p>If this assumption is true the acoustic  samples early in each experimental run should look very similar, even though the experiment's final time_to_failure will be different.   </p>\n\n<p>The 'time_to_failure' data is retrospective -not causal.  At the time point the experiment is started, no one knows the time_to_failure, since it has not yet physically failed in the run. </p>\n\n<p>That is what makes this problem hard. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 525841,
          "author_name": "aekoch95",
          "author_url": "",
          "post_date": "05/01/2019 20:06:13",
          "content": "<p>This also explains a really mysterious pattern I've been seeing in some of my cross-validation results.  On some folds, especially those which include longer ttf, I'm seeing this strange pattern of two distinct trendlines with slope of 1.  I've attached an example of my version of OP's results graph for one of my cross validation folds that exhibits this pattern.  </p>\n\n<p>You can clearly see two seperate trendlines, which after reading Ilu's hypothesis I suspect is actually evidence of the algorithm learning to correctly predict the ttf of minor earthquakes. These minor quakes occur within some but not all of our 16 quake samples, which is why this pattern only appears in some cross validation folds.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 525842,
          "author_name": "amjad85",
          "author_url": "",
          "post_date": "05/01/2019 20:09:44",
          "content": "<p>I wonder if the winners will be part of the nature paper the hosts are going to write after this competition</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 525847,
          "author_name": "amjad85",
          "author_url": "",
          "post_date": "05/01/2019 20:16:11",
          "content": "<p>totally agree with <a href=\"/filipmulier\">@filipmulier</a>. I wrote a similar explanation <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/89301#515975\">here</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 525854,
          "author_name": "mhviraf",
          "author_url": "",
          "post_date": "05/01/2019 20:44:06",
          "content": "<p><a href=\"/amjad85\">@amjad85</a>  I think I read somewhere that they are willing to include winners of this competition in their next papers. Not sure if it's gonna be a Nature paper though ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 525856,
          "author_name": "filipmulier",
          "author_url": "",
          "post_date": "05/01/2019 20:47:45",
          "content": "<p>Aaron, In your graph I see at least 2 bands, and maybe 4.  Could this be the experiment (quake ID) showing up as contributing to the variance -  See Amjad's post on CV above. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 525875,
          "author_name": "aekoch95",
          "author_url": "",
          "post_date": "05/01/2019 22:44:33",
          "content": "<p>Filip, the image I posted was taken from one fold of a straight 5-fold split.  If it is the quake ID that is responsible for that distinctive band of low prediction for high expected ttf, then that band should disappear when I switch to 'leave one quake out' CV.  However it does not.  In fact, the band shows up clearly, but only for quakes with high maximum ttf (&gt;10).  Quakes with low ttf, around 8-9, never show this distinctive banding pattern, though there may be other patterns in the error of individual quakes. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 525907,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "05/02/2019 01:39:07",
          "content": "<p>When I read the 4 documents that describe the previous experiments and publication of results I see that the \"data\" is probably grabbed after the process of making quakes has been established and fairly constant.  There is a sketch in one of the pubs that shows this clearly!</p>\n\n<p>The papers indicate to me that the definition of a quake is when the stress Gage goes to zero.  We don't have the stress data - if we do it might help explain the minor quakes.  In the pubs you can see dips in the stress plots for very brief time that probably are the minor quakes we are seeing.  </p>\n\n<p>It bugs me a bit that the quake end time of zero is after the big spikes in the acoustic signal.  I assume therefore that after the big acoustic signal the stress starts dropping and that bit of calm before zero time is related to the time it takes for stress to hit the cutoff point being used.  I am not a real fan to sponsors who leave out what I think is key data from our work - in this case the stress seems very important to me.</p>\n\n<p>Suggest all read the four pubs in the Welcome - they helped me either better understand (or continue to mis-understand the experiment and data collection)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 525912,
          "author_name": "bigironsphere",
          "author_url": "",
          "post_date": "05/02/2019 01:51:01",
          "content": "<p><a href=\"/pcjimmmy\">@pcjimmmy</a> The purpose of this competition is to see if our findings can be extrapolated to real world earthquake prediction. We cannot measure stress and strain deep underneath the Earth; we can only measure acoustic seismological waves, and in a very short timespan compared to geological scales. The data constraints are a problem, but they make sense in light of the researchers' aims.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 525925,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "05/02/2019 02:47:57",
          "content": "<p>OK - makes sense that stress needs to be left out - than my only remaining issue is that the zero time for quakes was based on stress rather than big signals.  Probably would have been too hard to use real world data and quakes - the folks who participate in these challenges are pretty smart about finding leaks in the data :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 525611,
      "author_name": "scirpus",
      "author_url": "",
      "post_date": "05/01/2019 11:54:03",
      "content": "<p>Have a look at this</p>\n\n<p>They don't seem to have dips either they have a very good model or they have cherry picked the plot! ;)</p>\n\n<p><a href=\"https://www.researchgate.net/figure/Time-remaining-before-the-next-failure-predicted-by-the-Random-Forest-As-in-Fig-1a_fig3_313858017\">Image</a></p>\n\n<p>Edit:  It looks like they count the \"Minor Quakes\" as Quakes whereas we have to predict only the major quakes.</p>",
      "votes": null,
      "replies": [
        {
          "id": 525623,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "05/01/2019 12:18:39",
          "content": "<p>&gt; Edit: It looks like they count the \"Minor Quakes\" as Quakes whereas we have to predict only the major quakes.</p>\n\n<p>this is exactly what makes it trivial to model. Here we need to find out the difference before a minor and a major quake.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 525630,
          "author_name": "scirpus",
          "author_url": "",
          "post_date": "05/01/2019 12:28:51",
          "content": "<p>Differentiation would be a lot easier if they hadn't shuffled test;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 525684,
          "author_name": "steubk",
          "author_url": "",
          "post_date": "05/01/2019 14:31:05",
          "content": "<p>They also use a time window of 1.8 sec (!) </p>\n\n<blockquote>\n  <p>To create a model that uncovers the physics of shear failure, we make predictions using moving time windows applied to the data. Each window is 1.8s, which is small compared to the time between fault gouge failures (8s on average).</p>\n</blockquote>\n\n<p><a href=\"https://agupubs.onlinelibrary.wiley.com/action/downloadSupplement?doi=10.1002%2F2017GL074677&amp;file=grl56367-sup-0001-supinfo.pdf\">grl56367-sup-0001-supinfo.pdf</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 525796,
          "author_name": "mhviraf",
          "author_url": "",
          "post_date": "05/01/2019 18:17:28",
          "content": "<p><a href=\"/scirpus\">@scirpus</a> <a href=\"/steubk\">@steubk</a>  <a href=\"/ilu000\">@ilu000</a>  Thank you for pointing out this. I didn't know Minor quakes are a thing. I will investigate it more and share if I found a way to enhance my models by this.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 525911,
          "author_name": "vettejeep",
          "author_url": "",
          "post_date": "05/02/2019 01:47:24",
          "content": "<p>Agree, 1.8 sec makes it a lot easier.  The BIG spike is then in a lot of signals rather than 16/4194.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 525621,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "05/01/2019 12:13:46",
      "content": "<p>My model doesn't fail this.  But there is a max value lower than 16 indeed.  This is probably because there are way less samples with very high ttf values.  Also, it may be that some ttf are very high because there are miniquakes later that relive the stress enough to postpone the real quake.  Given we only have a small snapshot, it is impossible to predict this situation.</p>\n\n<p>Edit: I see Ilu gave the same explanation.</p>",
      "votes": null,
      "replies": [
        {
          "id": 525627,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "05/01/2019 12:23:00",
          "content": "<p>That implies, you have found a way to predict minor quakes?\nIncredible work and that would explain your score. Did everyone up in the top 50 find that method?</p>\n\n<p>Would you mind clarifying your statement? Or are you afraid of another \"Santander\"?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 525631,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/01/2019 12:29:20",
          "content": "<p>No, my model is able to predict ttf &gt; 11 up to some value.  Not sure why others can't actually.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 525661,
          "author_name": "bigironsphere",
          "author_url": "",
          "post_date": "05/01/2019 13:37:18",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> I've found that a couple of features make predicting TTF up to about 14 trivial, but only as a result of K-fold leakage. I'll be very impressed if you've found a way to do this without any leakage. Like <a href=\"/ilu000\">@ilu000</a> said, it's contrary to my understanding of the physics at play here.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 525702,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/01/2019 15:14:09",
          "content": "<p>I don't think I have leakage.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 525793,
          "author_name": "mhviraf",
          "author_url": "",
          "post_date": "05/01/2019 18:14:12",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> there are also some datapoints where true ttf is less than about 0.5 but all of my models predict them somehow randomly. look at the bottom of the figure I posted earlier; there are about 1% of points that have a very different trend than the rest of them. Do you happen to have the same problem in your models?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 525800,
          "author_name": "scirpus",
          "author_url": "",
          "post_date": "05/01/2019 18:25:24",
          "content": "<p>Might be worth using the denoise kernel</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 525802,
          "author_name": "bigironsphere",
          "author_url": "",
          "post_date": "05/01/2019 18:29:05",
          "content": "<p><a href=\"/mhviraf\">@mhviraf</a> The models often have a hard time distinguishing between the start and end of earthquake periods as the signal profile is very similar.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 525815,
          "author_name": "mhviraf",
          "author_url": "",
          "post_date": "05/01/2019 19:00:55",
          "content": "<p><a href=\"/scirpus\">@scirpus</a> why do you think denoised data would solve this?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 525817,
          "author_name": "mhviraf",
          "author_url": "",
          "post_date": "05/01/2019 19:04:20",
          "content": "<p><a href=\"/bigironsphere\">@bigironsphere</a> that's a good point. based on this explanation I expect to see very low (or randomly distributed) predictions for very large TTFs too but this is not the case and models could surprisingly distinguish them. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 525836,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "05/01/2019 19:55:59",
          "content": "<p>I am with <a href=\"/bigironsphere\">@bigironsphere</a> here in terms of the physics in this experiment. Predicting TTF values above 11 sec implies an upcoming minor quake. So your model must be able to somehow predict those minor quakes <a href=\"/cpmpml\">@cpmpml</a> . \nIf this isn't some kind of leakage I am truly impressed and eagerly looking forward to your solution. \nIf you don't mind, you might play with your features a bit and find out which feature(s) is/are helping to identify the minor slips. Of course you can wait until after the competition with revealing that information ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 525871,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/01/2019 22:27:50",
          "content": "<p>I'll share my approach after competition end for sure.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 525889,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/02/2019 00:01:03",
          "content": "<p>Sharing one model with max predictions above 12.  The model is still fooled by the mini quake after the  7th quake.  </p>\n\n<p>As many others predicting ttf below 2 is an issue.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/525889/13136/pred.png\" alt=\"pred\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526033,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "05/02/2019 08:48:57",
          "content": "<p>Thank you for sharing <a href=\"/cpmpml\">@cpmpml</a> . Is this a LGB model?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526034,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/02/2019 08:49:42",
          "content": "<p>Lgb indeed.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526036,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "05/02/2019 08:58:48",
          "content": "<p>Very interesting. So there really seems to be a way to separate the quakes. \nYou claimed that you don't have a leak in your code. Does that include you removed the mean of the data? I saw that the model is able to cluster the quakes if the mean is used. \nIs there a way to cluster the test set (new unseen clusters) and make that information available to a LGB model?</p>\n\n<p>It seems to me that clustering/separating the quakes is the key to win this competition.</p>\n\n<p>Edit: I just saw that you have 2 negative TTF predictions. How is that possible?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526038,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/02/2019 09:02:38",
          "content": "<p>Mean of data leaks?  Interesting.  I am not using it anyway as I assumed that mean should be 0 for acoustic data and that any deviation is a measurement artifact.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526043,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "05/02/2019 09:07:17",
          "content": "<p>Yes, depending on how you split your folds, the model is able to say which quake corresponds to a data chunk by evaluating the mean. The mean is slowly drifting thoughout the train set. Of course that can only give an advantage if the model is trained with a part of the same quake. \nBut I have no idea how to use that in the test set. Thus, i substracted the mean in all data chunks. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526044,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "05/02/2019 09:07:48",
          "content": "<p>Negative predictions? I guess he predicted time to LAST failure and did some math with ttf? ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526048,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "05/02/2019 09:16:25",
          "content": "<p>That makes sense, <a href=\"/khahuras\">@khahuras</a> . Thanks.\nIn the papers the also tried that route and had some success. I might be giving it a try, too ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526050,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "05/02/2019 09:23:40",
          "content": "<p>Well just a joke, but they are linearly the same anyway... </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526061,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/02/2019 09:45:39",
          "content": "<blockquote>\n  <p>Negative predictions? I guess he predicted time to LAST failure and did some math with ttf? ;)</p>\n</blockquote>\n\n<p>No postprocessing or prepossessing of ttf here.  I don't know why there is some negative prediction.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526084,
          "author_name": "amjad85",
          "author_url": "",
          "post_date": "05/02/2019 10:20:44",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a>  your predictions are up to 12 yes, but you have higher prediction for the short quakes. It's always a trade-off.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526089,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/02/2019 10:25:43",
          "content": "<p>Higher than what?  I don't get your point.  Predictions for short ttf peaks are lower than the one shared by Ilu elsewhere in this topic discussion for instance.</p>\n\n<p>Maybe you should show yours to make your point.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526105,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "05/02/2019 11:05:22",
          "content": "<p>Please see my oof results from a LGB model using only one feature:</p>\n\n<p>for 3 folds (CV: 2.0932) we see all TTF ranges are about equal.\n<img src=\"https://i.imgur.com/2EsOJDP.png\" alt=\"3fold\"></p>\n\n<p>but for 4 folds (CV: 2.0813) the TTF ranges already differ quite a bit. Thus, the splits are very important to the model in this competition.\n<img src=\"https://i.imgur.com/oIoHv6p.png\" alt=\"4fold\"></p>\n\n<p>I am not sure if I should trust the 4fold model in this case even though it has the better CV.</p>\n\n<p>Depending on your CV setup, <a href=\"/cpmpml\">@cpmpml</a> , this might also be a reason for different TTF values right after the quakes. In my opinion standard KFold is pretty much the same way as they splitted train and test. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526214,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/02/2019 15:31:55",
          "content": "<p><a href=\"/ilu000\">@ilu000</a> the range of values per period seems rather constant within each of your folds.  It is not the case for my model. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526219,
          "author_name": "stecasasso",
          "author_url": "",
          "post_date": "05/02/2019 15:47:50",
          "content": "<p>It's totally normal to see some negative predictions using gradient boosting algorithms, even when target is positive definite. It should not happen with RandomForest though. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526224,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "05/02/2019 16:11:54",
          "content": "<p>Really? Then postprocessing of the prediction and capping the values at zero can only increase your LB score. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526233,
          "author_name": "stecasasso",
          "author_url": "",
          "post_date": "05/02/2019 16:40:01",
          "content": "<p>I expect a very small improvement but, yes, can only make it better</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526261,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/02/2019 17:22:15",
          "content": "<blockquote>\n  <p>I expect a very small improvemen</p>\n</blockquote>\n\n<p>Indeed, let's say there is 1 prediction at -1.  If we clip it at 0, the improvement on test would be 1/2624= 0.00038</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526298,
          "author_name": "mykper",
          "author_url": "",
          "post_date": "05/02/2019 18:34:04",
          "content": "<p>@cpmp Oversampling? Your plot suggests over 12K points.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526317,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/02/2019 19:49:10",
          "content": "<p>Right, this one was with some data augmentation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526362,
          "author_name": "amjad85",
          "author_url": "",
          "post_date": "05/02/2019 21:31:17",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a>  Mine has equal heights between quakes, just like the others. If I'm not mistaken, the variable height you have in your oof is a result of quake-wise early stopping. I wonder how is that working for you in the LB. Intriguing indeed.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526373,
          "author_name": "aekoch95",
          "author_url": "",
          "post_date": "05/02/2019 22:33:05",
          "content": "<p>Please forgive my ignorance, but what is an oof?  I gather that it's this specific type of graph that Ilu and CPMP have been using to compare their results, but I haven't had any luck googling it to get more context.  Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526375,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "05/02/2019 22:36:17",
          "content": "<p><a href=\"/aekoch95\">@aekoch95</a> <code>oof</code> stands for <code>out of fold</code> and includes the combined predictions of each validation set inside the cross validation folds. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526474,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/03/2019 06:16:20",
          "content": "<p><a href=\"/amjad85\">@amjad85</a> it works on my cv setting.  Lb improvement is a by product.</p>\n\n<p>Not sure why early stopping is or is not relevant.  Difference between using it or not is tiny in general.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526485,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "05/03/2019 06:34:10",
          "content": "<p>Early stopping if you split out single EQs should not work as you leak the length and it does not work for me. Really surprised that it does for you.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526497,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/03/2019 07:07:36",
          "content": "<p>I did not say I use it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526499,
          "author_name": "steubk",
          "author_url": "",
          "post_date": "05/03/2019 07:13:57",
          "content": "<p>So, you did not use it ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526508,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/03/2019 07:28:26",
          "content": "<p>I didn’t say that either.  What I say is that using it or not does not make much difference in general.  I am surprised to see it makes a difference for some here.  I guess it has to do with the cv setting.  Also I don’t get <a href=\"/philippsinger\">@philippsinger</a> comment on length:  which length is it?  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526514,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "05/03/2019 07:37:58",
          "content": "<p>@CPMP length of EQ</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526529,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/03/2019 08:26:32",
          "content": "<p>How would this leak?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526762,
          "author_name": "danijelk",
          "author_url": "",
          "post_date": "05/03/2019 17:26:15",
          "content": "<p>@CPMP in the graph with your oof predictions, do you use shuffled CV or unshuffled (or CV by EQ)? Would be interested to know...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526823,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/03/2019 21:10:23",
          "content": "<blockquote>\n  <p>do you use shuffled CV or unshuffled (or CV by EQ)? </p>\n</blockquote>\n\n<p>You'll know after competition end ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526830,
          "author_name": "danijelk",
          "author_url": "",
          "post_date": "05/03/2019 21:26:10",
          "content": "<p>Hehe 👍 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 527210,
          "author_name": "elvenmonk",
          "author_url": "",
          "post_date": "05/04/2019 20:19:53",
          "content": "<p>I've added classification of training data by type of earthquake (presence of intermediate release). For each class TTF can be predicted with much better score, than for mixed data.\nMy assumption was, that derivative of signal energy (or frequency) would be lower for long earthquakes (event escalates slowly giving chance for intermediate relieve).\nBut none of my features show more than 0.13 Pearson correlation with this classification, and models give about 34% error predicting correct class, which is not enough to use it to improve final score.\nSo I'm really curious to know what features could help reliably predict prolonged earthquake.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 532814,
          "author_name": "davids1992",
          "author_url": "",
          "post_date": "05/17/2019 18:39:31",
          "content": "<p><a href=\"/ilu000\">@ilu000</a> I assume your plots show CV predictions, but you are not using it \"as is\" for the final prediction - you either average it or retrain model on all the data. So these plots are a little bit misleading. 4fold model can be good for the ensemble after averaging. Also, assuming the CPMP plot shows oof, it will not be such accurate after averaging, it will rather show constant predictions (I mean predictions will be in a given range).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 525778,
      "author_name": "hmcranbercourt",
      "author_url": "",
      "post_date": "05/01/2019 17:34:22",
      "content": "<p>How does this all relate to a chunk size of 150000 as most kernels are using?\n1.8 seconds is much larger than 150000 datapoints.</p>",
      "votes": null,
      "replies": [
        {
          "id": 525791,
          "author_name": "mhviraf",
          "author_url": "",
          "post_date": "05/01/2019 18:08:53",
          "content": "<p>Would you explain more? I didn't get your point.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 525811,
          "author_name": "alexfir",
          "author_url": "",
          "post_date": "05/01/2019 18:50:24",
          "content": "<p>Hein is referring to <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91125#525684\">the comment</a> by steubk.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 525906,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "05/02/2019 01:26:41",
          "content": "<p>I think he means that if you selected a 150000 size chuck of data - the last row of that data set becomes the target value.   </p>\n\n<p>So model is not seeing any of the times for the first 149999 rows and will therefore not predict them.  </p>\n\n<p>You could confirm this by running your model and have the y value the first row of the segment rather than the last - I think in most kernels that's a one line simple change.  </p>\n\n<p>Hmm - not sure I made it any clearer but I did say it in a different way :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526291,
          "author_name": "mhviraf",
          "author_url": "",
          "post_date": "05/02/2019 18:18:36",
          "content": "<p><a href=\"/pcjimmmy\">@pcjimmmy</a> thanks for the clarification. I got his point refering back to <a href=\"/steubk\">@steubk</a>'s comment</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 526173,
      "author_name": "hmcranbercourt",
      "author_url": "",
      "post_date": "05/02/2019 13:55:25",
      "content": "",
      "votes": null,
      "replies": []
    },
    {
      "id": 536029,
      "author_name": "gerlav",
      "author_url": "",
      "post_date": "05/23/2019 20:45:40",
      "content": "<p>I'm not able to predict anything over 10 seconds.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "525533": "Attached is the result of my LightGBM model (x axis) vs. true TTF values (y axis). My model fails to predict anything above approx. 10.5 ~ 11 seconds. Those data points form 14.35% of the train data and since our metric is MAE, they may be the most contributing points due to their large errors. I also observed this trend in my other models including RNN, CNN and even linear regression models. I saw this in results of public kernels too.\n\nI would like to hear your thought on this. What do you think is the reason? How do you think I can enhance it?",
    "525553": "Regression models predict expected value of target, and it usually not close to edge values especially when number of observation with this extreme values are small.\nYou may try something like [winner of recent ELO competition did](https://www.kaggle.com/c/elo-merchant-category-recommendation/discussion/82036) - combine regression and classification.",
    "525582": "I believe this problem cannot be solved and is inherently anchored to the physics of the experiment:\nAfter each slip (quake) the crumbled parts solidify again and due to the constant speed of the piston we always expect the same TTF. But in this experiment we also have some minor slips that are kind of resetting back the state of the experiment. That adds a few seconds of TTF and it can not be predicted by a model if the coming quake is major or minor. The predictions between minor and major quakes are in line with the ground truth again. \n\n![minor quake](https://i.imgur.com/Htg77qk.png)\n\n\nIf anyone would be able to model this, expect not less then a nature paper from your results ;)\n\nI am very curious about the top solutions in this competition.",
    "525611": "Have a look at this\n\n\n\nThey don't seem to have dips either they have a very good model or they have cherry picked the plot! ;)\n\n[Image](https://www.researchgate.net/figure/Time-remaining-before-the-next-failure-predicted-by-the-Random-Forest-As-in-Fig-1a_fig3_313858017)\n\nEdit:  It looks like they count the \"Minor Quakes\" as Quakes whereas we have to predict only the major quakes.",
    "525621": "My model doesn't fail this.  But there is a max value lower than 16 indeed.  This is probably because there are way less samples with very high ttf values.  Also, it may be that some ttf are very high because there are miniquakes later that relive the stress enough to postpone the real quake.  Given we only have a small snapshot, it is impossible to predict this situation.\n\nEdit: I see Ilu gave the same explanation.",
    "525623": "&gt; Edit: It looks like they count the \"Minor Quakes\" as Quakes whereas we have to predict only the major quakes.\n\nthis is exactly what makes it trivial to model. Here we need to find out the difference before a minor and a major quake.",
    "525627": "That implies, you have found a way to predict minor quakes?\nIncredible work and that would explain your score. Did everyone up in the top 50 find that method?\n\nWould you mind clarifying your statement? Or are you afraid of another \"Santander\"?",
    "525630": "Differentiation would be a lot easier if they hadn't shuffled test;)",
    "525631": "No, my model is able to predict ttf &gt; 11 up to some value.  Not sure why others can't actually.",
    "525661": "cpmpml I've found that a couple of features make predicting TTF up to about 14 trivial, but only as a result of K-fold leakage. I'll be very impressed if you've found a way to do this without any leakage. Like @ilu000 said, it's contrary to my understanding of the physics at play here.",
    "525684": "They also use a time window of 1.8 sec (!) \n&gt; To create a model that uncovers the physics of shear failure, we make predictions using moving time windows applied to the data. Each window is 1.8s, which is small compared to the time between fault gouge failures (8s on average).\n\n[grl56367-sup-0001-supinfo.pdf](https://agupubs.onlinelibrary.wiley.com/action/downloadSupplement?doi=10.1002%2F2017GL074677&amp;file=grl56367-sup-0001-supinfo.pdf)",
    "525702": "I don't think I have leakage.",
    "525778": "How does this all relate to a chunk size of 150000 as most kernels are using?\n1.8 seconds is much larger than 150000 datapoints.",
    "525791": "Would you explain more? I didn't get your point.",
    "525793": "cpmpml there are also some datapoints where true ttf is less than about 0.5 but all of my models predict them somehow randomly. look at the bottom of the figure I posted earlier; there are about 1% of points that have a very different trend than the rest of them. Do you happen to have the same problem in your models?",
    "525796": "scirpus @steubk  @ilu000  Thank you for pointing out this. I didn't know Minor quakes are a thing. I will investigate it more and share if I found a way to enhance my models by this.",
    "525799": "Thanks @ilu000. This graph is very informative. I am curious to see top solutions too.",
    "525800": "Might be worth using the denoise kernel",
    "525801": "Thank you for your input. I think the count of those points (about 14% of total) is not too small for the model to completely ignore their effects. I had a glance at the link you posted; couldn't understand it :D Need to take more time to digest it.",
    "525802": "mhviraf The models often have a hard time distinguishing between the start and end of earthquake periods as the signal profile is very similar.",
    "525811": "Hein is referring to [the comment](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91125#525684) by steubk.",
    "525815": "scirpus why do you think denoised data would solve this?",
    "525817": "bigironsphere that's a good point. based on this explanation I expect to see very low (or randomly distributed) predictions for very large TTFs too but this is not the case and models could surprisingly distinguish them.",
    "525836": "I am with @bigironsphere here in terms of the physics in this experiment. Predicting TTF values above 11 sec implies an upcoming minor quake. So your model must be able to somehow predict those minor quakes @cpmpml . \nIf this isn't some kind of leakage I am truly impressed and eagerly looking forward to your solution. \nIf you don't mind, you might play with your features a bit and find out which feature(s) is/are helping to identify the minor slips. Of course you can wait until after the competition with revealing that information ;)",
    "525837": "I agree, I think it has to do with the physics of the problem.  \n\nThe data shows 14 complete experiments.  \n\nHere is my assumption -  The initial conditions of the lab earthquake system at the start of each experiment are controlled to be nearly identical by controlling the quantity and type of materials and setup procedure.  With time the system itself is influenced by its own random/chaotic behavior and its internal 'state' starts to move away from initial conditions,  and this reflection of state (hopefully) shows up in the acoustic data.\n\nIf this assumption is true the acoustic  samples early in each experimental run should look very similar, even though the experiment's final time_to_failure will be different.   \n\nThe 'time_to_failure' data is retrospective -not causal.  At the time point the experiment is started, no one knows the time_to_failure, since it has not yet physically failed in the run. \n\nThat is what makes this problem hard.",
    "525841": "This also explains a really mysterious pattern I've been seeing in some of my cross-validation results.  On some folds, especially those which include longer ttf, I'm seeing this strange pattern of two distinct trendlines with slope of 1.  I've attached an example of my version of OP's results graph for one of my cross validation folds that exhibits this pattern.  \n\nYou can clearly see two seperate trendlines, which after reading Ilu's hypothesis I suspect is actually evidence of the algorithm learning to correctly predict the ttf of minor earthquakes. These minor quakes occur within some but not all of our 16 quake samples, which is why this pattern only appears in some cross validation folds.",
    "525842": "I wonder if the winners will be part of the nature paper the hosts are going to write after this competition",
    "525847": "totally agree with @filipmulier. I wrote a similar explanation [here](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/89301#515975)",
    "525854": "amjad85  I think I read somewhere that they are willing to include winners of this competition in their next papers. Not sure if it's gonna be a Nature paper though ;)",
    "525856": "Aaron, In your graph I see at least 2 bands, and maybe 4.  Could this be the experiment (quake ID) showing up as contributing to the variance -  See Amjad's post on CV above.",
    "525871": "I'll share my approach after competition end for sure.",
    "525875": "Filip, the image I posted was taken from one fold of a straight 5-fold split.  If it is the quake ID that is responsible for that distinctive band of low prediction for high expected ttf, then that band should disappear when I switch to 'leave one quake out' CV.  However it does not.  In fact, the band shows up clearly, but only for quakes with high maximum ttf (&gt;10).  Quakes with low ttf, around 8-9, never show this distinctive banding pattern, though there may be other patterns in the error of individual quakes.",
    "525889": "Sharing one model with max predictions above 12.  The model is still fooled by the mini quake after the  7th quake.  \n\nAs many others predicting ttf below 2 is an issue.\n\n![pred](https://storage.googleapis.com/kaggle-forum-message-attachments/525889/13136/pred.png)",
    "525906": "I think he means that if you selected a 150000 size chuck of data - the last row of that data set becomes the target value.   \n\nSo model is not seeing any of the times for the first 149999 rows and will therefore not predict them.  \n\nYou could confirm this by running your model and have the y value the first row of the segment rather than the last - I think in most kernels that's a one line simple change.  \n\nHmm - not sure I made it any clearer but I did say it in a different way :)",
    "525907": "When I read the 4 documents that describe the previous experiments and publication of results I see that the \"data\" is probably grabbed after the process of making quakes has been established and fairly constant.  There is a sketch in one of the pubs that shows this clearly!\n\nThe papers indicate to me that the definition of a quake is when the stress Gage goes to zero.  We don't have the stress data - if we do it might help explain the minor quakes.  In the pubs you can see dips in the stress plots for very brief time that probably are the minor quakes we are seeing.  \n\nIt bugs me a bit that the quake end time of zero is after the big spikes in the acoustic signal.  I assume therefore that after the big acoustic signal the stress starts dropping and that bit of calm before zero time is related to the time it takes for stress to hit the cutoff point being used.  I am not a real fan to sponsors who leave out what I think is key data from our work - in this case the stress seems very important to me.\n\nSuggest all read the four pubs in the Welcome - they helped me either better understand (or continue to mis-understand the experiment and data collection)",
    "525911": "Agree, 1.8 sec makes it a lot easier.  The BIG spike is then in a lot of signals rather than 16/4194.",
    "525912": "pcjimmmy The purpose of this competition is to see if our findings can be extrapolated to real world earthquake prediction. We cannot measure stress and strain deep underneath the Earth; we can only measure acoustic seismological waves, and in a very short timespan compared to geological scales. The data constraints are a problem, but they make sense in light of the researchers' aims.",
    "525925": "OK - makes sense that stress needs to be left out - than my only remaining issue is that the zero time for quakes was based on stress rather than big signals.  Probably would have been too hard to use real world data and quakes - the folks who participate in these challenges are pretty smart about finding leaks in the data :)",
    "526033": "Thank you for sharing @cpmpml . Is this a LGB model?",
    "526034": "Lgb indeed.",
    "526036": "Very interesting. So there really seems to be a way to separate the quakes. \nYou claimed that you don't have a leak in your code. Does that include you removed the mean of the data? I saw that the model is able to cluster the quakes if the mean is used. \nIs there a way to cluster the test set (new unseen clusters) and make that information available to a LGB model?\n\nIt seems to me that clustering/separating the quakes is the key to win this competition.\n\nEdit: I just saw that you have 2 negative TTF predictions. How is that possible?",
    "526038": "Mean of data leaks?  Interesting.  I am not using it anyway as I assumed that mean should be 0 for acoustic data and that any deviation is a measurement artifact.",
    "526043": "Yes, depending on how you split your folds, the model is able to say which quake corresponds to a data chunk by evaluating the mean. The mean is slowly drifting thoughout the train set. Of course that can only give an advantage if the model is trained with a part of the same quake. \nBut I have no idea how to use that in the test set. Thus, i substracted the mean in all data chunks.",
    "526044": "Negative predictions? I guess he predicted time to LAST failure and did some math with ttf? ;)",
    "526048": "That makes sense, @khahuras . Thanks.\nIn the papers the also tried that route and had some success. I might be giving it a try, too ;)",
    "526050": "Well just a joke, but they are linearly the same anyway...",
    "526061": "&gt; Negative predictions? I guess he predicted time to LAST failure and did some math with ttf? ;)\n\nNo postprocessing or prepossessing of ttf here.  I don't know why there is some negative prediction.",
    "526084": "cpmpml  your predictions are up to 12 yes, but you have higher prediction for the short quakes. It's always a trade-off.",
    "526089": "Higher than what?  I don't get your point.  Predictions for short ttf peaks are lower than the one shared by Ilu elsewhere in this topic discussion for instance.\n\nMaybe you should show yours to make your point.",
    "526105": "Please see my oof results from a LGB model using only one feature:\n\nfor 3 folds (CV: 2.0932) we see all TTF ranges are about equal.\n![3fold](https://i.imgur.com/2EsOJDP.png)\n\nbut for 4 folds (CV: 2.0813) the TTF ranges already differ quite a bit. Thus, the splits are very important to the model in this competition.\n![4fold](https://i.imgur.com/oIoHv6p.png)\n\nI am not sure if I should trust the 4fold model in this case even though it has the better CV.\n\nDepending on your CV setup, @cpmpml , this might also be a reason for different TTF values right after the quakes. In my opinion standard KFold is pretty much the same way as they splitted train and test.",
    "526173": "",
    "526214": "ilu000 the range of values per period seems rather constant within each of your folds.  It is not the case for my model.",
    "526219": "It's totally normal to see some negative predictions using gradient boosting algorithms, even when target is positive definite. It should not happen with RandomForest though.",
    "526224": "Really? Then postprocessing of the prediction and capping the values at zero can only increase your LB score.",
    "526233": "I expect a very small improvement but, yes, can only make it better",
    "526261": "&gt; I expect a very small improvemen\n\nIndeed, let's say there is 1 prediction at -1.  If we clip it at 0, the improvement on test would be 1/2624= 0.00038",
    "526291": "pcjimmmy thanks for the clarification. I got his point refering back to @steubk's comment",
    "526298": "cpmp Oversampling? Your plot suggests over 12K points.",
    "526317": "Right, this one was with some data augmentation.",
    "526362": "cpmpml  Mine has equal heights between quakes, just like the others. If I'm not mistaken, the variable height you have in your oof is a result of quake-wise early stopping. I wonder how is that working for you in the LB. Intriguing indeed.",
    "526373": "Please forgive my ignorance, but what is an oof?  I gather that it's this specific type of graph that Ilu and CPMP have been using to compare their results, but I haven't had any luck googling it to get more context.  Thanks!",
    "526375": "aekoch95 `oof` stands for `out of fold` and includes the combined predictions of each validation set inside the cross validation folds.",
    "526474": "amjad85 it works on my cv setting.  Lb improvement is a by product.\n\nNot sure why early stopping is or is not relevant.  Difference between using it or not is tiny in general.",
    "526485": "Early stopping if you split out single EQs should not work as you leak the length and it does not work for me. Really surprised that it does for you.",
    "526497": "I did not say I use it.",
    "526499": "So, you did not use it ;)",
    "526508": "I didn’t say that either.  What I say is that using it or not does not make much difference in general.  I am surprised to see it makes a difference for some here.  I guess it has to do with the cv setting.  Also I don’t get @philippsinger comment on length:  which length is it?",
    "526514": "CPMP length of EQ",
    "526529": "How would this leak?",
    "526762": "CPMP in the graph with your oof predictions, do you use shuffled CV or unshuffled (or CV by EQ)? Would be interested to know...",
    "526823": "&gt; do you use shuffled CV or unshuffled (or CV by EQ)? \n\nYou'll know after competition end ;)",
    "526830": "Hehe 👍",
    "527210": "I've added classification of training data by type of earthquake (presence of intermediate release). For each class TTF can be predicted with much better score, than for mixed data.\nMy assumption was, that derivative of signal energy (or frequency) would be lower for long earthquakes (event escalates slowly giving chance for intermediate relieve).\nBut none of my features show more than 0.13 Pearson correlation with this classification, and models give about 34% error predicting correct class, which is not enough to use it to improve final score.\nSo I'm really curious to know what features could help reliably predict prolonged earthquake.",
    "532814": "ilu000 I assume your plots show CV predictions, but you are not using it \"as is\" for the final prediction - you either average it or retrain model on all the data. So these plots are a little bit misleading. 4fold model can be good for the ensemble after averaging. Also, assuming the CPMP plot shows oof, it will not be such accurate after averaging, it will rather show constant predictions (I mean predictions will be in a given range).",
    "536029": "I'm not able to predict anything over 10 seconds."
  },
  "source": "meta"
}