{
  "id": 94407,
  "title": "7th solution, CPMP view",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/94407",
  "author_name": "",
  "post_date": "2019-06-04T10:05:46.768029400Z",
  "votes": 111,
  "comment_count": 62,
  "views": 0,
  "content": "<p>First of all, I want to thank organizers for this interesting and tricky challenge, as well as my team mates <a href=\"/antoine\">@antoine</a> and <a href=\"/stecasasso\">@stecasasso</a> for their interesting views.  You can find <a href=\"/antoine\">@antoine</a>'s view <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94359#latest-542974\">here</a>, and <a href=\"/stecasasso\">@stecasasso</a> <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94408\">here</a>.  I also want to thank <a href=\"/mykper\">@mykper</a> for his great analysis of public test data.  </p>\n\n<p>Our solution is based on few pieces.  I'll start by what I did before teaming.</p>\n\n<p><strong>Cross Validation</strong></p>\n\n<p>We used 16 fold CV where each fold includes a time to failure (TTF) reset.  Lately we realized that a simple unshuffled  16 folds would be equivalent  to what we did.  Even with 16 folds, it was quickly clear that CV LB correlation was poor, and that public LB was not a reliable.  I decided to use a nested CV approach where one fold Ti is used as test data, and CV is run on the other 15 folds.  15 models are trained using the 15 fold cv and their predictions on Ti are averaged.  This mimics the submission process 16 times, and provides a reliable indicator of a model strength by averaging the mae of Ti folds.  Correlation with public LB was very good.</p>\n\n<p><strong>Data Sampling</strong></p>\n\n<p>I used 150k segments starting every 50k rows.  Every segment overlaps with 2 segments before it and 2 segments after it.  I tested more overlap, and less overlap, and this one was the best trade off between running time and result accuracy.  A potential issue could be that segments at fold boundary overlap.  I tried removing those overlap but result accuracy decreased a bit, hence I kept all segments.</p>\n\n<p><strong>Outliers</strong></p>\n\n<p>TTF is not reset at acoustic data peaks which makes the problem difficult.  I decided to treat the points between acoustic peak and the following TTF reset as outlier, and trained a binary model to predict outliers.  I also trained model where the target is TTF outside outliers, and TTF is reset at acoustic peaks.  </p>\n\n<p>Final prediction is 0.2 if binary prediction is above a threshold (0.45 to 0.7 depending on the runs), and an average of models trained on original TTF and modified TTF as above elsewhere.</p>\n\n<p><strong>Features</strong></p>\n\n<p>I used 7 MFCC from librosa.  To make it work I pretended than data was sampled at 40kHz.  Not doing so means the useful info is way higher, around the 20th MFCC.  I also used 4 features based on std of various signal quantiles.  Last, I used a binary feature indicating high variance segments that correspond to acoustic peaks very accurately.</p>\n\n<p><strong>Models</strong></p>\n\n<p>Mostly lgb with conservative settings, like 7 leaves only.  But also general additive models using pygam.  I also trained a knn model which was surprisingly good.  In hindsight, gam was better than lgb which was better than knn, the gam model would get a gold medal alone at 2.348.  For lgb I tried mse, gamma and huber objective.  With the above I got to 1.288 on public LB at which point I teamed.</p>\n\n<p>After teaming, progress was in many areas, but here are the three most important ones.  </p>\n\n<p><strong>Ensembling</strong></p>\n\n<p>First, model diversity as my team mates were using different data sets.  By averaging our top public LB <a href=\"/areveillon\">@areveillon</a> had a 1.286 public LB) we got to 1.275 and took the lead.  TMy team mates adapted their models to the binary vs regression models.    We also reused some features from each other dataset, which added to model diversity.  We then worked on stacking, ending with a stack based on a lgb, a gam, and a knn from me, a lgb from @Antoine and a NN from <a href=\"/stecasasso\">@stecasasso</a> .  lgb was used for the second level model.  We validated stacking using our nested CV which gives us some confidence that stacking was indeed working.  Private LB confirms that stacking was improving over base models.</p>\n\n<p><strong>Train / Test Difference</strong></p>\n\n<p>I must admit that I hate LB probing, and avoid it as plague.  But here, using pictures from academic papers could lead to a good estimate of test data, and it was the way to go given the significant difference between train and test data.  @Antoine had done an estimate before teaming and was using it already.  Few days before end he convinced me that we should really base our final sub on it and I measured precisely length of EQ cycles in that picture and could estimate their duration by running a linear regression on train data.  From it I estimated its mean to be 6.35 which is quite close to the actual 6.32.</p>\n\n<p>From this we can estimate density function of TTF for train and test.  We then used sample weights to map the train distribution to the test distribution.  I see that top 2 teams (at least) rather resampled train, and this may explain why they are ahead.  There are two ways to compute weights.  I tried to base them on density function, i.e. weights only depend on the TTF, while @Antoine was pushing for weights per eq cycle.  In hindsight he was right.  Given we could not agree before competition end we decided to produce two submissions, one that optimizes my weights while the second one was optimizing <a href=\"/areveillon\">@areveillon</a> 's weights.   <a href=\"/stecasasso\">@stecasasso</a>  then had the idea of using the weights not only for evaluating our models, but also for training models.  In the end, submissions optimizing <a href=\"/areveillon\">@areveillon</a> 's weights were the best ones.</p>\n\n<p><strong>Post-processing</strong></p>\n\n<p><a href=\"/stecasasso\">@stecasasso</a> found that we could set all high variance segments TTF to 0.31.  @Antoine looked at how to best combine binary models to fix segments TTF at 0.2.  Applying these as postprocessing improved submissions.  It also moved our best public LB from 2.75 to 2.45.</p>\n\n<p><strong>What did not work</strong></p>\n\n<p>I could not get NN models to work well enough to be useful in our stack.  <a href=\"/stecasasso\">@stecasasso</a> managed to get one, but it was the weakest model on the stack.  lgb and gam are better.  I think NNs could be useful on processed data, either sftt or Hilbert envelope, but I did not have the time to try seriously.</p>\n\n<p><strong>Answers</strong></p>\n\n<p>I was asked about what was the 'no magic' feature. It was the 4th MFCC I was using.  I also was asked how to overfit public LB.  Well, first answer is that it was easy given how many people had way worse results on private Lb than public LB.  But in our case it was just by capping test prediction to 9.6 then to 9, given <a href=\"/mykper\">@mykper</a> had shown public test TTF was below 10. This led us to 1.236 from 1.245.  I'm curious about how kopeyka team overfitted public test so effectively.  Last, I was asked about how I survive shakeups like this one or in previous competitions: it is because I only trust my CV and do not rely on LB probing.</p>",
  "messages": [
    {
      "id": "543068",
      "postDate": "06/04/2019 10:05:46",
      "content": "<p>First of all, I want to thank organizers for this interesting and tricky challenge, as well as my team mates <a href=\"/antoine\">@antoine</a> and <a href=\"/stecasasso\">@stecasasso</a> for their interesting views.  You can find <a href=\"/antoine\">@antoine</a>'s view <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94359#latest-542974\">here</a>, and <a href=\"/stecasasso\">@stecasasso</a> <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94408\">here</a>.  I also want to thank <a href=\"/mykper\">@mykper</a> for his great analysis of public test data.  </p>\n\n<p>Our solution is based on few pieces.  I'll start by what I did before teaming.</p>\n\n<p><strong>Cross Validation</strong></p>\n\n<p>We used 16 fold CV where each fold includes a time to failure (TTF) reset.  Lately we realized that a simple unshuffled  16 folds would be equivalent  to what we did.  Even with 16 folds, it was quickly clear that CV LB correlation was poor, and that public LB was not a reliable.  I decided to use a nested CV approach where one fold Ti is used as test data, and CV is run on the other 15 folds.  15 models are trained using the 15 fold cv and their predictions on Ti are averaged.  This mimics the submission process 16 times, and provides a reliable indicator of a model strength by averaging the mae of Ti folds.  Correlation with public LB was very good.</p>\n\n<p><strong>Data Sampling</strong></p>\n\n<p>I used 150k segments starting every 50k rows.  Every segment overlaps with 2 segments before it and 2 segments after it.  I tested more overlap, and less overlap, and this one was the best trade off between running time and result accuracy.  A potential issue could be that segments at fold boundary overlap.  I tried removing those overlap but result accuracy decreased a bit, hence I kept all segments.</p>\n\n<p><strong>Outliers</strong></p>\n\n<p>TTF is not reset at acoustic data peaks which makes the problem difficult.  I decided to treat the points between acoustic peak and the following TTF reset as outlier, and trained a binary model to predict outliers.  I also trained model where the target is TTF outside outliers, and TTF is reset at acoustic peaks.  </p>\n\n<p>Final prediction is 0.2 if binary prediction is above a threshold (0.45 to 0.7 depending on the runs), and an average of models trained on original TTF and modified TTF as above elsewhere.</p>\n\n<p><strong>Features</strong></p>\n\n<p>I used 7 MFCC from librosa.  To make it work I pretended than data was sampled at 40kHz.  Not doing so means the useful info is way higher, around the 20th MFCC.  I also used 4 features based on std of various signal quantiles.  Last, I used a binary feature indicating high variance segments that correspond to acoustic peaks very accurately.</p>\n\n<p><strong>Models</strong></p>\n\n<p>Mostly lgb with conservative settings, like 7 leaves only.  But also general additive models using pygam.  I also trained a knn model which was surprisingly good.  In hindsight, gam was better than lgb which was better than knn, the gam model would get a gold medal alone at 2.348.  For lgb I tried mse, gamma and huber objective.  With the above I got to 1.288 on public LB at which point I teamed.</p>\n\n<p>After teaming, progress was in many areas, but here are the three most important ones.  </p>\n\n<p><strong>Ensembling</strong></p>\n\n<p>First, model diversity as my team mates were using different data sets.  By averaging our top public LB <a href=\"/areveillon\">@areveillon</a> had a 1.286 public LB) we got to 1.275 and took the lead.  TMy team mates adapted their models to the binary vs regression models.    We also reused some features from each other dataset, which added to model diversity.  We then worked on stacking, ending with a stack based on a lgb, a gam, and a knn from me, a lgb from @Antoine and a NN from <a href=\"/stecasasso\">@stecasasso</a> .  lgb was used for the second level model.  We validated stacking using our nested CV which gives us some confidence that stacking was indeed working.  Private LB confirms that stacking was improving over base models.</p>\n\n<p><strong>Train / Test Difference</strong></p>\n\n<p>I must admit that I hate LB probing, and avoid it as plague.  But here, using pictures from academic papers could lead to a good estimate of test data, and it was the way to go given the significant difference between train and test data.  @Antoine had done an estimate before teaming and was using it already.  Few days before end he convinced me that we should really base our final sub on it and I measured precisely length of EQ cycles in that picture and could estimate their duration by running a linear regression on train data.  From it I estimated its mean to be 6.35 which is quite close to the actual 6.32.</p>\n\n<p>From this we can estimate density function of TTF for train and test.  We then used sample weights to map the train distribution to the test distribution.  I see that top 2 teams (at least) rather resampled train, and this may explain why they are ahead.  There are two ways to compute weights.  I tried to base them on density function, i.e. weights only depend on the TTF, while @Antoine was pushing for weights per eq cycle.  In hindsight he was right.  Given we could not agree before competition end we decided to produce two submissions, one that optimizes my weights while the second one was optimizing <a href=\"/areveillon\">@areveillon</a> 's weights.   <a href=\"/stecasasso\">@stecasasso</a>  then had the idea of using the weights not only for evaluating our models, but also for training models.  In the end, submissions optimizing <a href=\"/areveillon\">@areveillon</a> 's weights were the best ones.</p>\n\n<p><strong>Post-processing</strong></p>\n\n<p><a href=\"/stecasasso\">@stecasasso</a> found that we could set all high variance segments TTF to 0.31.  @Antoine looked at how to best combine binary models to fix segments TTF at 0.2.  Applying these as postprocessing improved submissions.  It also moved our best public LB from 2.75 to 2.45.</p>\n\n<p><strong>What did not work</strong></p>\n\n<p>I could not get NN models to work well enough to be useful in our stack.  <a href=\"/stecasasso\">@stecasasso</a> managed to get one, but it was the weakest model on the stack.  lgb and gam are better.  I think NNs could be useful on processed data, either sftt or Hilbert envelope, but I did not have the time to try seriously.</p>\n\n<p><strong>Answers</strong></p>\n\n<p>I was asked about what was the 'no magic' feature. It was the 4th MFCC I was using.  I also was asked how to overfit public LB.  Well, first answer is that it was easy given how many people had way worse results on private Lb than public LB.  But in our case it was just by capping test prediction to 9.6 then to 9, given <a href=\"/mykper\">@mykper</a> had shown public test TTF was below 10. This led us to 1.236 from 1.245.  I'm curious about how kopeyka team overfitted public test so effectively.  Last, I was asked about how I survive shakeups like this one or in previous competitions: it is because I only trust my CV and do not rely on LB probing.</p>",
      "rawMarkdown": "First of all, I want to thank organizers for this interesting and tricky challenge, as well as my team mates @antoine and @stecasasso for their interesting views.  You can find @antoine's view [here](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94359#latest-542974), and @stecasasso [here](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94408).  I also want to thank @mykper for his great analysis of public test data.  \n\nOur solution is based on few pieces.  I'll start by what I did before teaming.\n\n**Cross Validation**\n\nWe used 16 fold CV where each fold includes a time to failure (TTF) reset.  Lately we realized that a simple unshuffled  16 folds would be equivalent  to what we did.  Even with 16 folds, it was quickly clear that CV LB correlation was poor, and that public LB was not a reliable.  I decided to use a nested CV approach where one fold Ti is used as test data, and CV is run on the other 15 folds.  15 models are trained using the 15 fold cv and their predictions on Ti are averaged.  This mimics the submission process 16 times, and provides a reliable indicator of a model strength by averaging the mae of Ti folds.  Correlation with public LB was very good.\n\n**Data Sampling**\n\nI used 150k segments starting every 50k rows.  Every segment overlaps with 2 segments before it and 2 segments after it.  I tested more overlap, and less overlap, and this one was the best trade off between running time and result accuracy.  A potential issue could be that segments at fold boundary overlap.  I tried removing those overlap but result accuracy decreased a bit, hence I kept all segments.\n\n**Outliers**\n\nTTF is not reset at acoustic data peaks which makes the problem difficult.  I decided to treat the points between acoustic peak and the following TTF reset as outlier, and trained a binary model to predict outliers.  I also trained model where the target is TTF outside outliers, and TTF is reset at acoustic peaks.  \n\nFinal prediction is 0.2 if binary prediction is above a threshold (0.45 to 0.7 depending on the runs), and an average of models trained on original TTF and modified TTF as above elsewhere.\n\n**Features**\n\nI used 7 MFCC from librosa.  To make it work I pretended than data was sampled at 40kHz.  Not doing so means the useful info is way higher, around the 20th MFCC.  I also used 4 features based on std of various signal quantiles.  Last, I used a binary feature indicating high variance segments that correspond to acoustic peaks very accurately.\n\n**Models**\n\nMostly lgb with conservative settings, like 7 leaves only.  But also general additive models using pygam.  I also trained a knn model which was surprisingly good.  In hindsight, gam was better than lgb which was better than knn, the gam model would get a gold medal alone at 2.348.  For lgb I tried mse, gamma and huber objective.  With the above I got to 1.288 on public LB at which point I teamed.\n\nAfter teaming, progress was in many areas, but here are the three most important ones.  \n\n**Ensembling**\n\nFirst, model diversity as my team mates were using different data sets.  By averaging our top public LB @areveillon had a 1.286 public LB) we got to 1.275 and took the lead.  TMy team mates adapted their models to the binary vs regression models.    We also reused some features from each other dataset, which added to model diversity.  We then worked on stacking, ending with a stack based on a lgb, a gam, and a knn from me, a lgb from @Antoine and a NN from @stecasasso .  lgb was used for the second level model.  We validated stacking using our nested CV which gives us some confidence that stacking was indeed working.  Private LB confirms that stacking was improving over base models.\n\n**Train / Test Difference**\n\nI must admit that I hate LB probing, and avoid it as plague.  But here, using pictures from academic papers could lead to a good estimate of test data, and it was the way to go given the significant difference between train and test data.  @Antoine had done an estimate before teaming and was using it already.  Few days before end he convinced me that we should really base our final sub on it and I measured precisely length of EQ cycles in that picture and could estimate their duration by running a linear regression on train data.  From it I estimated its mean to be 6.35 which is quite close to the actual 6.32.\n\nFrom this we can estimate density function of TTF for train and test.  We then used sample weights to map the train distribution to the test distribution.  I see that top 2 teams (at least) rather resampled train, and this may explain why they are ahead.  There are two ways to compute weights.  I tried to base them on density function, i.e. weights only depend on the TTF, while @Antoine was pushing for weights per eq cycle.  In hindsight he was right.  Given we could not agree before competition end we decided to produce two submissions, one that optimizes my weights while the second one was optimizing @areveillon 's weights.   @stecasasso  then had the idea of using the weights not only for evaluating our models, but also for training models.  In the end, submissions optimizing @areveillon 's weights were the best ones.\n\n**Post-processing**\n\n@stecasasso found that we could set all high variance segments TTF to 0.31.  @Antoine looked at how to best combine binary models to fix segments TTF at 0.2.  Applying these as postprocessing improved submissions.  It also moved our best public LB from 2.75 to 2.45.\n\n**What did not work**\n\nI could not get NN models to work well enough to be useful in our stack.  @stecasasso managed to get one, but it was the weakest model on the stack.  lgb and gam are better.  I think NNs could be useful on processed data, either sftt or Hilbert envelope, but I did not have the time to try seriously.\n\n**Answers**\n\nI was asked about what was the 'no magic' feature. It was the 4th MFCC I was using.  I also was asked how to overfit public LB.  Well, first answer is that it was easy given how many people had way worse results on private Lb than public LB.  But in our case it was just by capping test prediction to 9.6 then to 9, given @mykper had shown public test TTF was below 10. This led us to 1.236 from 1.245.  I'm curious about how kopeyka team overfitted public test so effectively.  Last, I was asked about how I survive shakeups like this one or in previous competitions: it is because I only trust my CV and do not rely on LB probing.",
      "votes": null
    },
    {
      "id": "543085",
      "postDate": "06/04/2019 10:17:51",
      "content": "<p>Thank you so much <a href=\"/cpmpml\">@cpmpml</a> this is the best part of the competition, be able to learn from such great competitors like you, <a href=\"/stecasasso\">@stecasasso</a> and Antoine. Great solution, I'll try to do this by myself ;)</p>\n\n<h1>👍</h1>",
      "rawMarkdown": "Thank you so much @cpmpml this is the best part of the competition, be able to learn from such great competitors like you, @stecasasso and Antoine. Great solution, I'll try to do this by myself ;)\n# 👍",
      "votes": null
    },
    {
      "id": "543091",
      "postDate": "06/04/2019 10:21:34",
      "content": "<p>I was waiting for your solution. Thanks for sharing.</p>",
      "rawMarkdown": "I was waiting for your solution. Thanks for sharing.",
      "votes": null
    },
    {
      "id": "543102",
      "postDate": "06/04/2019 10:28:20",
      "content": "<p>It seems you are the only one able to predict ttf at <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91125#525889\">beginning of quake</a>.  Which feature(s) let you do that?</p>",
      "rawMarkdown": "It seems you are the only one able to predict ttf at [beginning of quake](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91125#525889).  Which feature(s) let you do that?",
      "votes": null
    },
    {
      "id": "543115",
      "postDate": "06/04/2019 10:38:17",
      "content": "<p>Congratulations! Interesting and complete explanations, as usual, some of them will keep me thinking for a while, thanks for sharing :-)</p>",
      "rawMarkdown": "Congratulations! Interesting and complete explanations, as usual, some of them will keep me thinking for a while, thanks for sharing :-)",
      "votes": null
    },
    {
      "id": "543141",
      "postDate": "06/04/2019 10:50:26",
      "content": "<p>The features I describe above.</p>",
      "rawMarkdown": "The features I describe above.",
      "votes": null
    },
    {
      "id": "543152",
      "postDate": "06/04/2019 10:58:47",
      "content": "<p>Congrats! </p>\n\n<p>&gt; I also trained model where the target is TTF outside outliers, and TTF is reset at acoustic peaks.</p>\n\n<p>Is this the modified TTF model? You modified the TTF such that outliers' TTF=0 and non-outliers' TTF is unchanged?</p>\n\n<p>&gt; Final prediction is 0.2 if binary prediction is above a threshold (0.45 to 0.7 depending on the runs), and an average of models trained on original TTF and modified TTF as above.</p>\n\n<p>I don't understand your method of setting the threshold, it is different for each fold? And it is selected manually? The weights of model average is decided by some lgb model?</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "Congrats! \n\n&gt; I also trained model where the target is TTF outside outliers, and TTF is reset at acoustic peaks.\n\nIs this the modified TTF model? You modified the TTF such that outliers' TTF=0 and non-outliers' TTF is unchanged?\n\n&gt; Final prediction is 0.2 if binary prediction is above a threshold (0.45 to 0.7 depending on the runs), and an average of models trained on original TTF and modified TTF as above.\n\nI don't understand your method of setting the threshold, it is different for each fold? And it is selected manually? The weights of model average is decided by some lgb model?\n\nThanks!",
      "votes": null
    },
    {
      "id": "543161",
      "postDate": "06/04/2019 11:04:06",
      "content": "<p>We have three targets:</p>\n\n<ol>\n<li>binary one for outliers</li>\n<li>original ttf</li>\n<li>modified ttf: same values as ttf outside outliers, and ttf + next ttf peak for outliers. </li>\n</ol>\n\n<p>The third target is reset at acoustic peak, then decreases at the same rate as ttf.  The reset  height is such that the modified target meets original ttf value outside outliers.</p>",
      "rawMarkdown": "We have three targets:\n\n1. binary one for outliers\n2. original ttf\n3. modified ttf: same values as ttf outside outliers, and ttf + next ttf peak for outliers. \n\nThe third target is reset at acoustic peak, then decreases at the same rate as ttf.  The reset  height is such that the modified target meets original ttf value outside outliers.",
      "votes": null
    },
    {
      "id": "543179",
      "postDate": "06/04/2019 11:16:38",
      "content": "<p>Thanks for the neat and comprehensive write-up! As usual ;)\nAgain, it was a pleasure to team up with you!</p>",
      "rawMarkdown": "Thanks for the neat and comprehensive write-up! As usual ;)\nAgain, it was a pleasure to team up with you!",
      "votes": null
    },
    {
      "id": "543190",
      "postDate": "06/04/2019 11:20:02",
      "content": "<p>By the way, is there any plan of open-sourcing your code (or part of). The method is really interesting and worth more time studying. But since it is quite complicated, so I guess the code could help us understand better. Thanks!</p>",
      "rawMarkdown": "By the way, is there any plan of open-sourcing your code (or part of). The method is really interesting and worth more time studying. But since it is quite complicated, so I guess the code could help us understand better. Thanks!",
      "votes": null
    },
    {
      "id": "543199",
      "postDate": "06/04/2019 11:22:40",
      "content": "<p>The dataprep is scattered in several places, and cleaning it would be a time consuming task.  I am glad to not be in prize zone for that ;)  </p>",
      "rawMarkdown": "The dataprep is scattered in several places, and cleaning it would be a time consuming task.  I am glad to not be in prize zone for that ;)",
      "votes": null
    },
    {
      "id": "543200",
      "postDate": "06/04/2019 11:22:53",
      "content": "<p>Thanks! Your modified TTF makes sense. If I remember correctly, peaks are start of failure, and the end of failure is only indicated by stress values.</p>",
      "rawMarkdown": "Thanks! Your modified TTF makes sense. If I remember correctly, peaks are start of failure, and the end of failure is only indicated by stress values.",
      "votes": null
    },
    {
      "id": "543201",
      "postDate": "06/04/2019 11:23:24",
      "content": "<p>Thanks, that is of course understandable.</p>",
      "rawMarkdown": "Thanks, that is of course understandable.",
      "votes": null
    },
    {
      "id": "543225",
      "postDate": "06/04/2019 11:50:42",
      "content": "<p>Lovely - if you write an academic paper about it just state your score as 1.2 and you are done! You will have to delete your validation stuff though;)</p>",
      "rawMarkdown": "Lovely - if you write an academic paper about it just state your score as 1.2 and you are done! You will have to delete your validation stuff though;)",
      "votes": null
    },
    {
      "id": "543231",
      "postDate": "06/04/2019 11:57:18",
      "content": "<p>Thanks for sharing CPMP!</p>",
      "rawMarkdown": "Thanks for sharing CPMP!",
      "votes": null
    },
    {
      "id": "543299",
      "postDate": "06/04/2019 13:00:55",
      "content": "<p>I used nested CV too, but only 4 Ti (and oof CV on the rest 12). It's my only small satisfaction that direction I took was at least a little bit similar to yours and Giba's. \nI was trying to learn as much as possible from you, thank you again for all the sharings.</p>",
      "rawMarkdown": "I used nested CV too, but only 4 Ti (and oof CV on the rest 12). It's my only small satisfaction that direction I took was at least a little bit similar to yours and Giba's. \nI was trying to learn as much as possible from you, thank you again for all the sharings.",
      "votes": null
    },
    {
      "id": "543337",
      "postDate": "06/04/2019 13:34:01",
      "content": "<p>Thanks ! My takeaway are those 3 brilliant ideas from <a href=\"/cpmpml\">@cpmpml</a> : \n-Nested CV Setup\n-Strong usage of sample Weight based on ttf. Some team discarded some EQ from train data which also was very effective. But I think using sample weight was a better option here.\n-Binary classification model combined with 2 others models, this really gave us a strong boost, our model was able to predict correctly a majority of those point where we had the biggest error before. (first team used this as well)</p>",
      "rawMarkdown": "Thanks ! My takeaway are those 3 brilliant ideas from @cpmpml : \n-Nested CV Setup\n-Strong usage of sample Weight based on ttf. Some team discarded some EQ from train data which also was very effective. But I think using sample weight was a better option here.\n-Binary classification model combined with 2 others models, this really gave us a strong boost, our model was able to predict correctly a majority of those point where we had the biggest error before. (first team used this as well)",
      "votes": null
    },
    {
      "id": "543343",
      "postDate": "06/04/2019 13:38:53",
      "content": "<p>You also had brilliant ideas, including the idea of weighting samples by EQ cycles!</p>\n\n<p>To be fair, the binary model idea was used in previous competitions to deal with outliers, in TGS and in ELO at least.  I didn't invented it ;)</p>",
      "rawMarkdown": "You also had brilliant ideas, including the idea of weighting samples by EQ cycles!\n\nTo be fair, the binary model idea was used in previous competitions to deal with outliers, in TGS and in ELO at least.  I didn't invented it ;)",
      "votes": null
    },
    {
      "id": "543357",
      "postDate": "06/04/2019 13:44:24",
      "content": "<p>Congrats again and thanks for sharing. The 40kHz trick looks like magic. Just kidding ;)</p>",
      "rawMarkdown": "Congrats again and thanks for sharing. The 40kHz trick looks like magic. Just kidding ;)",
      "votes": null
    },
    {
      "id": "543362",
      "postDate": "06/04/2019 13:51:15",
      "content": "<p>Thanks, congrats to you as well!  I'm sure this reminds you of recent shakeup in malware ;)</p>",
      "rawMarkdown": "Thanks, congrats to you as well!  I'm sure this reminds you of recent shakeup in malware ;)",
      "votes": null
    },
    {
      "id": "543366",
      "postDate": "06/04/2019 13:54:05",
      "content": "<p>Thank you for sharing! This is very interesting!</p>",
      "rawMarkdown": "Thank you for sharing! This is very interesting!",
      "votes": null
    },
    {
      "id": "543371",
      "postDate": "06/04/2019 13:56:48",
      "content": "<p>Congrats guys! You had the most consistent LB results.</p>",
      "rawMarkdown": "Congrats guys! You had the most consistent LB results.",
      "votes": null
    },
    {
      "id": "543414",
      "postDate": "06/04/2019 14:16:26",
      "content": "<p>Thank you for your sharing!</p>",
      "rawMarkdown": "Thank you for your sharing!",
      "votes": null
    },
    {
      "id": "543435",
      "postDate": "06/04/2019 14:26:50",
      "content": "<p>Thanks for sharing.  A numbers of solid ideas without overly gaming the system.   </p>",
      "rawMarkdown": "Thanks for sharing.  A numbers of solid ideas without overly gaming the system.",
      "votes": null
    },
    {
      "id": "543492",
      "postDate": "06/04/2019 15:02:01",
      "content": "<p>Thanks for sharing !</p>",
      "rawMarkdown": "Thanks for sharing !",
      "votes": null
    },
    {
      "id": "543496",
      "postDate": "06/04/2019 15:08:15",
      "content": "<p>Congrats! Thanks for sharing your solution!</p>",
      "rawMarkdown": "Congrats! Thanks for sharing your solution!",
      "votes": null
    },
    {
      "id": "543553",
      "postDate": "06/04/2019 15:37:20",
      "content": "<p>Yes, this reminds me Malware. Previous knowledge of Private testset survives the shakeup ;)</p>",
      "rawMarkdown": "Yes, this reminds me Malware. Previous knowledge of Private testset survives the shakeup ;)",
      "votes": null
    },
    {
      "id": "543581",
      "postDate": "06/04/2019 15:52:38",
      "content": "<p>cool, Congrats.</p>",
      "rawMarkdown": "cool, Congrats.",
      "votes": null
    },
    {
      "id": "543617",
      "postDate": "06/04/2019 16:27:07",
      "content": "<p>Thanks for sharing, I used a similar nested CV approach and the same data sampling, but did not find enough time to do proper FE. Congrats on the lead and (again) Gold medal :) </p>",
      "rawMarkdown": "Thanks for sharing, I used a similar nested CV approach and the same data sampling, but did not find enough time to do proper FE. Congrats on the lead and (again) Gold medal :)",
      "votes": null
    },
    {
      "id": "543628",
      "postDate": "06/04/2019 16:38:31",
      "content": "<p>Thanks for sharing. Interesting approach to treat outliers. Congrats!</p>",
      "rawMarkdown": "Thanks for sharing. Interesting approach to treat outliers. Congrats!",
      "votes": null
    },
    {
      "id": "543731",
      "postDate": "06/04/2019 18:22:27",
      "content": "<p>i believe you tagged the wrong person (Antoine is <a href=\"/areveillon\">@areveillon</a>) <a href=\"/cpmpml\">@cpmpml</a> ...\nThanks for the write-up!</p>",
      "rawMarkdown": "i believe you tagged the wrong person (Antoine is @areveillon) @cpmpml ...\nThanks for the write-up!",
      "votes": null
    },
    {
      "id": "543887",
      "postDate": "06/04/2019 23:06:23",
      "content": "<p>Thank you for sharing. I have one question, how did you proceed feature selection to obtain specific frequency of MFCC (4th was the best?). Did you try increasing the feature one frequency by one to see the difference??</p>",
      "rawMarkdown": "Thank you for sharing. I have one question, how did you proceed feature selection to obtain specific frequency of MFCC (4th was the best?). Did you try increasing the feature one frequency by one to see the difference??",
      "votes": null
    },
    {
      "id": "543919",
      "postDate": "06/04/2019 23:57:47",
      "content": "<p>Congrats! I can enjoy this competition with your discussions, thank you.</p>",
      "rawMarkdown": "Congrats! I can enjoy this competition with your discussions, thank you.",
      "votes": null
    },
    {
      "id": "543983",
      "postDate": "06/05/2019 02:18:13",
      "content": "<p>Oops will fix that asap</p>",
      "rawMarkdown": "Oops will fix that asap",
      "votes": null
    },
    {
      "id": "543984",
      "postDate": "06/05/2019 02:19:27",
      "content": "<p>I started from librosa default value then tuned it via cross validation.</p>",
      "rawMarkdown": "I started from librosa default value then tuned it via cross validation.",
      "votes": null
    },
    {
      "id": "544029",
      "postDate": "06/05/2019 03:54:56",
      "content": "<p>Congrats!\nCould you explain more about the binary model here?</p>",
      "rawMarkdown": "Congrats!\nCould you explain more about the binary model here?",
      "votes": null
    },
    {
      "id": "544032",
      "postDate": "06/05/2019 03:57:53",
      "content": "<p>Congratulations! Thanks for sharing :) </p>",
      "rawMarkdown": "Congratulations! Thanks for sharing :)",
      "votes": null
    },
    {
      "id": "544045",
      "postDate": "06/05/2019 04:07:59",
      "content": "<p>Congratulations and thanks for sharing! I always learn a lot from you, grand master!!</p>",
      "rawMarkdown": "Congratulations and thanks for sharing! I always learn a lot from you, grand master!!",
      "votes": null
    },
    {
      "id": "544169",
      "postDate": "06/05/2019 08:26:15",
      "content": "<p>I mean how did you find \"4th\" frequency is the best?</p>",
      "rawMarkdown": "I mean how did you find \"4th\" frequency is the best?",
      "votes": null
    },
    {
      "id": "544176",
      "postDate": "06/05/2019 08:33:14",
      "content": "<p>By impact on CV score.</p>",
      "rawMarkdown": "By impact on CV score.",
      "votes": null
    },
    {
      "id": "544180",
      "postDate": "06/05/2019 08:41:12",
      "content": "<p>It is lgb trained with the same features, only the target changes.  Target is 1 when the segment ends between an acoustic peak and the following TTF reset</p>",
      "rawMarkdown": "It is lgb trained with the same features, only the target changes.  Target is 1 when the segment ends between an acoustic peak and the following TTF reset",
      "votes": null
    },
    {
      "id": "544435",
      "postDate": "06/05/2019 14:05:47",
      "content": "<p>I see, so you really run a lot of experiments to determine best features. Thank you for the answer.</p>",
      "rawMarkdown": "I see, so you really run a lot of experiments to determine best features. Thank you for the answer.",
      "votes": null
    },
    {
      "id": "544443",
      "postDate": "06/05/2019 14:13:45",
      "content": "<p>Yes, a lot, and having a small set of features mean I can run experiments very quickly ;)</p>",
      "rawMarkdown": "Yes, a lot, and having a small set of features mean I can run experiments very quickly ;)",
      "votes": null
    },
    {
      "id": "544481",
      "postDate": "06/05/2019 14:56:17",
      "content": "<p><a href=\"/cpmpml\">@cpmpml</a> I'm curious about how your model performed in the long cycles. Is it possible that you show us a plot of the oof vs ttf? Thanks</p>",
      "rawMarkdown": "cpmpml I'm curious about how your model performed in the long cycles. Is it possible that you show us a plot of the oof vs ttf? Thanks",
      "votes": null
    },
    {
      "id": "544504",
      "postDate": "06/05/2019 15:24:17",
      "content": "<p>Here it is. Compared to yours we are lower, and probably more accurate on small ttf values.</p>\n\n<p>We have other models that look closer to yours, great for high ttf, but we blend them and I show the final result.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/544504/13413/oof.png\" alt=\"oof\"></p>\n\n<p>The spikes to 0 are due to our binary model fixing.</p>",
      "rawMarkdown": "Here it is. Compared to yours we are lower, and probably more accurate on small ttf values.\n\nWe have other models that look closer to yours, great for high ttf, but we blend them and I show the final result.\n\n![oof](https://storage.googleapis.com/kaggle-forum-message-attachments/544504/13413/oof.png)\n\nThe spikes to 0 are due to our binary model fixing.",
      "votes": null
    },
    {
      "id": "544539",
      "postDate": "06/05/2019 16:16:44",
      "content": "<p>Wow... the prediction for low ttf looks awesome! Thanks for sharing. Very good job.</p>",
      "rawMarkdown": "Wow... the prediction for low ttf looks awesome! Thanks for sharing. Very good job.",
      "votes": null
    },
    {
      "id": "544635",
      "postDate": "06/05/2019 18:12:34",
      "content": "<p>Between you and us we would have killed this competition maybe ;)  But a solo gold is invaluable these days!</p>",
      "rawMarkdown": "Between you and us we would have killed this competition maybe ;)  But a solo gold is invaluable these days!",
      "votes": null
    },
    {
      "id": "544660",
      "postDate": "06/05/2019 18:35:09",
      "content": "<p>are you posting your solution overview? :)</p>",
      "rawMarkdown": "are you posting your solution overview? :)",
      "votes": null
    },
    {
      "id": "544777",
      "postDate": "06/05/2019 22:51:33",
      "content": "<p>Thank you for sharing. Can you please offer your intuition on why did you pick <code>pygam</code> as a model? This is pretty new to me. Thanks!</p>",
      "rawMarkdown": "Thank you for sharing. Can you please offer your intuition on why did you pick `pygam` as a model? This is pretty new to me. Thanks!",
      "votes": null
    },
    {
      "id": "545970",
      "postDate": "06/06/2019 05:07:35",
      "content": "<p>I had a hunch but it is hard to say why.  I thought that feature interaction would not play a big role here.  Gam are just the addition of monovariate models.  </p>",
      "rawMarkdown": "I had a hunch but it is hard to say why.  I thought that feature interaction would not play a big role here.  Gam are just the addition of monovariate models.",
      "votes": null
    },
    {
      "id": "546009",
      "postDate": "06/06/2019 06:18:04",
      "content": "<p>Hey <a href=\"/cpmpml\">@cpmpml</a> \nThanks for sharing your approach to this competition and congrats with staying in top 10!</p>\n\n<p>Could you give me some rope to understand the Outliers part of your solution? Let's see how I understood it partly and then questions.</p>\n\n<p>You train a binary classifier to detect TTF outliers which are exactly how you defined them. You then use this prediction as a feature in training your regression model. </p>\n\n<p>Questions:</p>\n\n<blockquote>\n  <p>Final prediction is 0.2 if binary prediction is above a threshold (0.45 to 0.7 depending on the runs), and an average of models trained on original TTF and modified TTF as above elsewhere.</p>\n</blockquote>\n\n<p>Don't understand this sentence. What do you mean by final prediction being 0.2? Can you elaborate please?</p>",
      "rawMarkdown": "Hey @cpmpml \nThanks for sharing your approach to this competition and congrats with staying in top 10!\n\nCould you give me some rope to understand the Outliers part of your solution? Let's see how I understood it partly and then questions.\n\nYou train a binary classifier to detect TTF outliers which are exactly how you defined them. You then use this prediction as a feature in training your regression model. \n\nQuestions:\n&gt;Final prediction is 0.2 if binary prediction is above a threshold (0.45 to 0.7 depending on the runs), and an average of models trained on original TTF and modified TTF as above elsewhere.\n\nDon't understand this sentence. What do you mean by final prediction being 0.2? Can you elaborate please?",
      "votes": null
    },
    {
      "id": "546328",
      "postDate": "06/06/2019 13:26:59",
      "content": "<p>If the binary mode prediction is above, say 0.5, then we fix the prediction to be 0.2.  The threshold is tuned by cv, and it depends on the binary model.</p>",
      "rawMarkdown": "If the binary mode prediction is above, say 0.5, then we fix the prediction to be 0.2.  The threshold is tuned by cv, and it depends on the binary model.",
      "votes": null
    },
    {
      "id": "546370",
      "postDate": "06/06/2019 14:19:43",
      "content": "<p>Thank you! Learned a new tool from you.</p>",
      "rawMarkdown": "Thank you! Learned a new tool from you.",
      "votes": null
    },
    {
      "id": "546472",
      "postDate": "06/06/2019 15:55:44",
      "content": "<p>Thanks! For me it is a new tool too.</p>",
      "rawMarkdown": "Thanks! For me it is a new tool too.",
      "votes": null
    },
    {
      "id": "546578",
      "postDate": "06/06/2019 17:38:53",
      "content": "<p>Hi <a href=\"/michaeltam\">@michaeltam</a>, I already did it  <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94433#latest-543978\">here</a> ;)</p>",
      "rawMarkdown": "Hi @michaeltam, I already did it  [here](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94433#latest-543978) ;)",
      "votes": null
    },
    {
      "id": "546972",
      "postDate": "06/07/2019 04:36:41",
      "content": "<p>Clear now. Thanks for explaining. Looks like at the end you effectively trained on modified TTF, correct?</p>",
      "rawMarkdown": "Clear now. Thanks for explaining. Looks like at the end you effectively trained on modified TTF, correct?",
      "votes": null
    },
    {
      "id": "547017",
      "postDate": "06/07/2019 06:39:06",
      "content": "<p>We trained models on original ttf and we trained models on modified target.</p>",
      "rawMarkdown": "We trained models on original ttf and we trained models on modified target.",
      "votes": null
    },
    {
      "id": "547090",
      "postDate": "06/07/2019 09:15:09",
      "content": "<p><strong>Thanks for sharing <a href=\"/cpmpml\">@cpmpml</a></strong> </p>",
      "rawMarkdown": "**Thanks for sharing @cpmpml**",
      "votes": null
    },
    {
      "id": "547592",
      "postDate": "06/07/2019 22:56:45",
      "content": "<p>Congrats and thanks for sharing! \nThis is the first time I hear about weighting samples, could you please elaborate on how it works? Thanks!</p>",
      "rawMarkdown": "Congrats and thanks for sharing! \nThis is the first time I hear about weighting samples, could you please elaborate on how it works? Thanks!",
      "votes": null
    },
    {
      "id": "547827",
      "postDate": "06/08/2019 11:04:41",
      "content": "<p>I assume it works be weighting each sample term in the overall objective function. </p>",
      "rawMarkdown": "I assume it works be weighting each sample term in the overall objective function.",
      "votes": null
    },
    {
      "id": "547840",
      "postDate": "06/08/2019 11:23:00",
      "content": "<p>Thanky you ! It helped a lot :)</p>",
      "rawMarkdown": "Thanky you ! It helped a lot :)",
      "votes": null
    },
    {
      "id": "548305",
      "postDate": "06/09/2019 06:36:09",
      "content": "<p>Great write-up. Thanks for sharing the many insights. </p>",
      "rawMarkdown": "Great write-up. Thanks for sharing the many insights.",
      "votes": null
    },
    {
      "id": "555528",
      "postDate": "06/19/2019 02:49:50",
      "content": "<p>Wow, amazing.  Thanks for sharing!</p>",
      "rawMarkdown": "Wow, amazing.  Thanks for sharing!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 543085,
      "author_name": "jesucristo",
      "author_url": "",
      "post_date": "06/04/2019 10:17:51",
      "content": "<p>Thank you so much <a href=\"/cpmpml\">@cpmpml</a> this is the best part of the competition, be able to learn from such great competitors like you, <a href=\"/stecasasso\">@stecasasso</a> and Antoine. Great solution, I'll try to do this by myself ;)</p>\n\n<h1>👍</h1>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543091,
      "author_name": "karanjakhar",
      "author_url": "",
      "post_date": "06/04/2019 10:21:34",
      "content": "<p>I was waiting for your solution. Thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543102,
      "author_name": "frankw",
      "author_url": "",
      "post_date": "06/04/2019 10:28:20",
      "content": "<p>It seems you are the only one able to predict ttf at <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91125#525889\">beginning of quake</a>.  Which feature(s) let you do that?</p>",
      "votes": null,
      "replies": [
        {
          "id": 543141,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/04/2019 10:50:26",
          "content": "<p>The features I describe above.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 543115,
      "author_name": "miguelpm",
      "author_url": "",
      "post_date": "06/04/2019 10:38:17",
      "content": "<p>Congratulations! Interesting and complete explanations, as usual, some of them will keep me thinking for a while, thanks for sharing :-)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543152,
      "author_name": "lucaskg",
      "author_url": "",
      "post_date": "06/04/2019 10:58:47",
      "content": "<p>Congrats! </p>\n\n<p>&gt; I also trained model where the target is TTF outside outliers, and TTF is reset at acoustic peaks.</p>\n\n<p>Is this the modified TTF model? You modified the TTF such that outliers' TTF=0 and non-outliers' TTF is unchanged?</p>\n\n<p>&gt; Final prediction is 0.2 if binary prediction is above a threshold (0.45 to 0.7 depending on the runs), and an average of models trained on original TTF and modified TTF as above.</p>\n\n<p>I don't understand your method of setting the threshold, it is different for each fold? And it is selected manually? The weights of model average is decided by some lgb model?</p>\n\n<p>Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 543161,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/04/2019 11:04:06",
          "content": "<p>We have three targets:</p>\n\n<ol>\n<li>binary one for outliers</li>\n<li>original ttf</li>\n<li>modified ttf: same values as ttf outside outliers, and ttf + next ttf peak for outliers. </li>\n</ol>\n\n<p>The third target is reset at acoustic peak, then decreases at the same rate as ttf.  The reset  height is such that the modified target meets original ttf value outside outliers.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 543190,
          "author_name": "lucaskg",
          "author_url": "",
          "post_date": "06/04/2019 11:20:02",
          "content": "<p>By the way, is there any plan of open-sourcing your code (or part of). The method is really interesting and worth more time studying. But since it is quite complicated, so I guess the code could help us understand better. Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 543199,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/04/2019 11:22:40",
          "content": "<p>The dataprep is scattered in several places, and cleaning it would be a time consuming task.  I am glad to not be in prize zone for that ;)  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 543200,
          "author_name": "lucaskg",
          "author_url": "",
          "post_date": "06/04/2019 11:22:53",
          "content": "<p>Thanks! Your modified TTF makes sense. If I remember correctly, peaks are start of failure, and the end of failure is only indicated by stress values.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 543201,
          "author_name": "lucaskg",
          "author_url": "",
          "post_date": "06/04/2019 11:23:24",
          "content": "<p>Thanks, that is of course understandable.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 543179,
      "author_name": "stecasasso",
      "author_url": "",
      "post_date": "06/04/2019 11:16:38",
      "content": "<p>Thanks for the neat and comprehensive write-up! As usual ;)\nAgain, it was a pleasure to team up with you!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543225,
      "author_name": "scirpus",
      "author_url": "",
      "post_date": "06/04/2019 11:50:42",
      "content": "<p>Lovely - if you write an academic paper about it just state your score as 1.2 and you are done! You will have to delete your validation stuff though;)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543231,
      "author_name": "joaopmpeinado",
      "author_url": "",
      "post_date": "06/04/2019 11:57:18",
      "content": "<p>Thanks for sharing CPMP!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543299,
      "author_name": "davids1992",
      "author_url": "",
      "post_date": "06/04/2019 13:00:55",
      "content": "<p>I used nested CV too, but only 4 Ti (and oof CV on the rest 12). It's my only small satisfaction that direction I took was at least a little bit similar to yours and Giba's. \nI was trying to learn as much as possible from you, thank you again for all the sharings.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543337,
      "author_name": "areveillon",
      "author_url": "",
      "post_date": "06/04/2019 13:34:01",
      "content": "<p>Thanks ! My takeaway are those 3 brilliant ideas from <a href=\"/cpmpml\">@cpmpml</a> : \n-Nested CV Setup\n-Strong usage of sample Weight based on ttf. Some team discarded some EQ from train data which also was very effective. But I think using sample weight was a better option here.\n-Binary classification model combined with 2 others models, this really gave us a strong boost, our model was able to predict correctly a majority of those point where we had the biggest error before. (first team used this as well)</p>",
      "votes": null,
      "replies": [
        {
          "id": 543343,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/04/2019 13:38:53",
          "content": "<p>You also had brilliant ideas, including the idea of weighting samples by EQ cycles!</p>\n\n<p>To be fair, the binary model idea was used in previous competitions to deal with outliers, in TGS and in ELO at least.  I didn't invented it ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 544029,
          "author_name": "simon80743314",
          "author_url": "",
          "post_date": "06/05/2019 03:54:56",
          "content": "<p>Congrats!\nCould you explain more about the binary model here?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 544180,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/05/2019 08:41:12",
          "content": "<p>It is lgb trained with the same features, only the target changes.  Target is 1 when the segment ends between an acoustic peak and the following TTF reset</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 547592,
          "author_name": "delai50",
          "author_url": "",
          "post_date": "06/07/2019 22:56:45",
          "content": "<p>Congrats and thanks for sharing! \nThis is the first time I hear about weighting samples, could you please elaborate on how it works? Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 547827,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/08/2019 11:04:41",
          "content": "<p>I assume it works be weighting each sample term in the overall objective function. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 543357,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "06/04/2019 13:44:24",
      "content": "<p>Congrats again and thanks for sharing. The 40kHz trick looks like magic. Just kidding ;)</p>",
      "votes": null,
      "replies": [
        {
          "id": 543362,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/04/2019 13:51:15",
          "content": "<p>Thanks, congrats to you as well!  I'm sure this reminds you of recent shakeup in malware ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 543553,
          "author_name": "titericz",
          "author_url": "",
          "post_date": "06/04/2019 15:37:20",
          "content": "<p>Yes, this reminds me Malware. Previous knowledge of Private testset survives the shakeup ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 543366,
      "author_name": "artgor",
      "author_url": "",
      "post_date": "06/04/2019 13:54:05",
      "content": "<p>Thank you for sharing! This is very interesting!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543371,
      "author_name": "gaborfodor",
      "author_url": "",
      "post_date": "06/04/2019 13:56:48",
      "content": "<p>Congrats guys! You had the most consistent LB results.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543414,
      "author_name": "timmmmmms",
      "author_url": "",
      "post_date": "06/04/2019 14:16:26",
      "content": "<p>Thank you for your sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543435,
      "author_name": "joejeo1",
      "author_url": "",
      "post_date": "06/04/2019 14:26:50",
      "content": "<p>Thanks for sharing.  A numbers of solid ideas without overly gaming the system.   </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543492,
      "author_name": "songwonho",
      "author_url": "",
      "post_date": "06/04/2019 15:02:01",
      "content": "<p>Thanks for sharing !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543496,
      "author_name": "dhaqui",
      "author_url": "",
      "post_date": "06/04/2019 15:08:15",
      "content": "<p>Congrats! Thanks for sharing your solution!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543581,
      "author_name": "asterisk",
      "author_url": "",
      "post_date": "06/04/2019 15:52:38",
      "content": "<p>cool, Congrats.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543617,
      "author_name": "abrial",
      "author_url": "",
      "post_date": "06/04/2019 16:27:07",
      "content": "<p>Thanks for sharing, I used a similar nested CV approach and the same data sampling, but did not find enough time to do proper FE. Congrats on the lead and (again) Gold medal :) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543628,
      "author_name": "carlospk",
      "author_url": "",
      "post_date": "06/04/2019 16:38:31",
      "content": "<p>Thanks for sharing. Interesting approach to treat outliers. Congrats!</p>",
      "votes": null,
      "replies": [
        {
          "id": 544660,
          "author_name": "michaeltam",
          "author_url": "",
          "post_date": "06/05/2019 18:35:09",
          "content": "<p>are you posting your solution overview? :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 546578,
          "author_name": "carlospk",
          "author_url": "",
          "post_date": "06/06/2019 17:38:53",
          "content": "<p>Hi <a href=\"/michaeltam\">@michaeltam</a>, I already did it  <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94433#latest-543978\">here</a> ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 543731,
      "author_name": "adityaecdrid",
      "author_url": "",
      "post_date": "06/04/2019 18:22:27",
      "content": "<p>i believe you tagged the wrong person (Antoine is <a href=\"/areveillon\">@areveillon</a>) <a href=\"/cpmpml\">@cpmpml</a> ...\nThanks for the write-up!</p>",
      "votes": null,
      "replies": [
        {
          "id": 543983,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/05/2019 02:18:13",
          "content": "<p>Oops will fix that asap</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 543887,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "06/04/2019 23:06:23",
      "content": "<p>Thank you for sharing. I have one question, how did you proceed feature selection to obtain specific frequency of MFCC (4th was the best?). Did you try increasing the feature one frequency by one to see the difference??</p>",
      "votes": null,
      "replies": [
        {
          "id": 543984,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/05/2019 02:19:27",
          "content": "<p>I started from librosa default value then tuned it via cross validation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 544169,
          "author_name": "corochann",
          "author_url": "",
          "post_date": "06/05/2019 08:26:15",
          "content": "<p>I mean how did you find \"4th\" frequency is the best?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 544176,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/05/2019 08:33:14",
          "content": "<p>By impact on CV score.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 544435,
          "author_name": "corochann",
          "author_url": "",
          "post_date": "06/05/2019 14:05:47",
          "content": "<p>I see, so you really run a lot of experiments to determine best features. Thank you for the answer.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 544443,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/05/2019 14:13:45",
          "content": "<p>Yes, a lot, and having a small set of features mean I can run experiments very quickly ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 543919,
      "author_name": "sishihara",
      "author_url": "",
      "post_date": "06/04/2019 23:57:47",
      "content": "<p>Congrats! I can enjoy this competition with your discussions, thank you.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 544032,
      "author_name": "prashanththangavel",
      "author_url": "",
      "post_date": "06/05/2019 03:57:53",
      "content": "<p>Congratulations! Thanks for sharing :) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 544045,
      "author_name": "shinsei66",
      "author_url": "",
      "post_date": "06/05/2019 04:07:59",
      "content": "<p>Congratulations and thanks for sharing! I always learn a lot from you, grand master!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 544481,
      "author_name": "carlospk",
      "author_url": "",
      "post_date": "06/05/2019 14:56:17",
      "content": "<p><a href=\"/cpmpml\">@cpmpml</a> I'm curious about how your model performed in the long cycles. Is it possible that you show us a plot of the oof vs ttf? Thanks</p>",
      "votes": null,
      "replies": [
        {
          "id": 544504,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/05/2019 15:24:17",
          "content": "<p>Here it is. Compared to yours we are lower, and probably more accurate on small ttf values.</p>\n\n<p>We have other models that look closer to yours, great for high ttf, but we blend them and I show the final result.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/544504/13413/oof.png\" alt=\"oof\"></p>\n\n<p>The spikes to 0 are due to our binary model fixing.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 544539,
          "author_name": "carlospk",
          "author_url": "",
          "post_date": "06/05/2019 16:16:44",
          "content": "<p>Wow... the prediction for low ttf looks awesome! Thanks for sharing. Very good job.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 544635,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/05/2019 18:12:34",
          "content": "<p>Between you and us we would have killed this competition maybe ;)  But a solo gold is invaluable these days!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 544777,
      "author_name": "wjshenggggg",
      "author_url": "",
      "post_date": "06/05/2019 22:51:33",
      "content": "<p>Thank you for sharing. Can you please offer your intuition on why did you pick <code>pygam</code> as a model? This is pretty new to me. Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 545970,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/06/2019 05:07:35",
          "content": "<p>I had a hunch but it is hard to say why.  I thought that feature interaction would not play a big role here.  Gam are just the addition of monovariate models.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 546370,
          "author_name": "wjshenggggg",
          "author_url": "",
          "post_date": "06/06/2019 14:19:43",
          "content": "<p>Thank you! Learned a new tool from you.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 546472,
          "author_name": "ivanbagmut",
          "author_url": "",
          "post_date": "06/06/2019 15:55:44",
          "content": "<p>Thanks! For me it is a new tool too.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 546009,
      "author_name": "dronych",
      "author_url": "",
      "post_date": "06/06/2019 06:18:04",
      "content": "<p>Hey <a href=\"/cpmpml\">@cpmpml</a> \nThanks for sharing your approach to this competition and congrats with staying in top 10!</p>\n\n<p>Could you give me some rope to understand the Outliers part of your solution? Let's see how I understood it partly and then questions.</p>\n\n<p>You train a binary classifier to detect TTF outliers which are exactly how you defined them. You then use this prediction as a feature in training your regression model. </p>\n\n<p>Questions:</p>\n\n<blockquote>\n  <p>Final prediction is 0.2 if binary prediction is above a threshold (0.45 to 0.7 depending on the runs), and an average of models trained on original TTF and modified TTF as above elsewhere.</p>\n</blockquote>\n\n<p>Don't understand this sentence. What do you mean by final prediction being 0.2? Can you elaborate please?</p>",
      "votes": null,
      "replies": [
        {
          "id": 546328,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/06/2019 13:26:59",
          "content": "<p>If the binary mode prediction is above, say 0.5, then we fix the prediction to be 0.2.  The threshold is tuned by cv, and it depends on the binary model.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 546972,
          "author_name": "dronych",
          "author_url": "",
          "post_date": "06/07/2019 04:36:41",
          "content": "<p>Clear now. Thanks for explaining. Looks like at the end you effectively trained on modified TTF, correct?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 547017,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/07/2019 06:39:06",
          "content": "<p>We trained models on original ttf and we trained models on modified target.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 547090,
      "author_name": "akashravichandran",
      "author_url": "",
      "post_date": "06/07/2019 09:15:09",
      "content": "<p><strong>Thanks for sharing <a href=\"/cpmpml\">@cpmpml</a></strong> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 547840,
      "author_name": "shauryagupta06",
      "author_url": "",
      "post_date": "06/08/2019 11:23:00",
      "content": "<p>Thanky you ! It helped a lot :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 548305,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "06/09/2019 06:36:09",
      "content": "<p>Great write-up. Thanks for sharing the many insights. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 555528,
      "author_name": "vettejeep",
      "author_url": "",
      "post_date": "06/19/2019 02:49:50",
      "content": "<p>Wow, amazing.  Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "543068": "First of all, I want to thank organizers for this interesting and tricky challenge, as well as my team mates @antoine and @stecasasso for their interesting views.  You can find @antoine's view [here](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94359#latest-542974), and @stecasasso [here](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94408).  I also want to thank @mykper for his great analysis of public test data.  \n\nOur solution is based on few pieces.  I'll start by what I did before teaming.\n\n**Cross Validation**\n\nWe used 16 fold CV where each fold includes a time to failure (TTF) reset.  Lately we realized that a simple unshuffled  16 folds would be equivalent  to what we did.  Even with 16 folds, it was quickly clear that CV LB correlation was poor, and that public LB was not a reliable.  I decided to use a nested CV approach where one fold Ti is used as test data, and CV is run on the other 15 folds.  15 models are trained using the 15 fold cv and their predictions on Ti are averaged.  This mimics the submission process 16 times, and provides a reliable indicator of a model strength by averaging the mae of Ti folds.  Correlation with public LB was very good.\n\n**Data Sampling**\n\nI used 150k segments starting every 50k rows.  Every segment overlaps with 2 segments before it and 2 segments after it.  I tested more overlap, and less overlap, and this one was the best trade off between running time and result accuracy.  A potential issue could be that segments at fold boundary overlap.  I tried removing those overlap but result accuracy decreased a bit, hence I kept all segments.\n\n**Outliers**\n\nTTF is not reset at acoustic data peaks which makes the problem difficult.  I decided to treat the points between acoustic peak and the following TTF reset as outlier, and trained a binary model to predict outliers.  I also trained model where the target is TTF outside outliers, and TTF is reset at acoustic peaks.  \n\nFinal prediction is 0.2 if binary prediction is above a threshold (0.45 to 0.7 depending on the runs), and an average of models trained on original TTF and modified TTF as above elsewhere.\n\n**Features**\n\nI used 7 MFCC from librosa.  To make it work I pretended than data was sampled at 40kHz.  Not doing so means the useful info is way higher, around the 20th MFCC.  I also used 4 features based on std of various signal quantiles.  Last, I used a binary feature indicating high variance segments that correspond to acoustic peaks very accurately.\n\n**Models**\n\nMostly lgb with conservative settings, like 7 leaves only.  But also general additive models using pygam.  I also trained a knn model which was surprisingly good.  In hindsight, gam was better than lgb which was better than knn, the gam model would get a gold medal alone at 2.348.  For lgb I tried mse, gamma and huber objective.  With the above I got to 1.288 on public LB at which point I teamed.\n\nAfter teaming, progress was in many areas, but here are the three most important ones.  \n\n**Ensembling**\n\nFirst, model diversity as my team mates were using different data sets.  By averaging our top public LB @areveillon had a 1.286 public LB) we got to 1.275 and took the lead.  TMy team mates adapted their models to the binary vs regression models.    We also reused some features from each other dataset, which added to model diversity.  We then worked on stacking, ending with a stack based on a lgb, a gam, and a knn from me, a lgb from @Antoine and a NN from @stecasasso .  lgb was used for the second level model.  We validated stacking using our nested CV which gives us some confidence that stacking was indeed working.  Private LB confirms that stacking was improving over base models.\n\n**Train / Test Difference**\n\nI must admit that I hate LB probing, and avoid it as plague.  But here, using pictures from academic papers could lead to a good estimate of test data, and it was the way to go given the significant difference between train and test data.  @Antoine had done an estimate before teaming and was using it already.  Few days before end he convinced me that we should really base our final sub on it and I measured precisely length of EQ cycles in that picture and could estimate their duration by running a linear regression on train data.  From it I estimated its mean to be 6.35 which is quite close to the actual 6.32.\n\nFrom this we can estimate density function of TTF for train and test.  We then used sample weights to map the train distribution to the test distribution.  I see that top 2 teams (at least) rather resampled train, and this may explain why they are ahead.  There are two ways to compute weights.  I tried to base them on density function, i.e. weights only depend on the TTF, while @Antoine was pushing for weights per eq cycle.  In hindsight he was right.  Given we could not agree before competition end we decided to produce two submissions, one that optimizes my weights while the second one was optimizing @areveillon 's weights.   @stecasasso  then had the idea of using the weights not only for evaluating our models, but also for training models.  In the end, submissions optimizing @areveillon 's weights were the best ones.\n\n**Post-processing**\n\n@stecasasso found that we could set all high variance segments TTF to 0.31.  @Antoine looked at how to best combine binary models to fix segments TTF at 0.2.  Applying these as postprocessing improved submissions.  It also moved our best public LB from 2.75 to 2.45.\n\n**What did not work**\n\nI could not get NN models to work well enough to be useful in our stack.  @stecasasso managed to get one, but it was the weakest model on the stack.  lgb and gam are better.  I think NNs could be useful on processed data, either sftt or Hilbert envelope, but I did not have the time to try seriously.\n\n**Answers**\n\nI was asked about what was the 'no magic' feature. It was the 4th MFCC I was using.  I also was asked how to overfit public LB.  Well, first answer is that it was easy given how many people had way worse results on private Lb than public LB.  But in our case it was just by capping test prediction to 9.6 then to 9, given @mykper had shown public test TTF was below 10. This led us to 1.236 from 1.245.  I'm curious about how kopeyka team overfitted public test so effectively.  Last, I was asked about how I survive shakeups like this one or in previous competitions: it is because I only trust my CV and do not rely on LB probing.",
    "543085": "Thank you so much @cpmpml this is the best part of the competition, be able to learn from such great competitors like you, @stecasasso and Antoine. Great solution, I'll try to do this by myself ;)\n# 👍",
    "543091": "I was waiting for your solution. Thanks for sharing.",
    "543102": "It seems you are the only one able to predict ttf at [beginning of quake](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91125#525889).  Which feature(s) let you do that?",
    "543115": "Congratulations! Interesting and complete explanations, as usual, some of them will keep me thinking for a while, thanks for sharing :-)",
    "543141": "The features I describe above.",
    "543152": "Congrats! \n\n&gt; I also trained model where the target is TTF outside outliers, and TTF is reset at acoustic peaks.\n\nIs this the modified TTF model? You modified the TTF such that outliers' TTF=0 and non-outliers' TTF is unchanged?\n\n&gt; Final prediction is 0.2 if binary prediction is above a threshold (0.45 to 0.7 depending on the runs), and an average of models trained on original TTF and modified TTF as above.\n\nI don't understand your method of setting the threshold, it is different for each fold? And it is selected manually? The weights of model average is decided by some lgb model?\n\nThanks!",
    "543161": "We have three targets:\n\n1. binary one for outliers\n2. original ttf\n3. modified ttf: same values as ttf outside outliers, and ttf + next ttf peak for outliers. \n\nThe third target is reset at acoustic peak, then decreases at the same rate as ttf.  The reset  height is such that the modified target meets original ttf value outside outliers.",
    "543179": "Thanks for the neat and comprehensive write-up! As usual ;)\nAgain, it was a pleasure to team up with you!",
    "543190": "By the way, is there any plan of open-sourcing your code (or part of). The method is really interesting and worth more time studying. But since it is quite complicated, so I guess the code could help us understand better. Thanks!",
    "543199": "The dataprep is scattered in several places, and cleaning it would be a time consuming task.  I am glad to not be in prize zone for that ;)",
    "543200": "Thanks! Your modified TTF makes sense. If I remember correctly, peaks are start of failure, and the end of failure is only indicated by stress values.",
    "543201": "Thanks, that is of course understandable.",
    "543225": "Lovely - if you write an academic paper about it just state your score as 1.2 and you are done! You will have to delete your validation stuff though;)",
    "543231": "Thanks for sharing CPMP!",
    "543299": "I used nested CV too, but only 4 Ti (and oof CV on the rest 12). It's my only small satisfaction that direction I took was at least a little bit similar to yours and Giba's. \nI was trying to learn as much as possible from you, thank you again for all the sharings.",
    "543337": "Thanks ! My takeaway are those 3 brilliant ideas from @cpmpml : \n-Nested CV Setup\n-Strong usage of sample Weight based on ttf. Some team discarded some EQ from train data which also was very effective. But I think using sample weight was a better option here.\n-Binary classification model combined with 2 others models, this really gave us a strong boost, our model was able to predict correctly a majority of those point where we had the biggest error before. (first team used this as well)",
    "543343": "You also had brilliant ideas, including the idea of weighting samples by EQ cycles!\n\nTo be fair, the binary model idea was used in previous competitions to deal with outliers, in TGS and in ELO at least.  I didn't invented it ;)",
    "543357": "Congrats again and thanks for sharing. The 40kHz trick looks like magic. Just kidding ;)",
    "543362": "Thanks, congrats to you as well!  I'm sure this reminds you of recent shakeup in malware ;)",
    "543366": "Thank you for sharing! This is very interesting!",
    "543371": "Congrats guys! You had the most consistent LB results.",
    "543414": "Thank you for your sharing!",
    "543435": "Thanks for sharing.  A numbers of solid ideas without overly gaming the system.",
    "543492": "Thanks for sharing !",
    "543496": "Congrats! Thanks for sharing your solution!",
    "543553": "Yes, this reminds me Malware. Previous knowledge of Private testset survives the shakeup ;)",
    "543581": "cool, Congrats.",
    "543617": "Thanks for sharing, I used a similar nested CV approach and the same data sampling, but did not find enough time to do proper FE. Congrats on the lead and (again) Gold medal :)",
    "543628": "Thanks for sharing. Interesting approach to treat outliers. Congrats!",
    "543731": "i believe you tagged the wrong person (Antoine is @areveillon) @cpmpml ...\nThanks for the write-up!",
    "543887": "Thank you for sharing. I have one question, how did you proceed feature selection to obtain specific frequency of MFCC (4th was the best?). Did you try increasing the feature one frequency by one to see the difference??",
    "543919": "Congrats! I can enjoy this competition with your discussions, thank you.",
    "543983": "Oops will fix that asap",
    "543984": "I started from librosa default value then tuned it via cross validation.",
    "544029": "Congrats!\nCould you explain more about the binary model here?",
    "544032": "Congratulations! Thanks for sharing :)",
    "544045": "Congratulations and thanks for sharing! I always learn a lot from you, grand master!!",
    "544169": "I mean how did you find \"4th\" frequency is the best?",
    "544176": "By impact on CV score.",
    "544180": "It is lgb trained with the same features, only the target changes.  Target is 1 when the segment ends between an acoustic peak and the following TTF reset",
    "544435": "I see, so you really run a lot of experiments to determine best features. Thank you for the answer.",
    "544443": "Yes, a lot, and having a small set of features mean I can run experiments very quickly ;)",
    "544481": "cpmpml I'm curious about how your model performed in the long cycles. Is it possible that you show us a plot of the oof vs ttf? Thanks",
    "544504": "Here it is. Compared to yours we are lower, and probably more accurate on small ttf values.\n\nWe have other models that look closer to yours, great for high ttf, but we blend them and I show the final result.\n\n![oof](https://storage.googleapis.com/kaggle-forum-message-attachments/544504/13413/oof.png)\n\nThe spikes to 0 are due to our binary model fixing.",
    "544539": "Wow... the prediction for low ttf looks awesome! Thanks for sharing. Very good job.",
    "544635": "Between you and us we would have killed this competition maybe ;)  But a solo gold is invaluable these days!",
    "544660": "are you posting your solution overview? :)",
    "544777": "Thank you for sharing. Can you please offer your intuition on why did you pick `pygam` as a model? This is pretty new to me. Thanks!",
    "545970": "I had a hunch but it is hard to say why.  I thought that feature interaction would not play a big role here.  Gam are just the addition of monovariate models.",
    "546009": "Hey @cpmpml \nThanks for sharing your approach to this competition and congrats with staying in top 10!\n\nCould you give me some rope to understand the Outliers part of your solution? Let's see how I understood it partly and then questions.\n\nYou train a binary classifier to detect TTF outliers which are exactly how you defined them. You then use this prediction as a feature in training your regression model. \n\nQuestions:\n&gt;Final prediction is 0.2 if binary prediction is above a threshold (0.45 to 0.7 depending on the runs), and an average of models trained on original TTF and modified TTF as above elsewhere.\n\nDon't understand this sentence. What do you mean by final prediction being 0.2? Can you elaborate please?",
    "546328": "If the binary mode prediction is above, say 0.5, then we fix the prediction to be 0.2.  The threshold is tuned by cv, and it depends on the binary model.",
    "546370": "Thank you! Learned a new tool from you.",
    "546472": "Thanks! For me it is a new tool too.",
    "546578": "Hi @michaeltam, I already did it  [here](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94433#latest-543978) ;)",
    "546972": "Clear now. Thanks for explaining. Looks like at the end you effectively trained on modified TTF, correct?",
    "547017": "We trained models on original ttf and we trained models on modified target.",
    "547090": "**Thanks for sharing @cpmpml**",
    "547592": "Congrats and thanks for sharing! \nThis is the first time I hear about weighting samples, could you please elaborate on how it works? Thanks!",
    "547827": "I assume it works be weighting each sample term in the overall objective function.",
    "547840": "Thanky you ! It helped a lot :)",
    "548305": "Great write-up. Thanks for sharing the many insights.",
    "555528": "Wow, amazing.  Thanks for sharing!"
  },
  "source": "meta"
}