{
  "id": 94369,
  "title": "2nd place solution",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/94369",
  "author_name": "🐢 Jun Koda",
  "post_date": "2019-06-04T06:04:09.494000",
  "votes": 89,
  "comment_count": 30,
  "views": 0,
  "content": "<p>Update: Feature importance label fixed (5th June GMT 3 am) </p>\n\n<p>I select segments from the traning ​​data and create a private-set-like data based on the inter-earthquake times seen in the figure \"Are data from p4677?\"</p>\n\n<p>mykper: <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90664#latest-535844\">https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90664#latest-535844</a></p>\n\n<p><img src=\"https://junkoda.github.io/figs/quake/quake_final_private_like.png\" alt=\"quake_private_like_set\"></p>\n\n<p>The dashed line is 11.5 sec.</p>\n\n<p>The first segment is counted twice; the private set has 8 segments. The last segment is created from the 14-sec segment and shifted up by 2 seconds, and the last 4.5 second is cut, treating that as the public part.</p>\n\n<p>I use 32 features:</p>\n\n<ul>\n<li>standard deviation (std)</li>\n<li>std_nopeak: Std of data that are not part of peaks</li>\n<li>kurtosis_truncated: Kurtosis of data with abs(v - mean(v)) &lt; 20</li>\n<li>7 peak counts: Number of peaks with hight &gt; 50, 75, ..., 200</li>\n<li>5 percentiles: 95 percentile - 5 percentile, 80 - 20, 70 - 30, 60 - 40.</li>\n<li>trend: slope of robust linear regression to 30 sub chunks of std_truncated</li>\n<li>trend_error: Abs difference in the slope of RANSAC and Huber fit</li>\n<li>power spectrum; Fast-Fourier Transform the data and average the absolute value in 15 bins</li>\n</ul>\n\n<p>I choose combinations such as 95 percentile - 5 percentile to avoid direct dependence on the mean; which is drifting with time. Same for the peak height; the peak height is defined as (max - min)/2.</p>\n\n<p>The std_truncated (std instead of kurtosis in kurtosis_truncated) works almost as well as std_nopeak.</p>\n\n<p>I randomly select 2000x1000 training chunks of length 150_000 from my private-like set, which is 1000 times the number of independent/non-overlapping chunks and put all of them into CatBoost. I do not provide CV data to the regressor; eveything ​is the training set.</p>\n\n<p>I also tried to predict time since failure using all the training set and tried to stitch together with time to failure, but I was not able to do that successfully.</p>\n\n<p>This is the feature importance:</p>\n\n<p><img src=\"https://junkoda.github.io/figs/quake/quake_feature_importance.png\" alt=\"Feature importance\"></p>",
  "messages": [
    {
      "id": 542827,
      "postDate": "2019-06-04T06:04:09.493Z",
      "content": "<p>Update: Feature importance label fixed (5th June GMT 3 am) </p>\n\n<p>I select segments from the traning ​​data and create a private-set-like data based on the inter-earthquake times seen in the figure \"Are data from p4677?\"</p>\n\n<p>mykper: <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90664#latest-535844\">https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90664#latest-535844</a></p>\n\n<p><img src=\"https://junkoda.github.io/figs/quake/quake_final_private_like.png\" alt=\"quake_private_like_set\"></p>\n\n<p>The dashed line is 11.5 sec.</p>\n\n<p>The first segment is counted twice; the private set has 8 segments. The last segment is created from the 14-sec segment and shifted up by 2 seconds, and the last 4.5 second is cut, treating that as the public part.</p>\n\n<p>I use 32 features:</p>\n\n<ul>\n<li>standard deviation (std)</li>\n<li>std_nopeak: Std of data that are not part of peaks</li>\n<li>kurtosis_truncated: Kurtosis of data with abs(v - mean(v)) &lt; 20</li>\n<li>7 peak counts: Number of peaks with hight &gt; 50, 75, ..., 200</li>\n<li>5 percentiles: 95 percentile - 5 percentile, 80 - 20, 70 - 30, 60 - 40.</li>\n<li>trend: slope of robust linear regression to 30 sub chunks of std_truncated</li>\n<li>trend_error: Abs difference in the slope of RANSAC and Huber fit</li>\n<li>power spectrum; Fast-Fourier Transform the data and average the absolute value in 15 bins</li>\n</ul>\n\n<p>I choose combinations such as 95 percentile - 5 percentile to avoid direct dependence on the mean; which is drifting with time. Same for the peak height; the peak height is defined as (max - min)/2.</p>\n\n<p>The std_truncated (std instead of kurtosis in kurtosis_truncated) works almost as well as std_nopeak.</p>\n\n<p>I randomly select 2000x1000 training chunks of length 150_000 from my private-like set, which is 1000 times the number of independent/non-overlapping chunks and put all of them into CatBoost. I do not provide CV data to the regressor; eveything ​is the training set.</p>\n\n<p>I also tried to predict time since failure using all the training set and tried to stitch together with time to failure, but I was not able to do that successfully.</p>\n\n<p>This is the feature importance:</p>\n\n<p><img src=\"https://junkoda.github.io/figs/quake/quake_feature_importance.png\" alt=\"Feature importance\"></p>",
      "rawMarkdown": "Update: Feature importance label fixed (5th June GMT 3 am) \n\nI select segments from the traning ​​data and create a private-set-like data based on the inter-earthquake times seen in the figure \"Are data from p4677?\"\n\nmykper: https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90664#latest-535844\n\n![quake_private_like_set](https://junkoda.github.io/figs/quake/quake_final_private_like.png)\n\nThe dashed line is 11.5 sec.\n\nThe first segment is counted twice; the private set has 8 segments. The last segment is created from the 14-sec segment and shifted up by 2 seconds, and the last 4.5 second is cut, treating that as the public part.\n\nI use 32 features:\n\n- standard deviation (std)\n- std_nopeak: Std of data that are not part of peaks\n- kurtosis_truncated: Kurtosis of data with abs(v - mean(v)) &lt; 20\n- 7 peak counts: Number of peaks with hight &gt; 50, 75, ..., 200\n- 5 percentiles: 95 percentile - 5 percentile, 80 - 20, 70 - 30, 60 - 40.\n- trend: slope of robust linear regression to 30 sub chunks of std_truncated\n- trend_error: Abs difference in the slope of RANSAC and Huber fit\n- power spectrum; Fast-Fourier Transform the data and average the absolute value in 15 bins\n\nI choose combinations such as 95 percentile - 5 percentile to avoid direct dependence on the mean; which is drifting with time. Same for the peak height; the peak height is defined as (max - min)/2.\n\nThe std_truncated (std instead of kurtosis in kurtosis_truncated) works almost as well as std_nopeak.\n\n\nI randomly select 2000x1000 training chunks of length 150_000 from my private-like set, which is 1000 times the number of independent/non-overlapping chunks and put all of them into CatBoost. I do not provide CV data to the regressor; eveything ​is the training set.\n\n\nI also tried to predict time since failure using all the training set and tried to stitch together with time to failure, but I was not able to do that successfully.\n\nThis is the feature importance:\n\n![Feature importance](https://junkoda.github.io/figs/quake/quake_feature_importance.png)",
      "votes": 87
    },
    {
      "id": 564291,
      "postDate": "2019-06-29T08:42:26.190Z",
      "content": "<p>At first heartily congrats <a href=\"/junkoda\">@junkoda</a>  and thanks a ton for sharing the solution </p>",
      "rawMarkdown": "At first heartily congrats @junkoda  and thanks a ton for sharing the solution ",
      "votes": 4
    },
    {
      "id": 542964,
      "postDate": "2019-06-04T08:36:23.613Z",
      "content": "<p>Сongratulations!\nHow did you select these features? What was your feature selection/feature generation approach?</p>",
      "rawMarkdown": "Сongratulations!\nHow did you select these features? What was your feature selection/feature generation approach?",
      "votes": 4,
      "replies": [
        {
          "id": 542987,
          "postDate": "2019-06-04T08:53:18.253Z",
          "content": "<p>Umm, that's difficult to answer. I felt there aren'​t much information we can extract from the acustic ​data and I felt I have more than enough features to get all the information. This is intuitive, not logical. I did some trial and errors trying more percentiles, switching std_truncated/std_nopeak, but those were manual and not systematic. For FFT, I don't think real, imag, phase are translational invariant; that is, depends on where you choose time 0, so I only use magnitude (abs), but percentiles of abs are reasonable; simply, I forgot trying that. Short answer: intution ​and some trial and error 😅</p>",
          "rawMarkdown": "Umm, that's difficult to answer. I felt there aren'​t much information we can extract from the acustic ​data and I felt I have more than enough features to get all the information. This is intuitive, not logical. I did some trial and errors trying more percentiles, switching std_truncated/std_nopeak, but those were manual and not systematic. For FFT, I don't think real, imag, phase are translational invariant; that is, depends on where you choose time 0, so I only use magnitude (abs), but percentiles of abs are reasonable; simply, I forgot trying that. Short answer: intution ​and some trial and error 😅",
          "votes": 5
        },
        {
          "id": 542990,
          "postDate": "2019-06-04T08:56:07.647Z",
          "content": "<p>PS: FFT features are not ranking high in my feature importance, and that is also so in one of the organizers' random-forest paper, so I though ​FFT features are not effective (which is probably not true)</p>",
          "rawMarkdown": "PS: FFT features are not ranking high in my feature importance, and that is also so in one of the organizers' random-forest paper, so I though ​FFT features are not effective (which is probably not true)",
          "votes": 2
        },
        {
          "id": 543017,
          "postDate": "2019-06-04T09:20:48.627Z",
          "content": "<p>FFT features did not do much for my score, but then the score isn't very reliable :)</p>\n\n<p>Intuitively, I suspected using Wavelets is the right way to go but I wasn't sure how to apply them.</p>",
          "rawMarkdown": "FFT features did not do much for my score, but then the score isn't very reliable :)\n\nIntuitively, I suspected using Wavelets is the right way to go but I wasn't sure how to apply them.",
          "votes": 2
        },
        {
          "id": 544017,
          "postDate": "2019-06-05T03:25:00.280Z",
          "content": "<p>I fixed the feature importance ordering. FFT features (power spectrum power n) are working. I saw somewhere that low-frequency power spectrum is an intererting​ feature near the earthquake, and the high rank of power0 confirms that.</p>",
          "rawMarkdown": "I fixed the feature importance ordering. FFT features (power spectrum power n) are working. I saw somewhere that low-frequency power spectrum is an intererting​ feature near the earthquake, and the high rank of power0 confirms that.",
          "votes": 1
        }
      ]
    },
    {
      "id": 542912,
      "postDate": "2019-06-04T07:25:03.573Z",
      "content": "<p>Congrats! You did well also in <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection\">VSB Power Line Fault Detection</a>. I feel you're really great because you always make such robust model. </p>\n\n<p>I'd like to ask one thing: Have you tried other machine learning algorithm like LightGBM/XGboost? If there is a reason why you use Catboost, I want to know.</p>",
      "rawMarkdown": "Congrats! You did well also in [VSB Power Line Fault Detection](https://www.kaggle.com/c/vsb-power-line-fault-detection). I feel you're really great because you always make such robust model. \n\nI'd like to ask one thing: Have you tried other machine learning algorithm like LightGBM/XGboost? If there is a reason why you use Catboost, I want to know.",
      "votes": 4,
      "replies": [
        {
          "id": 542931,
          "postDate": "2019-06-04T08:00:53.263Z",
          "content": "<p>Thank you! Seems I have a strong will and not moved by the Public Leaderboard score. No, I didn't mix with other gradient-boosting results; I will do so next time because I would regret if I did not do so and lost by a score like 0.001. I did not expect such small differences in the prize zone. No reason using Catboost, it was simply the first gradient boosting library I saw in a public kernel. 🐣</p>",
          "rawMarkdown": "Thank you! Seems I have a strong will and not moved by the Public Leaderboard score. No, I didn't mix with other gradient-boosting results; I will do so next time because I would regret if I did not do so and lost by a score like 0.001. I did not expect such small differences in the prize zone. No reason using Catboost, it was simply the first gradient boosting library I saw in a public kernel. 🐣",
          "votes": 6
        },
        {
          "id": 542981,
          "postDate": "2019-06-04T08:50:37.843Z",
          "content": "<p>I tried a few gradient boosters with my kernel template. For reference, I used this to give insight into the different gradient boosters:\n<a href=\"https://towardsdatascience.com/catboost-vs-light-gbm-vs-xgboost-5f93620723db\">CatBoost vs Light GBM vs XGBoost</a>.</p>\n\n<p>XGBoost performed the best. With the same features, and only changing the booster, Cat Boost achieved 2.34994, while XGBoost managed a public score of 1.60521.</p>\n\n<p>Hope this helps!</p>",
          "rawMarkdown": "I tried a few gradient boosters with my kernel template. For reference, I used this to give insight into the different gradient boosters:\n[CatBoost vs Light GBM vs XGBoost](https://towardsdatascience.com/catboost-vs-light-gbm-vs-xgboost-5f93620723db).\n\nXGBoost performed the best. With the same features, and only changing the booster, Cat Boost achieved 2.34994, while XGBoost managed a public score of 1.60521.\n\nHope this helps!",
          "votes": -1
        }
      ]
    },
    {
      "id": 543990,
      "postDate": "2019-06-05T02:32:34.943Z",
      "content": "<p>Congratulation for the 2nd place and sharing your thoughts, this post truly gave me a deep understanding!</p>",
      "rawMarkdown": "Congratulation for the 2nd place and sharing your thoughts, this post truly gave me a deep understanding!",
      "votes": 1
    },
    {
      "id": 543864,
      "postDate": "2019-06-04T22:29:38.887Z",
      "content": "<p>Congrats on the result. I'm totally new to this and I really learn a lot from your post.</p>",
      "rawMarkdown": "Congrats on the result. I'm totally new to this and I really learn a lot from your post.",
      "votes": 1
    },
    {
      "id": 543081,
      "postDate": "2019-06-04T10:16:16.430Z",
      "content": "<p>Congrats on the result and the method!</p>",
      "rawMarkdown": "Congrats on the result and the method!",
      "votes": 1
    },
    {
      "id": 543580,
      "postDate": "2019-06-04T15:51:09.203Z",
      "content": "<p>Did you tried Lightgbm or Xgboost on your solution ?</p>",
      "rawMarkdown": "Did you tried Lightgbm or Xgboost on your solution ?",
      "votes": 2,
      "replies": [
        {
          "id": 543975,
          "postDate": "2019-06-05T02:11:27.957Z",
          "content": "<p>No. I have installed them..., but I mistakenly remember the deadline for 1 week and didn't have time for the stacking phase. Default parameter ​CatBoost single model 😅</p>",
          "rawMarkdown": "No. I have installed them..., but I mistakenly remember the deadline for 1 week and didn't have time for the stacking phase. Default parameter ​CatBoost single model 😅",
          "votes": 2
        }
      ]
    },
    {
      "id": 543227,
      "postDate": "2019-06-04T11:53:08.093Z",
      "content": "<p>Congrats! \nThe idea of tweaking the training data to create virtual test set is interesting.</p>",
      "rawMarkdown": "Congrats! \nThe idea of tweaking the training data to create virtual test set is interesting.",
      "votes": 2
    },
    {
      "id": 542895,
      "postDate": "2019-06-04T07:08:59.007Z",
      "content": "<p>Congratulations! Thank you for the update and for explaining your solution!</p>",
      "rawMarkdown": "Congratulations! Thank you for the update and for explaining your solution!",
      "votes": 2
    },
    {
      "id": 542869,
      "postDate": "2019-06-04T06:44:00.133Z",
      "content": "<p>Well done!</p>",
      "rawMarkdown": "Well done!",
      "votes": 2
    },
    {
      "id": 542850,
      "postDate": "2019-06-04T06:32:05.450Z",
      "content": "<p>Congrats!</p>",
      "rawMarkdown": "Congrats!",
      "votes": 2
    },
    {
      "id": 547184,
      "postDate": "2019-06-07T11:28:59.483Z",
      "content": "<p>Congratulation for the solution and 2nd place ! \nWhat exactly means that \"14-sec segment and shifted up by 2 seconds\" ? You just simply added 2 sec or did you do something more fancy ?</p>",
      "rawMarkdown": "Congratulation for the solution and 2nd place ! \nWhat exactly means that \"14-sec segment and shifted up by 2 seconds\" ? You just simply added 2 sec or did you do something more fancy ?",
      "replies": [
        {
          "id": 547200,
          "postDate": "2019-06-07T11:55:39.273Z",
          "content": "<p>Thanks. Exaclty​, just added 2.* something to the time to failure.</p>",
          "rawMarkdown": "Thanks. Exaclty​, just added 2.* something to the time to failure.",
          "votes": 1
        }
      ]
    },
    {
      "id": 545951,
      "postDate": "2019-06-06T04:42:03.473Z",
      "content": "<p>Congratulations on the position and thanks for sharing!</p>",
      "rawMarkdown": "Congratulations on the position and thanks for sharing!"
    },
    {
      "id": 543982,
      "postDate": "2019-06-05T02:16:21.763Z",
      "content": "<p>Congratulations <a href=\"/junkoda\">@junkoda</a>. Thank you for sharing. </p>",
      "rawMarkdown": "Congratulations @junkoda. Thank you for sharing. "
    },
    {
      "id": 543579,
      "postDate": "2019-06-04T15:50:12.727Z",
      "content": "<p>Congrats <a href=\"/junkoda\">@junkoda</a> and thanks for sharing. </p>",
      "rawMarkdown": "Congrats @junkoda and thanks for sharing. "
    },
    {
      "id": 544077,
      "postDate": "2019-06-05T05:14:58.200Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    },
    {
      "id": 542968,
      "postDate": "2019-06-04T08:40:52.090Z",
      "content": "<p>Thanks for sharing! Good job!</p>",
      "rawMarkdown": "Thanks for sharing! Good job!",
      "votes": 1
    },
    {
      "id": 542945,
      "postDate": "2019-06-04T08:11:57.570Z",
      "content": "<p>Congrats and thanks for sharing !</p>",
      "rawMarkdown": "Congrats and thanks for sharing !",
      "votes": 1
    },
    {
      "id": 543568,
      "postDate": "2019-06-04T15:45:04.120Z",
      "content": "<p>Great job!\nThanks for sharing.</p>",
      "rawMarkdown": "Great job!\nThanks for sharing."
    },
    {
      "id": 543124,
      "postDate": "2019-06-04T10:42:27.667Z",
      "content": "<p>Great work! Thanks for sharing.</p>",
      "rawMarkdown": "Great work! Thanks for sharing."
    },
    {
      "id": 543087,
      "postDate": "2019-06-04T10:19:20.613Z",
      "content": "<p>Thank you for sharing.\nCongratulations!</p>",
      "rawMarkdown": "Thank you for sharing.\nCongratulations!"
    },
    {
      "id": 543043,
      "postDate": "2019-06-04T09:46:06.583Z",
      "content": "<p>Congrats! Thanks for sharing :)</p>",
      "rawMarkdown": "Congrats! Thanks for sharing :)"
    }
  ],
  "comments": [
    {
      "id": 564291,
      "author_name": "Raju Kumar Mishra",
      "author_url": "",
      "post_date": "2019-06-29T08:42:26.190000",
      "content": "<p>At first heartily congrats <a href=\"/junkoda\">@junkoda</a>  and thanks a ton for sharing the solution </p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 542964,
      "author_name": "Nickolay Safronov",
      "author_url": "",
      "post_date": "2019-06-04T08:36:23.613000",
      "content": "<p>Сongratulations!\nHow did you select these features? What was your feature selection/feature generation approach?</p>",
      "votes": 4,
      "replies": [
        {
          "id": 542987,
          "author_name": "🐢 Jun Koda",
          "author_url": "",
          "post_date": "2019-06-04T08:53:18.253000",
          "content": "<p>Umm, that's difficult to answer. I felt there aren'​t much information we can extract from the acustic ​data and I felt I have more than enough features to get all the information. This is intuitive, not logical. I did some trial and errors trying more percentiles, switching std_truncated/std_nopeak, but those were manual and not systematic. For FFT, I don't think real, imag, phase are translational invariant; that is, depends on where you choose time 0, so I only use magnitude (abs), but percentiles of abs are reasonable; simply, I forgot trying that. Short answer: intution ​and some trial and error 😅</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 542990,
          "author_name": "🐢 Jun Koda",
          "author_url": "",
          "post_date": "2019-06-04T08:56:07.647000",
          "content": "<p>PS: FFT features are not ranking high in my feature importance, and that is also so in one of the organizers' random-forest paper, so I though ​FFT features are not effective (which is probably not true)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 543017,
          "author_name": "DevilEars",
          "author_url": "",
          "post_date": "2019-06-04T09:20:48.627000",
          "content": "<p>FFT features did not do much for my score, but then the score isn't very reliable :)</p>\n\n<p>Intuitively, I suspected using Wavelets is the right way to go but I wasn't sure how to apply them.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 544017,
          "author_name": "🐢 Jun Koda",
          "author_url": "",
          "post_date": "2019-06-05T03:25:00.280000",
          "content": "<p>I fixed the feature importance ordering. FFT features (power spectrum power n) are working. I saw somewhere that low-frequency power spectrum is an intererting​ feature near the earthquake, and the high rank of power0 confirms that.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 542912,
      "author_name": "u++",
      "author_url": "",
      "post_date": "2019-06-04T07:25:03.573000",
      "content": "<p>Congrats! You did well also in <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection\">VSB Power Line Fault Detection</a>. I feel you're really great because you always make such robust model. </p>\n\n<p>I'd like to ask one thing: Have you tried other machine learning algorithm like LightGBM/XGboost? If there is a reason why you use Catboost, I want to know.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 542931,
          "author_name": "🐢 Jun Koda",
          "author_url": "",
          "post_date": "2019-06-04T08:00:53.263000",
          "content": "<p>Thank you! Seems I have a strong will and not moved by the Public Leaderboard score. No, I didn't mix with other gradient-boosting results; I will do so next time because I would regret if I did not do so and lost by a score like 0.001. I did not expect such small differences in the prize zone. No reason using Catboost, it was simply the first gradient boosting library I saw in a public kernel. 🐣</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 542981,
          "author_name": "DevilEars",
          "author_url": "",
          "post_date": "2019-06-04T08:50:37.843000",
          "content": "<p>I tried a few gradient boosters with my kernel template. For reference, I used this to give insight into the different gradient boosters:\n<a href=\"https://towardsdatascience.com/catboost-vs-light-gbm-vs-xgboost-5f93620723db\">CatBoost vs Light GBM vs XGBoost</a>.</p>\n\n<p>XGBoost performed the best. With the same features, and only changing the booster, Cat Boost achieved 2.34994, while XGBoost managed a public score of 1.60521.</p>\n\n<p>Hope this helps!</p>",
          "votes": -1,
          "replies": []
        }
      ]
    },
    {
      "id": 543990,
      "author_name": "Beans",
      "author_url": "",
      "post_date": "2019-06-05T02:32:34.943000",
      "content": "<p>Congratulation for the 2nd place and sharing your thoughts, this post truly gave me a deep understanding!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 543864,
      "author_name": "Ziteng Pang",
      "author_url": "",
      "post_date": "2019-06-04T22:29:38.887000",
      "content": "<p>Congrats on the result. I'm totally new to this and I really learn a lot from your post.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 543081,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2019-06-04T10:16:16.430000",
      "content": "<p>Congrats on the result and the method!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 543580,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2019-06-04T15:51:09.203000",
      "content": "<p>Did you tried Lightgbm or Xgboost on your solution ?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 543975,
          "author_name": "🐢 Jun Koda",
          "author_url": "",
          "post_date": "2019-06-05T02:11:27.957000",
          "content": "<p>No. I have installed them..., but I mistakenly remember the deadline for 1 week and didn't have time for the stacking phase. Default parameter ​CatBoost single model 😅</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 543227,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2019-06-04T11:53:08.093000",
      "content": "<p>Congrats! \nThe idea of tweaking the training data to create virtual test set is interesting.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 542895,
      "author_name": "DevilEars",
      "author_url": "",
      "post_date": "2019-06-04T07:08:59.007000",
      "content": "<p>Congratulations! Thank you for the update and for explaining your solution!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 542869,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2019-06-04T06:44:00.133000",
      "content": "<p>Well done!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 542850,
      "author_name": "Timmmmmms",
      "author_url": "",
      "post_date": "2019-06-04T06:32:05.450000",
      "content": "<p>Congrats!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 547184,
      "author_name": "rafzy",
      "author_url": "",
      "post_date": "2019-06-07T11:28:59.483000",
      "content": "<p>Congratulation for the solution and 2nd place ! \nWhat exactly means that \"14-sec segment and shifted up by 2 seconds\" ? You just simply added 2 sec or did you do something more fancy ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 547200,
          "author_name": "🐢 Jun Koda",
          "author_url": "",
          "post_date": "2019-06-07T11:55:39.273000",
          "content": "<p>Thanks. Exaclty​, just added 2.* something to the time to failure.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 545951,
      "author_name": "Ricard Delgado",
      "author_url": "",
      "post_date": "2019-06-06T04:42:03.473000",
      "content": "<p>Congratulations on the position and thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 543982,
      "author_name": "Massoud Hosseinali",
      "author_url": "",
      "post_date": "2019-06-05T02:16:21.763000",
      "content": "<p>Congratulations <a href=\"/junkoda\">@junkoda</a>. Thank you for sharing. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 543579,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2019-06-04T15:50:12.727000",
      "content": "<p>Congrats <a href=\"/junkoda\">@junkoda</a> and thanks for sharing. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 544077,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-05T05:14:58.200000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 542968,
      "author_name": "Stanislav Blinov",
      "author_url": "",
      "post_date": "2019-06-04T08:40:52.090000",
      "content": "<p>Thanks for sharing! Good job!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 542945,
      "author_name": "Wonho Song",
      "author_url": "",
      "post_date": "2019-06-04T08:11:57.570000",
      "content": "<p>Congrats and thanks for sharing !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 543568,
      "author_name": "Eric Vos",
      "author_url": "",
      "post_date": "2019-06-04T15:45:04.120000",
      "content": "<p>Great job!\nThanks for sharing.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 543124,
      "author_name": "Karan Jakhar",
      "author_url": "",
      "post_date": "2019-06-04T10:42:27.667000",
      "content": "<p>Great work! Thanks for sharing.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 543087,
      "author_name": "mocobt",
      "author_url": "",
      "post_date": "2019-06-04T10:19:20.613000",
      "content": "<p>Thank you for sharing.\nCongratulations!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 543043,
      "author_name": "Prashanth Thangavel",
      "author_url": "",
      "post_date": "2019-06-04T09:46:06.583000",
      "content": "<p>Congrats! Thanks for sharing :)</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "542827": "Update: Feature importance label fixed (5th June GMT 3 am) \n\nI select segments from the traning ​​data and create a private-set-like data based on the inter-earthquake times seen in the figure \"Are data from p4677?\"\n\nmykper: https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90664#latest-535844\n\n![quake_private_like_set](https://junkoda.github.io/figs/quake/quake_final_private_like.png)\n\nThe dashed line is 11.5 sec.\n\nThe first segment is counted twice; the private set has 8 segments. The last segment is created from the 14-sec segment and shifted up by 2 seconds, and the last 4.5 second is cut, treating that as the public part.\n\nI use 32 features:\n\n- standard deviation (std)\n- std_nopeak: Std of data that are not part of peaks\n- kurtosis_truncated: Kurtosis of data with abs(v - mean(v)) &lt; 20\n- 7 peak counts: Number of peaks with hight &gt; 50, 75, ..., 200\n- 5 percentiles: 95 percentile - 5 percentile, 80 - 20, 70 - 30, 60 - 40.\n- trend: slope of robust linear regression to 30 sub chunks of std_truncated\n- trend_error: Abs difference in the slope of RANSAC and Huber fit\n- power spectrum; Fast-Fourier Transform the data and average the absolute value in 15 bins\n\nI choose combinations such as 95 percentile - 5 percentile to avoid direct dependence on the mean; which is drifting with time. Same for the peak height; the peak height is defined as (max - min)/2.\n\nThe std_truncated (std instead of kurtosis in kurtosis_truncated) works almost as well as std_nopeak.\n\n\nI randomly select 2000x1000 training chunks of length 150_000 from my private-like set, which is 1000 times the number of independent/non-overlapping chunks and put all of them into CatBoost. I do not provide CV data to the regressor; eveything ​is the training set.\n\n\nI also tried to predict time since failure using all the training set and tried to stitch together with time to failure, but I was not able to do that successfully.\n\nThis is the feature importance:\n\n![Feature importance](https://junkoda.github.io/figs/quake/quake_feature_importance.png)",
    "564291": "At first heartily congrats @junkoda  and thanks a ton for sharing the solution ",
    "542964": "Сongratulations!\nHow did you select these features? What was your feature selection/feature generation approach?",
    "542912": "Congrats! You did well also in [VSB Power Line Fault Detection](https://www.kaggle.com/c/vsb-power-line-fault-detection). I feel you're really great because you always make such robust model. \n\nI'd like to ask one thing: Have you tried other machine learning algorithm like LightGBM/XGboost? If there is a reason why you use Catboost, I want to know.",
    "543990": "Congratulation for the 2nd place and sharing your thoughts, this post truly gave me a deep understanding!",
    "543864": "Congrats on the result. I'm totally new to this and I really learn a lot from your post.",
    "543081": "Congrats on the result and the method!",
    "543580": "Did you tried Lightgbm or Xgboost on your solution ?",
    "543227": "Congrats! \nThe idea of tweaking the training data to create virtual test set is interesting.",
    "542895": "Congratulations! Thank you for the update and for explaining your solution!",
    "542869": "Well done!",
    "542850": "Congrats!",
    "547184": "Congratulation for the solution and 2nd place ! \nWhat exactly means that \"14-sec segment and shifted up by 2 seconds\" ? You just simply added 2 sec or did you do something more fancy ?",
    "545951": "Congratulations on the position and thanks for sharing!",
    "543982": "Congratulations @junkoda. Thank you for sharing. ",
    "543579": "Congrats @junkoda and thanks for sharing. ",
    "544077": "",
    "542968": "Thanks for sharing! Good job!",
    "542945": "Congrats and thanks for sharing !",
    "543568": "Great job!\nThanks for sharing.",
    "543124": "Great work! Thanks for sharing.",
    "543087": "Thank you for sharing.\nCongratulations!",
    "543043": "Congrats! Thanks for sharing :)"
  }
}