{
  "id": 94467,
  "title": "10th place solution (Amjad's view)",
  "url": "/competitions/LANL-Earthquake-Prediction/writeups/shakinglab-10th-place-solution-amjad-s-view",
  "author_name": "",
  "post_date": "2019-09-11T14:53:21.057Z",
  "votes": 19,
  "comment_count": 10,
  "views": 0,
  "content": "<p>First of all, I would like to thank the organizers for providing an interesting research problem (even though badly designed) and kaggle for providing the computational resources that I used throughout this competition. My biggest thanks of course go to Giba who put his magical touch on my models. By the way, take a look at his views <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94466#latest-543723\">here</a>. You can check out my kernel <a href=\"https://www.kaggle.com/amjad85/10th-place-feature-engineering?scriptVersionId=15179664\">here</a>, and Giba’s kernel that uses it <a href=\"https://www.kaggle.com/titericz/top-1-lb-2-254-private\">here</a>.</p>\n\n<p><strong>Preprocessing:</strong></p>\n\n<p>I used denoising with discrete wavelet transform (DWT). However, the wavelet length was important here. Short wavelets like Haar filter too much useful information. I used  the best localized wavelet 14 (l14 from wmtsa R package).</p>\n\n<p><strong>Feature generation</strong></p>\n\n<p>Since the early days, one feature stood out, that is zero-crossing. So I generalized this feature to compute the rate of crossing points at various quantile levels of acoustic data. I computed crossing points features on several levels of DWT decomposed signals (started with 5 levels and went up to 10). Then I expanded the one DWT decomposition to three wavelets with different length. In addition to that, I added features representing the mean and variance at rolling windows of the spectrum of the signal. I also had some features counting peaks, which I took from public kernels and translated to R, and some autocorrelation features from the tsfeatures R package. I had also statistical features computed on different levels of DWT, but we didn’t end up using those given their different distribution.</p>\n\n<p><strong>Cross validation</strong></p>\n\n<p>I used shuffled KFold stratified by both time and quakes to enhance the features. This had the lowest variance among different strategies, and correlated well with public LB. I kept an eye open on other cross-validation schemes to know what causes overfitting. Towards the end, I also implemented nested cross validation, which is more robust than a simple leave-one-quake out cross-validation.</p>\n\n<p><strong>Model training</strong></p>\n\n<p>I started with Catboost at the beginning and regretted it given how slow it is. Once I switched to LGB, I immediately jumped the public LB.</p>\n\n<p><strong>The twist</strong></p>\n\n<p>Transferring the prediction to test distribution was pulled out by Giba at the last moment. See his post for explanation.</p>\n\n<p><strong>Thing that didn’t work for me</strong></p>\n\n<ul>\n<li>Classifying long vs short quakes</li>\n<li>Augmentation with overlapping segments</li>\n<li>MFCC features, rolling features, skew, kurtosis</li>\n<li>Deep learning. My best model was 2-channel 1d-CNN (mean and sd) – bi-GRU scored 1.499 on public LB.</li>\n</ul>\n\n<p><strong>My tips to beginners</strong></p>\n\n<p>I’m not new to machine learning, but this competition was my first project with gradient boosting, so it was a great learning experience. Also I’m pretty much new to kaggle competitions. I participated in one some years ago and it was a struggle on my crappy laptop. Now it’s a much better experience. \nHere are some tips to other beginners:\n-   Spend some time on making your code runs faster and easy to develop\n-   Start with a ML algorithm that runs fast for development\n-       Spend some time implementing your cross-validation</p>",
  "messages": [
    {
      "id": "543680",
      "postDate": "06/04/2019 17:38:18",
      "content": "<p>First of all, I would like to thank the organizers for providing an interesting research problem (even though badly designed) and kaggle for providing the computational resources that I used throughout this competition. My biggest thanks of course go to Giba who put his magical touch on my models. By the way, take a look at his views <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94466#latest-543723\">here</a>. You can check out my kernel <a href=\"https://www.kaggle.com/amjad85/10th-place-feature-engineering?scriptVersionId=15179664\">here</a>, and Giba’s kernel that uses it <a href=\"https://www.kaggle.com/titericz/top-1-lb-2-254-private\">here</a>.</p>\n\n<p><strong>Preprocessing:</strong></p>\n\n<p>I used denoising with discrete wavelet transform (DWT). However, the wavelet length was important here. Short wavelets like Haar filter too much useful information. I used  the best localized wavelet 14 (l14 from wmtsa R package).</p>\n\n<p><strong>Feature generation</strong></p>\n\n<p>Since the early days, one feature stood out, that is zero-crossing. So I generalized this feature to compute the rate of crossing points at various quantile levels of acoustic data. I computed crossing points features on several levels of DWT decomposed signals (started with 5 levels and went up to 10). Then I expanded the one DWT decomposition to three wavelets with different length. In addition to that, I added features representing the mean and variance at rolling windows of the spectrum of the signal. I also had some features counting peaks, which I took from public kernels and translated to R, and some autocorrelation features from the tsfeatures R package. I had also statistical features computed on different levels of DWT, but we didn’t end up using those given their different distribution.</p>\n\n<p><strong>Cross validation</strong></p>\n\n<p>I used shuffled KFold stratified by both time and quakes to enhance the features. This had the lowest variance among different strategies, and correlated well with public LB. I kept an eye open on other cross-validation schemes to know what causes overfitting. Towards the end, I also implemented nested cross validation, which is more robust than a simple leave-one-quake out cross-validation.</p>\n\n<p><strong>Model training</strong></p>\n\n<p>I started with Catboost at the beginning and regretted it given how slow it is. Once I switched to LGB, I immediately jumped the public LB.</p>\n\n<p><strong>The twist</strong></p>\n\n<p>Transferring the prediction to test distribution was pulled out by Giba at the last moment. See his post for explanation.</p>\n\n<p><strong>Thing that didn’t work for me</strong></p>\n\n<ul>\n<li>Classifying long vs short quakes</li>\n<li>Augmentation with overlapping segments</li>\n<li>MFCC features, rolling features, skew, kurtosis</li>\n<li>Deep learning. My best model was 2-channel 1d-CNN (mean and sd) – bi-GRU scored 1.499 on public LB.</li>\n</ul>\n\n<p><strong>My tips to beginners</strong></p>\n\n<p>I’m not new to machine learning, but this competition was my first project with gradient boosting, so it was a great learning experience. Also I’m pretty much new to kaggle competitions. I participated in one some years ago and it was a struggle on my crappy laptop. Now it’s a much better experience. \nHere are some tips to other beginners:\n-   Spend some time on making your code runs faster and easy to develop\n-   Start with a ML algorithm that runs fast for development\n-       Spend some time implementing your cross-validation</p>",
      "rawMarkdown": "First of all, I would like to thank the organizers for providing an interesting research problem (even though badly designed) and kaggle for providing the computational resources that I used throughout this competition. My biggest thanks of course go to Giba who put his magical touch on my models. By the way, take a look at his views [here](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94466#latest-543723). You can check out my kernel [here](https://www.kaggle.com/amjad85/10th-place-feature-engineering?scriptVersionId=15179664), and Giba’s kernel that uses it [here](https://www.kaggle.com/titericz/top-1-lb-2-254-private).\n\n**Preprocessing:**\n\nI used denoising with discrete wavelet transform (DWT). However, the wavelet length was important here. Short wavelets like Haar filter too much useful information. I used  the best localized wavelet 14 (l14 from wmtsa R package).\n\n**Feature generation**\n\nSince the early days, one feature stood out, that is zero-crossing. So I generalized this feature to compute the rate of crossing points at various quantile levels of acoustic data. I computed crossing points features on several levels of DWT decomposed signals (started with 5 levels and went up to 10). Then I expanded the one DWT decomposition to three wavelets with different length. In addition to that, I added features representing the mean and variance at rolling windows of the spectrum of the signal. I also had some features counting peaks, which I took from public kernels and translated to R, and some autocorrelation features from the tsfeatures R package. I had also statistical features computed on different levels of DWT, but we didn’t end up using those given their different distribution.\n\n**Cross validation**\n\nI used shuffled KFold stratified by both time and quakes to enhance the features. This had the lowest variance among different strategies, and correlated well with public LB. I kept an eye open on other cross-validation schemes to know what causes overfitting. Towards the end, I also implemented nested cross validation, which is more robust than a simple leave-one-quake out cross-validation.\n\n**Model training**\n\nI started with Catboost at the beginning and regretted it given how slow it is. Once I switched to LGB, I immediately jumped the public LB.\n\n**The twist**\n\nTransferring the prediction to test distribution was pulled out by Giba at the last moment. See his post for explanation.\n\n**Thing that didn’t work for me**\n\n-\tClassifying long vs short quakes\n-\tAugmentation with overlapping segments\n-\tMFCC features, rolling features, skew, kurtosis\n-\tDeep learning. My best model was 2-channel 1d-CNN (mean and sd) – bi-GRU scored 1.499 on public LB.\n\n**My tips to beginners**\n\nI’m not new to machine learning, but this competition was my first project with gradient boosting, so it was a great learning experience. Also I’m pretty much new to kaggle competitions. I participated in one some years ago and it was a struggle on my crappy laptop. Now it’s a much better experience. \nHere are some tips to other beginners:\n-\tSpend some time on making your code runs faster and easy to develop\n-\tStart with a ML algorithm that runs fast for development\n-       Spend some time implementing your cross-validation",
      "votes": null
    },
    {
      "id": "543687",
      "postDate": "06/04/2019 17:43:20",
      "content": "<p>Thank you for you nerves of steel with that last minute 1.9 Public LB submission ;)</p>",
      "rawMarkdown": "Thank you for you nerves of steel with that last minute 1.9 Public LB submission ;)",
      "votes": null
    },
    {
      "id": "543700",
      "postDate": "06/04/2019 17:55:04",
      "content": "<p>Thanks for sharing, interesting that wavelet transforms worked so well.  Congrats on the result for  a first competition!</p>",
      "rawMarkdown": "Thanks for sharing, interesting that wavelet transforms worked so well.  Congrats on the result for  a first competition!",
      "votes": null
    },
    {
      "id": "543706",
      "postDate": "06/04/2019 17:59:50",
      "content": "<p>I survived the heart attack :D</p>",
      "rawMarkdown": "I survived the heart attack :D",
      "votes": null
    },
    {
      "id": "543707",
      "postDate": "06/04/2019 18:00:57",
      "content": "<p>I joined this competition almost from the beginning, and I have to say that when you joined, you enriched the discussion in this forum with many interesting thoughts. Thank you for that.</p>",
      "rawMarkdown": "I joined this competition almost from the beginning, and I have to say that when you joined, you enriched the discussion in this forum with many interesting thoughts. Thank you for that.",
      "votes": null
    },
    {
      "id": "543728",
      "postDate": "06/04/2019 18:18:14",
      "content": "<p>I tried to classify long vs short quakes or EQs with mini-quakes. Had no luck as well.</p>",
      "rawMarkdown": "I tried to classify long vs short quakes or EQs with mini-quakes. Had no luck as well.",
      "votes": null
    },
    {
      "id": "543740",
      "postDate": "06/04/2019 18:42:17",
      "content": "<p>Interesting view, and good tips. Thanks for sharing, and congratulations!</p>",
      "rawMarkdown": "Interesting view, and good tips. Thanks for sharing, and congratulations!",
      "votes": null
    },
    {
      "id": "543751",
      "postDate": "06/04/2019 19:06:38",
      "content": "<p>Congratulations! Not bad at all as first competition!\nAlso great advices for beginners! 👌🏻</p>",
      "rawMarkdown": "Congratulations! Not bad at all as first competition!\nAlso great advices for beginners! 👌🏻",
      "votes": null
    },
    {
      "id": "547392",
      "postDate": "06/07/2019 16:51:17",
      "content": "<p>Congratz to 10th place, i was actually scared that you relied purely on shuffled cv and could drop a lot :D</p>\n\n<p>My best features were also crossings, however they're based on the gradient of the signal. Have you tried gradient-based features? I'd be really interested if this could improve some models.</p>\n\n<p>Feature creation is something along the lines of:</p>\n\n<p>1) get raw signal chunk x of size 150.000\n2) normalize to zero mean: \n     <code>x_norm = x - np.mean(x)</code>, actually has no impact on gradient, might influence other features\n3) filter it with wavelet/bandpass\n    <code>x_filt = some_filter_function(x_norm)</code>\n4) calculate the gradient of normalized, filtered signal:\n     <code>gradient = np.gradient(x_filt)</code>\n5) calculate mask, where gradient jumps above/below a certain threshold:\n     <code>grad_diff = (np.diff(np.sign(gradient - 0.50), n=1) != 0)</code>\n     Let's look at the equation: The gradient is shifted by a defined threshold (here e.g. -0.50).\n     The sign looks where this shifted gradient is above or below 0 and outputs 1s for above, 0 for equal 0 and \n     -1 for below 0)\n     The difference then outputs, where this mask jumps from gradients above a certain value to values below it. \n     I'm sure this can be done nicer with some numpy function, hope it's clear what's happening.\n6) calculate the distance between positions where gradient_diff mask changes direction:\n     <code>np.quantile((np.diff(np.where(grad_diff != 0))), .95)</code>\n     The inner part gives us the temporal distances in the gradient signal, where it jumps over thresholds. Finally, you <br>\n     can throw any statistical measure on this signal, for me high quantiles worked best.</p>\n\n<p>I interpret it kind of as a better version of peak distances. It's also related to fft-features, since high change in signal --&gt; high gradient and high fft-components.</p>\n\n<p>So in the end, you have 3 tuneable parameters (filter function, shift of gradient, quantile value). Would be cool if you or someone else had the time to try and maybe compare it to own crossing-features.</p>",
      "rawMarkdown": "Congratz to 10th place, i was actually scared that you relied purely on shuffled cv and could drop a lot :D\n\nMy best features were also crossings, however they're based on the gradient of the signal. Have you tried gradient-based features? I'd be really interested if this could improve some models.\n\nFeature creation is something along the lines of:\n\n1) get raw signal chunk x of size 150.000\n2) normalize to zero mean: \n     `x_norm = x - np.mean(x)`, actually has no impact on gradient, might influence other features\n3) filter it with wavelet/bandpass\n    ` x_filt = some_filter_function(x_norm)`\n4) calculate the gradient of normalized, filtered signal:\n     `gradient = np.gradient(x_filt)`\n5) calculate mask, where gradient jumps above/below a certain threshold:\n     `grad_diff = (np.diff(np.sign(gradient - 0.50), n=1) != 0)`\n     Let's look at the equation: The gradient is shifted by a defined threshold (here e.g. -0.50).\n     The sign looks where this shifted gradient is above or below 0 and outputs 1s for above, 0 for equal 0 and \n     -1 for below 0)\n     The difference then outputs, where this mask jumps from gradients above a certain value to values below it. \n     I'm sure this can be done nicer with some numpy function, hope it's clear what's happening.\n6) calculate the distance between positions where gradient_diff mask changes direction:\n     `np.quantile((np.diff(np.where(grad_diff != 0))), .95)`\n     The inner part gives us the temporal distances in the gradient signal, where it jumps over thresholds. Finally, you   \n     can throw any statistical measure on this signal, for me high quantiles worked best.\n\nI interpret it kind of as a better version of peak distances. It's also related to fft-features, since high change in signal --&gt; high gradient and high fft-components.\n\nSo in the end, you have 3 tuneable parameters (filter function, shift of gradient, quantile value). Would be cool if you or someone else had the time to try and maybe compare it to own crossing-features.",
      "votes": null
    },
    {
      "id": "547433",
      "postDate": "06/07/2019 17:37:09",
      "content": "<p>Thanks. Shuffled CV was not bad, it only gives over-optimistic estimates of the error. \nRegarding your approach, I don't think I had similar features, but if you like you can compare with my features from my kernel: <a href=\"https://www.kaggle.com/amjad85/10th-place-feature-engineering\">https://www.kaggle.com/amjad85/10th-place-feature-engineering</a></p>",
      "rawMarkdown": "Thanks. Shuffled CV was not bad, it only gives over-optimistic estimates of the error. \nRegarding your approach, I don't think I had similar features, but if you like you can compare with my features from my kernel: https://www.kaggle.com/amjad85/10th-place-feature-engineering",
      "votes": null
    },
    {
      "id": "547452",
      "postDate": "06/07/2019 18:24:53",
      "content": "<p>Okay i'll have a look at it. I think i described it a litte too complex, it's basically a vector containing the distances between zero-crossing in the filtered gradient, only that you can shift it to be crossings with 1, 5, 10, ...\nI'm sure the is a function somewhere that can do this as a one-liner, i just didn't search for it when i initially came up with it.</p>\n\n<p>So if the signal is close to the quake, you'll have a lot of high gradients --&gt; smaller distances --&gt; smaller quantiles. It just worked better for me than than zero-crossings of the acoustic signal so i thought i'd share, but i had both in my models.</p>",
      "rawMarkdown": "Okay i'll have a look at it. I think i described it a litte too complex, it's basically a vector containing the distances between zero-crossing in the filtered gradient, only that you can shift it to be crossings with 1, 5, 10, ...\nI'm sure the is a function somewhere that can do this as a one-liner, i just didn't search for it when i initially came up with it.\n\nSo if the signal is close to the quake, you'll have a lot of high gradients --&gt; smaller distances --&gt; smaller quantiles. It just worked better for me than than zero-crossings of the acoustic signal so i thought i'd share, but i had both in my models.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 543687,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "06/04/2019 17:43:20",
      "content": "<p>Thank you for you nerves of steel with that last minute 1.9 Public LB submission ;)</p>",
      "votes": null,
      "replies": [
        {
          "id": 543706,
          "author_name": "amjad85",
          "author_url": "",
          "post_date": "06/04/2019 17:59:50",
          "content": "<p>I survived the heart attack :D</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 543700,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/04/2019 17:55:04",
      "content": "<p>Thanks for sharing, interesting that wavelet transforms worked so well.  Congrats on the result for  a first competition!</p>",
      "votes": null,
      "replies": [
        {
          "id": 543707,
          "author_name": "amjad85",
          "author_url": "",
          "post_date": "06/04/2019 18:00:57",
          "content": "<p>I joined this competition almost from the beginning, and I have to say that when you joined, you enriched the discussion in this forum with many interesting thoughts. Thank you for that.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 543728,
      "author_name": "teeyee314",
      "author_url": "",
      "post_date": "06/04/2019 18:18:14",
      "content": "<p>I tried to classify long vs short quakes or EQs with mini-quakes. Had no luck as well.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543740,
      "author_name": "miguelpm",
      "author_url": "",
      "post_date": "06/04/2019 18:42:17",
      "content": "<p>Interesting view, and good tips. Thanks for sharing, and congratulations!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543751,
      "author_name": "stecasasso",
      "author_url": "",
      "post_date": "06/04/2019 19:06:38",
      "content": "<p>Congratulations! Not bad at all as first competition!\nAlso great advices for beginners! 👌🏻</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 547392,
      "author_name": "svenhinderer",
      "author_url": "",
      "post_date": "06/07/2019 16:51:17",
      "content": "<p>Congratz to 10th place, i was actually scared that you relied purely on shuffled cv and could drop a lot :D</p>\n\n<p>My best features were also crossings, however they're based on the gradient of the signal. Have you tried gradient-based features? I'd be really interested if this could improve some models.</p>\n\n<p>Feature creation is something along the lines of:</p>\n\n<p>1) get raw signal chunk x of size 150.000\n2) normalize to zero mean: \n     <code>x_norm = x - np.mean(x)</code>, actually has no impact on gradient, might influence other features\n3) filter it with wavelet/bandpass\n    <code>x_filt = some_filter_function(x_norm)</code>\n4) calculate the gradient of normalized, filtered signal:\n     <code>gradient = np.gradient(x_filt)</code>\n5) calculate mask, where gradient jumps above/below a certain threshold:\n     <code>grad_diff = (np.diff(np.sign(gradient - 0.50), n=1) != 0)</code>\n     Let's look at the equation: The gradient is shifted by a defined threshold (here e.g. -0.50).\n     The sign looks where this shifted gradient is above or below 0 and outputs 1s for above, 0 for equal 0 and \n     -1 for below 0)\n     The difference then outputs, where this mask jumps from gradients above a certain value to values below it. \n     I'm sure this can be done nicer with some numpy function, hope it's clear what's happening.\n6) calculate the distance between positions where gradient_diff mask changes direction:\n     <code>np.quantile((np.diff(np.where(grad_diff != 0))), .95)</code>\n     The inner part gives us the temporal distances in the gradient signal, where it jumps over thresholds. Finally, you <br>\n     can throw any statistical measure on this signal, for me high quantiles worked best.</p>\n\n<p>I interpret it kind of as a better version of peak distances. It's also related to fft-features, since high change in signal --&gt; high gradient and high fft-components.</p>\n\n<p>So in the end, you have 3 tuneable parameters (filter function, shift of gradient, quantile value). Would be cool if you or someone else had the time to try and maybe compare it to own crossing-features.</p>",
      "votes": null,
      "replies": [
        {
          "id": 547433,
          "author_name": "amjad85",
          "author_url": "",
          "post_date": "06/07/2019 17:37:09",
          "content": "<p>Thanks. Shuffled CV was not bad, it only gives over-optimistic estimates of the error. \nRegarding your approach, I don't think I had similar features, but if you like you can compare with my features from my kernel: <a href=\"https://www.kaggle.com/amjad85/10th-place-feature-engineering\">https://www.kaggle.com/amjad85/10th-place-feature-engineering</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 547452,
          "author_name": "svenhinderer",
          "author_url": "",
          "post_date": "06/07/2019 18:24:53",
          "content": "<p>Okay i'll have a look at it. I think i described it a litte too complex, it's basically a vector containing the distances between zero-crossing in the filtered gradient, only that you can shift it to be crossings with 1, 5, 10, ...\nI'm sure the is a function somewhere that can do this as a one-liner, i just didn't search for it when i initially came up with it.</p>\n\n<p>So if the signal is close to the quake, you'll have a lot of high gradients --&gt; smaller distances --&gt; smaller quantiles. It just worked better for me than than zero-crossings of the acoustic signal so i thought i'd share, but i had both in my models.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "543680": "First of all, I would like to thank the organizers for providing an interesting research problem (even though badly designed) and kaggle for providing the computational resources that I used throughout this competition. My biggest thanks of course go to Giba who put his magical touch on my models. By the way, take a look at his views [here](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94466#latest-543723). You can check out my kernel [here](https://www.kaggle.com/amjad85/10th-place-feature-engineering?scriptVersionId=15179664), and Giba’s kernel that uses it [here](https://www.kaggle.com/titericz/top-1-lb-2-254-private).\n\n**Preprocessing:**\n\nI used denoising with discrete wavelet transform (DWT). However, the wavelet length was important here. Short wavelets like Haar filter too much useful information. I used  the best localized wavelet 14 (l14 from wmtsa R package).\n\n**Feature generation**\n\nSince the early days, one feature stood out, that is zero-crossing. So I generalized this feature to compute the rate of crossing points at various quantile levels of acoustic data. I computed crossing points features on several levels of DWT decomposed signals (started with 5 levels and went up to 10). Then I expanded the one DWT decomposition to three wavelets with different length. In addition to that, I added features representing the mean and variance at rolling windows of the spectrum of the signal. I also had some features counting peaks, which I took from public kernels and translated to R, and some autocorrelation features from the tsfeatures R package. I had also statistical features computed on different levels of DWT, but we didn’t end up using those given their different distribution.\n\n**Cross validation**\n\nI used shuffled KFold stratified by both time and quakes to enhance the features. This had the lowest variance among different strategies, and correlated well with public LB. I kept an eye open on other cross-validation schemes to know what causes overfitting. Towards the end, I also implemented nested cross validation, which is more robust than a simple leave-one-quake out cross-validation.\n\n**Model training**\n\nI started with Catboost at the beginning and regretted it given how slow it is. Once I switched to LGB, I immediately jumped the public LB.\n\n**The twist**\n\nTransferring the prediction to test distribution was pulled out by Giba at the last moment. See his post for explanation.\n\n**Thing that didn’t work for me**\n\n-\tClassifying long vs short quakes\n-\tAugmentation with overlapping segments\n-\tMFCC features, rolling features, skew, kurtosis\n-\tDeep learning. My best model was 2-channel 1d-CNN (mean and sd) – bi-GRU scored 1.499 on public LB.\n\n**My tips to beginners**\n\nI’m not new to machine learning, but this competition was my first project with gradient boosting, so it was a great learning experience. Also I’m pretty much new to kaggle competitions. I participated in one some years ago and it was a struggle on my crappy laptop. Now it’s a much better experience. \nHere are some tips to other beginners:\n-\tSpend some time on making your code runs faster and easy to develop\n-\tStart with a ML algorithm that runs fast for development\n-       Spend some time implementing your cross-validation",
    "543687": "Thank you for you nerves of steel with that last minute 1.9 Public LB submission ;)",
    "543700": "Thanks for sharing, interesting that wavelet transforms worked so well.  Congrats on the result for  a first competition!",
    "543706": "I survived the heart attack :D",
    "543707": "I joined this competition almost from the beginning, and I have to say that when you joined, you enriched the discussion in this forum with many interesting thoughts. Thank you for that.",
    "543728": "I tried to classify long vs short quakes or EQs with mini-quakes. Had no luck as well.",
    "543740": "Interesting view, and good tips. Thanks for sharing, and congratulations!",
    "543751": "Congratulations! Not bad at all as first competition!\nAlso great advices for beginners! 👌🏻",
    "547392": "Congratz to 10th place, i was actually scared that you relied purely on shuffled cv and could drop a lot :D\n\nMy best features were also crossings, however they're based on the gradient of the signal. Have you tried gradient-based features? I'd be really interested if this could improve some models.\n\nFeature creation is something along the lines of:\n\n1) get raw signal chunk x of size 150.000\n2) normalize to zero mean: \n     `x_norm = x - np.mean(x)`, actually has no impact on gradient, might influence other features\n3) filter it with wavelet/bandpass\n    ` x_filt = some_filter_function(x_norm)`\n4) calculate the gradient of normalized, filtered signal:\n     `gradient = np.gradient(x_filt)`\n5) calculate mask, where gradient jumps above/below a certain threshold:\n     `grad_diff = (np.diff(np.sign(gradient - 0.50), n=1) != 0)`\n     Let's look at the equation: The gradient is shifted by a defined threshold (here e.g. -0.50).\n     The sign looks where this shifted gradient is above or below 0 and outputs 1s for above, 0 for equal 0 and \n     -1 for below 0)\n     The difference then outputs, where this mask jumps from gradients above a certain value to values below it. \n     I'm sure this can be done nicer with some numpy function, hope it's clear what's happening.\n6) calculate the distance between positions where gradient_diff mask changes direction:\n     `np.quantile((np.diff(np.where(grad_diff != 0))), .95)`\n     The inner part gives us the temporal distances in the gradient signal, where it jumps over thresholds. Finally, you   \n     can throw any statistical measure on this signal, for me high quantiles worked best.\n\nI interpret it kind of as a better version of peak distances. It's also related to fft-features, since high change in signal --&gt; high gradient and high fft-components.\n\nSo in the end, you have 3 tuneable parameters (filter function, shift of gradient, quantile value). Would be cool if you or someone else had the time to try and maybe compare it to own crossing-features.",
    "547433": "Thanks. Shuffled CV was not bad, it only gives over-optimistic estimates of the error. \nRegarding your approach, I don't think I had similar features, but if you like you can compare with my features from my kernel: https://www.kaggle.com/amjad85/10th-place-feature-engineering",
    "547452": "Okay i'll have a look at it. I think i described it a litte too complex, it's basically a vector containing the distances between zero-crossing in the filtered gradient, only that you can shift it to be crossings with 1, 5, 10, ...\nI'm sure the is a function somewhere that can do this as a one-liner, i just didn't search for it when i initially came up with it.\n\nSo if the signal is close to the quake, you'll have a lot of high gradients --&gt; smaller distances --&gt; smaller quantiles. It just worked better for me than than zero-crossings of the acoustic signal so i thought i'd share, but i had both in my models."
  },
  "source": "meta"
}