{
  "id": 92367,
  "title": "Worst feature I found",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/92367",
  "author_name": "",
  "post_date": "2019-05-15T19:07:26.572630200Z",
  "votes": 14,
  "comment_count": 20,
  "views": 0,
  "content": "<p>While doing feature engineering, I stumbled upon this feature that may convince an ML model that it's useful (especially if you use shuffled KFold split), but when it comes to the test set, it will most likely lead to overfitting. </p>\n\n<p>The feature is a simple median value of the signal filtered with a low pass Butterworth filter. If you look at it, it has a noise component and a slowly changing component that an ML model may pick up and try using for making predictions. But it appears that both components change randomly without carrying any useful information, making this feature not just useless, but actually harmful. </p>\n\n<p>So, in case you are using something like that, it's probably a good idea to remove this feature.</p>\n\n<p><img src=\"https://i.ibb.co/MkS1CNQ/download-12.png\" alt=\"median of a low passed signal\"></p>",
  "messages": [
    {
      "id": "531906",
      "postDate": "05/15/2019 19:07:26",
      "content": "<p>While doing feature engineering, I stumbled upon this feature that may convince an ML model that it's useful (especially if you use shuffled KFold split), but when it comes to the test set, it will most likely lead to overfitting. </p>\n\n<p>The feature is a simple median value of the signal filtered with a low pass Butterworth filter. If you look at it, it has a noise component and a slowly changing component that an ML model may pick up and try using for making predictions. But it appears that both components change randomly without carrying any useful information, making this feature not just useless, but actually harmful. </p>\n\n<p>So, in case you are using something like that, it's probably a good idea to remove this feature.</p>\n\n<p><img src=\"https://i.ibb.co/MkS1CNQ/download-12.png\" alt=\"median of a low passed signal\"></p>",
      "rawMarkdown": "While doing feature engineering, I stumbled upon this feature that may convince an ML model that it's useful (especially if you use shuffled KFold split), but when it comes to the test set, it will most likely lead to overfitting. \n\nThe feature is a simple median value of the signal filtered with a low pass Butterworth filter. If you look at it, it has a noise component and a slowly changing component that an ML model may pick up and try using for making predictions. But it appears that both components change randomly without carrying any useful information, making this feature not just useless, but actually harmful. \n\nSo, in case you are using something like that, it's probably a good idea to remove this feature.\n\n![median of a low passed signal](https://i.ibb.co/MkS1CNQ/download-12.png)",
      "votes": null
    },
    {
      "id": "531914",
      "postDate": "05/15/2019 19:36:14",
      "content": "<p>What is the cut-off frequency of the low pass filter?</p>",
      "rawMarkdown": "What is the cut-off frequency of the low pass filter?",
      "votes": null
    },
    {
      "id": "531929",
      "postDate": "05/15/2019 20:36:22",
      "content": "<p>1000/75000. Here is the function I'm using for filtering (adapted from Vettejeep):\n```\ndef butter_filter(x, n=4, low=None, high=None, nyquist_f=75000):\n    if low is None and high is None:\n        return\n    elif low is None:\n        b, a = sg.butter(n, Wn=high/nyquist_f, btype='lowpass')\n    elif high is None:\n        b, a = sg.butter(n, Wn=low/nyquist_f, btype='highpass')\n    else:\n        b, a = sg.butter(4, Wn=(low/nyquist_f, high/nyquist_f), btype='bandpass')</p>\n\n<pre><code>return sg.lfilter(b, a, x)\n</code></pre>\n\n<p>```</p>\n\n<p>Basically, you can call this function on the 150000-long chunk of data with high=1000:\n<code>x_low_pass = butter_filter(x, high=1000)</code></p>",
      "rawMarkdown": "1000/75000. Here is the function I'm using for filtering (adapted from Vettejeep):\n```\ndef butter_filter(x, n=4, low=None, high=None, nyquist_f=75000):\n    if low is None and high is None:\n        return\n    elif low is None:\n        b, a = sg.butter(n, Wn=high/nyquist_f, btype='lowpass')\n    elif high is None:\n        b, a = sg.butter(n, Wn=low/nyquist_f, btype='highpass')\n    else:\n        b, a = sg.butter(4, Wn=(low/nyquist_f, high/nyquist_f), btype='bandpass')\n        \n    return sg.lfilter(b, a, x)\n```\n\nBasically, you can call this function on the 150000-long chunk of data with high=1000:\n`x_low_pass = butter_filter(x, high=1000)`",
      "votes": null
    },
    {
      "id": "531955",
      "postDate": "05/15/2019 23:16:25",
      "content": "<p>The mean and median are likely to be artefacts of the recording instrumentation. I would be very wary of any features that rely on them, particularly since the distribution for the test set is very different.</p>",
      "rawMarkdown": "The mean and median are likely to be artefacts of the recording instrumentation. I would be very wary of any features that rely on them, particularly since the distribution for the test set is very different.",
      "votes": null
    },
    {
      "id": "532059",
      "postDate": "05/16/2019 05:52:55",
      "content": "<p>Looks like you use a very complex way to compute the mean of acoustic data ;)  </p>\n\n<p>Its distribution is different between train and test.</p>",
      "rawMarkdown": "Looks like you use a very complex way to compute the mean of acoustic data ;)  \n\nIts distribution is different between train and test.",
      "votes": null
    },
    {
      "id": "532083",
      "postDate": "05/16/2019 07:04:02",
      "content": "<p>Using it degrades my CV score, my ML mdoels aren't fooled by it for some reason ;)</p>",
      "rawMarkdown": "Using it degrades my CV score, my ML mdoels aren't fooled by it for some reason ;)",
      "votes": null
    },
    {
      "id": "532085",
      "postDate": "05/16/2019 07:05:50",
      "content": "<p>Maybe energy within a narrow band is more useful.</p>",
      "rawMarkdown": "Maybe energy within a narrow band is more useful.",
      "votes": null
    },
    {
      "id": "532086",
      "postDate": "05/16/2019 07:08:54",
      "content": "<p>Independently from machine learning, I am really surprised that they did not calibrate their sensor.  Acoustic data should have a 0 mean.</p>",
      "rawMarkdown": "Independently from machine learning, I am really surprised that they did not calibrate their sensor.  Acoustic data should have a 0 mean.",
      "votes": null
    },
    {
      "id": "532095",
      "postDate": "05/16/2019 07:25:03",
      "content": "<p>yeah, I agree that it's a complex way to compute the mean, I realized it after posting this. I was just making some other features using frequency filters and stumbled upon this one. I put it in my model out of curiosity, and it was only mildly fooled by it, its permutation importance on the training set was positive (so it used it for predictions), but negative on the validation set. So, I would've gotten rid of it anyway during feature selection.</p>",
      "rawMarkdown": "yeah, I agree that it's a complex way to compute the mean, I realized it after posting this. I was just making some other features using frequency filters and stumbled upon this one. I put it in my model out of curiosity, and it was only mildly fooled by it, its permutation importance on the training set was positive (so it used it for predictions), but negative on the validation set. So, I would've gotten rid of it anyway during feature selection.",
      "votes": null
    },
    {
      "id": "532104",
      "postDate": "05/16/2019 08:00:06",
      "content": "<p>I think it may be calibrated. We probably have just the small part of the entire recording, that can be low-frequency oscillation. </p>",
      "rawMarkdown": "I think it may be calibrated. We probably have just the small part of the entire recording, that can be low-frequency oscillation.",
      "votes": null
    },
    {
      "id": "532388",
      "postDate": "05/16/2019 19:49:38",
      "content": "<p>Wasn't there a drift in mean caused by the degradation of the material during the experiments? I think i saw something like that in the additional data discussions. </p>",
      "rawMarkdown": "Wasn't there a drift in mean caused by the degradation of the material during the experiments? I think i saw something like that in the additional data discussions.",
      "votes": null
    },
    {
      "id": "532882",
      "postDate": "05/18/2019 00:44:51",
      "content": "<p>It indeed is the mean. The following is distribution of mean vs ttf from <a href=\"https://www.kaggle.com/gpreda/lanl-earthquake-new-approach-eda\">https://www.kaggle.com/gpreda/lanl-earthquake-new-approach-eda</a></p>\n\n<p><img src=\"https://www.kaggleusercontent.com/kf/14309733/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..ctuIlZ7NH_on0xUbSmnQOg.ANrkULpie5oKS9HTUaQoiOZPo7g7vrx1Wb8GS-yMqYPfsCNNOdkrU-NKUbsB4L54KxKx8ftwwLxRuhLHa_ya5kFRTuInGQ_Gp1sn7K_PwdUc6tsJB3YNjhwy0Gv3nrdVOrBuQnUhNdaPlqGy4ahh-8-7S8hbJm55d9GJn6hvEHc.Zcoiw8HtSz_A0O_EqjyuNg/__results___files/__results___34_0.png\" alt=\"\"></p>",
      "rawMarkdown": "It indeed is the mean. The following is distribution of mean vs ttf from https://www.kaggle.com/gpreda/lanl-earthquake-new-approach-eda\n\n![](https://www.kaggleusercontent.com/kf/14309733/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..ctuIlZ7NH_on0xUbSmnQOg.ANrkULpie5oKS9HTUaQoiOZPo7g7vrx1Wb8GS-yMqYPfsCNNOdkrU-NKUbsB4L54KxKx8ftwwLxRuhLHa_ya5kFRTuInGQ_Gp1sn7K_PwdUc6tsJB3YNjhwy0Gv3nrdVOrBuQnUhNdaPlqGy4ahh-8-7S8hbJm55d9GJn6hvEHc.Zcoiw8HtSz_A0O_EqjyuNg/__results___files/__results___34_0.png)",
      "votes": null
    },
    {
      "id": "532883",
      "postDate": "05/18/2019 00:52:46",
      "content": "<p><a href=\"/cpmpml\">@cpmpml</a> I think this isn't due to calibration but because of the degradation of shear stresses over time and consequently the generated acoustic signal as <a href=\"/davids1992\">@davids1992</a> and <a href=\"/svenhinderer\">@svenhinderer</a> mentioned. To add to their comments, I'm gonna refer you to this figure from LANL's papers:\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/532883/13232/Untitled.jpg\" alt=\"\"></p>",
      "rawMarkdown": "cpmpml I think this isn't due to calibration but because of the degradation of shear stresses over time and consequently the generated acoustic signal as @davids1992 and @svenhinderer mentioned. To add to their comments, I'm gonna refer you to this figure from LANL's papers:\n![](https://storage.googleapis.com/kaggle-forum-message-attachments/532883/13232/Untitled.jpg)",
      "votes": null
    },
    {
      "id": "532885",
      "postDate": "05/18/2019 01:02:37",
      "content": "<p><a href=\"/batalov\">@batalov</a> I'm glad you noticed this. What is more interesting for me, is that most of the public kernels not only use <code>mean</code> but also <code>sum</code> that is <code>= mean*150000</code> too =&gt; correlation = 1; there are also tons of features with correlation &gt;= 0.95. They also have some features that is always 0 in all instances such as <code>min(abs(signal))</code>.</p>\n\n<p>What you found is ONE of the MANY worst features :D</p>",
      "rawMarkdown": "batalov I'm glad you noticed this. What is more interesting for me, is that most of the public kernels not only use `mean` but also `sum` that is `= mean*150000` too =&gt; correlation = 1; there are also tons of features with correlation &gt;= 0.95. They also have some features that is always 0 in all instances such as `min(abs(signal))`.\n\nWhat you found is ONE of the MANY worst features :D",
      "votes": null
    },
    {
      "id": "532959",
      "postDate": "05/18/2019 05:44:40",
      "content": "<p>Sum or mean is the same for gbms .  Why have both?</p>",
      "rawMarkdown": "Sum or mean is the same for gbms .  Why have both?",
      "votes": null
    },
    {
      "id": "532961",
      "postDate": "05/18/2019 05:49:57",
      "content": "<p>That was my point. Because apparently somebody generated summary statistics at first without even thinking about it and everyone else blindly copied it.</p>",
      "rawMarkdown": "That was my point. Because apparently somebody generated summary statistics at first without even thinking about it and everyone else blindly copied it.",
      "votes": null
    },
    {
      "id": "539031",
      "postDate": "05/29/2019 12:22:21",
      "content": "<p>This article is underestimated.\nI have one feature that relies on mean. Without it, my CV is worse, and my public LB is also worse, by 0.02! My model outputs it as one of the top features. That's really confused here. From the plot it has no meaning, but what if it can be combined with other features and produces good predictions?  Instead of the mean value itself (which is bad) in the global scope that is different in train/test distributions, can it be good locally within a sample when combined with other features?</p>",
      "rawMarkdown": "This article is underestimated.\nI have one feature that relies on mean. Without it, my CV is worse, and my public LB is also worse, by 0.02! My model outputs it as one of the top features. That's really confused here. From the plot it has no meaning, but what if it can be combined with other features and produces good predictions?  Instead of the mean value itself (which is bad) in the global scope that is different in train/test distributions, can it be good locally within a sample when combined with other features?",
      "votes": null
    },
    {
      "id": "539094",
      "postDate": "05/29/2019 14:00:09",
      "content": "<p>I was a little bit unhappy with my score, but now, when I see here and in \"LANL Adversarial Validation - Shakeup is Coming\" topic, an excellent score comes with a cost and risk. Maybe the real talent is to be able to choose well being in that place.</p>",
      "rawMarkdown": "I was a little bit unhappy with my score, but now, when I see here and in \"LANL Adversarial Validation - Shakeup is Coming\" topic, an excellent score comes with a cost and risk. Maybe the real talent is to be able to choose well being in that place.",
      "votes": null
    },
    {
      "id": "539106",
      "postDate": "05/29/2019 14:28:05",
      "content": "<p>I have given up on entirely on the leaderboard, I just don't think it is informative enough. People who are forking the new Master's Approach kernel to get a &lt; 1.4 score are going to be disappointed, I think. </p>",
      "rawMarkdown": "I have given up on entirely on the leaderboard, I just don't think it is informative enough. People who are forking the new Master's Approach kernel to get a &lt; 1.4 score are going to be disappointed, I think.",
      "votes": null
    },
    {
      "id": "539113",
      "postDate": "05/29/2019 14:40:42",
      "content": "<p>The thing is, my CV now seems to correlate very well with LB. That’s really puzzled.</p>",
      "rawMarkdown": "The thing is, my CV now seems to correlate very well with LB. That’s really puzzled.",
      "votes": null
    },
    {
      "id": "539140",
      "postDate": "05/29/2019 15:30:08",
      "content": "<p>I've finally found a CV strategy that I'm happy with, I'll be testing my final features in the next couple of days. It sounds like you engineered some really useful ones and a robust CV strategy so congratulations and good luck! </p>",
      "rawMarkdown": "I've finally found a CV strategy that I'm happy with, I'll be testing my final features in the next couple of days. It sounds like you engineered some really useful ones and a robust CV strategy so congratulations and good luck!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 531914,
      "author_name": "lucaskg",
      "author_url": "",
      "post_date": "05/15/2019 19:36:14",
      "content": "<p>What is the cut-off frequency of the low pass filter?</p>",
      "votes": null,
      "replies": [
        {
          "id": 531929,
          "author_name": "batalov",
          "author_url": "",
          "post_date": "05/15/2019 20:36:22",
          "content": "<p>1000/75000. Here is the function I'm using for filtering (adapted from Vettejeep):\n```\ndef butter_filter(x, n=4, low=None, high=None, nyquist_f=75000):\n    if low is None and high is None:\n        return\n    elif low is None:\n        b, a = sg.butter(n, Wn=high/nyquist_f, btype='lowpass')\n    elif high is None:\n        b, a = sg.butter(n, Wn=low/nyquist_f, btype='highpass')\n    else:\n        b, a = sg.butter(4, Wn=(low/nyquist_f, high/nyquist_f), btype='bandpass')</p>\n\n<pre><code>return sg.lfilter(b, a, x)\n</code></pre>\n\n<p>```</p>\n\n<p>Basically, you can call this function on the 150000-long chunk of data with high=1000:\n<code>x_low_pass = butter_filter(x, high=1000)</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 531955,
      "author_name": "bigironsphere",
      "author_url": "",
      "post_date": "05/15/2019 23:16:25",
      "content": "<p>The mean and median are likely to be artefacts of the recording instrumentation. I would be very wary of any features that rely on them, particularly since the distribution for the test set is very different.</p>",
      "votes": null,
      "replies": [
        {
          "id": 532086,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/16/2019 07:08:54",
          "content": "<p>Independently from machine learning, I am really surprised that they did not calibrate their sensor.  Acoustic data should have a 0 mean.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 532104,
          "author_name": "davids1992",
          "author_url": "",
          "post_date": "05/16/2019 08:00:06",
          "content": "<p>I think it may be calibrated. We probably have just the small part of the entire recording, that can be low-frequency oscillation. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 532388,
          "author_name": "svenhinderer",
          "author_url": "",
          "post_date": "05/16/2019 19:49:38",
          "content": "<p>Wasn't there a drift in mean caused by the degradation of the material during the experiments? I think i saw something like that in the additional data discussions. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 532883,
          "author_name": "mhviraf",
          "author_url": "",
          "post_date": "05/18/2019 00:52:46",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> I think this isn't due to calibration but because of the degradation of shear stresses over time and consequently the generated acoustic signal as <a href=\"/davids1992\">@davids1992</a> and <a href=\"/svenhinderer\">@svenhinderer</a> mentioned. To add to their comments, I'm gonna refer you to this figure from LANL's papers:\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/532883/13232/Untitled.jpg\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 532059,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "05/16/2019 05:52:55",
      "content": "<p>Looks like you use a very complex way to compute the mean of acoustic data ;)  </p>\n\n<p>Its distribution is different between train and test.</p>",
      "votes": null,
      "replies": [
        {
          "id": 532083,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/16/2019 07:04:02",
          "content": "<p>Using it degrades my CV score, my ML mdoels aren't fooled by it for some reason ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 532085,
          "author_name": "lucaskg",
          "author_url": "",
          "post_date": "05/16/2019 07:05:50",
          "content": "<p>Maybe energy within a narrow band is more useful.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 532095,
          "author_name": "batalov",
          "author_url": "",
          "post_date": "05/16/2019 07:25:03",
          "content": "<p>yeah, I agree that it's a complex way to compute the mean, I realized it after posting this. I was just making some other features using frequency filters and stumbled upon this one. I put it in my model out of curiosity, and it was only mildly fooled by it, its permutation importance on the training set was positive (so it used it for predictions), but negative on the validation set. So, I would've gotten rid of it anyway during feature selection.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 532882,
          "author_name": "mhviraf",
          "author_url": "",
          "post_date": "05/18/2019 00:44:51",
          "content": "<p>It indeed is the mean. The following is distribution of mean vs ttf from <a href=\"https://www.kaggle.com/gpreda/lanl-earthquake-new-approach-eda\">https://www.kaggle.com/gpreda/lanl-earthquake-new-approach-eda</a></p>\n\n<p><img src=\"https://www.kaggleusercontent.com/kf/14309733/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..ctuIlZ7NH_on0xUbSmnQOg.ANrkULpie5oKS9HTUaQoiOZPo7g7vrx1Wb8GS-yMqYPfsCNNOdkrU-NKUbsB4L54KxKx8ftwwLxRuhLHa_ya5kFRTuInGQ_Gp1sn7K_PwdUc6tsJB3YNjhwy0Gv3nrdVOrBuQnUhNdaPlqGy4ahh-8-7S8hbJm55d9GJn6hvEHc.Zcoiw8HtSz_A0O_EqjyuNg/__results___files/__results___34_0.png\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 532885,
      "author_name": "mhviraf",
      "author_url": "",
      "post_date": "05/18/2019 01:02:37",
      "content": "<p><a href=\"/batalov\">@batalov</a> I'm glad you noticed this. What is more interesting for me, is that most of the public kernels not only use <code>mean</code> but also <code>sum</code> that is <code>= mean*150000</code> too =&gt; correlation = 1; there are also tons of features with correlation &gt;= 0.95. They also have some features that is always 0 in all instances such as <code>min(abs(signal))</code>.</p>\n\n<p>What you found is ONE of the MANY worst features :D</p>",
      "votes": null,
      "replies": [
        {
          "id": 532959,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/18/2019 05:44:40",
          "content": "<p>Sum or mean is the same for gbms .  Why have both?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 532961,
          "author_name": "mhviraf",
          "author_url": "",
          "post_date": "05/18/2019 05:49:57",
          "content": "<p>That was my point. Because apparently somebody generated summary statistics at first without even thinking about it and everyone else blindly copied it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 539031,
      "author_name": "khahuras",
      "author_url": "",
      "post_date": "05/29/2019 12:22:21",
      "content": "<p>This article is underestimated.\nI have one feature that relies on mean. Without it, my CV is worse, and my public LB is also worse, by 0.02! My model outputs it as one of the top features. That's really confused here. From the plot it has no meaning, but what if it can be combined with other features and produces good predictions?  Instead of the mean value itself (which is bad) in the global scope that is different in train/test distributions, can it be good locally within a sample when combined with other features?</p>",
      "votes": null,
      "replies": [
        {
          "id": 539094,
          "author_name": "davids1992",
          "author_url": "",
          "post_date": "05/29/2019 14:00:09",
          "content": "<p>I was a little bit unhappy with my score, but now, when I see here and in \"LANL Adversarial Validation - Shakeup is Coming\" topic, an excellent score comes with a cost and risk. Maybe the real talent is to be able to choose well being in that place.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 539106,
          "author_name": "bigironsphere",
          "author_url": "",
          "post_date": "05/29/2019 14:28:05",
          "content": "<p>I have given up on entirely on the leaderboard, I just don't think it is informative enough. People who are forking the new Master's Approach kernel to get a &lt; 1.4 score are going to be disappointed, I think. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 539113,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "05/29/2019 14:40:42",
          "content": "<p>The thing is, my CV now seems to correlate very well with LB. That’s really puzzled.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 539140,
          "author_name": "bigironsphere",
          "author_url": "",
          "post_date": "05/29/2019 15:30:08",
          "content": "<p>I've finally found a CV strategy that I'm happy with, I'll be testing my final features in the next couple of days. It sounds like you engineered some really useful ones and a robust CV strategy so congratulations and good luck! </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "531906": "While doing feature engineering, I stumbled upon this feature that may convince an ML model that it's useful (especially if you use shuffled KFold split), but when it comes to the test set, it will most likely lead to overfitting. \n\nThe feature is a simple median value of the signal filtered with a low pass Butterworth filter. If you look at it, it has a noise component and a slowly changing component that an ML model may pick up and try using for making predictions. But it appears that both components change randomly without carrying any useful information, making this feature not just useless, but actually harmful. \n\nSo, in case you are using something like that, it's probably a good idea to remove this feature.\n\n![median of a low passed signal](https://i.ibb.co/MkS1CNQ/download-12.png)",
    "531914": "What is the cut-off frequency of the low pass filter?",
    "531929": "1000/75000. Here is the function I'm using for filtering (adapted from Vettejeep):\n```\ndef butter_filter(x, n=4, low=None, high=None, nyquist_f=75000):\n    if low is None and high is None:\n        return\n    elif low is None:\n        b, a = sg.butter(n, Wn=high/nyquist_f, btype='lowpass')\n    elif high is None:\n        b, a = sg.butter(n, Wn=low/nyquist_f, btype='highpass')\n    else:\n        b, a = sg.butter(4, Wn=(low/nyquist_f, high/nyquist_f), btype='bandpass')\n        \n    return sg.lfilter(b, a, x)\n```\n\nBasically, you can call this function on the 150000-long chunk of data with high=1000:\n`x_low_pass = butter_filter(x, high=1000)`",
    "531955": "The mean and median are likely to be artefacts of the recording instrumentation. I would be very wary of any features that rely on them, particularly since the distribution for the test set is very different.",
    "532059": "Looks like you use a very complex way to compute the mean of acoustic data ;)  \n\nIts distribution is different between train and test.",
    "532083": "Using it degrades my CV score, my ML mdoels aren't fooled by it for some reason ;)",
    "532085": "Maybe energy within a narrow band is more useful.",
    "532086": "Independently from machine learning, I am really surprised that they did not calibrate their sensor.  Acoustic data should have a 0 mean.",
    "532095": "yeah, I agree that it's a complex way to compute the mean, I realized it after posting this. I was just making some other features using frequency filters and stumbled upon this one. I put it in my model out of curiosity, and it was only mildly fooled by it, its permutation importance on the training set was positive (so it used it for predictions), but negative on the validation set. So, I would've gotten rid of it anyway during feature selection.",
    "532104": "I think it may be calibrated. We probably have just the small part of the entire recording, that can be low-frequency oscillation.",
    "532388": "Wasn't there a drift in mean caused by the degradation of the material during the experiments? I think i saw something like that in the additional data discussions.",
    "532882": "It indeed is the mean. The following is distribution of mean vs ttf from https://www.kaggle.com/gpreda/lanl-earthquake-new-approach-eda\n\n![](https://www.kaggleusercontent.com/kf/14309733/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..ctuIlZ7NH_on0xUbSmnQOg.ANrkULpie5oKS9HTUaQoiOZPo7g7vrx1Wb8GS-yMqYPfsCNNOdkrU-NKUbsB4L54KxKx8ftwwLxRuhLHa_ya5kFRTuInGQ_Gp1sn7K_PwdUc6tsJB3YNjhwy0Gv3nrdVOrBuQnUhNdaPlqGy4ahh-8-7S8hbJm55d9GJn6hvEHc.Zcoiw8HtSz_A0O_EqjyuNg/__results___files/__results___34_0.png)",
    "532883": "cpmpml I think this isn't due to calibration but because of the degradation of shear stresses over time and consequently the generated acoustic signal as @davids1992 and @svenhinderer mentioned. To add to their comments, I'm gonna refer you to this figure from LANL's papers:\n![](https://storage.googleapis.com/kaggle-forum-message-attachments/532883/13232/Untitled.jpg)",
    "532885": "batalov I'm glad you noticed this. What is more interesting for me, is that most of the public kernels not only use `mean` but also `sum` that is `= mean*150000` too =&gt; correlation = 1; there are also tons of features with correlation &gt;= 0.95. They also have some features that is always 0 in all instances such as `min(abs(signal))`.\n\nWhat you found is ONE of the MANY worst features :D",
    "532959": "Sum or mean is the same for gbms .  Why have both?",
    "532961": "That was my point. Because apparently somebody generated summary statistics at first without even thinking about it and everyone else blindly copied it.",
    "539031": "This article is underestimated.\nI have one feature that relies on mean. Without it, my CV is worse, and my public LB is also worse, by 0.02! My model outputs it as one of the top features. That's really confused here. From the plot it has no meaning, but what if it can be combined with other features and produces good predictions?  Instead of the mean value itself (which is bad) in the global scope that is different in train/test distributions, can it be good locally within a sample when combined with other features?",
    "539094": "I was a little bit unhappy with my score, but now, when I see here and in \"LANL Adversarial Validation - Shakeup is Coming\" topic, an excellent score comes with a cost and risk. Maybe the real talent is to be able to choose well being in that place.",
    "539106": "I have given up on entirely on the leaderboard, I just don't think it is informative enough. People who are forking the new Master's Approach kernel to get a &lt; 1.4 score are going to be disappointed, I think.",
    "539113": "The thing is, my CV now seems to correlate very well with LB. That’s really puzzled.",
    "539140": "I've finally found a CV strategy that I'm happy with, I'll be testing my final features in the next couple of days. It sounds like you engineered some really useful ones and a robust CV strategy so congratulations and good luck!"
  },
  "source": "meta"
}