{
  "id": 90420,
  "title": "Are there really useful features in public kernels?",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/90420",
  "author_name": "DmitryS",
  "post_date": "2019-04-23T18:29:56.922000",
  "votes": 15,
  "comment_count": 27,
  "views": 0,
  "content": "<p>I’ve tried almost all features from public kernels, but nothing works for me. </p>\n\n<p>Methodology: train model on base features, then add new group of the features, retrain and compare results by folds. </p>\n\n<p>What doesn’t work for me:\n- librosa features. \n- tsfresh features. \n- statistics (mean, std, etc. on the whole signal / its parts, rolling windows, filtered signal, fft real and imag. parts ). \n- features on lta-sta. \n- features on HP filtered signal. </p>\n\n<p>Do you find something useful for you? </p>",
  "messages": [
    {
      "id": 521994,
      "postDate": "2019-04-23T18:29:56.923Z",
      "content": "<p>I’ve tried almost all features from public kernels, but nothing works for me. </p>\n\n<p>Methodology: train model on base features, then add new group of the features, retrain and compare results by folds. </p>\n\n<p>What doesn’t work for me:\n- librosa features. \n- tsfresh features. \n- statistics (mean, std, etc. on the whole signal / its parts, rolling windows, filtered signal, fft real and imag. parts ). \n- features on lta-sta. \n- features on HP filtered signal. </p>\n\n<p>Do you find something useful for you? </p>",
      "rawMarkdown": "I’ve tried almost all features from public kernels, but nothing works for me. \n\nMethodology: train model on base features, then add new group of the features, retrain and compare results by folds. \n\nWhat doesn’t work for me:\n- librosa features. \n- tsfresh features. \n- statistics (mean, std, etc. on the whole signal / its parts, rolling windows, filtered signal, fft real and imag. parts ). \n- features on lta-sta. \n- features on HP filtered signal. \n\nDo you find something useful for you? ",
      "votes": 15
    },
    {
      "id": 522801,
      "postDate": "2019-04-25T03:41:32.203Z",
      "content": "<p>One of the useful libraries is <a href=\"https://docs.obspy.org/contents.html\">ObsPy</a>,  but still figuring out how to use the competition datasets to plot &amp; analyze details.</p>",
      "rawMarkdown": "One of the useful libraries is [ObsPy](https://docs.obspy.org/contents.html),  but still figuring out how to use the competition datasets to plot &amp; analyze details.",
      "votes": 1
    },
    {
      "id": 522362,
      "postDate": "2019-04-24T10:08:07.813Z",
      "content": "<p>I also tried a lot of various combinations of all these features (except librosa), but CV is always about 2.01 +- 0.01.\nIt seems that secret lays somewhere outside of standard and almost-standard features</p>",
      "rawMarkdown": "I also tried a lot of various combinations of all these features (except librosa), but CV is always about 2.01 +- 0.01.\nIt seems that secret lays somewhere outside of standard and almost-standard features",
      "votes": 1
    },
    {
      "id": 523392,
      "postDate": "2019-04-26T06:39:26.937Z",
      "content": "<p>Has anyone looked at intensities from point processes such as Hawkes as a feature?</p>",
      "rawMarkdown": "Has anyone looked at intensities from point processes such as Hawkes as a feature?",
      "votes": 2
    },
    {
      "id": 523076,
      "postDate": "2019-04-25T14:37:16.793Z",
      "content": "<p>For me, most of the features I've implemented do not amount to anything. At the end, just 3 of the rolling features do most of the work. I expected quite a lot from the lta-sta features, but it didn't work for me.</p>",
      "rawMarkdown": "For me, most of the features I've implemented do not amount to anything. At the end, just 3 of the rolling features do most of the work. I expected quite a lot from the lta-sta features, but it didn't work for me.",
      "votes": 2,
      "replies": [
        {
          "id": 523394,
          "postDate": "2019-04-26T06:46:08.440Z",
          "content": "<p>&gt; I expected quite a lot from the lta-sta features, but it didn't work for me.</p>\n\n<p>I think the time window of 150k points is too small for it.  Detecting trends is near to impossible at that scale.</p>",
          "rawMarkdown": "&gt; I expected quite a lot from the lta-sta features, but it didn't work for me.\n\nI think the time window of 150k points is too small for it.  Detecting trends is near to impossible at that scale.",
          "votes": 1
        }
      ]
    },
    {
      "id": 522134,
      "postDate": "2019-04-23T22:38:24.363Z",
      "content": "<p>What works for you then?</p>\n\n<p>Anyway, I only found 3 features to be useful in kernels so far.</p>",
      "rawMarkdown": "What works for you then?\n\nAnyway, I only found 3 features to be useful in kernels so far.",
      "votes": 2,
      "replies": [
        {
          "id": 522139,
          "postDate": "2019-04-23T22:50:01.793Z",
          "content": "<p>So, let`s make things more interesting. Magic and some features inspired by audio frequency. </p>",
          "rawMarkdown": "So, let`s make things more interesting. Magic and some features inspired by audio frequency. ",
          "votes": 5
        },
        {
          "id": 522399,
          "postDate": "2019-04-24T11:46:41.850Z",
          "content": "<p>You don't have to answer but i'd still like to ask: Are the frequency features handcrafted or extracted by neural nets? </p>",
          "rawMarkdown": "You don't have to answer but i'd still like to ask: Are the frequency features handcrafted or extracted by neural nets? ",
          "votes": 1
        },
        {
          "id": 522402,
          "postDate": "2019-04-24T11:52:18.830Z",
          "content": "<p>I still do not use NN in my solution.</p>",
          "rawMarkdown": "I still do not use NN in my solution.",
          "votes": 2
        },
        {
          "id": 522408,
          "postDate": "2019-04-24T12:02:03.113Z",
          "content": "<p>Okay thanks. I thought about adding autoencoder-features from spectrograms again. Will probably try it anyway ^^</p>",
          "rawMarkdown": "Okay thanks. I thought about adding autoencoder-features from spectrograms again. Will probably try it anyway ^^"
        },
        {
          "id": 522512,
          "postDate": "2019-04-24T14:18:41.147Z",
          "content": "<p>I only got one serious feature so far ... The rest is just modeling noise, i guess. </p>",
          "rawMarkdown": "I only got one serious feature so far ... The rest is just modeling noise, i guess. "
        },
        {
          "id": 523809,
          "postDate": "2019-04-27T05:01:13.463Z",
          "content": "<p>May i ask you if you make your magic frequency feature by FFT ? </p>",
          "rawMarkdown": "May i ask you if you make your magic frequency feature by FFT ? "
        }
      ]
    },
    {
      "id": 522056,
      "postDate": "2019-04-23T19:58:17.340Z",
      "content": "<p>Your own features must be rather good (judging by your position on LB :-) I've managed to squeeze a bit from the ones present in public kernels - especially statistics.</p>",
      "rawMarkdown": "Your own features must be rather good (judging by your position on LB :-) I've managed to squeeze a bit from the ones present in public kernels - especially statistics.",
      "votes": 2
    },
    {
      "id": 523410,
      "postDate": "2019-04-26T07:32:21.340Z",
      "content": "<p>I think the first challenge for pre-processing is to create a robust seismic signal segmentation for the train set so we can apply those statistical features, filtering &amp; denoising signals and another feature engineering. I think 150000 samples per segment (in reference to the test set) does not apply in the train set. Once we figure out that correct signal segmentation, we can formulate a validation strategy based on the results of signal segmentation. Yes, I tried those released public kernels to integrate into my model. Unfortunately, most of them does not help. Like I stated earlier, I think the key is to have a correct signal segmentation and I'm still working on it.</p>",
      "rawMarkdown": "I think the first challenge for pre-processing is to create a robust seismic signal segmentation for the train set so we can apply those statistical features, filtering &amp; denoising signals and another feature engineering. I think 150000 samples per segment (in reference to the test set) does not apply in the train set. Once we figure out that correct signal segmentation, we can formulate a validation strategy based on the results of signal segmentation. Yes, I tried those released public kernels to integrate into my model. Unfortunately, most of them does not help. Like I stated earlier, I think the key is to have a correct signal segmentation and I'm still working on it.",
      "replies": [
        {
          "id": 528684,
          "postDate": "2019-05-08T11:25:14.537Z",
          "content": "<p>I don't understand your point. What do you mean with:\n\" I think 150000 samples per segment (in reference to the test set) does not apply in the train set\"?</p>\n\n<p>I agree the size of the chunks is not the ideal, because it is just an instant, but we are forced to train the alhorithms with no more than 150000 samples because of the size of the chunks of the test set.</p>",
          "rawMarkdown": "I don't understand your point. What do you mean with:\n\" I think 150000 samples per segment (in reference to the test set) does not apply in the train set\"?\n\nI agree the size of the chunks is not the ideal, because it is just an instant, but we are forced to train the alhorithms with no more than 150000 samples because of the size of the chunks of the test set."
        }
      ]
    },
    {
      "id": 523275,
      "postDate": "2019-04-25T22:02:05.687Z",
      "content": "<p>Did you applied a rolling window or divided into sections to the segments of the time series  while using tsfresh features? </p>",
      "rawMarkdown": "Did you applied a rolling window or divided into sections to the segments of the time series  while using tsfresh features? ",
      "replies": [
        {
          "id": 523277,
          "postDate": "2019-04-25T22:09:38.307Z",
          "content": "<p>Divide into sections.</p>",
          "rawMarkdown": "Divide into sections."
        }
      ]
    },
    {
      "id": 522672,
      "postDate": "2019-04-24T19:47:07.150Z",
      "content": "<p>None of my 10 best features are in the public kernels, but I haven't made them for the test set yet and so haven't evaluated them against the LB. I'm judging 'best' by how they improve my CV score, and how important the models consider them. One of them took three kernels to calculate for the training set even without oversampling, so I'm glad it proved helpful!</p>\n\n<p>As a side note, what are the strongest absolute correlations your features have with TTF? I have quite a few over 0.65 but only one which reaches 0.675.</p>",
      "rawMarkdown": "None of my 10 best features are in the public kernels, but I haven't made them for the test set yet and so haven't evaluated them against the LB. I'm judging 'best' by how they improve my CV score, and how important the models consider them. One of them took three kernels to calculate for the training set even without oversampling, so I'm glad it proved helpful!\n\nAs a side note, what are the strongest absolute correlations your features have with TTF? I have quite a few over 0.65 but only one which reaches 0.675.",
      "replies": [
        {
          "id": 522689,
          "postDate": "2019-04-24T20:27:42.577Z",
          "content": "<p>I don`t think that correlation is a good measure for this task, because I use nonlinear models.   Nevertheless, an answer on your question is 0.676.  I have three features out of 35 with correlation more than 0.675 and eight with correlation less than 0.1.</p>",
          "rawMarkdown": "I don`t think that correlation is a good measure for this task, because I use nonlinear models.   Nevertheless, an answer on your question is 0.676.  I have three features out of 35 with correlation more than 0.675 and eight with correlation less than 0.1."
        }
      ]
    },
    {
      "id": 522127,
      "postDate": "2019-04-23T22:17:22.790Z",
      "content": "<p><a href=\"/simakov\">@simakov</a>, do you not trust your results at #20 on the LB? What CV strategy are you using as I believe it is the most important question with this tricky data.</p>",
      "rawMarkdown": "@simakov, do you not trust your results at #20 on the LB? What CV strategy are you using as I believe it is the most important question with this tricky data.",
      "replies": [
        {
          "id": 522132,
          "postDate": "2019-04-23T22:27:11.323Z",
          "content": "<p>I trust my CV strategy:  I use three different CV with some improvement criteria. But I don`t know how other competitors perform on unseen private data. It is simple to see that model performs better on some observations and worse on others. You can slightly impove the model for the whole data set or well for particular groups of observations. What will work better on private lb? 😄 </p>",
          "rawMarkdown": "I trust my CV strategy:  I use three different CV with some improvement criteria. But I don`t know how other competitors perform on unseen private data. It is simple to see that model performs better on some observations and worse on others. You can slightly impove the model for the whole data set or well for particular groups of observations. What will work better on private lb? 😄 ",
          "votes": 1
        },
        {
          "id": 522136,
          "postDate": "2019-04-23T22:40:27.443Z",
          "content": "<blockquote>\n  <p>What will work better in private lb?</p>\n</blockquote>\n\n<p><a href=\"/simakov\">@simakov</a>, that is the main question. I have a couple of CV strategy myself but I cannot answer the above question which makes me doubt them.</p>",
          "rawMarkdown": "&gt;What will work better in private lb?\n\n@simakov, that is the main question. I have a couple of CV strategy myself but I cannot answer the above question which makes me doubt them."
        }
      ]
    },
    {
      "id": 522055,
      "postDate": "2019-04-23T19:57:57.850Z",
      "content": "<p>You're in 25th place so you've obviously found something useful lol. I have found the statistics from Andrew's script to be the most useful, those are the only ones I am using.</p>",
      "rawMarkdown": "You're in 25th place so you've obviously found something useful lol. I have found the statistics from Andrew's script to be the most useful, those are the only ones I am using."
    },
    {
      "id": 528490,
      "postDate": "2019-05-08T00:28:40.300Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 528689,
          "postDate": "2019-05-08T11:35:28.987Z",
          "content": "<p>Here, and in many competitions, I start with 0 features, and iterate the following: add some, train models, evaluate result, keep features if result is better than before.  Results are evaluated with cross validation (CV), not leaderboard.  You can also look at feature importance as provided by algorithms but this is not always reliable. Very important features can overfit training data.</p>\n\n<p>Whatever you do, always use the result of a model trained with your features to asses feature effectiveness.  I never understood why people keep using methods (eg correlation, information value, etc) that were designed specifically for general linear models.</p>\n\n<p>I guess I know why it is tempting to rely on feature statistics: it gives the false impression that one can assess features without having to worry about a sound CV setting.</p>",
          "rawMarkdown": "Here, and in many competitions, I start with 0 features, and iterate the following: add some, train models, evaluate result, keep features if result is better than before.  Results are evaluated with cross validation (CV), not leaderboard.  You can also look at feature importance as provided by algorithms but this is not always reliable. Very important features can overfit training data.\n\nWhatever you do, always use the result of a model trained with your features to asses feature effectiveness.  I never understood why people keep using methods (eg correlation, information value, etc) that were designed specifically for general linear models.\n\nI guess I know why it is tempting to rely on feature statistics: it gives the false impression that one can assess features without having to worry about a sound CV setting.",
          "votes": 8
        },
        {
          "id": 528728,
          "postDate": "2019-05-08T13:01:42.283Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 528786,
          "postDate": "2019-05-08T15:02:28.743Z",
          "content": "<p>shuffled KFold, while avoiding leakage from overlapping segments is one way to go. </p>\n\n<p>another way would be a low number of KFolds (higher number tends to overfit due to the early stopping peaking)</p>",
          "rawMarkdown": "shuffled KFold, while avoiding leakage from overlapping segments is one way to go. \n\nanother way would be a low number of KFolds (higher number tends to overfit due to the early stopping peaking)",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 522801,
      "author_name": "FGPC",
      "author_url": "",
      "post_date": "2019-04-25T03:41:32.203000",
      "content": "<p>One of the useful libraries is <a href=\"https://docs.obspy.org/contents.html\">ObsPy</a>,  but still figuring out how to use the competition datasets to plot &amp; analyze details.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 522362,
      "author_name": "Stanislav Blinov",
      "author_url": "",
      "post_date": "2019-04-24T10:08:07.813000",
      "content": "<p>I also tried a lot of various combinations of all these features (except librosa), but CV is always about 2.01 +- 0.01.\nIt seems that secret lays somewhere outside of standard and almost-standard features</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 523392,
      "author_name": "pavy bez",
      "author_url": "",
      "post_date": "2019-04-26T06:39:26.937000",
      "content": "<p>Has anyone looked at intensities from point processes such as Hawkes as a feature?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 523076,
      "author_name": "Ricard Delgado",
      "author_url": "",
      "post_date": "2019-04-25T14:37:16.793000",
      "content": "<p>For me, most of the features I've implemented do not amount to anything. At the end, just 3 of the rolling features do most of the work. I expected quite a lot from the lta-sta features, but it didn't work for me.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 523394,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-04-26T06:46:08.440000",
          "content": "<p>&gt; I expected quite a lot from the lta-sta features, but it didn't work for me.</p>\n\n<p>I think the time window of 150k points is too small for it.  Detecting trends is near to impossible at that scale.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 522134,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2019-04-23T22:38:24.363000",
      "content": "<p>What works for you then?</p>\n\n<p>Anyway, I only found 3 features to be useful in kernels so far.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 522139,
          "author_name": "DmitryS",
          "author_url": "",
          "post_date": "2019-04-23T22:50:01.793000",
          "content": "<p>So, let`s make things more interesting. Magic and some features inspired by audio frequency. </p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 522399,
          "author_name": "Sven Hinderer",
          "author_url": "",
          "post_date": "2019-04-24T11:46:41.850000",
          "content": "<p>You don't have to answer but i'd still like to ask: Are the frequency features handcrafted or extracted by neural nets? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 522402,
          "author_name": "DmitryS",
          "author_url": "",
          "post_date": "2019-04-24T11:52:18.830000",
          "content": "<p>I still do not use NN in my solution.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 522408,
          "author_name": "Sven Hinderer",
          "author_url": "",
          "post_date": "2019-04-24T12:02:03.113000",
          "content": "<p>Okay thanks. I thought about adding autoencoder-features from spectrograms again. Will probably try it anyway ^^</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 522512,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2019-04-24T14:18:41.147000",
          "content": "<p>I only got one serious feature so far ... The rest is just modeling noise, i guess. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 523809,
          "author_name": "dinosaur",
          "author_url": "",
          "post_date": "2019-04-27T05:01:13.463000",
          "content": "<p>May i ask you if you make your magic frequency feature by FFT ? </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 522056,
      "author_name": "Konrad Banachewicz",
      "author_url": "",
      "post_date": "2019-04-23T19:58:17.340000",
      "content": "<p>Your own features must be rather good (judging by your position on LB :-) I've managed to squeeze a bit from the ones present in public kernels - especially statistics.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 523410,
      "author_name": "FGPC",
      "author_url": "",
      "post_date": "2019-04-26T07:32:21.340000",
      "content": "<p>I think the first challenge for pre-processing is to create a robust seismic signal segmentation for the train set so we can apply those statistical features, filtering &amp; denoising signals and another feature engineering. I think 150000 samples per segment (in reference to the test set) does not apply in the train set. Once we figure out that correct signal segmentation, we can formulate a validation strategy based on the results of signal segmentation. Yes, I tried those released public kernels to integrate into my model. Unfortunately, most of them does not help. Like I stated earlier, I think the key is to have a correct signal segmentation and I'm still working on it.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 528684,
          "author_name": "Carlos Prades K.",
          "author_url": "",
          "post_date": "2019-05-08T11:25:14.537000",
          "content": "<p>I don't understand your point. What do you mean with:\n\" I think 150000 samples per segment (in reference to the test set) does not apply in the train set\"?</p>\n\n<p>I agree the size of the chunks is not the ideal, because it is just an instant, but we are forced to train the alhorithms with no more than 150000 samples because of the size of the chunks of the test set.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 523275,
      "author_name": "SUPERLUMINAL",
      "author_url": "",
      "post_date": "2019-04-25T22:02:05.687000",
      "content": "<p>Did you applied a rolling window or divided into sections to the segments of the time series  while using tsfresh features? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 523277,
          "author_name": "DmitryS",
          "author_url": "",
          "post_date": "2019-04-25T22:09:38.307000",
          "content": "<p>Divide into sections.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 522672,
      "author_name": "RNA",
      "author_url": "",
      "post_date": "2019-04-24T19:47:07.150000",
      "content": "<p>None of my 10 best features are in the public kernels, but I haven't made them for the test set yet and so haven't evaluated them against the LB. I'm judging 'best' by how they improve my CV score, and how important the models consider them. One of them took three kernels to calculate for the training set even without oversampling, so I'm glad it proved helpful!</p>\n\n<p>As a side note, what are the strongest absolute correlations your features have with TTF? I have quite a few over 0.65 but only one which reaches 0.675.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 522689,
          "author_name": "DmitryS",
          "author_url": "",
          "post_date": "2019-04-24T20:27:42.577000",
          "content": "<p>I don`t think that correlation is a good measure for this task, because I use nonlinear models.   Nevertheless, an answer on your question is 0.676.  I have three features out of 35 with correlation more than 0.675 and eight with correlation less than 0.1.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 522127,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2019-04-23T22:17:22.790000",
      "content": "<p><a href=\"/simakov\">@simakov</a>, do you not trust your results at #20 on the LB? What CV strategy are you using as I believe it is the most important question with this tricky data.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 522132,
          "author_name": "DmitryS",
          "author_url": "",
          "post_date": "2019-04-23T22:27:11.323000",
          "content": "<p>I trust my CV strategy:  I use three different CV with some improvement criteria. But I don`t know how other competitors perform on unseen private data. It is simple to see that model performs better on some observations and worse on others. You can slightly impove the model for the whole data set or well for particular groups of observations. What will work better on private lb? 😄 </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 522136,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2019-04-23T22:40:27.443000",
          "content": "<blockquote>\n  <p>What will work better in private lb?</p>\n</blockquote>\n\n<p><a href=\"/simakov\">@simakov</a>, that is the main question. I have a couple of CV strategy myself but I cannot answer the above question which makes me doubt them.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 522055,
      "author_name": "Dalton Hall",
      "author_url": "",
      "post_date": "2019-04-23T19:57:57.850000",
      "content": "<p>You're in 25th place so you've obviously found something useful lol. I have found the statistics from Andrew's script to be the most useful, those are the only ones I am using.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 528490,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-05-08T00:28:40.300000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 528689,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-08T11:35:28.987000",
          "content": "<p>Here, and in many competitions, I start with 0 features, and iterate the following: add some, train models, evaluate result, keep features if result is better than before.  Results are evaluated with cross validation (CV), not leaderboard.  You can also look at feature importance as provided by algorithms but this is not always reliable. Very important features can overfit training data.</p>\n\n<p>Whatever you do, always use the result of a model trained with your features to asses feature effectiveness.  I never understood why people keep using methods (eg correlation, information value, etc) that were designed specifically for general linear models.</p>\n\n<p>I guess I know why it is tempting to rely on feature statistics: it gives the false impression that one can assess features without having to worry about a sound CV setting.</p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 528728,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-05-08T13:01:42.283000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 528786,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2019-05-08T15:02:28.743000",
          "content": "<p>shuffled KFold, while avoiding leakage from overlapping segments is one way to go. </p>\n\n<p>another way would be a low number of KFolds (higher number tends to overfit due to the early stopping peaking)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "521994": "I’ve tried almost all features from public kernels, but nothing works for me. \n\nMethodology: train model on base features, then add new group of the features, retrain and compare results by folds. \n\nWhat doesn’t work for me:\n- librosa features. \n- tsfresh features. \n- statistics (mean, std, etc. on the whole signal / its parts, rolling windows, filtered signal, fft real and imag. parts ). \n- features on lta-sta. \n- features on HP filtered signal. \n\nDo you find something useful for you? ",
    "522801": "One of the useful libraries is [ObsPy](https://docs.obspy.org/contents.html),  but still figuring out how to use the competition datasets to plot &amp; analyze details.",
    "522362": "I also tried a lot of various combinations of all these features (except librosa), but CV is always about 2.01 +- 0.01.\nIt seems that secret lays somewhere outside of standard and almost-standard features",
    "523392": "Has anyone looked at intensities from point processes such as Hawkes as a feature?",
    "523076": "For me, most of the features I've implemented do not amount to anything. At the end, just 3 of the rolling features do most of the work. I expected quite a lot from the lta-sta features, but it didn't work for me.",
    "522134": "What works for you then?\n\nAnyway, I only found 3 features to be useful in kernels so far.",
    "522056": "Your own features must be rather good (judging by your position on LB :-) I've managed to squeeze a bit from the ones present in public kernels - especially statistics.",
    "523410": "I think the first challenge for pre-processing is to create a robust seismic signal segmentation for the train set so we can apply those statistical features, filtering &amp; denoising signals and another feature engineering. I think 150000 samples per segment (in reference to the test set) does not apply in the train set. Once we figure out that correct signal segmentation, we can formulate a validation strategy based on the results of signal segmentation. Yes, I tried those released public kernels to integrate into my model. Unfortunately, most of them does not help. Like I stated earlier, I think the key is to have a correct signal segmentation and I'm still working on it.",
    "523275": "Did you applied a rolling window or divided into sections to the segments of the time series  while using tsfresh features? ",
    "522672": "None of my 10 best features are in the public kernels, but I haven't made them for the test set yet and so haven't evaluated them against the LB. I'm judging 'best' by how they improve my CV score, and how important the models consider them. One of them took three kernels to calculate for the training set even without oversampling, so I'm glad it proved helpful!\n\nAs a side note, what are the strongest absolute correlations your features have with TTF? I have quite a few over 0.65 but only one which reaches 0.675.",
    "522127": "@simakov, do you not trust your results at #20 on the LB? What CV strategy are you using as I believe it is the most important question with this tricky data.",
    "522055": "You're in 25th place so you've obviously found something useful lol. I have found the statistics from Andrew's script to be the most useful, those are the only ones I am using.",
    "528490": ""
  }
}