{
  "id": 10835,
  "title": "Feature Engineering",
  "url": "/competitions/seizure-prediction/discussion/10835",
  "author_name": "",
  "post_date": "2014-11-04T17:48:19.020Z",
  "votes": 1,
  "comment_count": 8,
  "views": 2977,
  "content": "<p>I have&nbsp;tried extracting 15 unique types of features from the given time-series data. I would like to post what these features are and see if anyone has suggestions regarding useful features that I may have missed or any other advice. Would doing so be appropriate for this forum?&nbsp;</p>",
  "messages": [
    {
      "id": "57309",
      "postDate": "11/04/2014 17:48:19",
      "content": "<p>I have&nbsp;tried extracting 15 unique types of features from the given time-series data. I would like to post what these features are and see if anyone has suggestions regarding useful features that I may have missed or any other advice. Would doing so be appropriate for this forum?&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57425",
      "postDate": "11/06/2014 17:38:30",
      "content": "<p>yes</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57680",
      "postDate": "11/10/2014 21:24:43",
      "content": "<p>The features that I've extracted are listed below, anyone have comments or suggestions?&nbsp;(note that I've only tried one classifier at this point: random forest).</p>\n<ul>\n<li>fast fourier transform</li>\n<li>frequency correlation of&nbsp;channels</li>\n<li>time correlation of channels</li>\n<li>daubechies wavelet stats</li>\n<li>bin power for delta, theta, alpha and beta bands</li>\n<li>spectral entropy of&nbsp;channel&nbsp;</li>\n<li>hurst exponent of channel</li>\n<li>petrosian fractal dimension</li>\n<li>hjorth Fractal Dimension</li>\n<li>hjorth mobility and comlexity</li>\n<li>svd entropy</li>\n<li>fisher info</li>\n<li>approx. entropy</li>\n<li>sample entropy</li>\n<li>detrended fluctuation analysis</li>\n</ul>\n<p>For now I am moving on to optimizing my classifier, but I also am considering Hilber-Huang EMD.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57681",
      "postDate": "11/10/2014 21:40:05",
      "content": "<p>[quote=Mart&#237;n Molina;57680]</p>\n<ul>\n<li>fast fourier transform</li>\n</ul>\n<p>[/quote]</p>\n<p>What feature(s) are you taking from the fourier space?</p>\n<p>Also, have you tried auto-correlation? (That also gives a huge number of potential features. I personally haven't found anything useful, but I'm just throwing out ideas.)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57685",
      "postDate": "11/10/2014 21:55:18",
      "content": "<p>[quote=inversion;57681]</p>\n<p>What feature(s) are you taking from the fourier space?</p>\n<p>Also, have you tried auto-correlation? (That also gives a huge number of potential features. I personally haven't found anything useful, but I'm just throwing out ideas.)</p>\n<p>[/quote]</p>\n<p>So far the only feature i take from the fourier space is the correlation matrix (between channels) which I believe includes various autocorrelation features (each channel correlated with itself). This yielded poor results though, correlation matrix in the time-domain scored much better.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57720",
      "postDate": "11/11/2014 00:38:23",
      "content": "<p>Hi, may I know what is the size (sample X dimension) of features for 10mins EEG?</p>\n<p>[quote=Mart&#237;n Molina;57680]</p>\n<p>The features that I've extracted are listed below, anyone have comments or suggestions?&nbsp;(note that I've only tried one classifier at this point: random forest).</p>\n<ul>\n<li>fast fourier transform</li>\n<li>frequency correlation of&nbsp;channels</li>\n<li>time correlation of channels</li>\n<li>daubechies wavelet stats</li>\n<li>bin power for delta, theta, alpha and beta bands</li>\n<li>spectral entropy of&nbsp;channel&nbsp;</li>\n<li>hurst exponent of channel</li>\n<li>petrosian fractal dimension</li>\n<li>hjorth Fractal Dimension</li>\n<li>hjorth mobility and comlexity</li>\n<li>svd entropy</li>\n<li>fisher info</li>\n<li>approx. entropy</li>\n<li>sample entropy</li>\n<li>detrended fluctuation analysis</li>\n</ul>\n<p>For now I am moving on to optimizing my classifier, but I also am considering Hilber-Huang EMD.</p>\n<p>[/quote]</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57871",
      "postDate": "11/12/2014 03:37:43",
      "content": "<p>Hey Martin, with that many features, how do you prevent the model from overfitting? For some subject like Patient_1, which&nbsp;only has ~60 training samples, I suspect using anything&nbsp;with more than 10 features would lead to overfit.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57913",
      "postDate": "11/12/2014 15:13:49",
      "content": "<p>@Beck,</p>\n<p>I'm not using all of them at the same time. I started by using each one individually, then began trying&nbsp;different combinations to see what yields the best results.&nbsp;</p>\n<p>That said @Steven, my most successful combination thus far uses&nbsp;1016 features (mostly from&nbsp;the flattened time_correlation matrix, the other statistics&nbsp;only generate 1 or 2 features per channel so roughly 20 per sample)...which seems far&nbsp;too&nbsp;high, would you agree? I am considering using PCA or some other dimensionality reduction technique to get ride of noisy features, any recommendations here?</p>\n<p>Also, I do a lot of re-sampling (scipy's signal.resample method) of the original data before extracting features in order to make the run-time feasible. I get away with as little re-sampling as possible, but usually I sample down to between 400-4000 columns. This could be completely dumb, but this is my first ML project so I'm not sure. Please advise.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57954",
      "postDate": "11/12/2014 23:18:30",
      "content": "<p>Hi Martin may I ask whats the timings of your algorithms for the Approximate/Sample Entropy computations? (for a given size N ..lets say for 5000 data points, or whatever size you use). Thanks.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 57425,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "11/06/2014 17:38:30",
      "content": "<p>yes</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 57680,
      "author_name": "martinmolina",
      "author_url": "",
      "post_date": "11/10/2014 21:24:43",
      "content": "<p>The features that I've extracted are listed below, anyone have comments or suggestions?&nbsp;(note that I've only tried one classifier at this point: random forest).</p>\n<ul>\n<li>fast fourier transform</li>\n<li>frequency correlation of&nbsp;channels</li>\n<li>time correlation of channels</li>\n<li>daubechies wavelet stats</li>\n<li>bin power for delta, theta, alpha and beta bands</li>\n<li>spectral entropy of&nbsp;channel&nbsp;</li>\n<li>hurst exponent of channel</li>\n<li>petrosian fractal dimension</li>\n<li>hjorth Fractal Dimension</li>\n<li>hjorth mobility and comlexity</li>\n<li>svd entropy</li>\n<li>fisher info</li>\n<li>approx. entropy</li>\n<li>sample entropy</li>\n<li>detrended fluctuation analysis</li>\n</ul>\n<p>For now I am moving on to optimizing my classifier, but I also am considering Hilber-Huang EMD.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 57681,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "11/10/2014 21:40:05",
      "content": "<p>[quote=Mart&#237;n Molina;57680]</p>\n<ul>\n<li>fast fourier transform</li>\n</ul>\n<p>[/quote]</p>\n<p>What feature(s) are you taking from the fourier space?</p>\n<p>Also, have you tried auto-correlation? (That also gives a huge number of potential features. I personally haven't found anything useful, but I'm just throwing out ideas.)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 57685,
      "author_name": "martinmolina",
      "author_url": "",
      "post_date": "11/10/2014 21:55:18",
      "content": "<p>[quote=inversion;57681]</p>\n<p>What feature(s) are you taking from the fourier space?</p>\n<p>Also, have you tried auto-correlation? (That also gives a huge number of potential features. I personally haven't found anything useful, but I'm just throwing out ideas.)</p>\n<p>[/quote]</p>\n<p>So far the only feature i take from the fourier space is the correlation matrix (between channels) which I believe includes various autocorrelation features (each channel correlated with itself). This yielded poor results though, correlation matrix in the time-domain scored much better.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 57720,
      "author_name": "stevendu",
      "author_url": "",
      "post_date": "11/11/2014 00:38:23",
      "content": "<p>Hi, may I know what is the size (sample X dimension) of features for 10mins EEG?</p>\n<p>[quote=Mart&#237;n Molina;57680]</p>\n<p>The features that I've extracted are listed below, anyone have comments or suggestions?&nbsp;(note that I've only tried one classifier at this point: random forest).</p>\n<ul>\n<li>fast fourier transform</li>\n<li>frequency correlation of&nbsp;channels</li>\n<li>time correlation of channels</li>\n<li>daubechies wavelet stats</li>\n<li>bin power for delta, theta, alpha and beta bands</li>\n<li>spectral entropy of&nbsp;channel&nbsp;</li>\n<li>hurst exponent of channel</li>\n<li>petrosian fractal dimension</li>\n<li>hjorth Fractal Dimension</li>\n<li>hjorth mobility and comlexity</li>\n<li>svd entropy</li>\n<li>fisher info</li>\n<li>approx. entropy</li>\n<li>sample entropy</li>\n<li>detrended fluctuation analysis</li>\n</ul>\n<p>For now I am moving on to optimizing my classifier, but I also am considering Hilber-Huang EMD.</p>\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 57871,
      "author_name": "meditativeape",
      "author_url": "",
      "post_date": "11/12/2014 03:37:43",
      "content": "<p>Hey Martin, with that many features, how do you prevent the model from overfitting? For some subject like Patient_1, which&nbsp;only has ~60 training samples, I suspect using anything&nbsp;with more than 10 features would lead to overfit.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 57913,
      "author_name": "martinmolina",
      "author_url": "",
      "post_date": "11/12/2014 15:13:49",
      "content": "<p>@Beck,</p>\n<p>I'm not using all of them at the same time. I started by using each one individually, then began trying&nbsp;different combinations to see what yields the best results.&nbsp;</p>\n<p>That said @Steven, my most successful combination thus far uses&nbsp;1016 features (mostly from&nbsp;the flattened time_correlation matrix, the other statistics&nbsp;only generate 1 or 2 features per channel so roughly 20 per sample)...which seems far&nbsp;too&nbsp;high, would you agree? I am considering using PCA or some other dimensionality reduction technique to get ride of noisy features, any recommendations here?</p>\n<p>Also, I do a lot of re-sampling (scipy's signal.resample method) of the original data before extracting features in order to make the run-time feasible. I get away with as little re-sampling as possible, but usually I sample down to between 400-4000 columns. This could be completely dumb, but this is my first ML project so I'm not sure. Please advise.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 57954,
      "author_name": "chkaryst",
      "author_url": "",
      "post_date": "11/12/2014 23:18:30",
      "content": "<p>Hi Martin may I ask whats the timings of your algorithms for the Approximate/Sample Entropy computations? (for a given size N ..lets say for 5000 data points, or whatever size you use). Thanks.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "57309": "",
    "57425": "",
    "57680": "",
    "57681": "",
    "57685": "",
    "57720": "",
    "57871": "",
    "57913": "",
    "57954": ""
  },
  "source": "meta"
}