{
  "id": 10079,
  "title": "Features for seizure detection",
  "url": "/competitions/seizure-detection/discussion/10079",
  "author_name": "",
  "post_date": "2014-08-20T03:51:23.600Z",
  "votes": 4,
  "comment_count": 20,
  "views": 11548,
  "content": "<p>Hello, everyone,</p>\n<p>I'm wondering which features you have extracted&nbsp;for this dataset. I mainly have used the following features for each channel to achieve an accuracy of 0.9423:</p>\n<p>fractal dimension,&nbsp;mobility,&nbsp;complexity,&nbsp;skewness,&nbsp;kurotsis,&nbsp;variance,&nbsp;</p>\n<p>and frequency energy at following bands:</p>\n<p>delta(0.5-4Hz),&nbsp;theta(4-7Hz),&nbsp;alpha(7-14Hz),&nbsp;beta(14-30Hz),&nbsp;gamma(30-100Hz)</p>\n<p>Have fun~</p>",
  "messages": [
    {
      "id": "52208",
      "postDate": "08/20/2014 03:51:23",
      "content": "<p>Hello, everyone,</p>\n<p>I'm wondering which features you have extracted&nbsp;for this dataset. I mainly have used the following features for each channel to achieve an accuracy of 0.9423:</p>\n<p>fractal dimension,&nbsp;mobility,&nbsp;complexity,&nbsp;skewness,&nbsp;kurotsis,&nbsp;variance,&nbsp;</p>\n<p>and frequency energy at following bands:</p>\n<p>delta(0.5-4Hz),&nbsp;theta(4-7Hz),&nbsp;alpha(7-14Hz),&nbsp;beta(14-30Hz),&nbsp;gamma(30-100Hz)</p>\n<p>Have fun~</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "52212",
      "postDate": "08/20/2014 06:17:57",
      "content": "<p>The following text probably won't be quite precise, but the intuition is I just threw a kitchen sink of feature transformations at the problem in the limited time I had (about 2-3? weeks) without thinking about it too much.&nbsp;</p>\n<p>I tried various things at various windows and preprocessed with various filters not all of which was done within the box of conventional signal processing:</p>\n<p>R's summary, variance, sd, mad, IQR, range, R (moment's package) skewness, kurtosis, geary, fft(Arg()) (half of it), mean, median filter, gaussian filter, correlation (upper tri, pearson, of the channels), lagged differences (diff) at various lags, various combinations of the above in different orders.&nbsp;</p>\n<p>Then I ensembled gbm, brnn, glmnet, and randomForest. While brnn and glmnet did fairly poorly, these still helped overall. Everything was really slow since I did this all in R although multiple instances on AWS helped. Maybe some people had better ideas on handling the size of the data.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "52220",
      "postDate": "08/20/2014 08:53:35",
      "content": "<p>I have used the next features:</p>\n<p>maximal and minimal &nbsp;values for&nbsp;each channel;</p>\n<p>standard deviation by&nbsp;channel;</p>\n<p>I had an accuracy about &nbsp;0.92 with Random Forest, KNN and Logistic Regression.</p>\n<p>Then I've added:</p>\n<p>logarithm of the sum of absolute values for each channel;</p>\n<p>logarithm from the first 250 values of&nbsp;FFT for two&nbsp;channels</p>\n<p>and improve accuracy to 0.94&nbsp;.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "52222",
      "postDate": "08/20/2014 09:12:05",
      "content": "<p>Hi, I thought about hardware constrains and so I decided for 'cheap' time domain features, so I went with zero crossing rate (zcr) and average short time energy (aste). Extracting these features for all of the available channels and running them against a RF I've achieved a 0.89 score.</p>\n<p>This was a fun competition.</p>\n<p>PS: I like Serhii (min, Max, sd) approach :)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "52226",
      "postDate": "08/20/2014 11:24:27",
      "content": "<p>I hand-crafted features based on the following reference: D F Wulsin, J R Gupta, R Mani, J A Blanco and B Litt, Modeling electroencephalography waveforms with semi-supervised deep belief nets: fast classification and anomaly measurement, J. Neural Eng. 8 (2011) 036015</p>\n<p>From Appendix B I extracted the following features for each EEG channel seperately:</p>\n<p>-Normalized positive area under the curve</p>\n<p>-Normalized decay: the chance-corrected fraction of data that is decreasing or increasing</p>\n<p>-Frequency band power: the mean power spectral density in each of the frequency bands 1.5Hz-7Hz, 8Hz-15Hz and 16Hz-29Hz computed using Welch's method</p>\n<p>-Line length: sum of the absolute differences between successive data samples</p>\n<p>-Mean energy: mean of the square of samples across each channel</p>\n<p>-Average peak amplitude: base-10 log of the mean-squared amplitude of each peak in the channel</p>\n<p>-Average valley amplitude: as above but for the valleys</p>\n<p>-Normalized peak number: number of peaks in the channel normalized by the mean difference between adjacent data points</p>\n<p>-Zero crossings: subtracts the mean value (rather than the line of best fit as in the paper) from the channel and then counts how many times the zero-mean data crosses zero</p>\n<p>Looking at the scores achieved above with smaller/simpler feature sets it seems that it was wasted effort to extract some of these.&nbsp; I plateaued very early on in this competition and couldn't figure out how to make any significant improvements to my score.&nbsp; I trained using Scikit ExtraTreesClassifier for both feature selection and classification.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "52233",
      "postDate": "08/20/2014 14:47:11",
      "content": "<p>This competition was definitely all about feature engineering. Using an ensemble of models (random forest, glmnet, bagged MARS, gbm) I got to 0.95+ with the following:</p>\n<p>Measures of variability: Variance, max-min, 95th-5th percentile</p>\n<p>Average Spectral Power (using wavelets): Delta (0-4 Hz); Theta (4-8); Alpha (8-14); Beta (15-30); lowGamma (30-100); highGamma (100-200)</p>\n<p>Ratio of Spectral Powers to each other (e.g., theta&nbsp; / alpha; all combinations).</p>\n<p>Right at the end, I added the mean and variance of cross-correlation values (i.e. lagged correlation ranging from 1-20 samples), which improved the accuracy of some of my individual models, but in the end, didn't improve the final ensemble.</p>\n<p>No single model of mine had better prediction accuracy of 0.94, and some were as low as 0.87, so combining the predictions definitely helped here.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "52234",
      "postDate": "08/20/2014 15:13:23",
      "content": "<p>I mainly extracted frequency bins of about 3 Hz range and used a hamming window to improve the frequency resolution. Further marginal improvements came by splitting the time series into 'early' and 'late' pieces (crossover hamming windows) and extracting frequency content, as well as ratios between frequency bins.</p>\n<p>I used a random forest model to get scores above .95. If I had used my best public leaderboard score for submission, that would have been good enough for ~ 6th place. I was concerned about overfitting so relied on a 5 fold CV for choosing the best model (in terms of reducing unnecessary features), but this yielded ~ .94 in both. Perhaps a leave-on-out CV would have been better?</p>\n<p>I meant to spend more time at the end developing more features (both time and spectral) but didn't get around to it. Congrats to the winners! Fun competition!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "52235",
      "postDate": "08/20/2014 15:18:18",
      "content": "<p>I have evaluated the&nbsp;features individually with the label with AUC score described in the first post. I found that std, mobility, fractal dimensions and gamma are the most important ones. I have also included the 20%, 50%, 80% percentiles of &nbsp;cross correlation&nbsp;values, The attached is what I have got (the number in front indicates the channel):</p>\n<p>Dog_1<br>Seizure top 10: ['4_std' '3_std' '7_std' 'percentile0.2' '2_std' '11_std' '15_std' '12_std''13_std' '9_std']<br>Early seizure top 10: ['15_beta' '11_mobility' '15_gamma' '7_gamma' '15_fractals' '3_gamma''7_fractals' '11_gamma' '3_fractals' '11_fractals']</p>\n<p>Dog_2<br>Seizure top 10: ['14_mobility' '9_mobility' '14_gamma' '10_gamma' '7_mobility''11_mobility' '7_gamma' '15_mobility' '15_gamma' '10_mobility']<br>Early seizure top 10: ['12_fractals' '7_fractals' '11_fractals' '15_gamma' '14_gamma' '10_gamma''7_gamma' '10_fractals' '14_fractals' '15_fractals']</p>\n<p>Dog_3<br>Seizure top 10: ['9_std' '2_std' '3_std' '5_std' '7_std' '14_std' '10_std' '6_std' '12_std''13_std']<br>Early seizure top 10:['11_std' '2_std' '5_std' '3_std' '10_std' '14_std' '12_std' '7_std''13_std' '6_std']</p>\n<p>Dog_4<br>Seizure top 10: ['12_mobility' '7_std' '14_fractals' '12_complexity' '5_gamma' '5_fractals''6_fractals' '5_complexity' '5_mobility' '5_std']<br>Early seizure top 10: ['3_std' 'median' '15_std' '11_std' '13_std' '7_std' '4_std' '12_std''5_std' 'percentile0.2']</p>\n<p>Patient_1<br>Seizure top 10:&nbsp;['10_alpha' '9_fractals' '62_alpha' '19_fractals' '18_alpha' '13_alpha''13_fractals' '12_alpha' '20_alpha' '29_alpha']</p>\n<p>Early seizure top 10: ['18_complexity' '19_mobility' '29_fractals' '20_complexity' '18_fractals''18_mobility' '29_gamma' '29_mobility' '19_complexity' '20_mobility']</p>\n<p>Patient_2<br>Seizure top 10: ['2_std' '3_complexity' '1_complexity' '1_std' '2_beta' '1_beta' '3_beta''0_complexity' '0_std' '0_beta']<br>Early seizure top 10: ['0_std' '2_mobility' '3_complexity' '3_beta' '0_complexity' '0_beta''1_complexity' '2_complexity' '1_beta' '2_beta']</p>\n<p>Patient_3<br>Seizure top 10: ['8_theta' '6_std' '5_delta' '8_beta' '5_beta' '4_alpha' '5_complexity''54_fractals' '4_delta' '4_complexity']<br>Early seizure top 10: ['5_skew' '5_alpha' '54_fractals' '5_complexity' '5_beta' '4_mobility''4_beta' '4_alpha' '4_delta' '4_complexity']</p>\n<p>Patient_4<br>Seizure top 10: ['48_std' '33_gamma' '42_std' '34_std' '8_alpha' '56_fractals' '56_std''33_fractals' '33_std' '49_std']<br>Early seizure top 10: ['48_std' '33_gamma' '42_std' '34_std' '8_alpha' '56_fractals' '56_std''33_fractals' '33_std' '49_std']</p>\n<p>Patient_5<br>Seizure top 10: ['7_mobility' '0_gamma' '0_complexity' '31_fractals' '15_mobility''1_fractals' '7_fractals' '0_mobility' '15_fractals' '0_fractals']<br>Early seizure top 10: ['1_mobility' '7_mobility' '15_fractals' '7_fractals' '0_gamma''15_mobility' '1_fractals' '0_complexity' '0_mobility' '0_fractals']</p>\n<p>Patient_6<br>Seizure top 10: ['14_skew' '23_beta' '22_complexity' '15_alpha' '22_std' '14_delta''23_std' '23_complexity' '14_complexity' '14_alpha']<br>Early seizure top 10: ['22_complexity' '14_skew' '22_fractals' '14_complexity' '23_beta''23_complexity' '23_alpha' '14_alpha' '23_std' '22_std']</p>\n<p>Patient_7<br>Seizure top 10: ['35_theta' '27_theta' '27_std' '31_fractals' '34_fractals' '26_fractals''30_fractals' '27_fractals' '7_std' '27_alpha']<br>Early seizure top 10: ['25_fractals' '12_alpha' '2_beta' '35_theta' 'percentile0.5' '12_fractals''23_gamma' '12_gamma' '1_gamma' 'percentile0.2']</p>\n<p>Patient_8<br>Seizure top 10: ['9_std' '6_mobility' '5_mobility' '9_fractals' '14_mobility' '12_mobility''13_mobility' '11_fractals' '12_fractals' '10_fractals']<br>Early seizure top 10: ['9_beta' '13_delta' '13_complexity' '8_fractals' '14_mobility''15_mobility' '12_delta' '13_mobility' '12_complexity' '12_mobility']</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "52236",
      "postDate": "08/20/2014 15:30:30",
      "content": "<p>A quick summary of my model:</p>\n<p>Resample to 500 sps. Extract 0.5 second windows from the beginning, middle, and end of each segment. Apply Hanning windows and compute DFTs. Sum the power in bands 4-8, 8-13, 13-30, and 30-100 Hz and convert to log scale. Discard all but 16 channels (I did this because I was short on time and didn't want to search for a better way to incorporate additional channels). The channels that provided the greatest d-prime discrimination of ictal vs. interictal were retained and ordered by their d-prime values. This resulted in a feature vector of length 3(times)x4(bands)x16(channels) = 192 for each segment. I used SVMs with RBFs and gamma = 1.58. One SVM for each of the predictions we needed to make.</p>\n<p>Matt</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "52247",
      "postDate": "08/20/2014 18:02:39",
      "content": "<p>We simply used 5 frequency bands per channel, and&nbsp;used pwelch() for power-spectral-density estimation, using tons of overlap within each segment and a small window size. &nbsp;Using those features alone, the tricky part, from our&nbsp;point of view, seemed to be get rid of all those&nbsp;&quot;spikes&quot; which occurred randomly, and which we interpreted as measurement error. &nbsp;We didn't have enough time to write a robust enough spike detector to improve our score significantly, w.r.t. not&nbsp;removing&nbsp;the spikes.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "52248",
      "postDate": "08/20/2014 18:15:26",
      "content": "<p>Can people elaborate on computational speed and methods they used in terms of implementation? The data set was fairly large compared to the median Kaggle contest.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "52262",
      "postDate": "08/20/2014 21:10:51",
      "content": "<p>For feature extraction, I used Julia on an Amazon c3.xlarge. Runtimes generally on the order of an hour or less, but I was only looking at fairly simple features.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "52274",
      "postDate": "08/21/2014 00:48:27",
      "content": "<p>For features I used median, variance, and extracted 6 frequency bands from an FFT. Classification was done using a Random Forest.</p>\n\n<p>The way I handled the data import was quite inefficient. Reading in the data took the majority of the run time for building the model and generating predictions. I rented a EC2 r3.xlarge instance overnight to avoid spending time optimizing my implementation.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "52276",
      "postDate": "08/21/2014 03:17:40",
      "content": "<p>The dataset was only a few gb after downsampling the clips to O(100) time units&nbsp;and storing data as single precision (and it didn't seem like much&nbsp;information was lost doing this). &nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "52297",
      "postDate": "08/21/2014 20:01:28",
      "content": "<p>[quote=Mike Kim;52248]</p>\n<p>Can people elaborate on computational speed and methods they used in terms of implementation? The data set was fairly large compared to the median Kaggle contest.&nbsp;</p>\n<p>[/quote]</p>\n<p>Using joblib's Parallel function was a big help - I was able to run four subjects at once on my quad-core processor. I'm not sure what the exact speedup was, but it was certainly much faster, and pretty simple to implement.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "52301",
      "postDate": "08/21/2014 20:45:17",
      "content": "<p>I decided to avoid the whole problem of feature selection by turning the signals into images and feeding them to nolearn - the image classification neural net that comes pre-trained on a huge image database and which performed well on the Cats vs Dogs challenge. This strategy worked pretty well (92.1 on the final leaderboard), but not as well as the feature selection done by all you other clever folks.</p>\n<p>A description of my strategy is here: <a href=\"http://small-yellow-duck.github.io/seizure_detection.html\" target=\"_blank\">http://small-yellow-duck.github.io/seizure_detection.html</a></p>\n<p><img src=\"http://small-yellow-duck.github.io/seizure_detection/fft_dog_1_chan_15.png\" alt width=\"557\" height=\"268\"></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "52315",
      "postDate": "08/22/2014 16:17:51",
      "content": "<p>This Challenge was a great learning experience for me, and the forum provided a great platform for interactive education.</p>\n<p>I had very little time through and for the Challenge and very little computer resources. However, I had a&nbsp; desire to understand and contribute to this important topic (of seizures).</p>\n<p>I derived mean, median, ...., temporal correlation (of one file with the next file), and spatial correlation (of one channel with the next channel) for each file. Then I used Random Forest in R. The score of about 0.58 was not impressive.</p>\n<p>To improve the score, I started recomputing temporal correlation of a file's beginning part with the file's ending part. However, I could not finish this &quot;experiment&quot;.</p>\n<p>Thanks a lot to all of you for sharing your wisdom.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "52317",
      "postDate": "08/22/2014 16:31:19",
      "content": "<p>[quote=Mike Kim;52248]</p>\n<p>Can people elaborate on computational speed and methods they used in terms of implementation? The data set was fairly large compared to the median Kaggle contest.&nbsp;</p>\n<p>[/quote]</p>\n\n<p>I did all of my processing and modeling on my laptop: 16g ram, 4 core i7. Initial reading and pre-processing of the data took quite some time (i.e., 4+ hours), but once it was downsampled to 400 hz or so, things became more manageable. Once all the processing was done, my actual input files to my models were never larger than 5 mb or so. Some of the models I ran (e.g., bagged MARS) took 10+ hours to complete over all subjects.</p>\n<p>I used R for just about everything. All of my feature creation was done within <a href=\"http://datatable.r-forge.r-project.org/\">data.table</a>. Modeling through <a href=\"http://topepo.github.io/caret/index.html\">caret</a>, but I started to use <a href=\"http://0xdata.com/h2o/\">h2o </a>at the end, which has a more limited number of algorithms (for now), but is SUPER fast. </p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "52319",
      "postDate": "08/22/2014 16:45:19",
      "content": "<p>Thanks to everyone for participating in the competition and for your innovative work. The contest has been a tremendous success, and our congratulations to the winners.</p>\n<p>First, we would like to announce our next competition on kaggle will be on seizure forecasting - identifying features in EEG that may indicate a seizure-permissive brain state, making it more likely a patient will have a seizure in the near future. This contest will begin soon, and will have a larger prize pool. It will use data clips in the same format as this contest, so many of the tools you developed for this contest could be directly applicable.</p>\n<p>Second, we would invite all participants to log in to ieeg.org to continue the seizure detection effort on continuous data. Creating an account takes only a few minutes and allows you to save analyses and upload tools. We will also set up a forum and link area for github repositories for seizure detection. This is a great opportunity for experts in machine learning like yourselves to team up with epilepsy researchers to help advance our understanding of epilepsy.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "52320",
      "postDate": "08/22/2014 16:45:21",
      "content": "<p>[quote=Mike Kim;52248]</p>\n<p>Can people elaborate on computational speed and methods they used in terms of implementation? The data set was fairly large compared to the median Kaggle contest.</p>\n<p>[/quote]</p>\n<p>I processed the data on a workstation with two 3.47G 4-core CPUs with Python sklearn. Downsampling the data to&nbsp;400Hz finishes in minutes and calculating features takes the most time. But it can be paralleled for different subjects and takes about two hours. Then the classification model only needs to&nbsp;load the feature tables every time and runs in parallel with ensemble methods. The classification takes about&nbsp;between 1 to 1.5 hours.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "52441",
      "postDate": "08/25/2014 21:38:00",
      "content": "<p>Hey everyone, I've posted up my documentation and code in another thread, but I thought I'd post here as well as I also wanted to&nbsp;give my answer on the question of how did people handle the size of the data with respect to computational speed etc.</p>\n<p>For feature selection I used FFT 1-47Hz, concatenated with correlation coefficients (and their eigenvalues) of both the FFT output data, as well as the input time data. The data was then trained on per-patient Random Forest classifiers (3000 trees).</p>\n<p>My background is Computer Engineering, so I spent a lot of time optimising my development cycle to be fast. As an example using 4-cores I can do a cross-validation run against all patients on a 150-tree Random Forest using&nbsp;FFT 1-47Hz in under 4 minutes. The processed data is also cached for re-use if I wanted to try another classifier on it. You can check out my code in the other thread, but to summarise the techniques that I used to streamline the code:</p>\n<ul>\n<li>Process each piece of training data as it comes in rather than loading it all and transforming all of it at once. This keeps memory usage down. Before doing this I used to crash my laptop a LOT even&nbsp;with 16GB of RAM.</li>\n<li>Cache all&nbsp;processed data to save recomputing it again&nbsp;later</li>\n<li>For python numpy arrays use hickle (h5py) instead of pickle, loads most data in 0 seconds</li>\n<li>Never hold in memory data you don't need, e.g.&nbsp;for making predictions, first load training data to train the classifier, then drop that data and load the test data for classification. Loading data from hickle format is fast enough that you can forget about load times</li>\n<li>Dump scores to disk so I can re-run previous runs to get immediate scoring output if I forgot to copy and paste it out</li>\n<li>Can specify lists of different data pipelines to try and lists of classifiers to try. The code would run all combinations and print out the cross-validation scores in sorted order at the end. Additionally if I decided to cancel a&nbsp;run, I could run it again from roughly where it left off as any intermediate results were already cached to disk.</li>\n</ul>\n<p>For final classification though I would use 3000 estimators in my Random Forest which blew out run times to 2-3 hours.</p>\n<p>Having a fast SSD is also a huge benefit.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 52212,
      "author_name": "mikeskim",
      "author_url": "",
      "post_date": "08/20/2014 06:17:57",
      "content": "<p>The following text probably won't be quite precise, but the intuition is I just threw a kitchen sink of feature transformations at the problem in the limited time I had (about 2-3? weeks) without thinking about it too much.&nbsp;</p>\n<p>I tried various things at various windows and preprocessed with various filters not all of which was done within the box of conventional signal processing:</p>\n<p>R's summary, variance, sd, mad, IQR, range, R (moment's package) skewness, kurtosis, geary, fft(Arg()) (half of it), mean, median filter, gaussian filter, correlation (upper tri, pearson, of the channels), lagged differences (diff) at various lags, various combinations of the above in different orders.&nbsp;</p>\n<p>Then I ensembled gbm, brnn, glmnet, and randomForest. While brnn and glmnet did fairly poorly, these still helped overall. Everything was really slow since I did this all in R although multiple instances on AWS helped. Maybe some people had better ideas on handling the size of the data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 52220,
      "author_name": "skorsun",
      "author_url": "",
      "post_date": "08/20/2014 08:53:35",
      "content": "<p>I have used the next features:</p>\n<p>maximal and minimal &nbsp;values for&nbsp;each channel;</p>\n<p>standard deviation by&nbsp;channel;</p>\n<p>I had an accuracy about &nbsp;0.92 with Random Forest, KNN and Logistic Regression.</p>\n<p>Then I've added:</p>\n<p>logarithm of the sum of absolute values for each channel;</p>\n<p>logarithm from the first 250 values of&nbsp;FFT for two&nbsp;channels</p>\n<p>and improve accuracy to 0.94&nbsp;.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 52222,
      "author_name": "rybnik",
      "author_url": "",
      "post_date": "08/20/2014 09:12:05",
      "content": "<p>Hi, I thought about hardware constrains and so I decided for 'cheap' time domain features, so I went with zero crossing rate (zcr) and average short time energy (aste). Extracting these features for all of the available channels and running them against a RF I've achieved a 0.89 score.</p>\n<p>This was a fun competition.</p>\n<p>PS: I like Serhii (min, Max, sd) approach :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 52226,
      "author_name": "senecaur",
      "author_url": "",
      "post_date": "08/20/2014 11:24:27",
      "content": "<p>I hand-crafted features based on the following reference: D F Wulsin, J R Gupta, R Mani, J A Blanco and B Litt, Modeling electroencephalography waveforms with semi-supervised deep belief nets: fast classification and anomaly measurement, J. Neural Eng. 8 (2011) 036015</p>\n<p>From Appendix B I extracted the following features for each EEG channel seperately:</p>\n<p>-Normalized positive area under the curve</p>\n<p>-Normalized decay: the chance-corrected fraction of data that is decreasing or increasing</p>\n<p>-Frequency band power: the mean power spectral density in each of the frequency bands 1.5Hz-7Hz, 8Hz-15Hz and 16Hz-29Hz computed using Welch's method</p>\n<p>-Line length: sum of the absolute differences between successive data samples</p>\n<p>-Mean energy: mean of the square of samples across each channel</p>\n<p>-Average peak amplitude: base-10 log of the mean-squared amplitude of each peak in the channel</p>\n<p>-Average valley amplitude: as above but for the valleys</p>\n<p>-Normalized peak number: number of peaks in the channel normalized by the mean difference between adjacent data points</p>\n<p>-Zero crossings: subtracts the mean value (rather than the line of best fit as in the paper) from the channel and then counts how many times the zero-mean data crosses zero</p>\n<p>Looking at the scores achieved above with smaller/simpler feature sets it seems that it was wasted effort to extract some of these.&nbsp; I plateaued very early on in this competition and couldn't figure out how to make any significant improvements to my score.&nbsp; I trained using Scikit ExtraTreesClassifier for both feature selection and classification.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 52233,
      "author_name": "maineiac",
      "author_url": "",
      "post_date": "08/20/2014 14:47:11",
      "content": "<p>This competition was definitely all about feature engineering. Using an ensemble of models (random forest, glmnet, bagged MARS, gbm) I got to 0.95+ with the following:</p>\n<p>Measures of variability: Variance, max-min, 95th-5th percentile</p>\n<p>Average Spectral Power (using wavelets): Delta (0-4 Hz); Theta (4-8); Alpha (8-14); Beta (15-30); lowGamma (30-100); highGamma (100-200)</p>\n<p>Ratio of Spectral Powers to each other (e.g., theta&nbsp; / alpha; all combinations).</p>\n<p>Right at the end, I added the mean and variance of cross-correlation values (i.e. lagged correlation ranging from 1-20 samples), which improved the accuracy of some of my individual models, but in the end, didn't improve the final ensemble.</p>\n<p>No single model of mine had better prediction accuracy of 0.94, and some were as low as 0.87, so combining the predictions definitely helped here.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 52234,
      "author_name": "ohiojoe",
      "author_url": "",
      "post_date": "08/20/2014 15:13:23",
      "content": "<p>I mainly extracted frequency bins of about 3 Hz range and used a hamming window to improve the frequency resolution. Further marginal improvements came by splitting the time series into 'early' and 'late' pieces (crossover hamming windows) and extracting frequency content, as well as ratios between frequency bins.</p>\n<p>I used a random forest model to get scores above .95. If I had used my best public leaderboard score for submission, that would have been good enough for ~ 6th place. I was concerned about overfitting so relied on a 5 fold CV for choosing the best model (in terms of reducing unnecessary features), but this yielded ~ .94 in both. Perhaps a leave-on-out CV would have been better?</p>\n<p>I meant to spend more time at the end developing more features (both time and spectral) but didn't get around to it. Congrats to the winners! Fun competition!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 52235,
      "author_name": "yansoftware",
      "author_url": "",
      "post_date": "08/20/2014 15:18:18",
      "content": "<p>I have evaluated the&nbsp;features individually with the label with AUC score described in the first post. I found that std, mobility, fractal dimensions and gamma are the most important ones. I have also included the 20%, 50%, 80% percentiles of &nbsp;cross correlation&nbsp;values, The attached is what I have got (the number in front indicates the channel):</p>\n<p>Dog_1<br>Seizure top 10: ['4_std' '3_std' '7_std' 'percentile0.2' '2_std' '11_std' '15_std' '12_std''13_std' '9_std']<br>Early seizure top 10: ['15_beta' '11_mobility' '15_gamma' '7_gamma' '15_fractals' '3_gamma''7_fractals' '11_gamma' '3_fractals' '11_fractals']</p>\n<p>Dog_2<br>Seizure top 10: ['14_mobility' '9_mobility' '14_gamma' '10_gamma' '7_mobility''11_mobility' '7_gamma' '15_mobility' '15_gamma' '10_mobility']<br>Early seizure top 10: ['12_fractals' '7_fractals' '11_fractals' '15_gamma' '14_gamma' '10_gamma''7_gamma' '10_fractals' '14_fractals' '15_fractals']</p>\n<p>Dog_3<br>Seizure top 10: ['9_std' '2_std' '3_std' '5_std' '7_std' '14_std' '10_std' '6_std' '12_std''13_std']<br>Early seizure top 10:['11_std' '2_std' '5_std' '3_std' '10_std' '14_std' '12_std' '7_std''13_std' '6_std']</p>\n<p>Dog_4<br>Seizure top 10: ['12_mobility' '7_std' '14_fractals' '12_complexity' '5_gamma' '5_fractals''6_fractals' '5_complexity' '5_mobility' '5_std']<br>Early seizure top 10: ['3_std' 'median' '15_std' '11_std' '13_std' '7_std' '4_std' '12_std''5_std' 'percentile0.2']</p>\n<p>Patient_1<br>Seizure top 10:&nbsp;['10_alpha' '9_fractals' '62_alpha' '19_fractals' '18_alpha' '13_alpha''13_fractals' '12_alpha' '20_alpha' '29_alpha']</p>\n<p>Early seizure top 10: ['18_complexity' '19_mobility' '29_fractals' '20_complexity' '18_fractals''18_mobility' '29_gamma' '29_mobility' '19_complexity' '20_mobility']</p>\n<p>Patient_2<br>Seizure top 10: ['2_std' '3_complexity' '1_complexity' '1_std' '2_beta' '1_beta' '3_beta''0_complexity' '0_std' '0_beta']<br>Early seizure top 10: ['0_std' '2_mobility' '3_complexity' '3_beta' '0_complexity' '0_beta''1_complexity' '2_complexity' '1_beta' '2_beta']</p>\n<p>Patient_3<br>Seizure top 10: ['8_theta' '6_std' '5_delta' '8_beta' '5_beta' '4_alpha' '5_complexity''54_fractals' '4_delta' '4_complexity']<br>Early seizure top 10: ['5_skew' '5_alpha' '54_fractals' '5_complexity' '5_beta' '4_mobility''4_beta' '4_alpha' '4_delta' '4_complexity']</p>\n<p>Patient_4<br>Seizure top 10: ['48_std' '33_gamma' '42_std' '34_std' '8_alpha' '56_fractals' '56_std''33_fractals' '33_std' '49_std']<br>Early seizure top 10: ['48_std' '33_gamma' '42_std' '34_std' '8_alpha' '56_fractals' '56_std''33_fractals' '33_std' '49_std']</p>\n<p>Patient_5<br>Seizure top 10: ['7_mobility' '0_gamma' '0_complexity' '31_fractals' '15_mobility''1_fractals' '7_fractals' '0_mobility' '15_fractals' '0_fractals']<br>Early seizure top 10: ['1_mobility' '7_mobility' '15_fractals' '7_fractals' '0_gamma''15_mobility' '1_fractals' '0_complexity' '0_mobility' '0_fractals']</p>\n<p>Patient_6<br>Seizure top 10: ['14_skew' '23_beta' '22_complexity' '15_alpha' '22_std' '14_delta''23_std' '23_complexity' '14_complexity' '14_alpha']<br>Early seizure top 10: ['22_complexity' '14_skew' '22_fractals' '14_complexity' '23_beta''23_complexity' '23_alpha' '14_alpha' '23_std' '22_std']</p>\n<p>Patient_7<br>Seizure top 10: ['35_theta' '27_theta' '27_std' '31_fractals' '34_fractals' '26_fractals''30_fractals' '27_fractals' '7_std' '27_alpha']<br>Early seizure top 10: ['25_fractals' '12_alpha' '2_beta' '35_theta' 'percentile0.5' '12_fractals''23_gamma' '12_gamma' '1_gamma' 'percentile0.2']</p>\n<p>Patient_8<br>Seizure top 10: ['9_std' '6_mobility' '5_mobility' '9_fractals' '14_mobility' '12_mobility''13_mobility' '11_fractals' '12_fractals' '10_fractals']<br>Early seizure top 10: ['9_beta' '13_delta' '13_complexity' '8_fractals' '14_mobility''15_mobility' '12_delta' '13_mobility' '12_complexity' '12_mobility']</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 52236,
      "author_name": "matthewroos",
      "author_url": "",
      "post_date": "08/20/2014 15:30:30",
      "content": "<p>A quick summary of my model:</p>\n<p>Resample to 500 sps. Extract 0.5 second windows from the beginning, middle, and end of each segment. Apply Hanning windows and compute DFTs. Sum the power in bands 4-8, 8-13, 13-30, and 30-100 Hz and convert to log scale. Discard all but 16 channels (I did this because I was short on time and didn't want to search for a better way to incorporate additional channels). The channels that provided the greatest d-prime discrimination of ictal vs. interictal were retained and ordered by their d-prime values. This resulted in a feature vector of length 3(times)x4(bands)x16(channels) = 192 for each segment. I used SVMs with RBFs and gamma = 1.58. One SVM for each of the predictions we needed to make.</p>\n<p>Matt</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 52247,
      "author_name": "drewabbot",
      "author_url": "",
      "post_date": "08/20/2014 18:02:39",
      "content": "<p>We simply used 5 frequency bands per channel, and&nbsp;used pwelch() for power-spectral-density estimation, using tons of overlap within each segment and a small window size. &nbsp;Using those features alone, the tricky part, from our&nbsp;point of view, seemed to be get rid of all those&nbsp;&quot;spikes&quot; which occurred randomly, and which we interpreted as measurement error. &nbsp;We didn't have enough time to write a robust enough spike detector to improve our score significantly, w.r.t. not&nbsp;removing&nbsp;the spikes.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 52248,
      "author_name": "mikeskim",
      "author_url": "",
      "post_date": "08/20/2014 18:15:26",
      "content": "<p>Can people elaborate on computational speed and methods they used in terms of implementation? The data set was fairly large compared to the median Kaggle contest.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 52262,
      "author_name": "rbroberg",
      "author_url": "",
      "post_date": "08/20/2014 21:10:51",
      "content": "<p>For feature extraction, I used Julia on an Amazon c3.xlarge. Runtimes generally on the order of an hour or less, but I was only looking at fairly simple features.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 52274,
      "author_name": "jstreet",
      "author_url": "",
      "post_date": "08/21/2014 00:48:27",
      "content": "<p>For features I used median, variance, and extracted 6 frequency bands from an FFT. Classification was done using a Random Forest.</p>\n\n<p>The way I handled the data import was quite inefficient. Reading in the data took the majority of the run time for building the model and generating predictions. I rented a EC2 r3.xlarge instance overnight to avoid spending time optimizing my implementation.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 52276,
      "author_name": "george04",
      "author_url": "",
      "post_date": "08/21/2014 03:17:40",
      "content": "<p>The dataset was only a few gb after downsampling the clips to O(100) time units&nbsp;and storing data as single precision (and it didn't seem like much&nbsp;information was lost doing this). &nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 52297,
      "author_name": "dryrun",
      "author_url": "",
      "post_date": "08/21/2014 20:01:28",
      "content": "<p>[quote=Mike Kim;52248]</p>\n<p>Can people elaborate on computational speed and methods they used in terms of implementation? The data set was fairly large compared to the median Kaggle contest.&nbsp;</p>\n<p>[/quote]</p>\n<p>Using joblib's Parallel function was a big help - I was able to run four subjects at once on my quad-core processor. I'm not sure what the exact speedup was, but it was certainly much faster, and pretty simple to implement.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 52301,
      "author_name": "smallyellowduck",
      "author_url": "",
      "post_date": "08/21/2014 20:45:17",
      "content": "<p>I decided to avoid the whole problem of feature selection by turning the signals into images and feeding them to nolearn - the image classification neural net that comes pre-trained on a huge image database and which performed well on the Cats vs Dogs challenge. This strategy worked pretty well (92.1 on the final leaderboard), but not as well as the feature selection done by all you other clever folks.</p>\n<p>A description of my strategy is here: <a href=\"http://small-yellow-duck.github.io/seizure_detection.html\" target=\"_blank\">http://small-yellow-duck.github.io/seizure_detection.html</a></p>\n<p><img src=\"http://small-yellow-duck.github.io/seizure_detection/fft_dog_1_chan_15.png\" alt width=\"557\" height=\"268\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 52315,
      "author_name": "lalitapatel",
      "author_url": "",
      "post_date": "08/22/2014 16:17:51",
      "content": "<p>This Challenge was a great learning experience for me, and the forum provided a great platform for interactive education.</p>\n<p>I had very little time through and for the Challenge and very little computer resources. However, I had a&nbsp; desire to understand and contribute to this important topic (of seizures).</p>\n<p>I derived mean, median, ...., temporal correlation (of one file with the next file), and spatial correlation (of one channel with the next channel) for each file. Then I used Random Forest in R. The score of about 0.58 was not impressive.</p>\n<p>To improve the score, I started recomputing temporal correlation of a file's beginning part with the file's ending part. However, I could not finish this &quot;experiment&quot;.</p>\n<p>Thanks a lot to all of you for sharing your wisdom.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 52317,
      "author_name": "maineiac",
      "author_url": "",
      "post_date": "08/22/2014 16:31:19",
      "content": "<p>[quote=Mike Kim;52248]</p>\n<p>Can people elaborate on computational speed and methods they used in terms of implementation? The data set was fairly large compared to the median Kaggle contest.&nbsp;</p>\n<p>[/quote]</p>\n\n<p>I did all of my processing and modeling on my laptop: 16g ram, 4 core i7. Initial reading and pre-processing of the data took quite some time (i.e., 4+ hours), but once it was downsampled to 400 hz or so, things became more manageable. Once all the processing was done, my actual input files to my models were never larger than 5 mb or so. Some of the models I ran (e.g., bagged MARS) took 10+ hours to complete over all subjects.</p>\n<p>I used R for just about everything. All of my feature creation was done within <a href=\"http://datatable.r-forge.r-project.org/\">data.table</a>. Modeling through <a href=\"http://topepo.github.io/caret/index.html\">caret</a>, but I started to use <a href=\"http://0xdata.com/h2o/\">h2o </a>at the end, which has a more limited number of algorithms (for now), but is SUPER fast. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 52319,
      "author_name": "bbrinkm",
      "author_url": "",
      "post_date": "08/22/2014 16:45:19",
      "content": "<p>Thanks to everyone for participating in the competition and for your innovative work. The contest has been a tremendous success, and our congratulations to the winners.</p>\n<p>First, we would like to announce our next competition on kaggle will be on seizure forecasting - identifying features in EEG that may indicate a seizure-permissive brain state, making it more likely a patient will have a seizure in the near future. This contest will begin soon, and will have a larger prize pool. It will use data clips in the same format as this contest, so many of the tools you developed for this contest could be directly applicable.</p>\n<p>Second, we would invite all participants to log in to ieeg.org to continue the seizure detection effort on continuous data. Creating an account takes only a few minutes and allows you to save analyses and upload tools. We will also set up a forum and link area for github repositories for seizure detection. This is a great opportunity for experts in machine learning like yourselves to team up with epilepsy researchers to help advance our understanding of epilepsy.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 52320,
      "author_name": "yansoftware",
      "author_url": "",
      "post_date": "08/22/2014 16:45:21",
      "content": "<p>[quote=Mike Kim;52248]</p>\n<p>Can people elaborate on computational speed and methods they used in terms of implementation? The data set was fairly large compared to the median Kaggle contest.</p>\n<p>[/quote]</p>\n<p>I processed the data on a workstation with two 3.47G 4-core CPUs with Python sklearn. Downsampling the data to&nbsp;400Hz finishes in minutes and calculating features takes the most time. But it can be paralleled for different subjects and takes about two hours. Then the classification model only needs to&nbsp;load the feature tables every time and runs in parallel with ensemble methods. The classification takes about&nbsp;between 1 to 1.5 hours.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 52441,
      "author_name": "michaelhills",
      "author_url": "",
      "post_date": "08/25/2014 21:38:00",
      "content": "<p>Hey everyone, I've posted up my documentation and code in another thread, but I thought I'd post here as well as I also wanted to&nbsp;give my answer on the question of how did people handle the size of the data with respect to computational speed etc.</p>\n<p>For feature selection I used FFT 1-47Hz, concatenated with correlation coefficients (and their eigenvalues) of both the FFT output data, as well as the input time data. The data was then trained on per-patient Random Forest classifiers (3000 trees).</p>\n<p>My background is Computer Engineering, so I spent a lot of time optimising my development cycle to be fast. As an example using 4-cores I can do a cross-validation run against all patients on a 150-tree Random Forest using&nbsp;FFT 1-47Hz in under 4 minutes. The processed data is also cached for re-use if I wanted to try another classifier on it. You can check out my code in the other thread, but to summarise the techniques that I used to streamline the code:</p>\n<ul>\n<li>Process each piece of training data as it comes in rather than loading it all and transforming all of it at once. This keeps memory usage down. Before doing this I used to crash my laptop a LOT even&nbsp;with 16GB of RAM.</li>\n<li>Cache all&nbsp;processed data to save recomputing it again&nbsp;later</li>\n<li>For python numpy arrays use hickle (h5py) instead of pickle, loads most data in 0 seconds</li>\n<li>Never hold in memory data you don't need, e.g.&nbsp;for making predictions, first load training data to train the classifier, then drop that data and load the test data for classification. Loading data from hickle format is fast enough that you can forget about load times</li>\n<li>Dump scores to disk so I can re-run previous runs to get immediate scoring output if I forgot to copy and paste it out</li>\n<li>Can specify lists of different data pipelines to try and lists of classifiers to try. The code would run all combinations and print out the cross-validation scores in sorted order at the end. Additionally if I decided to cancel a&nbsp;run, I could run it again from roughly where it left off as any intermediate results were already cached to disk.</li>\n</ul>\n<p>For final classification though I would use 3000 estimators in my Random Forest which blew out run times to 2-3 hours.</p>\n<p>Having a fast SSD is also a huge benefit.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "52208": "",
    "52212": "",
    "52220": "",
    "52222": "",
    "52226": "",
    "52233": "",
    "52234": "",
    "52235": "",
    "52236": "",
    "52247": "",
    "52248": "",
    "52262": "",
    "52274": "",
    "52276": "",
    "52297": "",
    "52301": "",
    "52315": "",
    "52317": "",
    "52319": "",
    "52320": "",
    "52441": ""
  },
  "source": "meta"
}