{
  "id": 12604,
  "title": "Congratulations to the winners and top entries!",
  "url": "/competitions/inria-bci-challenge/discussion/12604",
  "author_name": "",
  "post_date": "2015-02-25T00:25:16.780Z",
  "votes": 4,
  "comment_count": 23,
  "views": 5502,
  "content": "<p>Congratulations to the winners, top entries and new Master Kagglers!</p>\n<p>What was everyone's approach?</p>",
  "messages": [
    {
      "id": "64830",
      "postDate": "02/25/2015 00:25:16",
      "content": "<p>Congratulations to the winners, top entries and new Master Kagglers!</p>\n<p>What was everyone's approach?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64833",
      "postDate": "02/25/2015 00:44:16",
      "content": "<p>My approach was to overfit continuously until I dropped around 50 maybe 60 places. Then I would drink a beer.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64834",
      "postDate": "02/25/2015 00:45:46",
      "content": "<p>Congrats to all!</p>\n<p>Coming off the end&nbsp;of the last competition, I didn't have a lot of&nbsp;time for this contest, and with the small test set&nbsp;size it's hard to tell whether what I did was actually good... but here's what I tried.</p>\n<p>1. Band-pass filtered and normalized the&nbsp;EEG signals.</p>\n<p>2. Trained&nbsp;stacked convolutional autoencoders with three hidden layers on all the EEG data. I guessed at some architectures and hyper-parameters that I thought might&nbsp;reasonably capture&nbsp;the EEG signal. There is&nbsp;probably room for optimization the architectures. I tried&nbsp;three different architectures of varying convolution sizes and hidden unit sizes.&nbsp;The auto encoder with the fewest parameters seemed to do best.</p>\n<p>3. I then fed the features&nbsp;at the top layer of the autoencoders into a logistic regression and&nbsp;randomized tree classifiers. My thought was that these algorithms&nbsp;would be better at reducing&nbsp;overfitting.</p>\n<p>4. Finally, I ensembled the different model predictions using a logistic regression.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64838",
      "postDate": "02/25/2015 01:09:31",
      "content": "<p>Congrats to the winners. I'm interested to hear about your approaches. This competition was a bit frustrating for me, both because the leaderboard was so clearly off, and because I found it difficult to get a grasp on good features.</p>\n<p>My approach:</p>\n<p>1. Standard pre-processing, such as down-sampling by half and baseline correction prior to the feedback.</p>\n<p>2. ICA on a subsample of (mostly) central electrodes, but done on error trials only.</p>\n<p>3. Upsampling the data to deal with class imbalance (in some models).</p>\n<p>4. Adding raw data from a few central electrodes (Cz, PCz, FCz) into the model with the ICA components (some models).</p>\n<p>5. Model averaging</p>\n<p>As far as models, nothing special: random forests, svms with a radial kernel (not so good, but added diversity), gbms. Tried some multi-layer NNs with dropout, but these didn't seem to be learning much.</p>\n\n<p>Things that didn't work for me (but maybe should have based on other efforts I read about):</p>\n<p>1. Features based on spectral power (either raw spectra, or derived features such as mean power in the theta range 4-8 hz)</p>\n<p>2. Amplitude-based features (e.g., peak-to-peak amplitudes around 200-400 and 400-600 ms to capture the Ne and Pe)</p>\n<p>3. Quantifying slow cortical potentials as linear changes in the signal over the trial or mean differences from the first 100 ms</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64847",
      "postDate": "02/25/2015 05:51:32",
      "content": "<p>If you like shakeups, this competition ranks up there close to MLSP2014 Schizophrenia and Big Data Combine:</p>\n<p>&gt; shakeup('http://www.kaggle.com/c/inria-bci-challenge/leaderboard')<br>Joining by: id<br>$shakeup.top<br>[1] 0.1610839</p>\n<p>$shakeup.all<br>[1] 0.2885489</p>\n<p>Source: https://www.kaggle.com/c/liberty-mutual-fire-peril/forums/t/10187/quantifying-leaderboard-shake-up/52910</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64856",
      "postDate": "02/25/2015 11:21:43",
      "content": "<p>I dropped an impressive 87 places so that's cool. My 10-fold AUC was 0.78 and public leaderboard wasn't far off but didn't generalize well.&nbsp;</p>\n<p>My approach was pretty simple.&nbsp;</p>\n<p>1. apply bandpass filter between 1 - 20 Hz</p>\n<p>2. For every feedback event, gather 1 sec trajectory for every channel</p>\n<p>3. Pool trajectories per channel and apply functional principal component analysis. FPCA provides a set of coefficients that summarize functional variation</p>\n<p>4. Apply random forest to data reduction (410 features for all channels (1 sec data)) + feedback order</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64859",
      "postDate": "02/25/2015 11:56:01",
      "content": "<p>Preprocessing:</p>\n<ul>\n<li>Linear de-trend</li>\n<li>Remove freqs &gt; 32 Hz</li>\n<li>Normalization</li>\n<li>Used window of 25-650 ms after event</li>\n</ul>\n<p>Model:</p>\n<ul>\n<li>PLS (usually 4-5 components performed best)</li>\n<li>Brute force testing of channel pairs that gave best CV result</li>\n<li>Continue to add&nbsp;next best channel until performance dropped.</li>\n</ul>\n<p>Best channels were:</p>\n<p>['Fp1', 'CP4', 'AF7', 'AF8', 'Fz', 'CP2']</p>\n<p>Repeat using single-layer perceptron and GBM, and average models.</p>\n<p>Only dropped 2 places from public to private LB, so CV was pretty good I guess.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64860",
      "postDate": "02/25/2015 13:07:23",
      "content": "<p>I adapted methods from <a href=\"http://eprints.pascal-network.org/archive/00008033/01/BlaLemTreHauMue10.pdf\">single trial ERP tutorial</a>, and paper on <a href=\"http://www.researchgate.net/profile/J_Farquhar/publication/234112682_FarquharHill_ERP_Classification_Best_Practice_Author_Pre-print/links/0fcfd50f46f47278f0000000.pdf\">preprocessing and classification of single trial ERP</a>.</p>\n<p>Preprocessing:</p>\n<ul>\n<li>Used pandas rolling window of size thirty to remove high frequency information</li>\n<li>Baseline corrected with mean of 200 ms before feedback onset for each channel</li>\n<li>Took 200ms to 1000ms after feedback and downsampled to 7 samples per channel</li>\n<li>Concatenated all channels for each trial into single feature</li>\n<li>Normalized with sklearn standard scaler</li>\n</ul>\n<p>Classification:</p>\n<ul>\n<li>Used Linear Discriminant analysis with shrinkage from sklearn 0.16 dev</li>\n<li>Additionally used Logistic Regression with L1 or L2 regularization</li>\n<li>averaged 6 classifiers with different parameters</li>\n</ul>\n<p>Private leaderboard score without label counts from leak was ~0.710. Adding the label counts from the leak increased my score by about .13 on private leaderboard, though it only increased my CV score by 0.05. I additionally added label counts from my own classifier and this added an additional 0.01 to my score.</p>\n<p>I several other things which I couldn't get to work correctly.</p>\n<ul>\n<li>Latency correction for ERP onset between subjects</li>\n<li>Standard bandpass gave higher score for two subjects but couldn't tune for others</li>\n<li>Channel selection. I overfit doing recursive channel elimination. It turned out that the best channels were quite different depending on which two subjects I left out.</li>\n<li>Clustering of subjects and training on similar subjects</li>\n<li>many other things I have forgotten</li>\n</ul>\n<p>I am interested to hear if any one used transfer learning techniques and/or explicitly attempted to calibrate inter-subject probabilities for ROC AUC.</p>\n<p>I am satisfied with my results but I feel like I missed some easy things to improve score. Interested to see what overfitting avengers did. I read a couple of Alaxandre Barachant's papers on MDM classifier using Riemann Geometry. Wanted to try to implement but seemed slightly intimidating. Is this what you guys used?</p>\n<p>Edit: links to the first paper is broken. the first paper is called &quot;<em>Single-trial analysis and classification of ERP components&#8212;a tutorial</em>&quot; by Blankertz et al,</p>\n<p>the second paper is called: &quot;<em>Interactions between Pre-Processing and Classification </em><em>Methods for Event-Related-Potential Classification: </em><em>Best-Practice Guidelines for Brain-Computer Interfacing</em>&quot; by Farquhar and Hill.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64862",
      "postDate": "02/25/2015 14:18:32",
      "content": "<p>[quote=Mike Kim;64833]</p>\n<p>My approach was to overfit continuously until I dropped around 50 maybe 60 places. Then I would drink a beer.</p>\n<p>[/quote]</p>\n<p>Did you have the beer?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64864",
      "postDate": "02/25/2015 14:25:06",
      "content": "<p>I'm really impressed at the final performance of the winning team. I'm also interested in their approach, because Alexandre already won the competition on MEG with a large margin.</p>\n<p>Phalaris, congratulation for your second place. I'm also impressed at your good result in spite of its simplicity (no Random Forest, no Convolutional NNs, etc). I think that in this kind of applications, a sensible preprocessing as the one you performed is more important than using fancy classification techniques (one question remains, though: did you use a <em>pool</em> approach, i.e. use data from all subjects to train a common classifier?)</p>\n<p>Also, keeping things simple is important to avoid overfitting. The fact that the public LB showed the performance on 2 of the 10 subjects suggested that there would be surprises on the private LB. In my case, I jumped from the 63th to the 5th place!</p>\n<p>---------------------------------------------------------</p>\n<p>My approach:</p>\n<p>Preprocessing:</p>\n<p>- Band-Pass filtering between 1 and 10 Hz with a FIR filter with order 512.&nbsp; Starting from the 30th sample, we keep 25 samples applying a down-sample factor of 4.</p>\n<p>Features (or 1st level classification):</p>\n<p>-For each of the 25 time-lags and for each of the 16 subjects,&nbsp; obtain two spatial filters using LDA and Log. Regression, respectively. Use 240 of the 340 examples to do this. The EOG channel is included here with the aim of reduce the contribution of artifacts.</p>\n<p>(2nd level) Classification</p>\n<p>- Using different data than those used for extracting the filters (the remaining 100 examples), we train an SVM with RBF kernel.</p>\n\n<p>The approach described is based on the Stacked Generalization procedure: in the first level we solve a set of per-subject classification problems, and in the second one we combine features obtained with filters trained on the 16 subjects; this way, we expect to capture some diversity.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64868",
      "postDate": "02/25/2015 14:59:25",
      "content": "<p>Nice to have some insight into the methods used by you all!</p>\n<p>Our&nbsp;approach was&nbsp;similar to that of Phalaris, although I did not use the info of the 'leakage'.</p>\n<p>On the Public board we got stuck around 0.71% corresponding to place 154.&nbsp;However, I was somewhat surprised to see the Private place to be 25, a jump up of 129 places.</p>\n<p>Features:</p>\n<ul>\n<li><span style=\"line-height: 1.4\">Bandpass filtered the data 1-20Hz.</span></li>\n<li>Feature reduction by averaging time-windows (different window lengths) per trial in 0-1000ms interval. (similar to downsample).</li>\n<li>Concatenated all channels and did a z-score normalization.</li>\n</ul>\n<p>Classification:</p>\n<ul>\n<li><span style=\"line-height: 1.4\">Linear Discriminant Analysis with shrinkage 0.10.</span></li>\n<li><span style=\"line-height: 1.4\">Combined with&nbsp;Logistic Regression</span></li>\n<li>Averaged the&nbsp;classifiers&nbsp;for different features/settings</li>\n</ul>\n<p>Explored but not beneficial for us:</p>\n<ul>\n<li><span style=\"line-height: 1.4\">Step-wise feature selection, in which features are added/removed (by setting a threshold in p-value) to form a robust feature-subspace model. (e.g. as applied in swLDA).</span></li>\n<li><span style=\"line-height: 1.4\">Explored Least Squares Support Vector Machines and the GBM - did not improve the overall.</span></li>\n</ul>\n<p>In the end we scored 0,69 - however, without the leakage information. Our result was quite stable and not prone to overfitting the public LB, as evident from the large increase in the final ranking.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64873",
      "postDate": "02/25/2015 16:14:03",
      "content": "<p>I did a model with only a few features. These features probably are not fully available at run-time and I myself consider them a bit &quot;leaky&quot;.</p>\n<p>A) Length after feedback event, before the next feedback event.</p>\n<p>B) Length after last feedback event, till current feedback event</p>\n<p>C) A - B</p>\n<p>D) Feedback event counter</p>\n<p>E) Session counter</p>\n<p>F) Average variance over all channels / average energy</p>\n<p>This scored 0.64 public LB and 0.56 private LB. Leave two-subjects out CV was 0.61. This model got ensembled with the other models made by team members. Luckily Phil&nbsp;convinced me that CV was far more important than LB in this competition, so we did not end up with only 0.76 public LB&nbsp;-&nbsp;0.47 private LB models.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64877",
      "postDate": "02/25/2015 16:40:44",
      "content": "<p>We applied our new BCI framework for joint optimization [1] that will be presented at NER15. It is based on the Alternating Direction Method of Multipliers [2] and jointly optimizes 3 core objectives (discriminativity, compactness, and robustness) using an objective function that consists of</p>\n<p>(1) least-squares regression term f(x) to model discriminative brain activity patterns,</p>\n<p>(2) sparsity inducing L1 penalty regularization term g(x) to generate compact models and</p>\n<p>(3) sum-of-squares regularization term h(x) for unsupervised transfer learning.</p>\n<p>To remove inter-subject influences, differences between the mean feature vectors of each of the 16 training subjects and each of the 10 testing subjects (160 vectors) were penalized by the sum-of-squares regularization h(x).</p>\n<p>Features: Combination of time-domain amplitudes (0-800ms), power spectral density features (1 Hz wide bins, Welch's method), and meta-data features (session numbers and delay to previous trial) into a 1570 dimensional feature vector.<br>We intentionally did not include the data leakage information in our system.</p>\n<p>In a post-processing step we adjusted the predictions for sessions 1-4 of each subject to integrate prior information about the two different types of copy spelling (fast and slow condition) by adding a constant to long delayed trials.</p>\n<p>Regularization parameters were optimized using a 4-fold cross-validation with splitting on subject bounds, which represented the results on the private LB well.&nbsp;</p>\n<p>Our system scored 0.81124 on the public LB and 0.7457 on the private LB, which is 6th place in the final ranking.</p>\n<p><br>[1] Dominic Heger, Christian Herff, Felix Putze, Tanja Schultz, &#8220;Joint Optimization for Discriminative, Compact and Robust Brain-Computer Interfacing,&#8221; in 7th International IEEE EMBS Neural Engineering Conference 2015 (accepted).<br>[2] Stephen Boyd, et al. &quot;Distributed optimization and statistical learning via the alternating direction method of multipliers.&quot; Foundations and Trends in Machine Learning 3.1 (2011): 1-122.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64882",
      "postDate": "02/25/2015 17:41:56",
      "content": "<p>@Jose M. I used single classifier for all subjects. Glad to see that using single subject classifiers worked so well. Your score is 0.03 better then what I got without information from leak.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64883",
      "postDate": "02/25/2015 18:21:02",
      "content": "<p>[quote=Triskelion;64873]</p>\n<p>&nbsp;Luckily Marios convinced me&nbsp;</p>\n<p>[/quote]</p>\n<p>My true name is revealed! I 've been hiding it for the last 15 years!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64901",
      "postDate": "02/26/2015 01:12:45",
      "content": "<p>Good job, everyone! <br><br>I joined this competition late, so much of my technique was influenced by time constraints.<br><br>My approach was to use figures 7 and 9 from&nbsp;http://www.hindawi.com/journals/ahci/2012/578295/ to determine a handful of channels that were likely to be most relevant. I calculated a running sum over each of those channels, and used both the raw data and the running sums as features. <br><br>I used PCA to decrease the number of features even further to about 100. I then did Linear regression with recursive feature elimination to get my final result. I tried some other methods (Random forest, Lasso, ElasticNet, etc.), but I found that the minor improvements were not worth the additional time requirements. I'm going to clean up my code and upload it to github in case anyone is interested.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64904",
      "postDate": "02/26/2015 02:35:36",
      "content": "<p>Is it just me, or have the Private LB scores of the individual models not yet been placed in the &quot;My Submissions&quot; tab? Without that info, how does anyone know &nbsp;which of their two selections gave the better result??&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64905",
      "postDate": "02/26/2015 02:37:36",
      "content": "<p>Bruce, you can send in submissions and they will show you how they would have scored.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64906",
      "postDate": "02/26/2015 02:38:59",
      "content": "<p>I see. Thanks!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64925",
      "postDate": "02/26/2015 10:20:14",
      "content": "<p>It's really interesting to learn of the other approaches taken.&nbsp;&nbsp;My approach was as follows</p>\n<p>Preprocessing:</p>\n<ul>\n<li><span style=\"line-height: 1.4\">Bandpass filter all channels to between 1 and 20Hz using the Butterworth filter.</span></li>\n</ul>\n<p>Feature extraction:</p>\n<ul>\n<li>Meta features (timestamp, session number etc.)</li>\n<li><span style=\"line-height: 1.4\">Mean of EEG values for each channel in windows of various lengths and lags after the feedback event. These were based on <a href=\"http://videolectures.net/bbci09_blankertz_muller_mlasp/\">this video lecture</a>.</span></li>\n<li>Template matching features. Here I averaged all ERP signals labelled as positive feedback in the training set to create a 'positive feedback template' and similarly to create a&nbsp;'negative feedback template'. This was done separately for each channel. I then used&nbsp;correlation, cross-correlation, covariance, cross-covariance and euclidean distance between the templates and test ERP signals as features. These features were inspired by <a href=\"http://www.sciencedirect.com/science/article/pii/0168559794900515\">this paper</a>.</li>\n</ul>\n<p>Model:</p>\n<p>I took a weighted average of a regularised SVM with a linear kernel (meta and mean EEG features) and a&nbsp;regularised SVM with a linear kernel (meta and template matching&nbsp;features). Putting all the features into one model didn't seem to improve predictions over just using either the mean EEG or template matching features separately. I'm still&nbsp;haven't got a handle on why this would be the case.</p>\n<p>Cross validation:</p>\n<p>I went for 4 fold CV cut on a per subject basis and calculated the AUC on the 4 subjects in the test fold only. To me, this seemed like the best way to 'mimic' the scoring used for the private leader-board. I'd be interested to learn what CV approach others took and if there was a more principled way which I missed. My best model had a CV score of 0.75, a private leader-board score of 0.76921 and public leader-board score of ~0.77, which I guess are all&nbsp;in the same ballpark.&nbsp;</p>\n<p>Things that didn't work or I didn't use</p>\n<ul>\n<li>Logistic regression with elastic net regularisation (gave slightly worse results than the linear SVMs)</li>\n<li>GBT, SVMs with RBF kernel (gave very poor CV scores)</li>\n<li>Spectral features</li>\n<li>Features based on peak picking</li>\n<li>Dynamic time warping distance between templates and ERP signal</li>\n<li>Clustering test subjects with similar training subjects. My intuition was that this would&nbsp;work well but I notice Phalaris also tried it and couldn't get it to work, so it wasn't just me!</li>\n</ul>\n<p>@Maineiac, I noticed you used ICA. How much did it help?&nbsp;It seems to get a good&nbsp;&nbsp;rep in the literature but&nbsp;in the end I didn't have the energy&nbsp;to implement&nbsp;it!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64947",
      "postDate": "02/26/2015 14:07:19",
      "content": "<p>In the interests of sharing, here's the approach I took:</p>\n<p>- For each word spelled, determine whether it was a fast or slow speed test, and use this as a constant feature value for the feedback events in that word (determined&nbsp;by comparing each word time, proxied through delta between first and last feedback events, to the min and max word times for that session, and accounting for the extra delay after each word is completed).&nbsp;</p>\n<p>-&nbsp;Look into the &quot;leak&quot; information that phalaris had mentioned, and create a time_ratio feature for each subject, based on the timings in their session 5 (skipping the first feedback event in each session, since there was some noise in the start timings), so essentially a gauge of each feedback event's post-feedback time to the min/max range of the feedback times observed for that session (the closer to the min time, the less error correction was done). This to me seemed more like a proxy for when the existing algorithm thought there was an error, not necessarily when there actually was an error - so using it as a feature was like boosting the existing algo.</p>\n<p>- I then labelled each user with this time ratio, and ran an OLS between the user's time ratios and their % accuracy overall, on fast, and on slow runs, and obtained a beta that could then be used to approximate the % accuracies of the test users based on their observed test_ratio values. Here's a snapshot of the actual vs estimated % accuracies. For the training set, the correlations between the time_ratios and the actual accuracies were 0.695 for fast, 0.643 for slow, 0.694 for combined.</p>\n<p>Training set</p>\n<p>Fit[02]: all=[0.65 0.77] fast=[0.60 0.71] slow=[0.74 0.88]<br>Fit[06]: all=[0.93 0.92] fast=[0.91 0.85] slow=[0.97 1.05]<br>Fit[07]: all=[0.90 0.88] fast=[0.85 0.81] slow=[0.99 1.00]<br>Fit[11]: all=[0.66 0.59] fast=[0.60 0.55] slow=[0.78 0.68]<br>Fit[12]: all=[0.56 0.62] fast=[0.46 0.57] slow=[0.73 0.71]<br>Fit[13]: all=[0.52 0.62] fast=[0.46 0.57] slow=[0.64 0.70]<br>Fit[14]: all=[0.63 0.78] fast=[0.54 0.72] slow=[0.80 0.89]<br>Fit[16]: all=[0.62 0.61] fast=[0.54 0.56] slow=[0.77 0.70]<br>Fit[17]: all=[0.66 0.60] fast=[0.63 0.56] slow=[0.73 0.69]<br>Fit[18]: all=[0.77 0.75] fast=[0.73 0.70] slow=[0.83 0.86]<br>Fit[20]: all=[0.69 0.66] fast=[0.61 0.61] slow=[0.85 0.75]<br>Fit[21]: all=[0.92 0.67] fast=[0.89 0.62] slow=[0.97 0.77]<br>Fit[22]: all=[0.92 0.82] fast=[0.90 0.76] slow=[0.97 0.94]<br>Fit[23]: all=[0.65 0.65] fast=[0.59 0.60] slow=[0.76 0.74]<br>Fit[24]: all=[0.70 0.68] fast=[0.61 0.63] slow=[0.87 0.78]<br>Fit[26]: all=[0.53 0.68] fast=[0.49 0.63] slow=[0.61 0.78]</p>\n<p>Test set</p>\n<p>Fit[01]: all=[0.00 0.80] fast=[0.00 0.74] slow=[0.00 0.92]<br>Fit[03]: all=[0.00 0.66] fast=[0.00 0.61] slow=[0.00 0.75]<br>Fit[04]: all=[0.00 0.87] fast=[0.00 0.81] slow=[0.00 1.00]<br>Fit[05]: all=[0.00 0.39] fast=[0.00 0.36] slow=[0.00 0.45]<br>Fit[08]: all=[0.00 0.67] fast=[0.00 0.62] slow=[0.00 0.77]<br>Fit[09]: all=[0.00 0.77] fast=[0.00 0.71] slow=[0.00 0.88]<br>Fit[10]: all=[0.00 0.84] fast=[0.00 0.78] slow=[0.00 0.96]<br>Fit[15]: all=[0.00 0.86] fast=[0.00 0.79] slow=[0.00 0.98]<br>Fit[19]: all=[0.00 0.74] fast=[0.00 0.69] slow=[0.00 0.85]<br>Fit[25]: all=[0.00 0.58] fast=[0.00 0.54] slow=[0.00 0.67]</p>\n<p><br>These became additional &quot;fitness&quot; features in the model, bucketed into deciles.</p>\n<p>- All channel values were normalized into z-scores, based on the 0-260 timeslots post feedback event, and baselined to zero at the first feedback event. These became features, along with the mean value for that subject / channel / timeslot, to gauge how different the current feedback run was compared to their usual, between the 50 and 150 timeslot bucket</p>\n<p>- Had a couple of features based on the position of the min and max zscore values for each channel, relative to the 100 timeslot&nbsp;</p>\n<p>- session id flag features</p>\n<p>- Also had channel correlation features, and volatility of correlation features pre and post feedback event</p>\n<p>- A bunch of other ideas that didn't pan out, including general subject features based on typical volatility of response variation etc</p>\n<p>In terms of model, I essentially used a logistic regression via a no hidden layer NN, and experimented with input dropout, various regularization parameter values, de-noising, different cost functions.</p>\n<p>The CV setup was as follows:</p>\n<p>-&nbsp;Randomly shuffle the set of all available feedback events&nbsp;</p>\n<p>- pick 15% of the shuffled events for an ensemble validation set</p>\n<p>- With the remaining set of events, split into K-fold CV by subject, picking 3 of each type of subject (i split them into good | medium | poor rating subjects) for training, and the remaining for CV</p>\n<p>- Average the resulting model from each fold progressively together, and validate its result improves the ensemble validation set</p>\n<p>- Run the above with multiple seed values and average the resulting models</p>\n<p>I was obviously disappointed with the reshuffle from 5th to 66th but turns out I managed to overfit somewhere along the way, since my ensemble CV result was pretty well&nbsp;tracking the public LB score and improvement. To try to improve and figure out what happened, I spent a couple of hours post competition backtracking features to figure out where the overfit took place, and noticed that even with just the fitness and session features, plus a couple of channels, I could get a&nbsp;0.79471 4th place private leaderboard submission. but tweaking the L1 up or down a decile ranged the rankings between mid twenties and mid sixties.</p>\n<p>tough dataset!&nbsp;</p>\n<p>appreciate any comments or suggestions for improvement to the CV process from those who had a more robust&nbsp;process.</p>\n<p>- one additional observation is that the _min_ time between feedback events seemed to fairly linearly increase with each subject, which made me wonder if there was some form of inherent bias in the test machine, since later subjects had more time in general than earlier subjects</p>\n<p>cheers</p>\n<p>ash</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65036",
      "postDate": "02/27/2015 15:15:53",
      "content": "<p>Will you guys post/share&nbsp;the code?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65040",
      "postDate": "02/27/2015 16:03:10",
      "content": "<p>I will post mine as soon as I clean it up. It ranked 57th on the private leaderboard.</p>\n<p>[quote=Asymptote;65036]</p>\n<p>Will you guys post/share&nbsp;the code?</p>\n<p>[/quote]</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65064",
      "postDate": "02/27/2015 19:43:21",
      "content": "<p>Here is some info about the 3rd place team.</p>\n<p>As expressed in <a href=\"https://www.kaggle.com/c/inria-bci-challenge/forums/t/12603/inter-subject-auc/64843#post64843\">another&nbsp;thread</a>, success of our solution is more to do with <a href=\"https://www.kaggle.com/wiki/Leakage\">leakage</a>&nbsp;mining than signal processing. My signal processing efforts alone were top 15% or so, though there were some good ideas I didn't get to because&nbsp;I focused on the leakage.</p>\n<p>I attempted to clean mine up a bit, and most of it is on my <a href=\"https://github.com/mlandry22/kaggle/tree/master/bci_challenge\">github repo</a>--enough that running that code will get third. It's very ugly as a result of (1) spending the last few days hammering out the leakage; and (2) I write ugly code.</p>\n<p><strong>Non-leakage efforts:</strong></p>\n<ul>\n<li>Min/max scaled every series of 0.0 - 1.0 seconds for Cz, used 15-period moving average for the basis of most hand-engineered features.</li>\n<li>Sent the above 0.0-1.0 range, used those values as is, but trimmed the timepoints used in modeling&nbsp;based on GBM importance values.</li>\n<ul>\n<li>The range around 70-95 was consistently favored by most models, including the unscaled python forum code.</li>\n</ul>\n<li>Based on the full 0.0-1.0 range from above calculated several features that did well in basic AUC testing</li>\n<ul>\n<li>sum(0.0 - 0.5) - sum(0.5 - 1.0). So basically subtract the sum of the scaled signal in the first half-second minus the second half-second.</li>\n<li>the location of the minimum point between 20 and 70 readings after feedback. I.e. the time point at which a local minimum was reached.</li>\n<li>several other simple calculations: raw range of sample, range of first to last point, standard deviation, volatility. These didn't have a lot of influence.</li>\n</ul>\n<li>Time frequency kernel to adjust for EOG</li>\n<ul>\n<li>Teammate had some nice prior experience with EMGs&nbsp;in matlab, so we used octave here</li>\n<li>Used a time frequency kernel to generate about 500 vectors per time point.</li>\n<li>TF kernel vectors were used in PLS regression with the EOG signal as the target</li>\n<li>Top 10 pls coefficients were&nbsp;used as features</li>\n<li>This had a small but reliable impact. The model consistently favored the 4th pls coefficient, which is a little peculiar.</li>\n</ul>\n<li>Other EOG reduction methods</li>\n<ul>\n<li>Scaled subtraction. Min/max scale the EOG signal as with the Cz signal, then scale the max of the EOG result to wherever it experienced the largest gap between the Cz signal.</li>\n<li>Straight differencing:&nbsp;Cz - EOG.</li>\n<li>Each of these two provided some small gains.</li>\n</ul>\n</ul>\n<p><strong>Leakage efforts:</strong></p>\n<ul>\n<li>Applied <a href=\"https://www.kaggle.com/c/inria-bci-challenge/forums/t/11178/possible-data-leakage\">phalaris's discovery</a>&nbsp;of the importance of timing in the fifth session. Note that this has been available&nbsp;for about a month.</li>\n<li>Identify the extended time by:</li>\n<ul>\n<li>accounting for the extra time at the end of a word</li>\n<li><span style=\"line-height: 1.4\">accounting for the &quot;drift&quot; in the cutline of what expected time and extra time looked like. I.e. the first few subjects had a cutpoint that won't work for later subjects, so you have to take that into account.</span></li>\n<li><span style=\"line-height: 1.4\">guessing at the cutpoint for S25 and S26. These subjects were far noisier than the rest. And I think that's a big part of what was going on with the public leaderboard. S25 is a &quot;bad speller&quot;, but it's tough to tell, even looking at the leakage.</span></li>\n</ul>\n<li>Used the overall count of the normal timings per subject as a feature, but in two ways</li>\n<ul>\n<li>Scale it for the GBM, so it doesn't&nbsp;overfit the range, and will scale to worse spellers. Two subjects in the test set were worse than any speller seen in the train set. I broke it up into great, good, average, bad.</li>\n<li>After the model is done, linearly add/subtract to the entire subject. Holdout testing for the train data showed that -0.05 - 0.05 is about all that is needed. I looked at the accuracy implied by the timings and the mean coming out of the GBM and adjusted accordingly to get the means mostly in line. Fortunately this last part was heplful, but unnecesary--the score would have been good for third anyway, but it did improve by about 0.03.</li>\n</ul>\n<li>Used the timing itself for each fifth-session point</li>\n<ul>\n<li>Provide that to the model at each point in the fifth session. I also encoded these in a great/good/average/bad, though the real gain was on the great spellers. If the extra timing occurred, it was a really strong signal of an error. For the rest of the spellers, it was more hit and miss, though still valuable to know. That's why you see the plots from the other thread have extremely condensed ranges for two spellers. They don't make many mistakes, and when they do, we feel very confident the error detection algorithm was correct (75% I think).</li>\n</ul>\n</ul>\n\n<p>I'll keep working to clean up the github code, though it's probably less useful than the general idea. It's been great to see the diversity of ideas already posted on this thread.</p>\n<p>On to some rain prediction!</p>\n<p>Mark, Rob, and David</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 64833,
      "author_name": "mikeskim",
      "author_url": "",
      "post_date": "02/25/2015 00:44:16",
      "content": "<p>My approach was to overfit continuously until I dropped around 50 maybe 60 places. Then I would drink a beer.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64834,
      "author_name": "danielyoo",
      "author_url": "",
      "post_date": "02/25/2015 00:45:46",
      "content": "<p>Congrats to all!</p>\n<p>Coming off the end&nbsp;of the last competition, I didn't have a lot of&nbsp;time for this contest, and with the small test set&nbsp;size it's hard to tell whether what I did was actually good... but here's what I tried.</p>\n<p>1. Band-pass filtered and normalized the&nbsp;EEG signals.</p>\n<p>2. Trained&nbsp;stacked convolutional autoencoders with three hidden layers on all the EEG data. I guessed at some architectures and hyper-parameters that I thought might&nbsp;reasonably capture&nbsp;the EEG signal. There is&nbsp;probably room for optimization the architectures. I tried&nbsp;three different architectures of varying convolution sizes and hidden unit sizes.&nbsp;The auto encoder with the fewest parameters seemed to do best.</p>\n<p>3. I then fed the features&nbsp;at the top layer of the autoencoders into a logistic regression and&nbsp;randomized tree classifiers. My thought was that these algorithms&nbsp;would be better at reducing&nbsp;overfitting.</p>\n<p>4. Finally, I ensembled the different model predictions using a logistic regression.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64838,
      "author_name": "maineiac",
      "author_url": "",
      "post_date": "02/25/2015 01:09:31",
      "content": "<p>Congrats to the winners. I'm interested to hear about your approaches. This competition was a bit frustrating for me, both because the leaderboard was so clearly off, and because I found it difficult to get a grasp on good features.</p>\n<p>My approach:</p>\n<p>1. Standard pre-processing, such as down-sampling by half and baseline correction prior to the feedback.</p>\n<p>2. ICA on a subsample of (mostly) central electrodes, but done on error trials only.</p>\n<p>3. Upsampling the data to deal with class imbalance (in some models).</p>\n<p>4. Adding raw data from a few central electrodes (Cz, PCz, FCz) into the model with the ICA components (some models).</p>\n<p>5. Model averaging</p>\n<p>As far as models, nothing special: random forests, svms with a radial kernel (not so good, but added diversity), gbms. Tried some multi-layer NNs with dropout, but these didn't seem to be learning much.</p>\n\n<p>Things that didn't work for me (but maybe should have based on other efforts I read about):</p>\n<p>1. Features based on spectral power (either raw spectra, or derived features such as mean power in the theta range 4-8 hz)</p>\n<p>2. Amplitude-based features (e.g., peak-to-peak amplitudes around 200-400 and 400-600 ms to capture the Ne and Pe)</p>\n<p>3. Quantifying slow cortical potentials as linear changes in the signal over the trial or mean differences from the first 100 ms</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64847,
      "author_name": "mikeskim",
      "author_url": "",
      "post_date": "02/25/2015 05:51:32",
      "content": "<p>If you like shakeups, this competition ranks up there close to MLSP2014 Schizophrenia and Big Data Combine:</p>\n<p>&gt; shakeup('http://www.kaggle.com/c/inria-bci-challenge/leaderboard')<br>Joining by: id<br>$shakeup.top<br>[1] 0.1610839</p>\n<p>$shakeup.all<br>[1] 0.2885489</p>\n<p>Source: https://www.kaggle.com/c/liberty-mutual-fire-peril/forums/t/10187/quantifying-leaderboard-shake-up/52910</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64856,
      "author_name": "bgeier",
      "author_url": "",
      "post_date": "02/25/2015 11:21:43",
      "content": "<p>I dropped an impressive 87 places so that's cool. My 10-fold AUC was 0.78 and public leaderboard wasn't far off but didn't generalize well.&nbsp;</p>\n<p>My approach was pretty simple.&nbsp;</p>\n<p>1. apply bandpass filter between 1 - 20 Hz</p>\n<p>2. For every feedback event, gather 1 sec trajectory for every channel</p>\n<p>3. Pool trajectories per channel and apply functional principal component analysis. FPCA provides a set of coefficients that summarize functional variation</p>\n<p>4. Apply random forest to data reduction (410 features for all channels (1 sec data)) + feedback order</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64859,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "02/25/2015 11:56:01",
      "content": "<p>Preprocessing:</p>\n<ul>\n<li>Linear de-trend</li>\n<li>Remove freqs &gt; 32 Hz</li>\n<li>Normalization</li>\n<li>Used window of 25-650 ms after event</li>\n</ul>\n<p>Model:</p>\n<ul>\n<li>PLS (usually 4-5 components performed best)</li>\n<li>Brute force testing of channel pairs that gave best CV result</li>\n<li>Continue to add&nbsp;next best channel until performance dropped.</li>\n</ul>\n<p>Best channels were:</p>\n<p>['Fp1', 'CP4', 'AF7', 'AF8', 'Fz', 'CP2']</p>\n<p>Repeat using single-layer perceptron and GBM, and average models.</p>\n<p>Only dropped 2 places from public to private LB, so CV was pretty good I guess.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64860,
      "author_name": "devinanzelmo",
      "author_url": "",
      "post_date": "02/25/2015 13:07:23",
      "content": "<p>I adapted methods from <a href=\"http://eprints.pascal-network.org/archive/00008033/01/BlaLemTreHauMue10.pdf\">single trial ERP tutorial</a>, and paper on <a href=\"http://www.researchgate.net/profile/J_Farquhar/publication/234112682_FarquharHill_ERP_Classification_Best_Practice_Author_Pre-print/links/0fcfd50f46f47278f0000000.pdf\">preprocessing and classification of single trial ERP</a>.</p>\n<p>Preprocessing:</p>\n<ul>\n<li>Used pandas rolling window of size thirty to remove high frequency information</li>\n<li>Baseline corrected with mean of 200 ms before feedback onset for each channel</li>\n<li>Took 200ms to 1000ms after feedback and downsampled to 7 samples per channel</li>\n<li>Concatenated all channels for each trial into single feature</li>\n<li>Normalized with sklearn standard scaler</li>\n</ul>\n<p>Classification:</p>\n<ul>\n<li>Used Linear Discriminant analysis with shrinkage from sklearn 0.16 dev</li>\n<li>Additionally used Logistic Regression with L1 or L2 regularization</li>\n<li>averaged 6 classifiers with different parameters</li>\n</ul>\n<p>Private leaderboard score without label counts from leak was ~0.710. Adding the label counts from the leak increased my score by about .13 on private leaderboard, though it only increased my CV score by 0.05. I additionally added label counts from my own classifier and this added an additional 0.01 to my score.</p>\n<p>I several other things which I couldn't get to work correctly.</p>\n<ul>\n<li>Latency correction for ERP onset between subjects</li>\n<li>Standard bandpass gave higher score for two subjects but couldn't tune for others</li>\n<li>Channel selection. I overfit doing recursive channel elimination. It turned out that the best channels were quite different depending on which two subjects I left out.</li>\n<li>Clustering of subjects and training on similar subjects</li>\n<li>many other things I have forgotten</li>\n</ul>\n<p>I am interested to hear if any one used transfer learning techniques and/or explicitly attempted to calibrate inter-subject probabilities for ROC AUC.</p>\n<p>I am satisfied with my results but I feel like I missed some easy things to improve score. Interested to see what overfitting avengers did. I read a couple of Alaxandre Barachant's papers on MDM classifier using Riemann Geometry. Wanted to try to implement but seemed slightly intimidating. Is this what you guys used?</p>\n<p>Edit: links to the first paper is broken. the first paper is called &quot;<em>Single-trial analysis and classification of ERP components&#8212;a tutorial</em>&quot; by Blankertz et al,</p>\n<p>the second paper is called: &quot;<em>Interactions between Pre-Processing and Classification </em><em>Methods for Event-Related-Potential Classification: </em><em>Best-Practice Guidelines for Brain-Computer Interfacing</em>&quot; by Farquhar and Hill.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64862,
      "author_name": "kazanova",
      "author_url": "",
      "post_date": "02/25/2015 14:18:32",
      "content": "<p>[quote=Mike Kim;64833]</p>\n<p>My approach was to overfit continuously until I dropped around 50 maybe 60 places. Then I would drink a beer.</p>\n<p>[/quote]</p>\n<p>Did you have the beer?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64864,
      "author_name": "joseleiva",
      "author_url": "",
      "post_date": "02/25/2015 14:25:06",
      "content": "<p>I'm really impressed at the final performance of the winning team. I'm also interested in their approach, because Alexandre already won the competition on MEG with a large margin.</p>\n<p>Phalaris, congratulation for your second place. I'm also impressed at your good result in spite of its simplicity (no Random Forest, no Convolutional NNs, etc). I think that in this kind of applications, a sensible preprocessing as the one you performed is more important than using fancy classification techniques (one question remains, though: did you use a <em>pool</em> approach, i.e. use data from all subjects to train a common classifier?)</p>\n<p>Also, keeping things simple is important to avoid overfitting. The fact that the public LB showed the performance on 2 of the 10 subjects suggested that there would be surprises on the private LB. In my case, I jumped from the 63th to the 5th place!</p>\n<p>---------------------------------------------------------</p>\n<p>My approach:</p>\n<p>Preprocessing:</p>\n<p>- Band-Pass filtering between 1 and 10 Hz with a FIR filter with order 512.&nbsp; Starting from the 30th sample, we keep 25 samples applying a down-sample factor of 4.</p>\n<p>Features (or 1st level classification):</p>\n<p>-For each of the 25 time-lags and for each of the 16 subjects,&nbsp; obtain two spatial filters using LDA and Log. Regression, respectively. Use 240 of the 340 examples to do this. The EOG channel is included here with the aim of reduce the contribution of artifacts.</p>\n<p>(2nd level) Classification</p>\n<p>- Using different data than those used for extracting the filters (the remaining 100 examples), we train an SVM with RBF kernel.</p>\n\n<p>The approach described is based on the Stacked Generalization procedure: in the first level we solve a set of per-subject classification problems, and in the second one we combine features obtained with filters trained on the 16 subjects; this way, we expect to capture some diversity.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64868,
      "author_name": "robzink",
      "author_url": "",
      "post_date": "02/25/2015 14:59:25",
      "content": "<p>Nice to have some insight into the methods used by you all!</p>\n<p>Our&nbsp;approach was&nbsp;similar to that of Phalaris, although I did not use the info of the 'leakage'.</p>\n<p>On the Public board we got stuck around 0.71% corresponding to place 154.&nbsp;However, I was somewhat surprised to see the Private place to be 25, a jump up of 129 places.</p>\n<p>Features:</p>\n<ul>\n<li><span style=\"line-height: 1.4\">Bandpass filtered the data 1-20Hz.</span></li>\n<li>Feature reduction by averaging time-windows (different window lengths) per trial in 0-1000ms interval. (similar to downsample).</li>\n<li>Concatenated all channels and did a z-score normalization.</li>\n</ul>\n<p>Classification:</p>\n<ul>\n<li><span style=\"line-height: 1.4\">Linear Discriminant Analysis with shrinkage 0.10.</span></li>\n<li><span style=\"line-height: 1.4\">Combined with&nbsp;Logistic Regression</span></li>\n<li>Averaged the&nbsp;classifiers&nbsp;for different features/settings</li>\n</ul>\n<p>Explored but not beneficial for us:</p>\n<ul>\n<li><span style=\"line-height: 1.4\">Step-wise feature selection, in which features are added/removed (by setting a threshold in p-value) to form a robust feature-subspace model. (e.g. as applied in swLDA).</span></li>\n<li><span style=\"line-height: 1.4\">Explored Least Squares Support Vector Machines and the GBM - did not improve the overall.</span></li>\n</ul>\n<p>In the end we scored 0,69 - however, without the leakage information. Our result was quite stable and not prone to overfitting the public LB, as evident from the large increase in the final ranking.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64873,
      "author_name": "triskelion",
      "author_url": "",
      "post_date": "02/25/2015 16:14:03",
      "content": "<p>I did a model with only a few features. These features probably are not fully available at run-time and I myself consider them a bit &quot;leaky&quot;.</p>\n<p>A) Length after feedback event, before the next feedback event.</p>\n<p>B) Length after last feedback event, till current feedback event</p>\n<p>C) A - B</p>\n<p>D) Feedback event counter</p>\n<p>E) Session counter</p>\n<p>F) Average variance over all channels / average energy</p>\n<p>This scored 0.64 public LB and 0.56 private LB. Leave two-subjects out CV was 0.61. This model got ensembled with the other models made by team members. Luckily Phil&nbsp;convinced me that CV was far more important than LB in this competition, so we did not end up with only 0.76 public LB&nbsp;-&nbsp;0.47 private LB models.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64877,
      "author_name": "dheger",
      "author_url": "",
      "post_date": "02/25/2015 16:40:44",
      "content": "<p>We applied our new BCI framework for joint optimization [1] that will be presented at NER15. It is based on the Alternating Direction Method of Multipliers [2] and jointly optimizes 3 core objectives (discriminativity, compactness, and robustness) using an objective function that consists of</p>\n<p>(1) least-squares regression term f(x) to model discriminative brain activity patterns,</p>\n<p>(2) sparsity inducing L1 penalty regularization term g(x) to generate compact models and</p>\n<p>(3) sum-of-squares regularization term h(x) for unsupervised transfer learning.</p>\n<p>To remove inter-subject influences, differences between the mean feature vectors of each of the 16 training subjects and each of the 10 testing subjects (160 vectors) were penalized by the sum-of-squares regularization h(x).</p>\n<p>Features: Combination of time-domain amplitudes (0-800ms), power spectral density features (1 Hz wide bins, Welch's method), and meta-data features (session numbers and delay to previous trial) into a 1570 dimensional feature vector.<br>We intentionally did not include the data leakage information in our system.</p>\n<p>In a post-processing step we adjusted the predictions for sessions 1-4 of each subject to integrate prior information about the two different types of copy spelling (fast and slow condition) by adding a constant to long delayed trials.</p>\n<p>Regularization parameters were optimized using a 4-fold cross-validation with splitting on subject bounds, which represented the results on the private LB well.&nbsp;</p>\n<p>Our system scored 0.81124 on the public LB and 0.7457 on the private LB, which is 6th place in the final ranking.</p>\n<p><br>[1] Dominic Heger, Christian Herff, Felix Putze, Tanja Schultz, &#8220;Joint Optimization for Discriminative, Compact and Robust Brain-Computer Interfacing,&#8221; in 7th International IEEE EMBS Neural Engineering Conference 2015 (accepted).<br>[2] Stephen Boyd, et al. &quot;Distributed optimization and statistical learning via the alternating direction method of multipliers.&quot; Foundations and Trends in Machine Learning 3.1 (2011): 1-122.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64882,
      "author_name": "devinanzelmo",
      "author_url": "",
      "post_date": "02/25/2015 17:41:56",
      "content": "<p>@Jose M. I used single classifier for all subjects. Glad to see that using single subject classifiers worked so well. Your score is 0.03 better then what I got without information from leak.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64883,
      "author_name": "kazanova",
      "author_url": "",
      "post_date": "02/25/2015 18:21:02",
      "content": "<p>[quote=Triskelion;64873]</p>\n<p>&nbsp;Luckily Marios convinced me&nbsp;</p>\n<p>[/quote]</p>\n<p>My true name is revealed! I 've been hiding it for the last 15 years!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64901,
      "author_name": "thekannman",
      "author_url": "",
      "post_date": "02/26/2015 01:12:45",
      "content": "<p>Good job, everyone! <br><br>I joined this competition late, so much of my technique was influenced by time constraints.<br><br>My approach was to use figures 7 and 9 from&nbsp;http://www.hindawi.com/journals/ahci/2012/578295/ to determine a handful of channels that were likely to be most relevant. I calculated a running sum over each of those channels, and used both the raw data and the running sums as features. <br><br>I used PCA to decrease the number of features even further to about 100. I then did Linear regression with recursive feature elimination to get my final result. I tried some other methods (Random forest, Lasso, ElasticNet, etc.), but I found that the minor improvements were not worth the additional time requirements. I'm going to clean up my code and upload it to github in case anyone is interested.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64904,
      "author_name": "bcragin",
      "author_url": "",
      "post_date": "02/26/2015 02:35:36",
      "content": "<p>Is it just me, or have the Private LB scores of the individual models not yet been placed in the &quot;My Submissions&quot; tab? Without that info, how does anyone know &nbsp;which of their two selections gave the better result??&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64905,
      "author_name": "mlandry",
      "author_url": "",
      "post_date": "02/26/2015 02:37:36",
      "content": "<p>Bruce, you can send in submissions and they will show you how they would have scored.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64906,
      "author_name": "bcragin",
      "author_url": "",
      "post_date": "02/26/2015 02:38:59",
      "content": "<p>I see. Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64925,
      "author_name": "duncanbarrack",
      "author_url": "",
      "post_date": "02/26/2015 10:20:14",
      "content": "<p>It's really interesting to learn of the other approaches taken.&nbsp;&nbsp;My approach was as follows</p>\n<p>Preprocessing:</p>\n<ul>\n<li><span style=\"line-height: 1.4\">Bandpass filter all channels to between 1 and 20Hz using the Butterworth filter.</span></li>\n</ul>\n<p>Feature extraction:</p>\n<ul>\n<li>Meta features (timestamp, session number etc.)</li>\n<li><span style=\"line-height: 1.4\">Mean of EEG values for each channel in windows of various lengths and lags after the feedback event. These were based on <a href=\"http://videolectures.net/bbci09_blankertz_muller_mlasp/\">this video lecture</a>.</span></li>\n<li>Template matching features. Here I averaged all ERP signals labelled as positive feedback in the training set to create a 'positive feedback template' and similarly to create a&nbsp;'negative feedback template'. This was done separately for each channel. I then used&nbsp;correlation, cross-correlation, covariance, cross-covariance and euclidean distance between the templates and test ERP signals as features. These features were inspired by <a href=\"http://www.sciencedirect.com/science/article/pii/0168559794900515\">this paper</a>.</li>\n</ul>\n<p>Model:</p>\n<p>I took a weighted average of a regularised SVM with a linear kernel (meta and mean EEG features) and a&nbsp;regularised SVM with a linear kernel (meta and template matching&nbsp;features). Putting all the features into one model didn't seem to improve predictions over just using either the mean EEG or template matching features separately. I'm still&nbsp;haven't got a handle on why this would be the case.</p>\n<p>Cross validation:</p>\n<p>I went for 4 fold CV cut on a per subject basis and calculated the AUC on the 4 subjects in the test fold only. To me, this seemed like the best way to 'mimic' the scoring used for the private leader-board. I'd be interested to learn what CV approach others took and if there was a more principled way which I missed. My best model had a CV score of 0.75, a private leader-board score of 0.76921 and public leader-board score of ~0.77, which I guess are all&nbsp;in the same ballpark.&nbsp;</p>\n<p>Things that didn't work or I didn't use</p>\n<ul>\n<li>Logistic regression with elastic net regularisation (gave slightly worse results than the linear SVMs)</li>\n<li>GBT, SVMs with RBF kernel (gave very poor CV scores)</li>\n<li>Spectral features</li>\n<li>Features based on peak picking</li>\n<li>Dynamic time warping distance between templates and ERP signal</li>\n<li>Clustering test subjects with similar training subjects. My intuition was that this would&nbsp;work well but I notice Phalaris also tried it and couldn't get it to work, so it wasn't just me!</li>\n</ul>\n<p>@Maineiac, I noticed you used ICA. How much did it help?&nbsp;It seems to get a good&nbsp;&nbsp;rep in the literature but&nbsp;in the end I didn't have the energy&nbsp;to implement&nbsp;it!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64947,
      "author_name": "ashhafez",
      "author_url": "",
      "post_date": "02/26/2015 14:07:19",
      "content": "<p>In the interests of sharing, here's the approach I took:</p>\n<p>- For each word spelled, determine whether it was a fast or slow speed test, and use this as a constant feature value for the feedback events in that word (determined&nbsp;by comparing each word time, proxied through delta between first and last feedback events, to the min and max word times for that session, and accounting for the extra delay after each word is completed).&nbsp;</p>\n<p>-&nbsp;Look into the &quot;leak&quot; information that phalaris had mentioned, and create a time_ratio feature for each subject, based on the timings in their session 5 (skipping the first feedback event in each session, since there was some noise in the start timings), so essentially a gauge of each feedback event's post-feedback time to the min/max range of the feedback times observed for that session (the closer to the min time, the less error correction was done). This to me seemed more like a proxy for when the existing algorithm thought there was an error, not necessarily when there actually was an error - so using it as a feature was like boosting the existing algo.</p>\n<p>- I then labelled each user with this time ratio, and ran an OLS between the user's time ratios and their % accuracy overall, on fast, and on slow runs, and obtained a beta that could then be used to approximate the % accuracies of the test users based on their observed test_ratio values. Here's a snapshot of the actual vs estimated % accuracies. For the training set, the correlations between the time_ratios and the actual accuracies were 0.695 for fast, 0.643 for slow, 0.694 for combined.</p>\n<p>Training set</p>\n<p>Fit[02]: all=[0.65 0.77] fast=[0.60 0.71] slow=[0.74 0.88]<br>Fit[06]: all=[0.93 0.92] fast=[0.91 0.85] slow=[0.97 1.05]<br>Fit[07]: all=[0.90 0.88] fast=[0.85 0.81] slow=[0.99 1.00]<br>Fit[11]: all=[0.66 0.59] fast=[0.60 0.55] slow=[0.78 0.68]<br>Fit[12]: all=[0.56 0.62] fast=[0.46 0.57] slow=[0.73 0.71]<br>Fit[13]: all=[0.52 0.62] fast=[0.46 0.57] slow=[0.64 0.70]<br>Fit[14]: all=[0.63 0.78] fast=[0.54 0.72] slow=[0.80 0.89]<br>Fit[16]: all=[0.62 0.61] fast=[0.54 0.56] slow=[0.77 0.70]<br>Fit[17]: all=[0.66 0.60] fast=[0.63 0.56] slow=[0.73 0.69]<br>Fit[18]: all=[0.77 0.75] fast=[0.73 0.70] slow=[0.83 0.86]<br>Fit[20]: all=[0.69 0.66] fast=[0.61 0.61] slow=[0.85 0.75]<br>Fit[21]: all=[0.92 0.67] fast=[0.89 0.62] slow=[0.97 0.77]<br>Fit[22]: all=[0.92 0.82] fast=[0.90 0.76] slow=[0.97 0.94]<br>Fit[23]: all=[0.65 0.65] fast=[0.59 0.60] slow=[0.76 0.74]<br>Fit[24]: all=[0.70 0.68] fast=[0.61 0.63] slow=[0.87 0.78]<br>Fit[26]: all=[0.53 0.68] fast=[0.49 0.63] slow=[0.61 0.78]</p>\n<p>Test set</p>\n<p>Fit[01]: all=[0.00 0.80] fast=[0.00 0.74] slow=[0.00 0.92]<br>Fit[03]: all=[0.00 0.66] fast=[0.00 0.61] slow=[0.00 0.75]<br>Fit[04]: all=[0.00 0.87] fast=[0.00 0.81] slow=[0.00 1.00]<br>Fit[05]: all=[0.00 0.39] fast=[0.00 0.36] slow=[0.00 0.45]<br>Fit[08]: all=[0.00 0.67] fast=[0.00 0.62] slow=[0.00 0.77]<br>Fit[09]: all=[0.00 0.77] fast=[0.00 0.71] slow=[0.00 0.88]<br>Fit[10]: all=[0.00 0.84] fast=[0.00 0.78] slow=[0.00 0.96]<br>Fit[15]: all=[0.00 0.86] fast=[0.00 0.79] slow=[0.00 0.98]<br>Fit[19]: all=[0.00 0.74] fast=[0.00 0.69] slow=[0.00 0.85]<br>Fit[25]: all=[0.00 0.58] fast=[0.00 0.54] slow=[0.00 0.67]</p>\n<p><br>These became additional &quot;fitness&quot; features in the model, bucketed into deciles.</p>\n<p>- All channel values were normalized into z-scores, based on the 0-260 timeslots post feedback event, and baselined to zero at the first feedback event. These became features, along with the mean value for that subject / channel / timeslot, to gauge how different the current feedback run was compared to their usual, between the 50 and 150 timeslot bucket</p>\n<p>- Had a couple of features based on the position of the min and max zscore values for each channel, relative to the 100 timeslot&nbsp;</p>\n<p>- session id flag features</p>\n<p>- Also had channel correlation features, and volatility of correlation features pre and post feedback event</p>\n<p>- A bunch of other ideas that didn't pan out, including general subject features based on typical volatility of response variation etc</p>\n<p>In terms of model, I essentially used a logistic regression via a no hidden layer NN, and experimented with input dropout, various regularization parameter values, de-noising, different cost functions.</p>\n<p>The CV setup was as follows:</p>\n<p>-&nbsp;Randomly shuffle the set of all available feedback events&nbsp;</p>\n<p>- pick 15% of the shuffled events for an ensemble validation set</p>\n<p>- With the remaining set of events, split into K-fold CV by subject, picking 3 of each type of subject (i split them into good | medium | poor rating subjects) for training, and the remaining for CV</p>\n<p>- Average the resulting model from each fold progressively together, and validate its result improves the ensemble validation set</p>\n<p>- Run the above with multiple seed values and average the resulting models</p>\n<p>I was obviously disappointed with the reshuffle from 5th to 66th but turns out I managed to overfit somewhere along the way, since my ensemble CV result was pretty well&nbsp;tracking the public LB score and improvement. To try to improve and figure out what happened, I spent a couple of hours post competition backtracking features to figure out where the overfit took place, and noticed that even with just the fitness and session features, plus a couple of channels, I could get a&nbsp;0.79471 4th place private leaderboard submission. but tweaking the L1 up or down a decile ranged the rankings between mid twenties and mid sixties.</p>\n<p>tough dataset!&nbsp;</p>\n<p>appreciate any comments or suggestions for improvement to the CV process from those who had a more robust&nbsp;process.</p>\n<p>- one additional observation is that the _min_ time between feedback events seemed to fairly linearly increase with each subject, which made me wonder if there was some form of inherent bias in the test machine, since later subjects had more time in general than earlier subjects</p>\n<p>cheers</p>\n<p>ash</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65036,
      "author_name": "asymptote",
      "author_url": "",
      "post_date": "02/27/2015 15:15:53",
      "content": "<p>Will you guys post/share&nbsp;the code?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65040,
      "author_name": "thekannman",
      "author_url": "",
      "post_date": "02/27/2015 16:03:10",
      "content": "<p>I will post mine as soon as I clean it up. It ranked 57th on the private leaderboard.</p>\n<p>[quote=Asymptote;65036]</p>\n<p>Will you guys post/share&nbsp;the code?</p>\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65064,
      "author_name": "mlandry",
      "author_url": "",
      "post_date": "02/27/2015 19:43:21",
      "content": "<p>Here is some info about the 3rd place team.</p>\n<p>As expressed in <a href=\"https://www.kaggle.com/c/inria-bci-challenge/forums/t/12603/inter-subject-auc/64843#post64843\">another&nbsp;thread</a>, success of our solution is more to do with <a href=\"https://www.kaggle.com/wiki/Leakage\">leakage</a>&nbsp;mining than signal processing. My signal processing efforts alone were top 15% or so, though there were some good ideas I didn't get to because&nbsp;I focused on the leakage.</p>\n<p>I attempted to clean mine up a bit, and most of it is on my <a href=\"https://github.com/mlandry22/kaggle/tree/master/bci_challenge\">github repo</a>--enough that running that code will get third. It's very ugly as a result of (1) spending the last few days hammering out the leakage; and (2) I write ugly code.</p>\n<p><strong>Non-leakage efforts:</strong></p>\n<ul>\n<li>Min/max scaled every series of 0.0 - 1.0 seconds for Cz, used 15-period moving average for the basis of most hand-engineered features.</li>\n<li>Sent the above 0.0-1.0 range, used those values as is, but trimmed the timepoints used in modeling&nbsp;based on GBM importance values.</li>\n<ul>\n<li>The range around 70-95 was consistently favored by most models, including the unscaled python forum code.</li>\n</ul>\n<li>Based on the full 0.0-1.0 range from above calculated several features that did well in basic AUC testing</li>\n<ul>\n<li>sum(0.0 - 0.5) - sum(0.5 - 1.0). So basically subtract the sum of the scaled signal in the first half-second minus the second half-second.</li>\n<li>the location of the minimum point between 20 and 70 readings after feedback. I.e. the time point at which a local minimum was reached.</li>\n<li>several other simple calculations: raw range of sample, range of first to last point, standard deviation, volatility. These didn't have a lot of influence.</li>\n</ul>\n<li>Time frequency kernel to adjust for EOG</li>\n<ul>\n<li>Teammate had some nice prior experience with EMGs&nbsp;in matlab, so we used octave here</li>\n<li>Used a time frequency kernel to generate about 500 vectors per time point.</li>\n<li>TF kernel vectors were used in PLS regression with the EOG signal as the target</li>\n<li>Top 10 pls coefficients were&nbsp;used as features</li>\n<li>This had a small but reliable impact. The model consistently favored the 4th pls coefficient, which is a little peculiar.</li>\n</ul>\n<li>Other EOG reduction methods</li>\n<ul>\n<li>Scaled subtraction. Min/max scale the EOG signal as with the Cz signal, then scale the max of the EOG result to wherever it experienced the largest gap between the Cz signal.</li>\n<li>Straight differencing:&nbsp;Cz - EOG.</li>\n<li>Each of these two provided some small gains.</li>\n</ul>\n</ul>\n<p><strong>Leakage efforts:</strong></p>\n<ul>\n<li>Applied <a href=\"https://www.kaggle.com/c/inria-bci-challenge/forums/t/11178/possible-data-leakage\">phalaris's discovery</a>&nbsp;of the importance of timing in the fifth session. Note that this has been available&nbsp;for about a month.</li>\n<li>Identify the extended time by:</li>\n<ul>\n<li>accounting for the extra time at the end of a word</li>\n<li><span style=\"line-height: 1.4\">accounting for the &quot;drift&quot; in the cutline of what expected time and extra time looked like. I.e. the first few subjects had a cutpoint that won't work for later subjects, so you have to take that into account.</span></li>\n<li><span style=\"line-height: 1.4\">guessing at the cutpoint for S25 and S26. These subjects were far noisier than the rest. And I think that's a big part of what was going on with the public leaderboard. S25 is a &quot;bad speller&quot;, but it's tough to tell, even looking at the leakage.</span></li>\n</ul>\n<li>Used the overall count of the normal timings per subject as a feature, but in two ways</li>\n<ul>\n<li>Scale it for the GBM, so it doesn't&nbsp;overfit the range, and will scale to worse spellers. Two subjects in the test set were worse than any speller seen in the train set. I broke it up into great, good, average, bad.</li>\n<li>After the model is done, linearly add/subtract to the entire subject. Holdout testing for the train data showed that -0.05 - 0.05 is about all that is needed. I looked at the accuracy implied by the timings and the mean coming out of the GBM and adjusted accordingly to get the means mostly in line. Fortunately this last part was heplful, but unnecesary--the score would have been good for third anyway, but it did improve by about 0.03.</li>\n</ul>\n<li>Used the timing itself for each fifth-session point</li>\n<ul>\n<li>Provide that to the model at each point in the fifth session. I also encoded these in a great/good/average/bad, though the real gain was on the great spellers. If the extra timing occurred, it was a really strong signal of an error. For the rest of the spellers, it was more hit and miss, though still valuable to know. That's why you see the plots from the other thread have extremely condensed ranges for two spellers. They don't make many mistakes, and when they do, we feel very confident the error detection algorithm was correct (75% I think).</li>\n</ul>\n</ul>\n\n<p>I'll keep working to clean up the github code, though it's probably less useful than the general idea. It's been great to see the diversity of ideas already posted on this thread.</p>\n<p>On to some rain prediction!</p>\n<p>Mark, Rob, and David</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "64830": "",
    "64833": "",
    "64834": "",
    "64838": "",
    "64847": "",
    "64856": "",
    "64859": "",
    "64860": "",
    "64862": "",
    "64864": "",
    "64868": "",
    "64873": "",
    "64877": "",
    "64882": "",
    "64883": "",
    "64901": "",
    "64904": "",
    "64905": "",
    "64906": "",
    "64925": "",
    "64947": "",
    "65036": "",
    "65040": "",
    "65064": ""
  },
  "source": "meta"
}