{
  "id": 12819,
  "title": " The Overfitting Avengers Solution & Code ",
  "url": "/competitions/inria-bci-challenge/writeups/the-overfitting-avengers-the-overfitting-avengers-",
  "author_name": "",
  "post_date": "2015-03-15T17:39:58.370Z",
  "votes": 14,
  "comment_count": 17,
  "views": 7315,
  "content": "<p>Hi,</p>\n<p>You can find our code and documentation on&nbsp;<a href=\"https://github.com/alexandrebarachant/bci-challenge-ner-2015\">github</a>.</p>\n<p>In short, we used a special form covariance matrices as feature and tools from Riemannian Geometry to manipulate them. A channel selection algorithm was involved, and the classification was achieved by an ElasticNet classifier.&nbsp;</p>\n<p>The above pipeline was&nbsp;applied on a number of random subsets of subjects, and predictions are averaged across bagged models.</p>\n<p>The complete&nbsp;explanation of our model is provided <a href=\"https://github.com/alexandrebarachant/bci-challenge-ner-2015/blob/master/README.md\">here</a>.</p>\n\n<p>We made two different submission. The first one does not make any use of the leakage information and satisfies an &quot;online processing&quot; constraint, which means that any trial performed by a subject can be classified without the need for future complementary data or information. The second model uses the leak, thus it is not online-compatible. The two models are built upon the same classification pipeline, but with parameters tuned independently to achieve the highest performance.</p>\n<ul>\n<li><span style=\"line-height: 1.4\">The online model (no leak) got&nbsp;0.85083 on Public LB and 0.84585 on Private LB</span></li>\n<li><span style=\"line-height: 1.4\">The leak model got&nbsp;</span><span style=\"line-height: 1.4\">0.85145 on Public LB and 0.87224 on Private LB.</span></li>\n</ul>\n\n<p><span style=\"line-height: 1.4\">Any comments are welcomes !</span></p>\n\n<p><span style=\"line-height: 1.4\">Cheers,</span></p>\n<p><span style=\"line-height: 1.4\">The overfitting Avengers :)</span></p>",
  "messages": [
    {
      "id": "66251",
      "postDate": "03/15/2015 17:39:58",
      "content": "<p>Hi,</p>\n<p>You can find our code and documentation on&nbsp;<a href=\"https://github.com/alexandrebarachant/bci-challenge-ner-2015\">github</a>.</p>\n<p>In short, we used a special form covariance matrices as feature and tools from Riemannian Geometry to manipulate them. A channel selection algorithm was involved, and the classification was achieved by an ElasticNet classifier.&nbsp;</p>\n<p>The above pipeline was&nbsp;applied on a number of random subsets of subjects, and predictions are averaged across bagged models.</p>\n<p>The complete&nbsp;explanation of our model is provided <a href=\"https://github.com/alexandrebarachant/bci-challenge-ner-2015/blob/master/README.md\">here</a>.</p>\n\n<p>We made two different submission. The first one does not make any use of the leakage information and satisfies an &quot;online processing&quot; constraint, which means that any trial performed by a subject can be classified without the need for future complementary data or information. The second model uses the leak, thus it is not online-compatible. The two models are built upon the same classification pipeline, but with parameters tuned independently to achieve the highest performance.</p>\n<ul>\n<li><span style=\"line-height: 1.4\">The online model (no leak) got&nbsp;0.85083 on Public LB and 0.84585 on Private LB</span></li>\n<li><span style=\"line-height: 1.4\">The leak model got&nbsp;</span><span style=\"line-height: 1.4\">0.85145 on Public LB and 0.87224 on Private LB.</span></li>\n</ul>\n\n<p><span style=\"line-height: 1.4\">Any comments are welcomes !</span></p>\n\n<p><span style=\"line-height: 1.4\">Cheers,</span></p>\n<p><span style=\"line-height: 1.4\">The overfitting Avengers :)</span></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "66291",
      "postDate": "03/15/2015 21:55:09",
      "content": "<p>Thanks for posting your teams solutions! I have been looking forward seeing your methodology since the competition ended. The python code looks quite usable. I was considering converting your matlab code from decMeg to experiment with but now I don't have to. I will post some questions when I have time to closely read the solution.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "66295",
      "postDate": "03/15/2015 22:48:57",
      "content": "<p>Great job! Are you finding Riemannian Geometry to be an indispensable tool for signal processing and BCI-related work?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "66296",
      "postDate": "03/15/2015 22:49:37",
      "content": "<p>You're welcome :)</p>\n<p>For me, it was a good opportunity to switch to python, and i have to say that i will never go back to matlab.</p>\n<p>Now that i have an equivalent of my <a href=\"https://github.com/alexandrebarachant/covariancetoolbox\">covariance toolbox</a>&nbsp;in python, I plan to release a package as soon as possible.</p>\n<p>We took our time before publishing the code, but we wanted to make a deeper analysis of our solution. The final solution is a bit computationally intensive, but you can speed of the process by turning off the bagging and the electrode selection.</p>\n<p>It turns out that by doing this you can achieve a score ~0.843 (for the model without the leak) and it will takes around 2 minutes to process.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "66301",
      "postDate": "03/15/2015 23:43:06",
      "content": "<p>[quote=TF;66295]</p>\n<p>Great job! Are you finding Riemannian Geometry to be an indispensable tool for signal processing and BCI-related work?</p>\n<p>[/quote]</p>\n<p>I would say definitely yes but it will be a bit unfair for all the other people who have done amazing work in the BCI community.</p>\n<p>At first, i developed these methods to get rid of the source separation step (i.e. spatial filtering) during online BCI experiments. While algorithm such Common Spatial Patterns (CSP) or XDAWN are really powerful, they lack of robustness and introduce extra parameters to tune. And if they are not used correctly, they can lead to a loss of spatial information.</p>\n<p>By using covariance matrices as features, you make sure you get all the spatial information in your classifier. Then, tools from Riemannian geometry provide you a convenient and robust way to manipulate them.&nbsp;More specifically, the tangent space mapping (which can be see as a kernel operation) is a very powerful tool. It allows to transform these matrices in euclidean vectors without breaking their particular structure &nbsp;(Symmetric and positive definite). Then, you can do whatever you want with these vectors.</p>\n<p>Now, i have to say that you can probably get similar results with other methods, but thanks to a bunch of properties of the Riemannian metric, it is very well suited for cross-subject classification (which was the case for DecMeg and this challenge).</p>\n<p>Covariance matrices are involved in wide range of&nbsp;signal processing problem (well, let say every time you have a multivariate problem), so&nbsp;riemannian geometry has a lot of potential. It is also used in <a href=\"http://ieeexplore.ieee.org/xpl/login.jsp?tp=&arnumber=6514112&url=http:%2F%2Fieeexplore.ieee.org%2Fxpls%2Fabs_all.jsp%3Farnumber%3D6514112\">Radar signal&nbsp;processing</a>&nbsp;and&nbsp;<a href=\"http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.336.4231&rep=rep1&type=pdf\">Image classification</a>.</p>\n<p>Finally, I wanted to say that the riemannian metric i use is a special case of something bigger called Information Geometry, a theory that study the geometric properties of probability distributions. In my case, the distribution of interest is the multivariate normal distribution with zero mean (only parametrized by it covariance matrix), but such metric can be extracted for any other distribution (but the math behind can be hardcore).</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "66327",
      "postDate": "03/16/2015 07:06:43",
      "content": "<p>Alexandre,&nbsp;</p>\n<p>many thanks for sharing your approach and your code!</p>\n<p>I have a few questions</p>\n<p>1) In DecMeg competition you used CSP filters in order to reduce dimensionality of the problem (DR seems to be an important preprocessing step for Riemannian geometry tools). Here you used&nbsp;a channel selection procedure instead of CSP. Could you please explain a little more about the motivation of this decision?</p>\n<p>2) Why did you obtain such a huge difference between&nbsp;Fold-wise CV (0.7294) and Private/Public LB (0.84585/0.85083)? As far as I understand you used the same criterion (joint AUC) for Fold-Wise CV calculation. The only difference I see is that you used 150 bagged models for CV and 500 bagged models for the final submission. But according to your post-competition analysis even 10 models give ~0.84 on the LB datasets.</p>\n<p>3) Could you please recommend useful books&nbsp;on Information Geometry with all necessary&nbsp;math concepts (even hardcore ones)?&nbsp;</p>\n<p>Thanks in advance,&nbsp;</p>\n<p>Mike</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "66469",
      "postDate": "03/16/2015 19:52:09",
      "content": "<p>Hi Mike,</p>\n<p>these are very relevant questions.</p>\n<p>1) There is multiple reason we took this decision, but first lets got back to the&nbsp;dimensionality reduction&nbsp;for Riemannian geometry (RG). Dealing with High dimensional data is the main problem when using RG. The main reasons are :</p>\n<ul>\n<ul>\n<li>Your matrices must be definite positive and this is not always the case when the number of channel is high. You can still regularize, but it can degrade perfs by affecting the structure of the matrix.</li>\n<li>Tangent space mapping is a local approximation of the manifold. it suppose that all your matrices are scattered in a relatively small part of the manifold. when you increase the dimension, you decrease the density of your data (or increase the sparsity), and it result in bad tangent space approximation.</li>\n<li>Everything become more and more computationally expensive. RG involve a lot of eigenvalue/eigenvector decomposition.</li>\n</ul>\n</ul>\n<p>My experience&nbsp;with EEG/MEG data is that if you keep everything bellow 64 channel, you will have no trouble.</p>\n<p>For the Decmeg Challenge, the dimensionality was simply to high to use any of my methods (including the channel selection) out of the box. So i had no choice and i decided to use spatial filtering.</p>\n<p>The problem with spatial filtering is that is is not very well suited for a cross-subject design. The basic hypothesis of&nbsp;source separation&nbsp;is that the training data and the test data&nbsp;are mixed by the same operator&nbsp;and are stationary. When you train on a subject, and apply on another, you take the risk to discard relevant information in the process. Again, when&nbsp;you aggregates data from different subject for training, you source separation becomes less effective because your algorithm try to estimate a unique unmixing matrix.&nbsp;</p>\n<p>For decMeg, i found a workaround. My classification was in two stage, the first was a cross-subject classifier, using stacked generalization (so spatial filters were not train on aggregated data), the second was an iterative retraining using the data on the test subject only (the case where spatial filtering is really efficient). As you see, the key is to know what are the advantages and the flaws of you methods in order to use them in the best possible conditions.</p>\n<p>Now for this challenge, the dimensionality was manageable, so spatial filtering was not necessary. Still, it is a good thing to reduce dimensionality, so we apply the channel selection procedure to discard irrelevant channels. The number of channels&nbsp;we kept&nbsp;is not very tight (35). it makes everything running more smoothly without having much risk to discard relevant information. EEG has a low spatial resolution, so close channels are highly correlated and often carry the same information. Discarding a few of them doesn't hurt</p>\n<p>So the next question is why we didn't apply the same retraining process for this challenge. Well, first because we wanted to keep thing as 'online' as possible. The second is simply because the convergence were more tricky. The method i used in Decmeg was based on the assumption&nbsp;that the classes were&nbsp;balanced (which was the case by design), in order to set the optimal threshold for the&nbsp;re-labeling between iteration. Here the classes balance were unknown (if you don't use the leak).</p>\n\n<p>2) Okay, now for the second question, i will let my teammate answer.</p>\n<p>3) I think the best book you can find is &quot;Methods of Information Geometry&quot; from S. Amari. The basic idea is that each probability distribution is a point of a manifold with the parameters of the distribution as coordinates. For example a normal distribution of mean u and variance s is a point of coordinate (u,s) in the manifold. The trick is, the Information geometry make the assumption that the natural metric for this manifold is the Fisher information. After a bunch of equation manipulation, it allow you to define a true distance between two distribution, and therefore unlock a lot of issue&nbsp;(for example, the mean of a distribution is not the distribution with the mean parameters :) &nbsp;)</p>\n\n<p>Alex</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "66509",
      "postDate": "03/16/2015 21:21:25",
      "content": "<p>Hey Mike, let me answer your 2nd question. The gap in performance between CV and privateLB arises from two factors:</p>\n<p>1) As also noted in the Results section of the readme, the model itself quite accurately inferred the class balance of subjects in privateLB, i.e. predictions of each subject got noticeably shifted w.r.t. one another, as according to their overall performances in the experiment. This caused a significant boost in global AUC. We think the effect was not (or vaguely) seen in CV due to a smaller test set size.</p>\n<p>2) The bias-variance tradeoff of CV. Concretely, the less folds you use in CV the more biased the performance will be (worse than the actual performance). In 4-fold CV you have 12 training subjects, whereas when preparing a submission you have 16 subjects, this is an addition of significant amount of information. You can get a &quot;boost&quot; in CV score by simply using 8 folds etc. On a side note, there were certain settings that brought a significant improvement in CV, but more rigorous tests (explained in the last paragraph <a href=\"https://www.kaggle.com/c/inria-bci-challenge/forums/t/12635/learning-from-this-competition/65122#post65122\">here</a>) revealed that they could potentially lead us to severe overfitting, despite promising publicLB scores.</p>\n<p>Cheers,</p>\n<p>Rafal</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "67689",
      "postDate": "03/23/2015 10:29:21",
      "content": "<p>Alexander, </p>\n<p>I'd like to congratulate you and your partner for this victory.</p>\n<p>Sometimes it looks like there's gap between academics and real world (i.e. Kaggle problems), and many times data engineers win competitions just using some 20 years-old techniques (such as Random Forest) and a sensible feature choice, while researchers publish complex papers in top machine-learning conferences claiming to improve the state-of-the-art.</p>\n<p>In you case, however, academic excellence meets performance excellence. Winning two different competitions using your Riemann-geometric stuff is definitely an incredibly outstanding achievement. </p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "67810",
      "postDate": "03/23/2015 22:54:26",
      "content": "<p>Hi Alexandre &amp; Rafal.</p>\n<p>Many thanks for your very detailed answers! I absolutely agree with Jose that your solution is fascinating. It combines both extremely&nbsp;interesting mathematical ideas and outstanding performance (and this combination as rare as unicorns:) ). Moreover you allows us to catch some internal&nbsp;details of your algorithms, so we should be (and we are!) much obliged to you.</p>\n<p>Alexandre,&nbsp;</p>\n<p>Thanks, I understand your motivation of using channel selection here (in fact you have to use CSP in DecMeg due to huge number of MEG sensors). We've played&nbsp;a little with your DecMeg code and noticed two issues you mentioned (off-line analysis only and a requirement of balanced classes). Not so evident result (for me) is that performance of retraining strongly depends on distance type. The best choice is riemann (loo accuracy is&nbsp;~0.756); logeuclid also shows competitive results (~0.744). Euclid doesn't work at all (~0.625). The most surprising result is&nbsp;that kullback slightly improves (~0.728)&nbsp;result of the generic classifier, but still shows worse accuracy than riemann. If anyone had asked me how to measure distance between two distributions I would have answered 'certainly&nbsp;kullback leibler'. Now I'm not so confident :)&nbsp;</p>\n<p>I understand that performance of different metrics depends on a particular dataset, so kullback can get own back and exceed riemann in another problem. But I have no idea why logeuclid works at all :) Both riemann and kullback have some understandable mathematical basis (thank you for the IG book, I'll definitely read it!), but what about logeuclid? Could you please give a short comment why does it work?</p>\n<p>Rafal,&nbsp;</p>\n<p>I agree, global AUC&nbsp;effect depends on sample size, so it boosts performance on the large Private set in comparison with single-subject AUC. However, the Public set is rather small, but your approach shows&nbsp;almost the same performance as for the Private set. Probably the reason is that there are &quot;easy&quot; subjects in the Public set.</p>\n<p>Thanks again,</p>\n<p>Mike</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72858",
      "postDate": "04/21/2015 06:17:08",
      "content": "<p>hi sir, is there anyway we can have your features(not the meta features),it would be great help.we need to examine how that works in our classifier,and the accuracy it could manage with your features..</p>\n<p>thanking you in advance,</p>\n<p>Sunu A</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72966",
      "postDate": "04/21/2015 17:02:43",
      "content": "<p>Hi Mike,</p>\n<p>Thanks for looking into the decMeg code and pointing the issue of the metric.&nbsp;</p>\n<ul>\n<li><span style=\"line-height: 1.4\">The Riemannian metric can be considered as a gold standard. It has numerous interesting properties. The most useful is the invariance by congruent transformation. Therefore, In a source separation point of view, the Riemannian distance in the sensor space is equivalent of the Riemannian distance in the source space. So under some circumstances, you can skip the source separation step and directly use the Riemannian distance without any loss of information.<br></span></li>\n<li><span style=\"line-height: 1.4\">Obviously, the Euclidean metric is a bad choice and lacks of sensitivity , because it does not take into account the particular structure of the covariance matrices (SPD). It could give good results when data are clean and classe are really separable (for example in EMG signal) but i would not recommend to use it.</span></li>\n<li><span style=\"line-height: 1.4\">The log-euclid metric is</span><span style=\"line-height: 1.4\">&nbsp;a good approximation of the Riemannian metric. Actually, the two metric are equivalent if all the matrices commute in multiplication. This is the case for diagonal matrices. This is why you can improve results by whitening your signal before using the log-euclid metric. However, this metric does not have the property of invariance by congruent transformation, which makes the results sensitive to preprocessing steps like spatial filtering. However, this metric is less computationally extensive, so i use it when i want to tune parameters.</span></li>\n<li><span style=\"line-height: 1.4\">The kullback-leibler divergence is not a metric because it does not have the symmetry property. You can use the symmetrised version, but it is an artificial workaround. There is also no definition of the mean according to the kullback-leibler divergence. However, used with the Riemannian mean, it can give surprisingly good results. After all, the kullback leibler is also based on the Fisher information, and is actually a measure of the manifold curvature (the second derivative of the Riemannian metric)</span></li>\n<li><span style=\"line-height: 1.4\">Another interesting metric is the log-det metric (based on alpha divergence), and described <a href=\"http://www.sciencedirect.com/science/article/pii/S002437951100783X#\">here</a>.</span></li>\n</ul>\n<p><span style=\"line-height: 1.4\">if i have to rank the metric by their efficiency <strong>on EEG data</strong>, i would do it like this :</span></p>\n<p><span style=\"line-height: 1.4\">Riemann &gt; Log-det &gt; Log-euclid &gt; kullback &gt; Euclidean</span></p>\n<p><span style=\"line-height: 1.4\">It may be different for other problem, as you said. The log-euclidean seems very popular in image processing and computer vision.&nbsp;</span></p>\n<p>As a side note, i just released a python package,&nbsp;<a href=\"https://github.com/alexandrebarachant/pyRiemann\">pyRiemann</a>. It is an extended version of the code developed for this challenge, and it support the different metrics I mention in this post (except the kullblack, but i accept pull request :) )</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72969",
      "postDate": "04/21/2015 17:10:44",
      "content": "<p>[quote=Sunu A;72858]</p>\n<p>hi sir, is there anyway we can have your features(not the meta features),it would be great help.we need to examine how that works in our classifier,and the accuracy it could manage with your features..</p>\n<p>thanking you in advance,</p>\n<p>Sunu A</p>\n<p>[/quote]</p>\n<p>Hi,</p>\n<p>The feature extraction we used include numerous supervised steps, like Xdawn spatial filtering, special form covariance estimation or channel selection. Giving you the raw feature will lead to biased results.</p>\n<p>However, we published the code with a easy mechanism to test your ideas. All of this is explained <a href=\"https://github.com/alexandrebarachant/bci-challenge-ner-2015#parameter-file\">here</a>. All you have to do is to build a python class that inherit from the sklearn base API (basically, two method, fit() and predict() ) and then declare it in the parameter file.</p>\n<p>This way, you will evaluates the performance of your classifier in the same scheme as our, and you will get comparable results.</p>\n\n<p>Alex</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "77258",
      "postDate": "05/06/2015 06:23:59",
      "content": "<p>Hello&nbsp; sir...</p>\n<p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; i have tried your code but i keep on ending up with memory error.As you have noted, i reduced my no of cores but still...In between,my system has only 4GB ram with four cores.Does your code works with my system spec?&nbsp; i'm very new to python</p>\n\n\n\n<p>thanking you in advance</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "77277",
      "postDate": "05/06/2015 08:30:51",
      "content": "<p>Hi Alexandre!&nbsp;</p>\n<p>As usual, many thanks for your very detailed answer&nbsp;which helps us to better understand your ideas!</p>\n<p>1) About metrics: I totally agree with you :) Some comments are below.</p>\n<ul>\n<li><span style=\"line-height: 1.4\">I've already read your papers, so link&nbsp;between Riemann and log-euclid became clear as well as euclid drawbacks. </span></li>\n<li><span style=\"line-height: 1.4\">As for Kullback-Leibler divergence, while its symmetric version isn't very elegant, it's rather popular metric in many algorithms.</span></li>\n<li><span style=\"line-height: 1.4\">I also agree that in practical&nbsp;data science it's impossible to predict which mathematical idea will be the best choice, so your insights into analysis of EEG data is very&nbsp;valuable!</span></li>\n<li><span style=\"line-height: 1.4\">Thanks for log-det, it's a new idea for me.</span></li>\n<li><span style=\"line-height: 1.4\">It seems that Informational Geometry approaches become rather popular in BCI&nbsp;</span>community. A few examples from BBCI group:&nbsp;<a href=\"http://papers.nips.cc/paper/4922-robust-spatial-filtering-with-beta-divergence\">Robust Spatial Filtering with Beta Divergence</a>,&nbsp;<a href=\"http://ieeexplore.ieee.org/xpl/login.jsp?tp=&arnumber=6782545&url=http:%2F%2Fieeexplore.ieee.org%2Fxpls%2Fabs_all.jsp%3Farnumber%3D6782545\">Information geometry meets BCI spatial filtering using divergences</a>,&nbsp;<a href=\"http://ieeexplore.ieee.org/xpl/login.jsp?tp=&arnumber=7073030&url=http:%2F%2Fieeexplore.ieee.org%2Fxpls%2Fabs_all.jsp%3Farnumber%3D7073030\"> Robust common spatial patterns based on Bhattacharyya distance and Gamma divergence</a>.</li>\n</ul>\n<p>2) I reckon&nbsp;there is a huge problem in algorithmic development for BCI: it's hard to compare results from different publications. There are three main issues:</p>\n<ul>\n<li><span style=\"line-height: 1.4\">lack&nbsp;of standard benchmarks (although BBCI competitions datasets are very popular). Moreover, there a lot of papers which show results on their private datasets. It would be great to have something like&nbsp;<a href=\"http://rodrigob.github.io/are_we_there_yet/build/classification_datasets_results.html\">this</a>.</span></li>\n<li><span style=\"line-height: 1.4\">lack of code sharing culture. Algorithms like CSP can be implemented straightforward, so it's rather easy to compare your-new-cutting-edge-algorithm with such&nbsp;classical algorithms. But if you want to compare with something complex and non published, you have to expend your time in coding or to compare with classical algorithms only. This issue leads to the next one:</span></li>\n<li><span style=\"line-height: 1.4\">lack of comparison with recent algorithms. It's not hard to defeat ancient simple algorithms, but a lot of people already did it, so you have to compete with them.</span></li>\n</ul>\n<p><span style=\"line-height: 1.4\">(I've started&nbsp;exploring BCI data science&nbsp;problems only few months ago, so I may&nbsp;be wrong. Please correct me in that case.)</span></p>\n<p><span style=\"line-height: 1.4\">So in my opinion&nbsp;pyRiemann release is a great news (as well as your Matlab covariance Toolbox)! </span></p>\n<p>I plan to finish my current project in aerospace in 1-2 months and to dive into analysis of BCI data. So I hope I&nbsp;will have&nbsp;something to commit in&nbsp;pyRiemann :)</p>\n\n<p>Best regards,&nbsp;</p>\n<p>Mike</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "77278",
      "postDate": "05/06/2015 08:33:50",
      "content": "<p>[quote=Sunu A;77258]</p>\n<p>Hello sir...</p>\n<p>i have tried your code but i keep on ending up with memory error.As you have noted, i reduced my no of cores but still...In between,my system has only 4GB ram with four cores.Does your code works with my system spec? i'm very new to python</p>\n<p>thanking you in advance</p>\n<p>[/quote]</p>\n<p>According to <a href=\"https://github.com/alexandrebarachant/bci-challenge-ner-2015#dependencies--requirements\">specification</a>&nbsp;you need at least 8GB of RAM. It's rather typical situation&nbsp;for EEG analysis due to huge size of raw data.</p>\n<p>Best regards,</p>\n<p>Mike</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "77342",
      "postDate": "05/06/2015 15:46:13",
      "content": "<p>Hi Mike,</p>\n<p>Thank you for carrying this interesting discussion.</p>\n<p>I could not agree more with you :)</p>\n<p>Information geometry is gaining &nbsp;momentum in a lot of different domain. We can see a lot of stuff in computer vision &nbsp;for example. I would explain that by the facts that :</p>\n<ul>\n<li>There is a strong theoretical background that make the methods elegant.</li>\n<li>It brings a new way to tacle problems, and thats refreshing.</li>\n<li>the two point above make the thing easier to publish. Even if you don't show cutting edge performances, the novelty of the approach is well appreciated by the reviewers.</li>\n</ul>\n\n<p>About BCI, i totally share you point of view. There is a lack of&nbsp;<a href=\"http://en.wikipedia.org/wiki/Reproducibility#Reproducible_research\">reproducible research</a>. The problem does not only comes from the methods implementation, but also from the environment it is applied. Results are so bad and datasets are so small that even changing the order of the frequential filter you apply as preprocessing step can change significantly the outcome of you classification.</p>\n<p>In addition, parameters are usually over-optimized for a specific&nbsp;dataset, and when you tried a new method on a new set of data, it require an high amount of domain specific knowledge to make it works properly. This partially explain the popularity of the 'old' methods like CSP. They does not give optimal results, but they are stable across a wide range of applications.</p>\n<p>The other problem is that EEG data are considered as medical data and are really hard to release publicly. You need consents to release from the subject, and it is rarely done before the experiment. There is a lot of data that are just sitting on the labs hard drive without any chance to be released even if you want to (i have tons of data like this).</p>\n<p>What we really need is a platform where you can submit your code and your method will be blindly evaluated on a wide range of dataset (you can even add realistic stress test, like adding noise on data or on labels, etc).</p>\n<p>Finally, thanks for your interest for <a href=\"https://github.com/alexandrebarachant/pyRiemann\">pyRiemann</a>. your commits will be greatly appreciated :)</p>\n<p>Best,</p>\n<p>Alexandre</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "91984",
      "postDate": "09/10/2015 06:53:54",
      "content": "<p>from where can i download data sets?</p>",
      "rawMarkdown": "from where can i download data sets?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 66291,
      "author_name": "devinanzelmo",
      "author_url": "",
      "post_date": "03/15/2015 21:55:09",
      "content": "<p>Thanks for posting your teams solutions! I have been looking forward seeing your methodology since the competition ended. The python code looks quite usable. I was considering converting your matlab code from decMeg to experiment with but now I don't have to. I will post some questions when I have time to closely read the solution.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 66295,
      "author_name": "tf6452",
      "author_url": "",
      "post_date": "03/15/2015 22:48:57",
      "content": "<p>Great job! Are you finding Riemannian Geometry to be an indispensable tool for signal processing and BCI-related work?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 66296,
      "author_name": "alexandrebarachant",
      "author_url": "",
      "post_date": "03/15/2015 22:49:37",
      "content": "<p>You're welcome :)</p>\n<p>For me, it was a good opportunity to switch to python, and i have to say that i will never go back to matlab.</p>\n<p>Now that i have an equivalent of my <a href=\"https://github.com/alexandrebarachant/covariancetoolbox\">covariance toolbox</a>&nbsp;in python, I plan to release a package as soon as possible.</p>\n<p>We took our time before publishing the code, but we wanted to make a deeper analysis of our solution. The final solution is a bit computationally intensive, but you can speed of the process by turning off the bagging and the electrode selection.</p>\n<p>It turns out that by doing this you can achieve a score ~0.843 (for the model without the leak) and it will takes around 2 minutes to process.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 66301,
      "author_name": "alexandrebarachant",
      "author_url": "",
      "post_date": "03/15/2015 23:43:06",
      "content": "<p>[quote=TF;66295]</p>\n<p>Great job! Are you finding Riemannian Geometry to be an indispensable tool for signal processing and BCI-related work?</p>\n<p>[/quote]</p>\n<p>I would say definitely yes but it will be a bit unfair for all the other people who have done amazing work in the BCI community.</p>\n<p>At first, i developed these methods to get rid of the source separation step (i.e. spatial filtering) during online BCI experiments. While algorithm such Common Spatial Patterns (CSP) or XDAWN are really powerful, they lack of robustness and introduce extra parameters to tune. And if they are not used correctly, they can lead to a loss of spatial information.</p>\n<p>By using covariance matrices as features, you make sure you get all the spatial information in your classifier. Then, tools from Riemannian geometry provide you a convenient and robust way to manipulate them.&nbsp;More specifically, the tangent space mapping (which can be see as a kernel operation) is a very powerful tool. It allows to transform these matrices in euclidean vectors without breaking their particular structure &nbsp;(Symmetric and positive definite). Then, you can do whatever you want with these vectors.</p>\n<p>Now, i have to say that you can probably get similar results with other methods, but thanks to a bunch of properties of the Riemannian metric, it is very well suited for cross-subject classification (which was the case for DecMeg and this challenge).</p>\n<p>Covariance matrices are involved in wide range of&nbsp;signal processing problem (well, let say every time you have a multivariate problem), so&nbsp;riemannian geometry has a lot of potential. It is also used in <a href=\"http://ieeexplore.ieee.org/xpl/login.jsp?tp=&arnumber=6514112&url=http:%2F%2Fieeexplore.ieee.org%2Fxpls%2Fabs_all.jsp%3Farnumber%3D6514112\">Radar signal&nbsp;processing</a>&nbsp;and&nbsp;<a href=\"http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.336.4231&rep=rep1&type=pdf\">Image classification</a>.</p>\n<p>Finally, I wanted to say that the riemannian metric i use is a special case of something bigger called Information Geometry, a theory that study the geometric properties of probability distributions. In my case, the distribution of interest is the multivariate normal distribution with zero mean (only parametrized by it covariance matrix), but such metric can be extracted for any other distribution (but the math behind can be hardcore).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 66327,
      "author_name": "mibell",
      "author_url": "",
      "post_date": "03/16/2015 07:06:43",
      "content": "<p>Alexandre,&nbsp;</p>\n<p>many thanks for sharing your approach and your code!</p>\n<p>I have a few questions</p>\n<p>1) In DecMeg competition you used CSP filters in order to reduce dimensionality of the problem (DR seems to be an important preprocessing step for Riemannian geometry tools). Here you used&nbsp;a channel selection procedure instead of CSP. Could you please explain a little more about the motivation of this decision?</p>\n<p>2) Why did you obtain such a huge difference between&nbsp;Fold-wise CV (0.7294) and Private/Public LB (0.84585/0.85083)? As far as I understand you used the same criterion (joint AUC) for Fold-Wise CV calculation. The only difference I see is that you used 150 bagged models for CV and 500 bagged models for the final submission. But according to your post-competition analysis even 10 models give ~0.84 on the LB datasets.</p>\n<p>3) Could you please recommend useful books&nbsp;on Information Geometry with all necessary&nbsp;math concepts (even hardcore ones)?&nbsp;</p>\n<p>Thanks in advance,&nbsp;</p>\n<p>Mike</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 66469,
      "author_name": "alexandrebarachant",
      "author_url": "",
      "post_date": "03/16/2015 19:52:09",
      "content": "<p>Hi Mike,</p>\n<p>these are very relevant questions.</p>\n<p>1) There is multiple reason we took this decision, but first lets got back to the&nbsp;dimensionality reduction&nbsp;for Riemannian geometry (RG). Dealing with High dimensional data is the main problem when using RG. The main reasons are :</p>\n<ul>\n<ul>\n<li>Your matrices must be definite positive and this is not always the case when the number of channel is high. You can still regularize, but it can degrade perfs by affecting the structure of the matrix.</li>\n<li>Tangent space mapping is a local approximation of the manifold. it suppose that all your matrices are scattered in a relatively small part of the manifold. when you increase the dimension, you decrease the density of your data (or increase the sparsity), and it result in bad tangent space approximation.</li>\n<li>Everything become more and more computationally expensive. RG involve a lot of eigenvalue/eigenvector decomposition.</li>\n</ul>\n</ul>\n<p>My experience&nbsp;with EEG/MEG data is that if you keep everything bellow 64 channel, you will have no trouble.</p>\n<p>For the Decmeg Challenge, the dimensionality was simply to high to use any of my methods (including the channel selection) out of the box. So i had no choice and i decided to use spatial filtering.</p>\n<p>The problem with spatial filtering is that is is not very well suited for a cross-subject design. The basic hypothesis of&nbsp;source separation&nbsp;is that the training data and the test data&nbsp;are mixed by the same operator&nbsp;and are stationary. When you train on a subject, and apply on another, you take the risk to discard relevant information in the process. Again, when&nbsp;you aggregates data from different subject for training, you source separation becomes less effective because your algorithm try to estimate a unique unmixing matrix.&nbsp;</p>\n<p>For decMeg, i found a workaround. My classification was in two stage, the first was a cross-subject classifier, using stacked generalization (so spatial filters were not train on aggregated data), the second was an iterative retraining using the data on the test subject only (the case where spatial filtering is really efficient). As you see, the key is to know what are the advantages and the flaws of you methods in order to use them in the best possible conditions.</p>\n<p>Now for this challenge, the dimensionality was manageable, so spatial filtering was not necessary. Still, it is a good thing to reduce dimensionality, so we apply the channel selection procedure to discard irrelevant channels. The number of channels&nbsp;we kept&nbsp;is not very tight (35). it makes everything running more smoothly without having much risk to discard relevant information. EEG has a low spatial resolution, so close channels are highly correlated and often carry the same information. Discarding a few of them doesn't hurt</p>\n<p>So the next question is why we didn't apply the same retraining process for this challenge. Well, first because we wanted to keep thing as 'online' as possible. The second is simply because the convergence were more tricky. The method i used in Decmeg was based on the assumption&nbsp;that the classes were&nbsp;balanced (which was the case by design), in order to set the optimal threshold for the&nbsp;re-labeling between iteration. Here the classes balance were unknown (if you don't use the leak).</p>\n\n<p>2) Okay, now for the second question, i will let my teammate answer.</p>\n<p>3) I think the best book you can find is &quot;Methods of Information Geometry&quot; from S. Amari. The basic idea is that each probability distribution is a point of a manifold with the parameters of the distribution as coordinates. For example a normal distribution of mean u and variance s is a point of coordinate (u,s) in the manifold. The trick is, the Information geometry make the assumption that the natural metric for this manifold is the Fisher information. After a bunch of equation manipulation, it allow you to define a true distance between two distribution, and therefore unlock a lot of issue&nbsp;(for example, the mean of a distribution is not the distribution with the mean parameters :) &nbsp;)</p>\n\n<p>Alex</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 66509,
      "author_name": "rafalcycon",
      "author_url": "",
      "post_date": "03/16/2015 21:21:25",
      "content": "<p>Hey Mike, let me answer your 2nd question. The gap in performance between CV and privateLB arises from two factors:</p>\n<p>1) As also noted in the Results section of the readme, the model itself quite accurately inferred the class balance of subjects in privateLB, i.e. predictions of each subject got noticeably shifted w.r.t. one another, as according to their overall performances in the experiment. This caused a significant boost in global AUC. We think the effect was not (or vaguely) seen in CV due to a smaller test set size.</p>\n<p>2) The bias-variance tradeoff of CV. Concretely, the less folds you use in CV the more biased the performance will be (worse than the actual performance). In 4-fold CV you have 12 training subjects, whereas when preparing a submission you have 16 subjects, this is an addition of significant amount of information. You can get a &quot;boost&quot; in CV score by simply using 8 folds etc. On a side note, there were certain settings that brought a significant improvement in CV, but more rigorous tests (explained in the last paragraph <a href=\"https://www.kaggle.com/c/inria-bci-challenge/forums/t/12635/learning-from-this-competition/65122#post65122\">here</a>) revealed that they could potentially lead us to severe overfitting, despite promising publicLB scores.</p>\n<p>Cheers,</p>\n<p>Rafal</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 67689,
      "author_name": "joseleiva",
      "author_url": "",
      "post_date": "03/23/2015 10:29:21",
      "content": "<p>Alexander, </p>\n<p>I'd like to congratulate you and your partner for this victory.</p>\n<p>Sometimes it looks like there's gap between academics and real world (i.e. Kaggle problems), and many times data engineers win competitions just using some 20 years-old techniques (such as Random Forest) and a sensible feature choice, while researchers publish complex papers in top machine-learning conferences claiming to improve the state-of-the-art.</p>\n<p>In you case, however, academic excellence meets performance excellence. Winning two different competitions using your Riemann-geometric stuff is definitely an incredibly outstanding achievement. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 67810,
      "author_name": "mibell",
      "author_url": "",
      "post_date": "03/23/2015 22:54:26",
      "content": "<p>Hi Alexandre &amp; Rafal.</p>\n<p>Many thanks for your very detailed answers! I absolutely agree with Jose that your solution is fascinating. It combines both extremely&nbsp;interesting mathematical ideas and outstanding performance (and this combination as rare as unicorns:) ). Moreover you allows us to catch some internal&nbsp;details of your algorithms, so we should be (and we are!) much obliged to you.</p>\n<p>Alexandre,&nbsp;</p>\n<p>Thanks, I understand your motivation of using channel selection here (in fact you have to use CSP in DecMeg due to huge number of MEG sensors). We've played&nbsp;a little with your DecMeg code and noticed two issues you mentioned (off-line analysis only and a requirement of balanced classes). Not so evident result (for me) is that performance of retraining strongly depends on distance type. The best choice is riemann (loo accuracy is&nbsp;~0.756); logeuclid also shows competitive results (~0.744). Euclid doesn't work at all (~0.625). The most surprising result is&nbsp;that kullback slightly improves (~0.728)&nbsp;result of the generic classifier, but still shows worse accuracy than riemann. If anyone had asked me how to measure distance between two distributions I would have answered 'certainly&nbsp;kullback leibler'. Now I'm not so confident :)&nbsp;</p>\n<p>I understand that performance of different metrics depends on a particular dataset, so kullback can get own back and exceed riemann in another problem. But I have no idea why logeuclid works at all :) Both riemann and kullback have some understandable mathematical basis (thank you for the IG book, I'll definitely read it!), but what about logeuclid? Could you please give a short comment why does it work?</p>\n<p>Rafal,&nbsp;</p>\n<p>I agree, global AUC&nbsp;effect depends on sample size, so it boosts performance on the large Private set in comparison with single-subject AUC. However, the Public set is rather small, but your approach shows&nbsp;almost the same performance as for the Private set. Probably the reason is that there are &quot;easy&quot; subjects in the Public set.</p>\n<p>Thanks again,</p>\n<p>Mike</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72858,
      "author_name": "sunu3848",
      "author_url": "",
      "post_date": "04/21/2015 06:17:08",
      "content": "<p>hi sir, is there anyway we can have your features(not the meta features),it would be great help.we need to examine how that works in our classifier,and the accuracy it could manage with your features..</p>\n<p>thanking you in advance,</p>\n<p>Sunu A</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72966,
      "author_name": "alexandrebarachant",
      "author_url": "",
      "post_date": "04/21/2015 17:02:43",
      "content": "<p>Hi Mike,</p>\n<p>Thanks for looking into the decMeg code and pointing the issue of the metric.&nbsp;</p>\n<ul>\n<li><span style=\"line-height: 1.4\">The Riemannian metric can be considered as a gold standard. It has numerous interesting properties. The most useful is the invariance by congruent transformation. Therefore, In a source separation point of view, the Riemannian distance in the sensor space is equivalent of the Riemannian distance in the source space. So under some circumstances, you can skip the source separation step and directly use the Riemannian distance without any loss of information.<br></span></li>\n<li><span style=\"line-height: 1.4\">Obviously, the Euclidean metric is a bad choice and lacks of sensitivity , because it does not take into account the particular structure of the covariance matrices (SPD). It could give good results when data are clean and classe are really separable (for example in EMG signal) but i would not recommend to use it.</span></li>\n<li><span style=\"line-height: 1.4\">The log-euclid metric is</span><span style=\"line-height: 1.4\">&nbsp;a good approximation of the Riemannian metric. Actually, the two metric are equivalent if all the matrices commute in multiplication. This is the case for diagonal matrices. This is why you can improve results by whitening your signal before using the log-euclid metric. However, this metric does not have the property of invariance by congruent transformation, which makes the results sensitive to preprocessing steps like spatial filtering. However, this metric is less computationally extensive, so i use it when i want to tune parameters.</span></li>\n<li><span style=\"line-height: 1.4\">The kullback-leibler divergence is not a metric because it does not have the symmetry property. You can use the symmetrised version, but it is an artificial workaround. There is also no definition of the mean according to the kullback-leibler divergence. However, used with the Riemannian mean, it can give surprisingly good results. After all, the kullback leibler is also based on the Fisher information, and is actually a measure of the manifold curvature (the second derivative of the Riemannian metric)</span></li>\n<li><span style=\"line-height: 1.4\">Another interesting metric is the log-det metric (based on alpha divergence), and described <a href=\"http://www.sciencedirect.com/science/article/pii/S002437951100783X#\">here</a>.</span></li>\n</ul>\n<p><span style=\"line-height: 1.4\">if i have to rank the metric by their efficiency <strong>on EEG data</strong>, i would do it like this :</span></p>\n<p><span style=\"line-height: 1.4\">Riemann &gt; Log-det &gt; Log-euclid &gt; kullback &gt; Euclidean</span></p>\n<p><span style=\"line-height: 1.4\">It may be different for other problem, as you said. The log-euclidean seems very popular in image processing and computer vision.&nbsp;</span></p>\n<p>As a side note, i just released a python package,&nbsp;<a href=\"https://github.com/alexandrebarachant/pyRiemann\">pyRiemann</a>. It is an extended version of the code developed for this challenge, and it support the different metrics I mention in this post (except the kullblack, but i accept pull request :) )</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72969,
      "author_name": "alexandrebarachant",
      "author_url": "",
      "post_date": "04/21/2015 17:10:44",
      "content": "<p>[quote=Sunu A;72858]</p>\n<p>hi sir, is there anyway we can have your features(not the meta features),it would be great help.we need to examine how that works in our classifier,and the accuracy it could manage with your features..</p>\n<p>thanking you in advance,</p>\n<p>Sunu A</p>\n<p>[/quote]</p>\n<p>Hi,</p>\n<p>The feature extraction we used include numerous supervised steps, like Xdawn spatial filtering, special form covariance estimation or channel selection. Giving you the raw feature will lead to biased results.</p>\n<p>However, we published the code with a easy mechanism to test your ideas. All of this is explained <a href=\"https://github.com/alexandrebarachant/bci-challenge-ner-2015#parameter-file\">here</a>. All you have to do is to build a python class that inherit from the sklearn base API (basically, two method, fit() and predict() ) and then declare it in the parameter file.</p>\n<p>This way, you will evaluates the performance of your classifier in the same scheme as our, and you will get comparable results.</p>\n\n<p>Alex</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 77258,
      "author_name": "sunu3848",
      "author_url": "",
      "post_date": "05/06/2015 06:23:59",
      "content": "<p>Hello&nbsp; sir...</p>\n<p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; i have tried your code but i keep on ending up with memory error.As you have noted, i reduced my no of cores but still...In between,my system has only 4GB ram with four cores.Does your code works with my system spec?&nbsp; i'm very new to python</p>\n\n\n\n<p>thanking you in advance</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 77277,
      "author_name": "mibell",
      "author_url": "",
      "post_date": "05/06/2015 08:30:51",
      "content": "<p>Hi Alexandre!&nbsp;</p>\n<p>As usual, many thanks for your very detailed answer&nbsp;which helps us to better understand your ideas!</p>\n<p>1) About metrics: I totally agree with you :) Some comments are below.</p>\n<ul>\n<li><span style=\"line-height: 1.4\">I've already read your papers, so link&nbsp;between Riemann and log-euclid became clear as well as euclid drawbacks. </span></li>\n<li><span style=\"line-height: 1.4\">As for Kullback-Leibler divergence, while its symmetric version isn't very elegant, it's rather popular metric in many algorithms.</span></li>\n<li><span style=\"line-height: 1.4\">I also agree that in practical&nbsp;data science it's impossible to predict which mathematical idea will be the best choice, so your insights into analysis of EEG data is very&nbsp;valuable!</span></li>\n<li><span style=\"line-height: 1.4\">Thanks for log-det, it's a new idea for me.</span></li>\n<li><span style=\"line-height: 1.4\">It seems that Informational Geometry approaches become rather popular in BCI&nbsp;</span>community. A few examples from BBCI group:&nbsp;<a href=\"http://papers.nips.cc/paper/4922-robust-spatial-filtering-with-beta-divergence\">Robust Spatial Filtering with Beta Divergence</a>,&nbsp;<a href=\"http://ieeexplore.ieee.org/xpl/login.jsp?tp=&arnumber=6782545&url=http:%2F%2Fieeexplore.ieee.org%2Fxpls%2Fabs_all.jsp%3Farnumber%3D6782545\">Information geometry meets BCI spatial filtering using divergences</a>,&nbsp;<a href=\"http://ieeexplore.ieee.org/xpl/login.jsp?tp=&arnumber=7073030&url=http:%2F%2Fieeexplore.ieee.org%2Fxpls%2Fabs_all.jsp%3Farnumber%3D7073030\"> Robust common spatial patterns based on Bhattacharyya distance and Gamma divergence</a>.</li>\n</ul>\n<p>2) I reckon&nbsp;there is a huge problem in algorithmic development for BCI: it's hard to compare results from different publications. There are three main issues:</p>\n<ul>\n<li><span style=\"line-height: 1.4\">lack&nbsp;of standard benchmarks (although BBCI competitions datasets are very popular). Moreover, there a lot of papers which show results on their private datasets. It would be great to have something like&nbsp;<a href=\"http://rodrigob.github.io/are_we_there_yet/build/classification_datasets_results.html\">this</a>.</span></li>\n<li><span style=\"line-height: 1.4\">lack of code sharing culture. Algorithms like CSP can be implemented straightforward, so it's rather easy to compare your-new-cutting-edge-algorithm with such&nbsp;classical algorithms. But if you want to compare with something complex and non published, you have to expend your time in coding or to compare with classical algorithms only. This issue leads to the next one:</span></li>\n<li><span style=\"line-height: 1.4\">lack of comparison with recent algorithms. It's not hard to defeat ancient simple algorithms, but a lot of people already did it, so you have to compete with them.</span></li>\n</ul>\n<p><span style=\"line-height: 1.4\">(I've started&nbsp;exploring BCI data science&nbsp;problems only few months ago, so I may&nbsp;be wrong. Please correct me in that case.)</span></p>\n<p><span style=\"line-height: 1.4\">So in my opinion&nbsp;pyRiemann release is a great news (as well as your Matlab covariance Toolbox)! </span></p>\n<p>I plan to finish my current project in aerospace in 1-2 months and to dive into analysis of BCI data. So I hope I&nbsp;will have&nbsp;something to commit in&nbsp;pyRiemann :)</p>\n\n<p>Best regards,&nbsp;</p>\n<p>Mike</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 77278,
      "author_name": "mibell",
      "author_url": "",
      "post_date": "05/06/2015 08:33:50",
      "content": "<p>[quote=Sunu A;77258]</p>\n<p>Hello sir...</p>\n<p>i have tried your code but i keep on ending up with memory error.As you have noted, i reduced my no of cores but still...In between,my system has only 4GB ram with four cores.Does your code works with my system spec? i'm very new to python</p>\n<p>thanking you in advance</p>\n<p>[/quote]</p>\n<p>According to <a href=\"https://github.com/alexandrebarachant/bci-challenge-ner-2015#dependencies--requirements\">specification</a>&nbsp;you need at least 8GB of RAM. It's rather typical situation&nbsp;for EEG analysis due to huge size of raw data.</p>\n<p>Best regards,</p>\n<p>Mike</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 77342,
      "author_name": "alexandrebarachant",
      "author_url": "",
      "post_date": "05/06/2015 15:46:13",
      "content": "<p>Hi Mike,</p>\n<p>Thank you for carrying this interesting discussion.</p>\n<p>I could not agree more with you :)</p>\n<p>Information geometry is gaining &nbsp;momentum in a lot of different domain. We can see a lot of stuff in computer vision &nbsp;for example. I would explain that by the facts that :</p>\n<ul>\n<li>There is a strong theoretical background that make the methods elegant.</li>\n<li>It brings a new way to tacle problems, and thats refreshing.</li>\n<li>the two point above make the thing easier to publish. Even if you don't show cutting edge performances, the novelty of the approach is well appreciated by the reviewers.</li>\n</ul>\n\n<p>About BCI, i totally share you point of view. There is a lack of&nbsp;<a href=\"http://en.wikipedia.org/wiki/Reproducibility#Reproducible_research\">reproducible research</a>. The problem does not only comes from the methods implementation, but also from the environment it is applied. Results are so bad and datasets are so small that even changing the order of the frequential filter you apply as preprocessing step can change significantly the outcome of you classification.</p>\n<p>In addition, parameters are usually over-optimized for a specific&nbsp;dataset, and when you tried a new method on a new set of data, it require an high amount of domain specific knowledge to make it works properly. This partially explain the popularity of the 'old' methods like CSP. They does not give optimal results, but they are stable across a wide range of applications.</p>\n<p>The other problem is that EEG data are considered as medical data and are really hard to release publicly. You need consents to release from the subject, and it is rarely done before the experiment. There is a lot of data that are just sitting on the labs hard drive without any chance to be released even if you want to (i have tons of data like this).</p>\n<p>What we really need is a platform where you can submit your code and your method will be blindly evaluated on a wide range of dataset (you can even add realistic stress test, like adding noise on data or on labels, etc).</p>\n<p>Finally, thanks for your interest for <a href=\"https://github.com/alexandrebarachant/pyRiemann\">pyRiemann</a>. your commits will be greatly appreciated :)</p>\n<p>Best,</p>\n<p>Alexandre</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 91984,
      "author_name": "qasimgilani",
      "author_url": "",
      "post_date": "09/10/2015 06:53:54",
      "content": "<p>from where can i download data sets?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "66251": "",
    "66291": "",
    "66295": "",
    "66296": "",
    "66301": "",
    "66327": "",
    "66469": "",
    "66509": "",
    "67689": "",
    "67810": "",
    "72858": "",
    "72966": "",
    "72969": "",
    "77258": "",
    "77277": "",
    "77278": "",
    "77342": "",
    "91984": "from where can i download data sets?"
  },
  "source": "meta"
}