{
  "id": 11056,
  "title": "Beating the Benchmark with GBM",
  "url": "/competitions/inria-bci-challenge/discussion/11056",
  "author_name": "",
  "post_date": "2014-11-26T00:28:19.163Z",
  "votes": 18,
  "comment_count": 41,
  "views": 14290,
  "content": "<p><strong>NOTE: </strong>The original version of this benchmark included subject number as a feature. This miss-calibrated subjects probability scores resulting in a .47 AUC on the private leader board. An updated version removing the subject column results in .6 AUC on the private leaderboard. I apologize to anyone who heavily modified this code and had decreased score due to the subject feature. See this&nbsp;<a href=\"https://www.kaggle.com/c/inria-bci-challenge/forums/t/12603/inter-subject-auc/64843#post64843\">post</a> by for discussion of inter-subject AUC.</p>\n\n<p><strong>Original post below</strong></p>\n<p>Here is a simple benchmark using sklearn and pandas. It uses the Cz channel for the 1.3 seconds after feedback for each example as training. Uses sklearn gradient boosting classifier with 500 estimators. It should take about 10 to 15 minutes to run.</p>\n<p>There is no cleaning or processing of the data, and it only uses 1 channel so there is plenty of room to expand. Leaderboard score ~.72</p>\n<p>Its also slow in some places so if you want to extract all of the data after feedbacks you should probably modify the code to use a more efficient extraction method. I used method of creating submission file from Abhishek's code as I never dealt with AUC before.</p>\n<p>Edit: Probably better to set max_features to default which is sqrt(n_features) or just delete that part, the current value was left in by accident, and probably is not great.</p>\n<p>EDIT: version two should not have numexpr dependency for pandas query</p>\n<p><strong>EDIT: Version three post competition. Fixes problem that caused low inter-subject AUC. </strong></p>",
  "messages": [
    {
      "id": "58964",
      "postDate": "11/26/2014 00:28:19",
      "content": "<p><strong>NOTE: </strong>The original version of this benchmark included subject number as a feature. This miss-calibrated subjects probability scores resulting in a .47 AUC on the private leader board. An updated version removing the subject column results in .6 AUC on the private leaderboard. I apologize to anyone who heavily modified this code and had decreased score due to the subject feature. See this&nbsp;<a href=\"https://www.kaggle.com/c/inria-bci-challenge/forums/t/12603/inter-subject-auc/64843#post64843\">post</a> by for discussion of inter-subject AUC.</p>\n\n<p><strong>Original post below</strong></p>\n<p>Here is a simple benchmark using sklearn and pandas. It uses the Cz channel for the 1.3 seconds after feedback for each example as training. Uses sklearn gradient boosting classifier with 500 estimators. It should take about 10 to 15 minutes to run.</p>\n<p>There is no cleaning or processing of the data, and it only uses 1 channel so there is plenty of room to expand. Leaderboard score ~.72</p>\n<p>Its also slow in some places so if you want to extract all of the data after feedbacks you should probably modify the code to use a more efficient extraction method. I used method of creating submission file from Abhishek's code as I never dealt with AUC before.</p>\n<p>Edit: Probably better to set max_features to default which is sqrt(n_features) or just delete that part, the current value was left in by accident, and probably is not great.</p>\n<p>EDIT: version two should not have numexpr dependency for pandas query</p>\n<p><strong>EDIT: Version three post competition. Fixes problem that caused low inter-subject AUC. </strong></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58973",
      "postDate": "11/26/2014 05:49:44",
      "content": "<p>i tried your code 'gbm_benchmark.py ' but i am getting the following error&nbsp;</p>\n<p>loading train data<br>Traceback (most recent call last):</p>\n<p><br> File &quot;GbmBenchMark.py&quot;, line 18, in &lt;module&gt;<br> fb = temp.query('FeedBackEvent == 1')['FeedBackEvent']</p>\n<p><br> File &quot;/usr/local/lib/python2.7/dist-packages/pandas/core/frame.py&quot;, line 1822, in query<br> res = self.eval(expr, **kwargs)</p>\n<p><br> File &quot;/usr/local/lib/python2.7/dist-packages/pandas/core/frame.py&quot;, line 1874, in eval<br> return _eval(expr, **kwargs)</p>\n<p><br> File &quot;/usr/local/lib/python2.7/dist-packages/pandas/computation/eval.py&quot;, line 218, in eval<br> _check_engine(engine)</p>\n<p><br> File &quot;/usr/local/lib/python2.7/dist-packages/pandas/computation/eval.py&quot;, line 40, in _check_engine<br> raise ImportError(&quot;'numexpr' not found. Cannot use &quot;</p>\n<p><br>ImportError: 'numexpr' not found. Cannot use engine='numexpr' for query/eval if 'numexpr' is not installed</p>\n<p>please help me to solve this issue</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58974",
      "postDate": "11/26/2014 06:03:34",
      "content": "<p>[quote=phalaris;58964]</p>\n<p>Here is a simple benchmark using sklearn and pandas. It uses the Cz channel for the 1.3 seconds after feedback for each example as training. Uses sklearn gradient boosting classifier with 500 estimators. It should take about 10 to 15 minutes to run.</p>\n<p>There is no cleaning or processing of the data, and it only uses 1 channel so there is plenty of room to expand. Leaderboard score ~.72</p>\n<p>Its also slow in some places so if you want to extract all of the data after feedbacks you should probably modify the code to use a more efficient extraction method. I used method of creating submission file from Abhishek's code as I never dealt with AUC before.</p>\n<p>Edit: Probably better to set max_features to default which is sqrt(n_features) or just delete that part, the current value was left in by accident, and probably is not great.</p>\n<p>[/quote]</p>\n\n<p>Thanks, so we&nbsp;have 2 benchmark now.</p>\n<p>+ rf</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58975",
      "postDate": "11/26/2014 06:32:34",
      "content": "<p>@Lijo numexpr is used by pandas query. I attached another file which should fix the problem. let me know if it works for you. I added engine='python' to the query call, but that line can be rewritten if you are still having problems.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58979",
      "postDate": "11/26/2014 07:32:20",
      "content": "<p>@phalaris the file named&nbsp;gbm_benchmark_v2.py is working.<br>Thanks for your support</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58982",
      "postDate": "11/26/2014 07:56:06",
      "content": "<p>or do this: sudo pip install numexpr</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58983",
      "postDate": "11/26/2014 08:03:23",
      "content": "<p>That will work. I'm&nbsp; not sure how difficult package installation is on windows or mac so I included the other file. I think anaconda python comes with numexpr.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "59027",
      "postDate": "11/26/2014 19:03:05",
      "content": "<p>LB 0.72545 after changing the max_features back to the default. The introduction of the lag is clearly very important as someone pointed out on another thread. So it seems like adding more channels and maybe multiple lags should give improvement. I'm a bit surprised that this approach of using the raw data works this well, I expected a good result would require constructing some good features.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "59033",
      "postDate": "11/26/2014 19:36:14",
      "content": "<p>@James King, By lag do you mean starting some point after the onset of the feedback signal to compensate for the time it takes for the brain respond to the signal?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "59040",
      "postDate": "11/26/2014 19:56:03",
      "content": "<p>[quote=phalaris;59033]</p>\n<p>@James King, By lag do you mean starting some point after the onset of the feedback signal to compensate for the time it takes for the brain respond to the signal?</p>\n<p>[/quote]</p>\n<p>Yes; probably I should have said lead. I meant your 261 observations for 1.3 seconds after the feedback; as opposed to looking at the EEG readings only at the moment of the feedback event, when the subject has not had time to react.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "59044",
      "postDate": "11/26/2014 20:32:47",
      "content": "<p>@James King, thanks for the clarification. 1.3 seconds could be cut down as most of the efforts from the paper for this competition focused on 0-600ms. I think there is a lot of room to improve the score with better features.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "59092",
      "postDate": "11/27/2014 12:38:36",
      "content": "<p>@phalaris , thanks for your support , i also have a doubt in your code.</p>\n<p>there are other columns rather than 'Cz' but you only used values from that column and you take next 261 data's. what is the logic behind 261? &nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "59096",
      "postDate": "11/27/2014 14:19:30",
      "content": "<p>[quote=Lijo Joseph;59092]</p>\n<p>@phalaris , thanks for your support , i also have a doubt in your code.</p>\n<p>there are other columns rather than 'Cz' but you only used values from that column and you take next 261 data's. what is the logic behind 261? &nbsp;</p>\n<p>[/quote]</p>\n\n<p>Data is downsampled to 200Hz, which means the time interval is 5ms. So if you want to get data in 1300ms after the event, you will select 1300/5 = 260 items. 260 or 261 doesn't differ a lot.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "59097",
      "postDate": "11/27/2014 14:36:48",
      "content": "<p>@Lijo, The data is given to us at 200 samples per second , so each row of the original data corresponds to 5ms. In the&nbsp; experiment&nbsp; this data came from an algorithm predicted which letter a person was thinking about, and then showed that letter for 1.3 seconds. We are trying to determine whether the letter shown on the screen is the one the human intended by signals detected with EEG, so I chose data happening when the letter was shown.&nbsp; 260 samples is 5ms*260 = 1300ms =1.3s.&nbsp; Its 261 because I messed up on array indexing, and ended up with an extra sample.</p>\n<p>I only used 'Cz' because it is easy and gives a decent result. You can add extra channels and see if it gives a better result. Looking at couple of the past EEG competitions on Kaggle, it seems most of the effort goes into feature engineering.&nbsp; Take a look at at this&nbsp;<a href=\"https://github.com/MichaelHills/seizure-prediction\">code</a> by Michael Hills for the seizure prediction challenge.&nbsp; Look for transform.py under the seizure-prediction folder. It has many transformations. These can be used to combine channels and extract useful information from channels. It may take a lot of time to find good features</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "59105",
      "postDate": "11/27/2014 17:19:00",
      "content": "<p>I can't cross-validate this method; for example training on subjects</p>\n<p>['02', '06', '16', '18', '20', '21', '24', '26']</p>\n<p>and testing on the others I get a much lower AUC. Possibly I've made some mistake.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "59106",
      "postDate": "11/27/2014 17:36:06",
      "content": "<p>@James King, you are probably not doing anything wrong. As discussed in another thread the LB is likely based on two test subjects. The <a href=\"http://www.hindawi.com/journals/ahci/2012/578295/\">paper</a> which describes the data for competition indicates that the test subjects can be separated into two clusters based on the specificity of an error detection classifier. There are other ways in which test subjects vary. I think we are seeing two 'easy' subjects on the leaderboard which is way CV scores are lower. For the benchmark I believe CV score is around 0.6. Changes in CV score do reflect changes on the leader board just shifted up because of the leaderboard test subject choice.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "59434",
      "postDate": "12/03/2014 03:24:20",
      "content": "<p>You've included subject and session in your model as covariates. This doesn't seem appropriate given that the test data is not really connected. Did you consider this when applying the GBM? I've submitted an adaboostm1 using only subject and session as continuous variables and score ~.62. Adding Cz to the mix boosts it to .7; in fact, any feature seems to provide same &quot;added&quot; value. The signal introduced by the subject session label seems artificial to me. Any thoughts?&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "59445",
      "postDate": "12/03/2014 07:41:17",
      "content": "<p>@Brian, I quickly checked the the contribution of Subject and Session. Taking out Subject decreased my CV score by about .003, and taking out Session decreased it by an additional .018. The increase in score due to including subject is not clear to me, as I am doing CV by subject, it shouldn't change anything.</p>\n<p>Including session information could improve the results if there is either learning over sessions, or alternatively fatigue over sessions. I don't remember the details of experiment but either idea could account for the usefulness of session information. </p>\n<p>In my opinion I don't see a reason not to include both as long as subjects are kept separate during CV</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "59489",
      "postDate": "12/03/2014 18:53:45",
      "content": "<p>Do you mean you are doing CV by splitting data for the same subject, or you are assigning subjects to be entirely in training or test? &nbsp;In the latter case, an indicator for subject may be capturing the ungeneralizable idiosyncrasies of that subject that might otherwise be (incorrectly) incorporated into other estimates, or it may be displacing a feature the model is overfitting on in the random variable subset selection.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "59552",
      "postDate": "12/04/2014 16:22:11",
      "content": "<p>@phalaris, try your CV with more than one subject held out for CV. The reason for doing this is that with one subject held out you are only analyzing&nbsp;your intra-subject AUC.&nbsp;But you also need to check that your inter-subject AUC is working as well. This matters because of the reason you mentioned from the paper--some people make far fewer&nbsp;errors than others, and you need to ensure your calibration picks that up, or at least is robust to it.&nbsp;I think you will see what James King is reporting when you expand your CV, and if you plot the distributions by subject for each CV set you try, you will see why you incur that drop.</p>\n<p>Additionally, when you don't mix subjects, you won't realize the contribution that including the subject ID in your model is making, since it is the same for all your predictions. And it isn't really possible that it's a good contribution unless the ID scheme of the subjects carries some information (which seems unreasonable).</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "59555",
      "postDate": "12/04/2014 17:45:54",
      "content": "<p>@mlandry, thanks for pointing these issues out. I should have clarified more. For the most part I am doing four fold CV. I have also tried leave one out and eight fold CV. I have noticed that the average AUC is different depending on how many subjects I leave out.</p>\n<p>I will try plotting the AUC for each subject based the size of the hold out set.&nbsp; I imagine figuring out which subjects in the test set make fewer errors would be very useful, because we could then train on similar subjects.</p>\n<p>I doubt the subject variable is useful I mainly used it for easily selecting subjects using pandas query(I am new to pandas and currently learning how to reshape the data quickly so I can efficiently try many techniques for generating features).</p>\n<p>Nothing like having ones assumptions pointed out to speed up learning, thanks!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "59562",
      "postDate": "12/04/2014 18:56:49",
      "content": "<p>@phalaris, my mistake on reading it as if you were leaving one out.</p>\n<p>OK, interesting that you still see it across subjects. I wasn't very clear on the plots--what I found interesting when looking through CV scores was the distribution of the predictions by subject. I was looking at box plots with 0-1 in the y axis and each subject on the X axis. I found intra-subject AUC can vary a lot, but it gets far worse when you mix subjects, if the prediction range doesn't line up very well. Of course the overall prediction range is the same, but if you have most of the predictions for a low-accuracy subject occurring just slightly higher than the predictions for a high-accuracy subject, the AUC score will pay a big price.</p>\n<p>That all said, if you're leaving out 2 and 4 subjects at a time and still seeing results comparable to the leaderboard...I am surprised, as my initial take on this was to agree with @James King. Now it is my turn to thankfully have my assumptions pointed out, and I will have some work to do to see the same progress realized with CV tweaks.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "59573",
      "postDate": "12/04/2014 21:25:30",
      "content": "<p>I have been doing leave-one-subject-out cross validation, as&nbsp;I like to have a&nbsp;clear&nbsp;indication of which subjects a method is performing well/poorly on.</p>\n<p>However, I also record the predictions generated for each subject and combine them at the end to calculate a global AUC. When doing this I see a consistent but minor decrease relative to the averaged&nbsp;single subject AUC (typically 0.02).</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "59588",
      "postDate": "12/05/2014 01:50:16",
      "content": "<p>To add to the current string.&nbsp;</p>\n<p>Preamble: Cz was filtered to be between [0.1,60] Hz. Kept 250 samples after feedback, feedback time point, feedback vector index.&nbsp;</p>\n<p>A 30% hold out was used to estimate model performance throughout boosting. Subject/Session included gives 0.709 AUC (consistent with public LB) while exclusion (Cz + feedback time point) lowers to 0.62.&nbsp;</p>\n<p>It's also cool to compare the slope during boosting; see graph</p>\n<p>Matlab code:&nbsp;</p>\n<p>ens = fitensemble(X,y,'adaboostm1',500,'tree',...<br>'prior','uniform','type','classification','learnrate',.05,'holdout',.3);</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "59589",
      "postDate": "12/05/2014 02:19:21",
      "content": "<p>Hi, Brian</p>\n<p>Do you mind post the code on how to extract features ?</p>\n<p>[quote=Brian Geier;59588]</p>\n<p>To add to the current string.&nbsp;</p>\n<p>Preamble: Cz was filtered to be between [0.1,60] Hz. Kept 250 samples after feedback, feedback time point, feedback vector index.&nbsp;</p>\n<p>A 30% hold out was used to estimate model performance throughout boosting. Subject/Session included gives 0.709 AUC (consistent with public LB) while exclusion (Cz + feedback time point) lowers to 0.62.&nbsp;</p>\n<p>It's also cool to compare the slope during boosting; see graph</p>\n<p>Matlab code:&nbsp;</p>\n<p>ens = fitensemble(X,y,'adaboostm1',500,'tree',...<br>'prior','uniform','type','classification','learnrate',.05,'holdout',.3);</p>\n<p>[/quote]</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "59591",
      "postDate": "12/05/2014 03:09:48",
      "content": "<p>Sure, happy to share. Let me know if you see any errors</p>\n<p>example_load_process.m is a script to do&nbsp;pre-processing, feature matrix build, and then model fit.&nbsp;</p>\n<p>I wrote parse_frame.m to make text parsing easier within matlab; everything goes into a structure array. pullname.m is also required (equivalent to fileparts.m in most cases)</p>\n<p>apologies for not pre-allocating arrays in the driver...</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "59600",
      "postDate": "12/05/2014 05:15:16",
      "content": "<p>Cool, thanks!</p>\n\n<p>[quote=Brian Geier;59591]</p>\n<p>Sure, happy to share. Let me know if you see any errors</p>\n<p>example_load_process.m is a script to do&nbsp;pre-processing, feature matrix build, and then model fit.&nbsp;</p>\n<p>I wrote parse_frame.m to make text parsing easier within matlab; everything goes into a structure array. pullname.m is also required (equivalent to fileparts.m in most cases)</p>\n<p>apologies for not pre-allocating arrays in the driver...</p>\n<p>[/quote]</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "59897",
      "postDate": "12/09/2014 18:29:03",
      "content": "<p>Hmm, isn't it kinda wrong to use session number, feedback timestamp and other such features in the model?</p>\n<p>I believe that the authors of the dataset are interested in a model, which can predict feedback from <em>brain data</em>&nbsp;alone, not from some meta-data...</p>\n<p>If we take away first 4 features from this dataset it will give&nbsp;0.54582 on LB.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "59898",
      "postDate": "12/09/2014 19:36:20",
      "content": "<p>I think you are right about what the authors are interested in. It hasn't been stated clearly enough what is acceptable and unacceptable information to use in this competition. In addition to session number, we can extract whether it is a short or long session(which have different baseline error rates), and even the labels for the authors error detection classifier. This information along with some unprocessed spectral information gets .72 on four fold CV with the same classifier as the benchmark. Specific rules about what we can and can't use would be useful.</p>\n<p>As far as the benchmark, my intention was to point people at least in the right direction for the which time frame of spectral features to use.&nbsp; It was mainly aimed at people who, like me, have no background in this area. Generally when I start a problem I use every feature I can easily extract. In this case some of these are probably not what the data providers are interested in.</p>\n<p>This competition seems to be about transfer learning. I have just started reading the literature and it seems there are a few methods based on modifying spatial filters to transfer across subjects. But over all it seems methods for this type of problem are still in the earlier stages of development within the field of BCI. Likely the authors are hoping people will develop new methods or adapt methods not yet applied within the BCI community.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "59908",
      "postDate": "12/09/2014 21:12:47",
      "content": "<p>@Ilya, I think using session number, trial, and or whether it is a short or long trial are justifiable because they are present during an online spelling task. If spelling accuracy changes with the number of completed trials either due to learning or fatigue this information can be incorporated as a prior. In the same line of thinking if the error is higher when the letters are flashed less times this should be incorporated into the classifier and is not really meta-data.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "59914",
      "postDate": "12/10/2014 00:25:20",
      "content": "<p>Yep, it depends on whether the task is to <em>&quot;predict when error happens&quot;</em> or <em>&quot;predict when there is an&nbsp;error-related potential in the test-subject's brain&quot;</em>. In the first case all means are good, in the second case we should look solely at the brain signal. I personally think that they had the second task in mind, since it is scientifically interesting. Just saying that the authors of the dataset could have been more specific about that.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "59919",
      "postDate": "12/10/2014 01:34:10",
      "content": "<p>[quote=Ilya Kuzovkin;59914]</p>\n<p>Yep, it depends on whether the task is to <em>&quot;predict when error happens&quot;</em> or <em>&quot;predict when there is an&nbsp;error-related potential in the test-subject's brain&quot;</em>.</p>\n<p>[/quote]</p>\n<p>By providing continuous EEG signal rather than short trials after the feedback, i would say that the hosts are not only interested by error related potential detection. The ErrPs, as for other kind of Evoked potential, are well studied in the literature. Scientifically speaking, it is very interesting to understand what affect BCI performances (and therefore, the origin of errors).</p>\n<p>Generally, using meta-features it's not a bad thing, as long as they are &quot;online&quot; compatible. After all, the goal is to improve the error detection, and therefore get a better BCI system. If we have a better detection, no matter how it's done, it will be beneficial for the BCI users.</p>\n<p>Saying that, i don't thing that Session number or trial are a kind of feature that really matters in a real BCI system. The problem here is that the performance criterion is badly designed. The use of a global AUC across subject, especially when the error probability is not the same across all the subjects, makes possible all kind of optimization that are not related to the true subject-specific performance. What i'm saying here is that the Session number indeed helps to achieve better global AUC by optimizing the dynamic of the predictions, but has only a little influence of the single trial detection.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "59937",
      "postDate": "12/10/2014 05:55:50",
      "content": "<p>[quote=Alexandre Barachant;59919]</p>\n<p>By providing continuous EEG signal rather than short trials after the feedback, i would say that the hosts are not only interested by error related potential detection. The ErrPs, as for other kind of Evoked potential, are well studied in the literature. Scientifically speaking, it is very interesting to understand what affect BCI performances (and therefore, the origin of errors).</p>\n<p>Saying that, i don't thing that Session number or trial are a kind of feature that really matters in a real BCI system. The problem here is that the performance criterion is badly designed. The use of a global AUC across subject, especially when the error probability is not the same across all the subjects, makes possible all kind of optimization that are not related to the true subject-specific performance. What i'm saying here is that the Session number indeed helps to achieve better global AUC by optimizing the dynamic of the predictions, but has only a little influence of the single trial detection.</p>\n<p>[/quote]</p>\n\n<p>I agree that use a classifier to judge a classifier is not an&nbsp;elegant&nbsp;choice if we could find the defects&nbsp;of the spelling algorithm directly. But I dont think they intend to give ous ability to rconstruct the whold speller and its algorithm. First there is a 2~4.5sec break between spelling and feedback. Second they have two spelling mode, fast and slow, each uses different time. At last, the most important thing is you dont know the order of letter flashing during the spelling, which is essential to the algorithm.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "59959",
      "postDate": "12/10/2014 17:42:20",
      "content": "<p>@Alexandre, Thanks for pointing this out about global AUC. I think this is also what mlandry was referring to earlier. I am now reading this&nbsp;<a href=\"https://cours.etsmtl.ca/sys828/REFS/A1/Fawcett_PRL2006.pdf\">paper</a> by Tom Fawcet linked in a seizure prediction thread to get a better understanding of ROC and AUC. I am still not quite clear on how seemingly irrelevant information is increasing AUC.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "61558",
      "postDate": "12/19/2014 23:22:25",
      "content": "<p>phalaris, &nbsp;thanks for the starter code. &nbsp;Having this as a foundation saved me tons of time during the initial&nbsp;build of my program. &nbsp;</p>\n<p>I have attached a slightly faster version&nbsp;that has the exact same functionality, but uses numpy arrays during feature extraction. &nbsp;On my computer the runtimes are:</p>\n<p>Original : ~8 minutes</p>\n<p>Numpy_speedup : ~6 minutes</p>\n<p>Nothing ground breaking, but hopefully it might save you some time during your experiments.</p>\n<p>(Tested in Python2.7 and Python3.4 on Ubuntu 14.04)</p>\n\n<p>Edit: Sorry for the multiple attachments. &nbsp;Annoyingly, once you click the attachment, there is no way to remove it, even if you haven't submitted your post yet... &nbsp;Anyways, just pick one, I think they all should work fine.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "61567",
      "postDate": "12/20/2014 10:48:06",
      "content": "<p>[quote=Brandon Veber;61558]</p>\n<p>phalaris, &nbsp;thanks for the starter code. &nbsp;Having this as a foundation saved me tons of time during the initial&nbsp;build of my program. &nbsp;</p>\n<p>I have attached a slightly faster version&nbsp;that has the exact same functionality, but uses numpy arrays during feature extraction. &nbsp;On my computer the runtimes are:</p>\n<p>Original : ~8 minutes</p>\n<p>Numpy_speedup : ~6 minutes</p>\n<p>Nothing ground breaking, but hopefully it might save you some time during your experiments.</p>\n<p>(Tested in Python2.7 and Python3.4 on Ubuntu 14.04)</p>\n<p>Edit: Sorry for the multiple attachments. &nbsp;Annoyingly, once you click the attachment, there is no way to remove it, even if you haven't submitted your post yet... &nbsp;Anyways, just pick one, I think they all should work fine.</p>\n<p>[/quote]</p>\n<p>Thanks phalaris and&nbsp;Brandon!</p>\n<p>Could you please point out which parts of code make it faster, Brandon?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "61568",
      "postDate": "12/20/2014 12:43:27",
      "content": "<p>[quote=tund;61567]</p>\n<p>Thanks phalaris and&nbsp;Brandon!</p>\n<p>Could you please point out which parts of code make it faster, Brandon?</p>\n<p>[/quote]</p>\n<p>In the original code the variables 'train' and 'test' are initialized as Pandas DataFrames. &nbsp;And in the sped up version they are initialized as Numpy arrays. &nbsp;Pandas DataFrames are great, they make it easy to find and save data, but it takes longer to iterate over and add new data to this type of structure. &nbsp;</p>\n<p>After the command &quot;for k in fb.index&quot; (line 38 in speedup version), you can see the chunk of code that loads in the data as a Numpy array, and then inserts it into the 'train' variable. &nbsp;This part of the code, combined with the array initialization (line 25) is what makes it run faster.</p>\n<p>Also, if you compare the two codes you will see some minor differences. &nbsp;I chose to pre-define the electrode name and the length of time after a feedback event at the top of the program (they were hard-coded in as 'Cz' and 260 respectively at a number of locations in the original code). &nbsp;Which, in my opinion, made some of the commands cleaner and more readable. &nbsp;Also,&nbsp;in the new code the title of&nbsp;the output csv files depend&nbsp;on the electrode you have chosen (i.e. 'train_Cz.csv'). &nbsp;I personally like this, but it also means your folder will be full of csv files&nbsp;if you try a bunch of different electrodes.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "61572",
      "postDate": "12/20/2014 13:43:56",
      "content": "<p>@Brandon, Glad you found the code helpful! As you said assigning rows to preallocated DataFrame is really slow. I tried using python dict but there were memory issues sometimes. This is better solution.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "61623",
      "postDate": "12/21/2014 16:38:38",
      "content": "<p>I've tried to run <strong>gbm_benchmark_v2.py</strong> script, and it generates <strong>gbm_benchmark.csv</strong> file, where <strong>Prediction</strong> column contains not 1 and 0, but values like 0.5533333. Is it correct?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "61626",
      "postDate": "12/21/2014 17:50:26",
      "content": "<p>Yes.</p>\n<p><a href=\"http://www.kaggle.com/c/inria-bci-challenge/details/evaluation\">http://www.kaggle.com/c/inria-bci-challenge/details/evaluation</a></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "61627",
      "postDate": "12/21/2014 17:52:15",
      "content": "<p>@Vitalii, Yes, it is correct. Using probabilities instead of discrete label will generally give you a higher AUC score.</p>\n<p>edit: oops piotrek already answered. </p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "61651",
      "postDate": "12/22/2014 08:11:07",
      "content": "<p>[quote=Brandon Veber;61568]</p>\n<p>[quote=tund;61567]</p>\n<p>Thanks phalaris and&nbsp;Brandon!</p>\n<p>Could you please point out which parts of code make it faster, Brandon?</p>\n<p>[/quote]</p>\n<p>In the original code the variables 'train' and 'test' are initialized as Pandas DataFrames. &nbsp;And in the sped up version they are initialized as Numpy arrays. &nbsp;Pandas DataFrames are great, they make it easy to find and save data, but it takes longer to iterate over and add new data to this type of structure. &nbsp;</p>\n<p>After the command &quot;for k in fb.index&quot; (line 38 in speedup version), you can see the chunk of code that loads in the data as a Numpy array, and then inserts it into the 'train' variable. &nbsp;This part of the code, combined with the array initialization (line 25) is what makes it run faster.</p>\n<p>Also, if you compare the two codes you will see some minor differences. &nbsp;I chose to pre-define the electrode name and the length of time after a feedback event at the top of the program (they were hard-coded in as 'Cz' and 260 respectively at a number of locations in the original code). &nbsp;Which, in my opinion, made some of the commands cleaner and more readable. &nbsp;Also,&nbsp;in the new code the title of&nbsp;the output csv files depend&nbsp;on the electrode you have chosen (i.e. 'train_Cz.csv'). &nbsp;I personally like this, but it also means your folder will be full of csv files&nbsp;if you try a bunch of different electrodes.</p>\n<p>[/quote]</p>\n\n<p>cool. Many thanks! Good to learn more about python.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 58973,
      "author_name": "lijojoseph1",
      "author_url": "",
      "post_date": "11/26/2014 05:49:44",
      "content": "<p>i tried your code 'gbm_benchmark.py ' but i am getting the following error&nbsp;</p>\n<p>loading train data<br>Traceback (most recent call last):</p>\n<p><br> File &quot;GbmBenchMark.py&quot;, line 18, in &lt;module&gt;<br> fb = temp.query('FeedBackEvent == 1')['FeedBackEvent']</p>\n<p><br> File &quot;/usr/local/lib/python2.7/dist-packages/pandas/core/frame.py&quot;, line 1822, in query<br> res = self.eval(expr, **kwargs)</p>\n<p><br> File &quot;/usr/local/lib/python2.7/dist-packages/pandas/core/frame.py&quot;, line 1874, in eval<br> return _eval(expr, **kwargs)</p>\n<p><br> File &quot;/usr/local/lib/python2.7/dist-packages/pandas/computation/eval.py&quot;, line 218, in eval<br> _check_engine(engine)</p>\n<p><br> File &quot;/usr/local/lib/python2.7/dist-packages/pandas/computation/eval.py&quot;, line 40, in _check_engine<br> raise ImportError(&quot;'numexpr' not found. Cannot use &quot;</p>\n<p><br>ImportError: 'numexpr' not found. Cannot use engine='numexpr' for query/eval if 'numexpr' is not installed</p>\n<p>please help me to solve this issue</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58974,
      "author_name": "stevendu",
      "author_url": "",
      "post_date": "11/26/2014 06:03:34",
      "content": "<p>[quote=phalaris;58964]</p>\n<p>Here is a simple benchmark using sklearn and pandas. It uses the Cz channel for the 1.3 seconds after feedback for each example as training. Uses sklearn gradient boosting classifier with 500 estimators. It should take about 10 to 15 minutes to run.</p>\n<p>There is no cleaning or processing of the data, and it only uses 1 channel so there is plenty of room to expand. Leaderboard score ~.72</p>\n<p>Its also slow in some places so if you want to extract all of the data after feedbacks you should probably modify the code to use a more efficient extraction method. I used method of creating submission file from Abhishek's code as I never dealt with AUC before.</p>\n<p>Edit: Probably better to set max_features to default which is sqrt(n_features) or just delete that part, the current value was left in by accident, and probably is not great.</p>\n<p>[/quote]</p>\n\n<p>Thanks, so we&nbsp;have 2 benchmark now.</p>\n<p>+ rf</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58975,
      "author_name": "devinanzelmo",
      "author_url": "",
      "post_date": "11/26/2014 06:32:34",
      "content": "<p>@Lijo numexpr is used by pandas query. I attached another file which should fix the problem. let me know if it works for you. I added engine='python' to the query call, but that line can be rewritten if you are still having problems.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58979,
      "author_name": "lijojoseph1",
      "author_url": "",
      "post_date": "11/26/2014 07:32:20",
      "content": "<p>@phalaris the file named&nbsp;gbm_benchmark_v2.py is working.<br>Thanks for your support</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58982,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "11/26/2014 07:56:06",
      "content": "<p>or do this: sudo pip install numexpr</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58983,
      "author_name": "devinanzelmo",
      "author_url": "",
      "post_date": "11/26/2014 08:03:23",
      "content": "<p>That will work. I'm&nbsp; not sure how difficult package installation is on windows or mac so I included the other file. I think anaconda python comes with numexpr.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 59027,
      "author_name": "jfkingiii",
      "author_url": "",
      "post_date": "11/26/2014 19:03:05",
      "content": "<p>LB 0.72545 after changing the max_features back to the default. The introduction of the lag is clearly very important as someone pointed out on another thread. So it seems like adding more channels and maybe multiple lags should give improvement. I'm a bit surprised that this approach of using the raw data works this well, I expected a good result would require constructing some good features.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 59033,
      "author_name": "devinanzelmo",
      "author_url": "",
      "post_date": "11/26/2014 19:36:14",
      "content": "<p>@James King, By lag do you mean starting some point after the onset of the feedback signal to compensate for the time it takes for the brain respond to the signal?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 59040,
      "author_name": "jfkingiii",
      "author_url": "",
      "post_date": "11/26/2014 19:56:03",
      "content": "<p>[quote=phalaris;59033]</p>\n<p>@James King, By lag do you mean starting some point after the onset of the feedback signal to compensate for the time it takes for the brain respond to the signal?</p>\n<p>[/quote]</p>\n<p>Yes; probably I should have said lead. I meant your 261 observations for 1.3 seconds after the feedback; as opposed to looking at the EEG readings only at the moment of the feedback event, when the subject has not had time to react.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 59044,
      "author_name": "devinanzelmo",
      "author_url": "",
      "post_date": "11/26/2014 20:32:47",
      "content": "<p>@James King, thanks for the clarification. 1.3 seconds could be cut down as most of the efforts from the paper for this competition focused on 0-600ms. I think there is a lot of room to improve the score with better features.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 59092,
      "author_name": "lijojoseph1",
      "author_url": "",
      "post_date": "11/27/2014 12:38:36",
      "content": "<p>@phalaris , thanks for your support , i also have a doubt in your code.</p>\n<p>there are other columns rather than 'Cz' but you only used values from that column and you take next 261 data's. what is the logic behind 261? &nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 59096,
      "author_name": "luyitian",
      "author_url": "",
      "post_date": "11/27/2014 14:19:30",
      "content": "<p>[quote=Lijo Joseph;59092]</p>\n<p>@phalaris , thanks for your support , i also have a doubt in your code.</p>\n<p>there are other columns rather than 'Cz' but you only used values from that column and you take next 261 data's. what is the logic behind 261? &nbsp;</p>\n<p>[/quote]</p>\n\n<p>Data is downsampled to 200Hz, which means the time interval is 5ms. So if you want to get data in 1300ms after the event, you will select 1300/5 = 260 items. 260 or 261 doesn't differ a lot.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 59097,
      "author_name": "devinanzelmo",
      "author_url": "",
      "post_date": "11/27/2014 14:36:48",
      "content": "<p>@Lijo, The data is given to us at 200 samples per second , so each row of the original data corresponds to 5ms. In the&nbsp; experiment&nbsp; this data came from an algorithm predicted which letter a person was thinking about, and then showed that letter for 1.3 seconds. We are trying to determine whether the letter shown on the screen is the one the human intended by signals detected with EEG, so I chose data happening when the letter was shown.&nbsp; 260 samples is 5ms*260 = 1300ms =1.3s.&nbsp; Its 261 because I messed up on array indexing, and ended up with an extra sample.</p>\n<p>I only used 'Cz' because it is easy and gives a decent result. You can add extra channels and see if it gives a better result. Looking at couple of the past EEG competitions on Kaggle, it seems most of the effort goes into feature engineering.&nbsp; Take a look at at this&nbsp;<a href=\"https://github.com/MichaelHills/seizure-prediction\">code</a> by Michael Hills for the seizure prediction challenge.&nbsp; Look for transform.py under the seizure-prediction folder. It has many transformations. These can be used to combine channels and extract useful information from channels. It may take a lot of time to find good features</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 59105,
      "author_name": "jfkingiii",
      "author_url": "",
      "post_date": "11/27/2014 17:19:00",
      "content": "<p>I can't cross-validate this method; for example training on subjects</p>\n<p>['02', '06', '16', '18', '20', '21', '24', '26']</p>\n<p>and testing on the others I get a much lower AUC. Possibly I've made some mistake.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 59106,
      "author_name": "devinanzelmo",
      "author_url": "",
      "post_date": "11/27/2014 17:36:06",
      "content": "<p>@James King, you are probably not doing anything wrong. As discussed in another thread the LB is likely based on two test subjects. The <a href=\"http://www.hindawi.com/journals/ahci/2012/578295/\">paper</a> which describes the data for competition indicates that the test subjects can be separated into two clusters based on the specificity of an error detection classifier. There are other ways in which test subjects vary. I think we are seeing two 'easy' subjects on the leaderboard which is way CV scores are lower. For the benchmark I believe CV score is around 0.6. Changes in CV score do reflect changes on the leader board just shifted up because of the leaderboard test subject choice.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 59434,
      "author_name": "bgeier",
      "author_url": "",
      "post_date": "12/03/2014 03:24:20",
      "content": "<p>You've included subject and session in your model as covariates. This doesn't seem appropriate given that the test data is not really connected. Did you consider this when applying the GBM? I've submitted an adaboostm1 using only subject and session as continuous variables and score ~.62. Adding Cz to the mix boosts it to .7; in fact, any feature seems to provide same &quot;added&quot; value. The signal introduced by the subject session label seems artificial to me. Any thoughts?&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 59445,
      "author_name": "devinanzelmo",
      "author_url": "",
      "post_date": "12/03/2014 07:41:17",
      "content": "<p>@Brian, I quickly checked the the contribution of Subject and Session. Taking out Subject decreased my CV score by about .003, and taking out Session decreased it by an additional .018. The increase in score due to including subject is not clear to me, as I am doing CV by subject, it shouldn't change anything.</p>\n<p>Including session information could improve the results if there is either learning over sessions, or alternatively fatigue over sessions. I don't remember the details of experiment but either idea could account for the usefulness of session information. </p>\n<p>In my opinion I don't see a reason not to include both as long as subjects are kept separate during CV</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 59489,
      "author_name": "telser",
      "author_url": "",
      "post_date": "12/03/2014 18:53:45",
      "content": "<p>Do you mean you are doing CV by splitting data for the same subject, or you are assigning subjects to be entirely in training or test? &nbsp;In the latter case, an indicator for subject may be capturing the ungeneralizable idiosyncrasies of that subject that might otherwise be (incorrectly) incorporated into other estimates, or it may be displacing a feature the model is overfitting on in the random variable subset selection.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 59552,
      "author_name": "mlandry",
      "author_url": "",
      "post_date": "12/04/2014 16:22:11",
      "content": "<p>@phalaris, try your CV with more than one subject held out for CV. The reason for doing this is that with one subject held out you are only analyzing&nbsp;your intra-subject AUC.&nbsp;But you also need to check that your inter-subject AUC is working as well. This matters because of the reason you mentioned from the paper--some people make far fewer&nbsp;errors than others, and you need to ensure your calibration picks that up, or at least is robust to it.&nbsp;I think you will see what James King is reporting when you expand your CV, and if you plot the distributions by subject for each CV set you try, you will see why you incur that drop.</p>\n<p>Additionally, when you don't mix subjects, you won't realize the contribution that including the subject ID in your model is making, since it is the same for all your predictions. And it isn't really possible that it's a good contribution unless the ID scheme of the subjects carries some information (which seems unreasonable).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 59555,
      "author_name": "devinanzelmo",
      "author_url": "",
      "post_date": "12/04/2014 17:45:54",
      "content": "<p>@mlandry, thanks for pointing these issues out. I should have clarified more. For the most part I am doing four fold CV. I have also tried leave one out and eight fold CV. I have noticed that the average AUC is different depending on how many subjects I leave out.</p>\n<p>I will try plotting the AUC for each subject based the size of the hold out set.&nbsp; I imagine figuring out which subjects in the test set make fewer errors would be very useful, because we could then train on similar subjects.</p>\n<p>I doubt the subject variable is useful I mainly used it for easily selecting subjects using pandas query(I am new to pandas and currently learning how to reshape the data quickly so I can efficiently try many techniques for generating features).</p>\n<p>Nothing like having ones assumptions pointed out to speed up learning, thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 59562,
      "author_name": "mlandry",
      "author_url": "",
      "post_date": "12/04/2014 18:56:49",
      "content": "<p>@phalaris, my mistake on reading it as if you were leaving one out.</p>\n<p>OK, interesting that you still see it across subjects. I wasn't very clear on the plots--what I found interesting when looking through CV scores was the distribution of the predictions by subject. I was looking at box plots with 0-1 in the y axis and each subject on the X axis. I found intra-subject AUC can vary a lot, but it gets far worse when you mix subjects, if the prediction range doesn't line up very well. Of course the overall prediction range is the same, but if you have most of the predictions for a low-accuracy subject occurring just slightly higher than the predictions for a high-accuracy subject, the AUC score will pay a big price.</p>\n<p>That all said, if you're leaving out 2 and 4 subjects at a time and still seeing results comparable to the leaderboard...I am surprised, as my initial take on this was to agree with @James King. Now it is my turn to thankfully have my assumptions pointed out, and I will have some work to do to see the same progress realized with CV tweaks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 59573,
      "author_name": "emolson",
      "author_url": "",
      "post_date": "12/04/2014 21:25:30",
      "content": "<p>I have been doing leave-one-subject-out cross validation, as&nbsp;I like to have a&nbsp;clear&nbsp;indication of which subjects a method is performing well/poorly on.</p>\n<p>However, I also record the predictions generated for each subject and combine them at the end to calculate a global AUC. When doing this I see a consistent but minor decrease relative to the averaged&nbsp;single subject AUC (typically 0.02).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 59588,
      "author_name": "bgeier",
      "author_url": "",
      "post_date": "12/05/2014 01:50:16",
      "content": "<p>To add to the current string.&nbsp;</p>\n<p>Preamble: Cz was filtered to be between [0.1,60] Hz. Kept 250 samples after feedback, feedback time point, feedback vector index.&nbsp;</p>\n<p>A 30% hold out was used to estimate model performance throughout boosting. Subject/Session included gives 0.709 AUC (consistent with public LB) while exclusion (Cz + feedback time point) lowers to 0.62.&nbsp;</p>\n<p>It's also cool to compare the slope during boosting; see graph</p>\n<p>Matlab code:&nbsp;</p>\n<p>ens = fitensemble(X,y,'adaboostm1',500,'tree',...<br>'prior','uniform','type','classification','learnrate',.05,'holdout',.3);</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 59589,
      "author_name": "stevendu",
      "author_url": "",
      "post_date": "12/05/2014 02:19:21",
      "content": "<p>Hi, Brian</p>\n<p>Do you mind post the code on how to extract features ?</p>\n<p>[quote=Brian Geier;59588]</p>\n<p>To add to the current string.&nbsp;</p>\n<p>Preamble: Cz was filtered to be between [0.1,60] Hz. Kept 250 samples after feedback, feedback time point, feedback vector index.&nbsp;</p>\n<p>A 30% hold out was used to estimate model performance throughout boosting. Subject/Session included gives 0.709 AUC (consistent with public LB) while exclusion (Cz + feedback time point) lowers to 0.62.&nbsp;</p>\n<p>It's also cool to compare the slope during boosting; see graph</p>\n<p>Matlab code:&nbsp;</p>\n<p>ens = fitensemble(X,y,'adaboostm1',500,'tree',...<br>'prior','uniform','type','classification','learnrate',.05,'holdout',.3);</p>\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 59591,
      "author_name": "bgeier",
      "author_url": "",
      "post_date": "12/05/2014 03:09:48",
      "content": "<p>Sure, happy to share. Let me know if you see any errors</p>\n<p>example_load_process.m is a script to do&nbsp;pre-processing, feature matrix build, and then model fit.&nbsp;</p>\n<p>I wrote parse_frame.m to make text parsing easier within matlab; everything goes into a structure array. pullname.m is also required (equivalent to fileparts.m in most cases)</p>\n<p>apologies for not pre-allocating arrays in the driver...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 59600,
      "author_name": "stevendu",
      "author_url": "",
      "post_date": "12/05/2014 05:15:16",
      "content": "<p>Cool, thanks!</p>\n\n<p>[quote=Brian Geier;59591]</p>\n<p>Sure, happy to share. Let me know if you see any errors</p>\n<p>example_load_process.m is a script to do&nbsp;pre-processing, feature matrix build, and then model fit.&nbsp;</p>\n<p>I wrote parse_frame.m to make text parsing easier within matlab; everything goes into a structure array. pullname.m is also required (equivalent to fileparts.m in most cases)</p>\n<p>apologies for not pre-allocating arrays in the driver...</p>\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 59897,
      "author_name": "ilyakuzovkin",
      "author_url": "",
      "post_date": "12/09/2014 18:29:03",
      "content": "<p>Hmm, isn't it kinda wrong to use session number, feedback timestamp and other such features in the model?</p>\n<p>I believe that the authors of the dataset are interested in a model, which can predict feedback from <em>brain data</em>&nbsp;alone, not from some meta-data...</p>\n<p>If we take away first 4 features from this dataset it will give&nbsp;0.54582 on LB.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 59898,
      "author_name": "devinanzelmo",
      "author_url": "",
      "post_date": "12/09/2014 19:36:20",
      "content": "<p>I think you are right about what the authors are interested in. It hasn't been stated clearly enough what is acceptable and unacceptable information to use in this competition. In addition to session number, we can extract whether it is a short or long session(which have different baseline error rates), and even the labels for the authors error detection classifier. This information along with some unprocessed spectral information gets .72 on four fold CV with the same classifier as the benchmark. Specific rules about what we can and can't use would be useful.</p>\n<p>As far as the benchmark, my intention was to point people at least in the right direction for the which time frame of spectral features to use.&nbsp; It was mainly aimed at people who, like me, have no background in this area. Generally when I start a problem I use every feature I can easily extract. In this case some of these are probably not what the data providers are interested in.</p>\n<p>This competition seems to be about transfer learning. I have just started reading the literature and it seems there are a few methods based on modifying spatial filters to transfer across subjects. But over all it seems methods for this type of problem are still in the earlier stages of development within the field of BCI. Likely the authors are hoping people will develop new methods or adapt methods not yet applied within the BCI community.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 59908,
      "author_name": "devinanzelmo",
      "author_url": "",
      "post_date": "12/09/2014 21:12:47",
      "content": "<p>@Ilya, I think using session number, trial, and or whether it is a short or long trial are justifiable because they are present during an online spelling task. If spelling accuracy changes with the number of completed trials either due to learning or fatigue this information can be incorporated as a prior. In the same line of thinking if the error is higher when the letters are flashed less times this should be incorporated into the classifier and is not really meta-data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 59914,
      "author_name": "ilyakuzovkin",
      "author_url": "",
      "post_date": "12/10/2014 00:25:20",
      "content": "<p>Yep, it depends on whether the task is to <em>&quot;predict when error happens&quot;</em> or <em>&quot;predict when there is an&nbsp;error-related potential in the test-subject's brain&quot;</em>. In the first case all means are good, in the second case we should look solely at the brain signal. I personally think that they had the second task in mind, since it is scientifically interesting. Just saying that the authors of the dataset could have been more specific about that.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 59919,
      "author_name": "alexandrebarachant",
      "author_url": "",
      "post_date": "12/10/2014 01:34:10",
      "content": "<p>[quote=Ilya Kuzovkin;59914]</p>\n<p>Yep, it depends on whether the task is to <em>&quot;predict when error happens&quot;</em> or <em>&quot;predict when there is an&nbsp;error-related potential in the test-subject's brain&quot;</em>.</p>\n<p>[/quote]</p>\n<p>By providing continuous EEG signal rather than short trials after the feedback, i would say that the hosts are not only interested by error related potential detection. The ErrPs, as for other kind of Evoked potential, are well studied in the literature. Scientifically speaking, it is very interesting to understand what affect BCI performances (and therefore, the origin of errors).</p>\n<p>Generally, using meta-features it's not a bad thing, as long as they are &quot;online&quot; compatible. After all, the goal is to improve the error detection, and therefore get a better BCI system. If we have a better detection, no matter how it's done, it will be beneficial for the BCI users.</p>\n<p>Saying that, i don't thing that Session number or trial are a kind of feature that really matters in a real BCI system. The problem here is that the performance criterion is badly designed. The use of a global AUC across subject, especially when the error probability is not the same across all the subjects, makes possible all kind of optimization that are not related to the true subject-specific performance. What i'm saying here is that the Session number indeed helps to achieve better global AUC by optimizing the dynamic of the predictions, but has only a little influence of the single trial detection.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 59937,
      "author_name": "luyitian",
      "author_url": "",
      "post_date": "12/10/2014 05:55:50",
      "content": "<p>[quote=Alexandre Barachant;59919]</p>\n<p>By providing continuous EEG signal rather than short trials after the feedback, i would say that the hosts are not only interested by error related potential detection. The ErrPs, as for other kind of Evoked potential, are well studied in the literature. Scientifically speaking, it is very interesting to understand what affect BCI performances (and therefore, the origin of errors).</p>\n<p>Saying that, i don't thing that Session number or trial are a kind of feature that really matters in a real BCI system. The problem here is that the performance criterion is badly designed. The use of a global AUC across subject, especially when the error probability is not the same across all the subjects, makes possible all kind of optimization that are not related to the true subject-specific performance. What i'm saying here is that the Session number indeed helps to achieve better global AUC by optimizing the dynamic of the predictions, but has only a little influence of the single trial detection.</p>\n<p>[/quote]</p>\n\n<p>I agree that use a classifier to judge a classifier is not an&nbsp;elegant&nbsp;choice if we could find the defects&nbsp;of the spelling algorithm directly. But I dont think they intend to give ous ability to rconstruct the whold speller and its algorithm. First there is a 2~4.5sec break between spelling and feedback. Second they have two spelling mode, fast and slow, each uses different time. At last, the most important thing is you dont know the order of letter flashing during the spelling, which is essential to the algorithm.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 59959,
      "author_name": "devinanzelmo",
      "author_url": "",
      "post_date": "12/10/2014 17:42:20",
      "content": "<p>@Alexandre, Thanks for pointing this out about global AUC. I think this is also what mlandry was referring to earlier. I am now reading this&nbsp;<a href=\"https://cours.etsmtl.ca/sys828/REFS/A1/Fawcett_PRL2006.pdf\">paper</a> by Tom Fawcet linked in a seizure prediction thread to get a better understanding of ROC and AUC. I am still not quite clear on how seemingly irrelevant information is increasing AUC.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 61558,
      "author_name": "bveber",
      "author_url": "",
      "post_date": "12/19/2014 23:22:25",
      "content": "<p>phalaris, &nbsp;thanks for the starter code. &nbsp;Having this as a foundation saved me tons of time during the initial&nbsp;build of my program. &nbsp;</p>\n<p>I have attached a slightly faster version&nbsp;that has the exact same functionality, but uses numpy arrays during feature extraction. &nbsp;On my computer the runtimes are:</p>\n<p>Original : ~8 minutes</p>\n<p>Numpy_speedup : ~6 minutes</p>\n<p>Nothing ground breaking, but hopefully it might save you some time during your experiments.</p>\n<p>(Tested in Python2.7 and Python3.4 on Ubuntu 14.04)</p>\n\n<p>Edit: Sorry for the multiple attachments. &nbsp;Annoyingly, once you click the attachment, there is no way to remove it, even if you haven't submitted your post yet... &nbsp;Anyways, just pick one, I think they all should work fine.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 61567,
      "author_name": "tudinhnguyen",
      "author_url": "",
      "post_date": "12/20/2014 10:48:06",
      "content": "<p>[quote=Brandon Veber;61558]</p>\n<p>phalaris, &nbsp;thanks for the starter code. &nbsp;Having this as a foundation saved me tons of time during the initial&nbsp;build of my program. &nbsp;</p>\n<p>I have attached a slightly faster version&nbsp;that has the exact same functionality, but uses numpy arrays during feature extraction. &nbsp;On my computer the runtimes are:</p>\n<p>Original : ~8 minutes</p>\n<p>Numpy_speedup : ~6 minutes</p>\n<p>Nothing ground breaking, but hopefully it might save you some time during your experiments.</p>\n<p>(Tested in Python2.7 and Python3.4 on Ubuntu 14.04)</p>\n<p>Edit: Sorry for the multiple attachments. &nbsp;Annoyingly, once you click the attachment, there is no way to remove it, even if you haven't submitted your post yet... &nbsp;Anyways, just pick one, I think they all should work fine.</p>\n<p>[/quote]</p>\n<p>Thanks phalaris and&nbsp;Brandon!</p>\n<p>Could you please point out which parts of code make it faster, Brandon?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 61568,
      "author_name": "bveber",
      "author_url": "",
      "post_date": "12/20/2014 12:43:27",
      "content": "<p>[quote=tund;61567]</p>\n<p>Thanks phalaris and&nbsp;Brandon!</p>\n<p>Could you please point out which parts of code make it faster, Brandon?</p>\n<p>[/quote]</p>\n<p>In the original code the variables 'train' and 'test' are initialized as Pandas DataFrames. &nbsp;And in the sped up version they are initialized as Numpy arrays. &nbsp;Pandas DataFrames are great, they make it easy to find and save data, but it takes longer to iterate over and add new data to this type of structure. &nbsp;</p>\n<p>After the command &quot;for k in fb.index&quot; (line 38 in speedup version), you can see the chunk of code that loads in the data as a Numpy array, and then inserts it into the 'train' variable. &nbsp;This part of the code, combined with the array initialization (line 25) is what makes it run faster.</p>\n<p>Also, if you compare the two codes you will see some minor differences. &nbsp;I chose to pre-define the electrode name and the length of time after a feedback event at the top of the program (they were hard-coded in as 'Cz' and 260 respectively at a number of locations in the original code). &nbsp;Which, in my opinion, made some of the commands cleaner and more readable. &nbsp;Also,&nbsp;in the new code the title of&nbsp;the output csv files depend&nbsp;on the electrode you have chosen (i.e. 'train_Cz.csv'). &nbsp;I personally like this, but it also means your folder will be full of csv files&nbsp;if you try a bunch of different electrodes.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 61572,
      "author_name": "devinanzelmo",
      "author_url": "",
      "post_date": "12/20/2014 13:43:56",
      "content": "<p>@Brandon, Glad you found the code helpful! As you said assigning rows to preallocated DataFrame is really slow. I tried using python dict but there were memory issues sometimes. This is better solution.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 61623,
      "author_name": "rootua",
      "author_url": "",
      "post_date": "12/21/2014 16:38:38",
      "content": "<p>I've tried to run <strong>gbm_benchmark_v2.py</strong> script, and it generates <strong>gbm_benchmark.csv</strong> file, where <strong>Prediction</strong> column contains not 1 and 0, but values like 0.5533333. Is it correct?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 61626,
      "author_name": "piotrw",
      "author_url": "",
      "post_date": "12/21/2014 17:50:26",
      "content": "<p>Yes.</p>\n<p><a href=\"http://www.kaggle.com/c/inria-bci-challenge/details/evaluation\">http://www.kaggle.com/c/inria-bci-challenge/details/evaluation</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 61627,
      "author_name": "devinanzelmo",
      "author_url": "",
      "post_date": "12/21/2014 17:52:15",
      "content": "<p>@Vitalii, Yes, it is correct. Using probabilities instead of discrete label will generally give you a higher AUC score.</p>\n<p>edit: oops piotrek already answered. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 61651,
      "author_name": "tudinhnguyen",
      "author_url": "",
      "post_date": "12/22/2014 08:11:07",
      "content": "<p>[quote=Brandon Veber;61568]</p>\n<p>[quote=tund;61567]</p>\n<p>Thanks phalaris and&nbsp;Brandon!</p>\n<p>Could you please point out which parts of code make it faster, Brandon?</p>\n<p>[/quote]</p>\n<p>In the original code the variables 'train' and 'test' are initialized as Pandas DataFrames. &nbsp;And in the sped up version they are initialized as Numpy arrays. &nbsp;Pandas DataFrames are great, they make it easy to find and save data, but it takes longer to iterate over and add new data to this type of structure. &nbsp;</p>\n<p>After the command &quot;for k in fb.index&quot; (line 38 in speedup version), you can see the chunk of code that loads in the data as a Numpy array, and then inserts it into the 'train' variable. &nbsp;This part of the code, combined with the array initialization (line 25) is what makes it run faster.</p>\n<p>Also, if you compare the two codes you will see some minor differences. &nbsp;I chose to pre-define the electrode name and the length of time after a feedback event at the top of the program (they were hard-coded in as 'Cz' and 260 respectively at a number of locations in the original code). &nbsp;Which, in my opinion, made some of the commands cleaner and more readable. &nbsp;Also,&nbsp;in the new code the title of&nbsp;the output csv files depend&nbsp;on the electrode you have chosen (i.e. 'train_Cz.csv'). &nbsp;I personally like this, but it also means your folder will be full of csv files&nbsp;if you try a bunch of different electrodes.</p>\n<p>[/quote]</p>\n\n<p>cool. Many thanks! Good to learn more about python.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "58964": "",
    "58973": "",
    "58974": "",
    "58975": "",
    "58979": "",
    "58982": "",
    "58983": "",
    "59027": "",
    "59033": "",
    "59040": "",
    "59044": "",
    "59092": "",
    "59096": "",
    "59097": "",
    "59105": "",
    "59106": "",
    "59434": "",
    "59445": "",
    "59489": "",
    "59552": "",
    "59555": "",
    "59562": "",
    "59573": "",
    "59588": "",
    "59589": "",
    "59591": "",
    "59600": "",
    "59897": "",
    "59898": "",
    "59908": "",
    "59914": "",
    "59919": "",
    "59937": "",
    "59959": "",
    "61558": "",
    "61567": "",
    "61568": "",
    "61572": "",
    "61623": "",
    "61626": "",
    "61627": "",
    "61651": ""
  },
  "source": "meta"
}