{
  "id": 12635,
  "title": "Learning from this competition",
  "url": "/competitions/inria-bci-challenge/discussion/12635",
  "author_name": "",
  "post_date": "2015-02-28T04:38:34.013Z",
  "votes": null,
  "comment_count": 5,
  "views": 1940,
  "content": "<p>Hi guys,</p>\n<p>It was a great experience participating in this competition, and surely a lot to &nbsp;learn reading through the&nbsp;forum-posts. Considering it was my first participation, it was a bit disappointing in the end to see a drop of 111 in my leader board score. Thus, I would appreciate certain insights as to how I should approach future competitions by pointing out errors I made in the current one. &nbsp;</p>\n<p>My approach was(in python):</p>\n<ol>\n<li>Extracting Epochs of raw data as given by phalaris's benchmark code</li>\n<li>Channel Selection - I found <em>channel 46</em> to give the best boost in cross validation. Any other individual or combination of channels were detrimental</li>\n<li>Applying a <em>Butterworth Filter with cutoff of 0.1 to 15 Hz and order 6</em>.<span style=\"line-height: 1.4\">I tested a combination of features</span></li>\n</ol>\n\n<ul>\n<ul style=\"line-height: 1.4\">\n<li><span style=\"line-height: 1.4\">Time Series</span></li>\n<li><span style=\"line-height: 1.4\">Hjorth Params</span></li>\n<li><span style=\"line-height: 1.4\">FFT</span></li>\n<li><span style=\"line-height: 1.4\">Time Series Differential</span></li>\n<li><span style=\"line-height: 1.4\">Subject and Trial related Info</span></li>\n<li><span style=\"line-height: 1.4\">Skewness/Kurtosis and basic statistical features</span></li>\n<li><span style=\"line-height: 1.4\"><span style=\"line-height: 1.4\">Wavelet Transform</span></span></li>\n<li><span style=\"line-height: 1.4\"><span style=\"line-height: 1.4\">EOG values during trial-downsampled</span></span></li>\n</ul>\n</ul>\n<p>Finally I found that using Time Series, Subject Trial and EOG info gave best cross validation results<br>4. Classifiers: I tried RFs,GBM,LDA,SVMs. Found GBM to give best values</p>\n<p>CV method : I had posted a question about the same and according to the responses followed<em> Leave One Subject Out CV</em> over the training subjects</p>\n<p>However my model quite obviously badly overfit the data. I am still unable to understand where exactly my method failed? Any other suggestions as to how I should approach future competitions?</p>",
  "messages": [
    {
      "id": "65103",
      "postDate": "02/28/2015 04:38:34",
      "content": "<p>Hi guys,</p>\n<p>It was a great experience participating in this competition, and surely a lot to &nbsp;learn reading through the&nbsp;forum-posts. Considering it was my first participation, it was a bit disappointing in the end to see a drop of 111 in my leader board score. Thus, I would appreciate certain insights as to how I should approach future competitions by pointing out errors I made in the current one. &nbsp;</p>\n<p>My approach was(in python):</p>\n<ol>\n<li>Extracting Epochs of raw data as given by phalaris's benchmark code</li>\n<li>Channel Selection - I found <em>channel 46</em> to give the best boost in cross validation. Any other individual or combination of channels were detrimental</li>\n<li>Applying a <em>Butterworth Filter with cutoff of 0.1 to 15 Hz and order 6</em>.<span style=\"line-height: 1.4\">I tested a combination of features</span></li>\n</ol>\n\n<ul>\n<ul style=\"line-height: 1.4\">\n<li><span style=\"line-height: 1.4\">Time Series</span></li>\n<li><span style=\"line-height: 1.4\">Hjorth Params</span></li>\n<li><span style=\"line-height: 1.4\">FFT</span></li>\n<li><span style=\"line-height: 1.4\">Time Series Differential</span></li>\n<li><span style=\"line-height: 1.4\">Subject and Trial related Info</span></li>\n<li><span style=\"line-height: 1.4\">Skewness/Kurtosis and basic statistical features</span></li>\n<li><span style=\"line-height: 1.4\"><span style=\"line-height: 1.4\">Wavelet Transform</span></span></li>\n<li><span style=\"line-height: 1.4\"><span style=\"line-height: 1.4\">EOG values during trial-downsampled</span></span></li>\n</ul>\n</ul>\n<p>Finally I found that using Time Series, Subject Trial and EOG info gave best cross validation results<br>4. Classifiers: I tried RFs,GBM,LDA,SVMs. Found GBM to give best values</p>\n<p>CV method : I had posted a question about the same and according to the responses followed<em> Leave One Subject Out CV</em> over the training subjects</p>\n<p>However my model quite obviously badly overfit the data. I am still unable to understand where exactly my method failed? Any other suggestions as to how I should approach future competitions?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65111",
      "postDate": "02/28/2015 06:05:42",
      "content": "<p>This competition was really challenging. Yann LeCun once said&nbsp;problems where you have&nbsp;few training examples and lots of&nbsp;features is&nbsp;machine learning hell. I think this problem falls in that space. With so many features (i.e. so many channels with multiple time points), it is very&nbsp;easy to overfit.</p>\n<p>I think you used the right cross-validation method, but you might not be aware that you can actually also overfit to your cross-validation performance. I would guess that overfitting to your CV set&nbsp;is what happened&nbsp;for your solution. By&nbsp;trying lots of models and features and picking&nbsp;the one combination that performed best on your CV, you likely overfit to your&nbsp;CV performance. Imagine that you had a thousand EEG channels&nbsp;or even an infinite number. If you were to pick the EEG channel that performed the best on your CV set, you would likely be&nbsp;picking&nbsp;a channel that performed best just by random chance. I hope that makes sense. Trying to pick the &quot;best&quot; feature through CV probably led you towards overfitting. I would guess that if you had picked the &quot;best&quot; 5 features through CV&nbsp;and then averaged the predictions, you would have done better.</p>\n<p>I think for this type of problem, where you have many features for relatively few&nbsp;examples,&nbsp;there are three pretty&nbsp;useful methods:</p>\n<ol>\n<li><span style=\"line-height: 1.4\">Ensembling. Ensembling multiple models can help you reduce&nbsp;variance/overfitting. So instead of picking the best single model or feature, choose&nbsp;several good models or features and combine your prediction.</span></li>\n<li><span style=\"line-height: 1.4\">Data Reduction. Try reducing the number of features you are examining through some kind of data reduction. I used a stacked unsupervised auto encoder, and it seemed to help&nbsp;with essentially no&nbsp;hyper parameter tuning. PCA or downsampling are other options.&nbsp;By reducing the dimension of the feature space, you are reducing the&nbsp;potential for&nbsp;your model to overfit.</span></li>\n<li><span style=\"line-height: 1.4\">Get more data :-).</span></li>\n</ol>\n<p>Finally, there is a small chance that your model actually wouldn't do that badly with a larger&nbsp;test set. I&nbsp;am still ambivalent about&nbsp;whether I actually did well (7th in private leaderboard without using the data leak) or if I just got lucky due to the small test set :-). Also, keep in mind that your general method of picking the best model and best channel would have been more successful&nbsp;with more data.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65122",
      "postDate": "02/28/2015 11:34:54",
      "content": "<p>Hey,</p>\n<p>I agree with Daniel Yoo, good points. But I think that LOSO CV was quite easy to overfit to and wasn't a very reliable measure. When you are doing LOSO CV you get 16 AUCs that you average or you can decide to concatenate all predictions and calculate a single AUC. In the former case you get a score with a relatively high variance, due to high inter-subject variations. Each AUC is, obviously, a measure of performance on a single subject, whereas the leaderboard measures joint performance on a set of subjects - these AUCs can differ greatly. In the latter case you are actually calculating an AUC from concatenated predictions of 16 different models, this can be very unstable and is dissimilar to how you are preparing your submissions for the LB.</p>\n<p>In my opinion a more robust method of testing stuff and locally measuring your performance was K-fold cross-validation with data randomly divided into folds <strong>subject-wise</strong>. The whole procedure should be repeated several times and obtained AUCs averaged to get a single measure of performance. We used 4-fold CV, 10 repetitions for most tests (score is then an average of 40 AUCs), and 100 repetitions if we wanted to be very confident with certain decisions. The K-fold setting is also more challenging for the model, due to less training data available.</p>\n<p>Another thing, which I find very important, is to validate your methodology of making crucial decisions on the basis of data. @Tangy, f.eg. you said that you selected a channel on the basis of CV procedure. What I believe would have revealed your overfit (and subsequently guide you to better solutions) was to split the data into 2 sets subject-wise, let's say setA and setB. Apply your procedure of channel selection on setA (i.e. do CV on setA to determine the best channel), and test the performance with obtained channel on setB. You could repeat this many times to get a feeling of how good that methodology is. I'm writing specifically about channel selection, because in our case this methodology revealed that adding top1-5 channels or removing worst1-5 channels on the basis of CV was gambling, i.e. there was ~50% chance that a new set of channels &quot;suggested&quot; by training subjects would actually decrease the AUC on validation subjects.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65124",
      "postDate": "02/28/2015 11:54:15",
      "content": "<p>I'm not totally convinced leave one subject out CV was the best approach with this competition. There appear to be two ways of using LOO CV. You can calculate a separate AUC score for &nbsp;each subject and take the average (i.e. average of 16 scores), However, this will certainly give you misleading results as, for this competition, the AUC is calculated in one big group. Because the ordering of predictions is important for the AUC score, the average of 16 scores won't&nbsp;give you the same result as the 'global' AUC unless you calibrate your probabilities. <a href=\"https://cours.etsmtl.ca/sys828/REFS/A1/Fawcett_PRL2006.pdf\">This paper</a>&nbsp;which has done the rounds on various Kaggle forums explains this well.</p>\n<p>The second LOO option is to obtain posterior probabilities for each subject and then calculate the 'global' AUC for the group. As Daniel points out you can overfit to your CV score with this approach. Also, I think because you are using slightly different models to predict each&nbsp;CV subject (i.e. the subjects&nbsp;which make up each&nbsp;model&nbsp;will be slightly different each time) you should also be careful that the posterior probabilities are calibrated otherwise you may fall foul of the same problem as outlined above.&nbsp;</p>\n<p>To circumvent these issues I went for 4 fold CV&nbsp;and calculated the AUC for the 4 subjects in the test fold only. This gives&nbsp;4 AUC scores (16 subjects divided by 4 folds=4). Because you're not blending probabilities from different models (as in LOO option 2) there isn't the danger of &nbsp;uncalibrated probabilities in the AUC calculation. Furthermore, to avoid the danger of over fitting your CV score, you can calculate the AUC score for several different splits,&nbsp;not just one. I calculated AUC scores for 5 different splits giving a total of 20 AUC scores (5X4) and then took the average of these. I also made note of the standard deviation of the AUC sores and used it as a proxy for model stability. For my final model the CV score was ~0.75 (it also had the lowest standard deviation), the public leaderboard score was ~0.77 and private ~0.77 as well. The public and private scores are very close which I guess could just be luck. I put the different between these scores and my CV score down to the fact that in the CV procedure you use less data and less data=poorer results.&nbsp;</p>\n\n<p>Edit: It looks like Blaine posted his response as I was typing mine and makes many of the same points. Apologies for the repetition.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65139",
      "postDate": "02/28/2015 17:16:49",
      "content": "<p>@blaine and @barrack_d,&nbsp;that's really&nbsp;important about the across-subject AUC. I hadn't realized that the AUC was calculated across the board! Nice explanations. Thanks!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65240",
      "postDate": "03/02/2015 14:40:08",
      "content": "<p>[quote=barrack_d;65124]</p>\n<p>Edit: It looks like Blaine posted his response as I was typing mine and makes many of the same points. Apologies for the repetition.</p>\n<p>[/quote]</p>\n<p>I think that through this coincidence our reasoning was validated, it's a good thing :-)</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 65111,
      "author_name": "danielyoo",
      "author_url": "",
      "post_date": "02/28/2015 06:05:42",
      "content": "<p>This competition was really challenging. Yann LeCun once said&nbsp;problems where you have&nbsp;few training examples and lots of&nbsp;features is&nbsp;machine learning hell. I think this problem falls in that space. With so many features (i.e. so many channels with multiple time points), it is very&nbsp;easy to overfit.</p>\n<p>I think you used the right cross-validation method, but you might not be aware that you can actually also overfit to your cross-validation performance. I would guess that overfitting to your CV set&nbsp;is what happened&nbsp;for your solution. By&nbsp;trying lots of models and features and picking&nbsp;the one combination that performed best on your CV, you likely overfit to your&nbsp;CV performance. Imagine that you had a thousand EEG channels&nbsp;or even an infinite number. If you were to pick the EEG channel that performed the best on your CV set, you would likely be&nbsp;picking&nbsp;a channel that performed best just by random chance. I hope that makes sense. Trying to pick the &quot;best&quot; feature through CV probably led you towards overfitting. I would guess that if you had picked the &quot;best&quot; 5 features through CV&nbsp;and then averaged the predictions, you would have done better.</p>\n<p>I think for this type of problem, where you have many features for relatively few&nbsp;examples,&nbsp;there are three pretty&nbsp;useful methods:</p>\n<ol>\n<li><span style=\"line-height: 1.4\">Ensembling. Ensembling multiple models can help you reduce&nbsp;variance/overfitting. So instead of picking the best single model or feature, choose&nbsp;several good models or features and combine your prediction.</span></li>\n<li><span style=\"line-height: 1.4\">Data Reduction. Try reducing the number of features you are examining through some kind of data reduction. I used a stacked unsupervised auto encoder, and it seemed to help&nbsp;with essentially no&nbsp;hyper parameter tuning. PCA or downsampling are other options.&nbsp;By reducing the dimension of the feature space, you are reducing the&nbsp;potential for&nbsp;your model to overfit.</span></li>\n<li><span style=\"line-height: 1.4\">Get more data :-).</span></li>\n</ol>\n<p>Finally, there is a small chance that your model actually wouldn't do that badly with a larger&nbsp;test set. I&nbsp;am still ambivalent about&nbsp;whether I actually did well (7th in private leaderboard without using the data leak) or if I just got lucky due to the small test set :-). Also, keep in mind that your general method of picking the best model and best channel would have been more successful&nbsp;with more data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65122,
      "author_name": "rafalcycon",
      "author_url": "",
      "post_date": "02/28/2015 11:34:54",
      "content": "<p>Hey,</p>\n<p>I agree with Daniel Yoo, good points. But I think that LOSO CV was quite easy to overfit to and wasn't a very reliable measure. When you are doing LOSO CV you get 16 AUCs that you average or you can decide to concatenate all predictions and calculate a single AUC. In the former case you get a score with a relatively high variance, due to high inter-subject variations. Each AUC is, obviously, a measure of performance on a single subject, whereas the leaderboard measures joint performance on a set of subjects - these AUCs can differ greatly. In the latter case you are actually calculating an AUC from concatenated predictions of 16 different models, this can be very unstable and is dissimilar to how you are preparing your submissions for the LB.</p>\n<p>In my opinion a more robust method of testing stuff and locally measuring your performance was K-fold cross-validation with data randomly divided into folds <strong>subject-wise</strong>. The whole procedure should be repeated several times and obtained AUCs averaged to get a single measure of performance. We used 4-fold CV, 10 repetitions for most tests (score is then an average of 40 AUCs), and 100 repetitions if we wanted to be very confident with certain decisions. The K-fold setting is also more challenging for the model, due to less training data available.</p>\n<p>Another thing, which I find very important, is to validate your methodology of making crucial decisions on the basis of data. @Tangy, f.eg. you said that you selected a channel on the basis of CV procedure. What I believe would have revealed your overfit (and subsequently guide you to better solutions) was to split the data into 2 sets subject-wise, let's say setA and setB. Apply your procedure of channel selection on setA (i.e. do CV on setA to determine the best channel), and test the performance with obtained channel on setB. You could repeat this many times to get a feeling of how good that methodology is. I'm writing specifically about channel selection, because in our case this methodology revealed that adding top1-5 channels or removing worst1-5 channels on the basis of CV was gambling, i.e. there was ~50% chance that a new set of channels &quot;suggested&quot; by training subjects would actually decrease the AUC on validation subjects.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65124,
      "author_name": "duncanbarrack",
      "author_url": "",
      "post_date": "02/28/2015 11:54:15",
      "content": "<p>I'm not totally convinced leave one subject out CV was the best approach with this competition. There appear to be two ways of using LOO CV. You can calculate a separate AUC score for &nbsp;each subject and take the average (i.e. average of 16 scores), However, this will certainly give you misleading results as, for this competition, the AUC is calculated in one big group. Because the ordering of predictions is important for the AUC score, the average of 16 scores won't&nbsp;give you the same result as the 'global' AUC unless you calibrate your probabilities. <a href=\"https://cours.etsmtl.ca/sys828/REFS/A1/Fawcett_PRL2006.pdf\">This paper</a>&nbsp;which has done the rounds on various Kaggle forums explains this well.</p>\n<p>The second LOO option is to obtain posterior probabilities for each subject and then calculate the 'global' AUC for the group. As Daniel points out you can overfit to your CV score with this approach. Also, I think because you are using slightly different models to predict each&nbsp;CV subject (i.e. the subjects&nbsp;which make up each&nbsp;model&nbsp;will be slightly different each time) you should also be careful that the posterior probabilities are calibrated otherwise you may fall foul of the same problem as outlined above.&nbsp;</p>\n<p>To circumvent these issues I went for 4 fold CV&nbsp;and calculated the AUC for the 4 subjects in the test fold only. This gives&nbsp;4 AUC scores (16 subjects divided by 4 folds=4). Because you're not blending probabilities from different models (as in LOO option 2) there isn't the danger of &nbsp;uncalibrated probabilities in the AUC calculation. Furthermore, to avoid the danger of over fitting your CV score, you can calculate the AUC score for several different splits,&nbsp;not just one. I calculated AUC scores for 5 different splits giving a total of 20 AUC scores (5X4) and then took the average of these. I also made note of the standard deviation of the AUC sores and used it as a proxy for model stability. For my final model the CV score was ~0.75 (it also had the lowest standard deviation), the public leaderboard score was ~0.77 and private ~0.77 as well. The public and private scores are very close which I guess could just be luck. I put the different between these scores and my CV score down to the fact that in the CV procedure you use less data and less data=poorer results.&nbsp;</p>\n\n<p>Edit: It looks like Blaine posted his response as I was typing mine and makes many of the same points. Apologies for the repetition.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65139,
      "author_name": "danielyoo",
      "author_url": "",
      "post_date": "02/28/2015 17:16:49",
      "content": "<p>@blaine and @barrack_d,&nbsp;that's really&nbsp;important about the across-subject AUC. I hadn't realized that the AUC was calculated across the board! Nice explanations. Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65240,
      "author_name": "rafalcycon",
      "author_url": "",
      "post_date": "03/02/2015 14:40:08",
      "content": "<p>[quote=barrack_d;65124]</p>\n<p>Edit: It looks like Blaine posted his response as I was typing mine and makes many of the same points. Apologies for the repetition.</p>\n<p>[/quote]</p>\n<p>I think that through this coincidence our reasoning was validated, it's a good thing :-)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "65103": "",
    "65111": "",
    "65122": "",
    "65124": "",
    "65139": "",
    "65240": ""
  },
  "source": "meta"
}