{
  "id": 10945,
  "title": "Congratulations to the winners!",
  "url": "/competitions/seizure-prediction/discussion/10945",
  "author_name": "",
  "post_date": "2014-11-18T01:00:26.133Z",
  "votes": 6,
  "comment_count": 91,
  "views": 20404,
  "content": "<p>This was a tough competition on a lot of levels. Hats off to the winners.</p>\n<p>I'm looking forward to hearing about the different approaches that were taken.</p>",
  "messages": [
    {
      "id": "58249",
      "postDate": "11/18/2014 01:00:26",
      "content": "<p>This was a tough competition on a lot of levels. Hats off to the winners.</p>\n<p>I'm looking forward to hearing about the different approaches that were taken.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58250",
      "postDate": "11/18/2014 01:05:58",
      "content": "<p>I was also very impressed by Jonathan Tapson taking the lead with&nbsp;so few submissions earlier on in the competition&nbsp;and&nbsp;similarly impressed by Medrr's jump to 0.90!&nbsp;An amazing effort.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58251",
      "postDate": "11/18/2014 01:22:33",
      "content": "<p>This was my first ML implementation and I found the discourse on the forums highly stimulating. Thank to everyone who partook!</p>\n<p>I'm looking forward to dissecting the winning models.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58252",
      "postDate": "11/18/2014 01:23:16",
      "content": "<p>Yeah congrats everyone especially the winners, that was a hell of a challenge.</p>\n<p>I can't wait to find out what strategy everybody else used, would be really interested to know the subject AUC breakdown people were getting too.</p>\n<p>For pretty much every feature we tried the t-SNE plot of Patient 2 showed the test and training data as totally different! &nbsp;Our best subject specific classifier for Patient 2 tended be those trained on the noisiest untransformed features or the least cleaned datasets which kind of makes me think they weren't actually working but the noise meant we&nbsp;weren't as misleadingly training that model.</p>\n<p>Ah well, it was great fun, we are over the moon with our result&nbsp;as it stands!</p>\n\n<p>(We were only 20th but if anyone is interested&nbsp;we&nbsp;used a standard SVC with RBF kernel with random forest&nbsp;feature selection on a feature set composed of cleaned (harmonics and high-pass filter) Common Spatial Patterns basis transform correlation coefficient eigenvalues, and cleaned Independent Component Analysis transformed&nbsp;(FastICA):</p>\n<p>Power spectral density logf correlation&nbsp;coefficients<br>Low-gamma phase sync<br>Power in band<br>Power spectral density logf Broad Band<br>Partial Directed Coherence of the coefficients for a&nbsp;Multivariate Autoregression model&nbsp;&nbsp;<br>We also did some meta-bagging of this classifier and a couple of others with slightly different feature sets.)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58255",
      "postDate": "11/18/2014 02:07:50",
      "content": "<p>That was fun but very frustrating towards the end.&nbsp; For what it's worth, I used only spectral features (fft over 30s and 60s intervals, overlapped 50%).&nbsp; The secret weapon was linear regression - you could get to about 0.85 on the leaderboard with straight LR, post scaled through a logistic function to (0,1) interval.&nbsp; No networks, trees, SVMs, RBMs, etc. required.&nbsp; Because LR is superfast and can be inverted (i.e. you can back-process to see what features it is scaling up, and what features it ignores)&nbsp; I found some good feature sets.&nbsp; It is also very hard to overtrain.&nbsp; The general best featuresets were 1Hz bands from 0-50 Hz and then 5-10Hz bands up to 180Hz.&nbsp; All data was filtered for 60Hz + harmonics.&nbsp; I did use some networks (ELM ensembles) to get a couple of extra points.&nbsp; The final solution was computed on my laptop&nbsp; (a 2010 MacBook Air).&nbsp; No AWS required.&nbsp; Now I just hope I can reproduce the damn result for the organizers!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58259",
      "postDate": "11/18/2014 02:38:07",
      "content": "<p>[quote=Jonathan Tapson;58255]</p>\n<p>The final solution was computed on my laptop&nbsp; (a 2010 MacBook Air).&nbsp; No AWS required.&nbsp; Now I just hope I can reproduce the damn result for the organizers!</p>\n<p>[/quote]</p>\n<p>This makes the best commercial of Mac Air :D</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58260",
      "postDate": "11/18/2014 02:48:01",
      "content": "<p>It sounds crazy but actually the Mac Air has a solid state drive, and in these computations with big data sets sometimes disk read/write is what dominates the processing time.&nbsp; Towards the end the feature sets were taking a couple of hours to compute though, it was dumb to carry on with the Air at that point.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58261",
      "postDate": "11/18/2014 02:49:05",
      "content": "<p>Looks like simplicity is key. :) I was getting stressed out&nbsp;in this competition with how much of a complete mess my solution had become!&nbsp;Can you elaborate on your linear regression with logistic function setup?&nbsp;What did you use for training labels?</p>\n<p>Late in the competition I realised that my main problem was properly tuning SVM parameters to work with my features. Using one set of C/gamma I could use random feature selection (e.g. randomly mask 50%) and then ensemble the results of 10 such masks to see an improvement (from ~0.796 to ~0.829 leaderboard). Further improvement could be made using genetic algorithm to do that selection (0.84+), but results varied wildly with minor changes in runs and was highly unstable. A while later I discovered that I could match the random mask ensembling by using a different set of C/gamma and just using the features in whole. With those parameters ensembling random masks made it worse.</p>\n<p>Then on the last day I tried Logistic Regression for Dog_5 instead of SVM and got another huge improvement, but we aren't allowed &quot;if Dog_5&quot; so couldn't use it.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58263",
      "postDate": "11/18/2014 03:00:00",
      "content": "<p>I used 0's for interictal and 1's for preictal and those were the target values for the LR; the LR weights were computed using a regularized pseudoinverse.&nbsp; The test set values were then normalised by subtracting the mean and dividing by standard deviation (for the whole test set per subject).&nbsp; That gave a set of values with mean 0, so then just put the values into 1/(1+e^-k.values)&nbsp; where k was a scaling factor (k=0.5, and I can't remember why I used that, probably legacy code).&nbsp; That gave values compressed between 0 and 1 and was pretty close to the normalization methods recommended in a couple of the papers that were posted somewhere on the forum.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58264",
      "postDate": "11/18/2014 03:06:12",
      "content": "<p>So to let me check I understand, if I wanted to implement this I would do the following:</p>\n<ol>\n<li>Train linear regression (with feature scaling or without?) with target labels 0 and 1</li>\n<li>Make predictions on cv or test set</li>\n<li>For the array of predictions, subtract mean and divide by standard deviation</li>\n<li>Transform predictions with logistic function</li>\n</ol>\n<p>Does that sound right?</p>\n<p>Edit: I attempted this with my existing features and scored&nbsp;0.83618 on public leaderboard. Wow that beat all my whole feature no ensembling&nbsp;attempts with SVM etc. Then&nbsp;0.84208 if average reference montage is applied to Patient 1 and 2.</p>\n<p>Interestingly the distribution of predictions is completely different to my other submissions yet scores similarly. I wrote a script that counts how many predictions are &lt; 0.5 and how many &gt;1.0 so I could see what is happening in my test set predictions. The test % is how much of the test set was predicted preictal, and train % is how much of the training set consisted of preictal as a rough guide for what I might expect the data split to be like. E.g. in the first one, Dog_1 has 247 interictal predictions and 255 preictal predictions.</p>\n<p>From linear regression run:</p>\n<p>Dog_1 [247, 255] test 50.8% train 4.8%<br>Dog_2 [514, 486] test 48.6% train 7.7%<br>Dog_3 [483, 424] test 46.7% train 4.8%<br>Dog_4 [578, 412] test 41.6% train 10.8%<br>Dog_5 [150, 41] test 21.5% train 6.2%<br>Patient_1 [107, 88] test 45.1% train 26.5%<br>Patient_2 [51, 99] test 66.0% train 30.0%</p>\n<p>From SVM with genetic algorithm feature selection:</p>\n<p>Dog_1 [481, 21] test 4.2% train 4.8%<br>Dog_2 [923, 77] test 7.7% train 7.7%<br>Dog_3 [835, 72] test 7.9% train 4.8%<br>Dog_4 [930, 60] test 6.1% train 10.8%<br>Dog_5 [161, 30] test 15.7% train 6.2%<br>Patient_1 [92, 103] test 52.8% train 26.5%<br>Patient_2 [135, 15] test 10.0% train 30.0%</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58268",
      "postDate": "11/18/2014 03:49:26",
      "content": "<p>Great work Jonathan, I'm absolutely kicking myself for not trying something as simple as linear regression.</p>\n\n<p>looking forward to the write up and code.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58281",
      "postDate": "11/18/2014 06:07:34",
      "content": "<p>Congratulation to all and especially to the winners. The competition is harder than I expected. I started the competition a bit late, made some push in the final two weeks, touched the LB top ten in the final 24 hours. I was struggling in selecting a classifier that can balance bias and variance. I chose SVM for it has some control in setting&nbsp; c and gamma. But I didn't have time to fine tune c and gamma for individual subject.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58282",
      "postDate": "11/18/2014 06:54:03",
      "content": "<p>@Michael Hills - yes, as you describe. &nbsp;I generally calculated a separate regression on each permutation (fold) of interictal/preictal data and then summed the test set predictions generated by those weights - I also just summed the predictions for each 30s/60s segment to get the outcome for a 10 minute file (which may have been your method from the previous competition?) &nbsp;it is essentially a simplistic logistic regression. &nbsp;I wasn't planning to do much more than some basic feature selection with it, but it turned out to be a good match for this weird situation with very skew training sets and no knowledge of the test set stats (priors). &nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58283",
      "postDate": "11/18/2014 07:06:09",
      "content": "<p>[quote=Michael Hills;58264]</p>\n<p>Interestingly the distribution of predictions is completely different to my other submissions yet scores similarly. I wrote a script that counts how many predictions are &lt; 0.5 and how many &gt;1.0 so I could see what is happening in my test set predictions. The test % is how much of the test set was predicted preictal, and train % is how much of the training set consisted of preictal as a rough guide for what I might expect the data split to be like. E.g. in the first one, Dog_1 has 247 interictal predictions and 255 preictal predictions.</p>\n<p>[/quote]</p>\n<p>As I understand the AUC score doesn't care too much (at all) about the absolute values of predictions (as long as all of them are calibrated across the targets).&nbsp; So, &lt;0.5 and &gt;0.5 split doesn't make too much sense.</p>\n<p>I think it did work for SVM just because the skilearn's SVM implementation has the built-in Platt calibration (when called with the probability=True) which pushes the predicted probabilities to absolute values of 0 and 1. On the other hand the LR implementation doesn't perform the calibration and the optimal cut-off value (threshold) is not 0.5.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58286",
      "postDate": "11/18/2014 07:50:14",
      "content": "<p>[quote=Sergey Korotkov;58283]</p>\n<p>As I understand the AUC score doesn't care too much (at all) about the absolute values of predictions (as long as all of them are calibrated across the targets).&nbsp; So, &lt;0.5 and &gt;0.5 split doesn't make too much sense.[/quote]</p>\n<p>Indeed, only rank counts. Essentially, the competition was about predicting not probability but ranking. I.e. predicting fewer incorrectly swapped pairs 0/1 as possible.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58288",
      "postDate": "11/18/2014 08:47:53",
      "content": "<p>I don't have a full handle on ROC AUC yet,&nbsp;so let me see if I've got this right. As long as every prediction for preictal has a higher value than every prediction for interictal, you will score 1.0? I.e all about ranking as you say. I remember reading something like that one way to think of ROC AUC was, given a random preictal/interictal pair, the value for the AUC is the chance of correctly identifying which is which. Is that the right way to think about it?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58289",
      "postDate": "11/18/2014 09:01:48",
      "content": "<p>I think yes.&nbsp; A simple sample:</p>\n<p><code>import numpy as np</code><code><code></code></code></p>\n<p><code>from sklearn.metrics.metrics import roc_auc_score</code></p>\n<p><code><code>p = np.array([0.1, 0.2, 0.3, 0.4, 0.7, 0.8, 0.9])</code></code></p>\n<p><code><code>y</code></code><code> = np.array([0, 0, 1, 1, 1, 1, 1])</code><code></code></p>\n<p><code>roc_auc_score(y, p)<br></code></p>\n<p><code>Out[5]: 1.0<br></code></p>\n\n<p>Note, the cut-off is not 0.5</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58290",
      "postDate": "11/18/2014 09:23:22",
      "content": "<p>[quote=Michael Hills;58288]</p>\n<p>I don't have a full handle on ROC AUC yet, so let me see if I've got this right. As long as every prediction for preictal has a higher value than every prediction for interictal, you will score 1.0? [/quote]</p>\n<p>Yes.</p>\n<p>And from my experience, although it needs verification, AUC == fraction of correctly ordered pairs out of all 0/1 pairs.</p>\n<p>You may try your features with some ranking machine, e.g. SVMrank or rank booster. I tried both and got good results on per subject basis. Unfortunately didn't figure out a good method to merge them.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58300",
      "postDate": "11/18/2014 10:40:28",
      "content": "<p>First of all, congratulations to the winners :-) and to the organization of Kaggle and American society of epilepsy.&nbsp;This has been a very nice competition, with very&nbsp;encouraging and motivational discussions in the forum.</p>\n<p>Regarding the final results, I think the key is in the features, and FFT filters. We decided to use standard filtering, as in the PLOS ONE paper referenced by competition host. Using logistic regression over this features leads to a very bad LB result (0.678), but very high in CV (0.934). However, a deep ANN using same features (with 5 layers) achieves a LB result of 0.794 and CV of 0.928.&nbsp;</p>\n<p>We didn't received any prize (just 7th place), but our results were&nbsp;very consistent between public and private LB (0.82488 and 0.79347). I think it worth the effort to explain briefly our system, but I just advance that we tried a lot of different features and models, and finally a little gain was obtained by making a linear combination of our ideas.</p>\n<p>The key to improve results is the combination of several&nbsp;models with different features and different hyper-parameters, and to use PCA or ICA for decorrelation of data. We feel that&nbsp;consistency of our results has been increased by system combination approach. We observe that AUC variance in LB using different combinations of models was very similar.</p>\n<p>In brief, the pipeline is composed of the following feature sets:</p>\n<p>1) FFT features: computed by using hamming windows of 60s with an overlapping of 50%.&nbsp;Over FFT the filter bank of PLOS ONE paper has been computed. In order to decorrelate data, we use a PCA (and also ICA) transformation, which allows to win 0.04 points of AUC (from 0.7489 to 0.78153 in an ANN with 2 layers).</p>\n<p>2) Eigen values of correlation matrix between channels, and computed&nbsp;for the same 60s with overlap windows (similar to what Michael Hills has been done in previous competition).</p>\n<p>3)&nbsp;Eigen values of correlation matrix between channels for the entire 10 minutes signal.</p>\n<p>4)&nbsp;Eigen values of correlation matrix between channels for the entire 10 minutes of the signal after differentiation.</p>\n<p>5) A bunch of selected features (mean and variances) over the 10 minutes original signal.</p>\n<p>By using this feature sets, different models have been estimated:</p>\n<p>a)&nbsp;ANN with 2 layers over features 1 (using PCA decorrelation) and features 2</p>\n<p>b) Deep&nbsp;ANN with 5 layers over features 1 (using PCA decorrelation) and features 2</p>\n<p>c)&nbsp;ANN with 2 layers over features 1 (using ICA&nbsp;decorrelation) and features 2</p>\n<p>d)&nbsp;K-nearest-neighbor over&nbsp; features 1 (using PCA decorrelation) and features 2</p>\n<p>e)&nbsp;K-nearest-neighbor over&nbsp; features 1 (using ICA&nbsp;decorrelation) and features 2</p>\n<p>f) K-nearest-neighbor over features 3 and 4</p>\n<p>g) K-nearest-neighbor over features 5</p>\n<p>We find that KNNs achieves better results than logistic regression, and that ANNs were even better if you can estimate properly&nbsp;the learning rate, momentum and regularization. Also, the difference between ICA and PCA wasn't significant, both approaches obtain similar results in AUC, but they were important to improve convergence of ANN models.</p>\n<p>Finally, we decided to use a Bayesian model combination (BMC)&nbsp;of all the seven models (a-g). The BMC increases our AUC by 0.03 points in both private and public LB. It is possible to observe that BMC is not achieving a huge improvement, but&nbsp;we feel more comfortable and sure about the consistency of our result.</p>\n<p>I hope someone find this explanation interesting. We will try to publish our system&nbsp;pipeline, just&nbsp;for reproducibility and disemination.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58302",
      "postDate": "11/18/2014 11:01:19",
      "content": "<p>[quote=Jonathan Tapson;58282]</p>\n<p>I generally calculated a separate regression on each permutation (fold) of interictal/preictal data and then summed the test set predictions generated by those weights</p>\n<p>[/quote]</p>\n<p>Jonathan, how many permutations of interictal/preictal data did you use for a given subject and how did you choose these permutations? Does it make difference to calculate several regressions instead of doing single one using all data&nbsp; for a given subject? In the end you average them via summation?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58322",
      "postDate": "11/18/2014 16:45:57",
      "content": "<p>Congratulation to all. I learned a lots from you guys.</p>\n<p>My approach:</p>\n<p>1)Features:</p>\n<p>Michael Hills's feature pipe which gives&nbsp; 900+ features per analysis window</p>\n<p>Each 2s get one analysis window</p>\n<p>Each analysis window contains 4s's data(4*400*Channels for Dog_n and 4*5000*Channels for human),</p>\n<p>Eg for Dog_1&nbsp;&nbsp;&nbsp; EEG&nbsp; 600*400*16&nbsp; =&gt; 298*1024</p>\n<p>(Samples*FeaturesSize)</p>\n<p>2)Models: 1 layer neural network with 20 nodes,same setup for all subjects</p>\n<p>3) Training: Training&nbsp; 50 to 70 models with [positive-class, negative-class(random and shuffle) ] data pairs for each subject</p>\n<p>4)Scoring: majority vote</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58337",
      "postDate": "11/18/2014 19:06:48",
      "content": "<p>@rakhlin: I used the fact that the preictal data needed to be separated in the sequence blocks of 6 to determine the folding method. &nbsp;The method goes like this:</p>\n<p>1. &nbsp;Find number of preictal blocks of 6 (e.g. dog 1 has 24/6 = 4 blocks)&nbsp;</p>\n<p>2. &nbsp;Divide interictal into the same number of blocks (dog 1 has 4 blocks of 120 each)</p>\n<p>3. perform leave-one-block-out folding for all permutations of interictal and preictal blocks (16 permutations here). &nbsp;Generate a solution for each permutation and sum (average) the results.</p>\n<p>This method has two weaknesses: it leaves out some data for subjects where the numbers are not divisible by 6. &nbsp;It also mixes up preictal blocks in dog 4 where the sequences are not complete blocks of 6. &nbsp;I am kicking myself for not doing that subject properly, it might have given me another .005...</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58342",
      "postDate": "11/18/2014 20:03:18",
      "content": "<p>Thank you &nbsp;Jonathan.</p>\n<p>I don't fully understand the goal of these permutations. Leave-one-block-out folding makes a big sense for cross-validation. But for ultimate solution - because you feed all data to deterministic(?) linear model equally and average the result in the end - is it in principal different from training single LR on the whole data set?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58350",
      "postDate": "11/18/2014 21:36:23",
      "content": "<p>Thank you all. It was not easy. I will be happy to give some information about my strategy for this competition in the near future.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58351",
      "postDate": "11/18/2014 21:41:49",
      "content": "<p>[quote=Jonathan Tapson;58255]</p>\n<p>That was fun but very frustrating towards the end.&nbsp; For what it's worth, I used only spectral features (fft over 30s and 60s intervals, overlapped 50%).&nbsp; The secret weapon was linear regression - you could get to about 0.85 on the leaderboard with straight LR, post scaled through a logistic function to (0,1) interval.&nbsp; No networks, trees, SVMs, RBMs, etc. required.&nbsp; Because LR is superfast and can be inverted (i.e. you can back-process to see what features it is scaling up, and what features it ignores)&nbsp; I found some good feature sets.&nbsp; It is also very hard to overtrain.&nbsp; The general best featuresets were 1Hz bands from 0-50 Hz and then 5-10Hz bands up to 180Hz.&nbsp; All data was filtered for 60Hz + harmonics.&nbsp; I did use some networks (ELM ensembles) to get a couple of extra points.&nbsp; The final solution was computed on my laptop&nbsp; (a 2010 MacBook Air).&nbsp; No AWS required.&nbsp; Now I just hope I can reproduce the damn result for the organizers!</p>\n<p>[/quote]</p>\n\n<p>I know these questions might seem dumb, but it would be nice if you could clarify a bit more as I am trying to learn about this better. When you say you overlapped 50% what do you mean? &nbsp;Secondly what do you mean by &quot;1Hz bands from 0-50 Hz and then 5-10Hz bands up to 180Hz. All data was filtered for 60Hz + harmonics&quot; . Just a little more clarification on this would do me wonders.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58357",
      "postDate": "11/19/2014 00:00:55",
      "content": "<p>If using say 60s intervals, the 600s of data were partitioned into 10 blocks (with partitions at n.60s for n=1:9).&nbsp; Then another 10 blocks were generated by starting 30s into the data, with partitions at 30s+n.60s. (the first 30s was joined with the last 30s to make the tenth block here).</p>\n<p>The FFT generates lines at intervals of Fs/N where Fs is the sampling frequency and N is the data length (say 400 Hz and 24k samples for a 60s block).&nbsp; So there would be 24k/400=60 fft lines in any 1Hz interval.&nbsp; So to generate a band for 0-1Hz you add the power (log(abs(fft)) for those 60 lines in that interval.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58358",
      "postDate": "11/19/2014 00:04:10",
      "content": "<p>@rahklin: In principle one model should be as good as the folded model.&nbsp; In practice it seems that the leave-one-out ensemble gives better generalisation, I think because you are effectively bagging.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58364",
      "postDate": "11/19/2014 03:41:06",
      "content": "<p>Congratulations everyone. This has been an excellent competition on a very difficult problem, and you've come up with a great set of solutions.&nbsp;</p>\n<p>On behalf of the organizers and sponsors, thank you all for your great work on this competition.</p>\n<p>I realize there are a few outstanding questions, and I promise detailed answers will be forthcoming once I've confirmed some details with the group here.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58368",
      "postDate": "11/19/2014 03:57:30",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58371",
      "postDate": "11/19/2014 05:23:25",
      "content": "<p>Could the admins (or anyone else knowledgeable about the performance metrics used in this competition) comment on how the competition results compare to prior &quot;state of the art&quot; algorithms that have previously been published; e.g.:</p>\n<p style=\"padding-left: 30px\">&#8226; Howbert JJ, Patterson EE, Stead SM, Brinkmann B, Vasoli V, Crepeau D, Vite CH, Sturges B, Ruedebusch V, Mavoori J, Leyde K, Sheffield WD, Litt B, Worrell GA (2014) Forecasting seizures in dogs with naturally occurring epilepsy. PLoS One 9(1):e81920.<br>&#8226; Cook MJ, O'Brien TJ, Berkovic SF, Murphy M, Morokoff A, Fabinyi G, D'Souza W, Yerra R, Archer J, Litewka L, Hosking S, Lightfoot P, Ruedebusch V, Sheffield WD, Snyder D, Leyde K, Himes D (2013) Prediction of seizure likelihood with a long-term, implanted seizure advisory system in patients with drug-resistant epilepsy: a first-in-man study. LANCET NEUROL 12:563-571.</p>\n<p>The subjects and estimators of performance are different in this competition versus these papers, so comparing&nbsp;the different systems may be like comparing apples to oranges. However,&nbsp;it would be very&nbsp;impressive&nbsp;if this competition's winners have outperformed the best that academic research has had to offer these last few years.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58372",
      "postDate": "11/19/2014 05:35:39",
      "content": "<p>Second @idaniboy's request. I was thinking about asking the same question later on.</p>\n<p>It would be interesting to see the performance comparison of the best of the new/fresh/outsider approaches vs. the results from well known approaches from domain experts.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58373",
      "postDate": "11/19/2014 05:46:41",
      "content": "<p>I am also interested. I've read a lot of research papers over these two competitions, with many talking about non-linear features or other things like phase synchrony, bispectrum, Lyapunov exponent, autoregressive mode and so on. My own experience with these competitions has shown that spectral features have worked the best. FFT bins and spectral entropy combined with cross correlation in time and frequency domains are my core features. I will make some submissions later to compare the strength of the multivariate correlation features and the univariate spectral features.</p>\n<p>I did try AR model, detrended fluctuation analysis and phase synchrony and none&nbsp;improved my model. I appeared to get a very very small bonus from petrosian fractal dimension, hurst exponent, and higuchi fractal dimension, but need to do more testing.</p>\n<p>However it's hard to be conclusive on anything as I felt my setup was very sensitive to SVM parameter tuning and throwing in more features might have required some retuning to squeeze the performance out. Additionally I implemented some of these things myself because performance was too slow for some libraries I tried to use. So there is also the question of whether I implemented it correctly and I didn't have much way of unit testing my maths. I will try some experiments a bit later using Jonathan Tapson's linear regression technique to get a rough gauge of the individual predictive power of my features.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58384",
      "postDate": "11/19/2014 09:15:31",
      "content": "<p>Congratulations to Nir, Jonathan and Andronicus1000 ! Nice work!</p>\n<p>Also Congratulations to the other top-ten finishers as they win a chance to publish their solutions.</p>\n<p>I have one question for the top-ten finishers:</p>\n<p>How much time did you spend roughly on this competition?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58389",
      "postDate": "11/19/2014 10:29:39",
      "content": "<p>[quote=idaniboy;58371]</p>\n<p>Could the admins (or anyone else knowledgeable about the performance metrics used in this competition) comment on how the competition results compare to prior &quot;state of the art&quot; algorithms that have previously been published;&nbsp;</p>\n<p>[/quote]</p>\n<p>Everyone is happy, so I'd like to dilute it with some criticism :) There is a fundamental difference. Usage of test data. No offence, I congratulate&nbsp;Jonathan and other, in the given pre-specified framework he did an excellent work, but unfortunately I have to 'congratulate' also organisers and data providers, in particular, 'cause apparently in the follow-up paper 10 genuine technical contributions will be reported :))) In the area with several decade history such as seizure prediction, logistic regression on FFT features is a definite step forward. A jump :)) If that's what data providers wanted, to trade solution complexity for an ability to report 83% performance instead of 70-75. I hope they will not occidentally 'forget' to mention that in the paper and the consequences of that decision. &nbsp;I just wonder what was the thinking behind? Did they realise the practical consequences of that decision? That the solutions will be practically useless. Do organiser understand the difference between unlabelled data and unseen data, 'cause they report it as allowing 'semi-supervised learning' :) It has nothing to do with semi-supervised learning. Unseen data imply not just unseen labels. Personally, I realised that the testing data were allowed to be used only when I saw the winner's jump from 86 to 90. Looking forward to seeing other winners' ways to normalise the output probabilities :)&nbsp;</p>\n\n<p>Jonathan, can you report the performance of your method if testing data were not allowed to be used? Just out of curiosity.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58395",
      "postDate": "11/19/2014 12:40:17",
      "content": "<p>Congratulations to everyone and thank you for sharing your approaches! I used convolutional neural networks and finished in 10th place, which makes me euphoric:) I regret of not trying something simpler, although I think convnets are quite elegant. Did anybody else use them? I posted my code and description&nbsp;<a href=\"https://github.com/IraKorshunova/kaggle-seizure-prediction#seizure-prediction-using-convolutional-neural-networks\">here</a></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58396",
      "postDate": "11/19/2014 12:51:47",
      "content": "<p>congratulations to the winners !</p>\n<p>my code for taking an existing submission and making from it a new submission (with &nbsp;higher score) using information I found in the test data</p>\n<p>https://github.com/udibr/seizure-detection-boost</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58400",
      "postDate": "11/19/2014 13:23:45",
      "content": "<p>Hi @golondrina. We tried CNNs at the begining but just throw away the idea. I feel that first layer filters can be well estimated, but output probabilities not. However, our best model was a&nbsp;deep neural network with 5 hidden layers, 128 relus in every layer, and one logistic output layer.&nbsp;So, the ANN was trained using a sliding window over input data, that is, the whole ANN is a non-linear kernel which is applied over temporal axis. The outputs of ths kernel are probabilities and we combine them using just a geometric mean of the complementary probability. We are preparing the code,&nbsp;cleaning, removing dead code parts (not used), etc. You can find a more extend explanation of our system in the forum:</p>\n<p>https://www.kaggle.com/c/seizure-prediction/forums/t/10945/congratulations-to-the-winners/58300#post58300</p>\n<p>And we will publish the code as soon as possible.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58401",
      "postDate": "11/19/2014 13:25:23",
      "content": "<p>Andy, quite the contrary, the organizers needlessly complicated the problem. Other works don't report performance across subjects putting them on a common scale like here - it is meaningless. Add small and highly unbalanced data, particularly for 2 humans, and the problem can not be generalized well even on per subject basis. I think without post-calibration performance would not outbid 60%. Finally add quite meaningless metric. In practice you're not interested in AUC. For perfect classifier it should be enough to produce no false positives and at least&nbsp;<strong><em>one</em></strong> true positive for every preictal period.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58402",
      "postDate": "11/19/2014 13:39:53",
      "content": "<p>I tried CNN similar to the one described in LeCun et al &quot;Classification of Patterns of EEG Synchronization for Seizure Prediction&quot; but it failed estimate 1st layer filters - I suspect due to huge amount of channel pairs I used (120) in contrast with original work (only 15)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58403",
      "postDate": "11/19/2014 13:43:29",
      "content": "<p>@Francisco, it's impressive what you've done! &nbsp;I'm looking forward to see your code, because some details in your explanation I don't understand.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58404",
      "postDate": "11/19/2014 13:45:38",
      "content": "<p>@rakhlin, it is surprising to me that first layer filters didn't converge. Our approach is basically a&nbsp;deep one filter, and a geometric mean combination of multiple probability ouputs for each file. Depending in&nbsp;your temporal resolution, it is possible to increase the number of samples by 20, or even more,&nbsp;it is a&nbsp;number of samples large enough to train the first layer filters. Which activation function do you used?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58405",
      "postDate": "11/19/2014 13:47:59",
      "content": "<p>@golondrina sorry for the complexity of the explanation. We tried a lot of different kind of features and different model combinations. We&nbsp;will&nbsp;to put it together in a brief and&nbsp;clear way. However, my post in the forum is not as easy to follow as I would wanted :(</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58408",
      "postDate": "11/19/2014 13:57:52",
      "content": "<p>Francisco, now I suspect too little number of samples was another problem. I used tanh activation and L1-regularizer. &nbsp;I'm too looking forward to see your code.</p>\n\n<p>By the way. If you kept 1st layer filters of your CNN it would be instructive to see patterns they found.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58415",
      "postDate": "11/19/2014 14:29:10",
      "content": "<p>@rakhlin my intuition is that if you want to compute one output probability for each segment file&nbsp;by using a CNN, the major problem relies&nbsp;in the last layer. It is important a clever selection of how the last layer is computed due to lack of data.</p>\n<p>For example, Dog_1 only has around 500 samples, where only 24 are positive. This scenario difficults the&nbsp;learning of this&nbsp;map from CNN top layer features into preictal probability for each segment, IMHO.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58417",
      "postDate": "11/19/2014 14:56:32",
      "content": "<p>Francisco, handling unbalanced data is an interesting topic. I don't feel this was a problem in case of CNN as I tried several strategies like oversampling with noise, assigning proportional weights to positive examples in the output layer, training many nets with balanced chunks. It failed generalize not because of unbalanced data but over little data or high dimensionality. In the end I abandoned CNN, &nbsp;then tried denoising autoencoders, and finally concentrated on linear approaches. It's amazing how simple linear models worked out for Jonathan Tapson.</p>\n<table border=\"0\" width=\"100%\" cellspacing=\"0\" cellpadding=\"15\">\n<tbody>\n<tr>\n<td>\n<table align=\"center\">\n<tbody>\n<tr>\n<td width=\"500\">\n<table width=\"100%\" align=\"center\">\n<tbody>\n<tr>\n<td align=\"left\">&nbsp;</td>\n</tr>\n</tbody>\n</table>\n</td>\n</tr>\n</tbody>\n</table>\n</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58420",
      "postDate": "11/19/2014 16:04:39",
      "content": "<p>Congratulations to all the toppers!</p>\n<p>@Jonathan, wonderful to read about your approach, and more importantly the simplicity of your solution. I wonder if you can share what the outcome of your approach was when you used shorter sliding window lengths and smaller time steps? A natural start is to use much smaller window lengths, and I suspect that by using&nbsp; longer window lengths you've effectively filtered out (or muted) spectral variations which accompany natural&nbsp; working, which could potentially hinder discrimination, and for these subjects, perhaps it is even filtering out shorter ictal activity (ie if sufficiently many bursts occur in short period, they are captured, or else not). Again, great job on the approach &amp; win!</p>\n<p>@bbrinkm, If I may, perhaps a naive question... did the preictal clips lead up to clinical seizures only? Or were there instances of subclinical seizures also? And related, were all the interictal clips completely free of any ictal-like onset signature (however sparse)? Many thanks for hosting such an interesting competition.</p>\n<p>Thanks.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58425",
      "postDate": "11/19/2014 16:31:07",
      "content": "<p>@Cy - great question. They were electrographic seizures. We didn't have access to video monitoring on the dogs, so we identified seizures in the EEG. The interictal clips were well clear of any seizures. Philosophically, it is well known that epilepsy patients have interictal signatures (spikes, high frequency oscillations, etc.) of epilepsy in their EEG - so no patient's interictal EEG is totally normal. But there shouldn't have been any seizure-like activity in the interictal clips. Artifacts, perhaps - but no seizures.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58430",
      "postDate": "11/19/2014 17:28:52",
      "content": "<p>@bbrinkm, Thank you for the quick response.To say the least, must have been a laborious task to build this dataset! May I further inquire whether the electrographic patterns were identified in the time-domain or frequency domain? And, if classified in the frequency domain, may I inquire what signature was used for the classification? </p>\n<p>I used much shorter sliding windows to identify the spectral signatures, and wondered if I was picking up short ictal activities in the interictal clips. Hence my question earlier... However, as you note, perhaps the algorithm was conflating artifacts with interictal/preictal activity (though, I'd have thought that the artifacts would results in a different broadband spectral response; different enough to be differentiated). Then again, that's a supposition on my part... I will take a closer look this weekend.</p>\n\n<p>Again, many thanks for addressing our questions! </p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58449",
      "postDate": "11/19/2014 22:00:46",
      "content": "<p>Congratulations to the winners. This competition turned out to be an interesting and fun competition for me unexpectedly.<br>I have learned and used many tips from the discussions from other Kagglers in this and other competitions. Even though&nbsp;I did not squeeze into top 10, I still would like to share some of my thoughts and approaches which may provide some aspects that have not been&nbsp;used/mentioned by others.&nbsp;The final private LB scores of this competition is so tight that if I had chose my best private&nbsp;score submission, (I chose my 2nd and 3rd best private scores, which is not too bad. The 2nd best is only 0.006&nbsp;less than my best score.), I could have jumped to No. 6 instead of the current No. 11.&nbsp;I hope someone can find some of the points useful. The sharing and learning environment at&nbsp;Kaggle is the best thing I like about Kaggle.</p>\n<p>I did not plan to participate this competition as I saw the data set is pretty big seem like require a lot of time.&nbsp;Just about two weeks before the end, I saw that I could have time in my schedule to participate a competition and at around the same&nbsp;time it was announced that the daily submission limit was increased to 10 which make my late entrance more feasible.&nbsp;I thought I might take a crack at this.</p>\n<p>I used R for this competition. The runs were mostly on my Home PC with i7 CPU and 12GB memory.</p>\n<p>Overall consideration and approach:</p>\n<p>a. Because the limited time available and the significant turn-around time to generate new set of features.&nbsp;I figured that I need to chose a set of features and stick to it. No time to do much feature selection/reduction.</p>\n<p>b. I have no domain knowledge and did not participate in the first Seizue competition. Did not have much time to dive into the literatures. All the features I chose was from browsing through the forum of this and previous Seizure&nbsp;competition. And only choose items that I can understand and quickly implement.</p>\n<p>c. One of the problem I see that I need to deal with at the beginning is the conflict of a performance metrics that&nbsp;includes all subjects and seperately trained models for each individual subjects. Seperately trained models for&nbsp;individual subjects may perform better for the individual subject, however, the probabilities generated from different&nbsp;models can have very different range, magnitude. And because for AUC metric, what matters is the ranking order of the predictions, when you combine results from individual models, because the differences in overall magnitudes or&nbsp;ranges of the predictions, simply combining the results from individual models may not yield best results. I decided&nbsp;at the beginning that for each type of model I use, I will have models trained for each subject using all the features available to that subject and a same type model that are global, trained using data from all subjects using a set of features common to all subjects. My plan was to later use the predictions from the global model to calibrate&nbsp;and combine the outputs of individual models. But I did not get time to explore this and at the end just simply ensemble both&nbsp;global and simply combined individual model results.</p>\n<p>I will discuss my approach on the the following topics: Feature generation, Models used, Ensemble of different models.</p>\n<p>Features:<br>a. To make features more comparable between subject and also to reduce the size, down sample signal data to 200Hz.</p>\n<p>b. There were 3 evolving stages in the progress of feature generation in my approach. Each stage saw significant improvement in LB results.</p>\n<p>c. Summary statistics only features: First, I saw in the post of the previous competition, some of the top winners&nbsp;only used summary statistics of the time series and FFT data. For each mat file, there are N number of channels.&nbsp;For each channel, besides the original signal series, I also generated delta and delta of delta series.&nbsp;For each of these series in time domain, I also generated FFT series. Then for each data series (time-domain and FFT)&nbsp;I generated a few basic summary statistics (mean, max, Stdev). Plus Frequency of FFT peak.&nbsp;</p>\n<p>Then also the overall summary statistics. First, generated statistic for a series of average of all channels.&nbsp;mean, max, stdev for all the quantities across all channels. Then the mean, max , SD and max,2nd max, 3rd max of&nbsp;elements of covariance matrix across all channel for both the time-domain series and FFT series.</p>\n<p>Total 600+ features for dogs, little less number of common features, much more for patient_2. The best public LB results I could get with&nbsp;these features were: 0.657 for single model, 0.687 from ensemble</p>\n<p>d. Summary statistics plus FFT: Adding FFT series is the natural next step. To reduce the size of FFT series,&nbsp;borrowed from the thinking of @ruai at <a href=\"https://www.kaggle.com/c/seizure-prediction/forums/t/10884/what-type-of-models-are-people-using/57672#post57672\">&quot;what-type-of-models-are-people-using&quot;</a>,&nbsp;somewhat different parameters, using only the front portion (0.2) and smoothing the raw FFT to get only 24 points&nbsp;per FFT series. To arrive at the value of these parameters, I just simply plot the FFT series of a few files with&nbsp;different cutoff and averaging settings to get what I feel that are manageable and still kept enough details of&nbsp;the FFT series.<br>2000+ features for dogs, best single model: 0.69284, ensemble: 0.717<br>e. With the tips from this discussion&nbsp;<a href=\"https://www.kaggle.com/c/seizure-prediction/forums/t/10909/increase-number-of-samples-by-using-short-time-segments/57972\">&quot;increase-number-of-samples-by-using-short-time-segments&quot;</a>:</p>\n<p>Break each of the original 10Min series into 10 non-overlap 60s series. The train data increased 10-fold.&nbsp;Used the same features as above for these shorter series. The eventual prediction for each original 10Min series&nbsp;was taken the average or certain percentile of the 10 predictions from the shorter series.<br>Best single model: 0.80514, Best Ensemble: 0.81803</p>\n<p>Model used:<br>a. After the Higgs-Boson competition, xgboost become my favorite tool. Fast, effective.<br>b. For this data set SVM performed very well.<br>c. glmnet is competitive<br>d. also tried gbm and randomForest. These models had much worse results and very slow. I wish I did not waste too much time&nbsp;at the end to run these models instead I should have put more efforts into better tuning the better performed models.<br>My thought was to use as many types of fundamentally different models to ensemble as possible. In hindsight, I probably should ensemble several runs of the better performed models with different parameters. (Not tested.)<br>e. Each model type,&nbsp;both individual model and global model were constructed and tested.</p>\n<p>Ensemble method:<br>a. AUC only depends on the order of the predictions. From previous competitions, I found the most effective way&nbsp;to ensemble a set of results from different models for AUC metric is to convert each solution to ranks and take the&nbsp;average ranks of all the solutions.<br>b. This ensemble method provides about 0.01-0.03 improvements.</p>\n<p>Cross-Validation</p>\n<p>2-fold CV with 10 shuffles, based on what the data description page used as &quot;series&quot;, i.e. mat files with different sequence number but belong to the same series are in and out of folds together. The resulted AUC higher than public LB, but direction wise are pretty good.</p>\n<p>My public and private LB scores are pretty much in sync. My public LB position was 25, private LB position jump to 11.&nbsp;I think my approach of constructing models trained both globally and individually helped in giving the stability of my&nbsp;results.</p>\n<p>This is my experience. Hope you can find something useful.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58450",
      "postDate": "11/19/2014 22:09:41",
      "content": "<p>@Wei Wu, I would love to see your R code. Is it possible if you put it up?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58451",
      "postDate": "11/19/2014 22:13:31",
      "content": "<p>Oh no, I was looking forward to not having to worry about a writeup / code-cleanup.&nbsp; Andronicusssss!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58456",
      "postDate": "11/19/2014 23:07:49",
      "content": "<p>Haha when Andronicus first appeared on the leaderboard it seemed very suspicious, 7 days old account and an immediate 3rd place. I figured it would jump me 2 places on the leaderboard at the end.</p>\n<p>For those discussing unbalanced data problem, I too tried many different techniques like oversampling preictal segments, or training many smaller classifiers on more balanced subsets, e.g. if there's 42 preictal and 500 interictal, split the interictal into groups of 42, train N classifiers and mean their probability outputs. Always made it worse.</p>\n<p>I'm not sure I stated it anywhere so the features I used. I used a non-overlapping 75s window size to generate 8x training samples. Features are:</p>\n<p><span style=\"line-height: 1.4\"><span style=\"line-height: 1.4\"></span></span></p>\n<ol>\n<li><span style=\"line-height: 1.4\">cross correlation in time domain</span></li>\n<li><span style=\"line-height: 1.4\">cross correlation in frequency domain</span></li>\n<li><span style=\"line-height: 1.4\">hand-picked frequency bins&nbsp;[0.5, 2.25, 4, 5.5, 7, 9.5, 12, 21, 30, 39, 48], where the frequency magnitude is taken as a mean for each hz-range pair, then apply log10</span></li>\n<li><span style=\"line-height: 1.4\">power-in-band spectral entropy for various bands in the range of 0.5-24Hz</span></li>\n<li><span style=\"line-height: 1.4\">higuchi fractal dimension</span></li>\n<li><span style=\"line-height: 1.4\">petrosian fractal dimension</span></li>\n<li><span style=\"line-height: 1.4\">hurst exponent</span></li>\n</ol>\n<p><span style=\"line-height: 1.4\"><span style=\"line-height: 1.4\"></span></span></p>\n<p><span style=\"line-height: 1.4\"><span style=\"line-height: 1.4\"><span style=\"line-height: 1.4\">The first 3 seemed to work well. This is basically same features as I used for first competition. The spectral entropy seemed to benefit Dogs 3 and 4 more. The last 3 seemed to help only very slightly.</span></span></span></p>\n<p><span style=\"line-height: 1.4\"><span style=\"line-height: 1.4\"></span></span></p>\n<p><span style=\"line-height: 1.4\"><span style=\"line-height: 1.4\"><span style=\"line-height: 1.4\"><span style=\"line-height: 1.4\">Random feature selection was used on 1, 2 and 3 where 52.5% of features were selected randomly to form a feature mask.&nbsp;I&nbsp;ran genetic algorithm optimising for local CV scores on 3 separate feature groups, the spectral entropies with higuchi fractal dimension, and then petrosian fractal dimension and hurst exponent in their own GA runs. The GA was instructed to seed 30 population with 55% features used. It was run for 10 generations.</span></span></span></span></p>\n<p><span style=\"line-height: 1.4\"><span style=\"line-height: 1.4\"></span></span></p>\n<p>Finally the best 2 feature masks from each GA feature group, and 2 randomly generated feature marks for the random groups, were combined. With all the features combined, I then had 2 different masks across the features. I trained each one, and then averaged the predictions. Also because I was using 75s window size, I averaged the predictions for the 8 windows in each segment.</p>\n<p><span style=\"line-height: 1.4\"><span style=\"line-height: 1.4\"></span></span></p>\n<p><span style=\"line-height: 1.4\">Later I learned that all this boost I was getting from GA (from 0.79 to 0.84) was mostly because my SVM parameters were not tuned correctly for Dog 3 and 4 and I could get 0.83 without any genetic algorithm or random feature selection, but plain simple all-features one SVM with different parameters. Then I tried Tapson's approach post-competition and got 0.84 easily using my features with linear regression. :)</span></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58458",
      "postDate": "11/19/2014 23:51:05",
      "content": "<p>[quote=Wei Wu;58449]</p>\n<p>Ensemble method:</p>\n<p>a. AUC only depends on the order of the predictions. From previous competitions, I found the most effective way&nbsp;to ensemble a set of results from different models for AUC metric is to convert each solution to ranks and take the&nbsp;average ranks of all the solutions.</p>\n<p>....</p>\n<p>This is my experience. Hope you can find something useful.</p>\n<p>[/quote]</p>\n\n<p>Thank you, very interesting. Could you please elaborate how do you combine individual ranks.&nbsp;Particularly what does it mean:&nbsp;&quot;....&nbsp;and take the average ranks of all the solutions.&quot; From my experience&nbsp;different&nbsp;subjects demonstrate&nbsp;large&nbsp;difference in rank&nbsp;deviation&nbsp;and mean.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58468",
      "postDate": "11/20/2014 01:54:56",
      "content": "<p>If anyone wants to try Jonathan's approach with sklearn I wrote a wrapper for LinearRegression that implements predict_proba using his method so you can just slot it in. I'm not 100% sure it's exactly what he did, we can wait for his code to verify. I do see the easy high scores he mentioned it gave although not as high as his but that's probably to do with my windowing and features than the classifier.</p>\n\n<p>class SimpleLogisticRegression(LinearRegression):<br>&nbsp; &nbsp; def predict_proba(self, X):<br>&nbsp; &nbsp; &nbsp; &nbsp; predictions = self.predict(X)<br>&nbsp; &nbsp; &nbsp; &nbsp; predictions = sklearn.preprocessing.scale(predictions)<br>&nbsp; &nbsp; &nbsp; &nbsp; predictions = 1.0 / (1.0 + np.exp(-0.5 * predictions))<br>&nbsp; &nbsp; &nbsp; &nbsp; return np.vstack((1.0 - predictions, predictions)).T</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58480",
      "postDate": "11/20/2014 05:34:43",
      "content": "<p>[quote=Francisco Zamora-Martinez;58415]</p>\n<p>For example, Dog_1 only has around 500 samples, where only 24 are positive. This scenario difficults the&nbsp;learning of this&nbsp;map from CNN top layer features into preictal probability for each segment, IMHO.</p>\n<p>[/quote]</p>\n<p>I tried some training where I used sets of 50:50 interictal:preictal with neural networks (re-using the small number of preictals several times). &nbsp;They worked OK but no better than anything else...</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58489",
      "postDate": "11/20/2014 05:45:39",
      "content": "<p>[quote=Andy;58389]</p>\n<p>Everyone is happy, so I'd like to dilute it with some criticism :) There is a fundamental difference. Usage of test data. No offence, I congratulate&nbsp;Jonathan and other, in the given pre-specified framework he did an excellent work, but unfortunately I have to 'congratulate' also organisers and data providers, in particular, 'cause apparently in the follow-up paper 10 genuine technical contributions will be reported :))) In the area with several decade history such as seizure prediction, logistic regression on FFT features is a definite step forward. A jump :)) If that's what data providers wanted, to trade solution complexity for an ability to report 83% performance instead of 70-75. I hope they will not occidentally 'forget' to mention that in the paper and the consequences of that decision. &nbsp;I just wonder what was the thinking behind? Did they realise the practical consequences of that decision? That the solutions will be practically useless. Do organiser understand the difference between unlabelled data and unseen data, 'cause they report it as allowing 'semi-supervised learning' :) It has nothing to do with semi-supervised learning. Unseen data imply not just unseen labels. Personally, I realised that the testing data were allowed to be used only when I saw the winner's jump from 86 to 90. Looking forward to seeing other winners' ways to normalise the output probabilities :)&nbsp;</p>\n<p>Jonathan, can you report the performance of your method if testing data were not allowed to be used? Just out of curiosity.&nbsp;</p>\n<p>[/quote]</p>\n<p>I think you may be overestimating the effect of using the test data - in my case it had no effect on the individual (per-subject) outcomes, it was only used to fix the mean and variance of all subjects to be the same, because that was critical for the use of the combined AUC. &nbsp;That (the combined AUC metric) has nothing to do with the effectiveness of the method in practical terms. &nbsp;My method worked because the data is, for once, quite linearly separable and doesn't require to be nonlinearly transformed in a higher dimensional space. &nbsp;That characteristic&nbsp;is independent of whether you use test data or not (and that is why the CNN methods didn't work any better).</p>\n<p>&nbsp;To answer the question: the best I got without using the test data was 0.81806. &nbsp;That was quite early on and I think one could do better, maybe others did?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58492",
      "postDate": "11/20/2014 06:28:15",
      "content": "<p>Jonathan when I used your approach I did not use any test data and achieved in the 0.83-0.84 range (public LB) &nbsp;with my features. I cleaned up most of my code so will try to release it as soon as I can but probably won't have much time to finish it up until Sunday night or Monday probably.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58512",
      "postDate": "11/20/2014 11:24:31",
      "content": "<p>[quote=Jonathan Tapson;58489]</p>\n<p>I think you may be overestimating the effect of using the test data - in my case it had no effect on the individual (per-subject) outcomes, it was only used to fix the mean and variance of all subjects to be the same, because that was critical for the use of the combined AUC. &nbsp;That (the combined AUC metric) has nothing to do with the effectiveness of the method in practical terms. &nbsp;My method worked because the data is, for once, quite linearly separable and doesn't require to be nonlinearly transformed in a higher dimensional space. &nbsp;That characteristic&nbsp;is independent of whether you use test data or not (and that is why the CNN methods didn't work any better).</p>\n<p>&nbsp;To answer the question: the best I got without using the test data was 0.81806. &nbsp;That was quite early on and I think one could do better, maybe others did?</p>\n<p>[/quote]</p>\n<p>Jonathan thanks for your reply. I assume&nbsp;0.81806 is on private LB, right? I might be overestimating the effect of using test data for this competition score but for practical terms it can not be overestimated. Forget the madness I wrote before explaining the idea on the model-level, lets just consider level of probabilities for simplicity. To simulate a real-life situation the dataset should be say 10-20 patients, with only 1-2 of them with representation of preictal activity and this representation is say 1-2 precital events. So the vast majority of data are background EEG. That's a real-life situation in any EEG-based brain injury problem. What's gonna happen if you apply score normalisation on the test data where there is no preictal activity? Exactly, the normalisation will 'create' this activity from nowhere. This testing-data score normalisation works a) offline and b) with the assumption that preictal activity is present. Neither is pertaining to&nbsp;a real-life SeizPred problem. I'm not criticising the engineering solution perse, just its relation to the real-world problem.&nbsp;</p>\n<p>Michael, what is the&nbsp;0.83-0.84 public LB converted to private LB in your case? There were guys on the fourth place public with 86 falling to 75 in private LB.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58515",
      "postDate": "11/20/2014 12:26:32",
      "content": "<p>[quote=Andy;58389]</p>\n<p>Everyone is happy, so I'd like to dilute it with some criticism :) There is a fundamental difference. Usage of test data. No offence, I congratulate&nbsp;Jonathan and other, in the given pre-specified framework he did an excellent work, but unfortunately I have to 'congratulate' also organisers and data providers, in particular, 'cause apparently in the follow-up paper 10 genuine technical contributions will be reported :))) In the area with several decade history such as seizure prediction, logistic regression on FFT features is a definite step forward. A jump :)) If that's what data providers wanted, to trade solution complexity for an ability to report 83% performance instead of 70-75. I hope they will not occidentally 'forget' to mention that in the paper and the consequences of that decision. &nbsp;I just wonder what was the thinking behind? Did they realise the practical consequences of that decision? That the solutions will be practically useless. Do organiser understand the difference between unlabelled data and unseen data, 'cause they report it as allowing 'semi-supervised learning' :) It has nothing to do with semi-supervised learning. Unseen data imply not just unseen labels. Personally, I realised that the testing data were allowed to be used only when I saw the winner's jump from 86 to 90. Looking forward to seeing other winners' ways to normalise the output probabilities :)&nbsp;</p>\n<p>Jonathan, can you report the performance of your method if testing data were not allowed to be used? Just out of curiosity.&nbsp;</p>\n<p>[/quote]</p>\n\n<p>Andy,</p>\n\n<p>I think your post is interesting and can be extended to a more general issue: increasingly complex machine learning methods and frameworks are proposed in the top conferences and research journals, but the competitions are&nbsp; won by people using simple, well established methods, together with a convenient feature design.</p>\n<p>Although I'm in the academia, I'm increasingly skeptical about how the experiments are performed in machine learning papers, where the most convenient dataset and experimental set-up is chosen to give the proposed method an edge to be published (I try to be honest in my own work, though).</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58517",
      "postDate": "11/20/2014 12:37:15",
      "content": "<p>Jose M. I don't agree&nbsp;with your opinion about machine learning community. For instance, you can look into the following competitions, were complex algorithms win:</p>\n<p>- Merck competition:&nbsp;http://blog.kaggle.com/2012/11/01/deep-learning-how-i-did-it-merck-1st-place-interview/</p>\n<p>- Galaxy zoo compeition:&nbsp;http://benanne.github.io/2014/04/05/galaxy-zoo.html</p>\n<p>- Higgs Boson:&nbsp;http://www.kaggle.com/c/higgs-boson/forums/t/10344/winning-methodology-sharing?page=2</p>\n<p>Outside Kaggle there are more competitions where complex models show better performance.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58518",
      "postDate": "11/20/2014 12:45:04",
      "content": "<p>Francisco,</p>\n<p>Yes, you are right. Deep methods and convolutional NNs have won some competitions. But still, you could find many competitions won by random forests and boosting.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58519",
      "postDate": "11/20/2014 12:50:46",
      "content": "<p>Of course,&nbsp;it is not possible to have a general method to solve all the problems. And this is true for every model, so, it is not an statement which allow you to questioning about machine learning community and its research ;-)</p>\n<p>So, I think that we can agree with that it is necessary to try different approaches, in order to check what works better ;-)&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58520",
      "postDate": "11/20/2014 13:04:52",
      "content": "<p>[quote=Jose M.;58515]</p>\n<p>I think your post is interesting and can be extended to a more general issue: increasingly complex machine learning methods and frameworks are proposed in the top conferences and research journals, but the competitions are&nbsp; won by people using simple, well established methods, together with a convenient feature design.</p>\n<p>Although I'm in the academia, I'm increasingly skeptical about how the experiments are performed in machine learning papers, where the most convenient dataset and experimental set-up is chosen to give the proposed method an edge to be published (I try to be honest in my own work, though).</p>\n<p>[/quote]</p>\n<p>I'm in the academia as well but I'd not be that categoric in judging PR&nbsp;literature. The competitions are very simplified real-world problems. And sometimes (in this case) some decisions are made to make it even less connected to real-life. A nice contribution to the follow-up paper would be to run the best solutions including logreg (fft) on say publicly available Freiburg intracranial dataset. I bet they would not stand a chance. This study is more on early days statistical significant difference between various features for e.g. clinical neurophysiology journal. Simple approach + simple features + best results = &nbsp;data problem (easy data~=real-life, representation problem train~=test in terms of artifacts, etc). Only organisers/data providers can analyse that. In this case the results are not particularly good, true. But I agree with you in the sense that, it'd be nice to have some platform so that every submitted research paper on the EEG-based SP/SD topic would first have to obtain results on the fixed DB with the fixed performance assessment routine, etc. something similar to UCL&nbsp;Machine Learning Repository.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58533",
      "postDate": "11/20/2014 15:52:13",
      "content": "<p>These are good observations, and we will take into account your suggestions for future competitions. As you can tell from other posts on this forum the use of test data for calibration was a controversial and difficult issue. In the real world problem a seizure prediction device would not have access to future data during training. However, for a competition such as this one, there is no way to hide the test set entirely from the contestants. We have to give you that data a priori in order for the contest to work. We removed the restriction of using test data to calibrate because we realized the prohibition was not really enforceable.</p>\n<p>While not perfect, the results of this competition will be far from useless. Calibrating models on the test data perhaps compensates for the disadvantage contestants have of not being able to do ongoing EEG baseline correction, given the lack of time stamps in the test data. For the paper we plan to run contenders' algorithms on held out data segments, data from entirely new dogs and humans if possible, and this will provide the ultimate test of these approaches. </p>\n<p>Thank you for pointing out the admitted weaknesses of this format, and we do plan to disclose all of this in our paper. We would not want anyone to think we have entirely solved the seizure forecasting problem with this competition - clearly this problem is very difficult and more work is needed. Generally in academic papers better comparison between methods is needed, and this contest is part of that effort via the IEEG project for sharing data and algorithms. Researchers (and we hope manuscript reviewers) will be able to run competing algorithms on the same data sets and compare results directly. </p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58541",
      "postDate": "11/20/2014 16:37:47",
      "content": "<p>[quote=bbrinkm;58533]</p>\n<p>While not perfect, the results of this competition will be far from useless.&nbsp;</p>\n<p>[/quote]</p>\n<p>No doubt in it. In fact, plenty of room for analysis and message formulation.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58561",
      "postDate": "11/20/2014 19:16:18",
      "content": "<p>@Mahi Karim:</p>\n<p>My code is pretty messy and ugly. I will need to clean it up a little bit.</p>\n<p>@rakhlin:</p>\n<p>I may not have explained it clearly in English. A few lines of code probably can explain it better, if you know R a little:</p>\n<p>a=read.csv('submission1.csv')</p>\n<p>b=read.csv('submission2.csv')</p>\n<p>a$rank[order(a[,2])]=(1:nrow(a))/nrow(a)</p>\n<p>b$rank[order(b[,2])]=(1:nrow(b))/nrow(b)</p>\n<p>c=data.frame(clip=a[,1], preictal=(a$rank+b$rank)/2</p>\n<p>write.csv(c, row.names=F, quote=F, file='mixed_submission.csv')</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58563",
      "postDate": "11/20/2014 19:26:54",
      "content": "<p>[quote=bbrinkm;58533]</p>\n<p>... However, for a competition such as this one, there is no way to hide the test set entirely from the contestants. We have to give you that data ...</p>\n<p>[/quote]</p>\n<p>You can insert noise or fake samples to test data, eg 50%true test 50% dummy, then 20% for pubic lb 30% private lb....</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58568",
      "postDate": "11/20/2014 20:20:39",
      "content": "<p>[quote=Andy;58520]</p>\n<p>[quote=Jose M.;58515]</p>\n<p>I think your post is interesting and can be extended to a more general issue: increasingly complex machine learning methods and frameworks are proposed in the top conferences and research journals, but the competitions are&nbsp; won by people using simple, well established methods, together with a convenient feature design.</p>\n<p>Although I'm in the academia, I'm increasingly skeptical about how the experiments are performed in machine learning papers, where the most convenient dataset and experimental set-up is chosen to give the proposed method an edge to be published (I try to be honest in my own work, though).</p>\n<p>[/quote]</p>\n<p>I'm in the academia as well but I'd not be that categoric in judging PR&nbsp;literature. The competitions are very simplified real-world problems. And sometimes (in this case) some decisions are made to make it even less connected to real-life. A nice contribution to the follow-up paper would be to run the best solutions including logreg (fft) on say publicly available Freiburg intracranial dataset. I bet they would not stand a chance. This study is more on early days statistical significant difference between various features for e.g. clinical neurophysiology journal. Simple approach + simple features + best results = &nbsp;data problem (easy data~=real-life, representation problem train~=test in terms of artifacts, etc).&nbsp;</p>\n<p>[/quote]</p>\n<p>I'm also in academia and have published ML methods in the journals. &nbsp;My view is that competitions like this are a whole lot more realistic than most of the ML benchmarks, because it's generally real-world data (messy) that someone actually cares about, and the test set is genuinely unseen. &nbsp;If you consider how many thousands of papers have been written on MNIST, where the data is already a significant modification of the original raw data, and there are published algorithms that have obviously been written to address a few recognized problem cases in the test set, then this (Kaggle) is a much more realistic test of methods. &nbsp;Of course, abstracting the real-world problem to a competition requires the organizers to make some simplifications and artificial constructions (like the test set and combined AUC in this competition) but that's unavoidable.</p>\n<p>I came into this competition to demonstrate the value of a particular type of neural network (LSHDI or ELM type), and defeated myself with linear regression. &nbsp;You don't get those outcomes from the regular benchmarks. &nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58570",
      "postDate": "11/20/2014 20:29:12",
      "content": "<p>Geez, you guys are professors?! I can't imagine my boss competing with me (a lousy phd student)..</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58572",
      "postDate": "11/20/2014 21:13:24",
      "content": "<p>[quote=rcarson;58570]</p>\n<p>Geez, you guys are professors?! I can't imagine my boss competing with me (a lousy phd student)..</p>\n<p>[/quote]</p>\n<p>It is better not. 'Cause he may accidentally send you one of his submissions, you might accidentally submit it and you both will be disqualified :)) or not. depends on whether you end up in the money. If yes then it's ok :)))</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58574",
      "postDate": "11/20/2014 21:38:17",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58609",
      "postDate": "11/21/2014 12:29:01",
      "content": "<p>I read a post from a guy who claimed to have helped both Andronicus1000 and Jonathan (he called him Jon). I think it was in this forum (but I could be wrong).<br>He talked about holding the ladder and stuf and also about how Jonathan (Jon) gave a submission (or code files) to Andronicus which Andronicus submitted before he got around to merge with 'Jon'.<br>Does anybody know where this post went?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58639",
      "postDate": "11/21/2014 21:27:53",
      "content": "<p>Not sure if asking for a phd position is on topic here :)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58655",
      "postDate": "11/21/2014 22:43:28",
      "content": "<p>Congratulations QMSDP and Birchwood with your 2nd and 3rd place!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58656",
      "postDate": "11/21/2014 22:53:23",
      "content": "<p>I inadvertently shared methods with someone else in my research group, who I did not know was competing at the time. &nbsp;When I realized that, I invited him to join a team (which would have made it legitimate), but we missed the deadline for the merger. &nbsp;You will note that he made no submissions after the deadline, and withdrew his submission when it came in high, so we acted in good faith; but I guess the rules don't allow for that. &nbsp;We are both out of it now. &nbsp;Congratulations to the winners and I will be interested to see their methods.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58659",
      "postDate": "11/21/2014 23:11:18",
      "content": "<p>Jonathan please accept my sympathy and please share your solution anyway. Its lucidity fascinates me!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58660",
      "postDate": "11/21/2014 23:18:29",
      "content": "<p>[quote=rakhlin;58659]</p>\n<p>Jonathan please accept my sympathy and please share your solution anyway. Its lucidity fascinates me!</p>\n<p>[/quote]</p>\n<p>+1. It is sad some top 10 score is removed and we may lose those great models. None of you do it on purpose. Please share your wonderful work and let it help the epilepsy society.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58661",
      "postDate": "11/21/2014 23:44:48",
      "content": "<p>+2. Indeed very sad that technicalities like merger deadlines should cause removal of what from this thread seems to have been one of the best solutions in&nbsp;the competition.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58663",
      "postDate": "11/22/2014 00:45:58",
      "content": "<p>[quote=Jonathan Tapson;58656]</p>\n<p>I inadvertently shared methods with someone else in my research group, who I did not know was competing at the time. &nbsp;When I realized that, I invited him to join a team (which would have made it legitimate), but we missed the deadline for the merger. &nbsp;You will note that he made no submissions after the deadline, and withdrew his submission when it came in high, so we acted in good faith; but I guess the rules don't allow for that. &nbsp;We are both out of it now. &nbsp;Congratulations to the winners and I will be interested to see their methods.</p>\n<p>[/quote]</p>\n<p>I just wanted to clarify that my 21 submissions were <em>from my own work</em>. As Jonathan and I had discussions about his methods it was necessary that we formed a team. I accepted Jonathan&#8217;s merge request at around 23:58 UTC on the day of the team merger deadline. However the Kaggle server was unresponsive at the time (presumably from everyone making their submissions just prior to the deadline). So my merge request was denied because by the time the server got to my request it was past the deadline. I&#8217;m sure someone at Kaggle can corroborate this if they look at their server logs!</p>\n<p>Anyway that&#8217;s why I made no further submissions past the team merger deadline. Unlike others, I was not&nbsp;disqualified because I was caught cheating. <em>There was no cheating by&nbsp;us</em>. I emailed Kaggle and let them know what happened and requested to be removed. I didn&#8217;t do so because I was in the money. Simply, I did not know whom to email prior to the competition closing. It was only after William posted the 'Cheaters Removed' thread, when I finally found a suitable email address.</p>\n<p>&#8230;.What a mess&#8230;. &nbsp;</p>\n<p>:S</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58666",
      "postDate": "11/22/2014 03:05:31",
      "content": "<p>Jonathan and Andronicus, both of you have my empathy (although being in 2nd place now,&nbsp;admittedly I'm not complaining). &nbsp;That said, I would love to see both&nbsp;of your methods. &nbsp;It would be sad for all your work to go to waste.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58674",
      "postDate": "11/22/2014 08:24:02",
      "content": "<p>For some of you, is your private score better than public score? or is it always less than your public score?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58789",
      "postDate": "11/24/2014 16:56:45",
      "content": "<p>[quote=rakhlin;58401]</p>\n<p>Andy, quite the contrary, the organizers needlessly complicated the problem. Other works don't report performance across subjects putting them on a common scale like here - it is meaningless. Add small and highly unbalanced data, particularly for 2 humans, and the problem can not be generalized well even on per subject basis. I think without post-calibration performance would not outbid 60%. Finally add quite meaningless metric. In practice you're not interested in AUC. For perfect classifier it should be enough to produce no false positives and at least&nbsp;<strong><em>one</em></strong> true positive for every preictal period.&nbsp;</p>\n<p>[/quote]</p>\n<p>I'd disagree. AUC is quite a good metric for this task. If you think of a real application and for ethical reasons it will never be an automated system but a decision support system, then for a decision support tool a probabilistic trend is a reasonable output. AUC measures an overlap of the two distributions, which is indicative of a perceived difference between probabilistic levels of ictal and preictal activity for intended end-users. Computing one AUC across all patients or average of AUCs per patient is a choice connected to the level of robustness required from the tool. There was a lack of consistency from organisers about these aspects, high robustness by the metric but low robustness by test data usage.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58834",
      "postDate": "11/25/2014 00:15:14",
      "content": "<p>[quote=rakhlin;58401]</p>\n<p>I think without post-calibration performance would not outbid 60%.</p>\n<p>[/quote]</p>\n<p>With one of our models, we achieved over 80% on private LB without any post-calibration, use of test data, or model ensembling. &nbsp;We look forward to reporting our results soon.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58854",
      "postDate": "11/25/2014 06:01:20",
      "content": "<p>[quote=Andy;58789]</p>\n<p>AUC measures an overlap of the two distributions, which is indicative of a perceived difference between probabilistic levels of ictal and preictal activity for intended end-users.[/quote]</p>\n<p>Indeed, AUC measures an overlap of the two distributions. But it does not tell whether absolute value of probability prediction has any good use at all unless&nbsp;<em>properly</em> calibrated - another task not related to AUC metric. You can reach AUC=1 and still be unable practically interpret the output because AUC does not care about true boundary between the distributions. See <a href=\"https://www.kaggle.com/c/seizure-prediction/forums/t/10945/congratulations-to-the-winners/58289#post58289\">here</a>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58865",
      "postDate": "11/25/2014 09:41:38",
      "content": "<p>Well, if you talk about an operating point, then it does not matter. Have you ever observed a trend of say probability in time? What you perceive is not an absolute value, but decays and rises with respect to background.</p>\n<p>A separate problem that you may have 0.499999 and 0.511111, with the AUC = 1 but the boundary will not be perceivable. It is true in theory, but given a real problem your scores are either normally distributed&nbsp;likelihoods or gamma distributed posteriors.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58870",
      "postDate": "11/25/2014 09:47:12",
      "content": "<p>[quote=Drew Abbot;58834]</p>\n<p>[quote=rakhlin;58401]</p>\n<p>I think without post-calibration performance would not outbid 60%.</p>\n<p>[/quote]</p>\n<p>With one of our models, we achieved over 80% on private LB without any post-calibration, use of test data, or model ensembling. &nbsp;We look forward to reporting our results soon.</p>\n<p>[/quote]</p>\n<p>We&nbsp;didn't perform test calibration in any of our models, and our result is between 79% and 80%</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58944",
      "postDate": "11/25/2014 19:06:40",
      "content": "<p>[quote=Andy;58865]</p>\n<p>Well, if you talk about an operating point, then it does not matter. Have you ever observed a trend of say probability in time? What you perceive is not an absolute value, but decays and rises with respect to background.[/quote]</p>\n<p>Imagine a model that scores all available data [0...0.1] (interictal) or [0.9...1] (preictal). For a new data it returns 0.2. We'll have no idea how to interpret that. Moreover, label can be anything, preictal or interictal, it won't change previous AUC=1. Given limited data absolutely possible scenario, particularly &nbsp;for a problem like this competition. The problem becomes even more general if a model's score isn't restricted to [0 1]. This is why binary metric makes a sense.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "59336",
      "postDate": "12/01/2014 16:40:43",
      "content": "<p>Summary of my solution:</p>\n<p>Feature Models:</p>\n<p>My best submission according to the LB score is based on single window model. All the data is first re-sample to 100Hz to reduce high frequency noise. Then every data file is split into 12 parts, about 50 seconds each. For each part of the split data, FFT is applied to transform the data to frequency domain. The power magnitudes in the frequency band from 1 to 50 Hz were selected and converted to logarithmic scale, then resample the frequency band to 18 bins to further reduce noise. The covariance and eigenvalues of the reduced frequency band across channels are also added as features, along with the covariance and eigenvalues in time domain.</p>\n<p>Classifier Models</p>\n<p><br>Several common classifiers in scikit-learn package have been tested, such as random forest, gradient tree boosting, support vector machine. Most of them had really good CV score for individual subject. But did not get good score in LB. The gaps between CV score and LB score were very big. One of the reason is that LB score is across all subject. Other possible reason is due to overfitting. Platt scaling was also added to calibrate the prediction across subjects. It improved LB score slightly.<br>My best submissions according to the LB score were based on support vector machine with RBF kernel, which produced better results because of more control in balancing bias and variance. The classifier gave an estimate for each 50-second window. An averaging method was used to combine the results and provide estimate for the whole 10-minute period. Several averaging methods, such as arithmetic average, geometric average and harmonic average, have been tested. Arithmetic average was best suited for evenly distributed estimate while harmonic average was best suited for oddly distribution. Therefore, a combined averaging method was used. A percentage projection was also used to align results across subjects based on the assumption that the test dataset was similar to the training dataset.</p>\n<p>The repository is available at <a href=\"https://github.com/jlnh/SeizurePrediction\">https://github.com/jlnh/SeizurePrediction</a></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "59348",
      "postDate": "12/01/2014 18:30:10",
      "content": "<p>Thanks @Birchwood! Please also attach the repo&nbsp;to&nbsp;your team's Github section&nbsp;(https://www.kaggle.com/c/seizure-prediction/github).</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "59374",
      "postDate": "12/02/2014 04:51:01",
      "content": "<p>@Birchwood: I want to implement your model.</p>\n<p>Can you please provide me more details like.</p>\n<p>1.&nbsp;What each python code of yours is doing.</p>\n<p>2. Can you represent the entire flow of code graphically, so that we get to know more about it.</p>\n<p>Sorry, if i am asking for more.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "164846",
      "postDate": "03/02/2017 14:46:59",
      "content": "<p>Hello people, i so details about past competitions and i am interested when will it be a new competition! I have epilepsy so i have to search for <a href=\"https://pharmacyreviews.md/\">pharmacy reviews</a> often because of my needs. but thanks anyway and have a good day!! </p>",
      "rawMarkdown": "Hello people, i so details about past competitions and i am interested when will it be a new competition! I have epilepsy so i have to search for <a href=\"https://pharmacyreviews.md/\">pharmacy reviews</a> often because of my needs. but thanks anyway and have a good day!!",
      "votes": null
    },
    {
      "id": "593000",
      "postDate": "08/06/2019 05:10:28",
      "content": "<p>I'm curious if this repo has moved to somewhere? Did anyone make a backup/mirror of the winners solutions? Thanks.</p>",
      "rawMarkdown": "I'm curious if this repo has moved to somewhere? Did anyone make a backup/mirror of the winners solutions? Thanks.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 58250,
      "author_name": "michaelhills",
      "author_url": "",
      "post_date": "11/18/2014 01:05:58",
      "content": "<p>I was also very impressed by Jonathan Tapson taking the lead with&nbsp;so few submissions earlier on in the competition&nbsp;and&nbsp;similarly impressed by Medrr's jump to 0.90!&nbsp;An amazing effort.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58251,
      "author_name": "martinmolina",
      "author_url": "",
      "post_date": "11/18/2014 01:22:33",
      "content": "<p>This was my first ML implementation and I found the discourse on the forums highly stimulating. Thank to everyone who partook!</p>\n<p>I'm looking forward to dissecting the winning models.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58252,
      "author_name": "finlaymaguire",
      "author_url": "",
      "post_date": "11/18/2014 01:23:16",
      "content": "<p>Yeah congrats everyone especially the winners, that was a hell of a challenge.</p>\n<p>I can't wait to find out what strategy everybody else used, would be really interested to know the subject AUC breakdown people were getting too.</p>\n<p>For pretty much every feature we tried the t-SNE plot of Patient 2 showed the test and training data as totally different! &nbsp;Our best subject specific classifier for Patient 2 tended be those trained on the noisiest untransformed features or the least cleaned datasets which kind of makes me think they weren't actually working but the noise meant we&nbsp;weren't as misleadingly training that model.</p>\n<p>Ah well, it was great fun, we are over the moon with our result&nbsp;as it stands!</p>\n\n<p>(We were only 20th but if anyone is interested&nbsp;we&nbsp;used a standard SVC with RBF kernel with random forest&nbsp;feature selection on a feature set composed of cleaned (harmonics and high-pass filter) Common Spatial Patterns basis transform correlation coefficient eigenvalues, and cleaned Independent Component Analysis transformed&nbsp;(FastICA):</p>\n<p>Power spectral density logf correlation&nbsp;coefficients<br>Low-gamma phase sync<br>Power in band<br>Power spectral density logf Broad Band<br>Partial Directed Coherence of the coefficients for a&nbsp;Multivariate Autoregression model&nbsp;&nbsp;<br>We also did some meta-bagging of this classifier and a couple of others with slightly different feature sets.)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58255,
      "author_name": "jontapson",
      "author_url": "",
      "post_date": "11/18/2014 02:07:50",
      "content": "<p>That was fun but very frustrating towards the end.&nbsp; For what it's worth, I used only spectral features (fft over 30s and 60s intervals, overlapped 50%).&nbsp; The secret weapon was linear regression - you could get to about 0.85 on the leaderboard with straight LR, post scaled through a logistic function to (0,1) interval.&nbsp; No networks, trees, SVMs, RBMs, etc. required.&nbsp; Because LR is superfast and can be inverted (i.e. you can back-process to see what features it is scaling up, and what features it ignores)&nbsp; I found some good feature sets.&nbsp; It is also very hard to overtrain.&nbsp; The general best featuresets were 1Hz bands from 0-50 Hz and then 5-10Hz bands up to 180Hz.&nbsp; All data was filtered for 60Hz + harmonics.&nbsp; I did use some networks (ELM ensembles) to get a couple of extra points.&nbsp; The final solution was computed on my laptop&nbsp; (a 2010 MacBook Air).&nbsp; No AWS required.&nbsp; Now I just hope I can reproduce the damn result for the organizers!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58259,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "11/18/2014 02:38:07",
      "content": "<p>[quote=Jonathan Tapson;58255]</p>\n<p>The final solution was computed on my laptop&nbsp; (a 2010 MacBook Air).&nbsp; No AWS required.&nbsp; Now I just hope I can reproduce the damn result for the organizers!</p>\n<p>[/quote]</p>\n<p>This makes the best commercial of Mac Air :D</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58260,
      "author_name": "jontapson",
      "author_url": "",
      "post_date": "11/18/2014 02:48:01",
      "content": "<p>It sounds crazy but actually the Mac Air has a solid state drive, and in these computations with big data sets sometimes disk read/write is what dominates the processing time.&nbsp; Towards the end the feature sets were taking a couple of hours to compute though, it was dumb to carry on with the Air at that point.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58261,
      "author_name": "michaelhills",
      "author_url": "",
      "post_date": "11/18/2014 02:49:05",
      "content": "<p>Looks like simplicity is key. :) I was getting stressed out&nbsp;in this competition with how much of a complete mess my solution had become!&nbsp;Can you elaborate on your linear regression with logistic function setup?&nbsp;What did you use for training labels?</p>\n<p>Late in the competition I realised that my main problem was properly tuning SVM parameters to work with my features. Using one set of C/gamma I could use random feature selection (e.g. randomly mask 50%) and then ensemble the results of 10 such masks to see an improvement (from ~0.796 to ~0.829 leaderboard). Further improvement could be made using genetic algorithm to do that selection (0.84+), but results varied wildly with minor changes in runs and was highly unstable. A while later I discovered that I could match the random mask ensembling by using a different set of C/gamma and just using the features in whole. With those parameters ensembling random masks made it worse.</p>\n<p>Then on the last day I tried Logistic Regression for Dog_5 instead of SVM and got another huge improvement, but we aren't allowed &quot;if Dog_5&quot; so couldn't use it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58263,
      "author_name": "jontapson",
      "author_url": "",
      "post_date": "11/18/2014 03:00:00",
      "content": "<p>I used 0's for interictal and 1's for preictal and those were the target values for the LR; the LR weights were computed using a regularized pseudoinverse.&nbsp; The test set values were then normalised by subtracting the mean and dividing by standard deviation (for the whole test set per subject).&nbsp; That gave a set of values with mean 0, so then just put the values into 1/(1+e^-k.values)&nbsp; where k was a scaling factor (k=0.5, and I can't remember why I used that, probably legacy code).&nbsp; That gave values compressed between 0 and 1 and was pretty close to the normalization methods recommended in a couple of the papers that were posted somewhere on the forum.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58264,
      "author_name": "michaelhills",
      "author_url": "",
      "post_date": "11/18/2014 03:06:12",
      "content": "<p>So to let me check I understand, if I wanted to implement this I would do the following:</p>\n<ol>\n<li>Train linear regression (with feature scaling or without?) with target labels 0 and 1</li>\n<li>Make predictions on cv or test set</li>\n<li>For the array of predictions, subtract mean and divide by standard deviation</li>\n<li>Transform predictions with logistic function</li>\n</ol>\n<p>Does that sound right?</p>\n<p>Edit: I attempted this with my existing features and scored&nbsp;0.83618 on public leaderboard. Wow that beat all my whole feature no ensembling&nbsp;attempts with SVM etc. Then&nbsp;0.84208 if average reference montage is applied to Patient 1 and 2.</p>\n<p>Interestingly the distribution of predictions is completely different to my other submissions yet scores similarly. I wrote a script that counts how many predictions are &lt; 0.5 and how many &gt;1.0 so I could see what is happening in my test set predictions. The test % is how much of the test set was predicted preictal, and train % is how much of the training set consisted of preictal as a rough guide for what I might expect the data split to be like. E.g. in the first one, Dog_1 has 247 interictal predictions and 255 preictal predictions.</p>\n<p>From linear regression run:</p>\n<p>Dog_1 [247, 255] test 50.8% train 4.8%<br>Dog_2 [514, 486] test 48.6% train 7.7%<br>Dog_3 [483, 424] test 46.7% train 4.8%<br>Dog_4 [578, 412] test 41.6% train 10.8%<br>Dog_5 [150, 41] test 21.5% train 6.2%<br>Patient_1 [107, 88] test 45.1% train 26.5%<br>Patient_2 [51, 99] test 66.0% train 30.0%</p>\n<p>From SVM with genetic algorithm feature selection:</p>\n<p>Dog_1 [481, 21] test 4.2% train 4.8%<br>Dog_2 [923, 77] test 7.7% train 7.7%<br>Dog_3 [835, 72] test 7.9% train 4.8%<br>Dog_4 [930, 60] test 6.1% train 10.8%<br>Dog_5 [161, 30] test 15.7% train 6.2%<br>Patient_1 [92, 103] test 52.8% train 26.5%<br>Patient_2 [135, 15] test 10.0% train 30.0%</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58268,
      "author_name": "franklyn",
      "author_url": "",
      "post_date": "11/18/2014 03:49:26",
      "content": "<p>Great work Jonathan, I'm absolutely kicking myself for not trying something as simple as linear regression.</p>\n\n<p>looking forward to the write up and code.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58281,
      "author_name": "jialunhe",
      "author_url": "",
      "post_date": "11/18/2014 06:07:34",
      "content": "<p>Congratulation to all and especially to the winners. The competition is harder than I expected. I started the competition a bit late, made some push in the final two weeks, touched the LB top ten in the final 24 hours. I was struggling in selecting a classifier that can balance bias and variance. I chose SVM for it has some control in setting&nbsp; c and gamma. But I didn't have time to fine tune c and gamma for individual subject.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58282,
      "author_name": "jontapson",
      "author_url": "",
      "post_date": "11/18/2014 06:54:03",
      "content": "<p>@Michael Hills - yes, as you describe. &nbsp;I generally calculated a separate regression on each permutation (fold) of interictal/preictal data and then summed the test set predictions generated by those weights - I also just summed the predictions for each 30s/60s segment to get the outcome for a 10 minute file (which may have been your method from the previous competition?) &nbsp;it is essentially a simplistic logistic regression. &nbsp;I wasn't planning to do much more than some basic feature selection with it, but it turned out to be a good match for this weird situation with very skew training sets and no knowledge of the test set stats (priors). &nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58283,
      "author_name": "skorotkov",
      "author_url": "",
      "post_date": "11/18/2014 07:06:09",
      "content": "<p>[quote=Michael Hills;58264]</p>\n<p>Interestingly the distribution of predictions is completely different to my other submissions yet scores similarly. I wrote a script that counts how many predictions are &lt; 0.5 and how many &gt;1.0 so I could see what is happening in my test set predictions. The test % is how much of the test set was predicted preictal, and train % is how much of the training set consisted of preictal as a rough guide for what I might expect the data split to be like. E.g. in the first one, Dog_1 has 247 interictal predictions and 255 preictal predictions.</p>\n<p>[/quote]</p>\n<p>As I understand the AUC score doesn't care too much (at all) about the absolute values of predictions (as long as all of them are calibrated across the targets).&nbsp; So, &lt;0.5 and &gt;0.5 split doesn't make too much sense.</p>\n<p>I think it did work for SVM just because the skilearn's SVM implementation has the built-in Platt calibration (when called with the probability=True) which pushes the predicted probabilities to absolute values of 0 and 1. On the other hand the LR implementation doesn't perform the calibration and the optimal cut-off value (threshold) is not 0.5.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58286,
      "author_name": "rakhlin",
      "author_url": "",
      "post_date": "11/18/2014 07:50:14",
      "content": "<p>[quote=Sergey Korotkov;58283]</p>\n<p>As I understand the AUC score doesn't care too much (at all) about the absolute values of predictions (as long as all of them are calibrated across the targets).&nbsp; So, &lt;0.5 and &gt;0.5 split doesn't make too much sense.[/quote]</p>\n<p>Indeed, only rank counts. Essentially, the competition was about predicting not probability but ranking. I.e. predicting fewer incorrectly swapped pairs 0/1 as possible.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58288,
      "author_name": "michaelhills",
      "author_url": "",
      "post_date": "11/18/2014 08:47:53",
      "content": "<p>I don't have a full handle on ROC AUC yet,&nbsp;so let me see if I've got this right. As long as every prediction for preictal has a higher value than every prediction for interictal, you will score 1.0? I.e all about ranking as you say. I remember reading something like that one way to think of ROC AUC was, given a random preictal/interictal pair, the value for the AUC is the chance of correctly identifying which is which. Is that the right way to think about it?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58289,
      "author_name": "skorotkov",
      "author_url": "",
      "post_date": "11/18/2014 09:01:48",
      "content": "<p>I think yes.&nbsp; A simple sample:</p>\n<p><code>import numpy as np</code><code><code></code></code></p>\n<p><code>from sklearn.metrics.metrics import roc_auc_score</code></p>\n<p><code><code>p = np.array([0.1, 0.2, 0.3, 0.4, 0.7, 0.8, 0.9])</code></code></p>\n<p><code><code>y</code></code><code> = np.array([0, 0, 1, 1, 1, 1, 1])</code><code></code></p>\n<p><code>roc_auc_score(y, p)<br></code></p>\n<p><code>Out[5]: 1.0<br></code></p>\n\n<p>Note, the cut-off is not 0.5</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58290,
      "author_name": "rakhlin",
      "author_url": "",
      "post_date": "11/18/2014 09:23:22",
      "content": "<p>[quote=Michael Hills;58288]</p>\n<p>I don't have a full handle on ROC AUC yet, so let me see if I've got this right. As long as every prediction for preictal has a higher value than every prediction for interictal, you will score 1.0? [/quote]</p>\n<p>Yes.</p>\n<p>And from my experience, although it needs verification, AUC == fraction of correctly ordered pairs out of all 0/1 pairs.</p>\n<p>You may try your features with some ranking machine, e.g. SVMrank or rank booster. I tried both and got good results on per subject basis. Unfortunately didn't figure out a good method to merge them.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58300,
      "author_name": "pakozm",
      "author_url": "",
      "post_date": "11/18/2014 10:40:28",
      "content": "<p>First of all, congratulations to the winners :-) and to the organization of Kaggle and American society of epilepsy.&nbsp;This has been a very nice competition, with very&nbsp;encouraging and motivational discussions in the forum.</p>\n<p>Regarding the final results, I think the key is in the features, and FFT filters. We decided to use standard filtering, as in the PLOS ONE paper referenced by competition host. Using logistic regression over this features leads to a very bad LB result (0.678), but very high in CV (0.934). However, a deep ANN using same features (with 5 layers) achieves a LB result of 0.794 and CV of 0.928.&nbsp;</p>\n<p>We didn't received any prize (just 7th place), but our results were&nbsp;very consistent between public and private LB (0.82488 and 0.79347). I think it worth the effort to explain briefly our system, but I just advance that we tried a lot of different features and models, and finally a little gain was obtained by making a linear combination of our ideas.</p>\n<p>The key to improve results is the combination of several&nbsp;models with different features and different hyper-parameters, and to use PCA or ICA for decorrelation of data. We feel that&nbsp;consistency of our results has been increased by system combination approach. We observe that AUC variance in LB using different combinations of models was very similar.</p>\n<p>In brief, the pipeline is composed of the following feature sets:</p>\n<p>1) FFT features: computed by using hamming windows of 60s with an overlapping of 50%.&nbsp;Over FFT the filter bank of PLOS ONE paper has been computed. In order to decorrelate data, we use a PCA (and also ICA) transformation, which allows to win 0.04 points of AUC (from 0.7489 to 0.78153 in an ANN with 2 layers).</p>\n<p>2) Eigen values of correlation matrix between channels, and computed&nbsp;for the same 60s with overlap windows (similar to what Michael Hills has been done in previous competition).</p>\n<p>3)&nbsp;Eigen values of correlation matrix between channels for the entire 10 minutes signal.</p>\n<p>4)&nbsp;Eigen values of correlation matrix between channels for the entire 10 minutes of the signal after differentiation.</p>\n<p>5) A bunch of selected features (mean and variances) over the 10 minutes original signal.</p>\n<p>By using this feature sets, different models have been estimated:</p>\n<p>a)&nbsp;ANN with 2 layers over features 1 (using PCA decorrelation) and features 2</p>\n<p>b) Deep&nbsp;ANN with 5 layers over features 1 (using PCA decorrelation) and features 2</p>\n<p>c)&nbsp;ANN with 2 layers over features 1 (using ICA&nbsp;decorrelation) and features 2</p>\n<p>d)&nbsp;K-nearest-neighbor over&nbsp; features 1 (using PCA decorrelation) and features 2</p>\n<p>e)&nbsp;K-nearest-neighbor over&nbsp; features 1 (using ICA&nbsp;decorrelation) and features 2</p>\n<p>f) K-nearest-neighbor over features 3 and 4</p>\n<p>g) K-nearest-neighbor over features 5</p>\n<p>We find that KNNs achieves better results than logistic regression, and that ANNs were even better if you can estimate properly&nbsp;the learning rate, momentum and regularization. Also, the difference between ICA and PCA wasn't significant, both approaches obtain similar results in AUC, but they were important to improve convergence of ANN models.</p>\n<p>Finally, we decided to use a Bayesian model combination (BMC)&nbsp;of all the seven models (a-g). The BMC increases our AUC by 0.03 points in both private and public LB. It is possible to observe that BMC is not achieving a huge improvement, but&nbsp;we feel more comfortable and sure about the consistency of our result.</p>\n<p>I hope someone find this explanation interesting. We will try to publish our system&nbsp;pipeline, just&nbsp;for reproducibility and disemination.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58302,
      "author_name": "rakhlin",
      "author_url": "",
      "post_date": "11/18/2014 11:01:19",
      "content": "<p>[quote=Jonathan Tapson;58282]</p>\n<p>I generally calculated a separate regression on each permutation (fold) of interictal/preictal data and then summed the test set predictions generated by those weights</p>\n<p>[/quote]</p>\n<p>Jonathan, how many permutations of interictal/preictal data did you use for a given subject and how did you choose these permutations? Does it make difference to calculate several regressions instead of doing single one using all data&nbsp; for a given subject? In the end you average them via summation?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58322,
      "author_name": "stevendu",
      "author_url": "",
      "post_date": "11/18/2014 16:45:57",
      "content": "<p>Congratulation to all. I learned a lots from you guys.</p>\n<p>My approach:</p>\n<p>1)Features:</p>\n<p>Michael Hills's feature pipe which gives&nbsp; 900+ features per analysis window</p>\n<p>Each 2s get one analysis window</p>\n<p>Each analysis window contains 4s's data(4*400*Channels for Dog_n and 4*5000*Channels for human),</p>\n<p>Eg for Dog_1&nbsp;&nbsp;&nbsp; EEG&nbsp; 600*400*16&nbsp; =&gt; 298*1024</p>\n<p>(Samples*FeaturesSize)</p>\n<p>2)Models: 1 layer neural network with 20 nodes,same setup for all subjects</p>\n<p>3) Training: Training&nbsp; 50 to 70 models with [positive-class, negative-class(random and shuffle) ] data pairs for each subject</p>\n<p>4)Scoring: majority vote</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58337,
      "author_name": "jontapson",
      "author_url": "",
      "post_date": "11/18/2014 19:06:48",
      "content": "<p>@rakhlin: I used the fact that the preictal data needed to be separated in the sequence blocks of 6 to determine the folding method. &nbsp;The method goes like this:</p>\n<p>1. &nbsp;Find number of preictal blocks of 6 (e.g. dog 1 has 24/6 = 4 blocks)&nbsp;</p>\n<p>2. &nbsp;Divide interictal into the same number of blocks (dog 1 has 4 blocks of 120 each)</p>\n<p>3. perform leave-one-block-out folding for all permutations of interictal and preictal blocks (16 permutations here). &nbsp;Generate a solution for each permutation and sum (average) the results.</p>\n<p>This method has two weaknesses: it leaves out some data for subjects where the numbers are not divisible by 6. &nbsp;It also mixes up preictal blocks in dog 4 where the sequences are not complete blocks of 6. &nbsp;I am kicking myself for not doing that subject properly, it might have given me another .005...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58342,
      "author_name": "rakhlin",
      "author_url": "",
      "post_date": "11/18/2014 20:03:18",
      "content": "<p>Thank you &nbsp;Jonathan.</p>\n<p>I don't fully understand the goal of these permutations. Leave-one-block-out folding makes a big sense for cross-validation. But for ultimate solution - because you feed all data to deterministic(?) linear model equally and average the result in the end - is it in principal different from training single LR on the whole data set?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58350,
      "author_name": "medial",
      "author_url": "",
      "post_date": "11/18/2014 21:36:23",
      "content": "<p>Thank you all. It was not easy. I will be happy to give some information about my strategy for this competition in the near future.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58351,
      "author_name": "mahi83",
      "author_url": "",
      "post_date": "11/18/2014 21:41:49",
      "content": "<p>[quote=Jonathan Tapson;58255]</p>\n<p>That was fun but very frustrating towards the end.&nbsp; For what it's worth, I used only spectral features (fft over 30s and 60s intervals, overlapped 50%).&nbsp; The secret weapon was linear regression - you could get to about 0.85 on the leaderboard with straight LR, post scaled through a logistic function to (0,1) interval.&nbsp; No networks, trees, SVMs, RBMs, etc. required.&nbsp; Because LR is superfast and can be inverted (i.e. you can back-process to see what features it is scaling up, and what features it ignores)&nbsp; I found some good feature sets.&nbsp; It is also very hard to overtrain.&nbsp; The general best featuresets were 1Hz bands from 0-50 Hz and then 5-10Hz bands up to 180Hz.&nbsp; All data was filtered for 60Hz + harmonics.&nbsp; I did use some networks (ELM ensembles) to get a couple of extra points.&nbsp; The final solution was computed on my laptop&nbsp; (a 2010 MacBook Air).&nbsp; No AWS required.&nbsp; Now I just hope I can reproduce the damn result for the organizers!</p>\n<p>[/quote]</p>\n\n<p>I know these questions might seem dumb, but it would be nice if you could clarify a bit more as I am trying to learn about this better. When you say you overlapped 50% what do you mean? &nbsp;Secondly what do you mean by &quot;1Hz bands from 0-50 Hz and then 5-10Hz bands up to 180Hz. All data was filtered for 60Hz + harmonics&quot; . Just a little more clarification on this would do me wonders.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58357,
      "author_name": "jontapson",
      "author_url": "",
      "post_date": "11/19/2014 00:00:55",
      "content": "<p>If using say 60s intervals, the 600s of data were partitioned into 10 blocks (with partitions at n.60s for n=1:9).&nbsp; Then another 10 blocks were generated by starting 30s into the data, with partitions at 30s+n.60s. (the first 30s was joined with the last 30s to make the tenth block here).</p>\n<p>The FFT generates lines at intervals of Fs/N where Fs is the sampling frequency and N is the data length (say 400 Hz and 24k samples for a 60s block).&nbsp; So there would be 24k/400=60 fft lines in any 1Hz interval.&nbsp; So to generate a band for 0-1Hz you add the power (log(abs(fft)) for those 60 lines in that interval.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58358,
      "author_name": "jontapson",
      "author_url": "",
      "post_date": "11/19/2014 00:04:10",
      "content": "<p>@rahklin: In principle one model should be as good as the folded model.&nbsp; In practice it seems that the leave-one-out ensemble gives better generalisation, I think because you are effectively bagging.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58364,
      "author_name": "bbrinkm",
      "author_url": "",
      "post_date": "11/19/2014 03:41:06",
      "content": "<p>Congratulations everyone. This has been an excellent competition on a very difficult problem, and you've come up with a great set of solutions.&nbsp;</p>\n<p>On behalf of the organizers and sponsors, thank you all for your great work on this competition.</p>\n<p>I realize there are a few outstanding questions, and I promise detailed answers will be forthcoming once I've confirmed some details with the group here.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58368,
      "author_name": "jontapson",
      "author_url": "",
      "post_date": "11/19/2014 03:57:30",
      "content": "",
      "votes": null,
      "replies": []
    },
    {
      "id": 58371,
      "author_name": "idaniboy",
      "author_url": "",
      "post_date": "11/19/2014 05:23:25",
      "content": "<p>Could the admins (or anyone else knowledgeable about the performance metrics used in this competition) comment on how the competition results compare to prior &quot;state of the art&quot; algorithms that have previously been published; e.g.:</p>\n<p style=\"padding-left: 30px\">&#8226; Howbert JJ, Patterson EE, Stead SM, Brinkmann B, Vasoli V, Crepeau D, Vite CH, Sturges B, Ruedebusch V, Mavoori J, Leyde K, Sheffield WD, Litt B, Worrell GA (2014) Forecasting seizures in dogs with naturally occurring epilepsy. PLoS One 9(1):e81920.<br>&#8226; Cook MJ, O'Brien TJ, Berkovic SF, Murphy M, Morokoff A, Fabinyi G, D'Souza W, Yerra R, Archer J, Litewka L, Hosking S, Lightfoot P, Ruedebusch V, Sheffield WD, Snyder D, Leyde K, Himes D (2013) Prediction of seizure likelihood with a long-term, implanted seizure advisory system in patients with drug-resistant epilepsy: a first-in-man study. LANCET NEUROL 12:563-571.</p>\n<p>The subjects and estimators of performance are different in this competition versus these papers, so comparing&nbsp;the different systems may be like comparing apples to oranges. However,&nbsp;it would be very&nbsp;impressive&nbsp;if this competition's winners have outperformed the best that academic research has had to offer these last few years.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58372,
      "author_name": "weiwunyc",
      "author_url": "",
      "post_date": "11/19/2014 05:35:39",
      "content": "<p>Second @idaniboy's request. I was thinking about asking the same question later on.</p>\n<p>It would be interesting to see the performance comparison of the best of the new/fresh/outsider approaches vs. the results from well known approaches from domain experts.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58373,
      "author_name": "michaelhills",
      "author_url": "",
      "post_date": "11/19/2014 05:46:41",
      "content": "<p>I am also interested. I've read a lot of research papers over these two competitions, with many talking about non-linear features or other things like phase synchrony, bispectrum, Lyapunov exponent, autoregressive mode and so on. My own experience with these competitions has shown that spectral features have worked the best. FFT bins and spectral entropy combined with cross correlation in time and frequency domains are my core features. I will make some submissions later to compare the strength of the multivariate correlation features and the univariate spectral features.</p>\n<p>I did try AR model, detrended fluctuation analysis and phase synchrony and none&nbsp;improved my model. I appeared to get a very very small bonus from petrosian fractal dimension, hurst exponent, and higuchi fractal dimension, but need to do more testing.</p>\n<p>However it's hard to be conclusive on anything as I felt my setup was very sensitive to SVM parameter tuning and throwing in more features might have required some retuning to squeeze the performance out. Additionally I implemented some of these things myself because performance was too slow for some libraries I tried to use. So there is also the question of whether I implemented it correctly and I didn't have much way of unit testing my maths. I will try some experiments a bit later using Jonathan Tapson's linear regression technique to get a rough gauge of the individual predictive power of my features.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58384,
      "author_name": "julesvanligtenberg",
      "author_url": "",
      "post_date": "11/19/2014 09:15:31",
      "content": "<p>Congratulations to Nir, Jonathan and Andronicus1000 ! Nice work!</p>\n<p>Also Congratulations to the other top-ten finishers as they win a chance to publish their solutions.</p>\n<p>I have one question for the top-ten finishers:</p>\n<p>How much time did you spend roughly on this competition?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58389,
      "author_name": "andtem2000",
      "author_url": "",
      "post_date": "11/19/2014 10:29:39",
      "content": "<p>[quote=idaniboy;58371]</p>\n<p>Could the admins (or anyone else knowledgeable about the performance metrics used in this competition) comment on how the competition results compare to prior &quot;state of the art&quot; algorithms that have previously been published;&nbsp;</p>\n<p>[/quote]</p>\n<p>Everyone is happy, so I'd like to dilute it with some criticism :) There is a fundamental difference. Usage of test data. No offence, I congratulate&nbsp;Jonathan and other, in the given pre-specified framework he did an excellent work, but unfortunately I have to 'congratulate' also organisers and data providers, in particular, 'cause apparently in the follow-up paper 10 genuine technical contributions will be reported :))) In the area with several decade history such as seizure prediction, logistic regression on FFT features is a definite step forward. A jump :)) If that's what data providers wanted, to trade solution complexity for an ability to report 83% performance instead of 70-75. I hope they will not occidentally 'forget' to mention that in the paper and the consequences of that decision. &nbsp;I just wonder what was the thinking behind? Did they realise the practical consequences of that decision? That the solutions will be practically useless. Do organiser understand the difference between unlabelled data and unseen data, 'cause they report it as allowing 'semi-supervised learning' :) It has nothing to do with semi-supervised learning. Unseen data imply not just unseen labels. Personally, I realised that the testing data were allowed to be used only when I saw the winner's jump from 86 to 90. Looking forward to seeing other winners' ways to normalise the output probabilities :)&nbsp;</p>\n\n<p>Jonathan, can you report the performance of your method if testing data were not allowed to be used? Just out of curiosity.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58395,
      "author_name": "golondrina",
      "author_url": "",
      "post_date": "11/19/2014 12:40:17",
      "content": "<p>Congratulations to everyone and thank you for sharing your approaches! I used convolutional neural networks and finished in 10th place, which makes me euphoric:) I regret of not trying something simpler, although I think convnets are quite elegant. Did anybody else use them? I posted my code and description&nbsp;<a href=\"https://github.com/IraKorshunova/kaggle-seizure-prediction#seizure-prediction-using-convolutional-neural-networks\">here</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58396,
      "author_name": "udibr1",
      "author_url": "",
      "post_date": "11/19/2014 12:51:47",
      "content": "<p>congratulations to the winners !</p>\n<p>my code for taking an existing submission and making from it a new submission (with &nbsp;higher score) using information I found in the test data</p>\n<p>https://github.com/udibr/seizure-detection-boost</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58400,
      "author_name": "pakozm",
      "author_url": "",
      "post_date": "11/19/2014 13:23:45",
      "content": "<p>Hi @golondrina. We tried CNNs at the begining but just throw away the idea. I feel that first layer filters can be well estimated, but output probabilities not. However, our best model was a&nbsp;deep neural network with 5 hidden layers, 128 relus in every layer, and one logistic output layer.&nbsp;So, the ANN was trained using a sliding window over input data, that is, the whole ANN is a non-linear kernel which is applied over temporal axis. The outputs of ths kernel are probabilities and we combine them using just a geometric mean of the complementary probability. We are preparing the code,&nbsp;cleaning, removing dead code parts (not used), etc. You can find a more extend explanation of our system in the forum:</p>\n<p>https://www.kaggle.com/c/seizure-prediction/forums/t/10945/congratulations-to-the-winners/58300#post58300</p>\n<p>And we will publish the code as soon as possible.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58401,
      "author_name": "rakhlin",
      "author_url": "",
      "post_date": "11/19/2014 13:25:23",
      "content": "<p>Andy, quite the contrary, the organizers needlessly complicated the problem. Other works don't report performance across subjects putting them on a common scale like here - it is meaningless. Add small and highly unbalanced data, particularly for 2 humans, and the problem can not be generalized well even on per subject basis. I think without post-calibration performance would not outbid 60%. Finally add quite meaningless metric. In practice you're not interested in AUC. For perfect classifier it should be enough to produce no false positives and at least&nbsp;<strong><em>one</em></strong> true positive for every preictal period.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58402,
      "author_name": "rakhlin",
      "author_url": "",
      "post_date": "11/19/2014 13:39:53",
      "content": "<p>I tried CNN similar to the one described in LeCun et al &quot;Classification of Patterns of EEG Synchronization for Seizure Prediction&quot; but it failed estimate 1st layer filters - I suspect due to huge amount of channel pairs I used (120) in contrast with original work (only 15)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58403,
      "author_name": "golondrina",
      "author_url": "",
      "post_date": "11/19/2014 13:43:29",
      "content": "<p>@Francisco, it's impressive what you've done! &nbsp;I'm looking forward to see your code, because some details in your explanation I don't understand.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58404,
      "author_name": "pakozm",
      "author_url": "",
      "post_date": "11/19/2014 13:45:38",
      "content": "<p>@rakhlin, it is surprising to me that first layer filters didn't converge. Our approach is basically a&nbsp;deep one filter, and a geometric mean combination of multiple probability ouputs for each file. Depending in&nbsp;your temporal resolution, it is possible to increase the number of samples by 20, or even more,&nbsp;it is a&nbsp;number of samples large enough to train the first layer filters. Which activation function do you used?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58405,
      "author_name": "pakozm",
      "author_url": "",
      "post_date": "11/19/2014 13:47:59",
      "content": "<p>@golondrina sorry for the complexity of the explanation. We tried a lot of different kind of features and different model combinations. We&nbsp;will&nbsp;to put it together in a brief and&nbsp;clear way. However, my post in the forum is not as easy to follow as I would wanted :(</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58408,
      "author_name": "rakhlin",
      "author_url": "",
      "post_date": "11/19/2014 13:57:52",
      "content": "<p>Francisco, now I suspect too little number of samples was another problem. I used tanh activation and L1-regularizer. &nbsp;I'm too looking forward to see your code.</p>\n\n<p>By the way. If you kept 1st layer filters of your CNN it would be instructive to see patterns they found.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58415,
      "author_name": "pakozm",
      "author_url": "",
      "post_date": "11/19/2014 14:29:10",
      "content": "<p>@rakhlin my intuition is that if you want to compute one output probability for each segment file&nbsp;by using a CNN, the major problem relies&nbsp;in the last layer. It is important a clever selection of how the last layer is computed due to lack of data.</p>\n<p>For example, Dog_1 only has around 500 samples, where only 24 are positive. This scenario difficults the&nbsp;learning of this&nbsp;map from CNN top layer features into preictal probability for each segment, IMHO.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58417,
      "author_name": "rakhlin",
      "author_url": "",
      "post_date": "11/19/2014 14:56:32",
      "content": "<p>Francisco, handling unbalanced data is an interesting topic. I don't feel this was a problem in case of CNN as I tried several strategies like oversampling with noise, assigning proportional weights to positive examples in the output layer, training many nets with balanced chunks. It failed generalize not because of unbalanced data but over little data or high dimensionality. In the end I abandoned CNN, &nbsp;then tried denoising autoencoders, and finally concentrated on linear approaches. It's amazing how simple linear models worked out for Jonathan Tapson.</p>\n<table border=\"0\" width=\"100%\" cellspacing=\"0\" cellpadding=\"15\">\n<tbody>\n<tr>\n<td>\n<table align=\"center\">\n<tbody>\n<tr>\n<td width=\"500\">\n<table width=\"100%\" align=\"center\">\n<tbody>\n<tr>\n<td align=\"left\">&nbsp;</td>\n</tr>\n</tbody>\n</table>\n</td>\n</tr>\n</tbody>\n</table>\n</td>\n</tr>\n</tbody>\n</table>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58420,
      "author_name": "cygnids",
      "author_url": "",
      "post_date": "11/19/2014 16:04:39",
      "content": "<p>Congratulations to all the toppers!</p>\n<p>@Jonathan, wonderful to read about your approach, and more importantly the simplicity of your solution. I wonder if you can share what the outcome of your approach was when you used shorter sliding window lengths and smaller time steps? A natural start is to use much smaller window lengths, and I suspect that by using&nbsp; longer window lengths you've effectively filtered out (or muted) spectral variations which accompany natural&nbsp; working, which could potentially hinder discrimination, and for these subjects, perhaps it is even filtering out shorter ictal activity (ie if sufficiently many bursts occur in short period, they are captured, or else not). Again, great job on the approach &amp; win!</p>\n<p>@bbrinkm, If I may, perhaps a naive question... did the preictal clips lead up to clinical seizures only? Or were there instances of subclinical seizures also? And related, were all the interictal clips completely free of any ictal-like onset signature (however sparse)? Many thanks for hosting such an interesting competition.</p>\n<p>Thanks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58425,
      "author_name": "bbrinkm",
      "author_url": "",
      "post_date": "11/19/2014 16:31:07",
      "content": "<p>@Cy - great question. They were electrographic seizures. We didn't have access to video monitoring on the dogs, so we identified seizures in the EEG. The interictal clips were well clear of any seizures. Philosophically, it is well known that epilepsy patients have interictal signatures (spikes, high frequency oscillations, etc.) of epilepsy in their EEG - so no patient's interictal EEG is totally normal. But there shouldn't have been any seizure-like activity in the interictal clips. Artifacts, perhaps - but no seizures.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58430,
      "author_name": "cygnids",
      "author_url": "",
      "post_date": "11/19/2014 17:28:52",
      "content": "<p>@bbrinkm, Thank you for the quick response.To say the least, must have been a laborious task to build this dataset! May I further inquire whether the electrographic patterns were identified in the time-domain or frequency domain? And, if classified in the frequency domain, may I inquire what signature was used for the classification? </p>\n<p>I used much shorter sliding windows to identify the spectral signatures, and wondered if I was picking up short ictal activities in the interictal clips. Hence my question earlier... However, as you note, perhaps the algorithm was conflating artifacts with interictal/preictal activity (though, I'd have thought that the artifacts would results in a different broadband spectral response; different enough to be differentiated). Then again, that's a supposition on my part... I will take a closer look this weekend.</p>\n\n<p>Again, many thanks for addressing our questions! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58449,
      "author_name": "weiwunyc",
      "author_url": "",
      "post_date": "11/19/2014 22:00:46",
      "content": "<p>Congratulations to the winners. This competition turned out to be an interesting and fun competition for me unexpectedly.<br>I have learned and used many tips from the discussions from other Kagglers in this and other competitions. Even though&nbsp;I did not squeeze into top 10, I still would like to share some of my thoughts and approaches which may provide some aspects that have not been&nbsp;used/mentioned by others.&nbsp;The final private LB scores of this competition is so tight that if I had chose my best private&nbsp;score submission, (I chose my 2nd and 3rd best private scores, which is not too bad. The 2nd best is only 0.006&nbsp;less than my best score.), I could have jumped to No. 6 instead of the current No. 11.&nbsp;I hope someone can find some of the points useful. The sharing and learning environment at&nbsp;Kaggle is the best thing I like about Kaggle.</p>\n<p>I did not plan to participate this competition as I saw the data set is pretty big seem like require a lot of time.&nbsp;Just about two weeks before the end, I saw that I could have time in my schedule to participate a competition and at around the same&nbsp;time it was announced that the daily submission limit was increased to 10 which make my late entrance more feasible.&nbsp;I thought I might take a crack at this.</p>\n<p>I used R for this competition. The runs were mostly on my Home PC with i7 CPU and 12GB memory.</p>\n<p>Overall consideration and approach:</p>\n<p>a. Because the limited time available and the significant turn-around time to generate new set of features.&nbsp;I figured that I need to chose a set of features and stick to it. No time to do much feature selection/reduction.</p>\n<p>b. I have no domain knowledge and did not participate in the first Seizue competition. Did not have much time to dive into the literatures. All the features I chose was from browsing through the forum of this and previous Seizure&nbsp;competition. And only choose items that I can understand and quickly implement.</p>\n<p>c. One of the problem I see that I need to deal with at the beginning is the conflict of a performance metrics that&nbsp;includes all subjects and seperately trained models for each individual subjects. Seperately trained models for&nbsp;individual subjects may perform better for the individual subject, however, the probabilities generated from different&nbsp;models can have very different range, magnitude. And because for AUC metric, what matters is the ranking order of the predictions, when you combine results from individual models, because the differences in overall magnitudes or&nbsp;ranges of the predictions, simply combining the results from individual models may not yield best results. I decided&nbsp;at the beginning that for each type of model I use, I will have models trained for each subject using all the features available to that subject and a same type model that are global, trained using data from all subjects using a set of features common to all subjects. My plan was to later use the predictions from the global model to calibrate&nbsp;and combine the outputs of individual models. But I did not get time to explore this and at the end just simply ensemble both&nbsp;global and simply combined individual model results.</p>\n<p>I will discuss my approach on the the following topics: Feature generation, Models used, Ensemble of different models.</p>\n<p>Features:<br>a. To make features more comparable between subject and also to reduce the size, down sample signal data to 200Hz.</p>\n<p>b. There were 3 evolving stages in the progress of feature generation in my approach. Each stage saw significant improvement in LB results.</p>\n<p>c. Summary statistics only features: First, I saw in the post of the previous competition, some of the top winners&nbsp;only used summary statistics of the time series and FFT data. For each mat file, there are N number of channels.&nbsp;For each channel, besides the original signal series, I also generated delta and delta of delta series.&nbsp;For each of these series in time domain, I also generated FFT series. Then for each data series (time-domain and FFT)&nbsp;I generated a few basic summary statistics (mean, max, Stdev). Plus Frequency of FFT peak.&nbsp;</p>\n<p>Then also the overall summary statistics. First, generated statistic for a series of average of all channels.&nbsp;mean, max, stdev for all the quantities across all channels. Then the mean, max , SD and max,2nd max, 3rd max of&nbsp;elements of covariance matrix across all channel for both the time-domain series and FFT series.</p>\n<p>Total 600+ features for dogs, little less number of common features, much more for patient_2. The best public LB results I could get with&nbsp;these features were: 0.657 for single model, 0.687 from ensemble</p>\n<p>d. Summary statistics plus FFT: Adding FFT series is the natural next step. To reduce the size of FFT series,&nbsp;borrowed from the thinking of @ruai at <a href=\"https://www.kaggle.com/c/seizure-prediction/forums/t/10884/what-type-of-models-are-people-using/57672#post57672\">&quot;what-type-of-models-are-people-using&quot;</a>,&nbsp;somewhat different parameters, using only the front portion (0.2) and smoothing the raw FFT to get only 24 points&nbsp;per FFT series. To arrive at the value of these parameters, I just simply plot the FFT series of a few files with&nbsp;different cutoff and averaging settings to get what I feel that are manageable and still kept enough details of&nbsp;the FFT series.<br>2000+ features for dogs, best single model: 0.69284, ensemble: 0.717<br>e. With the tips from this discussion&nbsp;<a href=\"https://www.kaggle.com/c/seizure-prediction/forums/t/10909/increase-number-of-samples-by-using-short-time-segments/57972\">&quot;increase-number-of-samples-by-using-short-time-segments&quot;</a>:</p>\n<p>Break each of the original 10Min series into 10 non-overlap 60s series. The train data increased 10-fold.&nbsp;Used the same features as above for these shorter series. The eventual prediction for each original 10Min series&nbsp;was taken the average or certain percentile of the 10 predictions from the shorter series.<br>Best single model: 0.80514, Best Ensemble: 0.81803</p>\n<p>Model used:<br>a. After the Higgs-Boson competition, xgboost become my favorite tool. Fast, effective.<br>b. For this data set SVM performed very well.<br>c. glmnet is competitive<br>d. also tried gbm and randomForest. These models had much worse results and very slow. I wish I did not waste too much time&nbsp;at the end to run these models instead I should have put more efforts into better tuning the better performed models.<br>My thought was to use as many types of fundamentally different models to ensemble as possible. In hindsight, I probably should ensemble several runs of the better performed models with different parameters. (Not tested.)<br>e. Each model type,&nbsp;both individual model and global model were constructed and tested.</p>\n<p>Ensemble method:<br>a. AUC only depends on the order of the predictions. From previous competitions, I found the most effective way&nbsp;to ensemble a set of results from different models for AUC metric is to convert each solution to ranks and take the&nbsp;average ranks of all the solutions.<br>b. This ensemble method provides about 0.01-0.03 improvements.</p>\n<p>Cross-Validation</p>\n<p>2-fold CV with 10 shuffles, based on what the data description page used as &quot;series&quot;, i.e. mat files with different sequence number but belong to the same series are in and out of folds together. The resulted AUC higher than public LB, but direction wise are pretty good.</p>\n<p>My public and private LB scores are pretty much in sync. My public LB position was 25, private LB position jump to 11.&nbsp;I think my approach of constructing models trained both globally and individually helped in giving the stability of my&nbsp;results.</p>\n<p>This is my experience. Hope you can find something useful.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58450,
      "author_name": "mahi83",
      "author_url": "",
      "post_date": "11/19/2014 22:09:41",
      "content": "<p>@Wei Wu, I would love to see your R code. Is it possible if you put it up?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58451,
      "author_name": "phillipadkins",
      "author_url": "",
      "post_date": "11/19/2014 22:13:31",
      "content": "<p>Oh no, I was looking forward to not having to worry about a writeup / code-cleanup.&nbsp; Andronicusssss!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58456,
      "author_name": "michaelhills",
      "author_url": "",
      "post_date": "11/19/2014 23:07:49",
      "content": "<p>Haha when Andronicus first appeared on the leaderboard it seemed very suspicious, 7 days old account and an immediate 3rd place. I figured it would jump me 2 places on the leaderboard at the end.</p>\n<p>For those discussing unbalanced data problem, I too tried many different techniques like oversampling preictal segments, or training many smaller classifiers on more balanced subsets, e.g. if there's 42 preictal and 500 interictal, split the interictal into groups of 42, train N classifiers and mean their probability outputs. Always made it worse.</p>\n<p>I'm not sure I stated it anywhere so the features I used. I used a non-overlapping 75s window size to generate 8x training samples. Features are:</p>\n<p><span style=\"line-height: 1.4\"><span style=\"line-height: 1.4\"></span></span></p>\n<ol>\n<li><span style=\"line-height: 1.4\">cross correlation in time domain</span></li>\n<li><span style=\"line-height: 1.4\">cross correlation in frequency domain</span></li>\n<li><span style=\"line-height: 1.4\">hand-picked frequency bins&nbsp;[0.5, 2.25, 4, 5.5, 7, 9.5, 12, 21, 30, 39, 48], where the frequency magnitude is taken as a mean for each hz-range pair, then apply log10</span></li>\n<li><span style=\"line-height: 1.4\">power-in-band spectral entropy for various bands in the range of 0.5-24Hz</span></li>\n<li><span style=\"line-height: 1.4\">higuchi fractal dimension</span></li>\n<li><span style=\"line-height: 1.4\">petrosian fractal dimension</span></li>\n<li><span style=\"line-height: 1.4\">hurst exponent</span></li>\n</ol>\n<p><span style=\"line-height: 1.4\"><span style=\"line-height: 1.4\"></span></span></p>\n<p><span style=\"line-height: 1.4\"><span style=\"line-height: 1.4\"><span style=\"line-height: 1.4\">The first 3 seemed to work well. This is basically same features as I used for first competition. The spectral entropy seemed to benefit Dogs 3 and 4 more. The last 3 seemed to help only very slightly.</span></span></span></p>\n<p><span style=\"line-height: 1.4\"><span style=\"line-height: 1.4\"></span></span></p>\n<p><span style=\"line-height: 1.4\"><span style=\"line-height: 1.4\"><span style=\"line-height: 1.4\"><span style=\"line-height: 1.4\">Random feature selection was used on 1, 2 and 3 where 52.5% of features were selected randomly to form a feature mask.&nbsp;I&nbsp;ran genetic algorithm optimising for local CV scores on 3 separate feature groups, the spectral entropies with higuchi fractal dimension, and then petrosian fractal dimension and hurst exponent in their own GA runs. The GA was instructed to seed 30 population with 55% features used. It was run for 10 generations.</span></span></span></span></p>\n<p><span style=\"line-height: 1.4\"><span style=\"line-height: 1.4\"></span></span></p>\n<p>Finally the best 2 feature masks from each GA feature group, and 2 randomly generated feature marks for the random groups, were combined. With all the features combined, I then had 2 different masks across the features. I trained each one, and then averaged the predictions. Also because I was using 75s window size, I averaged the predictions for the 8 windows in each segment.</p>\n<p><span style=\"line-height: 1.4\"><span style=\"line-height: 1.4\"></span></span></p>\n<p><span style=\"line-height: 1.4\">Later I learned that all this boost I was getting from GA (from 0.79 to 0.84) was mostly because my SVM parameters were not tuned correctly for Dog 3 and 4 and I could get 0.83 without any genetic algorithm or random feature selection, but plain simple all-features one SVM with different parameters. Then I tried Tapson's approach post-competition and got 0.84 easily using my features with linear regression. :)</span></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58458,
      "author_name": "rakhlin",
      "author_url": "",
      "post_date": "11/19/2014 23:51:05",
      "content": "<p>[quote=Wei Wu;58449]</p>\n<p>Ensemble method:</p>\n<p>a. AUC only depends on the order of the predictions. From previous competitions, I found the most effective way&nbsp;to ensemble a set of results from different models for AUC metric is to convert each solution to ranks and take the&nbsp;average ranks of all the solutions.</p>\n<p>....</p>\n<p>This is my experience. Hope you can find something useful.</p>\n<p>[/quote]</p>\n\n<p>Thank you, very interesting. Could you please elaborate how do you combine individual ranks.&nbsp;Particularly what does it mean:&nbsp;&quot;....&nbsp;and take the average ranks of all the solutions.&quot; From my experience&nbsp;different&nbsp;subjects demonstrate&nbsp;large&nbsp;difference in rank&nbsp;deviation&nbsp;and mean.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58468,
      "author_name": "michaelhills",
      "author_url": "",
      "post_date": "11/20/2014 01:54:56",
      "content": "<p>If anyone wants to try Jonathan's approach with sklearn I wrote a wrapper for LinearRegression that implements predict_proba using his method so you can just slot it in. I'm not 100% sure it's exactly what he did, we can wait for his code to verify. I do see the easy high scores he mentioned it gave although not as high as his but that's probably to do with my windowing and features than the classifier.</p>\n\n<p>class SimpleLogisticRegression(LinearRegression):<br>&nbsp; &nbsp; def predict_proba(self, X):<br>&nbsp; &nbsp; &nbsp; &nbsp; predictions = self.predict(X)<br>&nbsp; &nbsp; &nbsp; &nbsp; predictions = sklearn.preprocessing.scale(predictions)<br>&nbsp; &nbsp; &nbsp; &nbsp; predictions = 1.0 / (1.0 + np.exp(-0.5 * predictions))<br>&nbsp; &nbsp; &nbsp; &nbsp; return np.vstack((1.0 - predictions, predictions)).T</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58480,
      "author_name": "jontapson",
      "author_url": "",
      "post_date": "11/20/2014 05:34:43",
      "content": "<p>[quote=Francisco Zamora-Martinez;58415]</p>\n<p>For example, Dog_1 only has around 500 samples, where only 24 are positive. This scenario difficults the&nbsp;learning of this&nbsp;map from CNN top layer features into preictal probability for each segment, IMHO.</p>\n<p>[/quote]</p>\n<p>I tried some training where I used sets of 50:50 interictal:preictal with neural networks (re-using the small number of preictals several times). &nbsp;They worked OK but no better than anything else...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58489,
      "author_name": "jontapson",
      "author_url": "",
      "post_date": "11/20/2014 05:45:39",
      "content": "<p>[quote=Andy;58389]</p>\n<p>Everyone is happy, so I'd like to dilute it with some criticism :) There is a fundamental difference. Usage of test data. No offence, I congratulate&nbsp;Jonathan and other, in the given pre-specified framework he did an excellent work, but unfortunately I have to 'congratulate' also organisers and data providers, in particular, 'cause apparently in the follow-up paper 10 genuine technical contributions will be reported :))) In the area with several decade history such as seizure prediction, logistic regression on FFT features is a definite step forward. A jump :)) If that's what data providers wanted, to trade solution complexity for an ability to report 83% performance instead of 70-75. I hope they will not occidentally 'forget' to mention that in the paper and the consequences of that decision. &nbsp;I just wonder what was the thinking behind? Did they realise the practical consequences of that decision? That the solutions will be practically useless. Do organiser understand the difference between unlabelled data and unseen data, 'cause they report it as allowing 'semi-supervised learning' :) It has nothing to do with semi-supervised learning. Unseen data imply not just unseen labels. Personally, I realised that the testing data were allowed to be used only when I saw the winner's jump from 86 to 90. Looking forward to seeing other winners' ways to normalise the output probabilities :)&nbsp;</p>\n<p>Jonathan, can you report the performance of your method if testing data were not allowed to be used? Just out of curiosity.&nbsp;</p>\n<p>[/quote]</p>\n<p>I think you may be overestimating the effect of using the test data - in my case it had no effect on the individual (per-subject) outcomes, it was only used to fix the mean and variance of all subjects to be the same, because that was critical for the use of the combined AUC. &nbsp;That (the combined AUC metric) has nothing to do with the effectiveness of the method in practical terms. &nbsp;My method worked because the data is, for once, quite linearly separable and doesn't require to be nonlinearly transformed in a higher dimensional space. &nbsp;That characteristic&nbsp;is independent of whether you use test data or not (and that is why the CNN methods didn't work any better).</p>\n<p>&nbsp;To answer the question: the best I got without using the test data was 0.81806. &nbsp;That was quite early on and I think one could do better, maybe others did?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58492,
      "author_name": "michaelhills",
      "author_url": "",
      "post_date": "11/20/2014 06:28:15",
      "content": "<p>Jonathan when I used your approach I did not use any test data and achieved in the 0.83-0.84 range (public LB) &nbsp;with my features. I cleaned up most of my code so will try to release it as soon as I can but probably won't have much time to finish it up until Sunday night or Monday probably.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58512,
      "author_name": "andtem2000",
      "author_url": "",
      "post_date": "11/20/2014 11:24:31",
      "content": "<p>[quote=Jonathan Tapson;58489]</p>\n<p>I think you may be overestimating the effect of using the test data - in my case it had no effect on the individual (per-subject) outcomes, it was only used to fix the mean and variance of all subjects to be the same, because that was critical for the use of the combined AUC. &nbsp;That (the combined AUC metric) has nothing to do with the effectiveness of the method in practical terms. &nbsp;My method worked because the data is, for once, quite linearly separable and doesn't require to be nonlinearly transformed in a higher dimensional space. &nbsp;That characteristic&nbsp;is independent of whether you use test data or not (and that is why the CNN methods didn't work any better).</p>\n<p>&nbsp;To answer the question: the best I got without using the test data was 0.81806. &nbsp;That was quite early on and I think one could do better, maybe others did?</p>\n<p>[/quote]</p>\n<p>Jonathan thanks for your reply. I assume&nbsp;0.81806 is on private LB, right? I might be overestimating the effect of using test data for this competition score but for practical terms it can not be overestimated. Forget the madness I wrote before explaining the idea on the model-level, lets just consider level of probabilities for simplicity. To simulate a real-life situation the dataset should be say 10-20 patients, with only 1-2 of them with representation of preictal activity and this representation is say 1-2 precital events. So the vast majority of data are background EEG. That's a real-life situation in any EEG-based brain injury problem. What's gonna happen if you apply score normalisation on the test data where there is no preictal activity? Exactly, the normalisation will 'create' this activity from nowhere. This testing-data score normalisation works a) offline and b) with the assumption that preictal activity is present. Neither is pertaining to&nbsp;a real-life SeizPred problem. I'm not criticising the engineering solution perse, just its relation to the real-world problem.&nbsp;</p>\n<p>Michael, what is the&nbsp;0.83-0.84 public LB converted to private LB in your case? There were guys on the fourth place public with 86 falling to 75 in private LB.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58515,
      "author_name": "joseleiva",
      "author_url": "",
      "post_date": "11/20/2014 12:26:32",
      "content": "<p>[quote=Andy;58389]</p>\n<p>Everyone is happy, so I'd like to dilute it with some criticism :) There is a fundamental difference. Usage of test data. No offence, I congratulate&nbsp;Jonathan and other, in the given pre-specified framework he did an excellent work, but unfortunately I have to 'congratulate' also organisers and data providers, in particular, 'cause apparently in the follow-up paper 10 genuine technical contributions will be reported :))) In the area with several decade history such as seizure prediction, logistic regression on FFT features is a definite step forward. A jump :)) If that's what data providers wanted, to trade solution complexity for an ability to report 83% performance instead of 70-75. I hope they will not occidentally 'forget' to mention that in the paper and the consequences of that decision. &nbsp;I just wonder what was the thinking behind? Did they realise the practical consequences of that decision? That the solutions will be practically useless. Do organiser understand the difference between unlabelled data and unseen data, 'cause they report it as allowing 'semi-supervised learning' :) It has nothing to do with semi-supervised learning. Unseen data imply not just unseen labels. Personally, I realised that the testing data were allowed to be used only when I saw the winner's jump from 86 to 90. Looking forward to seeing other winners' ways to normalise the output probabilities :)&nbsp;</p>\n<p>Jonathan, can you report the performance of your method if testing data were not allowed to be used? Just out of curiosity.&nbsp;</p>\n<p>[/quote]</p>\n\n<p>Andy,</p>\n\n<p>I think your post is interesting and can be extended to a more general issue: increasingly complex machine learning methods and frameworks are proposed in the top conferences and research journals, but the competitions are&nbsp; won by people using simple, well established methods, together with a convenient feature design.</p>\n<p>Although I'm in the academia, I'm increasingly skeptical about how the experiments are performed in machine learning papers, where the most convenient dataset and experimental set-up is chosen to give the proposed method an edge to be published (I try to be honest in my own work, though).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58517,
      "author_name": "pakozm",
      "author_url": "",
      "post_date": "11/20/2014 12:37:15",
      "content": "<p>Jose M. I don't agree&nbsp;with your opinion about machine learning community. For instance, you can look into the following competitions, were complex algorithms win:</p>\n<p>- Merck competition:&nbsp;http://blog.kaggle.com/2012/11/01/deep-learning-how-i-did-it-merck-1st-place-interview/</p>\n<p>- Galaxy zoo compeition:&nbsp;http://benanne.github.io/2014/04/05/galaxy-zoo.html</p>\n<p>- Higgs Boson:&nbsp;http://www.kaggle.com/c/higgs-boson/forums/t/10344/winning-methodology-sharing?page=2</p>\n<p>Outside Kaggle there are more competitions where complex models show better performance.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58518,
      "author_name": "joseleiva",
      "author_url": "",
      "post_date": "11/20/2014 12:45:04",
      "content": "<p>Francisco,</p>\n<p>Yes, you are right. Deep methods and convolutional NNs have won some competitions. But still, you could find many competitions won by random forests and boosting.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58519,
      "author_name": "pakozm",
      "author_url": "",
      "post_date": "11/20/2014 12:50:46",
      "content": "<p>Of course,&nbsp;it is not possible to have a general method to solve all the problems. And this is true for every model, so, it is not an statement which allow you to questioning about machine learning community and its research ;-)</p>\n<p>So, I think that we can agree with that it is necessary to try different approaches, in order to check what works better ;-)&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58520,
      "author_name": "andtem2000",
      "author_url": "",
      "post_date": "11/20/2014 13:04:52",
      "content": "<p>[quote=Jose M.;58515]</p>\n<p>I think your post is interesting and can be extended to a more general issue: increasingly complex machine learning methods and frameworks are proposed in the top conferences and research journals, but the competitions are&nbsp; won by people using simple, well established methods, together with a convenient feature design.</p>\n<p>Although I'm in the academia, I'm increasingly skeptical about how the experiments are performed in machine learning papers, where the most convenient dataset and experimental set-up is chosen to give the proposed method an edge to be published (I try to be honest in my own work, though).</p>\n<p>[/quote]</p>\n<p>I'm in the academia as well but I'd not be that categoric in judging PR&nbsp;literature. The competitions are very simplified real-world problems. And sometimes (in this case) some decisions are made to make it even less connected to real-life. A nice contribution to the follow-up paper would be to run the best solutions including logreg (fft) on say publicly available Freiburg intracranial dataset. I bet they would not stand a chance. This study is more on early days statistical significant difference between various features for e.g. clinical neurophysiology journal. Simple approach + simple features + best results = &nbsp;data problem (easy data~=real-life, representation problem train~=test in terms of artifacts, etc). Only organisers/data providers can analyse that. In this case the results are not particularly good, true. But I agree with you in the sense that, it'd be nice to have some platform so that every submitted research paper on the EEG-based SP/SD topic would first have to obtain results on the fixed DB with the fixed performance assessment routine, etc. something similar to UCL&nbsp;Machine Learning Repository.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58533,
      "author_name": "bbrinkm",
      "author_url": "",
      "post_date": "11/20/2014 15:52:13",
      "content": "<p>These are good observations, and we will take into account your suggestions for future competitions. As you can tell from other posts on this forum the use of test data for calibration was a controversial and difficult issue. In the real world problem a seizure prediction device would not have access to future data during training. However, for a competition such as this one, there is no way to hide the test set entirely from the contestants. We have to give you that data a priori in order for the contest to work. We removed the restriction of using test data to calibrate because we realized the prohibition was not really enforceable.</p>\n<p>While not perfect, the results of this competition will be far from useless. Calibrating models on the test data perhaps compensates for the disadvantage contestants have of not being able to do ongoing EEG baseline correction, given the lack of time stamps in the test data. For the paper we plan to run contenders' algorithms on held out data segments, data from entirely new dogs and humans if possible, and this will provide the ultimate test of these approaches. </p>\n<p>Thank you for pointing out the admitted weaknesses of this format, and we do plan to disclose all of this in our paper. We would not want anyone to think we have entirely solved the seizure forecasting problem with this competition - clearly this problem is very difficult and more work is needed. Generally in academic papers better comparison between methods is needed, and this contest is part of that effort via the IEEG project for sharing data and algorithms. Researchers (and we hope manuscript reviewers) will be able to run competing algorithms on the same data sets and compare results directly. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58541,
      "author_name": "andtem2000",
      "author_url": "",
      "post_date": "11/20/2014 16:37:47",
      "content": "<p>[quote=bbrinkm;58533]</p>\n<p>While not perfect, the results of this competition will be far from useless.&nbsp;</p>\n<p>[/quote]</p>\n<p>No doubt in it. In fact, plenty of room for analysis and message formulation.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58561,
      "author_name": "weiwunyc",
      "author_url": "",
      "post_date": "11/20/2014 19:16:18",
      "content": "<p>@Mahi Karim:</p>\n<p>My code is pretty messy and ugly. I will need to clean it up a little bit.</p>\n<p>@rakhlin:</p>\n<p>I may not have explained it clearly in English. A few lines of code probably can explain it better, if you know R a little:</p>\n<p>a=read.csv('submission1.csv')</p>\n<p>b=read.csv('submission2.csv')</p>\n<p>a$rank[order(a[,2])]=(1:nrow(a))/nrow(a)</p>\n<p>b$rank[order(b[,2])]=(1:nrow(b))/nrow(b)</p>\n<p>c=data.frame(clip=a[,1], preictal=(a$rank+b$rank)/2</p>\n<p>write.csv(c, row.names=F, quote=F, file='mixed_submission.csv')</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58563,
      "author_name": "stevendu",
      "author_url": "",
      "post_date": "11/20/2014 19:26:54",
      "content": "<p>[quote=bbrinkm;58533]</p>\n<p>... However, for a competition such as this one, there is no way to hide the test set entirely from the contestants. We have to give you that data ...</p>\n<p>[/quote]</p>\n<p>You can insert noise or fake samples to test data, eg 50%true test 50% dummy, then 20% for pubic lb 30% private lb....</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58568,
      "author_name": "jontapson",
      "author_url": "",
      "post_date": "11/20/2014 20:20:39",
      "content": "<p>[quote=Andy;58520]</p>\n<p>[quote=Jose M.;58515]</p>\n<p>I think your post is interesting and can be extended to a more general issue: increasingly complex machine learning methods and frameworks are proposed in the top conferences and research journals, but the competitions are&nbsp; won by people using simple, well established methods, together with a convenient feature design.</p>\n<p>Although I'm in the academia, I'm increasingly skeptical about how the experiments are performed in machine learning papers, where the most convenient dataset and experimental set-up is chosen to give the proposed method an edge to be published (I try to be honest in my own work, though).</p>\n<p>[/quote]</p>\n<p>I'm in the academia as well but I'd not be that categoric in judging PR&nbsp;literature. The competitions are very simplified real-world problems. And sometimes (in this case) some decisions are made to make it even less connected to real-life. A nice contribution to the follow-up paper would be to run the best solutions including logreg (fft) on say publicly available Freiburg intracranial dataset. I bet they would not stand a chance. This study is more on early days statistical significant difference between various features for e.g. clinical neurophysiology journal. Simple approach + simple features + best results = &nbsp;data problem (easy data~=real-life, representation problem train~=test in terms of artifacts, etc).&nbsp;</p>\n<p>[/quote]</p>\n<p>I'm also in academia and have published ML methods in the journals. &nbsp;My view is that competitions like this are a whole lot more realistic than most of the ML benchmarks, because it's generally real-world data (messy) that someone actually cares about, and the test set is genuinely unseen. &nbsp;If you consider how many thousands of papers have been written on MNIST, where the data is already a significant modification of the original raw data, and there are published algorithms that have obviously been written to address a few recognized problem cases in the test set, then this (Kaggle) is a much more realistic test of methods. &nbsp;Of course, abstracting the real-world problem to a competition requires the organizers to make some simplifications and artificial constructions (like the test set and combined AUC in this competition) but that's unavoidable.</p>\n<p>I came into this competition to demonstrate the value of a particular type of neural network (LSHDI or ELM type), and defeated myself with linear regression. &nbsp;You don't get those outcomes from the regular benchmarks. &nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58570,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "11/20/2014 20:29:12",
      "content": "<p>Geez, you guys are professors?! I can't imagine my boss competing with me (a lousy phd student)..</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58572,
      "author_name": "andtem2000",
      "author_url": "",
      "post_date": "11/20/2014 21:13:24",
      "content": "<p>[quote=rcarson;58570]</p>\n<p>Geez, you guys are professors?! I can't imagine my boss competing with me (a lousy phd student)..</p>\n<p>[/quote]</p>\n<p>It is better not. 'Cause he may accidentally send you one of his submissions, you might accidentally submit it and you both will be disqualified :)) or not. depends on whether you end up in the money. If yes then it's ok :)))</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58574,
      "author_name": "golondrina",
      "author_url": "",
      "post_date": "11/20/2014 21:38:17",
      "content": "",
      "votes": null,
      "replies": []
    },
    {
      "id": 58609,
      "author_name": "julesvanligtenberg",
      "author_url": "",
      "post_date": "11/21/2014 12:29:01",
      "content": "<p>I read a post from a guy who claimed to have helped both Andronicus1000 and Jonathan (he called him Jon). I think it was in this forum (but I could be wrong).<br>He talked about holding the ladder and stuf and also about how Jonathan (Jon) gave a submission (or code files) to Andronicus which Andronicus submitted before he got around to merge with 'Jon'.<br>Does anybody know where this post went?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58639,
      "author_name": "goosemonkey1234",
      "author_url": "",
      "post_date": "11/21/2014 21:27:53",
      "content": "<p>Not sure if asking for a phd position is on topic here :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58655,
      "author_name": "julesvanligtenberg",
      "author_url": "",
      "post_date": "11/21/2014 22:43:28",
      "content": "<p>Congratulations QMSDP and Birchwood with your 2nd and 3rd place!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58656,
      "author_name": "jontapson",
      "author_url": "",
      "post_date": "11/21/2014 22:53:23",
      "content": "<p>I inadvertently shared methods with someone else in my research group, who I did not know was competing at the time. &nbsp;When I realized that, I invited him to join a team (which would have made it legitimate), but we missed the deadline for the merger. &nbsp;You will note that he made no submissions after the deadline, and withdrew his submission when it came in high, so we acted in good faith; but I guess the rules don't allow for that. &nbsp;We are both out of it now. &nbsp;Congratulations to the winners and I will be interested to see their methods.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58659,
      "author_name": "rakhlin",
      "author_url": "",
      "post_date": "11/21/2014 23:11:18",
      "content": "<p>Jonathan please accept my sympathy and please share your solution anyway. Its lucidity fascinates me!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58660,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "11/21/2014 23:18:29",
      "content": "<p>[quote=rakhlin;58659]</p>\n<p>Jonathan please accept my sympathy and please share your solution anyway. Its lucidity fascinates me!</p>\n<p>[/quote]</p>\n<p>+1. It is sad some top 10 score is removed and we may lose those great models. None of you do it on purpose. Please share your wonderful work and let it help the epilepsy society.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58661,
      "author_name": "goosemonkey1234",
      "author_url": "",
      "post_date": "11/21/2014 23:44:48",
      "content": "<p>+2. Indeed very sad that technicalities like merger deadlines should cause removal of what from this thread seems to have been one of the best solutions in&nbsp;the competition.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58663,
      "author_name": "andronicus1000",
      "author_url": "",
      "post_date": "11/22/2014 00:45:58",
      "content": "<p>[quote=Jonathan Tapson;58656]</p>\n<p>I inadvertently shared methods with someone else in my research group, who I did not know was competing at the time. &nbsp;When I realized that, I invited him to join a team (which would have made it legitimate), but we missed the deadline for the merger. &nbsp;You will note that he made no submissions after the deadline, and withdrew his submission when it came in high, so we acted in good faith; but I guess the rules don't allow for that. &nbsp;We are both out of it now. &nbsp;Congratulations to the winners and I will be interested to see their methods.</p>\n<p>[/quote]</p>\n<p>I just wanted to clarify that my 21 submissions were <em>from my own work</em>. As Jonathan and I had discussions about his methods it was necessary that we formed a team. I accepted Jonathan&#8217;s merge request at around 23:58 UTC on the day of the team merger deadline. However the Kaggle server was unresponsive at the time (presumably from everyone making their submissions just prior to the deadline). So my merge request was denied because by the time the server got to my request it was past the deadline. I&#8217;m sure someone at Kaggle can corroborate this if they look at their server logs!</p>\n<p>Anyway that&#8217;s why I made no further submissions past the team merger deadline. Unlike others, I was not&nbsp;disqualified because I was caught cheating. <em>There was no cheating by&nbsp;us</em>. I emailed Kaggle and let them know what happened and requested to be removed. I didn&#8217;t do so because I was in the money. Simply, I did not know whom to email prior to the competition closing. It was only after William posted the 'Cheaters Removed' thread, when I finally found a suitable email address.</p>\n<p>&#8230;.What a mess&#8230;. &nbsp;</p>\n<p>:S</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58666,
      "author_name": "drewabbot",
      "author_url": "",
      "post_date": "11/22/2014 03:05:31",
      "content": "<p>Jonathan and Andronicus, both of you have my empathy (although being in 2nd place now,&nbsp;admittedly I'm not complaining). &nbsp;That said, I would love to see both&nbsp;of your methods. &nbsp;It would be sad for all your work to go to waste.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58674,
      "author_name": "",
      "author_url": "",
      "post_date": "11/22/2014 08:24:02",
      "content": "<p>For some of you, is your private score better than public score? or is it always less than your public score?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58789,
      "author_name": "andtem2000",
      "author_url": "",
      "post_date": "11/24/2014 16:56:45",
      "content": "<p>[quote=rakhlin;58401]</p>\n<p>Andy, quite the contrary, the organizers needlessly complicated the problem. Other works don't report performance across subjects putting them on a common scale like here - it is meaningless. Add small and highly unbalanced data, particularly for 2 humans, and the problem can not be generalized well even on per subject basis. I think without post-calibration performance would not outbid 60%. Finally add quite meaningless metric. In practice you're not interested in AUC. For perfect classifier it should be enough to produce no false positives and at least&nbsp;<strong><em>one</em></strong> true positive for every preictal period.&nbsp;</p>\n<p>[/quote]</p>\n<p>I'd disagree. AUC is quite a good metric for this task. If you think of a real application and for ethical reasons it will never be an automated system but a decision support system, then for a decision support tool a probabilistic trend is a reasonable output. AUC measures an overlap of the two distributions, which is indicative of a perceived difference between probabilistic levels of ictal and preictal activity for intended end-users. Computing one AUC across all patients or average of AUCs per patient is a choice connected to the level of robustness required from the tool. There was a lack of consistency from organisers about these aspects, high robustness by the metric but low robustness by test data usage.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58834,
      "author_name": "drewabbot",
      "author_url": "",
      "post_date": "11/25/2014 00:15:14",
      "content": "<p>[quote=rakhlin;58401]</p>\n<p>I think without post-calibration performance would not outbid 60%.</p>\n<p>[/quote]</p>\n<p>With one of our models, we achieved over 80% on private LB without any post-calibration, use of test data, or model ensembling. &nbsp;We look forward to reporting our results soon.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58854,
      "author_name": "rakhlin",
      "author_url": "",
      "post_date": "11/25/2014 06:01:20",
      "content": "<p>[quote=Andy;58789]</p>\n<p>AUC measures an overlap of the two distributions, which is indicative of a perceived difference between probabilistic levels of ictal and preictal activity for intended end-users.[/quote]</p>\n<p>Indeed, AUC measures an overlap of the two distributions. But it does not tell whether absolute value of probability prediction has any good use at all unless&nbsp;<em>properly</em> calibrated - another task not related to AUC metric. You can reach AUC=1 and still be unable practically interpret the output because AUC does not care about true boundary between the distributions. See <a href=\"https://www.kaggle.com/c/seizure-prediction/forums/t/10945/congratulations-to-the-winners/58289#post58289\">here</a>&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58865,
      "author_name": "andtem2000",
      "author_url": "",
      "post_date": "11/25/2014 09:41:38",
      "content": "<p>Well, if you talk about an operating point, then it does not matter. Have you ever observed a trend of say probability in time? What you perceive is not an absolute value, but decays and rises with respect to background.</p>\n<p>A separate problem that you may have 0.499999 and 0.511111, with the AUC = 1 but the boundary will not be perceivable. It is true in theory, but given a real problem your scores are either normally distributed&nbsp;likelihoods or gamma distributed posteriors.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58870,
      "author_name": "pakozm",
      "author_url": "",
      "post_date": "11/25/2014 09:47:12",
      "content": "<p>[quote=Drew Abbot;58834]</p>\n<p>[quote=rakhlin;58401]</p>\n<p>I think without post-calibration performance would not outbid 60%.</p>\n<p>[/quote]</p>\n<p>With one of our models, we achieved over 80% on private LB without any post-calibration, use of test data, or model ensembling. &nbsp;We look forward to reporting our results soon.</p>\n<p>[/quote]</p>\n<p>We&nbsp;didn't perform test calibration in any of our models, and our result is between 79% and 80%</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58944,
      "author_name": "rakhlin",
      "author_url": "",
      "post_date": "11/25/2014 19:06:40",
      "content": "<p>[quote=Andy;58865]</p>\n<p>Well, if you talk about an operating point, then it does not matter. Have you ever observed a trend of say probability in time? What you perceive is not an absolute value, but decays and rises with respect to background.[/quote]</p>\n<p>Imagine a model that scores all available data [0...0.1] (interictal) or [0.9...1] (preictal). For a new data it returns 0.2. We'll have no idea how to interpret that. Moreover, label can be anything, preictal or interictal, it won't change previous AUC=1. Given limited data absolutely possible scenario, particularly &nbsp;for a problem like this competition. The problem becomes even more general if a model's score isn't restricted to [0 1]. This is why binary metric makes a sense.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 59336,
      "author_name": "jialunhe",
      "author_url": "",
      "post_date": "12/01/2014 16:40:43",
      "content": "<p>Summary of my solution:</p>\n<p>Feature Models:</p>\n<p>My best submission according to the LB score is based on single window model. All the data is first re-sample to 100Hz to reduce high frequency noise. Then every data file is split into 12 parts, about 50 seconds each. For each part of the split data, FFT is applied to transform the data to frequency domain. The power magnitudes in the frequency band from 1 to 50 Hz were selected and converted to logarithmic scale, then resample the frequency band to 18 bins to further reduce noise. The covariance and eigenvalues of the reduced frequency band across channels are also added as features, along with the covariance and eigenvalues in time domain.</p>\n<p>Classifier Models</p>\n<p><br>Several common classifiers in scikit-learn package have been tested, such as random forest, gradient tree boosting, support vector machine. Most of them had really good CV score for individual subject. But did not get good score in LB. The gaps between CV score and LB score were very big. One of the reason is that LB score is across all subject. Other possible reason is due to overfitting. Platt scaling was also added to calibrate the prediction across subjects. It improved LB score slightly.<br>My best submissions according to the LB score were based on support vector machine with RBF kernel, which produced better results because of more control in balancing bias and variance. The classifier gave an estimate for each 50-second window. An averaging method was used to combine the results and provide estimate for the whole 10-minute period. Several averaging methods, such as arithmetic average, geometric average and harmonic average, have been tested. Arithmetic average was best suited for evenly distributed estimate while harmonic average was best suited for oddly distribution. Therefore, a combined averaging method was used. A percentage projection was also used to align results across subjects based on the assumption that the test dataset was similar to the training dataset.</p>\n<p>The repository is available at <a href=\"https://github.com/jlnh/SeizurePrediction\">https://github.com/jlnh/SeizurePrediction</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 59348,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "12/01/2014 18:30:10",
      "content": "<p>Thanks @Birchwood! Please also attach the repo&nbsp;to&nbsp;your team's Github section&nbsp;(https://www.kaggle.com/c/seizure-prediction/github).</p>",
      "votes": null,
      "replies": [
        {
          "id": 593000,
          "author_name": "aikudo",
          "author_url": "",
          "post_date": "08/06/2019 05:10:28",
          "content": "<p>I'm curious if this repo has moved to somewhere? Did anyone make a backup/mirror of the winners solutions? Thanks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 59374,
      "author_name": "saikumarallaka",
      "author_url": "",
      "post_date": "12/02/2014 04:51:01",
      "content": "<p>@Birchwood: I want to implement your model.</p>\n<p>Can you please provide me more details like.</p>\n<p>1.&nbsp;What each python code of yours is doing.</p>\n<p>2. Can you represent the entire flow of code graphically, so that we get to know more about it.</p>\n<p>Sorry, if i am asking for more.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 164846,
      "author_name": "haltoorroust1950",
      "author_url": "",
      "post_date": "03/02/2017 14:46:59",
      "content": "<p>Hello people, i so details about past competitions and i am interested when will it be a new competition! I have epilepsy so i have to search for <a href=\"https://pharmacyreviews.md/\">pharmacy reviews</a> often because of my needs. but thanks anyway and have a good day!! </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "58249": "",
    "58250": "",
    "58251": "",
    "58252": "",
    "58255": "",
    "58259": "",
    "58260": "",
    "58261": "",
    "58263": "",
    "58264": "",
    "58268": "",
    "58281": "",
    "58282": "",
    "58283": "",
    "58286": "",
    "58288": "",
    "58289": "",
    "58290": "",
    "58300": "",
    "58302": "",
    "58322": "",
    "58337": "",
    "58342": "",
    "58350": "",
    "58351": "",
    "58357": "",
    "58358": "",
    "58364": "",
    "58368": "",
    "58371": "",
    "58372": "",
    "58373": "",
    "58384": "",
    "58389": "",
    "58395": "",
    "58396": "",
    "58400": "",
    "58401": "",
    "58402": "",
    "58403": "",
    "58404": "",
    "58405": "",
    "58408": "",
    "58415": "",
    "58417": "",
    "58420": "",
    "58425": "",
    "58430": "",
    "58449": "",
    "58450": "",
    "58451": "",
    "58456": "",
    "58458": "",
    "58468": "",
    "58480": "",
    "58489": "",
    "58492": "",
    "58512": "",
    "58515": "",
    "58517": "",
    "58518": "",
    "58519": "",
    "58520": "",
    "58533": "",
    "58541": "",
    "58561": "",
    "58563": "",
    "58568": "",
    "58570": "",
    "58572": "",
    "58574": "",
    "58609": "",
    "58639": "",
    "58655": "",
    "58656": "",
    "58659": "",
    "58660": "",
    "58661": "",
    "58663": "",
    "58666": "",
    "58674": "",
    "58789": "",
    "58834": "",
    "58854": "",
    "58865": "",
    "58870": "",
    "58944": "",
    "59336": "",
    "59348": "",
    "59374": "",
    "164846": "Hello people, i so details about past competitions and i am interested when will it be a new competition! I have epilepsy so i have to search for <a href=\"https://pharmacyreviews.md/\">pharmacy reviews</a> often because of my needs. but thanks anyway and have a good day!!",
    "593000": "I'm curious if this repo has moved to somewhere? Did anyone make a backup/mirror of the winners solutions? Thanks."
  },
  "source": "meta"
}