{
  "id": 10935,
  "title": "Not long left now",
  "url": "/competitions/seizure-prediction/discussion/10935",
  "author_name": "",
  "post_date": "2014-11-16T13:17:15.670Z",
  "votes": 2,
  "comment_count": 34,
  "views": 4887,
  "content": "<p>34 hours to go and I am exhausted. How is everybody else doing? I'm really keen for the post-competition discussion, can't wait. :) The first thing I would really like to know is Medrr's secret for the sudden jump from 0.86 to 0.90!</p>",
  "messages": [
    {
      "id": "58133",
      "postDate": "11/16/2014 13:17:15",
      "content": "<p>34 hours to go and I am exhausted. How is everybody else doing? I'm really keen for the post-competition discussion, can't wait. :) The first thing I would really like to know is Medrr's secret for the sudden jump from 0.86 to 0.90!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58135",
      "postDate": "11/16/2014 14:10:27",
      "content": "<p>I am exhausted too, and I cant get over 0.64. What is your secret?!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58136",
      "postDate": "11/16/2014 14:29:24",
      "content": "<p>Don't worry. Maybe some shake up!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58139",
      "postDate": "11/16/2014 16:03:21",
      "content": "<p>I expect lots of shakeup with so many submissions and a fairly small public test set.</p>\n<p>I decided&nbsp;to&nbsp;wait until&nbsp;yesterday to start - if I don't like my results I can blame it on shortage of time and still feel OK about it.</p>\n<p>&nbsp;I was also expecting to ensemble a bunch of &quot;beating the benchmark&quot; code but I don't see much out there.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58140",
      "postDate": "11/16/2014 16:44:15",
      "content": "<p>[quote=James King;58139]</p>\n<p>I decided&nbsp;to&nbsp;wait until&nbsp;yesterday to start - if I don't like my results I can blame it on shortage of time and still feel OK about it.</p>\n<p>[/quote]</p>\n<p>Wow, you started yesterday and are already at 0.7. It took me almost a week just to generate my features :(</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58143",
      "postDate": "11/16/2014 17:06:03",
      "content": "<p>I generated the features on an Amazon 32 core&nbsp;r3.8xlarge with an SSD volume to store the data. Otherwise I'd still be starting at the screen waiting for the first run to finish.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58162",
      "postDate": "11/16/2014 23:14:46",
      "content": "<p>Mahi you should be able to achieve&nbsp;a better score&nbsp;using only cross correlation features&nbsp;using&nbsp;my code from the previous competition (cross correlation coefficients upper right triangle and eigenvalues). You can also&nbsp;break the 600s into smaller windows to generate more training samples.</p>\n<p>I struggled a lot at the start because I forgot to scale my features when using with SVM. By scale I mean subtract mean and divide by standard deviation for each feature i.e. StandardScaler() in sklearn. I was using RandomForest previously which doesn't care about feature scaling and so forgot to update that when trying out other classifiers.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58165",
      "postDate": "11/17/2014 00:33:45",
      "content": "<p>[quote=Michael Hills;58133]</p>\n<p>34 hours to go and I am exhausted. How is everybody else doing? I'm really keen for the post-competition discussion, can't wait. :) The first thing I would really like to know is Medrr's secret for the sudden jump from 0.86 to 0.90!</p>\n<p>[/quote]</p>\n<p>I have some crazy/stupid hypothesis.</p>\n<p>We have some statistics:<br>test_file, probability, auc</p>\n<p>Auc is general for each post. Total 3935*[number of posts] rows.</p>\n<p>Can it be used to find function<br>auc = some_function(test_file, probability)?</p>\n<p>Or even<br>probability = some_function(test_file, auc) ?</p>\n<p>If you put auc as 1 then &#8230;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58167",
      "postDate": "11/17/2014 01:22:42",
      "content": "<p>[quote=ruai;58165]</p>\n<p>Can it be used to find function<br>auc = some_function(test_file, probability)?</p>\n<p>[/quote]</p>\n<p>Suppose it is found. Does it generalize to private LB?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58168",
      "postDate": "11/17/2014 03:04:13",
      "content": "<p>[quote=Michael Hills;58162]</p>\n<p>Mahi you should be able to achieve&nbsp;a better score&nbsp;using only cross correlation features&nbsp;using&nbsp;my code from the previous competition (cross correlation coefficients upper right triangle and eigenvalues). You can also&nbsp;break the 600s into smaller windows to generate more training samples.</p>\n<p>[/quote]</p>\n<p>Thank you so much for your code! I really hope you can win this time!&nbsp;Without your code we can't even start this competition.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58170",
      "postDate": "11/17/2014 03:30:53",
      "content": "<p>No problem. :) Unfortunately it doesn't take advantage of multiple cores except for training RandomForest with n_jobs parameter. This time around I rewrote everything to handle the sheer size of the data better, although the code is not as clean. So now uses all cores for data processing, as well as queuing up and running N different cross-validation folds in parallel. Maximise CPU usage. :)</p>\n<p>I'll be releasing my code for this competition win or lose, spent far too much time on this to let all that effort go to waste!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58171",
      "postDate": "11/17/2014 03:43:50",
      "content": "<p>I have hit a wall of 0.64 but have learned so much in the process...much more than what my Data Science internship taught me., However, the results have been disappointing. I have tried so many features, read a few papers. Nothing has helped to the extent that i wished for.</p>\n<p>I wanted to get my AUC above 0.7 at least. I have just a few hours to go, I will not sleep tonight and until the deadline , i have exhausted most approaches but will be trying a few last remaining ones. </p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58173",
      "postDate": "11/17/2014 03:56:06",
      "content": "<p>[quote=Kushank Raghav;58171]</p>\n<p>I have hit a wall of 0.64 but have learned so much in the process...much more than what my Data Science internship taught me., However, the results have been disappointing. I have tried so many features, read a few papers. Nothing has helped to the extent that i wished for.</p>\n<p>I wanted to get my AUC above 0.7 at least. I have just a few hours to go, I will not sleep tonight and until the deadline , i have exhausted most approaches but will be trying a few last remaining ones.</p>\n<p>[/quote]</p>\n\n<p>Dude, I am on the same boat as you. Stuck at that number too. Michael Hills gave a suggestion to do. I will try that sometime tonight and see if my score increases. Best of luck man.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58190",
      "postDate": "11/17/2014 10:03:20",
      "content": "<p>One of my very poorly performing public LB entries is getting 0.92-0.96 AUC via 10x+ fold local cross validation with 10x+ random seeds. Was the split non-random? Or am I messing up somewhere along the pipeline? It's very possible to make mistakes given the size of the dataset and length of the pipeline. However, I'd guess the vast majority of people are overfitting.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58191",
      "postDate": "11/17/2014 10:32:25",
      "content": "<p>Ive hit the wall. Cannot improve anymore, no matter what I try.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58192",
      "postDate": "11/17/2014 10:33:27",
      "content": "<p>@Abhishek: i can totally relate to that :-(</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58193",
      "postDate": "11/17/2014 11:09:12",
      "content": "<p>@Mike, look into the forum. Several times has been stated that it is important to make the cross validation taking into account the sequence numbers in the training. It is expected higher similarity between samples in the same sequence, therefore&nbsp;a random split can be a bad idea.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58204",
      "postDate": "11/17/2014 14:45:38",
      "content": "<p>[quote=Francisco Zamora-Martinez;58193]</p>\n<p>@Mike, look into the forum. Several times has been stated that it is important to make the cross validation taking into account the sequence numbers in the training. It is expected higher similarity between samples in the same sequence, therefore&nbsp;a random split can be a bad idea.</p>\n<p>[/quote]</p>\n<p>Can you please point us the these posts?</p>\n<p>I was looking for a good way to evaluate but nothing worked for me.&nbsp;</p>\n<p>My CV score is always much much higher.&nbsp;</p>\n<p>Thanks,</p>\n<p>C</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58205",
      "postDate": "11/17/2014 14:50:56",
      "content": "<p>[quote=clustifier;58204]</p>\n<p>Can you please point us the these posts?</p>\n<p>[/quote]</p>\n<p>something like this&nbsp;https://www.kaggle.com/c/seizure-prediction/forums/t/10405/python-code-for-cv-splitting/54390#post54390</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58208",
      "postDate": "11/17/2014 15:56:29",
      "content": "<p>[quote=rcarson;58205]</p>\n<p>[quote=clustifier;58204]</p>\n<p>Can you please point us the these posts?</p>\n<p>[/quote]</p>\n<p>something like this&nbsp;https://www.kaggle.com/c/seizure-prediction/forums/t/10405/python-code-for-cv-splitting/54390#post54390</p>\n<p>[/quote]</p>\n<p>I have tried this (I think, I've tried&nbsp;StratifiedShuffleSplit).</p>\n<p>Does this code do more than sklearn&nbsp;StratifiedShuffleSplit?</p>\n<p>How does it&nbsp;help?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58210",
      "postDate": "11/17/2014 16:36:36",
      "content": "<p>If the leaderboard train/test split is done via segment sequence, then you will want to split your train and test locally by that exact same rule. I haven't seen any definite confirmation on how the split was conducted. It might be on the forums or some webpage.</p>\n<p>This rule will lower CV compared to pure SRS but not enough to get the LB scores I see. There are&nbsp;629 sequences in the training dataset. The testing dataset is a bit smaller in raw row size compared to the train. So basically you're looking at about 300 public / 300 private leaderboard test sequences (private should be slightly bigger). When I do 50/50 split repeated local CV based upon sequences I get IQRs on the range of 0.06 AUC and one SD is a bit lower (this is still huge). This means the potential shakeup can be quite large. My estimate is around the order of the Africa competition.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58213",
      "postDate": "11/17/2014 17:00:19",
      "content": "<p>[quote=clustifier;58208]</p>\n<p>I have tried this (I think, I've tried&nbsp;StratifiedShuffleSplit).</p>\n<p>Does this code do more than sklearn&nbsp;StratifiedShuffleSplit?</p>\n<p>[/quote]</p>\n<p>You need to ensure that you don't split a sequence, since clips within a sequence are similar and so will introduce bias. If you do a StratifiedShuffleSplit on the grouped clips then it should&nbsp;amount to the same thing.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58217",
      "postDate": "11/17/2014 18:28:16",
      "content": "<p>[quote=Michael Hills;58162]</p>\n<p>Mahi you should be able to achieve&nbsp;a better score&nbsp;using only cross correlation features&nbsp;using&nbsp;my code from the previous competition (cross correlation coefficients upper right triangle and eigenvalues). You can also&nbsp;break the 600s into smaller windows to generate more training samples.</p>\n<p>I struggled a lot at the start because I forgot to scale my features when using with SVM. By scale I mean subtract mean and divide by standard deviation for each feature i.e. StandardScaler() in sklearn. I was using RandomForest previously which doesn't care about feature scaling and so forgot to update that when trying out other classifiers.</p>\n<p>[/quote]</p>\n<p>Hi, Michael</p>\n<p>Just let you know I am using your code,&nbsp; hope you can in top3 then you have a chance to post your code again&nbsp;&nbsp; :)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58218",
      "postDate": "11/17/2014 18:36:09",
      "content": "<p>[quote=Vilen Jumutc;58216]</p>\n<p>I see no scientific and ML use of such competitions if people are optimizing over LB scores. Test set should have been provided few days before the end with only couple of possible submissions. All major and &quot;good&quot; old ML competitions were doing so... providing only validation set in advance (not the whole test set)!!!</p>\n<p>[/quote]</p>\n<p>Hi, Vilen</p>\n<p>I think, there is something to do with &quot;user engagement&quot;.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58224",
      "postDate": "11/17/2014 20:37:19",
      "content": "<p>I find that feature engineering has a lot to do with the CV LB difference as well as respecting the sequences. For example I got massive CV scores when i forgot to filter out the DC/low freq components in my features which didn't generalize to the LB.</p>\n<p>One thing that i wanted to try that i'm afraid i do not have the time for is to rank the features based on Mutual information with the output. Then select the K highest ranked features and supply thse to your classifier.</p>\n<p>if anyone is struggling against&nbsp; a wall, you might want to give that a shot.</p>\n\n<p>edit: another thing you can try is to use Kernel PCA to clean up your feature matrix before supplying it to your classifier.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58233",
      "postDate": "11/17/2014 22:26:56",
      "content": "<p>Michael Hills,</p>\n\n<p>Is it possible if I collaborated with you on other competitions so I learn from you. The biggest issue I feel like I had with this problem was the approach. I did not know how to deal with such big data set. I believe It would be good if I could see someone's thought process and learn from it.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58235",
      "postDate": "11/17/2014 22:47:02",
      "content": "<p>[quote=Abhishek;58191]</p>\n<p>Ive hit the wall. Cannot improve anymore, no matter what I try.</p>\n<p>[/quote]</p>\n<p>Very reassuring to someone new to hear an expert say that! &nbsp;Thanks for that.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58240",
      "postDate": "11/18/2014 00:21:54",
      "content": "<p>Steven don't worry I'll be posting my code either way.</p>\n<p>Mahi, to be honest I am not much of an expert, I just spent an extensive amount of time banging my head against the problem until I saw a result. Learning the hard way really. I can give you a step by step of what I did though?</p>\n<p>Handling the huge dataset was quite a big problem for me as well. I didn't want to resort to renting hardware (Amazon EC2 etc) so I tried to make it work on my Macbook Pro (quad i7, 16gb ram). I had&nbsp;to rewrite all the data-loading code from scratch. First step is to convert the original files into a more convenient format (from mat to hdf5). For data-format I tested several approaches, using hickle (python pickle/hdf5 thing I used in previous competition), using mat format, and using h5py directly. It turned out hickle was dreadfully slow for some unknown reason, and both mat/hdf5 formats seemed to offer the same fast performance.</p>\n<p>At the same time reduce their size by decimating the original time signals down to 200Hz which was then only 29GB on disk (using int16). I chose 200Hz&nbsp;because my efforts on&nbsp;the previous competition seemed to indicate this was a good tradeoff as it gives you up to 100Hz for frequency analysis. However decimating down to 100Hz might have been a good idea too.</p>\n<p>Next was windowing the data, I used 75s windows because it seemed like a good balance between number of training samples (increase by factor of 8) and leaderboard submissions seemed to do better on it (possibly overfitting though). It took a bit to get this code working properly, I think it was more than just doing a numpy reshape.</p>\n<p>Another major win was my Pipeline, InputSource and FeatureConcatPipeline concepts. Pipeline is like from my previous code, just a series of data transformations e.g. Pipeline(Windower(75), FFT(), Magnitude(), Log10(), FlattenChannels()). However recalculating FFT all the time was really, really slow. I used a lot of spectral features so I didn't want to be redoing this calculation every time. I wrote InputSource to solve this problem, a Pipeline takes an InputSource to say where to source the data from so previously processed data could be reused. e.g. Pipeline(InputSource(Windower(75), FFT(), Magnitude()), SpectralEntropy()) loads the previously calculated FFT data from disk and then pipes it into SpectralEntropy. Finally FeatureConcatPipeline let me mix and match different features very easily. It lets you specify multiple pipelines to group together, e.g. time correlation is one pipeline, frequency correlation is another pipeline, you put them together in the FeatureConcatPipeline, and both pipelines will be loaded and their features concatenated together.</p>\n<p>The actual processing of the pipeline uses all cores. I used python multiprocessing Pool so each process gets a fraction of the data to process. It loads in one segment, processes it, and then writes it out. This is to minimise memory usage. Loading all the data in for processing uses too much memory. So one segment in, process it, one segment out. Then afterwards all these individual segments are collected and merged into one big hdf5 file because this loads much faster the next time you need it (milliseconds). The whole process is also stoppable/restartable. I never wanted to have to worry about killing my program and corrupting data. So data is first written to temp files marked with the process id, and then when it's finalised it is renamed to the final name. Temp files can be cleaned up as the parsed process id will no longer be alive. The processing one segment at a time also meant each segment is a&nbsp;finished piece of work, and would skip over them if you restarted the program.</p>\n<p>On top of all of this, I used a python multiprocessing Pool for training classifiers. I used 3 folds, and often would try out different classifiers too.&nbsp;Trying out 10 different classifiers on the same data only processes the data once, then loads it 10 different times. Fast. A cross-validation run for a specific pipeline and classifier is also saved to disk so I can pull the scores in next time for comparison.</p>\n<p>The biggest caveat was not having enough disk space. I only had around 150GB free on my SSD. Storing large datasets like the FFT chewed up a lot of space and made it difficult to try more things and I would have to delete from&nbsp;the data cache to free up space.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58241",
      "postDate": "11/18/2014 00:25:07",
      "content": "<p>Well done! Michael. I learn a lot from your code and approach :D</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58243",
      "postDate": "11/18/2014 00:36:09",
      "content": "<p>Thank you so much Michael. It's great that you shared your approach with all of us.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58244",
      "postDate": "11/18/2014 00:37:41",
      "content": "<p>Michael can you post your CV scores distribution here?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58245",
      "postDate": "11/18/2014 00:41:47",
      "content": "<p>Do you mean my public leaderboard scores for various submissions or my local cross-validation scores? I think my local scores are not all that accurate.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58246",
      "postDate": "11/18/2014 00:43:44",
      "content": "<p>Local CV scores. I think it's possible the train/test split was not random or even stratified random. However, I didn't organize the contest so I'm not sure.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58248",
      "postDate": "11/18/2014 00:56:32",
      "content": "<p>mean=0.877 std=0.102 [0.709,0.996,0.907,0.906,0.998,0.863,0.759] c=2 p=0<br>mean=0.898 std=0.085 [0.869,0.976,0.888,0.813,0.996,0.986,0.759] c=0 p=0<br>mean=0.900 std=0.084 [0.854,0.977,0.896,0.818,0.998,0.991,0.769] c=1 p=0</p>\n<p>c=0 is svm rbf gamma=0.0079 C=2.7<br>c=1 is svm rbf gamma=0.0068 C=2.0<br>c=2 is logistic regression C=0.04</p>\n<p>This is using 3 random folds against sequence groups. Random because I was lazy but hand-picked random_state values to get 'good enough' diversity in the folds for the preictal sequences. Later on I did write a more reliable k-fold setup with manually-enforced diversity in the fold choices but it still didn't seem all that great.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58253",
      "postDate": "11/18/2014 01:33:02",
      "content": "<p>Thanks.</p>\n<p>I computed the shake up using the code posted here: http://www.kaggle.com/c/liberty-mutual-fire-peril/forums/t/10187/quantifying-leaderboard-shake-up</p>\n<p>&gt; shakeup('http://www.kaggle.com/c/seizure-prediction/leaderboard')<br>Joining by: id<br>$shakeup.top<br>[1] 0.02324478</p>\n<p>$shakeup.all<br>[1] 0.05887034</p>\n<p>For comparison http://www.kaggle.com/c/higgs-boson/forums/t/10320/quantifying-leaderboard-shake-up and:</p>\n<p>&gt; shakeup('http://www.kaggle.com/c/afsis-soil-properties/leaderboard')<br>Joining by: id<br>$shakeup.top<br>[1] 0.1806223</p>\n<p>$shakeup.all<br>[1] 0.1187472</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 58135,
      "author_name": "mahi83",
      "author_url": "",
      "post_date": "11/16/2014 14:10:27",
      "content": "<p>I am exhausted too, and I cant get over 0.64. What is your secret?!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58136,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "11/16/2014 14:29:24",
      "content": "<p>Don't worry. Maybe some shake up!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58139,
      "author_name": "jfkingiii",
      "author_url": "",
      "post_date": "11/16/2014 16:03:21",
      "content": "<p>I expect lots of shakeup with so many submissions and a fairly small public test set.</p>\n<p>I decided&nbsp;to&nbsp;wait until&nbsp;yesterday to start - if I don't like my results I can blame it on shortage of time and still feel OK about it.</p>\n<p>&nbsp;I was also expecting to ensemble a bunch of &quot;beating the benchmark&quot; code but I don't see much out there.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58140,
      "author_name": "nissanpow",
      "author_url": "",
      "post_date": "11/16/2014 16:44:15",
      "content": "<p>[quote=James King;58139]</p>\n<p>I decided&nbsp;to&nbsp;wait until&nbsp;yesterday to start - if I don't like my results I can blame it on shortage of time and still feel OK about it.</p>\n<p>[/quote]</p>\n<p>Wow, you started yesterday and are already at 0.7. It took me almost a week just to generate my features :(</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58143,
      "author_name": "jfkingiii",
      "author_url": "",
      "post_date": "11/16/2014 17:06:03",
      "content": "<p>I generated the features on an Amazon 32 core&nbsp;r3.8xlarge with an SSD volume to store the data. Otherwise I'd still be starting at the screen waiting for the first run to finish.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58162,
      "author_name": "michaelhills",
      "author_url": "",
      "post_date": "11/16/2014 23:14:46",
      "content": "<p>Mahi you should be able to achieve&nbsp;a better score&nbsp;using only cross correlation features&nbsp;using&nbsp;my code from the previous competition (cross correlation coefficients upper right triangle and eigenvalues). You can also&nbsp;break the 600s into smaller windows to generate more training samples.</p>\n<p>I struggled a lot at the start because I forgot to scale my features when using with SVM. By scale I mean subtract mean and divide by standard deviation for each feature i.e. StandardScaler() in sklearn. I was using RandomForest previously which doesn't care about feature scaling and so forgot to update that when trying out other classifiers.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58165,
      "author_name": "ruai00",
      "author_url": "",
      "post_date": "11/17/2014 00:33:45",
      "content": "<p>[quote=Michael Hills;58133]</p>\n<p>34 hours to go and I am exhausted. How is everybody else doing? I'm really keen for the post-competition discussion, can't wait. :) The first thing I would really like to know is Medrr's secret for the sudden jump from 0.86 to 0.90!</p>\n<p>[/quote]</p>\n<p>I have some crazy/stupid hypothesis.</p>\n<p>We have some statistics:<br>test_file, probability, auc</p>\n<p>Auc is general for each post. Total 3935*[number of posts] rows.</p>\n<p>Can it be used to find function<br>auc = some_function(test_file, probability)?</p>\n<p>Or even<br>probability = some_function(test_file, auc) ?</p>\n<p>If you put auc as 1 then &#8230;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58167,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "11/17/2014 01:22:42",
      "content": "<p>[quote=ruai;58165]</p>\n<p>Can it be used to find function<br>auc = some_function(test_file, probability)?</p>\n<p>[/quote]</p>\n<p>Suppose it is found. Does it generalize to private LB?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58168,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "11/17/2014 03:04:13",
      "content": "<p>[quote=Michael Hills;58162]</p>\n<p>Mahi you should be able to achieve&nbsp;a better score&nbsp;using only cross correlation features&nbsp;using&nbsp;my code from the previous competition (cross correlation coefficients upper right triangle and eigenvalues). You can also&nbsp;break the 600s into smaller windows to generate more training samples.</p>\n<p>[/quote]</p>\n<p>Thank you so much for your code! I really hope you can win this time!&nbsp;Without your code we can't even start this competition.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58170,
      "author_name": "michaelhills",
      "author_url": "",
      "post_date": "11/17/2014 03:30:53",
      "content": "<p>No problem. :) Unfortunately it doesn't take advantage of multiple cores except for training RandomForest with n_jobs parameter. This time around I rewrote everything to handle the sheer size of the data better, although the code is not as clean. So now uses all cores for data processing, as well as queuing up and running N different cross-validation folds in parallel. Maximise CPU usage. :)</p>\n<p>I'll be releasing my code for this competition win or lose, spent far too much time on this to let all that effort go to waste!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58171,
      "author_name": "",
      "author_url": "",
      "post_date": "11/17/2014 03:43:50",
      "content": "<p>I have hit a wall of 0.64 but have learned so much in the process...much more than what my Data Science internship taught me., However, the results have been disappointing. I have tried so many features, read a few papers. Nothing has helped to the extent that i wished for.</p>\n<p>I wanted to get my AUC above 0.7 at least. I have just a few hours to go, I will not sleep tonight and until the deadline , i have exhausted most approaches but will be trying a few last remaining ones. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58173,
      "author_name": "mahi83",
      "author_url": "",
      "post_date": "11/17/2014 03:56:06",
      "content": "<p>[quote=Kushank Raghav;58171]</p>\n<p>I have hit a wall of 0.64 but have learned so much in the process...much more than what my Data Science internship taught me., However, the results have been disappointing. I have tried so many features, read a few papers. Nothing has helped to the extent that i wished for.</p>\n<p>I wanted to get my AUC above 0.7 at least. I have just a few hours to go, I will not sleep tonight and until the deadline , i have exhausted most approaches but will be trying a few last remaining ones.</p>\n<p>[/quote]</p>\n\n<p>Dude, I am on the same boat as you. Stuck at that number too. Michael Hills gave a suggestion to do. I will try that sometime tonight and see if my score increases. Best of luck man.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58190,
      "author_name": "mikeskim",
      "author_url": "",
      "post_date": "11/17/2014 10:03:20",
      "content": "<p>One of my very poorly performing public LB entries is getting 0.92-0.96 AUC via 10x+ fold local cross validation with 10x+ random seeds. Was the split non-random? Or am I messing up somewhere along the pipeline? It's very possible to make mistakes given the size of the dataset and length of the pipeline. However, I'd guess the vast majority of people are overfitting.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58191,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "11/17/2014 10:32:25",
      "content": "<p>Ive hit the wall. Cannot improve anymore, no matter what I try.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58192,
      "author_name": "konradb",
      "author_url": "",
      "post_date": "11/17/2014 10:33:27",
      "content": "<p>@Abhishek: i can totally relate to that :-(</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58193,
      "author_name": "pakozm",
      "author_url": "",
      "post_date": "11/17/2014 11:09:12",
      "content": "<p>@Mike, look into the forum. Several times has been stated that it is important to make the cross validation taking into account the sequence numbers in the training. It is expected higher similarity between samples in the same sequence, therefore&nbsp;a random split can be a bad idea.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58204,
      "author_name": "clustifier",
      "author_url": "",
      "post_date": "11/17/2014 14:45:38",
      "content": "<p>[quote=Francisco Zamora-Martinez;58193]</p>\n<p>@Mike, look into the forum. Several times has been stated that it is important to make the cross validation taking into account the sequence numbers in the training. It is expected higher similarity between samples in the same sequence, therefore&nbsp;a random split can be a bad idea.</p>\n<p>[/quote]</p>\n<p>Can you please point us the these posts?</p>\n<p>I was looking for a good way to evaluate but nothing worked for me.&nbsp;</p>\n<p>My CV score is always much much higher.&nbsp;</p>\n<p>Thanks,</p>\n<p>C</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58205,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "11/17/2014 14:50:56",
      "content": "<p>[quote=clustifier;58204]</p>\n<p>Can you please point us the these posts?</p>\n<p>[/quote]</p>\n<p>something like this&nbsp;https://www.kaggle.com/c/seizure-prediction/forums/t/10405/python-code-for-cv-splitting/54390#post54390</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58208,
      "author_name": "clustifier",
      "author_url": "",
      "post_date": "11/17/2014 15:56:29",
      "content": "<p>[quote=rcarson;58205]</p>\n<p>[quote=clustifier;58204]</p>\n<p>Can you please point us the these posts?</p>\n<p>[/quote]</p>\n<p>something like this&nbsp;https://www.kaggle.com/c/seizure-prediction/forums/t/10405/python-code-for-cv-splitting/54390#post54390</p>\n<p>[/quote]</p>\n<p>I have tried this (I think, I've tried&nbsp;StratifiedShuffleSplit).</p>\n<p>Does this code do more than sklearn&nbsp;StratifiedShuffleSplit?</p>\n<p>How does it&nbsp;help?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58210,
      "author_name": "mikeskim",
      "author_url": "",
      "post_date": "11/17/2014 16:36:36",
      "content": "<p>If the leaderboard train/test split is done via segment sequence, then you will want to split your train and test locally by that exact same rule. I haven't seen any definite confirmation on how the split was conducted. It might be on the forums or some webpage.</p>\n<p>This rule will lower CV compared to pure SRS but not enough to get the LB scores I see. There are&nbsp;629 sequences in the training dataset. The testing dataset is a bit smaller in raw row size compared to the train. So basically you're looking at about 300 public / 300 private leaderboard test sequences (private should be slightly bigger). When I do 50/50 split repeated local CV based upon sequences I get IQRs on the range of 0.06 AUC and one SD is a bit lower (this is still huge). This means the potential shakeup can be quite large. My estimate is around the order of the Africa competition.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58213,
      "author_name": "nissanpow",
      "author_url": "",
      "post_date": "11/17/2014 17:00:19",
      "content": "<p>[quote=clustifier;58208]</p>\n<p>I have tried this (I think, I've tried&nbsp;StratifiedShuffleSplit).</p>\n<p>Does this code do more than sklearn&nbsp;StratifiedShuffleSplit?</p>\n<p>[/quote]</p>\n<p>You need to ensure that you don't split a sequence, since clips within a sequence are similar and so will introduce bias. If you do a StratifiedShuffleSplit on the grouped clips then it should&nbsp;amount to the same thing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58217,
      "author_name": "stevendu",
      "author_url": "",
      "post_date": "11/17/2014 18:28:16",
      "content": "<p>[quote=Michael Hills;58162]</p>\n<p>Mahi you should be able to achieve&nbsp;a better score&nbsp;using only cross correlation features&nbsp;using&nbsp;my code from the previous competition (cross correlation coefficients upper right triangle and eigenvalues). You can also&nbsp;break the 600s into smaller windows to generate more training samples.</p>\n<p>I struggled a lot at the start because I forgot to scale my features when using with SVM. By scale I mean subtract mean and divide by standard deviation for each feature i.e. StandardScaler() in sklearn. I was using RandomForest previously which doesn't care about feature scaling and so forgot to update that when trying out other classifiers.</p>\n<p>[/quote]</p>\n<p>Hi, Michael</p>\n<p>Just let you know I am using your code,&nbsp; hope you can in top3 then you have a chance to post your code again&nbsp;&nbsp; :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58218,
      "author_name": "stevendu",
      "author_url": "",
      "post_date": "11/17/2014 18:36:09",
      "content": "<p>[quote=Vilen Jumutc;58216]</p>\n<p>I see no scientific and ML use of such competitions if people are optimizing over LB scores. Test set should have been provided few days before the end with only couple of possible submissions. All major and &quot;good&quot; old ML competitions were doing so... providing only validation set in advance (not the whole test set)!!!</p>\n<p>[/quote]</p>\n<p>Hi, Vilen</p>\n<p>I think, there is something to do with &quot;user engagement&quot;.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58224,
      "author_name": "franklyn",
      "author_url": "",
      "post_date": "11/17/2014 20:37:19",
      "content": "<p>I find that feature engineering has a lot to do with the CV LB difference as well as respecting the sequences. For example I got massive CV scores when i forgot to filter out the DC/low freq components in my features which didn't generalize to the LB.</p>\n<p>One thing that i wanted to try that i'm afraid i do not have the time for is to rank the features based on Mutual information with the output. Then select the K highest ranked features and supply thse to your classifier.</p>\n<p>if anyone is struggling against&nbsp; a wall, you might want to give that a shot.</p>\n\n<p>edit: another thing you can try is to use Kernel PCA to clean up your feature matrix before supplying it to your classifier.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58233,
      "author_name": "mahi83",
      "author_url": "",
      "post_date": "11/17/2014 22:26:56",
      "content": "<p>Michael Hills,</p>\n\n<p>Is it possible if I collaborated with you on other competitions so I learn from you. The biggest issue I feel like I had with this problem was the approach. I did not know how to deal with such big data set. I believe It would be good if I could see someone's thought process and learn from it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58235,
      "author_name": "mmyers",
      "author_url": "",
      "post_date": "11/17/2014 22:47:02",
      "content": "<p>[quote=Abhishek;58191]</p>\n<p>Ive hit the wall. Cannot improve anymore, no matter what I try.</p>\n<p>[/quote]</p>\n<p>Very reassuring to someone new to hear an expert say that! &nbsp;Thanks for that.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58240,
      "author_name": "michaelhills",
      "author_url": "",
      "post_date": "11/18/2014 00:21:54",
      "content": "<p>Steven don't worry I'll be posting my code either way.</p>\n<p>Mahi, to be honest I am not much of an expert, I just spent an extensive amount of time banging my head against the problem until I saw a result. Learning the hard way really. I can give you a step by step of what I did though?</p>\n<p>Handling the huge dataset was quite a big problem for me as well. I didn't want to resort to renting hardware (Amazon EC2 etc) so I tried to make it work on my Macbook Pro (quad i7, 16gb ram). I had&nbsp;to rewrite all the data-loading code from scratch. First step is to convert the original files into a more convenient format (from mat to hdf5). For data-format I tested several approaches, using hickle (python pickle/hdf5 thing I used in previous competition), using mat format, and using h5py directly. It turned out hickle was dreadfully slow for some unknown reason, and both mat/hdf5 formats seemed to offer the same fast performance.</p>\n<p>At the same time reduce their size by decimating the original time signals down to 200Hz which was then only 29GB on disk (using int16). I chose 200Hz&nbsp;because my efforts on&nbsp;the previous competition seemed to indicate this was a good tradeoff as it gives you up to 100Hz for frequency analysis. However decimating down to 100Hz might have been a good idea too.</p>\n<p>Next was windowing the data, I used 75s windows because it seemed like a good balance between number of training samples (increase by factor of 8) and leaderboard submissions seemed to do better on it (possibly overfitting though). It took a bit to get this code working properly, I think it was more than just doing a numpy reshape.</p>\n<p>Another major win was my Pipeline, InputSource and FeatureConcatPipeline concepts. Pipeline is like from my previous code, just a series of data transformations e.g. Pipeline(Windower(75), FFT(), Magnitude(), Log10(), FlattenChannels()). However recalculating FFT all the time was really, really slow. I used a lot of spectral features so I didn't want to be redoing this calculation every time. I wrote InputSource to solve this problem, a Pipeline takes an InputSource to say where to source the data from so previously processed data could be reused. e.g. Pipeline(InputSource(Windower(75), FFT(), Magnitude()), SpectralEntropy()) loads the previously calculated FFT data from disk and then pipes it into SpectralEntropy. Finally FeatureConcatPipeline let me mix and match different features very easily. It lets you specify multiple pipelines to group together, e.g. time correlation is one pipeline, frequency correlation is another pipeline, you put them together in the FeatureConcatPipeline, and both pipelines will be loaded and their features concatenated together.</p>\n<p>The actual processing of the pipeline uses all cores. I used python multiprocessing Pool so each process gets a fraction of the data to process. It loads in one segment, processes it, and then writes it out. This is to minimise memory usage. Loading all the data in for processing uses too much memory. So one segment in, process it, one segment out. Then afterwards all these individual segments are collected and merged into one big hdf5 file because this loads much faster the next time you need it (milliseconds). The whole process is also stoppable/restartable. I never wanted to have to worry about killing my program and corrupting data. So data is first written to temp files marked with the process id, and then when it's finalised it is renamed to the final name. Temp files can be cleaned up as the parsed process id will no longer be alive. The processing one segment at a time also meant each segment is a&nbsp;finished piece of work, and would skip over them if you restarted the program.</p>\n<p>On top of all of this, I used a python multiprocessing Pool for training classifiers. I used 3 folds, and often would try out different classifiers too.&nbsp;Trying out 10 different classifiers on the same data only processes the data once, then loads it 10 different times. Fast. A cross-validation run for a specific pipeline and classifier is also saved to disk so I can pull the scores in next time for comparison.</p>\n<p>The biggest caveat was not having enough disk space. I only had around 150GB free on my SSD. Storing large datasets like the FFT chewed up a lot of space and made it difficult to try more things and I would have to delete from&nbsp;the data cache to free up space.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58241,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "11/18/2014 00:25:07",
      "content": "<p>Well done! Michael. I learn a lot from your code and approach :D</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58243,
      "author_name": "feixia",
      "author_url": "",
      "post_date": "11/18/2014 00:36:09",
      "content": "<p>Thank you so much Michael. It's great that you shared your approach with all of us.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58244,
      "author_name": "mikeskim",
      "author_url": "",
      "post_date": "11/18/2014 00:37:41",
      "content": "<p>Michael can you post your CV scores distribution here?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58245,
      "author_name": "michaelhills",
      "author_url": "",
      "post_date": "11/18/2014 00:41:47",
      "content": "<p>Do you mean my public leaderboard scores for various submissions or my local cross-validation scores? I think my local scores are not all that accurate.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58246,
      "author_name": "mikeskim",
      "author_url": "",
      "post_date": "11/18/2014 00:43:44",
      "content": "<p>Local CV scores. I think it's possible the train/test split was not random or even stratified random. However, I didn't organize the contest so I'm not sure.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58248,
      "author_name": "michaelhills",
      "author_url": "",
      "post_date": "11/18/2014 00:56:32",
      "content": "<p>mean=0.877 std=0.102 [0.709,0.996,0.907,0.906,0.998,0.863,0.759] c=2 p=0<br>mean=0.898 std=0.085 [0.869,0.976,0.888,0.813,0.996,0.986,0.759] c=0 p=0<br>mean=0.900 std=0.084 [0.854,0.977,0.896,0.818,0.998,0.991,0.769] c=1 p=0</p>\n<p>c=0 is svm rbf gamma=0.0079 C=2.7<br>c=1 is svm rbf gamma=0.0068 C=2.0<br>c=2 is logistic regression C=0.04</p>\n<p>This is using 3 random folds against sequence groups. Random because I was lazy but hand-picked random_state values to get 'good enough' diversity in the folds for the preictal sequences. Later on I did write a more reliable k-fold setup with manually-enforced diversity in the fold choices but it still didn't seem all that great.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58253,
      "author_name": "mikeskim",
      "author_url": "",
      "post_date": "11/18/2014 01:33:02",
      "content": "<p>Thanks.</p>\n<p>I computed the shake up using the code posted here: http://www.kaggle.com/c/liberty-mutual-fire-peril/forums/t/10187/quantifying-leaderboard-shake-up</p>\n<p>&gt; shakeup('http://www.kaggle.com/c/seizure-prediction/leaderboard')<br>Joining by: id<br>$shakeup.top<br>[1] 0.02324478</p>\n<p>$shakeup.all<br>[1] 0.05887034</p>\n<p>For comparison http://www.kaggle.com/c/higgs-boson/forums/t/10320/quantifying-leaderboard-shake-up and:</p>\n<p>&gt; shakeup('http://www.kaggle.com/c/afsis-soil-properties/leaderboard')<br>Joining by: id<br>$shakeup.top<br>[1] 0.1806223</p>\n<p>$shakeup.all<br>[1] 0.1187472</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "58133": "",
    "58135": "",
    "58136": "",
    "58139": "",
    "58140": "",
    "58143": "",
    "58162": "",
    "58165": "",
    "58167": "",
    "58168": "",
    "58170": "",
    "58171": "",
    "58173": "",
    "58190": "",
    "58191": "",
    "58192": "",
    "58193": "",
    "58204": "",
    "58205": "",
    "58208": "",
    "58210": "",
    "58213": "",
    "58217": "",
    "58218": "",
    "58224": "",
    "58233": "",
    "58235": "",
    "58240": "",
    "58241": "",
    "58243": "",
    "58244": "",
    "58245": "",
    "58246": "",
    "58248": "",
    "58253": ""
  },
  "source": "meta"
}