{
  "id": 10312,
  "title": "Getting Started Code by Elliot Dawson",
  "url": "/competitions/seizure-prediction/discussion/10312",
  "author_name": "",
  "post_date": "2014-09-13T07:21:44.587Z",
  "votes": 26,
  "comment_count": 23,
  "views": 10912,
  "content": "<p>Hey guys,<br><br>So I noticed that roughly 33% of the leaderboard consists of people who have yet to beat the random benchmark of 0.5 and thought that since these people have made the effort to download the considerably large dataset I would lend a hand in providing some starter code. Not really sure if I'm going to have a great deal of time to work on this competition with my University schedule at the moment and given the positive externality that this competition intends to have I hope that this might provide someone else with a platform to produce a great algorithm.<br><br>This method achieves a public leaderboard score of around 0.635 and takes about 25 minutes to run using 6 workers in MATLAB (courtesy of the parallel computing toolbox) on a quad core i7 processor. My machine has 16GB of RAM and didn't have any memory issues when running the code. Having said that I'm unsure of how an 8GB RAM machine might cope. At the moment this approach works as follows:<br>- For each subject clip calculate : The variance of each channel as well as the correlation coefficient between each channel<br>- Take these values and use them as your predictor matrix<br>- For each subject train a decision tree using MATLAB's treeBagger with 1000 trees using the subjects' respective predictor matrix calculated from the clips<br>- Predict on each subjects' test data using their respective trained decision tree</p>\n<p>Just run BenchmarkCode once you've added FeatureEngineer2.m to the path and changed the directories from where to write and read the data to and from.<br><br>On a separate note I used the 'table' data type to import the sample submission and to write the resulting submission to file which I believe was only introduced in MATLAB version 2013b but of course you can modify the code to write the result however you like :)</p>\n<p>Let me know if you're having any issues with the code.</p>",
  "messages": [
    {
      "id": "53617",
      "postDate": "09/13/2014 07:21:44",
      "content": "<p>Hey guys,<br><br>So I noticed that roughly 33% of the leaderboard consists of people who have yet to beat the random benchmark of 0.5 and thought that since these people have made the effort to download the considerably large dataset I would lend a hand in providing some starter code. Not really sure if I'm going to have a great deal of time to work on this competition with my University schedule at the moment and given the positive externality that this competition intends to have I hope that this might provide someone else with a platform to produce a great algorithm.<br><br>This method achieves a public leaderboard score of around 0.635 and takes about 25 minutes to run using 6 workers in MATLAB (courtesy of the parallel computing toolbox) on a quad core i7 processor. My machine has 16GB of RAM and didn't have any memory issues when running the code. Having said that I'm unsure of how an 8GB RAM machine might cope. At the moment this approach works as follows:<br>- For each subject clip calculate : The variance of each channel as well as the correlation coefficient between each channel<br>- Take these values and use them as your predictor matrix<br>- For each subject train a decision tree using MATLAB's treeBagger with 1000 trees using the subjects' respective predictor matrix calculated from the clips<br>- Predict on each subjects' test data using their respective trained decision tree</p>\n<p>Just run BenchmarkCode once you've added FeatureEngineer2.m to the path and changed the directories from where to write and read the data to and from.<br><br>On a separate note I used the 'table' data type to import the sample submission and to write the resulting submission to file which I believe was only introduced in MATLAB version 2013b but of course you can modify the code to write the result however you like :)</p>\n<p>Let me know if you're having any issues with the code.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "53660",
      "postDate": "09/13/2014 19:21:32",
      "content": "<p>A lot of those people also just have one entry, so i'm not too worried. Thanks for posting your approach.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "53692",
      "postDate": "09/14/2014 14:36:37",
      "content": "<p>Elliot,</p>\n\n<p>Thanks a lot for sharing.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "54114",
      "postDate": "09/17/2014 02:42:02",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "54491",
      "postDate": "09/23/2014 01:17:36",
      "content": "<p>Thanks for sharing. Similarly, a&nbsp;lasso logistic regression model with only variance of each feature does equally well at ~0.64. Try cvglmnet.m from glmnet toolbox. This is an open source version of Matlab's lassoglm.m &nbsp;I ran on 8GB 2008 MacPro and had no problems with completion between&nbsp;45-60 minutes&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55201",
      "postDate": "09/25/2014 06:24:53",
      "content": "<p>Thank you very much. I don't have the patience to wait for the download and don't want to bother with feature engineering today. Would it be possible to upload the hopefully smaller pre-processed data set here?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55213",
      "postDate": "09/25/2014 09:57:25",
      "content": "<p>[quote=Damian Koesters;55201]</p>\n<p>Thank you very much. I don't have the patience to wait for the download and don't want to bother with feature engineering today. Would it be possible to upload the hopefully smaller pre-processed data set here?</p>\n<p>[/quote]</p>\n<p>@Damian</p>\n<p>Hi Damian,</p>\n<p>I voted your post down -1. And I don't want to do this anonymously. (for the admins: it is maybe a good idea to have 2 lines under each post: 'voted up by' and 'voted down by' to avoid teammates giving up-votes and trolling)</p>\n<p>You first have to accept the rules before you can download (if the interface works as I think it works you haven't done that yet).</p>\n<p>Furthermore there is a minimum of work involved getting the promised leaderboard score:</p>\n<p>Namely downloading the data and running the shared script.</p>\n<p>Sharing is disruptive enough as it is, if the new work modus will change to down and uploading preprocessed data, why not give everybody who enters the challenge the promised LB score.</p>\n<p>I hope you don't take this personally, everybody could have asked the same question, it is just that I want to preserve some of the meaning the score on the private LB has.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55216",
      "postDate": "09/25/2014 11:21:40",
      "content": "<p>Hi Jules, ugh, downvoting really hurts but I won't take it personally. I just thought that it would be great for each challenge to have a set of reasonable pre-processing steps or an alternative pre-processed dataset to have a starting point.</p>\n<p>I want to try a simple idea on the classifier side without worrying about the feature engineering steps. The feature engineering is quite a hurdle as I have no knowledge about EEG-data.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55218",
      "postDate": "09/25/2014 12:31:12",
      "content": "<p>In the spirit of openness inspired by Jules: Damian, I downvoted you as well. While I understand your approach (you want to get to the fun part, i.e. modeling), feature engineering and data cleaning are a core part of the job. Elliot's code does quite a bit of heavy lifting for you, so saying &quot;<em>you can't be bothered</em>&quot; to do some of the extra work required (it's nothing more than downloading and running his code, for crying out loud) is quite inappropriate in my view.&nbsp;</p>\n\n<p>And as for lack of domain expertise:</p>\n<p>1. i don't think a lot of top contenders know a lot about EEG</p>\n<p>2. in the Avito contest, which dealt with text mining of data in Russian, hardly any of the top contenders spoke Russian&nbsp;</p>\n<p>so I wouldn't worry that much :-)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55219",
      "postDate": "09/25/2014 12:52:27",
      "content": "<p>I agree that data cleaning and feature engineering is part of the challenge in most Kaggle competitions, and, in general, data science projects.</p>\n<p>BTW, regarding the up and down voting, am I the only one who preferred the previous forum with &quot;Thanks&quot; feature? I thought it was quite original and nice, as it was promoting development of friendly relationships between contestants. Personally I find the up and down buttons a bit dull, in contrast.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55221",
      "postDate": "09/25/2014 13:07:57",
      "content": "<p>Elliot, thank you for sharing this code. I learn a lot.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55223",
      "postDate": "09/25/2014 13:24:32",
      "content": "<p>Hi Damian, in response to your request, I agree with everyone else who has responded (however I have not downvoted your comment on the basis that you've applied the theory of <em>If you don't ask the answer is always no</em>) and given there is no prominent precedent of people providing preprocessed data directly I would rather not make it a trend. As for domain expertise, I would have to say that mine would total to 0, my current studies are in Actuarial science has given me no working knowledge of an EEG&nbsp; (unless I missed that lecture haha) yet this hasn't stopped me from creating some competent code to get started with.<br><br>When I posted this two weeks I had hoped that somebody down the bottom of the leaderboard who had a hard time getting started would modify both the preprocessing script and the modelling component of the code to create something a lot better. So I hope you do indeed download the data and I wish you the best of luck with the competition!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55236",
      "postDate": "09/25/2014 16:22:22",
      "content": "<p>@Elliot: Thanks for the code and your spirit of helping others. The code will definitely help me and others &quot;to rise by being on the shoulder s of elders&quot;.</p>\n<p>@Damian: Just as a &quot;streaming video media player&quot; can show a video while it is being downloaded, a &quot;streaming dataset data player&quot; can show a dataset while it is being downloaded. I don't know whether it exists, but it is possible.</p>\n<p>Thanks.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55243",
      "postDate": "09/25/2014 17:55:59",
      "content": "<p>[quote=Piotr Kuchta;55219]</p>\n<p>BTW, regarding the up and down voting, am I the only one who preferred the previous forum with &quot;Thanks&quot; feature? I thought it was quite original and nice, as it was promoting development of friendly relationships between contestants. Personally I find the up and down buttons a bit dull, in contrast.</p>\n<p>[/quote]</p>\n<p>I agree, I liked the &quot;thank&quot; implementation&nbsp;and the sense of respect that it implied by attaching your name to it. An anonymous system doesn't seem to make a lot of sense in this context.</p>\n<p>I'm also writing this post now as opposed to just thanking your idea before.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55751",
      "postDate": "10/07/2014 21:16:23",
      "content": "<p>How similar is the classifier in this benchmark to sklearn's random forest classifier?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55783",
      "postDate": "10/08/2014 12:58:25",
      "content": "<p>[quote=rcarson;55751]</p>\n<p>How similar is the classifier in this benchmark to sklearn's random forest classifier?</p>\n<p>[/quote]</p>\n\n<p>Not exactly sure of the precise differences as I haven't ever coded in Python but I'm sure that the general idea of the decision tree is conveyed in the algorithm. Is there anyone else who might be able to shed some light on the precise differences as I know that in other competitions different implementations have produces different results?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55784",
      "postDate": "10/08/2014 13:05:17",
      "content": "<p>according to the documentation about NVarToSample: &quot;<em>Setting this argument to any valid value but 'all' invokes Breiman's 'random forest' algorithm</em>.&quot;:</p>\n<p>http://www.mathworks.nl/help/stats/treebagger.html</p>\n\n<p>Main advantage of R that I am aware of (and have used in the past) is the &quot;experimental&quot; setting for corr.bias = TRUE, which does seem to help in regression context.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "56045",
      "postDate": "10/14/2014 10:43:28",
      "content": "<p>[quote=Shiraz University|Neda;56042]</p>\n<p>Thank U for sharing</p>\n<p>I run the code, in this code, we <strong>train a separate model for each subject</strong> and then predict by constructed model for identical&nbsp;subject. is it legal regarding to competition rules?&nbsp;</p>\n<p>[/quote]</p>\n\n<p>Yes, this is correct. From my interpretation of the rules this is a legal approach, I'm sure someone would have piped up by now if it had been illegal.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "56483",
      "postDate": "10/22/2014 12:03:39",
      "content": "<p>Thank you very much Elliot!</p>\n<p>Does it also work with octave?</p>\n<p>Thanks,</p>\n<p>C.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "56725",
      "postDate": "10/25/2014 10:41:58",
      "content": "<p>[quote=Elliot Dawson;53617]</p>\n<p>Hey guys,<br><br>So I noticed that roughly 33% of the leaderboard consists of people who have yet to beat the random benchmark of 0.5 and thought that since these people have made the effort to download the considerably large dataset I would lend a hand in providing some starter code. Not really sure if I'm going to have a great deal of time to work on this competition with my University schedule at the moment and given the positive externality that this competition intends to have I hope that this might provide someone else with a platform to produce a great algorithm.<br><br>This method achieves a public leaderboard score of around 0.635 and takes about 25 minutes to run using 6 workers in MATLAB (courtesy of the parallel computing toolbox) on a quad core i7 processor. My machine has 16GB of RAM and didn't have any memory issues when running the code. Having said that I'm unsure of how an 8GB RAM machine might cope. At the moment this approach works as follows:<br>- For each subject clip calculate : The variance of each channel as well as the correlation coefficient between each channel<br>- Take these values and use them as your predictor matrix<br>- For each subject train a decision tree using MATLAB's treeBagger with 1000 trees using the subjects' respective predictor matrix calculated from the clips<br>- Predict on each subjects' test data using their respective trained decision tree</p>\n<p>Just run BenchmarkCode once you've added FeatureEngineer2.m to the path and changed the directories from where to write and read the data to and from.<br><br>On a separate note I used the 'table' data type to import the sample submission and to write the resulting submission to file which I believe was only introduced in MATLAB version 2013b but of course you can modify the code to write the result however you like :)</p>\n<p>Let me know if you're having any issues with the code.</p>\n<p>[/quote]</p>\n<p>Hi,thanks for posting the code.</p>\n<p>I just downloaded it and yet still downloading dog2.</p>\n<p>Why the training feature for Dog1, trainDog1 ..3,4&nbsp; is 137 dimensions and for trainDog5 and trainHuman1 is 121 and for trainHuman2 is 301?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "56743",
      "postDate": "10/25/2014 13:17:57",
      "content": "<p>[quote=Steven Du;56725]</p>\n\n<p>Why the training feature for Dog1, trainDog1 ..3,4&nbsp; is 137 dimensions and for trainDog5 and trainHuman1 is 121 and for trainHuman2 is 301?</p>\n<p>[/quote]</p>\n\n<p>Because the number of channels is different: 16 channels for dogs 1 to 4, 15 for Dog 5 and Patient 1, 24 for Patient 2.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "56745",
      "postDate": "10/25/2014 13:26:49",
      "content": "<p>[quote=clustifier;56483]</p>\n<p>Thank you very much Elliot!</p>\n<p>Does it also work with octave?</p>\n<p>Thanks,</p>\n<p>C.</p>\n<p>[/quote]</p>\n<p>I'm not an Octave user myself so I'm not really sure, I think the syntax principles are the same so if there are any functions that are MATLAB exclusive you can always change/replace them to be compatible with Octave.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57307",
      "postDate": "11/04/2014 17:38:26",
      "content": "<p>Hi&nbsp;Elliot,</p>\n<p>Your code is producing this error when i am trying to run after following the instructions:</p>\n\n<p>Benchmarkcode<br>Index exceeds matrix dimensions.</p>\n<p>Error in FeatureEngineer2 (line 6)<br>file = load([x preictalClips(1).name]);</p>\n<p>Error in Benchmarkcode (line 10)<br>[preictalTrainDog1, interIctalTrainDog1, testDog1] = FeatureEngineer2('E:\\Dog_1');</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57333",
      "postDate": "11/05/2014 03:23:55",
      "content": "<p>Changing the file directory paths should fix the error.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 53660,
      "author_name": "franklyn",
      "author_url": "",
      "post_date": "09/13/2014 19:21:32",
      "content": "<p>A lot of those people also just have one entry, so i'm not too worried. Thanks for posting your approach.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 53692,
      "author_name": "joseleiva",
      "author_url": "",
      "post_date": "09/14/2014 14:36:37",
      "content": "<p>Elliot,</p>\n\n<p>Thanks a lot for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 54114,
      "author_name": "mela213006",
      "author_url": "",
      "post_date": "09/17/2014 02:42:02",
      "content": "<p>Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 54491,
      "author_name": "bgeier",
      "author_url": "",
      "post_date": "09/23/2014 01:17:36",
      "content": "<p>Thanks for sharing. Similarly, a&nbsp;lasso logistic regression model with only variance of each feature does equally well at ~0.64. Try cvglmnet.m from glmnet toolbox. This is an open source version of Matlab's lassoglm.m &nbsp;I ran on 8GB 2008 MacPro and had no problems with completion between&nbsp;45-60 minutes&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55201,
      "author_name": "damian4",
      "author_url": "",
      "post_date": "09/25/2014 06:24:53",
      "content": "<p>Thank you very much. I don't have the patience to wait for the download and don't want to bother with feature engineering today. Would it be possible to upload the hopefully smaller pre-processed data set here?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55213,
      "author_name": "julesvanligtenberg",
      "author_url": "",
      "post_date": "09/25/2014 09:57:25",
      "content": "<p>[quote=Damian Koesters;55201]</p>\n<p>Thank you very much. I don't have the patience to wait for the download and don't want to bother with feature engineering today. Would it be possible to upload the hopefully smaller pre-processed data set here?</p>\n<p>[/quote]</p>\n<p>@Damian</p>\n<p>Hi Damian,</p>\n<p>I voted your post down -1. And I don't want to do this anonymously. (for the admins: it is maybe a good idea to have 2 lines under each post: 'voted up by' and 'voted down by' to avoid teammates giving up-votes and trolling)</p>\n<p>You first have to accept the rules before you can download (if the interface works as I think it works you haven't done that yet).</p>\n<p>Furthermore there is a minimum of work involved getting the promised leaderboard score:</p>\n<p>Namely downloading the data and running the shared script.</p>\n<p>Sharing is disruptive enough as it is, if the new work modus will change to down and uploading preprocessed data, why not give everybody who enters the challenge the promised LB score.</p>\n<p>I hope you don't take this personally, everybody could have asked the same question, it is just that I want to preserve some of the meaning the score on the private LB has.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55216,
      "author_name": "damian4",
      "author_url": "",
      "post_date": "09/25/2014 11:21:40",
      "content": "<p>Hi Jules, ugh, downvoting really hurts but I won't take it personally. I just thought that it would be great for each challenge to have a set of reasonable pre-processing steps or an alternative pre-processed dataset to have a starting point.</p>\n<p>I want to try a simple idea on the classifier side without worrying about the feature engineering steps. The feature engineering is quite a hurdle as I have no knowledge about EEG-data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55218,
      "author_name": "konradb",
      "author_url": "",
      "post_date": "09/25/2014 12:31:12",
      "content": "<p>In the spirit of openness inspired by Jules: Damian, I downvoted you as well. While I understand your approach (you want to get to the fun part, i.e. modeling), feature engineering and data cleaning are a core part of the job. Elliot's code does quite a bit of heavy lifting for you, so saying &quot;<em>you can't be bothered</em>&quot; to do some of the extra work required (it's nothing more than downloading and running his code, for crying out loud) is quite inappropriate in my view.&nbsp;</p>\n\n<p>And as for lack of domain expertise:</p>\n<p>1. i don't think a lot of top contenders know a lot about EEG</p>\n<p>2. in the Avito contest, which dealt with text mining of data in Russian, hardly any of the top contenders spoke Russian&nbsp;</p>\n<p>so I wouldn't worry that much :-)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55219,
      "author_name": "piotrkuchta",
      "author_url": "",
      "post_date": "09/25/2014 12:52:27",
      "content": "<p>I agree that data cleaning and feature engineering is part of the challenge in most Kaggle competitions, and, in general, data science projects.</p>\n<p>BTW, regarding the up and down voting, am I the only one who preferred the previous forum with &quot;Thanks&quot; feature? I thought it was quite original and nice, as it was promoting development of friendly relationships between contestants. Personally I find the up and down buttons a bit dull, in contrast.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55221,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "09/25/2014 13:07:57",
      "content": "<p>Elliot, thank you for sharing this code. I learn a lot.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55223,
      "author_name": "elliotdawson",
      "author_url": "",
      "post_date": "09/25/2014 13:24:32",
      "content": "<p>Hi Damian, in response to your request, I agree with everyone else who has responded (however I have not downvoted your comment on the basis that you've applied the theory of <em>If you don't ask the answer is always no</em>) and given there is no prominent precedent of people providing preprocessed data directly I would rather not make it a trend. As for domain expertise, I would have to say that mine would total to 0, my current studies are in Actuarial science has given me no working knowledge of an EEG&nbsp; (unless I missed that lecture haha) yet this hasn't stopped me from creating some competent code to get started with.<br><br>When I posted this two weeks I had hoped that somebody down the bottom of the leaderboard who had a hard time getting started would modify both the preprocessing script and the modelling component of the code to create something a lot better. So I hope you do indeed download the data and I wish you the best of luck with the competition!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55236,
      "author_name": "lalitapatel",
      "author_url": "",
      "post_date": "09/25/2014 16:22:22",
      "content": "<p>@Elliot: Thanks for the code and your spirit of helping others. The code will definitely help me and others &quot;to rise by being on the shoulder s of elders&quot;.</p>\n<p>@Damian: Just as a &quot;streaming video media player&quot; can show a video while it is being downloaded, a &quot;streaming dataset data player&quot; can show a dataset while it is being downloaded. I don't know whether it exists, but it is possible.</p>\n<p>Thanks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55243,
      "author_name": "dkaylor",
      "author_url": "",
      "post_date": "09/25/2014 17:55:59",
      "content": "<p>[quote=Piotr Kuchta;55219]</p>\n<p>BTW, regarding the up and down voting, am I the only one who preferred the previous forum with &quot;Thanks&quot; feature? I thought it was quite original and nice, as it was promoting development of friendly relationships between contestants. Personally I find the up and down buttons a bit dull, in contrast.</p>\n<p>[/quote]</p>\n<p>I agree, I liked the &quot;thank&quot; implementation&nbsp;and the sense of respect that it implied by attaching your name to it. An anonymous system doesn't seem to make a lot of sense in this context.</p>\n<p>I'm also writing this post now as opposed to just thanking your idea before.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55751,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "10/07/2014 21:16:23",
      "content": "<p>How similar is the classifier in this benchmark to sklearn's random forest classifier?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55783,
      "author_name": "elliotdawson",
      "author_url": "",
      "post_date": "10/08/2014 12:58:25",
      "content": "<p>[quote=rcarson;55751]</p>\n<p>How similar is the classifier in this benchmark to sklearn's random forest classifier?</p>\n<p>[/quote]</p>\n\n<p>Not exactly sure of the precise differences as I haven't ever coded in Python but I'm sure that the general idea of the decision tree is conveyed in the algorithm. Is there anyone else who might be able to shed some light on the precise differences as I know that in other competitions different implementations have produces different results?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55784,
      "author_name": "konradb",
      "author_url": "",
      "post_date": "10/08/2014 13:05:17",
      "content": "<p>according to the documentation about NVarToSample: &quot;<em>Setting this argument to any valid value but 'all' invokes Breiman's 'random forest' algorithm</em>.&quot;:</p>\n<p>http://www.mathworks.nl/help/stats/treebagger.html</p>\n\n<p>Main advantage of R that I am aware of (and have used in the past) is the &quot;experimental&quot; setting for corr.bias = TRUE, which does seem to help in regression context.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 56045,
      "author_name": "elliotdawson",
      "author_url": "",
      "post_date": "10/14/2014 10:43:28",
      "content": "<p>[quote=Shiraz University|Neda;56042]</p>\n<p>Thank U for sharing</p>\n<p>I run the code, in this code, we <strong>train a separate model for each subject</strong> and then predict by constructed model for identical&nbsp;subject. is it legal regarding to competition rules?&nbsp;</p>\n<p>[/quote]</p>\n\n<p>Yes, this is correct. From my interpretation of the rules this is a legal approach, I'm sure someone would have piped up by now if it had been illegal.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 56483,
      "author_name": "clustifier",
      "author_url": "",
      "post_date": "10/22/2014 12:03:39",
      "content": "<p>Thank you very much Elliot!</p>\n<p>Does it also work with octave?</p>\n<p>Thanks,</p>\n<p>C.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 56725,
      "author_name": "stevendu",
      "author_url": "",
      "post_date": "10/25/2014 10:41:58",
      "content": "<p>[quote=Elliot Dawson;53617]</p>\n<p>Hey guys,<br><br>So I noticed that roughly 33% of the leaderboard consists of people who have yet to beat the random benchmark of 0.5 and thought that since these people have made the effort to download the considerably large dataset I would lend a hand in providing some starter code. Not really sure if I'm going to have a great deal of time to work on this competition with my University schedule at the moment and given the positive externality that this competition intends to have I hope that this might provide someone else with a platform to produce a great algorithm.<br><br>This method achieves a public leaderboard score of around 0.635 and takes about 25 minutes to run using 6 workers in MATLAB (courtesy of the parallel computing toolbox) on a quad core i7 processor. My machine has 16GB of RAM and didn't have any memory issues when running the code. Having said that I'm unsure of how an 8GB RAM machine might cope. At the moment this approach works as follows:<br>- For each subject clip calculate : The variance of each channel as well as the correlation coefficient between each channel<br>- Take these values and use them as your predictor matrix<br>- For each subject train a decision tree using MATLAB's treeBagger with 1000 trees using the subjects' respective predictor matrix calculated from the clips<br>- Predict on each subjects' test data using their respective trained decision tree</p>\n<p>Just run BenchmarkCode once you've added FeatureEngineer2.m to the path and changed the directories from where to write and read the data to and from.<br><br>On a separate note I used the 'table' data type to import the sample submission and to write the resulting submission to file which I believe was only introduced in MATLAB version 2013b but of course you can modify the code to write the result however you like :)</p>\n<p>Let me know if you're having any issues with the code.</p>\n<p>[/quote]</p>\n<p>Hi,thanks for posting the code.</p>\n<p>I just downloaded it and yet still downloading dog2.</p>\n<p>Why the training feature for Dog1, trainDog1 ..3,4&nbsp; is 137 dimensions and for trainDog5 and trainHuman1 is 121 and for trainHuman2 is 301?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 56743,
      "author_name": "joseleiva",
      "author_url": "",
      "post_date": "10/25/2014 13:17:57",
      "content": "<p>[quote=Steven Du;56725]</p>\n\n<p>Why the training feature for Dog1, trainDog1 ..3,4&nbsp; is 137 dimensions and for trainDog5 and trainHuman1 is 121 and for trainHuman2 is 301?</p>\n<p>[/quote]</p>\n\n<p>Because the number of channels is different: 16 channels for dogs 1 to 4, 15 for Dog 5 and Patient 1, 24 for Patient 2.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 56745,
      "author_name": "elliotdawson",
      "author_url": "",
      "post_date": "10/25/2014 13:26:49",
      "content": "<p>[quote=clustifier;56483]</p>\n<p>Thank you very much Elliot!</p>\n<p>Does it also work with octave?</p>\n<p>Thanks,</p>\n<p>C.</p>\n<p>[/quote]</p>\n<p>I'm not an Octave user myself so I'm not really sure, I think the syntax principles are the same so if there are any functions that are MATLAB exclusive you can always change/replace them to be compatible with Octave.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 57307,
      "author_name": "equester",
      "author_url": "",
      "post_date": "11/04/2014 17:38:26",
      "content": "<p>Hi&nbsp;Elliot,</p>\n<p>Your code is producing this error when i am trying to run after following the instructions:</p>\n\n<p>Benchmarkcode<br>Index exceeds matrix dimensions.</p>\n<p>Error in FeatureEngineer2 (line 6)<br>file = load([x preictalClips(1).name]);</p>\n<p>Error in Benchmarkcode (line 10)<br>[preictalTrainDog1, interIctalTrainDog1, testDog1] = FeatureEngineer2('E:\\Dog_1');</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 57333,
      "author_name": "markcheung",
      "author_url": "",
      "post_date": "11/05/2014 03:23:55",
      "content": "<p>Changing the file directory paths should fix the error.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "53617": "",
    "53660": "",
    "53692": "",
    "54114": "",
    "54491": "",
    "55201": "",
    "55213": "",
    "55216": "",
    "55218": "",
    "55219": "",
    "55221": "",
    "55223": "",
    "55236": "",
    "55243": "",
    "55751": "",
    "55783": "",
    "55784": "",
    "56045": "",
    "56483": "",
    "56725": "",
    "56743": "",
    "56745": "",
    "57307": "",
    "57333": ""
  },
  "source": "meta"
}