{
  "id": 4415,
  "title": "SVM Benchmark Code (R)",
  "url": "/competitions/challenges-in-representation-learning-the-black-box-learning-challenge/discussion/4415",
  "author_name": "",
  "post_date": "2013-04-23T18:26:13.330Z",
  "votes": 8,
  "comment_count": 8,
  "views": 3906,
  "content": "I don't think I'm going to have the time or energy to get pylearn2 going, and I don't think I'll be competitive without it. Given that, I decided to try and win the karma game instead of the real contest. Please see the attached analysis notebook to see\r\n simple code that trains a simple SVM for this contest. The displayed CV-error is fairly accurate; this lands you just below the worst pylearn2 benchmark. If you skip the grid-search part of the code, this runs in no time. I hope someone learns something or\r\n uses this in a nice ensemble. Best of luck!",
  "messages": [
    {
      "id": "23371",
      "postDate": "04/23/2013 18:26:13",
      "content": "I don't think I'm going to have the time or energy to get pylearn2 going, and I don't think I'll be competitive without it. Given that, I decided to try and win the karma game instead of the real contest. Please see the attached analysis notebook to see\r\n simple code that trains a simple SVM for this contest. The displayed CV-error is fairly accurate; this lands you just below the worst pylearn2 benchmark. If you skip the grid-search part of the code, this runs in no time. I hope someone learns something or\r\n uses this in a nice ensemble. Best of luck!",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23396",
      "postDate": "04/24/2013 01:45:09",
      "content": "<p>[quote=Shea Parkes;23371]</p>\r\n<p>I don't think I'm going to have the time or energy to get pylearn2 going, and I don't think I'll be competitive without it. Given that, I decided to try and win the karma game instead of the real contest. Please see the attached analysis notebook to see\r\n simple code that trains a simple SVM for this contest. The displayed CV-error is fairly accurate; this lands you just below the worst pylearn2 benchmark. If you skip the grid-search part of the code, this runs in no time. I hope someone learns something or\r\n uses this in a nice ensemble. Best of luck!</p>\r\n<p>[/quote]</p>\r\n<p>&nbsp;</p>\r\n<p>so far we managed to get past the best py2learn benchmark without a single line in python! So the game is not over yet for you...</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23401",
      "postDate": "04/24/2013 06:13:12",
      "content": "<p>Shea,</p>\r\n<p>Given that the sizes of the training and testing data sets is different, were you able to get the predictions for all the 10000 test samples .&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23413",
      "postDate": "04/24/2013 14:16:07",
      "content": "The vast majority of learning algorithms do not require equal sized training and testing samples. They require the same features/columns to be present on each sample, but not the same number of samples/rows. Since the test data has the same number of features\r\n and they have a similar enough distribution, it is very smooth to make predictions on all 10k testing samples. Of more concern is the fact that the training data has more features than observations. This means you will need to be extra careful to not overfit\r\n (learn too much). In my example code, the C parameter helps to regularize and encourage stable learning. You'll notice that C=0.1 might have given slightly better cv-error than C=1, but I still went with C=1. The testing feature distribution is not actually\r\n the same as the training feature distribution, so I'd like to extrapolate with a slightly simpler model. Also, C=1 is often a natural fit for SVMs and I put a strong subjective prior on it that requires strong evidence to deviate from it.",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23422",
      "postDate": "04/24/2013 17:36:15",
      "content": "<p>Just for fun (and comparison purposes) I grabbed a copy of libsvm and, with an hour of hunt-and-peck-style searching for good parameters, ended up with this:</p>\r\n<pre style=\"font-family:monospace; font-size:9pt\"><em>(after converting train.csv and test.csv into svmlight format yielding train.dat and test.dat)</em><br><br>% svm-scale -s scale.dat train.dat &gt;train.scaled.dat<br>% svm-scale -r scale.dat test.dat &gt;test.scaled.dat<br>% svm-train -s 1 -t 1 -d 2 -e 0.0005 -n 0.25 train.scaled.dat model.dat<br>% svm-predict test.scaled.dat model.dat test.predict.txt</pre>\r\n<p>It didn't do quite as well as Shea's R SVM, most likely because I didn't spend much time optimizing the training parameters, but still managed to score above 0.4.&nbsp; No doubt it could be improved considerably.</p>\r\n<p>Anybody want to give <a title=\"svmlin\" href=\"http://vikas.sindhwani.org/svmlin.html\">\r\nsvmlin</a> (semi-supervised SVM) a shot?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23423",
      "postDate": "04/24/2013 17:39:06",
      "content": "<p>My personal favorite SVM implementation, at least for dense data like this, is the matlab implementation of the L2-SVM that my friend Adam Coates wrote for his ICML paper on dictionary learning methods a while back:&nbsp;http://www.stanford.edu/~acoates/papers/kmeans_demo.tgz</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23436",
      "postDate": "04/24/2013 21:22:47",
      "content": "<p>[quote=Leustagos;23396]</p>\r\n<p>[quote=Shea Parkes;23371]</p>\r\n<p>I don't think I'm going to have the time or energy to get pylearn2 going, and I don't think I'll be competitive without it. Given that, I decided to try and win the karma game instead of the real contest. Please see the attached analysis notebook to see\r\n simple code that trains a simple SVM for this contest. The displayed CV-error is fairly accurate; this lands you just below the worst pylearn2 benchmark. If you skip the grid-search part of the code, this runs in no time. I hope someone learns something or\r\n uses this in a nice ensemble. Best of luck!</p>\r\n<p>[/quote]</p>\r\n<p>&nbsp;</p>\r\n<p>so far we managed to get past the best py2learn benchmark without a single line in python! So the game is not over yet for you...</p>\r\n<p>[/quote]</p>\r\n<p>I'm a bit confused. In mlp.yaml say the output submission get a 36.6% accuracy but &nbsp;leaderboard of this is 0.5164, and you say a line more to the code get 0.5396.</p>\r\n<p>Is this the code in the pylearn2/scripts directory in github or is other one?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23440",
      "postDate": "04/24/2013 23:06:37",
      "content": "<p>[quote=José A. Guerrero;23436]</p>\r\n<p>[quote=Leustagos;23396]</p>\r\n<p>[quote=Shea Parkes;23371]</p>\r\n<p>I don't think I'm going to have the time or energy to get pylearn2 going, and I don't think I'll be competitive without it. Given that, I decided to try and win the karma game instead of the real contest. Please see the attached analysis notebook to see\r\n simple code that trains a simple SVM for this contest. The displayed CV-error is fairly accurate; this lands you just below the worst pylearn2 benchmark. If you skip the grid-search part of the code, this runs in no time. I hope someone learns something or\r\n uses this in a nice ensemble. Best of luck!</p>\r\n<p>[/quote]</p>\r\n<p>&nbsp;</p>\r\n<p>so far we managed to get past the best py2learn benchmark without a single line in python! So the game is not over yet for you...</p>\r\n<p>[/quote]</p>\r\n<p>I'm a bit confused. In mlp.yaml say the output submission get a 36.6% accuracy but &nbsp;leaderboard of this is 0.5164, and you say a line more to the code get 0.5396.</p>\r\n<p>Is this the code in the pylearn2/scripts directory in github or is other one?</p>\r\n<p>[/quote]</p>\r\n<p>&nbsp;</p>\r\n<p>You misundestood me. I said that we got 0.53 WITHOUT using python. we didn't use pylearn2 or any other python library. Just R and matlab.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23632",
      "postDate": "04/29/2013 16:11:40",
      "content": "<p>José is right that there was a problem in the README. It should have said 51.5% accuracy like on the benchmark on the website, not 36.6%. 36.6% was the accuracy of an earlier baseline that I got rid of, but I forgot to update the README. I just submitted\r\n a pull request to pylearn2 to fix the README.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 23396,
      "author_name": "leustagos",
      "author_url": "",
      "post_date": "04/24/2013 01:45:09",
      "content": "<p>[quote=Shea Parkes;23371]</p>\r\n<p>I don't think I'm going to have the time or energy to get pylearn2 going, and I don't think I'll be competitive without it. Given that, I decided to try and win the karma game instead of the real contest. Please see the attached analysis notebook to see\r\n simple code that trains a simple SVM for this contest. The displayed CV-error is fairly accurate; this lands you just below the worst pylearn2 benchmark. If you skip the grid-search part of the code, this runs in no time. I hope someone learns something or\r\n uses this in a nice ensemble. Best of luck!</p>\r\n<p>[/quote]</p>\r\n<p>&nbsp;</p>\r\n<p>so far we managed to get past the best py2learn benchmark without a single line in python! So the game is not over yet for you...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23401,
      "author_name": "nsankar",
      "author_url": "",
      "post_date": "04/24/2013 06:13:12",
      "content": "<p>Shea,</p>\r\n<p>Given that the sizes of the training and testing data sets is different, were you able to get the predictions for all the 10000 test samples .&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23413,
      "author_name": "biasvariance",
      "author_url": "",
      "post_date": "04/24/2013 14:16:07",
      "content": "The vast majority of learning algorithms do not require equal sized training and testing samples. They require the same features/columns to be present on each sample, but not the same number of samples/rows. Since the test data has the same number of features\r\n and they have a similar enough distribution, it is very smooth to make predictions on all 10k testing samples. Of more concern is the fact that the training data has more features than observations. This means you will need to be extra careful to not overfit\r\n (learn too much). In my example code, the C parameter helps to regularize and encourage stable learning. You'll notice that C=0.1 might have given slightly better cv-error than C=1, but I still went with C=1. The testing feature distribution is not actually\r\n the same as the training feature distribution, so I'd like to extrapolate with a slightly simpler model. Also, C=1 is often a natural fit for SVMs and I put a strong subjective prior on it that requires strong evidence to deviate from it.",
      "votes": null,
      "replies": []
    },
    {
      "id": 23422,
      "author_name": "yetiman",
      "author_url": "",
      "post_date": "04/24/2013 17:36:15",
      "content": "<p>Just for fun (and comparison purposes) I grabbed a copy of libsvm and, with an hour of hunt-and-peck-style searching for good parameters, ended up with this:</p>\r\n<pre style=\"font-family:monospace; font-size:9pt\"><em>(after converting train.csv and test.csv into svmlight format yielding train.dat and test.dat)</em><br><br>% svm-scale -s scale.dat train.dat &gt;train.scaled.dat<br>% svm-scale -r scale.dat test.dat &gt;test.scaled.dat<br>% svm-train -s 1 -t 1 -d 2 -e 0.0005 -n 0.25 train.scaled.dat model.dat<br>% svm-predict test.scaled.dat model.dat test.predict.txt</pre>\r\n<p>It didn't do quite as well as Shea's R SVM, most likely because I didn't spend much time optimizing the training parameters, but still managed to score above 0.4.&nbsp; No doubt it could be improved considerably.</p>\r\n<p>Anybody want to give <a title=\"svmlin\" href=\"http://vikas.sindhwani.org/svmlin.html\">\r\nsvmlin</a> (semi-supervised SVM) a shot?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23423,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "04/24/2013 17:39:06",
      "content": "<p>My personal favorite SVM implementation, at least for dense data like this, is the matlab implementation of the L2-SVM that my friend Adam Coates wrote for his ICML paper on dictionary learning methods a while back:&nbsp;http://www.stanford.edu/~acoates/papers/kmeans_demo.tgz</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23436,
      "author_name": "blindape",
      "author_url": "",
      "post_date": "04/24/2013 21:22:47",
      "content": "<p>[quote=Leustagos;23396]</p>\r\n<p>[quote=Shea Parkes;23371]</p>\r\n<p>I don't think I'm going to have the time or energy to get pylearn2 going, and I don't think I'll be competitive without it. Given that, I decided to try and win the karma game instead of the real contest. Please see the attached analysis notebook to see\r\n simple code that trains a simple SVM for this contest. The displayed CV-error is fairly accurate; this lands you just below the worst pylearn2 benchmark. If you skip the grid-search part of the code, this runs in no time. I hope someone learns something or\r\n uses this in a nice ensemble. Best of luck!</p>\r\n<p>[/quote]</p>\r\n<p>&nbsp;</p>\r\n<p>so far we managed to get past the best py2learn benchmark without a single line in python! So the game is not over yet for you...</p>\r\n<p>[/quote]</p>\r\n<p>I'm a bit confused. In mlp.yaml say the output submission get a 36.6% accuracy but &nbsp;leaderboard of this is 0.5164, and you say a line more to the code get 0.5396.</p>\r\n<p>Is this the code in the pylearn2/scripts directory in github or is other one?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23440,
      "author_name": "leustagos",
      "author_url": "",
      "post_date": "04/24/2013 23:06:37",
      "content": "<p>[quote=José A. Guerrero;23436]</p>\r\n<p>[quote=Leustagos;23396]</p>\r\n<p>[quote=Shea Parkes;23371]</p>\r\n<p>I don't think I'm going to have the time or energy to get pylearn2 going, and I don't think I'll be competitive without it. Given that, I decided to try and win the karma game instead of the real contest. Please see the attached analysis notebook to see\r\n simple code that trains a simple SVM for this contest. The displayed CV-error is fairly accurate; this lands you just below the worst pylearn2 benchmark. If you skip the grid-search part of the code, this runs in no time. I hope someone learns something or\r\n uses this in a nice ensemble. Best of luck!</p>\r\n<p>[/quote]</p>\r\n<p>&nbsp;</p>\r\n<p>so far we managed to get past the best py2learn benchmark without a single line in python! So the game is not over yet for you...</p>\r\n<p>[/quote]</p>\r\n<p>I'm a bit confused. In mlp.yaml say the output submission get a 36.6% accuracy but &nbsp;leaderboard of this is 0.5164, and you say a line more to the code get 0.5396.</p>\r\n<p>Is this the code in the pylearn2/scripts directory in github or is other one?</p>\r\n<p>[/quote]</p>\r\n<p>&nbsp;</p>\r\n<p>You misundestood me. I said that we got 0.53 WITHOUT using python. we didn't use pylearn2 or any other python library. Just R and matlab.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23632,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "04/29/2013 16:11:40",
      "content": "<p>José is right that there was a problem in the README. It should have said 51.5% accuracy like on the benchmark on the website, not 36.6%. 36.6% was the accuracy of an earlier baseline that I got rid of, but I forgot to update the README. I just submitted\r\n a pull request to pylearn2 to fix the README.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "23371": "",
    "23396": "",
    "23401": "",
    "23413": "",
    "23422": "",
    "23423": "",
    "23436": "",
    "23440": "",
    "23632": ""
  },
  "source": "meta"
}