{
  "id": 4302,
  "title": "Methods and Algorithms",
  "url": "/competitions/challenges-in-representation-learning-the-black-box-learning-challenge/discussion/4302",
  "author_name": "",
  "post_date": "2013-04-14T00:36:30.353Z",
  "votes": null,
  "comment_count": 35,
  "views": 9385,
  "content": "<p>Just for fun I thought I'd see if anybody wants to talk about what methods and algorithms they've tried (or are working on).&nbsp; GIven the moniker of the early leader (i.e. RBM) I think the early leader's strategy is clear - and I may resurrect my own Restricted\r\n Boltzmann Machine implementation (circa 2008/2009) and give it a shot.</p>\r\n<p>But personally I wanted to try simpler things first - just to add to the &quot;benchmark&quot; collection.</p>\r\n<p>My first attempt was a plain vanilla multiclass random forest, followed by a 9-way attempt (making each feature binary).&nbsp; The latter scored higher: 0.36540, better than I thought it would.</p>\r\n<p>So far I haven't attempted to do anything (unsupervised or otherwise) with the unlabeled data, and I'm not sure I will.&nbsp; Depends on how much time I can scrape together for this.</p>",
  "messages": [
    {
      "id": "22757",
      "postDate": "04/14/2013 00:36:30",
      "content": "<p>Just for fun I thought I'd see if anybody wants to talk about what methods and algorithms they've tried (or are working on).&nbsp; GIven the moniker of the early leader (i.e. RBM) I think the early leader's strategy is clear - and I may resurrect my own Restricted\r\n Boltzmann Machine implementation (circa 2008/2009) and give it a shot.</p>\r\n<p>But personally I wanted to try simpler things first - just to add to the &quot;benchmark&quot; collection.</p>\r\n<p>My first attempt was a plain vanilla multiclass random forest, followed by a 9-way attempt (making each feature binary).&nbsp; The latter scored higher: 0.36540, better than I thought it would.</p>\r\n<p>So far I haven't attempted to do anything (unsupervised or otherwise) with the unlabeled data, and I'm not sure I will.&nbsp; Depends on how much time I can scrape together for this.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22758",
      "postDate": "04/14/2013 00:46:09",
      "content": "<p>Since I'm an organizer of the challenge, I won't be giving any opinions on what algorithm to use. But I think it's OK for me to point out that the baseline code provided for this challenge can be modified pretty easily to include RBMs if they're what you\r\n want to try.</p>\r\n<p>The baseline code for this challenge is here:</p>\r\n<p>http://github.com/lisa-lab/pylearn2/tree/master/pylearn2/scripts/icml_2013_wrepl/black_box</p>\r\n<p>The pylearn2 RBM is here:</p>\r\n<p>http://github.com/lisa-lab/pylearn2/blob/master/pylearn2/models/rbm.py</p>\r\n<p>It should be reasonably easy to modify the baseline demo script to incorporate RBMs, using the RBM_Layer class of the MLP:</p>\r\n<p>http://github.com/lisa-lab/pylearn2/blob/master/pylearn2/models/mlp.py</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22760",
      "postDate": "04/14/2013 00:52:42",
      "content": "<p>[quote=Ian Goodfellow;22758]</p>\r\n<p>...I think it's OK for me to point out that the baseline code provided for this challenge can be modified pretty easily to include RBMs if they're what you want to try...</p>\r\n<p>[/quote]</p>\r\n<p>Very true, but I have a &quot;thing&quot; about using as much of my own code as possible.&nbsp; That way nothing is really a black box.&nbsp; Just a personal quirk, though one that tends to fall by the wayside when it gets to crunch time.&nbsp; Pragmatism often trumps philosophy.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22801",
      "postDate": "04/14/2013 19:09:26",
      "content": "<p>I've put in a few submissions continuing my &quot;play with vw&quot; exploration... Using quadratic features gets me up into the range of the top three benchmarks. Whoo!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22869",
      "postDate": "04/15/2013 23:28:49",
      "content": "<p>[quote=Aaron Schumacher;22801]</p>\r\n<p>I've put in a few submissions continuing my &quot;play with vw&quot; exploration... Using quadratic features gets me up into the range of the top three benchmarks. Whoo!</p>\r\n<p>[/quote]</p>\r\n<p>Nice!&nbsp; I don't know much about vw (yet - it's on my list of things to learn), but from what I can tell it doesn't have any semi-supervised learning methods built in.&nbsp; So were you using just the labeled data?</p>\r\n<p>I thought I'd try some old-school semi-supervised methods next - gaussian mixture stuff most likely - just to see if there's any clearly discernable value in the unlabeled data.&nbsp; Memories of a prior (and similar) Kaggle competition (https://www.kaggle.com/c/SemiSupervisedFeatureLearning),\r\n where the winner didn't use the unlabeled data at all, have made me a bit gun-shy.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22874",
      "postDate": "04/16/2013 01:18:44",
      "content": "<p>This challenge has 1000 labeled examples, covering 9 classes which are not balanced either. This makes getting 0.99 AUC (or 99% accuracy) significantly more difficult (the underlying problem is also not particularly trivial, which adds to the complexity\r\n of course).</p>\r\n<p>Nonetheless, part of the point of this workshop is to add yet another data point to the debate on whether unsupervised learning is indeed helpful. A\r\n<em>hypothesis</em>&nbsp;is that it <em>is </em>helpful, and this competition will provide data to support it (or not).</p>\r\n<p>Dumitru</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22897",
      "postDate": "04/16/2013 14:42:59",
      "content": "<p>@YetiMan, yes, I don't know of any unsupervised/semi-supervised stuff in vw, so I was just using the labeled data. I'm not sure yet what I want to do with the unlabeled data. I thought about trying to label it based on my label-trained model and then train\r\n further using those guesses, but I don't know that that would be helpful. Thanks for the reference to the old comp! Interesting...</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22978",
      "postDate": "04/17/2013 15:24:12",
      "content": "<p>[quote=Aaron Schumacher;22897]</p>\r\n<p>@YetiMan, yes, I don't know of any unsupervised/semi-supervised stuff in vw, so I was just using the labeled data. I'm not sure yet what I want to do with the unlabeled data. I thought about trying to label it based on my label-trained model and then train\r\n further using those guesses, but I don't know that that would be helpful. Thanks for the reference to the old comp! Interesting...</p>\r\n<p>[/quote]</p>\r\n<p>I've used vw quite a bit, and i don't think they have unsupervised stuff. But o do know that this one has:&nbsp;<a href=\"https://code.google.com/p/sofia-ml/wiki/SofiaKMeans\">https://code.google.com/p/sofia-ml/wiki/SofiaKMeans</a></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22983",
      "postDate": "04/17/2013 16:21:24",
      "content": "<p>Neat! Yeah, I wasn't suggesting that vw could do anything unsupervised, I was suggesting the probably bad idea of using a label-trained model to label the unlabeled data and then training more off of that. Using a real unsupervised technique is likely to\r\n be better, I imagine. :)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22985",
      "postDate": "04/17/2013 17:00:13",
      "content": "<p>[quote=Aaron Schumacher;22983]</p>\r\n<p>Neat! Yeah, I wasn't suggesting that vw could do anything unsupervised, I was suggesting the probably bad idea of using a label-trained model to label the unlabeled data and then training more off of that. Using a real unsupervised technique is likely to\r\n be better, I imagine. :)</p>\r\n<p>[/quote]</p>\r\n<p>I tried that a couple of days ago, with labels imputed via gradient boosted decision trees (chosen only because I had the code handy).&nbsp; As you surmised, the results were horrible.</p>\r\n<p>To define &quot;horrible&quot; more concretely:</p>\r\n<ul>\r\n<li>Score with test data predicted directly via GBDT: 0.34 </li><li>Score with training data tripled (to 3000 samples) via GBDT imputation of labels: 0.21 (Ouch! But hardly surprising.)\r\n</li></ul>\r\n<p>Last night's attempt at semi-supervised learning (via TSVM) was also an abysmal failure.&nbsp; Much worse than either GBDT or &quot;random forest&quot; on only training data.&nbsp; To be fair, though, I only used 20% of the unlabeled data in order to decrease training time.&nbsp;\r\n So I haven't completely given up on using the unlabeled data, but thus far it's looking more harmful than helpful - which could simply be a sign that I'm choosing poor methods (or that there are bugs in my code, or that there's too much &quot;noise&quot; in the data\r\n overall, or...).</p>\r\n<p>Unfortunately the time I have available to spend on this is very limited, so I need to decide which path to take: supervised using only labeled data, semi-supervised, or unsupervised.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23238",
      "postDate": "04/20/2013 19:20:45",
      "content": "<p>[quote=Aaron Schumacher;22801]</p>\r\n<p>I've put in a few submissions continuing my &quot;play with vw&quot; exploration... Using quadratic features gets me up into the range of the top three benchmarks. Whoo!</p>\r\n<p>[/quote]</p>\r\n<p>By the way, what is vw?</p>\r\n<p>I tried to google it, and only results were related to volkswagen.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23239",
      "postDate": "04/20/2013 19:22:03",
      "content": "<p>Vowpal Wabbit</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23561",
      "postDate": "04/28/2013 01:40:44",
      "content": "<p>I've been wondering for a while whether the test set has the same distribution of classes as the training set.&nbsp; I've just submitted a probe full of 10000 &quot;8.0&quot;'s -- the result was 0.06520 which is pretty close to the 0.073 of the training set.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23624",
      "postDate": "04/29/2013 14:58:23",
      "content": "<p>Yop,</p>\r\n<p>I tried sciki-learn's <a href=\"http://scikit-learn.org/dev/modules/generated/sklearn.ensemble.RandomForestClassifier.html#sklearn.ensemble.RandomForestClassifier\">\r\nRandom Forests</a> w. n_estimators=100 and scored 0.34 using just the labelled data.</p>\r\n<p>GBDT, SVM, kNN, SGD, etc did perform horribly.</p>\r\n<p>I tried to learn it as a regression problem (as the class names might imply), but with no luck either using Random Forest Regressor.</p>\r\n<p>I also tried the <a href=\"http://scikit-learn.org/dev/modules/generated/sklearn.semi_supervised.LabelSpreading.html#sklearn.semi_supervised.LabelSpreading\">\r\nLabelPropagation</a>&nbsp;/ LabelSpreading implementations in sklearn to make use of unsupervised data but was quite unlucky so far (</p>\r\n<p>I'm wondering if this problem was not utterly designed to be tackled with (modern) neural networks (aka deep learning) &nbsp;stuff...</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23630",
      "postDate": "04/29/2013 16:04:25",
      "content": "<p>The problem was designed to be difficult, and we chose the amount of labeled and unlabeled data to give an advantage to algorithms that can run in semi-supervised mode. Beyond that we were not trying to make it work well with any particular kind of learning\r\n algorithm. I think the problem you're observing is that the learning algorithms you're trying to run only work well if they have good features as input, but here you need to learn the features.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23635",
      "postDate": "04/29/2013 16:30:17",
      "content": "<p>I am new to data crunching and meanwhile in the process of shifting from R to Python. I am an ordinary R user.</p>\r\n<p><span style=\"line-height:1.4em\">May I ask you guys a quick question that how much time does it take for you to get a result on your computer, given that algorithms and codes have been written? Does it take several hours running? Thank you!</span></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23638",
      "postDate": "04/29/2013 16:38:37",
      "content": "<p>Mmm... what i was trying to say was that deep learning seems a paradigm of choice for this problem given its popularity in the recent years. Or so it seems from my limited knowledge of the semi supervised field.&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>Of course if you or anyone else has pointers to other ss approaches that might work, i'm all ears ;-)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23639",
      "postDate": "04/29/2013 16:41:09",
      "content": "<p>Since I'm organizing the contest I probably shouldn't provide advice for how to compete in it, beyond helping people troubleshoot pylearn2 stuff.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23645",
      "postDate": "04/29/2013 17:00:06",
      "content": "<p>Of course that wasn't what I was asking !!!</p>\r\n<p>It's just that I was seeking new fields to explore and ideas ...</p>\r\n<p>if s.o. tells me I could look at the, say, manifold learning or co-training or self-training or ... he's not necessarily telling me how to solve the problem but rather educating me about existing, but possibly totally irrelevant approaches out there.\r\n</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23683",
      "postDate": "04/30/2013 04:32:10",
      "content": "<p>is there any example on how to use Restricted boltzmann machines (rbm.py) with a sample .yaml file</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23684",
      "postDate": "04/30/2013 04:35:19",
      "content": "<p>pylearn2/scripts/tutorials/grbm_smd and&nbsp;pylearn2/scripts/tutorials/deep_trainer both use RBMs.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23687",
      "postDate": "04/30/2013 07:44:07",
      "content": "<p>that is for a binary RBM</p>\r\n<p>Can you let us know which function in rbm.py is a multi-class rbm?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23698",
      "postDate": "04/30/2013 15:23:50",
      "content": "<p>Well, the hidden units are binary, but the visible units are real-valued.</p>\r\n<p>I don't think that rbm.py supports softmax variables, but dbm.py does. See pylearn2/scripts/tutorials/dbm_demo for a basic example of the interface.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23709",
      "postDate": "04/30/2013 16:19:39",
      "content": "<p>Thanks Ian.</p>\r\n<p>I guess the &quot;y&quot; layer must be softmax.</p>\r\n<p>rbm.yaml in the dbm_demo tutorial has no mention of the y interface</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23711",
      "postDate": "04/30/2013 16:23:56",
      "content": "<p>pylearn2.models.dbm.Softmax. There's not a nice flashy demo script for every feature of the library. If someone would like to make a pull request with more demos, that'd be great.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23712",
      "postDate": "04/30/2013 16:26:57",
      "content": "<p>ok, difficult for seasoned R programmers!</p>\r\n<p>visible_layer: !obj:pylearn2.models.dbm.BinaryVector</p>\r\n<p>to be replaced by</p>\r\n<p>visible_layer: !obj:pylearn2.models.dbm.Softmax</p>\r\n<p>&nbsp;</p>\r\n<p>Correct?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23713",
      "postDate": "04/30/2013 16:30:39",
      "content": "<p>Set the cost's supervised flag to True and use Softmax as the last hidden layer.</p>\r\n<p>The DBM code generally doesn't consider the class labels as &quot;visible&quot; because they're unobserved for test set data.</p>\r\n<p>The visible_layer for this challenge should be real-valued. I'm not sure if there's a real-valued DBM layer in the library right now.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23715",
      "postDate": "04/30/2013 17:03:33",
      "content": "<p>https://github.com/lisa-lab/pylearn2/tree/master/pylearn2/models/dbm</p>\r\n<p>There is no softmax here on github</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23716",
      "postDate": "04/30/2013 17:16:10",
      "content": "<p>There is no softmax function in dbm.py also.</p>\r\n<p>It is only in mlp.py - Using same syntax gives following error:</p>\r\n<p>TypeError: Softmax does not support the following keywords: init_bias_target_marginals. Did you mean irange?</p>\r\n<p><em>!obj:pylearn2.models.dbm.Softmax {</em><br>\r\n<em>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; layer_name: 'h1',</em><br>\r\n<em>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; init_bias_target_marginals: *raw_train,</em><br>\r\n<em>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; # Initialize the weights to all 0s</em><br>\r\n<em>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; irange: .0,</em><br>\r\n<em>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; n_classes: 9</em><br>\r\n<em>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; }</em></p>\r\n<p><br>\r\n<br>\r\n</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23717",
      "postDate": "04/30/2013 17:17:48",
      "content": "<p>https://github.com/lisa-lab/pylearn2/blob/master/pylearn2/models/dbm/__init__.py line 1612.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23920",
      "postDate": "05/05/2013 18:49:17",
      "content": "<p>My current algorithm (using python):</p>\r\n<ul>\r\n<li>20 minutes to pickle all data; </li><li>make a new train.pickle using PCA (30-100 factors in result set); </li><li>300 trees in random forest classifier; </li></ul>\r\n<p>Finally got 0.31 accuracy</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23929",
      "postDate": "05/06/2013 04:17:22",
      "content": "<ul>\r\n<li>select top 10% of factors using Anova F-value; </li><li>PCA to 30 factors on labeled and nonlabeled data; </li><li>binning all factors into 10 bins; </li><li>150 trees random forest; </li></ul>\r\n<p>0.35 accuracy</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23937",
      "postDate": "05/06/2013 12:16:35",
      "content": "<p>Just using Matlab NN toolbox without any pre processing.&nbsp;</p>\r\n<ul>\r\n<li>1875 inputs and 9 outputs with&nbsp;1 hidden layer with 450 neurons </li></ul>\r\n<p>0.40&nbsp;accuracy</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23988",
      "postDate": "05/07/2013 18:00:46",
      "content": "<p>[quote=Ulrich;23937]</p>\r\n<p>Just using Matlab NN toolbox without any pre processing.&nbsp;</p>\r\n<ul>\r\n<li>1875 inputs and 9 outputs with&nbsp;1 hidden layer with 450 neurons </li></ul>\r\n<p>0.40&nbsp;accuracy</p>\r\n<p>[/quote]</p>\r\n<p>I run the Matlab NN toolbox and for more than 100 neurons I have 'out of memory' problem on 4GB machine. Did you face any problem as such?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23994",
      "postDate": "05/07/2013 19:15:50",
      "content": "<p>Hi Rafael,</p>\r\n<p>&nbsp;Try using&nbsp;net.trainFcn = 'traingdx' as the matlab train function instead of the default 'trainlm' and you can have more than 100 neurons... also try to use a small initial learning rate if you use 2 hidden layers...</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "24013",
      "postDate": "05/08/2013 14:01:55",
      "content": "<p>Also&nbsp;&nbsp;&nbsp;&nbsp; net.efficiency.memoryReduction&nbsp;&nbsp;&nbsp;&nbsp; parameter may help.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 22758,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "04/14/2013 00:46:09",
      "content": "<p>Since I'm an organizer of the challenge, I won't be giving any opinions on what algorithm to use. But I think it's OK for me to point out that the baseline code provided for this challenge can be modified pretty easily to include RBMs if they're what you\r\n want to try.</p>\r\n<p>The baseline code for this challenge is here:</p>\r\n<p>http://github.com/lisa-lab/pylearn2/tree/master/pylearn2/scripts/icml_2013_wrepl/black_box</p>\r\n<p>The pylearn2 RBM is here:</p>\r\n<p>http://github.com/lisa-lab/pylearn2/blob/master/pylearn2/models/rbm.py</p>\r\n<p>It should be reasonably easy to modify the baseline demo script to incorporate RBMs, using the RBM_Layer class of the MLP:</p>\r\n<p>http://github.com/lisa-lab/pylearn2/blob/master/pylearn2/models/mlp.py</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22760,
      "author_name": "yetiman",
      "author_url": "",
      "post_date": "04/14/2013 00:52:42",
      "content": "<p>[quote=Ian Goodfellow;22758]</p>\r\n<p>...I think it's OK for me to point out that the baseline code provided for this challenge can be modified pretty easily to include RBMs if they're what you want to try...</p>\r\n<p>[/quote]</p>\r\n<p>Very true, but I have a &quot;thing&quot; about using as much of my own code as possible.&nbsp; That way nothing is really a black box.&nbsp; Just a personal quirk, though one that tends to fall by the wayside when it gets to crunch time.&nbsp; Pragmatism often trumps philosophy.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22801,
      "author_name": "ajschumacher",
      "author_url": "",
      "post_date": "04/14/2013 19:09:26",
      "content": "<p>I've put in a few submissions continuing my &quot;play with vw&quot; exploration... Using quadratic features gets me up into the range of the top three benchmarks. Whoo!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22869,
      "author_name": "yetiman",
      "author_url": "",
      "post_date": "04/15/2013 23:28:49",
      "content": "<p>[quote=Aaron Schumacher;22801]</p>\r\n<p>I've put in a few submissions continuing my &quot;play with vw&quot; exploration... Using quadratic features gets me up into the range of the top three benchmarks. Whoo!</p>\r\n<p>[/quote]</p>\r\n<p>Nice!&nbsp; I don't know much about vw (yet - it's on my list of things to learn), but from what I can tell it doesn't have any semi-supervised learning methods built in.&nbsp; So were you using just the labeled data?</p>\r\n<p>I thought I'd try some old-school semi-supervised methods next - gaussian mixture stuff most likely - just to see if there's any clearly discernable value in the unlabeled data.&nbsp; Memories of a prior (and similar) Kaggle competition (https://www.kaggle.com/c/SemiSupervisedFeatureLearning),\r\n where the winner didn't use the unlabeled data at all, have made me a bit gun-shy.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22874,
      "author_name": "dumitru0",
      "author_url": "",
      "post_date": "04/16/2013 01:18:44",
      "content": "<p>This challenge has 1000 labeled examples, covering 9 classes which are not balanced either. This makes getting 0.99 AUC (or 99% accuracy) significantly more difficult (the underlying problem is also not particularly trivial, which adds to the complexity\r\n of course).</p>\r\n<p>Nonetheless, part of the point of this workshop is to add yet another data point to the debate on whether unsupervised learning is indeed helpful. A\r\n<em>hypothesis</em>&nbsp;is that it <em>is </em>helpful, and this competition will provide data to support it (or not).</p>\r\n<p>Dumitru</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22897,
      "author_name": "ajschumacher",
      "author_url": "",
      "post_date": "04/16/2013 14:42:59",
      "content": "<p>@YetiMan, yes, I don't know of any unsupervised/semi-supervised stuff in vw, so I was just using the labeled data. I'm not sure yet what I want to do with the unlabeled data. I thought about trying to label it based on my label-trained model and then train\r\n further using those guesses, but I don't know that that would be helpful. Thanks for the reference to the old comp! Interesting...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22978,
      "author_name": "leustagos",
      "author_url": "",
      "post_date": "04/17/2013 15:24:12",
      "content": "<p>[quote=Aaron Schumacher;22897]</p>\r\n<p>@YetiMan, yes, I don't know of any unsupervised/semi-supervised stuff in vw, so I was just using the labeled data. I'm not sure yet what I want to do with the unlabeled data. I thought about trying to label it based on my label-trained model and then train\r\n further using those guesses, but I don't know that that would be helpful. Thanks for the reference to the old comp! Interesting...</p>\r\n<p>[/quote]</p>\r\n<p>I've used vw quite a bit, and i don't think they have unsupervised stuff. But o do know that this one has:&nbsp;<a href=\"https://code.google.com/p/sofia-ml/wiki/SofiaKMeans\">https://code.google.com/p/sofia-ml/wiki/SofiaKMeans</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22983,
      "author_name": "ajschumacher",
      "author_url": "",
      "post_date": "04/17/2013 16:21:24",
      "content": "<p>Neat! Yeah, I wasn't suggesting that vw could do anything unsupervised, I was suggesting the probably bad idea of using a label-trained model to label the unlabeled data and then training more off of that. Using a real unsupervised technique is likely to\r\n be better, I imagine. :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22985,
      "author_name": "yetiman",
      "author_url": "",
      "post_date": "04/17/2013 17:00:13",
      "content": "<p>[quote=Aaron Schumacher;22983]</p>\r\n<p>Neat! Yeah, I wasn't suggesting that vw could do anything unsupervised, I was suggesting the probably bad idea of using a label-trained model to label the unlabeled data and then training more off of that. Using a real unsupervised technique is likely to\r\n be better, I imagine. :)</p>\r\n<p>[/quote]</p>\r\n<p>I tried that a couple of days ago, with labels imputed via gradient boosted decision trees (chosen only because I had the code handy).&nbsp; As you surmised, the results were horrible.</p>\r\n<p>To define &quot;horrible&quot; more concretely:</p>\r\n<ul>\r\n<li>Score with test data predicted directly via GBDT: 0.34 </li><li>Score with training data tripled (to 3000 samples) via GBDT imputation of labels: 0.21 (Ouch! But hardly surprising.)\r\n</li></ul>\r\n<p>Last night's attempt at semi-supervised learning (via TSVM) was also an abysmal failure.&nbsp; Much worse than either GBDT or &quot;random forest&quot; on only training data.&nbsp; To be fair, though, I only used 20% of the unlabeled data in order to decrease training time.&nbsp;\r\n So I haven't completely given up on using the unlabeled data, but thus far it's looking more harmful than helpful - which could simply be a sign that I'm choosing poor methods (or that there are bugs in my code, or that there's too much &quot;noise&quot; in the data\r\n overall, or...).</p>\r\n<p>Unfortunately the time I have available to spend on this is very limited, so I need to decide which path to take: supervised using only labeled data, semi-supervised, or unsupervised.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23238,
      "author_name": "benoitplante",
      "author_url": "",
      "post_date": "04/20/2013 19:20:45",
      "content": "<p>[quote=Aaron Schumacher;22801]</p>\r\n<p>I've put in a few submissions continuing my &quot;play with vw&quot; exploration... Using quadratic features gets me up into the range of the top three benchmarks. Whoo!</p>\r\n<p>[/quote]</p>\r\n<p>By the way, what is vw?</p>\r\n<p>I tried to google it, and only results were related to volkswagen.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23239,
      "author_name": "dumitru0",
      "author_url": "",
      "post_date": "04/20/2013 19:22:03",
      "content": "<p>Vowpal Wabbit</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23561,
      "author_name": "johnclarke",
      "author_url": "",
      "post_date": "04/28/2013 01:40:44",
      "content": "<p>I've been wondering for a while whether the test set has the same distribution of classes as the training set.&nbsp; I've just submitted a probe full of 10000 &quot;8.0&quot;'s -- the result was 0.06520 which is pretty close to the 0.073 of the training set.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23624,
      "author_name": "eustache",
      "author_url": "",
      "post_date": "04/29/2013 14:58:23",
      "content": "<p>Yop,</p>\r\n<p>I tried sciki-learn's <a href=\"http://scikit-learn.org/dev/modules/generated/sklearn.ensemble.RandomForestClassifier.html#sklearn.ensemble.RandomForestClassifier\">\r\nRandom Forests</a> w. n_estimators=100 and scored 0.34 using just the labelled data.</p>\r\n<p>GBDT, SVM, kNN, SGD, etc did perform horribly.</p>\r\n<p>I tried to learn it as a regression problem (as the class names might imply), but with no luck either using Random Forest Regressor.</p>\r\n<p>I also tried the <a href=\"http://scikit-learn.org/dev/modules/generated/sklearn.semi_supervised.LabelSpreading.html#sklearn.semi_supervised.LabelSpreading\">\r\nLabelPropagation</a>&nbsp;/ LabelSpreading implementations in sklearn to make use of unsupervised data but was quite unlucky so far (</p>\r\n<p>I'm wondering if this problem was not utterly designed to be tackled with (modern) neural networks (aka deep learning) &nbsp;stuff...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23630,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "04/29/2013 16:04:25",
      "content": "<p>The problem was designed to be difficult, and we chose the amount of labeled and unlabeled data to give an advantage to algorithms that can run in semi-supervised mode. Beyond that we were not trying to make it work well with any particular kind of learning\r\n algorithm. I think the problem you're observing is that the learning algorithms you're trying to run only work well if they have good features as input, but here you need to learn the features.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23635,
      "author_name": "weiden",
      "author_url": "",
      "post_date": "04/29/2013 16:30:17",
      "content": "<p>I am new to data crunching and meanwhile in the process of shifting from R to Python. I am an ordinary R user.</p>\r\n<p><span style=\"line-height:1.4em\">May I ask you guys a quick question that how much time does it take for you to get a result on your computer, given that algorithms and codes have been written? Does it take several hours running? Thank you!</span></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23638,
      "author_name": "eustache",
      "author_url": "",
      "post_date": "04/29/2013 16:38:37",
      "content": "<p>Mmm... what i was trying to say was that deep learning seems a paradigm of choice for this problem given its popularity in the recent years. Or so it seems from my limited knowledge of the semi supervised field.&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>Of course if you or anyone else has pointers to other ss approaches that might work, i'm all ears ;-)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23639,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "04/29/2013 16:41:09",
      "content": "<p>Since I'm organizing the contest I probably shouldn't provide advice for how to compete in it, beyond helping people troubleshoot pylearn2 stuff.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23645,
      "author_name": "eustache",
      "author_url": "",
      "post_date": "04/29/2013 17:00:06",
      "content": "<p>Of course that wasn't what I was asking !!!</p>\r\n<p>It's just that I was seeking new fields to explore and ideas ...</p>\r\n<p>if s.o. tells me I could look at the, say, manifold learning or co-training or self-training or ... he's not necessarily telling me how to solve the problem but rather educating me about existing, but possibly totally irrelevant approaches out there.\r\n</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23683,
      "author_name": "rkirana",
      "author_url": "",
      "post_date": "04/30/2013 04:32:10",
      "content": "<p>is there any example on how to use Restricted boltzmann machines (rbm.py) with a sample .yaml file</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23684,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "04/30/2013 04:35:19",
      "content": "<p>pylearn2/scripts/tutorials/grbm_smd and&nbsp;pylearn2/scripts/tutorials/deep_trainer both use RBMs.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23687,
      "author_name": "rkirana",
      "author_url": "",
      "post_date": "04/30/2013 07:44:07",
      "content": "<p>that is for a binary RBM</p>\r\n<p>Can you let us know which function in rbm.py is a multi-class rbm?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23698,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "04/30/2013 15:23:50",
      "content": "<p>Well, the hidden units are binary, but the visible units are real-valued.</p>\r\n<p>I don't think that rbm.py supports softmax variables, but dbm.py does. See pylearn2/scripts/tutorials/dbm_demo for a basic example of the interface.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23709,
      "author_name": "rkirana",
      "author_url": "",
      "post_date": "04/30/2013 16:19:39",
      "content": "<p>Thanks Ian.</p>\r\n<p>I guess the &quot;y&quot; layer must be softmax.</p>\r\n<p>rbm.yaml in the dbm_demo tutorial has no mention of the y interface</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23711,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "04/30/2013 16:23:56",
      "content": "<p>pylearn2.models.dbm.Softmax. There's not a nice flashy demo script for every feature of the library. If someone would like to make a pull request with more demos, that'd be great.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23712,
      "author_name": "rkirana",
      "author_url": "",
      "post_date": "04/30/2013 16:26:57",
      "content": "<p>ok, difficult for seasoned R programmers!</p>\r\n<p>visible_layer: !obj:pylearn2.models.dbm.BinaryVector</p>\r\n<p>to be replaced by</p>\r\n<p>visible_layer: !obj:pylearn2.models.dbm.Softmax</p>\r\n<p>&nbsp;</p>\r\n<p>Correct?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23713,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "04/30/2013 16:30:39",
      "content": "<p>Set the cost's supervised flag to True and use Softmax as the last hidden layer.</p>\r\n<p>The DBM code generally doesn't consider the class labels as &quot;visible&quot; because they're unobserved for test set data.</p>\r\n<p>The visible_layer for this challenge should be real-valued. I'm not sure if there's a real-valued DBM layer in the library right now.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23715,
      "author_name": "rkirana",
      "author_url": "",
      "post_date": "04/30/2013 17:03:33",
      "content": "<p>https://github.com/lisa-lab/pylearn2/tree/master/pylearn2/models/dbm</p>\r\n<p>There is no softmax here on github</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23716,
      "author_name": "rkirana",
      "author_url": "",
      "post_date": "04/30/2013 17:16:10",
      "content": "<p>There is no softmax function in dbm.py also.</p>\r\n<p>It is only in mlp.py - Using same syntax gives following error:</p>\r\n<p>TypeError: Softmax does not support the following keywords: init_bias_target_marginals. Did you mean irange?</p>\r\n<p><em>!obj:pylearn2.models.dbm.Softmax {</em><br>\r\n<em>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; layer_name: 'h1',</em><br>\r\n<em>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; init_bias_target_marginals: *raw_train,</em><br>\r\n<em>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; # Initialize the weights to all 0s</em><br>\r\n<em>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; irange: .0,</em><br>\r\n<em>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; n_classes: 9</em><br>\r\n<em>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; }</em></p>\r\n<p><br>\r\n<br>\r\n</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23717,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "04/30/2013 17:17:48",
      "content": "<p>https://github.com/lisa-lab/pylearn2/blob/master/pylearn2/models/dbm/__init__.py line 1612.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23920,
      "author_name": "vnefedov",
      "author_url": "",
      "post_date": "05/05/2013 18:49:17",
      "content": "<p>My current algorithm (using python):</p>\r\n<ul>\r\n<li>20 minutes to pickle all data; </li><li>make a new train.pickle using PCA (30-100 factors in result set); </li><li>300 trees in random forest classifier; </li></ul>\r\n<p>Finally got 0.31 accuracy</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23929,
      "author_name": "vnefedov",
      "author_url": "",
      "post_date": "05/06/2013 04:17:22",
      "content": "<ul>\r\n<li>select top 10% of factors using Anova F-value; </li><li>PCA to 30 factors on labeled and nonlabeled data; </li><li>binning all factors into 10 bins; </li><li>150 trees random forest; </li></ul>\r\n<p>0.35 accuracy</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23937,
      "author_name": "ulrich2",
      "author_url": "",
      "post_date": "05/06/2013 12:16:35",
      "content": "<p>Just using Matlab NN toolbox without any pre processing.&nbsp;</p>\r\n<ul>\r\n<li>1875 inputs and 9 outputs with&nbsp;1 hidden layer with 450 neurons </li></ul>\r\n<p>0.40&nbsp;accuracy</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23988,
      "author_name": "potamitis",
      "author_url": "",
      "post_date": "05/07/2013 18:00:46",
      "content": "<p>[quote=Ulrich;23937]</p>\r\n<p>Just using Matlab NN toolbox without any pre processing.&nbsp;</p>\r\n<ul>\r\n<li>1875 inputs and 9 outputs with&nbsp;1 hidden layer with 450 neurons </li></ul>\r\n<p>0.40&nbsp;accuracy</p>\r\n<p>[/quote]</p>\r\n<p>I run the Matlab NN toolbox and for more than 100 neurons I have 'out of memory' problem on 4GB machine. Did you face any problem as such?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23994,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "05/07/2013 19:15:50",
      "content": "<p>Hi Rafael,</p>\r\n<p>&nbsp;Try using&nbsp;net.trainFcn = 'traingdx' as the matlab train function instead of the default 'trainlm' and you can have more than 100 neurons... also try to use a small initial learning rate if you use 2 hidden layers...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 24013,
      "author_name": "ccccat",
      "author_url": "",
      "post_date": "05/08/2013 14:01:55",
      "content": "<p>Also&nbsp;&nbsp;&nbsp;&nbsp; net.efficiency.memoryReduction&nbsp;&nbsp;&nbsp;&nbsp; parameter may help.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "22757": "",
    "22758": "",
    "22760": "",
    "22801": "",
    "22869": "",
    "22874": "",
    "22897": "",
    "22978": "",
    "22983": "",
    "22985": "",
    "23238": "",
    "23239": "",
    "23561": "",
    "23624": "",
    "23630": "",
    "23635": "",
    "23638": "",
    "23639": "",
    "23645": "",
    "23683": "",
    "23684": "",
    "23687": "",
    "23698": "",
    "23709": "",
    "23711": "",
    "23712": "",
    "23713": "",
    "23715": "",
    "23716": "",
    "23717": "",
    "23920": "",
    "23929": "",
    "23937": "",
    "23988": "",
    "23994": "",
    "24013": ""
  },
  "source": "meta"
}