{
  "id": 4674,
  "title": "models",
  "url": "/competitions/challenges-in-representation-learning-the-black-box-learning-challenge/writeups/bing-xu-models",
  "author_name": "",
  "post_date": "2013-05-25T03:11:42.483Z",
  "votes": 5,
  "comment_count": 24,
  "views": 11818,
  "content": "<p>I use a 10 layers network as base, ensemble 400 such networks, and other svm/random forest</p>\r\n<p></p>\r\n<p>but overfit</p>\r\n<p></p>\r\n<p>The structure is like:</p>\r\n<p>INPUT-Autoencoder-Autoencoder-Autoencoder-Autoencoder-Maxout-Rectified Linear-Maxout-Softmax-Argmax</p>\r\n<p>Unsupervised Learning help the&nbsp;Maxout-Rectified Linear-Maxout-Softmax-Argmax network a lot, but contribute to overfit.</p>\r\n<p></p>\r\n<p>My classmate AuroraXie doesn't make complex ensemble only independently use one deep network, and got similar score in both public/private board.</p>\r\n<p>I ensemble 1000&#43; different complex models, and drop from 3rd in public to &nbsp;7th in private</p>\r\n<p></p>\r\n<p>Congratulations to the winners!</p>\r\n<p>And does anyone else, eg 3rd or 4th use deep network? It seems deep network fails in this time.</p>\r\n<p></p>\r\n<p></p>",
  "messages": [
    {
      "id": "24773",
      "postDate": "05/25/2013 00:48:29",
      "content": "<p>I use a 10 layers network as base, ensemble 400 such networks, and other svm/random forest</p>\r\n<p></p>\r\n<p>but overfit</p>\r\n<p></p>\r\n<p>The structure is like:</p>\r\n<p>INPUT-Autoencoder-Autoencoder-Autoencoder-Autoencoder-Maxout-Rectified Linear-Maxout-Softmax-Argmax</p>\r\n<p>Unsupervised Learning help the&nbsp;Maxout-Rectified Linear-Maxout-Softmax-Argmax network a lot, but contribute to overfit.</p>\r\n<p></p>\r\n<p>My classmate AuroraXie doesn't make complex ensemble only independently use one deep network, and got similar score in both public/private board.</p>\r\n<p>I ensemble 1000&#43; different complex models, and drop from 3rd in public to &nbsp;7th in private</p>\r\n<p></p>\r\n<p>Congratulations to the winners!</p>\r\n<p>And does anyone else, eg 3rd or 4th use deep network? It seems deep network fails in this time.</p>\r\n<p></p>\r\n<p></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "24776",
      "postDate": "05/25/2013 01:19:56",
      "content": "<p>I'll write more later, but briefly, my approach was:</p>\r\n<p>- run sparse filtering several times to generate several sets of features.</p>\r\n<p>- run a feature selection process to find the good features in each set</p>\r\n<p>- combine the good features in one set</p>\r\n<p>- run an svm on that</p>\r\n<p>So I did use unsupervised feature learning, but it wasn't deep learning.</p>\r\n<p>Strictly speaking, my highest-scoring entry (0.7022) was a small voted ensemble of outputs from that process. I had two entries that tested at 0.7018 and 0.7016 (probably 2 and 3 predictions back on the 5000 element private set) that were exactly as described\r\n above.</p>\r\n<p>All of my scored submissions made use of the unlabeled data for feature learning.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "24778",
      "postDate": "05/25/2013 01:38:32",
      "content": "<p>How big is the combined feature set that you ended up using?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "24783",
      "postDate": "05/25/2013 01:50:13",
      "content": "<p>They were disconcertingly large. The models that went into that ensemble were size 400, 300 and 274. I also scored the size 300 and 400 models individually. They scored at 0.6980 and 0.7018. The other set (0.7016) was size 366.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "24789",
      "postDate": "05/25/2013 02:37:13",
      "content": "<p>Congrats to 1st winner ! I almost made it, but I missed it :)</p>\r\n<p>My approach was (<span>I'll write more later) :</span></p>\r\n<p>- 1 hidden layer neural net with&nbsp;rectified linear hidden unit (8000 units) and sigmoid output unit.</p>\r\n<p>- pseudo-label for unlabeled data : just picking up the class that has maximal network output every weights update.</p>\r\n<p><span>- (semi-)supervised learning (CE cost) with labeled data and unlabeled data (with pseudo-label)&nbsp;simultaneously&nbsp;</span></p>\r\n<p><span>- dropout SGD training (without weight regularization)</span></p>\r\n<p>- 2 stage training using polarity&nbsp;splitting (4000 -&gt; 8000 units)&nbsp;</p>\r\n<hr>\r\n<p>pure supervised learning with 1000 labeled data : 0.53 ~ 0.57</p>\r\n<p>&#43; unlabeled data with pseudo-labels : ~ 0.65</p>\r\n<p>&#43; with dropout : ~ 0.6844</p>\r\n<p>&#43; 2 stage training using polarity&nbsp;splitting : ~ 0.6958 (max score)</p>\r\n<hr>\r\n<p>I never used unsupervised feature learning and network ensembles. And my MATLAB code is very simple. only 62 lines.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "24790",
      "postDate": "05/25/2013 02:48:55",
      "content": "<p>Interesting. What is CE cost? For the rectifier with polarity splitting, did you have any biases or just weights?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "24791",
      "postDate": "05/25/2013 03:08:48",
      "content": "<p>Cross Entropy costs with 1-of-K codes of labels and sigmoid outputs. (popular setting but not using softmax)</p>\r\n<p>For using polarity splitting, after training neural net (1st stage), I trained 1 more started with W and -W, same biases (2nd stage). MATLAB code is</p>\r\n<p>W1 = [W1;-W1]; B1 = [B1;B1]; <br>\r\nW2 = 0.01*grandn(nH2,nH1); B2(:)=0;</p>\r\n<p>W1 - weight matrix from visible to hidden, W2 - weight matrix&nbsp;from hidden to output</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "24792",
      "postDate": "05/25/2013 04:15:03",
      "content": "<p>And what was the best result for this data before the competition?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "24795",
      "postDate": "05/25/2013 05:00:02",
      "content": "<p>All methods that have been applied to it before depended heavily on knowledge of what the data was. The best result I know of is just over 98% accuracy.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "24796",
      "postDate": "05/25/2013 05:08:12",
      "content": "<p></p>\r\n<p>using all data to train or just 1000? 98% is awesome..... if so we have a great gap between it</p>\r\n<p></p>\r\n<p>[quote=Ian Goodfellow;24795]</p>\r\n<p>All methods that have been applied to it before depended heavily on knowledge of what the data was. The best result I know of is just over 98% accuracy.</p>\r\n<p>[/quote]</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "24799",
      "postDate": "05/25/2013 10:25:32",
      "content": "<p>I used a 6-layer pretrained contractive autoencoder with sigmoid units (couldn't get rectified linear to work). Finetuned it with a somewhat modifed dropout procedure, all in all a pretty straightforward approach that was easy to implement using pylearn2.\r\n Didn't have a lot of time on this competition, so basically spent all my time finetuning hyperparameters.</p>\r\n<p>Thanks to the organizers for their efforts and their quick and helpful replies!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "24800",
      "postDate": "05/25/2013 12:18:07",
      "content": "<p>I used Matlab Neural Network from DeepLearnToolbox with 2 hidden layers and 450 neurons in each. The model has sigmoid activation and softmax outputs. The best dropout I found was 33,34%. I did 10 fold cross validation on the 1000 instances dataset. With\r\n that simple approach I got 0,614% in the public leaderboard.</p>\r\n<p>Then I used that model to estimate the labels of the extra dataset picking just the max network output class. Retrainned the model with that extra data gives me about 0,632 in the public leaderboard.</p>\r\n<p></p>\r\n<p></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "24801",
      "postDate": "05/25/2013 13:35:51",
      "content": "<p>My method is similar to yours : pseudo-label for unlabeled data. But I trained the network with labeled data and unlabeled data simultaneously and estimated the labels every weights update. And I used sigmoid outputs instead of softmax. There are reasons\r\n for this (I'm not sure that these are reasonable yet)</p>\r\n<p>- using saturation regions of sigmoid unit : such as Contractive Autoencoder, I wanted for my network to be robust against small change of inputs.</p>\r\n<p>- I thought that contractive regularization &#43; reducing reconstruction cost ~=&nbsp;contractive regularization &#43; reducing supervised cost of some labeled data</p>\r\n<p>Anyway, pseudo-label is very simple but&nbsp;relatively good. In pilot test on MNIST dataset, the results is not bad as compared with conventional methods.</p>\r\n<p></p>\r\n<p>[quote=Gilberto Titericz Junior;24800]</p>\r\n<p>I used Matlab Neural Network from DeepLearnToolbox with 2 hidden layers and 450 neurons in each. The model has sigmoid activation and softmax outputs. The best dropout I found was 33,34%. I did 10 fold cross validation on the 1000 instances dataset. With\r\n that simple approach I got 0,614% in the public leaderboard.</p>\r\n<p>Then I used that model to estimate the labels of the extra dataset picking just the max network output class. Retrainned the model with that extra data gives me about 0,632 in the public leaderboard.</p>\r\n<p>[/quote]</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "24805",
      "postDate": "05/25/2013 13:52:24",
      "content": "<p>[quote=binghsu;24796]</p>\r\n<p></p>\r\n<p>using all data to train or just 1000? 98% is awesome..... if so we have a great gap between it</p>\r\n<p></p>\r\n<p>[quote=Ian Goodfellow;24795]</p>\r\n<p>All methods that have been applied to it before depended heavily on knowledge of what the data was. The best result I know of is just over 98% accuracy.</p>\r\n<p>[/quote]</p>\r\n<p>[/quote]</p>\r\n<p></p>\r\n<p>There's also a different amount of data available for the non-black box version of the task. I'm still avoiding saying how much exactly because it would help guess what the task is, and I think we still want it to be a surprise at the workshop.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "24809",
      "postDate": "05/25/2013 16:33:56",
      "content": "<p>I used sparse filtering (python) with softAbs activation to encode the (train&#43;test&#43;extra) data. I did this for 400, 800, and 1200 codes. These codes were the input for a number of one-hidden-layer and two-hidden-layer dropout networks with maxout activation\r\n (pylearn2). I found 10 such models with good (~0.6) validation scores, then averaged the posteriors from them with uniform weighting to get a score of ~0.67.</p>\r\n<p>I'm curious what doubleshot meant by &quot;run a feature selection process to find the good features in each set&quot;. Care to elaborate?</p>\r\n<p>I'm going to guess that the 1875 features are 3 channels (e.g. RGB) for 25x25 images. Somebody commented that this couldn't be true, because one could have easily reverse-engineered and discovered what the images were, but with the random ordering of the\r\n features that sounds pretty difficult to me.</p>\r\n<p>Thanks for the competition - I learned a lot.</p>\r\n<p></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "24811",
      "postDate": "05/25/2013 16:41:49",
      "content": "<p>Here's a write-up of my approach, with easy-to-reproduce-results code:</p>\r\n<p><a href=\"http://fastml.com/more-on-sparse-filtering-and-the-black-box-competition/\">http://fastml.com/more-on-sparse-filtering-and-the-black-box-competition/</a></p>\r\n<p></p>\r\n<p>Basically, one-layer sparse filtering &#43; a linear model for 0.634,&nbsp;</p>\r\n<p>or the same one-layer sparse filtering &#43; mrmr feature selection &#43; a small neural network for 0.645.</p>\r\n<p></p>\r\n<p></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "24813",
      "postDate": "05/25/2013 17:04:30",
      "content": "<p>My model&nbsp; was a blend of multiple 800x100 NNs with dropout &nbsp;and RFs trained on Sparse Filtering features. The most interesting and significant gain (from ~0.6 to ~0.66)&nbsp; I got when I used RF as a blending method using each class probabilities from different\r\n models &nbsp;as inputs.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "24824",
      "postDate": "05/25/2013 22:59:02",
      "content": "<p></p>\r\n<p>that's interesting. I also use 1000&#43; random forest for blending.</p>\r\n<p>but contribute to extremely overfit</p>\r\n<p></p>\r\n<p>your blending is awesome!</p>\r\n<p>[quote=Sergey Yurgenson;24813]</p>\r\n<p>My model&nbsp; was a blend of multiple 800x100 NNs with dropout &nbsp;and RFs trained on Sparse Filtering features. The most interesting and significant gain (from ~0.6 to ~0.66)&nbsp; I got when I used RF as a blending method using each class probabilities from different\r\n models &nbsp;as inputs.</p>\r\n<p>[/quote]</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "24969",
      "postDate": "05/31/2013 02:46:52",
      "content": "<p>This is an update post for anybody who is watching this thread (models), but not the whole forum. On Friday, I posted a short description of my method here. I just posted the long version in a new thread, with a link to the code.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "25110",
      "postDate": "06/03/2013 22:53:58",
      "content": "<p>I'm interested how to do feature extraction with pylearn2. I have a DAE trained with the labeled/unlabeled data and I want extract the features for the labeled training set and use then other tools like RF.</p>\r\n<p>I think TransformerDataset should be usefull but don't know how use it exactly. Any thought?</p>\r\n<p></p>\r\n<p></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "25115",
      "postDate": "06/03/2013 23:43:17",
      "content": "<p>Usually I just write my own scripts to write out the file format that I want. I think someone else wrote a class to dump features during training though. If you write to pylearn-dev@googlegroups.com you'll probably get a response.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "25118",
      "postDate": "06/04/2013 03:11:33",
      "content": "<p>You can modify the train.py, add what you want in the main_loop</p>\r\n<p>the other method is using serial.load to load a pkl file, then get its layers, use fprop function get the result of each layer.</p>\r\n<p></p>\r\n<p>Indeed, I have done both, but it is very dirty. So I think it's better to keep it privately :(</p>\r\n<p></p>\r\n<p>Here is two pictures I post in my Chinese weibo, and I think it may be helpful to you:</p>\r\n<p>First is using high level feature to train SVM.</p>\r\n<p>The other picture is two autoencoders, independently mapping training data to 2-dimension.</p>\r\n<p>The first can achieve 0.65&#43; with random forest, the second can only achieve 0.30 with same config of random forest.</p>\r\n<p></p>\r\n<p>[quote=José A. Guerrero;25110]</p>\r\n<p>I'm interested how to do feature extraction with pylearn2. I have a DAE trained with the labeled/unlabeled data and I want extract the features for the labeled training set and use then other tools like RF.</p>\r\n<p>I think TransformerDataset should be usefull but don't know how use it exactly. Any thought?</p>\r\n<p></p>\r\n<p></p>\r\n<p>[/quote]</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "25123",
      "postDate": "06/04/2013 12:18:38",
      "content": "<p>Thanks, Ian, Binghsu,</p>\r\n<p>I'm playing with&nbsp;<a href=\"https://github.com/lisa-lab/pylearn2/blob/master/pylearn2/scripts/tutorials/deep_trainer/run_deep_trainer.py\">https://github.com/lisa-lab/pylearn2/blob/master/pylearn2/scripts/tutorials/deep_trainer/run_deep_trainer.py</a></p>\r\n<p>to get the feature extraction (I'm finding run_deep_trainer.py file very didactic)</p>\r\n<p></p>\r\n<p>After unsupervised training has finished what I have in trainset[3] is a matrix of n (num of samples) &nbsp;x &nbsp;m (num of features) correspond to the transform of original trainset through the layers.</p>\r\n<p>If I save it I could use this features in other models. Is this correct?</p>\r\n<p></p>\r\n<p></p>\r\n<p></p>\r\n<p></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "25128",
      "postDate": "06/04/2013 13:17:41",
      "content": "<p>Yes. You don't even necessarily have to save it first; you could train the other models in the same script if you wanted.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "25130",
      "postDate": "06/04/2013 14:20:34",
      "content": "<p>[quote=Ian Goodfellow;25128]</p>\r\n<p>Yes. You don't even necessarily have to save it first; you could train the other models in the same script if you wanted.</p>\r\n<p>[/quote]</p>\r\n<p>I refer to save trainset[3] in csv format for use in an external tool like randomforest or general additive model</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 24776,
      "author_name": "davidthaler",
      "author_url": "",
      "post_date": "05/25/2013 01:19:56",
      "content": "<p>I'll write more later, but briefly, my approach was:</p>\r\n<p>- run sparse filtering several times to generate several sets of features.</p>\r\n<p>- run a feature selection process to find the good features in each set</p>\r\n<p>- combine the good features in one set</p>\r\n<p>- run an svm on that</p>\r\n<p>So I did use unsupervised feature learning, but it wasn't deep learning.</p>\r\n<p>Strictly speaking, my highest-scoring entry (0.7022) was a small voted ensemble of outputs from that process. I had two entries that tested at 0.7018 and 0.7016 (probably 2 and 3 predictions back on the 5000 element private set) that were exactly as described\r\n above.</p>\r\n<p>All of my scored submissions made use of the unlabeled data for feature learning.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 24778,
      "author_name": "dumitru0",
      "author_url": "",
      "post_date": "05/25/2013 01:38:32",
      "content": "<p>How big is the combined feature set that you ended up using?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 24783,
      "author_name": "davidthaler",
      "author_url": "",
      "post_date": "05/25/2013 01:50:13",
      "content": "<p>They were disconcertingly large. The models that went into that ensemble were size 400, 300 and 274. I also scored the size 300 and 400 models individually. They scored at 0.6980 and 0.7018. The other set (0.7016) was size 366.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 24789,
      "author_name": "donghyun",
      "author_url": "",
      "post_date": "05/25/2013 02:37:13",
      "content": "<p>Congrats to 1st winner ! I almost made it, but I missed it :)</p>\r\n<p>My approach was (<span>I'll write more later) :</span></p>\r\n<p>- 1 hidden layer neural net with&nbsp;rectified linear hidden unit (8000 units) and sigmoid output unit.</p>\r\n<p>- pseudo-label for unlabeled data : just picking up the class that has maximal network output every weights update.</p>\r\n<p><span>- (semi-)supervised learning (CE cost) with labeled data and unlabeled data (with pseudo-label)&nbsp;simultaneously&nbsp;</span></p>\r\n<p><span>- dropout SGD training (without weight regularization)</span></p>\r\n<p>- 2 stage training using polarity&nbsp;splitting (4000 -&gt; 8000 units)&nbsp;</p>\r\n<hr>\r\n<p>pure supervised learning with 1000 labeled data : 0.53 ~ 0.57</p>\r\n<p>&#43; unlabeled data with pseudo-labels : ~ 0.65</p>\r\n<p>&#43; with dropout : ~ 0.6844</p>\r\n<p>&#43; 2 stage training using polarity&nbsp;splitting : ~ 0.6958 (max score)</p>\r\n<hr>\r\n<p>I never used unsupervised feature learning and network ensembles. And my MATLAB code is very simple. only 62 lines.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 24790,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "05/25/2013 02:48:55",
      "content": "<p>Interesting. What is CE cost? For the rectifier with polarity splitting, did you have any biases or just weights?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 24791,
      "author_name": "donghyun",
      "author_url": "",
      "post_date": "05/25/2013 03:08:48",
      "content": "<p>Cross Entropy costs with 1-of-K codes of labels and sigmoid outputs. (popular setting but not using softmax)</p>\r\n<p>For using polarity splitting, after training neural net (1st stage), I trained 1 more started with W and -W, same biases (2nd stage). MATLAB code is</p>\r\n<p>W1 = [W1;-W1]; B1 = [B1;B1]; <br>\r\nW2 = 0.01*grandn(nH2,nH1); B2(:)=0;</p>\r\n<p>W1 - weight matrix from visible to hidden, W2 - weight matrix&nbsp;from hidden to output</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 24792,
      "author_name": "",
      "author_url": "",
      "post_date": "05/25/2013 04:15:03",
      "content": "<p>And what was the best result for this data before the competition?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 24795,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "05/25/2013 05:00:02",
      "content": "<p>All methods that have been applied to it before depended heavily on knowledge of what the data was. The best result I know of is just over 98% accuracy.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 24796,
      "author_name": "binghsu",
      "author_url": "",
      "post_date": "05/25/2013 05:08:12",
      "content": "<p></p>\r\n<p>using all data to train or just 1000? 98% is awesome..... if so we have a great gap between it</p>\r\n<p></p>\r\n<p>[quote=Ian Goodfellow;24795]</p>\r\n<p>All methods that have been applied to it before depended heavily on knowledge of what the data was. The best result I know of is just over 98% accuracy.</p>\r\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 24799,
      "author_name": "njdevos",
      "author_url": "",
      "post_date": "05/25/2013 10:25:32",
      "content": "<p>I used a 6-layer pretrained contractive autoencoder with sigmoid units (couldn't get rectified linear to work). Finetuned it with a somewhat modifed dropout procedure, all in all a pretty straightforward approach that was easy to implement using pylearn2.\r\n Didn't have a lot of time on this competition, so basically spent all my time finetuning hyperparameters.</p>\r\n<p>Thanks to the organizers for their efforts and their quick and helpful replies!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 24800,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "05/25/2013 12:18:07",
      "content": "<p>I used Matlab Neural Network from DeepLearnToolbox with 2 hidden layers and 450 neurons in each. The model has sigmoid activation and softmax outputs. The best dropout I found was 33,34%. I did 10 fold cross validation on the 1000 instances dataset. With\r\n that simple approach I got 0,614% in the public leaderboard.</p>\r\n<p>Then I used that model to estimate the labels of the extra dataset picking just the max network output class. Retrainned the model with that extra data gives me about 0,632 in the public leaderboard.</p>\r\n<p></p>\r\n<p></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 24801,
      "author_name": "donghyun",
      "author_url": "",
      "post_date": "05/25/2013 13:35:51",
      "content": "<p>My method is similar to yours : pseudo-label for unlabeled data. But I trained the network with labeled data and unlabeled data simultaneously and estimated the labels every weights update. And I used sigmoid outputs instead of softmax. There are reasons\r\n for this (I'm not sure that these are reasonable yet)</p>\r\n<p>- using saturation regions of sigmoid unit : such as Contractive Autoencoder, I wanted for my network to be robust against small change of inputs.</p>\r\n<p>- I thought that contractive regularization &#43; reducing reconstruction cost ~=&nbsp;contractive regularization &#43; reducing supervised cost of some labeled data</p>\r\n<p>Anyway, pseudo-label is very simple but&nbsp;relatively good. In pilot test on MNIST dataset, the results is not bad as compared with conventional methods.</p>\r\n<p></p>\r\n<p>[quote=Gilberto Titericz Junior;24800]</p>\r\n<p>I used Matlab Neural Network from DeepLearnToolbox with 2 hidden layers and 450 neurons in each. The model has sigmoid activation and softmax outputs. The best dropout I found was 33,34%. I did 10 fold cross validation on the 1000 instances dataset. With\r\n that simple approach I got 0,614% in the public leaderboard.</p>\r\n<p>Then I used that model to estimate the labels of the extra dataset picking just the max network output class. Retrainned the model with that extra data gives me about 0,632 in the public leaderboard.</p>\r\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 24805,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "05/25/2013 13:52:24",
      "content": "<p>[quote=binghsu;24796]</p>\r\n<p></p>\r\n<p>using all data to train or just 1000? 98% is awesome..... if so we have a great gap between it</p>\r\n<p></p>\r\n<p>[quote=Ian Goodfellow;24795]</p>\r\n<p>All methods that have been applied to it before depended heavily on knowledge of what the data was. The best result I know of is just over 98% accuracy.</p>\r\n<p>[/quote]</p>\r\n<p>[/quote]</p>\r\n<p></p>\r\n<p>There's also a different amount of data available for the non-black box version of the task. I'm still avoiding saying how much exactly because it would help guess what the task is, and I think we still want it to be a surprise at the workshop.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 24809,
      "author_name": "rkeisler",
      "author_url": "",
      "post_date": "05/25/2013 16:33:56",
      "content": "<p>I used sparse filtering (python) with softAbs activation to encode the (train&#43;test&#43;extra) data. I did this for 400, 800, and 1200 codes. These codes were the input for a number of one-hidden-layer and two-hidden-layer dropout networks with maxout activation\r\n (pylearn2). I found 10 such models with good (~0.6) validation scores, then averaged the posteriors from them with uniform weighting to get a score of ~0.67.</p>\r\n<p>I'm curious what doubleshot meant by &quot;run a feature selection process to find the good features in each set&quot;. Care to elaborate?</p>\r\n<p>I'm going to guess that the 1875 features are 3 channels (e.g. RGB) for 25x25 images. Somebody commented that this couldn't be true, because one could have easily reverse-engineered and discovered what the images were, but with the random ordering of the\r\n features that sounds pretty difficult to me.</p>\r\n<p>Thanks for the competition - I learned a lot.</p>\r\n<p></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 24811,
      "author_name": "zygmunt",
      "author_url": "",
      "post_date": "05/25/2013 16:41:49",
      "content": "<p>Here's a write-up of my approach, with easy-to-reproduce-results code:</p>\r\n<p><a href=\"http://fastml.com/more-on-sparse-filtering-and-the-black-box-competition/\">http://fastml.com/more-on-sparse-filtering-and-the-black-box-competition/</a></p>\r\n<p></p>\r\n<p>Basically, one-layer sparse filtering &#43; a linear model for 0.634,&nbsp;</p>\r\n<p>or the same one-layer sparse filtering &#43; mrmr feature selection &#43; a small neural network for 0.645.</p>\r\n<p></p>\r\n<p></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 24813,
      "author_name": "ccccat",
      "author_url": "",
      "post_date": "05/25/2013 17:04:30",
      "content": "<p>My model&nbsp; was a blend of multiple 800x100 NNs with dropout &nbsp;and RFs trained on Sparse Filtering features. The most interesting and significant gain (from ~0.6 to ~0.66)&nbsp; I got when I used RF as a blending method using each class probabilities from different\r\n models &nbsp;as inputs.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 24824,
      "author_name": "binghsu",
      "author_url": "",
      "post_date": "05/25/2013 22:59:02",
      "content": "<p></p>\r\n<p>that's interesting. I also use 1000&#43; random forest for blending.</p>\r\n<p>but contribute to extremely overfit</p>\r\n<p></p>\r\n<p>your blending is awesome!</p>\r\n<p>[quote=Sergey Yurgenson;24813]</p>\r\n<p>My model&nbsp; was a blend of multiple 800x100 NNs with dropout &nbsp;and RFs trained on Sparse Filtering features. The most interesting and significant gain (from ~0.6 to ~0.66)&nbsp; I got when I used RF as a blending method using each class probabilities from different\r\n models &nbsp;as inputs.</p>\r\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 24969,
      "author_name": "davidthaler",
      "author_url": "",
      "post_date": "05/31/2013 02:46:52",
      "content": "<p>This is an update post for anybody who is watching this thread (models), but not the whole forum. On Friday, I posted a short description of my method here. I just posted the long version in a new thread, with a link to the code.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 25110,
      "author_name": "blindape",
      "author_url": "",
      "post_date": "06/03/2013 22:53:58",
      "content": "<p>I'm interested how to do feature extraction with pylearn2. I have a DAE trained with the labeled/unlabeled data and I want extract the features for the labeled training set and use then other tools like RF.</p>\r\n<p>I think TransformerDataset should be usefull but don't know how use it exactly. Any thought?</p>\r\n<p></p>\r\n<p></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 25115,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "06/03/2013 23:43:17",
      "content": "<p>Usually I just write my own scripts to write out the file format that I want. I think someone else wrote a class to dump features during training though. If you write to pylearn-dev@googlegroups.com you'll probably get a response.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 25118,
      "author_name": "binghsu",
      "author_url": "",
      "post_date": "06/04/2013 03:11:33",
      "content": "<p>You can modify the train.py, add what you want in the main_loop</p>\r\n<p>the other method is using serial.load to load a pkl file, then get its layers, use fprop function get the result of each layer.</p>\r\n<p></p>\r\n<p>Indeed, I have done both, but it is very dirty. So I think it's better to keep it privately :(</p>\r\n<p></p>\r\n<p>Here is two pictures I post in my Chinese weibo, and I think it may be helpful to you:</p>\r\n<p>First is using high level feature to train SVM.</p>\r\n<p>The other picture is two autoencoders, independently mapping training data to 2-dimension.</p>\r\n<p>The first can achieve 0.65&#43; with random forest, the second can only achieve 0.30 with same config of random forest.</p>\r\n<p></p>\r\n<p>[quote=José A. Guerrero;25110]</p>\r\n<p>I'm interested how to do feature extraction with pylearn2. I have a DAE trained with the labeled/unlabeled data and I want extract the features for the labeled training set and use then other tools like RF.</p>\r\n<p>I think TransformerDataset should be usefull but don't know how use it exactly. Any thought?</p>\r\n<p></p>\r\n<p></p>\r\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 25123,
      "author_name": "blindape",
      "author_url": "",
      "post_date": "06/04/2013 12:18:38",
      "content": "<p>Thanks, Ian, Binghsu,</p>\r\n<p>I'm playing with&nbsp;<a href=\"https://github.com/lisa-lab/pylearn2/blob/master/pylearn2/scripts/tutorials/deep_trainer/run_deep_trainer.py\">https://github.com/lisa-lab/pylearn2/blob/master/pylearn2/scripts/tutorials/deep_trainer/run_deep_trainer.py</a></p>\r\n<p>to get the feature extraction (I'm finding run_deep_trainer.py file very didactic)</p>\r\n<p></p>\r\n<p>After unsupervised training has finished what I have in trainset[3] is a matrix of n (num of samples) &nbsp;x &nbsp;m (num of features) correspond to the transform of original trainset through the layers.</p>\r\n<p>If I save it I could use this features in other models. Is this correct?</p>\r\n<p></p>\r\n<p></p>\r\n<p></p>\r\n<p></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 25128,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "06/04/2013 13:17:41",
      "content": "<p>Yes. You don't even necessarily have to save it first; you could train the other models in the same script if you wanted.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 25130,
      "author_name": "blindape",
      "author_url": "",
      "post_date": "06/04/2013 14:20:34",
      "content": "<p>[quote=Ian Goodfellow;25128]</p>\r\n<p>Yes. You don't even necessarily have to save it first; you could train the other models in the same script if you wanted.</p>\r\n<p>[/quote]</p>\r\n<p>I refer to save trainset[3] in csv format for use in an external tool like randomforest or general additive model</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "24773": "",
    "24776": "",
    "24778": "",
    "24783": "",
    "24789": "",
    "24790": "",
    "24791": "",
    "24792": "",
    "24795": "",
    "24796": "",
    "24799": "",
    "24800": "",
    "24801": "",
    "24805": "",
    "24809": "",
    "24811": "",
    "24813": "",
    "24824": "",
    "24969": "",
    "25110": "",
    "25115": "",
    "25118": "",
    "25123": "",
    "25128": "",
    "25130": ""
  },
  "source": "meta"
}