{
  "id": 4491,
  "title": "Dropout, Maxout, and Deep Neural Networks",
  "url": "/competitions/challenges-in-representation-learning-the-black-box-learning-challenge/discussion/4491",
  "author_name": "",
  "post_date": "2013-05-02T21:13:16.063Z",
  "votes": 1,
  "comment_count": 32,
  "views": 36288,
  "content": "<p>These questions aren't about the competition per se, but since Ian is a first author on some of these papers, I was hoping he (or anyone else for that matter) would be kind enough to answer a few lingering questions I have. I'm pretty new to neural nets,\r\n so if any of these questions are too simplistic, I apologize.</p>\r\n<p><span style=\"line-height:1.4em\">The first concerns the use of dropout. In the original dropout paper and in Ian's paper on maxout, it is said that training with dropout produces exactly the geometric mean of all 2^j possible achitectures, where j is the\r\n number of hidden units. I follow the intuition that dropout is doing some sort of model averaging that greatly stabilizes predictions, but I can't follow how the geometric mean pops out of this. The original dropout paper says the following:</span></p>\r\n<p><a href=\"http://arxiv.org/abs/1207.0580\">http://arxiv.org/abs/1207.0580</a></p>\r\n<p>&quot;<span style=\"line-height:1.4em\">In networks with a single hidden layer of N units and a “softmax” output layer for</span><span style=\"line-height:1.4em\">computing the probabilities of the class labels, using the mean network is exactly equivalent&nbsp;</span><span style=\"line-height:1.4em\">to\r\n taking the geometric mean of the probability distributions over labels predicted by all 2^</span><span style=\"line-height:1.4em\">N&nbsp;</span><span style=\"line-height:1.4em\">possible networks. Assuming the dropout networks do not all make identical predictions,\r\n the&nbsp;</span><span style=\"line-height:1.4em\">prediction of the mean network is guaranteed to assign a higher log probability to the correct&nbsp;</span><span style=\"line-height:1.4em\">answer than the mean of the log probabilities assigned by the individual dropout\r\n networks</span><span style=\"line-height:1.4em\">&quot;</span></p>\r\n<p><span style=\"line-height:1.4em\">I can't find any proof of this and the references don't appear to be included in the tech report. Can you point me to a place where this is proven? As a follow-up, has there been anywork, either empirical or theoretical, in\r\n comparing the effectiveness of this geometric mean? It would be interesting to see, even for a small toy problem, how directly taking the geometric mean compares to the more common arthimetic mean. Perhaps a comparison to Neal's Bayesian neural nets with appropriate\r\n priors would also be a fair comparison between the geometric and arithmetic means.</span></p>\r\n<p><span style=\"line-height:1.4em\">The next question is about how to train a maxout network. The maxout function from my reading of your paper, should not be differentiable. I thought this is why people used softmax, because it was an approximation of the maximum\r\n function while still having a derivative? Can you still train networks with maxout activations using backprop or are they trained in some other manner?</span></p>\r\n<p><span style=\"line-height:1.4em\"><a href=\"http://arxiv.org/abs/1302.4389\">http://arxiv.org/abs/1302.4389</a>&nbsp;</span></p>\r\n<p>Thanks in advance and congratulations on the great work you're doing!</p>",
  "messages": [
    {
      "id": "23836",
      "postDate": "05/02/2013 21:13:16",
      "content": "<p>These questions aren't about the competition per se, but since Ian is a first author on some of these papers, I was hoping he (or anyone else for that matter) would be kind enough to answer a few lingering questions I have. I'm pretty new to neural nets,\r\n so if any of these questions are too simplistic, I apologize.</p>\r\n<p><span style=\"line-height:1.4em\">The first concerns the use of dropout. In the original dropout paper and in Ian's paper on maxout, it is said that training with dropout produces exactly the geometric mean of all 2^j possible achitectures, where j is the\r\n number of hidden units. I follow the intuition that dropout is doing some sort of model averaging that greatly stabilizes predictions, but I can't follow how the geometric mean pops out of this. The original dropout paper says the following:</span></p>\r\n<p><a href=\"http://arxiv.org/abs/1207.0580\">http://arxiv.org/abs/1207.0580</a></p>\r\n<p>&quot;<span style=\"line-height:1.4em\">In networks with a single hidden layer of N units and a “softmax” output layer for</span><span style=\"line-height:1.4em\">computing the probabilities of the class labels, using the mean network is exactly equivalent&nbsp;</span><span style=\"line-height:1.4em\">to\r\n taking the geometric mean of the probability distributions over labels predicted by all 2^</span><span style=\"line-height:1.4em\">N&nbsp;</span><span style=\"line-height:1.4em\">possible networks. Assuming the dropout networks do not all make identical predictions,\r\n the&nbsp;</span><span style=\"line-height:1.4em\">prediction of the mean network is guaranteed to assign a higher log probability to the correct&nbsp;</span><span style=\"line-height:1.4em\">answer than the mean of the log probabilities assigned by the individual dropout\r\n networks</span><span style=\"line-height:1.4em\">&quot;</span></p>\r\n<p><span style=\"line-height:1.4em\">I can't find any proof of this and the references don't appear to be included in the tech report. Can you point me to a place where this is proven? As a follow-up, has there been anywork, either empirical or theoretical, in\r\n comparing the effectiveness of this geometric mean? It would be interesting to see, even for a small toy problem, how directly taking the geometric mean compares to the more common arthimetic mean. Perhaps a comparison to Neal's Bayesian neural nets with appropriate\r\n priors would also be a fair comparison between the geometric and arithmetic means.</span></p>\r\n<p><span style=\"line-height:1.4em\">The next question is about how to train a maxout network. The maxout function from my reading of your paper, should not be differentiable. I thought this is why people used softmax, because it was an approximation of the maximum\r\n function while still having a derivative? Can you still train networks with maxout activations using backprop or are they trained in some other manner?</span></p>\r\n<p><span style=\"line-height:1.4em\"><a href=\"http://arxiv.org/abs/1302.4389\">http://arxiv.org/abs/1302.4389</a>&nbsp;</span></p>\r\n<p>Thanks in advance and congratulations on the great work you're doing!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23837",
      "postDate": "05/02/2013 21:33:07",
      "content": "<p>The geometric mean pops out because you're doing an (approximate) arithmetic mean in the log domain (i.e. on the pre-softmax values). Another way of accomplishing this is by taking the geometric mean in the softmax space and then renormalizing.</p>\r\n<p>Maxout activations are non-differentiable on a finite set of points, but so are rectifier units. In either case, when doing SGD (or dropout SGD), it works perfectly well to simply ignore these non-differentiable points, as the unit basically never fires\r\n with its activation at <em>exactly</em> that point, and so one filter or the other always has non-zero gradient.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23840",
      "postDate": "05/02/2013 23:34:51",
      "content": "<p>I actually do not understand this claim from the dropout paper that Andrew Beam has quoted: &quot;<span>Assuming the dropout networks do not all make identical predictions, the&nbsp;</span><span>prediction of the mean network is guaranteed to assign a higher log probability\r\n to the correct&nbsp;</span><span>answer than the mean of the log probabilities assigned by the individual dropout networks</span><span>&quot;</span></p>\r\n<p><span>I remember being confused by that when I first read the paper. The way that I am parsing it, it does not seem true to me. But probably I am parsing it differently from how Geoff intended. I looked at the product of experts paper that is cited right\r\n afterward, and couldn't figure out which part was meant to be relevant.</span></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23841",
      "postDate": "05/02/2013 23:37:40",
      "content": "<p>&quot;<span>&quot;</span><span>In networks with a single hidden layer of N units and a “softmax” output layer for</span><span>computing the probabilities of the class labels, using the mean network is exactly equivalent&nbsp;</span><span>to taking the geometric mean of\r\n the probability distributions over labels predicted by all 2^</span><span>N&nbsp;</span><span>possible networks. &quot;</span></p>\r\n<p><span>This isn't proven anywhere, but it's just pretty easy algebra and probability theory. You just need to use a few exponent / logarithm identities and the fact that a probability distribution sums to 1. If you just start by writing down the definition\r\n of the renormalized geometric mean and push through the algebra you should get the weights / 2 rule.</span></p>\r\n<p>Actually, it's pretty easy to show this not just for a softmax layer, but also for an MLP that has identity activation functions on all the hidden units.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23842",
      "postDate": "05/02/2013 23:38:50",
      "content": "<p>&quot;<span>As a follow-up, has there been anywork, either empirical or theoretical, in comparing the effectiveness of this geometric mean? It would be interesting to see, even for a small toy problem, how directly taking the geometric mean compares to the more\r\n common arthimetic mean. &quot;</span></p>\r\n<p><span>I agree this would be interesting. I don't know of any such work off the top of my head. In the maxout paper we were more concerned with doing empirical work to see how accurately dropout applied to a deep network with non-linearities reproduces the\r\n geometric mean.</span></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23843",
      "postDate": "05/02/2013 23:40:27",
      "content": "<p>I think what they meant is that this is true in AVERAGE, i.e., the error of the mean is guaranteed to smaller than the mean of the errors. This has been proven a long time ago in the early 90's for the case of squared error, and the proof could conceivably\r\n be generalized to the log-linear loss. The gist of the original proof is that the mean of the error equals the error of the mean plus the variance (of the outputs).&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23844",
      "postDate": "05/02/2013 23:43:25",
      "content": "<p>That makes a lot of sense, thanks.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23851",
      "postDate": "05/03/2013 03:00:22",
      "content": "<p>I've tried using dropout for this competition (and the facial expression competition) and my experience so far is that it has made my validation errors worse :(</p>\r\n<p><span style=\"line-height:1.4em\">If others have similar experience or successfully used dropout to improve their model, I'd love to hear about them...</span></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23852",
      "postDate": "05/03/2013 03:28:53",
      "content": "<p>[quote=Ian Goodfellow;23841]</p>\r\n<p>&quot;<span>&quot;</span><span>In networks with a single hidden layer of N units and a “softmax” output layer for</span><span>computing the probabilities of the class labels, using the mean network is exactly equivalent&nbsp;</span><span>to taking the geometric mean of\r\n the probability distributions over labels predicted by all 2^</span><span>N&nbsp;</span><span>possible networks. &quot;</span></p>\r\n<p><span>This isn't proven anywhere, but it's just pretty easy algebra and probability theory. You just need to use a few exponent / logarithm identities and the fact that a probability distribution sums to 1. If you just start by writing down the definition\r\n of the renormalized geometric mean and push through the algebra you should get the weights / 2 rule.</span></p>\r\n<p>Actually, it's pretty easy to show this not just for a softmax layer, but also for an MLP that has identity activation functions on all the hidden units.</p>\r\n<p>[/quote]</p>\r\n<p>Could you set this up for me, because I'm not sure where to start? Do you start with log p(y|x,theta) and work backwards, or do you start from the feed-forward perspective? I'm also not sure if the claim that the predicted class probability is the geometric\r\n mean of the output of all possible networks or if the individual weights are the geometric mean (with the zero values obviously excluded). &nbsp;Either way, you can use something like Jensen's inequality to make a statement about the relative magnitudes of predictions\r\n between an approach using a geometric mean and an arithmetic mean. The geometric mean of the output will be less than or equal to the arthimetic mean (with equality when all values are the same).</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23853",
      "postDate": "05/03/2013 03:32:40",
      "content": "<p>[quote=shiggles;23851]</p>\r\n<p>I've tried using dropout for this competition (and the facial expression competition) and my experience so far is that it has made my validation errors worse :(</p>\r\n<p><span style=\"line-height:1.4em\">If others have similar experience or successfully used dropout to improve their model, I'd love to hear about them...</span></p>\r\n<p>[/quote]</p>\r\n<p>I used dropout in this competition and for another project, and I will say that so far, it has been as advertised. I have been able to train arbitrarily long without seeing an increase in validation error. My submission that scored a 0.57 was a decently\r\n large NN trained with dropout without the use of any unlabeled data.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23854",
      "postDate": "05/03/2013 04:10:45",
      "content": "<p>[quote=Andrew Beam;23852]</p>\r\n<p>&nbsp; Either way, you can use something like Jensen's inequality to make a statement about the relative magnitudes of predictions between an approach using a geometric mean and an arithmetic mean. The geometric mean of the output will be less than or equal to\r\n the arthimetic mean (with equality when all values are the same).</p>\r\n<p>[/quote]</p>\r\n<p>Keep in mind that for the output to be a probability, you need to renormalize it. The geometric mean of all the individual predictions is going to be tiny compared to the arithmetic mean, but then you scale it up so that it sums to 1. When you throw that\r\n in, Jensen's inequality doesn't apply anymore.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23856",
      "postDate": "05/03/2013 04:22:59",
      "content": "<p>[quote=Andrew Beam;23852]</p>\r\n<p>[quote=Ian Goodfellow;23841]</p>\r\n<p>&quot;<span>&quot;</span><span>In networks with a single hidden layer of N units and a “softmax” output layer for</span><span>computing the probabilities of the class labels, using the mean network is exactly equivalent&nbsp;</span><span>to taking the geometric mean of\r\n the probability distributions over labels predicted by all 2^</span><span>N&nbsp;</span><span>possible networks. &quot;</span></p>\r\n<p><span>This isn't proven anywhere, but it's just pretty easy algebra and probability theory. You just need to use a few exponent / logarithm identities and the fact that a probability distribution sums to 1. If you just start by writing down the definition\r\n of the renormalized geometric mean and push through the algebra you should get the weights / 2 rule.</span></p>\r\n<p>Actually, it's pretty easy to show this not just for a softmax layer, but also for an MLP that has identity activation functions on all the hidden units.</p>\r\n<p>[/quote]</p>\r\n<p>Could you set this up for me, because I'm not sure where to start? Do you start with log p(y|x,theta) and work backwards, or do you start from the feed-forward perspective?</p>\r\n<p>[/quote]</p>\r\n<p>Let's say p_e(y|x) is the prediction of the &quot;ensemble&quot; using the geometric mean. p_e is just a name, I'm not using e as an index. Now let's say p_d(y|x) is the prediction of a single submodel. Here I am using &quot;d&quot; as a variable that indexes into different\r\n possible distributions. d should be a binary vector saying which inputs to the softmax classifier to include.</p>\r\n<p>p_d(y|x) = softmax( W * (x .* d))[y] &nbsp; &nbsp; &nbsp; &nbsp;(I'm using matlab notation, where .* is elementwise multiplication, and * is matrix multiplication)</p>\r\n<p>Suppose there are N different units. Then there are 2^N possible assignments to d, and</p>\r\n<p>p_e(y|x) = (product_d &nbsp;p_d(y|x) )^(1/2^N) / sum_y'&nbsp;(product_d &nbsp;p_d(y'|x) )^(1/2^N)</p>\r\n<p>That division by the sum is needed to make sure that the output p_e is still a probability.</p>\r\n<p>But we can ignore it for now, and just saw we'll renormalize at at the end:</p>\r\n<p>p_e(y|x) \\propto (product_d &nbsp;p_d(y|x) )^(1/2^N)</p>\r\n<p>=&nbsp;&nbsp;(product_d &nbsp;softmax( W * (x .* d))[y]&nbsp; )^(1/2^N) &nbsp; by definition of p_d</p>\r\n<p>=&nbsp;&nbsp;(product_d &nbsp;exp( W * (x .* d))[y] / sum_y' exp( W * (x.d))[y'] &nbsp;)^(1/2^N) by definition of softmax</p>\r\n<p>=&nbsp;&nbsp;(product_d &nbsp;exp( W * (x .* d))[y]) ^(1/2^N)&nbsp;/ ( product_d sum_y' exp( W * (x.d))[y'] &nbsp;)^(1/2^N)</p>\r\n<p>\\propto&nbsp;<span style=\"line-height:1.4em\">&nbsp; (product_d &nbsp;exp( W * (x .* d))[y]) ^(1/2^N)&nbsp;</span></p>\r\n<p><span style=\"line-height:1.4em\">=&nbsp;<span>product_d &nbsp;exp( (1/2^N) W * (x .* d))[y])</span></span></p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23857",
      "postDate": "05/03/2013 04:25:59",
      "content": "<p>It looks like I hit a comment length limit. The rest is</p>\r\n<p>product_d &nbsp;exp( (1/2^N) W * (x .* d))[y]</p>\r\n<p>= &nbsp;exp( (1/2^N) &nbsp;sum_d W * (x .* d))[y]</p>\r\n<p>= &nbsp;exp( (1/2) W * x)[y]</p>\r\n<p>So the predicted probability must be proportional to this. To renormalize it, we divide by sum_y' exp( (1/2) W x)[y'].</p>\r\n<p>But that means our predicted distribution is just softmax((1/2) W * x).</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23887",
      "postDate": "05/04/2013 06:33:13",
      "content": "<p>Andrew Beam said, &quot;<span>I used dropout in this competition and for another project, and I will say that so far, it has been as advertised. I have been able to train arbitrarily long without seeing an increase in validation error. My submission that scored\r\n a 0.57 was a decently large NN trained with dropout without the use of any unlabeled data. &quot;</span></p>\r\n<p><span style=\"font-size:14px; line-height:1.4em\">I'm curious - did you implement dropout using pylearn2 or something else?</span></p>\r\n<p>I trained a NN in pylearn2 supposedly using costs.mlp.dropout rather than the default cost, but it didn't seem to make any difference. &nbsp;But I wouldn't be surprised if I was doing something dumb.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23893",
      "postDate": "05/04/2013 13:45:34",
      "content": "<p>[quote=wweight;23887]</p>\r\n<p>Andrew Beam said, &quot;<span>I used dropout in this competition and for another project, and I will say that so far, it has been as advertised. I have been able to train arbitrarily long without seeing an increase in validation error. My submission that scored\r\n a 0.57 was a decently large NN trained with dropout without the use of any unlabeled data. &quot;</span></p>\r\n<p><span style=\"font-size:14px; line-height:1.4em\">I'm curious - did you implement dropout using pylearn2 or something else?</span></p>\r\n<p>I trained a NN in pylearn2 supposedly using costs.mlp.dropout rather than the default cost, but it didn't seem to make any difference. &nbsp;But I wouldn't be surprised if I was doing something dumb.</p>\r\n<p>[/quote]</p>\r\n<p>I've been using this toolbox for Matlab to get up to speed on all of these deep learning techniques:<br>\r\n<br>\r\nhttps://github.com/rasmusbergpalm/DeepLearnToolbox</p>\r\n<p>So far I have nothing but good things to say about it.</p>\r\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23896",
      "postDate": "05/04/2013 16:09:51",
      "content": "<p>[quote=Ian Goodfellow;23857]</p>\r\n<p>It looks like I hit a comment length limit. The rest is</p>\r\n<p>product_d &nbsp;exp( (1/2^N) W * (x .* d))[y]</p>\r\n<p>= &nbsp;exp( (1/2^N) &nbsp;sum_d W * (x .* d))[y]</p>\r\n<p>= &nbsp;exp( (1/2) W * x)[y]</p>\r\n<p>So the predicted probability must be proportional to this. To renormalize it, we divide by sum_y' exp( (1/2) W x)[y'].</p>\r\n<p>But that means our predicted distribution is just softmax((1/2) W * x).</p>\r\n<p>[/quote]</p>\r\n<p>I follow that and thanks for the explanation. Maybe I'm still missing something, but this doesn't seem to be a proof that dropout is taking the geometric mean of all possible models. Given your definitions, you showed that:</p>\r\n<p>p_d(y|x) = softmax(W*(X*.d))y&nbsp;</p>\r\n<p>and</p>\r\n<p>p_e = softmax(1/2*W*X)[y]</p>\r\n<p>but this seems a little backwards to me. You started with a fixed W, and then showed how for a given W, taking the geometric mean of all possible sub-models will produce the same output (up to a constant of 2) as the full model. I don't think this is surprising,\r\n and this doesn't seem to be what dropout is doing. Dropout is a way to train the weights, i.e. a way to obtain W.&nbsp;</p>\r\n<p><span style=\"line-height:1.4em\">For example you can start with the same definitons and define p_e as the arithmetic mean, p_e(y|x) = 1/(2^N)*sum_d(p_d(y|x)) and show the unnormalized relation between the full model's output and the average is, p_e(y|x) =\r\n y*sum_d(exp(W * (X_d)). I haven't simplified past that point yet, because sums of exponentials are obviously harder to work with than products. The point is, this just defines the relationship of the outputs between the full model and a function of sub-models\r\n for a given W. I do not think it means that if I trained all possible models independently and took the geometric mean of their outputs, I would observe similar behavior between this ensemble and the dropout-like ensemble. I think what would be more ineteresting\r\n is to calculate the variance of both models you defined. I would expect the geometric mean to be estimating nearly the same thing as the full model, but I would expect the geometric mean model to have lower variance.</span></p>\r\n<p><span style=\"line-height:1.4em\">However, when you train with dropout, I think what you are actually doing is estimating the weights via some bagging-like procedure. Instead of bagging on features, you are bagging on latent features in the hidden layer and\r\n instead of averaging the output, you endup averaging over gradient steps, and thus over possible weights. This will have a stabilizing effect and prevent overfitting, but from my understanding, is not the same as the geometric mean of the output of all possible\r\n models.&nbsp;</span></p>\r\n<p>Many thanks for the explanations,</p>\r\n<p>Andrew</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23897",
      "postDate": "05/04/2013 16:33:38",
      "content": "Dropout is really two things--a trick for training all submodels with a bagging like criterion, and a trick for averaging all of those model's predictions together. The proof I showed you was for the second trick. The first trick doesn't really have any\r\n proofs associated with it. If you ignore the fact that the submodels share parameters, it's just obviously bagging by construction. I don't know of any proof that it's ok for them to share parameters; it just works well empirically.",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23898",
      "postDate": "05/04/2013 16:42:06",
      "content": "<p>[quote=Ian Goodfellow;23897]</p>\r\n<p>Dropout is really two things--a trick for training all submodels with a bagging like criterion, and a trick for averaging all of those model's predictions together. The proof I showed you was for the second trick. The first trick doesn't really have any\r\n proofs associated with it. If you ignore the fact that the submodels share parameters, it's just obviously bagging by construction. I don't know of any proof that it's ok for them to share parameters; it just works well empirically.</p>\r\n<p>[/quote]</p>\r\n<p>Again, I don't think that is what you showed. You showed that an output produced using the geometric mean of all possible submodels has the same\r\n<strong>expected value</strong>&nbsp;as the full model, up to a constant. In certain scenarios, I would expect the average of many bagged sub-models to be close in expectation to the full model. What I'm saying is when you predict you are actually using the full\r\n model and so while you can expect them to have close to the same output in expectation, the full model is likely to have higher variance than if you had actually done the full geometric average. Dropout's strength appears to be from the bagging like style\r\n of obtaining the weights, and not from any sort of model averaging at the output level.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23899",
      "postDate": "05/04/2013 16:52:04",
      "content": "It's important to use both tricks--I'm not saying anything about which is more important. It doesn't make any sense to use the second trick if you aren't already using the first one. I think you're confused about my algebra. When say &quot;expected value&quot; what\r\n random variable do you mean for the expectation to be over? The only expectation I did was over the choice of which model to run. I'm not claiming anything controversial there--I'm just proving exactly what is in Geoff Hinton's paper.",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23900",
      "postDate": "05/04/2013 17:01:30",
      "content": "<p>Let's say we've trained and have access to our W matrix. We then compute the geometric mean of all possible sub-models using this W and compare it to prediction made by using the full model with this W and no dropout mask. You showed that these two will\r\n be the same. This isn't ensembling in any meaningful way though, this is just a fact of algebra. It is interesting to know that the geometric mean of all possible sub-models is the same as the full model, but that doesn't get us any of the benefits of actually\r\n training the sub-models seperately and then averaging over them.</p>\r\n<p>Maybe this is my fault for reading too much into the paper and maybe I still don't get it, but the statement I posted on the first page sounds like it promises a little more than that.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23901",
      "postDate": "05/04/2013 17:17:30",
      "content": "The second half of the statement you posted is only true in expectation, as Yoshua said. The first half is true, and &quot;just a fact of algebra.&quot;",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23910",
      "postDate": "05/05/2013 01:04:49",
      "content": "<pre>Someone has done work?<br><br>I've included the maxout layer, with randomize_pools given the features haven't a spatial order,<br><br>            !obj:pylearn2.models.maxout.Maxout {<br>                layer_name: 'h3',<br>                irange: .05,<br>                num_units: 200,<br>                num_pieces: 40,<br>                randomize_pools,<br>                max_col_norm: 2.<br>                },</pre>\r\n<pre><br>and I have specified the dropout in the cost:<br><br></pre>\r\n<pre>        cost: !obj:pylearn2.costs.mlp.dropout.Dropout {<br>            input_include_probs: { 'h0' : .8 , 'h1' : .8 },<br>            input_scales: { 'h0': 1. , 'h1': 1. }<br>        },<br><br></pre>\r\n<pre>but I can't get convergence. I tried several include_probs, num_units, num_pieces and learning_rates but the model is erratic.<br>I think this trick is specially important in cases with few cases labeled. like this. The training error go fast to 0. If I understand it correctly this is like a<br>random feature selection (analogue at the mtry parameter in randomForest).</pre>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23921",
      "postDate": "05/05/2013 21:17:42",
      "content": "<p>[quote=Andrew Beam;23893]</p>\r\n<p><span style=\"line-height:1.4em\">I've been using this toolbox for Matlab to get up to speed on all of these deep learning techniques:</span></p>\r\n<p>https://github.com/rasmusbergpalm/DeepLearnToolbox</p>\r\n<p>So far I have nothing but good things to say about it.</p>\r\n<p><span style=\"line-height:1.4em\">[/quote]</span></p>\r\n<p>Well I use the same toolbox but I must be doing something way wrong. I adapted the test_example_DBN by just putting the data (no scaling) and the results were quite bad. I can get it work efficiently in another database than the mnist. Can you share any\r\n info on the momemntum,batchsize,no_layers and nodes? Do all approaches work for you (CNN,SAE etc)?</p>\r\n<p>Any help appreciated !</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23923",
      "postDate": "05/05/2013 22:05:27",
      "content": "<p>I've had good luck with the SAE and NN, both trained with dropout. If you give those a try, I'm sure you'll have more luck.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23935",
      "postDate": "05/06/2013 08:50:06",
      "content": "<p>I use dropout to be a bit better than you ~~~</p>\r\n<p>:-)</p>\r\n<p>[quote=shiggles;23851]</p>\r\n<p>I've tried using dropout for this competition (and the facial expression competition) and my experience so far is that it has made my validation errors worse :(</p>\r\n<p><span style=\"line-height:1.4em\">If others have similar experience or successfully used dropout to improve their model, I'd love to hear about them...</span></p>\r\n<p>[/quote]</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "24191",
      "postDate": "05/11/2013 16:01:52",
      "content": "<p>Hi</p>\r\n<p>I'm hitting a problem in pylearn2.&nbsp; I'm trying to add the Standardize preprocessor to a maxout MLP.</p>\r\n<p>Standardize:</p>\r\n<blockquote>\r\n<p>from pylearn2.datasets.preprocessing import Standardize<br>\r\nfrom pylearn2.utils import serial</p>\r\n<p>from black_box_dataset import BlackBoxDataset</p>\r\n<p>extra = BlackBoxDataset('extra')</p>\r\n<p>std = Standardize()</p>\r\n<p>std.apply(extra, can_fit=True)</p>\r\n<p>serial.save('std.pkl', std)</p>\r\n</blockquote>\r\n<p></p>\r\n<p>I've verified that the _mean and _std fields are correct, but when I run the following .yaml I get numeric overflows and crash out with a NaN.</p>\r\n<p>YAML:</p>\r\n<blockquote>!obj:pylearn2.train.Train {<br>\r\n&nbsp;&nbsp;&nbsp; # Here we specify the dataset to train on. We train on only the first 900 of the examples, so<br>\r\n&nbsp;&nbsp;&nbsp; # that the rest may be used as a validation set.<br>\r\n&nbsp;&nbsp;&nbsp; # The &quot;&amp;train&quot; syntax lets us refer back to this object as &quot;*train&quot; elsewhere in the yaml file<br>\r\n&nbsp;&nbsp;&nbsp; dataset: &amp;train !obj:pylearn2.scripts.icml_2013_wrepl.black_box.black_box_dataset.BlackBoxDataset {<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; which_set: 'train',<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; start: 0,<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; stop: 900,<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; preprocessor: &amp;preprocessor !pkl: &quot;std.pkl&quot;<br>\r\n&nbsp;&nbsp;&nbsp; },<br>\r\n&nbsp;&nbsp;&nbsp; # Here we specify the model to train as being an MLP<br>\r\n&nbsp;&nbsp;&nbsp; model: !obj:pylearn2.models.mlp.MLP {<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; batch_size: 100,<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; layers : [<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; # We use two hidden layers with maxout activations<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; !obj:pylearn2.models.maxout.Maxout {<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; layer_name: 'h0',<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; num_units: 1875,<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; num_pieces: 2,<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; irange: .05,<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; # Rather than using weight decay, we constrain the norms of the weight vectors<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; # max_col_norm: 2.<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; },<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; !obj:pylearn2.models.maxout.Maxout {<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; layer_name: 'h1',<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; num_units: 469,<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; num_pieces: 2,<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; irange: .05,<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; # Rather than using weight decay, we constrain the norms of the weight vectors<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; # max_col_norm: 2.<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; },<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; !obj:pylearn2.models.mlp.Softmax {<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; layer_name: 'y',<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; init_bias_target_marginals: *train,<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; # Initialize the weights to all 0s<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; irange: .0,<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; n_classes: 9<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; }<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; ],<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; nvis: 1875,<br>\r\n&nbsp;&nbsp;&nbsp; },<br>\r\n&nbsp;&nbsp;&nbsp; # We train using SGD and momentum<br>\r\n&nbsp;&nbsp;&nbsp; algorithm: !obj:pylearn2.training_algorithms.sgd.SGD {<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; learning_rate: .1,<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; init_momentum: .5,<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; # We monitor how well we're doing during training on a validation set<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; monitoring_dataset:<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; {<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 'train' : *train,<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 'valid' : !obj:pylearn2.scripts.icml_2013_wrepl.black_box.black_box_dataset.BlackBoxDataset {<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; preprocessor: *preprocessor,<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; which_set: 'train',<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; start: 900,<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; stop: 1000,<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; }<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; },<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; cost: !obj:pylearn2.costs.mlp.dropout.Dropout {<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; # input_include_probs: { 'h0' : .8 },<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; # input_scales: { 'h0': 1. }<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; },<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; # We stop when validation set classification error hasn't decreased for 100 epochs<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; termination_criterion: !obj:pylearn2.termination_criteria.MonitorBased {<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; channel_name: &quot;valid_y_misclass&quot;,<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; prop_decrease: 0.,<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; N: 100<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; },<br>\r\n&nbsp;&nbsp;&nbsp; },<br>\r\n&nbsp;&nbsp;&nbsp; # We save the model whenever we improve on the validation set classification error<br>\r\n&nbsp;&nbsp;&nbsp; extensions: [<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; !obj:pylearn2.train_extensions.best_params.MonitorBasedSaveBest {<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; channel_name: 'valid_y_misclass',<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; save_path: &quot;${PYLEARN2_TRAIN_FILE_FULL_STEM}_best.pkl&quot;<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; },<br>\r\n&nbsp;&nbsp;&nbsp; ],<br>\r\n&nbsp;&nbsp;&nbsp; save_path: &quot;mlp.pkl&quot;,<br>\r\n&nbsp;&nbsp;&nbsp; save_freq: 5<br>\r\n}<br>\r\n<br>\r\n</blockquote>\r\n<p>Any help for a python newbie much appreciated.</p>\r\n<p>Thanks</p>\r\n<p>John</p>\r\n<p></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "24193",
      "postDate": "05/11/2013 16:18:11",
      "content": "<p>Probably a log or an exp somewhere is getting too extreme of a value. This might not be due directly to the preprocessing. The issue could be that the preprocessing increased the range of values in the dataset and thus made the gradient steps bigger. You\r\n can fix that by reducing the learning rate, momentum, and irange. This is mostly a trial and error thing.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "24232",
      "postDate": "05/12/2013 11:55:04",
      "content": "<p>Thanks Ian, will give it a try.</p>\r\n<p></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "24240",
      "postDate": "05/12/2013 16:09:19",
      "content": "<p>Hi Ian</p>\r\n<p></p>\r\n<p>Yep, you were right.&nbsp; Setting {learning_rate: .01,&nbsp; init_momentum: 0.0,} fixed the NaNs.&nbsp; If anyone is interested that 2-layer NN scores 0.55240 which is a victory for the maxout neuron!</p>\r\n<p></p>\r\n<p>~John</p>\r\n<p></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "27641",
      "postDate": "07/25/2013 08:47:22",
      "content": "<p>Sorry but I have a question. In the Maxout class what are num_pieces? The number of inputs to each unit?&nbsp;</p>\n<p>If I want to create a neural network that has 2 hidden layers should I use two Maxout hidden layers and one Softmax representing the output?</p>\n<p>Do I have to create a layer representing the input layer or it's not necessary?&nbsp;</p>\n<p>Thanks!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "90433",
      "postDate": "08/26/2015 14:44:15",
      "content": "<p>@John Clarke \nI am a little bit confused about how to back propagate errors in Maxout layer,\nDo I necessary need to use there dimension matrix to represent the weights at that linear.</p>\n\n<p>Thanks,\nNahed </p>",
      "rawMarkdown": "John Clarke \r\nI am a little bit confused about how to back propagate errors in Maxout layer,\r\nDo I necessary need to use there dimension matrix to represent the weights at that linear.\r\n\r\nThanks,\r\nNahed",
      "votes": null
    },
    {
      "id": "90447",
      "postDate": "08/26/2015 16:12:58",
      "content": "<p>Wow Nahed!  It's been an age since I looked at this.  Can't help you I'm afraid.</p>\n\n<p>Maybe try caffe now?</p>",
      "rawMarkdown": "Wow Nahed!  It's been an age since I looked at this.  Can't help you I'm afraid.\r\n\r\nMaybe try caffe now?",
      "votes": null
    },
    {
      "id": "368592",
      "postDate": "08/10/2018 08:15:47",
      "content": "<p>While validation and testing you have to switch of your dropout errors and use full networks...\nAs mentioned in Video lectures of Stanford CS231n by Justin Johnson. </p>",
      "rawMarkdown": "While validation and testing you have to switch of your dropout errors and use full networks...\nAs mentioned in Video lectures of Stanford CS231n by Justin Johnson.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 23837,
      "author_name": "dwf95285",
      "author_url": "",
      "post_date": "05/02/2013 21:33:07",
      "content": "<p>The geometric mean pops out because you're doing an (approximate) arithmetic mean in the log domain (i.e. on the pre-softmax values). Another way of accomplishing this is by taking the geometric mean in the softmax space and then renormalizing.</p>\r\n<p>Maxout activations are non-differentiable on a finite set of points, but so are rectifier units. In either case, when doing SGD (or dropout SGD), it works perfectly well to simply ignore these non-differentiable points, as the unit basically never fires\r\n with its activation at <em>exactly</em> that point, and so one filter or the other always has non-zero gradient.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23840,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "05/02/2013 23:34:51",
      "content": "<p>I actually do not understand this claim from the dropout paper that Andrew Beam has quoted: &quot;<span>Assuming the dropout networks do not all make identical predictions, the&nbsp;</span><span>prediction of the mean network is guaranteed to assign a higher log probability\r\n to the correct&nbsp;</span><span>answer than the mean of the log probabilities assigned by the individual dropout networks</span><span>&quot;</span></p>\r\n<p><span>I remember being confused by that when I first read the paper. The way that I am parsing it, it does not seem true to me. But probably I am parsing it differently from how Geoff intended. I looked at the product of experts paper that is cited right\r\n afterward, and couldn't figure out which part was meant to be relevant.</span></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23841,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "05/02/2013 23:37:40",
      "content": "<p>&quot;<span>&quot;</span><span>In networks with a single hidden layer of N units and a “softmax” output layer for</span><span>computing the probabilities of the class labels, using the mean network is exactly equivalent&nbsp;</span><span>to taking the geometric mean of\r\n the probability distributions over labels predicted by all 2^</span><span>N&nbsp;</span><span>possible networks. &quot;</span></p>\r\n<p><span>This isn't proven anywhere, but it's just pretty easy algebra and probability theory. You just need to use a few exponent / logarithm identities and the fact that a probability distribution sums to 1. If you just start by writing down the definition\r\n of the renormalized geometric mean and push through the algebra you should get the weights / 2 rule.</span></p>\r\n<p>Actually, it's pretty easy to show this not just for a softmax layer, but also for an MLP that has identity activation functions on all the hidden units.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23842,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "05/02/2013 23:38:50",
      "content": "<p>&quot;<span>As a follow-up, has there been anywork, either empirical or theoretical, in comparing the effectiveness of this geometric mean? It would be interesting to see, even for a small toy problem, how directly taking the geometric mean compares to the more\r\n common arthimetic mean. &quot;</span></p>\r\n<p><span>I agree this would be interesting. I don't know of any such work off the top of my head. In the maxout paper we were more concerned with doing empirical work to see how accurately dropout applied to a deep network with non-linearities reproduces the\r\n geometric mean.</span></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23843,
      "author_name": "yoshuabengio",
      "author_url": "",
      "post_date": "05/02/2013 23:40:27",
      "content": "<p>I think what they meant is that this is true in AVERAGE, i.e., the error of the mean is guaranteed to smaller than the mean of the errors. This has been proven a long time ago in the early 90's for the case of squared error, and the proof could conceivably\r\n be generalized to the log-linear loss. The gist of the original proof is that the mean of the error equals the error of the mean plus the variance (of the outputs).&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23844,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "05/02/2013 23:43:25",
      "content": "<p>That makes a lot of sense, thanks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23851,
      "author_name": "shiggles",
      "author_url": "",
      "post_date": "05/03/2013 03:00:22",
      "content": "<p>I've tried using dropout for this competition (and the facial expression competition) and my experience so far is that it has made my validation errors worse :(</p>\r\n<p><span style=\"line-height:1.4em\">If others have similar experience or successfully used dropout to improve their model, I'd love to hear about them...</span></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23852,
      "author_name": "albeam",
      "author_url": "",
      "post_date": "05/03/2013 03:28:53",
      "content": "<p>[quote=Ian Goodfellow;23841]</p>\r\n<p>&quot;<span>&quot;</span><span>In networks with a single hidden layer of N units and a “softmax” output layer for</span><span>computing the probabilities of the class labels, using the mean network is exactly equivalent&nbsp;</span><span>to taking the geometric mean of\r\n the probability distributions over labels predicted by all 2^</span><span>N&nbsp;</span><span>possible networks. &quot;</span></p>\r\n<p><span>This isn't proven anywhere, but it's just pretty easy algebra and probability theory. You just need to use a few exponent / logarithm identities and the fact that a probability distribution sums to 1. If you just start by writing down the definition\r\n of the renormalized geometric mean and push through the algebra you should get the weights / 2 rule.</span></p>\r\n<p>Actually, it's pretty easy to show this not just for a softmax layer, but also for an MLP that has identity activation functions on all the hidden units.</p>\r\n<p>[/quote]</p>\r\n<p>Could you set this up for me, because I'm not sure where to start? Do you start with log p(y|x,theta) and work backwards, or do you start from the feed-forward perspective? I'm also not sure if the claim that the predicted class probability is the geometric\r\n mean of the output of all possible networks or if the individual weights are the geometric mean (with the zero values obviously excluded). &nbsp;Either way, you can use something like Jensen's inequality to make a statement about the relative magnitudes of predictions\r\n between an approach using a geometric mean and an arithmetic mean. The geometric mean of the output will be less than or equal to the arthimetic mean (with equality when all values are the same).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23853,
      "author_name": "albeam",
      "author_url": "",
      "post_date": "05/03/2013 03:32:40",
      "content": "<p>[quote=shiggles;23851]</p>\r\n<p>I've tried using dropout for this competition (and the facial expression competition) and my experience so far is that it has made my validation errors worse :(</p>\r\n<p><span style=\"line-height:1.4em\">If others have similar experience or successfully used dropout to improve their model, I'd love to hear about them...</span></p>\r\n<p>[/quote]</p>\r\n<p>I used dropout in this competition and for another project, and I will say that so far, it has been as advertised. I have been able to train arbitrarily long without seeing an increase in validation error. My submission that scored a 0.57 was a decently\r\n large NN trained with dropout without the use of any unlabeled data.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23854,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "05/03/2013 04:10:45",
      "content": "<p>[quote=Andrew Beam;23852]</p>\r\n<p>&nbsp; Either way, you can use something like Jensen's inequality to make a statement about the relative magnitudes of predictions between an approach using a geometric mean and an arithmetic mean. The geometric mean of the output will be less than or equal to\r\n the arthimetic mean (with equality when all values are the same).</p>\r\n<p>[/quote]</p>\r\n<p>Keep in mind that for the output to be a probability, you need to renormalize it. The geometric mean of all the individual predictions is going to be tiny compared to the arithmetic mean, but then you scale it up so that it sums to 1. When you throw that\r\n in, Jensen's inequality doesn't apply anymore.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23856,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "05/03/2013 04:22:59",
      "content": "<p>[quote=Andrew Beam;23852]</p>\r\n<p>[quote=Ian Goodfellow;23841]</p>\r\n<p>&quot;<span>&quot;</span><span>In networks with a single hidden layer of N units and a “softmax” output layer for</span><span>computing the probabilities of the class labels, using the mean network is exactly equivalent&nbsp;</span><span>to taking the geometric mean of\r\n the probability distributions over labels predicted by all 2^</span><span>N&nbsp;</span><span>possible networks. &quot;</span></p>\r\n<p><span>This isn't proven anywhere, but it's just pretty easy algebra and probability theory. You just need to use a few exponent / logarithm identities and the fact that a probability distribution sums to 1. If you just start by writing down the definition\r\n of the renormalized geometric mean and push through the algebra you should get the weights / 2 rule.</span></p>\r\n<p>Actually, it's pretty easy to show this not just for a softmax layer, but also for an MLP that has identity activation functions on all the hidden units.</p>\r\n<p>[/quote]</p>\r\n<p>Could you set this up for me, because I'm not sure where to start? Do you start with log p(y|x,theta) and work backwards, or do you start from the feed-forward perspective?</p>\r\n<p>[/quote]</p>\r\n<p>Let's say p_e(y|x) is the prediction of the &quot;ensemble&quot; using the geometric mean. p_e is just a name, I'm not using e as an index. Now let's say p_d(y|x) is the prediction of a single submodel. Here I am using &quot;d&quot; as a variable that indexes into different\r\n possible distributions. d should be a binary vector saying which inputs to the softmax classifier to include.</p>\r\n<p>p_d(y|x) = softmax( W * (x .* d))[y] &nbsp; &nbsp; &nbsp; &nbsp;(I'm using matlab notation, where .* is elementwise multiplication, and * is matrix multiplication)</p>\r\n<p>Suppose there are N different units. Then there are 2^N possible assignments to d, and</p>\r\n<p>p_e(y|x) = (product_d &nbsp;p_d(y|x) )^(1/2^N) / sum_y'&nbsp;(product_d &nbsp;p_d(y'|x) )^(1/2^N)</p>\r\n<p>That division by the sum is needed to make sure that the output p_e is still a probability.</p>\r\n<p>But we can ignore it for now, and just saw we'll renormalize at at the end:</p>\r\n<p>p_e(y|x) \\propto (product_d &nbsp;p_d(y|x) )^(1/2^N)</p>\r\n<p>=&nbsp;&nbsp;(product_d &nbsp;softmax( W * (x .* d))[y]&nbsp; )^(1/2^N) &nbsp; by definition of p_d</p>\r\n<p>=&nbsp;&nbsp;(product_d &nbsp;exp( W * (x .* d))[y] / sum_y' exp( W * (x.d))[y'] &nbsp;)^(1/2^N) by definition of softmax</p>\r\n<p>=&nbsp;&nbsp;(product_d &nbsp;exp( W * (x .* d))[y]) ^(1/2^N)&nbsp;/ ( product_d sum_y' exp( W * (x.d))[y'] &nbsp;)^(1/2^N)</p>\r\n<p>\\propto&nbsp;<span style=\"line-height:1.4em\">&nbsp; (product_d &nbsp;exp( W * (x .* d))[y]) ^(1/2^N)&nbsp;</span></p>\r\n<p><span style=\"line-height:1.4em\">=&nbsp;<span>product_d &nbsp;exp( (1/2^N) W * (x .* d))[y])</span></span></p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23857,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "05/03/2013 04:25:59",
      "content": "<p>It looks like I hit a comment length limit. The rest is</p>\r\n<p>product_d &nbsp;exp( (1/2^N) W * (x .* d))[y]</p>\r\n<p>= &nbsp;exp( (1/2^N) &nbsp;sum_d W * (x .* d))[y]</p>\r\n<p>= &nbsp;exp( (1/2) W * x)[y]</p>\r\n<p>So the predicted probability must be proportional to this. To renormalize it, we divide by sum_y' exp( (1/2) W x)[y'].</p>\r\n<p>But that means our predicted distribution is just softmax((1/2) W * x).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23887,
      "author_name": "rkeisler",
      "author_url": "",
      "post_date": "05/04/2013 06:33:13",
      "content": "<p>Andrew Beam said, &quot;<span>I used dropout in this competition and for another project, and I will say that so far, it has been as advertised. I have been able to train arbitrarily long without seeing an increase in validation error. My submission that scored\r\n a 0.57 was a decently large NN trained with dropout without the use of any unlabeled data. &quot;</span></p>\r\n<p><span style=\"font-size:14px; line-height:1.4em\">I'm curious - did you implement dropout using pylearn2 or something else?</span></p>\r\n<p>I trained a NN in pylearn2 supposedly using costs.mlp.dropout rather than the default cost, but it didn't seem to make any difference. &nbsp;But I wouldn't be surprised if I was doing something dumb.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23893,
      "author_name": "albeam",
      "author_url": "",
      "post_date": "05/04/2013 13:45:34",
      "content": "<p>[quote=wweight;23887]</p>\r\n<p>Andrew Beam said, &quot;<span>I used dropout in this competition and for another project, and I will say that so far, it has been as advertised. I have been able to train arbitrarily long without seeing an increase in validation error. My submission that scored\r\n a 0.57 was a decently large NN trained with dropout without the use of any unlabeled data. &quot;</span></p>\r\n<p><span style=\"font-size:14px; line-height:1.4em\">I'm curious - did you implement dropout using pylearn2 or something else?</span></p>\r\n<p>I trained a NN in pylearn2 supposedly using costs.mlp.dropout rather than the default cost, but it didn't seem to make any difference. &nbsp;But I wouldn't be surprised if I was doing something dumb.</p>\r\n<p>[/quote]</p>\r\n<p>I've been using this toolbox for Matlab to get up to speed on all of these deep learning techniques:<br>\r\n<br>\r\nhttps://github.com/rasmusbergpalm/DeepLearnToolbox</p>\r\n<p>So far I have nothing but good things to say about it.</p>\r\n<p>&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23896,
      "author_name": "albeam",
      "author_url": "",
      "post_date": "05/04/2013 16:09:51",
      "content": "<p>[quote=Ian Goodfellow;23857]</p>\r\n<p>It looks like I hit a comment length limit. The rest is</p>\r\n<p>product_d &nbsp;exp( (1/2^N) W * (x .* d))[y]</p>\r\n<p>= &nbsp;exp( (1/2^N) &nbsp;sum_d W * (x .* d))[y]</p>\r\n<p>= &nbsp;exp( (1/2) W * x)[y]</p>\r\n<p>So the predicted probability must be proportional to this. To renormalize it, we divide by sum_y' exp( (1/2) W x)[y'].</p>\r\n<p>But that means our predicted distribution is just softmax((1/2) W * x).</p>\r\n<p>[/quote]</p>\r\n<p>I follow that and thanks for the explanation. Maybe I'm still missing something, but this doesn't seem to be a proof that dropout is taking the geometric mean of all possible models. Given your definitions, you showed that:</p>\r\n<p>p_d(y|x) = softmax(W*(X*.d))y&nbsp;</p>\r\n<p>and</p>\r\n<p>p_e = softmax(1/2*W*X)[y]</p>\r\n<p>but this seems a little backwards to me. You started with a fixed W, and then showed how for a given W, taking the geometric mean of all possible sub-models will produce the same output (up to a constant of 2) as the full model. I don't think this is surprising,\r\n and this doesn't seem to be what dropout is doing. Dropout is a way to train the weights, i.e. a way to obtain W.&nbsp;</p>\r\n<p><span style=\"line-height:1.4em\">For example you can start with the same definitons and define p_e as the arithmetic mean, p_e(y|x) = 1/(2^N)*sum_d(p_d(y|x)) and show the unnormalized relation between the full model's output and the average is, p_e(y|x) =\r\n y*sum_d(exp(W * (X_d)). I haven't simplified past that point yet, because sums of exponentials are obviously harder to work with than products. The point is, this just defines the relationship of the outputs between the full model and a function of sub-models\r\n for a given W. I do not think it means that if I trained all possible models independently and took the geometric mean of their outputs, I would observe similar behavior between this ensemble and the dropout-like ensemble. I think what would be more ineteresting\r\n is to calculate the variance of both models you defined. I would expect the geometric mean to be estimating nearly the same thing as the full model, but I would expect the geometric mean model to have lower variance.</span></p>\r\n<p><span style=\"line-height:1.4em\">However, when you train with dropout, I think what you are actually doing is estimating the weights via some bagging-like procedure. Instead of bagging on features, you are bagging on latent features in the hidden layer and\r\n instead of averaging the output, you endup averaging over gradient steps, and thus over possible weights. This will have a stabilizing effect and prevent overfitting, but from my understanding, is not the same as the geometric mean of the output of all possible\r\n models.&nbsp;</span></p>\r\n<p>Many thanks for the explanations,</p>\r\n<p>Andrew</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23897,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "05/04/2013 16:33:38",
      "content": "Dropout is really two things--a trick for training all submodels with a bagging like criterion, and a trick for averaging all of those model's predictions together. The proof I showed you was for the second trick. The first trick doesn't really have any\r\n proofs associated with it. If you ignore the fact that the submodels share parameters, it's just obviously bagging by construction. I don't know of any proof that it's ok for them to share parameters; it just works well empirically.",
      "votes": null,
      "replies": []
    },
    {
      "id": 23898,
      "author_name": "albeam",
      "author_url": "",
      "post_date": "05/04/2013 16:42:06",
      "content": "<p>[quote=Ian Goodfellow;23897]</p>\r\n<p>Dropout is really two things--a trick for training all submodels with a bagging like criterion, and a trick for averaging all of those model's predictions together. The proof I showed you was for the second trick. The first trick doesn't really have any\r\n proofs associated with it. If you ignore the fact that the submodels share parameters, it's just obviously bagging by construction. I don't know of any proof that it's ok for them to share parameters; it just works well empirically.</p>\r\n<p>[/quote]</p>\r\n<p>Again, I don't think that is what you showed. You showed that an output produced using the geometric mean of all possible submodels has the same\r\n<strong>expected value</strong>&nbsp;as the full model, up to a constant. In certain scenarios, I would expect the average of many bagged sub-models to be close in expectation to the full model. What I'm saying is when you predict you are actually using the full\r\n model and so while you can expect them to have close to the same output in expectation, the full model is likely to have higher variance than if you had actually done the full geometric average. Dropout's strength appears to be from the bagging like style\r\n of obtaining the weights, and not from any sort of model averaging at the output level.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23899,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "05/04/2013 16:52:04",
      "content": "It's important to use both tricks--I'm not saying anything about which is more important. It doesn't make any sense to use the second trick if you aren't already using the first one. I think you're confused about my algebra. When say &quot;expected value&quot; what\r\n random variable do you mean for the expectation to be over? The only expectation I did was over the choice of which model to run. I'm not claiming anything controversial there--I'm just proving exactly what is in Geoff Hinton's paper.",
      "votes": null,
      "replies": []
    },
    {
      "id": 23900,
      "author_name": "albeam",
      "author_url": "",
      "post_date": "05/04/2013 17:01:30",
      "content": "<p>Let's say we've trained and have access to our W matrix. We then compute the geometric mean of all possible sub-models using this W and compare it to prediction made by using the full model with this W and no dropout mask. You showed that these two will\r\n be the same. This isn't ensembling in any meaningful way though, this is just a fact of algebra. It is interesting to know that the geometric mean of all possible sub-models is the same as the full model, but that doesn't get us any of the benefits of actually\r\n training the sub-models seperately and then averaging over them.</p>\r\n<p>Maybe this is my fault for reading too much into the paper and maybe I still don't get it, but the statement I posted on the first page sounds like it promises a little more than that.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23901,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "05/04/2013 17:17:30",
      "content": "The second half of the statement you posted is only true in expectation, as Yoshua said. The first half is true, and &quot;just a fact of algebra.&quot;",
      "votes": null,
      "replies": []
    },
    {
      "id": 23910,
      "author_name": "blindape",
      "author_url": "",
      "post_date": "05/05/2013 01:04:49",
      "content": "<pre>Someone has done work?<br><br>I've included the maxout layer, with randomize_pools given the features haven't a spatial order,<br><br>            !obj:pylearn2.models.maxout.Maxout {<br>                layer_name: 'h3',<br>                irange: .05,<br>                num_units: 200,<br>                num_pieces: 40,<br>                randomize_pools,<br>                max_col_norm: 2.<br>                },</pre>\r\n<pre><br>and I have specified the dropout in the cost:<br><br></pre>\r\n<pre>        cost: !obj:pylearn2.costs.mlp.dropout.Dropout {<br>            input_include_probs: { 'h0' : .8 , 'h1' : .8 },<br>            input_scales: { 'h0': 1. , 'h1': 1. }<br>        },<br><br></pre>\r\n<pre>but I can't get convergence. I tried several include_probs, num_units, num_pieces and learning_rates but the model is erratic.<br>I think this trick is specially important in cases with few cases labeled. like this. The training error go fast to 0. If I understand it correctly this is like a<br>random feature selection (analogue at the mtry parameter in randomForest).</pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23921,
      "author_name": "potamitis",
      "author_url": "",
      "post_date": "05/05/2013 21:17:42",
      "content": "<p>[quote=Andrew Beam;23893]</p>\r\n<p><span style=\"line-height:1.4em\">I've been using this toolbox for Matlab to get up to speed on all of these deep learning techniques:</span></p>\r\n<p>https://github.com/rasmusbergpalm/DeepLearnToolbox</p>\r\n<p>So far I have nothing but good things to say about it.</p>\r\n<p><span style=\"line-height:1.4em\">[/quote]</span></p>\r\n<p>Well I use the same toolbox but I must be doing something way wrong. I adapted the test_example_DBN by just putting the data (no scaling) and the results were quite bad. I can get it work efficiently in another database than the mnist. Can you share any\r\n info on the momemntum,batchsize,no_layers and nodes? Do all approaches work for you (CNN,SAE etc)?</p>\r\n<p>Any help appreciated !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23923,
      "author_name": "albeam",
      "author_url": "",
      "post_date": "05/05/2013 22:05:27",
      "content": "<p>I've had good luck with the SAE and NN, both trained with dropout. If you give those a try, I'm sure you'll have more luck.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23935,
      "author_name": "binghsu",
      "author_url": "",
      "post_date": "05/06/2013 08:50:06",
      "content": "<p>I use dropout to be a bit better than you ~~~</p>\r\n<p>:-)</p>\r\n<p>[quote=shiggles;23851]</p>\r\n<p>I've tried using dropout for this competition (and the facial expression competition) and my experience so far is that it has made my validation errors worse :(</p>\r\n<p><span style=\"line-height:1.4em\">If others have similar experience or successfully used dropout to improve their model, I'd love to hear about them...</span></p>\r\n<p>[/quote]</p>",
      "votes": null,
      "replies": [
        {
          "id": 368592,
          "author_name": "infernop",
          "author_url": "",
          "post_date": "08/10/2018 08:15:47",
          "content": "<p>While validation and testing you have to switch of your dropout errors and use full networks...\nAs mentioned in Video lectures of Stanford CS231n by Justin Johnson. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 24191,
      "author_name": "johnclarke",
      "author_url": "",
      "post_date": "05/11/2013 16:01:52",
      "content": "<p>Hi</p>\r\n<p>I'm hitting a problem in pylearn2.&nbsp; I'm trying to add the Standardize preprocessor to a maxout MLP.</p>\r\n<p>Standardize:</p>\r\n<blockquote>\r\n<p>from pylearn2.datasets.preprocessing import Standardize<br>\r\nfrom pylearn2.utils import serial</p>\r\n<p>from black_box_dataset import BlackBoxDataset</p>\r\n<p>extra = BlackBoxDataset('extra')</p>\r\n<p>std = Standardize()</p>\r\n<p>std.apply(extra, can_fit=True)</p>\r\n<p>serial.save('std.pkl', std)</p>\r\n</blockquote>\r\n<p></p>\r\n<p>I've verified that the _mean and _std fields are correct, but when I run the following .yaml I get numeric overflows and crash out with a NaN.</p>\r\n<p>YAML:</p>\r\n<blockquote>!obj:pylearn2.train.Train {<br>\r\n&nbsp;&nbsp;&nbsp; # Here we specify the dataset to train on. We train on only the first 900 of the examples, so<br>\r\n&nbsp;&nbsp;&nbsp; # that the rest may be used as a validation set.<br>\r\n&nbsp;&nbsp;&nbsp; # The &quot;&amp;train&quot; syntax lets us refer back to this object as &quot;*train&quot; elsewhere in the yaml file<br>\r\n&nbsp;&nbsp;&nbsp; dataset: &amp;train !obj:pylearn2.scripts.icml_2013_wrepl.black_box.black_box_dataset.BlackBoxDataset {<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; which_set: 'train',<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; start: 0,<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; stop: 900,<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; preprocessor: &amp;preprocessor !pkl: &quot;std.pkl&quot;<br>\r\n&nbsp;&nbsp;&nbsp; },<br>\r\n&nbsp;&nbsp;&nbsp; # Here we specify the model to train as being an MLP<br>\r\n&nbsp;&nbsp;&nbsp; model: !obj:pylearn2.models.mlp.MLP {<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; batch_size: 100,<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; layers : [<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; # We use two hidden layers with maxout activations<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; !obj:pylearn2.models.maxout.Maxout {<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; layer_name: 'h0',<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; num_units: 1875,<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; num_pieces: 2,<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; irange: .05,<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; # Rather than using weight decay, we constrain the norms of the weight vectors<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; # max_col_norm: 2.<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; },<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; !obj:pylearn2.models.maxout.Maxout {<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; layer_name: 'h1',<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; num_units: 469,<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; num_pieces: 2,<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; irange: .05,<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; # Rather than using weight decay, we constrain the norms of the weight vectors<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; # max_col_norm: 2.<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; },<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; !obj:pylearn2.models.mlp.Softmax {<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; layer_name: 'y',<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; init_bias_target_marginals: *train,<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; # Initialize the weights to all 0s<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; irange: .0,<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; n_classes: 9<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; }<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; ],<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; nvis: 1875,<br>\r\n&nbsp;&nbsp;&nbsp; },<br>\r\n&nbsp;&nbsp;&nbsp; # We train using SGD and momentum<br>\r\n&nbsp;&nbsp;&nbsp; algorithm: !obj:pylearn2.training_algorithms.sgd.SGD {<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; learning_rate: .1,<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; init_momentum: .5,<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; # We monitor how well we're doing during training on a validation set<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; monitoring_dataset:<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; {<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 'train' : *train,<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 'valid' : !obj:pylearn2.scripts.icml_2013_wrepl.black_box.black_box_dataset.BlackBoxDataset {<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; preprocessor: *preprocessor,<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; which_set: 'train',<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; start: 900,<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; stop: 1000,<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; }<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; },<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; cost: !obj:pylearn2.costs.mlp.dropout.Dropout {<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; # input_include_probs: { 'h0' : .8 },<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; # input_scales: { 'h0': 1. }<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; },<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; # We stop when validation set classification error hasn't decreased for 100 epochs<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; termination_criterion: !obj:pylearn2.termination_criteria.MonitorBased {<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; channel_name: &quot;valid_y_misclass&quot;,<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; prop_decrease: 0.,<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; N: 100<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; },<br>\r\n&nbsp;&nbsp;&nbsp; },<br>\r\n&nbsp;&nbsp;&nbsp; # We save the model whenever we improve on the validation set classification error<br>\r\n&nbsp;&nbsp;&nbsp; extensions: [<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; !obj:pylearn2.train_extensions.best_params.MonitorBasedSaveBest {<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; channel_name: 'valid_y_misclass',<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; save_path: &quot;${PYLEARN2_TRAIN_FILE_FULL_STEM}_best.pkl&quot;<br>\r\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; },<br>\r\n&nbsp;&nbsp;&nbsp; ],<br>\r\n&nbsp;&nbsp;&nbsp; save_path: &quot;mlp.pkl&quot;,<br>\r\n&nbsp;&nbsp;&nbsp; save_freq: 5<br>\r\n}<br>\r\n<br>\r\n</blockquote>\r\n<p>Any help for a python newbie much appreciated.</p>\r\n<p>Thanks</p>\r\n<p>John</p>\r\n<p></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 24193,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "05/11/2013 16:18:11",
      "content": "<p>Probably a log or an exp somewhere is getting too extreme of a value. This might not be due directly to the preprocessing. The issue could be that the preprocessing increased the range of values in the dataset and thus made the gradient steps bigger. You\r\n can fix that by reducing the learning rate, momentum, and irange. This is mostly a trial and error thing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 24232,
      "author_name": "johnclarke",
      "author_url": "",
      "post_date": "05/12/2013 11:55:04",
      "content": "<p>Thanks Ian, will give it a try.</p>\r\n<p></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 24240,
      "author_name": "johnclarke",
      "author_url": "",
      "post_date": "05/12/2013 16:09:19",
      "content": "<p>Hi Ian</p>\r\n<p></p>\r\n<p>Yep, you were right.&nbsp; Setting {learning_rate: .01,&nbsp; init_momentum: 0.0,} fixed the NaNs.&nbsp; If anyone is interested that 2-layer NN scores 0.55240 which is a victory for the maxout neuron!</p>\r\n<p></p>\r\n<p>~John</p>\r\n<p></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 27641,
      "author_name": "marcortiztorres0",
      "author_url": "",
      "post_date": "07/25/2013 08:47:22",
      "content": "<p>Sorry but I have a question. In the Maxout class what are num_pieces? The number of inputs to each unit?&nbsp;</p>\n<p>If I want to create a neural network that has 2 hidden layers should I use two Maxout hidden layers and one Softmax representing the output?</p>\n<p>Do I have to create a layer representing the input layer or it's not necessary?&nbsp;</p>\n<p>Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 90433,
      "author_name": "nahedmahmoud",
      "author_url": "",
      "post_date": "08/26/2015 14:44:15",
      "content": "<p>@John Clarke \nI am a little bit confused about how to back propagate errors in Maxout layer,\nDo I necessary need to use there dimension matrix to represent the weights at that linear.</p>\n\n<p>Thanks,\nNahed </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 90447,
      "author_name": "johnclarke",
      "author_url": "",
      "post_date": "08/26/2015 16:12:58",
      "content": "<p>Wow Nahed!  It's been an age since I looked at this.  Can't help you I'm afraid.</p>\n\n<p>Maybe try caffe now?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "23836": "",
    "23837": "",
    "23840": "",
    "23841": "",
    "23842": "",
    "23843": "",
    "23844": "",
    "23851": "",
    "23852": "",
    "23853": "",
    "23854": "",
    "23856": "",
    "23857": "",
    "23887": "",
    "23893": "",
    "23896": "",
    "23897": "",
    "23898": "",
    "23899": "",
    "23900": "",
    "23901": "",
    "23910": "",
    "23921": "",
    "23923": "",
    "23935": "",
    "24191": "",
    "24193": "",
    "24232": "",
    "24240": "",
    "27641": "",
    "90433": "John Clarke \r\nI am a little bit confused about how to back propagate errors in Maxout layer,\r\nDo I necessary need to use there dimension matrix to represent the weights at that linear.\r\n\r\nThanks,\r\nNahed",
    "90447": "Wow Nahed!  It's been an age since I looked at this.  Can't help you I'm afraid.\r\n\r\nMaybe try caffe now?",
    "368592": "While validation and testing you have to switch of your dropout errors and use full networks...\nAs mentioned in Video lectures of Stanford CS231n by Justin Johnson."
  },
  "source": "meta"
}