{
  "id": 4602,
  "title": "Autoencoders",
  "url": "/competitions/challenges-in-representation-learning-the-black-box-learning-challenge/discussion/4602",
  "author_name": "",
  "post_date": "2013-05-15T18:11:07.317Z",
  "votes": null,
  "comment_count": 7,
  "views": 5063,
  "content": "<p>A methodological question about autoencoders. I’ve seen in tutorials always are referred to [0,1]^d space.</p>\r\n<p>How compatibilize this in cases X are in [-1,1]^d or [-Inf,Inf]^d?</p>\r\n<p>Would be necessary a preprocessing step for mapping [-Inf,Inf] in [0,1]?</p>\r\n<p></p>\r\n<p>If X is real values [-Inf,Inf] and act_enc=act_dec=sigmoid, then</p>\r\n<p>[-Inf,Inf] -&gt; [0,1] -&gt; [0,1] and can't match the original [-Inf,Inf] values.</p>\r\n<p></p>\r\n<p>I know I've a conceptual problem here.</p>\r\n<p></p>\r\n<p></p>\r\n<p></p>\r\n<p></p>",
  "messages": [
    {
      "id": "24355",
      "postDate": "05/15/2013 18:11:07",
      "content": "<p>A methodological question about autoencoders. I’ve seen in tutorials always are referred to [0,1]^d space.</p>\r\n<p>How compatibilize this in cases X are in [-1,1]^d or [-Inf,Inf]^d?</p>\r\n<p>Would be necessary a preprocessing step for mapping [-Inf,Inf] in [0,1]?</p>\r\n<p></p>\r\n<p>If X is real values [-Inf,Inf] and act_enc=act_dec=sigmoid, then</p>\r\n<p>[-Inf,Inf] -&gt; [0,1] -&gt; [0,1] and can't match the original [-Inf,Inf] values.</p>\r\n<p></p>\r\n<p>I know I've a conceptual problem here.</p>\r\n<p></p>\r\n<p></p>\r\n<p></p>\r\n<p></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "24356",
      "postDate": "05/15/2013 18:49:03",
      "content": "<p>For [-1,1] inputs you can use tanh activations for your decoder and modify the loss function accordingly (relatively simple). For general &quot;[-inf, inf]&quot; inputs you will likely want to simply use linear activation for the decoder &#43; mean squared error as the\r\n loss/objective.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "24357",
      "postDate": "05/15/2013 18:58:28",
      "content": "<p>[quote=Dumitru;24356]</p>\r\n<p>For [-1,1] inputs you can use tanh activations for your decoder and modify the loss function accordingly (relatively simple). For general &quot;[-inf, inf]&quot; inputs you will likely want to simply use linear activation for the decoder &#43; mean squared error as the\r\n loss/objective.</p>\r\n<p>[/quote]</p>\r\n<p></p>\r\n<p>Thank you, It's simpler than I though.</p>\r\n<p>I try with linear activation but always obtain:</p>\r\n<div>pylearn2\\training_algorithms\\sgd.py &nbsp;line 338 in train</div>\r\n<div>raise Exception (&quot;NaN in &quot; &#43; param.name)</div>\r\n<div>Exception: Nan in vb</div>\r\n<div></div>\r\n<div>With sigmoid or tanh my autoencoders works fine but with linear don't do it.</div>\r\n<div></div>\r\n<div>Some tip?</div>\r\n<div></div>\r\n<div>EDIT: Workaround</div>\r\n<div></div>\r\n<div>With preprocessing (preprocessor: &amp;preprocessor !pkl: &quot;std.pkl&quot;), the process fail, but without it work fine.</div>\r\n<div></div>\r\n<div>Only happened with linear activation.</div>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "24694",
      "postDate": "05/23/2013 01:36:21",
      "content": "<p>Actually I am getting the exact same error as before even without preprocessing the data ... Did you ever find a way to use linear activation with preprocessing?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "24695",
      "postDate": "05/23/2013 01:48:43",
      "content": "<p>Never mind, after getting the latest pylearn source code my problem went away.</p>\r\n<p>[edit] actually, even with the latest the problem is still there for linear decoders.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "24798",
      "postDate": "05/25/2013 10:15:35",
      "content": "<p>Anybody have any idea what the reason for the NaNs in the autoencoders with rectified linear units is, and how to avoid them?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "24802",
      "postDate": "05/25/2013 13:41:18",
      "content": "<p>I haven't tracked it down well enough to prove this, but I think the issue is that if you have several rectified layers composed together, then the wrong configuration of the weights essentially gives you exponentiation. Suppose you have a deep, narrow network\r\n with one unit in each layer. If the input is 1 and the weight for each layer is w and w &gt; 0, then the output of layer L is w^L. If your output layer is softmax, when you compute the denominator you must take e^input and I suspect that turns into Inf and then\r\n gets converted to NaN when you try to do arithmetic with it to compute the gradient.</p>\r\n<p>In practice, the problem usually seems to go away if you reduce the momentum coefficient or learning rate.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "24808",
      "postDate": "05/25/2013 16:16:33",
      "content": "<p>For me it did indeed go away by reducing the momentum. But later, I ended up preprocssing the data with a min-max scaler into the [0,1] range just so I could use a sigmoid decoder rather than a linear one.</p>\r\n<p>[quote=Ian Goodfellow;24802]</p>\r\n<p>I haven't tracked it down well enough to prove this, but I think the issue is that if you have several rectified layers composed together, then the wrong configuration of the weights essentially gives you exponentiation. Suppose you have a deep, narrow network\r\n with one unit in each layer. If the input is 1 and the weight for each layer is w and w &gt; 0, then the output of layer L is w^L. If your output layer is softmax, when you compute the denominator you must take e^input and I suspect that turns into Inf and then\r\n gets converted to NaN when you try to do arithmetic with it to compute the gradient.</p>\r\n<p>In practice, the problem usually seems to go away if you reduce the momentum coefficient or learning rate.</p>\r\n<p>[/quote]</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 24356,
      "author_name": "dumitru0",
      "author_url": "",
      "post_date": "05/15/2013 18:49:03",
      "content": "<p>For [-1,1] inputs you can use tanh activations for your decoder and modify the loss function accordingly (relatively simple). For general &quot;[-inf, inf]&quot; inputs you will likely want to simply use linear activation for the decoder &#43; mean squared error as the\r\n loss/objective.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 24357,
      "author_name": "blindape",
      "author_url": "",
      "post_date": "05/15/2013 18:58:28",
      "content": "<p>[quote=Dumitru;24356]</p>\r\n<p>For [-1,1] inputs you can use tanh activations for your decoder and modify the loss function accordingly (relatively simple). For general &quot;[-inf, inf]&quot; inputs you will likely want to simply use linear activation for the decoder &#43; mean squared error as the\r\n loss/objective.</p>\r\n<p>[/quote]</p>\r\n<p></p>\r\n<p>Thank you, It's simpler than I though.</p>\r\n<p>I try with linear activation but always obtain:</p>\r\n<div>pylearn2\\training_algorithms\\sgd.py &nbsp;line 338 in train</div>\r\n<div>raise Exception (&quot;NaN in &quot; &#43; param.name)</div>\r\n<div>Exception: Nan in vb</div>\r\n<div></div>\r\n<div>With sigmoid or tanh my autoencoders works fine but with linear don't do it.</div>\r\n<div></div>\r\n<div>Some tip?</div>\r\n<div></div>\r\n<div>EDIT: Workaround</div>\r\n<div></div>\r\n<div>With preprocessing (preprocessor: &amp;preprocessor !pkl: &quot;std.pkl&quot;), the process fail, but without it work fine.</div>\r\n<div></div>\r\n<div>Only happened with linear activation.</div>",
      "votes": null,
      "replies": []
    },
    {
      "id": 24694,
      "author_name": "seylom",
      "author_url": "",
      "post_date": "05/23/2013 01:36:21",
      "content": "<p>Actually I am getting the exact same error as before even without preprocessing the data ... Did you ever find a way to use linear activation with preprocessing?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 24695,
      "author_name": "seylom",
      "author_url": "",
      "post_date": "05/23/2013 01:48:43",
      "content": "<p>Never mind, after getting the latest pylearn source code my problem went away.</p>\r\n<p>[edit] actually, even with the latest the problem is still there for linear decoders.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 24798,
      "author_name": "njdevos",
      "author_url": "",
      "post_date": "05/25/2013 10:15:35",
      "content": "<p>Anybody have any idea what the reason for the NaNs in the autoencoders with rectified linear units is, and how to avoid them?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 24802,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "05/25/2013 13:41:18",
      "content": "<p>I haven't tracked it down well enough to prove this, but I think the issue is that if you have several rectified layers composed together, then the wrong configuration of the weights essentially gives you exponentiation. Suppose you have a deep, narrow network\r\n with one unit in each layer. If the input is 1 and the weight for each layer is w and w &gt; 0, then the output of layer L is w^L. If your output layer is softmax, when you compute the denominator you must take e^input and I suspect that turns into Inf and then\r\n gets converted to NaN when you try to do arithmetic with it to compute the gradient.</p>\r\n<p>In practice, the problem usually seems to go away if you reduce the momentum coefficient or learning rate.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 24808,
      "author_name": "seylom",
      "author_url": "",
      "post_date": "05/25/2013 16:16:33",
      "content": "<p>For me it did indeed go away by reducing the momentum. But later, I ended up preprocssing the data with a min-max scaler into the [0,1] range just so I could use a sigmoid decoder rather than a linear one.</p>\r\n<p>[quote=Ian Goodfellow;24802]</p>\r\n<p>I haven't tracked it down well enough to prove this, but I think the issue is that if you have several rectified layers composed together, then the wrong configuration of the weights essentially gives you exponentiation. Suppose you have a deep, narrow network\r\n with one unit in each layer. If the input is 1 and the weight for each layer is w and w &gt; 0, then the output of layer L is w^L. If your output layer is softmax, when you compute the denominator you must take e^input and I suspect that turns into Inf and then\r\n gets converted to NaN when you try to do arithmetic with it to compute the gradient.</p>\r\n<p>In practice, the problem usually seems to go away if you reduce the momentum coefficient or learning rate.</p>\r\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "24355": "",
    "24356": "",
    "24357": "",
    "24694": "",
    "24695": "",
    "24798": "",
    "24802": "",
    "24808": ""
  },
  "source": "meta"
}