{
  "id": 4352,
  "title": "Sigmoid units in PyLearn2",
  "url": "/competitions/challenges-in-representation-learning-the-black-box-learning-challenge/discussion/4352",
  "author_name": "",
  "post_date": "2013-04-17T21:55:12.933Z",
  "votes": null,
  "comment_count": 11,
  "views": 5349,
  "content": "<p>I'm experimenting with Sigmoid units by modifying the supplied benchmark code for 2 layer networks with rectifier units, but I'm getting a surprising result. &nbsp;My validation error keeps getting stuck at 0.800000011921.</p>\r\n<p>&nbsp;</p>\r\n<p>I've tried a wide range of parameters to the SIgmoid layer, but it doesn't appear to affect my output. &nbsp;I'm currently initializing it in the yaml file as</p>\r\n<pre>            !obj:pylearn2.models.mlp.Sigmoid {<br>                layer_name: 'h1',<br>                dim: 1875,<br>                monitor_style: 'detection',<br>                sparse_init: True<br>            },</pre>\r\n<p>&nbsp;</p>\r\n<p><span style=\"font-size:14px; line-height:1.4em\">though I've tried many other ways to initialize it. &nbsp;</span></p>\r\n<p><span style=\"font-size:14px; line-height:1.4em\">Is the Sigmoid benchmark publicly available? Anyone (and by &quot;anyone&quot;, I'm guessing I'm just asking Ian:) have any advice?&nbsp;</span><span style=\"font-size:14px; line-height:1.4em\">Also, is there a better venue\r\n where I should ask pylearn2 questions in the future?</span></p>\r\n<p><span style=\"font-size:14px; line-height:1.4em\">Thanks!</span></p>",
  "messages": [
    {
      "id": "23006",
      "postDate": "04/17/2013 21:55:12",
      "content": "<p>I'm experimenting with Sigmoid units by modifying the supplied benchmark code for 2 layer networks with rectifier units, but I'm getting a surprising result. &nbsp;My validation error keeps getting stuck at 0.800000011921.</p>\r\n<p>&nbsp;</p>\r\n<p>I've tried a wide range of parameters to the SIgmoid layer, but it doesn't appear to affect my output. &nbsp;I'm currently initializing it in the yaml file as</p>\r\n<pre>            !obj:pylearn2.models.mlp.Sigmoid {<br>                layer_name: 'h1',<br>                dim: 1875,<br>                monitor_style: 'detection',<br>                sparse_init: True<br>            },</pre>\r\n<p>&nbsp;</p>\r\n<p><span style=\"font-size:14px; line-height:1.4em\">though I've tried many other ways to initialize it. &nbsp;</span></p>\r\n<p><span style=\"font-size:14px; line-height:1.4em\">Is the Sigmoid benchmark publicly available? Anyone (and by &quot;anyone&quot;, I'm guessing I'm just asking Ian:) have any advice?&nbsp;</span><span style=\"font-size:14px; line-height:1.4em\">Also, is there a better venue\r\n where I should ask pylearn2 questions in the future?</span></p>\r\n<p><span style=\"font-size:14px; line-height:1.4em\">Thanks!</span></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23008",
      "postDate": "04/17/2013 22:01:12",
      "content": "<p>I'm guessing the problem is with sparse_init. That argument should actually be an integer, not a bool. It tells it how many weights should be initialized to be non-zero per hidden unit. Usually you should just set this to 15 and forget about it. I think\r\n it was James Martens that first advocated this initialization scheme. The idea is that the scale of the initial input to each unit doesn't depend on the size of the previous layer, so it's easier to initialize big nets.</p>\r\n<p>It's fine to ask pylearn2 questions on the forum here, or you can e-mail pylearn-dev@googlegroups.com.</p>\r\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23444",
      "postDate": "04/25/2013 05:34:20",
      "content": "<p>Hi, I have same issue. Did you fix it?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23454",
      "postDate": "04/25/2013 12:22:15",
      "content": "<p>There shouldn't be anything to fix in pylearn2, just with your own script for passing arguments to it.</p>\r\n<p>Have you tried using sparse_init: 15 , or using irange : .05 instead of using sparse_init?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23461",
      "postDate": "04/25/2013 14:58:00",
      "content": "<p>With regards to my original issue, Ian's proposed solution fixed it.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23496",
      "postDate": "04/26/2013 01:33:39",
      "content": "<p>Thank you. It works, but sigmoid seems still terrible.</p>\r\n<p>~0.81 for single leayer.</p>\r\n<p>@DanB</p>\r\n<p>@Ian</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23521",
      "postDate": "04/26/2013 12:37:35",
      "content": "<p>When you first start an MLP, it wonders around for a some iterations before it starts making progress. &nbsp;The amount of wandering depends to some extent on the initialization parameters (the irange or sparse_init), though the initial momentum setting and learning\r\n rate settings play a huge role too.</p>\r\n<p>You've likely also set the mlp to stop whenever it goes sufficiently long without making progress. &nbsp;If your final score is 0.81, then it's tripping this condition during the wandering period. &nbsp;</p>\r\n<p>In some cases, I've played with the momentum and learning rate to get an mlp that starts walking towards the goal sooner, and that has solve the problem. &nbsp;A more reliable solution would be to set the termination criteria to be more patient, and hope the\r\n mlp starts making progress before it hits the termination criteria. &nbsp;</p>\r\n<p>To allow more patience before termination, you need to increase N in the termination criteria. &nbsp;So that might look like</p>\r\n<pre>            !obj:pylearn2.termination_criteria.MonitorBased {<br>                prop_decrease: 0.,<br>                channel_name: &quot;valid_y_misclass&quot;,<br>                N: 50,<br>            },</pre>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23794",
      "postDate": "05/02/2013 01:32:34",
      "content": "<p>@Ian</p>\r\n<p>Can you post here the yaml you used for the 3 layered sigmoid benchmark? I'm findind it hard to tune the learn rate and momentum and get past the 30% acuracy on the validation set with 3 layer sigmoid units.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23806",
      "postDate": "05/02/2013 14:26:43",
      "content": "<p>I didn't make that one; Dumitru did, and I think he didn't use pylearn2.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23811",
      "postDate": "05/02/2013 15:05:51",
      "content": "<p>Yeah, sorry, didn't use pylearn2 indeed. But the code is simple: essentially, if you transform the training data into the format that&nbsp;http://deeplearning.net/tutorial/SdA.html#sda expects, you'll be able to replicate my results. In the test_SdA() function\r\n I used finetune_lr=0.1, 0 pretraining epochs (though you can change this obviously), batch_size of 20 and training_epochs of 1000. It converged before that however.</p>\r\n<p>For hidden_layer_sizes I used 1000 each.</p>\r\n<p>Hope that helps! I'd post the file itself, but it's moderately useful since you'd need to generate a new dataset (I used the actual test labels to evaluate the performance) and it might be instructive for you to debug doing that in order to understand what\r\n the tutorial code does.</p>\r\n<p>Hope that helps!</p>\r\n<p>Dumitru</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23819",
      "postDate": "05/02/2013 16:11:42",
      "content": "<p>@Ian,</p>\r\n<p>&nbsp; &nbsp; How could a configure the yaml to do an RBM transformation?&nbsp;</p>\r\n<p>&nbsp; &nbsp; Say i want to build a DBN here, with 2 hidden layers. How would i configure it? Can i do it with one yaml, or i need to construct an RBM presentation of the data first and then use it as a tranformation to another NNet? How would i configure the yaml\r\n to a tranformation using a rbm.pkl (pretty much what you did with the zca example)?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23821",
      "postDate": "05/02/2013 16:30:56",
      "content": "<p>Look at pylearn2/scripts/tutorials/deep_trainer.</p>\r\n<p>I don't personally train DBNs myself, but one of my colleagues put together that tutorial on it.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 23008,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "04/17/2013 22:01:12",
      "content": "<p>I'm guessing the problem is with sparse_init. That argument should actually be an integer, not a bool. It tells it how many weights should be initialized to be non-zero per hidden unit. Usually you should just set this to 15 and forget about it. I think\r\n it was James Martens that first advocated this initialization scheme. The idea is that the scale of the initial input to each unit doesn't depend on the size of the previous layer, so it's easier to initialize big nets.</p>\r\n<p>It's fine to ask pylearn2 questions on the forum here, or you can e-mail pylearn-dev@googlegroups.com.</p>\r\n<p>&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23444,
      "author_name": "binghsu",
      "author_url": "",
      "post_date": "04/25/2013 05:34:20",
      "content": "<p>Hi, I have same issue. Did you fix it?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23454,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "04/25/2013 12:22:15",
      "content": "<p>There shouldn't be anything to fix in pylearn2, just with your own script for passing arguments to it.</p>\r\n<p>Have you tried using sparse_init: 15 , or using irange : .05 instead of using sparse_init?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23461,
      "author_name": "dansbecker",
      "author_url": "",
      "post_date": "04/25/2013 14:58:00",
      "content": "<p>With regards to my original issue, Ian's proposed solution fixed it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23496,
      "author_name": "binghsu",
      "author_url": "",
      "post_date": "04/26/2013 01:33:39",
      "content": "<p>Thank you. It works, but sigmoid seems still terrible.</p>\r\n<p>~0.81 for single leayer.</p>\r\n<p>@DanB</p>\r\n<p>@Ian</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23521,
      "author_name": "dansbecker",
      "author_url": "",
      "post_date": "04/26/2013 12:37:35",
      "content": "<p>When you first start an MLP, it wonders around for a some iterations before it starts making progress. &nbsp;The amount of wandering depends to some extent on the initialization parameters (the irange or sparse_init), though the initial momentum setting and learning\r\n rate settings play a huge role too.</p>\r\n<p>You've likely also set the mlp to stop whenever it goes sufficiently long without making progress. &nbsp;If your final score is 0.81, then it's tripping this condition during the wandering period. &nbsp;</p>\r\n<p>In some cases, I've played with the momentum and learning rate to get an mlp that starts walking towards the goal sooner, and that has solve the problem. &nbsp;A more reliable solution would be to set the termination criteria to be more patient, and hope the\r\n mlp starts making progress before it hits the termination criteria. &nbsp;</p>\r\n<p>To allow more patience before termination, you need to increase N in the termination criteria. &nbsp;So that might look like</p>\r\n<pre>            !obj:pylearn2.termination_criteria.MonitorBased {<br>                prop_decrease: 0.,<br>                channel_name: &quot;valid_y_misclass&quot;,<br>                N: 50,<br>            },</pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23794,
      "author_name": "leustagos",
      "author_url": "",
      "post_date": "05/02/2013 01:32:34",
      "content": "<p>@Ian</p>\r\n<p>Can you post here the yaml you used for the 3 layered sigmoid benchmark? I'm findind it hard to tune the learn rate and momentum and get past the 30% acuracy on the validation set with 3 layer sigmoid units.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23806,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "05/02/2013 14:26:43",
      "content": "<p>I didn't make that one; Dumitru did, and I think he didn't use pylearn2.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23811,
      "author_name": "dumitru0",
      "author_url": "",
      "post_date": "05/02/2013 15:05:51",
      "content": "<p>Yeah, sorry, didn't use pylearn2 indeed. But the code is simple: essentially, if you transform the training data into the format that&nbsp;http://deeplearning.net/tutorial/SdA.html#sda expects, you'll be able to replicate my results. In the test_SdA() function\r\n I used finetune_lr=0.1, 0 pretraining epochs (though you can change this obviously), batch_size of 20 and training_epochs of 1000. It converged before that however.</p>\r\n<p>For hidden_layer_sizes I used 1000 each.</p>\r\n<p>Hope that helps! I'd post the file itself, but it's moderately useful since you'd need to generate a new dataset (I used the actual test labels to evaluate the performance) and it might be instructive for you to debug doing that in order to understand what\r\n the tutorial code does.</p>\r\n<p>Hope that helps!</p>\r\n<p>Dumitru</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23819,
      "author_name": "leustagos",
      "author_url": "",
      "post_date": "05/02/2013 16:11:42",
      "content": "<p>@Ian,</p>\r\n<p>&nbsp; &nbsp; How could a configure the yaml to do an RBM transformation?&nbsp;</p>\r\n<p>&nbsp; &nbsp; Say i want to build a DBN here, with 2 hidden layers. How would i configure it? Can i do it with one yaml, or i need to construct an RBM presentation of the data first and then use it as a tranformation to another NNet? How would i configure the yaml\r\n to a tranformation using a rbm.pkl (pretty much what you did with the zca example)?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23821,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "05/02/2013 16:30:56",
      "content": "<p>Look at pylearn2/scripts/tutorials/deep_trainer.</p>\r\n<p>I don't personally train DBNs myself, but one of my colleagues put together that tutorial on it.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "23006": "",
    "23008": "",
    "23444": "",
    "23454": "",
    "23461": "",
    "23496": "",
    "23521": "",
    "23794": "",
    "23806": "",
    "23811": "",
    "23819": "",
    "23821": ""
  },
  "source": "meta"
}