{
  "id": 4447,
  "title": "How to use other preprocessor in PyLearn2?",
  "url": "/competitions/challenges-in-representation-learning-the-black-box-learning-challenge/discussion/4447",
  "author_name": "",
  "post_date": "2013-04-26T12:51:13.643Z",
  "votes": null,
  "comment_count": 3,
  "views": 2307,
  "content": "<p>I have trained other model, aka, RBM, and output is rbm.pkl.</p>\r\n<p>I change the code in the dataset.py to:</p>\r\n<p>preprocessor.input_to_h_from_v(X)</p>\r\n<p>and I have seen the code in rbm.py that it is&nbsp;</p>\r\n<p><span class=\"x_k\">return</span><span class=\"x_p\">[</span><span class=\"x_bp\">self</span><span class=\"x_o\">.</span><span class=\"x_n\">input_to_h_from_v</span><span class=\"x_p\">(</span><span class=\"x_n\">vis</span><span class=\"x_p\">)&nbsp;</span><span class=\"x_k\">for&nbsp;</span><span class=\"x_n\">vis&nbsp;</span><span class=\"x_ow\">in&nbsp;</span><span class=\"x_n\">v</span><span class=\"x_p\">]</span></p>\r\n<p><span class=\"x_p\">But I got the error:</span></p>\r\n<p>return [self.input_to_h_from_v(vis) for vis in v]</p>\r\n<p>TypeError: 'numpy.float32' object is not iterable</p>\r\n<p>Can any one who is familiar with PyLearn2 tell me where I was wrong?</p>\r\n<p>I am quite confused and there is so limited documents.</p>\r\n<p>Bow</p>\r\n<p><span class=\"x_p\"><br>\r\n</span></p>",
  "messages": [
    {
      "id": "23523",
      "postDate": "04/26/2013 12:51:13",
      "content": "<p>I have trained other model, aka, RBM, and output is rbm.pkl.</p>\r\n<p>I change the code in the dataset.py to:</p>\r\n<p>preprocessor.input_to_h_from_v(X)</p>\r\n<p>and I have seen the code in rbm.py that it is&nbsp;</p>\r\n<p><span class=\"x_k\">return</span><span class=\"x_p\">[</span><span class=\"x_bp\">self</span><span class=\"x_o\">.</span><span class=\"x_n\">input_to_h_from_v</span><span class=\"x_p\">(</span><span class=\"x_n\">vis</span><span class=\"x_p\">)&nbsp;</span><span class=\"x_k\">for&nbsp;</span><span class=\"x_n\">vis&nbsp;</span><span class=\"x_ow\">in&nbsp;</span><span class=\"x_n\">v</span><span class=\"x_p\">]</span></p>\r\n<p><span class=\"x_p\">But I got the error:</span></p>\r\n<p>return [self.input_to_h_from_v(vis) for vis in v]</p>\r\n<p>TypeError: 'numpy.float32' object is not iterable</p>\r\n<p>Can any one who is familiar with PyLearn2 tell me where I was wrong?</p>\r\n<p>I am quite confused and there is so limited documents.</p>\r\n<p>Bow</p>\r\n<p><span class=\"x_p\"><br>\r\n</span></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23527",
      "postDate": "04/26/2013 14:36:55",
      "content": "<p>I'm not sure which line in dataset.py you've changed to&nbsp;</p>\r\n<p><span>preprocessor.input_to_h_from_v(X)</span></p>\r\n<p>so it's hard to debug it. &nbsp;Pylearn2 is built on theano, which adds a lot of complexity to normal python programming. &nbsp;Fortunately, pylearn2 saves you from having to muck about too much in the raw python/theano code by using separate model specification files.</p>\r\n<p>You write model specifications in yaml files and then call train.py on that yaml file. &nbsp;If you've already trained a model, you may have already done that. &nbsp;Otherwise, start by<span style=\"font-size:14px; line-height:1.4em\">&nbsp;looking at the yaml files in the&nbsp;scripts/icml_2013_wrepl/black_box\r\n and modifying those model specifications as desired. &nbsp;There is also a README in that directory with other relevant details.</span></p>\r\n<p>It's conceivable to do everything directly by modifying the library (the py files), but this is making things much harder on yourself.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23529",
      "postDate": "04/26/2013 16:07:17",
      "content": "<p>I could be mistaken, but it sounds like what you want to do is simply use your existing RBM to tranform the dataset for use in another layer (say to create a deep rbm)? &nbsp;If that is the case all you need to do is use the TransformerData set, which take a\r\n 'raw' dataset (your original) and a a 'transformer' which would be your rbm.pkl</p>\r\n<p>If for some other reason you want to actually manipulate the underlying numpy matrix you can get the data from the dataset's 'X' attribute and can call the model's 'perform' method. &nbsp;The output of that will be as if you had run that matrix through a layer\r\n of the rbm (or whatever model you're using).</p>\r\n<p>If you just want to use one of the existing preprocessors there's actually a 'preprocessor' argument to DataSet which will allow you to add one (you can add many by using the Pipeline preprocessor)</p>\r\n<p>There should be no reaon to edit dataset.py from what I've seen so far.</p>\r\n<p>Hope that helps!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23584",
      "postDate": "04/28/2013 19:31:43",
      "content": "<p>DanB and willkurt are correct.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 23527,
      "author_name": "dansbecker",
      "author_url": "",
      "post_date": "04/26/2013 14:36:55",
      "content": "<p>I'm not sure which line in dataset.py you've changed to&nbsp;</p>\r\n<p><span>preprocessor.input_to_h_from_v(X)</span></p>\r\n<p>so it's hard to debug it. &nbsp;Pylearn2 is built on theano, which adds a lot of complexity to normal python programming. &nbsp;Fortunately, pylearn2 saves you from having to muck about too much in the raw python/theano code by using separate model specification files.</p>\r\n<p>You write model specifications in yaml files and then call train.py on that yaml file. &nbsp;If you've already trained a model, you may have already done that. &nbsp;Otherwise, start by<span style=\"font-size:14px; line-height:1.4em\">&nbsp;looking at the yaml files in the&nbsp;scripts/icml_2013_wrepl/black_box\r\n and modifying those model specifications as desired. &nbsp;There is also a README in that directory with other relevant details.</span></p>\r\n<p>It's conceivable to do everything directly by modifying the library (the py files), but this is making things much harder on yourself.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23529,
      "author_name": "willkurt",
      "author_url": "",
      "post_date": "04/26/2013 16:07:17",
      "content": "<p>I could be mistaken, but it sounds like what you want to do is simply use your existing RBM to tranform the dataset for use in another layer (say to create a deep rbm)? &nbsp;If that is the case all you need to do is use the TransformerData set, which take a\r\n 'raw' dataset (your original) and a a 'transformer' which would be your rbm.pkl</p>\r\n<p>If for some other reason you want to actually manipulate the underlying numpy matrix you can get the data from the dataset's 'X' attribute and can call the model's 'perform' method. &nbsp;The output of that will be as if you had run that matrix through a layer\r\n of the rbm (or whatever model you're using).</p>\r\n<p>If you just want to use one of the existing preprocessors there's actually a 'preprocessor' argument to DataSet which will allow you to add one (you can add many by using the Pipeline preprocessor)</p>\r\n<p>There should be no reaon to edit dataset.py from what I've seen so far.</p>\r\n<p>Hope that helps!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23584,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "04/28/2013 19:31:43",
      "content": "<p>DanB and willkurt are correct.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "23523": "",
    "23527": "",
    "23529": "",
    "23584": ""
  },
  "source": "meta"
}