{
  "id": 4514,
  "title": "Understanding black_box_dataset.py",
  "url": "/competitions/challenges-in-representation-learning-the-black-box-learning-challenge/discussion/4514",
  "author_name": "",
  "post_date": "2013-05-06T19:19:18.887Z",
  "votes": null,
  "comment_count": 1,
  "views": 896,
  "content": "<p>I'm trying understand this script (this is my first approach to python).</p>\r\n<p>What is the use of preprocessor, fit_preprocessor and fit_test_preprocessor?. Where them are effectively used?&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>Other thing I can't find is relative a batch_size. This parameter is used in every yaml with values 100,20,6... but I don't know if its justification is just for memory manage or has other effect on the model and if so, what is the rule, if any, to choose\r\n it?</p>",
  "messages": [
    {
      "id": "23949",
      "postDate": "05/06/2013 19:19:18",
      "content": "<p>I'm trying understand this script (this is my first approach to python).</p>\r\n<p>What is the use of preprocessor, fit_preprocessor and fit_test_preprocessor?. Where them are effectively used?&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>Other thing I can't find is relative a batch_size. This parameter is used in every yaml with values 100,20,6... but I don't know if its justification is just for memory manage or has other effect on the model and if so, what is the rule, if any, to choose\r\n it?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23950",
      "postDate": "05/06/2013 19:30:06",
      "content": "<p>preprocessor: the argument is a Preprocessor object. It will be executed to preprocess your data.</p>\r\n<p>fit_preprocessor: the argument is a boolean flag, saying whether the preprocessor should be fit to the data or not. Some preprocessors, like ZCA, need to learn something about the data distribution before they can be applied. But you might want to prevent\r\n them from fitting when you run this script, for example, if you're loading a preprocessor from disk that has already been fit to a different dataset.</p>\r\n<p>fit_test_preprocessor: another boolean flag. If true, the preprocessor should be re-fit on the test set. If false, it will keep using the same parameters as were learned on the train set. Usually you want this to be false, but sometimes if you're doing some\r\n kind of transductive reasoning you might want to re-fit on the test set.</p>\r\n<p>The batch size has several different effects on the runtime, memory use, final training set error, and generalization error, not all of which are easy to predict. Like most parameters with neural networks, there are lots of rules of thumb for setting it,\r\n but no hard theory that I'm aware of. The only real way to set it is cross-validation.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 23950,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "05/06/2013 19:30:06",
      "content": "<p>preprocessor: the argument is a Preprocessor object. It will be executed to preprocess your data.</p>\r\n<p>fit_preprocessor: the argument is a boolean flag, saying whether the preprocessor should be fit to the data or not. Some preprocessors, like ZCA, need to learn something about the data distribution before they can be applied. But you might want to prevent\r\n them from fitting when you run this script, for example, if you're loading a preprocessor from disk that has already been fit to a different dataset.</p>\r\n<p>fit_test_preprocessor: another boolean flag. If true, the preprocessor should be re-fit on the test set. If false, it will keep using the same parameters as were learned on the train set. Usually you want this to be false, but sometimes if you're doing some\r\n kind of transductive reasoning you might want to re-fit on the test set.</p>\r\n<p>The batch size has several different effects on the runtime, memory use, final training set error, and generalization error, not all of which are easy to predict. Like most parameters with neural networks, there are lots of rules of thumb for setting it,\r\n but no hard theory that I'm aware of. The only real way to set it is cross-validation.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "23949": "",
    "23950": ""
  },
  "source": "meta"
}