{
  "id": 20073,
  "title": "Learning rate decay for SGD",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/20073",
  "author_name": "",
  "post_date": "2016-04-11T20:31:03.080Z",
  "votes": null,
  "comment_count": 2,
  "views": 1417,
  "content": "<p>Is there any documentation or general practice for setting the learning rate decay if I want to take multiple passes over the training set? </p>\n\n<p>in Keras I just copied in the out of the box code:</p>\n\n<p>sgd = SGD(lr=0.1, decay=1e-6, momentum=0.9, nesterov=True)</p>\n\n<p>but I think the decay only happens if I have more than one epoch? or is it happening with each batch? This seems like a pretty small number, can I have another epoch with this rate of decay?  </p>",
  "messages": [
    {
      "id": "114545",
      "postDate": "04/11/2016 20:31:03",
      "content": "<p>Is there any documentation or general practice for setting the learning rate decay if I want to take multiple passes over the training set? </p>\n\n<p>in Keras I just copied in the out of the box code:</p>\n\n<p>sgd = SGD(lr=0.1, decay=1e-6, momentum=0.9, nesterov=True)</p>\n\n<p>but I think the decay only happens if I have more than one epoch? or is it happening with each batch? This seems like a pretty small number, can I have another epoch with this rate of decay?  </p>",
      "rawMarkdown": "Is there any documentation or general practice for setting the learning rate decay if I want to take multiple passes over the training set? \r\n\r\nin Keras I just copied in the out of the box code:\r\n\r\nsgd = SGD(lr=0.1, decay=1e-6, momentum=0.9, nesterov=True)\r\n\r\nbut I think the decay only happens if I have more than one epoch? or is it happening with each batch? This seems like a pretty small number, can I have another epoch with this rate of decay?",
      "votes": null
    },
    {
      "id": "114593",
      "postDate": "04/12/2016 09:47:31",
      "content": "<p>In my previous experience, it happens with each batch.\nYou can get the exact learning rate of each batch by doing this:</p>\n\n<pre><code>opt = model.optimizer\nexact_lr = opt.lr.get_value() * (1.0 / (1.0 + opt.decay.get_value() * opt.iterations.get_value()))\n</code></pre>\n\n<p>opt.iterations is the number of updates, i.e., number of batches trained, and opt.decay is the one you pass into the constructor of SGD.</p>",
      "rawMarkdown": "In my previous experience, it happens with each batch.\r\nYou can get the exact learning rate of each batch by doing this:\r\n\r\n    opt = model.optimizer\r\n    exact_lr = opt.lr.get_value() * (1.0 / (1.0 + opt.decay.get_value() * opt.iterations.get_value()))\r\n\r\nopt.iterations is the number of updates, i.e., number of batches trained, and opt.decay is the one you pass into the constructor of SGD.",
      "votes": null
    },
    {
      "id": "119981",
      "postDate": "05/14/2016 09:18:43",
      "content": "<p>yes it happens for each batch. see \n<a href=\"http://stats.stackexchange.com/questions/211334/keras-how-does-sgd-learning-rate-decay-work\">http://stats.stackexchange.com/questions/211334/keras-how-does-sgd-learning-rate-decay-work</a></p>",
      "rawMarkdown": "yes it happens for each batch. see \r\nhttp://stats.stackexchange.com/questions/211334/keras-how-does-sgd-learning-rate-decay-work",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 114593,
      "author_name": "ianlini",
      "author_url": "",
      "post_date": "04/12/2016 09:47:31",
      "content": "<p>In my previous experience, it happens with each batch.\nYou can get the exact learning rate of each batch by doing this:</p>\n\n<pre><code>opt = model.optimizer\nexact_lr = opt.lr.get_value() * (1.0 / (1.0 + opt.decay.get_value() * opt.iterations.get_value()))\n</code></pre>\n\n<p>opt.iterations is the number of updates, i.e., number of batches trained, and opt.decay is the one you pass into the constructor of SGD.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119981,
      "author_name": "srinathperera",
      "author_url": "",
      "post_date": "05/14/2016 09:18:43",
      "content": "<p>yes it happens for each batch. see \n<a href=\"http://stats.stackexchange.com/questions/211334/keras-how-does-sgd-learning-rate-decay-work\">http://stats.stackexchange.com/questions/211334/keras-how-does-sgd-learning-rate-decay-work</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "114545": "Is there any documentation or general practice for setting the learning rate decay if I want to take multiple passes over the training set? \r\n\r\nin Keras I just copied in the out of the box code:\r\n\r\nsgd = SGD(lr=0.1, decay=1e-6, momentum=0.9, nesterov=True)\r\n\r\nbut I think the decay only happens if I have more than one epoch? or is it happening with each batch? This seems like a pretty small number, can I have another epoch with this rate of decay?",
    "114593": "In my previous experience, it happens with each batch.\r\nYou can get the exact learning rate of each batch by doing this:\r\n\r\n    opt = model.optimizer\r\n    exact_lr = opt.lr.get_value() * (1.0 / (1.0 + opt.decay.get_value() * opt.iterations.get_value()))\r\n\r\nopt.iterations is the number of updates, i.e., number of batches trained, and opt.decay is the one you pass into the constructor of SGD.",
    "119981": "yes it happens for each batch. see \r\nhttp://stats.stackexchange.com/questions/211334/keras-how-does-sgd-learning-rate-decay-work"
  },
  "source": "meta"
}