{
  "id": 125614,
  "title": "Some questions about the learning rate",
  "url": "/competitions/pku-autonomous-driving/discussion/125614",
  "author_name": "",
  "post_date": "2020-01-12T08:36:00.588356900Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>explrscheduler = lrscheduler.StepLR(optimizer, stepsize=max(nepochs, 10) * len(trainloader) // 3, gamma=0.1)\nI see that some baselines are set like this when setting the update step size of the learning rate. My understanding is that when the batchsize is small (gpu memory is insufficient, or other reasons), the update step size of the learning rate is very long, so the learning rate will not change during the training. I don't understand very well, can someone tell me the meaning behind this setup?</p>",
  "messages": [
    {
      "id": "716781",
      "postDate": "01/12/2020 08:36:00",
      "content": "<p>explrscheduler = lrscheduler.StepLR(optimizer, stepsize=max(nepochs, 10) * len(trainloader) // 3, gamma=0.1)\nI see that some baselines are set like this when setting the update step size of the learning rate. My understanding is that when the batchsize is small (gpu memory is insufficient, or other reasons), the update step size of the learning rate is very long, so the learning rate will not change during the training. I don't understand very well, can someone tell me the meaning behind this setup?</p>",
      "rawMarkdown": "explrscheduler = lrscheduler.StepLR(optimizer, stepsize=max(nepochs, 10) * len(trainloader) // 3, gamma=0.1)\nI see that some baselines are set like this when setting the update step size of the learning rate. My understanding is that when the batchsize is small (gpu memory is insufficient, or other reasons), the update step size of the learning rate is very long, so the learning rate will not change during the training. I don't understand very well, can someone tell me the meaning behind this setup?",
      "votes": null
    },
    {
      "id": "718886",
      "postDate": "01/14/2020 22:46:02",
      "content": "<p>Note that it uses len(trainloader) which is equivalent term for epoch, so the batchsize does not affect the lr update epoch-wise, which is the main unit used here. Think of lr = scheduler(epoch). The example u said just reduces lr by 0.1 every 3rd portion of the training e.g. epoch=12 then reduce lr by 0.1 factor every 4 epochs</p>",
      "rawMarkdown": "Note that it uses len(trainloader) which is equivalent term for epoch, so the batchsize does not affect the lr update epoch-wise, which is the main unit used here. Think of lr = scheduler(epoch). The example u said just reduces lr by 0.1 every 3rd portion of the training e.g. epoch=12 then reduce lr by 0.1 factor every 4 epochs",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 718886,
      "author_name": "joonl04",
      "author_url": "",
      "post_date": "01/14/2020 22:46:02",
      "content": "<p>Note that it uses len(trainloader) which is equivalent term for epoch, so the batchsize does not affect the lr update epoch-wise, which is the main unit used here. Think of lr = scheduler(epoch). The example u said just reduces lr by 0.1 every 3rd portion of the training e.g. epoch=12 then reduce lr by 0.1 factor every 4 epochs</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "716781": "explrscheduler = lrscheduler.StepLR(optimizer, stepsize=max(nepochs, 10) * len(trainloader) // 3, gamma=0.1)\nI see that some baselines are set like this when setting the update step size of the learning rate. My understanding is that when the batchsize is small (gpu memory is insufficient, or other reasons), the update step size of the learning rate is very long, so the learning rate will not change during the training. I don't understand very well, can someone tell me the meaning behind this setup?",
    "718886": "Note that it uses len(trainloader) which is equivalent term for epoch, so the batchsize does not affect the lr update epoch-wise, which is the main unit used here. Think of lr = scheduler(epoch). The example u said just reduces lr by 0.1 every 3rd portion of the training e.g. epoch=12 then reduce lr by 0.1 factor every 4 epochs"
  },
  "source": "meta"
}