{
  "id": 93156,
  "title": "Learning rate policy",
  "url": "/competitions/imet-2019-fgvc6/discussion/93156",
  "author_name": "",
  "post_date": "2019-05-23T19:36:53.550892700Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Which learning rate policy have you guys been using?\nI've been trying to go with ReduceLROnPlateau, but I couldn't get any significant improvement beyond my initial learning rate :( \nI'm using Adam with initial lr 3e-4, batch size 128. Augmentation - hflip and gaussian noise. Resize to 320x320. \nAny tips?</p>",
  "messages": [
    {
      "id": "536008",
      "postDate": "05/23/2019 19:36:53",
      "content": "<p>Which learning rate policy have you guys been using?\nI've been trying to go with ReduceLROnPlateau, but I couldn't get any significant improvement beyond my initial learning rate :( \nI'm using Adam with initial lr 3e-4, batch size 128. Augmentation - hflip and gaussian noise. Resize to 320x320. \nAny tips?</p>",
      "rawMarkdown": "Which learning rate policy have you guys been using?\nI've been trying to go with ReduceLROnPlateau, but I couldn't get any significant improvement beyond my initial learning rate :( \nI'm using Adam with initial lr 3e-4, batch size 128. Augmentation - hflip and gaussian noise. Resize to 320x320. \nAny tips?",
      "votes": null
    },
    {
      "id": "536112",
      "postDate": "05/24/2019 01:32:46",
      "content": "<p>I think cyclic learning rate will have a much faster convergence speed than  ReduceLROnPlateau\nYou can find some code on Github</p>",
      "rawMarkdown": "I think cyclic learning rate will have a much faster convergence speed than  ReduceLROnPlateau\nYou can find some code on Github",
      "votes": null
    },
    {
      "id": "536852",
      "postDate": "05/25/2019 12:39:52",
      "content": "<p>I use ReduceLROnPlateau with Adam and init learning rate: 1e-4. Generally it has an advantage over a static learning, but it overfits after 1-2 lr decays. \nI also think, that Cyclic Learning rates will do a better job in terms of convergence, and warm restarts could take the model out of the local minimas. (<a href=\"https://arxiv.org/abs/1608.03983\">https://arxiv.org/abs/1608.03983</a>)</p>",
      "rawMarkdown": "I use ReduceLROnPlateau with Adam and init learning rate: 1e-4. Generally it has an advantage over a static learning, but it overfits after 1-2 lr decays. \nI also think, that Cyclic Learning rates will do a better job in terms of convergence, and warm restarts could take the model out of the local minimas. (https://arxiv.org/abs/1608.03983)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 536112,
      "author_name": "dilapsky",
      "author_url": "",
      "post_date": "05/24/2019 01:32:46",
      "content": "<p>I think cyclic learning rate will have a much faster convergence speed than  ReduceLROnPlateau\nYou can find some code on Github</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 536852,
      "author_name": "deyury",
      "author_url": "",
      "post_date": "05/25/2019 12:39:52",
      "content": "<p>I use ReduceLROnPlateau with Adam and init learning rate: 1e-4. Generally it has an advantage over a static learning, but it overfits after 1-2 lr decays. \nI also think, that Cyclic Learning rates will do a better job in terms of convergence, and warm restarts could take the model out of the local minimas. (<a href=\"https://arxiv.org/abs/1608.03983\">https://arxiv.org/abs/1608.03983</a>)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "536008": "Which learning rate policy have you guys been using?\nI've been trying to go with ReduceLROnPlateau, but I couldn't get any significant improvement beyond my initial learning rate :( \nI'm using Adam with initial lr 3e-4, batch size 128. Augmentation - hflip and gaussian noise. Resize to 320x320. \nAny tips?",
    "536112": "I think cyclic learning rate will have a much faster convergence speed than  ReduceLROnPlateau\nYou can find some code on Github",
    "536852": "I use ReduceLROnPlateau with Adam and init learning rate: 1e-4. Generally it has an advantage over a static learning, but it overfits after 1-2 lr decays. \nI also think, that Cyclic Learning rates will do a better job in terms of convergence, and warm restarts could take the model out of the local minimas. (https://arxiv.org/abs/1608.03983)"
  },
  "source": "meta"
}