{
  "id": 172218,
  "title": "Pytorch equivalent of Keras scheduler everyone uses?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/172218",
  "author_name": "Signal",
  "post_date": "2020-08-04T06:56:20.739000",
  "votes": 0,
  "comment_count": 10,
  "views": 0,
  "content": "<p>It seems most of the TF/Keras notebooks are using a similar or same scheduler, which looks like this:</p>\n\n<p>```\ndef get_lr_callback(cfg):\n    lr_start   = cfg['LR_START']\n    lr_max     = cfg['LR_MAX'] * strategy.num_replicas_in_sync\n    lr_min     = cfg['LR_MIN']\n    lr_ramp_ep = cfg['LR_RAMPUP_EPOCHS']\n    lr_sus_ep  = cfg['LR_SUSTAIN_EPOCHS']\n    lr_decay   = cfg['LR_EXP_DECAY']</p>\n\n<pre><code>def lrfn(epoch):\n    if epoch &amp;lt; lr_ramp_ep:\n        lr = (lr_max - lr_start) / lr_ramp_ep * epoch + lr_start\n\n    elif epoch &amp;lt; lr_ramp_ep + lr_sus_ep:\n        lr = lr_max\n\n    else:\n        lr = (lr_max - lr_min) * lr_decay**(epoch - lr_ramp_ep - lr_sus_ep) + lr_min\n\n    return lr\n\nlr_callback = tf.keras.callbacks.LearningRateScheduler(lrfn, verbose=False)\nreturn lr_callback\n</code></pre>\n\n<p>```</p>\n\n<p>Does anyone know if something similar is for PyTorch?  It looks like <a href=\"https://pytorch.org/docs/stable/optim.html\">OneCycleLR</a>.</p>\n\n<p>OneCycleLR is set up a bit differently, as you set a <code>MaxLR</code> and then can figure out what you want it to start at by figuring out a <code>div_factor</code>, as well as a <code>MinLR</code> by figuring out a <code>div_factor_final</code>.  It has pct_start which looks to be similar to <code>lr_ramp_ep</code>.  So similar in some ways.  But I don't believe there is any concept of \"sustain\", it just cycles up and down.</p>",
  "messages": [
    {
      "id": 957459,
      "postDate": "2020-08-04T10:26:22.967Z",
      "content": "<p><a href=\"https://github.com/ildoonet/pytorch-gradual-warmup-lr\">https://github.com/ildoonet/pytorch-gradual-warmup-lr</a>\nThis does not have sustain part but It seems good for me</p>",
      "rawMarkdown": "https://github.com/ildoonet/pytorch-gradual-warmup-lr\nThis does not have sustain part but It seems good for me",
      "votes": 1,
      "replies": [
        {
          "id": 958058,
          "postDate": "2020-08-04T18:22:14.957Z",
          "content": "<p>Thanks, I will give that a try.  I really have had the best results with ReduceLROnPlateau, which I see is the base of that scheduler.</p>",
          "rawMarkdown": "Thanks, I will give that a try.  I really have had the best results with ReduceLROnPlateau, which I see is the base of that scheduler."
        }
      ]
    },
    {
      "id": 957231,
      "postDate": "2020-08-04T06:56:20.740Z",
      "content": "<p>It seems most of the TF/Keras notebooks are using a similar or same scheduler, which looks like this:</p>\n\n<p>```\ndef get_lr_callback(cfg):\n    lr_start   = cfg['LR_START']\n    lr_max     = cfg['LR_MAX'] * strategy.num_replicas_in_sync\n    lr_min     = cfg['LR_MIN']\n    lr_ramp_ep = cfg['LR_RAMPUP_EPOCHS']\n    lr_sus_ep  = cfg['LR_SUSTAIN_EPOCHS']\n    lr_decay   = cfg['LR_EXP_DECAY']</p>\n\n<pre><code>def lrfn(epoch):\n    if epoch &amp;lt; lr_ramp_ep:\n        lr = (lr_max - lr_start) / lr_ramp_ep * epoch + lr_start\n\n    elif epoch &amp;lt; lr_ramp_ep + lr_sus_ep:\n        lr = lr_max\n\n    else:\n        lr = (lr_max - lr_min) * lr_decay**(epoch - lr_ramp_ep - lr_sus_ep) + lr_min\n\n    return lr\n\nlr_callback = tf.keras.callbacks.LearningRateScheduler(lrfn, verbose=False)\nreturn lr_callback\n</code></pre>\n\n<p>```</p>\n\n<p>Does anyone know if something similar is for PyTorch?  It looks like <a href=\"https://pytorch.org/docs/stable/optim.html\">OneCycleLR</a>.</p>\n\n<p>OneCycleLR is set up a bit differently, as you set a <code>MaxLR</code> and then can figure out what you want it to start at by figuring out a <code>div_factor</code>, as well as a <code>MinLR</code> by figuring out a <code>div_factor_final</code>.  It has pct_start which looks to be similar to <code>lr_ramp_ep</code>.  So similar in some ways.  But I don't believe there is any concept of \"sustain\", it just cycles up and down.</p>",
      "rawMarkdown": "It seems most of the TF/Keras notebooks are using a similar or same scheduler, which looks like this:\n\n```\ndef get_lr_callback(cfg):\n    lr_start   = cfg['LR_START']\n    lr_max     = cfg['LR_MAX'] * strategy.num_replicas_in_sync\n    lr_min     = cfg['LR_MIN']\n    lr_ramp_ep = cfg['LR_RAMPUP_EPOCHS']\n    lr_sus_ep  = cfg['LR_SUSTAIN_EPOCHS']\n    lr_decay   = cfg['LR_EXP_DECAY']\n   \n    def lrfn(epoch):\n        if epoch &lt; lr_ramp_ep:\n            lr = (lr_max - lr_start) / lr_ramp_ep * epoch + lr_start\n            \n        elif epoch &lt; lr_ramp_ep + lr_sus_ep:\n            lr = lr_max\n            \n        else:\n            lr = (lr_max - lr_min) * lr_decay**(epoch - lr_ramp_ep - lr_sus_ep) + lr_min\n            \n        return lr\n\n    lr_callback = tf.keras.callbacks.LearningRateScheduler(lrfn, verbose=False)\n    return lr_callback\n```\n\nDoes anyone know if something similar is for PyTorch?  It looks like [OneCycleLR](https://pytorch.org/docs/stable/optim.html).\n\nOneCycleLR is set up a bit differently, as you set a `MaxLR` and then can figure out what you want it to start at by figuring out a `div_factor`, as well as a `MinLR` by figuring out a `div_factor_final`.  It has pct_start which looks to be similar to `lr_ramp_ep`.  So similar in some ways.  But I don't believe there is any concept of \"sustain\", it just cycles up and down."
    },
    {
      "id": 957732,
      "postDate": "2020-08-04T14:16:06.430Z",
      "content": "<p>I created a kind of custom scheduler where I can switch between a custom implementation, implemented in CustomSchedulerLR (you can code whatever logic you need here), and standard ReduceLROnPlateau.</p>\n\n<p>```\nclass CustomSchedulerLR:\n  def lrfn(self, epoch): <br>\n    if epoch &lt; self.lr_ramp_ep:\n        lr = (self.lr_max - self.lr_start) / self.lr_ramp_ep * epoch + self.lr_start <br>\n    elif epoch &lt; self.lr_ramp_ep + self.lr_sus_ep:\n        lr = self.lr_max\n    else:\n        lr = (self.lr_max - self.lr_min) * self.lr_decay**(epoch - self.lr_ramp_ep - self.lr_sus_ep) + self.lr_min</p>\n\n<pre><code>return lr   \n</code></pre>\n\n<p>def <strong>init</strong>(self, optimizer, epoch, batch_size):\n    self.lr_start = 0.000005\n    self.lr_min = 0.000001\n    self.lr_ramp_ep = 5\n    self.lr_sus_ep = 0\n    self.lr_decay = 0.8 <br>\n    self.lr_max = 0.00000125 * 8 * batch_size\n    self.optimizer = optimizer</p>\n\n<pre><code>lr = self.lrfn(epoch)    \nfor param_group in self.optimizer.param_groups:\n    param_group['lr'] = lr\n</code></pre>\n\n<p>def step(self, epoch):\n    lr = self.lrfn(epoch) <br>\n    for param_group in self.optimizer.param_groups:\n        param_group['lr'] = lr\n```</p>\n\n<p><code>\nif SCHEDULER_NAME == 'CustomLR': \n        scheduler = CustomSchedulerLR(optimizer=optimizer, epoch=0, batch_size=train_bs)\n    elif SCHEDULER_NAME == 'ReduceLROnPlateau':\n        scheduler = torch.optim.lr_scheduler.ReduceLROnPlateau(\n          optimizer,\n          patience=3, \n          threshold=0.001,\n          verbose=True,\n          mode=\"max\"\n        )\n    else:\n        scheduler = CustomSchedulerLR(optimizer=optimizer, epoch=0, batch_size=train_bs)\n</code></p>\n\n<p>```\nfor epoch in range(epochs):\n        &lt;&gt;</p>\n\n<pre><code>    if SCHEDULER_NAME == 'CustomLR':\n        scheduler.step(epoch+1)\n    elif SCHEDULER_NAME == 'ReduceLROnPlateau':\n        scheduler.step(auc)\n    else:\n        scheduler.step(epoch) \n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "I created a kind of custom scheduler where I can switch between a custom implementation, implemented in CustomSchedulerLR (you can code whatever logic you need here), and standard ReduceLROnPlateau.\n\n```\nclass CustomSchedulerLR:\n  def lrfn(self, epoch):      \n    if epoch &lt; self.lr_ramp_ep:\n        lr = (self.lr_max - self.lr_start) / self.lr_ramp_ep * epoch + self.lr_start           \n    elif epoch &lt; self.lr_ramp_ep + self.lr_sus_ep:\n        lr = self.lr_max\n    else:\n        lr = (self.lr_max - self.lr_min) * self.lr_decay**(epoch - self.lr_ramp_ep - self.lr_sus_ep) + self.lr_min\n            \n    return lr   \n\n  def __init__(self, optimizer, epoch, batch_size):\n    self.lr_start = 0.000005\n    self.lr_min = 0.000001\n    self.lr_ramp_ep = 5\n    self.lr_sus_ep = 0\n    self.lr_decay = 0.8  \n    self.lr_max = 0.00000125 * 8 * batch_size\n    self.optimizer = optimizer\n\n    lr = self.lrfn(epoch)    \n    for param_group in self.optimizer.param_groups:\n        param_group['lr'] = lr\n\n  def step(self, epoch):\n    lr = self.lrfn(epoch)  \n    for param_group in self.optimizer.param_groups:\n        param_group['lr'] = lr\n```\n\n```\nif SCHEDULER_NAME == 'CustomLR': \n        scheduler = CustomSchedulerLR(optimizer=optimizer, epoch=0, batch_size=train_bs)\n    elif SCHEDULER_NAME == 'ReduceLROnPlateau':\n        scheduler = torch.optim.lr_scheduler.ReduceLROnPlateau(\n          optimizer,\n          patience=3, \n          threshold=0.001,\n          verbose=True,\n          mode=\"max\"\n        )\n    else:\n        scheduler = CustomSchedulerLR(optimizer=optimizer, epoch=0, batch_size=train_bs)\n```\n\n```\nfor epoch in range(epochs):\n        &lt;&gt;\n\n        if SCHEDULER_NAME == 'CustomLR':\n            scheduler.step(epoch+1)\n        elif SCHEDULER_NAME == 'ReduceLROnPlateau':\n            scheduler.step(auc)\n        else:\n            scheduler.step(epoch) \n```",
      "replies": [
        {
          "id": 958080,
          "postDate": "2020-08-04T18:50:13.657Z",
          "content": "<p>What about this part?\n<code>self.lr_max = 0.00000125 * 8 * batch_size\n</code>\nThe 8 is the number of replicas right? which in TF is basically number of GPU's when doing distributed training correct?  And so would batch_size be the size of the batch on each GPU or the total size of the batch?  Example, I run batch size 220 in training, which ends up putting 55 on each of my 4 GPUs</p>",
          "rawMarkdown": "What about this part?\n`    self.lr_max = 0.00000125 * 8 * batch_size\n`\nThe 8 is the number of replicas right? which in TF is basically number of GPU's when doing distributed training correct?  And so would batch_size be the size of the batch on each GPU or the total size of the batch?  Example, I run batch size 220 in training, which ends up putting 55 on each of my 4 GPUs"
        },
        {
          "id": 958195,
          "postDate": "2020-08-04T20:41:40.777Z",
          "content": "<p>In original implementation  - yes, a number of replicas. In case of one GPU you can play out with this parameter (8,4,2 or 1) and see which value is better for your model. Speaking of batch size, I assume it should be \"the size of the batch on each GPU\" but not the total.</p>",
          "rawMarkdown": "In original implementation  - yes, a number of replicas. In case of one GPU you can play out with this parameter (8,4,2 or 1) and see which value is better for your model. Speaking of batch size, I assume it should be \"the size of the batch on each GPU\" but not the total."
        },
        {
          "id": 958224,
          "postDate": "2020-08-04T21:13:09.357Z",
          "content": "<p>Yes, batch size is a scaling factor and there is much research to support that.  I am going to run some tests using these parameters everyone seems to like and see how it does vs what I normally do.</p>",
          "rawMarkdown": "Yes, batch size is a scaling factor and there is much research to support that.  I am going to run some tests using these parameters everyone seems to like and see how it does vs what I normally do."
        }
      ]
    },
    {
      "id": 957345,
      "postDate": "2020-08-04T08:29:26.760Z",
      "content": "<p>Pytorch has torch.optim.lr_scheduler.LambdaLR which can be used to wrap any custom function to generate LR.</p>\n\n<p>Inspired by the discussion at <a href=\"https://github.com/PyTorchLightning/pytorch-lightning/issues/328\">https://github.com/PyTorchLightning/pytorch-lightning/issues/328</a>, I Tried using Lambda in Pytorch Lightning to implement above mentioned code, but it didn't work (throwing warning messages). </p>\n\n<p>Following is my implementation. If You are using Pytorch (without lightning), try implementing  and see if it works for You -</p>\n\n<ul>\n<li><p>def configure_optimizers(self):</p>\n\n<pre><code>optimizer = torch.optim.Adam(self.parameters(), lr=lr)\n\ndef lrfn(epoch):\n    lr_start   = 0.000003\n    lr_max     = 0.00003\n    lr_min     = 0.000001\n    lr_ramp_ep = 5\n    lr_sus_ep  = 2\n    lr_decay   = 0.8\n\n    if epoch &lt; lr_ramp_ep:\n        lr = (lr_max - lr_start) / lr_ramp_ep * epoch + lr_start\n\n    elif epoch &lt; lr_ramp_ep + lr_sus_ep:\n        lr = lr_max\n\n    else:\n        lr = (lr_max - lr_min) * lr_decay**(epoch - lr_ramp_ep - lr_sus_ep) + lr_min\n\n    return lr\n\nscheduler = torch.optim.lr_scheduler.LambdaLR(optimizer, lr_lambda=lrfn)\n\nreturn [optimizer],[scheduler]*\n</code></pre></li>\n</ul>",
      "rawMarkdown": "Pytorch has torch.optim.lr_scheduler.LambdaLR which can be used to wrap any custom function to generate LR.\n\nInspired by the discussion at https://github.com/PyTorchLightning/pytorch-lightning/issues/328, I Tried using Lambda in Pytorch Lightning to implement above mentioned code, but it didn't work (throwing warning messages). \n\nFollowing is my implementation. If You are using Pytorch (without lightning), try implementing  and see if it works for You -\n\n*   def configure_optimizers(self):\n\n        optimizer = torch.optim.Adam(self.parameters(), lr=lr)\n        \n        def lrfn(epoch):\n            lr_start   = 0.000003\n            lr_max     = 0.00003\n            lr_min     = 0.000001\n            lr_ramp_ep = 5\n            lr_sus_ep  = 2\n            lr_decay   = 0.8\n\n            if epoch &lt; lr_ramp_ep:\n                lr = (lr_max - lr_start) / lr_ramp_ep * epoch + lr_start\n\n            elif epoch &lt; lr_ramp_ep + lr_sus_ep:\n                lr = lr_max\n\n            else:\n                lr = (lr_max - lr_min) * lr_decay**(epoch - lr_ramp_ep - lr_sus_ep) + lr_min\n\n            return lr\n        \n        scheduler = torch.optim.lr_scheduler.LambdaLR(optimizer, lr_lambda=lrfn)\n\n        return [optimizer],[scheduler]*\n\n",
      "replies": [
        {
          "id": 957376,
          "postDate": "2020-08-04T08:54:34.913Z",
          "content": "<p>In TF/Keras using that scheduler, they all seem to use such small values.  Typically the start is .000005, max is around .000001 - .000002, min is around .00000125.</p>\n\n<p>With PyTorch in this competition m my loss seems to drop fastest with around ..000125 - .001</p>",
          "rawMarkdown": "In TF/Keras using that scheduler, they all seem to use such small values.  Typically the start is .000005, max is around .000001 - .000002, min is around .00000125.\n\nWith PyTorch in this competition m my loss seems to drop fastest with around ..000125 - .001\n"
        },
        {
          "id": 958087,
          "postDate": "2020-08-04T18:57:29.280Z",
          "content": "<p>Just looking at some of the TF code it looks to be the device batch size not the global batch size</p>",
          "rawMarkdown": "Just looking at some of the TF code it looks to be the device batch size not the global batch size"
        }
      ]
    },
    {
      "id": 957715,
      "postDate": "2020-08-04T14:05:33.310Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 957459,
      "author_name": "cp_t2",
      "author_url": "",
      "post_date": "2020-08-04T10:26:22.967000",
      "content": "<p><a href=\"https://github.com/ildoonet/pytorch-gradual-warmup-lr\">https://github.com/ildoonet/pytorch-gradual-warmup-lr</a>\nThis does not have sustain part but It seems good for me</p>",
      "votes": 1,
      "replies": [
        {
          "id": 958058,
          "author_name": "Signal",
          "author_url": "",
          "post_date": "2020-08-04T18:22:14.957000",
          "content": "<p>Thanks, I will give that a try.  I really have had the best results with ReduceLROnPlateau, which I see is the base of that scheduler.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 957732,
      "author_name": "ELEVEN",
      "author_url": "",
      "post_date": "2020-08-04T14:16:06.430000",
      "content": "<p>I created a kind of custom scheduler where I can switch between a custom implementation, implemented in CustomSchedulerLR (you can code whatever logic you need here), and standard ReduceLROnPlateau.</p>\n\n<p>```\nclass CustomSchedulerLR:\n  def lrfn(self, epoch): <br>\n    if epoch &lt; self.lr_ramp_ep:\n        lr = (self.lr_max - self.lr_start) / self.lr_ramp_ep * epoch + self.lr_start <br>\n    elif epoch &lt; self.lr_ramp_ep + self.lr_sus_ep:\n        lr = self.lr_max\n    else:\n        lr = (self.lr_max - self.lr_min) * self.lr_decay**(epoch - self.lr_ramp_ep - self.lr_sus_ep) + self.lr_min</p>\n\n<pre><code>return lr   \n</code></pre>\n\n<p>def <strong>init</strong>(self, optimizer, epoch, batch_size):\n    self.lr_start = 0.000005\n    self.lr_min = 0.000001\n    self.lr_ramp_ep = 5\n    self.lr_sus_ep = 0\n    self.lr_decay = 0.8 <br>\n    self.lr_max = 0.00000125 * 8 * batch_size\n    self.optimizer = optimizer</p>\n\n<pre><code>lr = self.lrfn(epoch)    \nfor param_group in self.optimizer.param_groups:\n    param_group['lr'] = lr\n</code></pre>\n\n<p>def step(self, epoch):\n    lr = self.lrfn(epoch) <br>\n    for param_group in self.optimizer.param_groups:\n        param_group['lr'] = lr\n```</p>\n\n<p><code>\nif SCHEDULER_NAME == 'CustomLR': \n        scheduler = CustomSchedulerLR(optimizer=optimizer, epoch=0, batch_size=train_bs)\n    elif SCHEDULER_NAME == 'ReduceLROnPlateau':\n        scheduler = torch.optim.lr_scheduler.ReduceLROnPlateau(\n          optimizer,\n          patience=3, \n          threshold=0.001,\n          verbose=True,\n          mode=\"max\"\n        )\n    else:\n        scheduler = CustomSchedulerLR(optimizer=optimizer, epoch=0, batch_size=train_bs)\n</code></p>\n\n<p>```\nfor epoch in range(epochs):\n        &lt;&gt;</p>\n\n<pre><code>    if SCHEDULER_NAME == 'CustomLR':\n        scheduler.step(epoch+1)\n    elif SCHEDULER_NAME == 'ReduceLROnPlateau':\n        scheduler.step(auc)\n    else:\n        scheduler.step(epoch) \n</code></pre>\n\n<p>```</p>",
      "votes": 0,
      "replies": [
        {
          "id": 958080,
          "author_name": "Signal",
          "author_url": "",
          "post_date": "2020-08-04T18:50:13.657000",
          "content": "<p>What about this part?\n<code>self.lr_max = 0.00000125 * 8 * batch_size\n</code>\nThe 8 is the number of replicas right? which in TF is basically number of GPU's when doing distributed training correct?  And so would batch_size be the size of the batch on each GPU or the total size of the batch?  Example, I run batch size 220 in training, which ends up putting 55 on each of my 4 GPUs</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 958195,
          "author_name": "ELEVEN",
          "author_url": "",
          "post_date": "2020-08-04T20:41:40.777000",
          "content": "<p>In original implementation  - yes, a number of replicas. In case of one GPU you can play out with this parameter (8,4,2 or 1) and see which value is better for your model. Speaking of batch size, I assume it should be \"the size of the batch on each GPU\" but not the total.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 958224,
          "author_name": "Signal",
          "author_url": "",
          "post_date": "2020-08-04T21:13:09.357000",
          "content": "<p>Yes, batch size is a scaling factor and there is much research to support that.  I am going to run some tests using these parameters everyone seems to like and see how it does vs what I normally do.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 957345,
      "author_name": "OldMonk",
      "author_url": "",
      "post_date": "2020-08-04T08:29:26.760000",
      "content": "<p>Pytorch has torch.optim.lr_scheduler.LambdaLR which can be used to wrap any custom function to generate LR.</p>\n\n<p>Inspired by the discussion at <a href=\"https://github.com/PyTorchLightning/pytorch-lightning/issues/328\">https://github.com/PyTorchLightning/pytorch-lightning/issues/328</a>, I Tried using Lambda in Pytorch Lightning to implement above mentioned code, but it didn't work (throwing warning messages). </p>\n\n<p>Following is my implementation. If You are using Pytorch (without lightning), try implementing  and see if it works for You -</p>\n\n<ul>\n<li><p>def configure_optimizers(self):</p>\n\n<pre><code>optimizer = torch.optim.Adam(self.parameters(), lr=lr)\n\ndef lrfn(epoch):\n    lr_start   = 0.000003\n    lr_max     = 0.00003\n    lr_min     = 0.000001\n    lr_ramp_ep = 5\n    lr_sus_ep  = 2\n    lr_decay   = 0.8\n\n    if epoch &lt; lr_ramp_ep:\n        lr = (lr_max - lr_start) / lr_ramp_ep * epoch + lr_start\n\n    elif epoch &lt; lr_ramp_ep + lr_sus_ep:\n        lr = lr_max\n\n    else:\n        lr = (lr_max - lr_min) * lr_decay**(epoch - lr_ramp_ep - lr_sus_ep) + lr_min\n\n    return lr\n\nscheduler = torch.optim.lr_scheduler.LambdaLR(optimizer, lr_lambda=lrfn)\n\nreturn [optimizer],[scheduler]*\n</code></pre></li>\n</ul>",
      "votes": 0,
      "replies": [
        {
          "id": 957376,
          "author_name": "Signal",
          "author_url": "",
          "post_date": "2020-08-04T08:54:34.913000",
          "content": "<p>In TF/Keras using that scheduler, they all seem to use such small values.  Typically the start is .000005, max is around .000001 - .000002, min is around .00000125.</p>\n\n<p>With PyTorch in this competition m my loss seems to drop fastest with around ..000125 - .001</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 958087,
          "author_name": "Signal",
          "author_url": "",
          "post_date": "2020-08-04T18:57:29.280000",
          "content": "<p>Just looking at some of the TF code it looks to be the device batch size not the global batch size</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 957715,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-04T14:05:33.310000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "957459": "https://github.com/ildoonet/pytorch-gradual-warmup-lr\nThis does not have sustain part but It seems good for me",
    "957231": "It seems most of the TF/Keras notebooks are using a similar or same scheduler, which looks like this:\n\n```\ndef get_lr_callback(cfg):\n    lr_start   = cfg['LR_START']\n    lr_max     = cfg['LR_MAX'] * strategy.num_replicas_in_sync\n    lr_min     = cfg['LR_MIN']\n    lr_ramp_ep = cfg['LR_RAMPUP_EPOCHS']\n    lr_sus_ep  = cfg['LR_SUSTAIN_EPOCHS']\n    lr_decay   = cfg['LR_EXP_DECAY']\n   \n    def lrfn(epoch):\n        if epoch &lt; lr_ramp_ep:\n            lr = (lr_max - lr_start) / lr_ramp_ep * epoch + lr_start\n            \n        elif epoch &lt; lr_ramp_ep + lr_sus_ep:\n            lr = lr_max\n            \n        else:\n            lr = (lr_max - lr_min) * lr_decay**(epoch - lr_ramp_ep - lr_sus_ep) + lr_min\n            \n        return lr\n\n    lr_callback = tf.keras.callbacks.LearningRateScheduler(lrfn, verbose=False)\n    return lr_callback\n```\n\nDoes anyone know if something similar is for PyTorch?  It looks like [OneCycleLR](https://pytorch.org/docs/stable/optim.html).\n\nOneCycleLR is set up a bit differently, as you set a `MaxLR` and then can figure out what you want it to start at by figuring out a `div_factor`, as well as a `MinLR` by figuring out a `div_factor_final`.  It has pct_start which looks to be similar to `lr_ramp_ep`.  So similar in some ways.  But I don't believe there is any concept of \"sustain\", it just cycles up and down.",
    "957732": "I created a kind of custom scheduler where I can switch between a custom implementation, implemented in CustomSchedulerLR (you can code whatever logic you need here), and standard ReduceLROnPlateau.\n\n```\nclass CustomSchedulerLR:\n  def lrfn(self, epoch):      \n    if epoch &lt; self.lr_ramp_ep:\n        lr = (self.lr_max - self.lr_start) / self.lr_ramp_ep * epoch + self.lr_start           \n    elif epoch &lt; self.lr_ramp_ep + self.lr_sus_ep:\n        lr = self.lr_max\n    else:\n        lr = (self.lr_max - self.lr_min) * self.lr_decay**(epoch - self.lr_ramp_ep - self.lr_sus_ep) + self.lr_min\n            \n    return lr   \n\n  def __init__(self, optimizer, epoch, batch_size):\n    self.lr_start = 0.000005\n    self.lr_min = 0.000001\n    self.lr_ramp_ep = 5\n    self.lr_sus_ep = 0\n    self.lr_decay = 0.8  \n    self.lr_max = 0.00000125 * 8 * batch_size\n    self.optimizer = optimizer\n\n    lr = self.lrfn(epoch)    \n    for param_group in self.optimizer.param_groups:\n        param_group['lr'] = lr\n\n  def step(self, epoch):\n    lr = self.lrfn(epoch)  \n    for param_group in self.optimizer.param_groups:\n        param_group['lr'] = lr\n```\n\n```\nif SCHEDULER_NAME == 'CustomLR': \n        scheduler = CustomSchedulerLR(optimizer=optimizer, epoch=0, batch_size=train_bs)\n    elif SCHEDULER_NAME == 'ReduceLROnPlateau':\n        scheduler = torch.optim.lr_scheduler.ReduceLROnPlateau(\n          optimizer,\n          patience=3, \n          threshold=0.001,\n          verbose=True,\n          mode=\"max\"\n        )\n    else:\n        scheduler = CustomSchedulerLR(optimizer=optimizer, epoch=0, batch_size=train_bs)\n```\n\n```\nfor epoch in range(epochs):\n        &lt;&gt;\n\n        if SCHEDULER_NAME == 'CustomLR':\n            scheduler.step(epoch+1)\n        elif SCHEDULER_NAME == 'ReduceLROnPlateau':\n            scheduler.step(auc)\n        else:\n            scheduler.step(epoch) \n```",
    "957345": "Pytorch has torch.optim.lr_scheduler.LambdaLR which can be used to wrap any custom function to generate LR.\n\nInspired by the discussion at https://github.com/PyTorchLightning/pytorch-lightning/issues/328, I Tried using Lambda in Pytorch Lightning to implement above mentioned code, but it didn't work (throwing warning messages). \n\nFollowing is my implementation. If You are using Pytorch (without lightning), try implementing  and see if it works for You -\n\n*   def configure_optimizers(self):\n\n        optimizer = torch.optim.Adam(self.parameters(), lr=lr)\n        \n        def lrfn(epoch):\n            lr_start   = 0.000003\n            lr_max     = 0.00003\n            lr_min     = 0.000001\n            lr_ramp_ep = 5\n            lr_sus_ep  = 2\n            lr_decay   = 0.8\n\n            if epoch &lt; lr_ramp_ep:\n                lr = (lr_max - lr_start) / lr_ramp_ep * epoch + lr_start\n\n            elif epoch &lt; lr_ramp_ep + lr_sus_ep:\n                lr = lr_max\n\n            else:\n                lr = (lr_max - lr_min) * lr_decay**(epoch - lr_ramp_ep - lr_sus_ep) + lr_min\n\n            return lr\n        \n        scheduler = torch.optim.lr_scheduler.LambdaLR(optimizer, lr_lambda=lrfn)\n\n        return [optimizer],[scheduler]*\n\n",
    "957715": ""
  }
}