{
  "id": 247574,
  "title": "[Question] How do we resume training by using the last LR?",
  "url": "/competitions/seti-breakthrough-listen/discussion/247574",
  "author_name": "gao-hongnan",
  "post_date": "2021-06-20T07:59:33.439000",
  "votes": 7,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Now, I am unsure if this has been answered, I remember my buddy telling me that anecdotally, one can \"reset\" the LR a little when one resumes training from a model's checkpoint.</p>\n<p>It is a bit counter-intuitive to me, but say I trained a model for 16 epochs, with a custom scheduler, say <code>OneCycleLr</code> + <code>Adam</code> or something, then when you resume training, should we reset the initial learning rate? I would think resetting is counter-intuitive as the purpose of the learning rate is to tune it such that your model can slowly converge to the minima (global one if the function is convex). </p>\n<p>So the question is: </p>\n<ol>\n<li><p>Should I reset the learning rate when resume training, if yes, reset to what?</p></li>\n<li><p>If we should not reset, what is a good way to \"extract\" the last learning rate, as some scheduler depends on factors like epochs…</p></li>\n</ol>",
  "messages": [
    {
      "id": 1358199,
      "postDate": "2021-06-20T09:05:36.080Z",
      "content": "<p>You have to save your optimizer parameters as you save your model weights when you checkpoint.</p>\n<p>Then when you resume you load them.</p>\n<p>For instance here are functions I use to save and restore checkpoints with pytorch.  Of course you need to adjust it to your code.  It is important that you create your model, optimizer, etc in the load checkpoint function exactly as you create them in your training loop.  The scaler is only useful if you use mixed precision (torch.cuda.amp).</p>\n<pre><code>def save_checkpoint(model, optimizer, scheduler, scaler, epoch, fold, seed, fname=fname):\n    checkpoint = {\n        'model': model.state_dict(),\n        'optimizer': optimizer.state_dict(),\n        'scheduler': scheduler.state_dict(),\n        'scaler': scaler.state_dict(),\n        'epoch': epoch,\n        'fold':fold,\n        'seed':seed,\n        }\n    torch.save(checkpoint, '../checkpoints/%s/%s_%d_%d.pt' % (fname, fname, fold, seed))\n\ndef load_checkpoint(fold, seed, fname):\n    model = create_model().to(device)\n    optimizer = optimizer = torch.optim.Adam(model.parameters(), lr=MAX_LR)\n    scheduler = torch.optim.lr_scheduler.OneCycleLR(optimizer=optimizer, \n                                              pct_start=PCT_START, \n                                              div_factor=DIV_FACTOR \n                                              max_lr=MAX_LR, \n                                              epochs=EPOCHS, \n                                              steps_per_epoch=int(np.ceil(len(train_data_loader)/GRADIENT_ACCUMULATION)))\n    scaler = GradScaler()\n    checkpoint = torch.load('../checkpoints/%s/%s_%d_%d.pt' % (fname, fname, fold, seed))\n    model.load_state_dict(checkpoint['model'])\n    optimizer.load_state_dict(checkpoint['optimizer'])\n    scheduler.load_state_dict(checkpoint['scheduler'])\n    scaler.load_state_dict(checkpoint['scaler'])\n    return model, optimizer, scheduler, scaler, epoch\n</code></pre>",
      "rawMarkdown": "You have to save your optimizer parameters as you save your model weights when you checkpoint.\n\nThen when you resume you load them.\n\nFor instance here are functions I use to save and restore checkpoints with pytorch.  Of course you need to adjust it to your code.  It is important that you create your model, optimizer, etc in the load checkpoint function exactly as you create them in your training loop.  The scaler is only useful if you use mixed precision (torch.cuda.amp).\n\n```\n\ndef save_checkpoint(model, optimizer, scheduler, scaler, epoch, fold, seed, fname=fname):\n    checkpoint = {\n        'model': model.state_dict(),\n        'optimizer': optimizer.state_dict(),\n        'scheduler': scheduler.state_dict(),\n        'scaler': scaler.state_dict(),\n        'epoch': epoch,\n        'fold':fold,\n        'seed':seed,\n        }\n    torch.save(checkpoint, '../checkpoints/%s/%s_%d_%d.pt' % (fname, fname, fold, seed))\n\ndef load_checkpoint(fold, seed, fname):\n    model = create_model().to(device)\n    optimizer = optimizer = torch.optim.Adam(model.parameters(), lr=MAX_LR)\n    scheduler = torch.optim.lr_scheduler.OneCycleLR(optimizer=optimizer, \n                                              pct_start=PCT_START, \n                                              div_factor=DIV_FACTOR \n                                              max_lr=MAX_LR, \n                                              epochs=EPOCHS, \n                                              steps_per_epoch=int(np.ceil(len(train_data_loader)/GRADIENT_ACCUMULATION)))\n    scaler = GradScaler()\n    checkpoint = torch.load('../checkpoints/%s/%s_%d_%d.pt' % (fname, fname, fold, seed))\n    model.load_state_dict(checkpoint['model'])\n    optimizer.load_state_dict(checkpoint['optimizer'])\n    scheduler.load_state_dict(checkpoint['scheduler'])\n    scaler.load_state_dict(checkpoint['scaler'])\n    return model, optimizer, scheduler, scaler, epoch\n\n```",
      "votes": 18,
      "replies": [
        {
          "id": 1358237,
          "postDate": "2021-06-20T09:36:09.393Z",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> Thanks, this clears the air, so we still resume in an almost verbatim manner when we resume training. I got it now :)</p>",
          "rawMarkdown": "@cpmpml Thanks, this clears the air, so we still resume in an almost verbatim manner when we resume training. I got it now :)",
          "votes": 1
        },
        {
          "id": 1358809,
          "postDate": "2021-06-20T18:50:57.543Z",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> Thanks, quite useful. One question on the seed, we would need to checkpoint the last state of random generator right? so when we resume training fresh random samples are created and not what were already created in the previous epochs (using same seed again).</p>",
          "rawMarkdown": "@cpmpml Thanks, quite useful. One question on the seed, we would need to checkpoint the last state of random generator right? so when we resume training fresh random samples are created and not what were already created in the previous epochs (using same seed again).",
          "votes": 3
        },
        {
          "id": 1358863,
          "postDate": "2021-06-20T20:36:54.110Z",
          "content": "<p><a href=\"https://www.kaggle.com/ankitsajwan\" target=\"_blank\">@ankitsajwan</a> I guess you're right.  Initializing the pytorch random generator would ensure path would be even closer to what it would have been without the restart.</p>",
          "rawMarkdown": "@ankitsajwan I guess you're right.  Initializing the pytorch random generator would ensure path would be even closer to what it would have been without the restart.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1358123,
      "postDate": "2021-06-20T07:59:33.440Z",
      "content": "<p>Now, I am unsure if this has been answered, I remember my buddy telling me that anecdotally, one can \"reset\" the LR a little when one resumes training from a model's checkpoint.</p>\n<p>It is a bit counter-intuitive to me, but say I trained a model for 16 epochs, with a custom scheduler, say <code>OneCycleLr</code> + <code>Adam</code> or something, then when you resume training, should we reset the initial learning rate? I would think resetting is counter-intuitive as the purpose of the learning rate is to tune it such that your model can slowly converge to the minima (global one if the function is convex). </p>\n<p>So the question is: </p>\n<ol>\n<li><p>Should I reset the learning rate when resume training, if yes, reset to what?</p></li>\n<li><p>If we should not reset, what is a good way to \"extract\" the last learning rate, as some scheduler depends on factors like epochs…</p></li>\n</ol>",
      "rawMarkdown": "Now, I am unsure if this has been answered, I remember my buddy telling me that anecdotally, one can \"reset\" the LR a little when one resumes training from a model's checkpoint.\n\nIt is a bit counter-intuitive to me, but say I trained a model for 16 epochs, with a custom scheduler, say `OneCycleLr` + `Adam` or something, then when you resume training, should we reset the initial learning rate? I would think resetting is counter-intuitive as the purpose of the learning rate is to tune it such that your model can slowly converge to the minima (global one if the function is convex). \n\nSo the question is: \n\n1. Should I reset the learning rate when resume training, if yes, reset to what?\n\n2. If we should not reset, what is a good way to \"extract\" the last learning rate, as some scheduler depends on factors like epochs...\n",
      "votes": 7
    },
    {
      "id": 1358255,
      "postDate": "2021-06-20T10:01:22.527Z",
      "content": "<p>i hope this discussion post of <a href=\"https://www.kaggle.com/raghaw\" target=\"_blank\">@raghaw</a> is useful : <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/129323\" target=\"_blank\">https://www.kaggle.com/c/bengaliai-cv19/discussion/129323</a></p>",
      "rawMarkdown": "i hope this discussion post of @raghaw is useful : https://www.kaggle.com/c/bengaliai-cv19/discussion/129323",
      "votes": 3,
      "replies": [
        {
          "id": 1358342,
          "postDate": "2021-06-20T11:33:06.703Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> , I also encountered this before, I think <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> code should solve this issue as well. It would be a bummer when \"resume\" training is not really a verbatim resumption.</p>",
          "rawMarkdown": "Thanks @mobassir , I also encountered this before, I think @cpmpml code should solve this issue as well. It would be a bummer when \"resume\" training is not really a verbatim resumption.",
          "votes": 2
        },
        {
          "id": 1395800,
          "postDate": "2021-07-21T14:44:05.997Z",
          "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> <a href=\"https://www.kaggle.com/reighns\" target=\"_blank\">@reighns</a>, The steps given by CPMP will solve your problem and it will work. I also solved my problem that I encountered in bengali AI competition, by a minor modification.</p>\n<p>Problematic sequences that I used in bengali AI competition:</p>\n<ol>\n<li>Initilize optimizer</li>\n<li>Load state_dict of optimizer</li>\n<li>Initilize scheduler</li>\n<li>Load state_dict of scheduler</li>\n</ol>\n<p>Problem in above sequence was at step 3. Once scheduler initlized, it resets the current learning rate in optimizer and initializes it to the learning rate of epoch 0. So I swapped step 2 and 3 that solved my problem. So the correct step that is working for all optimizers and schedulers that I have used till now is:</p>\n<ol>\n<li>Initilize optimizer</li>\n<li>Initilize scheduler</li>\n<li>Load state_dict of optimizer</li>\n<li>Load state_dict of scheduler</li>\n</ol>",
          "rawMarkdown": "@mobassir @reighns, The steps given by CPMP will solve your problem and it will work. I also solved my problem that I encountered in bengali AI competition, by a minor modification.\n\nProblematic sequences that I used in bengali AI competition:\n\n1. Initilize optimizer\n2. Load state_dict of optimizer\n3. Initilize scheduler\n4. Load state_dict of scheduler\n\nProblem in above sequence was at step 3. Once scheduler initlized, it resets the current learning rate in optimizer and initializes it to the learning rate of epoch 0. So I swapped step 2 and 3 that solved my problem. So the correct step that is working for all optimizers and schedulers that I have used till now is:\n\n1. Initilize optimizer\n2. Initilize scheduler\n3. Load state_dict of optimizer\n4. Load state_dict of scheduler",
          "votes": 3
        }
      ]
    },
    {
      "id": 1359403,
      "postDate": "2021-06-21T09:19:24.753Z",
      "content": "<p>Thanks Gao. I learned a lot from your work.<br>\nI usually training CNN model 5 epoch / 1day.<br>\nWhen I reload the model's state, the training rate was unstable sometimes.</p>\n<p>I didn't think I need to save optimizer state. Thank you.</p>",
      "rawMarkdown": "Thanks Gao. I learned a lot from your work.\nI usually training CNN model 5 epoch / 1day.\nWhen I reload the model's state, the training rate was unstable sometimes.\n \nI didn't think I need to save optimizer state. Thank you.",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1358199,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2021-06-20T09:05:36.080000",
      "content": "<p>You have to save your optimizer parameters as you save your model weights when you checkpoint.</p>\n<p>Then when you resume you load them.</p>\n<p>For instance here are functions I use to save and restore checkpoints with pytorch.  Of course you need to adjust it to your code.  It is important that you create your model, optimizer, etc in the load checkpoint function exactly as you create them in your training loop.  The scaler is only useful if you use mixed precision (torch.cuda.amp).</p>\n<pre><code>def save_checkpoint(model, optimizer, scheduler, scaler, epoch, fold, seed, fname=fname):\n    checkpoint = {\n        'model': model.state_dict(),\n        'optimizer': optimizer.state_dict(),\n        'scheduler': scheduler.state_dict(),\n        'scaler': scaler.state_dict(),\n        'epoch': epoch,\n        'fold':fold,\n        'seed':seed,\n        }\n    torch.save(checkpoint, '../checkpoints/%s/%s_%d_%d.pt' % (fname, fname, fold, seed))\n\ndef load_checkpoint(fold, seed, fname):\n    model = create_model().to(device)\n    optimizer = optimizer = torch.optim.Adam(model.parameters(), lr=MAX_LR)\n    scheduler = torch.optim.lr_scheduler.OneCycleLR(optimizer=optimizer, \n                                              pct_start=PCT_START, \n                                              div_factor=DIV_FACTOR \n                                              max_lr=MAX_LR, \n                                              epochs=EPOCHS, \n                                              steps_per_epoch=int(np.ceil(len(train_data_loader)/GRADIENT_ACCUMULATION)))\n    scaler = GradScaler()\n    checkpoint = torch.load('../checkpoints/%s/%s_%d_%d.pt' % (fname, fname, fold, seed))\n    model.load_state_dict(checkpoint['model'])\n    optimizer.load_state_dict(checkpoint['optimizer'])\n    scheduler.load_state_dict(checkpoint['scheduler'])\n    scaler.load_state_dict(checkpoint['scaler'])\n    return model, optimizer, scheduler, scaler, epoch\n</code></pre>",
      "votes": 18,
      "replies": [
        {
          "id": 1358237,
          "author_name": "gao-hongnan",
          "author_url": "",
          "post_date": "2021-06-20T09:36:09.393000",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> Thanks, this clears the air, so we still resume in an almost verbatim manner when we resume training. I got it now :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1358809,
          "author_name": "sajwankit",
          "author_url": "",
          "post_date": "2021-06-20T18:50:57.543000",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> Thanks, quite useful. One question on the seed, we would need to checkpoint the last state of random generator right? so when we resume training fresh random samples are created and not what were already created in the previous epochs (using same seed again).</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1358863,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-20T20:36:54.110000",
          "content": "<p><a href=\"https://www.kaggle.com/ankitsajwan\" target=\"_blank\">@ankitsajwan</a> I guess you're right.  Initializing the pytorch random generator would ensure path would be even closer to what it would have been without the restart.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1358255,
      "author_name": "Mobassir",
      "author_url": "",
      "post_date": "2021-06-20T10:01:22.527000",
      "content": "<p>i hope this discussion post of <a href=\"https://www.kaggle.com/raghaw\" target=\"_blank\">@raghaw</a> is useful : <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/129323\" target=\"_blank\">https://www.kaggle.com/c/bengaliai-cv19/discussion/129323</a></p>",
      "votes": 3,
      "replies": [
        {
          "id": 1358342,
          "author_name": "gao-hongnan",
          "author_url": "",
          "post_date": "2021-06-20T11:33:06.703000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> , I also encountered this before, I think <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> code should solve this issue as well. It would be a bummer when \"resume\" training is not really a verbatim resumption.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1395800,
          "author_name": "Raghawendra Singh",
          "author_url": "",
          "post_date": "2021-07-21T14:44:05.997000",
          "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> <a href=\"https://www.kaggle.com/reighns\" target=\"_blank\">@reighns</a>, The steps given by CPMP will solve your problem and it will work. I also solved my problem that I encountered in bengali AI competition, by a minor modification.</p>\n<p>Problematic sequences that I used in bengali AI competition:</p>\n<ol>\n<li>Initilize optimizer</li>\n<li>Load state_dict of optimizer</li>\n<li>Initilize scheduler</li>\n<li>Load state_dict of scheduler</li>\n</ol>\n<p>Problem in above sequence was at step 3. Once scheduler initlized, it resets the current learning rate in optimizer and initializes it to the learning rate of epoch 0. So I swapped step 2 and 3 that solved my problem. So the correct step that is working for all optimizers and schedulers that I have used till now is:</p>\n<ol>\n<li>Initilize optimizer</li>\n<li>Initilize scheduler</li>\n<li>Load state_dict of optimizer</li>\n<li>Load state_dict of scheduler</li>\n</ol>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1359403,
      "author_name": "WOOSUNG YOON",
      "author_url": "",
      "post_date": "2021-06-21T09:19:24.753000",
      "content": "<p>Thanks Gao. I learned a lot from your work.<br>\nI usually training CNN model 5 epoch / 1day.<br>\nWhen I reload the model's state, the training rate was unstable sometimes.</p>\n<p>I didn't think I need to save optimizer state. Thank you.</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1358199": "You have to save your optimizer parameters as you save your model weights when you checkpoint.\n\nThen when you resume you load them.\n\nFor instance here are functions I use to save and restore checkpoints with pytorch.  Of course you need to adjust it to your code.  It is important that you create your model, optimizer, etc in the load checkpoint function exactly as you create them in your training loop.  The scaler is only useful if you use mixed precision (torch.cuda.amp).\n\n```\n\ndef save_checkpoint(model, optimizer, scheduler, scaler, epoch, fold, seed, fname=fname):\n    checkpoint = {\n        'model': model.state_dict(),\n        'optimizer': optimizer.state_dict(),\n        'scheduler': scheduler.state_dict(),\n        'scaler': scaler.state_dict(),\n        'epoch': epoch,\n        'fold':fold,\n        'seed':seed,\n        }\n    torch.save(checkpoint, '../checkpoints/%s/%s_%d_%d.pt' % (fname, fname, fold, seed))\n\ndef load_checkpoint(fold, seed, fname):\n    model = create_model().to(device)\n    optimizer = optimizer = torch.optim.Adam(model.parameters(), lr=MAX_LR)\n    scheduler = torch.optim.lr_scheduler.OneCycleLR(optimizer=optimizer, \n                                              pct_start=PCT_START, \n                                              div_factor=DIV_FACTOR \n                                              max_lr=MAX_LR, \n                                              epochs=EPOCHS, \n                                              steps_per_epoch=int(np.ceil(len(train_data_loader)/GRADIENT_ACCUMULATION)))\n    scaler = GradScaler()\n    checkpoint = torch.load('../checkpoints/%s/%s_%d_%d.pt' % (fname, fname, fold, seed))\n    model.load_state_dict(checkpoint['model'])\n    optimizer.load_state_dict(checkpoint['optimizer'])\n    scheduler.load_state_dict(checkpoint['scheduler'])\n    scaler.load_state_dict(checkpoint['scaler'])\n    return model, optimizer, scheduler, scaler, epoch\n\n```",
    "1358123": "Now, I am unsure if this has been answered, I remember my buddy telling me that anecdotally, one can \"reset\" the LR a little when one resumes training from a model's checkpoint.\n\nIt is a bit counter-intuitive to me, but say I trained a model for 16 epochs, with a custom scheduler, say `OneCycleLr` + `Adam` or something, then when you resume training, should we reset the initial learning rate? I would think resetting is counter-intuitive as the purpose of the learning rate is to tune it such that your model can slowly converge to the minima (global one if the function is convex). \n\nSo the question is: \n\n1. Should I reset the learning rate when resume training, if yes, reset to what?\n\n2. If we should not reset, what is a good way to \"extract\" the last learning rate, as some scheduler depends on factors like epochs...\n",
    "1358255": "i hope this discussion post of @raghaw is useful : https://www.kaggle.com/c/bengaliai-cv19/discussion/129323",
    "1359403": "Thanks Gao. I learned a lot from your work.\nI usually training CNN model 5 epoch / 1day.\nWhen I reload the model's state, the training rate was unstable sometimes.\n \nI didn't think I need to save optimizer state. Thank you."
  }
}