{
  "id": 193879,
  "title": "Question: Pytorch saving and loading models for resuming the training process",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/193879",
  "author_name": "",
  "post_date": "2020-10-29T11:13:25.127295300Z",
  "votes": 4,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hello, <br>\nI am stuck in this dilemma of saving the weights along with the optimizer state to either resume training or go on with the inference. However, I'm not sure if I'm doing this correctly.</p>\n<p>This is my current code for saving and loading the model :-</p>\n<h6>saving</h6>\n<pre><code>    if i % cfg['train_params']['checkpoint_every_n_steps'] == 0:\n        state = {\n          'state_dict': model.state_dict(),\n          'optimizer': optimizer.state_dict(),\n        }\n        torch.save(state,f'{model_name}_stage2_{i}.pth')\n</code></pre>\n<h6>loading</h6>\n<p><code>device = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")</code><br>\n<code>model = LyftResNetModel(cfg)</code><br>\n<code>optimizer = optim.Adam(model.parameters(), lr=cfg[\"model_params\"][\"lr\"])</code></p>\n<p><code>checkpoint = torch.load(weight_path)</code><br>\n<code>model.load_state_dict(checkpoint[state_dict'])</code><br>\n<code>optimizer.load_state_dict(checkpoint['optimizer'])</code></p>\n<p>But I was going through the Pytorch documentation and it stated that </p>\n<blockquote>\n  <p>When saving a general checkpoint, to be used for either inference or resuming training, you must save more than just the model’s state_dict. It is important to also save the optimizer’s state_dict, as this contains buffers and parameters that are updated as the model trains. <strong>Other items that you may want to save are the epoch you left off on, the latest recorded training loss</strong>, external torch.nn.Embedding layers, etc. As a result, such a checkpoint is often 2~3 times larger than the model alone.  </p>\n</blockquote>\n<p>So, my question is : Should I also save the epochs and losses while saving the weights<br>\nlike this<br>\n<code>epoch = checkpoint['epoch']</code><br>\n<code>loss = checkpoint['loss']</code><br>\n or continue with my current code and what's the difference between these two ? </p>",
  "messages": [
    {
      "id": "1063809",
      "postDate": "10/29/2020 11:13:25",
      "content": "<p>Hello, <br>\nI am stuck in this dilemma of saving the weights along with the optimizer state to either resume training or go on with the inference. However, I'm not sure if I'm doing this correctly.</p>\n<p>This is my current code for saving and loading the model :-</p>\n<h6>saving</h6>\n<pre><code>    if i % cfg['train_params']['checkpoint_every_n_steps'] == 0:\n        state = {\n          'state_dict': model.state_dict(),\n          'optimizer': optimizer.state_dict(),\n        }\n        torch.save(state,f'{model_name}_stage2_{i}.pth')\n</code></pre>\n<h6>loading</h6>\n<p><code>device = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")</code><br>\n<code>model = LyftResNetModel(cfg)</code><br>\n<code>optimizer = optim.Adam(model.parameters(), lr=cfg[\"model_params\"][\"lr\"])</code></p>\n<p><code>checkpoint = torch.load(weight_path)</code><br>\n<code>model.load_state_dict(checkpoint[state_dict'])</code><br>\n<code>optimizer.load_state_dict(checkpoint['optimizer'])</code></p>\n<p>But I was going through the Pytorch documentation and it stated that </p>\n<blockquote>\n  <p>When saving a general checkpoint, to be used for either inference or resuming training, you must save more than just the model’s state_dict. It is important to also save the optimizer’s state_dict, as this contains buffers and parameters that are updated as the model trains. <strong>Other items that you may want to save are the epoch you left off on, the latest recorded training loss</strong>, external torch.nn.Embedding layers, etc. As a result, such a checkpoint is often 2~3 times larger than the model alone.  </p>\n</blockquote>\n<p>So, my question is : Should I also save the epochs and losses while saving the weights<br>\nlike this<br>\n<code>epoch = checkpoint['epoch']</code><br>\n<code>loss = checkpoint['loss']</code><br>\n or continue with my current code and what's the difference between these two ? </p>",
      "rawMarkdown": "Hello, \nI am stuck in this dilemma of saving the weights along with the optimizer state to either resume training or go on with the inference. However, I'm not sure if I'm doing this correctly.\n\nThis is my current code for saving and loading the model :-\n###### saving\n        if i % cfg['train_params']['checkpoint_every_n_steps'] == 0:\n            state = {\n              'state_dict': model.state_dict(),\n              'optimizer': optimizer.state_dict(),\n            }\n            torch.save(state,f'{model_name}_stage2_{i}.pth')\n\n###### loading\n`device = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")`\n`model = LyftResNetModel(cfg)`\n`optimizer = optim.Adam(model.parameters(), lr=cfg[\"model_params\"][\"lr\"])`\n\n`checkpoint = torch.load(weight_path)`\n`model.load_state_dict(checkpoint[state_dict'])`\n`optimizer.load_state_dict(checkpoint['optimizer'])`\n\nBut I was going through the Pytorch documentation and it stated that \n> When saving a general checkpoint, to be used for either inference or resuming training, you must save more than just the model’s state_dict. It is important to also save the optimizer’s state_dict, as this contains buffers and parameters that are updated as the model trains. **Other items that you may want to save are the epoch you left off on, the latest recorded training loss**, external torch.nn.Embedding layers, etc. As a result, such a checkpoint is often 2~3 times larger than the model alone.  \n\nSo, my question is : Should I also save the epochs and losses while saving the weights\nlike this\n`epoch = checkpoint['epoch']`\n`loss = checkpoint['loss']`\n or continue with my current code and what's the difference between these two ?",
      "votes": null
    },
    {
      "id": "1064243",
      "postDate": "10/29/2020 22:15:27",
      "content": "<p>There will be no difference if you save or don't save epoch and losses.The only purpose of saving is to make it look as the training had never stopped</p>",
      "rawMarkdown": "There will be no difference if you save or don't save epoch and losses.The only purpose of saving is to make it look as the training had never stopped",
      "votes": null
    },
    {
      "id": "1065068",
      "postDate": "10/30/2020 20:08:54",
      "content": "<p>I think you only need to save optimizer and model states.</p>",
      "rawMarkdown": "I think you only need to save optimizer and model states.",
      "votes": null
    },
    {
      "id": "1065864",
      "postDate": "10/31/2020 23:16:10",
      "content": "<p>Your current code is well enough for this task. I'm also doing it the same.</p>\n<p>You might need to save the dict of your LR-Scheduler too, if you use one, just add:<br>\n<code>\n...\n 'scheduler_state_dict': lr_scheduler.state_dict()</code></p>\n<p>to your dict.</p>",
      "rawMarkdown": "Your current code is well enough for this task. I'm also doing it the same.\n\nYou might need to save the dict of your LR-Scheduler too, if you use one, just add:\n`\n...\n 'scheduler_state_dict': lr_scheduler.state_dict()`\n\nto your dict.",
      "votes": null
    },
    {
      "id": "1066471",
      "postDate": "11/01/2020 19:57:12",
      "content": "<p>Thanks, I'll try implementing that too.</p>",
      "rawMarkdown": "Thanks, I'll try implementing that too.",
      "votes": null
    },
    {
      "id": "1067250",
      "postDate": "11/02/2020 12:13:11",
      "content": "<p>i am using pytorch lightning and it can resume the tranining pipeline very quickly</p>",
      "rawMarkdown": "i am using pytorch lightning and it can resume the tranining pipeline very quickly",
      "votes": null
    },
    {
      "id": "1067881",
      "postDate": "11/02/2020 19:33:32",
      "content": "<p>Seems like a good choice! I should try that!</p>",
      "rawMarkdown": "Seems like a good choice! I should try that!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1064243,
      "author_name": "",
      "author_url": "",
      "post_date": "10/29/2020 22:15:27",
      "content": "<p>There will be no difference if you save or don't save epoch and losses.The only purpose of saving is to make it look as the training had never stopped</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1065068,
      "author_name": "louis925",
      "author_url": "",
      "post_date": "10/30/2020 20:08:54",
      "content": "<p>I think you only need to save optimizer and model states.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1065864,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "10/31/2020 23:16:10",
      "content": "<p>Your current code is well enough for this task. I'm also doing it the same.</p>\n<p>You might need to save the dict of your LR-Scheduler too, if you use one, just add:<br>\n<code>\n...\n 'scheduler_state_dict': lr_scheduler.state_dict()</code></p>\n<p>to your dict.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1066471,
          "author_name": "thakurudit",
          "author_url": "",
          "post_date": "11/01/2020 19:57:12",
          "content": "<p>Thanks, I'll try implementing that too.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1067250,
      "author_name": "doanquanvietnamca",
      "author_url": "",
      "post_date": "11/02/2020 12:13:11",
      "content": "<p>i am using pytorch lightning and it can resume the tranining pipeline very quickly</p>",
      "votes": null,
      "replies": [
        {
          "id": 1067881,
          "author_name": "louis925",
          "author_url": "",
          "post_date": "11/02/2020 19:33:32",
          "content": "<p>Seems like a good choice! I should try that!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1063809": "Hello, \nI am stuck in this dilemma of saving the weights along with the optimizer state to either resume training or go on with the inference. However, I'm not sure if I'm doing this correctly.\n\nThis is my current code for saving and loading the model :-\n###### saving\n        if i % cfg['train_params']['checkpoint_every_n_steps'] == 0:\n            state = {\n              'state_dict': model.state_dict(),\n              'optimizer': optimizer.state_dict(),\n            }\n            torch.save(state,f'{model_name}_stage2_{i}.pth')\n\n###### loading\n`device = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")`\n`model = LyftResNetModel(cfg)`\n`optimizer = optim.Adam(model.parameters(), lr=cfg[\"model_params\"][\"lr\"])`\n\n`checkpoint = torch.load(weight_path)`\n`model.load_state_dict(checkpoint[state_dict'])`\n`optimizer.load_state_dict(checkpoint['optimizer'])`\n\nBut I was going through the Pytorch documentation and it stated that \n> When saving a general checkpoint, to be used for either inference or resuming training, you must save more than just the model’s state_dict. It is important to also save the optimizer’s state_dict, as this contains buffers and parameters that are updated as the model trains. **Other items that you may want to save are the epoch you left off on, the latest recorded training loss**, external torch.nn.Embedding layers, etc. As a result, such a checkpoint is often 2~3 times larger than the model alone.  \n\nSo, my question is : Should I also save the epochs and losses while saving the weights\nlike this\n`epoch = checkpoint['epoch']`\n`loss = checkpoint['loss']`\n or continue with my current code and what's the difference between these two ?",
    "1064243": "There will be no difference if you save or don't save epoch and losses.The only purpose of saving is to make it look as the training had never stopped",
    "1065068": "I think you only need to save optimizer and model states.",
    "1065864": "Your current code is well enough for this task. I'm also doing it the same.\n\nYou might need to save the dict of your LR-Scheduler too, if you use one, just add:\n`\n...\n 'scheduler_state_dict': lr_scheduler.state_dict()`\n\nto your dict.",
    "1066471": "Thanks, I'll try implementing that too.",
    "1067250": "i am using pytorch lightning and it can resume the tranining pipeline very quickly",
    "1067881": "Seems like a good choice! I should try that!"
  },
  "source": "meta"
}