{
  "id": 129323,
  "title": "Retraining checkpoints increases train and validation loss in pytorch",
  "url": "/competitions/bengaliai-cv19/discussion/129323",
  "author_name": "Raghawendra Singh",
  "post_date": "2020-02-07T06:43:06.460000",
  "votes": 11,
  "comment_count": 16,
  "views": 0,
  "content": "<p>Since I don't have my own GPU, I am using google colab for training my model. I save checkpoint after every epoch and my checkpoints include model state_dict, optimizer state_dict and scheduler state_dict. For training ahead from my last checkpoint I load state_dict of model, optimizer and scheduler of last saved epoch before starting. It seems that training loss and validation loss is not in continuation of previous saved epoch, It increases by significant amount and usually requirs 3-4 epoch of training to reach the same values as in last checkpoint. After searching for my problem I found this post <a href=\"https://discuss.pytorch.org/t/loading-a-saved-model-for-continue-training/17244\">https://discuss.pytorch.org/t/loading-a-saved-model-for-continue-training/17244</a> and found that this problem is happening only in adaptive optimizers such as adam, adamW and it is not an issue in SDG. does anyone facing same issue?</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1490082%2F1e3ae887439ce35384499f4e399dd655%2FOptim_problem.png?generation=1581057855538474&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 738904,
      "postDate": "2020-02-07T06:43:06.460Z",
      "content": "<p>Since I don't have my own GPU, I am using google colab for training my model. I save checkpoint after every epoch and my checkpoints include model state_dict, optimizer state_dict and scheduler state_dict. For training ahead from my last checkpoint I load state_dict of model, optimizer and scheduler of last saved epoch before starting. It seems that training loss and validation loss is not in continuation of previous saved epoch, It increases by significant amount and usually requirs 3-4 epoch of training to reach the same values as in last checkpoint. After searching for my problem I found this post <a href=\"https://discuss.pytorch.org/t/loading-a-saved-model-for-continue-training/17244\">https://discuss.pytorch.org/t/loading-a-saved-model-for-continue-training/17244</a> and found that this problem is happening only in adaptive optimizers such as adam, adamW and it is not an issue in SDG. does anyone facing same issue?</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1490082%2F1e3ae887439ce35384499f4e399dd655%2FOptim_problem.png?generation=1581057855538474&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Since I don't have my own GPU, I am using google colab for training my model. I save checkpoint after every epoch and my checkpoints include model state_dict, optimizer state_dict and scheduler state_dict. For training ahead from my last checkpoint I load state_dict of model, optimizer and scheduler of last saved epoch before starting. It seems that training loss and validation loss is not in continuation of previous saved epoch, It increases by significant amount and usually requirs 3-4 epoch of training to reach the same values as in last checkpoint. After searching for my problem I found this post https://discuss.pytorch.org/t/loading-a-saved-model-for-continue-training/17244 and found that this problem is happening only in adaptive optimizers such as adam, adamW and it is not an issue in SDG. does anyone facing same issue?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1490082%2F1e3ae887439ce35384499f4e399dd655%2FOptim_problem.png?generation=1581057855538474&amp;alt=media)\n",
      "votes": 11
    },
    {
      "id": 741338,
      "postDate": "2020-02-10T13:57:16.433Z",
      "content": "<p>Few optimizers in Pytorch such as AdamW or RMSProp are not property reloaded sometimes. I have similar experiences but I didn't check in detail what the problem was.</p>",
      "rawMarkdown": "Few optimizers in Pytorch such as AdamW or RMSProp are not property reloaded sometimes. I have similar experiences but I didn't check in detail what the problem was.",
      "votes": 1
    },
    {
      "id": 740619,
      "postDate": "2020-02-09T15:37:09.100Z",
      "content": "<p><a href=\"/raghaw\">@raghaw</a> with SGD have you been able to get the exact same results as of the previously trained model?</p>",
      "rawMarkdown": "@raghaw with SGD have you been able to get the exact same results as of the previously trained model?",
      "votes": 1,
      "replies": [
        {
          "id": 740642,
          "postDate": "2020-02-09T15:55:09.807Z",
          "content": "<p><a href=\"/rohitagarwal\">@rohitagarwal</a> yes, I have trained few model with sgd in diferent sessions, it seems working.</p>",
          "rawMarkdown": "@rohitagarwal yes, I have trained few model with sgd in diferent sessions, it seems working.",
          "votes": 1
        },
        {
          "id": 740656,
          "postDate": "2020-02-09T16:10:58.077Z",
          "content": "<p>Thanks <a href=\"/raghaw\">@raghaw</a> I'll try that too</p>",
          "rawMarkdown": "Thanks @raghaw I'll try that too"
        }
      ]
    },
    {
      "id": 739839,
      "postDate": "2020-02-08T13:40:57.790Z",
      "content": "<p>This may due to the moving average term in adaptive optimizers such as Adam and its variants. SGD doesn't use any moving average term but SGDM (SGD with Momentum) has. But I don't know why loading optimizer state dict couldn't help. </p>",
      "rawMarkdown": "This may due to the moving average term in adaptive optimizers such as Adam and its variants. SGD doesn't use any moving average term but SGDM (SGD with Momentum) has. But I don't know why loading optimizer state dict couldn't help. ",
      "votes": 1,
      "replies": [
        {
          "id": 740643,
          "postDate": "2020-02-09T15:56:05.917Z",
          "content": "<p><a href=\"/syoya1997\">@syoya1997</a>  SGD with momentom is working fine.</p>",
          "rawMarkdown": "@syoya1997  SGD with momentom is working fine."
        },
        {
          "id": 740998,
          "postDate": "2020-02-10T04:24:39.517Z",
          "content": "<p>Then it becomes more confusing for me.</p>",
          "rawMarkdown": "Then it becomes more confusing for me."
        },
        {
          "id": 741355,
          "postDate": "2020-02-10T14:30:44.877Z",
          "content": "<p>Size of checkpoints with Adam or AdamW is around 298MB, but with SGD with momentum it is around 198MB. it means SGD with momentum uses lot less parameters than Adam or AdamW and its parameters may be very less depandent on previous states. thats why it is working. it is just my guess, I may be wrong.</p>",
          "rawMarkdown": "Size of checkpoints with Adam or AdamW is around 298MB, but with SGD with momentum it is around 198MB. it means SGD with momentum uses lot less parameters than Adam or AdamW and its parameters may be very less depandent on previous states. thats why it is working. it is just my guess, I may be wrong."
        },
        {
          "id": 741365,
          "postDate": "2020-02-10T14:40:01.217Z",
          "content": "<p>The size of Adam or AdamW is larger and it's simply because Adam and AdamW has two moving average terms: the first moment estimation and the second moment estimation of gradients. SGDM has  only one term, which is its momentum term. You can easily recognize this from the math equations of these optimizers. </p>",
          "rawMarkdown": "The size of Adam or AdamW is larger and it's simply because Adam and AdamW has two moving average terms: the first moment estimation and the second moment estimation of gradients. SGDM has  only one term, which is its momentum term. You can easily recognize this from the math equations of these optimizers. "
        },
        {
          "id": 741378,
          "postDate": "2020-02-10T14:52:14.713Z",
          "content": "<p>I know but difference of 100 MB is huge.</p>",
          "rawMarkdown": "I know but difference of 100 MB is huge."
        },
        {
          "id": 741405,
          "postDate": "2020-02-10T15:38:17.623Z",
          "content": "<p>Well, I think that's quite reasonable considering how your model is so large.</p>",
          "rawMarkdown": "Well, I think that's quite reasonable considering how your model is so large."
        }
      ]
    },
    {
      "id": 738987,
      "postDate": "2020-02-07T08:54:37.903Z",
      "content": "<p>yes, I am facing same issue </p>",
      "rawMarkdown": "yes, I am facing same issue ",
      "votes": 1
    },
    {
      "id": 901397,
      "postDate": "2020-06-25T12:49:58.067Z",
      "content": "<p>I am facing the same issue in SGD. Validation loss increases and reaches up to 97% once I load the previously saved model using ModelCheckpoint. Otherwise, the validation accuracy reaches up to 71% only.</p>",
      "rawMarkdown": "I am facing the same issue in SGD. Validation loss increases and reaches up to 97% once I load the previously saved model using ModelCheckpoint. Otherwise, the validation accuracy reaches up to 71% only."
    },
    {
      "id": 738984,
      "postDate": "2020-02-07T08:46:14.140Z",
      "content": "<p>Its very strange that this problem is reported on pytorch forum but no solution!!!</p>",
      "rawMarkdown": "Its very strange that this problem is reported on pytorch forum but no solution!!!"
    },
    {
      "id": 738911,
      "postDate": "2020-02-07T07:10:26.247Z",
      "content": "<p>We have the same error in fast ai , and it's taking us about ~15 epochs to get back to the previous better score...</p>",
      "rawMarkdown": "We have the same error in fast ai , and it's taking us about ~15 epochs to get back to the previous better score..."
    },
    {
      "id": 739834,
      "postDate": "2020-02-08T13:35:21.213Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 741338,
      "author_name": "Ildoo Kim",
      "author_url": "",
      "post_date": "2020-02-10T13:57:16.433000",
      "content": "<p>Few optimizers in Pytorch such as AdamW or RMSProp are not property reloaded sometimes. I have similar experiences but I didn't check in detail what the problem was.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 740619,
      "author_name": "Rohit Agarwal",
      "author_url": "",
      "post_date": "2020-02-09T15:37:09.100000",
      "content": "<p><a href=\"/raghaw\">@raghaw</a> with SGD have you been able to get the exact same results as of the previously trained model?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 740642,
          "author_name": "Raghawendra Singh",
          "author_url": "",
          "post_date": "2020-02-09T15:55:09.807000",
          "content": "<p><a href=\"/rohitagarwal\">@rohitagarwal</a> yes, I have trained few model with sgd in diferent sessions, it seems working.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 740656,
          "author_name": "Rohit Agarwal",
          "author_url": "",
          "post_date": "2020-02-09T16:10:58.077000",
          "content": "<p>Thanks <a href=\"/raghaw\">@raghaw</a> I'll try that too</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 739839,
      "author_name": "syoya",
      "author_url": "",
      "post_date": "2020-02-08T13:40:57.790000",
      "content": "<p>This may due to the moving average term in adaptive optimizers such as Adam and its variants. SGD doesn't use any moving average term but SGDM (SGD with Momentum) has. But I don't know why loading optimizer state dict couldn't help. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 740643,
          "author_name": "Raghawendra Singh",
          "author_url": "",
          "post_date": "2020-02-09T15:56:05.917000",
          "content": "<p><a href=\"/syoya1997\">@syoya1997</a>  SGD with momentom is working fine.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 740998,
          "author_name": "syoya",
          "author_url": "",
          "post_date": "2020-02-10T04:24:39.517000",
          "content": "<p>Then it becomes more confusing for me.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 741355,
          "author_name": "Raghawendra Singh",
          "author_url": "",
          "post_date": "2020-02-10T14:30:44.877000",
          "content": "<p>Size of checkpoints with Adam or AdamW is around 298MB, but with SGD with momentum it is around 198MB. it means SGD with momentum uses lot less parameters than Adam or AdamW and its parameters may be very less depandent on previous states. thats why it is working. it is just my guess, I may be wrong.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 741365,
          "author_name": "syoya",
          "author_url": "",
          "post_date": "2020-02-10T14:40:01.217000",
          "content": "<p>The size of Adam or AdamW is larger and it's simply because Adam and AdamW has two moving average terms: the first moment estimation and the second moment estimation of gradients. SGDM has  only one term, which is its momentum term. You can easily recognize this from the math equations of these optimizers. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 741378,
          "author_name": "Raghawendra Singh",
          "author_url": "",
          "post_date": "2020-02-10T14:52:14.713000",
          "content": "<p>I know but difference of 100 MB is huge.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 741405,
          "author_name": "syoya",
          "author_url": "",
          "post_date": "2020-02-10T15:38:17.623000",
          "content": "<p>Well, I think that's quite reasonable considering how your model is so large.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 738987,
      "author_name": "Manish Nayak",
      "author_url": "",
      "post_date": "2020-02-07T08:54:37.903000",
      "content": "<p>yes, I am facing same issue </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 901397,
      "author_name": "Anik Sen",
      "author_url": "",
      "post_date": "2020-06-25T12:49:58.067000",
      "content": "<p>I am facing the same issue in SGD. Validation loss increases and reaches up to 97% once I load the previously saved model using ModelCheckpoint. Otherwise, the validation accuracy reaches up to 71% only.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 738984,
      "author_name": "Raghawendra Singh",
      "author_url": "",
      "post_date": "2020-02-07T08:46:14.140000",
      "content": "<p>Its very strange that this problem is reported on pytorch forum but no solution!!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 738911,
      "author_name": "Satwik",
      "author_url": "",
      "post_date": "2020-02-07T07:10:26.247000",
      "content": "<p>We have the same error in fast ai , and it's taking us about ~15 epochs to get back to the previous better score...</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 739834,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-08T13:35:21.213000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "738904": "Since I don't have my own GPU, I am using google colab for training my model. I save checkpoint after every epoch and my checkpoints include model state_dict, optimizer state_dict and scheduler state_dict. For training ahead from my last checkpoint I load state_dict of model, optimizer and scheduler of last saved epoch before starting. It seems that training loss and validation loss is not in continuation of previous saved epoch, It increases by significant amount and usually requirs 3-4 epoch of training to reach the same values as in last checkpoint. After searching for my problem I found this post https://discuss.pytorch.org/t/loading-a-saved-model-for-continue-training/17244 and found that this problem is happening only in adaptive optimizers such as adam, adamW and it is not an issue in SDG. does anyone facing same issue?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1490082%2F1e3ae887439ce35384499f4e399dd655%2FOptim_problem.png?generation=1581057855538474&amp;alt=media)\n",
    "741338": "Few optimizers in Pytorch such as AdamW or RMSProp are not property reloaded sometimes. I have similar experiences but I didn't check in detail what the problem was.",
    "740619": "@raghaw with SGD have you been able to get the exact same results as of the previously trained model?",
    "739839": "This may due to the moving average term in adaptive optimizers such as Adam and its variants. SGD doesn't use any moving average term but SGDM (SGD with Momentum) has. But I don't know why loading optimizer state dict couldn't help. ",
    "738987": "yes, I am facing same issue ",
    "901397": "I am facing the same issue in SGD. Validation loss increases and reaches up to 97% once I load the previously saved model using ModelCheckpoint. Otherwise, the validation accuracy reaches up to 71% only.",
    "738984": "Its very strange that this problem is reported on pytorch forum but no solution!!!",
    "738911": "We have the same error in fast ai , and it's taking us about ~15 epochs to get back to the previous better score...",
    "739834": ""
  }
}