{
  "id": 207394,
  "title": "Performance Difference Between Learning Rate 1e-4 and 0.00025118",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/207394",
  "author_name": "",
  "post_date": "2020-12-29T14:39:29.925565300Z",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I was using learning rate 1e-4 for one run and then tried PytorchLightning's learning rate finder to find a better learning rate.<br>\nThe learning rate finder suggested me ~0.00025118 to use as a learning rate.</p>\n<p>My model architecture has resnet50 with an additional linear layer.</p>\n<p><strong>For Learning Rate 1e-4:</strong></p>\n<p>Afte 21 epochs,<br>\nloss=0.267<br>\ntrain_loss=0.442, train_accuracy=0.793,<br>\nvalid_loss=0.396, valid_accuracy=0.863,<br>\naccuracy=0.863</p>\n<p><strong>For Learning Rate 0.00025118:</strong></p>\n<p>After 25 epochs,<br>\nloss=0.691,<br>\ntrain_loss=0.891, train_accuracy=0.621,<br>\nvalid_loss=0.62, valid_accuracy=0.7,<br>\naccuracy=0.7</p>\n<p>Is this drastic performance change normal? Or does this indicate bug in the code?</p>\n<p>*EDIT-1: Added total training loss for both scenarios. </p>\n<p>*UPDATE-1: There was a change in the weight decay, which I reverted and this hasn't made any difference.</p>",
  "messages": [
    {
      "id": "1131089",
      "postDate": "12/29/2020 14:39:29",
      "content": "<p>I was using learning rate 1e-4 for one run and then tried PytorchLightning's learning rate finder to find a better learning rate.<br>\nThe learning rate finder suggested me ~0.00025118 to use as a learning rate.</p>\n<p>My model architecture has resnet50 with an additional linear layer.</p>\n<p><strong>For Learning Rate 1e-4:</strong></p>\n<p>Afte 21 epochs,<br>\nloss=0.267<br>\ntrain_loss=0.442, train_accuracy=0.793,<br>\nvalid_loss=0.396, valid_accuracy=0.863,<br>\naccuracy=0.863</p>\n<p><strong>For Learning Rate 0.00025118:</strong></p>\n<p>After 25 epochs,<br>\nloss=0.691,<br>\ntrain_loss=0.891, train_accuracy=0.621,<br>\nvalid_loss=0.62, valid_accuracy=0.7,<br>\naccuracy=0.7</p>\n<p>Is this drastic performance change normal? Or does this indicate bug in the code?</p>\n<p>*EDIT-1: Added total training loss for both scenarios. </p>\n<p>*UPDATE-1: There was a change in the weight decay, which I reverted and this hasn't made any difference.</p>",
      "rawMarkdown": "I was using learning rate 1e-4 for one run and then tried PytorchLightning's learning rate finder to find a better learning rate.\nThe learning rate finder suggested me ~0.00025118 to use as a learning rate.\n\nMy model architecture has resnet50 with an additional linear layer.\n\n**For Learning Rate 1e-4:**\n\nAfte 21 epochs,\nloss=0.267\ntrain_loss=0.442, train_accuracy=0.793,\nvalid_loss=0.396, valid_accuracy=0.863,\naccuracy=0.863\n\n**For Learning Rate 0.00025118:**\n\nAfter 25 epochs,\nloss=0.691,\ntrain_loss=0.891, train_accuracy=0.621,\nvalid_loss=0.62, valid_accuracy=0.7,\naccuracy=0.7\n\nIs this drastic performance change normal? Or does this indicate bug in the code?\n\n*EDIT-1: Added total training loss for both scenarios. \n\n*UPDATE-1: There was a change in the weight decay, which I reverted and this hasn't made any difference.",
      "votes": null
    },
    {
      "id": "1131183",
      "postDate": "12/29/2020 15:26:10",
      "content": "<p>This looks odd to me. With a higher learning rate of about 2.5e-4 and 25 epochs you end up getting a larger training loss than with a lower learning rate of 1e-4 and 21 epochs? That's the one thing I would have guessed I could predict (on the other hand, if the 2nd option did worse on the validation loss or accuracy, I would not necessarily think that was odd). To me that suggests something else is different/has gone wrong. E.g. your training-validation split might be different (or the optimizer? or the learning rate schedule?). The difference also seems quite striking to me, so I would guess it's more than what you'd expect due to randomness in the training, but it's hard to guess.</p>",
      "rawMarkdown": "This looks odd to me. With a higher learning rate of about 2.5e-4 and 25 epochs you end up getting a larger training loss than with a lower learning rate of 1e-4 and 21 epochs? That's the one thing I would have guessed I could predict (on the other hand, if the 2nd option did worse on the validation loss or accuracy, I would not necessarily think that was odd). To me that suggests something else is different/has gone wrong. E.g. your training-validation split might be different (or the optimizer? or the learning rate schedule?). The difference also seems quite striking to me, so I would guess it's more than what you'd expect due to randomness in the training, but it's hard to guess.",
      "votes": null
    },
    {
      "id": "1131299",
      "postDate": "12/29/2020 16:38:33",
      "content": "<p>Are you setting seeds? You have to make sure, that everything is exactly the same for a second run, check if you are maybe shuffeling your folds.</p>\n<p>Are you using early-stopping?</p>",
      "rawMarkdown": "Are you setting seeds? You have to make sure, that everything is exactly the same for a second run, check if you are maybe shuffeling your folds.\n\nAre you using early-stopping?",
      "votes": null
    },
    {
      "id": "1131329",
      "postDate": "12/29/2020 16:57:26",
      "content": "<p>Yes, I am using pytorch-lightnings <em>seed_everything</em> function. </p>",
      "rawMarkdown": "Yes, I am using pytorch-lightnings *seed_everything* function.",
      "votes": null
    },
    {
      "id": "1131351",
      "postDate": "12/29/2020 17:10:51",
      "content": "<p>I have updated my post, the wight decay was different in both run. <br>\nSo, now I am training again with the first run's weight decay. But the total training loss is not decreasing.<br>\nI am also guessing the issue might be in the train, test split.<br>\nAny idea on how can I check whether my seeding is working perfectly?</p>",
      "rawMarkdown": "I have updated my post, the wight decay was different in both run. \nSo, now I am training again with the first run's weight decay. But the total training loss is not decreasing.\nI am also guessing the issue might be in the train, test split.\nAny idea on how can I check whether my seeding is working perfectly?",
      "votes": null
    },
    {
      "id": "1131856",
      "postDate": "12/30/2020 03:17:46",
      "content": "<p>I can see similar results, but the difference is not so big as yours</p>",
      "rawMarkdown": "I can see similar results, but the difference is not so big as yours",
      "votes": null
    },
    {
      "id": "1132134",
      "postDate": "12/30/2020 07:28:03",
      "content": "<p>Running the same code twice with absolutely no changes (and seeing that results are then 100% identical) would check that everything like the train test split is really fixed.</p>",
      "rawMarkdown": "Running the same code twice with absolutely no changes (and seeing that results are then 100% identical) would check that everything like the train test split is really fixed.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1131183,
      "author_name": "bjoernholzhauer",
      "author_url": "",
      "post_date": "12/29/2020 15:26:10",
      "content": "<p>This looks odd to me. With a higher learning rate of about 2.5e-4 and 25 epochs you end up getting a larger training loss than with a lower learning rate of 1e-4 and 21 epochs? That's the one thing I would have guessed I could predict (on the other hand, if the 2nd option did worse on the validation loss or accuracy, I would not necessarily think that was odd). To me that suggests something else is different/has gone wrong. E.g. your training-validation split might be different (or the optimizer? or the learning rate schedule?). The difference also seems quite striking to me, so I would guess it's more than what you'd expect due to randomness in the training, but it's hard to guess.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1131351,
          "author_name": "jabertuhin",
          "author_url": "",
          "post_date": "12/29/2020 17:10:51",
          "content": "<p>I have updated my post, the wight decay was different in both run. <br>\nSo, now I am training again with the first run's weight decay. But the total training loss is not decreasing.<br>\nI am also guessing the issue might be in the train, test split.<br>\nAny idea on how can I check whether my seeding is working perfectly?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1132134,
          "author_name": "bjoernholzhauer",
          "author_url": "",
          "post_date": "12/30/2020 07:28:03",
          "content": "<p>Running the same code twice with absolutely no changes (and seeing that results are then 100% identical) would check that everything like the train test split is really fixed.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1131299,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "12/29/2020 16:38:33",
      "content": "<p>Are you setting seeds? You have to make sure, that everything is exactly the same for a second run, check if you are maybe shuffeling your folds.</p>\n<p>Are you using early-stopping?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1131329,
          "author_name": "jabertuhin",
          "author_url": "",
          "post_date": "12/29/2020 16:57:26",
          "content": "<p>Yes, I am using pytorch-lightnings <em>seed_everything</em> function. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1131856,
      "author_name": "plugin1689",
      "author_url": "",
      "post_date": "12/30/2020 03:17:46",
      "content": "<p>I can see similar results, but the difference is not so big as yours</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1131089": "I was using learning rate 1e-4 for one run and then tried PytorchLightning's learning rate finder to find a better learning rate.\nThe learning rate finder suggested me ~0.00025118 to use as a learning rate.\n\nMy model architecture has resnet50 with an additional linear layer.\n\n**For Learning Rate 1e-4:**\n\nAfte 21 epochs,\nloss=0.267\ntrain_loss=0.442, train_accuracy=0.793,\nvalid_loss=0.396, valid_accuracy=0.863,\naccuracy=0.863\n\n**For Learning Rate 0.00025118:**\n\nAfter 25 epochs,\nloss=0.691,\ntrain_loss=0.891, train_accuracy=0.621,\nvalid_loss=0.62, valid_accuracy=0.7,\naccuracy=0.7\n\nIs this drastic performance change normal? Or does this indicate bug in the code?\n\n*EDIT-1: Added total training loss for both scenarios. \n\n*UPDATE-1: There was a change in the weight decay, which I reverted and this hasn't made any difference.",
    "1131183": "This looks odd to me. With a higher learning rate of about 2.5e-4 and 25 epochs you end up getting a larger training loss than with a lower learning rate of 1e-4 and 21 epochs? That's the one thing I would have guessed I could predict (on the other hand, if the 2nd option did worse on the validation loss or accuracy, I would not necessarily think that was odd). To me that suggests something else is different/has gone wrong. E.g. your training-validation split might be different (or the optimizer? or the learning rate schedule?). The difference also seems quite striking to me, so I would guess it's more than what you'd expect due to randomness in the training, but it's hard to guess.",
    "1131299": "Are you setting seeds? You have to make sure, that everything is exactly the same for a second run, check if you are maybe shuffeling your folds.\n\nAre you using early-stopping?",
    "1131329": "Yes, I am using pytorch-lightnings *seed_everything* function.",
    "1131351": "I have updated my post, the wight decay was different in both run. \nSo, now I am training again with the first run's weight decay. But the total training loss is not decreasing.\nI am also guessing the issue might be in the train, test split.\nAny idea on how can I check whether my seeding is working perfectly?",
    "1131856": "I can see similar results, but the difference is not so big as yours",
    "1132134": "Running the same code twice with absolutely no changes (and seeing that results are then 100% identical) would check that everything like the train test split is really fixed."
  },
  "source": "meta"
}