{
  "id": 227513,
  "title": "A note on LR schedulers",
  "url": "/competitions/bms-molecular-translation/discussion/227513",
  "author_name": "",
  "post_date": "2021-03-20T21:55:42.096219900Z",
  "votes": 17,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I'm sure I'm not the only one using <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a>'s excellent PyTorch notebook utilizing ResNet and LSTM (<a href=\"https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-starter\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-starter</a>) as a starting point for further work. Given that, I thought I would share an observation about the LR scheduler setup.</p>\n<p>My runs kept giving their best result after epoch 5 and epoch 6 was always worse, at which point I stopped them. I looked more carefully at the code and found the reason was the LR scheduler. It is set to use CosineAnnealingLR and therefore the learning rate goes down and back up again. The script is setup to use either ReduceLROnPlateau, CosineAnnealingLR or CosineAnnealingWarmRestarts - which would be best? To experiment quickly I trimmed the training data to just 50K entries (40K train, 10K val) and ran 10 epochs with each of the 3 schedulers configured in the script.</p>\n<p>ReduceLROnPlateau turned out to be best, followed by CosineAnnealingWarmRestarts. CosineAnnealingLR did recover, giving better results after epoch 10 that at epoch 5 (I should have let my runs continue) but was the worst of the 3.</p>\n<p>Of course I cannot be sure that my little experiment is representative when run on the full training set. But I think I'll use ReduceLROnPlateau for my next full run.</p>",
  "messages": [
    {
      "id": "1246587",
      "postDate": "03/20/2021 21:55:42",
      "content": "<p>I'm sure I'm not the only one using <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a>'s excellent PyTorch notebook utilizing ResNet and LSTM (<a href=\"https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-starter\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-starter</a>) as a starting point for further work. Given that, I thought I would share an observation about the LR scheduler setup.</p>\n<p>My runs kept giving their best result after epoch 5 and epoch 6 was always worse, at which point I stopped them. I looked more carefully at the code and found the reason was the LR scheduler. It is set to use CosineAnnealingLR and therefore the learning rate goes down and back up again. The script is setup to use either ReduceLROnPlateau, CosineAnnealingLR or CosineAnnealingWarmRestarts - which would be best? To experiment quickly I trimmed the training data to just 50K entries (40K train, 10K val) and ran 10 epochs with each of the 3 schedulers configured in the script.</p>\n<p>ReduceLROnPlateau turned out to be best, followed by CosineAnnealingWarmRestarts. CosineAnnealingLR did recover, giving better results after epoch 10 that at epoch 5 (I should have let my runs continue) but was the worst of the 3.</p>\n<p>Of course I cannot be sure that my little experiment is representative when run on the full training set. But I think I'll use ReduceLROnPlateau for my next full run.</p>",
      "rawMarkdown": "I'm sure I'm not the only one using @yasufuminakama's excellent PyTorch notebook utilizing ResNet and LSTM (https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-starter) as a starting point for further work. Given that, I thought I would share an observation about the LR scheduler setup.\n\nMy runs kept giving their best result after epoch 5 and epoch 6 was always worse, at which point I stopped them. I looked more carefully at the code and found the reason was the LR scheduler. It is set to use CosineAnnealingLR and therefore the learning rate goes down and back up again. The script is setup to use either ReduceLROnPlateau, CosineAnnealingLR or CosineAnnealingWarmRestarts - which would be best? To experiment quickly I trimmed the training data to just 50K entries (40K train, 10K val) and ran 10 epochs with each of the 3 schedulers configured in the script.\n\nReduceLROnPlateau turned out to be best, followed by CosineAnnealingWarmRestarts. CosineAnnealingLR did recover, giving better results after epoch 10 that at epoch 5 (I should have let my runs continue) but was the worst of the 3.\n\nOf course I cannot be sure that my little experiment is representative when run on the full training set. But I think I'll use ReduceLROnPlateau for my next full run.",
      "votes": null
    },
    {
      "id": "1247320",
      "postDate": "03/21/2021 16:17:39",
      "content": "<p>I haven't tried in this competition. But based on my previous experience, I think you should choose the scheduler more carefully as hyper-parameters are different.</p>\n<p>For example, It's normal that epoch 6 is always worse than epoch 5 for CosineAnnealingLR if the T_max is 5. But you may get a better result in epoch 10, epoch 15. So you should use more epochs.<br>\nFor ReduceLROnPlateau, you may get a better result by using a large learning rate.</p>\n<p>For me, I will test the LR schedulers with different hyper-parameters first and select the best one among all the experiments😁.</p>",
      "rawMarkdown": "I haven't tried in this competition. But based on my previous experience, I think you should choose the scheduler more carefully as hyper-parameters are different.\n\nFor example, It's normal that epoch 6 is always worse than epoch 5 for CosineAnnealingLR if the T_max is 5. But you may get a better result in epoch 10, epoch 15. So you should use more epochs.\nFor ReduceLROnPlateau, you may get a better result by using a large learning rate.\n\nFor me, I will test the LR schedulers with different hyper-parameters first and select the best one among all the experiments😁.",
      "votes": null
    },
    {
      "id": "1248927",
      "postDate": "03/22/2021 23:30:36",
      "content": "<p>Thanks for sharing your experience. Yes, in my short experiment with a very small sub-set of the training data, I did get better results with CosineAnnealingLR by waiting until epoch 10 - it was definitely a mistake to stop my earlier runs early.</p>\n<p>Yes, I plan to run more experiments - although I need to find the right balance between enough training data to be realistic and experiments that run fast enough to be useful.</p>",
      "rawMarkdown": "Thanks for sharing your experience. Yes, in my short experiment with a very small sub-set of the training data, I did get better results with CosineAnnealingLR by waiting until epoch 10 - it was definitely a mistake to stop my earlier runs early.\n\nYes, I plan to run more experiments - although I need to find the right balance between enough training data to be realistic and experiments that run fast enough to be useful.",
      "votes": null
    },
    {
      "id": "1249049",
      "postDate": "03/23/2021 04:02:49",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/tonyxu\" target=\"_blank\">@tonyxu</a> and <a href=\"https://www.kaggle.com/andypenrose\" target=\"_blank\">@andypenrose</a>,<br>\nI have a few doubts related to your experiments.</p>\n<ol>\n<li>Approximately how much data do you pick for your experiments?</li>\n<li>Do the results of experimental data and the whole data match?</li>\n</ol>\n<p>Any insights will be really helpful…. </p>",
      "rawMarkdown": "Hi @tonyxu and @andypenrose,\nI have a few doubts related to your experiments.\n1. Approximately how much data do you pick for your experiments?\n2. Do the results of experimental data and the whole data match?\n\nAny insights will be really helpful....",
      "votes": null
    },
    {
      "id": "1250278",
      "postDate": "03/23/2021 22:50:01",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/nitindatta\" target=\"_blank\">@nitindatta</a>, for my quick experiment I trimmed the training data to just 50K entries (40K train, 10K val). Will the best LR scheduler for that tiny training set over 10 epochs be best for the full training set? I have my doubts as well! If I have runtime later, I may try an intermediate test with more data.</p>",
      "rawMarkdown": "Hi @nitindatta, for my quick experiment I trimmed the training data to just 50K entries (40K train, 10K val). Will the best LR scheduler for that tiny training set over 10 epochs be best for the full training set? I have my doubts as well! If I have runtime later, I may try an intermediate test with more data.",
      "votes": null
    },
    {
      "id": "1274230",
      "postDate": "04/15/2021 05:59:15",
      "content": "<p>Noob Question: Should we use schedulers in validation part?</p>",
      "rawMarkdown": "Noob Question: Should we use schedulers in validation part?",
      "votes": null
    },
    {
      "id": "1274848",
      "postDate": "04/15/2021 17:01:52",
      "content": "<p>No. There is no need during validation. The LR schedulers allow control over the learning rate during the course of training, and therefore are not relevant during validation.</p>",
      "rawMarkdown": "No. There is no need during validation. The LR schedulers allow control over the learning rate during the course of training, and therefore are not relevant during validation.",
      "votes": null
    },
    {
      "id": "1300101",
      "postDate": "05/10/2021 09:13:47",
      "content": "<p>Thanks for your work, I wonder the detailed setting of your ReduceLROnPlateau scheduler, In my case ReduceLROnPlateau is not as good as CosineAnnealingLR in the first 6 epoch up to now, I wonder the reasons.</p>",
      "rawMarkdown": "Thanks for your work, I wonder the detailed setting of your ReduceLROnPlateau scheduler, In my case ReduceLROnPlateau is not as good as CosineAnnealingLR in the first 6 epoch up to now, I wonder the reasons.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1247320,
      "author_name": "tonyxu",
      "author_url": "",
      "post_date": "03/21/2021 16:17:39",
      "content": "<p>I haven't tried in this competition. But based on my previous experience, I think you should choose the scheduler more carefully as hyper-parameters are different.</p>\n<p>For example, It's normal that epoch 6 is always worse than epoch 5 for CosineAnnealingLR if the T_max is 5. But you may get a better result in epoch 10, epoch 15. So you should use more epochs.<br>\nFor ReduceLROnPlateau, you may get a better result by using a large learning rate.</p>\n<p>For me, I will test the LR schedulers with different hyper-parameters first and select the best one among all the experiments😁.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1248927,
          "author_name": "andypenrose",
          "author_url": "",
          "post_date": "03/22/2021 23:30:36",
          "content": "<p>Thanks for sharing your experience. Yes, in my short experiment with a very small sub-set of the training data, I did get better results with CosineAnnealingLR by waiting until epoch 10 - it was definitely a mistake to stop my earlier runs early.</p>\n<p>Yes, I plan to run more experiments - although I need to find the right balance between enough training data to be realistic and experiments that run fast enough to be useful.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1249049,
          "author_name": "nitindatta",
          "author_url": "",
          "post_date": "03/23/2021 04:02:49",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/tonyxu\" target=\"_blank\">@tonyxu</a> and <a href=\"https://www.kaggle.com/andypenrose\" target=\"_blank\">@andypenrose</a>,<br>\nI have a few doubts related to your experiments.</p>\n<ol>\n<li>Approximately how much data do you pick for your experiments?</li>\n<li>Do the results of experimental data and the whole data match?</li>\n</ol>\n<p>Any insights will be really helpful…. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1250278,
          "author_name": "andypenrose",
          "author_url": "",
          "post_date": "03/23/2021 22:50:01",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/nitindatta\" target=\"_blank\">@nitindatta</a>, for my quick experiment I trimmed the training data to just 50K entries (40K train, 10K val). Will the best LR scheduler for that tiny training set over 10 epochs be best for the full training set? I have my doubts as well! If I have runtime later, I may try an intermediate test with more data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1274230,
      "author_name": "mrinath",
      "author_url": "",
      "post_date": "04/15/2021 05:59:15",
      "content": "<p>Noob Question: Should we use schedulers in validation part?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1274848,
          "author_name": "andypenrose",
          "author_url": "",
          "post_date": "04/15/2021 17:01:52",
          "content": "<p>No. There is no need during validation. The LR schedulers allow control over the learning rate during the course of training, and therefore are not relevant during validation.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1300101,
      "author_name": "huangsiyi",
      "author_url": "",
      "post_date": "05/10/2021 09:13:47",
      "content": "<p>Thanks for your work, I wonder the detailed setting of your ReduceLROnPlateau scheduler, In my case ReduceLROnPlateau is not as good as CosineAnnealingLR in the first 6 epoch up to now, I wonder the reasons.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1246587": "I'm sure I'm not the only one using @yasufuminakama's excellent PyTorch notebook utilizing ResNet and LSTM (https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-starter) as a starting point for further work. Given that, I thought I would share an observation about the LR scheduler setup.\n\nMy runs kept giving their best result after epoch 5 and epoch 6 was always worse, at which point I stopped them. I looked more carefully at the code and found the reason was the LR scheduler. It is set to use CosineAnnealingLR and therefore the learning rate goes down and back up again. The script is setup to use either ReduceLROnPlateau, CosineAnnealingLR or CosineAnnealingWarmRestarts - which would be best? To experiment quickly I trimmed the training data to just 50K entries (40K train, 10K val) and ran 10 epochs with each of the 3 schedulers configured in the script.\n\nReduceLROnPlateau turned out to be best, followed by CosineAnnealingWarmRestarts. CosineAnnealingLR did recover, giving better results after epoch 10 that at epoch 5 (I should have let my runs continue) but was the worst of the 3.\n\nOf course I cannot be sure that my little experiment is representative when run on the full training set. But I think I'll use ReduceLROnPlateau for my next full run.",
    "1247320": "I haven't tried in this competition. But based on my previous experience, I think you should choose the scheduler more carefully as hyper-parameters are different.\n\nFor example, It's normal that epoch 6 is always worse than epoch 5 for CosineAnnealingLR if the T_max is 5. But you may get a better result in epoch 10, epoch 15. So you should use more epochs.\nFor ReduceLROnPlateau, you may get a better result by using a large learning rate.\n\nFor me, I will test the LR schedulers with different hyper-parameters first and select the best one among all the experiments😁.",
    "1248927": "Thanks for sharing your experience. Yes, in my short experiment with a very small sub-set of the training data, I did get better results with CosineAnnealingLR by waiting until epoch 10 - it was definitely a mistake to stop my earlier runs early.\n\nYes, I plan to run more experiments - although I need to find the right balance between enough training data to be realistic and experiments that run fast enough to be useful.",
    "1249049": "Hi @tonyxu and @andypenrose,\nI have a few doubts related to your experiments.\n1. Approximately how much data do you pick for your experiments?\n2. Do the results of experimental data and the whole data match?\n\nAny insights will be really helpful....",
    "1250278": "Hi @nitindatta, for my quick experiment I trimmed the training data to just 50K entries (40K train, 10K val). Will the best LR scheduler for that tiny training set over 10 epochs be best for the full training set? I have my doubts as well! If I have runtime later, I may try an intermediate test with more data.",
    "1274230": "Noob Question: Should we use schedulers in validation part?",
    "1274848": "No. There is no need during validation. The LR schedulers allow control over the learning rate during the course of training, and therefore are not relevant during validation.",
    "1300101": "Thanks for your work, I wonder the detailed setting of your ReduceLROnPlateau scheduler, In my case ReduceLROnPlateau is not as good as CosineAnnealingLR in the first 6 epoch up to now, I wonder the reasons."
  },
  "source": "meta"
}