{
  "id": 464755,
  "title": "anyone with any luck on OneCycleLR?",
  "url": "/competitions/blood-vessel-segmentation/discussion/464755",
  "author_name": "Sumo",
  "post_date": "2024-01-01T11:07:04.500000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p>hi, I'm trying out the OneCycleLR where I initially got really excited about their concept of \"super convergence\". To do so I set up the following</p>\n<ul>\n<li>perform their Learning Rate Range Test (LRRT), plotting the loss vs the learning rate to pick the maximum learning rate.</li>\n<li>then from that plot, set the learning rate where the val surface dice is the highest as the max lr for my OneCycleLR scheduler</li>\n<li>I tried tinkering with a couple of settings (batch size, model architecture types) to make \"the amount of regularization must be balanced for each dataset<br>\nand architecture\" as per the paper's idea</li>\n</ul>\n<p>There are 2 main issues I find (this model is se_resnext101)</p>\n<ul>\n<li>the LRRT doesn't really follow the paper's nice curves, where the val surface dice kinds of just bounces back up when the lr was increased to insane values (like 2+)</li>\n<li>so I tried using just the first peak to be the max LR (which is already high: 1.38, but since the paper is called \"Super-Convergence: Very Fast Training of Neural<br>\nNetworks Using Large Learning Rates\" I figured I'd give it a chance). But it wasn't better than my baseline configurations.</li>\n</ul>\n<p>does anyone have some insights / past experiences on this? if so I have the following questions</p>\n<ul>\n<li>have anyone seen this scheduler gives meaningful improvements at all, is it something worth integrating + tuning over?</li>\n<li>how does one pick the max iterations to run the scheduler for? like in this case I pick some arbitrary values like 2 epochs, but if there are knowledge in how one knows when to increase/decrease this value it'd be nice</li>\n</ul>\n<p>thank you!</p>",
  "messages": [
    {
      "id": 2582096,
      "postDate": "2024-01-01T11:07:04.500Z",
      "content": "<p>hi, I'm trying out the OneCycleLR where I initially got really excited about their concept of \"super convergence\". To do so I set up the following</p>\n<ul>\n<li>perform their Learning Rate Range Test (LRRT), plotting the loss vs the learning rate to pick the maximum learning rate.</li>\n<li>then from that plot, set the learning rate where the val surface dice is the highest as the max lr for my OneCycleLR scheduler</li>\n<li>I tried tinkering with a couple of settings (batch size, model architecture types) to make \"the amount of regularization must be balanced for each dataset<br>\nand architecture\" as per the paper's idea</li>\n</ul>\n<p>There are 2 main issues I find (this model is se_resnext101)</p>\n<ul>\n<li>the LRRT doesn't really follow the paper's nice curves, where the val surface dice kinds of just bounces back up when the lr was increased to insane values (like 2+)</li>\n<li>so I tried using just the first peak to be the max LR (which is already high: 1.38, but since the paper is called \"Super-Convergence: Very Fast Training of Neural<br>\nNetworks Using Large Learning Rates\" I figured I'd give it a chance). But it wasn't better than my baseline configurations.</li>\n</ul>\n<p>does anyone have some insights / past experiences on this? if so I have the following questions</p>\n<ul>\n<li>have anyone seen this scheduler gives meaningful improvements at all, is it something worth integrating + tuning over?</li>\n<li>how does one pick the max iterations to run the scheduler for? like in this case I pick some arbitrary values like 2 epochs, but if there are knowledge in how one knows when to increase/decrease this value it'd be nice</li>\n</ul>\n<p>thank you!</p>",
      "rawMarkdown": "hi, I'm trying out the OneCycleLR where I initially got really excited about their concept of \"super convergence\". To do so I set up the following\n- perform their Learning Rate Range Test (LRRT), plotting the loss vs the learning rate to pick the maximum learning rate.\n- then from that plot, set the learning rate where the val surface dice is the highest as the max lr for my OneCycleLR scheduler\n- I tried tinkering with a couple of settings (batch size, model architecture types) to make \"the amount of regularization must be balanced for each dataset\nand architecture\" as per the paper's idea\n\nThere are 2 main issues I find (this model is se_resnext101)\n- the LRRT doesn't really follow the paper's nice curves, where the val surface dice kinds of just bounces back up when the lr was increased to insane values (like 2+)\n- so I tried using just the first peak to be the max LR (which is already high: 1.38, but since the paper is called \"Super-Convergence: Very Fast Training of Neural\nNetworks Using Large Learning Rates\" I figured I'd give it a chance). But it wasn't better than my baseline configurations.\n\ndoes anyone have some insights / past experiences on this? if so I have the following questions\n- have anyone seen this scheduler gives meaningful improvements at all, is it something worth integrating + tuning over?\n- how does one pick the max iterations to run the scheduler for? like in this case I pick some arbitrary values like 2 epochs, but if there are knowledge in how one knows when to increase/decrease this value it'd be nice\n\nthank you!"
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2582096": "hi, I'm trying out the OneCycleLR where I initially got really excited about their concept of \"super convergence\". To do so I set up the following\n- perform their Learning Rate Range Test (LRRT), plotting the loss vs the learning rate to pick the maximum learning rate.\n- then from that plot, set the learning rate where the val surface dice is the highest as the max lr for my OneCycleLR scheduler\n- I tried tinkering with a couple of settings (batch size, model architecture types) to make \"the amount of regularization must be balanced for each dataset\nand architecture\" as per the paper's idea\n\nThere are 2 main issues I find (this model is se_resnext101)\n- the LRRT doesn't really follow the paper's nice curves, where the val surface dice kinds of just bounces back up when the lr was increased to insane values (like 2+)\n- so I tried using just the first peak to be the max LR (which is already high: 1.38, but since the paper is called \"Super-Convergence: Very Fast Training of Neural\nNetworks Using Large Learning Rates\" I figured I'd give it a chance). But it wasn't better than my baseline configurations.\n\ndoes anyone have some insights / past experiences on this? if so I have the following questions\n- have anyone seen this scheduler gives meaningful improvements at all, is it something worth integrating + tuning over?\n- how does one pick the max iterations to run the scheduler for? like in this case I pick some arbitrary values like 2 epochs, but if there are knowledge in how one knows when to increase/decrease this value it'd be nice\n\nthank you!"
  }
}