{
  "id": 452581,
  "title": "How to replicate fastai's fit_one_cycle in pytorch?",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/452581",
  "author_name": "",
  "post_date": "2023-11-02T16:33:38.714513800Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>So, this code is from <a href=\"https://www.kaggle.com/code/iafoss/rna-starter-0-186-lb\" target=\"_blank\">this public notebook.</a></p>\n<pre><code>learn = Learner(data, model, loss_func=loss,cbs=[GradientClip()], metrics=[MAE()]).to_fp16() \nlearn.fit_one_cycle(, lr_max=, wd=, pct_start=)\n</code></pre>\n<p>I tried to convert this to pytorch like this:</p>\n<pre><code>optimizer = torch.optim.AdamW(model.parameters(), lr=lr, weight_decay=)\nscheduler = torch.optim.lr_scheduler.OneCycleLR(optimizer, max_lr=, steps_per_epoch=(train_loader), epochs=,  pct_start=, final_div_factor=)\n</code></pre>\n<p>But it produces much worse results. Any idea what I am doing wrong? I copy pasted the final_div_factor of 100000 from fastai's library code, because pytorch default was 10000.</p>\n<p>Thanks in advance. </p>",
  "messages": [
    {
      "id": "2509931",
      "postDate": "11/02/2023 16:33:38",
      "content": "<p>So, this code is from <a href=\"https://www.kaggle.com/code/iafoss/rna-starter-0-186-lb\" target=\"_blank\">this public notebook.</a></p>\n<pre><code>learn = Learner(data, model, loss_func=loss,cbs=[GradientClip()], metrics=[MAE()]).to_fp16() \nlearn.fit_one_cycle(, lr_max=, wd=, pct_start=)\n</code></pre>\n<p>I tried to convert this to pytorch like this:</p>\n<pre><code>optimizer = torch.optim.AdamW(model.parameters(), lr=lr, weight_decay=)\nscheduler = torch.optim.lr_scheduler.OneCycleLR(optimizer, max_lr=, steps_per_epoch=(train_loader), epochs=,  pct_start=, final_div_factor=)\n</code></pre>\n<p>But it produces much worse results. Any idea what I am doing wrong? I copy pasted the final_div_factor of 100000 from fastai's library code, because pytorch default was 10000.</p>\n<p>Thanks in advance. </p>",
      "rawMarkdown": "So, this code is from [this public notebook.](https://www.kaggle.com/code/iafoss/rna-starter-0-186-lb)\n\n```python\nlearn = Learner(data, model, loss_func=loss,cbs=[GradientClip(3.0)], metrics=[MAE()]).to_fp16() \nlearn.fit_one_cycle(32, lr_max=5e-4, wd=0.05, pct_start=0.02)\n```\n\nI tried to convert this to pytorch like this:\n\n```python\noptimizer = torch.optim.AdamW(model.parameters(), lr=lr, weight_decay=0.05)\nscheduler = torch.optim.lr_scheduler.OneCycleLR(optimizer, max_lr=5e-4, steps_per_epoch=len(train_loader), epochs=32,  pct_start=0.020, final_div_factor=100000)\n```\n\nBut it produces much worse results. Any idea what I am doing wrong? I copy pasted the final_div_factor of 100000 from fastai's library code, because pytorch default was 10000.\n\nThanks in advance.",
      "votes": null
    },
    {
      "id": "2509961",
      "postDate": "11/02/2023 16:50:36",
      "content": "<p>1:  Are you doing the gradient clipping at 3.0? Is it using the same clip method?<br>\n2: Are you running Pytorch in FP16? or another dtype?<br>\n3: Are you using AdamW in both repos?</p>",
      "rawMarkdown": "1:  Are you doing the gradient clipping at 3.0? Is it using the same clip method?\n2: Are you running Pytorch in FP16? or another dtype?\n3: Are you using AdamW in both repos?",
      "votes": null
    },
    {
      "id": "2510580",
      "postDate": "11/03/2023 05:40:58",
      "content": "<p>From the documentation - <br>\n\"The 1cycle learning rate policy changes the learning rate after every batch. step should be called after a batch has been used for training.\"  <br>\nSo scheduler.step() is after optimizer.step() under a <code>for batch in data_loader:</code>  Just in case yours is different.</p>\n<p>Also you could try set three_phase=True, <br>\n\"The default behaviour of this scheduler follows the fastai implementation of 1cycle, which claims that “unpublished work has shown even better results by using only two phases”. To mimic the behaviour of the original paper instead, set three_phase=True.\"</p>\n<p>Or see the code for <code>def lrfn</code> in <br>\n<a href=\"https://www.kaggle.com/code/shlomoron/srrf-transformer-tpu-training\" target=\"_blank\">https://www.kaggle.com/code/shlomoron/srrf-transformer-tpu-training</a><br>\nyou could try a custom learning rate scheduler similar to that. </p>",
      "rawMarkdown": "From the documentation - \n\"The 1cycle learning rate policy changes the learning rate after every batch. step should be called after a batch has been used for training.\"  \nSo scheduler.step() is after optimizer.step() under a ` for batch in data_loader:`  Just in case yours is different.\n\nAlso you could try set three_phase=True, \n\"The default behaviour of this scheduler follows the fastai implementation of 1cycle, which claims that “unpublished work has shown even better results by using only two phases”. To mimic the behaviour of the original paper instead, set three_phase=True.\"\n\nOr see the code for `def lrfn ` in \nhttps://www.kaggle.com/code/shlomoron/srrf-transformer-tpu-training\nyou could try a custom learning rate scheduler similar to that.",
      "votes": null
    },
    {
      "id": "2512704",
      "postDate": "11/04/2023 19:20:15",
      "content": "<p>As far as I remember, div_final in fastai has another meaning (in pytorch final_lr = max_lr / div / final_div, in fastai - final_lr = max_lr / final_div). <br>\nAlso, I observed some strange gradclipping behaviour in case of pytorch-lightning, if you use it, this might be an issue<br>\nNext, fastai automatically removes weight decay for biases and norms layers. Native pytorch and some pytorch-based frameworks leaves that up to you</p>",
      "rawMarkdown": "As far as I remember, div_final in fastai has another meaning (in pytorch final_lr = max_lr / div / final_div, in fastai - final_lr = max_lr / final_div). \nAlso, I observed some strange gradclipping behaviour in case of pytorch-lightning, if you use it, this might be an issue\nNext, fastai automatically removes weight decay for biases and norms layers. Native pytorch and some pytorch-based frameworks leaves that up to you",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2509961,
      "author_name": "themecheng",
      "author_url": "",
      "post_date": "11/02/2023 16:50:36",
      "content": "<p>1:  Are you doing the gradient clipping at 3.0? Is it using the same clip method?<br>\n2: Are you running Pytorch in FP16? or another dtype?<br>\n3: Are you using AdamW in both repos?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2510580,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "11/03/2023 05:40:58",
      "content": "<p>From the documentation - <br>\n\"The 1cycle learning rate policy changes the learning rate after every batch. step should be called after a batch has been used for training.\"  <br>\nSo scheduler.step() is after optimizer.step() under a <code>for batch in data_loader:</code>  Just in case yours is different.</p>\n<p>Also you could try set three_phase=True, <br>\n\"The default behaviour of this scheduler follows the fastai implementation of 1cycle, which claims that “unpublished work has shown even better results by using only two phases”. To mimic the behaviour of the original paper instead, set three_phase=True.\"</p>\n<p>Or see the code for <code>def lrfn</code> in <br>\n<a href=\"https://www.kaggle.com/code/shlomoron/srrf-transformer-tpu-training\" target=\"_blank\">https://www.kaggle.com/code/shlomoron/srrf-transformer-tpu-training</a><br>\nyou could try a custom learning rate scheduler similar to that. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2512704,
      "author_name": "dmitrypenzar1996",
      "author_url": "",
      "post_date": "11/04/2023 19:20:15",
      "content": "<p>As far as I remember, div_final in fastai has another meaning (in pytorch final_lr = max_lr / div / final_div, in fastai - final_lr = max_lr / final_div). <br>\nAlso, I observed some strange gradclipping behaviour in case of pytorch-lightning, if you use it, this might be an issue<br>\nNext, fastai automatically removes weight decay for biases and norms layers. Native pytorch and some pytorch-based frameworks leaves that up to you</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2509931": "So, this code is from [this public notebook.](https://www.kaggle.com/code/iafoss/rna-starter-0-186-lb)\n\n```python\nlearn = Learner(data, model, loss_func=loss,cbs=[GradientClip(3.0)], metrics=[MAE()]).to_fp16() \nlearn.fit_one_cycle(32, lr_max=5e-4, wd=0.05, pct_start=0.02)\n```\n\nI tried to convert this to pytorch like this:\n\n```python\noptimizer = torch.optim.AdamW(model.parameters(), lr=lr, weight_decay=0.05)\nscheduler = torch.optim.lr_scheduler.OneCycleLR(optimizer, max_lr=5e-4, steps_per_epoch=len(train_loader), epochs=32,  pct_start=0.020, final_div_factor=100000)\n```\n\nBut it produces much worse results. Any idea what I am doing wrong? I copy pasted the final_div_factor of 100000 from fastai's library code, because pytorch default was 10000.\n\nThanks in advance.",
    "2509961": "1:  Are you doing the gradient clipping at 3.0? Is it using the same clip method?\n2: Are you running Pytorch in FP16? or another dtype?\n3: Are you using AdamW in both repos?",
    "2510580": "From the documentation - \n\"The 1cycle learning rate policy changes the learning rate after every batch. step should be called after a batch has been used for training.\"  \nSo scheduler.step() is after optimizer.step() under a ` for batch in data_loader:`  Just in case yours is different.\n\nAlso you could try set three_phase=True, \n\"The default behaviour of this scheduler follows the fastai implementation of 1cycle, which claims that “unpublished work has shown even better results by using only two phases”. To mimic the behaviour of the original paper instead, set three_phase=True.\"\n\nOr see the code for `def lrfn ` in \nhttps://www.kaggle.com/code/shlomoron/srrf-transformer-tpu-training\nyou could try a custom learning rate scheduler similar to that.",
    "2512704": "As far as I remember, div_final in fastai has another meaning (in pytorch final_lr = max_lr / div / final_div, in fastai - final_lr = max_lr / final_div). \nAlso, I observed some strange gradclipping behaviour in case of pytorch-lightning, if you use it, this might be an issue\nNext, fastai automatically removes weight decay for biases and norms layers. Native pytorch and some pytorch-based frameworks leaves that up to you"
  },
  "source": "meta"
}