{
  "id": 207290,
  "title": "New Optimizers?",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/207290",
  "author_name": "Debarshi Chanda",
  "post_date": "2020-12-29T04:37:19.884000",
  "votes": 16,
  "comment_count": 22,
  "views": 0,
  "content": "<p>Has anyone tried optimizers other than Adam or SGD momentum? If yes, do share your experiments.<br>\n<a href=\"https://github.com/jettify/pytorch-optimizer\" target=\"_blank\">torch-optimizer</a> - This package contains implementation of a lot of optimizers</p>\n<p>Read about <a href=\"https://arxiv.org/abs/1902.09843\" target=\"_blank\">Adabound</a> some time back </p>\n<blockquote>\n  <p>An optimizer that trains as fast as Adam and as good as SGD for developing state-of-the-art deep learning models on a wide variety of popular tasks in the field of CV, NLP, etc.</p>\n</blockquote>\n<p>Official Repo - <a href=\"https://github.com/Luolc/AdaBound\" target=\"_blank\">https://github.com/Luolc/AdaBound</a></p>\n<p>I haven't tried any except Adam yet.</p>",
  "messages": [
    {
      "id": 1130463,
      "postDate": "2020-12-29T04:37:19.883Z",
      "content": "<p>Has anyone tried optimizers other than Adam or SGD momentum? If yes, do share your experiments.<br>\n<a href=\"https://github.com/jettify/pytorch-optimizer\" target=\"_blank\">torch-optimizer</a> - This package contains implementation of a lot of optimizers</p>\n<p>Read about <a href=\"https://arxiv.org/abs/1902.09843\" target=\"_blank\">Adabound</a> some time back </p>\n<blockquote>\n  <p>An optimizer that trains as fast as Adam and as good as SGD for developing state-of-the-art deep learning models on a wide variety of popular tasks in the field of CV, NLP, etc.</p>\n</blockquote>\n<p>Official Repo - <a href=\"https://github.com/Luolc/AdaBound\" target=\"_blank\">https://github.com/Luolc/AdaBound</a></p>\n<p>I haven't tried any except Adam yet.</p>",
      "rawMarkdown": "Has anyone tried optimizers other than Adam or SGD momentum? If yes, do share your experiments.\n[torch-optimizer](https://github.com/jettify/pytorch-optimizer) - This package contains implementation of a lot of optimizers\n\nRead about [Adabound](https://arxiv.org/abs/1902.09843) some time back \n> An optimizer that trains as fast as Adam and as good as SGD for developing state-of-the-art deep learning models on a wide variety of popular tasks in the field of CV, NLP, etc.\n\nOfficial Repo - [https://github.com/Luolc/AdaBound](https://github.com/Luolc/AdaBound)\n\nI haven't tried any except Adam yet.",
      "votes": 16
    },
    {
      "id": 1130684,
      "postDate": "2020-12-29T08:25:55.360Z",
      "content": "<p>yeah ranger is doing pretty good actually, i am using it too</p>",
      "rawMarkdown": "yeah ranger is doing pretty good actually, i am using it too",
      "votes": 3
    },
    {
      "id": 1130528,
      "postDate": "2020-12-29T05:42:54.257Z",
      "content": "<p>AdamW and Ranger are my go-to optimizers and here Ranger has done pretty decently.</p>",
      "rawMarkdown": "AdamW and Ranger are my go-to optimizers and here Ranger has done pretty decently.",
      "votes": 3,
      "replies": [
        {
          "id": 1130607,
          "postDate": "2020-12-29T06:57:14.860Z",
          "content": "<p>Never heard about Ranger! I Will surely give it a try</p>",
          "rawMarkdown": "Never heard about Ranger! I Will surely give it a try",
          "votes": 1
        }
      ]
    },
    {
      "id": 1147486,
      "postDate": "2021-01-10T14:38:44.817Z",
      "content": "<p>Try AdamP and SGDP, I haven’t tried it here yet but worked well in other problems.</p>\n<p><a href=\"https://arxiv.org/abs/2006.08217\" target=\"_blank\">https://arxiv.org/abs/2006.08217</a><br>\n<a href=\"https://github.com/clovaai/AdamP\" target=\"_blank\">https://github.com/clovaai/AdamP</a><br>\n<a href=\"https://catalyst-team.github.io/catalyst/api/contrib.html#module-catalyst.contrib.nn.optimizers.adamp\" target=\"_blank\">https://catalyst-team.github.io/catalyst/api/contrib.html#module-catalyst.contrib.nn.optimizers.adamp</a></p>",
      "rawMarkdown": "Try AdamP and SGDP, I haven’t tried it here yet but worked well in other problems.\n\nhttps://arxiv.org/abs/2006.08217\nhttps://github.com/clovaai/AdamP\nhttps://catalyst-team.github.io/catalyst/api/contrib.html#module-catalyst.contrib.nn.optimizers.adamp\n",
      "votes": 1,
      "replies": [
        {
          "id": 1147569,
          "postDate": "2021-01-10T15:31:54.503Z",
          "content": "<p>Thank you for sharing</p>",
          "rawMarkdown": "Thank you for sharing"
        }
      ]
    },
    {
      "id": 1136768,
      "postDate": "2021-01-03T12:06:13.570Z",
      "content": "<p>Thank you for share <a href=\"https://www.kaggle.com/debarshichanda\" target=\"_blank\">@debarshichanda</a> . For Keras implementation Check  <a href=\"https://github.com/titu1994/keras-adabound\" target=\"_blank\">This</a>.</p>",
      "rawMarkdown": "Thank you for share @debarshichanda . For Keras implementation Check  [This](https://github.com/titu1994/keras-adabound).",
      "votes": 1,
      "replies": [
        {
          "id": 1139438,
          "postDate": "2021-01-05T11:49:37.913Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1130778,
      "postDate": "2020-12-29T09:52:32.867Z",
      "content": "<p>Is there a faster method to experiment with different optimizers without training on whole data? That would speed up the process.</p>",
      "rawMarkdown": "Is there a faster method to experiment with different optimizers without training on whole data? That would speed up the process.",
      "votes": 1,
      "replies": [
        {
          "id": 1130875,
          "postDate": "2020-12-29T11:14:06.747Z",
          "content": "<p>There's the option of training in small networks (e.g. EfficientNet B0, B1 or B2). That's what I've done for experimentation and I am taking a bet that the same things will work for B5 to B7 that I might use at the end.</p>",
          "rawMarkdown": "There's the option of training in small networks (e.g. EfficientNet B0, B1 or B2). That's what I've done for experimentation and I am taking a bet that the same things will work for B5 to B7 that I might use at the end.",
          "votes": 2
        },
        {
          "id": 1130894,
          "postDate": "2020-12-29T11:50:04.647Z",
          "content": "<p>that's a good point, I will try it</p>",
          "rawMarkdown": "that's a good point, I will try it"
        },
        {
          "id": 1130897,
          "postDate": "2020-12-29T11:51:59.993Z",
          "content": "<p><a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a> Looks like a good idea!</p>",
          "rawMarkdown": "@bjoernholzhauer Looks like a good idea!"
        },
        {
          "id": 1131352,
          "postDate": "2020-12-29T17:12:04.453Z",
          "content": "<p>If you scroll down on the page mentioned in this topic: <a href=\"https://github.com/jettify/pytorch-optimizer\" target=\"_blank\">https://github.com/jettify/pytorch-optimizer</a></p>\n<p>The creator does also compare the optimizers based on a set amount of iteration (501 iterations)</p>\n<p>You can do the same for your model with different optimizers (set an amount of iterations and check the loss) or on certain <code>batch-sizes</code> train your model for N certain batches with an optimizer and check which has the lowest loss.</p>",
          "rawMarkdown": "If you scroll down on the page mentioned in this topic: https://github.com/jettify/pytorch-optimizer\n\nThe creator does also compare the optimizers based on a set amount of iteration (501 iterations)\n\nYou can do the same for your model with different optimizers (set an amount of iterations and check the loss) or on certain `batch-sizes` train your model for N certain batches with an optimizer and check which has the lowest loss.",
          "votes": 4
        }
      ]
    },
    {
      "id": 1137427,
      "postDate": "2021-01-03T23:21:45.727Z",
      "content": "<p>One optimizer claimed to be better than AdaBound is the AdaBelief <a href=\"https://juntang-zhuang.github.io/adabelief/\" target=\"_blank\">https://juntang-zhuang.github.io/adabelief/</a><br>\nIt just came out in NeurIPS 2020</p>",
      "rawMarkdown": "One optimizer claimed to be better than AdaBound is the AdaBelief https://juntang-zhuang.github.io/adabelief/\nIt just came out in NeurIPS 2020",
      "votes": 2,
      "replies": [
        {
          "id": 1137439,
          "postDate": "2021-01-03T23:42:09.473Z",
          "content": "<p>Thank you!<br>\nThis looks promising</p>",
          "rawMarkdown": "Thank you!\nThis looks promising"
        },
        {
          "id": 1140363,
          "postDate": "2021-01-06T00:33:44.353Z",
          "content": "<p>There is a Keras compatible update if anyone is looking for as per Author's response here: <a href=\"https://github.com/juntang-zhuang/Adabelief-Optimizer/issues/2#issuecomment-721498692\" target=\"_blank\">https://github.com/juntang-zhuang/Adabelief-Optimizer/issues/2#issuecomment-721498692</a></p>\n<p>Link: <a href=\"https://github.com/juntang-zhuang/Adabelief-Optimizer\" target=\"_blank\">https://github.com/juntang-zhuang/Adabelief-Optimizer</a></p>",
          "rawMarkdown": "There is a Keras compatible update if anyone is looking for as per Author's response here: https://github.com/juntang-zhuang/Adabelief-Optimizer/issues/2#issuecomment-721498692\n\nLink: https://github.com/juntang-zhuang/Adabelief-Optimizer",
          "votes": 1
        }
      ]
    },
    {
      "id": 1133761,
      "postDate": "2020-12-31T14:03:01.370Z",
      "content": "<p>For me, Adam work's well. Haven't tried others.<br>\nIn case anyone wants Keras code for <code>ranger</code> optimizer:</p>\n<pre><code>radam = tfa.optimizers.RectifiedAdam()\nranger = tfa.optimizers.Lookahead(radam, sync_period=6, slow_step_size=0.5)\n</code></pre>\n<p>You can find it <a href=\"https://www.tensorflow.org/addons/api_docs/python/tfa/optimizers/RectifiedAdam\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "For me, Adam work's well. Haven't tried others.\nIn case anyone wants Keras code for `ranger` optimizer:\n\n```\nradam = tfa.optimizers.RectifiedAdam()\nranger = tfa.optimizers.Lookahead(radam, sync_period=6, slow_step_size=0.5)\n```\nYou can find it [here](https://www.tensorflow.org/addons/api_docs/python/tfa/optimizers/RectifiedAdam)",
      "votes": 2
    },
    {
      "id": 1131435,
      "postDate": "2020-12-29T18:12:27.903Z",
      "content": "<p>any implementation of Ranger on Pytorch?<br>\nthanks</p>",
      "rawMarkdown": "any implementation of Ranger on Pytorch?\nthanks",
      "replies": [
        {
          "id": 1131444,
          "postDate": "2020-12-29T18:17:43.760Z",
          "content": "<p>You can find it under: <a href=\"https://github.com/lessw2020/Ranger-Deep-Learning-Optimizer\" target=\"_blank\">https://github.com/lessw2020/Ranger-Deep-Learning-Optimizer</a></p>",
          "rawMarkdown": "You can find it under: https://github.com/lessw2020/Ranger-Deep-Learning-Optimizer",
          "votes": 2
        },
        {
          "id": 1131452,
          "postDate": "2020-12-29T18:22:55.587Z",
          "content": "<p>This package also has its implementation <br>\n<a href=\"https://github.com/jettify/pytorch-optimizer\" target=\"_blank\">https://github.com/jettify/pytorch-optimizer</a></p>",
          "rawMarkdown": "This package also has its implementation \n[https://github.com/jettify/pytorch-optimizer](https://github.com/jettify/pytorch-optimizer)",
          "votes": 1
        }
      ]
    },
    {
      "id": 1130571,
      "postDate": "2020-12-29T06:07:37.090Z",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing.",
      "votes": 1
    },
    {
      "id": 1131506,
      "postDate": "2020-12-29T19:15:18.563Z",
      "content": "<p>nice. Thanks for the share.</p>",
      "rawMarkdown": "nice. Thanks for the share.",
      "votes": 2
    },
    {
      "id": 1133553,
      "postDate": "2020-12-31T09:56:30.953Z",
      "content": "<p>Thanks for the share.</p>",
      "rawMarkdown": "Thanks for the share."
    }
  ],
  "comments": [
    {
      "id": 1130684,
      "author_name": "Eren Tekin",
      "author_url": "",
      "post_date": "2020-12-29T08:25:55.360000",
      "content": "<p>yeah ranger is doing pretty good actually, i am using it too</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1130528,
      "author_name": "ilovescience",
      "author_url": "",
      "post_date": "2020-12-29T05:42:54.257000",
      "content": "<p>AdamW and Ranger are my go-to optimizers and here Ranger has done pretty decently.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1130607,
          "author_name": "Debarshi Chanda",
          "author_url": "",
          "post_date": "2020-12-29T06:57:14.860000",
          "content": "<p>Never heard about Ranger! I Will surely give it a try</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1147486,
      "author_name": "Kirderf",
      "author_url": "",
      "post_date": "2021-01-10T14:38:44.817000",
      "content": "<p>Try AdamP and SGDP, I haven’t tried it here yet but worked well in other problems.</p>\n<p><a href=\"https://arxiv.org/abs/2006.08217\" target=\"_blank\">https://arxiv.org/abs/2006.08217</a><br>\n<a href=\"https://github.com/clovaai/AdamP\" target=\"_blank\">https://github.com/clovaai/AdamP</a><br>\n<a href=\"https://catalyst-team.github.io/catalyst/api/contrib.html#module-catalyst.contrib.nn.optimizers.adamp\" target=\"_blank\">https://catalyst-team.github.io/catalyst/api/contrib.html#module-catalyst.contrib.nn.optimizers.adamp</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 1147569,
          "author_name": "Debarshi Chanda",
          "author_url": "",
          "post_date": "2021-01-10T15:31:54.503000",
          "content": "<p>Thank you for sharing</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1136768,
      "author_name": "AIFahim",
      "author_url": "",
      "post_date": "2021-01-03T12:06:13.570000",
      "content": "<p>Thank you for share <a href=\"https://www.kaggle.com/debarshichanda\" target=\"_blank\">@debarshichanda</a> . For Keras implementation Check  <a href=\"https://github.com/titu1994/keras-adabound\" target=\"_blank\">This</a>.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1139438,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-01-05T11:49:37.913000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1130778,
      "author_name": "Fares Lassoued",
      "author_url": "",
      "post_date": "2020-12-29T09:52:32.867000",
      "content": "<p>Is there a faster method to experiment with different optimizers without training on whole data? That would speed up the process.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1130875,
          "author_name": "Björn",
          "author_url": "",
          "post_date": "2020-12-29T11:14:06.747000",
          "content": "<p>There's the option of training in small networks (e.g. EfficientNet B0, B1 or B2). That's what I've done for experimentation and I am taking a bet that the same things will work for B5 to B7 that I might use at the end.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1130894,
          "author_name": "Fares Lassoued",
          "author_url": "",
          "post_date": "2020-12-29T11:50:04.647000",
          "content": "<p>that's a good point, I will try it</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1130897,
          "author_name": "Debarshi Chanda",
          "author_url": "",
          "post_date": "2020-12-29T11:51:59.993000",
          "content": "<p><a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a> Looks like a good idea!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1131352,
          "author_name": "Ali Abdin",
          "author_url": "",
          "post_date": "2020-12-29T17:12:04.453000",
          "content": "<p>If you scroll down on the page mentioned in this topic: <a href=\"https://github.com/jettify/pytorch-optimizer\" target=\"_blank\">https://github.com/jettify/pytorch-optimizer</a></p>\n<p>The creator does also compare the optimizers based on a set amount of iteration (501 iterations)</p>\n<p>You can do the same for your model with different optimizers (set an amount of iterations and check the loss) or on certain <code>batch-sizes</code> train your model for N certain batches with an optimizer and check which has the lowest loss.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1137427,
      "author_name": "Louis Yang",
      "author_url": "",
      "post_date": "2021-01-03T23:21:45.727000",
      "content": "<p>One optimizer claimed to be better than AdaBound is the AdaBelief <a href=\"https://juntang-zhuang.github.io/adabelief/\" target=\"_blank\">https://juntang-zhuang.github.io/adabelief/</a><br>\nIt just came out in NeurIPS 2020</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1137439,
          "author_name": "Debarshi Chanda",
          "author_url": "",
          "post_date": "2021-01-03T23:42:09.473000",
          "content": "<p>Thank you!<br>\nThis looks promising</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1140363,
          "author_name": "Dee Kay",
          "author_url": "",
          "post_date": "2021-01-06T00:33:44.353000",
          "content": "<p>There is a Keras compatible update if anyone is looking for as per Author's response here: <a href=\"https://github.com/juntang-zhuang/Adabelief-Optimizer/issues/2#issuecomment-721498692\" target=\"_blank\">https://github.com/juntang-zhuang/Adabelief-Optimizer/issues/2#issuecomment-721498692</a></p>\n<p>Link: <a href=\"https://github.com/juntang-zhuang/Adabelief-Optimizer\" target=\"_blank\">https://github.com/juntang-zhuang/Adabelief-Optimizer</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1133761,
      "author_name": "Prateek Mishra",
      "author_url": "",
      "post_date": "2020-12-31T14:03:01.370000",
      "content": "<p>For me, Adam work's well. Haven't tried others.<br>\nIn case anyone wants Keras code for <code>ranger</code> optimizer:</p>\n<pre><code>radam = tfa.optimizers.RectifiedAdam()\nranger = tfa.optimizers.Lookahead(radam, sync_period=6, slow_step_size=0.5)\n</code></pre>\n<p>You can find it <a href=\"https://www.tensorflow.org/addons/api_docs/python/tfa/optimizers/RectifiedAdam\" target=\"_blank\">here</a></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1131435,
      "author_name": "DeepUnderstanding",
      "author_url": "",
      "post_date": "2020-12-29T18:12:27.903000",
      "content": "<p>any implementation of Ranger on Pytorch?<br>\nthanks</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1131444,
          "author_name": "Ali Abdin",
          "author_url": "",
          "post_date": "2020-12-29T18:17:43.760000",
          "content": "<p>You can find it under: <a href=\"https://github.com/lessw2020/Ranger-Deep-Learning-Optimizer\" target=\"_blank\">https://github.com/lessw2020/Ranger-Deep-Learning-Optimizer</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1131452,
          "author_name": "Debarshi Chanda",
          "author_url": "",
          "post_date": "2020-12-29T18:22:55.587000",
          "content": "<p>This package also has its implementation <br>\n<a href=\"https://github.com/jettify/pytorch-optimizer\" target=\"_blank\">https://github.com/jettify/pytorch-optimizer</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1130571,
      "author_name": "Heroseo",
      "author_url": "",
      "post_date": "2020-12-29T06:07:37.090000",
      "content": "<p>Thanks for sharing.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1131506,
      "author_name": "Ashik M",
      "author_url": "",
      "post_date": "2020-12-29T19:15:18.563000",
      "content": "<p>nice. Thanks for the share.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1133553,
      "author_name": "Bhaskar Sarma",
      "author_url": "",
      "post_date": "2020-12-31T09:56:30.953000",
      "content": "<p>Thanks for the share.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1130463": "Has anyone tried optimizers other than Adam or SGD momentum? If yes, do share your experiments.\n[torch-optimizer](https://github.com/jettify/pytorch-optimizer) - This package contains implementation of a lot of optimizers\n\nRead about [Adabound](https://arxiv.org/abs/1902.09843) some time back \n> An optimizer that trains as fast as Adam and as good as SGD for developing state-of-the-art deep learning models on a wide variety of popular tasks in the field of CV, NLP, etc.\n\nOfficial Repo - [https://github.com/Luolc/AdaBound](https://github.com/Luolc/AdaBound)\n\nI haven't tried any except Adam yet.",
    "1130684": "yeah ranger is doing pretty good actually, i am using it too",
    "1130528": "AdamW and Ranger are my go-to optimizers and here Ranger has done pretty decently.",
    "1147486": "Try AdamP and SGDP, I haven’t tried it here yet but worked well in other problems.\n\nhttps://arxiv.org/abs/2006.08217\nhttps://github.com/clovaai/AdamP\nhttps://catalyst-team.github.io/catalyst/api/contrib.html#module-catalyst.contrib.nn.optimizers.adamp\n",
    "1136768": "Thank you for share @debarshichanda . For Keras implementation Check  [This](https://github.com/titu1994/keras-adabound).",
    "1130778": "Is there a faster method to experiment with different optimizers without training on whole data? That would speed up the process.",
    "1137427": "One optimizer claimed to be better than AdaBound is the AdaBelief https://juntang-zhuang.github.io/adabelief/\nIt just came out in NeurIPS 2020",
    "1133761": "For me, Adam work's well. Haven't tried others.\nIn case anyone wants Keras code for `ranger` optimizer:\n\n```\nradam = tfa.optimizers.RectifiedAdam()\nranger = tfa.optimizers.Lookahead(radam, sync_period=6, slow_step_size=0.5)\n```\nYou can find it [here](https://www.tensorflow.org/addons/api_docs/python/tfa/optimizers/RectifiedAdam)",
    "1131435": "any implementation of Ranger on Pytorch?\nthanks",
    "1130571": "Thanks for sharing.",
    "1131506": "nice. Thanks for the share.",
    "1133553": "Thanks for the share."
  }
}