{
  "id": 392942,
  "title": "The new Lion optimizer",
  "url": "/competitions/asl-signs/discussion/392942",
  "author_name": "",
  "post_date": "2023-03-07T11:18:00.015620600Z",
  "votes": 8,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi to all! <br>\nHere is an implementation of the new lion optimizer for tensorflow.<br>\n<a href=\"https://github.com/OrigamiDream/lion-tf\" target=\"_blank\">https://github.com/OrigamiDream/lion-tf</a>. <br>\nI tried - training really happens faster than with Adam. But the final result plus minus is the same. There is an implementation in public notebooks.<br>\nAt the end of the training, I do not see any improvements.<br>\nWhat are your thoughts on this optimizer?</p>",
  "messages": [
    {
      "id": "2172213",
      "postDate": "03/07/2023 11:18:00",
      "content": "<p>Hi to all! <br>\nHere is an implementation of the new lion optimizer for tensorflow.<br>\n<a href=\"https://github.com/OrigamiDream/lion-tf\" target=\"_blank\">https://github.com/OrigamiDream/lion-tf</a>. <br>\nI tried - training really happens faster than with Adam. But the final result plus minus is the same. There is an implementation in public notebooks.<br>\nAt the end of the training, I do not see any improvements.<br>\nWhat are your thoughts on this optimizer?</p>",
      "rawMarkdown": "Hi to all! \nHere is an implementation of the new lion optimizer for tensorflow.\nhttps://github.com/OrigamiDream/lion-tf. \nI tried - training really happens faster than with Adam. But the final result plus minus is the same. There is an implementation in public notebooks.\nAt the end of the training, I do not see any improvements.\nWhat are your thoughts on this optimizer?",
      "votes": null
    },
    {
      "id": "2172658",
      "postDate": "03/07/2023 17:40:22",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/aikhmelnytskyy\" target=\"_blank\">@aikhmelnytskyy</a>,</p>\n<p>Thanks for the head's up.  I had heard about the paper, but hadn't read it yet.  When you say…</p>\n<blockquote>\n  <p>training really happens faster than with Adam</p>\n</blockquote>\n<p>…do you mean fewer epochs or less time per epoch?</p>\n<p>I've just had a quick play with it myself.  My initial experience is that…</p>\n<ol>\n<li>It takes almost exactly the same amount of time per epoch (with a P100).</li>\n<li>It converges more slowly (i.e. takes more epochs).  For some concrete data points, Adam first achieved a loss of &lt;1.4 at epoch 16 whilst Lion first achieved the same at epoch 30.  Adam achieved &lt;1.0 at epoch 25 and Lion achieved it at epoch 53.</li>\n<li>It converges to a markedly worse end state (loss of 0.8114 with Lion vs 0.3965 with Adam) at the point where the validation loss is no longer improving.</li>\n</ol>\n<p>Doesn't feel like one I'm going to be sticking with.</p>\n<hr>\n<p>I also see from the paper abstract that it goes faster with bigger batch sizes.  Obviously pretty much all optimizers work faster with bigger batch sizes.  So, whilst there was an marked improvement in timings with bigger batch sizes, the same improvement was seen for both Adam and Lion.</p>\n<p>Additionally, I see that Lion claims to work better with a lower LR.  I reduced the initial LR to 1e-4 (Lion's default) and this improved the end state loss to 0.4968 which is clearly much better than before, but still a long way off Adam's performance.</p>",
      "rawMarkdown": "Hi @aikhmelnytskyy,\n\nThanks for the head's up.  I had heard about the paper, but hadn't read it yet.  When you say...\n\n> training really happens faster than with Adam\n\n...do you mean fewer epochs or less time per epoch?\n\nI've just had a quick play with it myself.  My initial experience is that...\n\n1. It takes almost exactly the same amount of time per epoch (with a P100).\n1. It converges more slowly (i.e. takes more epochs).  For some concrete data points, Adam first achieved a loss of <1.4 at epoch 16 whilst Lion first achieved the same at epoch 30.  Adam achieved <1.0 at epoch 25 and Lion achieved it at epoch 53.\n1. It converges to a markedly worse end state (loss of 0.8114 with Lion vs 0.3965 with Adam) at the point where the validation loss is no longer improving.\n\nDoesn't feel like one I'm going to be sticking with.\n\n---\n\nI also see from the paper abstract that it goes faster with bigger batch sizes.  Obviously pretty much all optimizers work faster with bigger batch sizes.  So, whilst there was an marked improvement in timings with bigger batch sizes, the same improvement was seen for both Adam and Lion.\n\nAdditionally, I see that Lion claims to work better with a lower LR.  I reduced the initial LR to 1e-4 (Lion's default) and this improved the end state loss to 0.4968 which is clearly much better than before, but still a long way off Adam's performance.",
      "votes": null
    },
    {
      "id": "2172681",
      "postDate": "03/07/2023 17:54:40",
      "content": "<p>Hi! I meant that he showed better results than adam in the first 20 epochs. I ran the tuning using KerasTuner. So lion showed better results at 20 epochs, but at the end of training (after 100 epochs) the results were worse. I ran some more experiments on more epochs on KerasTuner: now adam shows better results. Lion story reminds me of effv2 story: On paper effv2 is better than effv1 in practice no.</p>",
      "rawMarkdown": "Hi! I meant that he showed better results than adam in the first 20 epochs. I ran the tuning using KerasTuner. So lion showed better results at 20 epochs, but at the end of training (after 100 epochs) the results were worse. I ran some more experiments on more epochs on KerasTuner: now adam shows better results. Lion story reminds me of effv2 story: On paper effv2 is better than effv1 in practice no.",
      "votes": null
    },
    {
      "id": "2173053",
      "postDate": "03/08/2023 03:50:55",
      "content": "<p><a href=\"https://github.com/keras-team/keras/pull/17605\" target=\"_blank\">https://github.com/keras-team/keras/pull/17605</a></p>",
      "rawMarkdown": "https://github.com/keras-team/keras/pull/17605",
      "votes": null
    },
    {
      "id": "2255552",
      "postDate": "05/11/2023 19:41:41",
      "content": "<p>Doesn't work close to AdamW. Nevertheless, it's been introduced to packages like bitsandbytes. Maybe there are some use cases where it works</p>",
      "rawMarkdown": "Doesn't work close to AdamW. Nevertheless, it's been introduced to packages like bitsandbytes. Maybe there are some use cases where it works",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2172658,
      "author_name": "andrewrrose",
      "author_url": "",
      "post_date": "03/07/2023 17:40:22",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/aikhmelnytskyy\" target=\"_blank\">@aikhmelnytskyy</a>,</p>\n<p>Thanks for the head's up.  I had heard about the paper, but hadn't read it yet.  When you say…</p>\n<blockquote>\n  <p>training really happens faster than with Adam</p>\n</blockquote>\n<p>…do you mean fewer epochs or less time per epoch?</p>\n<p>I've just had a quick play with it myself.  My initial experience is that…</p>\n<ol>\n<li>It takes almost exactly the same amount of time per epoch (with a P100).</li>\n<li>It converges more slowly (i.e. takes more epochs).  For some concrete data points, Adam first achieved a loss of &lt;1.4 at epoch 16 whilst Lion first achieved the same at epoch 30.  Adam achieved &lt;1.0 at epoch 25 and Lion achieved it at epoch 53.</li>\n<li>It converges to a markedly worse end state (loss of 0.8114 with Lion vs 0.3965 with Adam) at the point where the validation loss is no longer improving.</li>\n</ol>\n<p>Doesn't feel like one I'm going to be sticking with.</p>\n<hr>\n<p>I also see from the paper abstract that it goes faster with bigger batch sizes.  Obviously pretty much all optimizers work faster with bigger batch sizes.  So, whilst there was an marked improvement in timings with bigger batch sizes, the same improvement was seen for both Adam and Lion.</p>\n<p>Additionally, I see that Lion claims to work better with a lower LR.  I reduced the initial LR to 1e-4 (Lion's default) and this improved the end state loss to 0.4968 which is clearly much better than before, but still a long way off Adam's performance.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2172681,
          "author_name": "aikhmelnytskyy",
          "author_url": "",
          "post_date": "03/07/2023 17:54:40",
          "content": "<p>Hi! I meant that he showed better results than adam in the first 20 epochs. I ran the tuning using KerasTuner. So lion showed better results at 20 epochs, but at the end of training (after 100 epochs) the results were worse. I ran some more experiments on more epochs on KerasTuner: now adam shows better results. Lion story reminds me of effv2 story: On paper effv2 is better than effv1 in practice no.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2173053,
      "author_name": "ipythonx",
      "author_url": "",
      "post_date": "03/08/2023 03:50:55",
      "content": "<p><a href=\"https://github.com/keras-team/keras/pull/17605\" target=\"_blank\">https://github.com/keras-team/keras/pull/17605</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2255552,
      "author_name": "anttip",
      "author_url": "",
      "post_date": "05/11/2023 19:41:41",
      "content": "<p>Doesn't work close to AdamW. Nevertheless, it's been introduced to packages like bitsandbytes. Maybe there are some use cases where it works</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2172213": "Hi to all! \nHere is an implementation of the new lion optimizer for tensorflow.\nhttps://github.com/OrigamiDream/lion-tf. \nI tried - training really happens faster than with Adam. But the final result plus minus is the same. There is an implementation in public notebooks.\nAt the end of the training, I do not see any improvements.\nWhat are your thoughts on this optimizer?",
    "2172658": "Hi @aikhmelnytskyy,\n\nThanks for the head's up.  I had heard about the paper, but hadn't read it yet.  When you say...\n\n> training really happens faster than with Adam\n\n...do you mean fewer epochs or less time per epoch?\n\nI've just had a quick play with it myself.  My initial experience is that...\n\n1. It takes almost exactly the same amount of time per epoch (with a P100).\n1. It converges more slowly (i.e. takes more epochs).  For some concrete data points, Adam first achieved a loss of <1.4 at epoch 16 whilst Lion first achieved the same at epoch 30.  Adam achieved <1.0 at epoch 25 and Lion achieved it at epoch 53.\n1. It converges to a markedly worse end state (loss of 0.8114 with Lion vs 0.3965 with Adam) at the point where the validation loss is no longer improving.\n\nDoesn't feel like one I'm going to be sticking with.\n\n---\n\nI also see from the paper abstract that it goes faster with bigger batch sizes.  Obviously pretty much all optimizers work faster with bigger batch sizes.  So, whilst there was an marked improvement in timings with bigger batch sizes, the same improvement was seen for both Adam and Lion.\n\nAdditionally, I see that Lion claims to work better with a lower LR.  I reduced the initial LR to 1e-4 (Lion's default) and this improved the end state loss to 0.4968 which is clearly much better than before, but still a long way off Adam's performance.",
    "2172681": "Hi! I meant that he showed better results than adam in the first 20 epochs. I ran the tuning using KerasTuner. So lion showed better results at 20 epochs, but at the end of training (after 100 epochs) the results were worse. I ran some more experiments on more epochs on KerasTuner: now adam shows better results. Lion story reminds me of effv2 story: On paper effv2 is better than effv1 in practice no.",
    "2173053": "https://github.com/keras-team/keras/pull/17605",
    "2255552": "Doesn't work close to AdamW. Nevertheless, it's been introduced to packages like bitsandbytes. Maybe there are some use cases where it works"
  },
  "source": "meta"
}