{
  "id": 104383,
  "title": "[off-topic] New State of the Art AI Optimizer: Rectified Adam (RAdam)",
  "url": "/competitions/aptos2019-blindness-detection/discussion/104383",
  "author_name": "Bibek",
  "post_date": "2019-08-16T13:54:51.248000",
  "votes": 17,
  "comment_count": 15,
  "views": 0,
  "content": "<p>This looks interesting!! mayb help ur models for better convergence :)\n<a href=\"https://medium.com/@lessw/new-state-of-the-art-ai-optimizer-rectified-adam-radam-5d854730807b\">Recified Adam</a>\n<a href=\"https://github.com/LiyuanLucasLiu/RAdam\">Code</a>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1528571%2F7097cd9adca9ec0d66f8ff5dff0daa06%2Flr.jpeg?generation=1565963628491533&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 600724,
      "postDate": "2019-08-16T13:54:51.247Z",
      "content": "<p>This looks interesting!! mayb help ur models for better convergence :)\n<a href=\"https://medium.com/@lessw/new-state-of-the-art-ai-optimizer-rectified-adam-radam-5d854730807b\">Recified Adam</a>\n<a href=\"https://github.com/LiyuanLucasLiu/RAdam\">Code</a>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1528571%2F7097cd9adca9ec0d66f8ff5dff0daa06%2Flr.jpeg?generation=1565963628491533&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "This looks interesting!! mayb help ur models for better convergence :)\n[Recified Adam](https://medium.com/@lessw/new-state-of-the-art-ai-optimizer-rectified-adam-radam-5d854730807b)\n[Code](https://github.com/LiyuanLucasLiu/RAdam)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1528571%2F7097cd9adca9ec0d66f8ff5dff0daa06%2Flr.jpeg?generation=1565963628491533&amp;alt=media)\n",
      "votes": 16
    },
    {
      "id": 600888,
      "postDate": "2019-08-16T17:35:40.627Z",
      "content": "<p>Thanks for posting, wouldn't have seen this otherwise :)</p>",
      "rawMarkdown": "Thanks for posting, wouldn't have seen this otherwise :)",
      "votes": 5
    },
    {
      "id": 604616,
      "postDate": "2019-08-21T15:21:26.217Z",
      "content": "<p>Out of curiosity, has anyone tried <a href=\"https://medium.com/@lessw/new-deep-learning-optimizer-ranger-synergistic-combination-of-radam-lookahead-for-the-best-of-2dc83f79a48d\">the Ranger optimizer instead</a>? The original medium poster added an update that they now prefer Ranger.</p>",
      "rawMarkdown": "Out of curiosity, has anyone tried [the Ranger optimizer instead](https://medium.com/@lessw/new-deep-learning-optimizer-ranger-synergistic-combination-of-radam-lookahead-for-the-best-of-2dc83f79a48d)? The original medium poster added an update that they now prefer Ranger.",
      "votes": 3,
      "replies": [
        {
          "id": 604694,
          "postDate": "2019-08-21T16:56:04.287Z",
          "content": "<p>Thank u so much for sharing!! will be interesting to see how the results vary with this new optimizer</p>",
          "rawMarkdown": "Thank u so much for sharing!! will be interesting to see how the results vary with this new optimizer",
          "votes": 1
        }
      ]
    },
    {
      "id": 602135,
      "postDate": "2019-08-18T16:24:49.070Z",
      "content": "<p>I've tested this new optimizer to my model and got slightly worse results, even after trying to optimize hyperparameters.</p>",
      "rawMarkdown": "I've tested this new optimizer to my model and got slightly worse results, even after trying to optimize hyperparameters.",
      "votes": 1,
      "replies": [
        {
          "id": 602151,
          "postDate": "2019-08-18T16:52:42.010Z",
          "content": "<p>Same here. I am using Fastai, which has mostly great default hyperparameters, so it might happen that the library is tuned for Adam, thus giving poor results for RAdam in my case.</p>",
          "rawMarkdown": "Same here. I am using Fastai, which has mostly great default hyperparameters, so it might happen that the library is tuned for Adam, thus giving poor results for RAdam in my case."
        },
        {
          "id": 602397,
          "postDate": "2019-08-19T03:32:08.753Z",
          "content": "<p>Same here. 😢\nThose recent optimizer like radam, amsgrad insist they are better, but it doesn’t work in my case. Maybe task specific</p>",
          "rawMarkdown": "Same here. 😢\nThose recent optimizer like radam, amsgrad insist they are better, but it doesn’t work in my case. Maybe task specific"
        }
      ]
    },
    {
      "id": 604128,
      "postDate": "2019-08-21T04:18:06.883Z",
      "content": "<p>I am still running some experiments with RAdam on this competition with Fastai. (Just giving some thoughts, nothing is firm yet). </p>\n\n<p>Couple things I noticed:</p>\n\n<ol>\n<li><p>I think if you use LR find, compare to the AdamW, you will see the learning rate curve shift to the right by a scale of 10. I try to use suggest lr by turning on suggestion= True, and use exactly min_loss divide by 10, this will give a much worse result</p></li>\n<li><p>Since in the paper they mentioned they still have a range of lr that will work (if you dont go over the limit), I tried use original lr (which is my baseline model with AdamW). It gives a 1% up in the local validation and about 0.02 boost in LB (local cv 92% to 93%, LB 0.798 to 0.800),  with rest of the hyper parameter / arugmentation same as baseline model.</p></li>\n</ol>\n\n<p>As you can see, I feel this is just some random weights change. </p>\n\n<p>Base on my observation so far</p>\n\n<ol>\n<li><p>With Efficient Net B4, if you use lr_find, don't increase the lr too much. Just as an example, with AdamW, the suggested lr is around 1e-4, and suggested lr with RAdam are in the range of (4e-3 to 8e-3). When using AdamW's lr with RAdam,  the result is similar in LB, but around 1% up in local validation </p></li>\n<li><p>I wonder if the reason RAdam is not working as expected is because I have fine tuned the model on 2015 dataset, so it didn't start from scratch. Maybe later I can try re-train the model on 2015 dataset with RAdam, maybe it will converge faster than AdamW?</p></li>\n</ol>\n\n<p>But yes, not significant improvement, not significant drop (if you don't up lr too much). But I will keep testing. </p>",
      "rawMarkdown": "I am still running some experiments with RAdam on this competition with Fastai. (Just giving some thoughts, nothing is firm yet). \n\nCouple things I noticed:\n\n1. I think if you use LR find, compare to the AdamW, you will see the learning rate curve shift to the right by a scale of 10. I try to use suggest lr by turning on suggestion= True, and use exactly min_loss divide by 10, this will give a much worse result\n\n2. Since in the paper they mentioned they still have a range of lr that will work (if you dont go over the limit), I tried use original lr (which is my baseline model with AdamW). It gives a 1% up in the local validation and about 0.02 boost in LB (local cv 92% to 93%, LB 0.798 to 0.800),  with rest of the hyper parameter / arugmentation same as baseline model.\n\nAs you can see, I feel this is just some random weights change. \n\nBase on my observation so far\n\n1. With Efficient Net B4, if you use lr_find, don't increase the lr too much. Just as an example, with AdamW, the suggested lr is around 1e-4, and suggested lr with RAdam are in the range of (4e-3 to 8e-3). When using AdamW's lr with RAdam,  the result is similar in LB, but around 1% up in local validation \n\n2. I wonder if the reason RAdam is not working as expected is because I have fine tuned the model on 2015 dataset, so it didn't start from scratch. Maybe later I can try re-train the model on 2015 dataset with RAdam, maybe it will converge faster than AdamW?\n\nBut yes, not significant improvement, not significant drop (if you don't up lr too much). But I will keep testing. ",
      "votes": 2
    },
    {
      "id": 602885,
      "postDate": "2019-08-19T16:00:19.323Z",
      "content": "<p>not converge faster than AdamW in fastai . Anyone get better with RAdam?</p>",
      "rawMarkdown": "not converge faster than AdamW in fastai . Anyone get better with RAdam?",
      "votes": 2
    },
    {
      "id": 602060,
      "postDate": "2019-08-18T14:07:47.150Z",
      "content": "<p>Yes, it's awesome! Thanks for sharing and not off-topic it all! This will likely become the new standard over Vanilla Adam.</p>\n\n<p>CyberZHG also shared an implementation of RAdam for Keras. Check it out!\n<a href=\"https://github.com/CyberZHG/keras-radam/blob/master/keras_radam/optimizers.py\">https://github.com/CyberZHG/keras-radam/blob/master/keras_radam/optimizers.py</a></p>\n\n<p>Example implementation in a Kaggle kernel (APTOS 2019):\n<a href=\"https://www.kaggle.com/carlolepelaars/efficientnetb5-with-keras-aptos-2019?scriptVersionId=19071596\">https://www.kaggle.com/carlolepelaars/efficientnetb5-with-keras-aptos-2019?scriptVersionId=19071596</a></p>",
      "rawMarkdown": "Yes, it's awesome! Thanks for sharing and not off-topic it all! This will likely become the new standard over Vanilla Adam.\n\nCyberZHG also shared an implementation of RAdam for Keras. Check it out!\n[https://github.com/CyberZHG/keras-radam/blob/master/keras_radam/optimizers.py](https://github.com/CyberZHG/keras-radam/blob/master/keras_radam/optimizers.py)\n\nExample implementation in a Kaggle kernel (APTOS 2019):\nhttps://www.kaggle.com/carlolepelaars/efficientnetb5-with-keras-aptos-2019?scriptVersionId=19071596",
      "votes": 2
    },
    {
      "id": 604076,
      "postDate": "2019-08-21T02:25:05.613Z",
      "content": "<p>For me, this is giving significantly worse results that standard Adam(with minimal hyp adjustment)for some reason</p>",
      "rawMarkdown": "For me, this is giving significantly worse results that standard Adam(with minimal hyp adjustment)for some reason"
    },
    {
      "id": 602649,
      "postDate": "2019-08-19T10:15:51.673Z",
      "content": "<p>Maybe I don't know how to use - but I got slightly worse results both on validation and LB compared to Adam.</p>",
      "rawMarkdown": "Maybe I don't know how to use - but I got slightly worse results both on validation and LB compared to Adam."
    },
    {
      "id": 601451,
      "postDate": "2019-08-17T16:19:48.880Z",
      "rawMarkdown": "",
      "votes": 3,
      "isDeleted": true
    },
    {
      "id": 600740,
      "postDate": "2019-08-16T14:07:00.680Z",
      "rawMarkdown": "",
      "votes": -6,
      "isDeleted": true,
      "replies": [
        {
          "id": 600750,
          "postDate": "2019-08-16T14:13:23.073Z",
          "content": "<p>Not everybody participates in all competitions!! I posted this for those who are not aware of this new optimizer. </p>",
          "rawMarkdown": "Not everybody participates in all competitions!! I posted this for those who are not aware of this new optimizer. ",
          "votes": 8
        },
        {
          "id": 600754,
          "postDate": "2019-08-16T14:17:26.390Z",
          "rawMarkdown": "",
          "votes": -5,
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 600888,
      "author_name": "Tom Aindow",
      "author_url": "",
      "post_date": "2019-08-16T17:35:40.627000",
      "content": "<p>Thanks for posting, wouldn't have seen this otherwise :)</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 604616,
      "author_name": "Sohier Dane",
      "author_url": "",
      "post_date": "2019-08-21T15:21:26.217000",
      "content": "<p>Out of curiosity, has anyone tried <a href=\"https://medium.com/@lessw/new-deep-learning-optimizer-ranger-synergistic-combination-of-radam-lookahead-for-the-best-of-2dc83f79a48d\">the Ranger optimizer instead</a>? The original medium poster added an update that they now prefer Ranger.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 604694,
          "author_name": "Bibek",
          "author_url": "",
          "post_date": "2019-08-21T16:56:04.287000",
          "content": "<p>Thank u so much for sharing!! will be interesting to see how the results vary with this new optimizer</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 602135,
      "author_name": "Ildoo Kim",
      "author_url": "",
      "post_date": "2019-08-18T16:24:49.070000",
      "content": "<p>I've tested this new optimizer to my model and got slightly worse results, even after trying to optimize hyperparameters.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 602151,
          "author_name": "Salil Mishra",
          "author_url": "",
          "post_date": "2019-08-18T16:52:42.010000",
          "content": "<p>Same here. I am using Fastai, which has mostly great default hyperparameters, so it might happen that the library is tuned for Adam, thus giving poor results for RAdam in my case.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 602397,
          "author_name": "hahaha",
          "author_url": "",
          "post_date": "2019-08-19T03:32:08.753000",
          "content": "<p>Same here. 😢\nThose recent optimizer like radam, amsgrad insist they are better, but it doesn’t work in my case. Maybe task specific</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 604128,
      "author_name": "Hao He",
      "author_url": "",
      "post_date": "2019-08-21T04:18:06.883000",
      "content": "<p>I am still running some experiments with RAdam on this competition with Fastai. (Just giving some thoughts, nothing is firm yet). </p>\n\n<p>Couple things I noticed:</p>\n\n<ol>\n<li><p>I think if you use LR find, compare to the AdamW, you will see the learning rate curve shift to the right by a scale of 10. I try to use suggest lr by turning on suggestion= True, and use exactly min_loss divide by 10, this will give a much worse result</p></li>\n<li><p>Since in the paper they mentioned they still have a range of lr that will work (if you dont go over the limit), I tried use original lr (which is my baseline model with AdamW). It gives a 1% up in the local validation and about 0.02 boost in LB (local cv 92% to 93%, LB 0.798 to 0.800),  with rest of the hyper parameter / arugmentation same as baseline model.</p></li>\n</ol>\n\n<p>As you can see, I feel this is just some random weights change. </p>\n\n<p>Base on my observation so far</p>\n\n<ol>\n<li><p>With Efficient Net B4, if you use lr_find, don't increase the lr too much. Just as an example, with AdamW, the suggested lr is around 1e-4, and suggested lr with RAdam are in the range of (4e-3 to 8e-3). When using AdamW's lr with RAdam,  the result is similar in LB, but around 1% up in local validation </p></li>\n<li><p>I wonder if the reason RAdam is not working as expected is because I have fine tuned the model on 2015 dataset, so it didn't start from scratch. Maybe later I can try re-train the model on 2015 dataset with RAdam, maybe it will converge faster than AdamW?</p></li>\n</ol>\n\n<p>But yes, not significant improvement, not significant drop (if you don't up lr too much). But I will keep testing. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 602885,
      "author_name": "CC Joshua ",
      "author_url": "",
      "post_date": "2019-08-19T16:00:19.323000",
      "content": "<p>not converge faster than AdamW in fastai . Anyone get better with RAdam?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 602060,
      "author_name": "Carlo",
      "author_url": "",
      "post_date": "2019-08-18T14:07:47.150000",
      "content": "<p>Yes, it's awesome! Thanks for sharing and not off-topic it all! This will likely become the new standard over Vanilla Adam.</p>\n\n<p>CyberZHG also shared an implementation of RAdam for Keras. Check it out!\n<a href=\"https://github.com/CyberZHG/keras-radam/blob/master/keras_radam/optimizers.py\">https://github.com/CyberZHG/keras-radam/blob/master/keras_radam/optimizers.py</a></p>\n\n<p>Example implementation in a Kaggle kernel (APTOS 2019):\n<a href=\"https://www.kaggle.com/carlolepelaars/efficientnetb5-with-keras-aptos-2019?scriptVersionId=19071596\">https://www.kaggle.com/carlolepelaars/efficientnetb5-with-keras-aptos-2019?scriptVersionId=19071596</a></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 604076,
      "author_name": "sh",
      "author_url": "",
      "post_date": "2019-08-21T02:25:05.613000",
      "content": "<p>For me, this is giving significantly worse results that standard Adam(with minimal hyp adjustment)for some reason</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 602649,
      "author_name": "Georgi Pamukov",
      "author_url": "",
      "post_date": "2019-08-19T10:15:51.673000",
      "content": "<p>Maybe I don't know how to use - but I got slightly worse results both on validation and LB compared to Adam.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 601451,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-08-17T16:19:48.880000",
      "content": "",
      "votes": 3,
      "replies": []
    },
    {
      "id": 600740,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-08-16T14:07:00.680000",
      "content": "",
      "votes": -6,
      "replies": [
        {
          "id": 600750,
          "author_name": "Bibek",
          "author_url": "",
          "post_date": "2019-08-16T14:13:23.073000",
          "content": "<p>Not everybody participates in all competitions!! I posted this for those who are not aware of this new optimizer. </p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 600754,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-08-16T14:17:26.390000",
          "content": "",
          "votes": -5,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "600724": "This looks interesting!! mayb help ur models for better convergence :)\n[Recified Adam](https://medium.com/@lessw/new-state-of-the-art-ai-optimizer-rectified-adam-radam-5d854730807b)\n[Code](https://github.com/LiyuanLucasLiu/RAdam)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1528571%2F7097cd9adca9ec0d66f8ff5dff0daa06%2Flr.jpeg?generation=1565963628491533&amp;alt=media)\n",
    "600888": "Thanks for posting, wouldn't have seen this otherwise :)",
    "604616": "Out of curiosity, has anyone tried [the Ranger optimizer instead](https://medium.com/@lessw/new-deep-learning-optimizer-ranger-synergistic-combination-of-radam-lookahead-for-the-best-of-2dc83f79a48d)? The original medium poster added an update that they now prefer Ranger.",
    "602135": "I've tested this new optimizer to my model and got slightly worse results, even after trying to optimize hyperparameters.",
    "604128": "I am still running some experiments with RAdam on this competition with Fastai. (Just giving some thoughts, nothing is firm yet). \n\nCouple things I noticed:\n\n1. I think if you use LR find, compare to the AdamW, you will see the learning rate curve shift to the right by a scale of 10. I try to use suggest lr by turning on suggestion= True, and use exactly min_loss divide by 10, this will give a much worse result\n\n2. Since in the paper they mentioned they still have a range of lr that will work (if you dont go over the limit), I tried use original lr (which is my baseline model with AdamW). It gives a 1% up in the local validation and about 0.02 boost in LB (local cv 92% to 93%, LB 0.798 to 0.800),  with rest of the hyper parameter / arugmentation same as baseline model.\n\nAs you can see, I feel this is just some random weights change. \n\nBase on my observation so far\n\n1. With Efficient Net B4, if you use lr_find, don't increase the lr too much. Just as an example, with AdamW, the suggested lr is around 1e-4, and suggested lr with RAdam are in the range of (4e-3 to 8e-3). When using AdamW's lr with RAdam,  the result is similar in LB, but around 1% up in local validation \n\n2. I wonder if the reason RAdam is not working as expected is because I have fine tuned the model on 2015 dataset, so it didn't start from scratch. Maybe later I can try re-train the model on 2015 dataset with RAdam, maybe it will converge faster than AdamW?\n\nBut yes, not significant improvement, not significant drop (if you don't up lr too much). But I will keep testing. ",
    "602885": "not converge faster than AdamW in fastai . Anyone get better with RAdam?",
    "602060": "Yes, it's awesome! Thanks for sharing and not off-topic it all! This will likely become the new standard over Vanilla Adam.\n\nCyberZHG also shared an implementation of RAdam for Keras. Check it out!\n[https://github.com/CyberZHG/keras-radam/blob/master/keras_radam/optimizers.py](https://github.com/CyberZHG/keras-radam/blob/master/keras_radam/optimizers.py)\n\nExample implementation in a Kaggle kernel (APTOS 2019):\nhttps://www.kaggle.com/carlolepelaars/efficientnetb5-with-keras-aptos-2019?scriptVersionId=19071596",
    "604076": "For me, this is giving significantly worse results that standard Adam(with minimal hyp adjustment)for some reason",
    "602649": "Maybe I don't know how to use - but I got slightly worse results both on validation and LB compared to Adam.",
    "601451": "",
    "600740": ""
  }
}