{
  "id": 38505,
  "title": "RMSProp vs SGD vs Adam vs ...",
  "url": "/competitions/carvana-image-masking-challenge/discussion/38505",
  "author_name": "jackkwok",
  "post_date": "2017-08-24T04:45:41.111000",
  "votes": 11,
  "comment_count": 22,
  "views": 0,
  "content": "<p>Which optimization algo works best for training your UNet?</p>",
  "messages": [
    {
      "id": 216040,
      "postDate": "2017-08-24T04:45:41.113Z",
      "content": "<p>Which optimization algo works best for training your UNet?</p>",
      "rawMarkdown": "Which optimization algo works best for training your UNet?",
      "votes": 11
    },
    {
      "id": 216084,
      "postDate": "2017-08-24T08:31:48.267Z",
      "content": "<p>Adam seems to work better compared to SGD for me. I've yet to try RMSProp.</p>",
      "rawMarkdown": "Adam seems to work better compared to SGD for me. I've yet to try RMSProp.",
      "votes": 4
    },
    {
      "id": 216057,
      "postDate": "2017-08-24T06:36:29.733Z",
      "content": "<p>I got better results with RMSProp and SGD than with Adam. </p>",
      "rawMarkdown": "I got better results with RMSProp and SGD than with Adam. ",
      "votes": 2
    },
    {
      "id": 218363,
      "postDate": "2017-09-04T00:59:49.173Z",
      "content": "<p>Adam was the best for me so far. SGD came in second, with RMSprop in 3rd. I might return to RMSprop later since I didn't give it that much time</p>",
      "rawMarkdown": "Adam was the best for me so far. SGD came in second, with RMSprop in 3rd. I might return to RMSprop later since I didn't give it that much time",
      "votes": 1,
      "replies": [
        {
          "id": 218382,
          "postDate": "2017-09-04T02:26:16.587Z",
          "content": "<p>Could you please show your batch size and input size?</p>",
          "rawMarkdown": "Could you please show your batch size and input size?"
        },
        {
          "id": 218388,
          "postDate": "2017-09-04T03:04:18.923Z",
          "content": "<p>Sure thing. My batch size is 2 for training and 8 for inference, using bn. My input size is 1024x1024, using hq images. My hypothesis is that smaller batch sizes actually perform well here since pixel positional information is important in this challenge</p>",
          "rawMarkdown": "Sure thing. My batch size is 2 for training and 8 for inference, using bn. My input size is 1024x1024, using hq images. My hypothesis is that smaller batch sizes actually perform well here since pixel positional information is important in this challenge"
        },
        {
          "id": 221131,
          "postDate": "2017-09-14T08:04:25.857Z",
          "content": "<p>Have you ever tried BR layer. In my net , BR layer got 0.0002 higher LB score than BN layer.</p>",
          "rawMarkdown": "Have you ever tried BR layer. In my net , BR layer got 0.0002 higher LB score than BN layer.",
          "votes": 2
        },
        {
          "id": 221410,
          "postDate": "2017-09-15T06:54:00.470Z",
          "content": "<p>No I haven't actually <a href=\"https://arxiv.org/pdf/1702.03275.pdf\">https://arxiv.org/pdf/1702.03275.pdf</a></p>\n\n<p>I think this is a pretty good idea for several image recognition competitions, not just this one, since it appears the higher resolution models that sacrifice batch size do better from my experience</p>",
          "rawMarkdown": "No I haven't actually https://arxiv.org/pdf/1702.03275.pdf\n\nI think this is a pretty good idea for several image recognition competitions, not just this one, since it appears the higher resolution models that sacrifice batch size do better from my experience"
        }
      ]
    },
    {
      "id": 216154,
      "postDate": "2017-08-24T14:49:50.467Z",
      "content": "<p>Depending on your hardware and batch_size. You try and test it on your system and you will find the answer lay within your notebook.</p>",
      "rawMarkdown": "Depending on your hardware and batch_size. You try and test it on your system and you will find the answer lay within your notebook."
    },
    {
      "id": 217872,
      "postDate": "2017-09-01T09:52:58.027Z",
      "content": "<p><a href=\"http://ruder.io/optimizing-gradient-descent/index.html#gradientdescentoptimizationalgorithms\">http://ruder.io/optimizing-gradient-descent/index.html#gradientdescentoptimizationalgorithms</a></p>\n\n<p>I hope this helps to clear your understanding on optimizers.</p>",
      "rawMarkdown": "http://ruder.io/optimizing-gradient-descent/index.html#gradientdescentoptimizationalgorithms\n\nI hope this helps to clear your understanding on optimizers.",
      "replies": [
        {
          "id": 218364,
          "postDate": "2017-09-04T01:05:24.550Z",
          "content": "<p>Dont think the OP was asking for clarification, but i do think this is a good post here. thanks </p>",
          "rawMarkdown": "Dont think the OP was asking for clarification, but i do think this is a good post here. thanks "
        }
      ]
    },
    {
      "id": 1531533,
      "postDate": "2021-10-02T05:22:36.100Z",
      "content": "<p>see this : <a href=\"https://ruder.io/optimizing-gradient-descent/index.html\" target=\"_blank\">An overview of gradient descent optimization algorithms</a></p>",
      "rawMarkdown": "see this : [An overview of gradient descent optimization algorithms](https://ruder.io/optimizing-gradient-descent/index.html)"
    },
    {
      "id": 887486,
      "postDate": "2020-06-15T17:38:15.223Z",
      "content": "<p>For me, SGD work pretty good, if you face any issue with SGD try with bit low learning rate and use the value of momentum around 0.85 to 0.95</p>",
      "rawMarkdown": "For me, SGD work pretty good, if you face any issue with SGD try with bit low learning rate and use the value of momentum around 0.85 to 0.95"
    },
    {
      "id": 218601,
      "postDate": "2017-09-05T05:08:27.343Z",
      "content": "<p>rmsprop uses only sign and not magnitude of the computed gradients for descend.  A intermediate solution is gradient capping:</p>\n\n<pre><code>        loss.backward()\n        # accumulate gradients\n        if (it+1)%num_grad_acc==0:\n            torch.nn.utils.clip_grad_norm(net.parameters(), 1)\n\n            optimizer.step()\n            optimizer.zero_grad()  # assume no effects on bn for accumulating grad\n</code></pre>\n\n<p>I  find this effective in preventing oscillation in loss.</p>",
      "rawMarkdown": "rmsprop uses only sign and not magnitude of the computed gradients for descend.  A intermediate solution is gradient capping:\n\n            loss.backward()\n            # accumulate gradients\n            if (it+1)%num_grad_acc==0:\n                torch.nn.utils.clip_grad_norm(net.parameters(), 1)\n\n                optimizer.step()\n                optimizer.zero_grad()  # assume no effects on bn for accumulating grad\n\nI  find this effective in preventing oscillation in loss.",
      "replies": [
        {
          "id": 218794,
          "postDate": "2017-09-05T22:42:09.947Z",
          "content": "<p>This is for pytorch right? How do you do that with keras?</p>",
          "rawMarkdown": "This is for pytorch right? How do you do that with keras?"
        },
        {
          "id": 218811,
          "postDate": "2017-09-06T00:09:55.147Z",
          "content": "<blockquote>\n  <p>rmsprop uses only sign and not magnitude of the computed gradients for descend.</p>\n</blockquote>\n\n<p>I don't understand this well, so could you elaborate this mean?</p>\n\n<p>@YaGana Sheriff-Hussaini\nProbably we can get by this code </p>\n\n<p><code>\noptimizer = RMSprop(lr=0.0001, clipnorm=1.)\n</code></p>\n\n<p>I'm worried about that this gradient capping causes worse effect to learning. </p>\n\n<p>To get more stability of RMSprop, enlarging epsilon (1e-8 is keras default. change 1e-6.) is also a way.</p>",
          "rawMarkdown": " \n\n&gt; rmsprop uses only sign and not magnitude of the computed gradients for descend.\n\nI don't understand this well, so could you elaborate this mean?\n\n@YaGana Sheriff-Hussaini\nProbably we can get by this code \n\n```\noptimizer = RMSprop(lr=0.0001, clipnorm=1.)\n```\n\nI'm worried about that this gradient capping causes worse effect to learning. \n\nTo get more stability of RMSprop, enlarging epsilon (1e-8 is keras default. change 1e-6.) is also a way.",
          "votes": 3
        },
        {
          "id": 218873,
          "postDate": "2017-09-06T06:27:36.520Z",
          "content": "<p>@Heng, I have not used pytorch so I do not know how RMSprop is implemented there. But I took a look at the implementation of RMSprop in keras, and I do not think the oscillations you mentioned would be a problem. By definition, RMSprop uses oscillation reduction and momentum in tandem to achieve the desired effect.</p>\n\n<p>Thanks @Iyakaap. I have not tried it, but for those using Keras, I am also of the belief that it may very well result in an unwanted effect.</p>\n\n<p>In Keras, RMSprop is implemented according to <a href=\"http://www.cs.toronto.edu/~tijmen/csc321/slides/lecture_slides_lec6.pdf\">http://www.cs.toronto.edu/~tijmen/csc321/slides/lecture_slides_lec6.pdf</a></p>",
          "rawMarkdown": "@Heng, I have not used pytorch so I do not know how RMSprop is implemented there. But I took a look at the implementation of RMSprop in keras, and I do not think the oscillations you mentioned would be a problem. By definition, RMSprop uses oscillation reduction and momentum in tandem to achieve the desired effect.\n\nThanks @Iyakaap. I have not tried it, but for those using Keras, I am also of the belief that it may very well result in an unwanted effect.\n\nIn Keras, RMSprop is implemented according to http://www.cs.toronto.edu/~tijmen/csc321/slides/lecture_slides_lec6.pdf"
        },
        {
          "id": 219707,
          "postDate": "2017-09-09T08:23:08.863Z",
          "content": "<p>Why enlarging epsilon will make RMSprop more stable? Thanks for more explaining~ <br>\nAnd Keras is not recommending to change the eplison value. Any benefits can get according to  your experiment?</p>\n\n<p>RMSprop</p>\n\n<p>keras.optimizers.RMSprop(lr=0.001, rho=0.9, epsilon=1e-08, decay=0.0)\nRMSProp optimizer.</p>\n\n<p>It is recommended to leave the parameters of this optimizer at their default values (except the learning rate, which can be freely tuned).\nupdate:\nsame that epsilon here is added for non-zero dividing and so... it makes sense</p>",
          "rawMarkdown": "Why enlarging epsilon will make RMSprop more stable? Thanks for more explaining~  \nAnd Keras is not recommending to change the eplison value. Any benefits can get according to  your experiment?\n\nRMSprop\n\nkeras.optimizers.RMSprop(lr=0.001, rho=0.9, epsilon=1e-08, decay=0.0)\nRMSProp optimizer.\n\nIt is recommended to leave the parameters of this optimizer at their default values (except the learning rate, which can be freely tuned).\nupdate:\nsame that epsilon here is added for non-zero dividing and so... it makes sense\n"
        },
        {
          "id": 221129,
          "postDate": "2017-09-14T07:51:18.800Z",
          "content": "<p>I also have the same question. Could you please show more details about it?</p>",
          "rawMarkdown": " I also have the same question. Could you please show more details about it?"
        },
        {
          "id": 221515,
          "postDate": "2017-09-15T14:25:14.777Z",
          "content": "<p>I reviewed the definition and found that epsilon does not matter so much. The most straight way is shrinking lr moderately. And if we want to do training model more conservatively, enlarging rho might help.</p>",
          "rawMarkdown": "I reviewed the definition and found that epsilon does not matter so much. The most straight way is shrinking lr moderately. And if we want to do training model more conservatively, enlarging rho might help."
        }
      ]
    },
    {
      "id": 217882,
      "postDate": "2017-09-01T11:03:48.353Z",
      "content": "<p>For me only SGD works right away... mhhm</p>",
      "rawMarkdown": "For me only SGD works right away... mhhm"
    },
    {
      "id": 217820,
      "postDate": "2017-09-01T04:43:39.703Z",
      "content": "<p>Many people reported RMSProp works better than Adam for this contest so I tried RMSProp after using Adam exclusively.  For me, my RMSProp result was slightly worse than Adam.  I am using batch_size = 1 if that matters.  I will stick with Adam for now.</p>",
      "rawMarkdown": "Many people reported RMSProp works better than Adam for this contest so I tried RMSProp after using Adam exclusively.  For me, my RMSProp result was slightly worse than Adam.  I am using batch_size = 1 if that matters.  I will stick with Adam for now."
    },
    {
      "id": 1292351,
      "postDate": "2021-05-03T21:08:53.590Z",
      "content": "<p>I prefer adam, but it is important to pay attention to Beta_2 h-params as it affects your momentum vector and has a huge influence.<br>\nRegarding RMSProp, one could use adam instead. As both of them are in the same category and consider the Exponential moving average for previous gradients.<br>\nSo in my opinion is a matter of ADAGRAD, SGD vanilla, or RMSPROP family!</p>",
      "rawMarkdown": "I prefer adam, but it is important to pay attention to Beta_2 h-params as it affects your momentum vector and has a huge influence.\nRegarding RMSProp, one could use adam instead. As both of them are in the same category and consider the Exponential moving average for previous gradients.\nSo in my opinion is a matter of ADAGRAD, SGD vanilla, or RMSPROP family!",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 216084,
      "author_name": "Ern",
      "author_url": "",
      "post_date": "2017-08-24T08:31:48.267000",
      "content": "<p>Adam seems to work better compared to SGD for me. I've yet to try RMSProp.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 216057,
      "author_name": "James Requa",
      "author_url": "",
      "post_date": "2017-08-24T06:36:29.733000",
      "content": "<p>I got better results with RMSProp and SGD than with Adam. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 218363,
      "author_name": "SpruceMoose",
      "author_url": "",
      "post_date": "2017-09-04T00:59:49.173000",
      "content": "<p>Adam was the best for me so far. SGD came in second, with RMSprop in 3rd. I might return to RMSprop later since I didn't give it that much time</p>",
      "votes": 1,
      "replies": [
        {
          "id": 218382,
          "author_name": "Will",
          "author_url": "",
          "post_date": "2017-09-04T02:26:16.587000",
          "content": "<p>Could you please show your batch size and input size?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 218388,
          "author_name": "SpruceMoose",
          "author_url": "",
          "post_date": "2017-09-04T03:04:18.923000",
          "content": "<p>Sure thing. My batch size is 2 for training and 8 for inference, using bn. My input size is 1024x1024, using hq images. My hypothesis is that smaller batch sizes actually perform well here since pixel positional information is important in this challenge</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 221131,
          "author_name": "Will",
          "author_url": "",
          "post_date": "2017-09-14T08:04:25.857000",
          "content": "<p>Have you ever tried BR layer. In my net , BR layer got 0.0002 higher LB score than BN layer.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 221410,
          "author_name": "SpruceMoose",
          "author_url": "",
          "post_date": "2017-09-15T06:54:00.470000",
          "content": "<p>No I haven't actually <a href=\"https://arxiv.org/pdf/1702.03275.pdf\">https://arxiv.org/pdf/1702.03275.pdf</a></p>\n\n<p>I think this is a pretty good idea for several image recognition competitions, not just this one, since it appears the higher resolution models that sacrifice batch size do better from my experience</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 216154,
      "author_name": "Kaggoo",
      "author_url": "",
      "post_date": "2017-08-24T14:49:50.467000",
      "content": "<p>Depending on your hardware and batch_size. You try and test it on your system and you will find the answer lay within your notebook.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 217872,
      "author_name": "remidi",
      "author_url": "",
      "post_date": "2017-09-01T09:52:58.027000",
      "content": "<p><a href=\"http://ruder.io/optimizing-gradient-descent/index.html#gradientdescentoptimizationalgorithms\">http://ruder.io/optimizing-gradient-descent/index.html#gradientdescentoptimizationalgorithms</a></p>\n\n<p>I hope this helps to clear your understanding on optimizers.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 218364,
          "author_name": "SpruceMoose",
          "author_url": "",
          "post_date": "2017-09-04T01:05:24.550000",
          "content": "<p>Dont think the OP was asking for clarification, but i do think this is a good post here. thanks </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1531533,
      "author_name": "M.Akbari42",
      "author_url": "",
      "post_date": "2021-10-02T05:22:36.100000",
      "content": "<p>see this : <a href=\"https://ruder.io/optimizing-gradient-descent/index.html\" target=\"_blank\">An overview of gradient descent optimization algorithms</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 887486,
      "author_name": "Daksh Bhatia",
      "author_url": "",
      "post_date": "2020-06-15T17:38:15.223000",
      "content": "<p>For me, SGD work pretty good, if you face any issue with SGD try with bit low learning rate and use the value of momentum around 0.85 to 0.95</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 218601,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2017-09-05T05:08:27.343000",
      "content": "<p>rmsprop uses only sign and not magnitude of the computed gradients for descend.  A intermediate solution is gradient capping:</p>\n\n<pre><code>        loss.backward()\n        # accumulate gradients\n        if (it+1)%num_grad_acc==0:\n            torch.nn.utils.clip_grad_norm(net.parameters(), 1)\n\n            optimizer.step()\n            optimizer.zero_grad()  # assume no effects on bn for accumulating grad\n</code></pre>\n\n<p>I  find this effective in preventing oscillation in loss.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 218794,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2017-09-05T22:42:09.947000",
          "content": "<p>This is for pytorch right? How do you do that with keras?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 218811,
          "author_name": "lyakaap",
          "author_url": "",
          "post_date": "2017-09-06T00:09:55.147000",
          "content": "<blockquote>\n  <p>rmsprop uses only sign and not magnitude of the computed gradients for descend.</p>\n</blockquote>\n\n<p>I don't understand this well, so could you elaborate this mean?</p>\n\n<p>@YaGana Sheriff-Hussaini\nProbably we can get by this code </p>\n\n<p><code>\noptimizer = RMSprop(lr=0.0001, clipnorm=1.)\n</code></p>\n\n<p>I'm worried about that this gradient capping causes worse effect to learning. </p>\n\n<p>To get more stability of RMSprop, enlarging epsilon (1e-8 is keras default. change 1e-6.) is also a way.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 218873,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2017-09-06T06:27:36.520000",
          "content": "<p>@Heng, I have not used pytorch so I do not know how RMSprop is implemented there. But I took a look at the implementation of RMSprop in keras, and I do not think the oscillations you mentioned would be a problem. By definition, RMSprop uses oscillation reduction and momentum in tandem to achieve the desired effect.</p>\n\n<p>Thanks @Iyakaap. I have not tried it, but for those using Keras, I am also of the belief that it may very well result in an unwanted effect.</p>\n\n<p>In Keras, RMSprop is implemented according to <a href=\"http://www.cs.toronto.edu/~tijmen/csc321/slides/lecture_slides_lec6.pdf\">http://www.cs.toronto.edu/~tijmen/csc321/slides/lecture_slides_lec6.pdf</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 219707,
          "author_name": "Canoe",
          "author_url": "",
          "post_date": "2017-09-09T08:23:08.863000",
          "content": "<p>Why enlarging epsilon will make RMSprop more stable? Thanks for more explaining~ <br>\nAnd Keras is not recommending to change the eplison value. Any benefits can get according to  your experiment?</p>\n\n<p>RMSprop</p>\n\n<p>keras.optimizers.RMSprop(lr=0.001, rho=0.9, epsilon=1e-08, decay=0.0)\nRMSProp optimizer.</p>\n\n<p>It is recommended to leave the parameters of this optimizer at their default values (except the learning rate, which can be freely tuned).\nupdate:\nsame that epsilon here is added for non-zero dividing and so... it makes sense</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 221129,
          "author_name": "Will",
          "author_url": "",
          "post_date": "2017-09-14T07:51:18.800000",
          "content": "<p>I also have the same question. Could you please show more details about it?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 221515,
          "author_name": "lyakaap",
          "author_url": "",
          "post_date": "2017-09-15T14:25:14.777000",
          "content": "<p>I reviewed the definition and found that epsilon does not matter so much. The most straight way is shrinking lr moderately. And if we want to do training model more conservatively, enlarging rho might help.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 217882,
      "author_name": "Tim Joseph",
      "author_url": "",
      "post_date": "2017-09-01T11:03:48.353000",
      "content": "<p>For me only SGD works right away... mhhm</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 217820,
      "author_name": "jackkwok",
      "author_url": "",
      "post_date": "2017-09-01T04:43:39.703000",
      "content": "<p>Many people reported RMSProp works better than Adam for this contest so I tried RMSProp after using Adam exclusively.  For me, my RMSProp result was slightly worse than Adam.  I am using batch_size = 1 if that matters.  I will stick with Adam for now.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1292351,
      "author_name": "Saman",
      "author_url": "",
      "post_date": "2021-05-03T21:08:53.590000",
      "content": "<p>I prefer adam, but it is important to pay attention to Beta_2 h-params as it affects your momentum vector and has a huge influence.<br>\nRegarding RMSProp, one could use adam instead. As both of them are in the same category and consider the Exponential moving average for previous gradients.<br>\nSo in my opinion is a matter of ADAGRAD, SGD vanilla, or RMSPROP family!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "216040": "Which optimization algo works best for training your UNet?",
    "216084": "Adam seems to work better compared to SGD for me. I've yet to try RMSProp.",
    "216057": "I got better results with RMSProp and SGD than with Adam. ",
    "218363": "Adam was the best for me so far. SGD came in second, with RMSprop in 3rd. I might return to RMSprop later since I didn't give it that much time",
    "216154": "Depending on your hardware and batch_size. You try and test it on your system and you will find the answer lay within your notebook.",
    "217872": "http://ruder.io/optimizing-gradient-descent/index.html#gradientdescentoptimizationalgorithms\n\nI hope this helps to clear your understanding on optimizers.",
    "1531533": "see this : [An overview of gradient descent optimization algorithms](https://ruder.io/optimizing-gradient-descent/index.html)",
    "887486": "For me, SGD work pretty good, if you face any issue with SGD try with bit low learning rate and use the value of momentum around 0.85 to 0.95",
    "218601": "rmsprop uses only sign and not magnitude of the computed gradients for descend.  A intermediate solution is gradient capping:\n\n            loss.backward()\n            # accumulate gradients\n            if (it+1)%num_grad_acc==0:\n                torch.nn.utils.clip_grad_norm(net.parameters(), 1)\n\n                optimizer.step()\n                optimizer.zero_grad()  # assume no effects on bn for accumulating grad\n\nI  find this effective in preventing oscillation in loss.",
    "217882": "For me only SGD works right away... mhhm",
    "217820": "Many people reported RMSProp works better than Adam for this contest so I tried RMSProp after using Adam exclusively.  For me, my RMSProp result was slightly worse than Adam.  I am using batch_size = 1 if that matters.  I will stick with Adam for now.",
    "1292351": "I prefer adam, but it is important to pay attention to Beta_2 h-params as it affects your momentum vector and has a huge influence.\nRegarding RMSProp, one could use adam instead. As both of them are in the same category and consider the Exponential moving average for previous gradients.\nSo in my opinion is a matter of ADAGRAD, SGD vanilla, or RMSPROP family!"
  }
}