{
  "id": 250666,
  "title": "[Info]: Gradient Centralization: A New Optimization Technique for Deep Neural Networks",
  "url": "/competitions/siim-covid19-detection/discussion/250666",
  "author_name": "",
  "post_date": "2021-07-03T18:42:05.997806Z",
  "votes": 26,
  "comment_count": 2,
  "views": 0,
  "content": "<p>A recent work on optimization techniques for deep network and seems quite interesting. Check out the paper: <a href=\"https://arxiv.org/abs/2004.01461\" target=\"_blank\">Gradient Centralization</a></p>\n<pre><code>Gradient Centralization (GC) is a simple and effective optimization technique for Deep Neural Networks (DNNs), which operate directly on gradients by centralizing the gradient vectors to have zero mean. It can both speed up the training process and improve the final generalization performance of DNNs.\n</code></pre>\n<h2>Pytorch Implementation [Official]</h2>\n<ul>\n<li><a href=\"https://github.com/Yonghongwei/Gradient-Centralization\" target=\"_blank\">https://github.com/Yonghongwei/Gradient-Centralization</a></li>\n</ul>\n<h2>TensorFlow/Keras Implementation [Un-official]</h2>\n<p>The <strong>GC</strong> implementation is demonstrated in the official keras <a href=\"https://keras.io/examples/vision/gradient_centralization/\" target=\"_blank\">code example</a>,  and here it's adopted with bit modification.</p>\n<pre><code>from tensorflow.keras.optimizers import Adam, SGD, RMSprop\nfrom tensorflow_addons.optimizers import RectifiedAdam, Lookahead\n\nclass GradientCentralization(RectifiedAdam):\n    def get_gradients(self, loss, params):\n        grads = []\n        gradients = super().get_gradients()\n        for grad in gradients:\n            grad_len = len(grad.shape)\n            if grad_len &gt; 1:\n                axis = list(range(grad_len - 1))\n                grad -= tf.reduce_mean(grad, axis=axis, keep_dims=True)\n            grads.append(grad)\n        return grads\n\n# option 1: no gradient centralization (gcz)\n# opt = Adam(learning_rate=1e-4)\n\n# option 2: with gradient centralization (gcz)\n# opt = GradientCentralization(learning_rate=1e-4)\n\n# option 3: with gcz + lookahead \ngcz = GradientCentralization(learning_rate=1e-4)\nopt = Lookahead(gcz, sync_period=6, slow_step_size=0.5)\n</code></pre>\n<pre><code>model = tf.keras.Model(... , ...)\nmodel.compile(\n          loss  = ...,\n          metrics = ...,\n          optimizer = opt)\nmodel.fit\n</code></pre>\n<p>FYI,  quickly test on mnist/cifar dataset; seems promising. </p>",
  "messages": [
    {
      "id": "1374975",
      "postDate": "07/03/2021 18:42:05",
      "content": "<p>A recent work on optimization techniques for deep network and seems quite interesting. Check out the paper: <a href=\"https://arxiv.org/abs/2004.01461\" target=\"_blank\">Gradient Centralization</a></p>\n<pre><code>Gradient Centralization (GC) is a simple and effective optimization technique for Deep Neural Networks (DNNs), which operate directly on gradients by centralizing the gradient vectors to have zero mean. It can both speed up the training process and improve the final generalization performance of DNNs.\n</code></pre>\n<h2>Pytorch Implementation [Official]</h2>\n<ul>\n<li><a href=\"https://github.com/Yonghongwei/Gradient-Centralization\" target=\"_blank\">https://github.com/Yonghongwei/Gradient-Centralization</a></li>\n</ul>\n<h2>TensorFlow/Keras Implementation [Un-official]</h2>\n<p>The <strong>GC</strong> implementation is demonstrated in the official keras <a href=\"https://keras.io/examples/vision/gradient_centralization/\" target=\"_blank\">code example</a>,  and here it's adopted with bit modification.</p>\n<pre><code>from tensorflow.keras.optimizers import Adam, SGD, RMSprop\nfrom tensorflow_addons.optimizers import RectifiedAdam, Lookahead\n\nclass GradientCentralization(RectifiedAdam):\n    def get_gradients(self, loss, params):\n        grads = []\n        gradients = super().get_gradients()\n        for grad in gradients:\n            grad_len = len(grad.shape)\n            if grad_len &gt; 1:\n                axis = list(range(grad_len - 1))\n                grad -= tf.reduce_mean(grad, axis=axis, keep_dims=True)\n            grads.append(grad)\n        return grads\n\n# option 1: no gradient centralization (gcz)\n# opt = Adam(learning_rate=1e-4)\n\n# option 2: with gradient centralization (gcz)\n# opt = GradientCentralization(learning_rate=1e-4)\n\n# option 3: with gcz + lookahead \ngcz = GradientCentralization(learning_rate=1e-4)\nopt = Lookahead(gcz, sync_period=6, slow_step_size=0.5)\n</code></pre>\n<pre><code>model = tf.keras.Model(... , ...)\nmodel.compile(\n          loss  = ...,\n          metrics = ...,\n          optimizer = opt)\nmodel.fit\n</code></pre>\n<p>FYI,  quickly test on mnist/cifar dataset; seems promising. </p>",
      "rawMarkdown": "A recent work on optimization techniques for deep network and seems quite interesting. Check out the paper: [Gradient Centralization](https://arxiv.org/abs/2004.01461)\n\n```\nGradient Centralization (GC) is a simple and effective optimization technique for Deep Neural Networks (DNNs), which operate directly on gradients by centralizing the gradient vectors to have zero mean. It can both speed up the training process and improve the final generalization performance of DNNs.\n```\n\n## Pytorch Implementation [Official]\n\n- https://github.com/Yonghongwei/Gradient-Centralization\n\n## TensorFlow/Keras Implementation [Un-official]\n\nThe **GC** implementation is demonstrated in the official keras [code example](https://keras.io/examples/vision/gradient_centralization/),  and here it's adopted with bit modification.\n\n```\nfrom tensorflow.keras.optimizers import Adam, SGD, RMSprop\nfrom tensorflow_addons.optimizers import RectifiedAdam, Lookahead\n\nclass GradientCentralization(RectifiedAdam):\n    def get_gradients(self, loss, params):\n        grads = []\n        gradients = super().get_gradients()\n        for grad in gradients:\n            grad_len = len(grad.shape)\n            if grad_len > 1:\n                axis = list(range(grad_len - 1))\n                grad -= tf.reduce_mean(grad, axis=axis, keep_dims=True)\n            grads.append(grad)\n        return grads\n\n# option 1: no gradient centralization (gcz)\n# opt = Adam(learning_rate=1e-4)\n\n# option 2: with gradient centralization (gcz)\n# opt = GradientCentralization(learning_rate=1e-4)\n\n# option 3: with gcz + lookahead \ngcz = GradientCentralization(learning_rate=1e-4)\nopt = Lookahead(gcz, sync_period=6, slow_step_size=0.5)\n```\n```\nmodel = tf.keras.Model(... , ...)\nmodel.compile(\n          loss  = ...,\n          metrics = ...,\n          optimizer = opt)\nmodel.fit\n```\n\nFYI,  quickly test on mnist/cifar dataset; seems promising.",
      "votes": null
    },
    {
      "id": "1377881",
      "postDate": "07/06/2021 07:38:46",
      "content": "<p>Thanks for sharing this </p>",
      "rawMarkdown": "Thanks for sharing this",
      "votes": null
    },
    {
      "id": "1378223",
      "postDate": "07/06/2021 11:39:34",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1377881,
      "author_name": "deepdream02",
      "author_url": "",
      "post_date": "07/06/2021 07:38:46",
      "content": "<p>Thanks for sharing this </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1378223,
      "author_name": "samarthgupta39",
      "author_url": "",
      "post_date": "07/06/2021 11:39:34",
      "content": "<p>Thanks for sharing</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1374975": "A recent work on optimization techniques for deep network and seems quite interesting. Check out the paper: [Gradient Centralization](https://arxiv.org/abs/2004.01461)\n\n```\nGradient Centralization (GC) is a simple and effective optimization technique for Deep Neural Networks (DNNs), which operate directly on gradients by centralizing the gradient vectors to have zero mean. It can both speed up the training process and improve the final generalization performance of DNNs.\n```\n\n## Pytorch Implementation [Official]\n\n- https://github.com/Yonghongwei/Gradient-Centralization\n\n## TensorFlow/Keras Implementation [Un-official]\n\nThe **GC** implementation is demonstrated in the official keras [code example](https://keras.io/examples/vision/gradient_centralization/),  and here it's adopted with bit modification.\n\n```\nfrom tensorflow.keras.optimizers import Adam, SGD, RMSprop\nfrom tensorflow_addons.optimizers import RectifiedAdam, Lookahead\n\nclass GradientCentralization(RectifiedAdam):\n    def get_gradients(self, loss, params):\n        grads = []\n        gradients = super().get_gradients()\n        for grad in gradients:\n            grad_len = len(grad.shape)\n            if grad_len > 1:\n                axis = list(range(grad_len - 1))\n                grad -= tf.reduce_mean(grad, axis=axis, keep_dims=True)\n            grads.append(grad)\n        return grads\n\n# option 1: no gradient centralization (gcz)\n# opt = Adam(learning_rate=1e-4)\n\n# option 2: with gradient centralization (gcz)\n# opt = GradientCentralization(learning_rate=1e-4)\n\n# option 3: with gcz + lookahead \ngcz = GradientCentralization(learning_rate=1e-4)\nopt = Lookahead(gcz, sync_period=6, slow_step_size=0.5)\n```\n```\nmodel = tf.keras.Model(... , ...)\nmodel.compile(\n          loss  = ...,\n          metrics = ...,\n          optimizer = opt)\nmodel.fit\n```\n\nFYI,  quickly test on mnist/cifar dataset; seems promising.",
    "1377881": "Thanks for sharing this",
    "1378223": "Thanks for sharing"
  },
  "source": "meta"
}