{
  "id": 102570,
  "title": "Gradient accumulation in Keras",
  "url": "/competitions/aptos2019-blindness-detection/discussion/102570",
  "author_name": "",
  "post_date": "2019-08-02T20:47:40.407780500Z",
  "votes": 5,
  "comment_count": 1,
  "views": 0,
  "content": "<p>It is very important to set a sufficiently large batch size solving the problem of multiclass classification. Also most of the images in the dataset have pretty high resolution and it would be better downsample it not too much. Training a model with a big batch on high-resolution images in a kernel contest can be challenging. In this case, gradient accumulation could help.\nRecently I found a wrapper for Keras optimizers on Github (<a href=\"https://github.com/bojone/accum_optimizer_for_keras\">https://github.com/bojone/accum_optimizer_for_keras</a>), which allows to update gradients once in few iterations.</p>\n\n<p><code>\nclass AccumOptimizer(Optimizer):\n    \"\"\"Inheriting Optimizer class, wrapping the original optimizer\n    to achieve a new corresponding optimizer of gradient accumulation.\n    # Arguments\n        optimizer: an instance of keras optimizer (supporting\n                    all keras optimizers currently available);\n        steps_per_update: the steps of gradient accumulation\n    # Returns\n        a new keras optimizer.\n    \"\"\"\n    def __init__(self, optimizer, steps_per_update=1, **kwargs):\n        super(AccumOptimizer, self).__init__(**kwargs)\n        self.optimizer = optimizer\n        with K.name_scope(self.__class__.__name__):\n            self.steps_per_update = steps_per_update\n            self.iterations = K.variable(0, dtype='int64', name='iterations')\n            self.cond = K.equal(self.iterations % self.steps_per_update, 0)\n            self.lr = self.optimizer.lr\n            self.optimizer.lr = K.switch(self.cond, self.optimizer.lr, 0.)\n            for attr in ['momentum', 'rho', 'beta_1', 'beta_2']:\n                if hasattr(self.optimizer, attr):\n                    value = getattr(self.optimizer, attr)\n                    setattr(self, attr, value)\n                    setattr(self.optimizer, attr, K.switch(self.cond, value, 1 - 1e-7))\n            for attr in self.optimizer.get_config():\n                if not hasattr(self, attr):\n                    value = getattr(self.optimizer, attr)\n                    setattr(self, attr, value)\n            # Cover the original get_gradients method with accumulative gradients.\n            def get_gradients(loss, params):\n                return [ag / self.steps_per_update for ag in self.accum_grads]\n            self.optimizer.get_gradients = get_gradients\n    def get_updates(self, loss, params):\n        self.updates = [\n            K.update_add(self.iterations, 1),\n            K.update_add(self.optimizer.iterations, K.cast(self.cond, 'int64')),\n        ]\n        # gradient accumulation\n        self.accum_grads = [K.zeros(K.int_shape(p), dtype=K.dtype(p)) for p in params]\n        grads = self.get_gradients(loss, params)\n        for g, ag in zip(grads, self.accum_grads):\n            self.updates.append(K.update(ag, K.switch(self.cond, ag * 0, ag + g)))\n        # inheriting updates of original optimizer\n        self.updates.extend(self.optimizer.get_updates(loss, params)[1:])\n        self.weights.extend(self.optimizer.weights)\n        return self.updates\n    def get_config(self):\n        iterations = K.eval(self.iterations)\n        K.set_value(self.iterations, 0)\n        config = self.optimizer.get_config()\n        K.set_value(self.iterations, iterations)\n        return config\n</code>\nIt can be used in the next way:\n<code>\nmodel.compile(loss='categorical_crossentropy', optimizer=AccumOptimizer(Adam(), 4), metrics=['categorical_accuracy'])\n</code></p>",
  "messages": [
    {
      "id": "590925",
      "postDate": "08/02/2019 20:47:40",
      "content": "<p>It is very important to set a sufficiently large batch size solving the problem of multiclass classification. Also most of the images in the dataset have pretty high resolution and it would be better downsample it not too much. Training a model with a big batch on high-resolution images in a kernel contest can be challenging. In this case, gradient accumulation could help.\nRecently I found a wrapper for Keras optimizers on Github (<a href=\"https://github.com/bojone/accum_optimizer_for_keras\">https://github.com/bojone/accum_optimizer_for_keras</a>), which allows to update gradients once in few iterations.</p>\n\n<p><code>\nclass AccumOptimizer(Optimizer):\n    \"\"\"Inheriting Optimizer class, wrapping the original optimizer\n    to achieve a new corresponding optimizer of gradient accumulation.\n    # Arguments\n        optimizer: an instance of keras optimizer (supporting\n                    all keras optimizers currently available);\n        steps_per_update: the steps of gradient accumulation\n    # Returns\n        a new keras optimizer.\n    \"\"\"\n    def __init__(self, optimizer, steps_per_update=1, **kwargs):\n        super(AccumOptimizer, self).__init__(**kwargs)\n        self.optimizer = optimizer\n        with K.name_scope(self.__class__.__name__):\n            self.steps_per_update = steps_per_update\n            self.iterations = K.variable(0, dtype='int64', name='iterations')\n            self.cond = K.equal(self.iterations % self.steps_per_update, 0)\n            self.lr = self.optimizer.lr\n            self.optimizer.lr = K.switch(self.cond, self.optimizer.lr, 0.)\n            for attr in ['momentum', 'rho', 'beta_1', 'beta_2']:\n                if hasattr(self.optimizer, attr):\n                    value = getattr(self.optimizer, attr)\n                    setattr(self, attr, value)\n                    setattr(self.optimizer, attr, K.switch(self.cond, value, 1 - 1e-7))\n            for attr in self.optimizer.get_config():\n                if not hasattr(self, attr):\n                    value = getattr(self.optimizer, attr)\n                    setattr(self, attr, value)\n            # Cover the original get_gradients method with accumulative gradients.\n            def get_gradients(loss, params):\n                return [ag / self.steps_per_update for ag in self.accum_grads]\n            self.optimizer.get_gradients = get_gradients\n    def get_updates(self, loss, params):\n        self.updates = [\n            K.update_add(self.iterations, 1),\n            K.update_add(self.optimizer.iterations, K.cast(self.cond, 'int64')),\n        ]\n        # gradient accumulation\n        self.accum_grads = [K.zeros(K.int_shape(p), dtype=K.dtype(p)) for p in params]\n        grads = self.get_gradients(loss, params)\n        for g, ag in zip(grads, self.accum_grads):\n            self.updates.append(K.update(ag, K.switch(self.cond, ag * 0, ag + g)))\n        # inheriting updates of original optimizer\n        self.updates.extend(self.optimizer.get_updates(loss, params)[1:])\n        self.weights.extend(self.optimizer.weights)\n        return self.updates\n    def get_config(self):\n        iterations = K.eval(self.iterations)\n        K.set_value(self.iterations, 0)\n        config = self.optimizer.get_config()\n        K.set_value(self.iterations, iterations)\n        return config\n</code>\nIt can be used in the next way:\n<code>\nmodel.compile(loss='categorical_crossentropy', optimizer=AccumOptimizer(Adam(), 4), metrics=['categorical_accuracy'])\n</code></p>",
      "rawMarkdown": "It is very important to set a sufficiently large batch size solving the problem of multiclass classification. Also most of the images in the dataset have pretty high resolution and it would be better downsample it not too much. Training a model with a big batch on high-resolution images in a kernel contest can be challenging. In this case, gradient accumulation could help.\nRecently I found a wrapper for Keras optimizers on Github (https://github.com/bojone/accum_optimizer_for_keras), which allows to update gradients once in few iterations.\n\n```\nclass AccumOptimizer(Optimizer):\n    \"\"\"Inheriting Optimizer class, wrapping the original optimizer\n    to achieve a new corresponding optimizer of gradient accumulation.\n    # Arguments\n        optimizer: an instance of keras optimizer (supporting\n                    all keras optimizers currently available);\n        steps_per_update: the steps of gradient accumulation\n    # Returns\n        a new keras optimizer.\n    \"\"\"\n    def __init__(self, optimizer, steps_per_update=1, **kwargs):\n        super(AccumOptimizer, self).__init__(**kwargs)\n        self.optimizer = optimizer\n        with K.name_scope(self.__class__.__name__):\n            self.steps_per_update = steps_per_update\n            self.iterations = K.variable(0, dtype='int64', name='iterations')\n            self.cond = K.equal(self.iterations % self.steps_per_update, 0)\n            self.lr = self.optimizer.lr\n            self.optimizer.lr = K.switch(self.cond, self.optimizer.lr, 0.)\n            for attr in ['momentum', 'rho', 'beta_1', 'beta_2']:\n                if hasattr(self.optimizer, attr):\n                    value = getattr(self.optimizer, attr)\n                    setattr(self, attr, value)\n                    setattr(self.optimizer, attr, K.switch(self.cond, value, 1 - 1e-7))\n            for attr in self.optimizer.get_config():\n                if not hasattr(self, attr):\n                    value = getattr(self.optimizer, attr)\n                    setattr(self, attr, value)\n            # Cover the original get_gradients method with accumulative gradients.\n            def get_gradients(loss, params):\n                return [ag / self.steps_per_update for ag in self.accum_grads]\n            self.optimizer.get_gradients = get_gradients\n    def get_updates(self, loss, params):\n        self.updates = [\n            K.update_add(self.iterations, 1),\n            K.update_add(self.optimizer.iterations, K.cast(self.cond, 'int64')),\n        ]\n        # gradient accumulation\n        self.accum_grads = [K.zeros(K.int_shape(p), dtype=K.dtype(p)) for p in params]\n        grads = self.get_gradients(loss, params)\n        for g, ag in zip(grads, self.accum_grads):\n            self.updates.append(K.update(ag, K.switch(self.cond, ag * 0, ag + g)))\n        # inheriting updates of original optimizer\n        self.updates.extend(self.optimizer.get_updates(loss, params)[1:])\n        self.weights.extend(self.optimizer.weights)\n        return self.updates\n    def get_config(self):\n        iterations = K.eval(self.iterations)\n        K.set_value(self.iterations, 0)\n        config = self.optimizer.get_config()\n        K.set_value(self.iterations, iterations)\n        return config\n```\nIt can be used in the next way:\n```\nmodel.compile(loss='categorical_crossentropy', optimizer=AccumOptimizer(Adam(), 4), metrics=['categorical_accuracy'])\n```",
      "votes": null
    },
    {
      "id": "590968",
      "postDate": "08/02/2019 23:52:58",
      "content": "<p>Nice, did you got nice results using this?</p>",
      "rawMarkdown": "Nice, did you got nice results using this?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 590968,
      "author_name": "dimitreoliveira",
      "author_url": "",
      "post_date": "08/02/2019 23:52:58",
      "content": "<p>Nice, did you got nice results using this?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "590925": "It is very important to set a sufficiently large batch size solving the problem of multiclass classification. Also most of the images in the dataset have pretty high resolution and it would be better downsample it not too much. Training a model with a big batch on high-resolution images in a kernel contest can be challenging. In this case, gradient accumulation could help.\nRecently I found a wrapper for Keras optimizers on Github (https://github.com/bojone/accum_optimizer_for_keras), which allows to update gradients once in few iterations.\n\n```\nclass AccumOptimizer(Optimizer):\n    \"\"\"Inheriting Optimizer class, wrapping the original optimizer\n    to achieve a new corresponding optimizer of gradient accumulation.\n    # Arguments\n        optimizer: an instance of keras optimizer (supporting\n                    all keras optimizers currently available);\n        steps_per_update: the steps of gradient accumulation\n    # Returns\n        a new keras optimizer.\n    \"\"\"\n    def __init__(self, optimizer, steps_per_update=1, **kwargs):\n        super(AccumOptimizer, self).__init__(**kwargs)\n        self.optimizer = optimizer\n        with K.name_scope(self.__class__.__name__):\n            self.steps_per_update = steps_per_update\n            self.iterations = K.variable(0, dtype='int64', name='iterations')\n            self.cond = K.equal(self.iterations % self.steps_per_update, 0)\n            self.lr = self.optimizer.lr\n            self.optimizer.lr = K.switch(self.cond, self.optimizer.lr, 0.)\n            for attr in ['momentum', 'rho', 'beta_1', 'beta_2']:\n                if hasattr(self.optimizer, attr):\n                    value = getattr(self.optimizer, attr)\n                    setattr(self, attr, value)\n                    setattr(self.optimizer, attr, K.switch(self.cond, value, 1 - 1e-7))\n            for attr in self.optimizer.get_config():\n                if not hasattr(self, attr):\n                    value = getattr(self.optimizer, attr)\n                    setattr(self, attr, value)\n            # Cover the original get_gradients method with accumulative gradients.\n            def get_gradients(loss, params):\n                return [ag / self.steps_per_update for ag in self.accum_grads]\n            self.optimizer.get_gradients = get_gradients\n    def get_updates(self, loss, params):\n        self.updates = [\n            K.update_add(self.iterations, 1),\n            K.update_add(self.optimizer.iterations, K.cast(self.cond, 'int64')),\n        ]\n        # gradient accumulation\n        self.accum_grads = [K.zeros(K.int_shape(p), dtype=K.dtype(p)) for p in params]\n        grads = self.get_gradients(loss, params)\n        for g, ag in zip(grads, self.accum_grads):\n            self.updates.append(K.update(ag, K.switch(self.cond, ag * 0, ag + g)))\n        # inheriting updates of original optimizer\n        self.updates.extend(self.optimizer.get_updates(loss, params)[1:])\n        self.weights.extend(self.optimizer.weights)\n        return self.updates\n    def get_config(self):\n        iterations = K.eval(self.iterations)\n        K.set_value(self.iterations, 0)\n        config = self.optimizer.get_config()\n        K.set_value(self.iterations, iterations)\n        return config\n```\nIt can be used in the next way:\n```\nmodel.compile(loss='categorical_crossentropy', optimizer=AccumOptimizer(Adam(), 4), metrics=['categorical_accuracy'])\n```",
    "590968": "Nice, did you got nice results using this?"
  },
  "source": "meta"
}