{
  "id": 126543,
  "title": "[TF2.0] Way to accumulate gradients in custom loop TF2.0",
  "url": "/competitions/tensorflow2-question-answering/discussion/126543",
  "author_name": "cfiken",
  "post_date": "2020-01-18T07:44:48.314000",
  "votes": 10,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I opened a sample notebook to use gradients accumulation in custom loop for TF2.0.\n<a href=\"https://www.kaggle.com/kentaronakanishi/tf2-0-way-to-accumulate-gradients-in-custom-loop\">https://www.kaggle.com/kentaronakanishi/tf2-0-way-to-accumulate-gradients-in-custom-loop</a>\nI think many kagglers are using Pytorch, so this is more familiar way to train by custom loop IMO. <br>\nPlease advise me if you know better way to implement.</p>\n\n<p>In implementing gradient accumulation in TF 2.0 custom loop (I mean using <code>tf.gradientTape</code> ), there is a trick in handling a gradient of <code>tf.gather</code> operation.\nIf you use <code>tf.gather</code> in your model (most of nlp models are using this), you need to convert the gradients type from tf.IndexSlices into tf.Tensor to take averaging them like below:</p>\n\n<p>```\ndef accumulated_gradients(gradients: Optional[List[tf.Tensor]],\n                          step_gradients: List[tf.Tensor],\n                          num_grad_accumulates: int) -&gt; tf.Tensor:\n    if gradients is None:\n        gradients = [flat_gradients(g) / num_grad_accumulates for g in step_gradients]\n    else:\n        for i, g in enumerate(step_gradients):\n            gradients[i] += flat_gradients(g) / num_grad_accumulates</p>\n\n<pre><code>return gradients\n</code></pre>\n\n<p>def flat_gradients(grads_or_idx_slices: tf.Tensor) -&gt; tf.Tensor:\n    '''Convert gradients if it's tf.IndexedSlices.\n    When computing gradients for operation concerning <code>tf.gather</code>, the type of gradients \n    '''\n    if type(grads_or_idx_slices) == tf.IndexedSlices:\n        return tf.scatter_nd(\n            tf.expand_dims(grads_or_idx_slices.indices, 1),\n            grads_or_idx_slices.values,\n            grads_or_idx_slices.dense_shape\n        )\n    return grads_or_idx_slices\n```</p>\n\n<p>I hope this can help someone.\nMore detail for Gradient Accumulation and implementation (Japanese): <a href=\"https://qiita.com/cfiken/items/1de519e741cbbc09818c\">https://qiita.com/cfiken/items/1de519e741cbbc09818c</a></p>",
  "messages": [
    {
      "id": 722165,
      "postDate": "2020-01-18T07:44:48.313Z",
      "content": "<p>I opened a sample notebook to use gradients accumulation in custom loop for TF2.0.\n<a href=\"https://www.kaggle.com/kentaronakanishi/tf2-0-way-to-accumulate-gradients-in-custom-loop\">https://www.kaggle.com/kentaronakanishi/tf2-0-way-to-accumulate-gradients-in-custom-loop</a>\nI think many kagglers are using Pytorch, so this is more familiar way to train by custom loop IMO. <br>\nPlease advise me if you know better way to implement.</p>\n\n<p>In implementing gradient accumulation in TF 2.0 custom loop (I mean using <code>tf.gradientTape</code> ), there is a trick in handling a gradient of <code>tf.gather</code> operation.\nIf you use <code>tf.gather</code> in your model (most of nlp models are using this), you need to convert the gradients type from tf.IndexSlices into tf.Tensor to take averaging them like below:</p>\n\n<p>```\ndef accumulated_gradients(gradients: Optional[List[tf.Tensor]],\n                          step_gradients: List[tf.Tensor],\n                          num_grad_accumulates: int) -&gt; tf.Tensor:\n    if gradients is None:\n        gradients = [flat_gradients(g) / num_grad_accumulates for g in step_gradients]\n    else:\n        for i, g in enumerate(step_gradients):\n            gradients[i] += flat_gradients(g) / num_grad_accumulates</p>\n\n<pre><code>return gradients\n</code></pre>\n\n<p>def flat_gradients(grads_or_idx_slices: tf.Tensor) -&gt; tf.Tensor:\n    '''Convert gradients if it's tf.IndexedSlices.\n    When computing gradients for operation concerning <code>tf.gather</code>, the type of gradients \n    '''\n    if type(grads_or_idx_slices) == tf.IndexedSlices:\n        return tf.scatter_nd(\n            tf.expand_dims(grads_or_idx_slices.indices, 1),\n            grads_or_idx_slices.values,\n            grads_or_idx_slices.dense_shape\n        )\n    return grads_or_idx_slices\n```</p>\n\n<p>I hope this can help someone.\nMore detail for Gradient Accumulation and implementation (Japanese): <a href=\"https://qiita.com/cfiken/items/1de519e741cbbc09818c\">https://qiita.com/cfiken/items/1de519e741cbbc09818c</a></p>",
      "rawMarkdown": "I opened a sample notebook to use gradients accumulation in custom loop for TF2.0.\nhttps://www.kaggle.com/kentaronakanishi/tf2-0-way-to-accumulate-gradients-in-custom-loop\nI think many kagglers are using Pytorch, so this is more familiar way to train by custom loop IMO.   \nPlease advise me if you know better way to implement.\n\nIn implementing gradient accumulation in TF 2.0 custom loop (I mean using `tf.gradientTape` ), there is a trick in handling a gradient of `tf.gather` operation.\nIf you use `tf.gather` in your model (most of nlp models are using this), you need to convert the gradients type from tf.IndexSlices into tf.Tensor to take averaging them like below:\n\n```\ndef accumulated_gradients(gradients: Optional[List[tf.Tensor]],\n                          step_gradients: List[tf.Tensor],\n                          num_grad_accumulates: int) -&gt; tf.Tensor:\n    if gradients is None:\n        gradients = [flat_gradients(g) / num_grad_accumulates for g in step_gradients]\n    else:\n        for i, g in enumerate(step_gradients):\n            gradients[i] += flat_gradients(g) / num_grad_accumulates\n        \n    return gradients\n\n\ndef flat_gradients(grads_or_idx_slices: tf.Tensor) -&gt; tf.Tensor:\n    '''Convert gradients if it's tf.IndexedSlices.\n    When computing gradients for operation concerning `tf.gather`, the type of gradients \n    '''\n    if type(grads_or_idx_slices) == tf.IndexedSlices:\n        return tf.scatter_nd(\n            tf.expand_dims(grads_or_idx_slices.indices, 1),\n            grads_or_idx_slices.values,\n            grads_or_idx_slices.dense_shape\n        )\n    return grads_or_idx_slices\n```\n\nI hope this can help someone.\nMore detail for Gradient Accumulation and implementation (Japanese): https://qiita.com/cfiken/items/1de519e741cbbc09818c\n",
      "votes": 9
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "722165": "I opened a sample notebook to use gradients accumulation in custom loop for TF2.0.\nhttps://www.kaggle.com/kentaronakanishi/tf2-0-way-to-accumulate-gradients-in-custom-loop\nI think many kagglers are using Pytorch, so this is more familiar way to train by custom loop IMO.   \nPlease advise me if you know better way to implement.\n\nIn implementing gradient accumulation in TF 2.0 custom loop (I mean using `tf.gradientTape` ), there is a trick in handling a gradient of `tf.gather` operation.\nIf you use `tf.gather` in your model (most of nlp models are using this), you need to convert the gradients type from tf.IndexSlices into tf.Tensor to take averaging them like below:\n\n```\ndef accumulated_gradients(gradients: Optional[List[tf.Tensor]],\n                          step_gradients: List[tf.Tensor],\n                          num_grad_accumulates: int) -&gt; tf.Tensor:\n    if gradients is None:\n        gradients = [flat_gradients(g) / num_grad_accumulates for g in step_gradients]\n    else:\n        for i, g in enumerate(step_gradients):\n            gradients[i] += flat_gradients(g) / num_grad_accumulates\n        \n    return gradients\n\n\ndef flat_gradients(grads_or_idx_slices: tf.Tensor) -&gt; tf.Tensor:\n    '''Convert gradients if it's tf.IndexedSlices.\n    When computing gradients for operation concerning `tf.gather`, the type of gradients \n    '''\n    if type(grads_or_idx_slices) == tf.IndexedSlices:\n        return tf.scatter_nd(\n            tf.expand_dims(grads_or_idx_slices.indices, 1),\n            grads_or_idx_slices.values,\n            grads_or_idx_slices.dense_shape\n        )\n    return grads_or_idx_slices\n```\n\nI hope this can help someone.\nMore detail for Gradient Accumulation and implementation (Japanese): https://qiita.com/cfiken/items/1de519e741cbbc09818c\n"
  }
}