{
  "id": 217133,
  "title": "Don't forget to accumulate the Gradient",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/217133",
  "author_name": "",
  "post_date": "2021-02-05T12:27:13.941766Z",
  "votes": 29,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Particularly when you use self-Attention based Models (like ViT) or CNN models which use Group Normalization + Weight Standardisation instead of BatchNorm.</p>\n<p>Grandient Accumulation may not be very effective when BatchNormalization is heavily used but can very be very helpful for the models above when you have limited GPU VRAM and forced to use small batch size. </p>\n<p>AMP + Gradient accumulation is even more effective </p>\n<pre><code>from  torch.cuda import amp\n\nscaler = amp.GradScaler()\nn_accumulate = 16  # n_accumulate = 1 means no gradient accumulation is applied\n\nfor epoch in epochs:\n    for step, (input, target) in enumerate(train_dataloader):\n        with amp.autocast(enabled=True):\n            output = model(input)\n            loss = loss_fn(output, target)\n            loss = loss / n_accumulate \n\n       # Accumulates gradients after scale.\n        scaler.scale(loss).backward()\n\n        if (step+ 1) % n_accumulate == 0:\n\n            scaler.step(optimizer)\n            scaler.update()\n            optimizer.zero_grad()\n</code></pre>",
  "messages": [
    {
      "id": "1187410",
      "postDate": "02/05/2021 12:27:13",
      "content": "<p>Particularly when you use self-Attention based Models (like ViT) or CNN models which use Group Normalization + Weight Standardisation instead of BatchNorm.</p>\n<p>Grandient Accumulation may not be very effective when BatchNormalization is heavily used but can very be very helpful for the models above when you have limited GPU VRAM and forced to use small batch size. </p>\n<p>AMP + Gradient accumulation is even more effective </p>\n<pre><code>from  torch.cuda import amp\n\nscaler = amp.GradScaler()\nn_accumulate = 16  # n_accumulate = 1 means no gradient accumulation is applied\n\nfor epoch in epochs:\n    for step, (input, target) in enumerate(train_dataloader):\n        with amp.autocast(enabled=True):\n            output = model(input)\n            loss = loss_fn(output, target)\n            loss = loss / n_accumulate \n\n       # Accumulates gradients after scale.\n        scaler.scale(loss).backward()\n\n        if (step+ 1) % n_accumulate == 0:\n\n            scaler.step(optimizer)\n            scaler.update()\n            optimizer.zero_grad()\n</code></pre>",
      "rawMarkdown": "Particularly when you use self-Attention based Models (like ViT) or CNN models which use Group Normalization + Weight Standardisation instead of BatchNorm.\n\n\nGrandient Accumulation may not be very effective when BatchNormalization is heavily used but can very be very helpful for the models above when you have limited GPU VRAM and forced to use small batch size. \n\nAMP + Gradient accumulation is even more effective \n\n```\n\nfrom  torch.cuda import amp\n\nscaler = amp.GradScaler()\nn_accumulate = 16  # n_accumulate = 1 means no gradient accumulation is applied\n\nfor epoch in epochs:\n    for step, (input, target) in enumerate(train_dataloader):\n        with amp.autocast(enabled=True):\n            output = model(input)\n            loss = loss_fn(output, target)\n            loss = loss / n_accumulate \n\n       # Accumulates gradients after scale.\n        scaler.scale(loss).backward()\n        \n        if (step+ 1) % n_accumulate == 0:\n\n            scaler.step(optimizer)\n            scaler.update()\n            optimizer.zero_grad()\n```",
      "votes": null
    },
    {
      "id": "1187777",
      "postDate": "02/05/2021 17:47:05",
      "content": "<p>Thanks for your post, - this is the first time I tried mixed-precision training, it accelerated the whole process by around 10% and more than that now my GPU can be utilized better, because thanks for mixed-precision that reduces tensor's size, I can feed it with larger batches :)</p>",
      "rawMarkdown": "Thanks for your post, - this is the first time I tried mixed-precision training, it accelerated the whole process by around 10% and more than that now my GPU can be utilized better, because thanks for mixed-precision that reduces tensor's size, I can feed it with larger batches :)",
      "votes": null
    },
    {
      "id": "1188375",
      "postDate": "02/06/2021 07:33:41",
      "content": "<p>This is a very useful tip if you are memory constrained and find there is value to larger batches, but just want to put it out there that larger batches doesn't necessarily always yield better results. </p>\n<p>That being said there has been some research showing that an increasing batch size rather than decaying learning rate can actually be useful sometimes. Would be curious if that might be the case here and could just simulate bigger and bigger batches with accumulation like this. </p>\n<p><a href=\"https://arxiv.org/abs/1711.00489\" target=\"_blank\">https://arxiv.org/abs/1711.00489</a></p>",
      "rawMarkdown": "This is a very useful tip if you are memory constrained and find there is value to larger batches, but just want to put it out there that larger batches doesn't necessarily always yield better results. \n\nThat being said there has been some research showing that an increasing batch size rather than decaying learning rate can actually be useful sometimes. Would be curious if that might be the case here and could just simulate bigger and bigger batches with accumulation like this. \n\nhttps://arxiv.org/abs/1711.00489",
      "votes": null
    },
    {
      "id": "1188554",
      "postDate": "02/06/2021 10:41:07",
      "content": "<p>Ty for the info, looks like i have been using it wrongly. So sorry for the beginner question, n_accumulate  always must be equal to the batch size?<br>\nThank you.</p>",
      "rawMarkdown": "Ty for the info, looks like i have been using it wrongly. So sorry for the beginner question, n_accumulate  always must be equal to the batch size?\nThank you.",
      "votes": null
    },
    {
      "id": "1188574",
      "postDate": "02/06/2021 10:58:52",
      "content": "<p>You're right </p>\n<p>There is trade-off  to be found between batch-size and accumulation steps. But some huge models force you to use very small Batch size when you are memory constrained. </p>\n<p>I didn't try (very) large batch size here. May be those using TPU can share feedback ^^</p>",
      "rawMarkdown": "You're right \n\nThere is trade-off  to be found between batch-size and accumulation steps. But some huge models force you to use very small Batch size when you are memory constrained. \n\nI didn't try (very) large batch size here. May be those using TPU can share feedback ^^",
      "votes": null
    },
    {
      "id": "1188594",
      "postDate": "02/06/2021 11:17:15",
      "content": "<p>Not really </p>\n<p>n_accumulate means the number of iterations (forward-backward passes: <code>scaler.scale(loss).backward()</code> ) before the model updates the parameters (<code>scaler.step(optimizer)</code>)</p>\n<p>For instance if you use <code>batch_size = 32</code> and no gradient accumulations then the parameters are updated once every batch-size i.e after 32 number of samples.</p>\n<p>If you are memory constrained and don't have other choice than using let's say <code>batch_size = 4</code>  . Then you can set for exemple <code>n_accumulate = 8</code> . Thus the parameters are updated after every 8 iterations and you'll have virtually <code>4x8 = 32</code> equivalent batch size</p>",
      "rawMarkdown": "Not really \n\n n_accumulate means the number of iterations (forward-backward passes: `scaler.scale(loss).backward()` ) before the model updates the parameters (` scaler.step(optimizer)`)\n\nFor instance if you use `batch_size = 32` and no gradient accumulations then the parameters are updated once every batch-size i.e after 32 number of samples.\n\nIf you are memory constrained and don't have other choice than using let's say `batch_size = 4`  . Then you can set for exemple `n_accumulate = 8` . Thus the parameters are updated after every 8 iterations and you'll have virtually `4x8 = 32` equivalent batch size",
      "votes": null
    },
    {
      "id": "1188732",
      "postDate": "02/06/2021 13:26:22",
      "content": "<p>Hi Serigne, do you think batchsize=32 is enough for vit? Thanks !</p>",
      "rawMarkdown": "Hi Serigne, do you think batchsize=32 is enough for vit? Thanks !",
      "votes": null
    },
    {
      "id": "1188840",
      "postDate": "02/06/2021 14:42:58",
      "content": "<p>Thank you man, really helpful!</p>",
      "rawMarkdown": "Thank you man, really helpful!",
      "votes": null
    },
    {
      "id": "1189111",
      "postDate": "02/06/2021 18:10:20",
      "content": "<p><a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> Thanks a lot for the clear explanation. </p>",
      "rawMarkdown": "serigne Thanks a lot for the clear explanation.",
      "votes": null
    },
    {
      "id": "1191663",
      "postDate": "02/08/2021 15:37:01",
      "content": "<p>\"Friends dont let friends use minibatches larger than 32.\" (Yann LeCun)<br>\n<a href=\"https://arxiv.org/abs/1804.07612\" target=\"_blank\">https://arxiv.org/abs/1804.07612</a></p>",
      "rawMarkdown": "\"Friends dont let friends use minibatches larger than 32.\" (Yann LeCun)\nhttps://arxiv.org/abs/1804.07612",
      "votes": null
    },
    {
      "id": "1193263",
      "postDate": "02/09/2021 14:39:20",
      "content": "<p>Cool tip, as always!</p>",
      "rawMarkdown": "Cool tip, as always!",
      "votes": null
    },
    {
      "id": "1195009",
      "postDate": "02/10/2021 14:08:27",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> :)</p>",
      "rawMarkdown": "Thank you @serigne :)",
      "votes": null
    },
    {
      "id": "2750740",
      "postDate": "04/13/2024 20:21:50",
      "content": "<p>I change my opinion about this. Depends on the case :)</p>",
      "rawMarkdown": "I change my opinion about this. Depends on the case :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1187777,
      "author_name": "martynoveduard",
      "author_url": "",
      "post_date": "02/05/2021 17:47:05",
      "content": "<p>Thanks for your post, - this is the first time I tried mixed-precision training, it accelerated the whole process by around 10% and more than that now my GPU can be utilized better, because thanks for mixed-precision that reduces tensor's size, I can feed it with larger batches :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1188375,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "02/06/2021 07:33:41",
      "content": "<p>This is a very useful tip if you are memory constrained and find there is value to larger batches, but just want to put it out there that larger batches doesn't necessarily always yield better results. </p>\n<p>That being said there has been some research showing that an increasing batch size rather than decaying learning rate can actually be useful sometimes. Would be curious if that might be the case here and could just simulate bigger and bigger batches with accumulation like this. </p>\n<p><a href=\"https://arxiv.org/abs/1711.00489\" target=\"_blank\">https://arxiv.org/abs/1711.00489</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1188574,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "02/06/2021 10:58:52",
          "content": "<p>You're right </p>\n<p>There is trade-off  to be found between batch-size and accumulation steps. But some huge models force you to use very small Batch size when you are memory constrained. </p>\n<p>I didn't try (very) large batch size here. May be those using TPU can share feedback ^^</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1191663,
          "author_name": "btrentini",
          "author_url": "",
          "post_date": "02/08/2021 15:37:01",
          "content": "<p>\"Friends dont let friends use minibatches larger than 32.\" (Yann LeCun)<br>\n<a href=\"https://arxiv.org/abs/1804.07612\" target=\"_blank\">https://arxiv.org/abs/1804.07612</a></p>",
          "votes": null,
          "replies": [
            {
              "id": 2750740,
              "author_name": "btrentini",
              "author_url": "",
              "post_date": "04/13/2024 20:21:50",
              "content": "<p>I change my opinion about this. Depends on the case :)</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 1188554,
      "author_name": "marjan1111",
      "author_url": "",
      "post_date": "02/06/2021 10:41:07",
      "content": "<p>Ty for the info, looks like i have been using it wrongly. So sorry for the beginner question, n_accumulate  always must be equal to the batch size?<br>\nThank you.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1188594,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "02/06/2021 11:17:15",
          "content": "<p>Not really </p>\n<p>n_accumulate means the number of iterations (forward-backward passes: <code>scaler.scale(loss).backward()</code> ) before the model updates the parameters (<code>scaler.step(optimizer)</code>)</p>\n<p>For instance if you use <code>batch_size = 32</code> and no gradient accumulations then the parameters are updated once every batch-size i.e after 32 number of samples.</p>\n<p>If you are memory constrained and don't have other choice than using let's say <code>batch_size = 4</code>  . Then you can set for exemple <code>n_accumulate = 8</code> . Thus the parameters are updated after every 8 iterations and you'll have virtually <code>4x8 = 32</code> equivalent batch size</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1188840,
          "author_name": "marjan1111",
          "author_url": "",
          "post_date": "02/06/2021 14:42:58",
          "content": "<p>Thank you man, really helpful!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1189111,
          "author_name": "durbin164",
          "author_url": "",
          "post_date": "02/06/2021 18:10:20",
          "content": "<p><a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> Thanks a lot for the clear explanation. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1195009,
          "author_name": "joshi98kishan",
          "author_url": "",
          "post_date": "02/10/2021 14:08:27",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1188732,
      "author_name": "clwwlc",
      "author_url": "",
      "post_date": "02/06/2021 13:26:22",
      "content": "<p>Hi Serigne, do you think batchsize=32 is enough for vit? Thanks !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1193263,
      "author_name": "drcodikpollonny",
      "author_url": "",
      "post_date": "02/09/2021 14:39:20",
      "content": "<p>Cool tip, as always!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1187410": "Particularly when you use self-Attention based Models (like ViT) or CNN models which use Group Normalization + Weight Standardisation instead of BatchNorm.\n\n\nGrandient Accumulation may not be very effective when BatchNormalization is heavily used but can very be very helpful for the models above when you have limited GPU VRAM and forced to use small batch size. \n\nAMP + Gradient accumulation is even more effective \n\n```\n\nfrom  torch.cuda import amp\n\nscaler = amp.GradScaler()\nn_accumulate = 16  # n_accumulate = 1 means no gradient accumulation is applied\n\nfor epoch in epochs:\n    for step, (input, target) in enumerate(train_dataloader):\n        with amp.autocast(enabled=True):\n            output = model(input)\n            loss = loss_fn(output, target)\n            loss = loss / n_accumulate \n\n       # Accumulates gradients after scale.\n        scaler.scale(loss).backward()\n        \n        if (step+ 1) % n_accumulate == 0:\n\n            scaler.step(optimizer)\n            scaler.update()\n            optimizer.zero_grad()\n```",
    "1187777": "Thanks for your post, - this is the first time I tried mixed-precision training, it accelerated the whole process by around 10% and more than that now my GPU can be utilized better, because thanks for mixed-precision that reduces tensor's size, I can feed it with larger batches :)",
    "1188375": "This is a very useful tip if you are memory constrained and find there is value to larger batches, but just want to put it out there that larger batches doesn't necessarily always yield better results. \n\nThat being said there has been some research showing that an increasing batch size rather than decaying learning rate can actually be useful sometimes. Would be curious if that might be the case here and could just simulate bigger and bigger batches with accumulation like this. \n\nhttps://arxiv.org/abs/1711.00489",
    "1188554": "Ty for the info, looks like i have been using it wrongly. So sorry for the beginner question, n_accumulate  always must be equal to the batch size?\nThank you.",
    "1188574": "You're right \n\nThere is trade-off  to be found between batch-size and accumulation steps. But some huge models force you to use very small Batch size when you are memory constrained. \n\nI didn't try (very) large batch size here. May be those using TPU can share feedback ^^",
    "1188594": "Not really \n\n n_accumulate means the number of iterations (forward-backward passes: `scaler.scale(loss).backward()` ) before the model updates the parameters (` scaler.step(optimizer)`)\n\nFor instance if you use `batch_size = 32` and no gradient accumulations then the parameters are updated once every batch-size i.e after 32 number of samples.\n\nIf you are memory constrained and don't have other choice than using let's say `batch_size = 4`  . Then you can set for exemple `n_accumulate = 8` . Thus the parameters are updated after every 8 iterations and you'll have virtually `4x8 = 32` equivalent batch size",
    "1188732": "Hi Serigne, do you think batchsize=32 is enough for vit? Thanks !",
    "1188840": "Thank you man, really helpful!",
    "1189111": "serigne Thanks a lot for the clear explanation.",
    "1191663": "\"Friends dont let friends use minibatches larger than 32.\" (Yann LeCun)\nhttps://arxiv.org/abs/1804.07612",
    "1193263": "Cool tip, as always!",
    "1195009": "Thank you @serigne :)",
    "2750740": "I change my opinion about this. Depends on the case :)"
  },
  "source": "meta"
}