{
  "id": 320534,
  "title": "Gradient Accumulation in PyTorch",
  "url": "/competitions/birdclef-2022/discussion/320534",
  "author_name": "Bala Baskar",
  "post_date": "2022-04-22T05:47:26.996000",
  "votes": 16,
  "comment_count": 3,
  "views": 0,
  "content": "<p>When training a neural network, we usually divide our data in mini-batches and go through them one by one. The network predicts batch labels, which are used to compute the loss with respect to the actual targets. Next, we perform a backward pass to compute gradients and update model weights in the direction of those gradients.</p>\n<p>Gradient accumulation modifies the last step of the training process. Instead of updating the network weights on every batch, we can save gradient values, proceed to the next batch and add up the new gradients. The weight update is then done only after several batches have been processed by the model.</p>\n<p>Gradient accumulation helps to imitate a larger batch size. Imagine you want to use 32 images in one batch, but your hardware crashes once you go beyond 8. In that case, you can use batches of 8 images and update weights once every 4 batches. If you accumulate gradients from every batch in between, the results will be (almost) the same and you will be able to perform training on a less expensive machine!</p>\n<p>References:</p>\n<ol>\n<li><a href=\"https://towardsdatascience.com/what-is-gradient-accumulation-in-deep-learning-ec034122cfa\" target=\"_blank\">Gradient Accumulation</a></li>\n<li><a href=\"https://colab.research.google.com/github/kozodoi/website/blob/master/_notebooks/2021-02-19-gradient-accumulation.ipynb#scrollTo=qquHcwRiEm58\" target=\"_blank\">Gradient Accumulation in PyTorch</a></li>\n</ol>",
  "messages": [
    {
      "id": 1764067,
      "postDate": "2022-04-22T05:47:26.997Z",
      "content": "<p>When training a neural network, we usually divide our data in mini-batches and go through them one by one. The network predicts batch labels, which are used to compute the loss with respect to the actual targets. Next, we perform a backward pass to compute gradients and update model weights in the direction of those gradients.</p>\n<p>Gradient accumulation modifies the last step of the training process. Instead of updating the network weights on every batch, we can save gradient values, proceed to the next batch and add up the new gradients. The weight update is then done only after several batches have been processed by the model.</p>\n<p>Gradient accumulation helps to imitate a larger batch size. Imagine you want to use 32 images in one batch, but your hardware crashes once you go beyond 8. In that case, you can use batches of 8 images and update weights once every 4 batches. If you accumulate gradients from every batch in between, the results will be (almost) the same and you will be able to perform training on a less expensive machine!</p>\n<p>References:</p>\n<ol>\n<li><a href=\"https://towardsdatascience.com/what-is-gradient-accumulation-in-deep-learning-ec034122cfa\" target=\"_blank\">Gradient Accumulation</a></li>\n<li><a href=\"https://colab.research.google.com/github/kozodoi/website/blob/master/_notebooks/2021-02-19-gradient-accumulation.ipynb#scrollTo=qquHcwRiEm58\" target=\"_blank\">Gradient Accumulation in PyTorch</a></li>\n</ol>",
      "rawMarkdown": "When training a neural network, we usually divide our data in mini-batches and go through them one by one. The network predicts batch labels, which are used to compute the loss with respect to the actual targets. Next, we perform a backward pass to compute gradients and update model weights in the direction of those gradients.\n\nGradient accumulation modifies the last step of the training process. Instead of updating the network weights on every batch, we can save gradient values, proceed to the next batch and add up the new gradients. The weight update is then done only after several batches have been processed by the model.\n\nGradient accumulation helps to imitate a larger batch size. Imagine you want to use 32 images in one batch, but your hardware crashes once you go beyond 8. In that case, you can use batches of 8 images and update weights once every 4 batches. If you accumulate gradients from every batch in between, the results will be (almost) the same and you will be able to perform training on a less expensive machine!\n\nReferences:\n1. [Gradient Accumulation](https://towardsdatascience.com/what-is-gradient-accumulation-in-deep-learning-ec034122cfa)\n2. [Gradient Accumulation in PyTorch](https://colab.research.google.com/github/kozodoi/website/blob/master/_notebooks/2021-02-19-gradient-accumulation.ipynb#scrollTo=qquHcwRiEm58)",
      "votes": 15
    },
    {
      "id": 1786139,
      "postDate": "2022-05-12T16:26:37.477Z",
      "content": "<p>This is a great explanation of gradient accumulation! I definitely recommend checking out the references for more details.</p>",
      "rawMarkdown": "This is a great explanation of gradient accumulation! I definitely recommend checking out the references for more details.",
      "votes": 1
    },
    {
      "id": 1766619,
      "postDate": "2022-04-24T17:13:33.730Z",
      "content": "<p>Intuitive explanation…I was looking for this answer , as I am new to torch</p>",
      "rawMarkdown": "Intuitive explanation...I was looking for this answer , as I am new to torch",
      "votes": 1
    },
    {
      "id": 1764948,
      "postDate": "2022-04-23T03:26:30.107Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1786139,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-12T16:26:37.477000",
      "content": "<p>This is a great explanation of gradient accumulation! I definitely recommend checking out the references for more details.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1766619,
      "author_name": "viraj kadam",
      "author_url": "",
      "post_date": "2022-04-24T17:13:33.730000",
      "content": "<p>Intuitive explanation…I was looking for this answer , as I am new to torch</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1764948,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-04-23T03:26:30.107000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1764067": "When training a neural network, we usually divide our data in mini-batches and go through them one by one. The network predicts batch labels, which are used to compute the loss with respect to the actual targets. Next, we perform a backward pass to compute gradients and update model weights in the direction of those gradients.\n\nGradient accumulation modifies the last step of the training process. Instead of updating the network weights on every batch, we can save gradient values, proceed to the next batch and add up the new gradients. The weight update is then done only after several batches have been processed by the model.\n\nGradient accumulation helps to imitate a larger batch size. Imagine you want to use 32 images in one batch, but your hardware crashes once you go beyond 8. In that case, you can use batches of 8 images and update weights once every 4 batches. If you accumulate gradients from every batch in between, the results will be (almost) the same and you will be able to perform training on a less expensive machine!\n\nReferences:\n1. [Gradient Accumulation](https://towardsdatascience.com/what-is-gradient-accumulation-in-deep-learning-ec034122cfa)\n2. [Gradient Accumulation in PyTorch](https://colab.research.google.com/github/kozodoi/website/blob/master/_notebooks/2021-02-19-gradient-accumulation.ipynb#scrollTo=qquHcwRiEm58)",
    "1786139": "This is a great explanation of gradient accumulation! I definitely recommend checking out the references for more details.",
    "1766619": "Intuitive explanation...I was looking for this answer , as I am new to torch",
    "1764948": ""
  }
}