{
  "id": 165651,
  "title": "Gradient accumulation - How to speed up convergence",
  "url": "/competitions/alaska2-image-steganalysis/discussion/165651",
  "author_name": "",
  "post_date": "2020-07-10T13:20:28.425373800Z",
  "votes": 13,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hello everyone,</p>\n\n<p>As the competition is drawing to a close, I'd like to share with you a technique that worked quite well for me when training big networks. </p>\n\n<p>As you have noticed, we need to train images in full-size while using computationally intensive encoders. While EfficientNetB0 or B1 allows you to use a batch size of 16 when using bigger networks like B5 or B7, you end up using a very small batch size of (usually 4 or maybe 8). </p>\n\n<p>This low batch size can slow down the convergence by making the training noisy. Indeed, Stochastic Gradient Descent computes the average of the gradients computed over mini-batches. A minibatch that is too low will yield gradients that are noisy and that makes the training unstable. Thus we'd like to have bigger minibatch to mitigate noise when updating our weights. One method to do so is <strong>gradient accumulation</strong>. It is usually used in NLP with Transformers but can be easily applied in a computer vision problem.</p>\n\n<p><strong>Please be aware that it won't speed up training since you still need to backpropagate the loss and compute all the gradients. It will speed up the convergence of your model.</strong></p>\n\n<p>Simply put, gradient accumulation allows you to sum gradients over large minibatches, and not updating the weights are every backpropagation.</p>\n\n<p>Please find below, a link for PyTorch and another one for Tensorflow.</p>\n\n<p><strong>PyTorch</strong>: <a href=\"https://medium.com/huggingface/training-larger-batches-practical-tips-on-1-gpu-multi-gpu-distributed-setups-ec88c3e51255\">https://medium.com/huggingface/training-larger-batches-practical-tips-on-1-gpu-multi-gpu-distributed-setups-ec88c3e51255</a>\n<strong>Tensorflow</strong>: <a href=\"https://gchlebus.github.io/2018/06/05/gradient-averaging.html\">https://gchlebus.github.io/2018/06/05/gradient-averaging.html</a></p>",
  "messages": [
    {
      "id": "922995",
      "postDate": "07/10/2020 13:20:28",
      "content": "<p>Hello everyone,</p>\n\n<p>As the competition is drawing to a close, I'd like to share with you a technique that worked quite well for me when training big networks. </p>\n\n<p>As you have noticed, we need to train images in full-size while using computationally intensive encoders. While EfficientNetB0 or B1 allows you to use a batch size of 16 when using bigger networks like B5 or B7, you end up using a very small batch size of (usually 4 or maybe 8). </p>\n\n<p>This low batch size can slow down the convergence by making the training noisy. Indeed, Stochastic Gradient Descent computes the average of the gradients computed over mini-batches. A minibatch that is too low will yield gradients that are noisy and that makes the training unstable. Thus we'd like to have bigger minibatch to mitigate noise when updating our weights. One method to do so is <strong>gradient accumulation</strong>. It is usually used in NLP with Transformers but can be easily applied in a computer vision problem.</p>\n\n<p><strong>Please be aware that it won't speed up training since you still need to backpropagate the loss and compute all the gradients. It will speed up the convergence of your model.</strong></p>\n\n<p>Simply put, gradient accumulation allows you to sum gradients over large minibatches, and not updating the weights are every backpropagation.</p>\n\n<p>Please find below, a link for PyTorch and another one for Tensorflow.</p>\n\n<p><strong>PyTorch</strong>: <a href=\"https://medium.com/huggingface/training-larger-batches-practical-tips-on-1-gpu-multi-gpu-distributed-setups-ec88c3e51255\">https://medium.com/huggingface/training-larger-batches-practical-tips-on-1-gpu-multi-gpu-distributed-setups-ec88c3e51255</a>\n<strong>Tensorflow</strong>: <a href=\"https://gchlebus.github.io/2018/06/05/gradient-averaging.html\">https://gchlebus.github.io/2018/06/05/gradient-averaging.html</a></p>",
      "rawMarkdown": "Hello everyone,\n\nAs the competition is drawing to a close, I'd like to share with you a technique that worked quite well for me when training big networks. \n\nAs you have noticed, we need to train images in full-size while using computationally intensive encoders. While EfficientNetB0 or B1 allows you to use a batch size of 16 when using bigger networks like B5 or B7, you end up using a very small batch size of (usually 4 or maybe 8). \n\nThis low batch size can slow down the convergence by making the training noisy. Indeed, Stochastic Gradient Descent computes the average of the gradients computed over mini-batches. A minibatch that is too low will yield gradients that are noisy and that makes the training unstable. Thus we'd like to have bigger minibatch to mitigate noise when updating our weights. One method to do so is **gradient accumulation**. It is usually used in NLP with Transformers but can be easily applied in a computer vision problem.\n\n**Please be aware that it won't speed up training since you still need to backpropagate the loss and compute all the gradients. It will speed up the convergence of your model.**\n\nSimply put, gradient accumulation allows you to sum gradients over large minibatches, and not updating the weights are every backpropagation.\n\nPlease find below, a link for PyTorch and another one for Tensorflow.\n\n**PyTorch**: [https://medium.com/huggingface/training-larger-batches-practical-tips-on-1-gpu-multi-gpu-distributed-setups-ec88c3e51255](https://medium.com/huggingface/training-larger-batches-practical-tips-on-1-gpu-multi-gpu-distributed-setups-ec88c3e51255)\n**Tensorflow**: [https://gchlebus.github.io/2018/06/05/gradient-averaging.html](https://gchlebus.github.io/2018/06/05/gradient-averaging.html)",
      "votes": null
    },
    {
      "id": "923195",
      "postDate": "07/10/2020 15:44:55",
      "content": "<p>Thanks for sharing. As someone who's trained up some of the larger models, can you offer any insight into the performance gain in this space between, e.g. B0 and B5, etc?</p>",
      "rawMarkdown": "Thanks for sharing. As someone who's trained up some of the larger models, can you offer any insight into the performance gain in this space between, e.g. B0 and B5, etc?",
      "votes": null
    },
    {
      "id": "925269",
      "postDate": "07/12/2020 00:48:27",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": null
    },
    {
      "id": "925335",
      "postDate": "07/12/2020 02:51:37",
      "content": "<p>wow.. Thanks <a href=\"/rftexas\">@rftexas</a>  for sharing the speeding up trick in this time sensitive moment.  Kudos ! 👍 </p>",
      "rawMarkdown": "wow.. Thanks @rftexas  for sharing the speeding up trick in this time sensitive moment.  Kudos ! 👍",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 923195,
      "author_name": "authman",
      "author_url": "",
      "post_date": "07/10/2020 15:44:55",
      "content": "<p>Thanks for sharing. As someone who's trained up some of the larger models, can you offer any insight into the performance gain in this space between, e.g. B0 and B5, etc?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 925269,
      "author_name": "mariapushkareva",
      "author_url": "",
      "post_date": "07/12/2020 00:48:27",
      "content": "<p>Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 925335,
      "author_name": "redwankarimsony",
      "author_url": "",
      "post_date": "07/12/2020 02:51:37",
      "content": "<p>wow.. Thanks <a href=\"/rftexas\">@rftexas</a>  for sharing the speeding up trick in this time sensitive moment.  Kudos ! 👍 </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "922995": "Hello everyone,\n\nAs the competition is drawing to a close, I'd like to share with you a technique that worked quite well for me when training big networks. \n\nAs you have noticed, we need to train images in full-size while using computationally intensive encoders. While EfficientNetB0 or B1 allows you to use a batch size of 16 when using bigger networks like B5 or B7, you end up using a very small batch size of (usually 4 or maybe 8). \n\nThis low batch size can slow down the convergence by making the training noisy. Indeed, Stochastic Gradient Descent computes the average of the gradients computed over mini-batches. A minibatch that is too low will yield gradients that are noisy and that makes the training unstable. Thus we'd like to have bigger minibatch to mitigate noise when updating our weights. One method to do so is **gradient accumulation**. It is usually used in NLP with Transformers but can be easily applied in a computer vision problem.\n\n**Please be aware that it won't speed up training since you still need to backpropagate the loss and compute all the gradients. It will speed up the convergence of your model.**\n\nSimply put, gradient accumulation allows you to sum gradients over large minibatches, and not updating the weights are every backpropagation.\n\nPlease find below, a link for PyTorch and another one for Tensorflow.\n\n**PyTorch**: [https://medium.com/huggingface/training-larger-batches-practical-tips-on-1-gpu-multi-gpu-distributed-setups-ec88c3e51255](https://medium.com/huggingface/training-larger-batches-practical-tips-on-1-gpu-multi-gpu-distributed-setups-ec88c3e51255)\n**Tensorflow**: [https://gchlebus.github.io/2018/06/05/gradient-averaging.html](https://gchlebus.github.io/2018/06/05/gradient-averaging.html)",
    "923195": "Thanks for sharing. As someone who's trained up some of the larger models, can you offer any insight into the performance gain in this space between, e.g. B0 and B5, etc?",
    "925269": "Thanks for sharing!",
    "925335": "wow.. Thanks @rftexas  for sharing the speeding up trick in this time sensitive moment.  Kudos ! 👍"
  },
  "source": "meta"
}