{
  "id": 105614,
  "title": "A trick to use bigger batches for training: gradient accumulation",
  "url": "/competitions/understanding_cloud_organization/discussion/105614",
  "author_name": "Andrey Lukyanenko",
  "post_date": "2019-08-24T17:22:12.419000",
  "votes": 81,
  "comment_count": 28,
  "views": 0,
  "content": "<p>In most cases (not all, for example in GANs) using bigger batches is better. But we usually have a limitation on our GPU. In this competition we have big images <code>1400 x 2100</code> and if we want to use the original size, then even P100 allows only small batches (~2 images per batch).</p>\n\n<p>Gradient accumulation is one of ways to circumvent it. Basically it involves making optimizer steps after several batches thus increasing effective batch size.</p>\n\n<p>Keras code could be found here: <a href=\"https://stackoverflow.com/questions/55268762/how-to-accumulate-gradients-for-large-batch-sizes-in-keras\">https://stackoverflow.com/questions/55268762/how-to-accumulate-gradients-for-large-batch-sizes-in-keras</a></p>\n\n<p>For PyTorch it is easier:</p>\n\n<p>```</p>\n\n<h1><a href=\"https://gist.github.com/thomwolf/ac7a7da6b1888c2eeac8ac8b9b05d3d3\">https://gist.github.com/thomwolf/ac7a7da6b1888c2eeac8ac8b9b05d3d3</a></h1>\n\n<p>for i, (inputs, labels) in enumerate(training_set):\n    predictions = model(inputs)                     # Forward pass\n    loss = loss_function(predictions, labels)       # Compute loss function\n    loss = loss / accumulation_steps                # Normalize our loss (if averaged)\n    loss.backward()                                 # Backward pass\n    if (i+1) % accumulation_steps == 0:             # Wait for several backward steps\n        optimizer.step()                            # Now we can do an optimizer step\n        model.zero_grad()                           # Reset gradients tensors\n        if (i+1) % evaluation_steps == 0:           # Evaluate the model when we...\n            evaluate_model() <br>\n```</p>\n\n<p>Also, if you use <a href=\"https://github.com/catalyst-team/catalyst\">Catalyst</a> , you can simply use a  callback: <code>OptimizerCallback(accumulation_steps=n)</code></p>",
  "messages": [
    {
      "id": 607158,
      "postDate": "2019-08-24T17:22:12.420Z",
      "content": "<p>In most cases (not all, for example in GANs) using bigger batches is better. But we usually have a limitation on our GPU. In this competition we have big images <code>1400 x 2100</code> and if we want to use the original size, then even P100 allows only small batches (~2 images per batch).</p>\n\n<p>Gradient accumulation is one of ways to circumvent it. Basically it involves making optimizer steps after several batches thus increasing effective batch size.</p>\n\n<p>Keras code could be found here: <a href=\"https://stackoverflow.com/questions/55268762/how-to-accumulate-gradients-for-large-batch-sizes-in-keras\">https://stackoverflow.com/questions/55268762/how-to-accumulate-gradients-for-large-batch-sizes-in-keras</a></p>\n\n<p>For PyTorch it is easier:</p>\n\n<p>```</p>\n\n<h1><a href=\"https://gist.github.com/thomwolf/ac7a7da6b1888c2eeac8ac8b9b05d3d3\">https://gist.github.com/thomwolf/ac7a7da6b1888c2eeac8ac8b9b05d3d3</a></h1>\n\n<p>for i, (inputs, labels) in enumerate(training_set):\n    predictions = model(inputs)                     # Forward pass\n    loss = loss_function(predictions, labels)       # Compute loss function\n    loss = loss / accumulation_steps                # Normalize our loss (if averaged)\n    loss.backward()                                 # Backward pass\n    if (i+1) % accumulation_steps == 0:             # Wait for several backward steps\n        optimizer.step()                            # Now we can do an optimizer step\n        model.zero_grad()                           # Reset gradients tensors\n        if (i+1) % evaluation_steps == 0:           # Evaluate the model when we...\n            evaluate_model() <br>\n```</p>\n\n<p>Also, if you use <a href=\"https://github.com/catalyst-team/catalyst\">Catalyst</a> , you can simply use a  callback: <code>OptimizerCallback(accumulation_steps=n)</code></p>",
      "rawMarkdown": "In most cases (not all, for example in GANs) using bigger batches is better. But we usually have a limitation on our GPU. In this competition we have big images `1400 x 2100` and if we want to use the original size, then even P100 allows only small batches (~2 images per batch).\n\nGradient accumulation is one of ways to circumvent it. Basically it involves making optimizer steps after several batches thus increasing effective batch size.\n\nKeras code could be found here: https://stackoverflow.com/questions/55268762/how-to-accumulate-gradients-for-large-batch-sizes-in-keras\n\nFor PyTorch it is easier:\n\n```\n# https://gist.github.com/thomwolf/ac7a7da6b1888c2eeac8ac8b9b05d3d3\nfor i, (inputs, labels) in enumerate(training_set):\n    predictions = model(inputs)                     # Forward pass\n    loss = loss_function(predictions, labels)       # Compute loss function\n    loss = loss / accumulation_steps                # Normalize our loss (if averaged)\n    loss.backward()                                 # Backward pass\n    if (i+1) % accumulation_steps == 0:             # Wait for several backward steps\n        optimizer.step()                            # Now we can do an optimizer step\n        model.zero_grad()                           # Reset gradients tensors\n        if (i+1) % evaluation_steps == 0:           # Evaluate the model when we...\n            evaluate_model()   \n```\n\nAlso, if you use [Catalyst](https://github.com/catalyst-team/catalyst) , you can simply use a  callback: `OptimizerCallback(accumulation_steps=n)`",
      "votes": 80
    },
    {
      "id": 1069558,
      "postDate": "2020-11-04T15:47:47.280Z",
      "content": "<p>This is also very easy to do using <a href=\"https://github.com/PyTorchLightning/pytorch-lightning\" target=\"_blank\">PyTorch Lightning</a>. Simply pass a trainer flag:</p>\n<p><code># accumulate every 4 batches (effective batch size is batch*4)\ntrainer = Trainer(accumulate_grad_batches=4)</code></p>\n<p>You can also change the batch size for different epochs</p>\n<p><code># no accumulation for epochs 1-4. accumulate 3 for epochs 5-10. accumulate 20 after that\ntrainer = Trainer(accumulate_grad_batches={5: 3, 10: 20})</code></p>",
      "rawMarkdown": "This is also very easy to do using [PyTorch Lightning](https://github.com/PyTorchLightning/pytorch-lightning). Simply pass a trainer flag:\n\n`# accumulate every 4 batches (effective batch size is batch*4)\ntrainer = Trainer(accumulate_grad_batches=4)`\n\nYou can also change the batch size for different epochs\n\n`# no accumulation for epochs 1-4. accumulate 3 for epochs 5-10. accumulate 20 after that\ntrainer = Trainer(accumulate_grad_batches={5: 3, 10: 20})`",
      "votes": 3,
      "replies": [
        {
          "id": 1070412,
          "postDate": "2020-11-05T18:17:44.033Z",
          "content": "<p>The PyTorch Lightning is a great thing! So many tricks shipped out of the box.</p>",
          "rawMarkdown": "The PyTorch Lightning is a great thing! So many tricks shipped out of the box."
        },
        {
          "id": 1101544,
          "postDate": "2020-12-04T02:15:06.427Z",
          "content": "<p>Woah! This configuration is crazy good!</p>",
          "rawMarkdown": "Woah! This configuration is crazy good!"
        }
      ]
    },
    {
      "id": 662286,
      "postDate": "2019-10-31T11:34:39.447Z",
      "content": "<p>Another way to increase <code>batch_size</code> is to train with crops and predict on full. I post a starter kernel <a href=\"https://www.kaggle.com/cdeotte/train-with-crops-cv-0-60\">here</a></p>",
      "rawMarkdown": "Another way to increase `batch_size` is to train with crops and predict on full. I post a starter kernel [here][1]\n\n[1]: https://www.kaggle.com/cdeotte/train-with-crops-cv-0-60",
      "votes": 4,
      "replies": [
        {
          "id": 662351,
          "postDate": "2019-10-31T13:16:04.740Z",
          "content": "<p>Thanks <a href=\"/cdeotte\">@cdeotte</a> !</p>\n\n<p>I had started this approach (training only, up to this time)</p>\n\n<p>But I liked to learn about Gradient Accumulation in this competition.</p>\n\n<p>Also, perhaps it's not the competition to learn about this topic.</p>\n\n<p>Are you getting a better model using a larger batch size in this competition?</p>",
          "rawMarkdown": "Thanks @cdeotte !\n\nI had started this approach (training only, up to this time)\n\nBut I liked to learn about Gradient Accumulation in this competition.\n\nAlso, perhaps it's not the competition to learn about this topic.\n\nAre you getting a better model using a larger batch size in this competition?"
        },
        {
          "id": 662360,
          "postDate": "2019-10-31T13:25:30.147Z",
          "content": "<p>I don't know what's better yet. I have only made a few simple models. As I try to improve my models and advance my LB score, I'll post here my observations.</p>\n\n<p>Mathematically, larger batch size is good with noisy labels because averages (large batch) are always less noisy than individuals (small batch). The formula is <code>variance_batch_mean = variance_individuals / batch_size</code></p>",
          "rawMarkdown": "I don't know what's better yet. I have only made a few simple models. As I try to improve my models and advance my LB score, I'll post here my observations.\n\nMathematically, larger batch size is good with noisy labels because averages (large batch) are always less noisy than individuals (small batch). The formula is `variance_batch_mean = variance_individuals / batch_size`",
          "votes": 4
        }
      ]
    },
    {
      "id": 614787,
      "postDate": "2019-09-01T04:28:26.673Z",
      "content": "<p>Thanks for the share, it is very helpful. But if there are BatchNormalization layers in our NN, seems they will lose their effect under gradient accumulation situation, right?</p>",
      "rawMarkdown": "Thanks for the share, it is very helpful. But if there are BatchNormalization layers in our NN, seems they will lose their effect under gradient accumulation situation, right?",
      "votes": 3,
      "replies": [
        {
          "id": 614887,
          "postDate": "2019-09-01T07:44:46.070Z",
          "content": "<p>As far as I understand gradient accumulation is better than training with small batch, but if we can train with batch equal to efficient batch size with gradient accumulation - it would be better to train with bigger batch.</p>",
          "rawMarkdown": "As far as I understand gradient accumulation is better than training with small batch, but if we can train with batch equal to efficient batch size with gradient accumulation - it would be better to train with bigger batch.",
          "votes": 2
        },
        {
          "id": 643946,
          "postDate": "2019-10-08T06:40:52.903Z",
          "content": "<p>I think you can use weight decay for these BN layers.  Or change them by Group Norm...</p>\n\n<p>I haven't tested it yet.</p>\n\n<p>There is a large thread at fastai forums where the topic on BN in gradient accumulation is also taken into account</p>\n\n<p><a href=\"https://forums.fast.ai/t/accumulating-gradients/33219/89\">https://forums.fast.ai/t/accumulating-gradients/33219/89</a></p>",
          "rawMarkdown": "I think you can use weight decay for these BN layers.  Or change them by Group Norm...\n\nI haven't tested it yet.\n\nThere is a large thread at fastai forums where the topic on BN in gradient accumulation is also taken into account\n\nhttps://forums.fast.ai/t/accumulating-gradients/33219/89",
          "votes": 3
        }
      ]
    },
    {
      "id": 616265,
      "postDate": "2019-09-02T23:18:28.963Z",
      "content": "<p>When I tried to use <code>OptimizerCallback(accumulation_steps=n)</code>, it returns <code>'NoneType' object has no attribute 'backward'</code>.</p>",
      "rawMarkdown": "When I tried to use `OptimizerCallback(accumulation_steps=n)`, it returns `'NoneType' object has no attribute 'backward'`.",
      "votes": 4,
      "replies": [
        {
          "id": 616317,
          "postDate": "2019-09-03T02:23:23.790Z",
          "content": "<p>In the current version of Catalyst it is important to pass Callbacks in the correct order and in case of passing optimizer callback, it is necessary to also pass a placeholder <code>CriterionCallback()</code>, like this:</p>\n\n<p><code>\nrunner.train(\n    model=model,\n    criterion=criterion,\n    optimizer=optimizer,\n    scheduler=scheduler,\n    loaders=loaders,\n    callbacks=[DiceCallback(), EarlyStoppingCallback(patience=5, min_delta=0.001), CriterionCallback(), OptimizerCallback(accumulation_steps=4)],\n    logdir=logdir,\n    num_epochs=num_epochs,\n    verbose=True\n)\n</code></p>",
          "rawMarkdown": "In the current version of Catalyst it is important to pass Callbacks in the correct order and in case of passing optimizer callback, it is necessary to also pass a placeholder `CriterionCallback()`, like this:\n\n```\nrunner.train(\n    model=model,\n    criterion=criterion,\n    optimizer=optimizer,\n    scheduler=scheduler,\n    loaders=loaders,\n    callbacks=[DiceCallback(), EarlyStoppingCallback(patience=5, min_delta=0.001), CriterionCallback(), OptimizerCallback(accumulation_steps=4)],\n    logdir=logdir,\n    num_epochs=num_epochs,\n    verbose=True\n)\n```",
          "votes": 6
        },
        {
          "id": 616436,
          "postDate": "2019-09-03T06:11:17.407Z",
          "content": "<p>Thanks!</p>",
          "rawMarkdown": "Thanks!",
          "votes": 3
        },
        {
          "id": 616439,
          "postDate": "2019-09-03T06:14:33.620Z",
          "content": "<p>BTW, i think it is better to also adjust the learning rate after using gradient accumulation. Is it correct?</p>",
          "rawMarkdown": "BTW, i think it is better to also adjust the learning rate after using gradient accumulation. Is it correct?",
          "votes": 3
        },
        {
          "id": 616444,
          "postDate": "2019-09-03T06:20:27.990Z",
          "content": "<p>I agree, usually we adjust learning rate when we increase batch size. But we can treat it like any hyperparameter and optimize it.</p>",
          "rawMarkdown": "I agree, usually we adjust learning rate when we increase batch size. But we can treat it like any hyperparameter and optimize it.",
          "votes": 4
        },
        {
          "id": 638324,
          "postDate": "2019-10-01T19:19:44.397Z",
          "content": "<p>I also came across this problem but was unable to fix it (until I saw this post). How did you know to apply this fix?</p>",
          "rawMarkdown": "I also came across this problem but was unable to fix it (until I saw this post). How did you know to apply this fix?",
          "votes": 1
        },
        {
          "id": 638328,
          "postDate": "2019-10-01T19:23:09.220Z",
          "content": "<p>I asked the developers of catalyst :)\nThey are active in channel #tool_catalyst in ods.ai slack or you could write them here: <a href=\"https://gitter.im/catalyst-team/community?utm_source=badge&amp;utm_medium=badge&amp;utm_campaign=pr-badge\">https://gitter.im/catalyst-team/community?utm_source=badge&amp;utm_medium=badge&amp;utm_campaign=pr-badge</a>\nOf course, you can also create an issue on github.</p>",
          "rawMarkdown": "I asked the developers of catalyst :)\nThey are active in channel #tool_catalyst in ods.ai slack or you could write them here: https://gitter.im/catalyst-team/community?utm_source=badge&amp;utm_medium=badge&amp;utm_campaign=pr-badge\nOf course, you can also create an issue on github.",
          "votes": 1
        },
        {
          "id": 638350,
          "postDate": "2019-10-01T19:51:36.340Z",
          "content": "<p>nice...</p>",
          "rawMarkdown": "nice...",
          "votes": 1
        }
      ]
    },
    {
      "id": 643678,
      "postDate": "2019-10-07T19:15:32.647Z",
      "content": "<p>Thank you!\nI had a batch size of 64 and my kaggle kernel was dying because of high RAM usage/CUDA out of memory on GPU. \nThis might be very helpful. </p>",
      "rawMarkdown": "Thank you!\nI had a batch size of 64 and my kaggle kernel was dying because of high RAM usage/CUDA out of memory on GPU. \nThis might be very helpful. ",
      "votes": 1
    },
    {
      "id": 636737,
      "postDate": "2019-09-30T06:01:41.260Z",
      "content": "<p>Thank you for your sharing.\nI found keras's <a href=\"https://pypi.org/project/keras-gradient-accumulation/\">gradient accumulation package</a></p>",
      "rawMarkdown": "Thank you for your sharing.\nI found keras's [gradient accumulation package](https://pypi.org/project/keras-gradient-accumulation/)",
      "votes": 1
    },
    {
      "id": 607198,
      "postDate": "2019-08-24T19:06:28.653Z",
      "content": "<p>Thanks for the information. Really helpful. Got to learn something new!!</p>",
      "rawMarkdown": "Thanks for the information. Really helpful. Got to learn something new!!",
      "votes": 1
    },
    {
      "id": 607573,
      "postDate": "2019-08-25T14:03:10.920Z",
      "content": "<p>awesome stuff!</p>",
      "rawMarkdown": "awesome stuff!",
      "votes": 2
    },
    {
      "id": 607216,
      "postDate": "2019-08-24T20:08:16.660Z",
      "content": "<p>At first glance, the Keras implementation looks incredibly involved :|\nMaybe with low-level TensorFlow API it is more straightforward? Like, with gradient tape or something?</p>",
      "rawMarkdown": "At first glance, the Keras implementation looks incredibly involved :|\nMaybe with low-level TensorFlow API it is more straightforward? Like, with gradient tape or something?",
      "votes": 2
    },
    {
      "id": 1101546,
      "postDate": "2020-12-04T02:18:45Z",
      "content": "<p>Thanks for initiating this thread. </p>\n<p>I was wondering should we keep on adding up the losses calculated during the accumulation steps before performing the updates? </p>",
      "rawMarkdown": "Thanks for initiating this thread. \n\nI was wondering should we keep on adding up the losses calculated during the accumulation steps before performing the updates? "
    },
    {
      "id": 925967,
      "postDate": "2020-07-12T12:01:20.687Z",
      "content": "<p>👍 👍 </p>",
      "rawMarkdown": "👍 👍 "
    },
    {
      "id": 661798,
      "postDate": "2019-10-30T18:13:39.997Z",
      "content": "<p>I'm in a mess about what should I do with BatchNorm layers</p>\n\n<p>Using Pytorch/Catalyst.  Thanks <a href=\"/artgor\">@artgor</a> !</p>\n\n<p>I tried replacing them by Group Norm, but the results are worse.</p>\n\n<p>Naive gradient accumulation (no replacements) is even worst.</p>\n\n<p>I tried with <a href=\"/sgugger\">@sgugger</a>  implementation of AccumulateBN <a href=\"https://forums.fast.ai/t/accumulating-gradients/33219/74\">https://forums.fast.ai/t/accumulating-gradients/33219/74</a>\nBut there must be a memory leakage.</p>\n\n<p>Any help would be very welcome</p>",
      "rawMarkdown": "I'm in a mess about what should I do with BatchNorm layers\n\nUsing Pytorch/Catalyst.  Thanks @artgor !\n\nI tried replacing them by Group Norm, but the results are worse.\n\nNaive gradient accumulation (no replacements) is even worst.\n\nI tried with @sgugger  implementation of AccumulateBN https://forums.fast.ai/t/accumulating-gradients/33219/74\nBut there must be a memory leakage.\n\nAny help would be very welcome\n\n"
    },
    {
      "id": 638529,
      "postDate": "2019-10-02T02:29:57.720Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 615361,
      "postDate": "2019-09-01T20:21:50.180Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 607214,
      "postDate": "2019-08-24T20:06:40.857Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1069558,
      "author_name": "edenaf",
      "author_url": "",
      "post_date": "2020-11-04T15:47:47.280000",
      "content": "<p>This is also very easy to do using <a href=\"https://github.com/PyTorchLightning/pytorch-lightning\" target=\"_blank\">PyTorch Lightning</a>. Simply pass a trainer flag:</p>\n<p><code># accumulate every 4 batches (effective batch size is batch*4)\ntrainer = Trainer(accumulate_grad_batches=4)</code></p>\n<p>You can also change the batch size for different epochs</p>\n<p><code># no accumulation for epochs 1-4. accumulate 3 for epochs 5-10. accumulate 20 after that\ntrainer = Trainer(accumulate_grad_batches={5: 3, 10: 20})</code></p>",
      "votes": 3,
      "replies": [
        {
          "id": 1070412,
          "author_name": "Ilia Zaitsev",
          "author_url": "",
          "post_date": "2020-11-05T18:17:44.033000",
          "content": "<p>The PyTorch Lightning is a great thing! So many tricks shipped out of the box.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1101544,
          "author_name": "Sayak",
          "author_url": "",
          "post_date": "2020-12-04T02:15:06.427000",
          "content": "<p>Woah! This configuration is crazy good!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 662286,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2019-10-31T11:34:39.447000",
      "content": "<p>Another way to increase <code>batch_size</code> is to train with crops and predict on full. I post a starter kernel <a href=\"https://www.kaggle.com/cdeotte/train-with-crops-cv-0-60\">here</a></p>",
      "votes": 4,
      "replies": [
        {
          "id": 662351,
          "author_name": "Virilo Tejedor Aguilera",
          "author_url": "",
          "post_date": "2019-10-31T13:16:04.740000",
          "content": "<p>Thanks <a href=\"/cdeotte\">@cdeotte</a> !</p>\n\n<p>I had started this approach (training only, up to this time)</p>\n\n<p>But I liked to learn about Gradient Accumulation in this competition.</p>\n\n<p>Also, perhaps it's not the competition to learn about this topic.</p>\n\n<p>Are you getting a better model using a larger batch size in this competition?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 662360,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2019-10-31T13:25:30.147000",
          "content": "<p>I don't know what's better yet. I have only made a few simple models. As I try to improve my models and advance my LB score, I'll post here my observations.</p>\n\n<p>Mathematically, larger batch size is good with noisy labels because averages (large batch) are always less noisy than individuals (small batch). The formula is <code>variance_batch_mean = variance_individuals / batch_size</code></p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 614787,
      "author_name": "Tsai29",
      "author_url": "",
      "post_date": "2019-09-01T04:28:26.673000",
      "content": "<p>Thanks for the share, it is very helpful. But if there are BatchNormalization layers in our NN, seems they will lose their effect under gradient accumulation situation, right?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 614887,
          "author_name": "Andrey Lukyanenko",
          "author_url": "",
          "post_date": "2019-09-01T07:44:46.070000",
          "content": "<p>As far as I understand gradient accumulation is better than training with small batch, but if we can train with batch equal to efficient batch size with gradient accumulation - it would be better to train with bigger batch.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 643946,
          "author_name": "Virilo Tejedor Aguilera",
          "author_url": "",
          "post_date": "2019-10-08T06:40:52.903000",
          "content": "<p>I think you can use weight decay for these BN layers.  Or change them by Group Norm...</p>\n\n<p>I haven't tested it yet.</p>\n\n<p>There is a large thread at fastai forums where the topic on BN in gradient accumulation is also taken into account</p>\n\n<p><a href=\"https://forums.fast.ai/t/accumulating-gradients/33219/89\">https://forums.fast.ai/t/accumulating-gradients/33219/89</a></p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 616265,
      "author_name": "Yirun Zhang",
      "author_url": "",
      "post_date": "2019-09-02T23:18:28.963000",
      "content": "<p>When I tried to use <code>OptimizerCallback(accumulation_steps=n)</code>, it returns <code>'NoneType' object has no attribute 'backward'</code>.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 616317,
          "author_name": "Andrey Lukyanenko",
          "author_url": "",
          "post_date": "2019-09-03T02:23:23.790000",
          "content": "<p>In the current version of Catalyst it is important to pass Callbacks in the correct order and in case of passing optimizer callback, it is necessary to also pass a placeholder <code>CriterionCallback()</code>, like this:</p>\n\n<p><code>\nrunner.train(\n    model=model,\n    criterion=criterion,\n    optimizer=optimizer,\n    scheduler=scheduler,\n    loaders=loaders,\n    callbacks=[DiceCallback(), EarlyStoppingCallback(patience=5, min_delta=0.001), CriterionCallback(), OptimizerCallback(accumulation_steps=4)],\n    logdir=logdir,\n    num_epochs=num_epochs,\n    verbose=True\n)\n</code></p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 616436,
          "author_name": "Yirun Zhang",
          "author_url": "",
          "post_date": "2019-09-03T06:11:17.407000",
          "content": "<p>Thanks!</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 616439,
          "author_name": "Yirun Zhang",
          "author_url": "",
          "post_date": "2019-09-03T06:14:33.620000",
          "content": "<p>BTW, i think it is better to also adjust the learning rate after using gradient accumulation. Is it correct?</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 616444,
          "author_name": "Andrey Lukyanenko",
          "author_url": "",
          "post_date": "2019-09-03T06:20:27.990000",
          "content": "<p>I agree, usually we adjust learning rate when we increase batch size. But we can treat it like any hyperparameter and optimize it.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 638324,
          "author_name": "pete",
          "author_url": "",
          "post_date": "2019-10-01T19:19:44.397000",
          "content": "<p>I also came across this problem but was unable to fix it (until I saw this post). How did you know to apply this fix?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 638328,
          "author_name": "Andrey Lukyanenko",
          "author_url": "",
          "post_date": "2019-10-01T19:23:09.220000",
          "content": "<p>I asked the developers of catalyst :)\nThey are active in channel #tool_catalyst in ods.ai slack or you could write them here: <a href=\"https://gitter.im/catalyst-team/community?utm_source=badge&amp;utm_medium=badge&amp;utm_campaign=pr-badge\">https://gitter.im/catalyst-team/community?utm_source=badge&amp;utm_medium=badge&amp;utm_campaign=pr-badge</a>\nOf course, you can also create an issue on github.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 638350,
          "author_name": "pete",
          "author_url": "",
          "post_date": "2019-10-01T19:51:36.340000",
          "content": "<p>nice...</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 643678,
      "author_name": "timetraveller",
      "author_url": "",
      "post_date": "2019-10-07T19:15:32.647000",
      "content": "<p>Thank you!\nI had a batch size of 64 and my kaggle kernel was dying because of high RAM usage/CUDA out of memory on GPU. \nThis might be very helpful. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 636737,
      "author_name": "ikki1111",
      "author_url": "",
      "post_date": "2019-09-30T06:01:41.260000",
      "content": "<p>Thank you for your sharing.\nI found keras's <a href=\"https://pypi.org/project/keras-gradient-accumulation/\">gradient accumulation package</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 607198,
      "author_name": "Shweta Goyal",
      "author_url": "",
      "post_date": "2019-08-24T19:06:28.653000",
      "content": "<p>Thanks for the information. Really helpful. Got to learn something new!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 607573,
      "author_name": "Hamish",
      "author_url": "",
      "post_date": "2019-08-25T14:03:10.920000",
      "content": "<p>awesome stuff!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 607216,
      "author_name": "Ilia Zaitsev",
      "author_url": "",
      "post_date": "2019-08-24T20:08:16.660000",
      "content": "<p>At first glance, the Keras implementation looks incredibly involved :|\nMaybe with low-level TensorFlow API it is more straightforward? Like, with gradient tape or something?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1101546,
      "author_name": "Sayak",
      "author_url": "",
      "post_date": "2020-12-04T02:18:45",
      "content": "<p>Thanks for initiating this thread. </p>\n<p>I was wondering should we keep on adding up the losses calculated during the accumulation steps before performing the updates? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 925967,
      "author_name": "leilaSoleymani",
      "author_url": "",
      "post_date": "2020-07-12T12:01:20.687000",
      "content": "<p>👍 👍 </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 661798,
      "author_name": "Virilo Tejedor Aguilera",
      "author_url": "",
      "post_date": "2019-10-30T18:13:39.997000",
      "content": "<p>I'm in a mess about what should I do with BatchNorm layers</p>\n\n<p>Using Pytorch/Catalyst.  Thanks <a href=\"/artgor\">@artgor</a> !</p>\n\n<p>I tried replacing them by Group Norm, but the results are worse.</p>\n\n<p>Naive gradient accumulation (no replacements) is even worst.</p>\n\n<p>I tried with <a href=\"/sgugger\">@sgugger</a>  implementation of AccumulateBN <a href=\"https://forums.fast.ai/t/accumulating-gradients/33219/74\">https://forums.fast.ai/t/accumulating-gradients/33219/74</a>\nBut there must be a memory leakage.</p>\n\n<p>Any help would be very welcome</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 638529,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-10-02T02:29:57.720000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 615361,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-01T20:21:50.180000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 607214,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-08-24T20:06:40.857000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "607158": "In most cases (not all, for example in GANs) using bigger batches is better. But we usually have a limitation on our GPU. In this competition we have big images `1400 x 2100` and if we want to use the original size, then even P100 allows only small batches (~2 images per batch).\n\nGradient accumulation is one of ways to circumvent it. Basically it involves making optimizer steps after several batches thus increasing effective batch size.\n\nKeras code could be found here: https://stackoverflow.com/questions/55268762/how-to-accumulate-gradients-for-large-batch-sizes-in-keras\n\nFor PyTorch it is easier:\n\n```\n# https://gist.github.com/thomwolf/ac7a7da6b1888c2eeac8ac8b9b05d3d3\nfor i, (inputs, labels) in enumerate(training_set):\n    predictions = model(inputs)                     # Forward pass\n    loss = loss_function(predictions, labels)       # Compute loss function\n    loss = loss / accumulation_steps                # Normalize our loss (if averaged)\n    loss.backward()                                 # Backward pass\n    if (i+1) % accumulation_steps == 0:             # Wait for several backward steps\n        optimizer.step()                            # Now we can do an optimizer step\n        model.zero_grad()                           # Reset gradients tensors\n        if (i+1) % evaluation_steps == 0:           # Evaluate the model when we...\n            evaluate_model()   \n```\n\nAlso, if you use [Catalyst](https://github.com/catalyst-team/catalyst) , you can simply use a  callback: `OptimizerCallback(accumulation_steps=n)`",
    "1069558": "This is also very easy to do using [PyTorch Lightning](https://github.com/PyTorchLightning/pytorch-lightning). Simply pass a trainer flag:\n\n`# accumulate every 4 batches (effective batch size is batch*4)\ntrainer = Trainer(accumulate_grad_batches=4)`\n\nYou can also change the batch size for different epochs\n\n`# no accumulation for epochs 1-4. accumulate 3 for epochs 5-10. accumulate 20 after that\ntrainer = Trainer(accumulate_grad_batches={5: 3, 10: 20})`",
    "662286": "Another way to increase `batch_size` is to train with crops and predict on full. I post a starter kernel [here][1]\n\n[1]: https://www.kaggle.com/cdeotte/train-with-crops-cv-0-60",
    "614787": "Thanks for the share, it is very helpful. But if there are BatchNormalization layers in our NN, seems they will lose their effect under gradient accumulation situation, right?",
    "616265": "When I tried to use `OptimizerCallback(accumulation_steps=n)`, it returns `'NoneType' object has no attribute 'backward'`.",
    "643678": "Thank you!\nI had a batch size of 64 and my kaggle kernel was dying because of high RAM usage/CUDA out of memory on GPU. \nThis might be very helpful. ",
    "636737": "Thank you for your sharing.\nI found keras's [gradient accumulation package](https://pypi.org/project/keras-gradient-accumulation/)",
    "607198": "Thanks for the information. Really helpful. Got to learn something new!!",
    "607573": "awesome stuff!",
    "607216": "At first glance, the Keras implementation looks incredibly involved :|\nMaybe with low-level TensorFlow API it is more straightforward? Like, with gradient tape or something?",
    "1101546": "Thanks for initiating this thread. \n\nI was wondering should we keep on adding up the losses calculated during the accumulation steps before performing the updates? ",
    "925967": "👍 👍 ",
    "661798": "I'm in a mess about what should I do with BatchNorm layers\n\nUsing Pytorch/Catalyst.  Thanks @artgor !\n\nI tried replacing them by Group Norm, but the results are worse.\n\nNaive gradient accumulation (no replacements) is even worst.\n\nI tried with @sgugger  implementation of AccumulateBN https://forums.fast.ai/t/accumulating-gradients/33219/74\nBut there must be a memory leakage.\n\nAny help would be very welcome\n\n",
    "638529": "",
    "615361": "",
    "607214": ""
  }
}