{
  "id": 104686,
  "title": "Group Normalization",
  "url": "/competitions/aptos2019-blindness-detection/discussion/104686",
  "author_name": "Carlo",
  "post_date": "2019-08-18T13:08:18.102000",
  "votes": 34,
  "comment_count": 20,
  "views": 0,
  "content": "<p>Hi Kagglers,</p>\n\n<p>Did any of you try replacing the batch normalization layers in your models with group normalization when using a small batch size? It seems to be a great alternative, especially when batch sizes get really small. </p>\n\n<p>Did you encounter problems using batch normalization with small batch sizes and/or are you having success using group normalization layers.</p>\n\n<p>This image shows the problem with using batch normalization and small batch sizes:\n</p>\n\n<p></p>\n\n<p>I used the following loop to replace the layers in <a href=\"https://www.kaggle.com/carlolepelaars/efficientnetb5-with-keras-aptos-2019\">this Kaggle kernel</a>:</p>\n\n<p>```</p>\n\n<h1>Load in EfficientNetB5</h1>\n\n<p>effnet = EfficientNetB5(...)</p>\n\n<h1>Replace all Batch Normalization layers by Group Normalization layers</h1>\n\n<p>for i, layer in enumerate(effnet.layers):\n    if \"batch_normalization\" in layer.name:\n        effnet.layers[i] = GroupNormalization(groups=2, axis=-1, epsilon=0.1)\n```</p>\n\n<p>Group Normalization paper:\n<a href=\"https://arxiv.org/pdf/1803.08494.pdf\">https://arxiv.org/pdf/1803.08494.pdf</a></p>\n\n<p>Keras implementation:\n<a href=\"https://github.com/titu1994/Keras-Group-Normalization/blob/master/group_norm.py\">https://github.com/titu1994/Keras-Group-Normalization/blob/master/group_norm.py</a></p>",
  "messages": [
    {
      "id": 602012,
      "postDate": "2019-08-18T13:08:18.103Z",
      "content": "<p>Hi Kagglers,</p>\n\n<p>Did any of you try replacing the batch normalization layers in your models with group normalization when using a small batch size? It seems to be a great alternative, especially when batch sizes get really small. </p>\n\n<p>Did you encounter problems using batch normalization with small batch sizes and/or are you having success using group normalization layers.</p>\n\n<p>This image shows the problem with using batch normalization and small batch sizes:\n</p>\n\n<p></p>\n\n<p>I used the following loop to replace the layers in <a href=\"https://www.kaggle.com/carlolepelaars/efficientnetb5-with-keras-aptos-2019\">this Kaggle kernel</a>:</p>\n\n<p>```</p>\n\n<h1>Load in EfficientNetB5</h1>\n\n<p>effnet = EfficientNetB5(...)</p>\n\n<h1>Replace all Batch Normalization layers by Group Normalization layers</h1>\n\n<p>for i, layer in enumerate(effnet.layers):\n    if \"batch_normalization\" in layer.name:\n        effnet.layers[i] = GroupNormalization(groups=2, axis=-1, epsilon=0.1)\n```</p>\n\n<p>Group Normalization paper:\n<a href=\"https://arxiv.org/pdf/1803.08494.pdf\">https://arxiv.org/pdf/1803.08494.pdf</a></p>\n\n<p>Keras implementation:\n<a href=\"https://github.com/titu1994/Keras-Group-Normalization/blob/master/group_norm.py\">https://github.com/titu1994/Keras-Group-Normalization/blob/master/group_norm.py</a></p>",
      "rawMarkdown": "Hi Kagglers,\n\nDid any of you try replacing the batch normalization layers in your models with group normalization when using a small batch size? It seems to be a great alternative, especially when batch sizes get really small. \n\nDid you encounter problems using batch normalization with small batch sizes and/or are you having success using group normalization layers.\n\nThis image shows the problem with using batch normalization and small batch sizes:\n![GN_image](https://miro.medium.com/max/764/0*h7tx8LWRObeqqV53.)\n\n![GN_comparison](https://miro.medium.com/max/1400/0*P2BEbb-GFU1TCS0A.)\n\nI used the following loop to replace the layers in [this Kaggle kernel](https://www.kaggle.com/carlolepelaars/efficientnetb5-with-keras-aptos-2019):\n\n```\n# Load in EfficientNetB5\neffnet = EfficientNetB5(...)\n# Replace all Batch Normalization layers by Group Normalization layers\nfor i, layer in enumerate(effnet.layers):\n    if \"batch_normalization\" in layer.name:\n        effnet.layers[i] = GroupNormalization(groups=2, axis=-1, epsilon=0.1)\n```\n\nGroup Normalization paper:\nhttps://arxiv.org/pdf/1803.08494.pdf\n\nKeras implementation:\nhttps://github.com/titu1994/Keras-Group-Normalization/blob/master/group_norm.py\n\n",
      "votes": 32
    },
    {
      "id": 603684,
      "postDate": "2019-08-20T14:52:10.883Z",
      "content": "<p>Replacing can be done in PyTorch using the following function:</p>\n\n<p>````\ndef convert_layers(model, layer_type_old, layer_type_new, convert_weights=False, num_groups=None):\n    for name, module in reversed(model._modules.items()):\n        if len(list(module.children())) &gt; 0:\n            # recurse\n            model._modules[name] = convert_layers(module, layer_type_old, layer_type_new, convert_weights)</p>\n\n<pre><code>    if type(module) == layer_type_old:\n        layer_old = module\n        layer_new = layer_type_new(module.num_features if num_groups is None else num_groups, module.num_features, module.eps, module.affine) \n\n        if convert_weights:\n            layer_new.weight = layer_old.weight\n            layer_new.bias = layer_old.bias\n\n        model._modules[name] = layer_new\n\nreturn model\n</code></pre>\n\n<h1>Replace BatchNorm with GroupNorm</h1>\n\n<p>convert_layers(model, nn.BatchNorm2d, nn.GroupNorm, True, num_groups=2)\n````</p>\n\n<p>If num_groups is 1, GroupNorm turns into LayerNorm. If num_groups is None, GroupNorm turns into InstanceNorm</p>",
      "rawMarkdown": "Replacing can be done in PyTorch using the following function:\n\n````\ndef convert_layers(model, layer_type_old, layer_type_new, convert_weights=False, num_groups=None):\n    for name, module in reversed(model._modules.items()):\n        if len(list(module.children())) &gt; 0:\n            # recurse\n            model._modules[name] = convert_layers(module, layer_type_old, layer_type_new, convert_weights)\n\n        if type(module) == layer_type_old:\n            layer_old = module\n            layer_new = layer_type_new(module.num_features if num_groups is None else num_groups, module.num_features, module.eps, module.affine) \n\n            if convert_weights:\n                layer_new.weight = layer_old.weight\n                layer_new.bias = layer_old.bias\n\n            model._modules[name] = layer_new\n\n    return model\n\n\n# Replace BatchNorm with GroupNorm\nconvert_layers(model, nn.BatchNorm2d, nn.GroupNorm, True, num_groups=2)\n````\n\nIf num\\_groups is 1, GroupNorm turns into LayerNorm. If num\\_groups is None, GroupNorm turns into InstanceNorm",
      "votes": 13,
      "replies": [
        {
          "id": 603700,
          "postDate": "2019-08-20T15:10:41.857Z",
          "content": "<p>Nice! Thanks!</p>",
          "rawMarkdown": "Nice! Thanks!",
          "votes": 2
        },
        {
          "id": 904823,
          "postDate": "2020-06-28T01:21:48.990Z",
          "content": "<p>Thanks for the code!!! </p>\n\n<p>I found slight change, in <code>model._modules[name] = convert_layers</code>, the <code>num_groups</code> should pass in or the GroupNorm would be InstanceNorm. Can print the model to check the setting.</p>\n\n<p>Modified code is as follows:\n```</p>\n\n<h1>Replace BatchNorm with GroupNorm</h1>\n\n<p>def convert_layers(model, old_layer_type, new_layer_type, convert_weights=False, num_groups=None):</p>\n\n<pre><code>for name, module in reversed(model._modules.items()):\n\n    if len(list(module.children())) &gt; 0:\n        # recurse\n        model._modules[name] = convert_layers(module, old_layer_type, new_layer_type, convert_weights, num_groups=num_groups)\n\n    # single module\n    if type(module) == old_layer_type:\n        old_layer = module\n        new_layer = new_layer_type(module.num_features if num_groups is None else num_groups, module.num_features, module.eps, module.affine) \n\n        if convert_weights:\n            new_layer.weight = old_layer.weight\n            new_layer.bias = old_layer.bias\n\n        model._modules[name] = new_layer\n\nreturn model\n</code></pre>\n\n<p>```</p>",
          "rawMarkdown": "Thanks for the code!!! \n\nI found slight change, in `model._modules[name] = convert_layers`, the `num_groups` should pass in or the GroupNorm would be InstanceNorm. Can print the model to check the setting.\n\nModified code is as follows:\n```\n# Replace BatchNorm with GroupNorm\ndef convert_layers(model, old_layer_type, new_layer_type, convert_weights=False, num_groups=None):\n    \n    for name, module in reversed(model._modules.items()):\n    \n        if len(list(module.children())) &gt; 0:\n            # recurse\n            model._modules[name] = convert_layers(module, old_layer_type, new_layer_type, convert_weights, num_groups=num_groups)\n\n        # single module\n        if type(module) == old_layer_type:\n            old_layer = module\n            new_layer = new_layer_type(module.num_features if num_groups is None else num_groups, module.num_features, module.eps, module.affine) \n\n            if convert_weights:\n                new_layer.weight = old_layer.weight\n                new_layer.bias = old_layer.bias\n\n            model._modules[name] = new_layer\n\n    return model\n```",
          "votes": 2,
          "isDeleted": true
        },
        {
          "id": 991717,
          "postDate": "2020-08-30T16:00:34.570Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 602212,
      "postDate": "2019-08-18T19:17:42.397Z",
      "content": "<p>thanks a lot <a href=\"/carlolepelaars\">@carlolepelaars</a> !!!</p>\n\n<p>are you sure this code is working?  have you printed the summary of the CNN to check that these layers have been replaced as you liked?</p>\n\n<p>as far as I can see, you are creating a new GroupNormalization object and giving to \"layer\" a reference to it, but doing nothing with \"layer\"\nit shouldn't be enough to update the CNN </p>",
      "rawMarkdown": "thanks a lot @carlolepelaars !!!\n\nare you sure this code is working?  have you printed the summary of the CNN to check that these layers have been replaced as you liked?\n\nas far as I can see, you are creating a new GroupNormalization object and giving to \"layer\" a reference to it, but doing nothing with \"layer\"\nit shouldn't be enough to update the CNN ",
      "votes": 3,
      "replies": [
        {
          "id": 602231,
          "postDate": "2019-08-18T19:50:28.407Z",
          "content": "<p>Oops, you're right! Thank you! I updated the post with working code.</p>",
          "rawMarkdown": "Oops, you're right! Thank you! I updated the post with working code.",
          "votes": 1
        },
        {
          "id": 608144,
          "postDate": "2019-08-26T12:01:57.773Z",
          "content": "<p>I don't think that keras snippet physically replaces the layers. Plotting the model immediately after changing the layer name, shows the new layer but it doesn't look like its been built correctly. Plotting the model again after compiling and running a model summary it reverts back to the standard batch norm layer.</p>",
          "rawMarkdown": "I don't think that keras snippet physically replaces the layers. Plotting the model immediately after changing the layer name, shows the new layer but it doesn't look like its been built correctly. Plotting the model again after compiling and running a model summary it reverts back to the standard batch norm layer.",
          "votes": 1
        },
        {
          "id": 616027,
          "postDate": "2019-09-02T16:01:08.927Z",
          "content": "<p>I made a brief contrast exp in Keras, I agree with u so far</p>",
          "rawMarkdown": "I made a brief contrast exp in Keras, I agree with u so far"
        }
      ]
    },
    {
      "id": 603795,
      "postDate": "2019-08-20T17:06:44.590Z",
      "content": "<p>did you get any LB boost by using Group Norm</p>",
      "rawMarkdown": "did you get any LB boost by using Group Norm",
      "votes": 1,
      "replies": [
        {
          "id": 616121,
          "postDate": "2019-09-02T17:57:32.383Z",
          "content": "<p>Hi Sabbir, the main thing was that it allowed me to use small batch sizes (e.g. 4) without the training becoming unstable. Being able to use larger input sizes gave an LB boost. </p>\n\n<p>Hope this helps!</p>",
          "rawMarkdown": "Hi Sabbir, the main thing was that it allowed me to use small batch sizes (e.g. 4) without the training becoming unstable. Being able to use larger input sizes gave an LB boost. \n\nHope this helps!",
          "votes": 1
        }
      ]
    },
    {
      "id": 602922,
      "postDate": "2019-08-19T16:48:24.497Z",
      "content": "<p>thank you <a href=\"/carlolepelaars\">@carlolepelaars</a> .\nIt helps me lot</p>",
      "rawMarkdown": "thank you @carlolepelaars .\nIt helps me lot",
      "votes": 1,
      "replies": [
        {
          "id": 602925,
          "postDate": "2019-08-19T16:52:34.677Z",
          "content": "<p>Nice! You're welcome!</p>",
          "rawMarkdown": "Nice! You're welcome!",
          "votes": 2
        }
      ]
    },
    {
      "id": 602021,
      "postDate": "2019-08-18T13:14:54.033Z",
      "content": "<p><a href=\"/carlolepelaars\">@carlolepelaars</a> Insightful!</p>",
      "rawMarkdown": "@carlolepelaars Insightful!",
      "votes": 1,
      "replies": [
        {
          "id": 602067,
          "postDate": "2019-08-18T14:19:48.610Z",
          "content": "<p>👍 </p>",
          "rawMarkdown": "👍 "
        }
      ]
    },
    {
      "id": 607323,
      "postDate": "2019-08-25T03:21:13.627Z",
      "content": "<p>I was experimenting GroupNorm on  a segmentation problem (SIIM). It performed worse than normal BatchNorm. It looks GroupNorm can also be sensitive to the number of groups, and it varies for datasets. It would be intersting to know if it helps here.</p>",
      "rawMarkdown": "I was experimenting GroupNorm on  a segmentation problem (SIIM). It performed worse than normal BatchNorm. It looks GroupNorm can also be sensitive to the number of groups, and it varies for datasets. It would be intersting to know if it helps here.",
      "votes": 2
    },
    {
      "id": 606349,
      "postDate": "2019-08-23T13:21:53.073Z",
      "content": "<p>But, If you use GN, the pretrained weights with BN may be destroied. It can not take the advantages of transfer learning.</p>",
      "rawMarkdown": "But, If you use GN, the pretrained weights with BN may be destroied. It can not take the advantages of transfer learning.",
      "votes": 2,
      "replies": [
        {
          "id": 606842,
          "postDate": "2019-08-24T07:02:19.780Z",
          "content": "<p>If you use PyTorch, the weights can also be transfered. I'm not sure that it makes sense to transfer from different kind of layers, but it can be done.</p>",
          "rawMarkdown": "If you use PyTorch, the weights can also be transfered. I'm not sure that it makes sense to transfer from different kind of layers, but it can be done.",
          "votes": 1
        }
      ]
    },
    {
      "id": 602025,
      "postDate": "2019-08-18T13:20:06.037Z",
      "content": "<p>Is it possible to use gradient accumulation to solve the problem? You can compute the gradients for 4 epochs using a batch size of 4 and only update the network once. This should be equivalent to training with a batch size of 16.</p>",
      "rawMarkdown": "Is it possible to use gradient accumulation to solve the problem? You can compute the gradients for 4 epochs using a batch size of 4 and only update the network once. This should be equivalent to training with a batch size of 16.",
      "votes": 2,
      "replies": [
        {
          "id": 602058,
          "postDate": "2019-08-18T14:03:56.213Z",
          "content": "<p>Yes, good point! However, it seems like gradient accumulation doesn't prevent the calculation of inaccurate batch statistics in the batch normalization layers, since the normalization is still done over a batch size of 4. But, you're right that there are multiple methods to address the problem.</p>\n\n<p>The mention techniques like <a href=\"https://arxiv.org/pdf/1702.03275.pdf\">Batch Renormalization</a> and <a href=\"https://arxiv.org/pdf/1904.00346.pdf\">Group Convolutions</a> to address the problem of small batch sizes in the paper.</p>",
          "rawMarkdown": "Yes, good point! However, it seems like gradient accumulation doesn't prevent the calculation of inaccurate batch statistics in the batch normalization layers, since the normalization is still done over a batch size of 4. But, you're right that there are multiple methods to address the problem.\n\nThe mention techniques like [Batch Renormalization](https://arxiv.org/pdf/1702.03275.pdf) and [Group Convolutions](https://arxiv.org/pdf/1904.00346.pdf) to address the problem of small batch sizes in the paper.",
          "votes": 1
        }
      ]
    },
    {
      "id": 915220,
      "postDate": "2020-07-04T15:02:55.777Z",
      "content": "<p>For me, layer norm damage dramatically and insance norm damage a little. Neither work.</p>",
      "rawMarkdown": "For me, layer norm damage dramatically and insance norm damage a little. Neither work.",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 603684,
      "author_name": "Danylo Kasianenko",
      "author_url": "",
      "post_date": "2019-08-20T14:52:10.883000",
      "content": "<p>Replacing can be done in PyTorch using the following function:</p>\n\n<p>````\ndef convert_layers(model, layer_type_old, layer_type_new, convert_weights=False, num_groups=None):\n    for name, module in reversed(model._modules.items()):\n        if len(list(module.children())) &gt; 0:\n            # recurse\n            model._modules[name] = convert_layers(module, layer_type_old, layer_type_new, convert_weights)</p>\n\n<pre><code>    if type(module) == layer_type_old:\n        layer_old = module\n        layer_new = layer_type_new(module.num_features if num_groups is None else num_groups, module.num_features, module.eps, module.affine) \n\n        if convert_weights:\n            layer_new.weight = layer_old.weight\n            layer_new.bias = layer_old.bias\n\n        model._modules[name] = layer_new\n\nreturn model\n</code></pre>\n\n<h1>Replace BatchNorm with GroupNorm</h1>\n\n<p>convert_layers(model, nn.BatchNorm2d, nn.GroupNorm, True, num_groups=2)\n````</p>\n\n<p>If num_groups is 1, GroupNorm turns into LayerNorm. If num_groups is None, GroupNorm turns into InstanceNorm</p>",
      "votes": 13,
      "replies": [
        {
          "id": 603700,
          "author_name": "Carlo",
          "author_url": "",
          "post_date": "2019-08-20T15:10:41.857000",
          "content": "<p>Nice! Thanks!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 904823,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-28T01:21:48.990000",
          "content": "<p>Thanks for the code!!! </p>\n\n<p>I found slight change, in <code>model._modules[name] = convert_layers</code>, the <code>num_groups</code> should pass in or the GroupNorm would be InstanceNorm. Can print the model to check the setting.</p>\n\n<p>Modified code is as follows:\n```</p>\n\n<h1>Replace BatchNorm with GroupNorm</h1>\n\n<p>def convert_layers(model, old_layer_type, new_layer_type, convert_weights=False, num_groups=None):</p>\n\n<pre><code>for name, module in reversed(model._modules.items()):\n\n    if len(list(module.children())) &gt; 0:\n        # recurse\n        model._modules[name] = convert_layers(module, old_layer_type, new_layer_type, convert_weights, num_groups=num_groups)\n\n    # single module\n    if type(module) == old_layer_type:\n        old_layer = module\n        new_layer = new_layer_type(module.num_features if num_groups is None else num_groups, module.num_features, module.eps, module.affine) \n\n        if convert_weights:\n            new_layer.weight = old_layer.weight\n            new_layer.bias = old_layer.bias\n\n        model._modules[name] = new_layer\n\nreturn model\n</code></pre>\n\n<p>```</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 991717,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-30T16:00:34.570000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 602212,
      "author_name": "Virilo Tejedor Aguilera",
      "author_url": "",
      "post_date": "2019-08-18T19:17:42.397000",
      "content": "<p>thanks a lot <a href=\"/carlolepelaars\">@carlolepelaars</a> !!!</p>\n\n<p>are you sure this code is working?  have you printed the summary of the CNN to check that these layers have been replaced as you liked?</p>\n\n<p>as far as I can see, you are creating a new GroupNormalization object and giving to \"layer\" a reference to it, but doing nothing with \"layer\"\nit shouldn't be enough to update the CNN </p>",
      "votes": 3,
      "replies": [
        {
          "id": 602231,
          "author_name": "Carlo",
          "author_url": "",
          "post_date": "2019-08-18T19:50:28.407000",
          "content": "<p>Oops, you're right! Thank you! I updated the post with working code.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 608144,
          "author_name": "Gabriel",
          "author_url": "",
          "post_date": "2019-08-26T12:01:57.773000",
          "content": "<p>I don't think that keras snippet physically replaces the layers. Plotting the model immediately after changing the layer name, shows the new layer but it doesn't look like its been built correctly. Plotting the model again after compiling and running a model summary it reverts back to the standard batch norm layer.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 616027,
          "author_name": "CinKKKyo",
          "author_url": "",
          "post_date": "2019-09-02T16:01:08.927000",
          "content": "<p>I made a brief contrast exp in Keras, I agree with u so far</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 603795,
      "author_name": "Sabbir Ahmed",
      "author_url": "",
      "post_date": "2019-08-20T17:06:44.590000",
      "content": "<p>did you get any LB boost by using Group Norm</p>",
      "votes": 1,
      "replies": [
        {
          "id": 616121,
          "author_name": "Carlo",
          "author_url": "",
          "post_date": "2019-09-02T17:57:32.383000",
          "content": "<p>Hi Sabbir, the main thing was that it allowed me to use small batch sizes (e.g. 4) without the training becoming unstable. Being able to use larger input sizes gave an LB boost. </p>\n\n<p>Hope this helps!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 602922,
      "author_name": "Nitesh Yadav",
      "author_url": "",
      "post_date": "2019-08-19T16:48:24.497000",
      "content": "<p>thank you <a href=\"/carlolepelaars\">@carlolepelaars</a> .\nIt helps me lot</p>",
      "votes": 1,
      "replies": [
        {
          "id": 602925,
          "author_name": "Carlo",
          "author_url": "",
          "post_date": "2019-08-19T16:52:34.677000",
          "content": "<p>Nice! You're welcome!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 602021,
      "author_name": "Ritisha Jaiswal",
      "author_url": "",
      "post_date": "2019-08-18T13:14:54.033000",
      "content": "<p><a href=\"/carlolepelaars\">@carlolepelaars</a> Insightful!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 602067,
          "author_name": "Carlo",
          "author_url": "",
          "post_date": "2019-08-18T14:19:48.610000",
          "content": "<p>👍 </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 607323,
      "author_name": "Vishnu Subramanian",
      "author_url": "",
      "post_date": "2019-08-25T03:21:13.627000",
      "content": "<p>I was experimenting GroupNorm on  a segmentation problem (SIIM). It performed worse than normal BatchNorm. It looks GroupNorm can also be sensitive to the number of groups, and it varies for datasets. It would be intersting to know if it helps here.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 606349,
      "author_name": "seefun",
      "author_url": "",
      "post_date": "2019-08-23T13:21:53.073000",
      "content": "<p>But, If you use GN, the pretrained weights with BN may be destroied. It can not take the advantages of transfer learning.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 606842,
          "author_name": "Danylo Kasianenko",
          "author_url": "",
          "post_date": "2019-08-24T07:02:19.780000",
          "content": "<p>If you use PyTorch, the weights can also be transfered. I'm not sure that it makes sense to transfer from different kind of layers, but it can be done.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 602025,
      "author_name": "Mihai-Alexandru Badila",
      "author_url": "",
      "post_date": "2019-08-18T13:20:06.037000",
      "content": "<p>Is it possible to use gradient accumulation to solve the problem? You can compute the gradients for 4 epochs using a batch size of 4 and only update the network once. This should be equivalent to training with a batch size of 16.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 602058,
          "author_name": "Carlo",
          "author_url": "",
          "post_date": "2019-08-18T14:03:56.213000",
          "content": "<p>Yes, good point! However, it seems like gradient accumulation doesn't prevent the calculation of inaccurate batch statistics in the batch normalization layers, since the normalization is still done over a batch size of 4. But, you're right that there are multiple methods to address the problem.</p>\n\n<p>The mention techniques like <a href=\"https://arxiv.org/pdf/1702.03275.pdf\">Batch Renormalization</a> and <a href=\"https://arxiv.org/pdf/1904.00346.pdf\">Group Convolutions</a> to address the problem of small batch sizes in the paper.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 915220,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-04T15:02:55.777000",
      "content": "<p>For me, layer norm damage dramatically and insance norm damage a little. Neither work.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "602012": "Hi Kagglers,\n\nDid any of you try replacing the batch normalization layers in your models with group normalization when using a small batch size? It seems to be a great alternative, especially when batch sizes get really small. \n\nDid you encounter problems using batch normalization with small batch sizes and/or are you having success using group normalization layers.\n\nThis image shows the problem with using batch normalization and small batch sizes:\n![GN_image](https://miro.medium.com/max/764/0*h7tx8LWRObeqqV53.)\n\n![GN_comparison](https://miro.medium.com/max/1400/0*P2BEbb-GFU1TCS0A.)\n\nI used the following loop to replace the layers in [this Kaggle kernel](https://www.kaggle.com/carlolepelaars/efficientnetb5-with-keras-aptos-2019):\n\n```\n# Load in EfficientNetB5\neffnet = EfficientNetB5(...)\n# Replace all Batch Normalization layers by Group Normalization layers\nfor i, layer in enumerate(effnet.layers):\n    if \"batch_normalization\" in layer.name:\n        effnet.layers[i] = GroupNormalization(groups=2, axis=-1, epsilon=0.1)\n```\n\nGroup Normalization paper:\nhttps://arxiv.org/pdf/1803.08494.pdf\n\nKeras implementation:\nhttps://github.com/titu1994/Keras-Group-Normalization/blob/master/group_norm.py\n\n",
    "603684": "Replacing can be done in PyTorch using the following function:\n\n````\ndef convert_layers(model, layer_type_old, layer_type_new, convert_weights=False, num_groups=None):\n    for name, module in reversed(model._modules.items()):\n        if len(list(module.children())) &gt; 0:\n            # recurse\n            model._modules[name] = convert_layers(module, layer_type_old, layer_type_new, convert_weights)\n\n        if type(module) == layer_type_old:\n            layer_old = module\n            layer_new = layer_type_new(module.num_features if num_groups is None else num_groups, module.num_features, module.eps, module.affine) \n\n            if convert_weights:\n                layer_new.weight = layer_old.weight\n                layer_new.bias = layer_old.bias\n\n            model._modules[name] = layer_new\n\n    return model\n\n\n# Replace BatchNorm with GroupNorm\nconvert_layers(model, nn.BatchNorm2d, nn.GroupNorm, True, num_groups=2)\n````\n\nIf num\\_groups is 1, GroupNorm turns into LayerNorm. If num\\_groups is None, GroupNorm turns into InstanceNorm",
    "602212": "thanks a lot @carlolepelaars !!!\n\nare you sure this code is working?  have you printed the summary of the CNN to check that these layers have been replaced as you liked?\n\nas far as I can see, you are creating a new GroupNormalization object and giving to \"layer\" a reference to it, but doing nothing with \"layer\"\nit shouldn't be enough to update the CNN ",
    "603795": "did you get any LB boost by using Group Norm",
    "602922": "thank you @carlolepelaars .\nIt helps me lot",
    "602021": "@carlolepelaars Insightful!",
    "607323": "I was experimenting GroupNorm on  a segmentation problem (SIIM). It performed worse than normal BatchNorm. It looks GroupNorm can also be sensitive to the number of groups, and it varies for datasets. It would be intersting to know if it helps here.",
    "606349": "But, If you use GN, the pretrained weights with BN may be destroied. It can not take the advantages of transfer learning.",
    "602025": "Is it possible to use gradient accumulation to solve the problem? You can compute the gradients for 4 epochs using a batch size of 4 and only update the network once. This should be equivalent to training with a batch size of 16.",
    "915220": "For me, layer norm damage dramatically and insance norm damage a little. Neither work."
  }
}