{
  "id": 151769,
  "title": "ways to learn the NN that does not fit in GPU memory with reasonable minibatch.",
  "url": "/competitions/alaska2-image-steganalysis/discussion/151769",
  "author_name": "",
  "post_date": "2020-05-17T01:04:15.649286800Z",
  "votes": null,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I want to experiment with eb7 architecture and weights recently released. But unfortunately i can train it with at most minibatch of size 1-2 at my GPU and it is not big enough to update batchnorm layers (especially if we are talking about multiclass problem). How can i deal with it? Can you suggest an approach to learn such model? I have heard about accumulating the loss with several minibatches, but have not tried it yet. Is there any other suggestions?</p>",
  "messages": [
    {
      "id": "850739",
      "postDate": "05/17/2020 01:04:15",
      "content": "<p>I want to experiment with eb7 architecture and weights recently released. But unfortunately i can train it with at most minibatch of size 1-2 at my GPU and it is not big enough to update batchnorm layers (especially if we are talking about multiclass problem). How can i deal with it? Can you suggest an approach to learn such model? I have heard about accumulating the loss with several minibatches, but have not tried it yet. Is there any other suggestions?</p>",
      "rawMarkdown": "I want to experiment with eb7 architecture and weights recently released. But unfortunately i can train it with at most minibatch of size 1-2 at my GPU and it is not big enough to update batchnorm layers (especially if we are talking about multiclass problem). How can i deal with it? Can you suggest an approach to learn such model? I have heard about accumulating the loss with several minibatches, but have not tried it yet. Is there any other suggestions?",
      "votes": null
    },
    {
      "id": "850948",
      "postDate": "05/17/2020 07:37:49",
      "content": "<p><a href=\"/vovanf98\">@vovanf98</a> you need groupnormalization,batchnormalization doesn't do well for batch size less than 16,in such case groupnormalization works best.\nkeras implementation : </p>\n\n<p>```</p>\n\n<h1>Load in EfficientNetB7</h1>\n\n<p>effnet = EfficientNetB7(...)</p>\n\n<h1>Replace all Batch Normalization layers by Group Normalization layers</h1>\n\n<p>for i, layer in enumerate(effnet.layers):\n    if \"batch_normalization\" in layer.name:\n        effnet.layers[i] = GroupNormalization(groups=2, axis=-1, epsilon=0.1)\n```</p>\n\n<p>please check this image : \n</p>",
      "rawMarkdown": "vovanf98 you need groupnormalization,batchnormalization doesn't do well for batch size less than 16,in such case groupnormalization works best.\nkeras implementation : \n\n```\n# Load in EfficientNetB7\neffnet = EfficientNetB7(...)\n# Replace all Batch Normalization layers by Group Normalization layers\nfor i, layer in enumerate(effnet.layers):\n    if \"batch_normalization\" in layer.name:\n        effnet.layers[i] = GroupNormalization(groups=2, axis=-1, epsilon=0.1)\n```\n\nplease check this image : \n![](https://miro.medium.com/max/764/0*h7tx8LWRObeqqV53.)",
      "votes": null
    },
    {
      "id": "851045",
      "postDate": "05/17/2020 09:46:29",
      "content": "<p>Wow, thanks a lot! And still it is a good idea to use loss accumulation? </p>",
      "rawMarkdown": "Wow, thanks a lot! And still it is a good idea to use loss accumulation?",
      "votes": null
    },
    {
      "id": "851046",
      "postDate": "05/17/2020 09:48:31",
      "content": "<p>And where did you get this Image? Maybe there is a link to a useful arxiv material?) Thanks in advance! </p>",
      "rawMarkdown": "And where did you get this Image? Maybe there is a link to a useful arxiv material?) Thanks in advance!",
      "votes": null
    },
    {
      "id": "851056",
      "postDate": "05/17/2020 09:57:55",
      "content": "<p><a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/104686\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/104686</a> - found nice explaination</p>",
      "rawMarkdown": "https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/104686 - found nice explaination",
      "votes": null
    },
    {
      "id": "851062",
      "postDate": "05/17/2020 10:03:43",
      "content": "<p>yes <a href=\"/vovanf98\">@vovanf98</a>  gradient accumulation is good in this case but i have never got improvement with that in the  past,maybe i couldn't use that properly so you can try that and see what happens,,yeah i learnt that group normalization thing from carlo and there is another  thing that you might will like to try and that is Filter Response Normalization,,check it out here :<a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/120379\"> Filter Response Normalization</a></p>",
      "rawMarkdown": "yes @vovanf98  gradient accumulation is good in this case but i have never got improvement with that in the  past,maybe i couldn't use that properly so you can try that and see what happens,,yeah i learnt that group normalization thing from carlo and there is another  thing that you might will like to try and that is Filter Response Normalization,,check it out here :[ Filter Response Normalization](https://www.kaggle.com/c/pku-autonomous-driving/discussion/120379)",
      "votes": null
    },
    {
      "id": "851075",
      "postDate": "05/17/2020 10:17:14",
      "content": "<p>Thank you so much!</p>",
      "rawMarkdown": "Thank you so much!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 850948,
      "author_name": "mobassir",
      "author_url": "",
      "post_date": "05/17/2020 07:37:49",
      "content": "<p><a href=\"/vovanf98\">@vovanf98</a> you need groupnormalization,batchnormalization doesn't do well for batch size less than 16,in such case groupnormalization works best.\nkeras implementation : </p>\n\n<p>```</p>\n\n<h1>Load in EfficientNetB7</h1>\n\n<p>effnet = EfficientNetB7(...)</p>\n\n<h1>Replace all Batch Normalization layers by Group Normalization layers</h1>\n\n<p>for i, layer in enumerate(effnet.layers):\n    if \"batch_normalization\" in layer.name:\n        effnet.layers[i] = GroupNormalization(groups=2, axis=-1, epsilon=0.1)\n```</p>\n\n<p>please check this image : \n</p>",
      "votes": null,
      "replies": [
        {
          "id": 851045,
          "author_name": "vovanf98",
          "author_url": "",
          "post_date": "05/17/2020 09:46:29",
          "content": "<p>Wow, thanks a lot! And still it is a good idea to use loss accumulation? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 851046,
          "author_name": "vovanf98",
          "author_url": "",
          "post_date": "05/17/2020 09:48:31",
          "content": "<p>And where did you get this Image? Maybe there is a link to a useful arxiv material?) Thanks in advance! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 851056,
          "author_name": "vovanf98",
          "author_url": "",
          "post_date": "05/17/2020 09:57:55",
          "content": "<p><a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/104686\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/104686</a> - found nice explaination</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 851062,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "05/17/2020 10:03:43",
          "content": "<p>yes <a href=\"/vovanf98\">@vovanf98</a>  gradient accumulation is good in this case but i have never got improvement with that in the  past,maybe i couldn't use that properly so you can try that and see what happens,,yeah i learnt that group normalization thing from carlo and there is another  thing that you might will like to try and that is Filter Response Normalization,,check it out here :<a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/120379\"> Filter Response Normalization</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 851075,
          "author_name": "vovanf98",
          "author_url": "",
          "post_date": "05/17/2020 10:17:14",
          "content": "<p>Thank you so much!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "850739": "I want to experiment with eb7 architecture and weights recently released. But unfortunately i can train it with at most minibatch of size 1-2 at my GPU and it is not big enough to update batchnorm layers (especially if we are talking about multiclass problem). How can i deal with it? Can you suggest an approach to learn such model? I have heard about accumulating the loss with several minibatches, but have not tried it yet. Is there any other suggestions?",
    "850948": "vovanf98 you need groupnormalization,batchnormalization doesn't do well for batch size less than 16,in such case groupnormalization works best.\nkeras implementation : \n\n```\n# Load in EfficientNetB7\neffnet = EfficientNetB7(...)\n# Replace all Batch Normalization layers by Group Normalization layers\nfor i, layer in enumerate(effnet.layers):\n    if \"batch_normalization\" in layer.name:\n        effnet.layers[i] = GroupNormalization(groups=2, axis=-1, epsilon=0.1)\n```\n\nplease check this image : \n![](https://miro.medium.com/max/764/0*h7tx8LWRObeqqV53.)",
    "851045": "Wow, thanks a lot! And still it is a good idea to use loss accumulation?",
    "851046": "And where did you get this Image? Maybe there is a link to a useful arxiv material?) Thanks in advance!",
    "851056": "https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/104686 - found nice explaination",
    "851062": "yes @vovanf98  gradient accumulation is good in this case but i have never got improvement with that in the  past,maybe i couldn't use that properly so you can try that and see what happens,,yeah i learnt that group normalization thing from carlo and there is another  thing that you might will like to try and that is Filter Response Normalization,,check it out here :[ Filter Response Normalization](https://www.kaggle.com/c/pku-autonomous-driving/discussion/120379)",
    "851075": "Thank you so much!"
  },
  "source": "meta"
}