{
  "id": 226911,
  "title": "Did you use gradient accumulation to train your models?",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/226911",
  "author_name": "",
  "post_date": "2021-03-18T07:30:49.219866800Z",
  "votes": 4,
  "comment_count": 7,
  "views": 0,
  "content": "<p>First of all, congratulations to the winners, all the teams in the gold zone, and to everyone who took part and learnt something!</p>\n<p>Did you use gradient accumulation to train your models?</p>\n<p>How did it work?  How did you deal with BN layers (when present)?</p>\n<p>Which batch size and how many accumulation steps did you use?</p>\n<p>I didn't have time to try it, and it could have been important, given the high resolution input images</p>",
  "messages": [
    {
      "id": "1243400",
      "postDate": "03/18/2021 07:30:49",
      "content": "<p>First of all, congratulations to the winners, all the teams in the gold zone, and to everyone who took part and learnt something!</p>\n<p>Did you use gradient accumulation to train your models?</p>\n<p>How did it work?  How did you deal with BN layers (when present)?</p>\n<p>Which batch size and how many accumulation steps did you use?</p>\n<p>I didn't have time to try it, and it could have been important, given the high resolution input images</p>",
      "rawMarkdown": "First of all, congratulations to the winners, all the teams in the gold zone, and to everyone who took part and learnt something!\n\nDid you use gradient accumulation to train your models?\n\nHow did it work?  How did you deal with BN layers (when present)?\n\nWhich batch size and how many accumulation steps did you use?\n\nI didn't have time to try it, and it could have been important, given the high resolution input images",
      "votes": null
    },
    {
      "id": "1243454",
      "postDate": "03/18/2021 08:17:10",
      "content": "<p>Take this with a grain of salt since others managed to train these models better, but here's what I did:</p>\n<ul>\n<li>I <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/226684\" target=\"_blank\">used gradient accumulation</a> to train ResNet-200D with 640 by 640 images and SE-ResNet-152D with 704 by 704 on a GTX 1080 TI.</li>\n<li>My batch size was 4 and I accumulated over 8 batches (effective batch size of 32 - I did not experiment around that, but just went by the <a href=\"https://twitter.com/ylecun/status/989610208497360896?lang=en\" target=\"_blank\">\"Friends don't let friends use batch sizes over 32\" quote</a>). </li>\n<li>I froze all BatchNorm layers (i.e. set <code>.eval()</code> on the layer) after this was mentioned <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/224085\" target=\"_blank\">here</a> and it immediately improved my results.</li>\n<li>I had done some experimentation with smaller models, where more complex model heads worked pretty well (e.g. when I did not need to freeze BN layers, I used the head suggested by Jeremy Howard and Sylvain Gugger in <a href=\"https://www.amazon.com/Deep-Learning-Coders-fastai-PyTorch/dp/1492045527\" target=\"_blank\">their book</a> - see also the <a href=\"https://github.com/fastai/fastbook/blob/master/15_arch_details.ipynb\" target=\"_blank\">repository for the book</a>: <br>\nAdaptiveConcatPool2d - Flatten - BatchNorm1d - Dropout - FC - ReLU - BatchNorm1d - Dropout - FC to output). I just could not get that to work with small batch sizes. That was presumably because of the BN layers, so I tried to experiment with high momentum etc., but I did not find a way to get it to work well. In the end, I did just go straight to an output layer after pooling &amp; flattening.</li>\n</ul>\n<p>I'd love to hear whether other approaches worked for people/there's better ways of doing this. What you can do when working around hardware limitations is surely interesting to a lot of people.</p>",
      "rawMarkdown": "Take this with a grain of salt since others managed to train these models better, but here's what I did:\n* I [used gradient accumulation](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/226684) to train ResNet-200D with 640 by 640 images and SE-ResNet-152D with 704 by 704 on a GTX 1080 TI.\n* My batch size was 4 and I accumulated over 8 batches (effective batch size of 32 - I did not experiment around that, but just went by the [\"Friends don't let friends use batch sizes over 32\" quote](https://twitter.com/ylecun/status/989610208497360896?lang=en)). \n* I froze all BatchNorm layers (i.e. set `.eval()` on the layer) after this was mentioned [here](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/224085) and it immediately improved my results.\n* I had done some experimentation with smaller models, where more complex model heads worked pretty well (e.g. when I did not need to freeze BN layers, I used the head suggested by Jeremy Howard and Sylvain Gugger in [their book](https://www.amazon.com/Deep-Learning-Coders-fastai-PyTorch/dp/1492045527) - see also the [repository for the book](https://github.com/fastai/fastbook/blob/master/15_arch_details.ipynb): \nAdaptiveConcatPool2d - Flatten - BatchNorm1d - Dropout - FC - ReLU - BatchNorm1d - Dropout - FC to output). I just could not get that to work with small batch sizes. That was presumably because of the BN layers, so I tried to experiment with high momentum etc., but I did not find a way to get it to work well. In the end, I did just go straight to an output layer after pooling & flattening.\n\nI'd love to hear whether other approaches worked for people/there's better ways of doing this. What you can do when working around hardware limitations is surely interesting to a lot of people.",
      "votes": null
    },
    {
      "id": "1243530",
      "postDate": "03/18/2021 09:42:17",
      "content": "<p>I tried multiple times in the past and during this comp. It didn’t really work. It’s pretty hard to say if it works when the batchnorm presents.</p>",
      "rawMarkdown": "I tried multiple times in the past and during this comp. It didn’t really work. It’s pretty hard to say if it works when the batchnorm presents.",
      "votes": null
    },
    {
      "id": "1244888",
      "postDate": "03/19/2021 09:29:43",
      "content": "<p>I used gradient accumulation to train large model like Resnet200d. I have only 1080Ti GPU with 11GB RAM.</p>",
      "rawMarkdown": "I used gradient accumulation to train large model like Resnet200d. I have only 1080Ti GPU with 11GB RAM.",
      "votes": null
    },
    {
      "id": "1245941",
      "postDate": "03/20/2021 11:18:21",
      "content": "<p><a href=\"https://www.kaggle.com/yoshitaka1105\" target=\"_blank\">@yoshitaka1105</a> which batch size and accumulation steps did you use?  have you tested if it actually improved the results? did you replaced the BN layers by something, or freeze these layers or use a high momentum?  thanks!</p>",
      "rawMarkdown": "yoshitaka1105 which batch size and accumulation steps did you use?  have you tested if it actually improved the results? did you replaced the BN layers by something, or freeze these layers or use a high momentum?  thanks!",
      "votes": null
    },
    {
      "id": "1245958",
      "postDate": "03/20/2021 11:44:34",
      "content": "<p>thanks <a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a> <a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> <a href=\"https://www.kaggle.com/yoshitaka1105\" target=\"_blank\">@yoshitaka1105</a> !</p>\n<p><a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a> did you measured if it improved your results respect training without gradient accumulation and using the same simpler head?</p>\n<p>did you freeze BN layers to train models without gradient accumulation?  I suspect that it could be the reason of the improvement, instead of the gradient accumulation</p>",
      "rawMarkdown": "thanks @bjoernholzhauer @underwearfitting @yoshitaka1105 !\n\n@bjoernholzhauer did you measured if it improved your results respect training without gradient accumulation and using the same simpler head?\n\ndid you freeze BN layers to train models without gradient accumulation?  I suspect that it could be the reason of the improvement, instead of the gradient accumulation",
      "votes": null
    },
    {
      "id": "1246009",
      "postDate": "03/20/2021 12:26:48",
      "content": "<p>I didn't test without gradient accumulation, I suppose that might work fine with frozen BN layers (as long as one also reduces the learning rate accordingly). However, we know that with unfrozen BN layers largish batch sizes are good (for a start the models were originally trained like that), as long as your GPU(s) can handle it.</p>",
      "rawMarkdown": "I didn't test without gradient accumulation, I suppose that might work fine with frozen BN layers (as long as one also reduces the learning rate accordingly). However, we know that with unfrozen BN layers largish batch sizes are good (for a start the models were originally trained like that), as long as your GPU(s) can handle it.",
      "votes": null
    },
    {
      "id": "1246192",
      "postDate": "03/20/2021 14:23:31",
      "content": "<p>In my case, Batch size: 4 and Accumulation: 8 batches. I did not try other parameters, so I do not know how good my parameter is. And I froze all BN layers.</p>",
      "rawMarkdown": "In my case, Batch size: 4 and Accumulation: 8 batches. I did not try other parameters, so I do not know how good my parameter is. And I froze all BN layers.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1243454,
      "author_name": "bjoernholzhauer",
      "author_url": "",
      "post_date": "03/18/2021 08:17:10",
      "content": "<p>Take this with a grain of salt since others managed to train these models better, but here's what I did:</p>\n<ul>\n<li>I <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/226684\" target=\"_blank\">used gradient accumulation</a> to train ResNet-200D with 640 by 640 images and SE-ResNet-152D with 704 by 704 on a GTX 1080 TI.</li>\n<li>My batch size was 4 and I accumulated over 8 batches (effective batch size of 32 - I did not experiment around that, but just went by the <a href=\"https://twitter.com/ylecun/status/989610208497360896?lang=en\" target=\"_blank\">\"Friends don't let friends use batch sizes over 32\" quote</a>). </li>\n<li>I froze all BatchNorm layers (i.e. set <code>.eval()</code> on the layer) after this was mentioned <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/224085\" target=\"_blank\">here</a> and it immediately improved my results.</li>\n<li>I had done some experimentation with smaller models, where more complex model heads worked pretty well (e.g. when I did not need to freeze BN layers, I used the head suggested by Jeremy Howard and Sylvain Gugger in <a href=\"https://www.amazon.com/Deep-Learning-Coders-fastai-PyTorch/dp/1492045527\" target=\"_blank\">their book</a> - see also the <a href=\"https://github.com/fastai/fastbook/blob/master/15_arch_details.ipynb\" target=\"_blank\">repository for the book</a>: <br>\nAdaptiveConcatPool2d - Flatten - BatchNorm1d - Dropout - FC - ReLU - BatchNorm1d - Dropout - FC to output). I just could not get that to work with small batch sizes. That was presumably because of the BN layers, so I tried to experiment with high momentum etc., but I did not find a way to get it to work well. In the end, I did just go straight to an output layer after pooling &amp; flattening.</li>\n</ul>\n<p>I'd love to hear whether other approaches worked for people/there's better ways of doing this. What you can do when working around hardware limitations is surely interesting to a lot of people.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1243530,
      "author_name": "underwearfitting",
      "author_url": "",
      "post_date": "03/18/2021 09:42:17",
      "content": "<p>I tried multiple times in the past and during this comp. It didn’t really work. It’s pretty hard to say if it works when the batchnorm presents.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1244888,
      "author_name": "yoshitaka1105",
      "author_url": "",
      "post_date": "03/19/2021 09:29:43",
      "content": "<p>I used gradient accumulation to train large model like Resnet200d. I have only 1080Ti GPU with 11GB RAM.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1245941,
          "author_name": "virilo",
          "author_url": "",
          "post_date": "03/20/2021 11:18:21",
          "content": "<p><a href=\"https://www.kaggle.com/yoshitaka1105\" target=\"_blank\">@yoshitaka1105</a> which batch size and accumulation steps did you use?  have you tested if it actually improved the results? did you replaced the BN layers by something, or freeze these layers or use a high momentum?  thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1246192,
          "author_name": "yoshitaka1105",
          "author_url": "",
          "post_date": "03/20/2021 14:23:31",
          "content": "<p>In my case, Batch size: 4 and Accumulation: 8 batches. I did not try other parameters, so I do not know how good my parameter is. And I froze all BN layers.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1245958,
      "author_name": "virilo",
      "author_url": "",
      "post_date": "03/20/2021 11:44:34",
      "content": "<p>thanks <a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a> <a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> <a href=\"https://www.kaggle.com/yoshitaka1105\" target=\"_blank\">@yoshitaka1105</a> !</p>\n<p><a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a> did you measured if it improved your results respect training without gradient accumulation and using the same simpler head?</p>\n<p>did you freeze BN layers to train models without gradient accumulation?  I suspect that it could be the reason of the improvement, instead of the gradient accumulation</p>",
      "votes": null,
      "replies": [
        {
          "id": 1246009,
          "author_name": "bjoernholzhauer",
          "author_url": "",
          "post_date": "03/20/2021 12:26:48",
          "content": "<p>I didn't test without gradient accumulation, I suppose that might work fine with frozen BN layers (as long as one also reduces the learning rate accordingly). However, we know that with unfrozen BN layers largish batch sizes are good (for a start the models were originally trained like that), as long as your GPU(s) can handle it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1243400": "First of all, congratulations to the winners, all the teams in the gold zone, and to everyone who took part and learnt something!\n\nDid you use gradient accumulation to train your models?\n\nHow did it work?  How did you deal with BN layers (when present)?\n\nWhich batch size and how many accumulation steps did you use?\n\nI didn't have time to try it, and it could have been important, given the high resolution input images",
    "1243454": "Take this with a grain of salt since others managed to train these models better, but here's what I did:\n* I [used gradient accumulation](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/226684) to train ResNet-200D with 640 by 640 images and SE-ResNet-152D with 704 by 704 on a GTX 1080 TI.\n* My batch size was 4 and I accumulated over 8 batches (effective batch size of 32 - I did not experiment around that, but just went by the [\"Friends don't let friends use batch sizes over 32\" quote](https://twitter.com/ylecun/status/989610208497360896?lang=en)). \n* I froze all BatchNorm layers (i.e. set `.eval()` on the layer) after this was mentioned [here](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/224085) and it immediately improved my results.\n* I had done some experimentation with smaller models, where more complex model heads worked pretty well (e.g. when I did not need to freeze BN layers, I used the head suggested by Jeremy Howard and Sylvain Gugger in [their book](https://www.amazon.com/Deep-Learning-Coders-fastai-PyTorch/dp/1492045527) - see also the [repository for the book](https://github.com/fastai/fastbook/blob/master/15_arch_details.ipynb): \nAdaptiveConcatPool2d - Flatten - BatchNorm1d - Dropout - FC - ReLU - BatchNorm1d - Dropout - FC to output). I just could not get that to work with small batch sizes. That was presumably because of the BN layers, so I tried to experiment with high momentum etc., but I did not find a way to get it to work well. In the end, I did just go straight to an output layer after pooling & flattening.\n\nI'd love to hear whether other approaches worked for people/there's better ways of doing this. What you can do when working around hardware limitations is surely interesting to a lot of people.",
    "1243530": "I tried multiple times in the past and during this comp. It didn’t really work. It’s pretty hard to say if it works when the batchnorm presents.",
    "1244888": "I used gradient accumulation to train large model like Resnet200d. I have only 1080Ti GPU with 11GB RAM.",
    "1245941": "yoshitaka1105 which batch size and accumulation steps did you use?  have you tested if it actually improved the results? did you replaced the BN layers by something, or freeze these layers or use a high momentum?  thanks!",
    "1245958": "thanks @bjoernholzhauer @underwearfitting @yoshitaka1105 !\n\n@bjoernholzhauer did you measured if it improved your results respect training without gradient accumulation and using the same simpler head?\n\ndid you freeze BN layers to train models without gradient accumulation?  I suspect that it could be the reason of the improvement, instead of the gradient accumulation",
    "1246009": "I didn't test without gradient accumulation, I suppose that might work fine with frozen BN layers (as long as one also reduces the learning rate accordingly). However, we know that with unfrozen BN layers largish batch sizes are good (for a start the models were originally trained like that), as long as your GPU(s) can handle it.",
    "1246192": "In my case, Batch size: 4 and Accumulation: 8 batches. I did not try other parameters, so I do not know how good my parameter is. And I froze all BN layers."
  },
  "source": "meta"
}