{
  "id": 224085,
  "title": "Hardware to train ResNet200D or similar large models?",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/224085",
  "author_name": "",
  "post_date": "2021-03-06T22:05:12.546915400Z",
  "votes": 6,
  "comment_count": 10,
  "views": 0,
  "content": "<p>What hardware do those that have successfully trained a decently performing ResNet200D with large image size (say &gt;= 512 by 512) have/what size of node in the cloud did you get? It seems like even the Kaggle TPUs struggle and certainly the GPU notebooks and my slightly dated GPU at home do. Thus, I was curious what people have used for this.</p>",
  "messages": [
    {
      "id": "1228902",
      "postDate": "03/06/2021 22:05:12",
      "content": "<p>What hardware do those that have successfully trained a decently performing ResNet200D with large image size (say &gt;= 512 by 512) have/what size of node in the cloud did you get? It seems like even the Kaggle TPUs struggle and certainly the GPU notebooks and my slightly dated GPU at home do. Thus, I was curious what people have used for this.</p>",
      "rawMarkdown": "What hardware do those that have successfully trained a decently performing ResNet200D with large image size (say >= 512 by 512) have/what size of node in the cloud did you get? It seems like even the Kaggle TPUs struggle and certainly the GPU notebooks and my slightly dated GPU at home do. Thus, I was curious what people have used for this.",
      "votes": null
    },
    {
      "id": "1228973",
      "postDate": "03/06/2021 23:45:47",
      "content": "<p>you need something like 3090, basically large VRAM to squeeze in a decent batchsize</p>",
      "rawMarkdown": "you need something like 3090, basically large VRAM to squeeze in a decent batchsize",
      "votes": null
    },
    {
      "id": "1229228",
      "postDate": "03/07/2021 07:10:11",
      "content": "<p>Training it on RTX 3080 works fine, 10GB VRAM seems to be enough.<br>\nI use Batchsize=4, mixed precision and image_size=640.</p>",
      "rawMarkdown": "Training it on RTX 3080 works fine, 10GB VRAM seems to be enough.\nI use Batchsize=4, mixed precision and image_size=640.",
      "votes": null
    },
    {
      "id": "1229265",
      "postDate": "03/07/2021 07:57:55",
      "content": "<p>Interesting. And that got you a good model performance? I was worried that such small batch sizes - even with gradient accumulation - would be problematic e.g. for the BatchNorm layers. Our was there some other clever tricks like freezing those layers/changing their momentum?</p>",
      "rawMarkdown": "Interesting. And that got you a good model performance? I was worried that such small batch sizes - even with gradient accumulation - would be problematic e.g. for the BatchNorm layers. Our was there some other clever tricks like freezing those layers/changing their momentum?",
      "votes": null
    },
    {
      "id": "1229355",
      "postDate": "03/07/2021 09:16:45",
      "content": "<p>Yeah, about perfomance (GroupKFold):</p>\n<p>Resnet200d<br>\n0.9679<br>\n0.9644<br>\n0.9665<br>\n0.9686<br>\n0.9691<br>\nAvg: 0.96724<br>\nOOF (TTA): 0.96857 </p>\n<p>Resnest200e (same bs and image_size)<br>\nFolds:<br>\n0.9656<br>\n0.9624<br>\n0.9655<br>\n0.9678<br>\n0.9669<br>\nAvg: 0.96564<br>\nOOF (TTA): 0.96653</p>\n<p>OOF ensemble (TTA): 0.96949 </p>\n<p>I think this can be considered good</p>\n<p>I also use gradient accumulation with frozen batchnorm layers. This is usually a good combination.</p>",
      "rawMarkdown": "Yeah, about perfomance (GroupKFold):\n\nResnet200d\n0.9679\n0.9644\n0.9665\n0.9686\n0.9691\nAvg: 0.96724\nOOF (TTA): 0.96857 \n\nResnest200e (same bs and image_size)\nFolds:\n0.9656\n0.9624\n0.9655\n0.9678\n0.9669\nAvg: 0.96564\nOOF (TTA): 0.96653\n\nOOF ensemble (TTA): 0.96949 \n\nI think this can be considered good\n\nI also use gradient accumulation with frozen batchnorm layers. This is usually a good combination.",
      "votes": null
    },
    {
      "id": "1229359",
      "postDate": "03/07/2021 09:30:30",
      "content": "<p>Freezing the BN layers is probably what I'd missed out on. It was one of the things I wondered about, but failed to try… I'll have to try that, thanks.</p>",
      "rawMarkdown": "Freezing the BN layers is probably what I'd missed out on. It was one of the things I wondered about, but failed to try... I'll have to try that, thanks.",
      "votes": null
    },
    {
      "id": "1230372",
      "postDate": "03/08/2021 04:40:43",
      "content": "<p>Hey thanks a lot for this info! Do you by any chance use PyTorch to train? If so, do you mind sharing the code for freezing batchnorm layers? <a href=\"https://www.kaggle.com/vadimtimakin\" target=\"_blank\">@vadimtimakin</a> </p>",
      "rawMarkdown": "Hey thanks a lot for this info! Do you by any chance use PyTorch to train? If so, do you mind sharing the code for freezing batchnorm layers? @vadimtimakin",
      "votes": null
    },
    {
      "id": "1230443",
      "postDate": "03/08/2021 06:44:02",
      "content": "<p>At the right place in the training loop (i.e. after setting the model to train), you want to call <code>.eval()</code> on the BatchNorm layers by looping through the children of the model and checking which ones are BNs. I'm using the <code>fastai</code> training loop, so I could just use <a href=\"https://github.com/fastai/fastai/blob/master/fastai/callback/training.py#L43\" target=\"_blank\">their callback (source is here)</a>.</p>\n<p>I now sort of also wonder whether using a really high momentum in the BN layers could also work as an alternative. I'm especially wondering, because one might use a model head with BN layers, which would then need some training if they have been randomly initiated.</p>",
      "rawMarkdown": "At the right place in the training loop (i.e. after setting the model to train), you want to call `.eval()` on the BatchNorm layers by looping through the children of the model and checking which ones are BNs. I'm using the `fastai` training loop, so I could just use [their callback (source is here)](https://github.com/fastai/fastai/blob/master/fastai/callback/training.py#L43).\n\nI now sort of also wonder whether using a really high momentum in the BN layers could also work as an alternative. I'm especially wondering, because one might use a model head with BN layers, which would then need some training if they have been randomly initiated.",
      "votes": null
    },
    {
      "id": "1230552",
      "postDate": "03/08/2021 08:28:41",
      "content": "<p>Yeah, I use PyTorch for my pipelines. Here is the code for freezing BatchNorm layers. You should add it after model.train() in the train loop (train step).</p>\n<pre><code>model.train()\n# Freezing BatchNorm layers\nfor name, child in (model.named_children()):\n    if name.find('BatchNorm') != -1:\n        for param in child.parameters():\n            param.requires_grad = False\n    else:\n        for param in child.parameters():\n            param.requires_grad = True\ntotalloss = 0.0\n</code></pre>",
      "rawMarkdown": "Yeah, I use PyTorch for my pipelines. Here is the code for freezing BatchNorm layers. You should add it after model.train() in the train loop (train step).\n```\nmodel.train()\n# Freezing BatchNorm layers\nfor name, child in (model.named_children()):\n    if name.find('BatchNorm') != -1:\n        for param in child.parameters():\n            param.requires_grad = False\n    else:\n        for param in child.parameters():\n            param.requires_grad = True\ntotalloss = 0.0\n```",
      "votes": null
    },
    {
      "id": "1230646",
      "postDate": "03/08/2021 10:14:35",
      "content": "<p>Make a note as well, that freezing BN with <code>param.requires_grad = False</code> will prevent on calculating gradients and updating the weights for those layers, but forward pass will still update moving average of statistics (mean, std) to your current dataset. Setting BN to <code>eval</code> will prevent both: computing gradients and moving average of statistics. Though, I don't know which way is better. Probably some \"trial and error\" could help to decide</p>",
      "rawMarkdown": "Make a note as well, that freezing BN with `param.requires_grad = False` will prevent on calculating gradients and updating the weights for those layers, but forward pass will still update moving average of statistics (mean, std) to your current dataset. Setting BN to `eval` will prevent both: computing gradients and moving average of statistics. Though, I don't know which way is better. Probably some \"trial and error\" could help to decide",
      "votes": null
    },
    {
      "id": "1230744",
      "postDate": "03/08/2021 12:18:05",
      "content": "<p>Awesome <a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a> <a href=\"https://www.kaggle.com/vadimtimakin\" target=\"_blank\">@vadimtimakin</a> appreciated!</p>",
      "rawMarkdown": "Awesome @bjoernholzhauer @vadimtimakin appreciated!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1228973,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/06/2021 23:45:47",
      "content": "<p>you need something like 3090, basically large VRAM to squeeze in a decent batchsize</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1229228,
      "author_name": "vadimtimakin",
      "author_url": "",
      "post_date": "03/07/2021 07:10:11",
      "content": "<p>Training it on RTX 3080 works fine, 10GB VRAM seems to be enough.<br>\nI use Batchsize=4, mixed precision and image_size=640.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1229265,
          "author_name": "bjoernholzhauer",
          "author_url": "",
          "post_date": "03/07/2021 07:57:55",
          "content": "<p>Interesting. And that got you a good model performance? I was worried that such small batch sizes - even with gradient accumulation - would be problematic e.g. for the BatchNorm layers. Our was there some other clever tricks like freezing those layers/changing their momentum?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1229355,
          "author_name": "vadimtimakin",
          "author_url": "",
          "post_date": "03/07/2021 09:16:45",
          "content": "<p>Yeah, about perfomance (GroupKFold):</p>\n<p>Resnet200d<br>\n0.9679<br>\n0.9644<br>\n0.9665<br>\n0.9686<br>\n0.9691<br>\nAvg: 0.96724<br>\nOOF (TTA): 0.96857 </p>\n<p>Resnest200e (same bs and image_size)<br>\nFolds:<br>\n0.9656<br>\n0.9624<br>\n0.9655<br>\n0.9678<br>\n0.9669<br>\nAvg: 0.96564<br>\nOOF (TTA): 0.96653</p>\n<p>OOF ensemble (TTA): 0.96949 </p>\n<p>I think this can be considered good</p>\n<p>I also use gradient accumulation with frozen batchnorm layers. This is usually a good combination.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1229359,
          "author_name": "bjoernholzhauer",
          "author_url": "",
          "post_date": "03/07/2021 09:30:30",
          "content": "<p>Freezing the BN layers is probably what I'd missed out on. It was one of the things I wondered about, but failed to try… I'll have to try that, thanks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1230372,
          "author_name": "reighns",
          "author_url": "",
          "post_date": "03/08/2021 04:40:43",
          "content": "<p>Hey thanks a lot for this info! Do you by any chance use PyTorch to train? If so, do you mind sharing the code for freezing batchnorm layers? <a href=\"https://www.kaggle.com/vadimtimakin\" target=\"_blank\">@vadimtimakin</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1230443,
          "author_name": "bjoernholzhauer",
          "author_url": "",
          "post_date": "03/08/2021 06:44:02",
          "content": "<p>At the right place in the training loop (i.e. after setting the model to train), you want to call <code>.eval()</code> on the BatchNorm layers by looping through the children of the model and checking which ones are BNs. I'm using the <code>fastai</code> training loop, so I could just use <a href=\"https://github.com/fastai/fastai/blob/master/fastai/callback/training.py#L43\" target=\"_blank\">their callback (source is here)</a>.</p>\n<p>I now sort of also wonder whether using a really high momentum in the BN layers could also work as an alternative. I'm especially wondering, because one might use a model head with BN layers, which would then need some training if they have been randomly initiated.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1230552,
          "author_name": "vadimtimakin",
          "author_url": "",
          "post_date": "03/08/2021 08:28:41",
          "content": "<p>Yeah, I use PyTorch for my pipelines. Here is the code for freezing BatchNorm layers. You should add it after model.train() in the train loop (train step).</p>\n<pre><code>model.train()\n# Freezing BatchNorm layers\nfor name, child in (model.named_children()):\n    if name.find('BatchNorm') != -1:\n        for param in child.parameters():\n            param.requires_grad = False\n    else:\n        for param in child.parameters():\n            param.requires_grad = True\ntotalloss = 0.0\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1230646,
          "author_name": "ademyanchuk",
          "author_url": "",
          "post_date": "03/08/2021 10:14:35",
          "content": "<p>Make a note as well, that freezing BN with <code>param.requires_grad = False</code> will prevent on calculating gradients and updating the weights for those layers, but forward pass will still update moving average of statistics (mean, std) to your current dataset. Setting BN to <code>eval</code> will prevent both: computing gradients and moving average of statistics. Though, I don't know which way is better. Probably some \"trial and error\" could help to decide</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1230744,
          "author_name": "reighns",
          "author_url": "",
          "post_date": "03/08/2021 12:18:05",
          "content": "<p>Awesome <a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a> <a href=\"https://www.kaggle.com/vadimtimakin\" target=\"_blank\">@vadimtimakin</a> appreciated!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1228902": "What hardware do those that have successfully trained a decently performing ResNet200D with large image size (say >= 512 by 512) have/what size of node in the cloud did you get? It seems like even the Kaggle TPUs struggle and certainly the GPU notebooks and my slightly dated GPU at home do. Thus, I was curious what people have used for this.",
    "1228973": "you need something like 3090, basically large VRAM to squeeze in a decent batchsize",
    "1229228": "Training it on RTX 3080 works fine, 10GB VRAM seems to be enough.\nI use Batchsize=4, mixed precision and image_size=640.",
    "1229265": "Interesting. And that got you a good model performance? I was worried that such small batch sizes - even with gradient accumulation - would be problematic e.g. for the BatchNorm layers. Our was there some other clever tricks like freezing those layers/changing their momentum?",
    "1229355": "Yeah, about perfomance (GroupKFold):\n\nResnet200d\n0.9679\n0.9644\n0.9665\n0.9686\n0.9691\nAvg: 0.96724\nOOF (TTA): 0.96857 \n\nResnest200e (same bs and image_size)\nFolds:\n0.9656\n0.9624\n0.9655\n0.9678\n0.9669\nAvg: 0.96564\nOOF (TTA): 0.96653\n\nOOF ensemble (TTA): 0.96949 \n\nI think this can be considered good\n\nI also use gradient accumulation with frozen batchnorm layers. This is usually a good combination.",
    "1229359": "Freezing the BN layers is probably what I'd missed out on. It was one of the things I wondered about, but failed to try... I'll have to try that, thanks.",
    "1230372": "Hey thanks a lot for this info! Do you by any chance use PyTorch to train? If so, do you mind sharing the code for freezing batchnorm layers? @vadimtimakin",
    "1230443": "At the right place in the training loop (i.e. after setting the model to train), you want to call `.eval()` on the BatchNorm layers by looping through the children of the model and checking which ones are BNs. I'm using the `fastai` training loop, so I could just use [their callback (source is here)](https://github.com/fastai/fastai/blob/master/fastai/callback/training.py#L43).\n\nI now sort of also wonder whether using a really high momentum in the BN layers could also work as an alternative. I'm especially wondering, because one might use a model head with BN layers, which would then need some training if they have been randomly initiated.",
    "1230552": "Yeah, I use PyTorch for my pipelines. Here is the code for freezing BatchNorm layers. You should add it after model.train() in the train loop (train step).\n```\nmodel.train()\n# Freezing BatchNorm layers\nfor name, child in (model.named_children()):\n    if name.find('BatchNorm') != -1:\n        for param in child.parameters():\n            param.requires_grad = False\n    else:\n        for param in child.parameters():\n            param.requires_grad = True\ntotalloss = 0.0\n```",
    "1230646": "Make a note as well, that freezing BN with `param.requires_grad = False` will prevent on calculating gradients and updating the weights for those layers, but forward pass will still update moving average of statistics (mean, std) to your current dataset. Setting BN to `eval` will prevent both: computing gradients and moving average of statistics. Though, I don't know which way is better. Probably some \"trial and error\" could help to decide",
    "1230744": "Awesome @bjoernholzhauer @vadimtimakin appreciated!"
  },
  "source": "meta"
}