{
  "id": 212991,
  "title": "Need some help about out of memory",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/212991",
  "author_name": "",
  "post_date": "2021-01-21T03:16:25.228733900Z",
  "votes": null,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I use PyTorch and I read some notebooks. I see most people use 512x512 import images and 32 batch size. However, when I try to use 512x512，it always says the GPU is out of memory.<br>\nI also tried to use apex to reduce memory, but it still uses more memory on GPU than available and causes a lot of overflows.<br>\nIs there anything important I missed? I was confused by this and really need some advice.<br>\nThank you very much.</p>",
  "messages": [
    {
      "id": "1162186",
      "postDate": "01/21/2021 03:16:25",
      "content": "<p>I use PyTorch and I read some notebooks. I see most people use 512x512 import images and 32 batch size. However, when I try to use 512x512，it always says the GPU is out of memory.<br>\nI also tried to use apex to reduce memory, but it still uses more memory on GPU than available and causes a lot of overflows.<br>\nIs there anything important I missed? I was confused by this and really need some advice.<br>\nThank you very much.</p>",
      "rawMarkdown": "I use PyTorch and I read some notebooks. I see most people use 512x512 import images and 32 batch size. However, when I try to use 512x512，it always says the GPU is out of memory.\nI also tried to use apex to reduce memory, but it still uses more memory on GPU than available and causes a lot of overflows.\nIs there anything important I missed? I was confused by this and really need some advice.\nThank you very much.",
      "votes": null
    },
    {
      "id": "1162327",
      "postDate": "01/21/2021 05:23:00",
      "content": "<p>It depends upon your model<br>\nIf you are using Effnet b4 or Effnet b5 batch size of 32 isn't possible, the maximum you can reach is around 24 with Effnet b4<br>\nI think the notebooks that you are referring to are using small models like b0 or b1</p>",
      "rawMarkdown": "It depends upon your model\nIf you are using Effnet b4 or Effnet b5 batch size of 32 isn't possible, the maximum you can reach is around 24 with Effnet b4\nI think the notebooks that you are referring to are using small models like b0 or b1",
      "votes": null
    },
    {
      "id": "1163768",
      "postDate": "01/21/2021 23:00:40",
      "content": "<p>I was gonna say the same thing, for instance VIT max batch size is around ~4. The bigger the model, the more RAM it needs, so if you choose for example 8 batch size, you'll have more RAM alocated as you go from EffNetB0 -&gt; B1 -&gt; B2 and so on, so either you choose a simpler, less intensive model or just a smaller batch size, which usually you can compensate more than well enough by adjusting the learning rate proportionally.</p>",
      "rawMarkdown": "I was gonna say the same thing, for instance VIT max batch size is around ~4. The bigger the model, the more RAM it needs, so if you choose for example 8 batch size, you'll have more RAM alocated as you go from EffNetB0 -> B1 -> B2 and so on, so either you choose a simpler, less intensive model or just a smaller batch size, which usually you can compensate more than well enough by adjusting the learning rate proportionally.",
      "votes": null
    },
    {
      "id": "1163773",
      "postDate": "01/21/2021 23:16:50",
      "content": "<p>You can also use gradient accumulation.<br>\nExample: <a href=\"https://gist.github.com/thomwolf/ac7a7da6b1888c2eeac8ac8b9b05d3d3\" target=\"_blank\">https://gist.github.com/thomwolf/ac7a7da6b1888c2eeac8ac8b9b05d3d3</a></p>",
      "rawMarkdown": "You can also use gradient accumulation.\nExample: https://gist.github.com/thomwolf/ac7a7da6b1888c2eeac8ac8b9b05d3d3",
      "votes": null
    },
    {
      "id": "1163839",
      "postDate": "01/22/2021 01:27:33",
      "content": "<p>Thank you, I have already realized that different models use different memory. It really helps.</p>",
      "rawMarkdown": "Thank you, I have already realized that different models use different memory. It really helps.",
      "votes": null
    },
    {
      "id": "1163840",
      "postDate": "01/22/2021 01:27:55",
      "content": "<p>Thank you for the advice!</p>",
      "rawMarkdown": "Thank you for the advice!",
      "votes": null
    },
    {
      "id": "1164235",
      "postDate": "01/22/2021 08:31:39",
      "content": "<p><a href=\"https://www.kaggle.com/felix26\" target=\"_blank\">@felix26</a><br>\nIf you want to use large models with bigger batch sizes then one way is to save checkpoints for the model, optimizer, scheduler, and load after restarting the kernel then continue the training from where you left. Although, this is time-consuming and requires patients.</p>",
      "rawMarkdown": "felix26\nIf you want to use large models with bigger batch sizes then one way is to save checkpoints for the model, optimizer, scheduler, and load after restarting the kernel then continue the training from where you left. Although, this is time-consuming and requires patients.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1162327,
      "author_name": "debarshichanda",
      "author_url": "",
      "post_date": "01/21/2021 05:23:00",
      "content": "<p>It depends upon your model<br>\nIf you are using Effnet b4 or Effnet b5 batch size of 32 isn't possible, the maximum you can reach is around 24 with Effnet b4<br>\nI think the notebooks that you are referring to are using small models like b0 or b1</p>",
      "votes": null,
      "replies": [
        {
          "id": 1163768,
          "author_name": "capiru",
          "author_url": "",
          "post_date": "01/21/2021 23:00:40",
          "content": "<p>I was gonna say the same thing, for instance VIT max batch size is around ~4. The bigger the model, the more RAM it needs, so if you choose for example 8 batch size, you'll have more RAM alocated as you go from EffNetB0 -&gt; B1 -&gt; B2 and so on, so either you choose a simpler, less intensive model or just a smaller batch size, which usually you can compensate more than well enough by adjusting the learning rate proportionally.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1163839,
          "author_name": "felix26",
          "author_url": "",
          "post_date": "01/22/2021 01:27:33",
          "content": "<p>Thank you, I have already realized that different models use different memory. It really helps.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1163773,
      "author_name": "sparakhin",
      "author_url": "",
      "post_date": "01/21/2021 23:16:50",
      "content": "<p>You can also use gradient accumulation.<br>\nExample: <a href=\"https://gist.github.com/thomwolf/ac7a7da6b1888c2eeac8ac8b9b05d3d3\" target=\"_blank\">https://gist.github.com/thomwolf/ac7a7da6b1888c2eeac8ac8b9b05d3d3</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1163840,
          "author_name": "felix26",
          "author_url": "",
          "post_date": "01/22/2021 01:27:55",
          "content": "<p>Thank you for the advice!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1164235,
      "author_name": "vatsalmavani",
      "author_url": "",
      "post_date": "01/22/2021 08:31:39",
      "content": "<p><a href=\"https://www.kaggle.com/felix26\" target=\"_blank\">@felix26</a><br>\nIf you want to use large models with bigger batch sizes then one way is to save checkpoints for the model, optimizer, scheduler, and load after restarting the kernel then continue the training from where you left. Although, this is time-consuming and requires patients.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1162186": "I use PyTorch and I read some notebooks. I see most people use 512x512 import images and 32 batch size. However, when I try to use 512x512，it always says the GPU is out of memory.\nI also tried to use apex to reduce memory, but it still uses more memory on GPU than available and causes a lot of overflows.\nIs there anything important I missed? I was confused by this and really need some advice.\nThank you very much.",
    "1162327": "It depends upon your model\nIf you are using Effnet b4 or Effnet b5 batch size of 32 isn't possible, the maximum you can reach is around 24 with Effnet b4\nI think the notebooks that you are referring to are using small models like b0 or b1",
    "1163768": "I was gonna say the same thing, for instance VIT max batch size is around ~4. The bigger the model, the more RAM it needs, so if you choose for example 8 batch size, you'll have more RAM alocated as you go from EffNetB0 -> B1 -> B2 and so on, so either you choose a simpler, less intensive model or just a smaller batch size, which usually you can compensate more than well enough by adjusting the learning rate proportionally.",
    "1163773": "You can also use gradient accumulation.\nExample: https://gist.github.com/thomwolf/ac7a7da6b1888c2eeac8ac8b9b05d3d3",
    "1163839": "Thank you, I have already realized that different models use different memory. It really helps.",
    "1163840": "Thank you for the advice!",
    "1164235": "felix26\nIf you want to use large models with bigger batch sizes then one way is to save checkpoints for the model, optimizer, scheduler, and load after restarting the kernel then continue the training from where you left. Although, this is time-consuming and requires patients."
  },
  "source": "meta"
}