{
  "id": 267076,
  "title": "GPU runs out of memory",
  "url": "/competitions/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/267076",
  "author_name": "",
  "post_date": "2021-08-21T15:43:34.910593100Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I can run the following on the CPU, but the code quickly uses up the GPU memory and errors out if I send it to the GPU.  I'd like to leverage the GPU tech if it makes sense to do so.  I'm using the FLAIR training data at 200x200x70 resolution, batch size 8.  Am I just asking too much from my NVidia 1070 8gb GPU? Or is my NN architecture badly designed for this data set?  This is effectively my first competition using NN's so I'm thinking there may be something big I'm missing.  Pardon the R syntax. Note, I got the cuda installed and R torch confirms GPU is available.</p>\n<p>net &lt;- nn_module(</p>\n<p>\"corr-cnn\",</p>\n<p>initialize = function() {</p>\n<pre><code>self$conv1 &lt;- nn_conv3d(in_channels = 1, out_channels = 16, kernel_size = c(3,3,3))\nself$conv2 &lt;- nn_conv3d(in_channels = 16, out_channels = 32, kernel_size = c(3,3,3))\nself$conv3 &lt;- nn_conv3d(in_channels = 32, out_channels = 16, kernel_size = c(3,3,3))\n\nself$fc1 &lt;- nn_linear(in_features = 161* 16*23, out_features = 16)\nself$fc2 &lt;- nn_linear(in_features = 16, out_features = 1)\n</code></pre>\n<p>},</p>\n<p>forward = function(x) {</p>\n<pre><code>x %&gt;% \n  self$conv1() %&gt;% \n  nnf_relu() %&gt;%\n  nnf_avg_pool3d(2) %&gt;%\n\n  self$conv2() %&gt;%\n  nnf_relu() %&gt;%\n  nnf_avg_pool3d(2) %&gt;%\n\n  self$conv3() %&gt;%\n  nnf_relu() %&gt;%\n  nnf_avg_pool3d(2) %&gt;%\n\n  torch_flatten(start_dim = 2) %&gt;%\n  self$fc1() %&gt;%\n  nnf_relu() %&gt;%\n\n  self$fc2()\n</code></pre>\n<p>}<br>\n)</p>\n<p>model &lt;- net()</p>\n<p>optimizer &lt;- optim_adam(model$parameters)</p>\n<p>train_batch &lt;- function(b) {</p>\n<p>optimizer$zero_grad()</p>\n<p># get predictions<br>\n  output &lt;- model(b$x)</p>\n<p># calculate loss<br>\n  loss &lt;- nnf_binary_cross_entropy_with_logits(output, b$y$unsqueeze(2))</p>\n<p># have gradients get calculated        <br>\n  loss$backward()</p>\n<p># have gradients get applied<br>\n  optimizer$step()</p>\n<p>loss$item()</p>\n<p>}</p>\n<p>valid_batch &lt;- function(b) {</p>\n<p>output &lt;- model(b$x)<br>\n  loss &lt;- nnf_binary_cross_entropy_with_logits(output, b$y$unsqueeze(2))<br>\n  loss$item()</p>\n<p>}</p>",
  "messages": [
    {
      "id": "1484836",
      "postDate": "08/21/2021 15:43:34",
      "content": "<p>I can run the following on the CPU, but the code quickly uses up the GPU memory and errors out if I send it to the GPU.  I'd like to leverage the GPU tech if it makes sense to do so.  I'm using the FLAIR training data at 200x200x70 resolution, batch size 8.  Am I just asking too much from my NVidia 1070 8gb GPU? Or is my NN architecture badly designed for this data set?  This is effectively my first competition using NN's so I'm thinking there may be something big I'm missing.  Pardon the R syntax. Note, I got the cuda installed and R torch confirms GPU is available.</p>\n<p>net &lt;- nn_module(</p>\n<p>\"corr-cnn\",</p>\n<p>initialize = function() {</p>\n<pre><code>self$conv1 &lt;- nn_conv3d(in_channels = 1, out_channels = 16, kernel_size = c(3,3,3))\nself$conv2 &lt;- nn_conv3d(in_channels = 16, out_channels = 32, kernel_size = c(3,3,3))\nself$conv3 &lt;- nn_conv3d(in_channels = 32, out_channels = 16, kernel_size = c(3,3,3))\n\nself$fc1 &lt;- nn_linear(in_features = 161* 16*23, out_features = 16)\nself$fc2 &lt;- nn_linear(in_features = 16, out_features = 1)\n</code></pre>\n<p>},</p>\n<p>forward = function(x) {</p>\n<pre><code>x %&gt;% \n  self$conv1() %&gt;% \n  nnf_relu() %&gt;%\n  nnf_avg_pool3d(2) %&gt;%\n\n  self$conv2() %&gt;%\n  nnf_relu() %&gt;%\n  nnf_avg_pool3d(2) %&gt;%\n\n  self$conv3() %&gt;%\n  nnf_relu() %&gt;%\n  nnf_avg_pool3d(2) %&gt;%\n\n  torch_flatten(start_dim = 2) %&gt;%\n  self$fc1() %&gt;%\n  nnf_relu() %&gt;%\n\n  self$fc2()\n</code></pre>\n<p>}<br>\n)</p>\n<p>model &lt;- net()</p>\n<p>optimizer &lt;- optim_adam(model$parameters)</p>\n<p>train_batch &lt;- function(b) {</p>\n<p>optimizer$zero_grad()</p>\n<p># get predictions<br>\n  output &lt;- model(b$x)</p>\n<p># calculate loss<br>\n  loss &lt;- nnf_binary_cross_entropy_with_logits(output, b$y$unsqueeze(2))</p>\n<p># have gradients get calculated        <br>\n  loss$backward()</p>\n<p># have gradients get applied<br>\n  optimizer$step()</p>\n<p>loss$item()</p>\n<p>}</p>\n<p>valid_batch &lt;- function(b) {</p>\n<p>output &lt;- model(b$x)<br>\n  loss &lt;- nnf_binary_cross_entropy_with_logits(output, b$y$unsqueeze(2))<br>\n  loss$item()</p>\n<p>}</p>",
      "rawMarkdown": "I can run the following on the CPU, but the code quickly uses up the GPU memory and errors out if I send it to the GPU.  I'd like to leverage the GPU tech if it makes sense to do so.  I'm using the FLAIR training data at 200x200x70 resolution, batch size 8.  Am I just asking too much from my NVidia 1070 8gb GPU? Or is my NN architecture badly designed for this data set?  This is effectively my first competition using NN's so I'm thinking there may be something big I'm missing.  Pardon the R syntax. Note, I got the cuda installed and R torch confirms GPU is available.\n\nnet <- nn_module(\n  \n  \"corr-cnn\",\n  \n  initialize = function() {\n    \n    self$conv1 <- nn_conv3d(in_channels = 1, out_channels = 16, kernel_size = c(3,3,3))\n    self$conv2 <- nn_conv3d(in_channels = 16, out_channels = 32, kernel_size = c(3,3,3))\n    self$conv3 <- nn_conv3d(in_channels = 32, out_channels = 16, kernel_size = c(3,3,3))\n    \n    self$fc1 <- nn_linear(in_features = 161* 16*23, out_features = 16)\n    self$fc2 <- nn_linear(in_features = 16, out_features = 1)\n    \n  },\n  \n  forward = function(x) {\n    \n    x %>% \n      self$conv1() %>% \n      nnf_relu() %>%\n      nnf_avg_pool3d(2) %>%\n\n      self$conv2() %>%\n      nnf_relu() %>%\n      nnf_avg_pool3d(2) %>%\n\n      self$conv3() %>%\n      nnf_relu() %>%\n      nnf_avg_pool3d(2) %>%\n\n      torch_flatten(start_dim = 2) %>%\n      self$fc1() %>%\n      nnf_relu() %>%\n\n      self$fc2()\n  }\n)\n\nmodel <- net()\n \n  optimizer <- optim_adam(model$parameters)\n\ntrain_batch <- function(b) {\n  \n  optimizer$zero_grad()\n  \n  # get predictions\n  output <- model(b$x)\n  \n  # calculate loss\n  loss <- nnf_binary_cross_entropy_with_logits(output, b$y$unsqueeze(2))\n  \n  # have gradients get calculated        \n  loss$backward()\n  \n  # have gradients get applied\n  optimizer$step()\n  \n  loss$item()\n  \n}\n\nvalid_batch <- function(b) {\n  \n  output <- model(b$x)\n  loss <- nnf_binary_cross_entropy_with_logits(output, b$y$unsqueeze(2))\n  loss$item()\n  \n}",
      "votes": null
    },
    {
      "id": "1485320",
      "postDate": "08/22/2021 02:11:09",
      "content": "<p>With 8 gigs of gpu memory i the maximum batch size you can do per iteration i think should be 2. <br>\nTo confirm this try running a batch size of 1 and check nvidia-smi to see how much memory is being using. I was able to run only with a batch size of 1 on my 2080(8 gigs).</p>\n<p>However with pytorch you can run with a batch size of 1 but you can accumulate gradients which and do backpass after 8 batches have passed. Check out <a href=\"https://discuss.pytorch.org/t/accumulating-gradients/30020\" target=\"_blank\">https://discuss.pytorch.org/t/accumulating-gradients/30020</a> if you wanna try out gradient accumulation.</p>",
      "rawMarkdown": "With 8 gigs of gpu memory i the maximum batch size you can do per iteration i think should be 2. \nTo confirm this try running a batch size of 1 and check nvidia-smi to see how much memory is being using. I was able to run only with a batch size of 1 on my 2080(8 gigs).\n\nHowever with pytorch you can run with a batch size of 1 but you can accumulate gradients which and do backpass after 8 batches have passed. Check out https://discuss.pytorch.org/t/accumulating-gradients/30020 if you wanna try out gradient accumulation.",
      "votes": null
    },
    {
      "id": "1485361",
      "postDate": "08/22/2021 03:20:45",
      "content": "<p>Thanks Aryaman, I'll give a batch size of 1 or 2 a shot.  I'll follow that example of accumulating gradients for a number of iterations before I take the step.</p>\n<p>It's helpful to know that I'm running into a GPU hardware limitation.</p>\n<p>Would reducing the output channels of the CNN layers reduce some of the memory requirements as well?</p>",
      "rawMarkdown": "Thanks Aryaman, I'll give a batch size of 1 or 2 a shot.  I'll follow that example of accumulating gradients for a number of iterations before I take the step.\n\nIt's helpful to know that I'm running into a GPU hardware limitation.\n\nWould reducing the output channels of the CNN layers reduce some of the memory requirements as well?",
      "votes": null
    },
    {
      "id": "1485376",
      "postDate": "08/22/2021 04:06:05",
      "content": "<p>It will but not significantly. It's more about how much memory the images take on the gpu. </p>",
      "rawMarkdown": "It will but not significantly. It's more about how much memory the images take on the gpu.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1485320,
      "author_name": "aryamansharma47",
      "author_url": "",
      "post_date": "08/22/2021 02:11:09",
      "content": "<p>With 8 gigs of gpu memory i the maximum batch size you can do per iteration i think should be 2. <br>\nTo confirm this try running a batch size of 1 and check nvidia-smi to see how much memory is being using. I was able to run only with a batch size of 1 on my 2080(8 gigs).</p>\n<p>However with pytorch you can run with a batch size of 1 but you can accumulate gradients which and do backpass after 8 batches have passed. Check out <a href=\"https://discuss.pytorch.org/t/accumulating-gradients/30020\" target=\"_blank\">https://discuss.pytorch.org/t/accumulating-gradients/30020</a> if you wanna try out gradient accumulation.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1485361,
          "author_name": "andyatkinson",
          "author_url": "",
          "post_date": "08/22/2021 03:20:45",
          "content": "<p>Thanks Aryaman, I'll give a batch size of 1 or 2 a shot.  I'll follow that example of accumulating gradients for a number of iterations before I take the step.</p>\n<p>It's helpful to know that I'm running into a GPU hardware limitation.</p>\n<p>Would reducing the output channels of the CNN layers reduce some of the memory requirements as well?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1485376,
          "author_name": "aryamansharma47",
          "author_url": "",
          "post_date": "08/22/2021 04:06:05",
          "content": "<p>It will but not significantly. It's more about how much memory the images take on the gpu. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1484836": "I can run the following on the CPU, but the code quickly uses up the GPU memory and errors out if I send it to the GPU.  I'd like to leverage the GPU tech if it makes sense to do so.  I'm using the FLAIR training data at 200x200x70 resolution, batch size 8.  Am I just asking too much from my NVidia 1070 8gb GPU? Or is my NN architecture badly designed for this data set?  This is effectively my first competition using NN's so I'm thinking there may be something big I'm missing.  Pardon the R syntax. Note, I got the cuda installed and R torch confirms GPU is available.\n\nnet <- nn_module(\n  \n  \"corr-cnn\",\n  \n  initialize = function() {\n    \n    self$conv1 <- nn_conv3d(in_channels = 1, out_channels = 16, kernel_size = c(3,3,3))\n    self$conv2 <- nn_conv3d(in_channels = 16, out_channels = 32, kernel_size = c(3,3,3))\n    self$conv3 <- nn_conv3d(in_channels = 32, out_channels = 16, kernel_size = c(3,3,3))\n    \n    self$fc1 <- nn_linear(in_features = 161* 16*23, out_features = 16)\n    self$fc2 <- nn_linear(in_features = 16, out_features = 1)\n    \n  },\n  \n  forward = function(x) {\n    \n    x %>% \n      self$conv1() %>% \n      nnf_relu() %>%\n      nnf_avg_pool3d(2) %>%\n\n      self$conv2() %>%\n      nnf_relu() %>%\n      nnf_avg_pool3d(2) %>%\n\n      self$conv3() %>%\n      nnf_relu() %>%\n      nnf_avg_pool3d(2) %>%\n\n      torch_flatten(start_dim = 2) %>%\n      self$fc1() %>%\n      nnf_relu() %>%\n\n      self$fc2()\n  }\n)\n\nmodel <- net()\n \n  optimizer <- optim_adam(model$parameters)\n\ntrain_batch <- function(b) {\n  \n  optimizer$zero_grad()\n  \n  # get predictions\n  output <- model(b$x)\n  \n  # calculate loss\n  loss <- nnf_binary_cross_entropy_with_logits(output, b$y$unsqueeze(2))\n  \n  # have gradients get calculated        \n  loss$backward()\n  \n  # have gradients get applied\n  optimizer$step()\n  \n  loss$item()\n  \n}\n\nvalid_batch <- function(b) {\n  \n  output <- model(b$x)\n  loss <- nnf_binary_cross_entropy_with_logits(output, b$y$unsqueeze(2))\n  loss$item()\n  \n}",
    "1485320": "With 8 gigs of gpu memory i the maximum batch size you can do per iteration i think should be 2. \nTo confirm this try running a batch size of 1 and check nvidia-smi to see how much memory is being using. I was able to run only with a batch size of 1 on my 2080(8 gigs).\n\nHowever with pytorch you can run with a batch size of 1 but you can accumulate gradients which and do backpass after 8 batches have passed. Check out https://discuss.pytorch.org/t/accumulating-gradients/30020 if you wanna try out gradient accumulation.",
    "1485361": "Thanks Aryaman, I'll give a batch size of 1 or 2 a shot.  I'll follow that example of accumulating gradients for a number of iterations before I take the step.\n\nIt's helpful to know that I'm running into a GPU hardware limitation.\n\nWould reducing the output channels of the CNN layers reduce some of the memory requirements as well?",
    "1485376": "It will but not significantly. It's more about how much memory the images take on the gpu."
  },
  "source": "meta"
}