{
  "id": 133119,
  "title": "Google Colab and pytorch - CUDA out of memory",
  "url": "/competitions/bengaliai-cv19/discussion/133119",
  "author_name": "",
  "post_date": "2020-02-29T21:18:44.313036900Z",
  "votes": 1,
  "comment_count": 15,
  "views": 0,
  "content": "<p>RuntimeError: CUDA out of memory. Tried to allocate 2.00 MiB (GPU 0; 15.90 GiB total capacity; 15.18 GiB already allocated; 1.88 MiB free; 15.19 GiB reserved in total by PyTorch)</p>\n\n<p>Am I the only one with this problem?\nDoes someone know how to solve this?</p>\n\n<p>Thank you in advance</p>",
  "messages": [
    {
      "id": "760092",
      "postDate": "02/29/2020 21:18:44",
      "content": "<p>RuntimeError: CUDA out of memory. Tried to allocate 2.00 MiB (GPU 0; 15.90 GiB total capacity; 15.18 GiB already allocated; 1.88 MiB free; 15.19 GiB reserved in total by PyTorch)</p>\n\n<p>Am I the only one with this problem?\nDoes someone know how to solve this?</p>\n\n<p>Thank you in advance</p>",
      "rawMarkdown": "RuntimeError: CUDA out of memory. Tried to allocate 2.00 MiB (GPU 0; 15.90 GiB total capacity; 15.18 GiB already allocated; 1.88 MiB free; 15.19 GiB reserved in total by PyTorch)\n\nAm I the only one with this problem?\nDoes someone know how to solve this?\n\nThank you in advance",
      "votes": null
    },
    {
      "id": "760103",
      "postDate": "02/29/2020 21:36:38",
      "content": "<p>You must use less GPU VRAM by either (1) resizing your training images smaller (2) decreasing your batch size, (3) using mixed precision, or (4) using a model with fewer parameters, etc etc.</p>",
      "rawMarkdown": "You must use less GPU VRAM by either (1) resizing your training images smaller (2) decreasing your batch size, (3) using mixed precision, or (4) using a model with fewer parameters, etc etc.",
      "votes": null
    },
    {
      "id": "760197",
      "postDate": "03/01/2020 00:32:10",
      "content": "<p>Are you using PyTorch? If yes, and this problem happens during validation, then be sure to keep the validation loop inside <code>with torch.no_grad()</code>. Also, it is a good practice to have <code>model.eval()</code> before validation begins. I hope this answers your question.</p>",
      "rawMarkdown": "Are you using PyTorch? If yes, and this problem happens during validation, then be sure to keep the validation loop inside `with torch.no_grad()`. Also, it is a good practice to have `model.eval()` before validation begins. I hope this answers your question.",
      "votes": null
    },
    {
      "id": "760210",
      "postDate": "03/01/2020 01:12:39",
      "content": "<p>Thank you for the reply. \nI have already reduced the batch size... It doesn't seem to matter (it happens always around ~7%).\nI don't know how to use mixed precision.</p>",
      "rawMarkdown": "Thank you for the reply. \nI have already reduced the batch size... It doesn't seem to matter (it happens always around ~7%).\nI don't know how to use mixed precision.",
      "votes": null
    },
    {
      "id": "760211",
      "postDate": "03/01/2020 01:13:55",
      "content": "<p>Thank you for the reply.\nYes, I am using PyTorch. \nThe problem happens during validation, but torch.no_grad() doesn't seem to solve it. I was already using model.eval().</p>",
      "rawMarkdown": "Thank you for the reply.\nYes, I am using PyTorch. \nThe problem happens during validation, but torch.no_grad() doesn't seem to solve it. I was already using model.eval().",
      "votes": null
    },
    {
      "id": "760218",
      "postDate": "03/01/2020 01:47:11",
      "content": "<p>I just tried this one... same problem.</p>\n\n<p>Don't know what to do.</p>",
      "rawMarkdown": "I just tried this one... same problem.\n\nDon't know what to do.",
      "votes": null
    },
    {
      "id": "761742",
      "postDate": "03/02/2020 22:17:50",
      "content": "<p>I had the same problem. Reducing batch size solved the issue for me.</p>",
      "rawMarkdown": "I had the same problem. Reducing batch size solved the issue for me.",
      "votes": null
    },
    {
      "id": "764701",
      "postDate": "03/05/2020 19:00:36",
      "content": "<p>Shoud I put torch.no_grad() after model.eval()? </p>",
      "rawMarkdown": "Shoud I put torch.no_grad() after model.eval()?",
      "votes": null
    },
    {
      "id": "764876",
      "postDate": "03/06/2020 02:22:21",
      "content": "<p>I hope that you are using a function for evaluation. Your evaluation loop, that is iterating over the dataloaders should be within the <code>with torch.no_grad()</code> block. Something like the following:\n```\ndef validate(NUM_EPOCHS, epoch, model, testloader):\n    model.eval()\n    loss = 0.0\n    acc = 0\n    running_loss = 0.0\n    running_correct = 0\n    with torch.no_grad():\n        for i, data in enumerate(testloader):\n            img, labels = data[0].to(device), data[1].to(device)\n            new_img = new_img.to(device)\n            outputs = model(new_img)\n            loss = criterion(outputs, labels)\n            _, preds = torch.max(outputs.data, 1)\n            running_loss += loss.item()\n            running_correct += (preds == labels).sum().item()</p>\n\n<pre><code>loss = running_loss / len(testset)\nacc = 100. * running_correct / len(testset)\nreturn loss, acc \n</code></pre>\n\n<p>```\nThis will most probably work. If not, try reducing the batch size along with it.</p>",
      "rawMarkdown": "I hope that you are using a function for evaluation. Your evaluation loop, that is iterating over the dataloaders should be within the `with torch.no_grad()` block. Something like the following:\n```\ndef validate(NUM_EPOCHS, epoch, model, testloader):\n    model.eval()\n    loss = 0.0\n    acc = 0\n    running_loss = 0.0\n    running_correct = 0\n    with torch.no_grad():\n        for i, data in enumerate(testloader):\n            img, labels = data[0].to(device), data[1].to(device)\n            new_img = new_img.to(device)\n            outputs = model(new_img)\n            loss = criterion(outputs, labels)\n            _, preds = torch.max(outputs.data, 1)\n            running_loss += loss.item()\n            running_correct += (preds == labels).sum().item()\n \n    loss = running_loss / len(testset)\n    acc = 100. * running_correct / len(testset)\n    return loss, acc \n```\nThis will most probably work. If not, try reducing the batch size along with it.",
      "votes": null
    },
    {
      "id": "764981",
      "postDate": "03/06/2020 06:44:08",
      "content": "<p>Thanks man. I would surely give it a try. I am already using batch size of 32. Should I reduce it more? Right now, my first training epoch runs successfully and takes only 1. 8gb gpu ram and immediately after training step as it enters into validation step, CUDA goes out of memory. </p>",
      "rawMarkdown": "Thanks man. I would surely give it a try. I am already using batch size of 32. Should I reduce it more? Right now, my first training epoch runs successfully and takes only 1. 8gb gpu ram and immediately after training step as it enters into validation step, CUDA goes out of memory.",
      "votes": null
    },
    {
      "id": "765182",
      "postDate": "03/06/2020 10:37:33",
      "content": "<p>The GPU VRAM usage seems reasonable. I don't think that you need to reduce the batch size. Try the <code>with torch.no_grad()</code> solution.</p>",
      "rawMarkdown": "The GPU VRAM usage seems reasonable. I don't think that you need to reduce the batch size. Try the `with torch.no_grad()` solution.",
      "votes": null
    },
    {
      "id": "765548",
      "postDate": "03/06/2020 19:30:47",
      "content": "<p>Yup, it worked for me. Thanks.</p>",
      "rawMarkdown": "Yup, it worked for me. Thanks.",
      "votes": null
    },
    {
      "id": "765718",
      "postDate": "03/07/2020 01:55:46",
      "content": "<p>Glad to here that.</p>",
      "rawMarkdown": "Glad to here that.",
      "votes": null
    },
    {
      "id": "857796",
      "postDate": "05/22/2020 23:35:51",
      "content": "<p>i have same problem. I use google colab.  at menu runtime, I choose \"reset the process to factory settings\" to solve this  issue. </p>",
      "rawMarkdown": "i have same problem. I use google colab.  at menu runtime, I choose \"reset the process to factory settings\" to solve this  issue.",
      "votes": null
    },
    {
      "id": "1064669",
      "postDate": "10/30/2020 11:48:45",
      "content": "<p>Thanks man, you saved my code from turning into crap. Just had to push one button and it worked!! thanks you solved my issue.</p>",
      "rawMarkdown": "Thanks man, you saved my code from turning into crap. Just had to push one button and it worked!! thanks you solved my issue.",
      "votes": null
    },
    {
      "id": "1179170",
      "postDate": "01/31/2021 11:27:03",
      "content": "<p>Thanks, worked for me as well</p>",
      "rawMarkdown": "Thanks, worked for me as well",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 760103,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "02/29/2020 21:36:38",
      "content": "<p>You must use less GPU VRAM by either (1) resizing your training images smaller (2) decreasing your batch size, (3) using mixed precision, or (4) using a model with fewer parameters, etc etc.</p>",
      "votes": null,
      "replies": [
        {
          "id": 760210,
          "author_name": "joovasco",
          "author_url": "",
          "post_date": "03/01/2020 01:12:39",
          "content": "<p>Thank you for the reply. \nI have already reduced the batch size... It doesn't seem to matter (it happens always around ~7%).\nI don't know how to use mixed precision.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 760218,
          "author_name": "joovasco",
          "author_url": "",
          "post_date": "03/01/2020 01:47:11",
          "content": "<p>I just tried this one... same problem.</p>\n\n<p>Don't know what to do.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 760197,
      "author_name": "sovitrath",
      "author_url": "",
      "post_date": "03/01/2020 00:32:10",
      "content": "<p>Are you using PyTorch? If yes, and this problem happens during validation, then be sure to keep the validation loop inside <code>with torch.no_grad()</code>. Also, it is a good practice to have <code>model.eval()</code> before validation begins. I hope this answers your question.</p>",
      "votes": null,
      "replies": [
        {
          "id": 760211,
          "author_name": "joovasco",
          "author_url": "",
          "post_date": "03/01/2020 01:13:55",
          "content": "<p>Thank you for the reply.\nYes, I am using PyTorch. \nThe problem happens during validation, but torch.no_grad() doesn't seem to solve it. I was already using model.eval().</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 764701,
          "author_name": "abdurrehman245",
          "author_url": "",
          "post_date": "03/05/2020 19:00:36",
          "content": "<p>Shoud I put torch.no_grad() after model.eval()? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 764876,
          "author_name": "sovitrath",
          "author_url": "",
          "post_date": "03/06/2020 02:22:21",
          "content": "<p>I hope that you are using a function for evaluation. Your evaluation loop, that is iterating over the dataloaders should be within the <code>with torch.no_grad()</code> block. Something like the following:\n```\ndef validate(NUM_EPOCHS, epoch, model, testloader):\n    model.eval()\n    loss = 0.0\n    acc = 0\n    running_loss = 0.0\n    running_correct = 0\n    with torch.no_grad():\n        for i, data in enumerate(testloader):\n            img, labels = data[0].to(device), data[1].to(device)\n            new_img = new_img.to(device)\n            outputs = model(new_img)\n            loss = criterion(outputs, labels)\n            _, preds = torch.max(outputs.data, 1)\n            running_loss += loss.item()\n            running_correct += (preds == labels).sum().item()</p>\n\n<pre><code>loss = running_loss / len(testset)\nacc = 100. * running_correct / len(testset)\nreturn loss, acc \n</code></pre>\n\n<p>```\nThis will most probably work. If not, try reducing the batch size along with it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 764981,
          "author_name": "abdurrehman245",
          "author_url": "",
          "post_date": "03/06/2020 06:44:08",
          "content": "<p>Thanks man. I would surely give it a try. I am already using batch size of 32. Should I reduce it more? Right now, my first training epoch runs successfully and takes only 1. 8gb gpu ram and immediately after training step as it enters into validation step, CUDA goes out of memory. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 765182,
          "author_name": "sovitrath",
          "author_url": "",
          "post_date": "03/06/2020 10:37:33",
          "content": "<p>The GPU VRAM usage seems reasonable. I don't think that you need to reduce the batch size. Try the <code>with torch.no_grad()</code> solution.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 765548,
          "author_name": "abdurrehman245",
          "author_url": "",
          "post_date": "03/06/2020 19:30:47",
          "content": "<p>Yup, it worked for me. Thanks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 765718,
          "author_name": "sovitrath",
          "author_url": "",
          "post_date": "03/07/2020 01:55:46",
          "content": "<p>Glad to here that.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 761742,
      "author_name": "mhviraf",
      "author_url": "",
      "post_date": "03/02/2020 22:17:50",
      "content": "<p>I had the same problem. Reducing batch size solved the issue for me.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 857796,
      "author_name": "musgan",
      "author_url": "",
      "post_date": "05/22/2020 23:35:51",
      "content": "<p>i have same problem. I use google colab.  at menu runtime, I choose \"reset the process to factory settings\" to solve this  issue. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1064669,
          "author_name": "harshs21",
          "author_url": "",
          "post_date": "10/30/2020 11:48:45",
          "content": "<p>Thanks man, you saved my code from turning into crap. Just had to push one button and it worked!! thanks you solved my issue.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1179170,
          "author_name": "ladanmehran",
          "author_url": "",
          "post_date": "01/31/2021 11:27:03",
          "content": "<p>Thanks, worked for me as well</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "760092": "RuntimeError: CUDA out of memory. Tried to allocate 2.00 MiB (GPU 0; 15.90 GiB total capacity; 15.18 GiB already allocated; 1.88 MiB free; 15.19 GiB reserved in total by PyTorch)\n\nAm I the only one with this problem?\nDoes someone know how to solve this?\n\nThank you in advance",
    "760103": "You must use less GPU VRAM by either (1) resizing your training images smaller (2) decreasing your batch size, (3) using mixed precision, or (4) using a model with fewer parameters, etc etc.",
    "760197": "Are you using PyTorch? If yes, and this problem happens during validation, then be sure to keep the validation loop inside `with torch.no_grad()`. Also, it is a good practice to have `model.eval()` before validation begins. I hope this answers your question.",
    "760210": "Thank you for the reply. \nI have already reduced the batch size... It doesn't seem to matter (it happens always around ~7%).\nI don't know how to use mixed precision.",
    "760211": "Thank you for the reply.\nYes, I am using PyTorch. \nThe problem happens during validation, but torch.no_grad() doesn't seem to solve it. I was already using model.eval().",
    "760218": "I just tried this one... same problem.\n\nDon't know what to do.",
    "761742": "I had the same problem. Reducing batch size solved the issue for me.",
    "764701": "Shoud I put torch.no_grad() after model.eval()?",
    "764876": "I hope that you are using a function for evaluation. Your evaluation loop, that is iterating over the dataloaders should be within the `with torch.no_grad()` block. Something like the following:\n```\ndef validate(NUM_EPOCHS, epoch, model, testloader):\n    model.eval()\n    loss = 0.0\n    acc = 0\n    running_loss = 0.0\n    running_correct = 0\n    with torch.no_grad():\n        for i, data in enumerate(testloader):\n            img, labels = data[0].to(device), data[1].to(device)\n            new_img = new_img.to(device)\n            outputs = model(new_img)\n            loss = criterion(outputs, labels)\n            _, preds = torch.max(outputs.data, 1)\n            running_loss += loss.item()\n            running_correct += (preds == labels).sum().item()\n \n    loss = running_loss / len(testset)\n    acc = 100. * running_correct / len(testset)\n    return loss, acc \n```\nThis will most probably work. If not, try reducing the batch size along with it.",
    "764981": "Thanks man. I would surely give it a try. I am already using batch size of 32. Should I reduce it more? Right now, my first training epoch runs successfully and takes only 1. 8gb gpu ram and immediately after training step as it enters into validation step, CUDA goes out of memory.",
    "765182": "The GPU VRAM usage seems reasonable. I don't think that you need to reduce the batch size. Try the `with torch.no_grad()` solution.",
    "765548": "Yup, it worked for me. Thanks.",
    "765718": "Glad to here that.",
    "857796": "i have same problem. I use google colab.  at menu runtime, I choose \"reset the process to factory settings\" to solve this  issue.",
    "1064669": "Thanks man, you saved my code from turning into crap. Just had to push one button and it worked!! thanks you solved my issue.",
    "1179170": "Thanks, worked for me as well"
  },
  "source": "meta"
}