{
  "id": 38207,
  "title": "\"all CUDA-capable devices are busy or unavailable\" when use heng's script",
  "url": "/competitions/carvana-image-masking-challenge/discussion/38207",
  "author_name": "",
  "post_date": "2017-08-16T18:46:58.459587200Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I had asked the question under Heng's starter kits but not got any answer for the problem.</p>\n\n<p>When I try to execute the code of Heng's script ( I slight ly modified the file path), I got the following messages:</p>\n\n<pre><code>THCudaCheck FAIL file=/pytorch/torch/lib/THC/generic/THCStorage.cu line=66 error=46 : all CUDA-capable devices are busy or unavailable\nTraceback (most recent call last):\n  File \"/mnt/home/dunan/Learn/Kaggle/carvana/carvana/train_seg_net.py\", line 608, in &lt;module&gt;\n    run_train()\n  File \"/mnt/home/dunan/Learn/Kaggle/carvana/carvana/train_seg_net.py\", line 296, in run_train\n    net.cuda()\n  File \"/mnt/home/dunan/.local/lib/python3.6/site-packages/torch/nn/modules/module.py\", line 147, in cuda\n    return self._apply(lambda t: t.cuda(device_id))\n  File \"/mnt/home/dunan/.local/lib/python3.6/site-packages/torch/nn/modules/module.py\", line 118, in _apply\n    module._apply(fn)\n  File \"/mnt/home/dunan/.local/lib/python3.6/site-packages/torch/nn/modules/module.py\", line 118, in _apply\n    module._apply(fn)\n  File \"/mnt/home/dunan/.local/lib/python3.6/site-packages/torch/nn/modules/module.py\", line 124, in _apply\n    param.data = fn(param.data)\n  File \"/mnt/home/dunan/.local/lib/python3.6/site-packages/torch/nn/modules/module.py\", line 147, in &lt;lambda&gt;\n    return self._apply(lambda t: t.cuda(device_id))\n  File \"/mnt/home/dunan/.local/lib/python3.6/site-packages/torch/_utils.py\", line 66, in _cuda\n    return new_type(self.size()).copy_(self, async)\n  File \"/mnt/home/dunan/.local/lib/python3.6/site-packages/torch/cuda/__init__.py\", line 269, in _lazy_new\n    return super(_CudaBase, cls).__new__(cls, *args, **kwargs)\nRuntimeError: cuda runtime error (46) : all CUDA-capable devices are busy or unavailable at /pytorch/torch/lib/THC/generic/THCStorage.cu:66\n</code></pre>\n\n<p>Noticed that I also try a simple test script:</p>\n\n<pre><code>from net.segmentation.my_unet import UNet_double_1024_5 as Net\nimport torch\n\nnet = Net(in_shape=(3, 1024, 1024), num_classes=1)\nnet.cuda()\n</code></pre>\n\n<p>But there is no error for this simple script. So seems there are some operation before the Net() stuff cause GPU unavailable? Just wonder any one met this problem and how to solve it?</p>",
  "messages": [
    {
      "id": "214390",
      "postDate": "08/16/2017 18:46:58",
      "content": "<p>I had asked the question under Heng's starter kits but not got any answer for the problem.</p>\n\n<p>When I try to execute the code of Heng's script ( I slight ly modified the file path), I got the following messages:</p>\n\n<pre><code>THCudaCheck FAIL file=/pytorch/torch/lib/THC/generic/THCStorage.cu line=66 error=46 : all CUDA-capable devices are busy or unavailable\nTraceback (most recent call last):\n  File \"/mnt/home/dunan/Learn/Kaggle/carvana/carvana/train_seg_net.py\", line 608, in &lt;module&gt;\n    run_train()\n  File \"/mnt/home/dunan/Learn/Kaggle/carvana/carvana/train_seg_net.py\", line 296, in run_train\n    net.cuda()\n  File \"/mnt/home/dunan/.local/lib/python3.6/site-packages/torch/nn/modules/module.py\", line 147, in cuda\n    return self._apply(lambda t: t.cuda(device_id))\n  File \"/mnt/home/dunan/.local/lib/python3.6/site-packages/torch/nn/modules/module.py\", line 118, in _apply\n    module._apply(fn)\n  File \"/mnt/home/dunan/.local/lib/python3.6/site-packages/torch/nn/modules/module.py\", line 118, in _apply\n    module._apply(fn)\n  File \"/mnt/home/dunan/.local/lib/python3.6/site-packages/torch/nn/modules/module.py\", line 124, in _apply\n    param.data = fn(param.data)\n  File \"/mnt/home/dunan/.local/lib/python3.6/site-packages/torch/nn/modules/module.py\", line 147, in &lt;lambda&gt;\n    return self._apply(lambda t: t.cuda(device_id))\n  File \"/mnt/home/dunan/.local/lib/python3.6/site-packages/torch/_utils.py\", line 66, in _cuda\n    return new_type(self.size()).copy_(self, async)\n  File \"/mnt/home/dunan/.local/lib/python3.6/site-packages/torch/cuda/__init__.py\", line 269, in _lazy_new\n    return super(_CudaBase, cls).__new__(cls, *args, **kwargs)\nRuntimeError: cuda runtime error (46) : all CUDA-capable devices are busy or unavailable at /pytorch/torch/lib/THC/generic/THCStorage.cu:66\n</code></pre>\n\n<p>Noticed that I also try a simple test script:</p>\n\n<pre><code>from net.segmentation.my_unet import UNet_double_1024_5 as Net\nimport torch\n\nnet = Net(in_shape=(3, 1024, 1024), num_classes=1)\nnet.cuda()\n</code></pre>\n\n<p>But there is no error for this simple script. So seems there are some operation before the Net() stuff cause GPU unavailable? Just wonder any one met this problem and how to solve it?</p>",
      "rawMarkdown": "I had asked the question under Heng's starter kits but not got any answer for the problem.\n\nWhen I try to execute the code of Heng's script ( I slight ly modified the file path), I got the following messages:\n\n    THCudaCheck FAIL file=/pytorch/torch/lib/THC/generic/THCStorage.cu line=66 error=46 : all CUDA-capable devices are busy or unavailable\n    Traceback (most recent call last):\n      File \"/mnt/home/dunan/Learn/Kaggle/carvana/carvana/train_seg_net.py\", line 608, in",
      "votes": null
    },
    {
      "id": "214455",
      "postDate": "08/17/2017 00:12:26",
      "content": "<p>Just hope anyone provide some idea to debug</p>",
      "rawMarkdown": "Just hope anyone provide some idea to debug",
      "votes": null
    },
    {
      "id": "214591",
      "postDate": "08/17/2017 13:33:48",
      "content": "<p>Can someone help...</p>",
      "rawMarkdown": "Can someone help...",
      "votes": null
    },
    {
      "id": "216365",
      "postDate": "08/25/2017 12:40:44",
      "content": "<p>the code of your run_train() function would be very heplful and your machine config (OS, number and type of GPUs etc)</p>",
      "rawMarkdown": "the code of your run_train() function would be very heplful and your machine config (OS, number and type of GPUs etc)",
      "votes": null
    },
    {
      "id": "217843",
      "postDate": "09/01/2017 07:34:13",
      "content": "<p>Probably some other application on your machine use GPU and not allow you to execute script. You can check which applications utilize GPU via typing <code>nvidia-smi</code> into terminal. Try to kill all of them and run script again.</p>",
      "rawMarkdown": "Probably some other application on your machine use GPU and not allow you to execute script. You can check which applications utilize GPU via typing `nvidia-smi` into terminal. Try to kill all of them and run script again.",
      "votes": null
    },
    {
      "id": "287441",
      "postDate": "02/24/2018 19:59:19",
      "content": "<p>CSAdu, were you able to find a fix for this? </p>",
      "rawMarkdown": "CSAdu, were you able to find a fix for this?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 214455,
      "author_name": "strideradu",
      "author_url": "",
      "post_date": "08/17/2017 00:12:26",
      "content": "<p>Just hope anyone provide some idea to debug</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 214591,
      "author_name": "strideradu",
      "author_url": "",
      "post_date": "08/17/2017 13:33:48",
      "content": "<p>Can someone help...</p>",
      "votes": null,
      "replies": [
        {
          "id": 216365,
          "author_name": "pama328",
          "author_url": "",
          "post_date": "08/25/2017 12:40:44",
          "content": "<p>the code of your run_train() function would be very heplful and your machine config (OS, number and type of GPUs etc)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 217843,
      "author_name": "jiy3ahko",
      "author_url": "",
      "post_date": "09/01/2017 07:34:13",
      "content": "<p>Probably some other application on your machine use GPU and not allow you to execute script. You can check which applications utilize GPU via typing <code>nvidia-smi</code> into terminal. Try to kill all of them and run script again.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 287441,
      "author_name": "avsanjay",
      "author_url": "",
      "post_date": "02/24/2018 19:59:19",
      "content": "<p>CSAdu, were you able to find a fix for this? </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "214390": "I had asked the question under Heng's starter kits but not got any answer for the problem.\n\nWhen I try to execute the code of Heng's script ( I slight ly modified the file path), I got the following messages:\n\n    THCudaCheck FAIL file=/pytorch/torch/lib/THC/generic/THCStorage.cu line=66 error=46 : all CUDA-capable devices are busy or unavailable\n    Traceback (most recent call last):\n      File \"/mnt/home/dunan/Learn/Kaggle/carvana/carvana/train_seg_net.py\", line 608, in",
    "214455": "Just hope anyone provide some idea to debug",
    "214591": "Can someone help...",
    "216365": "the code of your run_train() function would be very heplful and your machine config (OS, number and type of GPUs etc)",
    "217843": "Probably some other application on your machine use GPU and not allow you to execute script. You can check which applications utilize GPU via typing `nvidia-smi` into terminal. Try to kill all of them and run script again.",
    "287441": "CSAdu, were you able to find a fix for this?"
  },
  "source": "meta"
}