{
  "id": 47252,
  "title": "GCloud NVidia error?",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/47252",
  "author_name": "",
  "post_date": "2018-01-11T00:30:49.641725400Z",
  "votes": 2,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n\n<p>I have been using a virtual machine on Google Cloud (Ubuntu 16.04 LTS) for a while now with a NVidia Tesla K80 without any problems. All of a sudden, I cannot use the the GPU anymore.</p>\n\n<p>I get the following error at training time:</p>\n\n<p>2018-01-11 00:23:57.759222: E tensorflow/stream_executor/cuda/cuda_driver.cc:406] failed call to cuInit: CUDA_ERROR_UNKNOWN\n2018-01-11 00:23:57.759293: I tensorflow/stream_executor/cuda/cuda_diagnostics.cc:145] kernel driver does not appear to be running on this host (gpu-1): /proc/driver/nvidia/version does not exist</p>\n\n<p>When I run nvidia-smi, I get the following error:</p>\n\n<p>NVIDIA-SMI has failed because it couldn't communicate with the NVIDIA driver. Make sure that the latest NVIDIA driver is installed and running.</p>\n\n<p>Nothing has changed since last night.  I even tried reverting creating a new instance with a fully functional snapshot of the machine that I took  a few days ago with no luck.</p>\n\n<p>Anyone else having the same issue? Any ideas of what it could be?</p>\n\n<p>Thanks!</p>",
  "messages": [
    {
      "id": "267287",
      "postDate": "01/11/2018 00:30:49",
      "content": "<p>Hi everyone,</p>\n\n<p>I have been using a virtual machine on Google Cloud (Ubuntu 16.04 LTS) for a while now with a NVidia Tesla K80 without any problems. All of a sudden, I cannot use the the GPU anymore.</p>\n\n<p>I get the following error at training time:</p>\n\n<p>2018-01-11 00:23:57.759222: E tensorflow/stream_executor/cuda/cuda_driver.cc:406] failed call to cuInit: CUDA_ERROR_UNKNOWN\n2018-01-11 00:23:57.759293: I tensorflow/stream_executor/cuda/cuda_diagnostics.cc:145] kernel driver does not appear to be running on this host (gpu-1): /proc/driver/nvidia/version does not exist</p>\n\n<p>When I run nvidia-smi, I get the following error:</p>\n\n<p>NVIDIA-SMI has failed because it couldn't communicate with the NVIDIA driver. Make sure that the latest NVIDIA driver is installed and running.</p>\n\n<p>Nothing has changed since last night.  I even tried reverting creating a new instance with a fully functional snapshot of the machine that I took  a few days ago with no luck.</p>\n\n<p>Anyone else having the same issue? Any ideas of what it could be?</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "Hi everyone,\n\nI have been using a virtual machine on Google Cloud (Ubuntu 16.04 LTS) for a while now with a NVidia Tesla K80 without any problems. All of a sudden, I cannot use the the GPU anymore.\n\nI get the following error at training time:\n\n2018-01-11 00:23:57.759222: E tensorflow/stream_executor/cuda/cuda_driver.cc:406] failed call to cuInit: CUDA_ERROR_UNKNOWN\n2018-01-11 00:23:57.759293: I tensorflow/stream_executor/cuda/cuda_diagnostics.cc:145] kernel driver does not appear to be running on this host (gpu-1): /proc/driver/nvidia/version does not exist\n\nWhen I run nvidia-smi, I get the following error:\n\nNVIDIA-SMI has failed because it couldn't communicate with the NVIDIA driver. Make sure that the latest NVIDIA driver is installed and running.\n\nNothing has changed since last night.  I even tried reverting creating a new instance with a fully functional snapshot of the machine that I took  a few days ago with no luck.\n\nAnyone else having the same issue? Any ideas of what it could be?\n\nThanks!",
      "votes": null
    },
    {
      "id": "267339",
      "postDate": "01/11/2018 03:50:34",
      "content": "<p>same, as many others I suppose. \nreinstall cuda+drivers <a href=\"https://cloud.google.com/compute/docs/gpus/add-gpus#install-gpu-driver\">https://cloud.google.com/compute/docs/gpus/add-gpus#install-gpu-driver</a>\nand export LD_LIBRARY_PATH=/usr/local/cuda/lib64:/usr/lib/nvidia-387 into your .bashrc</p>",
      "rawMarkdown": "same, as many others I suppose. \nreinstall cuda+drivers https://cloud.google.com/compute/docs/gpus/add-gpus#install-gpu-driver\nand export LD_LIBRARY_PATH=/usr/local/cuda/lib64:/usr/lib/nvidia-387 into your .bashrc",
      "votes": null
    },
    {
      "id": "267532",
      "postDate": "01/11/2018 16:35:43",
      "content": "<p>I reinstalled CUDA using apt-get and the error remains. Then I download the 'runfile(local)' from the NVIDIA website, use apt-get to install nvidia-384, and add the following to the ~/.bashrc :</p>\n\n<pre><code>export PATH=\"$PATH:/usr/local/cuda-8.0/bin\" \nexport LD_LIBRARY_PATH=\"/usr/local/cuda-8.0/lib64\"\n</code></pre>\n\n<p>It just works again.</p>",
      "rawMarkdown": "I reinstalled CUDA using apt-get and the error remains. Then I download the 'runfile(local)' from the NVIDIA website, use apt-get to install nvidia-384, and add the following to the ~/.bashrc :\n\n    export PATH=\"$PATH:/usr/local/cuda-8.0/bin\" \n    export LD_LIBRARY_PATH=\"/usr/local/cuda-8.0/lib64\"\n\nIt just works again.",
      "votes": null
    },
    {
      "id": "267707",
      "postDate": "01/12/2018 04:50:30",
      "content": "<p>Thanks guys, I did not have the same luck. I suspect this has to do something with the latest kernel update due to Spectre/Meltdown. Switching to Centos 7 solved the problem in 5 minutes.</p>",
      "rawMarkdown": "Thanks guys, I did not have the same luck. I suspect this has to do something with the latest kernel update due to Spectre/Meltdown. Switching to Centos 7 solved the problem in 5 minutes.",
      "votes": null
    },
    {
      "id": "267749",
      "postDate": "01/12/2018 07:18:21",
      "content": "<p>Even though CentOS was working, I wanted to fix my Ubuntu. Installing a different kernel did it for me:</p>\n\n<p>mkdir newkernel</p>\n\n<p>cd newkernel</p>\n\n<p>wget <a href=\"http://kernel.ubuntu.com/~kernel-ppa/mainline/v4.14/linux-headers-4.14.0-041400_4.14.0-041400.201711122031_all.deb\">http://kernel.ubuntu.com/~kernel-ppa/mainline/v4.14/linux-headers-4.14.0-041400_4.14.0-041400.201711122031_all.deb</a></p>\n\n<p>wget <a href=\"http://kernel.ubuntu.com/~kernel-ppa/mainline/v4.14/linux-headers-4.14.0-041400-generic_4.14.0-041400.201711122031_amd64.deb\">http://kernel.ubuntu.com/~kernel-ppa/mainline/v4.14/linux-headers-4.14.0-041400-generic_4.14.0-041400.201711122031_amd64.deb</a></p>\n\n<p>wget <a href=\"http://kernel.ubuntu.com/~kernel-ppa/mainline/v4.14/linux-image-4.14.0-041400-generic_4.14.0-041400.201711122031_amd64.deb\">http://kernel.ubuntu.com/~kernel-ppa/mainline/v4.14/linux-image-4.14.0-041400-generic_4.14.0-041400.201711122031_amd64.deb</a></p>\n\n<p>sudo dpkg -i *.deb</p>\n\n<p>sudo reboot</p>\n\n<p>Before:</p>\n\n<p>$ uname -a</p>\n\n<p>Linux gpu-1 4.13.0-1006-gcp #9-Ubuntu SMP Mon Jan 8 21:13:15 UTC 2018 x86_64 x86_64 x86_64 GNU/Linux</p>\n\n<p>$ nvidia-smi</p>\n\n<p>NVIDIA-SMI has failed because it couldn't communicate with the NVIDIA driver. Make sure that the latest NVIDIA driver is installed and running.</p>\n\n<p>After:</p>\n\n<p>$ uname -a</p>\n\n<p>Linux gpu-1 4.14.0-041400-generic #201711122031 SMP Sun Nov 12 20:32:29 UTC 2017 x86_64 x86_64 x86_64 GNU/Linux</p>\n\n<p>$ nvidia-smi</p>\n\n<p>Fri Jan 12 07:17:44 2018 <br>\n+-----------------------------------------------------------------------------+\n| NVIDIA-SMI 387.26                 Driver Version: 387.26                    |\n|-------------------------------+----------------------+----------------------+\n| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |\n| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |\n|===============================+======================+======================|\n|   0  Tesla K80           Off  | 00000000:00:04.0 Off |                    0 |\n| N/A   31C    P8    26W / 149W |     16MiB / 11439MiB |      0%      Default |\n+-------------------------------+----------------------+----------------------+</p>\n\n<p>+-----------------------------------------------------------------------------+\n| Processes:                                                       GPU Memory |\n|  GPU       PID   Type   Process name                             Usage      |\n|=============================================================================|\n|    0      1639      G   /usr/lib/xorg/Xorg                            15MiB |\n+-----------------------------------------------------------------------------+</p>",
      "rawMarkdown": "Even though CentOS was working, I wanted to fix my Ubuntu. Installing a different kernel did it for me:\n\nmkdir newkernel\n\ncd newkernel\n\nwget http://kernel.ubuntu.com/~kernel-ppa/mainline/v4.14/linux-headers-4.14.0-041400_4.14.0-041400.201711122031_all.deb\n\nwget http://kernel.ubuntu.com/~kernel-ppa/mainline/v4.14/linux-headers-4.14.0-041400-generic_4.14.0-041400.201711122031_amd64.deb\n\nwget http://kernel.ubuntu.com/~kernel-ppa/mainline/v4.14/linux-image-4.14.0-041400-generic_4.14.0-041400.201711122031_amd64.deb\n\nsudo dpkg -i *.deb\n\nsudo reboot\n\n\nBefore:\n\n$ uname -a\n\nLinux gpu-1 4.13.0-1006-gcp #9-Ubuntu SMP Mon Jan 8 21:13:15 UTC 2018 x86_64 x86_64 x86_64 GNU/Linux\n\n$ nvidia-smi\n\nNVIDIA-SMI has failed because it couldn't communicate with the NVIDIA driver. Make sure that the latest NVIDIA driver is installed and running.\n\nAfter:\n\n$ uname -a\n\nLinux gpu-1 4.14.0-041400-generic #201711122031 SMP Sun Nov 12 20:32:29 UTC 2017 x86_64 x86_64 x86_64 GNU/Linux\n\n$ nvidia-smi\n\nFri Jan 12 07:17:44 2018       \n+-----------------------------------------------------------------------------+\n| NVIDIA-SMI 387.26                 Driver Version: 387.26                    |\n|-------------------------------+----------------------+----------------------+\n| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |\n| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |\n|===============================+======================+======================|\n|   0  Tesla K80           Off  | 00000000:00:04.0 Off |                    0 |\n| N/A   31C    P8    26W / 149W |     16MiB / 11439MiB |      0%      Default |\n+-------------------------------+----------------------+----------------------+\n                                                                               \n+-----------------------------------------------------------------------------+\n| Processes:                                                       GPU Memory |\n|  GPU       PID   Type   Process name                             Usage      |\n|=============================================================================|\n|    0      1639      G   /usr/lib/xorg/Xorg                            15MiB |\n+-----------------------------------------------------------------------------+",
      "votes": null
    },
    {
      "id": "267757",
      "postDate": "01/12/2018 07:42:42",
      "content": "<p>Marco, thanks for sharing this solution. </p>\n\n<p>For me reinstalling the Nvidia drivers didn't solve the problem. But installing a new kernel also did the trick for me (running Ubuntu 16.04.3 LTS).</p>",
      "rawMarkdown": "Marco, thanks for sharing this solution. \n\nFor me reinstalling the Nvidia drivers didn't solve the problem. But installing a new kernel also did the trick for me (running Ubuntu 16.04.3 LTS).",
      "votes": null
    },
    {
      "id": "267788",
      "postDate": "01/12/2018 09:50:52",
      "content": "<p>I followed this <a href=\"http://www.python36.com/install-tensorflow141-gpu/\">guide</a> and it worked well (with updated nvidia drivers and CUDA9.1). Since TF 1.4 does not support directly cuda 9.1 the guide suggests to build tf from source. (There is also a fix in the comments for a missing lib)</p>",
      "rawMarkdown": "I followed this [guide][1] and it worked well (with updated nvidia drivers and CUDA9.1). Since TF 1.4 does not support directly cuda 9.1 the guide suggests to build tf from source. (There is also a fix in the comments for a missing lib)\n\n\n  [1]: http://www.python36.com/install-tensorflow141-gpu/",
      "votes": null
    },
    {
      "id": "267820",
      "postDate": "01/12/2018 12:34:28",
      "content": "<p>I faced the same problem with ubuntu 16.04 LTS. I ended up booting a new machine with ubuntu 17.04</p>",
      "rawMarkdown": "I faced the same problem with ubuntu 16.04 LTS. I ended up booting a new machine with ubuntu 17.04",
      "votes": null
    },
    {
      "id": "267824",
      "postDate": "01/12/2018 12:59:27",
      "content": "<p>My team had the same problem, we were on Ubuntu 16.04 and Tesla K80. We think it's related to google deploying in the background a security patch for the kernel against meltdown/spectre attacks. The solution was to roll it back to its previous unsafe state.</p>",
      "rawMarkdown": "My team had the same problem, we were on Ubuntu 16.04 and Tesla K80. We think it's related to google deploying in the background a security patch for the kernel against meltdown/spectre attacks. The solution was to roll it back to its previous unsafe state.",
      "votes": null
    },
    {
      "id": "267907",
      "postDate": "01/12/2018 17:48:06",
      "content": "<p>You're welcome, Peter! Glad I could help.</p>",
      "rawMarkdown": "You're welcome, Peter! Glad I could help.",
      "votes": null
    },
    {
      "id": "267920",
      "postDate": "01/12/2018 18:51:01",
      "content": "<p>I had same issue, but just 'sudo apt-get install nvidia-384' (thanks to WeiWei) and rebooting the machine worked for me.\nBTW, K80 is much slower than P100 (1/6 speed in my experience), so I think P100 is more cost &amp; time effective for training purpose. Though GCP admin required some 'deposit' for increasing quota.</p>",
      "rawMarkdown": "I had same issue, but just 'sudo apt-get install nvidia-384' (thanks to WeiWei) and rebooting the machine worked for me.\nBTW, K80 is much slower than P100 (1/6 speed in my experience), so I think P100 is more cost &amp; time effective for training purpose. Though GCP admin required some 'deposit' for increasing quota.",
      "votes": null
    },
    {
      "id": "267929",
      "postDate": "01/12/2018 19:53:10",
      "content": "<p>I had this problem on P100 GCP instances, but not on K80 instances.  Super frustrating.  I finally solved it by installing the latest 390 Nvidia drivers:</p>\n\n<p>sudo apt-get -y purge nvidia*</p>\n\n<p>sudo add-apt-repository ppa:graphics-drivers</p>\n\n<p>sudo apt-get -y update</p>\n\n<p>sudo apt-get -y install nvidia-390</p>\n\n<p>sudo reboot</p>\n\n<p>After you do that, you can't install cuda with apt-get because that will revert the driver.  instead you have to use the run scripts from nvidia <a href=\"https://developer.nvidia.com/cuda-80-ga2-download-archive\">https://developer.nvidia.com/cuda-80-ga2-download-archive</a> and follow the \"runfile\" directions.  That script will ask if you want to revert the nvidia drivers. say \"no\".</p>",
      "rawMarkdown": "I had this problem on P100 GCP instances, but not on K80 instances.  Super frustrating.  I finally solved it by installing the latest 390 Nvidia drivers:\n\n\nsudo apt-get -y purge nvidia*\n\nsudo add-apt-repository ppa:graphics-drivers\n\nsudo apt-get -y update\n\nsudo apt-get -y install nvidia-390\n\nsudo reboot\n\n\nAfter you do that, you can't install cuda with apt-get because that will revert the driver.  instead you have to use the run scripts from nvidia https://developer.nvidia.com/cuda-80-ga2-download-archive and follow the \"runfile\" directions.  That script will ask if you want to revert the nvidia drivers. say \"no\".",
      "votes": null
    },
    {
      "id": "268531",
      "postDate": "01/14/2018 19:20:09",
      "content": "<p>I had the same problem on Ubuntu 16.04.3 with K80. This solution worked for me, thank you for sharing.</p>",
      "rawMarkdown": "I had the same problem on Ubuntu 16.04.3 with K80. This solution worked for me, thank you for sharing.",
      "votes": null
    },
    {
      "id": "268573",
      "postDate": "01/14/2018 23:01:43",
      "content": "<p>I followed that guide. One thing I need to add is a symbolic link </p>\n\n<p>ln -s /usr/local/cuda/include/crt/math_functions.hpp /usr/local/cuda/include/math_functions.hpp</p>\n\n<p>Other than that the guide works great. </p>",
      "rawMarkdown": "I followed that guide. One thing I need to add is a symbolic link \n\nln -s /usr/local/cuda/include/crt/math_functions.hpp /usr/local/cuda/include/math_functions.hpp\n\nOther than that the guide works great.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 267339,
      "author_name": "left13",
      "author_url": "",
      "post_date": "01/11/2018 03:50:34",
      "content": "<p>same, as many others I suppose. \nreinstall cuda+drivers <a href=\"https://cloud.google.com/compute/docs/gpus/add-gpus#install-gpu-driver\">https://cloud.google.com/compute/docs/gpus/add-gpus#install-gpu-driver</a>\nand export LD_LIBRARY_PATH=/usr/local/cuda/lib64:/usr/lib/nvidia-387 into your .bashrc</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 267532,
      "author_name": "bencww",
      "author_url": "",
      "post_date": "01/11/2018 16:35:43",
      "content": "<p>I reinstalled CUDA using apt-get and the error remains. Then I download the 'runfile(local)' from the NVIDIA website, use apt-get to install nvidia-384, and add the following to the ~/.bashrc :</p>\n\n<pre><code>export PATH=\"$PATH:/usr/local/cuda-8.0/bin\" \nexport LD_LIBRARY_PATH=\"/usr/local/cuda-8.0/lib64\"\n</code></pre>\n\n<p>It just works again.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 267707,
      "author_name": "malr87",
      "author_url": "",
      "post_date": "01/12/2018 04:50:30",
      "content": "<p>Thanks guys, I did not have the same luck. I suspect this has to do something with the latest kernel update due to Spectre/Meltdown. Switching to Centos 7 solved the problem in 5 minutes.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 267749,
      "author_name": "malr87",
      "author_url": "",
      "post_date": "01/12/2018 07:18:21",
      "content": "<p>Even though CentOS was working, I wanted to fix my Ubuntu. Installing a different kernel did it for me:</p>\n\n<p>mkdir newkernel</p>\n\n<p>cd newkernel</p>\n\n<p>wget <a href=\"http://kernel.ubuntu.com/~kernel-ppa/mainline/v4.14/linux-headers-4.14.0-041400_4.14.0-041400.201711122031_all.deb\">http://kernel.ubuntu.com/~kernel-ppa/mainline/v4.14/linux-headers-4.14.0-041400_4.14.0-041400.201711122031_all.deb</a></p>\n\n<p>wget <a href=\"http://kernel.ubuntu.com/~kernel-ppa/mainline/v4.14/linux-headers-4.14.0-041400-generic_4.14.0-041400.201711122031_amd64.deb\">http://kernel.ubuntu.com/~kernel-ppa/mainline/v4.14/linux-headers-4.14.0-041400-generic_4.14.0-041400.201711122031_amd64.deb</a></p>\n\n<p>wget <a href=\"http://kernel.ubuntu.com/~kernel-ppa/mainline/v4.14/linux-image-4.14.0-041400-generic_4.14.0-041400.201711122031_amd64.deb\">http://kernel.ubuntu.com/~kernel-ppa/mainline/v4.14/linux-image-4.14.0-041400-generic_4.14.0-041400.201711122031_amd64.deb</a></p>\n\n<p>sudo dpkg -i *.deb</p>\n\n<p>sudo reboot</p>\n\n<p>Before:</p>\n\n<p>$ uname -a</p>\n\n<p>Linux gpu-1 4.13.0-1006-gcp #9-Ubuntu SMP Mon Jan 8 21:13:15 UTC 2018 x86_64 x86_64 x86_64 GNU/Linux</p>\n\n<p>$ nvidia-smi</p>\n\n<p>NVIDIA-SMI has failed because it couldn't communicate with the NVIDIA driver. Make sure that the latest NVIDIA driver is installed and running.</p>\n\n<p>After:</p>\n\n<p>$ uname -a</p>\n\n<p>Linux gpu-1 4.14.0-041400-generic #201711122031 SMP Sun Nov 12 20:32:29 UTC 2017 x86_64 x86_64 x86_64 GNU/Linux</p>\n\n<p>$ nvidia-smi</p>\n\n<p>Fri Jan 12 07:17:44 2018 <br>\n+-----------------------------------------------------------------------------+\n| NVIDIA-SMI 387.26                 Driver Version: 387.26                    |\n|-------------------------------+----------------------+----------------------+\n| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |\n| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |\n|===============================+======================+======================|\n|   0  Tesla K80           Off  | 00000000:00:04.0 Off |                    0 |\n| N/A   31C    P8    26W / 149W |     16MiB / 11439MiB |      0%      Default |\n+-------------------------------+----------------------+----------------------+</p>\n\n<p>+-----------------------------------------------------------------------------+\n| Processes:                                                       GPU Memory |\n|  GPU       PID   Type   Process name                             Usage      |\n|=============================================================================|\n|    0      1639      G   /usr/lib/xorg/Xorg                            15MiB |\n+-----------------------------------------------------------------------------+</p>",
      "votes": null,
      "replies": [
        {
          "id": 267757,
          "author_name": "peterdekkers101",
          "author_url": "",
          "post_date": "01/12/2018 07:42:42",
          "content": "<p>Marco, thanks for sharing this solution. </p>\n\n<p>For me reinstalling the Nvidia drivers didn't solve the problem. But installing a new kernel also did the trick for me (running Ubuntu 16.04.3 LTS).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 267907,
          "author_name": "malr87",
          "author_url": "",
          "post_date": "01/12/2018 17:48:06",
          "content": "<p>You're welcome, Peter! Glad I could help.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 268531,
          "author_name": "serdata",
          "author_url": "",
          "post_date": "01/14/2018 19:20:09",
          "content": "<p>I had the same problem on Ubuntu 16.04.3 with K80. This solution worked for me, thank you for sharing.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 267788,
      "author_name": "nicomon",
      "author_url": "",
      "post_date": "01/12/2018 09:50:52",
      "content": "<p>I followed this <a href=\"http://www.python36.com/install-tensorflow141-gpu/\">guide</a> and it worked well (with updated nvidia drivers and CUDA9.1). Since TF 1.4 does not support directly cuda 9.1 the guide suggests to build tf from source. (There is also a fix in the comments for a missing lib)</p>",
      "votes": null,
      "replies": [
        {
          "id": 268573,
          "author_name": "ryanzhang",
          "author_url": "",
          "post_date": "01/14/2018 23:01:43",
          "content": "<p>I followed that guide. One thing I need to add is a symbolic link </p>\n\n<p>ln -s /usr/local/cuda/include/crt/math_functions.hpp /usr/local/cuda/include/math_functions.hpp</p>\n\n<p>Other than that the guide works great. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 267820,
      "author_name": "princerk",
      "author_url": "",
      "post_date": "01/12/2018 12:34:28",
      "content": "<p>I faced the same problem with ubuntu 16.04 LTS. I ended up booting a new machine with ubuntu 17.04</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 267824,
      "author_name": "ekonishi",
      "author_url": "",
      "post_date": "01/12/2018 12:59:27",
      "content": "<p>My team had the same problem, we were on Ubuntu 16.04 and Tesla K80. We think it's related to google deploying in the background a security patch for the kernel against meltdown/spectre attacks. The solution was to roll it back to its previous unsafe state.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 267920,
      "author_name": "jandjenter",
      "author_url": "",
      "post_date": "01/12/2018 18:51:01",
      "content": "<p>I had same issue, but just 'sudo apt-get install nvidia-384' (thanks to WeiWei) and rebooting the machine worked for me.\nBTW, K80 is much slower than P100 (1/6 speed in my experience), so I think P100 is more cost &amp; time effective for training purpose. Though GCP admin required some 'deposit' for increasing quota.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 267929,
      "author_name": "gte620v",
      "author_url": "",
      "post_date": "01/12/2018 19:53:10",
      "content": "<p>I had this problem on P100 GCP instances, but not on K80 instances.  Super frustrating.  I finally solved it by installing the latest 390 Nvidia drivers:</p>\n\n<p>sudo apt-get -y purge nvidia*</p>\n\n<p>sudo add-apt-repository ppa:graphics-drivers</p>\n\n<p>sudo apt-get -y update</p>\n\n<p>sudo apt-get -y install nvidia-390</p>\n\n<p>sudo reboot</p>\n\n<p>After you do that, you can't install cuda with apt-get because that will revert the driver.  instead you have to use the run scripts from nvidia <a href=\"https://developer.nvidia.com/cuda-80-ga2-download-archive\">https://developer.nvidia.com/cuda-80-ga2-download-archive</a> and follow the \"runfile\" directions.  That script will ask if you want to revert the nvidia drivers. say \"no\".</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "267287": "Hi everyone,\n\nI have been using a virtual machine on Google Cloud (Ubuntu 16.04 LTS) for a while now with a NVidia Tesla K80 without any problems. All of a sudden, I cannot use the the GPU anymore.\n\nI get the following error at training time:\n\n2018-01-11 00:23:57.759222: E tensorflow/stream_executor/cuda/cuda_driver.cc:406] failed call to cuInit: CUDA_ERROR_UNKNOWN\n2018-01-11 00:23:57.759293: I tensorflow/stream_executor/cuda/cuda_diagnostics.cc:145] kernel driver does not appear to be running on this host (gpu-1): /proc/driver/nvidia/version does not exist\n\nWhen I run nvidia-smi, I get the following error:\n\nNVIDIA-SMI has failed because it couldn't communicate with the NVIDIA driver. Make sure that the latest NVIDIA driver is installed and running.\n\nNothing has changed since last night.  I even tried reverting creating a new instance with a fully functional snapshot of the machine that I took  a few days ago with no luck.\n\nAnyone else having the same issue? Any ideas of what it could be?\n\nThanks!",
    "267339": "same, as many others I suppose. \nreinstall cuda+drivers https://cloud.google.com/compute/docs/gpus/add-gpus#install-gpu-driver\nand export LD_LIBRARY_PATH=/usr/local/cuda/lib64:/usr/lib/nvidia-387 into your .bashrc",
    "267532": "I reinstalled CUDA using apt-get and the error remains. Then I download the 'runfile(local)' from the NVIDIA website, use apt-get to install nvidia-384, and add the following to the ~/.bashrc :\n\n    export PATH=\"$PATH:/usr/local/cuda-8.0/bin\" \n    export LD_LIBRARY_PATH=\"/usr/local/cuda-8.0/lib64\"\n\nIt just works again.",
    "267707": "Thanks guys, I did not have the same luck. I suspect this has to do something with the latest kernel update due to Spectre/Meltdown. Switching to Centos 7 solved the problem in 5 minutes.",
    "267749": "Even though CentOS was working, I wanted to fix my Ubuntu. Installing a different kernel did it for me:\n\nmkdir newkernel\n\ncd newkernel\n\nwget http://kernel.ubuntu.com/~kernel-ppa/mainline/v4.14/linux-headers-4.14.0-041400_4.14.0-041400.201711122031_all.deb\n\nwget http://kernel.ubuntu.com/~kernel-ppa/mainline/v4.14/linux-headers-4.14.0-041400-generic_4.14.0-041400.201711122031_amd64.deb\n\nwget http://kernel.ubuntu.com/~kernel-ppa/mainline/v4.14/linux-image-4.14.0-041400-generic_4.14.0-041400.201711122031_amd64.deb\n\nsudo dpkg -i *.deb\n\nsudo reboot\n\n\nBefore:\n\n$ uname -a\n\nLinux gpu-1 4.13.0-1006-gcp #9-Ubuntu SMP Mon Jan 8 21:13:15 UTC 2018 x86_64 x86_64 x86_64 GNU/Linux\n\n$ nvidia-smi\n\nNVIDIA-SMI has failed because it couldn't communicate with the NVIDIA driver. Make sure that the latest NVIDIA driver is installed and running.\n\nAfter:\n\n$ uname -a\n\nLinux gpu-1 4.14.0-041400-generic #201711122031 SMP Sun Nov 12 20:32:29 UTC 2017 x86_64 x86_64 x86_64 GNU/Linux\n\n$ nvidia-smi\n\nFri Jan 12 07:17:44 2018       \n+-----------------------------------------------------------------------------+\n| NVIDIA-SMI 387.26                 Driver Version: 387.26                    |\n|-------------------------------+----------------------+----------------------+\n| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |\n| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |\n|===============================+======================+======================|\n|   0  Tesla K80           Off  | 00000000:00:04.0 Off |                    0 |\n| N/A   31C    P8    26W / 149W |     16MiB / 11439MiB |      0%      Default |\n+-------------------------------+----------------------+----------------------+\n                                                                               \n+-----------------------------------------------------------------------------+\n| Processes:                                                       GPU Memory |\n|  GPU       PID   Type   Process name                             Usage      |\n|=============================================================================|\n|    0      1639      G   /usr/lib/xorg/Xorg                            15MiB |\n+-----------------------------------------------------------------------------+",
    "267757": "Marco, thanks for sharing this solution. \n\nFor me reinstalling the Nvidia drivers didn't solve the problem. But installing a new kernel also did the trick for me (running Ubuntu 16.04.3 LTS).",
    "267788": "I followed this [guide][1] and it worked well (with updated nvidia drivers and CUDA9.1). Since TF 1.4 does not support directly cuda 9.1 the guide suggests to build tf from source. (There is also a fix in the comments for a missing lib)\n\n\n  [1]: http://www.python36.com/install-tensorflow141-gpu/",
    "267820": "I faced the same problem with ubuntu 16.04 LTS. I ended up booting a new machine with ubuntu 17.04",
    "267824": "My team had the same problem, we were on Ubuntu 16.04 and Tesla K80. We think it's related to google deploying in the background a security patch for the kernel against meltdown/spectre attacks. The solution was to roll it back to its previous unsafe state.",
    "267907": "You're welcome, Peter! Glad I could help.",
    "267920": "I had same issue, but just 'sudo apt-get install nvidia-384' (thanks to WeiWei) and rebooting the machine worked for me.\nBTW, K80 is much slower than P100 (1/6 speed in my experience), so I think P100 is more cost &amp; time effective for training purpose. Though GCP admin required some 'deposit' for increasing quota.",
    "267929": "I had this problem on P100 GCP instances, but not on K80 instances.  Super frustrating.  I finally solved it by installing the latest 390 Nvidia drivers:\n\n\nsudo apt-get -y purge nvidia*\n\nsudo add-apt-repository ppa:graphics-drivers\n\nsudo apt-get -y update\n\nsudo apt-get -y install nvidia-390\n\nsudo reboot\n\n\nAfter you do that, you can't install cuda with apt-get because that will revert the driver.  instead you have to use the run scripts from nvidia https://developer.nvidia.com/cuda-80-ga2-download-archive and follow the \"runfile\" directions.  That script will ask if you want to revert the nvidia drivers. say \"no\".",
    "268531": "I had the same problem on Ubuntu 16.04.3 with K80. This solution worked for me, thank you for sharing.",
    "268573": "I followed that guide. One thing I need to add is a symbolic link \n\nln -s /usr/local/cuda/include/crt/math_functions.hpp /usr/local/cuda/include/math_functions.hpp\n\nOther than that the guide works great."
  },
  "source": "meta"
}