{
  "id": 70425,
  "title": "GPU computing advice?",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/70425",
  "author_name": "Dawid Dabkowski",
  "post_date": "2018-11-03T13:45:35.733000",
  "votes": 1,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Hello Kagglers!</p>\n\n<p>I try to build my keras NN on Kaggle kernel servers. Recently I've turned on the GPU computing. But it gave me no more than 15% speedup. Am I doing something wrong? What can I do to get more of it, maybe some architectures benefit more? I'd love to get something like this 13x speedup posted\n<a href=\"https://www.kaggle.com/dansbecker/running-kaggle-kernels-with-a-gpu\">here</a>.</p>\n\n<p>Thank you for any advice!</p>",
  "messages": [
    {
      "id": 415360,
      "postDate": "2018-11-05T01:35:41.933Z",
      "content": "<p>I am using fastai with a GPU on linux that I \"built\" (actually just installed hardware in a box). My development platform is a MacBook pro. The basic on both, including ssh tunnel, are contained in:</p>\n\n<p><a href=\"https://blog.slavv.com/the-1700-great-deep-learning-box-assembly-setup-and-benchmarks-148c5ebe6415\">https://blog.slavv.com/the-1700-great-deep-learning-box-assembly-setup-and-benchmarks-148c5ebe6415</a></p>\n\n<p>This is all fine but recently I had a surprise in that I can run two models at once on one graphics card. One model in a jupyter notebook is run directly on the linux box - this is used to run long models. The second is run from the mac using the deployment option in Pycharm - this is used for testing.</p>\n\n<p>Currently, I'm running full size images in the notebook and testing on the mac with smaller images. nvidia-smi shows two processes on the GPU and both are undating displays, so I think it works.</p>\n\n<p>It was surprising because, prior to the Pycharm deployment, running two processes invariably crashed the second</p>\n\n<p>Does this jibe with others experience? Maybe it's just obvious.</p>",
      "rawMarkdown": "I am using fastai with a GPU on linux that I \"built\" (actually just installed hardware in a box). My development platform is a MacBook pro. The basic on both, including ssh tunnel, are contained in:\n\nhttps://blog.slavv.com/the-1700-great-deep-learning-box-assembly-setup-and-benchmarks-148c5ebe6415\n\nThis is all fine but recently I had a surprise in that I can run two models at once on one graphics card. One model in a jupyter notebook is run directly on the linux box - this is used to run long models. The second is run from the mac using the deployment option in Pycharm - this is used for testing.\n\nCurrently, I'm running full size images in the notebook and testing on the mac with smaller images. nvidia-smi shows two processes on the GPU and both are undating displays, so I think it works.\n\nIt was surprising because, prior to the Pycharm deployment, running two processes invariably crashed the second\n\nDoes this jibe with others experience? Maybe it's just obvious.\n",
      "votes": 1,
      "replies": [
        {
          "id": 416591,
          "postDate": "2018-11-06T23:37:22.087Z",
          "content": "<p>I think that from some version of cuda it is possible as long as you have enough memory on the card for both (which you usually don't as you try to have the batches as large as possible...) </p>",
          "rawMarkdown": "I think that from some version of cuda it is possible as long as you have enough memory on the card for both (which you usually don't as you try to have the batches as large as possible...) "
        },
        {
          "id": 417707,
          "postDate": "2018-11-08T17:02:45.970Z",
          "content": "<p>I've run a few models at once with Keras on windows with minimal issue. The only thing I had to do was reset the session to not take up all memory by default. My dev system is similar to your linux box. I've got an asus x99 deluxe, 64gb ram, i7 5930k, samsung 970 ssd, 1500W power supply using a Titan X maxwell and a gtx 970.</p>\n\n<pre>from keras import backend as K\nimport tensorflow as tf\nconfig = tf.ConfigProto()\nconfig.gpu_options.allow_growth=True\nsess = tf.Session(config=config)\nK.set_session(sess)\n</pre>",
          "rawMarkdown": "I've run a few models at once with Keras on windows with minimal issue. The only thing I had to do was reset the session to not take up all memory by default. My dev system is similar to your linux box. I've got an asus x99 deluxe, 64gb ram, i7 5930k, samsung 970 ssd, 1500W power supply using a Titan X maxwell and a gtx 970.\n\n<pre>from keras import backend as K\nimport tensorflow as tf\nconfig = tf.ConfigProto()\nconfig.gpu_options.allow_growth=True\nsess = tf.Session(config=config)\nK.set_session(sess)\n</pre>"
        },
        {
          "id": 417880,
          "postDate": "2018-11-08T23:31:25.327Z",
          "content": "<p>Wow 1500W. I thought my power supply was big. I plan at some point to get another graphics card - I hope it’s big enough.</p>",
          "rawMarkdown": "Wow 1500W. I thought my power supply was big. I plan at some point to get another graphics card - I hope it’s big enough."
        },
        {
          "id": 418005,
          "postDate": "2018-11-09T05:18:38.863Z",
          "content": "<p>I ran 2 cards with a 750W just fine for a long time. I want to be able to get 3 cards in here and got a good deal on this one, used not new. This X99 board can run 3 cards at x16, x16 and x8 and I've been thinking of picking up a 2070 next.</p>",
          "rawMarkdown": "I ran 2 cards with a 750W just fine for a long time. I want to be able to get 3 cards in here and got a good deal on this one, used not new. This X99 board can run 3 cards at x16, x16 and x8 and I've been thinking of picking up a 2070 next."
        }
      ]
    },
    {
      "id": 414862,
      "postDate": "2018-11-03T18:44:37.703Z",
      "content": "<p>If you have one GPU it should do it \"automatically\"</p>\n\n<p>I am running on Windows, here is my process:\n* Install and use Anaconda for python environment\n* install nvidia cuda tools 9.0, 9.2, 10.0\n* use anaconda, install tensorflow-gpu\n* install keras, all the rest</p>\n\n<p>If you have more than one GPU, before you import keras you can select one or multiple to use:</p>\n\n<pre>import os\nos.environ['CUDA_VISIBLE_DEVICES'] = '0'  #0 = first device, can be multiple: 0,1\n</pre>",
      "rawMarkdown": "If you have one GPU it should do it \"automatically\"\n\nI am running on Windows, here is my process:\n* Install and use Anaconda for python environment\n* install nvidia cuda tools 9.0, 9.2, 10.0\n* use anaconda, install tensorflow-gpu\n* install keras, all the rest\n\nIf you have more than one GPU, before you import keras you can select one or multiple to use:\n<pre>import os\nos.environ['CUDA_VISIBLE_DEVICES'] = '0'  #0 = first device, can be multiple: 0,1\n</pre>\n",
      "votes": 1,
      "replies": [
        {
          "id": 415052,
          "postDate": "2018-11-04T09:07:06.443Z",
          "content": "<p>Thanks for the reply. Correct me if I'm wrong but the computing is done on Kaggle servers, I thought that my pc configuration has nothing to do with it?</p>",
          "rawMarkdown": "Thanks for the reply. Correct me if I'm wrong but the computing is done on Kaggle servers, I thought that my pc configuration has nothing to do with it?"
        },
        {
          "id": 415241,
          "postDate": "2018-11-04T18:26:42.830Z",
          "content": "<p>My mistake,  I thought you were running locally. The online kernels were annoying me so I've switched to local only. </p>",
          "rawMarkdown": "My mistake,  I thought you were running locally. The online kernels were annoying me so I've switched to local only. "
        }
      ]
    },
    {
      "id": 414740,
      "postDate": "2018-11-03T13:45:35.733Z",
      "content": "<p>Hello Kagglers!</p>\n\n<p>I try to build my keras NN on Kaggle kernel servers. Recently I've turned on the GPU computing. But it gave me no more than 15% speedup. Am I doing something wrong? What can I do to get more of it, maybe some architectures benefit more? I'd love to get something like this 13x speedup posted\n<a href=\"https://www.kaggle.com/dansbecker/running-kaggle-kernels-with-a-gpu\">here</a>.</p>\n\n<p>Thank you for any advice!</p>",
      "rawMarkdown": "Hello Kagglers!\n\nI try to build my keras NN on Kaggle kernel servers. Recently I've turned on the GPU computing. But it gave me no more than 15% speedup. Am I doing something wrong? What can I do to get more of it, maybe some architectures benefit more? I'd love to get something like this 13x speedup posted\n[here][1].\n\nThank you for any advice!\n\n\n  [1]: https://www.kaggle.com/dansbecker/running-kaggle-kernels-with-a-gpu",
      "votes": 1
    },
    {
      "id": 415313,
      "postDate": "2018-11-04T22:09:55.943Z",
      "content": "<p>I can notice a great speed increase with Keras and Kaggle GPU. </p>\n\n<p>Which library are you using? Notice that if you're using recurrent layers, they will not get very fast even with the GPU.</p>",
      "rawMarkdown": "I can notice a great speed increase with Keras and Kaggle GPU. \n\nWhich library are you using? Notice that if you're using recurrent layers, they will not get very fast even with the GPU.",
      "replies": [
        {
          "id": 417640,
          "postDate": "2018-11-08T15:06:56.207Z",
          "content": "<p>Thanks for a response. I use keras with some convolutions, dense layers etc (no recurrent ones).</p>",
          "rawMarkdown": "Thanks for a response. I use keras with some convolutions, dense layers etc (no recurrent ones)."
        },
        {
          "id": 417648,
          "postDate": "2018-11-08T15:25:08.680Z",
          "content": "<p>Are you using a generator? It may be holding the model speed. </p>",
          "rawMarkdown": "Are you using a generator? It may be holding the model speed. "
        }
      ]
    },
    {
      "id": 414743,
      "postDate": "2018-11-03T14:13:33.697Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 414756,
          "postDate": "2018-11-03T14:40:20.230Z",
          "content": "<p>No, I did not, neither did author of the cited kernel. Could you explain further?</p>",
          "rawMarkdown": "No, I did not, neither did author of the cited kernel. Could you explain further?"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 415360,
      "author_name": "pete",
      "author_url": "",
      "post_date": "2018-11-05T01:35:41.933000",
      "content": "<p>I am using fastai with a GPU on linux that I \"built\" (actually just installed hardware in a box). My development platform is a MacBook pro. The basic on both, including ssh tunnel, are contained in:</p>\n\n<p><a href=\"https://blog.slavv.com/the-1700-great-deep-learning-box-assembly-setup-and-benchmarks-148c5ebe6415\">https://blog.slavv.com/the-1700-great-deep-learning-box-assembly-setup-and-benchmarks-148c5ebe6415</a></p>\n\n<p>This is all fine but recently I had a surprise in that I can run two models at once on one graphics card. One model in a jupyter notebook is run directly on the linux box - this is used to run long models. The second is run from the mac using the deployment option in Pycharm - this is used for testing.</p>\n\n<p>Currently, I'm running full size images in the notebook and testing on the mac with smaller images. nvidia-smi shows two processes on the GPU and both are undating displays, so I think it works.</p>\n\n<p>It was surprising because, prior to the Pycharm deployment, running two processes invariably crashed the second</p>\n\n<p>Does this jibe with others experience? Maybe it's just obvious.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 416591,
          "author_name": "Moshel",
          "author_url": "",
          "post_date": "2018-11-06T23:37:22.087000",
          "content": "<p>I think that from some version of cuda it is possible as long as you have enough memory on the card for both (which you usually don't as you try to have the batches as large as possible...) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 417707,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-11-08T17:02:45.970000",
          "content": "<p>I've run a few models at once with Keras on windows with minimal issue. The only thing I had to do was reset the session to not take up all memory by default. My dev system is similar to your linux box. I've got an asus x99 deluxe, 64gb ram, i7 5930k, samsung 970 ssd, 1500W power supply using a Titan X maxwell and a gtx 970.</p>\n\n<pre>from keras import backend as K\nimport tensorflow as tf\nconfig = tf.ConfigProto()\nconfig.gpu_options.allow_growth=True\nsess = tf.Session(config=config)\nK.set_session(sess)\n</pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 417880,
          "author_name": "pete",
          "author_url": "",
          "post_date": "2018-11-08T23:31:25.327000",
          "content": "<p>Wow 1500W. I thought my power supply was big. I plan at some point to get another graphics card - I hope it’s big enough.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 418005,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-11-09T05:18:38.863000",
          "content": "<p>I ran 2 cards with a 750W just fine for a long time. I want to be able to get 3 cards in here and got a good deal on this one, used not new. This X99 board can run 3 cards at x16, x16 and x8 and I've been thinking of picking up a 2070 next.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 414862,
      "author_name": "Brian",
      "author_url": "",
      "post_date": "2018-11-03T18:44:37.703000",
      "content": "<p>If you have one GPU it should do it \"automatically\"</p>\n\n<p>I am running on Windows, here is my process:\n* Install and use Anaconda for python environment\n* install nvidia cuda tools 9.0, 9.2, 10.0\n* use anaconda, install tensorflow-gpu\n* install keras, all the rest</p>\n\n<p>If you have more than one GPU, before you import keras you can select one or multiple to use:</p>\n\n<pre>import os\nos.environ['CUDA_VISIBLE_DEVICES'] = '0'  #0 = first device, can be multiple: 0,1\n</pre>",
      "votes": 1,
      "replies": [
        {
          "id": 415052,
          "author_name": "Dawid Dabkowski",
          "author_url": "",
          "post_date": "2018-11-04T09:07:06.443000",
          "content": "<p>Thanks for the reply. Correct me if I'm wrong but the computing is done on Kaggle servers, I thought that my pc configuration has nothing to do with it?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 415241,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-11-04T18:26:42.830000",
          "content": "<p>My mistake,  I thought you were running locally. The online kernels were annoying me so I've switched to local only. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 415313,
      "author_name": "Daniel Möller",
      "author_url": "",
      "post_date": "2018-11-04T22:09:55.943000",
      "content": "<p>I can notice a great speed increase with Keras and Kaggle GPU. </p>\n\n<p>Which library are you using? Notice that if you're using recurrent layers, they will not get very fast even with the GPU.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 417640,
          "author_name": "Dawid Dabkowski",
          "author_url": "",
          "post_date": "2018-11-08T15:06:56.207000",
          "content": "<p>Thanks for a response. I use keras with some convolutions, dense layers etc (no recurrent ones).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 417648,
          "author_name": "Daniel Möller",
          "author_url": "",
          "post_date": "2018-11-08T15:25:08.680000",
          "content": "<p>Are you using a generator? It may be holding the model speed. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 414743,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-03T14:13:33.697000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 414756,
          "author_name": "Dawid Dabkowski",
          "author_url": "",
          "post_date": "2018-11-03T14:40:20.230000",
          "content": "<p>No, I did not, neither did author of the cited kernel. Could you explain further?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "415360": "I am using fastai with a GPU on linux that I \"built\" (actually just installed hardware in a box). My development platform is a MacBook pro. The basic on both, including ssh tunnel, are contained in:\n\nhttps://blog.slavv.com/the-1700-great-deep-learning-box-assembly-setup-and-benchmarks-148c5ebe6415\n\nThis is all fine but recently I had a surprise in that I can run two models at once on one graphics card. One model in a jupyter notebook is run directly on the linux box - this is used to run long models. The second is run from the mac using the deployment option in Pycharm - this is used for testing.\n\nCurrently, I'm running full size images in the notebook and testing on the mac with smaller images. nvidia-smi shows two processes on the GPU and both are undating displays, so I think it works.\n\nIt was surprising because, prior to the Pycharm deployment, running two processes invariably crashed the second\n\nDoes this jibe with others experience? Maybe it's just obvious.\n",
    "414862": "If you have one GPU it should do it \"automatically\"\n\nI am running on Windows, here is my process:\n* Install and use Anaconda for python environment\n* install nvidia cuda tools 9.0, 9.2, 10.0\n* use anaconda, install tensorflow-gpu\n* install keras, all the rest\n\nIf you have more than one GPU, before you import keras you can select one or multiple to use:\n<pre>import os\nos.environ['CUDA_VISIBLE_DEVICES'] = '0'  #0 = first device, can be multiple: 0,1\n</pre>\n",
    "414740": "Hello Kagglers!\n\nI try to build my keras NN on Kaggle kernel servers. Recently I've turned on the GPU computing. But it gave me no more than 15% speedup. Am I doing something wrong? What can I do to get more of it, maybe some architectures benefit more? I'd love to get something like this 13x speedup posted\n[here][1].\n\nThank you for any advice!\n\n\n  [1]: https://www.kaggle.com/dansbecker/running-kaggle-kernels-with-a-gpu",
    "415313": "I can notice a great speed increase with Keras and Kaggle GPU. \n\nWhich library are you using? Notice that if you're using recurrent layers, they will not get very fast even with the GPU.",
    "414743": ""
  }
}