{
  "id": 37087,
  "title": "Configuring GPUs in Cloud Machine Learning Engine",
  "url": "/competitions/passenger-screening-algorithm-challenge/discussion/37087",
  "author_name": "",
  "post_date": "2017-07-26T23:15:02.061751700Z",
  "votes": 9,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Running <a href=\"https://cloud.google.com/ml-engine/docs/how-tos/using-gpus\">Cloud Machine Learning Engine (CMLE) with GPU</a> resources is as simple as using one of the GPU-enabled machine types or scale tiers in your yaml configuration file that you pass into your training job. These configuration files usually use either predefined \"scale tiers\", or define a custom scale tier with a custom set of machines, chosen by you. Let's look at each of these in turn, and why you might want to use them. </p>\n\n<h1>GPUs by configuration</h1>\n\n<p>When first starting out and experimenting with GPUs, you can use the scale tier <code>BASIC_GPU</code>, which is a single worker machine with a GPU (basically the same as a single machine with a GPU attached). You will also want to only use a small subset of your data, since the purpose of this configuration is to set up your workflow. Your config.yaml might look something like this:</p>\n\n<pre><code>trainingInput:\n  region: us-central1\n  scaleTier: BASIC_GPU\n</code></pre>\n\n<p>Once you have your workflow down, you can graduate to <a href=\"https://cloud.google.com/ml-engine/docs/how-tos/using-gpus#requesting_gpu-enabled_machines\">more powerful machine types</a> using a custom scale tier: </p>\n\n<ul>\n<li><code>standard_gpu</code>: single GPU </li>\n<li><code>complex_model_m_gpu</code>: four GPUs</li>\n<li><code>complex_model_l_gpu</code>: eight GPUs</li>\n</ul>\n\n<p>Keep in mind that when you are configuring your infrastructure, each of these is <em>one machine</em>, so the complex_model_m_gpu configuration is four GPUs attached to the same machine. In a distributed system, you might have several of these machines acting as worker machines, along with some parameter servers. \nSuch a <code>config.yaml</code> file might look something like this:</p>\n\n<pre><code>trainingInput:\n  region: us-central1\n  scaleTier: CUSTOM\n  masterType: complex_model_m\n  workerType: complex_model_m_gpu\n  parameterServerType: large_model\n  workerCount: 4\n  parameterServerCount: 2\n</code></pre>\n\n<p>Notice that only the workers have GPUs attached -- the other machine types do not need GPUs attached. </p>\n\n<p><img src=\"https://storage.googleapis.com/image-uploader/kaggle_posts/guide1.png\" alt=\"distributed architecture diagram\" title=\"\"></p>\n\n<h1>Understanding ML Units and Quota</h1>\n\n<p>Different machines consume differing amounts of \"ML Units\". You should go to <a href=\"https://cloud.google.com/ml-engine/docs/concepts/training-overview#scale_tier\">this page</a> detailing each type of machine and how many ML Units it consumes. The current limit on concurrent training units is 25 ML Units. </p>\n\n<p>To compute how many ML Units we used in the example above, we first look up how many units each of the machine types we used account for. The master, <code>complex_model_m</code>, is 3 ML Units. The worker, a <code>complex_model_m_gpu</code>, is 12 ML Units. Finally, the parameter server, a <code>large_model</code>, is 3 ML Units. \nSo our total count is 1x <code>complex_model_m</code> (3 ML Units) + 4x <code>complex_model_m_gpu</code> (12 ML Units) + 2x <code>large_model</code> (3 ML Units) = 3 + 48 + 6 = 57 ML Units! </p>\n\n<p><img src=\"https://storage.googleapis.com/image-uploader/kaggle_posts/mlunittable.png\" alt=\"table of ML Unit examples\" title=\"\"></p>\n\n<p>If we had used <code>standard_gpu</code> instead of <code>complex_model_m_gpu</code>, we would have dramatically reduced our ML Unit usage (it would also cut our GPU count by a factor of 4), since the <code>standard_gpu</code> is only 3 ML Units (which when compared to the ML Unit values of some of the other machine types really makes it quite a good value!). It would have reduced our consumption to only 21 ML Units.</p>\n\n<p>Let's say you want to run that big job, with 57 ML Units consumed. You'd need to increase your quota. However, the quota for ML Units is not in the cloud console's IAM page, so you should head on over to the <a href=\"https://docs.google.com/a/google.com/forms/d/e/1FAIpQLScjm2W208wdf0iyQxBIESZ2_UvTS9xVPKleTQg9r_Uc1uPACQ/viewform\">CMLE Quota Request Form</a> and detail your needs there.</p>\n\n<p>Current quota limits can be found on our <a href=\"https://cloud.google.com/ml-engine/quotas#limits_on_concurrent_ml_training_units_and_prediction_nodeshttp://\">Quotas page</a>. Another limit to be aware of is the number of concurrent GPUs, which is 10 by default, so if you need more than 10 GPUs, be sure to request it using the <a href=\"https://docs.google.com/a/google.com/forms/d/e/1FAIpQLScjm2W208wdf0iyQxBIESZ2_UvTS9xVPKleTQg9r_Uc1uPACQ/viewform\">Quota Request Form</a>. In the example above, we had 4x complex_model_m_gpu machines, which each contain 4 GPUs, totaling 16 GPUs, so it would definitely need a quota increase for GPUs as well.</p>\n\n<h1>Accessing GPUs in TensorFlow code</h1>\n\n<p>The specifics of how you go about accessing GPUs in TensorFlow are a bit out of scope of this guide, but broadly speaking, you have two options: using the high-level APIs (Estimator) or using the low-level APIs. The \"Assigning ops to GPUs\" section of <a href=\"https://cloud.google.com/ml-engine/docs/how-tos/using-gpus#assigning_ops_to_gpus\">this article</a> goes into more detail.</p>\n\n<h1>Additional Resources and Reading</h1>\n\n<p><a href=\"https://www.tensorflow.org/api_docs/python/tf/train/ClusterSpec\">TensorFlow ClusterSpec</a>: This will be useful for setting up your training to access your GPU resources</p>\n\n<p>Using GPUs On Compute Engine\n<a href=\"https://cloud.google.com/compute/docs/gpus/\">https://cloud.google.com/compute/docs/gpus/</a> </p>\n\n<p>Documentation for using GPUs On Cloud Machine Learning Engine\n<a href=\"https://cloud.google.com/ml-engine/docs/how-tos/using-gpus\">https://cloud.google.com/ml-engine/docs/how-tos/using-gpus</a></p>\n\n<p>Reading in the TSA data:\nWill has some start code for reading in the images here: <a href=\"https://www.kaggle.com/wcukierski/reading-images\">https://www.kaggle.com/wcukierski/reading-images</a></p>",
  "messages": [
    {
      "id": "207580",
      "postDate": "07/26/2017 23:15:02",
      "content": "<p>Running <a href=\"https://cloud.google.com/ml-engine/docs/how-tos/using-gpus\">Cloud Machine Learning Engine (CMLE) with GPU</a> resources is as simple as using one of the GPU-enabled machine types or scale tiers in your yaml configuration file that you pass into your training job. These configuration files usually use either predefined \"scale tiers\", or define a custom scale tier with a custom set of machines, chosen by you. Let's look at each of these in turn, and why you might want to use them. </p>\n\n<h1>GPUs by configuration</h1>\n\n<p>When first starting out and experimenting with GPUs, you can use the scale tier <code>BASIC_GPU</code>, which is a single worker machine with a GPU (basically the same as a single machine with a GPU attached). You will also want to only use a small subset of your data, since the purpose of this configuration is to set up your workflow. Your config.yaml might look something like this:</p>\n\n<pre><code>trainingInput:\n  region: us-central1\n  scaleTier: BASIC_GPU\n</code></pre>\n\n<p>Once you have your workflow down, you can graduate to <a href=\"https://cloud.google.com/ml-engine/docs/how-tos/using-gpus#requesting_gpu-enabled_machines\">more powerful machine types</a> using a custom scale tier: </p>\n\n<ul>\n<li><code>standard_gpu</code>: single GPU </li>\n<li><code>complex_model_m_gpu</code>: four GPUs</li>\n<li><code>complex_model_l_gpu</code>: eight GPUs</li>\n</ul>\n\n<p>Keep in mind that when you are configuring your infrastructure, each of these is <em>one machine</em>, so the complex_model_m_gpu configuration is four GPUs attached to the same machine. In a distributed system, you might have several of these machines acting as worker machines, along with some parameter servers. \nSuch a <code>config.yaml</code> file might look something like this:</p>\n\n<pre><code>trainingInput:\n  region: us-central1\n  scaleTier: CUSTOM\n  masterType: complex_model_m\n  workerType: complex_model_m_gpu\n  parameterServerType: large_model\n  workerCount: 4\n  parameterServerCount: 2\n</code></pre>\n\n<p>Notice that only the workers have GPUs attached -- the other machine types do not need GPUs attached. </p>\n\n<p><img src=\"https://storage.googleapis.com/image-uploader/kaggle_posts/guide1.png\" alt=\"distributed architecture diagram\" title=\"\"></p>\n\n<h1>Understanding ML Units and Quota</h1>\n\n<p>Different machines consume differing amounts of \"ML Units\". You should go to <a href=\"https://cloud.google.com/ml-engine/docs/concepts/training-overview#scale_tier\">this page</a> detailing each type of machine and how many ML Units it consumes. The current limit on concurrent training units is 25 ML Units. </p>\n\n<p>To compute how many ML Units we used in the example above, we first look up how many units each of the machine types we used account for. The master, <code>complex_model_m</code>, is 3 ML Units. The worker, a <code>complex_model_m_gpu</code>, is 12 ML Units. Finally, the parameter server, a <code>large_model</code>, is 3 ML Units. \nSo our total count is 1x <code>complex_model_m</code> (3 ML Units) + 4x <code>complex_model_m_gpu</code> (12 ML Units) + 2x <code>large_model</code> (3 ML Units) = 3 + 48 + 6 = 57 ML Units! </p>\n\n<p><img src=\"https://storage.googleapis.com/image-uploader/kaggle_posts/mlunittable.png\" alt=\"table of ML Unit examples\" title=\"\"></p>\n\n<p>If we had used <code>standard_gpu</code> instead of <code>complex_model_m_gpu</code>, we would have dramatically reduced our ML Unit usage (it would also cut our GPU count by a factor of 4), since the <code>standard_gpu</code> is only 3 ML Units (which when compared to the ML Unit values of some of the other machine types really makes it quite a good value!). It would have reduced our consumption to only 21 ML Units.</p>\n\n<p>Let's say you want to run that big job, with 57 ML Units consumed. You'd need to increase your quota. However, the quota for ML Units is not in the cloud console's IAM page, so you should head on over to the <a href=\"https://docs.google.com/a/google.com/forms/d/e/1FAIpQLScjm2W208wdf0iyQxBIESZ2_UvTS9xVPKleTQg9r_Uc1uPACQ/viewform\">CMLE Quota Request Form</a> and detail your needs there.</p>\n\n<p>Current quota limits can be found on our <a href=\"https://cloud.google.com/ml-engine/quotas#limits_on_concurrent_ml_training_units_and_prediction_nodeshttp://\">Quotas page</a>. Another limit to be aware of is the number of concurrent GPUs, which is 10 by default, so if you need more than 10 GPUs, be sure to request it using the <a href=\"https://docs.google.com/a/google.com/forms/d/e/1FAIpQLScjm2W208wdf0iyQxBIESZ2_UvTS9xVPKleTQg9r_Uc1uPACQ/viewform\">Quota Request Form</a>. In the example above, we had 4x complex_model_m_gpu machines, which each contain 4 GPUs, totaling 16 GPUs, so it would definitely need a quota increase for GPUs as well.</p>\n\n<h1>Accessing GPUs in TensorFlow code</h1>\n\n<p>The specifics of how you go about accessing GPUs in TensorFlow are a bit out of scope of this guide, but broadly speaking, you have two options: using the high-level APIs (Estimator) or using the low-level APIs. The \"Assigning ops to GPUs\" section of <a href=\"https://cloud.google.com/ml-engine/docs/how-tos/using-gpus#assigning_ops_to_gpus\">this article</a> goes into more detail.</p>\n\n<h1>Additional Resources and Reading</h1>\n\n<p><a href=\"https://www.tensorflow.org/api_docs/python/tf/train/ClusterSpec\">TensorFlow ClusterSpec</a>: This will be useful for setting up your training to access your GPU resources</p>\n\n<p>Using GPUs On Compute Engine\n<a href=\"https://cloud.google.com/compute/docs/gpus/\">https://cloud.google.com/compute/docs/gpus/</a> </p>\n\n<p>Documentation for using GPUs On Cloud Machine Learning Engine\n<a href=\"https://cloud.google.com/ml-engine/docs/how-tos/using-gpus\">https://cloud.google.com/ml-engine/docs/how-tos/using-gpus</a></p>\n\n<p>Reading in the TSA data:\nWill has some start code for reading in the images here: <a href=\"https://www.kaggle.com/wcukierski/reading-images\">https://www.kaggle.com/wcukierski/reading-images</a></p>",
      "rawMarkdown": "Running [Cloud Machine Learning Engine (CMLE) with GPU][1] resources is as simple as using one of the GPU-enabled machine types or scale tiers in your yaml configuration file that you pass into your training job. These configuration files usually use either predefined \"scale tiers\", or define a custom scale tier with a custom set of machines, chosen by you. Let's look at each of these in turn, and why you might want to use them. \n\nGPUs by configuration\n================================\n\nWhen first starting out and experimenting with GPUs, you can use the scale tier `BASIC_GPU`, which is a single worker machine with a GPU (basically the same as a single machine with a GPU attached). You will also want to only use a small subset of your data, since the purpose of this configuration is to set up your workflow. Your config.yaml might look something like this:\n\n    trainingInput:\n      region: us-central1\n      scaleTier: BASIC_GPU\n\nOnce you have your workflow down, you can graduate to [more powerful machine types][2] using a custom scale tier: \n\n - `standard_gpu`: single GPU \n - `complex_model_m_gpu`: four GPUs\n - `complex_model_l_gpu`: eight GPUs\n\nKeep in mind that when you are configuring your infrastructure, each of these is *one machine*, so the complex_model_m_gpu configuration is four GPUs attached to the same machine. In a distributed system, you might have several of these machines acting as worker machines, along with some parameter servers. \nSuch a `config.yaml` file might look something like this:\n\n    trainingInput:\n      region: us-central1\n      scaleTier: CUSTOM\n      masterType: complex_model_m\n      workerType: complex_model_m_gpu\n      parameterServerType: large_model\n      workerCount: 4\n      parameterServerCount: 2\n\nNotice that only the workers have GPUs attached -- the other machine types do not need GPUs attached. \n\n![distributed architecture diagram][3]\n\nUnderstanding ML Units and Quota\n================================\n\nDifferent machines consume differing amounts of \"ML Units\". You should go to [this page][4] detailing each type of machine and how many ML Units it consumes. The current limit on concurrent training units is 25 ML Units. \n\nTo compute how many ML Units we used in the example above, we first look up how many units each of the machine types we used account for. The master, `complex_model_m`, is 3 ML Units. The worker, a `complex_model_m_gpu`, is 12 ML Units. Finally, the parameter server, a `large_model`, is 3 ML Units. \nSo our total count is 1x `complex_model_m` (3 ML Units) + 4x `complex_model_m_gpu` (12 ML Units) + 2x `large_model` (3 ML Units) = 3 + 48 + 6 = 57 ML Units! \n\n![table of ML Unit examples][5]\n\nIf we had used `standard_gpu` instead of `complex_model_m_gpu`, we would have dramatically reduced our ML Unit usage (it would also cut our GPU count by a factor of 4), since the `standard_gpu` is only 3 ML Units (which when compared to the ML Unit values of some of the other machine types really makes it quite a good value!). It would have reduced our consumption to only 21 ML Units.\n\nLet's say you want to run that big job, with 57 ML Units consumed. You'd need to increase your quota. However, the quota for ML Units is not in the cloud console's IAM page, so you should head on over to the [CMLE Quota Request Form][6] and detail your needs there.\n\nCurrent quota limits can be found on our [Quotas page][7]. Another limit to be aware of is the number of concurrent GPUs, which is 10 by default, so if you need more than 10 GPUs, be sure to request it using the [Quota Request Form][8]. In the example above, we had 4x complex_model_m_gpu machines, which each contain 4 GPUs, totaling 16 GPUs, so it would definitely need a quota increase for GPUs as well.\n\nAccessing GPUs in TensorFlow code\n===\n\nThe specifics of how you go about accessing GPUs in TensorFlow are a bit out of scope of this guide, but broadly speaking, you have two options: using the high-level APIs (Estimator) or using the low-level APIs. The \"Assigning ops to GPUs\" section of [this article][9] goes into more detail.\n\nAdditional Resources and Reading\n===\n\n[TensorFlow ClusterSpec][10]: This will be useful for setting up your training to access your GPU resources\n\nUsing GPUs On Compute Engine\nhttps://cloud.google.com/compute/docs/gpus/ \n\nDocumentation for using GPUs On Cloud Machine Learning Engine\nhttps://cloud.google.com/ml-engine/docs/how-tos/using-gpus\n\nReading in the TSA data:\nWill has some start code for reading in the images here: https://www.kaggle.com/wcukierski/reading-images\n\n\n  [1]: https://cloud.google.com/ml-engine/docs/how-tos/using-gpus\n  [2]: https://cloud.google.com/ml-engine/docs/how-tos/using-gpus#requesting_gpu-enabled_machines\n  [3]: https://storage.googleapis.com/image-uploader/kaggle_posts/guide1.png\n  [4]: https://cloud.google.com/ml-engine/docs/concepts/training-overview#scale_tier\n  [5]: https://storage.googleapis.com/image-uploader/kaggle_posts/mlunittable.png\n  [6]: https://docs.google.com/a/google.com/forms/d/e/1FAIpQLScjm2W208wdf0iyQxBIESZ2_UvTS9xVPKleTQg9r_Uc1uPACQ/viewform\n  [7]: https://cloud.google.com/ml-engine/quotas#limits_on_concurrent_ml_training_units_and_prediction_nodeshttp://\n  [8]: https://docs.google.com/a/google.com/forms/d/e/1FAIpQLScjm2W208wdf0iyQxBIESZ2_UvTS9xVPKleTQg9r_Uc1uPACQ/viewform\n  [9]: https://cloud.google.com/ml-engine/docs/how-tos/using-gpus#assigning_ops_to_gpus\n  [10]: https://www.tensorflow.org/api_docs/python/tf/train/ClusterSpec",
      "votes": null
    },
    {
      "id": "233695",
      "postDate": "10/20/2017 21:24:23",
      "content": "<p>Could you please mention how much does it cost to use various GPU configuration?</p>",
      "rawMarkdown": "Could you please mention how much does it cost to use various GPU configuration?",
      "votes": null
    },
    {
      "id": "234848",
      "postDate": "10/24/2017 07:33:13",
      "content": "<p>See <a href=\"https://cloud.google.com/ml-engine/docs/training-overview#comparing_machine_types\">this table</a> with all the machine types, and use the number of \"ML units\" to compute <a href=\"https://cloud.google.com/ml-engine/pricing\">pricing</a>.\nThe quick version is that as of this writing, it costs $0.49 per hour, per ML training unit, for training in the US, and $0.54 in Europe/Asia. </p>",
      "rawMarkdown": "See [this table](https://cloud.google.com/ml-engine/docs/training-overview#comparing_machine_types) with all the machine types, and use the number of \"ML units\" to compute [pricing](https://cloud.google.com/ml-engine/pricing).\nThe quick version is that as of this writing, it costs $0.49 per hour, per ML training unit, for training in the US, and $0.54 in Europe/Asia.",
      "votes": null
    },
    {
      "id": "253648",
      "postDate": "12/05/2017 11:37:18",
      "content": "<p>How do I run this on google cloud ?</p>\n\n<pre><code>import tensorflow as tf;\nhello = tf.constant('Hello, TensorFlow!');\nsess = tf.Session();\nprint(sess.run(hello));\na = tf.constant([1.0, 2.0, 3.0, 4.0], shape=[2, 2], name='a');\nc = tf.matmul(a,a)\nwith tf.device('/device:GPU:0'):\n   print(sess.run(c))\n</code></pre>",
      "rawMarkdown": "How do I run this on google cloud ?\n\n    import tensorflow as tf;\n    hello = tf.constant('Hello, TensorFlow!');\n    sess = tf.Session();\n    print(sess.run(hello));\n    a = tf.constant([1.0, 2.0, 3.0, 4.0], shape=[2, 2], name='a');\n    c = tf.matmul(a,a)\n    with tf.device('/device:GPU:0'):\n       print(sess.run(c))",
      "votes": null
    },
    {
      "id": "253783",
      "postDate": "12/05/2017 16:42:55",
      "content": "<p>I put the above code in a hi.py file, add the app.yaml below, run <code>gcloud app deploy</code> but get error:  hi.py not found.</p>\n\n<pre><code>runtime: python\nenv: flex\nentrypoint: hi.py\nruntime_config:\n  python_version: 3.5\nmanual_scaling:\n  instances: 1\nresources:\n  cpu: 8\n  memory_gb: 51.6\n  disk_size_gb: 100\nreadiness_check:\napp_start_timeout_sec: 3600\n</code></pre>",
      "rawMarkdown": "I put the above code in a hi.py file, add the app.yaml below, run `gcloud app deploy` but get error:  hi.py not found.\n\n    runtime: python\n    env: flex\n    entrypoint: hi.py\n    runtime_config:\n      python_version: 3.5\n    manual_scaling:\n      instances: 1\n    resources:\n      cpu: 8\n      memory_gb: 51.6\n      disk_size_gb: 100\n    readiness_check:\n    app_start_timeout_sec: 3600",
      "votes": null
    },
    {
      "id": "254149",
      "postDate": "12/06/2017 09:57:49",
      "content": "<p>I found this example, that works:  <a href=\"https://cloud.google.com/ml-engine/docs/command-line\">https://cloud.google.com/ml-engine/docs/command-line</a></p>",
      "rawMarkdown": "I found this example, that works:  https://cloud.google.com/ml-engine/docs/command-line",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 233695,
      "author_name": "nikhilper",
      "author_url": "",
      "post_date": "10/20/2017 21:24:23",
      "content": "<p>Could you please mention how much does it cost to use various GPU configuration?</p>",
      "votes": null,
      "replies": [
        {
          "id": 234848,
          "author_name": "yufengg",
          "author_url": "",
          "post_date": "10/24/2017 07:33:13",
          "content": "<p>See <a href=\"https://cloud.google.com/ml-engine/docs/training-overview#comparing_machine_types\">this table</a> with all the machine types, and use the number of \"ML units\" to compute <a href=\"https://cloud.google.com/ml-engine/pricing\">pricing</a>.\nThe quick version is that as of this writing, it costs $0.49 per hour, per ML training unit, for training in the US, and $0.54 in Europe/Asia. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 253648,
      "author_name": "jduffy",
      "author_url": "",
      "post_date": "12/05/2017 11:37:18",
      "content": "<p>How do I run this on google cloud ?</p>\n\n<pre><code>import tensorflow as tf;\nhello = tf.constant('Hello, TensorFlow!');\nsess = tf.Session();\nprint(sess.run(hello));\na = tf.constant([1.0, 2.0, 3.0, 4.0], shape=[2, 2], name='a');\nc = tf.matmul(a,a)\nwith tf.device('/device:GPU:0'):\n   print(sess.run(c))\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 253783,
          "author_name": "jduffy",
          "author_url": "",
          "post_date": "12/05/2017 16:42:55",
          "content": "<p>I put the above code in a hi.py file, add the app.yaml below, run <code>gcloud app deploy</code> but get error:  hi.py not found.</p>\n\n<pre><code>runtime: python\nenv: flex\nentrypoint: hi.py\nruntime_config:\n  python_version: 3.5\nmanual_scaling:\n  instances: 1\nresources:\n  cpu: 8\n  memory_gb: 51.6\n  disk_size_gb: 100\nreadiness_check:\napp_start_timeout_sec: 3600\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 254149,
          "author_name": "jduffy",
          "author_url": "",
          "post_date": "12/06/2017 09:57:49",
          "content": "<p>I found this example, that works:  <a href=\"https://cloud.google.com/ml-engine/docs/command-line\">https://cloud.google.com/ml-engine/docs/command-line</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "207580": "Running [Cloud Machine Learning Engine (CMLE) with GPU][1] resources is as simple as using one of the GPU-enabled machine types or scale tiers in your yaml configuration file that you pass into your training job. These configuration files usually use either predefined \"scale tiers\", or define a custom scale tier with a custom set of machines, chosen by you. Let's look at each of these in turn, and why you might want to use them. \n\nGPUs by configuration\n================================\n\nWhen first starting out and experimenting with GPUs, you can use the scale tier `BASIC_GPU`, which is a single worker machine with a GPU (basically the same as a single machine with a GPU attached). You will also want to only use a small subset of your data, since the purpose of this configuration is to set up your workflow. Your config.yaml might look something like this:\n\n    trainingInput:\n      region: us-central1\n      scaleTier: BASIC_GPU\n\nOnce you have your workflow down, you can graduate to [more powerful machine types][2] using a custom scale tier: \n\n - `standard_gpu`: single GPU \n - `complex_model_m_gpu`: four GPUs\n - `complex_model_l_gpu`: eight GPUs\n\nKeep in mind that when you are configuring your infrastructure, each of these is *one machine*, so the complex_model_m_gpu configuration is four GPUs attached to the same machine. In a distributed system, you might have several of these machines acting as worker machines, along with some parameter servers. \nSuch a `config.yaml` file might look something like this:\n\n    trainingInput:\n      region: us-central1\n      scaleTier: CUSTOM\n      masterType: complex_model_m\n      workerType: complex_model_m_gpu\n      parameterServerType: large_model\n      workerCount: 4\n      parameterServerCount: 2\n\nNotice that only the workers have GPUs attached -- the other machine types do not need GPUs attached. \n\n![distributed architecture diagram][3]\n\nUnderstanding ML Units and Quota\n================================\n\nDifferent machines consume differing amounts of \"ML Units\". You should go to [this page][4] detailing each type of machine and how many ML Units it consumes. The current limit on concurrent training units is 25 ML Units. \n\nTo compute how many ML Units we used in the example above, we first look up how many units each of the machine types we used account for. The master, `complex_model_m`, is 3 ML Units. The worker, a `complex_model_m_gpu`, is 12 ML Units. Finally, the parameter server, a `large_model`, is 3 ML Units. \nSo our total count is 1x `complex_model_m` (3 ML Units) + 4x `complex_model_m_gpu` (12 ML Units) + 2x `large_model` (3 ML Units) = 3 + 48 + 6 = 57 ML Units! \n\n![table of ML Unit examples][5]\n\nIf we had used `standard_gpu` instead of `complex_model_m_gpu`, we would have dramatically reduced our ML Unit usage (it would also cut our GPU count by a factor of 4), since the `standard_gpu` is only 3 ML Units (which when compared to the ML Unit values of some of the other machine types really makes it quite a good value!). It would have reduced our consumption to only 21 ML Units.\n\nLet's say you want to run that big job, with 57 ML Units consumed. You'd need to increase your quota. However, the quota for ML Units is not in the cloud console's IAM page, so you should head on over to the [CMLE Quota Request Form][6] and detail your needs there.\n\nCurrent quota limits can be found on our [Quotas page][7]. Another limit to be aware of is the number of concurrent GPUs, which is 10 by default, so if you need more than 10 GPUs, be sure to request it using the [Quota Request Form][8]. In the example above, we had 4x complex_model_m_gpu machines, which each contain 4 GPUs, totaling 16 GPUs, so it would definitely need a quota increase for GPUs as well.\n\nAccessing GPUs in TensorFlow code\n===\n\nThe specifics of how you go about accessing GPUs in TensorFlow are a bit out of scope of this guide, but broadly speaking, you have two options: using the high-level APIs (Estimator) or using the low-level APIs. The \"Assigning ops to GPUs\" section of [this article][9] goes into more detail.\n\nAdditional Resources and Reading\n===\n\n[TensorFlow ClusterSpec][10]: This will be useful for setting up your training to access your GPU resources\n\nUsing GPUs On Compute Engine\nhttps://cloud.google.com/compute/docs/gpus/ \n\nDocumentation for using GPUs On Cloud Machine Learning Engine\nhttps://cloud.google.com/ml-engine/docs/how-tos/using-gpus\n\nReading in the TSA data:\nWill has some start code for reading in the images here: https://www.kaggle.com/wcukierski/reading-images\n\n\n  [1]: https://cloud.google.com/ml-engine/docs/how-tos/using-gpus\n  [2]: https://cloud.google.com/ml-engine/docs/how-tos/using-gpus#requesting_gpu-enabled_machines\n  [3]: https://storage.googleapis.com/image-uploader/kaggle_posts/guide1.png\n  [4]: https://cloud.google.com/ml-engine/docs/concepts/training-overview#scale_tier\n  [5]: https://storage.googleapis.com/image-uploader/kaggle_posts/mlunittable.png\n  [6]: https://docs.google.com/a/google.com/forms/d/e/1FAIpQLScjm2W208wdf0iyQxBIESZ2_UvTS9xVPKleTQg9r_Uc1uPACQ/viewform\n  [7]: https://cloud.google.com/ml-engine/quotas#limits_on_concurrent_ml_training_units_and_prediction_nodeshttp://\n  [8]: https://docs.google.com/a/google.com/forms/d/e/1FAIpQLScjm2W208wdf0iyQxBIESZ2_UvTS9xVPKleTQg9r_Uc1uPACQ/viewform\n  [9]: https://cloud.google.com/ml-engine/docs/how-tos/using-gpus#assigning_ops_to_gpus\n  [10]: https://www.tensorflow.org/api_docs/python/tf/train/ClusterSpec",
    "233695": "Could you please mention how much does it cost to use various GPU configuration?",
    "234848": "See [this table](https://cloud.google.com/ml-engine/docs/training-overview#comparing_machine_types) with all the machine types, and use the number of \"ML units\" to compute [pricing](https://cloud.google.com/ml-engine/pricing).\nThe quick version is that as of this writing, it costs $0.49 per hour, per ML training unit, for training in the US, and $0.54 in Europe/Asia.",
    "253648": "How do I run this on google cloud ?\n\n    import tensorflow as tf;\n    hello = tf.constant('Hello, TensorFlow!');\n    sess = tf.Session();\n    print(sess.run(hello));\n    a = tf.constant([1.0, 2.0, 3.0, 4.0], shape=[2, 2], name='a');\n    c = tf.matmul(a,a)\n    with tf.device('/device:GPU:0'):\n       print(sess.run(c))",
    "253783": "I put the above code in a hi.py file, add the app.yaml below, run `gcloud app deploy` but get error:  hi.py not found.\n\n    runtime: python\n    env: flex\n    entrypoint: hi.py\n    runtime_config:\n      python_version: 3.5\n    manual_scaling:\n      instances: 1\n    resources:\n      cpu: 8\n      memory_gb: 51.6\n      disk_size_gb: 100\n    readiness_check:\n    app_start_timeout_sec: 3600",
    "254149": "I found this example, that works:  https://cloud.google.com/ml-engine/docs/command-line"
  },
  "source": "meta"
}