{
  "id": 36337,
  "title": "Sharing advice and tips for google cloud",
  "url": "/competitions/passenger-screening-algorithm-challenge/discussion/36337",
  "author_name": "",
  "post_date": "2017-07-14T04:45:51.014376100Z",
  "votes": 7,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hello! I doing this competition for fun and using it as a learning opportunity. I thought I would try using google cloud for this as my local machine does not have much space left. I assume some people here also do not have much experience with the cloud. I am making this post for people to help, share resources, or give any useful advice on how best to use google cloud. Please do share if  you have anything useful for newcomers.</p>",
  "messages": [
    {
      "id": "203148",
      "postDate": "07/14/2017 04:45:51",
      "content": "<p>Hello! I doing this competition for fun and using it as a learning opportunity. I thought I would try using google cloud for this as my local machine does not have much space left. I assume some people here also do not have much experience with the cloud. I am making this post for people to help, share resources, or give any useful advice on how best to use google cloud. Please do share if  you have anything useful for newcomers.</p>",
      "rawMarkdown": "Hello! I doing this competition for fun and using it as a learning opportunity. I thought I would try using google cloud for this as my local machine does not have much space left. I assume some people here also do not have much experience with the cloud. I am making this post for people to help, share resources, or give any useful advice on how best to use google cloud. Please do share if  you have anything useful for newcomers.",
      "votes": null
    },
    {
      "id": "204279",
      "postDate": "07/18/2017 05:46:13",
      "content": "<p>Hi Akshay,</p>\n\n<p>I'm Kaz, Developer Advocate at Google Cloud team. Thanks for having interest in using Google Cloud for this competition. As a starter, I'd recommend you taking look at <a href=\"https://cloud.google.com/datalab/\">Cloud Datalab</a>. The tool is a Jupyter Notebook integrated with Google Cloud. You can easily setup all-in-one toolset including numpy/scipy/sklearn/matplotlib/pandas and also TensorFlow. Also, with Datalab you can easily access Google Cloud services such as BigQuery (data warehouse), Cloud Storage (object storage), Cloud Dataflow (batch+stream processing), and Cloud Dataproc (managed Hadoop/Spark).</p>\n\n<p>For a few months, a few people from our team will take a look at this discussion group for finding anything we can help. So please feel free to post any Google Cloud/TensorFlow related questions on the group.</p>",
      "rawMarkdown": "Hi Akshay,\n\nI'm Kaz, Developer Advocate at Google Cloud team. Thanks for having interest in using Google Cloud for this competition. As a starter, I'd recommend you taking look at [Cloud Datalab](https://cloud.google.com/datalab/). The tool is a Jupyter Notebook integrated with Google Cloud. You can easily setup all-in-one toolset including numpy/scipy/sklearn/matplotlib/pandas and also TensorFlow. Also, with Datalab you can easily access Google Cloud services such as BigQuery (data warehouse), Cloud Storage (object storage), Cloud Dataflow (batch+stream processing), and Cloud Dataproc (managed Hadoop/Spark).\n\nFor a few months, a few people from our team will take a look at this discussion group for finding anything we can help. So please feel free to post any Google Cloud/TensorFlow related questions on the group.",
      "votes": null
    },
    {
      "id": "204499",
      "postDate": "07/18/2017 18:56:35",
      "content": "<p>Hi, So I noticed that in the <em>datalab create</em> command, there is no flag to add accelerators to the VM instance, whereas the <em>gcloud compute instances create</em> command has the <em>--accelerator</em> flag. Is there no ability to add a GPU to the datalab VM instance or am I missing something?</p>\n\n<p>Thank you</p>",
      "rawMarkdown": "Hi, So I noticed that in the *datalab create* command, there is no flag to add accelerators to the VM instance, whereas the *gcloud compute instances create* command has the *--accelerator* flag. Is there no ability to add a GPU to the datalab VM instance or am I missing something?\n\nThank you",
      "votes": null
    },
    {
      "id": "204682",
      "postDate": "07/19/2017 10:09:09",
      "content": "<p>Hi Poseidon,</p>\n\n<p>Cloud Datalab itself doesn't support GPU, but usually you may don't want to run a large training on Datalab. Instead, you may run training job on <a href=\"https://cloud.google.com/ml-engine/docs/how-tos/using-gpus\">Cloud ML Engine with GPUs</a>  so that you can control large scale distributed training from Datalab.</p>",
      "rawMarkdown": "Hi Poseidon,\n\nCloud Datalab itself doesn't support GPU, but usually you may don't want to run a large training on Datalab. Instead, you may run training job on [Cloud ML Engine with GPUs](https://cloud.google.com/ml-engine/docs/how-tos/using-gpus)  so that you can control large scale distributed training from Datalab.",
      "votes": null
    },
    {
      "id": "204711",
      "postDate": "07/19/2017 12:01:10",
      "content": "<p>I highly recommend anyone who wants to learn how to use gcloud to checkout the official tensorflow/object_detection tutorial on how to run the object detection api on gcloud. It introduces the basic pipeline to use google cloud. Here's the <a href=\"https://github.com/tensorflow/models/blob/master/object_detection/g3doc/running_pets.md\">link</a>. </p>\n\n<p>Also just as a general tip, the 'gsutil' command prepended to any command on your console, functions exactly like standard unix terminal commands i.e. gsutil mv foo.txt gs://test-bucket-3/data/utils moves foo.txt from your current working directory to your bucket named 'test-bucket-3' under /data/utils. You can also use tab auto-complete on local files, so something like 'gsutil cp fo' can auto-complete to 'gsutil cp foo.txt' after pressing tab, just like you normally can in console.</p>\n\n<p>Another important tip is to appropriately configure your custom YAML file according to <a href=\"https://cloud.google.com/ml-engine/docs/concepts/training-overview\">documentation</a> (Go to the section 'Scale Tier'). Gcloud doesn't really let you pick instances like DigitalOcean or AWS does, so this is where you decide your scale tier. </p>",
      "rawMarkdown": "I highly recommend anyone who wants to learn how to use gcloud to checkout the official tensorflow/object_detection tutorial on how to run the object detection api on gcloud. It introduces the basic pipeline to use google cloud. Here's the [link](https://github.com/tensorflow/models/blob/master/object_detection/g3doc/running_pets.md). \n\nAlso just as a general tip, the 'gsutil' command prepended to any command on your console, functions exactly like standard unix terminal commands i.e. gsutil mv foo.txt gs://test-bucket-3/data/utils moves foo.txt from your current working directory to your bucket named 'test-bucket-3' under /data/utils. You can also use tab auto-complete on local files, so something like 'gsutil cp fo' can auto-complete to 'gsutil cp foo.txt' after pressing tab, just like you normally can in console.\n\nAnother important tip is to appropriately configure your custom YAML file according to [documentation](https://cloud.google.com/ml-engine/docs/concepts/training-overview) (Go to the section 'Scale Tier'). Gcloud doesn't really let you pick instances like DigitalOcean or AWS does, so this is where you decide your scale tier.",
      "votes": null
    },
    {
      "id": "205249",
      "postDate": "07/21/2017 00:23:14",
      "content": "<p>Thank you, Kaz! This is great!</p>",
      "rawMarkdown": "Thank you, Kaz! This is great!",
      "votes": null
    },
    {
      "id": "208986",
      "postDate": "07/31/2017 20:00:39",
      "content": "<p>Hi Poseidon,</p>\n\n<p>To follow up on Kaz's comment below, did you have any luck training on GCP?</p>\n\n<p>Just in case, <a href=\"https://github.com/GoogleCloudPlatform/cloudml-samples\">this repository</a> contains many examples of how you can submit your training jobs to ML Engine and how you can subsequently perform batch predictions on your test data.</p>",
      "rawMarkdown": "Hi Poseidon,\n\nTo follow up on Kaz's comment below, did you have any luck training on GCP?\n\nJust in case, [this repository][1] contains many examples of how you can submit your training jobs to ML Engine and how you can subsequently perform batch predictions on your test data.\n\n\n  [1]: https://github.com/GoogleCloudPlatform/cloudml-samples",
      "votes": null
    },
    {
      "id": "218815",
      "postDate": "09/06/2017 00:22:24",
      "content": "<p>Some beginners questions:</p>\n\n<p>1) How do I read an image in datalab from the storage bucket? Using a python function that expects a filename, e.g. pyplot.imread()?</p>\n\n<p>2) Why can I not read files in datalab directly from the original kaggle-tsa-stage1 bucket? Why do I first have to copy the the whole thing to my own bucket and then read it? I presume there is some performance hit but I may choose for that rather than paying for 3TB of storage space initially. I can browse through the files with a browser,...</p>\n\n<p>3) The \"Kernels\" don't copy 1:1 to Google datalab, neither do they work in Kaggle's internal private Kernel mode due to data confidentiality and above cloud implications. Am I missing something?</p>",
      "rawMarkdown": "Some beginners questions:\n\n1) How do I read an image in datalab from the storage bucket? Using a python function that expects a filename, e.g. pyplot.imread()?\n\n2) Why can I not read files in datalab directly from the original kaggle-tsa-stage1 bucket? Why do I first have to copy the the whole thing to my own bucket and then read it? I presume there is some performance hit but I may choose for that rather than paying for 3TB of storage space initially. I can browse through the files with a browser,...\n    \n3) The \"Kernels\" don't copy 1:1 to Google datalab, neither do they work in Kaggle's internal private Kernel mode due to data confidentiality and above cloud implications. Am I missing something?",
      "votes": null
    },
    {
      "id": "218842",
      "postDate": "09/06/2017 02:34:51",
      "content": "<p>Hi Bastiaan,</p>\n\n<p>1) Yes, you can use the <a href=\"http://googledatalab.github.io/pydatalab/google.datalab.storage.html\">Python API</a> of Datalab for Storage access.</p>\n\n<p>2) Unfortunately it's the current limitation with Datalab. In general, it's not recommended to use your  personal account to access any cloud services (in this case it's Storage) from a program code, for security and portability reason. So Datalab requires to use its service account instead.</p>\n\n<p>3) Kaggle Kernels and Cloud Datalab are two different services and not compatible. You may try porting Kernels to Datalab, but you'd need to modify some parts manually.</p>",
      "rawMarkdown": "Hi Bastiaan,\n\n1) Yes, you can use the [Python API][1] of Datalab for Storage access.\n\n2) Unfortunately it's the current limitation with Datalab. In general, it's not recommended to use your  personal account to access any cloud services (in this case it's Storage) from a program code, for security and portability reason. So Datalab requires to use its service account instead.\n\n3) Kaggle Kernels and Cloud Datalab are two different services and not compatible. You may try porting Kernels to Datalab, but you'd need to modify some parts manually.\n\n\n  [1]: http://googledatalab.github.io/pydatalab/google.datalab.storage.html",
      "votes": null
    },
    {
      "id": "218862",
      "postDate": "09/06/2017 05:28:33",
      "content": "<p>Thanks Kaz, that makes sense now.</p>",
      "rawMarkdown": "Thanks Kaz, that makes sense now.",
      "votes": null
    },
    {
      "id": "394339",
      "postDate": "09/26/2018 17:38:26",
      "content": "<p>how to upload kaggle data using kaggle api on google cloud?</p>",
      "rawMarkdown": "how to upload kaggle data using kaggle api on google cloud?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 204279,
      "author_name": "kazunori279",
      "author_url": "",
      "post_date": "07/18/2017 05:46:13",
      "content": "<p>Hi Akshay,</p>\n\n<p>I'm Kaz, Developer Advocate at Google Cloud team. Thanks for having interest in using Google Cloud for this competition. As a starter, I'd recommend you taking look at <a href=\"https://cloud.google.com/datalab/\">Cloud Datalab</a>. The tool is a Jupyter Notebook integrated with Google Cloud. You can easily setup all-in-one toolset including numpy/scipy/sklearn/matplotlib/pandas and also TensorFlow. Also, with Datalab you can easily access Google Cloud services such as BigQuery (data warehouse), Cloud Storage (object storage), Cloud Dataflow (batch+stream processing), and Cloud Dataproc (managed Hadoop/Spark).</p>\n\n<p>For a few months, a few people from our team will take a look at this discussion group for finding anything we can help. So please feel free to post any Google Cloud/TensorFlow related questions on the group.</p>",
      "votes": null,
      "replies": [
        {
          "id": 205249,
          "author_name": "greymatter8654",
          "author_url": "",
          "post_date": "07/21/2017 00:23:14",
          "content": "<p>Thank you, Kaz! This is great!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 394339,
          "author_name": "aaryapatel",
          "author_url": "",
          "post_date": "09/26/2018 17:38:26",
          "content": "<p>how to upload kaggle data using kaggle api on google cloud?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 204499,
      "author_name": "posiedon",
      "author_url": "",
      "post_date": "07/18/2017 18:56:35",
      "content": "<p>Hi, So I noticed that in the <em>datalab create</em> command, there is no flag to add accelerators to the VM instance, whereas the <em>gcloud compute instances create</em> command has the <em>--accelerator</em> flag. Is there no ability to add a GPU to the datalab VM instance or am I missing something?</p>\n\n<p>Thank you</p>",
      "votes": null,
      "replies": [
        {
          "id": 208986,
          "author_name": "fuzzyfroghunter",
          "author_url": "",
          "post_date": "07/31/2017 20:00:39",
          "content": "<p>Hi Poseidon,</p>\n\n<p>To follow up on Kaz's comment below, did you have any luck training on GCP?</p>\n\n<p>Just in case, <a href=\"https://github.com/GoogleCloudPlatform/cloudml-samples\">this repository</a> contains many examples of how you can submit your training jobs to ML Engine and how you can subsequently perform batch predictions on your test data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 204682,
      "author_name": "kazunori279",
      "author_url": "",
      "post_date": "07/19/2017 10:09:09",
      "content": "<p>Hi Poseidon,</p>\n\n<p>Cloud Datalab itself doesn't support GPU, but usually you may don't want to run a large training on Datalab. Instead, you may run training job on <a href=\"https://cloud.google.com/ml-engine/docs/how-tos/using-gpus\">Cloud ML Engine with GPUs</a>  so that you can control large scale distributed training from Datalab.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 204711,
      "author_name": "int64cast",
      "author_url": "",
      "post_date": "07/19/2017 12:01:10",
      "content": "<p>I highly recommend anyone who wants to learn how to use gcloud to checkout the official tensorflow/object_detection tutorial on how to run the object detection api on gcloud. It introduces the basic pipeline to use google cloud. Here's the <a href=\"https://github.com/tensorflow/models/blob/master/object_detection/g3doc/running_pets.md\">link</a>. </p>\n\n<p>Also just as a general tip, the 'gsutil' command prepended to any command on your console, functions exactly like standard unix terminal commands i.e. gsutil mv foo.txt gs://test-bucket-3/data/utils moves foo.txt from your current working directory to your bucket named 'test-bucket-3' under /data/utils. You can also use tab auto-complete on local files, so something like 'gsutil cp fo' can auto-complete to 'gsutil cp foo.txt' after pressing tab, just like you normally can in console.</p>\n\n<p>Another important tip is to appropriately configure your custom YAML file according to <a href=\"https://cloud.google.com/ml-engine/docs/concepts/training-overview\">documentation</a> (Go to the section 'Scale Tier'). Gcloud doesn't really let you pick instances like DigitalOcean or AWS does, so this is where you decide your scale tier. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 218815,
      "author_name": "bastiaanbergman",
      "author_url": "",
      "post_date": "09/06/2017 00:22:24",
      "content": "<p>Some beginners questions:</p>\n\n<p>1) How do I read an image in datalab from the storage bucket? Using a python function that expects a filename, e.g. pyplot.imread()?</p>\n\n<p>2) Why can I not read files in datalab directly from the original kaggle-tsa-stage1 bucket? Why do I first have to copy the the whole thing to my own bucket and then read it? I presume there is some performance hit but I may choose for that rather than paying for 3TB of storage space initially. I can browse through the files with a browser,...</p>\n\n<p>3) The \"Kernels\" don't copy 1:1 to Google datalab, neither do they work in Kaggle's internal private Kernel mode due to data confidentiality and above cloud implications. Am I missing something?</p>",
      "votes": null,
      "replies": [
        {
          "id": 218842,
          "author_name": "kazunori279",
          "author_url": "",
          "post_date": "09/06/2017 02:34:51",
          "content": "<p>Hi Bastiaan,</p>\n\n<p>1) Yes, you can use the <a href=\"http://googledatalab.github.io/pydatalab/google.datalab.storage.html\">Python API</a> of Datalab for Storage access.</p>\n\n<p>2) Unfortunately it's the current limitation with Datalab. In general, it's not recommended to use your  personal account to access any cloud services (in this case it's Storage) from a program code, for security and portability reason. So Datalab requires to use its service account instead.</p>\n\n<p>3) Kaggle Kernels and Cloud Datalab are two different services and not compatible. You may try porting Kernels to Datalab, but you'd need to modify some parts manually.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 218862,
          "author_name": "bastiaanbergman",
          "author_url": "",
          "post_date": "09/06/2017 05:28:33",
          "content": "<p>Thanks Kaz, that makes sense now.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "203148": "Hello! I doing this competition for fun and using it as a learning opportunity. I thought I would try using google cloud for this as my local machine does not have much space left. I assume some people here also do not have much experience with the cloud. I am making this post for people to help, share resources, or give any useful advice on how best to use google cloud. Please do share if  you have anything useful for newcomers.",
    "204279": "Hi Akshay,\n\nI'm Kaz, Developer Advocate at Google Cloud team. Thanks for having interest in using Google Cloud for this competition. As a starter, I'd recommend you taking look at [Cloud Datalab](https://cloud.google.com/datalab/). The tool is a Jupyter Notebook integrated with Google Cloud. You can easily setup all-in-one toolset including numpy/scipy/sklearn/matplotlib/pandas and also TensorFlow. Also, with Datalab you can easily access Google Cloud services such as BigQuery (data warehouse), Cloud Storage (object storage), Cloud Dataflow (batch+stream processing), and Cloud Dataproc (managed Hadoop/Spark).\n\nFor a few months, a few people from our team will take a look at this discussion group for finding anything we can help. So please feel free to post any Google Cloud/TensorFlow related questions on the group.",
    "204499": "Hi, So I noticed that in the *datalab create* command, there is no flag to add accelerators to the VM instance, whereas the *gcloud compute instances create* command has the *--accelerator* flag. Is there no ability to add a GPU to the datalab VM instance or am I missing something?\n\nThank you",
    "204682": "Hi Poseidon,\n\nCloud Datalab itself doesn't support GPU, but usually you may don't want to run a large training on Datalab. Instead, you may run training job on [Cloud ML Engine with GPUs](https://cloud.google.com/ml-engine/docs/how-tos/using-gpus)  so that you can control large scale distributed training from Datalab.",
    "204711": "I highly recommend anyone who wants to learn how to use gcloud to checkout the official tensorflow/object_detection tutorial on how to run the object detection api on gcloud. It introduces the basic pipeline to use google cloud. Here's the [link](https://github.com/tensorflow/models/blob/master/object_detection/g3doc/running_pets.md). \n\nAlso just as a general tip, the 'gsutil' command prepended to any command on your console, functions exactly like standard unix terminal commands i.e. gsutil mv foo.txt gs://test-bucket-3/data/utils moves foo.txt from your current working directory to your bucket named 'test-bucket-3' under /data/utils. You can also use tab auto-complete on local files, so something like 'gsutil cp fo' can auto-complete to 'gsutil cp foo.txt' after pressing tab, just like you normally can in console.\n\nAnother important tip is to appropriately configure your custom YAML file according to [documentation](https://cloud.google.com/ml-engine/docs/concepts/training-overview) (Go to the section 'Scale Tier'). Gcloud doesn't really let you pick instances like DigitalOcean or AWS does, so this is where you decide your scale tier.",
    "205249": "Thank you, Kaz! This is great!",
    "208986": "Hi Poseidon,\n\nTo follow up on Kaz's comment below, did you have any luck training on GCP?\n\nJust in case, [this repository][1] contains many examples of how you can submit your training jobs to ML Engine and how you can subsequently perform batch predictions on your test data.\n\n\n  [1]: https://github.com/GoogleCloudPlatform/cloudml-samples",
    "218815": "Some beginners questions:\n\n1) How do I read an image in datalab from the storage bucket? Using a python function that expects a filename, e.g. pyplot.imread()?\n\n2) Why can I not read files in datalab directly from the original kaggle-tsa-stage1 bucket? Why do I first have to copy the the whole thing to my own bucket and then read it? I presume there is some performance hit but I may choose for that rather than paying for 3TB of storage space initially. I can browse through the files with a browser,...\n    \n3) The \"Kernels\" don't copy 1:1 to Google datalab, neither do they work in Kaggle's internal private Kernel mode due to data confidentiality and above cloud implications. Am I missing something?",
    "218842": "Hi Bastiaan,\n\n1) Yes, you can use the [Python API][1] of Datalab for Storage access.\n\n2) Unfortunately it's the current limitation with Datalab. In general, it's not recommended to use your  personal account to access any cloud services (in this case it's Storage) from a program code, for security and portability reason. So Datalab requires to use its service account instead.\n\n3) Kaggle Kernels and Cloud Datalab are two different services and not compatible. You may try porting Kernels to Datalab, but you'd need to modify some parts manually.\n\n\n  [1]: http://googledatalab.github.io/pydatalab/google.datalab.storage.html",
    "218862": "Thanks Kaz, that makes sense now.",
    "394339": "how to upload kaggle data using kaggle api on google cloud?"
  },
  "source": "meta"
}