{
  "id": 37088,
  "title": "Accessing the dataset for use on Google Cloud",
  "url": "/competitions/passenger-screening-algorithm-challenge/discussion/37088",
  "author_name": "Yufeng Guo",
  "post_date": "2017-07-26T23:16:07.438000",
  "votes": 22,
  "comment_count": 43,
  "views": 0,
  "content": "<p>Due to access restrictions, you will need to copy your data from the kaggle-tsa-stage1 bucket to your own <a href=\"https://cloud.google.com/storage/\">Google Cloud Storage</a> (GCS) bucket in order to access the data using Google Cloud. </p>\n\n<p>The easiest way to do this is to use the <a href=\"https://cloud.google.com/sdk/docs/\">gcloud command-line tool</a>, which includes a tool called <a href=\"https://cloud.google.com/storage/docs/gsutil\">gsutil</a>, which is used for managing your Google Cloud Storage data. You can also accomplish this copy using the UI if you are more comfortable with that. Check out this YouTube video for details: <a href=\"https://youtu.be/wo--uiuPsiU\">https://youtu.be/wo--uiuPsiU</a></p>\n\n<h1>Command line access </h1>\n\n<p>GCS uses buckets to represent data and access. \nIf you don't have an appropriate GCS bucket to put your data into, you should <a href=\"https://cloud.google.com/storage/docs/gsutil/commands/mb\">create one</a> using </p>\n\n<pre><code>gsutil mb -c regional -l us-east1 gs://my-kaggle-data\n</code></pre>\n\n<p>(\"mb\" is short for \"make bucket\") </p>\n\n<p>Bucket names are globally unique across all of gcs, so you will need to change the command to use your own unique bucket name. </p>\n\n<p>For the <code>-l</code> argument, you should choose to place your data in a region that contains GPUs if you plan to use GPUs in your training. Currently GPUs are only available in the following regions:</p>\n\n<ul>\n<li>us-east1 </li>\n<li>us-central1 </li>\n<li>asia-east1 </li>\n<li>europe-west1</li>\n</ul>\n\n<p>You can learn more about regions <a href=\"https://cloud.google.com/compute/docs/regions-zones/regions-zones\">here</a>. </p>\n\n<p>The command we will be interested in using is the copy command, which in this case we will use with a recursive copy flag (<code>-r</code>) and run in parallel (<code>-m</code>):</p>\n\n<pre><code>gsutil -m cp -r gs://kaggle-tsa-stage1/*  gs://my-kaggle-data\n</code></pre>\n\n<p>(the \"gs\" filesystem prefix is short for \"Google Storage\") </p>\n\n<p>This will grab all the data from the kaggle-tsa-stage1 bucket and put it into your own bucket. Since the dataset is quite large, this copy operation will take some time. In our experience, it has taken on the order of 30 minutes. </p>\n\n<p>Once the copy operation is complete, use gsutil ls to confirm the successful copy operation:</p>\n\n<pre><code>gsutil ls gs://my-kaggle-data\n</code></pre>\n\n<h1>Wrapping up </h1>\n\n<p>Now you are ready to use Cloud Machine Learning Engine to access your data, and utilize the power of distributed GPUs! Check out our guide on using GPUs with Cloud Machine Learning Engine here. </p>",
  "messages": [
    {
      "id": 207582,
      "postDate": "2017-07-26T23:16:07.437Z",
      "content": "<p>Due to access restrictions, you will need to copy your data from the kaggle-tsa-stage1 bucket to your own <a href=\"https://cloud.google.com/storage/\">Google Cloud Storage</a> (GCS) bucket in order to access the data using Google Cloud. </p>\n\n<p>The easiest way to do this is to use the <a href=\"https://cloud.google.com/sdk/docs/\">gcloud command-line tool</a>, which includes a tool called <a href=\"https://cloud.google.com/storage/docs/gsutil\">gsutil</a>, which is used for managing your Google Cloud Storage data. You can also accomplish this copy using the UI if you are more comfortable with that. Check out this YouTube video for details: <a href=\"https://youtu.be/wo--uiuPsiU\">https://youtu.be/wo--uiuPsiU</a></p>\n\n<h1>Command line access </h1>\n\n<p>GCS uses buckets to represent data and access. \nIf you don't have an appropriate GCS bucket to put your data into, you should <a href=\"https://cloud.google.com/storage/docs/gsutil/commands/mb\">create one</a> using </p>\n\n<pre><code>gsutil mb -c regional -l us-east1 gs://my-kaggle-data\n</code></pre>\n\n<p>(\"mb\" is short for \"make bucket\") </p>\n\n<p>Bucket names are globally unique across all of gcs, so you will need to change the command to use your own unique bucket name. </p>\n\n<p>For the <code>-l</code> argument, you should choose to place your data in a region that contains GPUs if you plan to use GPUs in your training. Currently GPUs are only available in the following regions:</p>\n\n<ul>\n<li>us-east1 </li>\n<li>us-central1 </li>\n<li>asia-east1 </li>\n<li>europe-west1</li>\n</ul>\n\n<p>You can learn more about regions <a href=\"https://cloud.google.com/compute/docs/regions-zones/regions-zones\">here</a>. </p>\n\n<p>The command we will be interested in using is the copy command, which in this case we will use with a recursive copy flag (<code>-r</code>) and run in parallel (<code>-m</code>):</p>\n\n<pre><code>gsutil -m cp -r gs://kaggle-tsa-stage1/*  gs://my-kaggle-data\n</code></pre>\n\n<p>(the \"gs\" filesystem prefix is short for \"Google Storage\") </p>\n\n<p>This will grab all the data from the kaggle-tsa-stage1 bucket and put it into your own bucket. Since the dataset is quite large, this copy operation will take some time. In our experience, it has taken on the order of 30 minutes. </p>\n\n<p>Once the copy operation is complete, use gsutil ls to confirm the successful copy operation:</p>\n\n<pre><code>gsutil ls gs://my-kaggle-data\n</code></pre>\n\n<h1>Wrapping up </h1>\n\n<p>Now you are ready to use Cloud Machine Learning Engine to access your data, and utilize the power of distributed GPUs! Check out our guide on using GPUs with Cloud Machine Learning Engine here. </p>",
      "rawMarkdown": "Due to access restrictions, you will need to copy your data from the kaggle-tsa-stage1 bucket to your own [Google Cloud Storage][1] (GCS) bucket in order to access the data using Google Cloud. \n\nThe easiest way to do this is to use the [gcloud command-line tool][2], which includes a tool called [gsutil][3], which is used for managing your Google Cloud Storage data. You can also accomplish this copy using the UI if you are more comfortable with that. Check out this YouTube video for details: https://youtu.be/wo--uiuPsiU\n\nCommand line access \n===\n\nGCS uses buckets to represent data and access. \nIf you don't have an appropriate GCS bucket to put your data into, you should [create one][4] using \n\n    gsutil mb -c regional -l us-east1 gs://my-kaggle-data\n(\"mb\" is short for \"make bucket\") \n\nBucket names are globally unique across all of gcs, so you will need to change the command to use your own unique bucket name. \n\nFor the `-l` argument, you should choose to place your data in a region that contains GPUs if you plan to use GPUs in your training. Currently GPUs are only available in the following regions:\n\n - us-east1 \n - us-central1 \n - asia-east1 \n - europe-west1\n\nYou can learn more about regions [here][5]. \n\nThe command we will be interested in using is the copy command, which in this case we will use with a recursive copy flag (`-r`) and run in parallel (`-m`):\n\n    gsutil -m cp -r gs://kaggle-tsa-stage1/*  gs://my-kaggle-data\n(the \"gs\" filesystem prefix is short for \"Google Storage\") \n\nThis will grab all the data from the kaggle-tsa-stage1 bucket and put it into your own bucket. Since the dataset is quite large, this copy operation will take some time. In our experience, it has taken on the order of 30 minutes. \n\n\nOnce the copy operation is complete, use gsutil ls to confirm the successful copy operation:\n\n    gsutil ls gs://my-kaggle-data\n\nWrapping up \n====\n\nNow you are ready to use Cloud Machine Learning Engine to access your data, and utilize the power of distributed GPUs! Check out our guide on using GPUs with Cloud Machine Learning Engine here. \n\n\n  [1]: https://cloud.google.com/storage/\n  [2]: https://cloud.google.com/sdk/docs/\n  [3]: https://cloud.google.com/storage/docs/gsutil\n  [4]: https://cloud.google.com/storage/docs/gsutil/commands/mb\n  [5]: https://cloud.google.com/compute/docs/regions-zones/regions-zones",
      "votes": 22
    },
    {
      "id": 219432,
      "postDate": "2017-09-08T06:25:08.510Z",
      "content": "<p>It adds a lot to the costs for each participant to copy the entire dataset into their own buckets. Why not give us Read access as recommended here for 'when a large dataset is shared across multiple projects':</p>\n\n<p><a href=\"https://cloud.google.com/ml-engine/docs/how-tos/working-with-data#using_a_cloud_storage_bucket_from_a_different_project\">Using a Cloud Storage Bucket from a Different Project</a></p>\n\n<p>All we would need is to share our 'service account information' string with you.</p>",
      "rawMarkdown": "It adds a lot to the costs for each participant to copy the entire dataset into their own buckets. Why not give us Read access as recommended here for 'when a large dataset is shared across multiple projects':\n\n[Using a Cloud Storage Bucket from a Different Project][1]\n\nAll we would need is to share our 'service account information' string with you.\n\n\n  [1]: https://cloud.google.com/ml-engine/docs/how-tos/working-with-data#using_a_cloud_storage_bucket_from_a_different_project",
      "votes": 2,
      "replies": [
        {
          "id": 219465,
          "postDate": "2017-09-08T09:39:03.303Z",
          "content": "<p>Hi Muneeb,</p>\n\n<p>In general use case of Cloud Storage, yes you can do it. But in the case of this competition,  \"Access to the Google Cloud Storage bucket is controlled by membership in a Google group\" as explained at the Data section. So you need to use your Google account to be a member of the group to get access to the bucket. </p>",
          "rawMarkdown": "Hi Muneeb,\n\nIn general use case of Cloud Storage, yes you can do it. But in the case of this competition,  \"Access to the Google Cloud Storage bucket is controlled by membership in a Google group\" as explained at the Data section. So you need to use your Google account to be a member of the group to get access to the bucket. ",
          "votes": -2
        },
        {
          "id": 219668,
          "postDate": "2017-09-09T01:58:29.867Z",
          "content": "<p>Yes I already have 'access' to the bucket per se because of Group membership, but to use it in the Google Cloud ML Engine project for training purposes, you need special permissions as detailed in the link I shared above.</p>",
          "rawMarkdown": "Yes I already have 'access' to the bucket per se because of Group membership, but to use it in the Google Cloud ML Engine project for training purposes, you need special permissions as detailed in the link I shared above."
        },
        {
          "id": 219676,
          "postDate": "2017-09-09T02:48:47.267Z",
          "content": "<p>I see, were you able to use ML Engine by granting access to its service account?</p>",
          "rawMarkdown": "I see, were you able to use ML Engine by granting access to its service account?"
        },
        {
          "id": 219709,
          "postDate": "2017-09-09T08:53:27.873Z",
          "content": "<p>I don't have admin access to the tsa bucket so I cannot modify access control for that bucket to allow my ML Engine project read access. I assume only owners of kaggle-tsa-stage1 bucket can allow access. It could be implemented simply by making a Google form and sharing it on the Google Group for this bucket. Anyone requiring access can then submit their 'service account info' string via the form.</p>",
          "rawMarkdown": "I don't have admin access to the tsa bucket so I cannot modify access control for that bucket to allow my ML Engine project read access. I assume only owners of kaggle-tsa-stage1 bucket can allow access. It could be implemented simply by making a Google form and sharing it on the Google Group for this bucket. Anyone requiring access can then submit their 'service account info' string via the form."
        },
        {
          "id": 219733,
          "postDate": "2017-09-09T11:38:16.120Z",
          "content": "<p>I agree, it would be much easier if we had an operation practice that allows participants using service accounts for accessing the bucket directly.</p>",
          "rawMarkdown": "I agree, it would be much easier if we had an operation practice that allows participants using service accounts for accessing the bucket directly."
        },
        {
          "id": 227346,
          "postDate": "2017-10-04T07:09:28.970Z",
          "content": "<p>Hi Yufeng, you think you can help here on avoiding the copying process like above?</p>",
          "rawMarkdown": "Hi Yufeng, you think you can help here on avoiding the copying process like above?"
        },
        {
          "id": 227563,
          "postDate": "2017-10-04T17:13:01.723Z",
          "content": "<p>Unfortunately, this competition can't grant general read-access to the dataset, but by having access to the dataset (via the Google Group), you can read data from it in any situation where you are authenticated through that same account.</p>\n\n<p>For example, if your email is kaggle_user@gmail.com and you've joined the Google group with that email, you can access the data directly from a cloud account that is signed in as kaggle_user@gmail.com. This means that you <em>can</em> read the data from <a href=\"https://cloud.google.com/solutions/running-distributed-tensorflow-on-compute-engine\">Google Compute Engine</a> or <a href=\"https://medium.com/google-cloud/jupyter-tensorflow-nvidia-gpu-docker-google-compute-engine-4a146f085f17\">Google Container Engine</a>, so if run your TensorFlow training from there, then you won't need to copy the data. If you want to download the data to your own cluster directly, you can do so as well using <code>gsutil -m cp &lt;source&gt; &lt;destination&gt;</code> while ssh'd into your cluster.</p>\n\n<p>The problem arises <em>only</em> when you are using Cloud Machine Learning Engine, as it utilizes a service account to access the dataset rather than directly using your signed-in credentials. The reason for this is that this allows the service to manage your resources on your behalf (which is the whole point), rather than having it exposed to you to manage that infrastructure layer. For this, you will need to copy the data out into your own bucket once you have access via the Google Group, as the data was required to be secured to some degree by the competition rules.</p>",
          "rawMarkdown": "Unfortunately, this competition can't grant general read-access to the dataset, but by having access to the dataset (via the Google Group), you can read data from it in any situation where you are authenticated through that same account.\n\nFor example, if your email is kaggle_user@gmail.com and you've joined the Google group with that email, you can access the data directly from a cloud account that is signed in as kaggle_user@gmail.com. This means that you *can* read the data from [Google Compute Engine][1] or [Google Container Engine][2], so if run your TensorFlow training from there, then you won't need to copy the data. If you want to download the data to your own cluster directly, you can do so as well using `gsutil -m cp  ",
          "votes": 2
        },
        {
          "id": 228214,
          "postDate": "2017-10-06T05:09:10.757Z",
          "content": "<p>I understand, but there is a method detailed in the link I shared (<a href=\"https://cloud.google.com/ml-engine/docs/how-tos/working-with-data#using_a_cloud_storage_bucket_from_a_different_project\">https://cloud.google.com/ml-engine/docs/how-tos/working-with-data#using_a_cloud_storage_bucket_from_a_different_project</a>), to get access via service account. There would just be one extra step required to collect one service account string per Google Group member using a form or a thread etc.</p>",
          "rawMarkdown": "I understand, but there is a method detailed in the link I shared (https://cloud.google.com/ml-engine/docs/how-tos/working-with-data#using_a_cloud_storage_bucket_from_a_different_project), to get access via service account. There would just be one extra step required to collect one service account string per Google Group member using a form or a thread etc."
        },
        {
          "id": 233378,
          "postDate": "2017-10-20T01:28:36.980Z",
          "content": "<p>hi Yufeng, thank you very much. This is very helpful. Could you compare the pro and cons for :\n1. using  Google Compute Engine or Google Container Engine \nvs.\n2. using Cloud Machine Learning Engine</p>\n\n<p>Based on your comments it seems using option 1 saves you from copying the data. \nBut based on the names itself, seems option 2 has ML utilities build in already. \nAny other considerations? What which would you recommend?</p>\n\n<p>Thanks!</p>",
          "rawMarkdown": "hi Yufeng, thank you very much. This is very helpful. Could you compare the pro and cons for :\n1. using  Google Compute Engine or Google Container Engine \nvs.\n2. using Cloud Machine Learning Engine\n\nBased on your comments it seems using option 1 saves you from copying the data. \nBut based on the names itself, seems option 2 has ML utilities build in already. \nAny other considerations? What which would you recommend?\n\nThanks!"
        },
        {
          "id": 234842,
          "postDate": "2017-10-24T07:29:11.370Z",
          "content": "<p>Using Cloud ML Engine allows you to run your TF code in a serverless/on-demand manner. You only pay for what you use, and it's trivial to scale up your model to multiple machines/GPUs if you are using tf.estimator framework.</p>\n\n<p>Using GCE/GKE means you will be configuring everything yourself, and pay for machines you allocate for the time you have them running. You also would need to configure your own cluster for scaling up to more than one machine. See <a href=\"https://www.tensorflow.org/versions/r0.12/api_docs/python/train/distributed_execution\">here</a> for more info.\nIf this is not a concern, then there are no other notable downsides.</p>\n\n<p>One idea would be to copy a subset of the dataset to test with using either your local environment or using the Cloud ML Engine environment to see how you like it.</p>",
          "rawMarkdown": "Using Cloud ML Engine allows you to run your TF code in a serverless/on-demand manner. You only pay for what you use, and it's trivial to scale up your model to multiple machines/GPUs if you are using tf.estimator framework.\n\nUsing GCE/GKE means you will be configuring everything yourself, and pay for machines you allocate for the time you have them running. You also would need to configure your own cluster for scaling up to more than one machine. See [here](https://www.tensorflow.org/versions/r0.12/api_docs/python/train/distributed_execution) for more info.\nIf this is not a concern, then there are no other notable downsides.\n\nOne idea would be to copy a subset of the dataset to test with using either your local environment or using the Cloud ML Engine environment to see how you like it."
        },
        {
          "id": 234859,
          "postDate": "2017-10-24T08:06:59.597Z",
          "content": "<blockquote>\n  <p><strong>Muneeb wrote</strong></p>\n  \n  <blockquote>\n    <p>I understand, but there is a method detailed in the link I shared (<a href=\"https://cloud.google.com/ml-engine/docs/how-tos/working-with-data#using_a_cloud_storage_bucket_from_a_different_project\">https://cloud.google.com/ml-engine/docs/how-tos/working-with-data#using_a_cloud_storage_bucket_from_a_different_project</a>), to get access via service account. There would just be one extra step required to collect one service account string per Google Group member using a form or a thread etc.</p>\n  </blockquote>\n</blockquote>\n\n<p>Unfortunately, Google Groups does not allow service accounts to be added as members, otherwise this would be a great solution :)</p>",
          "rawMarkdown": "\n&gt; **Muneeb wrote**\n&gt; \n&gt; &gt; I understand, but there is a method detailed in the link I shared (https://cloud.google.com/ml-engine/docs/how-tos/working-with-data#using_a_cloud_storage_bucket_from_a_different_project), to get access via service account. There would just be one extra step required to collect one service account string per Google Group member using a form or a thread etc.\n\nUnfortunately, Google Groups does not allow service accounts to be added as members, otherwise this would be a great solution :)"
        }
      ]
    },
    {
      "id": 254817,
      "postDate": "2017-12-07T17:27:33.070Z",
      "content": "<p>how do you access the files in the bucket?  I get \"No such file or directory 'gs://tsa//stage1/aps/342656834021fa0186a369f9652e73eb.aps'  when calling:  <br>\n <code>open('gs://tsa//stage1/aps/342656834021fa0186a369f9652e73eb.aps', 'r+b')</code></p>",
      "rawMarkdown": "how do you access the files in the bucket?  I get \"No such file or directory 'gs://tsa//stage1/aps/342656834021fa0186a369f9652e73eb.aps'  when calling:  <br>\n `open('gs://tsa//stage1/aps/342656834021fa0186a369f9652e73eb.aps', 'r+b')`\n",
      "replies": [
        {
          "id": 254945,
          "postDate": "2017-12-07T22:44:48.400Z",
          "content": "<p>The code below gets the data, but the fromfile function wants an actual file.</p>\n\n<p>client = storage.Client()<br>\nbucket = client.get_bucket('tsa')<br>\nblob = storage.Blob( 'stage1/aps/342656834021fa0186a369f9652e73eb.aps', bucket )<br>\nfid = BytesIO(blob.download_as_string())<br></p>",
          "rawMarkdown": "The code below gets the data, but the fromfile function wants an actual file.\n\nclient = storage.Client()<br>\nbucket = client.get_bucket('tsa')<br>\nblob = storage.Blob( 'stage1/aps/342656834021fa0186a369f9652e73eb.aps', bucket )<br>\nfid = BytesIO(blob.download_as_string())<br>\n"
        }
      ]
    },
    {
      "id": 240555,
      "postDate": "2017-11-06T21:25:12.113Z",
      "content": "<p>Hi, Is there any reason that my request to join the google group denied!? I requested again so please give me access to the group!</p>\n\n<blockquote>\n  <p>Unfortunately, your request to join the Kaggle Tsa Screening Challenge group was denied by a group moderator. </p>\n</blockquote>",
      "rawMarkdown": "Hi, Is there any reason that my request to join the google group denied!? I requested again so please give me access to the group!\n\n&gt; Unfortunately, your request to join the Kaggle Tsa Screening Challenge group was denied by a group moderator. \n\n\n",
      "replies": [
        {
          "id": 240583,
          "postDate": "2017-11-06T22:55:27.563Z",
          "content": "<p>Hi Cyrus,</p>\n\n<p>Read the rules of acceptance carefully. Thanks!</p>",
          "rawMarkdown": "Hi Cyrus,\n\nRead the rules of acceptance carefully. Thanks!"
        },
        {
          "id": 241032,
          "postDate": "2017-11-07T21:53:34.313Z",
          "content": "<p>Addison, I read the rule and since I'm qualified joined to the competition but still getting denied to join to google group. Can you please refer me to the reason I'm being denied?</p>",
          "rawMarkdown": "Addison, I read the rule and since I'm qualified joined to the competition but still getting denied to join to google group. Can you please refer me to the reason I'm being denied?"
        },
        {
          "id": 241056,
          "postDate": "2017-11-07T23:41:03.650Z",
          "content": "<p>Hi Cyrus,</p>\n\n<p>The rules contain specific instructions to indicate you have fully read and acknowledged them when applying for the Google Group. I hope that helps!</p>",
          "rawMarkdown": "Hi Cyrus,\n\nThe rules contain specific instructions to indicate you have fully read and acknowledged them when applying for the Google Group. I hope that helps!"
        }
      ]
    },
    {
      "id": 234024,
      "postDate": "2017-10-22T03:39:51.147Z",
      "content": "<p>Yufeng, you mentioned that we can access 'kaggle-tsa-stage1' bucket directly without copying the data.  I created a compute instance and can browse and copy individual files using 'gsutil'. However if I tried mounting the bucket with:</p>\n\n<pre><code>gcsfuse kaggle-tsa-stage1 /home/serg14/kaggle-tsa-stage1\n</code></pre>\n\n<p>my mount location would contain only two files 'stage1_sample_submission.csv' and 'stage1_labels.csv' with no visibility of 'stage1' directory.</p>\n\n<p>I also tried using python interface to access 'kaggle-tsa-stage1' bucket, but got the same error message as DouglasBear earlier:\ngoogle.api.core.exceptions.Forbidden: 403 GET <a href=\"https://www.googleapis.com/storage/v1/b/kaggle-tsa-stage1?projection=noAcl\">https://www.googleapis.com/storage/v1/b/kaggle-tsa-stage1?projection=noAcl</a>: myaccount@gmail.com does not have storage.buckets.get access to kaggle-tsa-stage1</p>\n\n<p>Can you maybe give a step-by-step guide as to how to access the images from the bucket without copying them? Would greatly appreciate it.</p>",
      "rawMarkdown": "Yufeng, you mentioned that we can access 'kaggle-tsa-stage1' bucket directly without copying the data.  I created a compute instance and can browse and copy individual files using 'gsutil'. However if I tried mounting the bucket with:\n\n    gcsfuse kaggle-tsa-stage1 /home/serg14/kaggle-tsa-stage1\n\nmy mount location would contain only two files 'stage1_sample_submission.csv' and 'stage1_labels.csv' with no visibility of 'stage1' directory.\n\nI also tried using python interface to access 'kaggle-tsa-stage1' bucket, but got the same error message as DouglasBear earlier:\ngoogle.api.core.exceptions.Forbidden: 403 GET https://www.googleapis.com/storage/v1/b/kaggle-tsa-stage1?projection=noAcl: myaccount@gmail.com does not have storage.buckets.get access to kaggle-tsa-stage1\n\nCan you maybe give a step-by-step guide as to how to access the images from the bucket without copying them? Would greatly appreciate it.",
      "replies": [
        {
          "id": 234851,
          "postDate": "2017-10-24T07:43:57.807Z",
          "content": "<p>I'm not familiar with <code>gcsfuse</code> , but can you try mounting the stage1 subdirectory directly?</p>\n\n<p>Something like</p>\n\n<pre><code>gcsfuse kaggle-tsa-stage1/stage1 /home/serg14/kaggle-tsa-stage1\n</code></pre>\n\n<p>It's odd that you can see the 2 files in the top level dir but not any further. </p>\n\n<p>On using the python interface, it looks like you haven't successfully authenticated using your email. If you are doing this from a compute engine instance, it's possible that you are logged in through the service account, rather than the account that was added to the Google Group that controls the access. </p>\n\n<p>You should also confirm that the email address that is being denied is the same as that which you used to successfully access and copy out data originally. If you are getting a similar email address shown as DouglasBear, then the solution of doing <code>gcloud auth login</code> should do the trick. Make sure you use the account that is in the Google Group.</p>\n\n<p>Which Python client are you using? I would recommend using <a href=\"https://googlecloudplatform.github.io/google-cloud-python/\">google-cloud-python</a>, or if you only want the storage part, <a href=\"https://github.com/GoogleCloudPlatform/google-cloud-python/tree/master/storage\">google-cloud-storage</a>. This library is meant to help make auth as painless as possible (among other things). </p>",
          "rawMarkdown": "I'm not familiar with `gcsfuse` , but can you try mounting the stage1 subdirectory directly?\n\nSomething like\n\n    gcsfuse kaggle-tsa-stage1/stage1 /home/serg14/kaggle-tsa-stage1\n\nIt's odd that you can see the 2 files in the top level dir but not any further. \n\nOn using the python interface, it looks like you haven't successfully authenticated using your email. If you are doing this from a compute engine instance, it's possible that you are logged in through the service account, rather than the account that was added to the Google Group that controls the access. \n\nYou should also confirm that the email address that is being denied is the same as that which you used to successfully access and copy out data originally. If you are getting a similar email address shown as DouglasBear, then the solution of doing `gcloud auth login` should do the trick. Make sure you use the account that is in the Google Group.\n\nWhich Python client are you using? I would recommend using [google-cloud-python][1], or if you only want the storage part, [google-cloud-storage][2]. This library is meant to help make auth as painless as possible (among other things). \n\n\n  [1]: https://googlecloudplatform.github.io/google-cloud-python/\n  [2]: https://github.com/GoogleCloudPlatform/google-cloud-python/tree/master/storage"
        },
        {
          "id": 234978,
          "postDate": "2017-10-24T13:41:49.397Z",
          "content": "<p>Hi Yufeng. Thanks for looking into that! The problem  with 'gcsfuse' is that it can only accept bucket names as the first argument. I found somewhere on github that it is possible to mount with '--implicit-dirs' flag, but it led to completely different set of problems.</p>\n\n<p>I use generic ipython and played with 'gcloud auth login', with no much success. I admit I might not exhaust all options though. At the end I just copied the entire data set over to save time. I will give 'google-cloud-python' a try for the next project.</p>",
          "rawMarkdown": "Hi Yufeng. Thanks for looking into that! The problem  with 'gcsfuse' is that it can only accept bucket names as the first argument. I found somewhere on github that it is possible to mount with '--implicit-dirs' flag, but it led to completely different set of problems.\n\nI use generic ipython and played with 'gcloud auth login', with no much success. I admit I might not exhaust all options though. At the end I just copied the entire data set over to save time. I will give 'google-cloud-python' a try for the next project.",
          "votes": 1
        },
        {
          "id": 260104,
          "postDate": "2017-12-19T15:54:48.437Z",
          "content": "<p>The '--implicit-dirs' flag works but you need to set up all the directories in your GC bucket taht you want to use (read and write ).  After you copy the original tsadata into your own GC buckets run gcsfuse --- so if your GC bucket is named my-bucket-name and has directories stage1/aps and  tsa/stage2 you run: gcsfuse --implicit-dirs my-backet-name ./tsadata and it will mount the bucket into the directory \"/tsadata\" and the subdirectories will exist within /tsadata. A little late, but for next time - it took me some time to play with it and figure it out.</p>",
          "rawMarkdown": "The '--implicit-dirs' flag works but you need to set up all the directories in your GC bucket taht you want to use (read and write ).  After you copy the original tsadata into your own GC buckets run gcsfuse --- so if your GC bucket is named my-bucket-name and has directories stage1/aps and  tsa/stage2 you run: gcsfuse --implicit-dirs my-backet-name ./tsadata and it will mount the bucket into the directory \"/tsadata\" and the subdirectories will exist within /tsadata. A little late, but for next time - it took me some time to play with it and figure it out."
        }
      ]
    },
    {
      "id": 219347,
      "postDate": "2017-09-07T20:00:16.633Z",
      "content": "<p>I like it</p>",
      "rawMarkdown": "I like it"
    },
    {
      "id": 217138,
      "postDate": "2017-08-29T14:35:45.040Z",
      "content": "<p>Having similar issues to the other posts in this thread:</p>\n\n<p><code>sowl_burning@instance-1:~$ gsutil -m cp -r gs://kaggle-tsa-stage1/*  gs://sowl-burning-tsa\nAccessDeniedException: 403 270721026557-compute@developer.gserviceaccount.com does not have storage.objects.list access to bucket kaggle-tsa-stage1.\nCommandException: 1 file/object could not be transferred.</code></p>\n\n<p>I'm definitely in the Google Group and I can definitely see the data at <a href=\"https://storage.cloud.google.com/kaggle-tsa-stage1/\">https://storage.cloud.google.com/kaggle-tsa-stage1/</a>. I'm guessing I probably have my service account configured incorrectly (everything is default settings) but I have no idea how I would configure that stuff correctly. Any help would be greatly appreciated.</p>",
      "rawMarkdown": "Having similar issues to the other posts in this thread:\n\n```sowl_burning@instance-1:~$ gsutil -m cp -r gs://kaggle-tsa-stage1/*  gs://sowl-burning-tsa\nAccessDeniedException: 403 270721026557-compute@developer.gserviceaccount.com does not have storage.objects.list access to bucket kaggle-tsa-stage1.\nCommandException: 1 file/object could not be transferred.```\n\nI'm definitely in the Google Group and I can definitely see the data at https://storage.cloud.google.com/kaggle-tsa-stage1/. I'm guessing I probably have my service account configured incorrectly (everything is default settings) but I have no idea how I would configure that stuff correctly. Any help would be greatly appreciated.",
      "replies": [
        {
          "id": 217251,
          "postDate": "2017-08-29T23:12:36.463Z",
          "content": "<p>The email enroll in the google group is not your service account, but rather it is the email that you are logged into the Cloud Console with. You have 2 options:</p>\n\n<ol>\n<li>Either run <code>gsutil</code> from the command line on your local machine (if you have that set up) where you have the email address that you've signed in as set to the same one as the one in the google group, or</li>\n<li>Change the logged in user at the command line of the cloud console (It appears you are running the command currently from there). You can do this by running <code>gcloud auth login</code> and oauth-ing the account that you're enrolled into the google group with.</li>\n</ol>",
          "rawMarkdown": "The email enroll in the google group is not your service account, but rather it is the email that you are logged into the Cloud Console with. You have 2 options:\n\n1. Either run `gsutil` from the command line on your local machine (if you have that set up) where you have the email address that you've signed in as set to the same one as the one in the google group, or\n2. Change the logged in user at the command line of the cloud console (It appears you are running the command currently from there). You can do this by running `gcloud auth login` and oauth-ing the account that you're enrolled into the google group with.",
          "votes": 1
        },
        {
          "id": 218300,
          "postDate": "2017-09-03T14:26:57.280Z",
          "content": "<p>Thanks Yufeng. That worked! Now I'm having another problem that's only tangentially related, but I don't know if it's worth making a separate thread about. </p>\n\n<p>I'm having trouble creating a VM instance with a GPU attached. I initially tried to create an instance with a GPU but the process failed since my GPU quota was 0 as I was under a free trial. I upgraded my account from a free trial specifically for this purpose but the VM creation failed again since my quota was still 0. Looking the issue up, I found that I was supposed to go to the quotas page at <a href=\"https://console.cloud.google.com/iam-admin/quotas\">https://console.cloud.google.com/iam-admin/quotas</a> and request an increase for my GPU quota. However, when going to that page and looking at \"All quotas-all services-all metrics-all regions\", there is no option pertaining to GPUs. Ctrl+F \"gpu\", Ctrl+F \"nvidia\", Ctrl+F \"k80\" are all coming back empty. Any help resolving this would be greatly appreciated.</p>",
          "rawMarkdown": "Thanks Yufeng. That worked! Now I'm having another problem that's only tangentially related, but I don't know if it's worth making a separate thread about. \n\nI'm having trouble creating a VM instance with a GPU attached. I initially tried to create an instance with a GPU but the process failed since my GPU quota was 0 as I was under a free trial. I upgraded my account from a free trial specifically for this purpose but the VM creation failed again since my quota was still 0. Looking the issue up, I found that I was supposed to go to the quotas page at https://console.cloud.google.com/iam-admin/quotas and request an increase for my GPU quota. However, when going to that page and looking at \"All quotas-all services-all metrics-all regions\", there is no option pertaining to GPUs. Ctrl+F \"gpu\", Ctrl+F \"nvidia\", Ctrl+F \"k80\" are all coming back empty. Any help resolving this would be greatly appreciated."
        },
        {
          "id": 218771,
          "postDate": "2017-09-05T20:36:16.057Z",
          "content": "<p>Hi BurningOwl. I work at Google. So sorry! We're investigating this issue right now. I'll have an update tomorrow afternoon with details on how to set up GPUs on GCE.</p>",
          "rawMarkdown": "Hi BurningOwl. I work at Google. So sorry! We're investigating this issue right now. I'll have an update tomorrow afternoon with details on how to set up GPUs on GCE."
        },
        {
          "id": 219074,
          "postDate": "2017-09-06T22:13:00.797Z",
          "content": "<p>Hi BurningOwl, </p>\n\n<p>Can you check to see if those Quotas are visible on your IAM Admin page now? There was a problem on our side that has been fixed. </p>",
          "rawMarkdown": "Hi BurningOwl, \n\nCan you check to see if those Quotas are visible on your IAM Admin page now? There was a problem on our side that has been fixed. "
        },
        {
          "id": 219076,
          "postDate": "2017-09-06T22:22:10.293Z",
          "content": "<p>Yeah, it's there now. Thanks!</p>",
          "rawMarkdown": "Yeah, it's there now. Thanks!"
        }
      ]
    },
    {
      "id": 210993,
      "postDate": "2017-08-07T19:16:13.570Z",
      "content": "<p>I have access and can see the bucket's contents in my browser, yet:</p>\n\n<p><code>AccessDeniedException: 403 Caller does not have storage.objects.list access to bucket kaggle-tsa-stage1.\nCommandException: 1 file/object could not be transferred.</code></p>",
      "rawMarkdown": "I have access and can see the bucket's contents in my browser, yet:\n\n```AccessDeniedException: 403 Caller does not have storage.objects.list access to bucket kaggle-tsa-stage1.\nCommandException: 1 file/object could not be transferred.```",
      "replies": [
        {
          "id": 211755,
          "postDate": "2017-08-09T21:41:32.600Z",
          "content": "<p>Hi Joe, can you paste the command you ran?</p>",
          "rawMarkdown": "Hi Joe, can you paste the command you ran?"
        }
      ]
    },
    {
      "id": 209955,
      "postDate": "2017-08-03T21:24:27.443Z",
      "content": "<p>I got the error whey trying to copy the data, anything wrong?\nI have joined the group.</p>\n\n<p>gsutil -m cp -r gs://kaggle-tsa-stage1/*  gs://my-kaggle-data\nCopying gs://kaggle-tsa-stage1/stage1_labels.csv [Content-Type=text/csv]...\nCopying gs://kaggle-tsa-stage1/stage1_sample_submission.csv [Content-Type=text/csv]...\nCopying gs://kaggle-tsa-stage1/stage1/a3d/00360f79fd6e02781457eda48f85da90.a3d [Content-Type=application/octet-stream]...\nCopying gs://kaggle-tsa-stage1/stage1/a3d/0043db5e8c819bffc15261b1f1ac5e42.a3d [Content-Type=application/octet-stream]...\nCopying gs://kaggle-tsa-stage1/stage1/a3d/0050492f92e22eed3474ae3a6fc907fa.a3d [Content-Type=application/octet-stream]...\nAccessDeniedException: 403 Forbidden\nAccessDeniedException: 403 Forbidden\nAccessDeniedException: 403 Forbidden\nAccessDeniedException: 403 Forbidden</p>",
      "rawMarkdown": "I got the error whey trying to copy the data, anything wrong?\nI have joined the group.\n\n gsutil -m cp -r gs://kaggle-tsa-stage1/*  gs://my-kaggle-data\nCopying gs://kaggle-tsa-stage1/stage1_labels.csv [Content-Type=text/csv]...\nCopying gs://kaggle-tsa-stage1/stage1_sample_submission.csv [Content-Type=text/csv]...\nCopying gs://kaggle-tsa-stage1/stage1/a3d/00360f79fd6e02781457eda48f85da90.a3d [Content-Type=application/octet-stream]...\nCopying gs://kaggle-tsa-stage1/stage1/a3d/0043db5e8c819bffc15261b1f1ac5e42.a3d [Content-Type=application/octet-stream]...\nCopying gs://kaggle-tsa-stage1/stage1/a3d/0050492f92e22eed3474ae3a6fc907fa.a3d [Content-Type=application/octet-stream]...\nAccessDeniedException: 403 Forbidden\nAccessDeniedException: 403 Forbidden\nAccessDeniedException: 403 Forbidden\nAccessDeniedException: 403 Forbidden",
      "replies": [
        {
          "id": 209959,
          "postDate": "2017-08-03T21:40:02.030Z",
          "content": "<p>Hi Richard,</p>\n\n<p>I don't see your email in the Google Group - can you confirm you were accepted?</p>",
          "rawMarkdown": "Hi Richard,\n\nI don't see your email in the Google Group - can you confirm you were accepted?"
        },
        {
          "id": 210190,
          "postDate": "2017-08-04T16:04:34.367Z",
          "content": "<p>I used my gmail account to be added into the group, and I can see  I was approved and in the group:\n\"Kaggle Tsa Screening Challenge\nThis group grants access to the full Kaggle TSA dataset on Google Cloud Storage.\"</p>",
          "rawMarkdown": "I used my gmail account to be added into the group, and I can see  I was approved and in the group:\n\"Kaggle Tsa Screening Challenge\nThis group grants access to the full Kaggle TSA dataset on Google Cloud Storage.\"\n"
        },
        {
          "id": 210970,
          "postDate": "2017-08-07T18:20:25.033Z",
          "content": "<p>The command you're running is trying to copy it to the bucket called my-kaggle-data, which already exists (and is owned by me).\nYou'll need to point it to copy to your own bucket. </p>\n\n<p>Bucket names are <strong>globally</strong> unique (across all users of Cloud Storage), so gs://my-kaggle-data is already taken (by me when I made this guide, sorry). </p>\n\n<p>You should make a bucket in your region of choice and with the name of your choice. For example, something like this (might) work, assuming you want the data to live in the us-east1 zone:</p>\n\n<p><code>gsutil mb -c regional -l us-east1 gs://kaggle-data-richard</code></p>\n\n<p>Once your bucket is created, you should be able to copy the data into your newly created bucket (notice I've replaced <code>my-kaggle-data</code> with <code>kaggle-data-richard</code>):</p>\n\n<p><code>gsutil -m cp -r gs://kaggle-tsa-stage1/* gs://kaggle-data-richard</code></p>\n\n<p>Hope this works for you!</p>",
          "rawMarkdown": "The command you're running is trying to copy it to the bucket called my-kaggle-data, which already exists (and is owned by me).\nYou'll need to point it to copy to your own bucket. \n\nBucket names are **globally** unique (across all users of Cloud Storage), so gs://my-kaggle-data is already taken (by me when I made this guide, sorry). \n\nYou should make a bucket in your region of choice and with the name of your choice. For example, something like this (might) work, assuming you want the data to live in the us-east1 zone:\n\n`gsutil mb -c regional -l us-east1 gs://kaggle-data-richard`\n\nOnce your bucket is created, you should be able to copy the data into your newly created bucket (notice I've replaced `my-kaggle-data` with `kaggle-data-richard`):\n\n`gsutil -m cp -r gs://kaggle-tsa-stage1/* gs://kaggle-data-richard`\n\nHope this works for you!",
          "votes": 1
        },
        {
          "id": 225014,
          "postDate": "2017-09-28T03:50:53.113Z",
          "content": "<p>Hi, \nThis thread is a bit old, so I am looking for an update.\nQuestions: \n 1. Do we really have to copy all of the kaggle TSA data to our own Google Cloud bucket? If not, how to access it directly? \n 2. I have registered access to the Kaggle Cloud data on an email different from the one I use for Google Cloud. How can I re-register with the new email?</p>\n\n<p>Thanks!</p>",
          "rawMarkdown": "Hi, \nThis thread is a bit old, so I am looking for an update.\nQuestions: \n 1. Do we really have to copy all of the kaggle TSA data to our own Google Cloud bucket? If not, how to access it directly? \n 2. I have registered access to the Kaggle Cloud data on an email different from the one I use for Google Cloud. How can I re-register with the new email?\n\nThanks!",
          "votes": 1
        },
        {
          "id": 225248,
          "postDate": "2017-09-28T15:23:59.883Z",
          "content": "<p>You can see the files directly by going to the dataset while logged into the account that has been added to the Google Group. Of course, since you are using a different email, you should first join the group using the email you wish to use with Google Cloud. \nIt's fine to add your other account to the Google Group: just log in with that account and go to the group and join.</p>\n\n<p>If you wish you could also create a Google Cloud account using the email that you've already used to join the google group.</p>",
          "rawMarkdown": "You can see the files directly by going to the dataset while logged into the account that has been added to the Google Group. Of course, since you are using a different email, you should first join the group using the email you wish to use with Google Cloud. \nIt's fine to add your other account to the Google Group: just log in with that account and go to the group and join.\n\nIf you wish you could also create a Google Cloud account using the email that you've already used to join the google group."
        },
        {
          "id": 226374,
          "postDate": "2017-10-02T03:53:11.293Z",
          "content": "<p>Hi,\nI tried what you suggested. Became a member of the TSA data group on Google Cloud.\nRunning the following in Jupyter:</p>\n\n<p>from google.cloud import storage\nclient = storage.Client()\nbucket = client.get_bucket('kaggle-tsa-stage1')</p>\n\n<p>Gives me:\nForbidden: 403 832637061858-compute@developer.gserviceaccount.com does not have storage.buckets.get access to bucket kaggle-tsa-stage1. (GET <a href=\"https://www.googleapis.com/storage/v1/b/kaggle-tsa-stage1?projection=noAcl\">https://www.googleapis.com/storage/v1/b/kaggle-tsa-stage1?projection=noAcl</a>)</p>\n\n<p>What am I doing wrong? Does anyone have a working Python  code snippet I can compare to?\nThanks!</p>",
          "rawMarkdown": "Hi,\nI tried what you suggested. Became a member of the TSA data group on Google Cloud.\nRunning the following in Jupyter:\n\nfrom google.cloud import storage\nclient = storage.Client()\nbucket = client.get_bucket('kaggle-tsa-stage1')\n\nGives me:\nForbidden: 403 832637061858-compute@developer.gserviceaccount.com does not have storage.buckets.get access to bucket kaggle-tsa-stage1. (GET https://www.googleapis.com/storage/v1/b/kaggle-tsa-stage1?projection=noAcl)\n\nWhat am I doing wrong? Does anyone have a working Python  code snippet I can compare to?\nThanks!"
        },
        {
          "id": 227217,
          "postDate": "2017-10-04T00:13:19.977Z",
          "content": "<p>Hi Yufeng,\nI would really appreciate an answer to my question. You seem to be the expert here on accessing the Kaggle data on Google Cloud!  I have been in contact with Google support, and they have said that only the owner of the Kaggle Group can resolve my problem to access the data. Thanks.</p>",
          "rawMarkdown": "Hi Yufeng,\nI would really appreciate an answer to my question. You seem to be the expert here on accessing the Kaggle data on Google Cloud!  I have been in contact with Google support, and they have said that only the owner of the Kaggle Group can resolve my problem to access the data. Thanks."
        },
        {
          "id": 227218,
          "postDate": "2017-10-04T00:19:28.540Z",
          "content": "<p>it looks like you are executing the copy command from a GCE VM instance? As a result, it is logged in with the default service account (in your case, 832637061858-compute@developer.gserviceaccount.com). The service account is not (and cannot) be part of the google group.\nSince you are trying to access it through a jupyter notebook directly, I would recommend you authenticate with the account you used for membership to the google group for data access. \nYou can do this using <code>gcloud auth login</code> and sign in with your credentials that you used to join the Google Group.</p>\n\n<p>Hope that helps. Let me know if you run into any further issues!</p>",
          "rawMarkdown": "it looks like you are executing the copy command from a GCE VM instance? As a result, it is logged in with the default service account (in your case, 832637061858-compute@developer.gserviceaccount.com). The service account is not (and cannot) be part of the google group.\nSince you are trying to access it through a jupyter notebook directly, I would recommend you authenticate with the account you used for membership to the google group for data access. \nYou can do this using `gcloud auth login` and sign in with your credentials that you used to join the Google Group.\n\nHope that helps. Let me know if you run into any further issues!"
        }
      ]
    },
    {
      "id": 208403,
      "postDate": "2017-07-29T17:20:43.440Z",
      "content": "<p>gsutil -m cp -r gs://kaggle-tsa-stage1/* gs://pbwills1_tsa_bucket\nAccessDeniedException: 403 Caller does not have storage.objects.list access to bucket kaggle-tsa-stage1.\nCommandException: 1 file/object could not be transferred</p>\n\n<p>gsutil ls gs://kaggle-tsa-stage1\nAccessDeniedException: 403 Caller does not have storage.objects.list access to bucket kaggle-tsa-stage1.</p>\n\n<p>What am I doing wrong?</p>",
      "rawMarkdown": "gsutil -m cp -r gs://kaggle-tsa-stage1/* gs://pbwills1_tsa_bucket\nAccessDeniedException: 403 Caller does not have storage.objects.list access to bucket kaggle-tsa-stage1.\nCommandException: 1 file/object could not be transferred\n\ngsutil ls gs://kaggle-tsa-stage1\nAccessDeniedException: 403 Caller does not have storage.objects.list access to bucket kaggle-tsa-stage1.\n\nWhat am I doing wrong?",
      "replies": [
        {
          "id": 208416,
          "postDate": "2017-07-29T18:48:05.623Z",
          "content": "<p>OK, I see that one has to join the group. I will try this...</p>",
          "rawMarkdown": "OK, I see that one has to join the group. I will try this..."
        }
      ]
    },
    {
      "id": 387542,
      "postDate": "2018-09-15T06:30:01.320Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 219432,
      "author_name": "Muneeb",
      "author_url": "",
      "post_date": "2017-09-08T06:25:08.510000",
      "content": "<p>It adds a lot to the costs for each participant to copy the entire dataset into their own buckets. Why not give us Read access as recommended here for 'when a large dataset is shared across multiple projects':</p>\n\n<p><a href=\"https://cloud.google.com/ml-engine/docs/how-tos/working-with-data#using_a_cloud_storage_bucket_from_a_different_project\">Using a Cloud Storage Bucket from a Different Project</a></p>\n\n<p>All we would need is to share our 'service account information' string with you.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 219465,
          "author_name": "Kaz Sato",
          "author_url": "",
          "post_date": "2017-09-08T09:39:03.303000",
          "content": "<p>Hi Muneeb,</p>\n\n<p>In general use case of Cloud Storage, yes you can do it. But in the case of this competition,  \"Access to the Google Cloud Storage bucket is controlled by membership in a Google group\" as explained at the Data section. So you need to use your Google account to be a member of the group to get access to the bucket. </p>",
          "votes": -2,
          "replies": []
        },
        {
          "id": 219668,
          "author_name": "Muneeb",
          "author_url": "",
          "post_date": "2017-09-09T01:58:29.867000",
          "content": "<p>Yes I already have 'access' to the bucket per se because of Group membership, but to use it in the Google Cloud ML Engine project for training purposes, you need special permissions as detailed in the link I shared above.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 219676,
          "author_name": "Kaz Sato",
          "author_url": "",
          "post_date": "2017-09-09T02:48:47.267000",
          "content": "<p>I see, were you able to use ML Engine by granting access to its service account?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 219709,
          "author_name": "Muneeb",
          "author_url": "",
          "post_date": "2017-09-09T08:53:27.873000",
          "content": "<p>I don't have admin access to the tsa bucket so I cannot modify access control for that bucket to allow my ML Engine project read access. I assume only owners of kaggle-tsa-stage1 bucket can allow access. It could be implemented simply by making a Google form and sharing it on the Google Group for this bucket. Anyone requiring access can then submit their 'service account info' string via the form.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 219733,
          "author_name": "Kaz Sato",
          "author_url": "",
          "post_date": "2017-09-09T11:38:16.120000",
          "content": "<p>I agree, it would be much easier if we had an operation practice that allows participants using service accounts for accessing the bucket directly.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 227346,
          "author_name": "Muneeb",
          "author_url": "",
          "post_date": "2017-10-04T07:09:28.970000",
          "content": "<p>Hi Yufeng, you think you can help here on avoiding the copying process like above?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 227563,
          "author_name": "Yufeng Guo",
          "author_url": "",
          "post_date": "2017-10-04T17:13:01.723000",
          "content": "<p>Unfortunately, this competition can't grant general read-access to the dataset, but by having access to the dataset (via the Google Group), you can read data from it in any situation where you are authenticated through that same account.</p>\n\n<p>For example, if your email is kaggle_user@gmail.com and you've joined the Google group with that email, you can access the data directly from a cloud account that is signed in as kaggle_user@gmail.com. This means that you <em>can</em> read the data from <a href=\"https://cloud.google.com/solutions/running-distributed-tensorflow-on-compute-engine\">Google Compute Engine</a> or <a href=\"https://medium.com/google-cloud/jupyter-tensorflow-nvidia-gpu-docker-google-compute-engine-4a146f085f17\">Google Container Engine</a>, so if run your TensorFlow training from there, then you won't need to copy the data. If you want to download the data to your own cluster directly, you can do so as well using <code>gsutil -m cp &lt;source&gt; &lt;destination&gt;</code> while ssh'd into your cluster.</p>\n\n<p>The problem arises <em>only</em> when you are using Cloud Machine Learning Engine, as it utilizes a service account to access the dataset rather than directly using your signed-in credentials. The reason for this is that this allows the service to manage your resources on your behalf (which is the whole point), rather than having it exposed to you to manage that infrastructure layer. For this, you will need to copy the data out into your own bucket once you have access via the Google Group, as the data was required to be secured to some degree by the competition rules.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 228214,
          "author_name": "Muneeb",
          "author_url": "",
          "post_date": "2017-10-06T05:09:10.757000",
          "content": "<p>I understand, but there is a method detailed in the link I shared (<a href=\"https://cloud.google.com/ml-engine/docs/how-tos/working-with-data#using_a_cloud_storage_bucket_from_a_different_project\">https://cloud.google.com/ml-engine/docs/how-tos/working-with-data#using_a_cloud_storage_bucket_from_a_different_project</a>), to get access via service account. There would just be one extra step required to collect one service account string per Google Group member using a form or a thread etc.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 233378,
          "author_name": "Jake Shi",
          "author_url": "",
          "post_date": "2017-10-20T01:28:36.980000",
          "content": "<p>hi Yufeng, thank you very much. This is very helpful. Could you compare the pro and cons for :\n1. using  Google Compute Engine or Google Container Engine \nvs.\n2. using Cloud Machine Learning Engine</p>\n\n<p>Based on your comments it seems using option 1 saves you from copying the data. \nBut based on the names itself, seems option 2 has ML utilities build in already. \nAny other considerations? What which would you recommend?</p>\n\n<p>Thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 234842,
          "author_name": "Yufeng Guo",
          "author_url": "",
          "post_date": "2017-10-24T07:29:11.370000",
          "content": "<p>Using Cloud ML Engine allows you to run your TF code in a serverless/on-demand manner. You only pay for what you use, and it's trivial to scale up your model to multiple machines/GPUs if you are using tf.estimator framework.</p>\n\n<p>Using GCE/GKE means you will be configuring everything yourself, and pay for machines you allocate for the time you have them running. You also would need to configure your own cluster for scaling up to more than one machine. See <a href=\"https://www.tensorflow.org/versions/r0.12/api_docs/python/train/distributed_execution\">here</a> for more info.\nIf this is not a concern, then there are no other notable downsides.</p>\n\n<p>One idea would be to copy a subset of the dataset to test with using either your local environment or using the Cloud ML Engine environment to see how you like it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 234859,
          "author_name": "Yufeng Guo",
          "author_url": "",
          "post_date": "2017-10-24T08:06:59.597000",
          "content": "<blockquote>\n  <p><strong>Muneeb wrote</strong></p>\n  \n  <blockquote>\n    <p>I understand, but there is a method detailed in the link I shared (<a href=\"https://cloud.google.com/ml-engine/docs/how-tos/working-with-data#using_a_cloud_storage_bucket_from_a_different_project\">https://cloud.google.com/ml-engine/docs/how-tos/working-with-data#using_a_cloud_storage_bucket_from_a_different_project</a>), to get access via service account. There would just be one extra step required to collect one service account string per Google Group member using a form or a thread etc.</p>\n  </blockquote>\n</blockquote>\n\n<p>Unfortunately, Google Groups does not allow service accounts to be added as members, otherwise this would be a great solution :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 254817,
      "author_name": "John Duffy",
      "author_url": "",
      "post_date": "2017-12-07T17:27:33.070000",
      "content": "<p>how do you access the files in the bucket?  I get \"No such file or directory 'gs://tsa//stage1/aps/342656834021fa0186a369f9652e73eb.aps'  when calling:  <br>\n <code>open('gs://tsa//stage1/aps/342656834021fa0186a369f9652e73eb.aps', 'r+b')</code></p>",
      "votes": 0,
      "replies": [
        {
          "id": 254945,
          "author_name": "John Duffy",
          "author_url": "",
          "post_date": "2017-12-07T22:44:48.400000",
          "content": "<p>The code below gets the data, but the fromfile function wants an actual file.</p>\n\n<p>client = storage.Client()<br>\nbucket = client.get_bucket('tsa')<br>\nblob = storage.Blob( 'stage1/aps/342656834021fa0186a369f9652e73eb.aps', bucket )<br>\nfid = BytesIO(blob.download_as_string())<br></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 240555,
      "author_name": "Cyrus",
      "author_url": "",
      "post_date": "2017-11-06T21:25:12.113000",
      "content": "<p>Hi, Is there any reason that my request to join the google group denied!? I requested again so please give me access to the group!</p>\n\n<blockquote>\n  <p>Unfortunately, your request to join the Kaggle Tsa Screening Challenge group was denied by a group moderator. </p>\n</blockquote>",
      "votes": 0,
      "replies": [
        {
          "id": 240583,
          "author_name": "Addison Howard",
          "author_url": "",
          "post_date": "2017-11-06T22:55:27.563000",
          "content": "<p>Hi Cyrus,</p>\n\n<p>Read the rules of acceptance carefully. Thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 241032,
          "author_name": "Cyrus",
          "author_url": "",
          "post_date": "2017-11-07T21:53:34.313000",
          "content": "<p>Addison, I read the rule and since I'm qualified joined to the competition but still getting denied to join to google group. Can you please refer me to the reason I'm being denied?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 241056,
          "author_name": "Addison Howard",
          "author_url": "",
          "post_date": "2017-11-07T23:41:03.650000",
          "content": "<p>Hi Cyrus,</p>\n\n<p>The rules contain specific instructions to indicate you have fully read and acknowledged them when applying for the Google Group. I hope that helps!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 234024,
      "author_name": "serg14",
      "author_url": "",
      "post_date": "2017-10-22T03:39:51.147000",
      "content": "<p>Yufeng, you mentioned that we can access 'kaggle-tsa-stage1' bucket directly without copying the data.  I created a compute instance and can browse and copy individual files using 'gsutil'. However if I tried mounting the bucket with:</p>\n\n<pre><code>gcsfuse kaggle-tsa-stage1 /home/serg14/kaggle-tsa-stage1\n</code></pre>\n\n<p>my mount location would contain only two files 'stage1_sample_submission.csv' and 'stage1_labels.csv' with no visibility of 'stage1' directory.</p>\n\n<p>I also tried using python interface to access 'kaggle-tsa-stage1' bucket, but got the same error message as DouglasBear earlier:\ngoogle.api.core.exceptions.Forbidden: 403 GET <a href=\"https://www.googleapis.com/storage/v1/b/kaggle-tsa-stage1?projection=noAcl\">https://www.googleapis.com/storage/v1/b/kaggle-tsa-stage1?projection=noAcl</a>: myaccount@gmail.com does not have storage.buckets.get access to kaggle-tsa-stage1</p>\n\n<p>Can you maybe give a step-by-step guide as to how to access the images from the bucket without copying them? Would greatly appreciate it.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 234851,
          "author_name": "Yufeng Guo",
          "author_url": "",
          "post_date": "2017-10-24T07:43:57.807000",
          "content": "<p>I'm not familiar with <code>gcsfuse</code> , but can you try mounting the stage1 subdirectory directly?</p>\n\n<p>Something like</p>\n\n<pre><code>gcsfuse kaggle-tsa-stage1/stage1 /home/serg14/kaggle-tsa-stage1\n</code></pre>\n\n<p>It's odd that you can see the 2 files in the top level dir but not any further. </p>\n\n<p>On using the python interface, it looks like you haven't successfully authenticated using your email. If you are doing this from a compute engine instance, it's possible that you are logged in through the service account, rather than the account that was added to the Google Group that controls the access. </p>\n\n<p>You should also confirm that the email address that is being denied is the same as that which you used to successfully access and copy out data originally. If you are getting a similar email address shown as DouglasBear, then the solution of doing <code>gcloud auth login</code> should do the trick. Make sure you use the account that is in the Google Group.</p>\n\n<p>Which Python client are you using? I would recommend using <a href=\"https://googlecloudplatform.github.io/google-cloud-python/\">google-cloud-python</a>, or if you only want the storage part, <a href=\"https://github.com/GoogleCloudPlatform/google-cloud-python/tree/master/storage\">google-cloud-storage</a>. This library is meant to help make auth as painless as possible (among other things). </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 234978,
          "author_name": "serg14",
          "author_url": "",
          "post_date": "2017-10-24T13:41:49.397000",
          "content": "<p>Hi Yufeng. Thanks for looking into that! The problem  with 'gcsfuse' is that it can only accept bucket names as the first argument. I found somewhere on github that it is possible to mount with '--implicit-dirs' flag, but it led to completely different set of problems.</p>\n\n<p>I use generic ipython and played with 'gcloud auth login', with no much success. I admit I might not exhaust all options though. At the end I just copied the entire data set over to save time. I will give 'google-cloud-python' a try for the next project.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 260104,
          "author_name": "FrankWo",
          "author_url": "",
          "post_date": "2017-12-19T15:54:48.437000",
          "content": "<p>The '--implicit-dirs' flag works but you need to set up all the directories in your GC bucket taht you want to use (read and write ).  After you copy the original tsadata into your own GC buckets run gcsfuse --- so if your GC bucket is named my-bucket-name and has directories stage1/aps and  tsa/stage2 you run: gcsfuse --implicit-dirs my-backet-name ./tsadata and it will mount the bucket into the directory \"/tsadata\" and the subdirectories will exist within /tsadata. A little late, but for next time - it took me some time to play with it and figure it out.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 219347,
      "author_name": "",
      "author_url": "",
      "post_date": "2017-09-07T20:00:16.633000",
      "content": "<p>I like it</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 217138,
      "author_name": "BurningOwl",
      "author_url": "",
      "post_date": "2017-08-29T14:35:45.040000",
      "content": "<p>Having similar issues to the other posts in this thread:</p>\n\n<p><code>sowl_burning@instance-1:~$ gsutil -m cp -r gs://kaggle-tsa-stage1/*  gs://sowl-burning-tsa\nAccessDeniedException: 403 270721026557-compute@developer.gserviceaccount.com does not have storage.objects.list access to bucket kaggle-tsa-stage1.\nCommandException: 1 file/object could not be transferred.</code></p>\n\n<p>I'm definitely in the Google Group and I can definitely see the data at <a href=\"https://storage.cloud.google.com/kaggle-tsa-stage1/\">https://storage.cloud.google.com/kaggle-tsa-stage1/</a>. I'm guessing I probably have my service account configured incorrectly (everything is default settings) but I have no idea how I would configure that stuff correctly. Any help would be greatly appreciated.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 217251,
          "author_name": "Yufeng Guo",
          "author_url": "",
          "post_date": "2017-08-29T23:12:36.463000",
          "content": "<p>The email enroll in the google group is not your service account, but rather it is the email that you are logged into the Cloud Console with. You have 2 options:</p>\n\n<ol>\n<li>Either run <code>gsutil</code> from the command line on your local machine (if you have that set up) where you have the email address that you've signed in as set to the same one as the one in the google group, or</li>\n<li>Change the logged in user at the command line of the cloud console (It appears you are running the command currently from there). You can do this by running <code>gcloud auth login</code> and oauth-ing the account that you're enrolled into the google group with.</li>\n</ol>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 218300,
          "author_name": "BurningOwl",
          "author_url": "",
          "post_date": "2017-09-03T14:26:57.280000",
          "content": "<p>Thanks Yufeng. That worked! Now I'm having another problem that's only tangentially related, but I don't know if it's worth making a separate thread about. </p>\n\n<p>I'm having trouble creating a VM instance with a GPU attached. I initially tried to create an instance with a GPU but the process failed since my GPU quota was 0 as I was under a free trial. I upgraded my account from a free trial specifically for this purpose but the VM creation failed again since my quota was still 0. Looking the issue up, I found that I was supposed to go to the quotas page at <a href=\"https://console.cloud.google.com/iam-admin/quotas\">https://console.cloud.google.com/iam-admin/quotas</a> and request an increase for my GPU quota. However, when going to that page and looking at \"All quotas-all services-all metrics-all regions\", there is no option pertaining to GPUs. Ctrl+F \"gpu\", Ctrl+F \"nvidia\", Ctrl+F \"k80\" are all coming back empty. Any help resolving this would be greatly appreciated.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 218771,
          "author_name": "Steve Greenberg",
          "author_url": "",
          "post_date": "2017-09-05T20:36:16.057000",
          "content": "<p>Hi BurningOwl. I work at Google. So sorry! We're investigating this issue right now. I'll have an update tomorrow afternoon with details on how to set up GPUs on GCE.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 219074,
          "author_name": "Steve Greenberg",
          "author_url": "",
          "post_date": "2017-09-06T22:13:00.797000",
          "content": "<p>Hi BurningOwl, </p>\n\n<p>Can you check to see if those Quotas are visible on your IAM Admin page now? There was a problem on our side that has been fixed. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 219076,
          "author_name": "BurningOwl",
          "author_url": "",
          "post_date": "2017-09-06T22:22:10.293000",
          "content": "<p>Yeah, it's there now. Thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 210993,
      "author_name": "Joe Davison",
      "author_url": "",
      "post_date": "2017-08-07T19:16:13.570000",
      "content": "<p>I have access and can see the bucket's contents in my browser, yet:</p>\n\n<p><code>AccessDeniedException: 403 Caller does not have storage.objects.list access to bucket kaggle-tsa-stage1.\nCommandException: 1 file/object could not be transferred.</code></p>",
      "votes": 0,
      "replies": [
        {
          "id": 211755,
          "author_name": "Yufeng Guo",
          "author_url": "",
          "post_date": "2017-08-09T21:41:32.600000",
          "content": "<p>Hi Joe, can you paste the command you ran?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 209955,
      "author_name": "Richard",
      "author_url": "",
      "post_date": "2017-08-03T21:24:27.443000",
      "content": "<p>I got the error whey trying to copy the data, anything wrong?\nI have joined the group.</p>\n\n<p>gsutil -m cp -r gs://kaggle-tsa-stage1/*  gs://my-kaggle-data\nCopying gs://kaggle-tsa-stage1/stage1_labels.csv [Content-Type=text/csv]...\nCopying gs://kaggle-tsa-stage1/stage1_sample_submission.csv [Content-Type=text/csv]...\nCopying gs://kaggle-tsa-stage1/stage1/a3d/00360f79fd6e02781457eda48f85da90.a3d [Content-Type=application/octet-stream]...\nCopying gs://kaggle-tsa-stage1/stage1/a3d/0043db5e8c819bffc15261b1f1ac5e42.a3d [Content-Type=application/octet-stream]...\nCopying gs://kaggle-tsa-stage1/stage1/a3d/0050492f92e22eed3474ae3a6fc907fa.a3d [Content-Type=application/octet-stream]...\nAccessDeniedException: 403 Forbidden\nAccessDeniedException: 403 Forbidden\nAccessDeniedException: 403 Forbidden\nAccessDeniedException: 403 Forbidden</p>",
      "votes": 0,
      "replies": [
        {
          "id": 209959,
          "author_name": "Addison Howard",
          "author_url": "",
          "post_date": "2017-08-03T21:40:02.030000",
          "content": "<p>Hi Richard,</p>\n\n<p>I don't see your email in the Google Group - can you confirm you were accepted?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 210190,
          "author_name": "Richard",
          "author_url": "",
          "post_date": "2017-08-04T16:04:34.367000",
          "content": "<p>I used my gmail account to be added into the group, and I can see  I was approved and in the group:\n\"Kaggle Tsa Screening Challenge\nThis group grants access to the full Kaggle TSA dataset on Google Cloud Storage.\"</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 210970,
          "author_name": "Yufeng Guo",
          "author_url": "",
          "post_date": "2017-08-07T18:20:25.033000",
          "content": "<p>The command you're running is trying to copy it to the bucket called my-kaggle-data, which already exists (and is owned by me).\nYou'll need to point it to copy to your own bucket. </p>\n\n<p>Bucket names are <strong>globally</strong> unique (across all users of Cloud Storage), so gs://my-kaggle-data is already taken (by me when I made this guide, sorry). </p>\n\n<p>You should make a bucket in your region of choice and with the name of your choice. For example, something like this (might) work, assuming you want the data to live in the us-east1 zone:</p>\n\n<p><code>gsutil mb -c regional -l us-east1 gs://kaggle-data-richard</code></p>\n\n<p>Once your bucket is created, you should be able to copy the data into your newly created bucket (notice I've replaced <code>my-kaggle-data</code> with <code>kaggle-data-richard</code>):</p>\n\n<p><code>gsutil -m cp -r gs://kaggle-tsa-stage1/* gs://kaggle-data-richard</code></p>\n\n<p>Hope this works for you!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 225014,
          "author_name": "DouglasBear",
          "author_url": "",
          "post_date": "2017-09-28T03:50:53.113000",
          "content": "<p>Hi, \nThis thread is a bit old, so I am looking for an update.\nQuestions: \n 1. Do we really have to copy all of the kaggle TSA data to our own Google Cloud bucket? If not, how to access it directly? \n 2. I have registered access to the Kaggle Cloud data on an email different from the one I use for Google Cloud. How can I re-register with the new email?</p>\n\n<p>Thanks!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 225248,
          "author_name": "Yufeng Guo",
          "author_url": "",
          "post_date": "2017-09-28T15:23:59.883000",
          "content": "<p>You can see the files directly by going to the dataset while logged into the account that has been added to the Google Group. Of course, since you are using a different email, you should first join the group using the email you wish to use with Google Cloud. \nIt's fine to add your other account to the Google Group: just log in with that account and go to the group and join.</p>\n\n<p>If you wish you could also create a Google Cloud account using the email that you've already used to join the google group.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 226374,
          "author_name": "DouglasBear",
          "author_url": "",
          "post_date": "2017-10-02T03:53:11.293000",
          "content": "<p>Hi,\nI tried what you suggested. Became a member of the TSA data group on Google Cloud.\nRunning the following in Jupyter:</p>\n\n<p>from google.cloud import storage\nclient = storage.Client()\nbucket = client.get_bucket('kaggle-tsa-stage1')</p>\n\n<p>Gives me:\nForbidden: 403 832637061858-compute@developer.gserviceaccount.com does not have storage.buckets.get access to bucket kaggle-tsa-stage1. (GET <a href=\"https://www.googleapis.com/storage/v1/b/kaggle-tsa-stage1?projection=noAcl\">https://www.googleapis.com/storage/v1/b/kaggle-tsa-stage1?projection=noAcl</a>)</p>\n\n<p>What am I doing wrong? Does anyone have a working Python  code snippet I can compare to?\nThanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 227217,
          "author_name": "DouglasBear",
          "author_url": "",
          "post_date": "2017-10-04T00:13:19.977000",
          "content": "<p>Hi Yufeng,\nI would really appreciate an answer to my question. You seem to be the expert here on accessing the Kaggle data on Google Cloud!  I have been in contact with Google support, and they have said that only the owner of the Kaggle Group can resolve my problem to access the data. Thanks.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 227218,
          "author_name": "Yufeng Guo",
          "author_url": "",
          "post_date": "2017-10-04T00:19:28.540000",
          "content": "<p>it looks like you are executing the copy command from a GCE VM instance? As a result, it is logged in with the default service account (in your case, 832637061858-compute@developer.gserviceaccount.com). The service account is not (and cannot) be part of the google group.\nSince you are trying to access it through a jupyter notebook directly, I would recommend you authenticate with the account you used for membership to the google group for data access. \nYou can do this using <code>gcloud auth login</code> and sign in with your credentials that you used to join the Google Group.</p>\n\n<p>Hope that helps. Let me know if you run into any further issues!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 208403,
      "author_name": "pete",
      "author_url": "",
      "post_date": "2017-07-29T17:20:43.440000",
      "content": "<p>gsutil -m cp -r gs://kaggle-tsa-stage1/* gs://pbwills1_tsa_bucket\nAccessDeniedException: 403 Caller does not have storage.objects.list access to bucket kaggle-tsa-stage1.\nCommandException: 1 file/object could not be transferred</p>\n\n<p>gsutil ls gs://kaggle-tsa-stage1\nAccessDeniedException: 403 Caller does not have storage.objects.list access to bucket kaggle-tsa-stage1.</p>\n\n<p>What am I doing wrong?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 208416,
          "author_name": "pete",
          "author_url": "",
          "post_date": "2017-07-29T18:48:05.623000",
          "content": "<p>OK, I see that one has to join the group. I will try this...</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 387542,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-09-15T06:30:01.320000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "207582": "Due to access restrictions, you will need to copy your data from the kaggle-tsa-stage1 bucket to your own [Google Cloud Storage][1] (GCS) bucket in order to access the data using Google Cloud. \n\nThe easiest way to do this is to use the [gcloud command-line tool][2], which includes a tool called [gsutil][3], which is used for managing your Google Cloud Storage data. You can also accomplish this copy using the UI if you are more comfortable with that. Check out this YouTube video for details: https://youtu.be/wo--uiuPsiU\n\nCommand line access \n===\n\nGCS uses buckets to represent data and access. \nIf you don't have an appropriate GCS bucket to put your data into, you should [create one][4] using \n\n    gsutil mb -c regional -l us-east1 gs://my-kaggle-data\n(\"mb\" is short for \"make bucket\") \n\nBucket names are globally unique across all of gcs, so you will need to change the command to use your own unique bucket name. \n\nFor the `-l` argument, you should choose to place your data in a region that contains GPUs if you plan to use GPUs in your training. Currently GPUs are only available in the following regions:\n\n - us-east1 \n - us-central1 \n - asia-east1 \n - europe-west1\n\nYou can learn more about regions [here][5]. \n\nThe command we will be interested in using is the copy command, which in this case we will use with a recursive copy flag (`-r`) and run in parallel (`-m`):\n\n    gsutil -m cp -r gs://kaggle-tsa-stage1/*  gs://my-kaggle-data\n(the \"gs\" filesystem prefix is short for \"Google Storage\") \n\nThis will grab all the data from the kaggle-tsa-stage1 bucket and put it into your own bucket. Since the dataset is quite large, this copy operation will take some time. In our experience, it has taken on the order of 30 minutes. \n\n\nOnce the copy operation is complete, use gsutil ls to confirm the successful copy operation:\n\n    gsutil ls gs://my-kaggle-data\n\nWrapping up \n====\n\nNow you are ready to use Cloud Machine Learning Engine to access your data, and utilize the power of distributed GPUs! Check out our guide on using GPUs with Cloud Machine Learning Engine here. \n\n\n  [1]: https://cloud.google.com/storage/\n  [2]: https://cloud.google.com/sdk/docs/\n  [3]: https://cloud.google.com/storage/docs/gsutil\n  [4]: https://cloud.google.com/storage/docs/gsutil/commands/mb\n  [5]: https://cloud.google.com/compute/docs/regions-zones/regions-zones",
    "219432": "It adds a lot to the costs for each participant to copy the entire dataset into their own buckets. Why not give us Read access as recommended here for 'when a large dataset is shared across multiple projects':\n\n[Using a Cloud Storage Bucket from a Different Project][1]\n\nAll we would need is to share our 'service account information' string with you.\n\n\n  [1]: https://cloud.google.com/ml-engine/docs/how-tos/working-with-data#using_a_cloud_storage_bucket_from_a_different_project",
    "254817": "how do you access the files in the bucket?  I get \"No such file or directory 'gs://tsa//stage1/aps/342656834021fa0186a369f9652e73eb.aps'  when calling:  <br>\n `open('gs://tsa//stage1/aps/342656834021fa0186a369f9652e73eb.aps', 'r+b')`\n",
    "240555": "Hi, Is there any reason that my request to join the google group denied!? I requested again so please give me access to the group!\n\n&gt; Unfortunately, your request to join the Kaggle Tsa Screening Challenge group was denied by a group moderator. \n\n\n",
    "234024": "Yufeng, you mentioned that we can access 'kaggle-tsa-stage1' bucket directly without copying the data.  I created a compute instance and can browse and copy individual files using 'gsutil'. However if I tried mounting the bucket with:\n\n    gcsfuse kaggle-tsa-stage1 /home/serg14/kaggle-tsa-stage1\n\nmy mount location would contain only two files 'stage1_sample_submission.csv' and 'stage1_labels.csv' with no visibility of 'stage1' directory.\n\nI also tried using python interface to access 'kaggle-tsa-stage1' bucket, but got the same error message as DouglasBear earlier:\ngoogle.api.core.exceptions.Forbidden: 403 GET https://www.googleapis.com/storage/v1/b/kaggle-tsa-stage1?projection=noAcl: myaccount@gmail.com does not have storage.buckets.get access to kaggle-tsa-stage1\n\nCan you maybe give a step-by-step guide as to how to access the images from the bucket without copying them? Would greatly appreciate it.",
    "219347": "I like it",
    "217138": "Having similar issues to the other posts in this thread:\n\n```sowl_burning@instance-1:~$ gsutil -m cp -r gs://kaggle-tsa-stage1/*  gs://sowl-burning-tsa\nAccessDeniedException: 403 270721026557-compute@developer.gserviceaccount.com does not have storage.objects.list access to bucket kaggle-tsa-stage1.\nCommandException: 1 file/object could not be transferred.```\n\nI'm definitely in the Google Group and I can definitely see the data at https://storage.cloud.google.com/kaggle-tsa-stage1/. I'm guessing I probably have my service account configured incorrectly (everything is default settings) but I have no idea how I would configure that stuff correctly. Any help would be greatly appreciated.",
    "210993": "I have access and can see the bucket's contents in my browser, yet:\n\n```AccessDeniedException: 403 Caller does not have storage.objects.list access to bucket kaggle-tsa-stage1.\nCommandException: 1 file/object could not be transferred.```",
    "209955": "I got the error whey trying to copy the data, anything wrong?\nI have joined the group.\n\n gsutil -m cp -r gs://kaggle-tsa-stage1/*  gs://my-kaggle-data\nCopying gs://kaggle-tsa-stage1/stage1_labels.csv [Content-Type=text/csv]...\nCopying gs://kaggle-tsa-stage1/stage1_sample_submission.csv [Content-Type=text/csv]...\nCopying gs://kaggle-tsa-stage1/stage1/a3d/00360f79fd6e02781457eda48f85da90.a3d [Content-Type=application/octet-stream]...\nCopying gs://kaggle-tsa-stage1/stage1/a3d/0043db5e8c819bffc15261b1f1ac5e42.a3d [Content-Type=application/octet-stream]...\nCopying gs://kaggle-tsa-stage1/stage1/a3d/0050492f92e22eed3474ae3a6fc907fa.a3d [Content-Type=application/octet-stream]...\nAccessDeniedException: 403 Forbidden\nAccessDeniedException: 403 Forbidden\nAccessDeniedException: 403 Forbidden\nAccessDeniedException: 403 Forbidden",
    "208403": "gsutil -m cp -r gs://kaggle-tsa-stage1/* gs://pbwills1_tsa_bucket\nAccessDeniedException: 403 Caller does not have storage.objects.list access to bucket kaggle-tsa-stage1.\nCommandException: 1 file/object could not be transferred\n\ngsutil ls gs://kaggle-tsa-stage1\nAccessDeniedException: 403 Caller does not have storage.objects.list access to bucket kaggle-tsa-stage1.\n\nWhat am I doing wrong?",
    "387542": ""
  }
}