{
  "id": 56014,
  "title": "How to access data in gcloud bucket using jupyter notebook",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/56014",
  "author_name": "",
  "post_date": "2018-05-04T14:21:39.433669Z",
  "votes": 4,
  "comment_count": 33,
  "views": 0,
  "content": "<p>Hi there,</p>\n\n<p>I have connected my jupyter notebook server to gcloud. But I have problem loading csv file from gcloud bucket. Can anyone share some code about this? Thanks a lot.</p>\n\n<p>Best regards</p>",
  "messages": [
    {
      "id": "323159",
      "postDate": "05/04/2018 14:21:39",
      "content": "<p>Hi there,</p>\n\n<p>I have connected my jupyter notebook server to gcloud. But I have problem loading csv file from gcloud bucket. Can anyone share some code about this? Thanks a lot.</p>\n\n<p>Best regards</p>",
      "rawMarkdown": "Hi there,\n\nI have connected my jupyter notebook server to gcloud. But I have problem loading csv file from gcloud bucket. Can anyone share some code about this? Thanks a lot.\n\nBest regards",
      "votes": null
    },
    {
      "id": "323169",
      "postDate": "05/04/2018 14:42:52",
      "content": "<p>I have dealt with how to read a csv to dataframe. Now, the question becomes how to write a pandas dataframe to gcloud storage bucket. Thanks.</p>",
      "rawMarkdown": "I have dealt with how to read a csv to dataframe. Now, the question becomes how to write a pandas dataframe to gcloud storage bucket. Thanks.",
      "votes": null
    },
    {
      "id": "323178",
      "postDate": "05/04/2018 14:55:28",
      "content": "<p>Don't you prefer to install Jupyter on gcloud. Redirect everything to the port and view data in your browser on your computer?</p>",
      "rawMarkdown": "Don't you prefer to install Jupyter on gcloud. Redirect everything to the port and view data in your browser on your computer?",
      "votes": null
    },
    {
      "id": "323183",
      "postDate": "05/04/2018 14:59:10",
      "content": "<p>Hi SkalskiP,  thanks for reply. I don't perfectly understand what you said. But I think I have done something similar. I install Jupyter on gcloud and now I can run Jupyter notebook on gcloud. But I don't know how to write dataframe to my storage bucket.</p>",
      "rawMarkdown": "Hi SkalskiP,  thanks for reply. I don't perfectly understand what you said. But I think I have done something similar. I install Jupyter on gcloud and now I can run Jupyter notebook on gcloud. But I don't know how to write dataframe to my storage bucket.",
      "votes": null
    },
    {
      "id": "323196",
      "postDate": "05/04/2018 15:33:00",
      "content": "<p>I understand. Of course, I wanted something similar. Honestly, I do not know why you want to have it in bucket. Do you want to submit with gc?</p>",
      "rawMarkdown": "I understand. Of course, I wanted something similar. Honestly, I do not know why you want to have it in bucket. Do you want to submit with gc?",
      "votes": null
    },
    {
      "id": "323197",
      "postDate": "05/04/2018 15:37:04",
      "content": "<p>Because I assume after I upload my dataframe(e.g the prediction result dataframe in this competition) to bucket, I can download it to my local computer. Or is there any other way to accomplish this? Can I download the result dataframe directly to my local computer without uploading to bucket?</p>",
      "rawMarkdown": "Because I assume after I upload my dataframe(e.g the prediction result dataframe in this competition) to bucket, I can download it to my local computer. Or is there any other way to accomplish this? Can I download the result dataframe directly to my local computer without uploading to bucket?",
      "votes": null
    },
    {
      "id": "323202",
      "postDate": "05/04/2018 15:48:42",
      "content": "<p>Hm there are few ways of doing that. If you just want to upload submission to Kaggle, you can use Kaggle API. You can find it on GitHub. If you want to download it on your computer and you can use Linux as your local machina you can use scp command. Or you can save your file in some directory on server and than copy it to bucket and than download. </p>",
      "rawMarkdown": "Hm there are few ways of doing that. If you just want to upload submission to Kaggle, you can use Kaggle API. You can find it on GitHub. If you want to download it on your computer and you can use Linux as your local machina you can use scp command. Or you can save your file in some directory on server and than copy it to bucket and than download.",
      "votes": null
    },
    {
      "id": "323203",
      "postDate": "05/04/2018 15:51:27",
      "content": "<p>Thanks for reply! Can you explain more about the third way? I am a windows user. I really appreciate your help!</p>",
      "rawMarkdown": "Thanks for reply! Can you explain more about the third way? I am a windows user. I really appreciate your help!",
      "votes": null
    },
    {
      "id": "323207",
      "postDate": "05/04/2018 16:03:45",
      "content": "<p>How do you connect to your gc? Via SSH? </p>",
      "rawMarkdown": "How do you connect to your gc? Via SSH?",
      "votes": null
    },
    {
      "id": "323208",
      "postDate": "05/04/2018 16:05:34",
      "content": "<p>Yes, I connect Jupyter using <a href=\"http://cs231n.github.io/gce-tutorial/\">this guide</a> </p>",
      "rawMarkdown": "Yes, I connect Jupyter using [this guide][1] \n\n\n  [1]: http://cs231n.github.io/gce-tutorial/",
      "votes": null
    },
    {
      "id": "323212",
      "postDate": "05/04/2018 16:10:35",
      "content": "<p>Ok I also used it :) I have never used bucket. Is it essentialy a directory on server? </p>",
      "rawMarkdown": "Ok I also used it :) I have never used bucket. Is it essentialy a directory on server?",
      "votes": null
    },
    {
      "id": "323214",
      "postDate": "05/04/2018 16:13:07",
      "content": "<p>If you never use bucket, how can you read train and test data? You read them directly from your local machine?</p>",
      "rawMarkdown": "If you never use bucket, how can you read train and test data? You read them directly from your local machine?",
      "votes": null
    },
    {
      "id": "323217",
      "postDate": "05/04/2018 16:17:01",
      "content": "<p>I used Kaggle API and downloaded all data to server. Now I store it locally on server like any other file.</p>",
      "rawMarkdown": "I used Kaggle API and downloaded all data to server. Now I store it locally on server like any other file.",
      "votes": null
    },
    {
      "id": "323220",
      "postDate": "05/04/2018 16:18:09",
      "content": "<p>BTW Kaggle have done great job with this API. It works like magic. Good job Kaggle! </p>",
      "rawMarkdown": "BTW Kaggle have done great job with this API. It works like magic. Good job Kaggle!",
      "votes": null
    },
    {
      "id": "323242",
      "postDate": "05/04/2018 17:18:10",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    },
    {
      "id": "323244",
      "postDate": "05/04/2018 17:22:45",
      "content": "<p>Did it work? </p>",
      "rawMarkdown": "Did it work?",
      "votes": null
    },
    {
      "id": "323256",
      "postDate": "05/04/2018 18:02:38",
      "content": "<p>Yes, I upload data from Jupyter. So the data will be stored locally on server. I think It will work, though I have not tested it.</p>",
      "rawMarkdown": "Yes, I upload data from Jupyter. So the data will be stored locally on server. I think It will work, though I have not tested it.",
      "votes": null
    },
    {
      "id": "323259",
      "postDate": "05/04/2018 18:08:04",
      "content": "<p>Good luck. Happy kaggling. </p>",
      "rawMarkdown": "Good luck. Happy kaggling.",
      "votes": null
    },
    {
      "id": "323293",
      "postDate": "05/04/2018 19:07:16",
      "content": "<p>I'm also using gcloud, following code works for me.</p>\n\n<p><code>from google.cloud import storage</code></p>\n\n<p><code>from io import BytesIO</code>   </p>\n\n<p><code>client = storage.Client()</code></p>\n\n<p><code>bucket = \"your bucket name\"</code></p>\n\n<p><strong>For read</strong></p>\n\n<p><code>blob = storage.blob.Blob(\"train.csv\",bucket)</code></p>\n\n<p><code>content = blob.download_as_string()</code></p>\n\n<p><code>train  = pd.read_csv(BytesIO(content))</code></p>\n\n<p><strong>For write</strong></p>\n\n<p><code>train.to_pickle(\"train.pkl\")</code></p>\n\n<p><code>blob = bucket.blob('train.pkl')</code></p>\n\n<p><code>blob.upload_from_filename('train.pkl')</code></p>\n\n<p>or \n<code>train.to_csv()</code> if you like</p>",
      "rawMarkdown": "I'm also using gcloud, following code works for me.\n\n`from google.cloud import storage`\n\n`from io import BytesIO`   \n\n`client = storage.Client()`\n\n`bucket = \"your bucket name\"`\n\n**For read**\n\n`blob = storage.blob.Blob(\"train.csv\",bucket)`\n\n`content = blob.download_as_string()`\n\n`train  = pd.read_csv(BytesIO(content))`\n\n**For write**\n\n`train.to_pickle(\"train.pkl\")`\n\n`blob = bucket.blob('train.pkl')`\n\n`blob.upload_from_filename('train.pkl')`\n\nor \n`train.to_csv()` if you like",
      "votes": null
    },
    {
      "id": "323323",
      "postDate": "05/04/2018 20:41:16",
      "content": "<p>Thanks for reply! It works!</p>",
      "rawMarkdown": "Thanks for reply! It works!",
      "votes": null
    },
    {
      "id": "323324",
      "postDate": "05/04/2018 20:42:10",
      "content": "<p>May I ask you how many memory you set for training the whole data set</p>",
      "rawMarkdown": "May I ask you how many memory you set for training the whole data set",
      "votes": null
    },
    {
      "id": "323351",
      "postDate": "05/04/2018 22:33:11",
      "content": "<p>About 150GB-200GB.</p>",
      "rawMarkdown": "About 150GB-200GB.",
      "votes": null
    },
    {
      "id": "323354",
      "postDate": "05/04/2018 22:38:51",
      "content": "<p>Wow, I think 52 GB is the maximum memory I can get.....</p>",
      "rawMarkdown": "Wow, I think 52 GB is the maximum memory I can get.....",
      "votes": null
    },
    {
      "id": "323355",
      "postDate": "05/04/2018 22:45:31",
      "content": "<p>Upgrade your account to paid user gives you access up to 24 cores + more than 200GB RAM. Don't forget to close your paid account after competition :)  </p>",
      "rawMarkdown": "Upgrade your account to paid user gives you access up to 24 cores + more than 200GB RAM. Don't forget to close your paid account after competition :)",
      "votes": null
    },
    {
      "id": "323375",
      "postDate": "05/05/2018 01:10:27",
      "content": "<p>I personally use <code>gsutil</code> to transfer files around\n<code>\n!gsutil -m cp gs://name-of-cloud-storage/file-name.csv ../path-to-store-file\n</code>\nhope this helps</p>",
      "rawMarkdown": "I personally use `gsutil` to transfer files around\n```\n!gsutil -m cp gs://name-of-cloud-storage/file-name.csv ../path-to-store-file\n```\nhope this helps",
      "votes": null
    },
    {
      "id": "323407",
      "postDate": "05/05/2018 03:47:29",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    },
    {
      "id": "323461",
      "postDate": "05/05/2018 07:50:12",
      "content": "<p>If you use Jupyer Lab on gcloud then you can easily upload and download files from the browser. There is a size limit on uploads but good for moving scripts. I would always download the competition data from Kaggles API as Googles servers will pull it down very quickly. </p>\n\n<p>Sub files can easily be brought down at click of button (despite being a few hundred MB). </p>\n\n<p>For uploading larger files \"gcloud compute scp\" from command line. </p>",
      "rawMarkdown": "If you use Jupyer Lab on gcloud then you can easily upload and download files from the browser. There is a size limit on uploads but good for moving scripts. I would always download the competition data from Kaggles API as Googles servers will pull it down very quickly. \n\nSub files can easily be brought down at click of button (despite being a few hundred MB). \n\nFor uploading larger files \"gcloud compute scp\" from command line.",
      "votes": null
    },
    {
      "id": "323569",
      "postDate": "05/05/2018 14:33:30",
      "content": "<p>Got it! Than you, Stanwar!</p>",
      "rawMarkdown": "Got it! Than you, Stanwar!",
      "votes": null
    },
    {
      "id": "394194",
      "postDate": "09/26/2018 13:33:09",
      "content": "<p>Hi\nHow do u download all data to google cloud using the kaggle api?\nI'm working on diagnosing Diabetic Retinopathy but the data is too large, so I have decided to use GCP.\nThanks</p>",
      "rawMarkdown": "Hi\nHow do u download all data to google cloud using the kaggle api?\nI'm working on diagnosing Diabetic Retinopathy but the data is too large, so I have decided to use GCP.\nThanks",
      "votes": null
    },
    {
      "id": "572870",
      "postDate": "07/11/2019 14:11:18",
      "content": "<p>Hi Zishi,\n         I am in a situation trying to access a csv file from my cloud storage bucket in my python Jupyter notebook. I have tried using the code snippet you had provided but I am getting an error. Will you be able to help in resolving the problem?\n<code>\nfrom google.cloud import storage\nfrom io import BytesIO\nclient = storage.Client()\nbucket = \"rainwater\"\nln_blob = storage.blob.Blob(\"training\\train.csv\",bucket)\ncnt = ln_blob.download_as_string()\ntrain = pd.read_csv(BytesIO(cnt))</code>\n`---------------------------------------------------------------------------\nAttributeError                            Traceback (most recent call last)\n in ()\n      4 bucket = \"rainwater\"\n      5 ln_blob = storage.blob.Blob(\"training\\train.csv\",bucket)\n----&gt; 6 cnt = ln_blob.download_as_string()\n      7 train = pd.read_csv(BytesIO(cnt))</p>\n\n<p>/opt/conda/anaconda/lib/python3.6/site-packages/google/cloud/storage/blob.py in download_as_string(self, client, start, end)\n    597         string_buffer = BytesIO()\n    598         self.download_to_file(\n--&gt; 599             string_buffer, client=client, start=start, end=end)\n    600         return string_buffer.getvalue()\n    601 </p>\n\n<p>/opt/conda/anaconda/lib/python3.6/site-packages/google/cloud/storage/blob.py in download_to_file(self, file_obj, client, start, end)\n    526         :raises: :class:<code>google.cloud.exceptions.NotFound</code>\n    527         \"\"\"\n--&gt; 528         download_url = self._get_download_url()\n    529         headers = _get_encryption_headers(self._encryption_key)\n    530         headers['accept-encoding'] = 'gzip'</p>\n\n<p>/opt/conda/anaconda/lib/python3.6/site-packages/google/cloud/storage/blob.py in _get_download_url(self)\n    432         name_value_pairs = []\n    433         if self.media_link is None:\n--&gt; 434             base_url = _DOWNLOAD_URL_TEMPLATE.format(path=self.path)\n    435             if self.generation is not None:\n    436                 name_value_pairs.append(</p>\n\n<p>/opt/conda/anaconda/lib/python3.6/site-packages/google/cloud/storage/blob.py in path(self)\n    238             raise ValueError('Cannot determine path without a blob name.')\n    239 \n--&gt; 240         return self.path_helper(self.bucket.path, self.name)\n    241 \n    242     <a href=\"/property\">@property</a></p>\n\n<p>AttributeError: 'str' object has no attribute 'path'`</p>",
      "rawMarkdown": "Hi Zishi,\n         I am in a situation trying to access a csv file from my cloud storage bucket in my python Jupyter notebook. I have tried using the code snippet you had provided but I am getting an error. Will you be able to help in resolving the problem?\n`\nfrom google.cloud import storage\nfrom io import BytesIO\nclient = storage.Client()\nbucket = \"rainwater\"\nln_blob = storage.blob.Blob(\"training\\train.csv\",bucket)\ncnt = ln_blob.download_as_string()\ntrain = pd.read_csv(BytesIO(cnt))`\n`---------------------------------------------------------------------------\nAttributeError                            Traceback (most recent call last)",
      "votes": null
    },
    {
      "id": "588354",
      "postDate": "07/30/2019 12:59:03",
      "content": "<p>Hi, sorry for the delay, haven't been here for a while, just noticed your question.</p>\n\n<p>You need to declare project_id explicitly while creating storage.Client(). In my case I was using Jupyter notebook on Google Cloud so I didn't need to do that.\nTry following code, it should work.</p>\n\n<p>```\n    storage_client = storage.Client(project = project_id)\n    bucket = storage_client.get_bucket(bucket_name)\n    blob = storage.blob.Blob(file_path,bucket)</p>\n\n<p>```</p>",
      "rawMarkdown": "Hi, sorry for the delay, haven't been here for a while, just noticed your question.\n\nYou need to declare project_id explicitly while creating storage.Client(). In my case I was using Jupyter notebook on Google Cloud so I didn't need to do that.\nTry following code, it should work.\n\n```\n    storage_client = storage.Client(project = project_id)\n    bucket = storage_client.get_bucket(bucket_name)\n    blob = storage.blob.Blob(file_path,bucket)\n\n```",
      "votes": null
    },
    {
      "id": "590642",
      "postDate": "08/02/2019 12:36:19",
      "content": "<p>Hello guys,\nsomeone know how to give path to access bucket and jupyter notebook?</p>",
      "rawMarkdown": "Hello guys,\nsomeone know how to give path to access bucket and jupyter notebook?",
      "votes": null
    },
    {
      "id": "810878",
      "postDate": "04/17/2020 11:08:55",
      "content": "<p><a href=\"/qraina\">@qraina</a> you may need also pass <code>-r</code> to load multiple files located in a different directory. Though it takes good amount of time and sometimes the browser gets hanged.</p>",
      "rawMarkdown": "qraina you may need also pass `-r` to load multiple files located in a different directory. Though it takes good amount of time and sometimes the browser gets hanged.",
      "votes": null
    },
    {
      "id": "867748",
      "postDate": "05/30/2020 15:30:14",
      "content": "<p>I try to read not public data from google cloud bucket but I dont know how to authorize to bucket.\n I try to use:\n<code>\n!pip install google-colab\nfrom google.colab import auth\nauth.authenticate_user()\n</code></p>\n\n<p>but it doesn't work.  Does anybode have an idea what is wrong ?</p>",
      "rawMarkdown": "I try to read not public data from google cloud bucket but I dont know how to authorize to bucket.\n I try to use:\n```\n!pip install google-colab\nfrom google.colab import auth\nauth.authenticate_user()\n```\n\nbut it doesn't work.  Does anybode have an idea what is wrong ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 323169,
      "author_name": "fangao",
      "author_url": "",
      "post_date": "05/04/2018 14:42:52",
      "content": "<p>I have dealt with how to read a csv to dataframe. Now, the question becomes how to write a pandas dataframe to gcloud storage bucket. Thanks.</p>",
      "votes": null,
      "replies": [
        {
          "id": 323178,
          "author_name": "skalskip",
          "author_url": "",
          "post_date": "05/04/2018 14:55:28",
          "content": "<p>Don't you prefer to install Jupyter on gcloud. Redirect everything to the port and view data in your browser on your computer?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323183,
          "author_name": "fangao",
          "author_url": "",
          "post_date": "05/04/2018 14:59:10",
          "content": "<p>Hi SkalskiP,  thanks for reply. I don't perfectly understand what you said. But I think I have done something similar. I install Jupyter on gcloud and now I can run Jupyter notebook on gcloud. But I don't know how to write dataframe to my storage bucket.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323196,
          "author_name": "skalskip",
          "author_url": "",
          "post_date": "05/04/2018 15:33:00",
          "content": "<p>I understand. Of course, I wanted something similar. Honestly, I do not know why you want to have it in bucket. Do you want to submit with gc?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323197,
          "author_name": "fangao",
          "author_url": "",
          "post_date": "05/04/2018 15:37:04",
          "content": "<p>Because I assume after I upload my dataframe(e.g the prediction result dataframe in this competition) to bucket, I can download it to my local computer. Or is there any other way to accomplish this? Can I download the result dataframe directly to my local computer without uploading to bucket?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323202,
          "author_name": "skalskip",
          "author_url": "",
          "post_date": "05/04/2018 15:48:42",
          "content": "<p>Hm there are few ways of doing that. If you just want to upload submission to Kaggle, you can use Kaggle API. You can find it on GitHub. If you want to download it on your computer and you can use Linux as your local machina you can use scp command. Or you can save your file in some directory on server and than copy it to bucket and than download. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323203,
          "author_name": "fangao",
          "author_url": "",
          "post_date": "05/04/2018 15:51:27",
          "content": "<p>Thanks for reply! Can you explain more about the third way? I am a windows user. I really appreciate your help!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323207,
          "author_name": "skalskip",
          "author_url": "",
          "post_date": "05/04/2018 16:03:45",
          "content": "<p>How do you connect to your gc? Via SSH? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323208,
          "author_name": "fangao",
          "author_url": "",
          "post_date": "05/04/2018 16:05:34",
          "content": "<p>Yes, I connect Jupyter using <a href=\"http://cs231n.github.io/gce-tutorial/\">this guide</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323212,
          "author_name": "skalskip",
          "author_url": "",
          "post_date": "05/04/2018 16:10:35",
          "content": "<p>Ok I also used it :) I have never used bucket. Is it essentialy a directory on server? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323214,
          "author_name": "fangao",
          "author_url": "",
          "post_date": "05/04/2018 16:13:07",
          "content": "<p>If you never use bucket, how can you read train and test data? You read them directly from your local machine?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323217,
          "author_name": "skalskip",
          "author_url": "",
          "post_date": "05/04/2018 16:17:01",
          "content": "<p>I used Kaggle API and downloaded all data to server. Now I store it locally on server like any other file.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323220,
          "author_name": "skalskip",
          "author_url": "",
          "post_date": "05/04/2018 16:18:09",
          "content": "<p>BTW Kaggle have done great job with this API. It works like magic. Good job Kaggle! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323242,
          "author_name": "fangao",
          "author_url": "",
          "post_date": "05/04/2018 17:18:10",
          "content": "<p>Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323244,
          "author_name": "skalskip",
          "author_url": "",
          "post_date": "05/04/2018 17:22:45",
          "content": "<p>Did it work? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323256,
          "author_name": "fangao",
          "author_url": "",
          "post_date": "05/04/2018 18:02:38",
          "content": "<p>Yes, I upload data from Jupyter. So the data will be stored locally on server. I think It will work, though I have not tested it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323259,
          "author_name": "skalskip",
          "author_url": "",
          "post_date": "05/04/2018 18:08:04",
          "content": "<p>Good luck. Happy kaggling. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 394194,
          "author_name": "aaryapatel",
          "author_url": "",
          "post_date": "09/26/2018 13:33:09",
          "content": "<p>Hi\nHow do u download all data to google cloud using the kaggle api?\nI'm working on diagnosing Diabetic Retinopathy but the data is too large, so I have decided to use GCP.\nThanks</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 323293,
      "author_name": "zhongzishi",
      "author_url": "",
      "post_date": "05/04/2018 19:07:16",
      "content": "<p>I'm also using gcloud, following code works for me.</p>\n\n<p><code>from google.cloud import storage</code></p>\n\n<p><code>from io import BytesIO</code>   </p>\n\n<p><code>client = storage.Client()</code></p>\n\n<p><code>bucket = \"your bucket name\"</code></p>\n\n<p><strong>For read</strong></p>\n\n<p><code>blob = storage.blob.Blob(\"train.csv\",bucket)</code></p>\n\n<p><code>content = blob.download_as_string()</code></p>\n\n<p><code>train  = pd.read_csv(BytesIO(content))</code></p>\n\n<p><strong>For write</strong></p>\n\n<p><code>train.to_pickle(\"train.pkl\")</code></p>\n\n<p><code>blob = bucket.blob('train.pkl')</code></p>\n\n<p><code>blob.upload_from_filename('train.pkl')</code></p>\n\n<p>or \n<code>train.to_csv()</code> if you like</p>",
      "votes": null,
      "replies": [
        {
          "id": 323323,
          "author_name": "fangao",
          "author_url": "",
          "post_date": "05/04/2018 20:41:16",
          "content": "<p>Thanks for reply! It works!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323324,
          "author_name": "fangao",
          "author_url": "",
          "post_date": "05/04/2018 20:42:10",
          "content": "<p>May I ask you how many memory you set for training the whole data set</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323351,
          "author_name": "zhongzishi",
          "author_url": "",
          "post_date": "05/04/2018 22:33:11",
          "content": "<p>About 150GB-200GB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323354,
          "author_name": "fangao",
          "author_url": "",
          "post_date": "05/04/2018 22:38:51",
          "content": "<p>Wow, I think 52 GB is the maximum memory I can get.....</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323355,
          "author_name": "zhongzishi",
          "author_url": "",
          "post_date": "05/04/2018 22:45:31",
          "content": "<p>Upgrade your account to paid user gives you access up to 24 cores + more than 200GB RAM. Don't forget to close your paid account after competition :)  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 572870,
          "author_name": "dineshsun7",
          "author_url": "",
          "post_date": "07/11/2019 14:11:18",
          "content": "<p>Hi Zishi,\n         I am in a situation trying to access a csv file from my cloud storage bucket in my python Jupyter notebook. I have tried using the code snippet you had provided but I am getting an error. Will you be able to help in resolving the problem?\n<code>\nfrom google.cloud import storage\nfrom io import BytesIO\nclient = storage.Client()\nbucket = \"rainwater\"\nln_blob = storage.blob.Blob(\"training\\train.csv\",bucket)\ncnt = ln_blob.download_as_string()\ntrain = pd.read_csv(BytesIO(cnt))</code>\n`---------------------------------------------------------------------------\nAttributeError                            Traceback (most recent call last)\n in ()\n      4 bucket = \"rainwater\"\n      5 ln_blob = storage.blob.Blob(\"training\\train.csv\",bucket)\n----&gt; 6 cnt = ln_blob.download_as_string()\n      7 train = pd.read_csv(BytesIO(cnt))</p>\n\n<p>/opt/conda/anaconda/lib/python3.6/site-packages/google/cloud/storage/blob.py in download_as_string(self, client, start, end)\n    597         string_buffer = BytesIO()\n    598         self.download_to_file(\n--&gt; 599             string_buffer, client=client, start=start, end=end)\n    600         return string_buffer.getvalue()\n    601 </p>\n\n<p>/opt/conda/anaconda/lib/python3.6/site-packages/google/cloud/storage/blob.py in download_to_file(self, file_obj, client, start, end)\n    526         :raises: :class:<code>google.cloud.exceptions.NotFound</code>\n    527         \"\"\"\n--&gt; 528         download_url = self._get_download_url()\n    529         headers = _get_encryption_headers(self._encryption_key)\n    530         headers['accept-encoding'] = 'gzip'</p>\n\n<p>/opt/conda/anaconda/lib/python3.6/site-packages/google/cloud/storage/blob.py in _get_download_url(self)\n    432         name_value_pairs = []\n    433         if self.media_link is None:\n--&gt; 434             base_url = _DOWNLOAD_URL_TEMPLATE.format(path=self.path)\n    435             if self.generation is not None:\n    436                 name_value_pairs.append(</p>\n\n<p>/opt/conda/anaconda/lib/python3.6/site-packages/google/cloud/storage/blob.py in path(self)\n    238             raise ValueError('Cannot determine path without a blob name.')\n    239 \n--&gt; 240         return self.path_helper(self.bucket.path, self.name)\n    241 \n    242     <a href=\"/property\">@property</a></p>\n\n<p>AttributeError: 'str' object has no attribute 'path'`</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 588354,
          "author_name": "zhongzishi",
          "author_url": "",
          "post_date": "07/30/2019 12:59:03",
          "content": "<p>Hi, sorry for the delay, haven't been here for a while, just noticed your question.</p>\n\n<p>You need to declare project_id explicitly while creating storage.Client(). In my case I was using Jupyter notebook on Google Cloud so I didn't need to do that.\nTry following code, it should work.</p>\n\n<p>```\n    storage_client = storage.Client(project = project_id)\n    bucket = storage_client.get_bucket(bucket_name)\n    blob = storage.blob.Blob(file_path,bucket)</p>\n\n<p>```</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 323375,
      "author_name": "qraina",
      "author_url": "",
      "post_date": "05/05/2018 01:10:27",
      "content": "<p>I personally use <code>gsutil</code> to transfer files around\n<code>\n!gsutil -m cp gs://name-of-cloud-storage/file-name.csv ../path-to-store-file\n</code>\nhope this helps</p>",
      "votes": null,
      "replies": [
        {
          "id": 323407,
          "author_name": "fangao",
          "author_url": "",
          "post_date": "05/05/2018 03:47:29",
          "content": "<p>Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323461,
          "author_name": "stanwar",
          "author_url": "",
          "post_date": "05/05/2018 07:50:12",
          "content": "<p>If you use Jupyer Lab on gcloud then you can easily upload and download files from the browser. There is a size limit on uploads but good for moving scripts. I would always download the competition data from Kaggles API as Googles servers will pull it down very quickly. </p>\n\n<p>Sub files can easily be brought down at click of button (despite being a few hundred MB). </p>\n\n<p>For uploading larger files \"gcloud compute scp\" from command line. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323569,
          "author_name": "fangao",
          "author_url": "",
          "post_date": "05/05/2018 14:33:30",
          "content": "<p>Got it! Than you, Stanwar!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 810878,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "04/17/2020 11:08:55",
          "content": "<p><a href=\"/qraina\">@qraina</a> you may need also pass <code>-r</code> to load multiple files located in a different directory. Though it takes good amount of time and sometimes the browser gets hanged.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 590642,
      "author_name": "mkunik",
      "author_url": "",
      "post_date": "08/02/2019 12:36:19",
      "content": "<p>Hello guys,\nsomeone know how to give path to access bucket and jupyter notebook?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 867748,
      "author_name": "peterpirog",
      "author_url": "",
      "post_date": "05/30/2020 15:30:14",
      "content": "<p>I try to read not public data from google cloud bucket but I dont know how to authorize to bucket.\n I try to use:\n<code>\n!pip install google-colab\nfrom google.colab import auth\nauth.authenticate_user()\n</code></p>\n\n<p>but it doesn't work.  Does anybode have an idea what is wrong ?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "323159": "Hi there,\n\nI have connected my jupyter notebook server to gcloud. But I have problem loading csv file from gcloud bucket. Can anyone share some code about this? Thanks a lot.\n\nBest regards",
    "323169": "I have dealt with how to read a csv to dataframe. Now, the question becomes how to write a pandas dataframe to gcloud storage bucket. Thanks.",
    "323178": "Don't you prefer to install Jupyter on gcloud. Redirect everything to the port and view data in your browser on your computer?",
    "323183": "Hi SkalskiP,  thanks for reply. I don't perfectly understand what you said. But I think I have done something similar. I install Jupyter on gcloud and now I can run Jupyter notebook on gcloud. But I don't know how to write dataframe to my storage bucket.",
    "323196": "I understand. Of course, I wanted something similar. Honestly, I do not know why you want to have it in bucket. Do you want to submit with gc?",
    "323197": "Because I assume after I upload my dataframe(e.g the prediction result dataframe in this competition) to bucket, I can download it to my local computer. Or is there any other way to accomplish this? Can I download the result dataframe directly to my local computer without uploading to bucket?",
    "323202": "Hm there are few ways of doing that. If you just want to upload submission to Kaggle, you can use Kaggle API. You can find it on GitHub. If you want to download it on your computer and you can use Linux as your local machina you can use scp command. Or you can save your file in some directory on server and than copy it to bucket and than download.",
    "323203": "Thanks for reply! Can you explain more about the third way? I am a windows user. I really appreciate your help!",
    "323207": "How do you connect to your gc? Via SSH?",
    "323208": "Yes, I connect Jupyter using [this guide][1] \n\n\n  [1]: http://cs231n.github.io/gce-tutorial/",
    "323212": "Ok I also used it :) I have never used bucket. Is it essentialy a directory on server?",
    "323214": "If you never use bucket, how can you read train and test data? You read them directly from your local machine?",
    "323217": "I used Kaggle API and downloaded all data to server. Now I store it locally on server like any other file.",
    "323220": "BTW Kaggle have done great job with this API. It works like magic. Good job Kaggle!",
    "323242": "Thanks!",
    "323244": "Did it work?",
    "323256": "Yes, I upload data from Jupyter. So the data will be stored locally on server. I think It will work, though I have not tested it.",
    "323259": "Good luck. Happy kaggling.",
    "323293": "I'm also using gcloud, following code works for me.\n\n`from google.cloud import storage`\n\n`from io import BytesIO`   \n\n`client = storage.Client()`\n\n`bucket = \"your bucket name\"`\n\n**For read**\n\n`blob = storage.blob.Blob(\"train.csv\",bucket)`\n\n`content = blob.download_as_string()`\n\n`train  = pd.read_csv(BytesIO(content))`\n\n**For write**\n\n`train.to_pickle(\"train.pkl\")`\n\n`blob = bucket.blob('train.pkl')`\n\n`blob.upload_from_filename('train.pkl')`\n\nor \n`train.to_csv()` if you like",
    "323323": "Thanks for reply! It works!",
    "323324": "May I ask you how many memory you set for training the whole data set",
    "323351": "About 150GB-200GB.",
    "323354": "Wow, I think 52 GB is the maximum memory I can get.....",
    "323355": "Upgrade your account to paid user gives you access up to 24 cores + more than 200GB RAM. Don't forget to close your paid account after competition :)",
    "323375": "I personally use `gsutil` to transfer files around\n```\n!gsutil -m cp gs://name-of-cloud-storage/file-name.csv ../path-to-store-file\n```\nhope this helps",
    "323407": "Thanks!",
    "323461": "If you use Jupyer Lab on gcloud then you can easily upload and download files from the browser. There is a size limit on uploads but good for moving scripts. I would always download the competition data from Kaggles API as Googles servers will pull it down very quickly. \n\nSub files can easily be brought down at click of button (despite being a few hundred MB). \n\nFor uploading larger files \"gcloud compute scp\" from command line.",
    "323569": "Got it! Than you, Stanwar!",
    "394194": "Hi\nHow do u download all data to google cloud using the kaggle api?\nI'm working on diagnosing Diabetic Retinopathy but the data is too large, so I have decided to use GCP.\nThanks",
    "572870": "Hi Zishi,\n         I am in a situation trying to access a csv file from my cloud storage bucket in my python Jupyter notebook. I have tried using the code snippet you had provided but I am getting an error. Will you be able to help in resolving the problem?\n`\nfrom google.cloud import storage\nfrom io import BytesIO\nclient = storage.Client()\nbucket = \"rainwater\"\nln_blob = storage.blob.Blob(\"training\\train.csv\",bucket)\ncnt = ln_blob.download_as_string()\ntrain = pd.read_csv(BytesIO(cnt))`\n`---------------------------------------------------------------------------\nAttributeError                            Traceback (most recent call last)",
    "588354": "Hi, sorry for the delay, haven't been here for a while, just noticed your question.\n\nYou need to declare project_id explicitly while creating storage.Client(). In my case I was using Jupyter notebook on Google Cloud so I didn't need to do that.\nTry following code, it should work.\n\n```\n    storage_client = storage.Client(project = project_id)\n    bucket = storage_client.get_bucket(bucket_name)\n    blob = storage.blob.Blob(file_path,bucket)\n\n```",
    "590642": "Hello guys,\nsomeone know how to give path to access bucket and jupyter notebook?",
    "810878": "qraina you may need also pass `-r` to load multiple files located in a different directory. Though it takes good amount of time and sometimes the browser gets hanged.",
    "867748": "I try to read not public data from google cloud bucket but I dont know how to authorize to bucket.\n I try to use:\n```\n!pip install google-colab\nfrom google.colab import auth\nauth.authenticate_user()\n```\n\nbut it doesn't work.  Does anybode have an idea what is wrong ?"
  },
  "source": "meta"
}