{
  "id": 130326,
  "title": "How to use my own data source? ",
  "url": "/competitions/flower-classification-with-tpus/discussion/130326",
  "author_name": "",
  "post_date": "2020-02-13T14:10:44.080058800Z",
  "votes": 8,
  "comment_count": 18,
  "views": 0,
  "content": "<p>Hi,</p>\n\n<p>I try to add some other data source into my kernel with tpu.  Then my kernel has two datas.  e.g. \nGCS_DS_PATH =KaggleDatasets().get_gcs_path('flower-classification')\nMY_DATA_PATH=KaggleDatasets().get_gcs_path(' mydata/')</p>\n\n<p>When I ran the codes , I got the errors:\nBackendError: Unexpected response from the service. Response: {'errors': ['Only public datasets and competitions are currently supported'], 'error': {'code': 12}, 'wasSuccessful': False}.</p>\n\n<p>What should I do ? Should I make my private data source public ? </p>",
  "messages": [
    {
      "id": "745117",
      "postDate": "02/13/2020 14:10:44",
      "content": "<p>Hi,</p>\n\n<p>I try to add some other data source into my kernel with tpu.  Then my kernel has two datas.  e.g. \nGCS_DS_PATH =KaggleDatasets().get_gcs_path('flower-classification')\nMY_DATA_PATH=KaggleDatasets().get_gcs_path(' mydata/')</p>\n\n<p>When I ran the codes , I got the errors:\nBackendError: Unexpected response from the service. Response: {'errors': ['Only public datasets and competitions are currently supported'], 'error': {'code': 12}, 'wasSuccessful': False}.</p>\n\n<p>What should I do ? Should I make my private data source public ? </p>",
      "rawMarkdown": "Hi,\n\nI try to add some other data source into my kernel with tpu.  Then my kernel has two datas.  e.g. \nGCS_DS_PATH =KaggleDatasets().get_gcs_path('flower-classification')\nMY_DATA_PATH=KaggleDatasets().get_gcs_path(' mydata/')\n\nWhen I ran the codes , I got the errors:\nBackendError: Unexpected response from the service. Response: {'errors': ['Only public datasets and competitions are currently supported'], 'error': {'code': 12}, 'wasSuccessful': False}.\n\n\nWhat should I do ? Should I make my private data source public ?",
      "votes": null
    },
    {
      "id": "745366",
      "postDate": "02/13/2020 18:38:51",
      "content": "<p>Two possible ways:\n1. Use your own public GCS bucket. Access it with <code>gs://my-bucket-name/...</code>. Note that if this is your bucket you pay for it.\n2. Upload your data in a public Kaggale Dataset. The get the GCS path to it with: <code>MY_GCS_PATH=KaggleDatasets().get_gcs_path('my-kaggle-dataset')</code></p>",
      "rawMarkdown": "Two possible ways:\n1. Use your own public GCS bucket. Access it with `gs://my-bucket-name/...`. Note that if this is your bucket you pay for it.\n2. Upload your data in a public Kaggale Dataset. The get the GCS path to it with: `MY_GCS_PATH=KaggleDatasets().get_gcs_path('my-kaggle-dataset')`",
      "votes": null
    },
    {
      "id": "745367",
      "postDate": "02/13/2020 18:40:23",
      "content": "<p>The directory name of the Kaggle dataset (i.e. the name to pass into <code>get_gcs_dataset('name')</code>) can be looked up by attaching the dataset to your notebook then running this in a cell:\n<code>\n!ls /kaggle/input\n</code></p>",
      "rawMarkdown": "The directory name of the Kaggle dataset (i.e. the name to pass into `get_gcs_dataset('name')`) can be looked up by attaching the dataset to your notebook then running this in a cell:\n```\n!ls /kaggle/input\n```",
      "votes": null
    },
    {
      "id": "745574",
      "postDate": "02/14/2020 00:24:11",
      "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> Thanks.  It seems I have to make my dataset public.  :(</p>",
      "rawMarkdown": "mgornergoogle Thanks.  It seems I have to make my dataset public.  :(",
      "votes": null
    },
    {
      "id": "746161",
      "postDate": "02/14/2020 17:29:20",
      "content": "<p>At the moment yes, unfortunately.\nI would add that this is a friendly \"playground\" competition and knowledge sharing is the goal.</p>",
      "rawMarkdown": "At the moment yes, unfortunately.\nI would add that this is a friendly \"playground\" competition and knowledge sharing is the goal.",
      "votes": null
    },
    {
      "id": "746208",
      "postDate": "02/14/2020 18:32:20",
      "content": "<p>I tried to do as you suggested; however, I still get the same error. I added the data and named 'test set' when I load using GCS_PATH = KaggleDatasets().get_gcs_path('test-set'). I get the following error</p>\n\n<p>BackendError: Unexpected response from the service. Response: {'errors': [\"Dataset not found for directory 'test-set', please make sure you are passing a valid directory name under /kaggle/input\"], 'error': {'code': 5}, 'wasSuccessful': False}.I tried to do as you suggested, however, I still get the same error. </p>\n\n<p>please help me with this</p>",
      "rawMarkdown": "I tried to do as you suggested; however, I still get the same error. I added the data and named 'test set' when I load using GCS_PATH = KaggleDatasets().get_gcs_path('test-set'). I get the following error\n\nBackendError: Unexpected response from the service. Response: {'errors': [\"Dataset not found for directory 'test-set', please make sure you are passing a valid directory name under /kaggle/input\"], 'error': {'code': 5}, 'wasSuccessful': False}.I tried to do as you suggested, however, I still get the same error. \n\nplease help me with this",
      "votes": null
    },
    {
      "id": "746291",
      "postDate": "02/14/2020 20:10:29",
      "content": "<p>List the directories you have under /kaggle/input/ with:\n<code>\n!ls /kaggle/input\n</code>\nThen use the directory name of your dataset. I would bet it is something like 'test_set'</p>",
      "rawMarkdown": "List the directories you have under /kaggle/input/ with:\n```\n!ls /kaggle/input\n```\nThen use the directory name of your dataset. I would bet it is something like 'test_set'",
      "votes": null
    },
    {
      "id": "746914",
      "postDate": "02/15/2020 18:22:27",
      "content": "<p>Thank you. It worked, however, I had to change the name of the folder from data_set to data-set I don't know why it makes me do like this.Thank you it worked, however, I had to change the name of the folder from data_set to data-set I don't know why it makes me do like this.</p>",
      "rawMarkdown": "Thank you. It worked, however, I had to change the name of the folder from data_set to data-set I don't know why it makes me do like this.Thank you it worked, however, I had to change the name of the folder from data_set to data-set I don't know why it makes me do like this.",
      "votes": null
    },
    {
      "id": "746979",
      "postDate": "02/15/2020 20:00:16",
      "content": "<p>Hello Martin</p>\n\n<p>Could you share with me if you don't mind the kernel of how you generate tfrecords from images? I guess the kernel resizes the individual image with different size and format to jpg to either of the square dimensions 331,224, etc.,Hello Martin</p>\n\n<p>Could you share with me if you don't mind the kernel of how you generate tfrecords from images?</p>",
      "rawMarkdown": "Hello Martin\n\nCould you share with me if you don't mind the kernel of how you generate tfrecords from images? I guess the kernel resizes the individual image with different size and format to jpg to either of the square dimensions 331,224, etc.,Hello Martin\n\nCould you share with me if you don't mind the kernel of how you generate tfrecords from images?",
      "votes": null
    },
    {
      "id": "749574",
      "postDate": "02/18/2020 19:37:56",
      "content": "<p>The name to pass into the <code>KaggleDatasets().get_gcs_path(name)</code> function is the <strong>name of the directory</strong> where the dataset was mounted.</p>",
      "rawMarkdown": "The name to pass into the `KaggleDatasets().get_gcs_path(name)` function is the **name of the directory** where the dataset was mounted.",
      "votes": null
    },
    {
      "id": "749576",
      "postDate": "02/18/2020 19:40:03",
      "content": "<p>You will find all the details about how TFRecords work in this tutorial: <a href=\"https://codelabs.developers.google.com/codelabs/keras-flowers-data/#0\">TPU-speed data pipelines: tf.data.Dataset and TFRecords</a></p>\n\n<p>The notebook generating the TFRecord dataset in the tutorial is this one (Colab link):\n<a href=\"https://colab.research.google.com/github/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/03_Flower_pictures_to_TFRecords.ipynb\">Flower pictures to TFRecords.ipynb</a></p>",
      "rawMarkdown": "You will find all the details about how TFRecords work in this tutorial: [TPU-speed data pipelines: tf.data.Dataset and TFRecords](https://codelabs.developers.google.com/codelabs/keras-flowers-data/#0)\n\nThe notebook generating the TFRecord dataset in the tutorial is this one (Colab link):\n[Flower pictures to TFRecords.ipynb](https://colab.research.google.com/github/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/03_Flower_pictures_to_TFRecords.ipynb)",
      "votes": null
    },
    {
      "id": "757564",
      "postDate": "02/26/2020 23:02:28",
      "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> I've been trying to use my own GSC bucket. \nI have made my own bucket (based off <a href=\"https://www.kaggle.com/yufengg/automl-getting-started-notebook\">this</a> kernel) and uploaded the example tfrecords. I have used the Add-ons to add my Google Cloud access to the kernel.</p>\n\n<p>The next step is getting the kaggle kernel to use my GSC bucket instead of the kaggle one.\nIf I try the following \n<code>\nGCS_DS_PATH = 'gs://my-bucket-name'\nTRAINING_FILENAMES = tf.io.gfile.glob(GCS_PATH + '/tfrecords-jpeg-512x512/train/*.tfrec')\n</code>\nI get \n<code>\nPermissionDeniedError: Error executing an HTTP request: HTTP response code 401 with body '{\n  \"error\": {\n    \"code\": 401,\n    \"message\": \"Anonymous caller does not have storage.objects.list access to my-bucket-name.\",\n    \"errors\": [\n      {\n        \"message\": \"Anonymous caller does not have storage.objects.list access to my-bucket-name.\",\n        \"domain\": \"global\",\n        \"reason\": \"required\",\n        \"locationType\": \"header\",\n        \"location\": \"Authorization\"\n      }\n    ]\n  }\n}\n</code></p>\n\n<p>This is not too surprising as it is not a public bucket. However in the kernel I do have a bucket object\n<code>\nbucket = storage_client.get_bucket(bucket_name)\n</code>\nand I can get a list of my blobs in it using \n<code>\nblobs = storage_client.list_blobs(bucket_name)\nfor blob in blobs:\n        print(blob.name)\n</code></p>\n\n<p>My understanding is if I pass a list my full paths as <code>[gs://my-bucket-name/my.tfrec]</code> to <code>dataset = tf.data.TFRecordDataset(filenames, num_parallel_reads=AUTO)</code> in the <code>load_dataset</code> function it will also have the permission issue. </p>\n\n<p>My question is how do I use my bucket with the kaggle TPUs? </p>\n\n<p>Thanks </p>",
      "rawMarkdown": "mgornergoogle I've been trying to use my own GSC bucket. \nI have made my own bucket (based off [this](https://www.kaggle.com/yufengg/automl-getting-started-notebook) kernel) and uploaded the example tfrecords. I have used the Add-ons to add my Google Cloud access to the kernel.\n\nThe next step is getting the kaggle kernel to use my GSC bucket instead of the kaggle one.\nIf I try the following \n```\nGCS_DS_PATH = 'gs://my-bucket-name'\nTRAINING_FILENAMES = tf.io.gfile.glob(GCS_PATH + '/tfrecords-jpeg-512x512/train/*.tfrec')\n```\nI get \n```\nPermissionDeniedError: Error executing an HTTP request: HTTP response code 401 with body '{\n  \"error\": {\n    \"code\": 401,\n    \"message\": \"Anonymous caller does not have storage.objects.list access to my-bucket-name.\",\n    \"errors\": [\n      {\n        \"message\": \"Anonymous caller does not have storage.objects.list access to my-bucket-name.\",\n        \"domain\": \"global\",\n        \"reason\": \"required\",\n        \"locationType\": \"header\",\n        \"location\": \"Authorization\"\n      }\n    ]\n  }\n}\n```\n\nThis is not too surprising as it is not a public bucket. However in the kernel I do have a bucket object\n```\nbucket = storage_client.get_bucket(bucket_name)\n```\nand I can get a list of my blobs in it using \n```\nblobs = storage_client.list_blobs(bucket_name)\nfor blob in blobs:\n        print(blob.name)\n```\n\nMy understanding is if I pass a list my full paths as `[gs://my-bucket-name/my.tfrec]` to `dataset = tf.data.TFRecordDataset(filenames, num_parallel_reads=AUTO)` in the `load_dataset` function it will also have the permission issue. \n\nMy question is how do I use my bucket with the kaggle TPUs? \n\nThanks",
      "votes": null
    },
    {
      "id": "758575",
      "postDate": "02/28/2020 00:03:38",
      "content": "<p>We have done the integration with the GCS API <code>storage_client.*</code>, not with the Tensorflow GCS API <code>tf.io.gfile.*</code>. So when you add your Google credentials in add-ons, it makes the GCS API work but not the other one.</p>\n\n<p>I'm going to give you some hints. I'm not sure it is even possible to make it work and I think you'd better use a Notebook VM on GCP where authentication is automatic within the same project and between projects, you can use service accounts.</p>\n\n<p>(Please NEVER grant the Kaggle TPU service account  any permissions. That service account is owned by Kaggle, not you. If you do that, you have no idea who you are granting permissions to.)</p>\n\n<p>For Tensorflow access to GCS to work, you will need to to export a service account's key as JSON and \n configure:\n<code>export GOOGLE_APPLICATION_CREDENTIALS=/path/to/credential/json/file</code></p>\n\n<p>I do recommend you to use the Kaggle secrets service to store this JSON file securely.</p>\n\n<p>And that only gives your Kaggle VM access to your private GCS bucket. The TPU needs more work.</p>\n\n<p>The TPU is another machine and it will be accessing the bucket directly. It needs to be authorized too. This is the code that used to work in TF 1.x. It obviously does not work anymore:\n```\nTF_MASTER = 'grpc://{}'.format(os.environ['COLAB_TPU_ADDR'])</p>\n\n<h1>Upload credentials to TPU.</h1>\n\n<p>with tf.Session(TF_MASTER) as sess: <br>\n    with open('/content/adc.json', 'r') as f:\n        auth_info = json.load(f)\n    tf.contrib.cloud.configure_gcs(sess, credentials=auth_info)\n```\nThe configure_gcs function now lives in module <a href=\"https://www.tensorflow.org/io/api_docs/python/tfio/gcs\">tfio.gcs</a> </p>\n\n<p>From here, you're on your own. Good luck and share your solution if you manage to make it work (securely!)</p>",
      "rawMarkdown": "We have done the integration with the GCS API `storage_client.*`, not with the Tensorflow GCS API `tf.io.gfile.*`. So when you add your Google credentials in add-ons, it makes the GCS API work but not the other one.\n\nI'm going to give you some hints. I'm not sure it is even possible to make it work and I think you'd better use a Notebook VM on GCP where authentication is automatic within the same project and between projects, you can use service accounts.\n\n(Please NEVER grant the Kaggle TPU service account  any permissions. That service account is owned by Kaggle, not you. If you do that, you have no idea who you are granting permissions to.)\n\nFor Tensorflow access to GCS to work, you will need to to export a service account's key as JSON and \n configure:\n`export GOOGLE_APPLICATION_CREDENTIALS=/path/to/credential/json/file `\n\nI do recommend you to use the Kaggle secrets service to store this JSON file securely.\n\nAnd that only gives your Kaggle VM access to your private GCS bucket. The TPU needs more work.\n\nThe TPU is another machine and it will be accessing the bucket directly. It needs to be authorized too. This is the code that used to work in TF 1.x. It obviously does not work anymore:\n```\nTF_MASTER = 'grpc://{}'.format(os.environ['COLAB_TPU_ADDR'])\n# Upload credentials to TPU.\nwith tf.Session(TF_MASTER) as sess:    \n    with open('/content/adc.json', 'r') as f:\n        auth_info = json.load(f)\n    tf.contrib.cloud.configure_gcs(sess, credentials=auth_info)\n```\nThe configure_gcs function now lives in module [tfio.gcs](https://www.tensorflow.org/io/api_docs/python/tfio/gcs) \n\nFrom here, you're on your own. Good luck and share your solution if you manage to make it work (securely!)",
      "votes": null
    },
    {
      "id": "758587",
      "postDate": "02/28/2020 00:20:16",
      "content": "<p>Thank you. I'll give it a try</p>",
      "rawMarkdown": "Thank you. I'll give it a try",
      "votes": null
    },
    {
      "id": "811492",
      "postDate": "04/18/2020 01:22:51",
      "content": "<p>Thank you for the notebook it helped e a lot to understand on how to generate tfrecords. But I have a question on how to deal with the dt_uint8 which you used in your code. I failed to use TPU since it says the data type used for conversion image is not supported. I tried to cast to the image to float32 when I read image still I couldn't be able to run my code on TPU though it runs just fine on GPU.</p>",
      "rawMarkdown": "Thank you for the notebook it helped e a lot to understand on how to generate tfrecords. But I have a question on how to deal with the dt_uint8 which you used in your code. I failed to use TPU since it says the data type used for conversion image is not supported. I tried to cast to the image to float32 when I read image still I couldn't be able to run my code on TPU though it runs just fine on GPU.",
      "votes": null
    },
    {
      "id": "814671",
      "postDate": "04/20/2020 21:20:53",
      "content": "<p>The <a href=\"https://www.kaggle.com/mgornergoogle/five-flowers-with-keras-and-xception-on-tpu\">TPU sample in the documentation</a> loads image data as uint8 and trains a model on them on TPU. The conversion to float 32 happens in function <code>read_tfrecord</code>. The line of code that does the conversion is:\n```</p>\n\n<p>image = tf.cast(image, tf.float32) / 255.0  # convert image to floats in [0, 1] range\n```</p>",
      "rawMarkdown": "The [TPU sample in the documentation](https://www.kaggle.com/mgornergoogle/five-flowers-with-keras-and-xception-on-tpu) loads image data as uint8 and trains a model on them on TPU. The conversion to float 32 happens in function `read_tfrecord`. The line of code that does the conversion is:\n```\n\nimage = tf.cast(image, tf.float32) / 255.0  # convert image to floats in [0, 1] range\n```",
      "votes": null
    },
    {
      "id": "821221",
      "postDate": "04/26/2020 02:51:36",
      "content": "<p>Thank you for good information. I have an additional question. If we use own public GCS bucket, do we need to choose the location of the bucket to locate it on the same location of TPU?   And if it is true, how to check the location of kaggle TPU?</p>",
      "rawMarkdown": "Thank you for good information. I have an additional question. If we use own public GCS bucket, do we need to choose the location of the bucket to locate it on the same location of TPU?   And if it is true, how to check the location of kaggle TPU?",
      "votes": null
    },
    {
      "id": "940684",
      "postDate": "07/23/2020 04:50:13",
      "content": "<p>Thank you! This helped a lot!</p>",
      "rawMarkdown": "Thank you! This helped a lot!",
      "votes": null
    },
    {
      "id": "1379083",
      "postDate": "07/07/2021 05:14:39",
      "content": "<p>make your dataset public</p>",
      "rawMarkdown": "make your dataset public",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1379083,
      "author_name": "benedicths",
      "author_url": "",
      "post_date": "07/07/2021 05:14:39",
      "content": "<p>make your dataset public</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 745366,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "02/13/2020 18:38:51",
      "content": "<p>Two possible ways:\n1. Use your own public GCS bucket. Access it with <code>gs://my-bucket-name/...</code>. Note that if this is your bucket you pay for it.\n2. Upload your data in a public Kaggale Dataset. The get the GCS path to it with: <code>MY_GCS_PATH=KaggleDatasets().get_gcs_path('my-kaggle-dataset')</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 745367,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/13/2020 18:40:23",
          "content": "<p>The directory name of the Kaggle dataset (i.e. the name to pass into <code>get_gcs_dataset('name')</code>) can be looked up by attaching the dataset to your notebook then running this in a cell:\n<code>\n!ls /kaggle/input\n</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745574,
          "author_name": "qinhui1999",
          "author_url": "",
          "post_date": "02/14/2020 00:24:11",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> Thanks.  It seems I have to make my dataset public.  :(</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746161,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/14/2020 17:29:20",
          "content": "<p>At the moment yes, unfortunately.\nI would add that this is a friendly \"playground\" competition and knowledge sharing is the goal.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746208,
          "author_name": "mikemollel",
          "author_url": "",
          "post_date": "02/14/2020 18:32:20",
          "content": "<p>I tried to do as you suggested; however, I still get the same error. I added the data and named 'test set' when I load using GCS_PATH = KaggleDatasets().get_gcs_path('test-set'). I get the following error</p>\n\n<p>BackendError: Unexpected response from the service. Response: {'errors': [\"Dataset not found for directory 'test-set', please make sure you are passing a valid directory name under /kaggle/input\"], 'error': {'code': 5}, 'wasSuccessful': False}.I tried to do as you suggested, however, I still get the same error. </p>\n\n<p>please help me with this</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746291,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/14/2020 20:10:29",
          "content": "<p>List the directories you have under /kaggle/input/ with:\n<code>\n!ls /kaggle/input\n</code>\nThen use the directory name of your dataset. I would bet it is something like 'test_set'</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746979,
          "author_name": "mikemollel",
          "author_url": "",
          "post_date": "02/15/2020 20:00:16",
          "content": "<p>Hello Martin</p>\n\n<p>Could you share with me if you don't mind the kernel of how you generate tfrecords from images? I guess the kernel resizes the individual image with different size and format to jpg to either of the square dimensions 331,224, etc.,Hello Martin</p>\n\n<p>Could you share with me if you don't mind the kernel of how you generate tfrecords from images?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 749576,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/18/2020 19:40:03",
          "content": "<p>You will find all the details about how TFRecords work in this tutorial: <a href=\"https://codelabs.developers.google.com/codelabs/keras-flowers-data/#0\">TPU-speed data pipelines: tf.data.Dataset and TFRecords</a></p>\n\n<p>The notebook generating the TFRecord dataset in the tutorial is this one (Colab link):\n<a href=\"https://colab.research.google.com/github/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/03_Flower_pictures_to_TFRecords.ipynb\">Flower pictures to TFRecords.ipynb</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 757564,
          "author_name": "ghostskipper",
          "author_url": "",
          "post_date": "02/26/2020 23:02:28",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> I've been trying to use my own GSC bucket. \nI have made my own bucket (based off <a href=\"https://www.kaggle.com/yufengg/automl-getting-started-notebook\">this</a> kernel) and uploaded the example tfrecords. I have used the Add-ons to add my Google Cloud access to the kernel.</p>\n\n<p>The next step is getting the kaggle kernel to use my GSC bucket instead of the kaggle one.\nIf I try the following \n<code>\nGCS_DS_PATH = 'gs://my-bucket-name'\nTRAINING_FILENAMES = tf.io.gfile.glob(GCS_PATH + '/tfrecords-jpeg-512x512/train/*.tfrec')\n</code>\nI get \n<code>\nPermissionDeniedError: Error executing an HTTP request: HTTP response code 401 with body '{\n  \"error\": {\n    \"code\": 401,\n    \"message\": \"Anonymous caller does not have storage.objects.list access to my-bucket-name.\",\n    \"errors\": [\n      {\n        \"message\": \"Anonymous caller does not have storage.objects.list access to my-bucket-name.\",\n        \"domain\": \"global\",\n        \"reason\": \"required\",\n        \"locationType\": \"header\",\n        \"location\": \"Authorization\"\n      }\n    ]\n  }\n}\n</code></p>\n\n<p>This is not too surprising as it is not a public bucket. However in the kernel I do have a bucket object\n<code>\nbucket = storage_client.get_bucket(bucket_name)\n</code>\nand I can get a list of my blobs in it using \n<code>\nblobs = storage_client.list_blobs(bucket_name)\nfor blob in blobs:\n        print(blob.name)\n</code></p>\n\n<p>My understanding is if I pass a list my full paths as <code>[gs://my-bucket-name/my.tfrec]</code> to <code>dataset = tf.data.TFRecordDataset(filenames, num_parallel_reads=AUTO)</code> in the <code>load_dataset</code> function it will also have the permission issue. </p>\n\n<p>My question is how do I use my bucket with the kaggle TPUs? </p>\n\n<p>Thanks </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 758575,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/28/2020 00:03:38",
          "content": "<p>We have done the integration with the GCS API <code>storage_client.*</code>, not with the Tensorflow GCS API <code>tf.io.gfile.*</code>. So when you add your Google credentials in add-ons, it makes the GCS API work but not the other one.</p>\n\n<p>I'm going to give you some hints. I'm not sure it is even possible to make it work and I think you'd better use a Notebook VM on GCP where authentication is automatic within the same project and between projects, you can use service accounts.</p>\n\n<p>(Please NEVER grant the Kaggle TPU service account  any permissions. That service account is owned by Kaggle, not you. If you do that, you have no idea who you are granting permissions to.)</p>\n\n<p>For Tensorflow access to GCS to work, you will need to to export a service account's key as JSON and \n configure:\n<code>export GOOGLE_APPLICATION_CREDENTIALS=/path/to/credential/json/file</code></p>\n\n<p>I do recommend you to use the Kaggle secrets service to store this JSON file securely.</p>\n\n<p>And that only gives your Kaggle VM access to your private GCS bucket. The TPU needs more work.</p>\n\n<p>The TPU is another machine and it will be accessing the bucket directly. It needs to be authorized too. This is the code that used to work in TF 1.x. It obviously does not work anymore:\n```\nTF_MASTER = 'grpc://{}'.format(os.environ['COLAB_TPU_ADDR'])</p>\n\n<h1>Upload credentials to TPU.</h1>\n\n<p>with tf.Session(TF_MASTER) as sess: <br>\n    with open('/content/adc.json', 'r') as f:\n        auth_info = json.load(f)\n    tf.contrib.cloud.configure_gcs(sess, credentials=auth_info)\n```\nThe configure_gcs function now lives in module <a href=\"https://www.tensorflow.org/io/api_docs/python/tfio/gcs\">tfio.gcs</a> </p>\n\n<p>From here, you're on your own. Good luck and share your solution if you manage to make it work (securely!)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 758587,
          "author_name": "ghostskipper",
          "author_url": "",
          "post_date": "02/28/2020 00:20:16",
          "content": "<p>Thank you. I'll give it a try</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 811492,
          "author_name": "mikemollel",
          "author_url": "",
          "post_date": "04/18/2020 01:22:51",
          "content": "<p>Thank you for the notebook it helped e a lot to understand on how to generate tfrecords. But I have a question on how to deal with the dt_uint8 which you used in your code. I failed to use TPU since it says the data type used for conversion image is not supported. I tried to cast to the image to float32 when I read image still I couldn't be able to run my code on TPU though it runs just fine on GPU.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 814671,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "04/20/2020 21:20:53",
          "content": "<p>The <a href=\"https://www.kaggle.com/mgornergoogle/five-flowers-with-keras-and-xception-on-tpu\">TPU sample in the documentation</a> loads image data as uint8 and trains a model on them on TPU. The conversion to float 32 happens in function <code>read_tfrecord</code>. The line of code that does the conversion is:\n```</p>\n\n<p>image = tf.cast(image, tf.float32) / 255.0  # convert image to floats in [0, 1] range\n```</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 821221,
          "author_name": "sai11fkaneko",
          "author_url": "",
          "post_date": "04/26/2020 02:51:36",
          "content": "<p>Thank you for good information. I have an additional question. If we use own public GCS bucket, do we need to choose the location of the bucket to locate it on the same location of TPU?   And if it is true, how to check the location of kaggle TPU?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 746914,
      "author_name": "mikemollel",
      "author_url": "",
      "post_date": "02/15/2020 18:22:27",
      "content": "<p>Thank you. It worked, however, I had to change the name of the folder from data_set to data-set I don't know why it makes me do like this.Thank you it worked, however, I had to change the name of the folder from data_set to data-set I don't know why it makes me do like this.</p>",
      "votes": null,
      "replies": [
        {
          "id": 749574,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/18/2020 19:37:56",
          "content": "<p>The name to pass into the <code>KaggleDatasets().get_gcs_path(name)</code> function is the <strong>name of the directory</strong> where the dataset was mounted.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 940684,
      "author_name": "fireheart7",
      "author_url": "",
      "post_date": "07/23/2020 04:50:13",
      "content": "<p>Thank you! This helped a lot!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "745117": "Hi,\n\nI try to add some other data source into my kernel with tpu.  Then my kernel has two datas.  e.g. \nGCS_DS_PATH =KaggleDatasets().get_gcs_path('flower-classification')\nMY_DATA_PATH=KaggleDatasets().get_gcs_path(' mydata/')\n\nWhen I ran the codes , I got the errors:\nBackendError: Unexpected response from the service. Response: {'errors': ['Only public datasets and competitions are currently supported'], 'error': {'code': 12}, 'wasSuccessful': False}.\n\n\nWhat should I do ? Should I make my private data source public ?",
    "745366": "Two possible ways:\n1. Use your own public GCS bucket. Access it with `gs://my-bucket-name/...`. Note that if this is your bucket you pay for it.\n2. Upload your data in a public Kaggale Dataset. The get the GCS path to it with: `MY_GCS_PATH=KaggleDatasets().get_gcs_path('my-kaggle-dataset')`",
    "745367": "The directory name of the Kaggle dataset (i.e. the name to pass into `get_gcs_dataset('name')`) can be looked up by attaching the dataset to your notebook then running this in a cell:\n```\n!ls /kaggle/input\n```",
    "745574": "mgornergoogle Thanks.  It seems I have to make my dataset public.  :(",
    "746161": "At the moment yes, unfortunately.\nI would add that this is a friendly \"playground\" competition and knowledge sharing is the goal.",
    "746208": "I tried to do as you suggested; however, I still get the same error. I added the data and named 'test set' when I load using GCS_PATH = KaggleDatasets().get_gcs_path('test-set'). I get the following error\n\nBackendError: Unexpected response from the service. Response: {'errors': [\"Dataset not found for directory 'test-set', please make sure you are passing a valid directory name under /kaggle/input\"], 'error': {'code': 5}, 'wasSuccessful': False}.I tried to do as you suggested, however, I still get the same error. \n\nplease help me with this",
    "746291": "List the directories you have under /kaggle/input/ with:\n```\n!ls /kaggle/input\n```\nThen use the directory name of your dataset. I would bet it is something like 'test_set'",
    "746914": "Thank you. It worked, however, I had to change the name of the folder from data_set to data-set I don't know why it makes me do like this.Thank you it worked, however, I had to change the name of the folder from data_set to data-set I don't know why it makes me do like this.",
    "746979": "Hello Martin\n\nCould you share with me if you don't mind the kernel of how you generate tfrecords from images? I guess the kernel resizes the individual image with different size and format to jpg to either of the square dimensions 331,224, etc.,Hello Martin\n\nCould you share with me if you don't mind the kernel of how you generate tfrecords from images?",
    "749574": "The name to pass into the `KaggleDatasets().get_gcs_path(name)` function is the **name of the directory** where the dataset was mounted.",
    "749576": "You will find all the details about how TFRecords work in this tutorial: [TPU-speed data pipelines: tf.data.Dataset and TFRecords](https://codelabs.developers.google.com/codelabs/keras-flowers-data/#0)\n\nThe notebook generating the TFRecord dataset in the tutorial is this one (Colab link):\n[Flower pictures to TFRecords.ipynb](https://colab.research.google.com/github/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/03_Flower_pictures_to_TFRecords.ipynb)",
    "757564": "mgornergoogle I've been trying to use my own GSC bucket. \nI have made my own bucket (based off [this](https://www.kaggle.com/yufengg/automl-getting-started-notebook) kernel) and uploaded the example tfrecords. I have used the Add-ons to add my Google Cloud access to the kernel.\n\nThe next step is getting the kaggle kernel to use my GSC bucket instead of the kaggle one.\nIf I try the following \n```\nGCS_DS_PATH = 'gs://my-bucket-name'\nTRAINING_FILENAMES = tf.io.gfile.glob(GCS_PATH + '/tfrecords-jpeg-512x512/train/*.tfrec')\n```\nI get \n```\nPermissionDeniedError: Error executing an HTTP request: HTTP response code 401 with body '{\n  \"error\": {\n    \"code\": 401,\n    \"message\": \"Anonymous caller does not have storage.objects.list access to my-bucket-name.\",\n    \"errors\": [\n      {\n        \"message\": \"Anonymous caller does not have storage.objects.list access to my-bucket-name.\",\n        \"domain\": \"global\",\n        \"reason\": \"required\",\n        \"locationType\": \"header\",\n        \"location\": \"Authorization\"\n      }\n    ]\n  }\n}\n```\n\nThis is not too surprising as it is not a public bucket. However in the kernel I do have a bucket object\n```\nbucket = storage_client.get_bucket(bucket_name)\n```\nand I can get a list of my blobs in it using \n```\nblobs = storage_client.list_blobs(bucket_name)\nfor blob in blobs:\n        print(blob.name)\n```\n\nMy understanding is if I pass a list my full paths as `[gs://my-bucket-name/my.tfrec]` to `dataset = tf.data.TFRecordDataset(filenames, num_parallel_reads=AUTO)` in the `load_dataset` function it will also have the permission issue. \n\nMy question is how do I use my bucket with the kaggle TPUs? \n\nThanks",
    "758575": "We have done the integration with the GCS API `storage_client.*`, not with the Tensorflow GCS API `tf.io.gfile.*`. So when you add your Google credentials in add-ons, it makes the GCS API work but not the other one.\n\nI'm going to give you some hints. I'm not sure it is even possible to make it work and I think you'd better use a Notebook VM on GCP where authentication is automatic within the same project and between projects, you can use service accounts.\n\n(Please NEVER grant the Kaggle TPU service account  any permissions. That service account is owned by Kaggle, not you. If you do that, you have no idea who you are granting permissions to.)\n\nFor Tensorflow access to GCS to work, you will need to to export a service account's key as JSON and \n configure:\n`export GOOGLE_APPLICATION_CREDENTIALS=/path/to/credential/json/file `\n\nI do recommend you to use the Kaggle secrets service to store this JSON file securely.\n\nAnd that only gives your Kaggle VM access to your private GCS bucket. The TPU needs more work.\n\nThe TPU is another machine and it will be accessing the bucket directly. It needs to be authorized too. This is the code that used to work in TF 1.x. It obviously does not work anymore:\n```\nTF_MASTER = 'grpc://{}'.format(os.environ['COLAB_TPU_ADDR'])\n# Upload credentials to TPU.\nwith tf.Session(TF_MASTER) as sess:    \n    with open('/content/adc.json', 'r') as f:\n        auth_info = json.load(f)\n    tf.contrib.cloud.configure_gcs(sess, credentials=auth_info)\n```\nThe configure_gcs function now lives in module [tfio.gcs](https://www.tensorflow.org/io/api_docs/python/tfio/gcs) \n\nFrom here, you're on your own. Good luck and share your solution if you manage to make it work (securely!)",
    "758587": "Thank you. I'll give it a try",
    "811492": "Thank you for the notebook it helped e a lot to understand on how to generate tfrecords. But I have a question on how to deal with the dt_uint8 which you used in your code. I failed to use TPU since it says the data type used for conversion image is not supported. I tried to cast to the image to float32 when I read image still I couldn't be able to run my code on TPU though it runs just fine on GPU.",
    "814671": "The [TPU sample in the documentation](https://www.kaggle.com/mgornergoogle/five-flowers-with-keras-and-xception-on-tpu) loads image data as uint8 and trains a model on them on TPU. The conversion to float 32 happens in function `read_tfrecord`. The line of code that does the conversion is:\n```\n\nimage = tf.cast(image, tf.float32) / 255.0  # convert image to floats in [0, 1] range\n```",
    "821221": "Thank you for good information. I have an additional question. If we use own public GCS bucket, do we need to choose the location of the bucket to locate it on the same location of TPU?   And if it is true, how to check the location of kaggle TPU?",
    "940684": "Thank you! This helped a lot!",
    "1379083": "make your dataset public"
  },
  "source": "meta"
}