{
  "id": 129817,
  "title": "TPU in Kaggle kernel?",
  "url": "/competitions/deepfake-detection-challenge/discussion/129817",
  "author_name": "Darko Androcec",
  "post_date": "2020-02-10T22:06:35.525000",
  "votes": 11,
  "comment_count": 69,
  "views": 0,
  "content": "<p>I see now in my kernel the option to enable TPU. Is this a new Kaggle feature? Do anyone have more information, e.g. did you try it and has it time limitations as GPUs on Kaggle?</p>",
  "messages": [
    {
      "id": 741686,
      "postDate": "2020-02-10T22:06:35.527Z",
      "content": "<p>I see now in my kernel the option to enable TPU. Is this a new Kaggle feature? Do anyone have more information, e.g. did you try it and has it time limitations as GPUs on Kaggle?</p>",
      "rawMarkdown": "I see now in my kernel the option to enable TPU. Is this a new Kaggle feature? Do anyone have more information, e.g. did you try it and has it time limitations as GPUs on Kaggle?",
      "votes": 11
    },
    {
      "id": 741696,
      "postDate": "2020-02-10T22:22:58.280Z",
      "content": "<p>if any google overlords are following this, what zone are these tpus in? </p>",
      "rawMarkdown": "if any google overlords are following this, what zone are these tpus in? ",
      "votes": 1,
      "replies": [
        {
          "id": 741709,
          "postDate": "2020-02-10T22:42:23.097Z",
          "content": "<p>Any particular reason you need to know? We don't really make any guarantees about this since we use resources from many different zones.</p>",
          "rawMarkdown": "Any particular reason you need to know? We don't really make any guarantees about this since we use resources from many different zones."
        },
        {
          "id": 741717,
          "postDate": "2020-02-10T22:51:37.017Z",
          "content": "<p>Bucket access costs. All my tfrecords are in central1 so if these tpus are coming from europe I'm not sure how that would affect the bucket fees</p>",
          "rawMarkdown": "Bucket access costs. All my tfrecords are in central1 so if these tpus are coming from europe I'm not sure how that would affect the bucket fees",
          "votes": 2
        },
        {
          "id": 741722,
          "postDate": "2020-02-10T22:56:25.750Z",
          "content": "<p>The TPUs you get can be in various zones but we do take care of data locality. As you can see in the <a href=\"https://www.kaggle.com/mgornergoogle/five-flowers-with-keras-and-xception-on-tpu\">TPU sample</a>, kaggle will copy the dataset to GCS (Google Cloud Storage) when you ask for GCS access to it through <code>KaggleDatasets().get_gcs_path()</code>. The data will be copied to a cache bucket in the same location as the TPU you have been allocated.</p>",
          "rawMarkdown": "The TPUs you get can be in various zones but we do take care of data locality. As you can see in the [TPU sample](https://www.kaggle.com/mgornergoogle/five-flowers-with-keras-and-xception-on-tpu), kaggle will copy the dataset to GCS (Google Cloud Storage) when you ask for GCS access to it through `KaggleDatasets().get_gcs_path()`. The data will be copied to a cache bucket in the same location as the TPU you have been allocated.",
          "votes": 1
        },
        {
          "id": 741723,
          "postDate": "2020-02-10T22:57:32.040Z",
          "content": "<p>Bucket access cost is covered by the Kaggle integration as long as you use pblic Kaggle datasets. We have not integrated private Kaggle datasets with TPUs yet. </p>",
          "rawMarkdown": "Bucket access cost is covered by the Kaggle integration as long as you use pblic Kaggle datasets. We have not integrated private Kaggle datasets with TPUs yet. ",
          "votes": 2
        },
        {
          "id": 741724,
          "postDate": "2020-02-10T22:57:33.283Z",
          "content": "<p>Nice. Can I take it that if I'm using the tpu to train from a private bucket in central1, it won't spike my bucket access costs or it will be similar to if I had set up a v3-8 in central1 from a cost perspective? </p>",
          "rawMarkdown": "Nice. Can I take it that if I'm using the tpu to train from a private bucket in central1, it won't spike my bucket access costs or it will be similar to if I had set up a v3-8 in central1 from a cost perspective? "
        },
        {
          "id": 741746,
          "postDate": "2020-02-10T23:25:29.320Z",
          "content": "<p>Not quite, if you use a TPU to train on a <strong>public Kaggle dataset</strong>, the GCS costs (because data will be coming from GCS in that case as well) will be covered by Kaggle.\nTraining on your own private bucket is not yet officially supported. I expect it will work if you know how to set IAMs correctly. If you manage to make it work, you are still paying for usage of your own bucket.</p>\n\n<p>Also, if you do make private buckets work, tell us how you did it exactly. It is not a feature we focused on at launch but we still want to hear about the user experience, especially if it could be improved.</p>",
          "rawMarkdown": "Not quite, if you use a TPU to train on a **public Kaggle dataset**, the GCS costs (because data will be coming from GCS in that case as well) will be covered by Kaggle.\nTraining on your own private bucket is not yet officially supported. I expect it will work if you know how to set IAMs correctly. If you manage to make it work, you are still paying for usage of your own bucket.\n\nAlso, if you do make private buckets work, tell us how you did it exactly. It is not a feature we focused on at launch but we still want to hear about the user experience, especially if it could be improved."
        },
        {
          "id": 741782,
          "postDate": "2020-02-11T00:16:49.517Z",
          "content": "<p>lulz so I fed the filenames to the tfrecord using my bucket. So something like gs://mybucket/tfrecord. </p>\n\n<p>This created an error message when trying to fit the model. Below is a copy paste of the relevant part of the error. I then went to IAM and then just granted rights to the TPU service account. Needed to reinit the TPU, but otherwise it got access. </p>\n\n<p>PermissionDeniedError: Error executing an HTTP request: HTTP response code 403 with body '{\n\"error\": {\n\"code\": 403,\n\"message\": \"service-467472385656@cloud-tpu.iam.gserviceaccount.com does not have storage.objects.list access to deepfakestorage.\",\n\"errors\": [\n{\n\"message\": \"service-467472385656@cloud-tpu.iam.gserviceaccount.com does not have storage.objects.list access to deepfakestorage.\",\n\"domain\": \"global\",\n\"reason\": \"forbidden\"\n}\n]\n}\n}</p>",
          "rawMarkdown": "lulz so I fed the filenames to the tfrecord using my bucket. So something like gs://mybucket/tfrecord. \n\nThis created an error message when trying to fit the model. Below is a copy paste of the relevant part of the error. I then went to IAM and then just granted rights to the TPU service account. Needed to reinit the TPU, but otherwise it got access. \n\nPermissionDeniedError: Error executing an HTTP request: HTTP response code 403 with body '{\n\"error\": {\n\"code\": 403,\n\"message\": \"service-467472385656@cloud-tpu.iam.gserviceaccount.com does not have storage.objects.list access to deepfakestorage.\",\n\"errors\": [\n{\n\"message\": \"service-467472385656@cloud-tpu.iam.gserviceaccount.com does not have storage.objects.list access to deepfakestorage.\",\n\"domain\": \"global\",\n\"reason\": \"forbidden\"\n}\n]\n}\n}\n"
        },
        {
          "id": 741784,
          "postDate": "2020-02-11T00:19:17.687Z",
          "content": "<p>Also I made a kernel explaining TPUs and tfrecord. <a href=\"https://www.kaggle.com/hooong/how-to-tpu?scriptVersionId=28471599\">https://www.kaggle.com/hooong/how-to-tpu?scriptVersionId=28471599</a></p>\n\n<p>It pretty much desribes how I use TPUs in google cloud and kaggle </p>",
          "rawMarkdown": "Also I made a kernel explaining TPUs and tfrecord. https://www.kaggle.com/hooong/how-to-tpu?scriptVersionId=28471599\n\nIt pretty much desribes how I use TPUs in google cloud and kaggle "
        },
        {
          "id": 741796,
          "postDate": "2020-02-11T00:29:14.423Z",
          "content": "<p>So I had a plan to cycle through TPU sessions until I get one that's in zone us central1. </p>\n\n<p>Unfortuantely, kaggle TPUs do not have the normal TPU attributes. </p>\n\n<p>```\ntpu = tf.distribute.cluster_resolver.TPUClusterResolver()\ntpu.<strong>dict</strong></p>\n\n<p>```\nreturn \n{'_tpu': b'10.0.0.2:8470',\n 'task_type': 'worker',\n 'task_id': 0,\n '_environment': '',\n 'rpc_layer': 'grpc',\n '_should_resolve_override': False,\n '_service': None,\n '_credentials': 'default',\n '_project': None,\n '_zone': None,\n '_discovery_url': None,\n '_coordinator_name': None,\n '_coordinator_address': None} </p>\n\n<p>When I run the same thing in google cloud I get </p>\n\n<p>{'_coordinator_address': None,\n '_coordinator_name': None,\n '_credentials': None,\n '_discovery_url': None,\n '_environment': '',\n '_project': 'tpu-44747',\n '_service': None,\n '_should_resolve_override': None,\n '_tpu': b'node-1',\n '_zone': 'us-central1-b',\n 'rpc_layer': 'grpc',\n 'task_id': 0,\n 'task_type': 'worker'}</p>",
          "rawMarkdown": "So I had a plan to cycle through TPU sessions until I get one that's in zone us central1. \n\nUnfortuantely, kaggle TPUs do not have the normal TPU attributes. \n\n```\ntpu = tf.distribute.cluster_resolver.TPUClusterResolver()\ntpu.__dict__\n\n```\nreturn \n{'_tpu': b'10.0.0.2:8470',\n 'task_type': 'worker',\n 'task_id': 0,\n '_environment': '',\n 'rpc_layer': 'grpc',\n '_should_resolve_override': False,\n '_service': None,\n '_credentials': 'default',\n '_project': None,\n '_zone': None,\n '_discovery_url': None,\n '_coordinator_name': None,\n '_coordinator_address': None} \n\nWhen I run the same thing in google cloud I get \n\n{'_coordinator_address': None,\n '_coordinator_name': None,\n '_credentials': None,\n '_discovery_url': None,\n '_environment': '',\n '_project': 'tpu-44747',\n '_service': None,\n '_should_resolve_override': None,\n '_tpu': b'node-1',\n '_zone': 'us-central1-b',\n 'rpc_layer': 'grpc',\n 'task_id': 0,\n 'task_type': 'worker'}"
        },
        {
          "id": 741798,
          "postDate": "2020-02-11T00:30:23.997Z",
          "content": "<p>what were you looking for in there ?</p>",
          "rawMarkdown": "what were you looking for in there ?"
        },
        {
          "id": 741809,
          "postDate": "2020-02-11T00:41:29.360Z",
          "content": "<p>Thank you for the notebook with TPU explanation. I went through it and added some comments.</p>",
          "rawMarkdown": "Thank you for the notebook with TPU explanation. I went through it and added some comments."
        },
        {
          "id": 741892,
          "postDate": "2020-02-11T01:49:06.367Z",
          "content": "<p>I was trying to check the zone of the tpu hoping that the tpu would show up as zone 'us-central1-b or similar. I'm concerned about egress bc it isn't cheap ($0.01 per gb for not free cloud services, bunch of conditions to determine if free but generally free if same region). I'm actually not sure how kaggle tpu data access will be charged and I don't know if it will count as a google cloud service accessing my bucket. </p>",
          "rawMarkdown": "I was trying to check the zone of the tpu hoping that the tpu would show up as zone 'us-central1-b or similar. I'm concerned about egress bc it isn't cheap ($0.01 per gb for not free cloud services, bunch of conditions to determine if free but generally free if same region). I'm actually not sure how kaggle tpu data access will be charged and I don't know if it will count as a google cloud service accessing my bucket. "
        },
        {
          "id": 741903,
          "postDate": "2020-02-11T02:02:06.543Z",
          "content": "<p>By the way, if you care about the data in your GCS bucket, DO NOT give the Kaggle TPU service account access. You do not control this TPU and you do not know who can use it to access your data.\nIf you do not care much about the data in your GCS bucket, make it public. You won't have to worry about IAMs anymore.</p>",
          "rawMarkdown": "By the way, if you care about the data in your GCS bucket, DO NOT give the Kaggle TPU service account access. You do not control this TPU and you do not know who can use it to access your data.\nIf you do not care much about the data in your GCS bucket, make it public. You won't have to worry about IAMs anymore."
        },
        {
          "id": 741909,
          "postDate": "2020-02-11T02:04:10.173Z",
          "content": "<p>And yes I'm pretty sure you will be charged for egress from your GCS bucket. Even if it is a Google cloud service, those TPUs are owned by Kaggle, not your by your GCP project.</p>",
          "rawMarkdown": "And yes I'm pretty sure you will be charged for egress from your GCS bucket. Even if it is a Google cloud service, those TPUs are owned by Kaggle, not your by your GCP project."
        },
        {
          "id": 745394,
          "postDate": "2020-02-13T19:18:43.443Z",
          "content": "<p>so I ran a test and I didn't get any egress fees using a multi region US bucket. But even if I did get egress fees, it is quite affordable at less than a penny per GB.</p>",
          "rawMarkdown": "so I ran a test and I didn't get any egress fees using a multi region US bucket. But even if I did get egress fees, it is quite affordable at less than a penny per GB."
        }
      ]
    },
    {
      "id": 741695,
      "postDate": "2020-02-10T22:22:40.787Z",
      "content": "<p>WOW! 30hr of free TPU on Kaggle Kernels!!! </p>",
      "rawMarkdown": "WOW! 30hr of free TPU on Kaggle Kernels!!! ",
      "votes": 1,
      "replies": [
        {
          "id": 741699,
          "postDate": "2020-02-10T22:25:00.733Z",
          "content": "<p>its a v3-8 </p>\n\n<p>so thats the newer tpu architecture and you get 8 tpu units. </p>\n\n<p>For reference, I've been using v2-8 and it computes a batch of 512, 256, 256, 3 in 300ms on resnet50. That's 512 images at 256 by 256. </p>\n\n<p>It is INSANE how much faster it is than gpus. IMO, v3 is mostly improvements for RNNs, but still good for CNNs. </p>",
          "rawMarkdown": "its a v3-8 \n\nso thats the newer tpu architecture and you get 8 tpu units. \n\nFor reference, I've been using v2-8 and it computes a batch of 512, 256, 256, 3 in 300ms on resnet50. That's 512 images at 256 by 256. \n\nIt is INSANE how much faster it is than gpus. IMO, v3 is mostly improvements for RNNs, but still good for CNNs. \n\n\n\n",
          "votes": 6
        },
        {
          "id": 741728,
          "postDate": "2020-02-10T23:00:42.960Z",
          "content": "<p>V3s have roughly double the matrix multiplication hardware of V2s. Utilization of the additional hardware will depend on your model though.</p>",
          "rawMarkdown": "V3s have roughly double the matrix multiplication hardware of V2s. Utilization of the additional hardware will depend on your model though."
        },
        {
          "id": 741729,
          "postDate": "2020-02-10T23:02:22.737Z",
          "content": "<p>That's really great! I am using SageMaker so far with the lowest spec GPU... After your comment, I am now definitely gonna try the 60days v3-8 TPU credit I got (and I was lazy to setup)! + Kaggle 30hrs is there too... ^-^</p>",
          "rawMarkdown": "That's really great! I am using SageMaker so far with the lowest spec GPU... After your comment, I am now definitely gonna try the 60days v3-8 TPU credit I got (and I was lazy to setup)! + Kaggle 30hrs is there too... ^-^"
        },
        {
          "id": 741738,
          "postDate": "2020-02-10T23:13:22.840Z",
          "content": "<p>If you want to use TPUs on GCP with <a href=\"https://console.cloud.google.com/ai-platform/notebooks\">Cloud AI platform Notebooks</a>, they work, but the \"TPU accelerator\" option is not yet in the UI. You have to create the Notebook VM + TPU pair through the command line. I have a script here: <a href=\"https://github.com/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/README-TF2.1.md\">bit.ly/keras-tpu-tf21</a></p>",
          "rawMarkdown": "If you want to use TPUs on GCP with [Cloud AI platform Notebooks](https://console.cloud.google.com/ai-platform/notebooks), they work, but the \"TPU accelerator\" option is not yet in the UI. You have to create the Notebook VM + TPU pair through the command line. I have a script here: [bit.ly/keras-tpu-tf21](https://github.com/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/README-TF2.1.md)",
          "votes": 2
        },
        {
          "id": 741745,
          "postDate": "2020-02-10T23:25:00.753Z",
          "content": "<p>Thanks <a href=\"/mgornergoogle\">@mgornergoogle</a> !</p>",
          "rawMarkdown": "Thanks @mgornergoogle !"
        },
        {
          "id": 741755,
          "postDate": "2020-02-10T23:36:28.497Z",
          "content": "<p>The other significant difference between TPU v2 and TPU v3 is 64GB of HBM vs 128GB for a TPU v3. (HBM = High Bandwidth Memory)</p>",
          "rawMarkdown": "The other significant difference between TPU v2 and TPU v3 is 64GB of HBM vs 128GB for a TPU v3. (HBM = High Bandwidth Memory)",
          "votes": 1
        },
        {
          "id": 741774,
          "postDate": "2020-02-11T00:05:28.447Z",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> I am trying to create a new notebook instance in AI Platform, but I cannot create a custom notebook in the <code>europe-west4-a</code> zone from UI. Is this zone not supported or I have to set it up in the cloud shell? I am mentioning this because my TPU credits are only for this zone. Thanks in advance.</p>",
          "rawMarkdown": "@mgornergoogle I am trying to create a new notebook instance in AI Platform, but I cannot create a custom notebook in the ```europe-west4-a``` zone from UI. Is this zone not supported or I have to set it up in the cloud shell? I am mentioning this because my TPU credits are only for this zone. Thanks in advance."
        },
        {
          "id": 741794,
          "postDate": "2020-02-11T00:28:47.510Z",
          "content": "<p>To use TPUs in GCP's Cloud AI Notebooks, use this script: <a href=\"https://raw.githubusercontent.com/GoogleCloudPlatform/training-data-analyst/master/courses/fast-and-lean-data-science/create-tpu-deep-learning-vm.sh\">create-tpu-deep-learning-vm.sh</a></p>\n\n<p><code>gcloud init</code>\nthen\n<code>./create-tpu-deep-learning-vm.sh choose-a-name --tpu-type v3-8</code></p>\n\n<p>If you open the script, there is no magic, just the two gcloud commands that create the VM and the TPU. There is a tiny bit of magic to set up the TPU_NAME environment variable on the VM. That's what makes the TPUCusterResolver() call work without parameters in your Python code. The second tiny bit of magic is the setting that makes the VM work as a Notebook VM with a Jupyter proxy.</p>",
          "rawMarkdown": "To use TPUs in GCP's Cloud AI Notebooks, use this script: [create-tpu-deep-learning-vm.sh](https://raw.githubusercontent.com/GoogleCloudPlatform/training-data-analyst/master/courses/fast-and-lean-data-science/create-tpu-deep-learning-vm.sh)\n\n`gcloud init`\nthen\n`./create-tpu-deep-learning-vm.sh choose-a-name --tpu-type v3-8`\n\nIf you open the script, there is no magic, just the two gcloud commands that create the VM and the TPU. There is a tiny bit of magic to set up the TPU_NAME environment variable on the VM. That's what makes the TPUCusterResolver() call work without parameters in your Python code. The second tiny bit of magic is the setting that makes the VM work as a Notebook VM with a Jupyter proxy."
        },
        {
          "id": 741800,
          "postDate": "2020-02-11T00:30:51.093Z",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> great, will do it this way. Also, I will add <code>--zone=europe-west4-a</code></p>",
          "rawMarkdown": "@mgornergoogle great, will do it this way. Also, I will add ```--zone=europe-west4-a```"
        },
        {
          "id": 756839,
          "postDate": "2020-02-26T06:12:01.950Z",
          "content": "<p>Sir, can you share your code to define and train the model using TPUs. Getting some issues here!</p>",
          "rawMarkdown": "Sir, can you share your code to define and train the model using TPUs. Getting some issues here!"
        }
      ]
    },
    {
      "id": 742992,
      "postDate": "2020-02-11T17:27:26.023Z",
      "content": "<p>An update regarding kaggle tpu zones. </p>\n\n<p>It appears that kaggle tpus are most likely located somewhere in the united states. Yesterday, I ran a kaggle tpu that pulled data from my private us-central1 storage bucket. This incurred an egress fee. </p>\n\n<p>Cloud Storage Inter-region GCP Storage egress within NA: 1.73 </p>\n\n<p>I will attempt to clone my bucket to a multi region NA bucket and then test again to see if egress fees are charged. This will take a day bc transactions don't usually take a while to appear on the billing menu. </p>",
      "rawMarkdown": "An update regarding kaggle tpu zones. \n\nIt appears that kaggle tpus are most likely located somewhere in the united states. Yesterday, I ran a kaggle tpu that pulled data from my private us-central1 storage bucket. This incurred an egress fee. \n\nCloud Storage Inter-region GCP Storage egress within NA: 1.73 \n\nI will attempt to clone my bucket to a multi region NA bucket and then test again to see if egress fees are charged. This will take a day bc transactions don't usually take a while to appear on the billing menu. ",
      "votes": 2,
      "replies": [
        {
          "id": 742998,
          "postDate": "2020-02-11T17:33:32.230Z",
          "content": "<p>Our TPUs are not limited to the US. The only way to make sure your data is co-located to the TPU right now is to use Kaggle public datasets + GCS loader (see <a href=\"https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu#Competition-data-access\">https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu#Competition-data-access</a>). </p>\n\n<p>Private datasets are not officially supported due to technical/privacy limitations which we are working on.</p>",
          "rawMarkdown": "Our TPUs are not limited to the US. The only way to make sure your data is co-located to the TPU right now is to use Kaggle public datasets + GCS loader (see https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu#Competition-data-access). \n\nPrivate datasets are not officially supported due to technical/privacy limitations which we are working on."
        },
        {
          "id": 755892,
          "postDate": "2020-02-25T08:30:38.503Z",
          "content": "<p>Hello, I cannot get a hold of the dataset for the BengaliAI competition in GCP Bucket, Code I am using is \nGCS_DS_PATH = KaggleDatasets().get_gcs_path('/kaggle/input/bengaliai-cv19'). Getting error \"Dataset not found for directory '/kaggle/input/bengaliai-cv19',\". Can you help please. </p>",
          "rawMarkdown": "Hello, I cannot get a hold of the dataset for the BengaliAI competition in GCP Bucket, Code I am using is \nGCS_DS_PATH = KaggleDatasets().get_gcs_path('/kaggle/input/bengaliai-cv19'). Getting error \"Dataset not found for directory '/kaggle/input/bengaliai-cv19',\". Can you help please. "
        },
        {
          "id": 756487,
          "postDate": "2020-02-25T18:57:08.400Z",
          "content": "<p><a href=\"/anirbank\">@anirbank</a> You should only need to assign the path with\n<code>GCS_DS_PATH = KaggleDatasets().get_gcs_path()</code></p>\n\n<p>And then if you want to list the bucket, you can use:\n<code>!gsutil ls $GCS_DS_PATH</code></p>",
          "rawMarkdown": "@anirbank You should only need to assign the path with\n`GCS_DS_PATH = KaggleDatasets().get_gcs_path()`\n\nAnd then if you want to list the bucket, you can use:\n`!gsutil ls $GCS_DS_PATH`",
          "votes": 6
        },
        {
          "id": 756633,
          "postDate": "2020-02-25T22:55:35.827Z",
          "content": "<p>Or if you have multiple datasets in your Notebook:\n<code>\nGCS_DS_PATH = KaggleDatasets().get_gcs_path('bengaliai-cv19')\n</code>\nYou got it almost right but the function wants the name of the directory where the dataset has been mapped, not the full path.</p>",
          "rawMarkdown": "Or if you have multiple datasets in your Notebook:\n```\nGCS_DS_PATH = KaggleDatasets().get_gcs_path('bengaliai-cv19')\n```\nYou got it almost right but the function wants the name of the directory where the dataset has been mapped, not the full path.",
          "votes": 3
        },
        {
          "id": 756813,
          "postDate": "2020-02-26T05:27:15.840Z",
          "content": "<p>Excellent, it works fine now, thank you. Lets try TPU on my notebook now!</p>",
          "rawMarkdown": "Excellent, it works fine now, thank you. Lets try TPU on my notebook now!"
        },
        {
          "id": 756821,
          "postDate": "2020-02-26T05:48:07.137Z",
          "content": "<p>Although it worked for csv file, getting this error when reading feather file : \"failed to open local file 'gs://kds-87e1f7817c6764d20c2f2841fd9048ac1f7b9c89a1508dbd796f13b4/train_image_data_1.feather'. Detail: [errno 2] No such file or directory.\" This file is however listed when i do !gsutil ls $GCS_DS_PATH1, with GCS_DS_PATH1 = KaggleDatasets().get_gcs_path('bengaliaicv19feather') .  The same works fine, file is read for parquet files though. Any ideas here ? thanks</p>",
          "rawMarkdown": "Although it worked for csv file, getting this error when reading feather file : \"failed to open local file 'gs://kds-87e1f7817c6764d20c2f2841fd9048ac1f7b9c89a1508dbd796f13b4/train_image_data_1.feather'. Detail: [errno 2] No such file or directory.\" This file is however listed when i do !gsutil ls $GCS_DS_PATH1, with GCS_DS_PATH1 = KaggleDatasets().get_gcs_path('bengaliaicv19feather') .  The same works fine, file is read for parquet files though. Any ideas here ? thanks"
        },
        {
          "id": 756830,
          "postDate": "2020-02-26T05:57:38.477Z",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> Can read from parquet files (using GCP Bucket path), but now getting this error on model.fit :\" NotFoundError: {{function_node __inference_distributed_function_71636}} No registered 'PyFunc' OpKernel for 'CPU' devices compatible with node {{node PyFunc}}\n    .  Registered:  \" I have defined TPUStrategy and this is done fine, as given in the sample here. The TPU is visible in the output. I have defined the Model using strategy.scope() as shown. So not sure what the issue is. However, I am not using TFRecords, my input data is from numpy arrays, derived from the image files. I guess that should not be a issue.</p>",
          "rawMarkdown": "@mgornergoogle Can read from parquet files (using GCP Bucket path), but now getting this error on model.fit :\" NotFoundError: {{function_node __inference_distributed_function_71636}} No registered 'PyFunc' OpKernel for 'CPU' devices compatible with node {{node PyFunc}}\n\t.  Registered:  \" I have defined TPUStrategy and this is done fine, as given in the sample here. The TPU is visible in the output. I have defined the Model using strategy.scope() as shown. So not sure what the issue is. However, I am not using TFRecords, my input data is from numpy arrays, derived from the image files. I guess that should not be a issue.",
          "votes": 1
        },
        {
          "id": 911453,
          "postDate": "2020-07-01T18:33:22.730Z",
          "content": "<p>Have you solved it yet?</p>",
          "rawMarkdown": "Have you solved it yet?"
        },
        {
          "id": 911660,
          "postDate": "2020-07-01T23:15:04.250Z",
          "content": "<p>If the loading of parket files involves Python (non-tensorflow) code, then the TPU will not be able to run it. You have to store your data in a format readable by Tensorflow. Wrapping a pure Python function with PyFunc to turn it into a Tensorflow op unfortunately does not work on TPU.</p>",
          "rawMarkdown": "If the loading of parket files involves Python (non-tensorflow) code, then the TPU will not be able to run it. You have to store your data in a format readable by Tensorflow. Wrapping a pure Python function with PyFunc to turn it into a Tensorflow op unfortunately does not work on TPU."
        }
      ]
    },
    {
      "id": 741693,
      "postDate": "2020-02-10T22:21:01.930Z",
      "content": "<p>I've been using tpus the whole time on gcloud. </p>\n\n<p>I'm testing the kaggle tpu interface and will update. </p>\n\n<p>You will need to use the tfrecord format if you want to do tpu training to its fullest. </p>",
      "rawMarkdown": "I've been using tpus the whole time on gcloud. \n\nI'm testing the kaggle tpu interface and will update. \n\nYou will need to use the tfrecord format if you want to do tpu training to its fullest. \n",
      "votes": 2
    },
    {
      "id": 757649,
      "postDate": "2020-02-27T01:49:47.473Z",
      "content": "<p>Hi Everyone,\nI wanted to confirm if anyone was able to create dataset reader for TPU that doesn't require saving data to tfrecords. I basically want to try to use it for model training for this competition?\nI wasn't able to find a way to do so on kaggle.\nRegards,\nAmro</p>",
      "rawMarkdown": "Hi Everyone,\nI wanted to confirm if anyone was able to create dataset reader for TPU that doesn't require saving data to tfrecords. I basically want to try to use it for model training for this competition?\nI wasn't able to find a way to do so on kaggle.\nRegards,\nAmro",
      "replies": [
        {
          "id": 758580,
          "postDate": "2020-02-28T00:13:35.413Z",
          "content": "<p>Not a direct answer to your question but a clarification. TFRecords are not required for TPUs.</p>\n\n<p>The problem to solve is to feed the TPU with data fast enough. GCS can sustain the throughput but needs parallel reading from multiple files (16 shards for example) and has a per-file request penalty of a fraction of a second (as any web request). The penalty will get in the way if you have millions or 100Ks of files.</p>\n\n<p>So you need a streamable (streamable = NOT zip) container format to bunch your N files together into a reasonable number of shards (&gt;16 but not thousands).</p>\n\n<p>TFRecords are a solution to that problem but you are free to solve it in any way you like.</p>",
          "rawMarkdown": "Not a direct answer to your question but a clarification. TFRecords are not required for TPUs.\n\nThe problem to solve is to feed the TPU with data fast enough. GCS can sustain the throughput but needs parallel reading from multiple files (16 shards for example) and has a per-file request penalty of a fraction of a second (as any web request). The penalty will get in the way if you have millions or 100Ks of files.\n\nSo you need a streamable (streamable = NOT zip) container format to bunch your N files together into a reasonable number of shards (&gt;16 but not thousands).\n\nTFRecords are a solution to that problem but you are free to solve it in any way you like.",
          "votes": 2
        },
        {
          "id": 759087,
          "postDate": "2020-02-28T14:55:15.523Z",
          "content": "<p>Hi <a href=\"/mgornergoogle\">@mgornergoogle</a> \nThanks for your answer. Really appreciated. I believe the files here on Kaggle side are not readable. May be I should change my question, I could write up a generator similar to how it's written in Keras for training that could read from the files here and generate the tensors for tensorflow in Memory. Keras automated the part of reading using multiprocessing if I'm not mistaken. I find it hard to do the same using TPUs not GPUs?</p>\n\n<p>Any ideas about that? Basic sample code that could use a generator from Kaggle input folders? And we could take it from there.</p>\n\n<p>Thanks again for your input.</p>\n\n<p>Regards,\nAmro</p>",
          "rawMarkdown": "Hi @mgornergoogle \nThanks for your answer. Really appreciated. I believe the files here on Kaggle side are not readable. May be I should change my question, I could write up a generator similar to how it's written in Keras for training that could read from the files here and generate the tensors for tensorflow in Memory. Keras automated the part of reading using multiprocessing if I'm not mistaken. I find it hard to do the same using TPUs not GPUs?\n\nAny ideas about that? Basic sample code that could use a generator from Kaggle input folders? And we could take it from there.\n\nThanks again for your input.\n\nRegards,\nAmro"
        },
        {
          "id": 916069,
          "postDate": "2020-07-05T10:28:48.567Z",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> <br>\nCould you help me with some questions on effective TFRecord read/write and TPU usage, i can't piece together the numerous puzzle pieces even after going through this whole discussion.</p>\n\n<ol>\n<li><p>On the number of shards, i always see 16 shards (100-200mb each) suggested in tutorials kernels, while on Kaggle youtube last week, i hear \"If you have 100k images, use 16 shards\". Confusing me is my 100k images are 10GB, meaning 16 shards would be 600mb each, way more than 100-200mb. So is the sharding concept about number of shards, or file size per shard?</p></li>\n<li><p>What's the difference between the 2 locations, is writing/reading TFRecords to/from 1 of them faster? :  </p>\n\n<ol><li><code>../input/</code>  </li>\n<li><code>../working/</code> (UI shows it as output)</li></ol></li>\n<li><p>Why is my TFrecord writing speed much slower on kaggle compared to local laptop? and Why does the write time not scale linearly with shard size? (4-6x time increase when 2x shard size from 3k to 6k)\n```\nExtra Tests:\n<strong>1 shard writing speed on my local laptop 6 core 2.6GHz (off kaggle)</strong>:\n64 shards of 1647 images: 17 secs\n48 shards of 2196 images: 29 secs\n32 shards of 3294 images: 1 min\n16 shards of 6587 images: 6 min</p></li>\n</ol>\n\n<p><strong>1 shard writing speed on kaggle</strong>:\n64 shards of 1647 images: 2 min\n48 shards of 2196 images: 3 min 44\n32 shards of 3294 images: 8 min 28\n16 shards of 6587 images: 35 min\n```\nGuessing this experiment has something to do with my 12 logical cores CPU (operating at 4ghz when writing) vs 4 CPU(s) on kaggle? Not sure whether or which part of the input pipeline and TFRecord writing uses multithreading. I'm basically following the Flowers tutorial <a href=\"https://colab.research.google.com/github/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/03_Flower_pictures_to_TFRecords.ipynb\">https://colab.research.google.com/github/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/03_Flower_pictures_to_TFRecords.ipynb</a> </p>\n\n<ol>\n<li>Why does TPU training take 2000 seconds on 1st epoch, while &lt;60s on next few epochs? (Not using TFRecords here). Is the extra 1st epoch time spent on copying data from kaggle local to the kaggle GCS Bucket associated with KaggleDatasets()? In Colab, we can copy files from google drive to the temporary directory right where colab is to speed up IO significantly, any analogous ways in TPU kernels to decrease the 1st epoch time?</li>\n</ol>",
          "rawMarkdown": "@mgornergoogle  \nCould you help me with some questions on effective TFRecord read/write and TPU usage, i can't piece together the numerous puzzle pieces even after going through this whole discussion.\n\n1. On the number of shards, i always see 16 shards (100-200mb each) suggested in tutorials kernels, while on Kaggle youtube last week, i hear \"If you have 100k images, use 16 shards\". Confusing me is my 100k images are 10GB, meaning 16 shards would be 600mb each, way more than 100-200mb. So is the sharding concept about number of shards, or file size per shard?\n\n2. What's the difference between the 2 locations, is writing/reading TFRecords to/from 1 of them faster? :  \n      1. `../input/`  \n      2. `../working/` (UI shows it as output)\n\n3. Why is my TFrecord writing speed much slower on kaggle compared to local laptop? and Why does the write time not scale linearly with shard size? (4-6x time increase when 2x shard size from 3k to 6k)\n```\nExtra Tests:\n**1 shard writing speed on my local laptop 6 core 2.6GHz (off kaggle)**:\n64 shards of 1647 images: 17 secs\n48 shards of 2196 images: 29 secs\n32 shards of 3294 images: 1 min\n16 shards of 6587 images: 6 min\n\n**1 shard writing speed on kaggle**:\n64 shards of 1647 images: 2 min\n48 shards of 2196 images: 3 min 44\n32 shards of 3294 images: 8 min 28\n16 shards of 6587 images: 35 min\n```\nGuessing this experiment has something to do with my 12 logical cores CPU (operating at 4ghz when writing) vs 4 CPU(s) on kaggle? Not sure whether or which part of the input pipeline and TFRecord writing uses multithreading. I'm basically following the Flowers tutorial https://colab.research.google.com/github/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/03_Flower_pictures_to_TFRecords.ipynb \n\n4. Why does TPU training take 2000 seconds on 1st epoch, while &lt;60s on next few epochs? (Not using TFRecords here). Is the extra 1st epoch time spent on copying data from kaggle local to the kaggle GCS Bucket associated with KaggleDatasets()? In Colab, we can copy files from google drive to the temporary directory right where colab is to speed up IO significantly, any analogous ways in TPU kernels to decrease the 1st epoch time?"
        },
        {
          "id": 929662,
          "postDate": "2020-07-14T20:34:36.540Z",
          "content": "<p>hey <a href=\"/datahan\">@datahan</a>! I hope these answers are helpful:</p>\n\n<ol>\n<li>the guidelines provided are meant to cover general use cases, and depending on the size of your data you might need to adjust accordingly. in your use case, 16 shards may not be enough, but in speaking with the team here using more shards is unlikely to have a noticeable difference.</li>\n<li>TFRecords should be stored in a GCS bucket - TPUs can't read//write from either of the locations listed.</li>\n<li>this is likely due to a difference between your computer's specs and the Kaggle VM.</li>\n<li>the first epoch is spent performing XLA calculations, compiling the model, and performing memory optimizations. using TFRecords can help reduce the time of the first epoch.</li>\n</ol>",
          "rawMarkdown": "hey @datahan! I hope these answers are helpful:\n\n1. the guidelines provided are meant to cover general use cases, and depending on the size of your data you might need to adjust accordingly. in your use case, 16 shards may not be enough, but in speaking with the team here using more shards is unlikely to have a noticeable difference.\n2. TFRecords should be stored in a GCS bucket - TPUs can't read//write from either of the locations listed.\n3. this is likely due to a difference between your computer's specs and the Kaggle VM.\n4. the first epoch is spent performing XLA calculations, compiling the model, and performing memory optimizations. using TFRecords can help reduce the time of the first epoch."
        }
      ]
    },
    {
      "id": 747098,
      "postDate": "2020-02-16T00:34:19.543Z",
      "content": "<p>Did anyone actually sucessfully used TPU for training?</p>",
      "rawMarkdown": "Did anyone actually sucessfully used TPU for training?",
      "replies": [
        {
          "id": 749616,
          "postDate": "2020-02-18T20:12:02.003Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 755242,
          "postDate": "2020-02-24T15:33:04.993Z",
          "content": "<p>Yes it was very fast. It is also a different beast than pytorch training. A bit part of the difference comes from a massive batch size of 512 or 1024. </p>",
          "rawMarkdown": "Yes it was very fast. It is also a different beast than pytorch training. A bit part of the difference comes from a massive batch size of 512 or 1024. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 741951,
      "postDate": "2020-02-11T02:35:43.193Z",
      "content": "<p>What is the speed difference of GPU and TPU?</p>",
      "rawMarkdown": "What is the speed difference of GPU and TPU?",
      "replies": [
        {
          "id": 741957,
          "postDate": "2020-02-11T02:40:37.720Z",
          "content": "<p>Try for yourself using the TPU notebooks in the TPU docs: <a href=\"https://www.kaggle.com/mgornergoogle/five-flowers-with-keras-and-xception-on-tpu\">Five flowers with Keras and Xception on TPU</a></p>",
          "rawMarkdown": "Try for yourself using the TPU notebooks in the TPU docs: [Five flowers with Keras and Xception on TPU](https://www.kaggle.com/mgornergoogle/five-flowers-with-keras-and-xception-on-tpu)",
          "votes": 1
        },
        {
          "id": 741986,
          "postDate": "2020-02-11T02:56:29.367Z",
          "content": "<p>ok. Thanks.</p>",
          "rawMarkdown": "ok. Thanks."
        },
        {
          "id": 742094,
          "postDate": "2020-02-11T04:30:30.917Z",
          "content": "<p>Did anyone encountered this error?\n&gt; ValueError: The two structures don't have the same nested structure.\nValueError: Could not pack sequence. Structure had 32 elements, but flat_sequence had 1 elements.  Structure: [1, 0, 0, 1, 1, 0, 1, 1, 0, 0, 1, 0, 1, 1, 0, 1, 1, 0, 1, 1, 0, 0, 0, 1, 0, 0, 1, 0, 0, 0, 0, 1], flat_sequence: [None].</p>",
          "rawMarkdown": "Did anyone encountered this error?\n&gt; ValueError: The two structures don't have the same nested structure.\nValueError: Could not pack sequence. Structure had 32 elements, but flat_sequence had 1 elements.  Structure: [1, 0, 0, 1, 1, 0, 1, 1, 0, 0, 1, 0, 1, 1, 0, 1, 1, 0, 1, 1, 0, 0, 0, 1, 0, 0, 1, 0, 0, 0, 0, 1], flat_sequence: [None]."
        },
        {
          "id": 742106,
          "postDate": "2020-02-11T04:40:02.350Z",
          "content": "<p>paste bin:\n<a href=\"https://pastebin.com/KMuKmbfu\">link</a></p>",
          "rawMarkdown": "paste bin:\n[link](https://pastebin.com/KMuKmbfu)"
        },
        {
          "id": 742995,
          "postDate": "2020-02-11T17:29:34.427Z",
          "content": "<p>It looks like you are trying model.fit without using the tf.data.Dataset API. This is what I see in the stack trace:</p>\n\n<p><code>\nfit(self, model, x, y, batch_size, epochs, verbose, callbacks, validation_split, validation_data, shuffle, class_weight, sample_weight, initial_epoch, steps_per_epoch, validation_steps, validation_freq, max_queue_size, workers, use_multiprocessing, **kwargs)\n    233           max_queue_size=max_queue_size,\n    234           workers=workers,\n--&gt; 235           use_multiprocessing=use_multiprocessing)\n</code>\nI would need more code to understand exactly what you are trying to do and whether your are running this on TPU or GPU but in general, tf.data.Dataset is the recommended data pipeline API in Tensorflow and especially when using TPUs.</p>\n\n<p>That said, training on in-memory numpy arrays is also supported on TPUs. It is not going to be fast but for data that fits into memory anyway, this might not matter.</p>\n\n<p>If you were trying to use a generator for data input, I'm pretty sure that will not work. Let me double-check.</p>\n\n<p>Really, the recommended API in Tensorflow for out of memory datasets is tf.data.Dataset</p>",
          "rawMarkdown": "It looks like you are trying model.fit without using the tf.data.Dataset API. This is what I see in the stack trace:\n\n```\nfit(self, model, x, y, batch_size, epochs, verbose, callbacks, validation_split, validation_data, shuffle, class_weight, sample_weight, initial_epoch, steps_per_epoch, validation_steps, validation_freq, max_queue_size, workers, use_multiprocessing, **kwargs)\n    233           max_queue_size=max_queue_size,\n    234           workers=workers,\n--&gt; 235           use_multiprocessing=use_multiprocessing)\n```\nI would need more code to understand exactly what you are trying to do and whether your are running this on TPU or GPU but in general, tf.data.Dataset is the recommended data pipeline API in Tensorflow and especially when using TPUs.\n\nThat said, training on in-memory numpy arrays is also supported on TPUs. It is not going to be fast but for data that fits into memory anyway, this might not matter.\n\nIf you were trying to use a generator for data input, I'm pretty sure that will not work. Let me double-check.\n\nReally, the recommended API in Tensorflow for out of memory datasets is tf.data.Dataset"
        },
        {
          "id": 743250,
          "postDate": "2020-02-11T23:34:55.933Z",
          "content": "<p>[this comment is deleted]</p>",
          "rawMarkdown": "[this comment is deleted]"
        },
        {
          "id": 743271,
          "postDate": "2020-02-12T00:31:13.667Z",
          "content": "<p>Now I'm getting UnavailableError.</p>",
          "rawMarkdown": "Now I'm getting UnavailableError."
        },
        {
          "id": 756880,
          "postDate": "2020-02-26T06:58:30.660Z",
          "content": "<p>Using TF Dataset now, still getting error in model.fit, though the error msg changed : \"RuntimeError: No registered 'Identity' OpKernel for 'TPU' devices compatible with node {{node Identity}}\n     (OpKernel was found, but attributes didn't match) Requested Attributes: T=DT_UINT8\n    .  Registered:  device='XLA_GPU_JIT'; T in [DT_FLOAT, DT_DOUBLE, DT_INT32, DT_UINT8, DT_INT16, ..., DT_HALF, DT_UINT32, DT_UINT64, DT_RESOURCE, DT_VARIANT]\n  device='XLA_CPU_JIT'; T in [DT_FLOAT, DT_DOUBLE, DT_INT32, DT_UINT8, DT_INT16, ..., DT_HALF, DT_UINT32, DT_UINT64, DT_RESOURCE, DT_VARIANT]\n  device='XLA_TPU_JIT'; T in [DT_FLOAT, DT_DOUBLE, DT_INT32, DT_COMPLEX64, DT_INT64, DT_BOOL, DT_BFLOAT16, DT_UINT32, DT_UINT64, DT_RESOURCE, DT_VARIANT]\n  device='XLA_CPU'; T in [DT_UINT8, DT_QUINT8, DT_UINT16, DT_INT8, DT_QINT8, ..., DT_DOUBLE, DT_COMPLEX64, DT_COMPLEX128, DT_BOOL, DT_BFLOAT16]\n  device='TPU'; T in [DT_INT32, DT_UINT32, DT_BFLOAT16, DT_FLOAT, DT_DOUBLE, DT_BOOL, DT_COMPLEX64, DT_INT64, DT_UINT64]\n  device='TPU_SYSTEM'\n  device='GPU'; T in [DT_HALF]\n  device='GPU'; T in [DT_BFLOAT16]\n  device='GPU'; T in [DT_FLOAT]\n  device='GPU'; T in [DT_DOUBLE]\n  device='GPU'; T in [DT_INT64]\n  device='GPU'; T in [DT_UINT16]\n  device='GPU'; T in [DT_INT16]\n  device='GPU'; T in [DT_UINT8]\n  device='GPU'; T in [DT_INT8]\n  device='GPU'; T in [DT_COMPLEX64]\n  device='GPU'; T in [DT_COMPLEX128]\n  device='GPU'; T in [DT_VARIANT]\n  device='DEFAULT'; T in [DT_STRING]\n  device='DEFAULT'; T in [DT_VARIANT]\n  device='DEFAULT'; T in [DT_RESOURCE]\n  device='CPU'</p>\n\n<pre><code> [[Identity]] \"\n</code></pre>\n\n<p>Any suggestions ? My code is like this : \n    resized_image=np.array()\n    datagen=tf.keras.preprocessing.image.ImageDataGenerator (....)\n    datagen.fit(resized_image)\n    datagen = datagen.flow(resized_image, trainGraphemeY, batch_size=BS)</p>\n\n<pre><code>ds = tf.data.Dataset.from_generator(\n    lambda:datagen,\noutput_types=(tf.uint8, tf.uint8),\noutput_shapes=(resized_image.shape, trainGraphemeY.shape)\n)  \n\nhistory = model_root.fit_generator(ds,\n                                  epochs = EPOCHS, \n                                  steps_per_epoch=resized_image.shape[0] // BS, \n                                  callbacks=[es],verbose=2) \"\n</code></pre>",
          "rawMarkdown": "Using TF Dataset now, still getting error in model.fit, though the error msg changed : \"RuntimeError: No registered 'Identity' OpKernel for 'TPU' devices compatible with node {{node Identity}}\n\t (OpKernel was found, but attributes didn't match) Requested Attributes: T=DT_UINT8\n\t.  Registered:  device='XLA_GPU_JIT'; T in [DT_FLOAT, DT_DOUBLE, DT_INT32, DT_UINT8, DT_INT16, ..., DT_HALF, DT_UINT32, DT_UINT64, DT_RESOURCE, DT_VARIANT]\n  device='XLA_CPU_JIT'; T in [DT_FLOAT, DT_DOUBLE, DT_INT32, DT_UINT8, DT_INT16, ..., DT_HALF, DT_UINT32, DT_UINT64, DT_RESOURCE, DT_VARIANT]\n  device='XLA_TPU_JIT'; T in [DT_FLOAT, DT_DOUBLE, DT_INT32, DT_COMPLEX64, DT_INT64, DT_BOOL, DT_BFLOAT16, DT_UINT32, DT_UINT64, DT_RESOURCE, DT_VARIANT]\n  device='XLA_CPU'; T in [DT_UINT8, DT_QUINT8, DT_UINT16, DT_INT8, DT_QINT8, ..., DT_DOUBLE, DT_COMPLEX64, DT_COMPLEX128, DT_BOOL, DT_BFLOAT16]\n  device='TPU'; T in [DT_INT32, DT_UINT32, DT_BFLOAT16, DT_FLOAT, DT_DOUBLE, DT_BOOL, DT_COMPLEX64, DT_INT64, DT_UINT64]\n  device='TPU_SYSTEM'\n  device='GPU'; T in [DT_HALF]\n  device='GPU'; T in [DT_BFLOAT16]\n  device='GPU'; T in [DT_FLOAT]\n  device='GPU'; T in [DT_DOUBLE]\n  device='GPU'; T in [DT_INT64]\n  device='GPU'; T in [DT_UINT16]\n  device='GPU'; T in [DT_INT16]\n  device='GPU'; T in [DT_UINT8]\n  device='GPU'; T in [DT_INT8]\n  device='GPU'; T in [DT_COMPLEX64]\n  device='GPU'; T in [DT_COMPLEX128]\n  device='GPU'; T in [DT_VARIANT]\n  device='DEFAULT'; T in [DT_STRING]\n  device='DEFAULT'; T in [DT_VARIANT]\n  device='DEFAULT'; T in [DT_RESOURCE]\n  device='CPU'\n\n\t [[Identity]] \"\nAny suggestions ? My code is like this : \n    resized_image=np.array()\n    datagen=tf.keras.preprocessing.image.ImageDataGenerator (....)\n    datagen.fit(resized_image)\n    datagen = datagen.flow(resized_image, trainGraphemeY, batch_size=BS)\n\n    ds = tf.data.Dataset.from_generator(\n        lambda:datagen,\n    output_types=(tf.uint8, tf.uint8),\n    output_shapes=(resized_image.shape, trainGraphemeY.shape)\n    )  \n\n    history = model_root.fit_generator(ds,\n                                      epochs = EPOCHS, \n                                      steps_per_epoch=resized_image.shape[0] // BS, \n                                      callbacks=[es],verbose=2) \""
        },
        {
          "id": 757459,
          "postDate": "2020-02-26T19:35:00.817Z",
          "content": "<p>Yes, <code>tf.keras.preprocessing.image.ImageDataGenerator</code> does not work on TPU (yet).\n<a href=\"/cdeotte\">@cdeotte</a> posted TPU-compatible image transformatoin routines <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132191\">here</a>.</p>",
          "rawMarkdown": "Yes, `tf.keras.preprocessing.image.ImageDataGenerator` does not work on TPU (yet).\n@cdeotte posted TPU-compatible image transformatoin routines [here](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132191)."
        }
      ]
    },
    {
      "id": 741801,
      "postDate": "2020-02-11T00:31:07.483Z",
      "content": "<p>By the way, thank you for the Notebook with TPU explanations! I'm reading it now.</p>",
      "rawMarkdown": "By the way, thank you for the Notebook with TPU explanations! I'm reading it now."
    },
    {
      "id": 741726,
      "postDate": "2020-02-10T22:58:16.197Z",
      "content": "<p>If you are looking for more info, the TPU docs are here: <a href=\"https://www.kaggle.com/docs/tpu\">https://www.kaggle.com/docs/tpu</a></p>",
      "rawMarkdown": "If you are looking for more info, the TPU docs are here: https://www.kaggle.com/docs/tpu\n"
    },
    {
      "id": 741721,
      "postDate": "2020-02-10T22:55:26.210Z",
      "rawMarkdown": ""
    },
    {
      "id": 741700,
      "postDate": "2020-02-10T22:25:06.423Z",
      "content": "<p>according to this, 30h a week, 3h limit per session\n<a href=\"https://www.kaggle.com/docs/tpu\">https://www.kaggle.com/docs/tpu</a></p>\n\n<p>I would think they would be slightly more generous considering the amount of time tpu setup takes.</p>",
      "rawMarkdown": "according to this, 30h a week, 3h limit per session\nhttps://www.kaggle.com/docs/tpu\n\nI would think they would be slightly more generous considering the amount of time tpu setup takes.",
      "replies": [
        {
          "id": 741706,
          "postDate": "2020-02-10T22:31:33.950Z",
          "content": "<p>tbh 3 hours per session is probably enough for this dataset. </p>\n\n<p>It takes me 30 seconds to process an epoch of 32768 images. </p>\n\n<p>I train for around 40 epochs and that takes less than 30 minutes. At that point, if you don't decay the LR you will definitely over fit. </p>",
          "rawMarkdown": "tbh 3 hours per session is probably enough for this dataset. \n\nIt takes me 30 seconds to process an epoch of 32768 images. \n\nI train for around 40 epochs and that takes less than 30 minutes. At that point, if you don't decay the LR you will definitely over fit. ",
          "votes": 2
        },
        {
          "id": 741781,
          "postDate": "2020-02-11T00:16:03.690Z",
          "content": "<p>What do you mean by TPU setup ? In the <a href=\"https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu\">Getting started with 100+</a> flowers notebook, the first epoch takes 40s instead of 15s for all other epochs. So yes, there is a little bit of overhead (25s) there, but it does not seem to be too bad, does it ?</p>\n\n<p>By the way, the extra time is where your model gets compiled through XLA to TPU bytecode. Memory layout optimizations are performed there as well.</p>",
          "rawMarkdown": "What do you mean by TPU setup ? In the [Getting started with 100+](https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu) flowers notebook, the first epoch takes 40s instead of 15s for all other epochs. So yes, there is a little bit of overhead (25s) there, but it does not seem to be too bad, does it ?\n\nBy the way, the extra time is where your model gets compiled through XLA to TPU bytecode. Memory layout optimizations are performed there as well."
        }
      ]
    },
    {
      "id": 741688,
      "postDate": "2020-02-10T22:12:39.970Z",
      "content": "<p>I have found more information: Tensor Processing Units (TPUs) are Now Available on Kaggle - <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus\">https://www.kaggle.com/c/flower-classification-with-tpus</a></p>",
      "rawMarkdown": "I have found more information: Tensor Processing Units (TPUs) are Now Available on Kaggle - https://www.kaggle.com/c/flower-classification-with-tpus",
      "replies": [
        {
          "id": 741701,
          "postDate": "2020-02-10T22:25:12.483Z",
          "content": "<p>Yes, we just soft-launched this today. Keep in mind that submissions to the Deepfake Detection Challenge with notebooks using the TPU integration will not be permitted. However, those using TPUs, whether through the integration or offline training (i.e. through Cloud TPUs) can still make submissions through Kaggle notebooks, with GPU turned on.</p>",
          "rawMarkdown": "Yes, we just soft-launched this today. Keep in mind that submissions to the Deepfake Detection Challenge with notebooks using the TPU integration will not be permitted. However, those using TPUs, whether through the integration or offline training (i.e. through Cloud TPUs) can still make submissions through Kaggle notebooks, with GPU turned on.",
          "votes": 8
        },
        {
          "id": 742741,
          "postDate": "2020-02-11T14:17:10.067Z",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> Well, that's a shame. I was really looking forward to some <strong>practical TPU usage other than identifying flowers</strong>.</p>",
          "rawMarkdown": "@juliaelliott Well, that's a shame. I was really looking forward to some **practical TPU usage other than identifying flowers**."
        },
        {
          "id": 742834,
          "postDate": "2020-02-11T15:12:48.843Z",
          "content": "<p>Everyone is very welcome to try this problem using TPUs! There is certainly no harm in giving it a spin. Allow me to clarify — I wanted to make sure there was no misconception about being able to submit from a notebook to <em>this competition (Deepfake)</em> with TPU turned on. You have the option to use Cloud TPUs directly or train on Kaggle using TPUs, but then will have to upload your trained model as an external data source into your submission notebook to perform inference (with GPU or CPU on). We’re definitely interested to see what you are able to do on “real problems,” hence enabling it site-wide instead. Best of luck!</p>",
          "rawMarkdown": "Everyone is very welcome to try this problem using TPUs! There is certainly no harm in giving it a spin. Allow me to clarify — I wanted to make sure there was no misconception about being able to submit from a notebook to *this competition (Deepfake)* with TPU turned on. You have the option to use Cloud TPUs directly or train on Kaggle using TPUs, but then will have to upload your trained model as an external data source into your submission notebook to perform inference (with GPU or CPU on). We’re definitely interested to see what you are able to do on “real problems,” hence enabling it site-wide instead. Best of luck!",
          "votes": 2
        },
        {
          "id": 743032,
          "postDate": "2020-02-11T18:20:52.910Z",
          "content": "<p>And the main reason for not allowing TPUs in this specific competition is that the competition was launched before TPUs were available on Kaggle. Allowing them now would change the playing field which might be seen as unfair.</p>",
          "rawMarkdown": "And the main reason for not allowing TPUs in this specific competition is that the competition was launched before TPUs were available on Kaggle. Allowing them now would change the playing field which might be seen as unfair.",
          "votes": 3
        },
        {
          "id": 756877,
          "postDate": "2020-02-26T06:55:23.913Z",
          "content": "<p>Is it av in Bengali.AI competition ? </p>",
          "rawMarkdown": "Is it av in Bengali.AI competition ? "
        }
      ]
    },
    {
      "id": 746116,
      "postDate": "2020-02-14T16:39:12.377Z",
      "content": "<p>Thank you! This answers my question!</p>",
      "rawMarkdown": "Thank you! This answers my question!"
    }
  ],
  "comments": [
    {
      "id": 741696,
      "author_name": "hongy",
      "author_url": "",
      "post_date": "2020-02-10T22:22:58.280000",
      "content": "<p>if any google overlords are following this, what zone are these tpus in? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 741709,
          "author_name": "Dustin",
          "author_url": "",
          "post_date": "2020-02-10T22:42:23.097000",
          "content": "<p>Any particular reason you need to know? We don't really make any guarantees about this since we use resources from many different zones.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 741717,
          "author_name": "hongy",
          "author_url": "",
          "post_date": "2020-02-10T22:51:37.017000",
          "content": "<p>Bucket access costs. All my tfrecords are in central1 so if these tpus are coming from europe I'm not sure how that would affect the bucket fees</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 741722,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-02-10T22:56:25.750000",
          "content": "<p>The TPUs you get can be in various zones but we do take care of data locality. As you can see in the <a href=\"https://www.kaggle.com/mgornergoogle/five-flowers-with-keras-and-xception-on-tpu\">TPU sample</a>, kaggle will copy the dataset to GCS (Google Cloud Storage) when you ask for GCS access to it through <code>KaggleDatasets().get_gcs_path()</code>. The data will be copied to a cache bucket in the same location as the TPU you have been allocated.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 741723,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-02-10T22:57:32.040000",
          "content": "<p>Bucket access cost is covered by the Kaggle integration as long as you use pblic Kaggle datasets. We have not integrated private Kaggle datasets with TPUs yet. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 741724,
          "author_name": "hongy",
          "author_url": "",
          "post_date": "2020-02-10T22:57:33.283000",
          "content": "<p>Nice. Can I take it that if I'm using the tpu to train from a private bucket in central1, it won't spike my bucket access costs or it will be similar to if I had set up a v3-8 in central1 from a cost perspective? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 741746,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-02-10T23:25:29.320000",
          "content": "<p>Not quite, if you use a TPU to train on a <strong>public Kaggle dataset</strong>, the GCS costs (because data will be coming from GCS in that case as well) will be covered by Kaggle.\nTraining on your own private bucket is not yet officially supported. I expect it will work if you know how to set IAMs correctly. If you manage to make it work, you are still paying for usage of your own bucket.</p>\n\n<p>Also, if you do make private buckets work, tell us how you did it exactly. It is not a feature we focused on at launch but we still want to hear about the user experience, especially if it could be improved.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 741782,
          "author_name": "hongy",
          "author_url": "",
          "post_date": "2020-02-11T00:16:49.517000",
          "content": "<p>lulz so I fed the filenames to the tfrecord using my bucket. So something like gs://mybucket/tfrecord. </p>\n\n<p>This created an error message when trying to fit the model. Below is a copy paste of the relevant part of the error. I then went to IAM and then just granted rights to the TPU service account. Needed to reinit the TPU, but otherwise it got access. </p>\n\n<p>PermissionDeniedError: Error executing an HTTP request: HTTP response code 403 with body '{\n\"error\": {\n\"code\": 403,\n\"message\": \"service-467472385656@cloud-tpu.iam.gserviceaccount.com does not have storage.objects.list access to deepfakestorage.\",\n\"errors\": [\n{\n\"message\": \"service-467472385656@cloud-tpu.iam.gserviceaccount.com does not have storage.objects.list access to deepfakestorage.\",\n\"domain\": \"global\",\n\"reason\": \"forbidden\"\n}\n]\n}\n}</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 741784,
          "author_name": "hongy",
          "author_url": "",
          "post_date": "2020-02-11T00:19:17.687000",
          "content": "<p>Also I made a kernel explaining TPUs and tfrecord. <a href=\"https://www.kaggle.com/hooong/how-to-tpu?scriptVersionId=28471599\">https://www.kaggle.com/hooong/how-to-tpu?scriptVersionId=28471599</a></p>\n\n<p>It pretty much desribes how I use TPUs in google cloud and kaggle </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 741796,
          "author_name": "hongy",
          "author_url": "",
          "post_date": "2020-02-11T00:29:14.423000",
          "content": "<p>So I had a plan to cycle through TPU sessions until I get one that's in zone us central1. </p>\n\n<p>Unfortuantely, kaggle TPUs do not have the normal TPU attributes. </p>\n\n<p>```\ntpu = tf.distribute.cluster_resolver.TPUClusterResolver()\ntpu.<strong>dict</strong></p>\n\n<p>```\nreturn \n{'_tpu': b'10.0.0.2:8470',\n 'task_type': 'worker',\n 'task_id': 0,\n '_environment': '',\n 'rpc_layer': 'grpc',\n '_should_resolve_override': False,\n '_service': None,\n '_credentials': 'default',\n '_project': None,\n '_zone': None,\n '_discovery_url': None,\n '_coordinator_name': None,\n '_coordinator_address': None} </p>\n\n<p>When I run the same thing in google cloud I get </p>\n\n<p>{'_coordinator_address': None,\n '_coordinator_name': None,\n '_credentials': None,\n '_discovery_url': None,\n '_environment': '',\n '_project': 'tpu-44747',\n '_service': None,\n '_should_resolve_override': None,\n '_tpu': b'node-1',\n '_zone': 'us-central1-b',\n 'rpc_layer': 'grpc',\n 'task_id': 0,\n 'task_type': 'worker'}</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 741798,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-02-11T00:30:23.997000",
          "content": "<p>what were you looking for in there ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 741809,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-02-11T00:41:29.360000",
          "content": "<p>Thank you for the notebook with TPU explanation. I went through it and added some comments.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 741892,
          "author_name": "hongy",
          "author_url": "",
          "post_date": "2020-02-11T01:49:06.367000",
          "content": "<p>I was trying to check the zone of the tpu hoping that the tpu would show up as zone 'us-central1-b or similar. I'm concerned about egress bc it isn't cheap ($0.01 per gb for not free cloud services, bunch of conditions to determine if free but generally free if same region). I'm actually not sure how kaggle tpu data access will be charged and I don't know if it will count as a google cloud service accessing my bucket. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 741903,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-02-11T02:02:06.543000",
          "content": "<p>By the way, if you care about the data in your GCS bucket, DO NOT give the Kaggle TPU service account access. You do not control this TPU and you do not know who can use it to access your data.\nIf you do not care much about the data in your GCS bucket, make it public. You won't have to worry about IAMs anymore.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 741909,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-02-11T02:04:10.173000",
          "content": "<p>And yes I'm pretty sure you will be charged for egress from your GCS bucket. Even if it is a Google cloud service, those TPUs are owned by Kaggle, not your by your GCP project.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 745394,
          "author_name": "hongy",
          "author_url": "",
          "post_date": "2020-02-13T19:18:43.443000",
          "content": "<p>so I ran a test and I didn't get any egress fees using a multi region US bucket. But even if I did get egress fees, it is quite affordable at less than a penny per GB.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 741695,
      "author_name": "Debanga Raj Neog",
      "author_url": "",
      "post_date": "2020-02-10T22:22:40.787000",
      "content": "<p>WOW! 30hr of free TPU on Kaggle Kernels!!! </p>",
      "votes": 1,
      "replies": [
        {
          "id": 741699,
          "author_name": "hongy",
          "author_url": "",
          "post_date": "2020-02-10T22:25:00.733000",
          "content": "<p>its a v3-8 </p>\n\n<p>so thats the newer tpu architecture and you get 8 tpu units. </p>\n\n<p>For reference, I've been using v2-8 and it computes a batch of 512, 256, 256, 3 in 300ms on resnet50. That's 512 images at 256 by 256. </p>\n\n<p>It is INSANE how much faster it is than gpus. IMO, v3 is mostly improvements for RNNs, but still good for CNNs. </p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 741728,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-02-10T23:00:42.960000",
          "content": "<p>V3s have roughly double the matrix multiplication hardware of V2s. Utilization of the additional hardware will depend on your model though.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 741729,
          "author_name": "Debanga Raj Neog",
          "author_url": "",
          "post_date": "2020-02-10T23:02:22.737000",
          "content": "<p>That's really great! I am using SageMaker so far with the lowest spec GPU... After your comment, I am now definitely gonna try the 60days v3-8 TPU credit I got (and I was lazy to setup)! + Kaggle 30hrs is there too... ^-^</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 741738,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-02-10T23:13:22.840000",
          "content": "<p>If you want to use TPUs on GCP with <a href=\"https://console.cloud.google.com/ai-platform/notebooks\">Cloud AI platform Notebooks</a>, they work, but the \"TPU accelerator\" option is not yet in the UI. You have to create the Notebook VM + TPU pair through the command line. I have a script here: <a href=\"https://github.com/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/README-TF2.1.md\">bit.ly/keras-tpu-tf21</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 741745,
          "author_name": "Debanga Raj Neog",
          "author_url": "",
          "post_date": "2020-02-10T23:25:00.753000",
          "content": "<p>Thanks <a href=\"/mgornergoogle\">@mgornergoogle</a> !</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 741755,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-02-10T23:36:28.497000",
          "content": "<p>The other significant difference between TPU v2 and TPU v3 is 64GB of HBM vs 128GB for a TPU v3. (HBM = High Bandwidth Memory)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 741774,
          "author_name": "Debanga Raj Neog",
          "author_url": "",
          "post_date": "2020-02-11T00:05:28.447000",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> I am trying to create a new notebook instance in AI Platform, but I cannot create a custom notebook in the <code>europe-west4-a</code> zone from UI. Is this zone not supported or I have to set it up in the cloud shell? I am mentioning this because my TPU credits are only for this zone. Thanks in advance.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 741794,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-02-11T00:28:47.510000",
          "content": "<p>To use TPUs in GCP's Cloud AI Notebooks, use this script: <a href=\"https://raw.githubusercontent.com/GoogleCloudPlatform/training-data-analyst/master/courses/fast-and-lean-data-science/create-tpu-deep-learning-vm.sh\">create-tpu-deep-learning-vm.sh</a></p>\n\n<p><code>gcloud init</code>\nthen\n<code>./create-tpu-deep-learning-vm.sh choose-a-name --tpu-type v3-8</code></p>\n\n<p>If you open the script, there is no magic, just the two gcloud commands that create the VM and the TPU. There is a tiny bit of magic to set up the TPU_NAME environment variable on the VM. That's what makes the TPUCusterResolver() call work without parameters in your Python code. The second tiny bit of magic is the setting that makes the VM work as a Notebook VM with a Jupyter proxy.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 741800,
          "author_name": "Debanga Raj Neog",
          "author_url": "",
          "post_date": "2020-02-11T00:30:51.093000",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> great, will do it this way. Also, I will add <code>--zone=europe-west4-a</code></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 756839,
          "author_name": "AnirbanK",
          "author_url": "",
          "post_date": "2020-02-26T06:12:01.950000",
          "content": "<p>Sir, can you share your code to define and train the model using TPUs. Getting some issues here!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 742992,
      "author_name": "hongy",
      "author_url": "",
      "post_date": "2020-02-11T17:27:26.023000",
      "content": "<p>An update regarding kaggle tpu zones. </p>\n\n<p>It appears that kaggle tpus are most likely located somewhere in the united states. Yesterday, I ran a kaggle tpu that pulled data from my private us-central1 storage bucket. This incurred an egress fee. </p>\n\n<p>Cloud Storage Inter-region GCP Storage egress within NA: 1.73 </p>\n\n<p>I will attempt to clone my bucket to a multi region NA bucket and then test again to see if egress fees are charged. This will take a day bc transactions don't usually take a while to appear on the billing menu. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 742998,
          "author_name": "Dustin",
          "author_url": "",
          "post_date": "2020-02-11T17:33:32.230000",
          "content": "<p>Our TPUs are not limited to the US. The only way to make sure your data is co-located to the TPU right now is to use Kaggle public datasets + GCS loader (see <a href=\"https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu#Competition-data-access\">https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu#Competition-data-access</a>). </p>\n\n<p>Private datasets are not officially supported due to technical/privacy limitations which we are working on.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 755892,
          "author_name": "AnirbanK",
          "author_url": "",
          "post_date": "2020-02-25T08:30:38.503000",
          "content": "<p>Hello, I cannot get a hold of the dataset for the BengaliAI competition in GCP Bucket, Code I am using is \nGCS_DS_PATH = KaggleDatasets().get_gcs_path('/kaggle/input/bengaliai-cv19'). Getting error \"Dataset not found for directory '/kaggle/input/bengaliai-cv19',\". Can you help please. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 756487,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2020-02-25T18:57:08.400000",
          "content": "<p><a href=\"/anirbank\">@anirbank</a> You should only need to assign the path with\n<code>GCS_DS_PATH = KaggleDatasets().get_gcs_path()</code></p>\n\n<p>And then if you want to list the bucket, you can use:\n<code>!gsutil ls $GCS_DS_PATH</code></p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 756633,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-02-25T22:55:35.827000",
          "content": "<p>Or if you have multiple datasets in your Notebook:\n<code>\nGCS_DS_PATH = KaggleDatasets().get_gcs_path('bengaliai-cv19')\n</code>\nYou got it almost right but the function wants the name of the directory where the dataset has been mapped, not the full path.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 756813,
          "author_name": "AnirbanK",
          "author_url": "",
          "post_date": "2020-02-26T05:27:15.840000",
          "content": "<p>Excellent, it works fine now, thank you. Lets try TPU on my notebook now!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 756821,
          "author_name": "AnirbanK",
          "author_url": "",
          "post_date": "2020-02-26T05:48:07.137000",
          "content": "<p>Although it worked for csv file, getting this error when reading feather file : \"failed to open local file 'gs://kds-87e1f7817c6764d20c2f2841fd9048ac1f7b9c89a1508dbd796f13b4/train_image_data_1.feather'. Detail: [errno 2] No such file or directory.\" This file is however listed when i do !gsutil ls $GCS_DS_PATH1, with GCS_DS_PATH1 = KaggleDatasets().get_gcs_path('bengaliaicv19feather') .  The same works fine, file is read for parquet files though. Any ideas here ? thanks</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 756830,
          "author_name": "AnirbanK",
          "author_url": "",
          "post_date": "2020-02-26T05:57:38.477000",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> Can read from parquet files (using GCP Bucket path), but now getting this error on model.fit :\" NotFoundError: {{function_node __inference_distributed_function_71636}} No registered 'PyFunc' OpKernel for 'CPU' devices compatible with node {{node PyFunc}}\n    .  Registered:  \" I have defined TPUStrategy and this is done fine, as given in the sample here. The TPU is visible in the output. I have defined the Model using strategy.scope() as shown. So not sure what the issue is. However, I am not using TFRecords, my input data is from numpy arrays, derived from the image files. I guess that should not be a issue.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 911453,
          "author_name": "Mohamed Taha",
          "author_url": "",
          "post_date": "2020-07-01T18:33:22.730000",
          "content": "<p>Have you solved it yet?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 911660,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-07-01T23:15:04.250000",
          "content": "<p>If the loading of parket files involves Python (non-tensorflow) code, then the TPU will not be able to run it. You have to store your data in a format readable by Tensorflow. Wrapping a pure Python function with PyFunc to turn it into a Tensorflow op unfortunately does not work on TPU.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 741693,
      "author_name": "hongy",
      "author_url": "",
      "post_date": "2020-02-10T22:21:01.930000",
      "content": "<p>I've been using tpus the whole time on gcloud. </p>\n\n<p>I'm testing the kaggle tpu interface and will update. </p>\n\n<p>You will need to use the tfrecord format if you want to do tpu training to its fullest. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 757649,
      "author_name": "Amro",
      "author_url": "",
      "post_date": "2020-02-27T01:49:47.473000",
      "content": "<p>Hi Everyone,\nI wanted to confirm if anyone was able to create dataset reader for TPU that doesn't require saving data to tfrecords. I basically want to try to use it for model training for this competition?\nI wasn't able to find a way to do so on kaggle.\nRegards,\nAmro</p>",
      "votes": 0,
      "replies": [
        {
          "id": 758580,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-02-28T00:13:35.413000",
          "content": "<p>Not a direct answer to your question but a clarification. TFRecords are not required for TPUs.</p>\n\n<p>The problem to solve is to feed the TPU with data fast enough. GCS can sustain the throughput but needs parallel reading from multiple files (16 shards for example) and has a per-file request penalty of a fraction of a second (as any web request). The penalty will get in the way if you have millions or 100Ks of files.</p>\n\n<p>So you need a streamable (streamable = NOT zip) container format to bunch your N files together into a reasonable number of shards (&gt;16 but not thousands).</p>\n\n<p>TFRecords are a solution to that problem but you are free to solve it in any way you like.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 759087,
          "author_name": "Amro",
          "author_url": "",
          "post_date": "2020-02-28T14:55:15.523000",
          "content": "<p>Hi <a href=\"/mgornergoogle\">@mgornergoogle</a> \nThanks for your answer. Really appreciated. I believe the files here on Kaggle side are not readable. May be I should change my question, I could write up a generator similar to how it's written in Keras for training that could read from the files here and generate the tensors for tensorflow in Memory. Keras automated the part of reading using multiprocessing if I'm not mistaken. I find it hard to do the same using TPUs not GPUs?</p>\n\n<p>Any ideas about that? Basic sample code that could use a generator from Kaggle input folders? And we could take it from there.</p>\n\n<p>Thanks again for your input.</p>\n\n<p>Regards,\nAmro</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 916069,
          "author_name": "KaggleHan",
          "author_url": "",
          "post_date": "2020-07-05T10:28:48.567000",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> <br>\nCould you help me with some questions on effective TFRecord read/write and TPU usage, i can't piece together the numerous puzzle pieces even after going through this whole discussion.</p>\n\n<ol>\n<li><p>On the number of shards, i always see 16 shards (100-200mb each) suggested in tutorials kernels, while on Kaggle youtube last week, i hear \"If you have 100k images, use 16 shards\". Confusing me is my 100k images are 10GB, meaning 16 shards would be 600mb each, way more than 100-200mb. So is the sharding concept about number of shards, or file size per shard?</p></li>\n<li><p>What's the difference between the 2 locations, is writing/reading TFRecords to/from 1 of them faster? :  </p>\n\n<ol><li><code>../input/</code>  </li>\n<li><code>../working/</code> (UI shows it as output)</li></ol></li>\n<li><p>Why is my TFrecord writing speed much slower on kaggle compared to local laptop? and Why does the write time not scale linearly with shard size? (4-6x time increase when 2x shard size from 3k to 6k)\n```\nExtra Tests:\n<strong>1 shard writing speed on my local laptop 6 core 2.6GHz (off kaggle)</strong>:\n64 shards of 1647 images: 17 secs\n48 shards of 2196 images: 29 secs\n32 shards of 3294 images: 1 min\n16 shards of 6587 images: 6 min</p></li>\n</ol>\n\n<p><strong>1 shard writing speed on kaggle</strong>:\n64 shards of 1647 images: 2 min\n48 shards of 2196 images: 3 min 44\n32 shards of 3294 images: 8 min 28\n16 shards of 6587 images: 35 min\n```\nGuessing this experiment has something to do with my 12 logical cores CPU (operating at 4ghz when writing) vs 4 CPU(s) on kaggle? Not sure whether or which part of the input pipeline and TFRecord writing uses multithreading. I'm basically following the Flowers tutorial <a href=\"https://colab.research.google.com/github/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/03_Flower_pictures_to_TFRecords.ipynb\">https://colab.research.google.com/github/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/03_Flower_pictures_to_TFRecords.ipynb</a> </p>\n\n<ol>\n<li>Why does TPU training take 2000 seconds on 1st epoch, while &lt;60s on next few epochs? (Not using TFRecords here). Is the extra 1st epoch time spent on copying data from kaggle local to the kaggle GCS Bucket associated with KaggleDatasets()? In Colab, we can copy files from google drive to the temporary directory right where colab is to speed up IO significantly, any analogous ways in TPU kernels to decrease the 1st epoch time?</li>\n</ol>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 929662,
          "author_name": "Jesse Mostipak",
          "author_url": "",
          "post_date": "2020-07-14T20:34:36.540000",
          "content": "<p>hey <a href=\"/datahan\">@datahan</a>! I hope these answers are helpful:</p>\n\n<ol>\n<li>the guidelines provided are meant to cover general use cases, and depending on the size of your data you might need to adjust accordingly. in your use case, 16 shards may not be enough, but in speaking with the team here using more shards is unlikely to have a noticeable difference.</li>\n<li>TFRecords should be stored in a GCS bucket - TPUs can't read//write from either of the locations listed.</li>\n<li>this is likely due to a difference between your computer's specs and the Kaggle VM.</li>\n<li>the first epoch is spent performing XLA calculations, compiling the model, and performing memory optimizations. using TFRecords can help reduce the time of the first epoch.</li>\n</ol>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 747098,
      "author_name": "Shangqiu Li",
      "author_url": "",
      "post_date": "2020-02-16T00:34:19.543000",
      "content": "<p>Did anyone actually sucessfully used TPU for training?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 749616,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-18T20:12:02.003000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 755242,
          "author_name": "hongy",
          "author_url": "",
          "post_date": "2020-02-24T15:33:04.993000",
          "content": "<p>Yes it was very fast. It is also a different beast than pytorch training. A bit part of the difference comes from a massive batch size of 512 or 1024. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 741951,
      "author_name": "Shangqiu Li",
      "author_url": "",
      "post_date": "2020-02-11T02:35:43.193000",
      "content": "<p>What is the speed difference of GPU and TPU?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 741957,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-02-11T02:40:37.720000",
          "content": "<p>Try for yourself using the TPU notebooks in the TPU docs: <a href=\"https://www.kaggle.com/mgornergoogle/five-flowers-with-keras-and-xception-on-tpu\">Five flowers with Keras and Xception on TPU</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 741986,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-02-11T02:56:29.367000",
          "content": "<p>ok. Thanks.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 742094,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-02-11T04:30:30.917000",
          "content": "<p>Did anyone encountered this error?\n&gt; ValueError: The two structures don't have the same nested structure.\nValueError: Could not pack sequence. Structure had 32 elements, but flat_sequence had 1 elements.  Structure: [1, 0, 0, 1, 1, 0, 1, 1, 0, 0, 1, 0, 1, 1, 0, 1, 1, 0, 1, 1, 0, 0, 0, 1, 0, 0, 1, 0, 0, 0, 0, 1], flat_sequence: [None].</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 742106,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-02-11T04:40:02.350000",
          "content": "<p>paste bin:\n<a href=\"https://pastebin.com/KMuKmbfu\">link</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 742995,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-02-11T17:29:34.427000",
          "content": "<p>It looks like you are trying model.fit without using the tf.data.Dataset API. This is what I see in the stack trace:</p>\n\n<p><code>\nfit(self, model, x, y, batch_size, epochs, verbose, callbacks, validation_split, validation_data, shuffle, class_weight, sample_weight, initial_epoch, steps_per_epoch, validation_steps, validation_freq, max_queue_size, workers, use_multiprocessing, **kwargs)\n    233           max_queue_size=max_queue_size,\n    234           workers=workers,\n--&gt; 235           use_multiprocessing=use_multiprocessing)\n</code>\nI would need more code to understand exactly what you are trying to do and whether your are running this on TPU or GPU but in general, tf.data.Dataset is the recommended data pipeline API in Tensorflow and especially when using TPUs.</p>\n\n<p>That said, training on in-memory numpy arrays is also supported on TPUs. It is not going to be fast but for data that fits into memory anyway, this might not matter.</p>\n\n<p>If you were trying to use a generator for data input, I'm pretty sure that will not work. Let me double-check.</p>\n\n<p>Really, the recommended API in Tensorflow for out of memory datasets is tf.data.Dataset</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 743250,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-02-11T23:34:55.933000",
          "content": "<p>[this comment is deleted]</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 743271,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-02-12T00:31:13.667000",
          "content": "<p>Now I'm getting UnavailableError.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 756880,
          "author_name": "AnirbanK",
          "author_url": "",
          "post_date": "2020-02-26T06:58:30.660000",
          "content": "<p>Using TF Dataset now, still getting error in model.fit, though the error msg changed : \"RuntimeError: No registered 'Identity' OpKernel for 'TPU' devices compatible with node {{node Identity}}\n     (OpKernel was found, but attributes didn't match) Requested Attributes: T=DT_UINT8\n    .  Registered:  device='XLA_GPU_JIT'; T in [DT_FLOAT, DT_DOUBLE, DT_INT32, DT_UINT8, DT_INT16, ..., DT_HALF, DT_UINT32, DT_UINT64, DT_RESOURCE, DT_VARIANT]\n  device='XLA_CPU_JIT'; T in [DT_FLOAT, DT_DOUBLE, DT_INT32, DT_UINT8, DT_INT16, ..., DT_HALF, DT_UINT32, DT_UINT64, DT_RESOURCE, DT_VARIANT]\n  device='XLA_TPU_JIT'; T in [DT_FLOAT, DT_DOUBLE, DT_INT32, DT_COMPLEX64, DT_INT64, DT_BOOL, DT_BFLOAT16, DT_UINT32, DT_UINT64, DT_RESOURCE, DT_VARIANT]\n  device='XLA_CPU'; T in [DT_UINT8, DT_QUINT8, DT_UINT16, DT_INT8, DT_QINT8, ..., DT_DOUBLE, DT_COMPLEX64, DT_COMPLEX128, DT_BOOL, DT_BFLOAT16]\n  device='TPU'; T in [DT_INT32, DT_UINT32, DT_BFLOAT16, DT_FLOAT, DT_DOUBLE, DT_BOOL, DT_COMPLEX64, DT_INT64, DT_UINT64]\n  device='TPU_SYSTEM'\n  device='GPU'; T in [DT_HALF]\n  device='GPU'; T in [DT_BFLOAT16]\n  device='GPU'; T in [DT_FLOAT]\n  device='GPU'; T in [DT_DOUBLE]\n  device='GPU'; T in [DT_INT64]\n  device='GPU'; T in [DT_UINT16]\n  device='GPU'; T in [DT_INT16]\n  device='GPU'; T in [DT_UINT8]\n  device='GPU'; T in [DT_INT8]\n  device='GPU'; T in [DT_COMPLEX64]\n  device='GPU'; T in [DT_COMPLEX128]\n  device='GPU'; T in [DT_VARIANT]\n  device='DEFAULT'; T in [DT_STRING]\n  device='DEFAULT'; T in [DT_VARIANT]\n  device='DEFAULT'; T in [DT_RESOURCE]\n  device='CPU'</p>\n\n<pre><code> [[Identity]] \"\n</code></pre>\n\n<p>Any suggestions ? My code is like this : \n    resized_image=np.array()\n    datagen=tf.keras.preprocessing.image.ImageDataGenerator (....)\n    datagen.fit(resized_image)\n    datagen = datagen.flow(resized_image, trainGraphemeY, batch_size=BS)</p>\n\n<pre><code>ds = tf.data.Dataset.from_generator(\n    lambda:datagen,\noutput_types=(tf.uint8, tf.uint8),\noutput_shapes=(resized_image.shape, trainGraphemeY.shape)\n)  \n\nhistory = model_root.fit_generator(ds,\n                                  epochs = EPOCHS, \n                                  steps_per_epoch=resized_image.shape[0] // BS, \n                                  callbacks=[es],verbose=2) \"\n</code></pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 757459,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-02-26T19:35:00.817000",
          "content": "<p>Yes, <code>tf.keras.preprocessing.image.ImageDataGenerator</code> does not work on TPU (yet).\n<a href=\"/cdeotte\">@cdeotte</a> posted TPU-compatible image transformatoin routines <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132191\">here</a>.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 741801,
      "author_name": "Martin Görner",
      "author_url": "",
      "post_date": "2020-02-11T00:31:07.483000",
      "content": "<p>By the way, thank you for the Notebook with TPU explanations! I'm reading it now.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 741726,
      "author_name": "Martin Görner",
      "author_url": "",
      "post_date": "2020-02-10T22:58:16.197000",
      "content": "<p>If you are looking for more info, the TPU docs are here: <a href=\"https://www.kaggle.com/docs/tpu\">https://www.kaggle.com/docs/tpu</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 741721,
      "author_name": "Martin Görner",
      "author_url": "",
      "post_date": "2020-02-10T22:55:26.210000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 741700,
      "author_name": "Moshel",
      "author_url": "",
      "post_date": "2020-02-10T22:25:06.423000",
      "content": "<p>according to this, 30h a week, 3h limit per session\n<a href=\"https://www.kaggle.com/docs/tpu\">https://www.kaggle.com/docs/tpu</a></p>\n\n<p>I would think they would be slightly more generous considering the amount of time tpu setup takes.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 741706,
          "author_name": "hongy",
          "author_url": "",
          "post_date": "2020-02-10T22:31:33.950000",
          "content": "<p>tbh 3 hours per session is probably enough for this dataset. </p>\n\n<p>It takes me 30 seconds to process an epoch of 32768 images. </p>\n\n<p>I train for around 40 epochs and that takes less than 30 minutes. At that point, if you don't decay the LR you will definitely over fit. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 741781,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-02-11T00:16:03.690000",
          "content": "<p>What do you mean by TPU setup ? In the <a href=\"https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu\">Getting started with 100+</a> flowers notebook, the first epoch takes 40s instead of 15s for all other epochs. So yes, there is a little bit of overhead (25s) there, but it does not seem to be too bad, does it ?</p>\n\n<p>By the way, the extra time is where your model gets compiled through XLA to TPU bytecode. Memory layout optimizations are performed there as well.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 741688,
      "author_name": "Darko Androcec",
      "author_url": "",
      "post_date": "2020-02-10T22:12:39.970000",
      "content": "<p>I have found more information: Tensor Processing Units (TPUs) are Now Available on Kaggle - <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus\">https://www.kaggle.com/c/flower-classification-with-tpus</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 741701,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2020-02-10T22:25:12.483000",
          "content": "<p>Yes, we just soft-launched this today. Keep in mind that submissions to the Deepfake Detection Challenge with notebooks using the TPU integration will not be permitted. However, those using TPUs, whether through the integration or offline training (i.e. through Cloud TPUs) can still make submissions through Kaggle notebooks, with GPU turned on.</p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 742741,
          "author_name": "Trigram",
          "author_url": "",
          "post_date": "2020-02-11T14:17:10.067000",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> Well, that's a shame. I was really looking forward to some <strong>practical TPU usage other than identifying flowers</strong>.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 742834,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2020-02-11T15:12:48.843000",
          "content": "<p>Everyone is very welcome to try this problem using TPUs! There is certainly no harm in giving it a spin. Allow me to clarify — I wanted to make sure there was no misconception about being able to submit from a notebook to <em>this competition (Deepfake)</em> with TPU turned on. You have the option to use Cloud TPUs directly or train on Kaggle using TPUs, but then will have to upload your trained model as an external data source into your submission notebook to perform inference (with GPU or CPU on). We’re definitely interested to see what you are able to do on “real problems,” hence enabling it site-wide instead. Best of luck!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 743032,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-02-11T18:20:52.910000",
          "content": "<p>And the main reason for not allowing TPUs in this specific competition is that the competition was launched before TPUs were available on Kaggle. Allowing them now would change the playing field which might be seen as unfair.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 756877,
          "author_name": "AnirbanK",
          "author_url": "",
          "post_date": "2020-02-26T06:55:23.913000",
          "content": "<p>Is it av in Bengali.AI competition ? </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 746116,
      "author_name": "Carlo",
      "author_url": "",
      "post_date": "2020-02-14T16:39:12.377000",
      "content": "<p>Thank you! This answers my question!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "741686": "I see now in my kernel the option to enable TPU. Is this a new Kaggle feature? Do anyone have more information, e.g. did you try it and has it time limitations as GPUs on Kaggle?",
    "741696": "if any google overlords are following this, what zone are these tpus in? ",
    "741695": "WOW! 30hr of free TPU on Kaggle Kernels!!! ",
    "742992": "An update regarding kaggle tpu zones. \n\nIt appears that kaggle tpus are most likely located somewhere in the united states. Yesterday, I ran a kaggle tpu that pulled data from my private us-central1 storage bucket. This incurred an egress fee. \n\nCloud Storage Inter-region GCP Storage egress within NA: 1.73 \n\nI will attempt to clone my bucket to a multi region NA bucket and then test again to see if egress fees are charged. This will take a day bc transactions don't usually take a while to appear on the billing menu. ",
    "741693": "I've been using tpus the whole time on gcloud. \n\nI'm testing the kaggle tpu interface and will update. \n\nYou will need to use the tfrecord format if you want to do tpu training to its fullest. \n",
    "757649": "Hi Everyone,\nI wanted to confirm if anyone was able to create dataset reader for TPU that doesn't require saving data to tfrecords. I basically want to try to use it for model training for this competition?\nI wasn't able to find a way to do so on kaggle.\nRegards,\nAmro",
    "747098": "Did anyone actually sucessfully used TPU for training?",
    "741951": "What is the speed difference of GPU and TPU?",
    "741801": "By the way, thank you for the Notebook with TPU explanations! I'm reading it now.",
    "741726": "If you are looking for more info, the TPU docs are here: https://www.kaggle.com/docs/tpu\n",
    "741721": "",
    "741700": "according to this, 30h a week, 3h limit per session\nhttps://www.kaggle.com/docs/tpu\n\nI would think they would be slightly more generous considering the amount of time tpu setup takes.",
    "741688": "I have found more information: Tensor Processing Units (TPUs) are Now Available on Kaggle - https://www.kaggle.com/c/flower-classification-with-tpus",
    "746116": "Thank you! This answers my question!"
  }
}