{
  "id": 130194,
  "title": "How to use tpus with my own data?",
  "url": "/competitions/flower-classification-with-tpus/discussion/130194",
  "author_name": "",
  "post_date": "2020-02-12T16:50:27.260698Z",
  "votes": null,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hi! I'm trying to use in gcp jupyter notebooks (ai platform) the TPU api.\nI have my own data (actually is the NIH Chest X-ray dataset) with a lot of preprocess. I'm available to train a on gpus from the numpy arrays alocated in RAM.\nWhat transformations i should do to this arrays for be the inputs of my tpu model.</p>\n\n<p>PS: My conexion with the tpu api is ok and i can create my tpu model.</p>\n\n<p>Thanks!</p>",
  "messages": [
    {
      "id": "744210",
      "postDate": "02/12/2020 16:50:27",
      "content": "<p>Hi! I'm trying to use in gcp jupyter notebooks (ai platform) the TPU api.\nI have my own data (actually is the NIH Chest X-ray dataset) with a lot of preprocess. I'm available to train a on gpus from the numpy arrays alocated in RAM.\nWhat transformations i should do to this arrays for be the inputs of my tpu model.</p>\n\n<p>PS: My conexion with the tpu api is ok and i can create my tpu model.</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "Hi! I'm trying to use in gcp jupyter notebooks (ai platform) the TPU api.\nI have my own data (actually is the NIH Chest X-ray dataset) with a lot of preprocess. I'm available to train a on gpus from the numpy arrays alocated in RAM.\nWhat transformations i should do to this arrays for be the inputs of my tpu model.\n\nPS: My conexion with the tpu api is ok and i can create my tpu model.\n\nThanks!",
      "votes": null
    },
    {
      "id": "744296",
      "postDate": "02/12/2020 17:47:43",
      "content": "<p>Yes, you can use TPUs in AI Platform Notebooks. If you are on GCP already, your training data should go into a GCS (Google Cloud Storage) bucket located in the same region as your TPU (for optimal performance, avoid multi-region buckets).</p>\n\n<p>To spin up an AI Platform Jupyter VM + TPU pair, you gave to use <code>gcloud</code> on the command line since there isn't an option for that in the UI yet. I can share a script that does that: <a href=\"https://raw.githubusercontent.com/GoogleCloudPlatform/training-data-analyst/master/courses/fast-and-lean-data-science/create-tpu-deep-learning-vm.sh\">create-tpu-deep-learning-vm.sh</a>.\n<code>gcloud init # to select your project, account, zones, ...</code>\nthen\n<code>./create-tpu-deep-learning-vm.sh choose-a-name --tpu-type v3-8</code></p>\n\n<p>For reading data in, numpy arrays in memory will work for TPUs but might not be the most performant option. If your data fits in memory, it might be OK. For out of memory datasets, us the <code>tf.data.Dataset</code> API. Tutorial here: <a href=\"https://codelabs.developers.google.com/codelabs/keras-flowers-data/#0\">TPU-speed data pipelines: tf.data.Dataset and TFRecords</a></p>",
      "rawMarkdown": "Yes, you can use TPUs in AI Platform Notebooks. If you are on GCP already, your training data should go into a GCS (Google Cloud Storage) bucket located in the same region as your TPU (for optimal performance, avoid multi-region buckets).\n\nTo spin up an AI Platform Jupyter VM + TPU pair, you gave to use `gcloud` on the command line since there isn't an option for that in the UI yet. I can share a script that does that: [create-tpu-deep-learning-vm.sh](https://raw.githubusercontent.com/GoogleCloudPlatform/training-data-analyst/master/courses/fast-and-lean-data-science/create-tpu-deep-learning-vm.sh).\n`gcloud init # to select your project, account, zones, ...`\nthen\n`./create-tpu-deep-learning-vm.sh choose-a-name --tpu-type v3-8`\n\nFor reading data in, numpy arrays in memory will work for TPUs but might not be the most performant option. If your data fits in memory, it might be OK. For out of memory datasets, us the `tf.data.Dataset` API. Tutorial here: [TPU-speed data pipelines: tf.data.Dataset and TFRecords](https://codelabs.developers.google.com/codelabs/keras-flowers-data/#0)",
      "votes": null
    },
    {
      "id": "744903",
      "postDate": "02/13/2020 09:35:58",
      "content": "<p>Thank you Martin. I made the connexion with the tpu without gcloud command line, and its work. I just need to set the enviroment variable to the tpu node.</p>\n\n<p>In the other hand my dataset in arrays way fit in memory well, but when i tried pass it to the model the model didnt work.\nI tried to pass this arrays (i did a lot of preprocess in the data with the arrays) to tensorflow datasets with tensorflow.dataset.to_tensor, but still not working when i train the model (even passing this new dataset with the bacth option) </p>",
      "rawMarkdown": "Thank you Martin. I made the connexion with the tpu without gcloud command line, and its work. I just need to set the enviroment variable to the tpu node.\n\nIn the other hand my dataset in arrays way fit in memory well, but when i tried pass it to the model the model didnt work.\nI tried to pass this arrays (i did a lot of preprocess in the data with the arrays) to tensorflow datasets with tensorflow.dataset.to_tensor, but still not working when i train the model (even passing this new dataset with the bacth option)",
      "votes": null
    },
    {
      "id": "745357",
      "postDate": "02/13/2020 18:31:52",
      "content": "<p>If you just want to convert an in-memory array to a tf.data.Dataset, the correct function is <a href=\"https://www.tensorflow.org/api_docs/python/tf/data/Dataset#from_tensor_slices\"><code>Dataset.from_tensor_slices</code></a>.</p>\n\n<p>But passing a numpy array to model.fit() directly works, even on TPU.</p>",
      "rawMarkdown": "If you just want to convert an in-memory array to a tf.data.Dataset, the correct function is [`Dataset.from_tensor_slices`](https://www.tensorflow.org/api_docs/python/tf/data/Dataset#from_tensor_slices).\n\nBut passing a numpy array to model.fit() directly works, even on TPU.",
      "votes": null
    },
    {
      "id": "745360",
      "postDate": "02/13/2020 18:34:10",
      "content": "<p>In the script I sent earlier (<a href=\"https://raw.githubusercontent.com/GoogleCloudPlatform/training-data-analyst/master/courses/fast-and-lean-data-science/create-tpu-deep-learning-vm.sh\">create-tpu-deep-learning-vm.sh</a>), you will find a way of setting the TPU_NAME environment variable sot that it persists across reboots. And yes of course you can do this by hand.</p>",
      "rawMarkdown": "In the script I sent earlier ([create-tpu-deep-learning-vm.sh](https://raw.githubusercontent.com/GoogleCloudPlatform/training-data-analyst/master/courses/fast-and-lean-data-science/create-tpu-deep-learning-vm.sh)), you will find a way of setting the TPU_NAME environment variable sot that it persists across reboots. And yes of course you can do this by hand.",
      "votes": null
    },
    {
      "id": "745833",
      "postDate": "02/14/2020 09:02:48",
      "content": "<p>Pass arrays didn't work on my case. Gave me error \"failed to serialize\".\nI tried the functions dataset.from_tensor_slices, but my data was larger than 2GB and y couldn't create the tensor.\nSo im trying work with tfrecords now. \nThanks!  </p>",
      "rawMarkdown": "Pass arrays didn't work on my case. Gave me error \"failed to serialize\".\nI tried the functions dataset.from_tensor_slices, but my data was larger than 2GB and y couldn't create the tensor.\nSo im trying work with tfrecords now. \nThanks!",
      "votes": null
    },
    {
      "id": "746178",
      "postDate": "02/14/2020 17:56:33",
      "content": "<p>Not entirely surprised 2GB \"failed to serialize\" 😊</p>",
      "rawMarkdown": "Not entirely surprised 2GB \"failed to serialize\" 😊",
      "votes": null
    },
    {
      "id": "801610",
      "postDate": "04/08/2020 16:18:40",
      "content": "<p>Hi,\nI have a notebook I have only images. How can use them in TPU. Because in TPU I must use  tf.data.Dataset like in this kernel</p>\n\n<p><a href=\"https://www.kaggle.com/hatemamine/flowertpuwin\">https://www.kaggle.com/hatemamine/flowertpuwin</a></p>\n\n<p>My code is like that </p>\n\n<p><a href=\"https://www.kaggle.com/furkankati/person-mask-u-net-model-tpu\">https://www.kaggle.com/furkankati/person-mask-u-net-model-tpu</a></p>\n\n<p>I read data directly from images. I convert images to numpy arrays.</p>\n\n<p>How can I upload my files to google cloud storage and use them for tpu ?</p>",
      "rawMarkdown": "Hi,\nI have a notebook I have only images. How can use them in TPU. Because in TPU I must use  tf.data.Dataset like in this kernel\n\n[https://www.kaggle.com/hatemamine/flowertpuwin](https://www.kaggle.com/hatemamine/flowertpuwin)\n\n\nMy code is like that \n\n[https://www.kaggle.com/furkankati/person-mask-u-net-model-tpu](https://www.kaggle.com/furkankati/person-mask-u-net-model-tpu)\n\nI read data directly from images. I convert images to numpy arrays.\n\nHow can I upload my files to google cloud storage and use them for tpu ?",
      "votes": null
    },
    {
      "id": "920574",
      "postDate": "07/08/2020 17:20:07",
      "content": "<p>Hey Martin, </p>\n\n<p>Is there a way to use my own TPU on a Kaggle dataset directly? Would it be possible for me to provide my own personal TPU the GCS paths for the public dataset on Kaggle?</p>",
      "rawMarkdown": "Hey Martin, \n\nIs there a way to use my own TPU on a Kaggle dataset directly? Would it be possible for me to provide my own personal TPU the GCS paths for the public dataset on Kaggle?",
      "votes": null
    },
    {
      "id": "932296",
      "postDate": "07/16/2020 23:04:48",
      "content": "<p>Sure.<br>\nHere is the <a href=\"https://github.com/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/README-TF2.1.md\" target=\"_blank\">script to provision a TPU + Jupyter on GCP</a>.</p>\n<p>To access the data, run <code>KaggleDatasets().get_gcs_path()</code>  <strong>on Kaggle</strong> once. It will return a GCS bucket name like 'gs://xyz12234567890'. That's where Kaggle stores the attached dataset. You can use this bucket on GCP as well. Just be careful: the bucket is a cache. It will be good for a few days and then be wiped out. If you want something more stable, copy to your own bucekt.</p>",
      "rawMarkdown": "Sure.\nHere is the [script to provision a TPU + Jupyter on GCP](https://github.com/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/README-TF2.1.md).\n\nTo access the data, run `KaggleDatasets().get_gcs_path()`  **on Kaggle** once. It will return a GCS bucket name like 'gs://xyz12234567890'. That's where Kaggle stores the attached dataset. You can use this bucket on GCP as well. Just be careful: the bucket is a cache. It will be good for a few days and then be wiped out. If you want something more stable, copy to your own bucekt.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 744296,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "02/12/2020 17:47:43",
      "content": "<p>Yes, you can use TPUs in AI Platform Notebooks. If you are on GCP already, your training data should go into a GCS (Google Cloud Storage) bucket located in the same region as your TPU (for optimal performance, avoid multi-region buckets).</p>\n\n<p>To spin up an AI Platform Jupyter VM + TPU pair, you gave to use <code>gcloud</code> on the command line since there isn't an option for that in the UI yet. I can share a script that does that: <a href=\"https://raw.githubusercontent.com/GoogleCloudPlatform/training-data-analyst/master/courses/fast-and-lean-data-science/create-tpu-deep-learning-vm.sh\">create-tpu-deep-learning-vm.sh</a>.\n<code>gcloud init # to select your project, account, zones, ...</code>\nthen\n<code>./create-tpu-deep-learning-vm.sh choose-a-name --tpu-type v3-8</code></p>\n\n<p>For reading data in, numpy arrays in memory will work for TPUs but might not be the most performant option. If your data fits in memory, it might be OK. For out of memory datasets, us the <code>tf.data.Dataset</code> API. Tutorial here: <a href=\"https://codelabs.developers.google.com/codelabs/keras-flowers-data/#0\">TPU-speed data pipelines: tf.data.Dataset and TFRecords</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 744903,
          "author_name": "ianwilkinson",
          "author_url": "",
          "post_date": "02/13/2020 09:35:58",
          "content": "<p>Thank you Martin. I made the connexion with the tpu without gcloud command line, and its work. I just need to set the enviroment variable to the tpu node.</p>\n\n<p>In the other hand my dataset in arrays way fit in memory well, but when i tried pass it to the model the model didnt work.\nI tried to pass this arrays (i did a lot of preprocess in the data with the arrays) to tensorflow datasets with tensorflow.dataset.to_tensor, but still not working when i train the model (even passing this new dataset with the bacth option) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745357,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/13/2020 18:31:52",
          "content": "<p>If you just want to convert an in-memory array to a tf.data.Dataset, the correct function is <a href=\"https://www.tensorflow.org/api_docs/python/tf/data/Dataset#from_tensor_slices\"><code>Dataset.from_tensor_slices</code></a>.</p>\n\n<p>But passing a numpy array to model.fit() directly works, even on TPU.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745360,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/13/2020 18:34:10",
          "content": "<p>In the script I sent earlier (<a href=\"https://raw.githubusercontent.com/GoogleCloudPlatform/training-data-analyst/master/courses/fast-and-lean-data-science/create-tpu-deep-learning-vm.sh\">create-tpu-deep-learning-vm.sh</a>), you will find a way of setting the TPU_NAME environment variable sot that it persists across reboots. And yes of course you can do this by hand.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745833,
          "author_name": "ianwilkinson",
          "author_url": "",
          "post_date": "02/14/2020 09:02:48",
          "content": "<p>Pass arrays didn't work on my case. Gave me error \"failed to serialize\".\nI tried the functions dataset.from_tensor_slices, but my data was larger than 2GB and y couldn't create the tensor.\nSo im trying work with tfrecords now. \nThanks!  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746178,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/14/2020 17:56:33",
          "content": "<p>Not entirely surprised 2GB \"failed to serialize\" 😊</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 920574,
          "author_name": "hooong",
          "author_url": "",
          "post_date": "07/08/2020 17:20:07",
          "content": "<p>Hey Martin, </p>\n\n<p>Is there a way to use my own TPU on a Kaggle dataset directly? Would it be possible for me to provide my own personal TPU the GCS paths for the public dataset on Kaggle?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 932296,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "07/16/2020 23:04:48",
          "content": "<p>Sure.<br>\nHere is the <a href=\"https://github.com/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/README-TF2.1.md\" target=\"_blank\">script to provision a TPU + Jupyter on GCP</a>.</p>\n<p>To access the data, run <code>KaggleDatasets().get_gcs_path()</code>  <strong>on Kaggle</strong> once. It will return a GCS bucket name like 'gs://xyz12234567890'. That's where Kaggle stores the attached dataset. You can use this bucket on GCP as well. Just be careful: the bucket is a cache. It will be good for a few days and then be wiped out. If you want something more stable, copy to your own bucekt.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 801610,
      "author_name": "furkankati",
      "author_url": "",
      "post_date": "04/08/2020 16:18:40",
      "content": "<p>Hi,\nI have a notebook I have only images. How can use them in TPU. Because in TPU I must use  tf.data.Dataset like in this kernel</p>\n\n<p><a href=\"https://www.kaggle.com/hatemamine/flowertpuwin\">https://www.kaggle.com/hatemamine/flowertpuwin</a></p>\n\n<p>My code is like that </p>\n\n<p><a href=\"https://www.kaggle.com/furkankati/person-mask-u-net-model-tpu\">https://www.kaggle.com/furkankati/person-mask-u-net-model-tpu</a></p>\n\n<p>I read data directly from images. I convert images to numpy arrays.</p>\n\n<p>How can I upload my files to google cloud storage and use them for tpu ?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "744210": "Hi! I'm trying to use in gcp jupyter notebooks (ai platform) the TPU api.\nI have my own data (actually is the NIH Chest X-ray dataset) with a lot of preprocess. I'm available to train a on gpus from the numpy arrays alocated in RAM.\nWhat transformations i should do to this arrays for be the inputs of my tpu model.\n\nPS: My conexion with the tpu api is ok and i can create my tpu model.\n\nThanks!",
    "744296": "Yes, you can use TPUs in AI Platform Notebooks. If you are on GCP already, your training data should go into a GCS (Google Cloud Storage) bucket located in the same region as your TPU (for optimal performance, avoid multi-region buckets).\n\nTo spin up an AI Platform Jupyter VM + TPU pair, you gave to use `gcloud` on the command line since there isn't an option for that in the UI yet. I can share a script that does that: [create-tpu-deep-learning-vm.sh](https://raw.githubusercontent.com/GoogleCloudPlatform/training-data-analyst/master/courses/fast-and-lean-data-science/create-tpu-deep-learning-vm.sh).\n`gcloud init # to select your project, account, zones, ...`\nthen\n`./create-tpu-deep-learning-vm.sh choose-a-name --tpu-type v3-8`\n\nFor reading data in, numpy arrays in memory will work for TPUs but might not be the most performant option. If your data fits in memory, it might be OK. For out of memory datasets, us the `tf.data.Dataset` API. Tutorial here: [TPU-speed data pipelines: tf.data.Dataset and TFRecords](https://codelabs.developers.google.com/codelabs/keras-flowers-data/#0)",
    "744903": "Thank you Martin. I made the connexion with the tpu without gcloud command line, and its work. I just need to set the enviroment variable to the tpu node.\n\nIn the other hand my dataset in arrays way fit in memory well, but when i tried pass it to the model the model didnt work.\nI tried to pass this arrays (i did a lot of preprocess in the data with the arrays) to tensorflow datasets with tensorflow.dataset.to_tensor, but still not working when i train the model (even passing this new dataset with the bacth option)",
    "745357": "If you just want to convert an in-memory array to a tf.data.Dataset, the correct function is [`Dataset.from_tensor_slices`](https://www.tensorflow.org/api_docs/python/tf/data/Dataset#from_tensor_slices).\n\nBut passing a numpy array to model.fit() directly works, even on TPU.",
    "745360": "In the script I sent earlier ([create-tpu-deep-learning-vm.sh](https://raw.githubusercontent.com/GoogleCloudPlatform/training-data-analyst/master/courses/fast-and-lean-data-science/create-tpu-deep-learning-vm.sh)), you will find a way of setting the TPU_NAME environment variable sot that it persists across reboots. And yes of course you can do this by hand.",
    "745833": "Pass arrays didn't work on my case. Gave me error \"failed to serialize\".\nI tried the functions dataset.from_tensor_slices, but my data was larger than 2GB and y couldn't create the tensor.\nSo im trying work with tfrecords now. \nThanks!",
    "746178": "Not entirely surprised 2GB \"failed to serialize\" 😊",
    "801610": "Hi,\nI have a notebook I have only images. How can use them in TPU. Because in TPU I must use  tf.data.Dataset like in this kernel\n\n[https://www.kaggle.com/hatemamine/flowertpuwin](https://www.kaggle.com/hatemamine/flowertpuwin)\n\n\nMy code is like that \n\n[https://www.kaggle.com/furkankati/person-mask-u-net-model-tpu](https://www.kaggle.com/furkankati/person-mask-u-net-model-tpu)\n\nI read data directly from images. I convert images to numpy arrays.\n\nHow can I upload my files to google cloud storage and use them for tpu ?",
    "920574": "Hey Martin, \n\nIs there a way to use my own TPU on a Kaggle dataset directly? Would it be possible for me to provide my own personal TPU the GCS paths for the public dataset on Kaggle?",
    "932296": "Sure.\nHere is the [script to provision a TPU + Jupyter on GCP](https://github.com/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/README-TF2.1.md).\n\nTo access the data, run `KaggleDatasets().get_gcs_path()`  **on Kaggle** once. It will return a GCS bucket name like 'gs://xyz12234567890'. That's where Kaggle stores the attached dataset. You can use this bucket on GCP as well. Just be careful: the bucket is a cache. It will be good for a few days and then be wiped out. If you want something more stable, copy to your own bucekt."
  },
  "source": "meta"
}