{
  "id": 130003,
  "title": "Excellent tutorial on understanding TFRecords",
  "url": "/competitions/flower-classification-with-tpus/discussion/130003",
  "author_name": "",
  "post_date": "2020-02-11T16:28:10.979425600Z",
  "votes": 25,
  "comment_count": 13,
  "views": 0,
  "content": "<p>I just found out about this excellent (and short!) <a href=\"https://codelabs.developers.google.com/codelabs/keras-flowers-data/#0\">tutorial written by Martin</a> that goes over the data loading process with TFRecords. I think this will definitely be useful if you want to know how to run models with your own image data on the TPUs.</p>",
  "messages": [
    {
      "id": "742902",
      "postDate": "02/11/2020 16:28:10",
      "content": "<p>I just found out about this excellent (and short!) <a href=\"https://codelabs.developers.google.com/codelabs/keras-flowers-data/#0\">tutorial written by Martin</a> that goes over the data loading process with TFRecords. I think this will definitely be useful if you want to know how to run models with your own image data on the TPUs.</p>",
      "rawMarkdown": "I just found out about this excellent (and short!) [tutorial written by Martin](https://codelabs.developers.google.com/codelabs/keras-flowers-data/#0) that goes over the data loading process with TFRecords. I think this will definitely be useful if you want to know how to run models with your own image data on the TPUs.",
      "votes": null
    },
    {
      "id": "743066",
      "postDate": "02/11/2020 18:58:08",
      "content": "<p>Thanks!\nThe TFRecord format is just a container format. The only problem it solves here is the large number of files in the dataset coupled with GCS (Google Cloud Storage). As any internet-accessed storage, a GCS request has a roundtrip time similar to any HTTP request, measured in tenths of a second. For a dataset with 100,000 files, that does not work. On the other hand, GCS can sustain serious throughput if streaming from multiple files in parallel. Sharding your data across a reasonable number (10s or 100s) of reasonably large files (10s to 100s of MB) provides the best GCS throughput.</p>\n\n<p>By the way, the TFRecord tutorial is the first part of the <a href=\"https://codelabs.developers.google.com/codelabs/keras-flowers-tpu/#0\">Keras and modern convnets, on TPUs</a> code lab. And spoiler alert if you were planning to go through it, the <a href=\"https://www.kaggle.com/mgornergoogle/five-flowers-with-keras-and-xception-on-tpu\">TPU sample model</a> in the <a href=\"https://www.kaggle.com/docs/tpu\">Kaggle TPU documentation</a> is the end result and best model you obtain at the end of the code lab.</p>",
      "rawMarkdown": "Thanks!\nThe TFRecord format is just a container format. The only problem it solves here is the large number of files in the dataset coupled with GCS (Google Cloud Storage). As any internet-accessed storage, a GCS request has a roundtrip time similar to any HTTP request, measured in tenths of a second. For a dataset with 100,000 files, that does not work. On the other hand, GCS can sustain serious throughput if streaming from multiple files in parallel. Sharding your data across a reasonable number (10s or 100s) of reasonably large files (10s to 100s of MB) provides the best GCS throughput.\n\nBy the way, the TFRecord tutorial is the first part of the [Keras and modern convnets, on TPUs](https://codelabs.developers.google.com/codelabs/keras-flowers-tpu/#0) code lab. And spoiler alert if you were planning to go through it, the [TPU sample model](https://www.kaggle.com/mgornergoogle/five-flowers-with-keras-and-xception-on-tpu) in the [Kaggle TPU documentation](https://www.kaggle.com/docs/tpu) is the end result and best model you obtain at the end of the code lab.",
      "votes": null
    },
    {
      "id": "743718",
      "postDate": "02/12/2020 08:25:33",
      "content": "<p>I recall using this line made things faster:</p>\n\n<p><code>image, label, height, width = map(lambda x: x.numpy(), (image, label, height, width))</code></p>\n\n<p>And then remove the <code>numpy()</code> calls in the <code>for i</code> loop.</p>\n\n<p>Is there any tutorial to use <code>from_tensor_slices</code> and compressed images? I always get an error if I use <code>tf.string</code> and try to decode on the fly via <code>tf.image.decode_image</code>. It works with GPUs.</p>",
      "rawMarkdown": "I recall using this line made things faster:\n\n`image, label, height, width = map(lambda x: x.numpy(), (image, label, height, width))`\n\nAnd then remove the `numpy()` calls in the `for i` loop.\n\nIs there any tutorial to use `from_tensor_slices` and compressed images? I always get an error if I use `tf.string` and try to decode on the fly via `tf.image.decode_image`. It works with GPUs.",
      "votes": null
    },
    {
      "id": "744305",
      "postDate": "02/12/2020 18:08:08",
      "content": "<p>No .numpy() calls are needed in your data pipeline. Generally only for printing and debugging.</p>\n\n<p>Both the TPU generic sample and the flowers competition getting started notebook read compressed JPEG images from TFRecords. You can look at the code there:\n<a href=\"https://www.kaggle.com/mgornergoogle/five-flowers-with-keras-and-xception-on-tpu\">Five flowers with Keras and Xception on TPU</a>\n<a href=\"https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu/\">Getting started with 100+ flowers on TPU</a></p>\n\n<p>If you want to know everything about tf.data.Dataset and TFRecords, I have a tutorial here:\n<a href=\"https://codelabs.developers.google.com/codelabs/keras-flowers-data/#0\">TPU-speed data pipelines: tf.data.Dataset and TFRecords</a></p>",
      "rawMarkdown": "No .numpy() calls are needed in your data pipeline. Generally only for printing and debugging.\n\nBoth the TPU generic sample and the flowers competition getting started notebook read compressed JPEG images from TFRecords. You can look at the code there:\n[Five flowers with Keras and Xception on TPU](https://www.kaggle.com/mgornergoogle/five-flowers-with-keras-and-xception-on-tpu)\n[Getting started with 100+ flowers on TPU](https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu/)\n\nIf you want to know everything about tf.data.Dataset and TFRecords, I have a tutorial here:\n[TPU-speed data pipelines: tf.data.Dataset and TFRecords](https://codelabs.developers.google.com/codelabs/keras-flowers-data/#0)",
      "votes": null
    },
    {
      "id": "744380",
      "postDate": "02/12/2020 19:28:59",
      "content": "<p>It seems to be needed for <code>BytesList</code>. This matches the documentation: <a href=\"https://www.tensorflow.org/tutorials/load_data/tfrecord\">BytesList won't unpack a string from an EagerTensor.</a>.</p>",
      "rawMarkdown": "It seems to be needed for `BytesList`. This matches the documentation: [BytesList won't unpack a string from an EagerTensor.](https://www.tensorflow.org/tutorials/load_data/tfrecord).",
      "votes": null
    },
    {
      "id": "744389",
      "postDate": "02/12/2020 19:37:36",
      "content": "<p>Ah ok, you were talking about <strong>writing</strong> TFRecords. I though the question was about reading them.</p>",
      "rawMarkdown": "Ah ok, you were talking about **writing** TFRecords. I though the question was about reading them.",
      "votes": null
    },
    {
      "id": "744397",
      "postDate": "02/12/2020 19:47:00",
      "content": "<p>Yes, I was running <a href=\"https://colab.research.google.com/github/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/03_Flower_pictures_to_TFRecords.ipynb#scrollTo=cDIfPMGCjqLO\">your notebook</a>. It actually does both.</p>",
      "rawMarkdown": "Yes, I was running [your notebook](https://colab.research.google.com/github/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/03_Flower_pictures_to_TFRecords.ipynb#scrollTo=cDIfPMGCjqLO). It actually does both.",
      "votes": null
    },
    {
      "id": "778712",
      "postDate": "03/18/2020 17:04:28",
      "content": "<p><a href=\"/xhlulu\">@xhlulu</a> Thank you for sharing this tutorial</p>",
      "rawMarkdown": "xhlulu Thank you for sharing this tutorial",
      "votes": null
    },
    {
      "id": "825871",
      "postDate": "04/29/2020 09:46:13",
      "content": "<p>Thanks it is so useful !</p>",
      "rawMarkdown": "Thanks it is so useful !",
      "votes": null
    },
    {
      "id": "936067",
      "postDate": "07/20/2020 00:15:43",
      "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> \nIndeed great tutorial! </p>\n\n<p>I would like to ask:</p>\n\n<p>1) When converting images to TFRecords, do I need to turn the GPU on? Would it make the conversion faster?\n2) Do you know how can I view TPU \"MXU\" and \"Idle Time\" in Colab?</p>\n\n<p>Many thanks!</p>",
      "rawMarkdown": "mgornergoogle \nIndeed great tutorial! \n\nI would like to ask:\n\n1) When converting images to TFRecords, do I need to turn the GPU on? Would it make the conversion faster?\n2) Do you know how can I view TPU \"MXU\" and \"Idle Time\" in Colab?\n\nMany thanks!",
      "votes": null
    },
    {
      "id": "938708",
      "postDate": "07/21/2020 17:35:47",
      "content": "<p>1) I don't believe my script uses the GPU but maybe.<br>\n2) You cannot</p>",
      "rawMarkdown": "1) I don't believe my script uses the GPU but maybe.\n2) You cannot",
      "votes": null
    },
    {
      "id": "957005",
      "postDate": "08/04/2020 01:58:53",
      "content": "<p><a href=\"/xhlulu\">@xhlulu</a> you are the best man </p>",
      "rawMarkdown": "xhlulu you are the best man",
      "votes": null
    },
    {
      "id": "969853",
      "postDate": "08/14/2020 02:11:01",
      "content": "<p><a href=\"https://www.kaggle.com/mgornergoogle\" target=\"_blank\">@mgornergoogle</a> Great tutorials! I particularly like your tutorials where you explain how CNNs work.</p>",
      "rawMarkdown": "mgornergoogle Great tutorials! I particularly like your tutorials where you explain how CNNs work.",
      "votes": null
    },
    {
      "id": "972299",
      "postDate": "08/16/2020 12:48:16",
      "content": "<p>Thanks man!</p>",
      "rawMarkdown": "Thanks man!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 969853,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "08/14/2020 02:11:01",
      "content": "<p><a href=\"https://www.kaggle.com/mgornergoogle\" target=\"_blank\">@mgornergoogle</a> Great tutorials! I particularly like your tutorials where you explain how CNNs work.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 972299,
      "author_name": "delllectron",
      "author_url": "",
      "post_date": "08/16/2020 12:48:16",
      "content": "<p>Thanks man!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 743066,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "02/11/2020 18:58:08",
      "content": "<p>Thanks!\nThe TFRecord format is just a container format. The only problem it solves here is the large number of files in the dataset coupled with GCS (Google Cloud Storage). As any internet-accessed storage, a GCS request has a roundtrip time similar to any HTTP request, measured in tenths of a second. For a dataset with 100,000 files, that does not work. On the other hand, GCS can sustain serious throughput if streaming from multiple files in parallel. Sharding your data across a reasonable number (10s or 100s) of reasonably large files (10s to 100s of MB) provides the best GCS throughput.</p>\n\n<p>By the way, the TFRecord tutorial is the first part of the <a href=\"https://codelabs.developers.google.com/codelabs/keras-flowers-tpu/#0\">Keras and modern convnets, on TPUs</a> code lab. And spoiler alert if you were planning to go through it, the <a href=\"https://www.kaggle.com/mgornergoogle/five-flowers-with-keras-and-xception-on-tpu\">TPU sample model</a> in the <a href=\"https://www.kaggle.com/docs/tpu\">Kaggle TPU documentation</a> is the end result and best model you obtain at the end of the code lab.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 743718,
      "author_name": "seesee",
      "author_url": "",
      "post_date": "02/12/2020 08:25:33",
      "content": "<p>I recall using this line made things faster:</p>\n\n<p><code>image, label, height, width = map(lambda x: x.numpy(), (image, label, height, width))</code></p>\n\n<p>And then remove the <code>numpy()</code> calls in the <code>for i</code> loop.</p>\n\n<p>Is there any tutorial to use <code>from_tensor_slices</code> and compressed images? I always get an error if I use <code>tf.string</code> and try to decode on the fly via <code>tf.image.decode_image</code>. It works with GPUs.</p>",
      "votes": null,
      "replies": [
        {
          "id": 744305,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/12/2020 18:08:08",
          "content": "<p>No .numpy() calls are needed in your data pipeline. Generally only for printing and debugging.</p>\n\n<p>Both the TPU generic sample and the flowers competition getting started notebook read compressed JPEG images from TFRecords. You can look at the code there:\n<a href=\"https://www.kaggle.com/mgornergoogle/five-flowers-with-keras-and-xception-on-tpu\">Five flowers with Keras and Xception on TPU</a>\n<a href=\"https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu/\">Getting started with 100+ flowers on TPU</a></p>\n\n<p>If you want to know everything about tf.data.Dataset and TFRecords, I have a tutorial here:\n<a href=\"https://codelabs.developers.google.com/codelabs/keras-flowers-data/#0\">TPU-speed data pipelines: tf.data.Dataset and TFRecords</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 744380,
          "author_name": "seesee",
          "author_url": "",
          "post_date": "02/12/2020 19:28:59",
          "content": "<p>It seems to be needed for <code>BytesList</code>. This matches the documentation: <a href=\"https://www.tensorflow.org/tutorials/load_data/tfrecord\">BytesList won't unpack a string from an EagerTensor.</a>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 744389,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/12/2020 19:37:36",
          "content": "<p>Ah ok, you were talking about <strong>writing</strong> TFRecords. I though the question was about reading them.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 744397,
          "author_name": "seesee",
          "author_url": "",
          "post_date": "02/12/2020 19:47:00",
          "content": "<p>Yes, I was running <a href=\"https://colab.research.google.com/github/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/03_Flower_pictures_to_TFRecords.ipynb#scrollTo=cDIfPMGCjqLO\">your notebook</a>. It actually does both.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 778712,
      "author_name": "usharengaraju",
      "author_url": "",
      "post_date": "03/18/2020 17:04:28",
      "content": "<p><a href=\"/xhlulu\">@xhlulu</a> Thank you for sharing this tutorial</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 825871,
      "author_name": "salihsarp",
      "author_url": "",
      "post_date": "04/29/2020 09:46:13",
      "content": "<p>Thanks it is so useful !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 936067,
      "author_name": "chho1229",
      "author_url": "",
      "post_date": "07/20/2020 00:15:43",
      "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> \nIndeed great tutorial! </p>\n\n<p>I would like to ask:</p>\n\n<p>1) When converting images to TFRecords, do I need to turn the GPU on? Would it make the conversion faster?\n2) Do you know how can I view TPU \"MXU\" and \"Idle Time\" in Colab?</p>\n\n<p>Many thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 938708,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "07/21/2020 17:35:47",
          "content": "<p>1) I don't believe my script uses the GPU but maybe.<br>\n2) You cannot</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 957005,
      "author_name": "pkhorram",
      "author_url": "",
      "post_date": "08/04/2020 01:58:53",
      "content": "<p><a href=\"/xhlulu\">@xhlulu</a> you are the best man </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "742902": "I just found out about this excellent (and short!) [tutorial written by Martin](https://codelabs.developers.google.com/codelabs/keras-flowers-data/#0) that goes over the data loading process with TFRecords. I think this will definitely be useful if you want to know how to run models with your own image data on the TPUs.",
    "743066": "Thanks!\nThe TFRecord format is just a container format. The only problem it solves here is the large number of files in the dataset coupled with GCS (Google Cloud Storage). As any internet-accessed storage, a GCS request has a roundtrip time similar to any HTTP request, measured in tenths of a second. For a dataset with 100,000 files, that does not work. On the other hand, GCS can sustain serious throughput if streaming from multiple files in parallel. Sharding your data across a reasonable number (10s or 100s) of reasonably large files (10s to 100s of MB) provides the best GCS throughput.\n\nBy the way, the TFRecord tutorial is the first part of the [Keras and modern convnets, on TPUs](https://codelabs.developers.google.com/codelabs/keras-flowers-tpu/#0) code lab. And spoiler alert if you were planning to go through it, the [TPU sample model](https://www.kaggle.com/mgornergoogle/five-flowers-with-keras-and-xception-on-tpu) in the [Kaggle TPU documentation](https://www.kaggle.com/docs/tpu) is the end result and best model you obtain at the end of the code lab.",
    "743718": "I recall using this line made things faster:\n\n`image, label, height, width = map(lambda x: x.numpy(), (image, label, height, width))`\n\nAnd then remove the `numpy()` calls in the `for i` loop.\n\nIs there any tutorial to use `from_tensor_slices` and compressed images? I always get an error if I use `tf.string` and try to decode on the fly via `tf.image.decode_image`. It works with GPUs.",
    "744305": "No .numpy() calls are needed in your data pipeline. Generally only for printing and debugging.\n\nBoth the TPU generic sample and the flowers competition getting started notebook read compressed JPEG images from TFRecords. You can look at the code there:\n[Five flowers with Keras and Xception on TPU](https://www.kaggle.com/mgornergoogle/five-flowers-with-keras-and-xception-on-tpu)\n[Getting started with 100+ flowers on TPU](https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu/)\n\nIf you want to know everything about tf.data.Dataset and TFRecords, I have a tutorial here:\n[TPU-speed data pipelines: tf.data.Dataset and TFRecords](https://codelabs.developers.google.com/codelabs/keras-flowers-data/#0)",
    "744380": "It seems to be needed for `BytesList`. This matches the documentation: [BytesList won't unpack a string from an EagerTensor.](https://www.tensorflow.org/tutorials/load_data/tfrecord).",
    "744389": "Ah ok, you were talking about **writing** TFRecords. I though the question was about reading them.",
    "744397": "Yes, I was running [your notebook](https://colab.research.google.com/github/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/03_Flower_pictures_to_TFRecords.ipynb#scrollTo=cDIfPMGCjqLO). It actually does both.",
    "778712": "xhlulu Thank you for sharing this tutorial",
    "825871": "Thanks it is so useful !",
    "936067": "mgornergoogle \nIndeed great tutorial! \n\nI would like to ask:\n\n1) When converting images to TFRecords, do I need to turn the GPU on? Would it make the conversion faster?\n2) Do you know how can I view TPU \"MXU\" and \"Idle Time\" in Colab?\n\nMany thanks!",
    "938708": "1) I don't believe my script uses the GPU but maybe.\n2) You cannot",
    "957005": "xhlulu you are the best man",
    "969853": "mgornergoogle Great tutorials! I particularly like your tutorials where you explain how CNNs work.",
    "972299": "Thanks man!"
  },
  "source": "meta"
}