{"cells":[{"metadata":{"id":"uJ9H8w73J5T-","colab_type":"text"},"cell_type":"markdown","source":"TFRecord Format:\n===\nIn this competition we're classifying 104 types of flowers based on their images drawn from five different public datasets. This competition is different in that images are provided in TFRecord format.\n\nTo make things easier, it will be better if we have basic understanding of dataset and how to experiment with different methods faster even on Google Colab without limiting yourself."},{"metadata":{"id":"fOwJ0V_9FnBB","colab_type":"text"},"cell_type":"markdown","source":"# Overview\nThis notebook describes all the different elements of TFRecord format. Obviously, to cover everything, it has to be fairly long. I have only covered the basics on using TFRecord on image data. I will be updating it in few weeks and I encourage everyone to read Tensorflow core tutorials as it has covered all the concpets thoroughly."},{"metadata":{"_uuid":"7afad70768ff847794be2ab8312d7fa04ca1c571","collapsed":true,"id":"0-uqg9HJ10uY","colab_type":"text"},"cell_type":"markdown","source":"# What is TFRecord?\n> As per Tensorflow's documentation, \" ... approach is to convert whatever data you have into a supported format. This approach makes it easier to mix and match data sets and network architectures. The recommended format for TensorFlow is a TFRecords file containing tf.train.Example protocol buffers (which contain Features as a field).\"\n\nIn Layman's terms, The TFRecord format is a simple format for storing a sequence of binary records.\n\nSo, why we even need a container for our image dataset we can simply extract our image into folder, read them into ram and then run, any image classification techniques. Well, There are plenty of reasons:\n1.  Suppose you are experimenting with different image classification algorithms it is beneficial to do basic preprocessing and convert data into format that is faster to load. TFRecord makes that job easier. You can preprocess data and then convert it into TFRecord file. It will save your time and all you need to do read TFRecord file without any preprocessing and then, you can test your ideas on that dataset much faster without going to same step every time.\ne.g., for pandas it is hdf5 files and for numpy it is npy, similarly for tensorflow, it is tfrecord files.\n\n2. To dynamically shuffle at random places and also change the ratio of train:test:validate from the whole dataset. When you are working with an image dataset, what is the first thing you do? Split into Train, Test, Vaildate, sets, and then shuffle to not have any biased data distribution. In case of TFRecord, everything is in a single file and we can use that file to shuffle our dataset.\n"},{"metadata":{"colab_type":"text","id":"WkRreBf1eDVc"},"cell_type":"markdown","source":"## Setup"},{"metadata":{"id":"hP_rUdRS3l0n","colab_type":"code","outputId":"91cb20a5-7687-4bd8-b966-71ae05b026ec","colab":{"base_uri":"https://localhost:8080/","height":102},"trusted":true},"cell_type":"code","source":"#%tensorflow_version 2.x #only exists in Colab.","execution_count":null,"outputs":[]},{"metadata":{"id":"dYN3fjnX3rMI","colab_type":"code","outputId":"458f1822-3cf5-4bba-ef53-5798f55de4d6","colab":{"base_uri":"https://localhost:8080/","height":34},"trusted":true},"cell_type":"code","source":"# import libraries\nimport math, re, os\nimport tensorflow as tf\nimport numpy as np\nfrom matplotlib import pyplot as plt\nimport IPython.display as display\n#from kaggle_datasets import KaggleDatasets\nfrom sklearn.metrics import f1_score, precision_score, recall_score, confusion_matrix\nprint(\"Tensorflow version \" + tf.__version__)\nAUTO = tf.data.experimental.AUTOTUNE","execution_count":null,"outputs":[]},{"metadata":{"id":"O-aBSiaPmFVK","colab_type":"text"},"cell_type":"markdown","source":"Writing and Reading a TFRecord file\n===\nIn practice, There are different types of input data but the process of creating a TFRecord file will be the same for each type of dataset:\n\n1. Within each observation, each value needs to be converted to a `tf.train.Feature` containing one of the 3 compatible types, BytesList, FloaList,and Int64List. These are not python data types but tf.train.Feature types, that are used to store python data in compatible formats for TensorFlow operations.\n\n1. You create a map (dictionary) from the feature name string to the encoded feature value produced in #1.\n\n1. The map produced in step 2 is converted to a [`Features` message](https://github.com/tensorflow/tensorflow/blob/master/tensorflow/core/example/feature.proto#L85).\n\n1. Write the `tf.Example` observations to the TFRecord file"},{"metadata":{"id":"v3swGzoPx5YF","colab_type":"text"},"cell_type":"markdown","source":">**Aknowledgement**  \nTensorFlow core team did a great job sharing tutorials on TFRecord.  \nhttps://www.tensorflow.org/tutorials/load_data/tfrecord  \nhttps://codelabs.developers.google.com/codelabs/keras-flowers-data"},{"metadata":{"colab_type":"code","id":"YZt4DXC41BSJ","outputId":"fa2e3ef0-d8a8-4157-f59d-26de31289f05","colab":{"base_uri":"https://localhost:8080/","height":105},"trusted":true},"cell_type":"code","source":"# In this notebook, We have only considered image data so, two image files are used.\ncat_in_snow  = tf.keras.utils.get_file('320px-Felis_catus-cat_on_snow.jpg', 'https://storage.googleapis.com/download.tensorflow.org/example_images/320px-Felis_catus-cat_on_snow.jpg')\nwilliamsburg_bridge = tf.keras.utils.get_file('194px-New_East_River_Bridge_from_Brooklyn_det.4a09796u.jpg','https://storage.googleapis.com/download.tensorflow.org/example_images/194px-New_East_River_Bridge_from_Brooklyn_det.4a09796u.jpg')","execution_count":null,"outputs":[]},{"metadata":{"colab_type":"code","id":"y-RakwRD1BSS","outputId":"ca4d2f81-8118-45b3-a37c-362dab40fcb9","colab":{"base_uri":"https://localhost:8080/","height":247},"trusted":true},"cell_type":"code","source":"# checking the image file\ndisplay.display(display.Image(filename=cat_in_snow))\ndisplay.display(display.HTML('Image cc-by: <a \"href=https://commons.wikimedia.org/wiki/File:Felis_catus-cat_on_snow.jpg\">Von.grzanka</a>'))","execution_count":null,"outputs":[]},{"metadata":{"colab_type":"code","id":"3VB8sHLr1BSd","outputId":"03abeb34-3973-4294-8614-b2f690a5cb47","colab":{"base_uri":"https://localhost:8080/","height":273},"trusted":true},"cell_type":"code","source":"# checking the image file\ndisplay.display(display.Image(filename=williamsburg_bridge))\ndisplay.display(display.HTML('<a \"href=https://commons.wikimedia.org/wiki/File:New_East_River_Bridge_from_Brooklyn_det.4a09796u.jpg\">From Wikimedia</a>'))","execution_count":null,"outputs":[]},{"metadata":{"id":"tTG57Mk68azO","colab_type":"text"},"cell_type":"markdown","source":"In order to convert python data type to a standard TensorFlow type tf.train.Feature, we will use the functions below. Each of these functions takes scalar input value and returns a tf.train.Feature"},{"metadata":{"id":"OyiZ-6Pg6fqa","colab_type":"code","colab":{},"trusted":true},"cell_type":"code","source":"# The following functions can be used to convert a value to a type compatible\n# with tf.Example.\n\ndef _bytes_feature(value):\n  \"\"\"Returns a bytes_list from a string / byte.\"\"\"\n  if isinstance(value, type(tf.constant(0))):\n    value = value.numpy() # BytesList won't unpack a string from an EagerTensor.\n  return tf.train.Feature(bytes_list=tf.train.BytesList(value=[value]))\n\ndef _int64_feature(value):\n  \"\"\"Returns an int64_list from a bool / enum / int / uint.\"\"\"\n  return tf.train.Feature(int64_list=tf.train.Int64List(value=[value]))","execution_count":null,"outputs":[]},{"metadata":{"colab_type":"code","id":"TZUCLCyZ1BSi","colab":{},"trusted":true},"cell_type":"code","source":"# create a dictionary to map classes to images, \n# this will be helpful when we use image classification algorithm on dataset\nimage_labels = {\n    cat_in_snow : 0,\n    williamsburg_bridge : 1,\n}\n\n# Create a function to apply entire process to each element of dataset.\n# process the two images into 'tf.Example' messages.\ndef image_example(image_string, label):\n  \"\"\"\n  Creates a tf.Example message ready to be written to a file.\n  \"\"\"\n  # Create a dictionary mapping the feature name to the tf.Example-compatible\n  # data type.\n  image_feature_description = {\n      \"image\": _bytes_feature(image_string),\n      \"class\": _int64_feature(label),\n      }\n  # Create a Features message using tf.train.Example.\n  return tf.train.Example(features=tf.train.Features(feature=image_feature_description))","execution_count":null,"outputs":[]},{"metadata":{"id":"8KIogfTG6vHt","colab_type":"code","colab":{},"trusted":true},"cell_type":"code","source":"# define a filename to store preprocessed image data:\nrecord_file = 'images.tfrecords'\n# Write the `tf.Example` observations to the file.\nwith tf.io.TFRecordWriter(record_file) as writer:\n  for filename, label in image_labels.items():\n    image_string = open(filename, 'rb').read()\n    # storing all the features in the tf.Example message.\n    tf_example = image_example(image_string, label)\n    # write the example messages to a file named images.tfrecords\n    writer.write(tf_example.SerializeToString())","execution_count":null,"outputs":[]},{"metadata":{"id":"bQqjELp99IGn","colab_type":"code","outputId":"4474eeda-36b5-4824-d313-46f730d43ae0","colab":{"base_uri":"https://localhost:8080/","height":34},"trusted":true},"cell_type":"code","source":"# checking if file is written\n!du -sh {record_file}","execution_count":null,"outputs":[]},{"metadata":{"id":"CfdO1ufQ9PkS","colab_type":"code","colab":{},"trusted":true},"cell_type":"code","source":"# to read TFRecord file use TFRecordDataset\nraw_image_dataset = tf.data.TFRecordDataset(record_file)\n\n# Create a dictionary describing the features.\nimage_feature_description = {\n    \"image\": tf.io.FixedLenFeature([], tf.string), # tf.string means bytestring\n    \"class\": tf.io.FixedLenFeature([], tf.int64),  # shape [] means single element\n    }\n\n# create a function to apply image feature description to each observation\ndef _parse_image_function(example_proto):\n  # parse the input tf.Example proto using the dictionary above.\n  return tf.io.parse_single_example(example_proto, image_feature_description)\n\n# use map to apply this operation to each element of dataset\nparsed_image_dataset = raw_image_dataset.map(_parse_image_function)","execution_count":null,"outputs":[]},{"metadata":{"id":"-w3R8LVo91ry","colab_type":"code","outputId":"3578ce87-ff2b-4cac-f5d1-9f97f51f2708","colab":{"base_uri":"https://localhost:8080/","height":247},"trusted":true},"cell_type":"code","source":"# Use the .take method to only pull one example from the dataset.\nfor image_features in parsed_image_dataset.take(1):\n  image = image_features['image'].numpy()\n  display.display(display.Image(data=image))\n  classes = image_features['class'].numpy()\n  print('The label of image is', classes)","execution_count":null,"outputs":[]},{"metadata":{"id":"qL-zt-Dc0al2","colab_type":"text"},"cell_type":"markdown","source":"Hence, The process of serializing data into TFRecord format will be: \n**Data -> FeatureSet (a dictionary of features)-> Example -> Serialized Example -> TFRecord.**\n\nand to read it back, the process is reversed.\n**TFRecord -> SerializedExample -> Example -> FeatureSet -> Data**"},{"metadata":{"id":"-IR_j-ntEzhF","colab_type":"text"},"cell_type":"markdown","source":"### On Flower Classification with TPUs\n"},{"metadata":{"id":"_xEvIGJMEmMw","colab_type":"code","colab":{},"trusted":true},"cell_type":"code","source":"file = '../input/flower-classification-with-tpus/tfrecords-jpeg-192x192/train/00-192x192-798.tfrec'","execution_count":null,"outputs":[]},{"metadata":{"id":"_GT99g9lECFQ","colab_type":"code","colab":{},"trusted":true},"cell_type":"code","source":"# on \n# to read TFRecord file use TFRecordDataset\nraw_image_dataset = tf.data.TFRecordDataset(file)\n\n# Create a dictionary describing the features.\nimage_feature_description = {\n    \"image\": tf.io.FixedLenFeature([], tf.string), # tf.string means bytestring\n    \"class\": tf.io.FixedLenFeature([], tf.int64),  # shape [] means single element\n    }\n\n# create a function to apply image feature description to each observation\ndef _parse_image_function(example_proto):\n  # parse the input tf.Example proto using the dictionary above.\n  return tf.io.parse_single_example(example_proto, image_feature_description)\n\n# use map to apply this operation to each element of dataset\nparsed_image_dataset = raw_image_dataset.map(_parse_image_function)","execution_count":null,"outputs":[]},{"metadata":{"id":"Bia7Vz9rEucI","colab_type":"code","colab":{"base_uri":"https://localhost:8080/","height":226},"outputId":"b0196df4-d1b5-4e10-f67c-26ef681b06ce","trusted":true},"cell_type":"code","source":"# Use the .take method to only pull one example from the dataset.\nfor image_features in parsed_image_dataset.take(1):\n  image = image_features['image'].numpy()\n  display.display(display.Image(data=image))\n  classes = image_features['class'].numpy()\n  print('The label of image is', classes)","execution_count":null,"outputs":[]},{"metadata":{"id":"QcjVVdXZ5gcL","colab_type":"text"},"cell_type":"markdown","source":"Faster training on TPU:\n===\nAccoring to Google's team behind Colab's free TPU:\n>\"Artificial neural networks based on the AI applications used to train the TPUs are 15 and 30 times faster than CPUs and GPUs!\"\n\nSo, how can we do that; As per my experimentation, when you use CPU and GPU static shape is not that important, but incase of XLA/TPU static shape and batch size makes a very big difference.\n\nHence, if you use static input batch_size i.e., train the TPU model with static batch_size*8 (number of TPU cores). The epoch time reduced to 20%-50% as compared to training model on GPU. Since colab TPU has 8 TPU cores which operates as independent processing units."},{"metadata":{"id":"KOLIVTGEGZn1","colab_type":"text"},"cell_type":"markdown","source":"References:\n===\n1. https://www.dlology.com/blog/how-to-train-keras-model-x20-times-faster-with-tpu-for-free/\n\n1. https://codelabs.developers.google.com/codelabs/keras-flowers-data/#4\n\n1. https://www.tensorflow.org/tutorials/load_data/tfrecord\n\n1. https://www.skcript.com/svr/why-every-tensorflow-developer-should-know-about-tfrecord/"},{"metadata":{"id":"tqN-lQ9lFuXW","colab_type":"text"},"cell_type":"markdown","source":"Conclusion\n===\nThis is just a first cut notebook on using tfrecord format. Also more rigorous experiment can be done and heuristics can be explored for faster preprocessing and training. I hope this helped you in understanding of using TensorFlow's TFRecords.\n\nComments, suggestions, criticism are welcomed. Thanks\n\n### To be continued ................\n"}],"metadata":{"colab":{"name":"Understanding TFRecord format.ipynb","provenance":[],"collapsed_sections":[]},"kernelspec":{"name":"python3","display_name":"Python 3"}},"nbformat":4,"nbformat_minor":4}