{
  "id": 284513,
  "title": "TFRecord Dataset from Images and Their Masks (Faster training with TPUs)",
  "url": "/competitions/sartorius-cell-instance-segmentation/discussion/284513",
  "author_name": "",
  "post_date": "2021-11-01T08:58:56.630977400Z",
  "votes": 4,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Converting train annotation data and preparing the masks is a complicated task and time-consuming. For faster processing, I have converted all training samples and their masks into one <code>tfrec</code> file. </p>\n<p>You can find the dataset <a href=\"https://www.kaggle.com/kavehshahhosseini/sartorius-cell-instance-segmentation-in-tfrecord\" target=\"_blank\">here</a>. Also, <a href=\"https://www.kaggle.com/kavehshahhosseini/sartorius-convert-images-and-masks-to-tfrecord\" target=\"_blank\">here</a> is the notebook that I converted the data with. </p>\n<p>After adding <a href=\"https://www.kaggle.com/kavehshahhosseini/sartorius-cell-instance-segmentation-in-tfrecord\" target=\"_blank\">this dataset</a> to your notebook, you can easily load the data and create a dataset in TensorFlow or Keras like this:</p>\n<pre><code>def deserialize_example(serialized_string):\n    image_feature_description = {\n        'image': tf.io.FixedLenFeature([], tf.string),\n        'label': tf.io.FixedLenFeature([], tf.string)\n    }\n    parsed_record = tf.io.parse_single_example(serialized_string, image_feature_description)\n    image = tf.reshape(tf.io.decode_raw(parsed_record['image'], tf.float32),(IMAGE_HEIGHT, IMAGE_WIDTH, 3))\n    label = tf.reshape(tf.io.decode_raw(parsed_record['label'], tf.uint8),(IMAGE_HEIGHT, IMAGE_WIDTH, -1))\n    return image, label\n</code></pre>\n<p>Then creating <code>TFRecordDataset</code>:</p>\n<pre><code>train_set = tf.data.TFRecordDataset(os.path.join(outpath,\"sartorius.tfrec\"), compression_type=\"GZIP\").map(deserialize_example)\n</code></pre>\n<p>There are 606 images  with shape <code>(520,704,3)</code> and dtype <code>float32</code>. Masks are in shape <code>(520,704,n_instances)</code> and dtype <code>uint8</code>. </p>\n<h2>Data</h2>\n<pre><code>ds = train_set.take(1)\nfor image, label in ds:\n    print(image.shape)\n    print(label.shape)\n\nfig, axs = plt.subplots(1, 2,figsize=(20, 20))\naxs[0].imshow(image)\naxs[0].axis('off')\naxs[1].imshow(np.sum(label, axis=-1))\naxs[1].axis('off')\nplt.show()\n</code></pre>\n<p><img src=\"https://i.imgur.com/odZkChr.png\" alt=\"result\"><br>\nEnjoy!</p>",
  "messages": [
    {
      "id": "1566800",
      "postDate": "11/01/2021 08:58:56",
      "content": "<p>Converting train annotation data and preparing the masks is a complicated task and time-consuming. For faster processing, I have converted all training samples and their masks into one <code>tfrec</code> file. </p>\n<p>You can find the dataset <a href=\"https://www.kaggle.com/kavehshahhosseini/sartorius-cell-instance-segmentation-in-tfrecord\" target=\"_blank\">here</a>. Also, <a href=\"https://www.kaggle.com/kavehshahhosseini/sartorius-convert-images-and-masks-to-tfrecord\" target=\"_blank\">here</a> is the notebook that I converted the data with. </p>\n<p>After adding <a href=\"https://www.kaggle.com/kavehshahhosseini/sartorius-cell-instance-segmentation-in-tfrecord\" target=\"_blank\">this dataset</a> to your notebook, you can easily load the data and create a dataset in TensorFlow or Keras like this:</p>\n<pre><code>def deserialize_example(serialized_string):\n    image_feature_description = {\n        'image': tf.io.FixedLenFeature([], tf.string),\n        'label': tf.io.FixedLenFeature([], tf.string)\n    }\n    parsed_record = tf.io.parse_single_example(serialized_string, image_feature_description)\n    image = tf.reshape(tf.io.decode_raw(parsed_record['image'], tf.float32),(IMAGE_HEIGHT, IMAGE_WIDTH, 3))\n    label = tf.reshape(tf.io.decode_raw(parsed_record['label'], tf.uint8),(IMAGE_HEIGHT, IMAGE_WIDTH, -1))\n    return image, label\n</code></pre>\n<p>Then creating <code>TFRecordDataset</code>:</p>\n<pre><code>train_set = tf.data.TFRecordDataset(os.path.join(outpath,\"sartorius.tfrec\"), compression_type=\"GZIP\").map(deserialize_example)\n</code></pre>\n<p>There are 606 images  with shape <code>(520,704,3)</code> and dtype <code>float32</code>. Masks are in shape <code>(520,704,n_instances)</code> and dtype <code>uint8</code>. </p>\n<h2>Data</h2>\n<pre><code>ds = train_set.take(1)\nfor image, label in ds:\n    print(image.shape)\n    print(label.shape)\n\nfig, axs = plt.subplots(1, 2,figsize=(20, 20))\naxs[0].imshow(image)\naxs[0].axis('off')\naxs[1].imshow(np.sum(label, axis=-1))\naxs[1].axis('off')\nplt.show()\n</code></pre>\n<p><img src=\"https://i.imgur.com/odZkChr.png\" alt=\"result\"><br>\nEnjoy!</p>",
      "rawMarkdown": "Converting train annotation data and preparing the masks is a complicated task and time-consuming. For faster processing, I have converted all training samples and their masks into one `tfrec` file. \n\nYou can find the dataset [here](https://www.kaggle.com/kavehshahhosseini/sartorius-cell-instance-segmentation-in-tfrecord). Also, [here](https://www.kaggle.com/kavehshahhosseini/sartorius-convert-images-and-masks-to-tfrecord) is the notebook that I converted the data with. \n\nAfter adding [this dataset](https://www.kaggle.com/kavehshahhosseini/sartorius-cell-instance-segmentation-in-tfrecord) to your notebook, you can easily load the data and create a dataset in TensorFlow or Keras like this:\n```\ndef deserialize_example(serialized_string):\n    image_feature_description = {\n        'image': tf.io.FixedLenFeature([], tf.string),\n        'label': tf.io.FixedLenFeature([], tf.string)\n    }\n    parsed_record = tf.io.parse_single_example(serialized_string, image_feature_description)\n    image = tf.reshape(tf.io.decode_raw(parsed_record['image'], tf.float32),(IMAGE_HEIGHT, IMAGE_WIDTH, 3))\n    label = tf.reshape(tf.io.decode_raw(parsed_record['label'], tf.uint8),(IMAGE_HEIGHT, IMAGE_WIDTH, -1))\n    return image, label\n```\nThen creating `TFRecordDataset`:\n```\ntrain_set = tf.data.TFRecordDataset(os.path.join(outpath,\"sartorius.tfrec\"), compression_type=\"GZIP\").map(deserialize_example)\n```\n\nThere are 606 images  with shape `(520,704,3)` and dtype `float32`. Masks are in shape `(520,704,n_instances)` and dtype `uint8`. \n\n## Data\n```\nds = train_set.take(1)\nfor image, label in ds:\n    print(image.shape)\n    print(label.shape)\n\nfig, axs = plt.subplots(1, 2,figsize=(20, 20))\naxs[0].imshow(image)\naxs[0].axis('off')\naxs[1].imshow(np.sum(label, axis=-1))\naxs[1].axis('off')\nplt.show()\n```\n![result](https://i.imgur.com/odZkChr.png)\nEnjoy!",
      "votes": null
    },
    {
      "id": "1566865",
      "postDate": "11/01/2021 10:33:22",
      "content": "<p>thx for sharing. however now TPU is rarely available on kaggle.</p>",
      "rawMarkdown": "thx for sharing. however now TPU is rarely available on kaggle.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1566865,
      "author_name": "dragonzhang",
      "author_url": "",
      "post_date": "11/01/2021 10:33:22",
      "content": "<p>thx for sharing. however now TPU is rarely available on kaggle.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1566800": "Converting train annotation data and preparing the masks is a complicated task and time-consuming. For faster processing, I have converted all training samples and their masks into one `tfrec` file. \n\nYou can find the dataset [here](https://www.kaggle.com/kavehshahhosseini/sartorius-cell-instance-segmentation-in-tfrecord). Also, [here](https://www.kaggle.com/kavehshahhosseini/sartorius-convert-images-and-masks-to-tfrecord) is the notebook that I converted the data with. \n\nAfter adding [this dataset](https://www.kaggle.com/kavehshahhosseini/sartorius-cell-instance-segmentation-in-tfrecord) to your notebook, you can easily load the data and create a dataset in TensorFlow or Keras like this:\n```\ndef deserialize_example(serialized_string):\n    image_feature_description = {\n        'image': tf.io.FixedLenFeature([], tf.string),\n        'label': tf.io.FixedLenFeature([], tf.string)\n    }\n    parsed_record = tf.io.parse_single_example(serialized_string, image_feature_description)\n    image = tf.reshape(tf.io.decode_raw(parsed_record['image'], tf.float32),(IMAGE_HEIGHT, IMAGE_WIDTH, 3))\n    label = tf.reshape(tf.io.decode_raw(parsed_record['label'], tf.uint8),(IMAGE_HEIGHT, IMAGE_WIDTH, -1))\n    return image, label\n```\nThen creating `TFRecordDataset`:\n```\ntrain_set = tf.data.TFRecordDataset(os.path.join(outpath,\"sartorius.tfrec\"), compression_type=\"GZIP\").map(deserialize_example)\n```\n\nThere are 606 images  with shape `(520,704,3)` and dtype `float32`. Masks are in shape `(520,704,n_instances)` and dtype `uint8`. \n\n## Data\n```\nds = train_set.take(1)\nfor image, label in ds:\n    print(image.shape)\n    print(label.shape)\n\nfig, axs = plt.subplots(1, 2,figsize=(20, 20))\naxs[0].imshow(image)\naxs[0].axis('off')\naxs[1].imshow(np.sum(label, axis=-1))\naxs[1].axis('off')\nplt.show()\n```\n![result](https://i.imgur.com/odZkChr.png)\nEnjoy!",
    "1566865": "thx for sharing. however now TPU is rarely available on kaggle."
  },
  "source": "meta"
}