{
  "id": 211711,
  "title": "TFRecords tensorflow Features",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/211711",
  "author_name": "",
  "post_date": "2021-01-16T09:40:37.531711600Z",
  "votes": 2,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I have little understanding of how tfrecords work. Basically, its a structure that stacks data in a binary format, be it an array, file or image. For storing images in particular, we mention a dictionary of features to store. If I understand this correctly so far (feel free to correct me in comments), it seems that in order to read into a tfrecords file, I require the feaures dictionary ? If so where do I get the features dictionary in the Cassava Leaf Disease Classification ?</p>",
  "messages": [
    {
      "id": "1155205",
      "postDate": "01/16/2021 09:40:37",
      "content": "<p>I have little understanding of how tfrecords work. Basically, its a structure that stacks data in a binary format, be it an array, file or image. For storing images in particular, we mention a dictionary of features to store. If I understand this correctly so far (feel free to correct me in comments), it seems that in order to read into a tfrecords file, I require the feaures dictionary ? If so where do I get the features dictionary in the Cassava Leaf Disease Classification ?</p>",
      "rawMarkdown": "I have little understanding of how tfrecords work. Basically, its a structure that stacks data in a binary format, be it an array, file or image. For storing images in particular, we mention a dictionary of features to store. If I understand this correctly so far (feel free to correct me in comments), it seems that in order to read into a tfrecords file, I require the feaures dictionary ? If so where do I get the features dictionary in the Cassava Leaf Disease Classification ?",
      "votes": null
    },
    {
      "id": "1155225",
      "postDate": "01/16/2021 10:01:25",
      "content": "<p>If you want to use the official comp TFRecords you can have a look at the starter notebook <a href=\"https://www.kaggle.com/jessemostipak/getting-started-tpus-cassava-leaf-disease\" target=\"_blank\">here</a>, otherwise <a href=\"https://www.tensorflow.org/tutorials/load_data/tfrecord\" target=\"_blank\">here</a> on the TF website you can find everything you need about reading TFRecords.</p>",
      "rawMarkdown": "If you want to use the official comp TFRecords you can have a look at the starter notebook [here](https://www.kaggle.com/jessemostipak/getting-started-tpus-cassava-leaf-disease), otherwise [here](https://www.tensorflow.org/tutorials/load_data/tfrecord) on the TF website you can find everything you need about reading TFRecords.",
      "votes": null
    },
    {
      "id": "1155283",
      "postDate": "01/16/2021 10:52:59",
      "content": "<p>Thank you for your reply. I did read both. In the starter notebook they used image and label as the features whereas in the TF website they used the dimensions of the image, the image itself and the annotation for the image as its features. My question is how do I know the features dictionary that the dataset uses ? Is there any standard dictionary or is it made available by the dataset itself ? If a different dataset uses a different features dictionary, how will I know of it ? </p>",
      "rawMarkdown": "Thank you for your reply. I did read both. In the starter notebook they used image and label as the features whereas in the TF website they used the dimensions of the image, the image itself and the annotation for the image as its features. My question is how do I know the features dictionary that the dataset uses ? Is there any standard dictionary or is it made available by the dataset itself ? If a different dataset uses a different features dictionary, how will I know of it ?",
      "votes": null
    },
    {
      "id": "1155573",
      "postDate": "01/16/2021 13:22:09",
      "content": "<p>You can inspect a TFRecord with this code:</p>\n<pre><code>filenames = [filename]\nraw_dataset = tf.data.TFRecordDataset(filenames)\nfor raw_record in raw_dataset.take(1):\n    example = tf.train.Example()\n    example.ParseFromString(raw_record.numpy())\n    print(example)\n</code></pre>\n<p>You can probably do better and show only the keys of the dictionary, but honestly I don't know how to do it.</p>",
      "rawMarkdown": "You can inspect a TFRecord with this code:\n```\nfilenames = [filename]\nraw_dataset = tf.data.TFRecordDataset(filenames)\nfor raw_record in raw_dataset.take(1):\n    example = tf.train.Example()\n    example.ParseFromString(raw_record.numpy())\n    print(example)\n```\nYou can probably do better and show only the keys of the dictionary, but honestly I don't know how to do it.",
      "votes": null
    },
    {
      "id": "1156273",
      "postDate": "01/17/2021 04:05:52",
      "content": "<p>Thanks ! This is just what I needed. I probed further and found how to get just keys of the dictionary. Just convert the <code>example.features.feature</code> into a list and you'll get the keys of the dict.</p>\n<pre><code>filenames = [filename]\nraw_dataset = tf.data.TFRecordDataset(filenames)\nfor raw_record in raw_dataset.take(1):\n    example = tf.train.Example()\n    example.ParseFromString(raw_record.numpy())\n    keysList = list(example.features.feature)\n    print(keysList)\n</code></pre>",
      "rawMarkdown": "Thanks ! This is just what I needed. I probed further and found how to get just keys of the dictionary. Just convert the `example.features.feature` into a list and you'll get the keys of the dict.\n```\nfilenames = [filename]\nraw_dataset = tf.data.TFRecordDataset(filenames)\nfor raw_record in raw_dataset.take(1):\n    example = tf.train.Example()\n    example.ParseFromString(raw_record.numpy())\n    keysList = list(example.features.feature)\n    print(keysList)\n```",
      "votes": null
    },
    {
      "id": "1156391",
      "postDate": "01/17/2021 06:06:21",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1155225,
      "author_name": "mviola",
      "author_url": "",
      "post_date": "01/16/2021 10:01:25",
      "content": "<p>If you want to use the official comp TFRecords you can have a look at the starter notebook <a href=\"https://www.kaggle.com/jessemostipak/getting-started-tpus-cassava-leaf-disease\" target=\"_blank\">here</a>, otherwise <a href=\"https://www.tensorflow.org/tutorials/load_data/tfrecord\" target=\"_blank\">here</a> on the TF website you can find everything you need about reading TFRecords.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1155283,
          "author_name": "siddalore",
          "author_url": "",
          "post_date": "01/16/2021 10:52:59",
          "content": "<p>Thank you for your reply. I did read both. In the starter notebook they used image and label as the features whereas in the TF website they used the dimensions of the image, the image itself and the annotation for the image as its features. My question is how do I know the features dictionary that the dataset uses ? Is there any standard dictionary or is it made available by the dataset itself ? If a different dataset uses a different features dictionary, how will I know of it ? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1155573,
          "author_name": "mviola",
          "author_url": "",
          "post_date": "01/16/2021 13:22:09",
          "content": "<p>You can inspect a TFRecord with this code:</p>\n<pre><code>filenames = [filename]\nraw_dataset = tf.data.TFRecordDataset(filenames)\nfor raw_record in raw_dataset.take(1):\n    example = tf.train.Example()\n    example.ParseFromString(raw_record.numpy())\n    print(example)\n</code></pre>\n<p>You can probably do better and show only the keys of the dictionary, but honestly I don't know how to do it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1156273,
          "author_name": "siddalore",
          "author_url": "",
          "post_date": "01/17/2021 04:05:52",
          "content": "<p>Thanks ! This is just what I needed. I probed further and found how to get just keys of the dictionary. Just convert the <code>example.features.feature</code> into a list and you'll get the keys of the dict.</p>\n<pre><code>filenames = [filename]\nraw_dataset = tf.data.TFRecordDataset(filenames)\nfor raw_record in raw_dataset.take(1):\n    example = tf.train.Example()\n    example.ParseFromString(raw_record.numpy())\n    keysList = list(example.features.feature)\n    print(keysList)\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1156391,
          "author_name": "saurabhbagchi",
          "author_url": "",
          "post_date": "01/17/2021 06:06:21",
          "content": "<p>Thanks for sharing!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1155205": "I have little understanding of how tfrecords work. Basically, its a structure that stacks data in a binary format, be it an array, file or image. For storing images in particular, we mention a dictionary of features to store. If I understand this correctly so far (feel free to correct me in comments), it seems that in order to read into a tfrecords file, I require the feaures dictionary ? If so where do I get the features dictionary in the Cassava Leaf Disease Classification ?",
    "1155225": "If you want to use the official comp TFRecords you can have a look at the starter notebook [here](https://www.kaggle.com/jessemostipak/getting-started-tpus-cassava-leaf-disease), otherwise [here](https://www.tensorflow.org/tutorials/load_data/tfrecord) on the TF website you can find everything you need about reading TFRecords.",
    "1155283": "Thank you for your reply. I did read both. In the starter notebook they used image and label as the features whereas in the TF website they used the dimensions of the image, the image itself and the annotation for the image as its features. My question is how do I know the features dictionary that the dataset uses ? Is there any standard dictionary or is it made available by the dataset itself ? If a different dataset uses a different features dictionary, how will I know of it ?",
    "1155573": "You can inspect a TFRecord with this code:\n```\nfilenames = [filename]\nraw_dataset = tf.data.TFRecordDataset(filenames)\nfor raw_record in raw_dataset.take(1):\n    example = tf.train.Example()\n    example.ParseFromString(raw_record.numpy())\n    print(example)\n```\nYou can probably do better and show only the keys of the dictionary, but honestly I don't know how to do it.",
    "1156273": "Thanks ! This is just what I needed. I probed further and found how to get just keys of the dictionary. Just convert the `example.features.feature` into a list and you'll get the keys of the dict.\n```\nfilenames = [filename]\nraw_dataset = tf.data.TFRecordDataset(filenames)\nfor raw_record in raw_dataset.take(1):\n    example = tf.train.Example()\n    example.ParseFromString(raw_record.numpy())\n    keysList = list(example.features.feature)\n    print(keysList)\n```",
    "1156391": "Thanks for sharing!"
  },
  "source": "meta"
}