{
  "id": 272605,
  "title": "plot all labels of TFRecords",
  "url": "/competitions/tpu-getting-started/discussion/272605",
  "author_name": "michael scheinfeild",
  "post_date": "2021-09-16T11:27:57.953000",
  "votes": 0,
  "comment_count": 1,
  "views": null,
  "content": "<p><strong>hi i wanted to run on all files of TFRecords and plot in one histogram all labels</strong>.<br>\ni used current loop :</p>\n<p>all_label = []<br>\nfor image, label in ds_train.take(10):<br>\n    all_label.append(label)</p>\n<p>and  sns.distplot(all_label)</p>\n<p>it plots the distribution but first i need to loop on 10 records and i have image which is not needed for me.<br>\ni try to write only label extraction :</p>\n<pre><code>def **read_only_labeled_tfrecord**(example):\n    LABELED_TFREC_FORMAT = {\n        \"image\": tf.io.FixedLenFeature([], tf.string), # tf.string means bytestring\n        \"class\": tf.io.FixedLenFeature([], tf.int64),  # shape [] means single element\n    }\n    example = tf.io.parse_single_example(example, LABELED_TFREC_FORMAT)\n    #image = decode_image(example['image'])\n    label = tf.cast(example['class'], tf.int32)\n    return  label # returns a   label from (image, label) pairs\n\n\ndef **load_dataset_labels**(filenames, labeled=True, ordered=False):\n    # Read from TFRecords. For optimal performance, reading from multiple files at once and\n    # disregarding data order. Order does not matter since we will be shuffling the data anyway.\n    # this reads only labels\n\n    ignore_order = tf.data.Options()\n    if not ordered:\n        ignore_order.experimental_deterministic = False # disable order, increase speed\n\n    dataset = tf.data.TFRecordDataset(filenames, num_parallel_reads=AUTO) # automatically interleaves reads from multiple files\n    dataset = dataset.with_options(ignore_order) # uses data as soon as it streams in, rather than in its original order\n    labels_set = dataset.map(read_only_labeled_tfrecord, num_parallel_calls=AUTO)\n    # returns a dataset of (image, label) pairs if labeled=True or (image, id) pairs if labeled=False\n    return dataset\n\n\ndef **get_training_dataset_labels**():\n    # return only labels\n    dataset = load_dataset_labels(TRAINING_FILENAMES, labeled=True)\n    #dataset = dataset.map(data_augment, num_parallel_calls=AUTO)\n    dataset = dataset.repeat() # the training dataset must repeat for several epochs\n    #dataset = dataset.shuffle(2048)\n    #dataset = dataset.batch(BATCH_SIZE)\n    dataset = dataset.prefetch(AUTO) # prefetch next batch while training (autotune prefetch buffer size)\n    return dataset\n</code></pre>\n<pre><code>ds_train_labels = get_training_dataset_labels()\n</code></pre>\n<p>so i wanted to plot all labels use ds_train_labels   but got complicated with data set type </p>",
  "messages": [
    {
      "id": 1524548,
      "postDate": "2021-09-26T16:32:56.663Z",
      "content": "<p>Hi, if you mean to to get the list of all the labels in the dataset, you can use this line:<br>\nnext(iter(ds_train.unbatch().map(lambda image, label: label).batch(NUM_TRAINING_IMAGES))).numpy() <br>\nwith NUM_TRAINING_IMAGES the number of images in your training set</p>",
      "rawMarkdown": "Hi, if you mean to to get the list of all the labels in the dataset, you can use this line:\nnext(iter(ds_train.unbatch().map(lambda image, label: label).batch(NUM_TRAINING_IMAGES))).numpy() \nwith NUM_TRAINING_IMAGES the number of images in your training set",
      "votes": 2
    },
    {
      "id": 1514714,
      "postDate": "2021-09-16T11:27:57.953Z",
      "content": "<p><strong>hi i wanted to run on all files of TFRecords and plot in one histogram all labels</strong>.<br>\ni used current loop :</p>\n<p>all_label = []<br>\nfor image, label in ds_train.take(10):<br>\n    all_label.append(label)</p>\n<p>and  sns.distplot(all_label)</p>\n<p>it plots the distribution but first i need to loop on 10 records and i have image which is not needed for me.<br>\ni try to write only label extraction :</p>\n<pre><code>def **read_only_labeled_tfrecord**(example):\n    LABELED_TFREC_FORMAT = {\n        \"image\": tf.io.FixedLenFeature([], tf.string), # tf.string means bytestring\n        \"class\": tf.io.FixedLenFeature([], tf.int64),  # shape [] means single element\n    }\n    example = tf.io.parse_single_example(example, LABELED_TFREC_FORMAT)\n    #image = decode_image(example['image'])\n    label = tf.cast(example['class'], tf.int32)\n    return  label # returns a   label from (image, label) pairs\n\n\ndef **load_dataset_labels**(filenames, labeled=True, ordered=False):\n    # Read from TFRecords. For optimal performance, reading from multiple files at once and\n    # disregarding data order. Order does not matter since we will be shuffling the data anyway.\n    # this reads only labels\n\n    ignore_order = tf.data.Options()\n    if not ordered:\n        ignore_order.experimental_deterministic = False # disable order, increase speed\n\n    dataset = tf.data.TFRecordDataset(filenames, num_parallel_reads=AUTO) # automatically interleaves reads from multiple files\n    dataset = dataset.with_options(ignore_order) # uses data as soon as it streams in, rather than in its original order\n    labels_set = dataset.map(read_only_labeled_tfrecord, num_parallel_calls=AUTO)\n    # returns a dataset of (image, label) pairs if labeled=True or (image, id) pairs if labeled=False\n    return dataset\n\n\ndef **get_training_dataset_labels**():\n    # return only labels\n    dataset = load_dataset_labels(TRAINING_FILENAMES, labeled=True)\n    #dataset = dataset.map(data_augment, num_parallel_calls=AUTO)\n    dataset = dataset.repeat() # the training dataset must repeat for several epochs\n    #dataset = dataset.shuffle(2048)\n    #dataset = dataset.batch(BATCH_SIZE)\n    dataset = dataset.prefetch(AUTO) # prefetch next batch while training (autotune prefetch buffer size)\n    return dataset\n</code></pre>\n<pre><code>ds_train_labels = get_training_dataset_labels()\n</code></pre>\n<p>so i wanted to plot all labels use ds_train_labels   but got complicated with data set type </p>",
      "rawMarkdown": "**hi i wanted to run on all files of TFRecords and plot in one histogram all labels**.\ni used current loop :\n\nall_label = []\nfor image, label in ds_train.take(10):\n    all_label.append(label)\n\nand  sns.distplot(all_label)\n\nit plots the distribution but first i need to loop on 10 records and i have image which is not needed for me.\ni try to write only label extraction :\n\n```\ndef **read_only_labeled_tfrecord**(example):\n    LABELED_TFREC_FORMAT = {\n        \"image\": tf.io.FixedLenFeature([], tf.string), # tf.string means bytestring\n        \"class\": tf.io.FixedLenFeature([], tf.int64),  # shape [] means single element\n    }\n    example = tf.io.parse_single_example(example, LABELED_TFREC_FORMAT)\n    #image = decode_image(example['image'])\n    label = tf.cast(example['class'], tf.int32)\n    return  label # returns a   label from (image, label) pairs\n\n\ndef **load_dataset_labels**(filenames, labeled=True, ordered=False):\n    # Read from TFRecords. For optimal performance, reading from multiple files at once and\n    # disregarding data order. Order does not matter since we will be shuffling the data anyway.\n    # this reads only labels\n\n    ignore_order = tf.data.Options()\n    if not ordered:\n        ignore_order.experimental_deterministic = False # disable order, increase speed\n\n    dataset = tf.data.TFRecordDataset(filenames, num_parallel_reads=AUTO) # automatically interleaves reads from multiple files\n    dataset = dataset.with_options(ignore_order) # uses data as soon as it streams in, rather than in its original order\n    labels_set = dataset.map(read_only_labeled_tfrecord, num_parallel_calls=AUTO)\n    # returns a dataset of (image, label) pairs if labeled=True or (image, id) pairs if labeled=False\n    return dataset\n\n\ndef **get_training_dataset_labels**():\n    # return only labels\n    dataset = load_dataset_labels(TRAINING_FILENAMES, labeled=True)\n    #dataset = dataset.map(data_augment, num_parallel_calls=AUTO)\n    dataset = dataset.repeat() # the training dataset must repeat for several epochs\n    #dataset = dataset.shuffle(2048)\n    #dataset = dataset.batch(BATCH_SIZE)\n    dataset = dataset.prefetch(AUTO) # prefetch next batch while training (autotune prefetch buffer size)\n    return dataset\n```\n```\n\nds_train_labels = get_training_dataset_labels()\n```\n\nso i wanted to plot all labels use ds_train_labels   but got complicated with data set type "
    }
  ],
  "comments": [
    {
      "id": 1524548,
      "author_name": "rehaqds",
      "author_url": "",
      "post_date": "2021-09-26T16:32:56.663000",
      "content": "<p>Hi, if you mean to to get the list of all the labels in the dataset, you can use this line:<br>\nnext(iter(ds_train.unbatch().map(lambda image, label: label).batch(NUM_TRAINING_IMAGES))).numpy() <br>\nwith NUM_TRAINING_IMAGES the number of images in your training set</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1524548": "Hi, if you mean to to get the list of all the labels in the dataset, you can use this line:\nnext(iter(ds_train.unbatch().map(lambda image, label: label).batch(NUM_TRAINING_IMAGES))).numpy() \nwith NUM_TRAINING_IMAGES the number of images in your training set",
    "1514714": "**hi i wanted to run on all files of TFRecords and plot in one histogram all labels**.\ni used current loop :\n\nall_label = []\nfor image, label in ds_train.take(10):\n    all_label.append(label)\n\nand  sns.distplot(all_label)\n\nit plots the distribution but first i need to loop on 10 records and i have image which is not needed for me.\ni try to write only label extraction :\n\n```\ndef **read_only_labeled_tfrecord**(example):\n    LABELED_TFREC_FORMAT = {\n        \"image\": tf.io.FixedLenFeature([], tf.string), # tf.string means bytestring\n        \"class\": tf.io.FixedLenFeature([], tf.int64),  # shape [] means single element\n    }\n    example = tf.io.parse_single_example(example, LABELED_TFREC_FORMAT)\n    #image = decode_image(example['image'])\n    label = tf.cast(example['class'], tf.int32)\n    return  label # returns a   label from (image, label) pairs\n\n\ndef **load_dataset_labels**(filenames, labeled=True, ordered=False):\n    # Read from TFRecords. For optimal performance, reading from multiple files at once and\n    # disregarding data order. Order does not matter since we will be shuffling the data anyway.\n    # this reads only labels\n\n    ignore_order = tf.data.Options()\n    if not ordered:\n        ignore_order.experimental_deterministic = False # disable order, increase speed\n\n    dataset = tf.data.TFRecordDataset(filenames, num_parallel_reads=AUTO) # automatically interleaves reads from multiple files\n    dataset = dataset.with_options(ignore_order) # uses data as soon as it streams in, rather than in its original order\n    labels_set = dataset.map(read_only_labeled_tfrecord, num_parallel_calls=AUTO)\n    # returns a dataset of (image, label) pairs if labeled=True or (image, id) pairs if labeled=False\n    return dataset\n\n\ndef **get_training_dataset_labels**():\n    # return only labels\n    dataset = load_dataset_labels(TRAINING_FILENAMES, labeled=True)\n    #dataset = dataset.map(data_augment, num_parallel_calls=AUTO)\n    dataset = dataset.repeat() # the training dataset must repeat for several epochs\n    #dataset = dataset.shuffle(2048)\n    #dataset = dataset.batch(BATCH_SIZE)\n    dataset = dataset.prefetch(AUTO) # prefetch next batch while training (autotune prefetch buffer size)\n    return dataset\n```\n```\n\nds_train_labels = get_training_dataset_labels()\n```\n\nso i wanted to plot all labels use ds_train_labels   but got complicated with data set type "
  }
}