{
  "id": 315397,
  "title": "Get value when I use flat_map in tfrecords to repeat records .",
  "url": "/competitions/happy-whale-and-dolphin/discussion/315397",
  "author_name": "k_tomo",
  "post_date": "2022-03-28T04:07:08.246000",
  "votes": 4,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I want to repeat some records in dataset . And to repeat how many times, I'll use posting_id(name of imege).<br>\nTherefore,I wrote code like this. However, I completely don't know how to  access to property <br>\nin flat_map. Do you know how to access? Or do you know better way to repeat dataset?<br>\nIf you know, please help me！！ </p>\n<p>↓↓</p>\n<pre><code>dataset = tf.data.TFRecordDataset(filenames, num_parallel_reads = AUTO)\n\nignore_order = tf.data.Options()\n\ndataset = dataset.with_options(ignore_order)\n\ndef read_labeled_tfrecord(example):\n    LABELED_TFREC_FORMAT = {\n        \"image_name\": tf.io.FixedLenFeature([], tf.string),\n        \"image\": tf.io.FixedLenFeature([], tf.string),\n        \"target\": tf.io.FixedLenFeature([], tf.int64),\n#         \"matches\": tf.io.FixedLenFeature([], tf.string)\n    }\n\n    example = tf.io.parse_single_example(example, LABELED_TFREC_FORMAT)\n    posting_id = example['image_name']\n    image = example['image']\n#     label_group = tf.one_hot(tf.cast(example['label_group'], tf.int32), depth = N_CLASSES)\n    label_group = tf.cast(example['target'], tf.int32)\n#     matches = example['matches']\n    matches = 1\n    return posting_id, image, label_group, matches\n\n# sample some part of dataset\ndataset = dataset.take(12).map(read_labeled_tfrecord, num_parallel_calls = AUTO) \n\n\ndef flat_map_impl(posting_id, image, label_group, matches):\n    print(posting_id, image, label_group, matches)\n\n    # I want to know here\n    posting_id_string =  ???\n\n\n    ## function to convert posting_id to repeat_count\n    repeat_id_count = f(posting_id_string)\n\n\n    return tf.data.Dataset.from_tensors(tf_example).repeat(repeat_id_count)\n\ndataset = dataset.flat_map(flat_map_impl)\n</code></pre>",
  "messages": [
    {
      "id": 1737033,
      "postDate": "2022-03-28T04:07:08.247Z",
      "content": "<p>I want to repeat some records in dataset . And to repeat how many times, I'll use posting_id(name of imege).<br>\nTherefore,I wrote code like this. However, I completely don't know how to  access to property <br>\nin flat_map. Do you know how to access? Or do you know better way to repeat dataset?<br>\nIf you know, please help me！！ </p>\n<p>↓↓</p>\n<pre><code>dataset = tf.data.TFRecordDataset(filenames, num_parallel_reads = AUTO)\n\nignore_order = tf.data.Options()\n\ndataset = dataset.with_options(ignore_order)\n\ndef read_labeled_tfrecord(example):\n    LABELED_TFREC_FORMAT = {\n        \"image_name\": tf.io.FixedLenFeature([], tf.string),\n        \"image\": tf.io.FixedLenFeature([], tf.string),\n        \"target\": tf.io.FixedLenFeature([], tf.int64),\n#         \"matches\": tf.io.FixedLenFeature([], tf.string)\n    }\n\n    example = tf.io.parse_single_example(example, LABELED_TFREC_FORMAT)\n    posting_id = example['image_name']\n    image = example['image']\n#     label_group = tf.one_hot(tf.cast(example['label_group'], tf.int32), depth = N_CLASSES)\n    label_group = tf.cast(example['target'], tf.int32)\n#     matches = example['matches']\n    matches = 1\n    return posting_id, image, label_group, matches\n\n# sample some part of dataset\ndataset = dataset.take(12).map(read_labeled_tfrecord, num_parallel_calls = AUTO) \n\n\ndef flat_map_impl(posting_id, image, label_group, matches):\n    print(posting_id, image, label_group, matches)\n\n    # I want to know here\n    posting_id_string =  ???\n\n\n    ## function to convert posting_id to repeat_count\n    repeat_id_count = f(posting_id_string)\n\n\n    return tf.data.Dataset.from_tensors(tf_example).repeat(repeat_id_count)\n\ndataset = dataset.flat_map(flat_map_impl)\n</code></pre>",
      "rawMarkdown": "I want to repeat some records in dataset . And to repeat how many times, I'll use posting_id(name of imege).\nTherefore,I wrote code like this. However, I completely don't know how to  access to property \nin flat_map. Do you know how to access? Or do you know better way to repeat dataset?\nIf you know, please help me！！ \n\n↓↓\n\n```\ndataset = tf.data.TFRecordDataset(filenames, num_parallel_reads = AUTO)\n\nignore_order = tf.data.Options()\n\ndataset = dataset.with_options(ignore_order)\n\ndef read_labeled_tfrecord(example):\n    LABELED_TFREC_FORMAT = {\n        \"image_name\": tf.io.FixedLenFeature([], tf.string),\n        \"image\": tf.io.FixedLenFeature([], tf.string),\n        \"target\": tf.io.FixedLenFeature([], tf.int64),\n#         \"matches\": tf.io.FixedLenFeature([], tf.string)\n    }\n\n    example = tf.io.parse_single_example(example, LABELED_TFREC_FORMAT)\n    posting_id = example['image_name']\n    image = example['image']\n#     label_group = tf.one_hot(tf.cast(example['label_group'], tf.int32), depth = N_CLASSES)\n    label_group = tf.cast(example['target'], tf.int32)\n#     matches = example['matches']\n    matches = 1\n    return posting_id, image, label_group, matches\n\n# sample some part of dataset\ndataset = dataset.take(12).map(read_labeled_tfrecord, num_parallel_calls = AUTO) \n\n\ndef flat_map_impl(posting_id, image, label_group, matches):\n    print(posting_id, image, label_group, matches)\n  \n    # I want to know here\n    posting_id_string =  ???\n    \n    \n    ## function to convert posting_id to repeat_count\n    repeat_id_count = f(posting_id_string)\n\n\n    return tf.data.Dataset.from_tensors(tf_example).repeat(repeat_id_count)\n\ndataset = dataset.flat_map(flat_map_impl)\n```",
      "votes": 4
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1737033": "I want to repeat some records in dataset . And to repeat how many times, I'll use posting_id(name of imege).\nTherefore,I wrote code like this. However, I completely don't know how to  access to property \nin flat_map. Do you know how to access? Or do you know better way to repeat dataset?\nIf you know, please help me！！ \n\n↓↓\n\n```\ndataset = tf.data.TFRecordDataset(filenames, num_parallel_reads = AUTO)\n\nignore_order = tf.data.Options()\n\ndataset = dataset.with_options(ignore_order)\n\ndef read_labeled_tfrecord(example):\n    LABELED_TFREC_FORMAT = {\n        \"image_name\": tf.io.FixedLenFeature([], tf.string),\n        \"image\": tf.io.FixedLenFeature([], tf.string),\n        \"target\": tf.io.FixedLenFeature([], tf.int64),\n#         \"matches\": tf.io.FixedLenFeature([], tf.string)\n    }\n\n    example = tf.io.parse_single_example(example, LABELED_TFREC_FORMAT)\n    posting_id = example['image_name']\n    image = example['image']\n#     label_group = tf.one_hot(tf.cast(example['label_group'], tf.int32), depth = N_CLASSES)\n    label_group = tf.cast(example['target'], tf.int32)\n#     matches = example['matches']\n    matches = 1\n    return posting_id, image, label_group, matches\n\n# sample some part of dataset\ndataset = dataset.take(12).map(read_labeled_tfrecord, num_parallel_calls = AUTO) \n\n\ndef flat_map_impl(posting_id, image, label_group, matches):\n    print(posting_id, image, label_group, matches)\n  \n    # I want to know here\n    posting_id_string =  ???\n    \n    \n    ## function to convert posting_id to repeat_count\n    repeat_id_count = f(posting_id_string)\n\n\n    return tf.data.Dataset.from_tensors(tf_example).repeat(repeat_id_count)\n\ndataset = dataset.flat_map(flat_map_impl)\n```"
  }
}