{
  "id": 126974,
  "title": "[TF2.0 Question] Another way to feed data instead of TFRecords",
  "url": "/competitions/tensorflow2-question-answering/discussion/126974",
  "author_name": "",
  "post_date": "2020-01-21T13:57:19.456028Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Please tell me if you know about this topic!</p>\n\n<p>I (and some kagglers maybe) like pytorch-like dataloader, which make custom class by implementing <code>torch.utils.data.Dataset</code> and <code>torch.utils.data.DataLoader</code>, then one can check the input one by one or on the way of training.</p>\n\n<p>In many TF notebooks, however, TFRecord and tf.Example are used. And this is bit difficult for me for the first look. (<a href=\"https://www.tensorflow.org/tutorials/load_data/tfrecord?hl=en\">Here</a> is good for understanding.)\nI found an api <a href=\"https://www.tensorflow.org/api_docs/python/tf/data/Dataset#from_generator\"><code>from_generator</code> in <code>tf.data.Dataset</code></a>, which allows me to implement pytorch-like dataloader by defining generator.\nI write <a href=\"https://www.kaggle.com/kentaronakanishi/my-custom-dataloader-simplified\">my custom dataloader (sorry for simplified example)</a> to try this and use it for inference.</p>\n\n<p>But now I don't clearly understand the advantage of using generator way or using TFRecords. In my opinion:</p>\n\n<ul>\n<li>using <code>from_generator</code> case\n<ul><li>no need to pre-compile</li>\n<li>able to use object-base dataloader</li>\n<li>memory efficiency (since this is based on generator) (edited)</li>\n<li>cannot shuffle dataset (edited)</li></ul></li>\n<li>using TFRecords case\n<ul><li>need pre-compile </li>\n<li>maybe faster if there are many loops</li>\n<li>can shuffle </li></ul></li>\n</ul>\n\n<p>Do you try both ways to feed data, or do you know other ways? Is there any other differences?</p>",
  "messages": [
    {
      "id": "724810",
      "postDate": "01/21/2020 13:57:19",
      "content": "<p>Please tell me if you know about this topic!</p>\n\n<p>I (and some kagglers maybe) like pytorch-like dataloader, which make custom class by implementing <code>torch.utils.data.Dataset</code> and <code>torch.utils.data.DataLoader</code>, then one can check the input one by one or on the way of training.</p>\n\n<p>In many TF notebooks, however, TFRecord and tf.Example are used. And this is bit difficult for me for the first look. (<a href=\"https://www.tensorflow.org/tutorials/load_data/tfrecord?hl=en\">Here</a> is good for understanding.)\nI found an api <a href=\"https://www.tensorflow.org/api_docs/python/tf/data/Dataset#from_generator\"><code>from_generator</code> in <code>tf.data.Dataset</code></a>, which allows me to implement pytorch-like dataloader by defining generator.\nI write <a href=\"https://www.kaggle.com/kentaronakanishi/my-custom-dataloader-simplified\">my custom dataloader (sorry for simplified example)</a> to try this and use it for inference.</p>\n\n<p>But now I don't clearly understand the advantage of using generator way or using TFRecords. In my opinion:</p>\n\n<ul>\n<li>using <code>from_generator</code> case\n<ul><li>no need to pre-compile</li>\n<li>able to use object-base dataloader</li>\n<li>memory efficiency (since this is based on generator) (edited)</li>\n<li>cannot shuffle dataset (edited)</li></ul></li>\n<li>using TFRecords case\n<ul><li>need pre-compile </li>\n<li>maybe faster if there are many loops</li>\n<li>can shuffle </li></ul></li>\n</ul>\n\n<p>Do you try both ways to feed data, or do you know other ways? Is there any other differences?</p>",
      "rawMarkdown": "Please tell me if you know about this topic!\n\nI (and some kagglers maybe) like pytorch-like dataloader, which make custom class by implementing `torch.utils.data.Dataset` and `torch.utils.data.DataLoader`, then one can check the input one by one or on the way of training.\n\nIn many TF notebooks, however, TFRecord and tf.Example are used. And this is bit difficult for me for the first look. ([Here](https://www.tensorflow.org/tutorials/load_data/tfrecord?hl=en) is good for understanding.)\nI found an api [`from_generator` in `tf.data.Dataset`](https://www.tensorflow.org/api_docs/python/tf/data/Dataset#from_generator), which allows me to implement pytorch-like dataloader by defining generator.\nI write [my custom dataloader (sorry for simplified example)](https://www.kaggle.com/kentaronakanishi/my-custom-dataloader-simplified) to try this and use it for inference.\n\nBut now I don't clearly understand the advantage of using generator way or using TFRecords. In my opinion:\n\n- using `from_generator` case\n  - no need to pre-compile\n  - able to use object-base dataloader\n  - memory efficiency (since this is based on generator) (edited)\n  - cannot shuffle dataset (edited)\n- using TFRecords case\n  - need pre-compile \n  - maybe faster if there are many loops\n  - can shuffle \n\nDo you try both ways to feed data, or do you know other ways? Is there any other differences?",
      "votes": null
    },
    {
      "id": "724820",
      "postDate": "01/21/2020 14:14:00",
      "content": "<p>One other way would be using <code>from_tensor_slices</code> (<a href=\"https://www.tensorflow.org/api_docs/python/tf/data/Dataset#from_tensor_slices\">https://www.tensorflow.org/api_docs/python/tf/data/Dataset#from_tensor_slices</a>). The drawback is that it only works for small datasets (it says &lt;1GB, I am not sure this is still true). I don't like writing encoders and decoders for TFRecords either :P</p>\n\n<p><code>from_generator</code> does not work on TPUs but TFRecords do. According to <a href=\"https://github.com/tensorflow/tensorflow/issues/34346\">this issue</a> generator-like inputs for the TPU may be in the works.</p>",
      "rawMarkdown": "One other way would be using `from_tensor_slices` (https://www.tensorflow.org/api_docs/python/tf/data/Dataset#from_tensor_slices). The drawback is that it only works for small datasets (it says &lt;1GB, I am not sure this is still true). I don't like writing encoders and decoders for TFRecords either :P\n\n`from_generator` does not work on TPUs but TFRecords do. According to [this issue](https://github.com/tensorflow/tensorflow/issues/34346) generator-like inputs for the TPU may be in the works.",
      "votes": null
    },
    {
      "id": "724828",
      "postDate": "01/21/2020 14:22:54",
      "content": "<p>Yes, <code>from_tensor_slices</code> is another way but It's allowed in small dataset.\nOne more merit for <code>from_generator</code> is memory efficiency since it's based on generator.\nI didn't know about TPU compatibility, maybe that's why many notebooks use TFRecords. Thanks for information!</p>",
      "rawMarkdown": "Yes, `from_tensor_slices` is another way but It's allowed in small dataset.\nOne more merit for `from_generator` is memory efficiency since it's based on generator.\nI didn't know about TPU compatibility, maybe that's why many notebooks use TFRecords. Thanks for information!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 724820,
      "author_name": "seesee",
      "author_url": "",
      "post_date": "01/21/2020 14:14:00",
      "content": "<p>One other way would be using <code>from_tensor_slices</code> (<a href=\"https://www.tensorflow.org/api_docs/python/tf/data/Dataset#from_tensor_slices\">https://www.tensorflow.org/api_docs/python/tf/data/Dataset#from_tensor_slices</a>). The drawback is that it only works for small datasets (it says &lt;1GB, I am not sure this is still true). I don't like writing encoders and decoders for TFRecords either :P</p>\n\n<p><code>from_generator</code> does not work on TPUs but TFRecords do. According to <a href=\"https://github.com/tensorflow/tensorflow/issues/34346\">this issue</a> generator-like inputs for the TPU may be in the works.</p>",
      "votes": null,
      "replies": [
        {
          "id": 724828,
          "author_name": "kentaronakanishi",
          "author_url": "",
          "post_date": "01/21/2020 14:22:54",
          "content": "<p>Yes, <code>from_tensor_slices</code> is another way but It's allowed in small dataset.\nOne more merit for <code>from_generator</code> is memory efficiency since it's based on generator.\nI didn't know about TPU compatibility, maybe that's why many notebooks use TFRecords. Thanks for information!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "724810": "Please tell me if you know about this topic!\n\nI (and some kagglers maybe) like pytorch-like dataloader, which make custom class by implementing `torch.utils.data.Dataset` and `torch.utils.data.DataLoader`, then one can check the input one by one or on the way of training.\n\nIn many TF notebooks, however, TFRecord and tf.Example are used. And this is bit difficult for me for the first look. ([Here](https://www.tensorflow.org/tutorials/load_data/tfrecord?hl=en) is good for understanding.)\nI found an api [`from_generator` in `tf.data.Dataset`](https://www.tensorflow.org/api_docs/python/tf/data/Dataset#from_generator), which allows me to implement pytorch-like dataloader by defining generator.\nI write [my custom dataloader (sorry for simplified example)](https://www.kaggle.com/kentaronakanishi/my-custom-dataloader-simplified) to try this and use it for inference.\n\nBut now I don't clearly understand the advantage of using generator way or using TFRecords. In my opinion:\n\n- using `from_generator` case\n  - no need to pre-compile\n  - able to use object-base dataloader\n  - memory efficiency (since this is based on generator) (edited)\n  - cannot shuffle dataset (edited)\n- using TFRecords case\n  - need pre-compile \n  - maybe faster if there are many loops\n  - can shuffle \n\nDo you try both ways to feed data, or do you know other ways? Is there any other differences?",
    "724820": "One other way would be using `from_tensor_slices` (https://www.tensorflow.org/api_docs/python/tf/data/Dataset#from_tensor_slices). The drawback is that it only works for small datasets (it says &lt;1GB, I am not sure this is still true). I don't like writing encoders and decoders for TFRecords either :P\n\n`from_generator` does not work on TPUs but TFRecords do. According to [this issue](https://github.com/tensorflow/tensorflow/issues/34346) generator-like inputs for the TPU may be in the works.",
    "724828": "Yes, `from_tensor_slices` is another way but It's allowed in small dataset.\nOne more merit for `from_generator` is memory efficiency since it's based on generator.\nI didn't know about TPU compatibility, maybe that's why many notebooks use TFRecords. Thanks for information!"
  },
  "source": "meta"
}