{
  "id": 154347,
  "title": "Use Of TFrecords",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/154347",
  "author_name": "",
  "post_date": "2020-05-28T04:37:26.777639500Z",
  "votes": 10,
  "comment_count": 10,
  "views": 0,
  "content": "<p>What is the use of TFrecords? I read from the documentation from the TensorFlow, still am in confusion. Answer along with a notebook will be more helpful</p>",
  "messages": [
    {
      "id": "864580",
      "postDate": "05/28/2020 04:37:26",
      "content": "<p>What is the use of TFrecords? I read from the documentation from the TensorFlow, still am in confusion. Answer along with a notebook will be more helpful</p>",
      "rawMarkdown": "What is the use of TFrecords? I read from the documentation from the TensorFlow, still am in confusion. Answer along with a notebook will be more helpful",
      "votes": null
    },
    {
      "id": "864588",
      "postDate": "05/28/2020 04:48:03",
      "content": "<p>TFrecords are TensorFlow's format for save files to disk. They have advantages for fast loading when using data loaders to feed your model.</p>\n\n<p>This notebook <a href=\"https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu\">here</a> from Flower Comp shows how to work with them</p>",
      "rawMarkdown": "TFrecords are TensorFlow's format for save files to disk. They have advantages for fast loading when using data loaders to feed your model.\n\nThis notebook [here][1] from Flower Comp shows how to work with them\n\n[1]:https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu",
      "votes": null
    },
    {
      "id": "865396",
      "postDate": "05/28/2020 15:53:08",
      "content": "<p>As <a href=\"/cdeotte\">@cdeotte</a> pointed out, there are TensorFlow files. They are mostly used with the tf.Data API in order to build a fast pipeline. It is a great news since that will make it easier for us to use TPU, I guess ;)</p>",
      "rawMarkdown": "As @cdeotte pointed out, there are TensorFlow files. They are mostly used with the tf.Data API in order to build a fast pipeline. It is a great news since that will make it easier for us to use TPU, I guess ;)",
      "votes": null
    },
    {
      "id": "865665",
      "postDate": "05/28/2020 19:33:11",
      "content": "<p>thankyou <a href=\"/rftexas\">@rftexas</a>  and  <a href=\"/cdeotte\">@cdeotte</a> \nDo upvote my discussion</p>",
      "rawMarkdown": "thankyou @rftexas  and  @cdeotte \nDo upvote my discussion",
      "votes": null
    },
    {
      "id": "865867",
      "postDate": "05/29/2020 00:12:26",
      "content": "<p><a href=\"/debasish05\">@debasish05</a> - Great question! 👍 \n<a href=\"/cdeotte\">@cdeotte</a> - Thanks for the great answer!  Found this really helpful! 👍 🙌 💯 </p>",
      "rawMarkdown": "debasish05 - Great question! 👍 \n@cdeotte - Thanks for the great answer!  Found this really helpful! 👍 🙌 💯",
      "votes": null
    },
    {
      "id": "868070",
      "postDate": "05/30/2020 21:50:16",
      "content": "<p><a href=\"/rftexas\">@rftexas</a> I was trying to load the image data available in the form of .jpg file, but I am facing memory issues, the whole ram getting filled up. what should I do?  if I use DICOM format data or the tfrecord data will it be helpful in this case for avoiding the memory issue?</p>",
      "rawMarkdown": "rftexas I was trying to load the image data available in the form of .jpg file, but I am facing memory issues, the whole ram getting filled up. what should I do?  if I use DICOM format data or the tfrecord data will it be helpful in this case for avoiding the memory issue?",
      "votes": null
    },
    {
      "id": "869271",
      "postDate": "05/31/2020 22:07:48",
      "content": "<p>If you try to load the .jpg files all at once, then it is problematic. You won't have enough memory to store the whole dataset in your RAM. If you want to use TPU, use the Keras tf.data.Dataset API. Otherwise you might want to write a custom Image Generator with Keras but you'll be obliged to use GPU since it doesn't work with TPU.</p>",
      "rawMarkdown": "If you try to load the .jpg files all at once, then it is problematic. You won't have enough memory to store the whole dataset in your RAM. If you want to use TPU, use the Keras tf.data.Dataset API. Otherwise you might want to write a custom Image Generator with Keras but you'll be obliged to use GPU since it doesn't work with TPU.",
      "votes": null
    },
    {
      "id": "872205",
      "postDate": "06/03/2020 02:01:54",
      "content": "<p>Wether you use TPUs or not, I strongly recommend you go with the modern tf.data.Dataset API !\nHere is a tutorial: <a href=\"https://codelabs.developers.google.com/codelabs/keras-flowers-data/#0\">TPU-speed data pipelines: tf.data.Dataset and TFRecords</a></p>",
      "rawMarkdown": "Wether you use TPUs or not, I strongly recommend you go with the modern tf.data.Dataset API !\nHere is a tutorial: [TPU-speed data pipelines: tf.data.Dataset and TFRecords](https://codelabs.developers.google.com/codelabs/keras-flowers-data/#0)",
      "votes": null
    },
    {
      "id": "872207",
      "postDate": "06/03/2020 02:04:38",
      "content": "<p>TFRecord is a container format for data. The goal is to shard a dataset into a reasonable number of reasonably large files. You can then use the tf.data.Dataset API to stream the data from GCS during training and get really good throughput. The sharding allows tf.data.Dataset to stream from multiple files in parallel.\nTutorial here: <a href=\"https://codelabs.developers.google.com/codelabs/keras-flowers-data/#0\">TPU-speed data pipelines: tf.data.Dataset and TFRecords</a></p>",
      "rawMarkdown": "TFRecord is a container format for data. The goal is to shard a dataset into a reasonable number of reasonably large files. You can then use the tf.data.Dataset API to stream the data from GCS during training and get really good throughput. The sharding allows tf.data.Dataset to stream from multiple files in parallel.\nTutorial here: [TPU-speed data pipelines: tf.data.Dataset and TFRecords](https://codelabs.developers.google.com/codelabs/keras-flowers-data/#0)",
      "votes": null
    },
    {
      "id": "938064",
      "postDate": "07/21/2020 10:19:46",
      "content": "<p>I am having some trouble understanding TFRecords! Having checked the google TF documentation everywhere to access them they defined feature set in advance. So, how would one get the idea of the features present in TFRecord files?</p>",
      "rawMarkdown": "I am having some trouble understanding TFRecords! Having checked the google TF documentation everywhere to access them they defined feature set in advance. So, how would one get the idea of the features present in TFRecord files?",
      "votes": null
    },
    {
      "id": "949891",
      "postDate": "07/29/2020 03:38:58",
      "content": "<p>That is what I wanted to know. I read the data by tf.TFRecordDataset('tfrec files')</p>\n\n<p>And parse the result by repr(result.take(1)), but only two fields 'image' and 'image name' are both string. How could I read the TFREC files?</p>",
      "rawMarkdown": "That is what I wanted to know. I read the data by tf.TFRecordDataset('tfrec files')\n\nAnd parse the result by repr(result.take(1)), but only two fields 'image' and 'image name' are both string. How could I read the TFREC files?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 864588,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "05/28/2020 04:48:03",
      "content": "<p>TFrecords are TensorFlow's format for save files to disk. They have advantages for fast loading when using data loaders to feed your model.</p>\n\n<p>This notebook <a href=\"https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu\">here</a> from Flower Comp shows how to work with them</p>",
      "votes": null,
      "replies": [
        {
          "id": 865867,
          "author_name": "yeayates21",
          "author_url": "",
          "post_date": "05/29/2020 00:12:26",
          "content": "<p><a href=\"/debasish05\">@debasish05</a> - Great question! 👍 \n<a href=\"/cdeotte\">@cdeotte</a> - Thanks for the great answer!  Found this really helpful! 👍 🙌 💯 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 938064,
          "author_name": "fireheart7",
          "author_url": "",
          "post_date": "07/21/2020 10:19:46",
          "content": "<p>I am having some trouble understanding TFRecords! Having checked the google TF documentation everywhere to access them they defined feature set in advance. So, how would one get the idea of the features present in TFRecord files?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 949891,
          "author_name": "laurencelin",
          "author_url": "",
          "post_date": "07/29/2020 03:38:58",
          "content": "<p>That is what I wanted to know. I read the data by tf.TFRecordDataset('tfrec files')</p>\n\n<p>And parse the result by repr(result.take(1)), but only two fields 'image' and 'image name' are both string. How could I read the TFREC files?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 865396,
      "author_name": "rftexas",
      "author_url": "",
      "post_date": "05/28/2020 15:53:08",
      "content": "<p>As <a href=\"/cdeotte\">@cdeotte</a> pointed out, there are TensorFlow files. They are mostly used with the tf.Data API in order to build a fast pipeline. It is a great news since that will make it easier for us to use TPU, I guess ;)</p>",
      "votes": null,
      "replies": [
        {
          "id": 868070,
          "author_name": "soumya9977",
          "author_url": "",
          "post_date": "05/30/2020 21:50:16",
          "content": "<p><a href=\"/rftexas\">@rftexas</a> I was trying to load the image data available in the form of .jpg file, but I am facing memory issues, the whole ram getting filled up. what should I do?  if I use DICOM format data or the tfrecord data will it be helpful in this case for avoiding the memory issue?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 869271,
          "author_name": "rftexas",
          "author_url": "",
          "post_date": "05/31/2020 22:07:48",
          "content": "<p>If you try to load the .jpg files all at once, then it is problematic. You won't have enough memory to store the whole dataset in your RAM. If you want to use TPU, use the Keras tf.data.Dataset API. Otherwise you might want to write a custom Image Generator with Keras but you'll be obliged to use GPU since it doesn't work with TPU.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 872205,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "06/03/2020 02:01:54",
          "content": "<p>Wether you use TPUs or not, I strongly recommend you go with the modern tf.data.Dataset API !\nHere is a tutorial: <a href=\"https://codelabs.developers.google.com/codelabs/keras-flowers-data/#0\">TPU-speed data pipelines: tf.data.Dataset and TFRecords</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 865665,
      "author_name": "debasish05",
      "author_url": "",
      "post_date": "05/28/2020 19:33:11",
      "content": "<p>thankyou <a href=\"/rftexas\">@rftexas</a>  and  <a href=\"/cdeotte\">@cdeotte</a> \nDo upvote my discussion</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 872207,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "06/03/2020 02:04:38",
      "content": "<p>TFRecord is a container format for data. The goal is to shard a dataset into a reasonable number of reasonably large files. You can then use the tf.data.Dataset API to stream the data from GCS during training and get really good throughput. The sharding allows tf.data.Dataset to stream from multiple files in parallel.\nTutorial here: <a href=\"https://codelabs.developers.google.com/codelabs/keras-flowers-data/#0\">TPU-speed data pipelines: tf.data.Dataset and TFRecords</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "864580": "What is the use of TFrecords? I read from the documentation from the TensorFlow, still am in confusion. Answer along with a notebook will be more helpful",
    "864588": "TFrecords are TensorFlow's format for save files to disk. They have advantages for fast loading when using data loaders to feed your model.\n\nThis notebook [here][1] from Flower Comp shows how to work with them\n\n[1]:https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu",
    "865396": "As @cdeotte pointed out, there are TensorFlow files. They are mostly used with the tf.Data API in order to build a fast pipeline. It is a great news since that will make it easier for us to use TPU, I guess ;)",
    "865665": "thankyou @rftexas  and  @cdeotte \nDo upvote my discussion",
    "865867": "debasish05 - Great question! 👍 \n@cdeotte - Thanks for the great answer!  Found this really helpful! 👍 🙌 💯",
    "868070": "rftexas I was trying to load the image data available in the form of .jpg file, but I am facing memory issues, the whole ram getting filled up. what should I do?  if I use DICOM format data or the tfrecord data will it be helpful in this case for avoiding the memory issue?",
    "869271": "If you try to load the .jpg files all at once, then it is problematic. You won't have enough memory to store the whole dataset in your RAM. If you want to use TPU, use the Keras tf.data.Dataset API. Otherwise you might want to write a custom Image Generator with Keras but you'll be obliged to use GPU since it doesn't work with TPU.",
    "872205": "Wether you use TPUs or not, I strongly recommend you go with the modern tf.data.Dataset API !\nHere is a tutorial: [TPU-speed data pipelines: tf.data.Dataset and TFRecords](https://codelabs.developers.google.com/codelabs/keras-flowers-data/#0)",
    "872207": "TFRecord is a container format for data. The goal is to shard a dataset into a reasonable number of reasonably large files. You can then use the tf.data.Dataset API to stream the data from GCS during training and get really good throughput. The sharding allows tf.data.Dataset to stream from multiple files in parallel.\nTutorial here: [TPU-speed data pipelines: tf.data.Dataset and TFRecords](https://codelabs.developers.google.com/codelabs/keras-flowers-data/#0)",
    "938064": "I am having some trouble understanding TFRecords! Having checked the google TF documentation everywhere to access them they defined feature set in advance. So, how would one get the idea of the features present in TFRecord files?",
    "949891": "That is what I wanted to know. I read the data by tf.TFRecordDataset('tfrec files')\n\nAnd parse the result by repr(result.take(1)), but only two fields 'image' and 'image name' are both string. How could I read the TFREC files?"
  },
  "source": "meta"
}