{
  "id": 164996,
  "title": "Benefits of TFRecord Dataset?  ",
  "url": "/competitions/alaska2-image-steganalysis/discussion/164996",
  "author_name": "",
  "post_date": "2020-07-08T06:46:14.189996100Z",
  "votes": 3,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi, I knew TFRecord Dataset in this competition.\nIt is a new type of dataset module for me.  It's image data, but stored as string. \nLooking at docs, it said it's good for serialization. The reason is they divided the data into 100 to 200MB?\nWhat are the advantages of TFRecord, unlike Pytorch's dataloader module?</p>",
  "messages": [
    {
      "id": "919877",
      "postDate": "07/08/2020 06:46:14",
      "content": "<p>Hi, I knew TFRecord Dataset in this competition.\nIt is a new type of dataset module for me.  It's image data, but stored as string. \nLooking at docs, it said it's good for serialization. The reason is they divided the data into 100 to 200MB?\nWhat are the advantages of TFRecord, unlike Pytorch's dataloader module?</p>",
      "rawMarkdown": "Hi, I knew TFRecord Dataset in this competition.\nIt is a new type of dataset module for me.  It's image data, but stored as string. \nLooking at docs, it said it's good for serialization. The reason is they divided the data into 100 to 200MB?\nWhat are the advantages of TFRecord, unlike Pytorch's dataloader module?",
      "votes": null
    },
    {
      "id": "919940",
      "postDate": "07/08/2020 07:57:11",
      "content": "<p>Hi, \nfrom my understanding, these are the benefits of it\n1. Reading each individual image in the large dataset is often a slow process. So, to make it fast we store many images as tensor Tf records and read from there. \n2. At a time, we can feed data from many TF records to TPU for training without worrying about the order. This will make the data feeding process fast and hence improve the overall run time. </p>\n\n<p>These are the obvious one I know. There might be other benefits too. </p>",
      "rawMarkdown": "Hi, \nfrom my understanding, these are the benefits of it\n1. Reading each individual image in the large dataset is often a slow process. So, to make it fast we store many images as tensor Tf records and read from there. \n2. At a time, we can feed data from many TF records to TPU for training without worrying about the order. This will make the data feeding process fast and hence improve the overall run time. \n\nThese are the obvious one I know. There might be other benefits too.",
      "votes": null
    },
    {
      "id": "920205",
      "postDate": "07/08/2020 12:42:19",
      "content": "<p>Thanks a lot! Now it's clear to me</p>",
      "rawMarkdown": "Thanks a lot! Now it's clear to me",
      "votes": null
    },
    {
      "id": "920559",
      "postDate": "07/08/2020 17:09:48",
      "content": "<p>its clear thanks <a href=\"/urvishp80\">@urvishp80</a> </p>",
      "rawMarkdown": "its clear thanks @urvishp80",
      "votes": null
    },
    {
      "id": "920571",
      "postDate": "07/08/2020 17:16:47",
      "content": "<p>Happy to help.</p>",
      "rawMarkdown": "Happy to help.",
      "votes": null
    },
    {
      "id": "920605",
      "postDate": "07/08/2020 17:50:25",
      "content": "<p>100-200mb is the recommended size in the TFRecord article on Tensorflow. see <a href=\"https://www.tensorflow.org/tutorials/load_data/tfrecord#walkthrough_reading_and_writing_image_data\">https://www.tensorflow.org/tutorials/load_data/tfrecord#walkthrough_reading_and_writing_image_data</a></p>\n\n<p>I believe any size greater than 100mb is fine. This is because it takes time for the TPU to make a request for a file. </p>\n\n<p>For example, lets say you have 10k image files. Each time the TPU needs a new image, it has to make a request for that image file so it is making 10k requests to get data. That can really slow you down because it has to wait for the request to be filled by google cloud storage. </p>",
      "rawMarkdown": "100-200mb is the recommended size in the TFRecord article on Tensorflow. see https://www.tensorflow.org/tutorials/load_data/tfrecord#walkthrough_reading_and_writing_image_data\n\nI believe any size greater than 100mb is fine. This is because it takes time for the TPU to make a request for a file. \n\nFor example, lets say you have 10k image files. Each time the TPU needs a new image, it has to make a request for that image file so it is making 10k requests to get data. That can really slow you down because it has to wait for the request to be filled by google cloud storage.",
      "votes": null
    },
    {
      "id": "920667",
      "postDate": "07/08/2020 18:38:25",
      "content": "<p>Thanks for the insights <a href=\"/hooong\">@hooong</a> \nIt makes sense. We can increase performance by reducing the number of requests.</p>",
      "rawMarkdown": "Thanks for the insights @hooong \nIt makes sense. We can increase performance by reducing the number of requests.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 919940,
      "author_name": "urvishp80",
      "author_url": "",
      "post_date": "07/08/2020 07:57:11",
      "content": "<p>Hi, \nfrom my understanding, these are the benefits of it\n1. Reading each individual image in the large dataset is often a slow process. So, to make it fast we store many images as tensor Tf records and read from there. \n2. At a time, we can feed data from many TF records to TPU for training without worrying about the order. This will make the data feeding process fast and hence improve the overall run time. </p>\n\n<p>These are the obvious one I know. There might be other benefits too. </p>",
      "votes": null,
      "replies": [
        {
          "id": 920205,
          "author_name": "telljoy",
          "author_url": "",
          "post_date": "07/08/2020 12:42:19",
          "content": "<p>Thanks a lot! Now it's clear to me</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 920559,
          "author_name": "vishnurapps",
          "author_url": "",
          "post_date": "07/08/2020 17:09:48",
          "content": "<p>its clear thanks <a href=\"/urvishp80\">@urvishp80</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 920571,
          "author_name": "urvishp80",
          "author_url": "",
          "post_date": "07/08/2020 17:16:47",
          "content": "<p>Happy to help.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 920605,
          "author_name": "hooong",
          "author_url": "",
          "post_date": "07/08/2020 17:50:25",
          "content": "<p>100-200mb is the recommended size in the TFRecord article on Tensorflow. see <a href=\"https://www.tensorflow.org/tutorials/load_data/tfrecord#walkthrough_reading_and_writing_image_data\">https://www.tensorflow.org/tutorials/load_data/tfrecord#walkthrough_reading_and_writing_image_data</a></p>\n\n<p>I believe any size greater than 100mb is fine. This is because it takes time for the TPU to make a request for a file. </p>\n\n<p>For example, lets say you have 10k image files. Each time the TPU needs a new image, it has to make a request for that image file so it is making 10k requests to get data. That can really slow you down because it has to wait for the request to be filled by google cloud storage. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 920667,
          "author_name": "vishnurapps",
          "author_url": "",
          "post_date": "07/08/2020 18:38:25",
          "content": "<p>Thanks for the insights <a href=\"/hooong\">@hooong</a> \nIt makes sense. We can increase performance by reducing the number of requests.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "919877": "Hi, I knew TFRecord Dataset in this competition.\nIt is a new type of dataset module for me.  It's image data, but stored as string. \nLooking at docs, it said it's good for serialization. The reason is they divided the data into 100 to 200MB?\nWhat are the advantages of TFRecord, unlike Pytorch's dataloader module?",
    "919940": "Hi, \nfrom my understanding, these are the benefits of it\n1. Reading each individual image in the large dataset is often a slow process. So, to make it fast we store many images as tensor Tf records and read from there. \n2. At a time, we can feed data from many TF records to TPU for training without worrying about the order. This will make the data feeding process fast and hence improve the overall run time. \n\nThese are the obvious one I know. There might be other benefits too.",
    "920205": "Thanks a lot! Now it's clear to me",
    "920559": "its clear thanks @urvishp80",
    "920571": "Happy to help.",
    "920605": "100-200mb is the recommended size in the TFRecord article on Tensorflow. see https://www.tensorflow.org/tutorials/load_data/tfrecord#walkthrough_reading_and_writing_image_data\n\nI believe any size greater than 100mb is fine. This is because it takes time for the TPU to make a request for a file. \n\nFor example, lets say you have 10k image files. Each time the TPU needs a new image, it has to make a request for that image file so it is making 10k requests to get data. That can really slow you down because it has to wait for the request to be filled by google cloud storage.",
    "920667": "Thanks for the insights @hooong \nIt makes sense. We can increase performance by reducing the number of requests."
  },
  "source": "meta"
}