{
  "id": 228823,
  "title": "How to accelerate the speed of data-loading?",
  "url": "/competitions/plant-pathology-2021-fgvc8/discussion/228823",
  "author_name": "",
  "post_date": "2021-03-26T18:02:56.756064600Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>It seems my GPU is always hungry, because the I/O of the large images are too slow, is there any solutions?</p>\n<p>Many thanks!</p>",
  "messages": [
    {
      "id": "1253432",
      "postDate": "03/26/2021 18:02:56",
      "content": "<p>It seems my GPU is always hungry, because the I/O of the large images are too slow, is there any solutions?</p>\n<p>Many thanks!</p>",
      "rawMarkdown": "It seems my GPU is always hungry, because the I/O of the large images are too slow, is there any solutions?\n\nMany thanks!",
      "votes": null
    },
    {
      "id": "1254120",
      "postDate": "03/27/2021 11:10:54",
      "content": "<p>One way to boost I/O performance is to use TFrecords instead of the normal data generators like flow_from_dataframe etc due to the fact that TFrecords is a binary format where images are stored as raw bitmaps, which means the CPU doesn’t need to decode the jpeg files, every time it reads them.</p>\n<p>It'll be great if you could tell a bit more about your current data pipeline, for easier debugging!</p>",
      "rawMarkdown": "One way to boost I/O performance is to use TFrecords instead of the normal data generators like flow_from_dataframe etc due to the fact that TFrecords is a binary format where images are stored as raw bitmaps, which means the CPU doesn’t need to decode the jpeg files, every time it reads them.\n\nIt'll be great if you could tell a bit more about your current data pipeline, for easier debugging!",
      "votes": null
    },
    {
      "id": "1258094",
      "postDate": "03/31/2021 10:42:51",
      "content": "<p>I use the Dataloader of Pytorch to load training batches. Since the size of image is large, reading images from disk is reeeeally slow. <br>\nResizing images into smaller size helps for accelerating! But I think this might make the images lose details!</p>",
      "rawMarkdown": "I use the Dataloader of Pytorch to load training batches. Since the size of image is large, reading images from disk is reeeeally slow. \nResizing images into smaller size helps for accelerating! But I think this might make the images lose details!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1254120,
      "author_name": "ashish2001",
      "author_url": "",
      "post_date": "03/27/2021 11:10:54",
      "content": "<p>One way to boost I/O performance is to use TFrecords instead of the normal data generators like flow_from_dataframe etc due to the fact that TFrecords is a binary format where images are stored as raw bitmaps, which means the CPU doesn’t need to decode the jpeg files, every time it reads them.</p>\n<p>It'll be great if you could tell a bit more about your current data pipeline, for easier debugging!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1258094,
          "author_name": "crissallan",
          "author_url": "",
          "post_date": "03/31/2021 10:42:51",
          "content": "<p>I use the Dataloader of Pytorch to load training batches. Since the size of image is large, reading images from disk is reeeeally slow. <br>\nResizing images into smaller size helps for accelerating! But I think this might make the images lose details!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1253432": "It seems my GPU is always hungry, because the I/O of the large images are too slow, is there any solutions?\n\nMany thanks!",
    "1254120": "One way to boost I/O performance is to use TFrecords instead of the normal data generators like flow_from_dataframe etc due to the fact that TFrecords is a binary format where images are stored as raw bitmaps, which means the CPU doesn’t need to decode the jpeg files, every time it reads them.\n\nIt'll be great if you could tell a bit more about your current data pipeline, for easier debugging!",
    "1258094": "I use the Dataloader of Pytorch to load training batches. Since the size of image is large, reading images from disk is reeeeally slow. \nResizing images into smaller size helps for accelerating! But I think this might make the images lose details!"
  },
  "source": "meta"
}