{
  "id": 105695,
  "title": "Training Speed Up with Local Image Store using h5py package",
  "url": "/competitions/aptos2019-blindness-detection/discussion/105695",
  "author_name": "",
  "post_date": "2019-08-25T15:08:11.419315Z",
  "votes": 7,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I was able to obtain a 10x speed up in my CNN training by first pre-processing all images once and then storing them on the kernel's local file system in a single <strong>h5py</strong> file. During training the images are then read from the <strong>h5py</strong> file for each epoc of the training.</p>\n\n<p>I wanted to share this approach in case it may be useful to other participants in this competition.</p>\n\n<p>The <strong>h5py reference material</strong> is available here: \n<a href=\"http://docs.h5py.org/en/2.9.0/index.html\">http://docs.h5py.org/en/2.9.0/index.html</a></p>\n\n<p>Below is some sample code that I am using that might be a helpful reference.</p>\n\n<p><strong>Initialize the h5py file for later use:</strong>\n```\nimport h5py\nimport io\nimport numpy as np\nfrom PIL import Image</p>\n\n<h3>Initialize the H5PY file</h3>\n\n<p>hdf5_file = 'hdf5_data_file.hdf5'\nf = h5py.File(hdf5_file)\nf.close()\n```</p>\n\n<p><strong>Helper functions for writing images to, and reading images from, the h5py file.</strong>\n```\ndef write_image_file_to_hdf5(new_file_name, image):\n    img_np_array = np.asarray(image)\n    f = h5py.File(hdf5_file, 'r+')\n    dset = f.create_dataset(new_file_name, data=img_np_array)\n    f.flush()\n    f.close()\n    return()</p>\n\n<p>def read_image_file_from_hdf5(new_file_name):\n    f = h5py.File(hdf5_file, 'r')\n    dset_read = f.get(new_file_name)\n    dset_read_np = np.array(dset_read)\n    f.close()\n    return(dset_read_np)\n<code>\n**One additional point, I use the following image transformation on the data returned by the read function**\n</code>\ndset_read_as_np = read_image_file_from_hdf5(new_file_name)\nimage = transforms.ToPILImage()(dset_read_as_np)\n```</p>",
  "messages": [
    {
      "id": "607599",
      "postDate": "08/25/2019 15:08:11",
      "content": "<p>I was able to obtain a 10x speed up in my CNN training by first pre-processing all images once and then storing them on the kernel's local file system in a single <strong>h5py</strong> file. During training the images are then read from the <strong>h5py</strong> file for each epoc of the training.</p>\n\n<p>I wanted to share this approach in case it may be useful to other participants in this competition.</p>\n\n<p>The <strong>h5py reference material</strong> is available here: \n<a href=\"http://docs.h5py.org/en/2.9.0/index.html\">http://docs.h5py.org/en/2.9.0/index.html</a></p>\n\n<p>Below is some sample code that I am using that might be a helpful reference.</p>\n\n<p><strong>Initialize the h5py file for later use:</strong>\n```\nimport h5py\nimport io\nimport numpy as np\nfrom PIL import Image</p>\n\n<h3>Initialize the H5PY file</h3>\n\n<p>hdf5_file = 'hdf5_data_file.hdf5'\nf = h5py.File(hdf5_file)\nf.close()\n```</p>\n\n<p><strong>Helper functions for writing images to, and reading images from, the h5py file.</strong>\n```\ndef write_image_file_to_hdf5(new_file_name, image):\n    img_np_array = np.asarray(image)\n    f = h5py.File(hdf5_file, 'r+')\n    dset = f.create_dataset(new_file_name, data=img_np_array)\n    f.flush()\n    f.close()\n    return()</p>\n\n<p>def read_image_file_from_hdf5(new_file_name):\n    f = h5py.File(hdf5_file, 'r')\n    dset_read = f.get(new_file_name)\n    dset_read_np = np.array(dset_read)\n    f.close()\n    return(dset_read_np)\n<code>\n**One additional point, I use the following image transformation on the data returned by the read function**\n</code>\ndset_read_as_np = read_image_file_from_hdf5(new_file_name)\nimage = transforms.ToPILImage()(dset_read_as_np)\n```</p>",
      "rawMarkdown": "I was able to obtain a 10x speed up in my CNN training by first pre-processing all images once and then storing them on the kernel's local file system in a single **h5py** file. During training the images are then read from the **h5py** file for each epoc of the training.\n\nI wanted to share this approach in case it may be useful to other participants in this competition.\n\nThe **h5py reference material** is available here: \n[http://docs.h5py.org/en/2.9.0/index.html](http://docs.h5py.org/en/2.9.0/index.html)\n\nBelow is some sample code that I am using that might be a helpful reference.\n\n**Initialize the h5py file for later use:**\n```\nimport h5py\nimport io\nimport numpy as np\nfrom PIL import Image\n\n### Initialize the H5PY file\nhdf5_file = 'hdf5_data_file.hdf5'\nf = h5py.File(hdf5_file)\nf.close()\n```\n\n\n**Helper functions for writing images to, and reading images from, the h5py file.**\n```\ndef write_image_file_to_hdf5(new_file_name, image):\n    img_np_array = np.asarray(image)\n    f = h5py.File(hdf5_file, 'r+')\n    dset = f.create_dataset(new_file_name, data=img_np_array)\n    f.flush()\n    f.close()\n    return()\n\ndef read_image_file_from_hdf5(new_file_name):\n    f = h5py.File(hdf5_file, 'r')\n    dset_read = f.get(new_file_name)\n    dset_read_np = np.array(dset_read)\n    f.close()\n    return(dset_read_np)\n```\n**One additional point, I use the following image transformation on the data returned by the read function**\n```\ndset_read_as_np = read_image_file_from_hdf5(new_file_name)\nimage = transforms.ToPILImage()(dset_read_as_np)\n```",
      "votes": null
    },
    {
      "id": "611640",
      "postDate": "08/29/2019 11:10:14",
      "content": "<p>How are you reading it in keras ?</p>",
      "rawMarkdown": "How are you reading it in keras ?",
      "votes": null
    },
    {
      "id": "611761",
      "postDate": "08/29/2019 12:22:12",
      "content": "<p>I am using EfficientNet B3.</p>",
      "rawMarkdown": "I am using EfficientNet B3.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 611640,
      "author_name": "prachi1211",
      "author_url": "",
      "post_date": "08/29/2019 11:10:14",
      "content": "<p>How are you reading it in keras ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 611761,
          "author_name": "glaskosk",
          "author_url": "",
          "post_date": "08/29/2019 12:22:12",
          "content": "<p>I am using EfficientNet B3.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "607599": "I was able to obtain a 10x speed up in my CNN training by first pre-processing all images once and then storing them on the kernel's local file system in a single **h5py** file. During training the images are then read from the **h5py** file for each epoc of the training.\n\nI wanted to share this approach in case it may be useful to other participants in this competition.\n\nThe **h5py reference material** is available here: \n[http://docs.h5py.org/en/2.9.0/index.html](http://docs.h5py.org/en/2.9.0/index.html)\n\nBelow is some sample code that I am using that might be a helpful reference.\n\n**Initialize the h5py file for later use:**\n```\nimport h5py\nimport io\nimport numpy as np\nfrom PIL import Image\n\n### Initialize the H5PY file\nhdf5_file = 'hdf5_data_file.hdf5'\nf = h5py.File(hdf5_file)\nf.close()\n```\n\n\n**Helper functions for writing images to, and reading images from, the h5py file.**\n```\ndef write_image_file_to_hdf5(new_file_name, image):\n    img_np_array = np.asarray(image)\n    f = h5py.File(hdf5_file, 'r+')\n    dset = f.create_dataset(new_file_name, data=img_np_array)\n    f.flush()\n    f.close()\n    return()\n\ndef read_image_file_from_hdf5(new_file_name):\n    f = h5py.File(hdf5_file, 'r')\n    dset_read = f.get(new_file_name)\n    dset_read_np = np.array(dset_read)\n    f.close()\n    return(dset_read_np)\n```\n**One additional point, I use the following image transformation on the data returned by the read function**\n```\ndset_read_as_np = read_image_file_from_hdf5(new_file_name)\nimage = transforms.ToPILImage()(dset_read_as_np)\n```",
    "611640": "How are you reading it in keras ?",
    "611761": "I am using EfficientNet B3."
  },
  "source": "meta"
}