{
  "id": 109731,
  "title": "How do you load data?",
  "url": "/competitions/recursion-cellular-image-classification/discussion/109731",
  "author_name": "",
  "post_date": "2019-09-21T18:56:46.934482100Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Since we have 6 channels in this competition, I noticed that there are a few ways one can approach data loading:</p>\n\n<ol>\n<li>Convert all the images to RGB and train using just 3 channels (easiest, but you'll probably loose some information)</li>\n<li>Create 6 channel images \"on the fly\" from 6 .png files on the disk. It seems to be the most common approach in kernels.</li>\n<li>Save 6-channel numpy arrays to disk and load them during training.</li>\n<li>???</li>\n</ol>\n\n<p>I am currently with #2, but suppose that #3 is the fastest way and can give you boost in training speed, but when I tried saving numpy arrays to disk it took all my free space :)</p>\n\n<p>If anyone tried #3, is it really worth it? And how much additional space did it take you?</p>",
  "messages": [
    {
      "id": "631286",
      "postDate": "09/21/2019 18:56:46",
      "content": "<p>Since we have 6 channels in this competition, I noticed that there are a few ways one can approach data loading:</p>\n\n<ol>\n<li>Convert all the images to RGB and train using just 3 channels (easiest, but you'll probably loose some information)</li>\n<li>Create 6 channel images \"on the fly\" from 6 .png files on the disk. It seems to be the most common approach in kernels.</li>\n<li>Save 6-channel numpy arrays to disk and load them during training.</li>\n<li>???</li>\n</ol>\n\n<p>I am currently with #2, but suppose that #3 is the fastest way and can give you boost in training speed, but when I tried saving numpy arrays to disk it took all my free space :)</p>\n\n<p>If anyone tried #3, is it really worth it? And how much additional space did it take you?</p>",
      "rawMarkdown": "Since we have 6 channels in this competition, I noticed that there are a few ways one can approach data loading:\n\n1. Convert all the images to RGB and train using just 3 channels (easiest, but you'll probably loose some information)\n2. Create 6 channel images \"on the fly\" from 6 .png files on the disk. It seems to be the most common approach in kernels.\n3. Save 6-channel numpy arrays to disk and load them during training.\n4. ???\n\nI am currently with #2, but suppose that #3 is the fastest way and can give you boost in training speed, but when I tried saving numpy arrays to disk it took all my free space :)\n\nIf anyone tried #3, is it really worth it? And how much additional space did it take you?",
      "votes": null
    },
    {
      "id": "631292",
      "postDate": "09/21/2019 19:05:39",
      "content": "<p>I'm saving to HDF5 files (h5py); same as numpy array (#3) for disk space, but also easy to read part of the array this way. Test data (17G) becomes about 60G using uint8 (pixel values are in [0, 255]). I haven't compared the speed.​</p>",
      "rawMarkdown": "I'm saving to HDF5 files (h5py); same as numpy array (#3) for disk space, but also easy to read part of the array this way. Test data (17G) becomes about 60G using uint8 (pixel values are in [0, 255]). I haven't compared the speed.​",
      "votes": null
    },
    {
      "id": "631943",
      "postDate": "09/23/2019 01:58:04",
      "content": "<p>I build numpy 6x512x512 uint8 images, then pickle and compress them to a disk-backed cache.  Compression uses zstandard (\"pip install zstandard\"), which is easy to use and pretty fast.  The entire cache (training and test) costs 60G, compared to the 50G for the original unpacked PNG images.  Loading images from disk cache is about 20x faster than using pyplot.imread (I can cycle through the entire training and test set in maybe 10 minutes, whereas reading the originals is &gt; 3h).</p>\n\n<p>I'd assume HDF5 reads a lot quicker, I just thought it was about time I wrote a cache manager :)</p>",
      "rawMarkdown": "I build numpy 6x512x512 uint8 images, then pickle and compress them to a disk-backed cache.  Compression uses zstandard (\"pip install zstandard\"), which is easy to use and pretty fast.  The entire cache (training and test) costs 60G, compared to the 50G for the original unpacked PNG images.  Loading images from disk cache is about 20x faster than using pyplot.imread (I can cycle through the entire training and test set in maybe 10 minutes, whereas reading the originals is &gt; 3h).\n\nI'd assume HDF5 reads a lot quicker, I just thought it was about time I wrote a cache manager :)",
      "votes": null
    },
    {
      "id": "632032",
      "postDate": "09/23/2019 06:28:47",
      "content": "<p>I'm just using tf.data pipeline, so it seems to be second option. I guess the speedup you can achieve depends on your hardware, but in my experience, using tf.data is about 20-30% faster than keras generator. But all of this depends heavily on code structure, hardware, task.</p>",
      "rawMarkdown": "I'm just using tf.data pipeline, so it seems to be second option. I guess the speedup you can achieve depends on your hardware, but in my experience, using tf.data is about 20-30% faster than keras generator. But all of this depends heavily on code structure, hardware, task.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 631292,
      "author_name": "junkoda",
      "author_url": "",
      "post_date": "09/21/2019 19:05:39",
      "content": "<p>I'm saving to HDF5 files (h5py); same as numpy array (#3) for disk space, but also easy to read part of the array this way. Test data (17G) becomes about 60G using uint8 (pixel values are in [0, 255]). I haven't compared the speed.​</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 631943,
      "author_name": "snorkle",
      "author_url": "",
      "post_date": "09/23/2019 01:58:04",
      "content": "<p>I build numpy 6x512x512 uint8 images, then pickle and compress them to a disk-backed cache.  Compression uses zstandard (\"pip install zstandard\"), which is easy to use and pretty fast.  The entire cache (training and test) costs 60G, compared to the 50G for the original unpacked PNG images.  Loading images from disk cache is about 20x faster than using pyplot.imread (I can cycle through the entire training and test set in maybe 10 minutes, whereas reading the originals is &gt; 3h).</p>\n\n<p>I'd assume HDF5 reads a lot quicker, I just thought it was about time I wrote a cache manager :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 632032,
      "author_name": "cateek",
      "author_url": "",
      "post_date": "09/23/2019 06:28:47",
      "content": "<p>I'm just using tf.data pipeline, so it seems to be second option. I guess the speedup you can achieve depends on your hardware, but in my experience, using tf.data is about 20-30% faster than keras generator. But all of this depends heavily on code structure, hardware, task.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "631286": "Since we have 6 channels in this competition, I noticed that there are a few ways one can approach data loading:\n\n1. Convert all the images to RGB and train using just 3 channels (easiest, but you'll probably loose some information)\n2. Create 6 channel images \"on the fly\" from 6 .png files on the disk. It seems to be the most common approach in kernels.\n3. Save 6-channel numpy arrays to disk and load them during training.\n4. ???\n\nI am currently with #2, but suppose that #3 is the fastest way and can give you boost in training speed, but when I tried saving numpy arrays to disk it took all my free space :)\n\nIf anyone tried #3, is it really worth it? And how much additional space did it take you?",
    "631292": "I'm saving to HDF5 files (h5py); same as numpy array (#3) for disk space, but also easy to read part of the array this way. Test data (17G) becomes about 60G using uint8 (pixel values are in [0, 255]). I haven't compared the speed.​",
    "631943": "I build numpy 6x512x512 uint8 images, then pickle and compress them to a disk-backed cache.  Compression uses zstandard (\"pip install zstandard\"), which is easy to use and pretty fast.  The entire cache (training and test) costs 60G, compared to the 50G for the original unpacked PNG images.  Loading images from disk cache is about 20x faster than using pyplot.imread (I can cycle through the entire training and test set in maybe 10 minutes, whereas reading the originals is &gt; 3h).\n\nI'd assume HDF5 reads a lot quicker, I just thought it was about time I wrote a cache manager :)",
    "632032": "I'm just using tf.data pipeline, so it seems to be second option. I guess the speedup you can achieve depends on your hardware, but in my experience, using tf.data is about 20-30% faster than keras generator. But all of this depends heavily on code structure, hardware, task."
  },
  "source": "meta"
}