{
  "id": 68118,
  "title": "Data loading strategy",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/68118",
  "author_name": "",
  "post_date": "2018-10-09T14:50:30.073752700Z",
  "votes": 8,
  "comment_count": 13,
  "views": 0,
  "content": "<p>It seems to me, that one of the key problems in this task is handling the loading/processing of the input data.</p>\n\n<p>Numpy array of size (1, 512, 512, 4) as float32 takes around 4 MB of memory, which means you would need ~128 GB of RAM just to fit the whole training set in it.</p>\n\n<p>Also, DataGenerators I've seen in Kernels so far, seem to repeat the time consuming Image reading over and over (maybe I'm missing something, I never used any DG before).</p>\n\n<p>I tested several approaches and this is what I found so far:</p>\n\n<ol>\n<li>PIL.Image seems to be fastest for image loading to numpy arrays</li>\n<li>np.savez_compressed seems to be fastest for numpy arrays writing and reading (h5py, np.save and np.savez were slower)</li>\n<li>Reading preprocessed and saved numpy array from np.savez_compressed is ~10 times faster than reading image using PIL.Image</li>\n<li>Time of reading and writing numpy array using np.savez_compressed increases ~linearly with the number of images in it.</li>\n</ol>\n\n<p>Most promising strategy for dataloading seems to me to process the dataset once, divided into many batches (preferably using CPU multiprocessing), save it to hard drive and modify DataGenerator to read input data from these preprocessed files. And since I'm just a beginner, it's much easier to be said than to be done.\nIt also has the obvious setback, that to change batches (e.g. because of different seed), you need to repeat the whole processing task.</p>\n\n<p>I will appreciate any insights for this problem. I also appologize in case I'm stating something very obvious.</p>",
  "messages": [
    {
      "id": "401173",
      "postDate": "10/09/2018 14:50:30",
      "content": "<p>It seems to me, that one of the key problems in this task is handling the loading/processing of the input data.</p>\n\n<p>Numpy array of size (1, 512, 512, 4) as float32 takes around 4 MB of memory, which means you would need ~128 GB of RAM just to fit the whole training set in it.</p>\n\n<p>Also, DataGenerators I've seen in Kernels so far, seem to repeat the time consuming Image reading over and over (maybe I'm missing something, I never used any DG before).</p>\n\n<p>I tested several approaches and this is what I found so far:</p>\n\n<ol>\n<li>PIL.Image seems to be fastest for image loading to numpy arrays</li>\n<li>np.savez_compressed seems to be fastest for numpy arrays writing and reading (h5py, np.save and np.savez were slower)</li>\n<li>Reading preprocessed and saved numpy array from np.savez_compressed is ~10 times faster than reading image using PIL.Image</li>\n<li>Time of reading and writing numpy array using np.savez_compressed increases ~linearly with the number of images in it.</li>\n</ol>\n\n<p>Most promising strategy for dataloading seems to me to process the dataset once, divided into many batches (preferably using CPU multiprocessing), save it to hard drive and modify DataGenerator to read input data from these preprocessed files. And since I'm just a beginner, it's much easier to be said than to be done.\nIt also has the obvious setback, that to change batches (e.g. because of different seed), you need to repeat the whole processing task.</p>\n\n<p>I will appreciate any insights for this problem. I also appologize in case I'm stating something very obvious.</p>",
      "rawMarkdown": "It seems to me, that one of the key problems in this task is handling the loading/processing of the input data.\n\nNumpy array of size (1, 512, 512, 4) as float32 takes around 4 MB of memory, which means you would need ~128 GB of RAM just to fit the whole training set in it.\n\nAlso, DataGenerators I've seen in Kernels so far, seem to repeat the time consuming Image reading over and over (maybe I'm missing something, I never used any DG before).\n\nI tested several approaches and this is what I found so far:\n\n 1. PIL.Image seems to be fastest for image loading to numpy arrays\n 2. np.savez_compressed seems to be fastest for numpy arrays writing and reading (h5py, np.save and np.savez were slower)\n 3. Reading preprocessed and saved numpy array from np.savez_compressed is ~10 times faster than reading image using PIL.Image\n 4. Time of reading and writing numpy array using np.savez_compressed increases ~linearly with the number of images in it.\n\nMost promising strategy for dataloading seems to me to process the dataset once, divided into many batches (preferably using CPU multiprocessing), save it to hard drive and modify DataGenerator to read input data from these preprocessed files. And since I'm just a beginner, it's much easier to be said than to be done.\nIt also has the obvious setback, that to change batches (e.g. because of different seed), you need to repeat the whole processing task.\n\nI will appreciate any insights for this problem. I also appologize in case I'm stating something very obvious.",
      "votes": null
    },
    {
      "id": "401178",
      "postDate": "10/09/2018 14:59:54",
      "content": "<p>It wasn’t so bad. Nevertheless, I wrote a custom Keras generator. I also enabled multiprocessing (with 16 processes) and created a large queue (256 batches).</p>",
      "rawMarkdown": "It wasn’t so bad. Nevertheless, I wrote a custom Keras generator. I also enabled multiprocessing (with 16 processes) and created a large queue (256 batches).",
      "votes": null
    },
    {
      "id": "401297",
      "postDate": "10/09/2018 19:27:44",
      "content": "<p>Initially with Keras I had issues with performance. The key is to have enough workers, the default of 1 is too slow. I also have the input images on a SSD drive.</p>\n\n<pre>model = ModelMGPU(orig_model , 2)\nmodel.compile(optimizer=opt, loss=brian_loss, metrics=metrics)\n\nmodel_name = 'hpa_sub_4.h5'\nprint(\"LOAD: \" + model_name)\nmodel.load_weights(model_name,reshape=True)\n\ntrain_gen.batch_size = BATCH_SIZE\nresults = model.fit_generator(train_gen, \n                                steps_per_epoch = train_gen.samples//BATCH_SIZE,\n                                validation_data = (valid_x, valid_y), \n                                epochs = 1000, \n                                callbacks=[ reduce_lr,checkpointer,metrics_cb],\n                                workers=24)\n</pre>",
      "rawMarkdown": "Initially with Keras I had issues with performance. The key is to have enough workers, the default of 1 is too slow. I also have the input images on a SSD drive.\n\n<pre>model = ModelMGPU(orig_model , 2)\nmodel.compile(optimizer=opt, loss=brian_loss, metrics=metrics)\n\nmodel_name = 'hpa_sub_4.h5'\nprint(\"LOAD: \" + model_name)\nmodel.load_weights(model_name,reshape=True)\n\ntrain_gen.batch_size = BATCH_SIZE\nresults = model.fit_generator(train_gen, \n                                steps_per_epoch = train_gen.samples//BATCH_SIZE,\n                                validation_data = (valid_x, valid_y), \n                                epochs = 1000, \n                                callbacks=[ reduce_lr,checkpointer,metrics_cb],\n                                workers=24)\n</pre>",
      "votes": null
    },
    {
      "id": "401470",
      "postDate": "10/10/2018 06:13:11",
      "content": "<p>One thing that I did notice is that if I had too many transformations on the generator then the training becomes cpu limited.</p>",
      "rawMarkdown": "One thing that I did notice is that if I had too many transformations on the generator then the training becomes cpu limited.",
      "votes": null
    },
    {
      "id": "402983",
      "postDate": "10/12/2018 17:32:20",
      "content": "<p>OpenCV is faster than PIL for reading in images with PyTorch.</p>",
      "rawMarkdown": "OpenCV is faster than PIL for reading in images with PyTorch.",
      "votes": null
    },
    {
      "id": "404324",
      "postDate": "10/15/2018 15:07:30",
      "content": "<p>Maybe there could be version and system specific issues, but I also have better results with open cv.</p>",
      "rawMarkdown": "Maybe there could be version and system specific issues, but I also have better results with open cv.",
      "votes": null
    },
    {
      "id": "404425",
      "postDate": "10/15/2018 17:17:27",
      "content": "<p>Maybe I'm doing something wrong, but I cannot confirm. Tried on my kernel here: <a href=\"https://www.kaggle.com/rejpalcz/datagenerator-for-fast-data-loading\">https://www.kaggle.com/rejpalcz/datagenerator-for-fast-data-loading</a>, replacing Image.open for cv2.imread and both tests for 128 images takes similar time (around 3.5 seconds). (I didn't commit changes, but you can check with private fork)</p>",
      "rawMarkdown": "Maybe I'm doing something wrong, but I cannot confirm. Tried on my kernel here: https://www.kaggle.com/rejpalcz/datagenerator-for-fast-data-loading, replacing Image.open for cv2.imread and both tests for 128 images takes similar time (around 3.5 seconds). (I didn't commit changes, but you can check with private fork)",
      "votes": null
    },
    {
      "id": "404479",
      "postDate": "10/15/2018 19:26:52",
      "content": "<p>Hi Michal, I haven't tried this for this competition yet as I'm working on TGS but plan to take a look next week hopefully.</p>",
      "rawMarkdown": "Hi Michal, I haven't tried this for this competition yet as I'm working on TGS but plan to take a look next week hopefully.",
      "votes": null
    },
    {
      "id": "416987",
      "postDate": "11/07/2018 15:11:33",
      "content": "<p>You can keep numpy arrays entirely on the hdd and load slices required to GPU as needed by specifying mmap_mode in np.load assuming you have saved the data as npy files. The initial process of creating these files will probably have required the use of your RAM to initially load the images and re-save them, but you can do this into multiple files if necessary.... if that helps at all?</p>\n\n<p>I guess it doesn't help with speed though...</p>",
      "rawMarkdown": "You can keep numpy arrays entirely on the hdd and load slices required to GPU as needed by specifying mmap_mode in np.load assuming you have saved the data as npy files. The initial process of creating these files will probably have required the use of your RAM to initially load the images and re-save them, but you can do this into multiple files if necessary.... if that helps at all?\n\nI guess it doesn't help with speed though...",
      "votes": null
    },
    {
      "id": "417022",
      "postDate": "11/07/2018 16:23:21",
      "content": "<p>If you load big chunks, then from the chunks gather many batches in a generator, it could be interesting, as long as you have at least two workers, one for loading a chunk while the other produces batches. </p>",
      "rawMarkdown": "If you load big chunks, then from the chunks gather many batches in a generator, it could be interesting, as long as you have at least two workers, one for loading a chunk while the other produces batches.",
      "votes": null
    },
    {
      "id": "417114",
      "postDate": "11/07/2018 19:21:39",
      "content": "<p>You should also test different approaches when using multithreading.  You can get quite a bit of boost from multithreaded operations if underlying functions release GIL.</p>",
      "rawMarkdown": "You should also test different approaches when using multithreading.  You can get quite a bit of boost from multithreaded operations if underlying functions release GIL.",
      "votes": null
    },
    {
      "id": "417229",
      "postDate": "11/08/2018 00:58:32",
      "content": "<p>Load images using .npy files speed up the training process; it takes only 70% of the original time needed.</p>",
      "rawMarkdown": "Load images using .npy files speed up the training process; it takes only 70% of the original time needed.",
      "votes": null
    },
    {
      "id": "417252",
      "postDate": "11/08/2018 01:48:34",
      "content": "<p>You should check out joblib for saving and loading numpy arrays</p>",
      "rawMarkdown": "You should check out joblib for saving and loading numpy arrays",
      "votes": null
    },
    {
      "id": "417254",
      "postDate": "11/08/2018 01:56:11",
      "content": "<p>More info on performance here, though the benchmarks might be dated: <a href=\"http://gael-varoquaux.info/programming/new_low-overhead_persistence_in_joblib_for_big_data.html\">http://gael-varoquaux.info/programming/new_low-overhead_persistence_in_joblib_for_big_data.html</a></p>",
      "rawMarkdown": "More info on performance here, though the benchmarks might be dated: http://gael-varoquaux.info/programming/new_low-overhead_persistence_in_joblib_for_big_data.html",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 401178,
      "author_name": "agoodman",
      "author_url": "",
      "post_date": "10/09/2018 14:59:54",
      "content": "<p>It wasn’t so bad. Nevertheless, I wrote a custom Keras generator. I also enabled multiprocessing (with 16 processes) and created a large queue (256 batches).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 401297,
      "author_name": "ldm314",
      "author_url": "",
      "post_date": "10/09/2018 19:27:44",
      "content": "<p>Initially with Keras I had issues with performance. The key is to have enough workers, the default of 1 is too slow. I also have the input images on a SSD drive.</p>\n\n<pre>model = ModelMGPU(orig_model , 2)\nmodel.compile(optimizer=opt, loss=brian_loss, metrics=metrics)\n\nmodel_name = 'hpa_sub_4.h5'\nprint(\"LOAD: \" + model_name)\nmodel.load_weights(model_name,reshape=True)\n\ntrain_gen.batch_size = BATCH_SIZE\nresults = model.fit_generator(train_gen, \n                                steps_per_epoch = train_gen.samples//BATCH_SIZE,\n                                validation_data = (valid_x, valid_y), \n                                epochs = 1000, \n                                callbacks=[ reduce_lr,checkpointer,metrics_cb],\n                                workers=24)\n</pre>",
      "votes": null,
      "replies": [
        {
          "id": 401470,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "10/10/2018 06:13:11",
          "content": "<p>One thing that I did notice is that if I had too many transformations on the generator then the training becomes cpu limited.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 402983,
      "author_name": "maw501",
      "author_url": "",
      "post_date": "10/12/2018 17:32:20",
      "content": "<p>OpenCV is faster than PIL for reading in images with PyTorch.</p>",
      "votes": null,
      "replies": [
        {
          "id": 404324,
          "author_name": "danmoller",
          "author_url": "",
          "post_date": "10/15/2018 15:07:30",
          "content": "<p>Maybe there could be version and system specific issues, but I also have better results with open cv.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 404425,
          "author_name": "rejpalcz",
          "author_url": "",
          "post_date": "10/15/2018 17:17:27",
          "content": "<p>Maybe I'm doing something wrong, but I cannot confirm. Tried on my kernel here: <a href=\"https://www.kaggle.com/rejpalcz/datagenerator-for-fast-data-loading\">https://www.kaggle.com/rejpalcz/datagenerator-for-fast-data-loading</a>, replacing Image.open for cv2.imread and both tests for 128 images takes similar time (around 3.5 seconds). (I didn't commit changes, but you can check with private fork)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 404479,
          "author_name": "maw501",
          "author_url": "",
          "post_date": "10/15/2018 19:26:52",
          "content": "<p>Hi Michal, I haven't tried this for this competition yet as I'm working on TGS but plan to take a look next week hopefully.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 416987,
      "author_name": "dstjhb",
      "author_url": "",
      "post_date": "11/07/2018 15:11:33",
      "content": "<p>You can keep numpy arrays entirely on the hdd and load slices required to GPU as needed by specifying mmap_mode in np.load assuming you have saved the data as npy files. The initial process of creating these files will probably have required the use of your RAM to initially load the images and re-save them, but you can do this into multiple files if necessary.... if that helps at all?</p>\n\n<p>I guess it doesn't help with speed though...</p>",
      "votes": null,
      "replies": [
        {
          "id": 417022,
          "author_name": "danmoller",
          "author_url": "",
          "post_date": "11/07/2018 16:23:21",
          "content": "<p>If you load big chunks, then from the chunks gather many batches in a generator, it could be interesting, as long as you have at least two workers, one for loading a chunk while the other produces batches. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 417114,
      "author_name": "sakvaua",
      "author_url": "",
      "post_date": "11/07/2018 19:21:39",
      "content": "<p>You should also test different approaches when using multithreading.  You can get quite a bit of boost from multithreaded operations if underlying functions release GIL.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 417229,
      "author_name": "kokecacao",
      "author_url": "",
      "post_date": "11/08/2018 00:58:32",
      "content": "<p>Load images using .npy files speed up the training process; it takes only 70% of the original time needed.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 417252,
      "author_name": "hortonhearsafoo",
      "author_url": "",
      "post_date": "11/08/2018 01:48:34",
      "content": "<p>You should check out joblib for saving and loading numpy arrays</p>",
      "votes": null,
      "replies": [
        {
          "id": 417254,
          "author_name": "hortonhearsafoo",
          "author_url": "",
          "post_date": "11/08/2018 01:56:11",
          "content": "<p>More info on performance here, though the benchmarks might be dated: <a href=\"http://gael-varoquaux.info/programming/new_low-overhead_persistence_in_joblib_for_big_data.html\">http://gael-varoquaux.info/programming/new_low-overhead_persistence_in_joblib_for_big_data.html</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "401173": "It seems to me, that one of the key problems in this task is handling the loading/processing of the input data.\n\nNumpy array of size (1, 512, 512, 4) as float32 takes around 4 MB of memory, which means you would need ~128 GB of RAM just to fit the whole training set in it.\n\nAlso, DataGenerators I've seen in Kernels so far, seem to repeat the time consuming Image reading over and over (maybe I'm missing something, I never used any DG before).\n\nI tested several approaches and this is what I found so far:\n\n 1. PIL.Image seems to be fastest for image loading to numpy arrays\n 2. np.savez_compressed seems to be fastest for numpy arrays writing and reading (h5py, np.save and np.savez were slower)\n 3. Reading preprocessed and saved numpy array from np.savez_compressed is ~10 times faster than reading image using PIL.Image\n 4. Time of reading and writing numpy array using np.savez_compressed increases ~linearly with the number of images in it.\n\nMost promising strategy for dataloading seems to me to process the dataset once, divided into many batches (preferably using CPU multiprocessing), save it to hard drive and modify DataGenerator to read input data from these preprocessed files. And since I'm just a beginner, it's much easier to be said than to be done.\nIt also has the obvious setback, that to change batches (e.g. because of different seed), you need to repeat the whole processing task.\n\nI will appreciate any insights for this problem. I also appologize in case I'm stating something very obvious.",
    "401178": "It wasn’t so bad. Nevertheless, I wrote a custom Keras generator. I also enabled multiprocessing (with 16 processes) and created a large queue (256 batches).",
    "401297": "Initially with Keras I had issues with performance. The key is to have enough workers, the default of 1 is too slow. I also have the input images on a SSD drive.\n\n<pre>model = ModelMGPU(orig_model , 2)\nmodel.compile(optimizer=opt, loss=brian_loss, metrics=metrics)\n\nmodel_name = 'hpa_sub_4.h5'\nprint(\"LOAD: \" + model_name)\nmodel.load_weights(model_name,reshape=True)\n\ntrain_gen.batch_size = BATCH_SIZE\nresults = model.fit_generator(train_gen, \n                                steps_per_epoch = train_gen.samples//BATCH_SIZE,\n                                validation_data = (valid_x, valid_y), \n                                epochs = 1000, \n                                callbacks=[ reduce_lr,checkpointer,metrics_cb],\n                                workers=24)\n</pre>",
    "401470": "One thing that I did notice is that if I had too many transformations on the generator then the training becomes cpu limited.",
    "402983": "OpenCV is faster than PIL for reading in images with PyTorch.",
    "404324": "Maybe there could be version and system specific issues, but I also have better results with open cv.",
    "404425": "Maybe I'm doing something wrong, but I cannot confirm. Tried on my kernel here: https://www.kaggle.com/rejpalcz/datagenerator-for-fast-data-loading, replacing Image.open for cv2.imread and both tests for 128 images takes similar time (around 3.5 seconds). (I didn't commit changes, but you can check with private fork)",
    "404479": "Hi Michal, I haven't tried this for this competition yet as I'm working on TGS but plan to take a look next week hopefully.",
    "416987": "You can keep numpy arrays entirely on the hdd and load slices required to GPU as needed by specifying mmap_mode in np.load assuming you have saved the data as npy files. The initial process of creating these files will probably have required the use of your RAM to initially load the images and re-save them, but you can do this into multiple files if necessary.... if that helps at all?\n\nI guess it doesn't help with speed though...",
    "417022": "If you load big chunks, then from the chunks gather many batches in a generator, it could be interesting, as long as you have at least two workers, one for loading a chunk while the other produces batches.",
    "417114": "You should also test different approaches when using multithreading.  You can get quite a bit of boost from multithreaded operations if underlying functions release GIL.",
    "417229": "Load images using .npy files speed up the training process; it takes only 70% of the original time needed.",
    "417252": "You should check out joblib for saving and loading numpy arrays",
    "417254": "More info on performance here, though the benchmarks might be dated: http://gael-varoquaux.info/programming/new_low-overhead_persistence_in_joblib_for_big_data.html"
  },
  "source": "meta"
}