{
  "id": 104121,
  "title": "Slow training process",
  "url": "/competitions/aptos2019-blindness-detection/discussion/104121",
  "author_name": "",
  "post_date": "2019-08-14T13:10:10.923982300Z",
  "votes": 3,
  "comment_count": 27,
  "views": 0,
  "content": "<p>I'm using keras to design the model and fitting a dataset from a generator.</p>\n\n<p>The base model is RestNet50, without top.</p>\n\n<p>The generator is an extension of Keras Sequence and does (using multiprocessing.Pool) :\n-  read a batch of images (using opencv)\n- apply some preprocessing such as resize (using opencv)\n- augment the images using the Imgaug package with (a few transformations)</p>\n\n<p>The problem is it takes 800s to train for one epoch on kaggle with gpu on, no matter how many layers are frozen in the ResNet50 model.</p>\n\n<p>The number of steps per epoch is defined as the length of the generator.</p>\n\n<p>I would appreciate if someone has an idea of why the process takes so much time.</p>\n\n<p>Watching public kernels, I can see that with approximately the same configuration, the running time for one epoch should be about 120s.</p>",
  "messages": [
    {
      "id": "599074",
      "postDate": "08/14/2019 13:10:10",
      "content": "<p>I'm using keras to design the model and fitting a dataset from a generator.</p>\n\n<p>The base model is RestNet50, without top.</p>\n\n<p>The generator is an extension of Keras Sequence and does (using multiprocessing.Pool) :\n-  read a batch of images (using opencv)\n- apply some preprocessing such as resize (using opencv)\n- augment the images using the Imgaug package with (a few transformations)</p>\n\n<p>The problem is it takes 800s to train for one epoch on kaggle with gpu on, no matter how many layers are frozen in the ResNet50 model.</p>\n\n<p>The number of steps per epoch is defined as the length of the generator.</p>\n\n<p>I would appreciate if someone has an idea of why the process takes so much time.</p>\n\n<p>Watching public kernels, I can see that with approximately the same configuration, the running time for one epoch should be about 120s.</p>",
      "rawMarkdown": "I'm using keras to design the model and fitting a dataset from a generator.\n\nThe base model is RestNet50, without top.\n\nThe generator is an extension of Keras Sequence and does (using multiprocessing.Pool) :\n-  read a batch of images (using opencv)\n- apply some preprocessing such as resize (using opencv)\n- augment the images using the Imgaug package with (a few transformations)\n\nThe problem is it takes 800s to train for one epoch on kaggle with gpu on, no matter how many layers are frozen in the ResNet50 model.\n\nThe number of steps per epoch is defined as the length of the generator.\n\nI would appreciate if someone has an idea of why the process takes so much time.\n\nWatching public kernels, I can see that with approximately the same configuration, the running time for one epoch should be about 120s.",
      "votes": null
    },
    {
      "id": "599080",
      "postDate": "08/14/2019 13:15:50",
      "content": "<p>Here is how you can speed up:</p>\n\n<p>1) Create folder for train and test \n2) Resize all the images in to these folders \n3) Run your model using resized images\n4) After you are done at the end of notebook remove folders that you have created\n5) Win kaggle gold.</p>",
      "rawMarkdown": "Here is how you can speed up:\n\n\n1) Create folder for train and test \n2) Resize all the images in to these folders \n3) Run your model using resized images\n4) After you are done at the end of notebook remove folders that you have created\n5) Win kaggle gold.",
      "votes": null
    },
    {
      "id": "599087",
      "postDate": "08/14/2019 13:29:10",
      "content": "<p>Thank you very much for your response.</p>\n\n<p>I am not sure whether I understood all the steps correctly.</p>\n\n<p>Before step 3) fitting the model, should I try to open again all the resized images in RAM and then apply augmentation ?</p>",
      "rawMarkdown": "Thank you very much for your response.\n\nI am not sure whether I understood all the steps correctly.\n\nBefore step 3) fitting the model, should I try to open again all the resized images in RAM and then apply augmentation ?",
      "votes": null
    },
    {
      "id": "599092",
      "postDate": "08/14/2019 13:38:25",
      "content": "<p>Yes, Time limiting step is resizing every time images. After you resize images you can train as you want with all the augmentation =)  Make sure to remove folder that contain images. </p>",
      "rawMarkdown": "Yes, Time limiting step is resizing every time images. After you resize images you can train as you want with all the augmentation =)  Make sure to remove folder that contain images.",
      "votes": null
    },
    {
      "id": "599098",
      "postDate": "08/14/2019 13:41:21",
      "content": "<p>Thank you, that makes sense :)</p>",
      "rawMarkdown": "Thank you, that makes sense :)",
      "votes": null
    },
    {
      "id": "599127",
      "postDate": "08/14/2019 14:20:06",
      "content": "<p>For other people that are in my previous situation DrHB's tip allowed me to decrease an epoch running time by a factor of 10 :)</p>",
      "rawMarkdown": "For other people that are in my previous situation DrHB's tip allowed me to decrease an epoch running time by a factor of 10 :)",
      "votes": null
    },
    {
      "id": "599201",
      "postDate": "08/14/2019 16:43:27",
      "content": "<p>Is there enough disk space to also dump resized full test?</p>",
      "rawMarkdown": "Is there enough disk space to also dump resized full test?",
      "votes": null
    },
    {
      "id": "599205",
      "postDate": "08/14/2019 16:57:08",
      "content": "<p>I resized the images in the training path to have a shape(224, 224, 3) and saved them to the working directory and I still had about 3GB free space on the disk. So I assume it should be fine to dump the test set, as long as the image's shape is not too big.</p>",
      "rawMarkdown": "I resized the images in the training path to have a shape(224, 224, 3) and saved them to the working directory and I still had about 3GB free space on the disk. So I assume it should be fine to dump the test set, as long as the image's shape is not too big.",
      "votes": null
    },
    {
      "id": "599212",
      "postDate": "08/14/2019 17:00:16",
      "content": "<p>the full test set is much larger though, but for just the public test it should be enough indeed</p>",
      "rawMarkdown": "the full test set is much larger though, but for just the public test it should be enough indeed",
      "votes": null
    },
    {
      "id": "599224",
      "postDate": "08/14/2019 17:22:14",
      "content": "<p>With the same image dimension I said the preprocessed images of the training set takes up 272Mb and consists of 3362 images. From kaggle's data description the private data set contains about 13'000 images.</p>\n\n<p>272[MB]/3662 * 13000 ~= 966MB</p>\n\n<p>And we have 4.9 GB free space so it should be ok.</p>",
      "rawMarkdown": "With the same image dimension I said the preprocessed images of the training set takes up 272Mb and consists of 3362 images. From kaggle's data description the private data set contains about 13'000 images.\n\n272[MB]/3662 * 13000 ~= 966MB\n\nAnd we have 4.9 GB free space so it should be ok.",
      "votes": null
    },
    {
      "id": "599456",
      "postDate": "08/15/2019 01:49:34",
      "content": "<p>You also can use the method in this kernel.\n<a href=\"https://www.kaggle.com/nomadista/large-training-speed-boost\">https://www.kaggle.com/nomadista/large-training-speed-boost</a></p>",
      "rawMarkdown": "You also can use the method in this kernel.\nhttps://www.kaggle.com/nomadista/large-training-speed-boost",
      "votes": null
    },
    {
      "id": "599473",
      "postDate": "08/15/2019 02:47:15",
      "content": "<p>Do   you  use   Image.save()    or  cv2.Save()   directly  without any  parameter?  It  decrease  the    size   but    in  the   meanwhile   it  will    lost  some  information  in the  picture.</p>",
      "rawMarkdown": "Do   you  use   Image.save()    or  cv2.Save()   directly  without any  parameter?  It  decrease  the    size   but    in  the   meanwhile   it  will    lost  some  information  in the  picture.",
      "votes": null
    },
    {
      "id": "599686",
      "postDate": "08/15/2019 08:53:33",
      "content": "<p>Thank you I was not aware of that, I used cv2.imwrite(), but I saved them as png files, isn't png a lossless format ?</p>",
      "rawMarkdown": "Thank you I was not aware of that, I used cv2.imwrite(), but I saved them as png files, isn't png a lossless format ?",
      "votes": null
    },
    {
      "id": "599687",
      "postDate": "08/15/2019 08:54:54",
      "content": "<p>Thank you, I will have a look</p>",
      "rawMarkdown": "Thank you, I will have a look",
      "votes": null
    },
    {
      "id": "600256",
      "postDate": "08/15/2019 21:00:08",
      "content": "<p>Thanks, I tried this one and it works really well - 15x speedup per epoch in my case. It's essentially the same solution as DrHB proposes, but more complicated using a HDF5 file. However, it should handle a larger amount of images, which is especially useful if you are using a lot of training images (10000+).</p>\n\n<p>As a sidenote, it's only necessary to do this for test images if you're using some kind of TTA. Otherwise you will just do the resize one time anyway :)</p>",
      "rawMarkdown": "Thanks, I tried this one and it works really well - 15x speedup per epoch in my case. It's essentially the same solution as DrHB proposes, but more complicated using a HDF5 file. However, it should handle a larger amount of images, which is especially useful if you are using a lot of training images (10000+).\n\nAs a sidenote, it's only necessary to do this for test images if you're using some kind of TTA. Otherwise you will just do the resize one time anyway :)",
      "votes": null
    },
    {
      "id": "600326",
      "postDate": "08/16/2019 01:46:57",
      "content": "<p>Yes,PNG  should  have  been  a lossless  format ,but  it  occur  some  difference in  my  kernel ,I will  inform  you   if  I  have  any   new  thoughts.</p>",
      "rawMarkdown": "Yes,PNG  should  have  been  a lossless  format ,but  it  occur  some  difference in  my  kernel ,I will  inform  you   if  I  have  any   new  thoughts.",
      "votes": null
    },
    {
      "id": "601492",
      "postDate": "08/17/2019 17:14:49",
      "content": "<p>I am trying to first save the images, then use them but I am seeing that saving takes 8 gb of RAM. Does \nthe same thing happen to you?</p>",
      "rawMarkdown": "I am trying to first save the images, then use them but I am seeing that saving takes 8 gb of RAM. Does \nthe same thing happen to you?",
      "votes": null
    },
    {
      "id": "601501",
      "postDate": "08/17/2019 17:48:49",
      "content": "<p>The problem might be that you open all the images at once. In my case it does not take up that much RAM. It it helps, this is the code I'm using :</p>\n\n<p><code>def save_preprocessed(df, img_dir, prep_dir):\n     for img_path in tqdm(df[\"id_code\"].values):\n        img = preprocess_aptos(os.path.join(img_dir, img_path), IMG_SIZE)\n        cv2.imwrite(os.path.join(prep_dir, img_path), img)</code></p>\n\n<p>with preprocess_aptos being a function where I read an image, resize it and return it.</p>\n\n<p>If your problem occurs when you open your preprocessed images, you should try using a generator.</p>\n\n<p>(Sorry, I can't figure out how to add tabulations to the code sample)</p>",
      "rawMarkdown": "The problem might be that you open all the images at once. In my case it does not take up that much RAM. It it helps, this is the code I'm using :\n\n```def save_preprocessed(df, img_dir, prep_dir):\n     for img_path in tqdm(df[\"id_code\"].values):\n        img = preprocess_aptos(os.path.join(img_dir, img_path), IMG_SIZE)\n        cv2.imwrite(os.path.join(prep_dir, img_path), img)```\n\nwith preprocess_aptos being a function where I read an image, resize it and return it.\n\nIf your problem occurs when you open your preprocessed images, you should try using a generator.\n\n(Sorry, I can't figure out how to add tabulations to the code sample)",
      "votes": null
    },
    {
      "id": "601540",
      "postDate": "08/17/2019 19:07:06",
      "content": "<p>I am also using a similar code structure. </p>\n\n<p>```\ndef save_processed_image(image_name, size, directory_name_to_read, directory_name_to_store, extension):\n    path = os.path.join(directory_name_to_read,image_name+extension)\n    img = preprocess_v1(path,size)\n    # saving the img\n    save_path = os.path.join(directory_name_to_store,image_name+extension)\n    cv2.imwrite(save_path,img)</p>\n\n<pre><code># print(\"read at\",path,\"saved at\",save_path)\ngc.enable()\ndel img\ngc.collect()\n</code></pre>\n\n<p>def save_processed_data(df,img_dir,store_dir):\n    Parallel(n_jobs=4,backend=\"multiprocessing\")(\n                            delayed(save_processed_image)(name, n_pixels, img_dir, store_dir, extension) \n                            for name in df[id_column].values)\n```</p>\n\n<p>I don't know why it's taking so much of RAM and also if you don't mind can you tell me how much time it takes for preprocessing and saving in your case. In my case, it takes nearly 19 minutes for train data</p>",
      "rawMarkdown": "I am also using a similar code structure. \n\n```\ndef save_processed_image(image_name, size, directory_name_to_read, directory_name_to_store, extension):\n    path = os.path.join(directory_name_to_read,image_name+extension)\n    img = preprocess_v1(path,size)\n    # saving the img\n    save_path = os.path.join(directory_name_to_store,image_name+extension)\n    cv2.imwrite(save_path,img)\n    \n    # print(\"read at\",path,\"saved at\",save_path)\n    gc.enable()\n    del img\n    gc.collect()\n\ndef save_processed_data(df,img_dir,store_dir):\n    Parallel(n_jobs=4,backend=\"multiprocessing\")(\n                            delayed(save_processed_image)(name, n_pixels, img_dir, store_dir, extension) \n                            for name in df[id_column].values)\n```\n\nI don't know why it's taking so much of RAM and also if you don't mind can you tell me how much time it takes for preprocessing and saving in your case. In my case, it takes nearly 19 minutes for train data",
      "votes": null
    },
    {
      "id": "601545",
      "postDate": "08/17/2019 19:15:46",
      "content": "<p>For the training set it takes me exactly 19 minutes aswell, I am not sure why it takes up so much RAM for you. The only obvious difference I see is that I don't use any multiprocessing.</p>",
      "rawMarkdown": "For the training set it takes me exactly 19 minutes aswell, I am not sure why it takes up so much RAM for you. The only obvious difference I see is that I don't use any multiprocessing.",
      "votes": null
    },
    {
      "id": "601546",
      "postDate": "08/17/2019 19:22:19",
      "content": "<p>I tried with for loop also it takes nearly 20 min and also the RAM usage is 8 GB.</p>",
      "rawMarkdown": "I tried with for loop also it takes nearly 20 min and also the RAM usage is 8 GB.",
      "votes": null
    },
    {
      "id": "601550",
      "postDate": "08/17/2019 19:29:14",
      "content": "<p>Another thought, is to restart your kernel provided you don't losse your work. Maybe the RAM was taken from earlier code execution.</p>",
      "rawMarkdown": "Another thought, is to restart your kernel provided you don't losse your work. Maybe the RAM was taken from earlier code execution.",
      "votes": null
    },
    {
      "id": "601551",
      "postDate": "08/17/2019 19:30:51",
      "content": "<p>I tried that but no use.</p>",
      "rawMarkdown": "I tried that but no use.",
      "votes": null
    },
    {
      "id": "601584",
      "postDate": "08/17/2019 20:47:40",
      "content": "<p>I've had some problems with RAM too, it seems to be a bug with the kernel not releasing RAM after reading an image. Even if the image is deleted, dereferenced and whatnot, the RAM usage keeps rising for each image. However I'm not sure if the high RAM usage is real or just an interface bug, since I managed to reach max RAM without anything crashing. I made a post about it here: <a href=\"https://www.kaggle.com/product-feedback/104464\">https://www.kaggle.com/product-feedback/104464</a> </p>\n\n<p>As for a solution, it is possible to do the preprocessing in another kernel, zip the images and use the zip to create a dataset that can be accessed by your main kernel.</p>",
      "rawMarkdown": "I've had some problems with RAM too, it seems to be a bug with the kernel not releasing RAM after reading an image. Even if the image is deleted, dereferenced and whatnot, the RAM usage keeps rising for each image. However I'm not sure if the high RAM usage is real or just an interface bug, since I managed to reach max RAM without anything crashing. I made a post about it here: https://www.kaggle.com/product-feedback/104464 \n\nAs for a solution, it is possible to do the preprocessing in another kernel, zip the images and use the zip to create a dataset that can be accessed by your main kernel.",
      "votes": null
    },
    {
      "id": "601586",
      "postDate": "08/17/2019 20:59:06",
      "content": "<p>Hi! <a href=\"/jonasfreibs\">@jonasfreibs</a>, for me finding the optimal image size, was critical in accelerating the speed of my training and using Image Generators from Keras / Tensorflow </p>",
      "rawMarkdown": "Hi! @jonasfreibs, for me finding the optimal image size, was critical in accelerating the speed of my training and using Image Generators from Keras / Tensorflow",
      "votes": null
    },
    {
      "id": "601715",
      "postDate": "08/18/2019 04:48:23",
      "content": "<p>I have made a <a href=\"https://www.kaggle.com/naivelamb/process-test-images-parallelly\">notebook</a> to show how to preprocessing test images using multiple CPUs. </p>\n\n<p>It can be combined with your inference notebook to accelerate the inference time. With small modifications you can process the train data locally (or using kernel?). </p>",
      "rawMarkdown": "I have made a [notebook](https://www.kaggle.com/naivelamb/process-test-images-parallelly) to show how to preprocessing test images using multiple CPUs. \n\nIt can be combined with your inference notebook to accelerate the inference time. With small modifications you can process the train data locally (or using kernel?).",
      "votes": null
    },
    {
      "id": "601983",
      "postDate": "08/18/2019 12:45:22",
      "content": "<p>Hi <a href=\"/cv13j0\">@cv13j0</a>, thank you, yes it does speed up the process.</p>",
      "rawMarkdown": "Hi @cv13j0, thank you, yes it does speed up the process.",
      "votes": null
    },
    {
      "id": "601991",
      "postDate": "08/18/2019 12:46:47",
      "content": "<p>Awesome, thank you</p>",
      "rawMarkdown": "Awesome, thank you",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 599080,
      "author_name": "drhabib",
      "author_url": "",
      "post_date": "08/14/2019 13:15:50",
      "content": "<p>Here is how you can speed up:</p>\n\n<p>1) Create folder for train and test \n2) Resize all the images in to these folders \n3) Run your model using resized images\n4) After you are done at the end of notebook remove folders that you have created\n5) Win kaggle gold.</p>",
      "votes": null,
      "replies": [
        {
          "id": 599087,
          "author_name": "jonasfreibs",
          "author_url": "",
          "post_date": "08/14/2019 13:29:10",
          "content": "<p>Thank you very much for your response.</p>\n\n<p>I am not sure whether I understood all the steps correctly.</p>\n\n<p>Before step 3) fitting the model, should I try to open again all the resized images in RAM and then apply augmentation ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 599092,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "08/14/2019 13:38:25",
          "content": "<p>Yes, Time limiting step is resizing every time images. After you resize images you can train as you want with all the augmentation =)  Make sure to remove folder that contain images. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 599098,
          "author_name": "jonasfreibs",
          "author_url": "",
          "post_date": "08/14/2019 13:41:21",
          "content": "<p>Thank you, that makes sense :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 599201,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "08/14/2019 16:43:27",
          "content": "<p>Is there enough disk space to also dump resized full test?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 599205,
          "author_name": "jonasfreibs",
          "author_url": "",
          "post_date": "08/14/2019 16:57:08",
          "content": "<p>I resized the images in the training path to have a shape(224, 224, 3) and saved them to the working directory and I still had about 3GB free space on the disk. So I assume it should be fine to dump the test set, as long as the image's shape is not too big.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 599212,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "08/14/2019 17:00:16",
          "content": "<p>the full test set is much larger though, but for just the public test it should be enough indeed</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 599224,
          "author_name": "jonasfreibs",
          "author_url": "",
          "post_date": "08/14/2019 17:22:14",
          "content": "<p>With the same image dimension I said the preprocessed images of the training set takes up 272Mb and consists of 3362 images. From kaggle's data description the private data set contains about 13'000 images.</p>\n\n<p>272[MB]/3662 * 13000 ~= 966MB</p>\n\n<p>And we have 4.9 GB free space so it should be ok.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 599473,
          "author_name": "xujingzhao",
          "author_url": "",
          "post_date": "08/15/2019 02:47:15",
          "content": "<p>Do   you  use   Image.save()    or  cv2.Save()   directly  without any  parameter?  It  decrease  the    size   but    in  the   meanwhile   it  will    lost  some  information  in the  picture.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 599686,
          "author_name": "jonasfreibs",
          "author_url": "",
          "post_date": "08/15/2019 08:53:33",
          "content": "<p>Thank you I was not aware of that, I used cv2.imwrite(), but I saved them as png files, isn't png a lossless format ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 600326,
          "author_name": "xujingzhao",
          "author_url": "",
          "post_date": "08/16/2019 01:46:57",
          "content": "<p>Yes,PNG  should  have  been  a lossless  format ,but  it  occur  some  difference in  my  kernel ,I will  inform  you   if  I  have  any   new  thoughts.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 599127,
      "author_name": "jonasfreibs",
      "author_url": "",
      "post_date": "08/14/2019 14:20:06",
      "content": "<p>For other people that are in my previous situation DrHB's tip allowed me to decrease an epoch running time by a factor of 10 :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 601492,
          "author_name": "suchith0312",
          "author_url": "",
          "post_date": "08/17/2019 17:14:49",
          "content": "<p>I am trying to first save the images, then use them but I am seeing that saving takes 8 gb of RAM. Does \nthe same thing happen to you?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 601501,
          "author_name": "jonasfreibs",
          "author_url": "",
          "post_date": "08/17/2019 17:48:49",
          "content": "<p>The problem might be that you open all the images at once. In my case it does not take up that much RAM. It it helps, this is the code I'm using :</p>\n\n<p><code>def save_preprocessed(df, img_dir, prep_dir):\n     for img_path in tqdm(df[\"id_code\"].values):\n        img = preprocess_aptos(os.path.join(img_dir, img_path), IMG_SIZE)\n        cv2.imwrite(os.path.join(prep_dir, img_path), img)</code></p>\n\n<p>with preprocess_aptos being a function where I read an image, resize it and return it.</p>\n\n<p>If your problem occurs when you open your preprocessed images, you should try using a generator.</p>\n\n<p>(Sorry, I can't figure out how to add tabulations to the code sample)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 601540,
          "author_name": "suchith0312",
          "author_url": "",
          "post_date": "08/17/2019 19:07:06",
          "content": "<p>I am also using a similar code structure. </p>\n\n<p>```\ndef save_processed_image(image_name, size, directory_name_to_read, directory_name_to_store, extension):\n    path = os.path.join(directory_name_to_read,image_name+extension)\n    img = preprocess_v1(path,size)\n    # saving the img\n    save_path = os.path.join(directory_name_to_store,image_name+extension)\n    cv2.imwrite(save_path,img)</p>\n\n<pre><code># print(\"read at\",path,\"saved at\",save_path)\ngc.enable()\ndel img\ngc.collect()\n</code></pre>\n\n<p>def save_processed_data(df,img_dir,store_dir):\n    Parallel(n_jobs=4,backend=\"multiprocessing\")(\n                            delayed(save_processed_image)(name, n_pixels, img_dir, store_dir, extension) \n                            for name in df[id_column].values)\n```</p>\n\n<p>I don't know why it's taking so much of RAM and also if you don't mind can you tell me how much time it takes for preprocessing and saving in your case. In my case, it takes nearly 19 minutes for train data</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 601545,
          "author_name": "jonasfreibs",
          "author_url": "",
          "post_date": "08/17/2019 19:15:46",
          "content": "<p>For the training set it takes me exactly 19 minutes aswell, I am not sure why it takes up so much RAM for you. The only obvious difference I see is that I don't use any multiprocessing.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 601546,
          "author_name": "suchith0312",
          "author_url": "",
          "post_date": "08/17/2019 19:22:19",
          "content": "<p>I tried with for loop also it takes nearly 20 min and also the RAM usage is 8 GB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 601550,
          "author_name": "jonasfreibs",
          "author_url": "",
          "post_date": "08/17/2019 19:29:14",
          "content": "<p>Another thought, is to restart your kernel provided you don't losse your work. Maybe the RAM was taken from earlier code execution.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 601551,
          "author_name": "suchith0312",
          "author_url": "",
          "post_date": "08/17/2019 19:30:51",
          "content": "<p>I tried that but no use.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 601584,
          "author_name": "larlia",
          "author_url": "",
          "post_date": "08/17/2019 20:47:40",
          "content": "<p>I've had some problems with RAM too, it seems to be a bug with the kernel not releasing RAM after reading an image. Even if the image is deleted, dereferenced and whatnot, the RAM usage keeps rising for each image. However I'm not sure if the high RAM usage is real or just an interface bug, since I managed to reach max RAM without anything crashing. I made a post about it here: <a href=\"https://www.kaggle.com/product-feedback/104464\">https://www.kaggle.com/product-feedback/104464</a> </p>\n\n<p>As for a solution, it is possible to do the preprocessing in another kernel, zip the images and use the zip to create a dataset that can be accessed by your main kernel.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 599456,
      "author_name": "chanhu",
      "author_url": "",
      "post_date": "08/15/2019 01:49:34",
      "content": "<p>You also can use the method in this kernel.\n<a href=\"https://www.kaggle.com/nomadista/large-training-speed-boost\">https://www.kaggle.com/nomadista/large-training-speed-boost</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 599687,
          "author_name": "jonasfreibs",
          "author_url": "",
          "post_date": "08/15/2019 08:54:54",
          "content": "<p>Thank you, I will have a look</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 600256,
          "author_name": "larlia",
          "author_url": "",
          "post_date": "08/15/2019 21:00:08",
          "content": "<p>Thanks, I tried this one and it works really well - 15x speedup per epoch in my case. It's essentially the same solution as DrHB proposes, but more complicated using a HDF5 file. However, it should handle a larger amount of images, which is especially useful if you are using a lot of training images (10000+).</p>\n\n<p>As a sidenote, it's only necessary to do this for test images if you're using some kind of TTA. Otherwise you will just do the resize one time anyway :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 601586,
      "author_name": "cv13j0",
      "author_url": "",
      "post_date": "08/17/2019 20:59:06",
      "content": "<p>Hi! <a href=\"/jonasfreibs\">@jonasfreibs</a>, for me finding the optimal image size, was critical in accelerating the speed of my training and using Image Generators from Keras / Tensorflow </p>",
      "votes": null,
      "replies": [
        {
          "id": 601983,
          "author_name": "jonasfreibs",
          "author_url": "",
          "post_date": "08/18/2019 12:45:22",
          "content": "<p>Hi <a href=\"/cv13j0\">@cv13j0</a>, thank you, yes it does speed up the process.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 601715,
      "author_name": "naivelamb",
      "author_url": "",
      "post_date": "08/18/2019 04:48:23",
      "content": "<p>I have made a <a href=\"https://www.kaggle.com/naivelamb/process-test-images-parallelly\">notebook</a> to show how to preprocessing test images using multiple CPUs. </p>\n\n<p>It can be combined with your inference notebook to accelerate the inference time. With small modifications you can process the train data locally (or using kernel?). </p>",
      "votes": null,
      "replies": [
        {
          "id": 601991,
          "author_name": "jonasfreibs",
          "author_url": "",
          "post_date": "08/18/2019 12:46:47",
          "content": "<p>Awesome, thank you</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "599074": "I'm using keras to design the model and fitting a dataset from a generator.\n\nThe base model is RestNet50, without top.\n\nThe generator is an extension of Keras Sequence and does (using multiprocessing.Pool) :\n-  read a batch of images (using opencv)\n- apply some preprocessing such as resize (using opencv)\n- augment the images using the Imgaug package with (a few transformations)\n\nThe problem is it takes 800s to train for one epoch on kaggle with gpu on, no matter how many layers are frozen in the ResNet50 model.\n\nThe number of steps per epoch is defined as the length of the generator.\n\nI would appreciate if someone has an idea of why the process takes so much time.\n\nWatching public kernels, I can see that with approximately the same configuration, the running time for one epoch should be about 120s.",
    "599080": "Here is how you can speed up:\n\n\n1) Create folder for train and test \n2) Resize all the images in to these folders \n3) Run your model using resized images\n4) After you are done at the end of notebook remove folders that you have created\n5) Win kaggle gold.",
    "599087": "Thank you very much for your response.\n\nI am not sure whether I understood all the steps correctly.\n\nBefore step 3) fitting the model, should I try to open again all the resized images in RAM and then apply augmentation ?",
    "599092": "Yes, Time limiting step is resizing every time images. After you resize images you can train as you want with all the augmentation =)  Make sure to remove folder that contain images.",
    "599098": "Thank you, that makes sense :)",
    "599127": "For other people that are in my previous situation DrHB's tip allowed me to decrease an epoch running time by a factor of 10 :)",
    "599201": "Is there enough disk space to also dump resized full test?",
    "599205": "I resized the images in the training path to have a shape(224, 224, 3) and saved them to the working directory and I still had about 3GB free space on the disk. So I assume it should be fine to dump the test set, as long as the image's shape is not too big.",
    "599212": "the full test set is much larger though, but for just the public test it should be enough indeed",
    "599224": "With the same image dimension I said the preprocessed images of the training set takes up 272Mb and consists of 3362 images. From kaggle's data description the private data set contains about 13'000 images.\n\n272[MB]/3662 * 13000 ~= 966MB\n\nAnd we have 4.9 GB free space so it should be ok.",
    "599456": "You also can use the method in this kernel.\nhttps://www.kaggle.com/nomadista/large-training-speed-boost",
    "599473": "Do   you  use   Image.save()    or  cv2.Save()   directly  without any  parameter?  It  decrease  the    size   but    in  the   meanwhile   it  will    lost  some  information  in the  picture.",
    "599686": "Thank you I was not aware of that, I used cv2.imwrite(), but I saved them as png files, isn't png a lossless format ?",
    "599687": "Thank you, I will have a look",
    "600256": "Thanks, I tried this one and it works really well - 15x speedup per epoch in my case. It's essentially the same solution as DrHB proposes, but more complicated using a HDF5 file. However, it should handle a larger amount of images, which is especially useful if you are using a lot of training images (10000+).\n\nAs a sidenote, it's only necessary to do this for test images if you're using some kind of TTA. Otherwise you will just do the resize one time anyway :)",
    "600326": "Yes,PNG  should  have  been  a lossless  format ,but  it  occur  some  difference in  my  kernel ,I will  inform  you   if  I  have  any   new  thoughts.",
    "601492": "I am trying to first save the images, then use them but I am seeing that saving takes 8 gb of RAM. Does \nthe same thing happen to you?",
    "601501": "The problem might be that you open all the images at once. In my case it does not take up that much RAM. It it helps, this is the code I'm using :\n\n```def save_preprocessed(df, img_dir, prep_dir):\n     for img_path in tqdm(df[\"id_code\"].values):\n        img = preprocess_aptos(os.path.join(img_dir, img_path), IMG_SIZE)\n        cv2.imwrite(os.path.join(prep_dir, img_path), img)```\n\nwith preprocess_aptos being a function where I read an image, resize it and return it.\n\nIf your problem occurs when you open your preprocessed images, you should try using a generator.\n\n(Sorry, I can't figure out how to add tabulations to the code sample)",
    "601540": "I am also using a similar code structure. \n\n```\ndef save_processed_image(image_name, size, directory_name_to_read, directory_name_to_store, extension):\n    path = os.path.join(directory_name_to_read,image_name+extension)\n    img = preprocess_v1(path,size)\n    # saving the img\n    save_path = os.path.join(directory_name_to_store,image_name+extension)\n    cv2.imwrite(save_path,img)\n    \n    # print(\"read at\",path,\"saved at\",save_path)\n    gc.enable()\n    del img\n    gc.collect()\n\ndef save_processed_data(df,img_dir,store_dir):\n    Parallel(n_jobs=4,backend=\"multiprocessing\")(\n                            delayed(save_processed_image)(name, n_pixels, img_dir, store_dir, extension) \n                            for name in df[id_column].values)\n```\n\nI don't know why it's taking so much of RAM and also if you don't mind can you tell me how much time it takes for preprocessing and saving in your case. In my case, it takes nearly 19 minutes for train data",
    "601545": "For the training set it takes me exactly 19 minutes aswell, I am not sure why it takes up so much RAM for you. The only obvious difference I see is that I don't use any multiprocessing.",
    "601546": "I tried with for loop also it takes nearly 20 min and also the RAM usage is 8 GB.",
    "601550": "Another thought, is to restart your kernel provided you don't losse your work. Maybe the RAM was taken from earlier code execution.",
    "601551": "I tried that but no use.",
    "601584": "I've had some problems with RAM too, it seems to be a bug with the kernel not releasing RAM after reading an image. Even if the image is deleted, dereferenced and whatnot, the RAM usage keeps rising for each image. However I'm not sure if the high RAM usage is real or just an interface bug, since I managed to reach max RAM without anything crashing. I made a post about it here: https://www.kaggle.com/product-feedback/104464 \n\nAs for a solution, it is possible to do the preprocessing in another kernel, zip the images and use the zip to create a dataset that can be accessed by your main kernel.",
    "601586": "Hi! @jonasfreibs, for me finding the optimal image size, was critical in accelerating the speed of my training and using Image Generators from Keras / Tensorflow",
    "601715": "I have made a [notebook](https://www.kaggle.com/naivelamb/process-test-images-parallelly) to show how to preprocessing test images using multiple CPUs. \n\nIt can be combined with your inference notebook to accelerate the inference time. With small modifications you can process the train data locally (or using kernel?).",
    "601983": "Hi @cv13j0, thank you, yes it does speed up the process.",
    "601991": "Awesome, thank you"
  },
  "source": "meta"
}