{
  "id": 316410,
  "title": "What are your augmentations, and how fast?",
  "url": "/competitions/hotel-id-to-combat-human-trafficking-2022-fgvc9/discussion/316410",
  "author_name": "",
  "post_date": "2022-04-01T19:26:52.820757Z",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I am applying a lot of heavy augmentations to the training data and my batch size is now 128. The augmentation takes about 15-18 secs per batch. Is this normal? I am using pytorch for training. For augmentations, I am mainly using albumentations but among them a few key steps are in torchvision &amp; timm. Perhaps it's also related to how I read the images (I use cv2.imread()) into the data loader.</p>\n<p>What bothers me most is that the transformations from albumentations only take in numpy arrays. But in my notebook, those come after some torchvision transforms outputting tensors already.</p>\n<p>What general suggestions you would give for fast and efficient data augmentation? Perhaps which package(s) and what order should the transforms be in?</p>",
  "messages": [
    {
      "id": "1742372",
      "postDate": "04/01/2022 19:26:52",
      "content": "<p>I am applying a lot of heavy augmentations to the training data and my batch size is now 128. The augmentation takes about 15-18 secs per batch. Is this normal? I am using pytorch for training. For augmentations, I am mainly using albumentations but among them a few key steps are in torchvision &amp; timm. Perhaps it's also related to how I read the images (I use cv2.imread()) into the data loader.</p>\n<p>What bothers me most is that the transformations from albumentations only take in numpy arrays. But in my notebook, those come after some torchvision transforms outputting tensors already.</p>\n<p>What general suggestions you would give for fast and efficient data augmentation? Perhaps which package(s) and what order should the transforms be in?</p>",
      "rawMarkdown": "I am applying a lot of heavy augmentations to the training data and my batch size is now 128. The augmentation takes about 15-18 secs per batch. Is this normal? I am using pytorch for training. For augmentations, I am mainly using albumentations but among them a few key steps are in torchvision & timm. Perhaps it's also related to how I read the images (I use cv2.imread()) into the data loader.\n\nWhat bothers me most is that the transformations from albumentations only take in numpy arrays. But in my notebook, those come after some torchvision transforms outputting tensors already.\n\nWhat general suggestions you would give for fast and efficient data augmentation? Perhaps which package(s) and what order should the transforms be in?",
      "votes": null
    },
    {
      "id": "1743012",
      "postDate": "04/02/2022 13:59:25",
      "content": "<p>Well i would say that 15-18 secs per batch is not normal but it depends on what you do and on what size of images.</p>\n<p>There are some benchmarks of image loading libraries:<br>\n<a href=\"https://stackoverflow.com/questions/57663734/how-to-speed-up-image-loading-in-pillow-python\" target=\"_blank\">https://stackoverflow.com/questions/57663734/how-to-speed-up-image-loading-in-pillow-python</a><br>\n<a href=\"https://www.kaggle.com/code/zfturbo/benchmark-2019-speed-of-image-reading/notebook\" target=\"_blank\">https://www.kaggle.com/code/zfturbo/benchmark-2019-speed-of-image-reading/notebook</a><br>\n<a href=\"https://learnopencv.com/efficient-image-loading/\" target=\"_blank\">https://learnopencv.com/efficient-image-loading/</a><br>\nAnd cv2 doesn't seem to be that bad so it's probably related to your augmentations and image processing.</p>\n<p>If you provide your code it will be easier to give some advice. In general if there are steps that are done all the time it's better to do it once before training, save the preprocessed images and use them during training. <br>\nLike resizing and padding the images: <a href=\"https://www.kaggle.com/code/michaln/hotel-id-image-preprocessing-512x512\" target=\"_blank\">https://www.kaggle.com/code/michaln/hotel-id-image-preprocessing-512x512</a><br>\nyou can create dataset from the notebook output that can be used in the training notebook: <a href=\"https://www.kaggle.com/datasets/michaln/hotelid-2022-train-images-512x512\" target=\"_blank\">https://www.kaggle.com/datasets/michaln/hotelid-2022-train-images-512x512</a></p>\n<p>In case you want to do resizing or crop from the original image it's better to do it as first so the following augmentations are done on the smaller image. I would try to use no augmentations and then add them one by one to measure the impact on performance. Once you locate the problem you can try to optimize it.</p>\n<p>There is a list of albumentations equivalents for torchvision transforms: <a href=\"https://albumentations.ai/docs/examples/migrating_from_torchvision_to_albumentations\" target=\"_blank\">https://albumentations.ai/docs/examples/migrating_from_torchvision_to_albumentations</a> <br>\nSo maybe you can replace torchvision with albumentation completely which could also help.</p>\n<p>I currently use only augmentations like in this notebook: <a href=\"https://www.kaggle.com/code/michaln/hotel-id-starter-classification-traning\" target=\"_blank\">https://www.kaggle.com/code/michaln/hotel-id-starter-classification-traning</a><br>\nAlbumentations library, preprocessed images (padded and resized to 256x256 pixels), loading using PIL, batch size 64. And it can process 5 iterations per second with efficientnet b0.</p>",
      "rawMarkdown": "Well i would say that 15-18 secs per batch is not normal but it depends on what you do and on what size of images.\n\nThere are some benchmarks of image loading libraries:\nhttps://stackoverflow.com/questions/57663734/how-to-speed-up-image-loading-in-pillow-python\nhttps://www.kaggle.com/code/zfturbo/benchmark-2019-speed-of-image-reading/notebook\nhttps://learnopencv.com/efficient-image-loading/\nAnd cv2 doesn't seem to be that bad so it's probably related to your augmentations and image processing.\n\nIf you provide your code it will be easier to give some advice. In general if there are steps that are done all the time it's better to do it once before training, save the preprocessed images and use them during training. \nLike resizing and padding the images: https://www.kaggle.com/code/michaln/hotel-id-image-preprocessing-512x512\nyou can create dataset from the notebook output that can be used in the training notebook: https://www.kaggle.com/datasets/michaln/hotelid-2022-train-images-512x512\n\nIn case you want to do resizing or crop from the original image it's better to do it as first so the following augmentations are done on the smaller image. I would try to use no augmentations and then add them one by one to measure the impact on performance. Once you locate the problem you can try to optimize it.\n\nThere is a list of albumentations equivalents for torchvision transforms: https://albumentations.ai/docs/examples/migrating_from_torchvision_to_albumentations \nSo maybe you can replace torchvision with albumentation completely which could also help.\n\nI currently use only augmentations like in this notebook: https://www.kaggle.com/code/michaln/hotel-id-starter-classification-traning\nAlbumentations library, preprocessed images (padded and resized to 256x256 pixels), loading using PIL, batch size 64. And it can process 5 iterations per second with efficientnet b0.",
      "votes": null
    },
    {
      "id": "1743388",
      "postDate": "04/03/2022 00:12:23",
      "content": "<p>Thank you so much for the insights Michal. My latest notebook is on colab and I'll move it back here later. Now I made every augmentation done in Albumentations and re-arranged the order of the transforms so that the first one reduces the image size to 256x256. The time taken per batch is now 3 seconds with batch size 64. Great progress compared to where I was yesterday.</p>\n<p>And you are right, the deterministic transforms like resizing can be done once and for all before training. I am going to do that and leave the ones with randomness done per batch. I guess even if the first one is resizing, it still costs a lot of time because of the original image size so shouldn't be re-done every batch.</p>",
      "rawMarkdown": "Thank you so much for the insights Michal. My latest notebook is on colab and I'll move it back here later. Now I made every augmentation done in Albumentations and re-arranged the order of the transforms so that the first one reduces the image size to 256x256. The time taken per batch is now 3 seconds with batch size 64. Great progress compared to where I was yesterday.\n\nAnd you are right, the deterministic transforms like resizing can be done once and for all before training. I am going to do that and leave the ones with randomness done per batch. I guess even if the first one is resizing, it still costs a lot of time because of the original image size so shouldn't be re-done every batch.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1743012,
      "author_name": "michaln",
      "author_url": "",
      "post_date": "04/02/2022 13:59:25",
      "content": "<p>Well i would say that 15-18 secs per batch is not normal but it depends on what you do and on what size of images.</p>\n<p>There are some benchmarks of image loading libraries:<br>\n<a href=\"https://stackoverflow.com/questions/57663734/how-to-speed-up-image-loading-in-pillow-python\" target=\"_blank\">https://stackoverflow.com/questions/57663734/how-to-speed-up-image-loading-in-pillow-python</a><br>\n<a href=\"https://www.kaggle.com/code/zfturbo/benchmark-2019-speed-of-image-reading/notebook\" target=\"_blank\">https://www.kaggle.com/code/zfturbo/benchmark-2019-speed-of-image-reading/notebook</a><br>\n<a href=\"https://learnopencv.com/efficient-image-loading/\" target=\"_blank\">https://learnopencv.com/efficient-image-loading/</a><br>\nAnd cv2 doesn't seem to be that bad so it's probably related to your augmentations and image processing.</p>\n<p>If you provide your code it will be easier to give some advice. In general if there are steps that are done all the time it's better to do it once before training, save the preprocessed images and use them during training. <br>\nLike resizing and padding the images: <a href=\"https://www.kaggle.com/code/michaln/hotel-id-image-preprocessing-512x512\" target=\"_blank\">https://www.kaggle.com/code/michaln/hotel-id-image-preprocessing-512x512</a><br>\nyou can create dataset from the notebook output that can be used in the training notebook: <a href=\"https://www.kaggle.com/datasets/michaln/hotelid-2022-train-images-512x512\" target=\"_blank\">https://www.kaggle.com/datasets/michaln/hotelid-2022-train-images-512x512</a></p>\n<p>In case you want to do resizing or crop from the original image it's better to do it as first so the following augmentations are done on the smaller image. I would try to use no augmentations and then add them one by one to measure the impact on performance. Once you locate the problem you can try to optimize it.</p>\n<p>There is a list of albumentations equivalents for torchvision transforms: <a href=\"https://albumentations.ai/docs/examples/migrating_from_torchvision_to_albumentations\" target=\"_blank\">https://albumentations.ai/docs/examples/migrating_from_torchvision_to_albumentations</a> <br>\nSo maybe you can replace torchvision with albumentation completely which could also help.</p>\n<p>I currently use only augmentations like in this notebook: <a href=\"https://www.kaggle.com/code/michaln/hotel-id-starter-classification-traning\" target=\"_blank\">https://www.kaggle.com/code/michaln/hotel-id-starter-classification-traning</a><br>\nAlbumentations library, preprocessed images (padded and resized to 256x256 pixels), loading using PIL, batch size 64. And it can process 5 iterations per second with efficientnet b0.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1743388,
          "author_name": "jaredfeng",
          "author_url": "",
          "post_date": "04/03/2022 00:12:23",
          "content": "<p>Thank you so much for the insights Michal. My latest notebook is on colab and I'll move it back here later. Now I made every augmentation done in Albumentations and re-arranged the order of the transforms so that the first one reduces the image size to 256x256. The time taken per batch is now 3 seconds with batch size 64. Great progress compared to where I was yesterday.</p>\n<p>And you are right, the deterministic transforms like resizing can be done once and for all before training. I am going to do that and leave the ones with randomness done per batch. I guess even if the first one is resizing, it still costs a lot of time because of the original image size so shouldn't be re-done every batch.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1742372": "I am applying a lot of heavy augmentations to the training data and my batch size is now 128. The augmentation takes about 15-18 secs per batch. Is this normal? I am using pytorch for training. For augmentations, I am mainly using albumentations but among them a few key steps are in torchvision & timm. Perhaps it's also related to how I read the images (I use cv2.imread()) into the data loader.\n\nWhat bothers me most is that the transformations from albumentations only take in numpy arrays. But in my notebook, those come after some torchvision transforms outputting tensors already.\n\nWhat general suggestions you would give for fast and efficient data augmentation? Perhaps which package(s) and what order should the transforms be in?",
    "1743012": "Well i would say that 15-18 secs per batch is not normal but it depends on what you do and on what size of images.\n\nThere are some benchmarks of image loading libraries:\nhttps://stackoverflow.com/questions/57663734/how-to-speed-up-image-loading-in-pillow-python\nhttps://www.kaggle.com/code/zfturbo/benchmark-2019-speed-of-image-reading/notebook\nhttps://learnopencv.com/efficient-image-loading/\nAnd cv2 doesn't seem to be that bad so it's probably related to your augmentations and image processing.\n\nIf you provide your code it will be easier to give some advice. In general if there are steps that are done all the time it's better to do it once before training, save the preprocessed images and use them during training. \nLike resizing and padding the images: https://www.kaggle.com/code/michaln/hotel-id-image-preprocessing-512x512\nyou can create dataset from the notebook output that can be used in the training notebook: https://www.kaggle.com/datasets/michaln/hotelid-2022-train-images-512x512\n\nIn case you want to do resizing or crop from the original image it's better to do it as first so the following augmentations are done on the smaller image. I would try to use no augmentations and then add them one by one to measure the impact on performance. Once you locate the problem you can try to optimize it.\n\nThere is a list of albumentations equivalents for torchvision transforms: https://albumentations.ai/docs/examples/migrating_from_torchvision_to_albumentations \nSo maybe you can replace torchvision with albumentation completely which could also help.\n\nI currently use only augmentations like in this notebook: https://www.kaggle.com/code/michaln/hotel-id-starter-classification-traning\nAlbumentations library, preprocessed images (padded and resized to 256x256 pixels), loading using PIL, batch size 64. And it can process 5 iterations per second with efficientnet b0.",
    "1743388": "Thank you so much for the insights Michal. My latest notebook is on colab and I'll move it back here later. Now I made every augmentation done in Albumentations and re-arranged the order of the transforms so that the first one reduces the image size to 256x256. The time taken per batch is now 3 seconds with batch size 64. Great progress compared to where I was yesterday.\n\nAnd you are right, the deterministic transforms like resizing can be done once and for all before training. I am going to do that and leave the ones with randomness done per batch. I guess even if the first one is resizing, it still costs a lot of time because of the original image size so shouldn't be re-done every batch."
  },
  "source": "meta"
}