{
  "id": 627529,
  "title": "Training speed improves 10x if you convert TIFF dataset to NPY / RLE",
  "url": "/competitions/vesuvius-challenge-surface-detection/discussion/627529",
  "author_name": "Mayukh Bhattacharyya",
  "post_date": "2025-11-16T20:39:23.452000",
  "votes": 17,
  "comment_count": 9,
  "views": 0,
  "content": "<p>If you are doing deep learning like UNet etc. and your training feels very slow, the bottleneck is almost certainly I/O, not your model. TIFF loading in Python is extremely slow — especially for 3D volumes — and it can easily dominate your epoch time.</p>\n<p>I benchmarked several 3D image data storage formats on the Vesuvius dataset (<code>tiff</code>, <code>npy</code>, <code>npz</code>, <code>zarr</code>), measuring average load time and file size. <code>npy</code> is a well known format as part of numpy. <code>zarr</code> was used as the data in the <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification\" target=\"_blank\">CZII-CryoET</a> challenge, .. maybe in other too. <code>npz</code> is an unstructured and compressed version of <code>npy</code>.</p>\n<p><code>npy</code> is the best in terms of speed. Gives about ~25x speedup in image loading times in isolation and decreased my overall training time by about ~10x.</p>\n<p>For labels, plain <code>npy</code> wastes a ton of disk space (<code>tif</code> labels are &lt; 1MB whereas <code>npy</code> labels are ~32 MB), as the actual label information does not need us to store the whole array. I used RLE (Run-Length Encoding) to compress, it's not that fast as <code>npy</code> but still 2-3x better than loading from the <code>tif</code> label files.</p>\n<p>This <a href=\"https://www.kaggle.com/code/mayukh18/speed-up-data-load-npy-and-rle-data\" target=\"_blank\">notebook</a> does the <code>npy</code> and <code>rle</code> conversion. You can directly use the output of the notebook as a replacement of the competition data. However note, I downsampled the data from 320 to 256 pixels for some extra speed and space saving as I didn't feel there's much of a difference.</p>\n<p>If you feel the need to have the original 320px resolution, or want to save the labels in <code>npy</code>, it's pretty straightforward to do so too.</p>\n<p>Happy training!</p>",
  "messages": [
    {
      "id": 3332707,
      "postDate": "2025-11-16T20:39:23.453Z",
      "content": "<p>If you are doing deep learning like UNet etc. and your training feels very slow, the bottleneck is almost certainly I/O, not your model. TIFF loading in Python is extremely slow — especially for 3D volumes — and it can easily dominate your epoch time.</p>\n<p>I benchmarked several 3D image data storage formats on the Vesuvius dataset (<code>tiff</code>, <code>npy</code>, <code>npz</code>, <code>zarr</code>), measuring average load time and file size. <code>npy</code> is a well known format as part of numpy. <code>zarr</code> was used as the data in the <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification\" target=\"_blank\">CZII-CryoET</a> challenge, .. maybe in other too. <code>npz</code> is an unstructured and compressed version of <code>npy</code>.</p>\n<p><code>npy</code> is the best in terms of speed. Gives about ~25x speedup in image loading times in isolation and decreased my overall training time by about ~10x.</p>\n<p>For labels, plain <code>npy</code> wastes a ton of disk space (<code>tif</code> labels are &lt; 1MB whereas <code>npy</code> labels are ~32 MB), as the actual label information does not need us to store the whole array. I used RLE (Run-Length Encoding) to compress, it's not that fast as <code>npy</code> but still 2-3x better than loading from the <code>tif</code> label files.</p>\n<p>This <a href=\"https://www.kaggle.com/code/mayukh18/speed-up-data-load-npy-and-rle-data\" target=\"_blank\">notebook</a> does the <code>npy</code> and <code>rle</code> conversion. You can directly use the output of the notebook as a replacement of the competition data. However note, I downsampled the data from 320 to 256 pixels for some extra speed and space saving as I didn't feel there's much of a difference.</p>\n<p>If you feel the need to have the original 320px resolution, or want to save the labels in <code>npy</code>, it's pretty straightforward to do so too.</p>\n<p>Happy training!</p>",
      "rawMarkdown": "If you are doing deep learning like UNet etc. and your training feels very slow, the bottleneck is almost certainly I/O, not your model. TIFF loading in Python is extremely slow — especially for 3D volumes — and it can easily dominate your epoch time.\n\nI benchmarked several 3D image data storage formats on the Vesuvius dataset (`tiff`, `npy`, `npz`, `zarr`), measuring average load time and file size. `npy` is a well known format as part of numpy. `zarr` was used as the data in the [CZII-CryoET](https://www.kaggle.com/competitions/czii-cryo-et-object-identification) challenge, .. maybe in other too. `npz` is an unstructured and compressed version of `npy`.\n\n`npy` is the best in terms of speed. Gives about ~25x speedup in image loading times in isolation and decreased my overall training time by about ~10x.\n\nFor labels, plain `npy` wastes a ton of disk space (`tif` labels are < 1MB whereas `npy` labels are ~32 MB), as the actual label information does not need us to store the whole array. I used RLE (Run-Length Encoding) to compress, it's not that fast as `npy` but still 2-3x better than loading from the `tif` label files.\n\nThis [notebook](https://www.kaggle.com/code/mayukh18/speed-up-data-load-npy-and-rle-data) does the `npy` and `rle` conversion. You can directly use the output of the notebook as a replacement of the competition data. However note, I downsampled the data from 320 to 256 pixels for some extra speed and space saving as I didn't feel there's much of a difference.\n\nIf you feel the need to have the original 320px resolution, or want to save the labels in `npy`, it's pretty straightforward to do so too.\n\nHappy training!",
      "votes": 17
    },
    {
      "id": 3401577,
      "postDate": "2026-02-04T03:50:08.423Z",
      "content": "<p>has anyone tried this conversion method during inference? if so did it improve your runtime? I am working on increasing TTA and postprocessing steps but my notebook timed out during scoring (10+ hrs), I am wondering if doing this tiff conversion can help during inference. TIA.</p>",
      "rawMarkdown": "has anyone tried this conversion method during inference? if so did it improve your runtime? I am working on increasing TTA and postprocessing steps but my notebook timed out during scoring (10+ hrs), I am wondering if doing this tiff conversion can help during inference. TIA.",
      "votes": 1
    },
    {
      "id": 3336061,
      "postDate": "2025-11-18T10:06:54.610Z",
      "content": "<p>If you do cropping, load your npy arrays as memmap (<code>numpy.load(file, mmap_mode='r')</code>). This is even more efficient.</p>",
      "rawMarkdown": "If you do cropping, load your npy arrays as memmap (`numpy.load(file, mmap_mode='r')`). This is even more efficient.",
      "votes": 4
    },
    {
      "id": 3341890,
      "postDate": "2025-11-20T13:49:52.060Z",
      "content": "<p>Yes this can be much faster, even uncompressing the original tifs (and using tifffile to load instead of say, opencv) can lead to pretty large speed gains. For the most part, i do most of my training using Zarr, but for the purposes of the competition i wanted more generalized data formats just for ease of use. </p>",
      "rawMarkdown": "Yes this can be much faster, even uncompressing the original tifs (and using tifffile to load instead of say, opencv) can lead to pretty large speed gains. For the most part, i do most of my training using Zarr, but for the purposes of the competition i wanted more generalized data formats just for ease of use. ",
      "votes": 1
    },
    {
      "id": 3341367,
      "postDate": "2025-11-20T06:09:21.237Z",
      "content": "<p>Another bottleneck could be augmentation, in particular if you try to do 3D segmentation, so doing it on GPU instead of CPU is also almost 10x boost, see <a href=\"https://www.kaggle.com/code/jirkaborovec/surface-detect-segment-3d-gpu-augmentation\" target=\"_blank\">https://www.kaggle.com/code/jirkaborovec/surface-detect-segment-3d-gpu-augmentation</a></p>",
      "rawMarkdown": "Another bottleneck could be augmentation, in particular if you try to do 3D segmentation, so doing it on GPU instead of CPU is also almost 10x boost, see https://www.kaggle.com/code/jirkaborovec/surface-detect-segment-3d-gpu-augmentation",
      "votes": 1
    },
    {
      "id": 3341371,
      "postDate": "2025-11-20T06:11:49.307Z",
      "content": "<p>seems someone did this conversion, and it seems that the converted dataset is about 2.5x larger\n<a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/630100\" target=\"_blank\">https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/630100</a></p>",
      "rawMarkdown": "seems someone did this conversion, and it seems that the converted dataset is about 2.5x larger\nhttps://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/630100",
      "replies": [
        {
          "id": 3341424,
          "postDate": "2025-11-20T06:54:17.990Z",
          "content": "<p>That's not unexpected. The provided tiff files are compressed. .npy arrays are not. If you use compressed numpy arrays you save space again but then loading is slow again (+memmap won't work anymore). If you want the best of both worlds you need blosc2 (or hdf5, but I have not experimented with that extensively), but blosc2 is not easy to get a hang of</p>",
          "rawMarkdown": "That's not unexpected. The provided tiff files are compressed. .npy arrays are not. If you use compressed numpy arrays you save space again but then loading is slow again (+memmap won't work anymore). If you want the best of both worlds you need blosc2 (or hdf5, but I have not experimented with that extensively), but blosc2 is not easy to get a hang of",
          "votes": 1,
          "replies": [
            {
              "id": 3341504,
              "postDate": "2025-11-20T07:58:01.643Z",
              "content": "<p>Sure, just observation, not a complain… :) Have you also benchmarked npz?</p>",
              "rawMarkdown": "Sure, just observation, not a complain... :) Have you also benchmarked npz?"
            },
            {
              "id": 3341512,
              "postDate": "2025-11-20T08:06:55.853Z",
              "content": "<p>I have not done anything for this challenge yet, but I have benchmarked a variety of file formats for precisely this type of application in the past. The issue with npz is that it will force you uncompress and load an entire 3D image before you can crop. This is quite wasteful unless you train with a patch size that is roughly the same as the image shape. So ideally you want a file format that allows partial reading. Loading npy as memmap and blosc2 allow you to do that. npy with mmap_mode='r' is probably the easiest to get into but it requires more storage + I/O because of the lack of compression. This is why we switched to blosc2 in nnU-Net</p>",
              "rawMarkdown": "I have not done anything for this challenge yet, but I have benchmarked a variety of file formats for precisely this type of application in the past. The issue with npz is that it will force you uncompress and load an entire 3D image before you can crop. This is quite wasteful unless you train with a patch size that is roughly the same as the image shape. So ideally you want a file format that allows partial reading. Loading npy as memmap and blosc2 allow you to do that. npy with mmap_mode='r' is probably the easiest to get into but it requires more storage + I/O because of the lack of compression. This is why we switched to blosc2 in nnU-Net"
            }
          ]
        }
      ]
    },
    {
      "id": 3335997,
      "postDate": "2025-11-18T09:22:18.830Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3401577,
      "author_name": "kawaii",
      "author_url": "",
      "post_date": "2026-02-04T03:50:08.423000",
      "content": "<p>has anyone tried this conversion method during inference? if so did it improve your runtime? I am working on increasing TTA and postprocessing steps but my notebook timed out during scoring (10+ hrs), I am wondering if doing this tiff conversion can help during inference. TIA.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3336061,
      "author_name": "FabianIsensee",
      "author_url": "",
      "post_date": "2025-11-18T10:06:54.610000",
      "content": "<p>If you do cropping, load your npy arrays as memmap (<code>numpy.load(file, mmap_mode='r')</code>). This is even more efficient.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 3341890,
      "author_name": "Sean Johnson_SP",
      "author_url": "",
      "post_date": "2025-11-20T13:49:52.060000",
      "content": "<p>Yes this can be much faster, even uncompressing the original tifs (and using tifffile to load instead of say, opencv) can lead to pretty large speed gains. For the most part, i do most of my training using Zarr, but for the purposes of the competition i wanted more generalized data formats just for ease of use. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3341367,
      "author_name": "Jirka",
      "author_url": "",
      "post_date": "2025-11-20T06:09:21.237000",
      "content": "<p>Another bottleneck could be augmentation, in particular if you try to do 3D segmentation, so doing it on GPU instead of CPU is also almost 10x boost, see <a href=\"https://www.kaggle.com/code/jirkaborovec/surface-detect-segment-3d-gpu-augmentation\" target=\"_blank\">https://www.kaggle.com/code/jirkaborovec/surface-detect-segment-3d-gpu-augmentation</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3341371,
      "author_name": "Jirka",
      "author_url": "",
      "post_date": "2025-11-20T06:11:49.307000",
      "content": "<p>seems someone did this conversion, and it seems that the converted dataset is about 2.5x larger\n<a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/630100\" target=\"_blank\">https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/630100</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 3341424,
          "author_name": "FabianIsensee",
          "author_url": "",
          "post_date": "2025-11-20T06:54:17.990000",
          "content": "<p>That's not unexpected. The provided tiff files are compressed. .npy arrays are not. If you use compressed numpy arrays you save space again but then loading is slow again (+memmap won't work anymore). If you want the best of both worlds you need blosc2 (or hdf5, but I have not experimented with that extensively), but blosc2 is not easy to get a hang of</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3341504,
              "author_name": "Jirka",
              "author_url": "",
              "post_date": "2025-11-20T07:58:01.643000",
              "content": "<p>Sure, just observation, not a complain… :) Have you also benchmarked npz?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3341512,
              "author_name": "FabianIsensee",
              "author_url": "",
              "post_date": "2025-11-20T08:06:55.853000",
              "content": "<p>I have not done anything for this challenge yet, but I have benchmarked a variety of file formats for precisely this type of application in the past. The issue with npz is that it will force you uncompress and load an entire 3D image before you can crop. This is quite wasteful unless you train with a patch size that is roughly the same as the image shape. So ideally you want a file format that allows partial reading. Loading npy as memmap and blosc2 allow you to do that. npy with mmap_mode='r' is probably the easiest to get into but it requires more storage + I/O because of the lack of compression. This is why we switched to blosc2 in nnU-Net</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3335997,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-11-18T09:22:18.830000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3332707": "If you are doing deep learning like UNet etc. and your training feels very slow, the bottleneck is almost certainly I/O, not your model. TIFF loading in Python is extremely slow — especially for 3D volumes — and it can easily dominate your epoch time.\n\nI benchmarked several 3D image data storage formats on the Vesuvius dataset (`tiff`, `npy`, `npz`, `zarr`), measuring average load time and file size. `npy` is a well known format as part of numpy. `zarr` was used as the data in the [CZII-CryoET](https://www.kaggle.com/competitions/czii-cryo-et-object-identification) challenge, .. maybe in other too. `npz` is an unstructured and compressed version of `npy`.\n\n`npy` is the best in terms of speed. Gives about ~25x speedup in image loading times in isolation and decreased my overall training time by about ~10x.\n\nFor labels, plain `npy` wastes a ton of disk space (`tif` labels are < 1MB whereas `npy` labels are ~32 MB), as the actual label information does not need us to store the whole array. I used RLE (Run-Length Encoding) to compress, it's not that fast as `npy` but still 2-3x better than loading from the `tif` label files.\n\nThis [notebook](https://www.kaggle.com/code/mayukh18/speed-up-data-load-npy-and-rle-data) does the `npy` and `rle` conversion. You can directly use the output of the notebook as a replacement of the competition data. However note, I downsampled the data from 320 to 256 pixels for some extra speed and space saving as I didn't feel there's much of a difference.\n\nIf you feel the need to have the original 320px resolution, or want to save the labels in `npy`, it's pretty straightforward to do so too.\n\nHappy training!",
    "3401577": "has anyone tried this conversion method during inference? if so did it improve your runtime? I am working on increasing TTA and postprocessing steps but my notebook timed out during scoring (10+ hrs), I am wondering if doing this tiff conversion can help during inference. TIA.",
    "3336061": "If you do cropping, load your npy arrays as memmap (`numpy.load(file, mmap_mode='r')`). This is even more efficient.",
    "3341890": "Yes this can be much faster, even uncompressing the original tifs (and using tifffile to load instead of say, opencv) can lead to pretty large speed gains. For the most part, i do most of my training using Zarr, but for the purposes of the competition i wanted more generalized data formats just for ease of use. ",
    "3341367": "Another bottleneck could be augmentation, in particular if you try to do 3D segmentation, so doing it on GPU instead of CPU is also almost 10x boost, see https://www.kaggle.com/code/jirkaborovec/surface-detect-segment-3d-gpu-augmentation",
    "3341371": "seems someone did this conversion, and it seems that the converted dataset is about 2.5x larger\nhttps://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/630100",
    "3335997": ""
  }
}