{
  "id": 466046,
  "title": "rle_decode vs cv2.imread vs tifffile.imread ",
  "url": "/competitions/blood-vessel-segmentation/discussion/466046",
  "author_name": "",
  "post_date": "2024-01-06T21:56:35.982021500Z",
  "votes": 9,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>I have run a quick time run test for reading the competition data. <br>\nWhat do you think? </p>\n<pre><code> cv2\n timeit\n tifffile  tiff\n tqdm.auto  tqdm \n\n\n ():\n    rle, H, W = row[][[, , ]].values\n    seq = rle.split()\n    start = np.asarray(rle.split()[::]).astype()\n    length = np.asarray(rle.split()[::]).astype()\n    end = start + length\n\n    img = np.zeros((H*W,), dtype=np.uint8)\n     hi, lo  (start, end):\n        img[hi:lo] = \n    img.shape = (H, W)\n     img\n\n ():\n    masks1 = [rle_decode(row)  row  tqdm(mrgd.iterrows(), total=(mrgd))]\n\n ():\n    masks2 = [tiff.imread(p)  p  tqdm(mrgd.mask_path)]\n\n ():\n    masks3 = [cv2.imread(p, cv2.IMREAD_GRAYSCALE)  p  tqdm(mrgd.mask_path)]\n\ntime1 = timeit.timeit(rle_decode_read, number=)\ntime2 = timeit.timeit(tiff_read, number=)\ntime3 = timeit.timeit(cv2_read, number=)\n\n()\n()\n()\n</code></pre>\n<pre><code>RLE Decode Read  times:  seconds\nTIFF Read        times:   seconds\nCV2 Read         times:  seconds\n</code></pre>",
  "messages": [
    {
      "id": "2590085",
      "postDate": "01/06/2024 21:56:35",
      "content": "<p>Hi,</p>\n<p>I have run a quick time run test for reading the competition data. <br>\nWhat do you think? </p>\n<pre><code> cv2\n timeit\n tifffile  tiff\n tqdm.auto  tqdm \n\n\n ():\n    rle, H, W = row[][[, , ]].values\n    seq = rle.split()\n    start = np.asarray(rle.split()[::]).astype()\n    length = np.asarray(rle.split()[::]).astype()\n    end = start + length\n\n    img = np.zeros((H*W,), dtype=np.uint8)\n     hi, lo  (start, end):\n        img[hi:lo] = \n    img.shape = (H, W)\n     img\n\n ():\n    masks1 = [rle_decode(row)  row  tqdm(mrgd.iterrows(), total=(mrgd))]\n\n ():\n    masks2 = [tiff.imread(p)  p  tqdm(mrgd.mask_path)]\n\n ():\n    masks3 = [cv2.imread(p, cv2.IMREAD_GRAYSCALE)  p  tqdm(mrgd.mask_path)]\n\ntime1 = timeit.timeit(rle_decode_read, number=)\ntime2 = timeit.timeit(tiff_read, number=)\ntime3 = timeit.timeit(cv2_read, number=)\n\n()\n()\n()\n</code></pre>\n<pre><code>RLE Decode Read  times:  seconds\nTIFF Read        times:   seconds\nCV2 Read         times:  seconds\n</code></pre>",
      "rawMarkdown": "Hi,\n\nI have run a quick time run test for reading the competition data. \nWhat do you think? \n\n```python\nimport cv2\nimport timeit\nimport tifffile as tiff\nfrom tqdm.auto import tqdm \n\n\ndef rle_decode(row):\n    rle, H, W = row[1][['rle', 'H', 'W']].values\n    seq = rle.split()\n    start = np.asarray(rle.split()[0::2]).astype(int)\n    length = np.asarray(rle.split()[1::2]).astype(int)\n    end = start + length\n\n    img = np.zeros((H*W,), dtype=np.uint8)\n    for hi, lo in zip(start, end):\n        img[hi:lo] = 255\n    img.shape = (H, W)\n    return img\n\ndef rle_decode_read():\n    masks1 = [rle_decode(row) for row in tqdm(mrgd.iterrows(), total=len(mrgd))]\n\ndef tiff_read():\n    masks2 = [tiff.imread(p) for p in tqdm(mrgd.mask_path)]\n\ndef cv2_read():\n    masks3 = [cv2.imread(p, cv2.IMREAD_GRAYSCALE) for p in tqdm(mrgd.mask_path)]\n\ntime1 = timeit.timeit(rle_decode_read, number=3)\ntime2 = timeit.timeit(tiff_read, number=3)\ntime3 = timeit.timeit(cv2_read, number=3)\n\nprint(f\"RLE Decode Read 3 times: {time1:>6.3f} seconds\")\nprint(f\"TIFF Read       3 times: {time2:>6.3f} seconds\")\nprint(f\"CV2 Read        3 times: {time3:>6.3f} seconds\")\n```\n\n```python\nRLE Decode Read 3 times: 24.173 seconds\nTIFF Read       3 times:  9.205 seconds\nCV2 Read        3 times: 59.641 seconds\n```",
      "votes": null
    },
    {
      "id": "2590091",
      "postDate": "01/06/2024 22:13:29",
      "content": "<p>If I'm not mistaken tifffile reads are lazy and don't actually read any image data until you do something to it. I haven't noticed significant differences in read times between tifffile, cv2 and PIL.</p>",
      "rawMarkdown": "If I'm not mistaken tifffile reads are lazy and don't actually read any image data until you do something to it. I haven't noticed significant differences in read times between tifffile, cv2 and PIL.",
      "votes": null
    },
    {
      "id": "2590105",
      "postDate": "01/06/2024 22:33:57",
      "content": "<p>I cannot find a notion of lazy read. Both arrays shows the same memory footprint after the read as well.<br>\nIt says it returns NDArray[Any],  here is the doc <a href=\"https://github.com/cgohlke/tifffile/blob/5ced58caa74936ecc53f58ade94938e251e83138/tifffile/tifffile.py#L1002\" target=\"_blank\">link</a>. Am I missing something?</p>\n<pre><code> sys\n(sys.getsizeof(masks2))\n(sys.getsizeof(masks3))\n\n\n\n</code></pre>",
      "rawMarkdown": "I cannot find a notion of lazy read. Both arrays shows the same memory footprint after the read as well.\nIt says it returns NDArray[Any],  here is the doc [link](https://github.com/cgohlke/tifffile/blob/5ced58caa74936ecc53f58ade94938e251e83138/tifffile/tifffile.py#L1002). Am I missing something?\n\n```python\nimport sys\nprint(sys.getsizeof(masks2))\nprint(sys.getsizeof(masks3))\n\n>>> 59736\n>>> 59736\n```",
      "votes": null
    },
    {
      "id": "2590274",
      "postDate": "01/07/2024 04:42:43",
      "content": "<p>may I ask, if you train kidney_1_voi locally or on Kaggle?</p>\n<p>why I always get OOM on Kaggle to train it?</p>",
      "rawMarkdown": "may I ask, if you train kidney_1_voi locally or on Kaggle?\n\nwhy I always get OOM on Kaggle to train it?",
      "votes": null
    },
    {
      "id": "2590568",
      "postDate": "01/07/2024 09:46:30",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/dragonzhang\" target=\"_blank\">@dragonzhang</a>, I would say this topic addresses a different issue and your question is a bit outside the scope. Though, I would recommend you to open a separate discussion describing your problem in details and may be providing minimum code reproducible example. There might be various reasons: big batch size or creating duplicate variables.</p>",
      "rawMarkdown": "Hi @dragonzhang, I would say this topic addresses a different issue and your question is a bit outside the scope. Though, I would recommend you to open a separate discussion describing your problem in details and may be providing minimum code reproducible example. There might be various reasons: big batch size or creating duplicate variables.",
      "votes": null
    },
    {
      "id": "2594492",
      "postDate": "01/09/2024 23:21:06",
      "content": "<p>Based on the <a href=\"https://github.com/cgohlke/tifffile/blob/master/tifffile/tifffile.py#L961\" target=\"_blank\">source code</a> you are correct.<br>\nIf you call imread with <code>aszarr=True</code>, you get a Zarr store object which can support lazy loading, but the default is to return an np array</p>",
      "rawMarkdown": "Based on the [source code](https://github.com/cgohlke/tifffile/blob/master/tifffile/tifffile.py#L961) you are correct.\nIf you call imread with `aszarr=True`, you get a Zarr store object which can support lazy loading, but the default is to return an np array",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2590091,
      "author_name": "sakvaua",
      "author_url": "",
      "post_date": "01/06/2024 22:13:29",
      "content": "<p>If I'm not mistaken tifffile reads are lazy and don't actually read any image data until you do something to it. I haven't noticed significant differences in read times between tifffile, cv2 and PIL.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2590105,
          "author_name": "sergiosaharovskiy",
          "author_url": "",
          "post_date": "01/06/2024 22:33:57",
          "content": "<p>I cannot find a notion of lazy read. Both arrays shows the same memory footprint after the read as well.<br>\nIt says it returns NDArray[Any],  here is the doc <a href=\"https://github.com/cgohlke/tifffile/blob/5ced58caa74936ecc53f58ade94938e251e83138/tifffile/tifffile.py#L1002\" target=\"_blank\">link</a>. Am I missing something?</p>\n<pre><code> sys\n(sys.getsizeof(masks2))\n(sys.getsizeof(masks3))\n\n\n\n</code></pre>",
          "votes": null,
          "replies": [
            {
              "id": 2594492,
              "author_name": "reubenschmidt",
              "author_url": "",
              "post_date": "01/09/2024 23:21:06",
              "content": "<p>Based on the <a href=\"https://github.com/cgohlke/tifffile/blob/master/tifffile/tifffile.py#L961\" target=\"_blank\">source code</a> you are correct.<br>\nIf you call imread with <code>aszarr=True</code>, you get a Zarr store object which can support lazy loading, but the default is to return an np array</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2590274,
      "author_name": "dragonzhang",
      "author_url": "",
      "post_date": "01/07/2024 04:42:43",
      "content": "<p>may I ask, if you train kidney_1_voi locally or on Kaggle?</p>\n<p>why I always get OOM on Kaggle to train it?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2590568,
          "author_name": "sergiosaharovskiy",
          "author_url": "",
          "post_date": "01/07/2024 09:46:30",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/dragonzhang\" target=\"_blank\">@dragonzhang</a>, I would say this topic addresses a different issue and your question is a bit outside the scope. Though, I would recommend you to open a separate discussion describing your problem in details and may be providing minimum code reproducible example. There might be various reasons: big batch size or creating duplicate variables.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2590085": "Hi,\n\nI have run a quick time run test for reading the competition data. \nWhat do you think? \n\n```python\nimport cv2\nimport timeit\nimport tifffile as tiff\nfrom tqdm.auto import tqdm \n\n\ndef rle_decode(row):\n    rle, H, W = row[1][['rle', 'H', 'W']].values\n    seq = rle.split()\n    start = np.asarray(rle.split()[0::2]).astype(int)\n    length = np.asarray(rle.split()[1::2]).astype(int)\n    end = start + length\n\n    img = np.zeros((H*W,), dtype=np.uint8)\n    for hi, lo in zip(start, end):\n        img[hi:lo] = 255\n    img.shape = (H, W)\n    return img\n\ndef rle_decode_read():\n    masks1 = [rle_decode(row) for row in tqdm(mrgd.iterrows(), total=len(mrgd))]\n\ndef tiff_read():\n    masks2 = [tiff.imread(p) for p in tqdm(mrgd.mask_path)]\n\ndef cv2_read():\n    masks3 = [cv2.imread(p, cv2.IMREAD_GRAYSCALE) for p in tqdm(mrgd.mask_path)]\n\ntime1 = timeit.timeit(rle_decode_read, number=3)\ntime2 = timeit.timeit(tiff_read, number=3)\ntime3 = timeit.timeit(cv2_read, number=3)\n\nprint(f\"RLE Decode Read 3 times: {time1:>6.3f} seconds\")\nprint(f\"TIFF Read       3 times: {time2:>6.3f} seconds\")\nprint(f\"CV2 Read        3 times: {time3:>6.3f} seconds\")\n```\n\n```python\nRLE Decode Read 3 times: 24.173 seconds\nTIFF Read       3 times:  9.205 seconds\nCV2 Read        3 times: 59.641 seconds\n```",
    "2590091": "If I'm not mistaken tifffile reads are lazy and don't actually read any image data until you do something to it. I haven't noticed significant differences in read times between tifffile, cv2 and PIL.",
    "2590105": "I cannot find a notion of lazy read. Both arrays shows the same memory footprint after the read as well.\nIt says it returns NDArray[Any],  here is the doc [link](https://github.com/cgohlke/tifffile/blob/5ced58caa74936ecc53f58ade94938e251e83138/tifffile/tifffile.py#L1002). Am I missing something?\n\n```python\nimport sys\nprint(sys.getsizeof(masks2))\nprint(sys.getsizeof(masks3))\n\n>>> 59736\n>>> 59736\n```",
    "2590274": "may I ask, if you train kidney_1_voi locally or on Kaggle?\n\nwhy I always get OOM on Kaggle to train it?",
    "2590568": "Hi @dragonzhang, I would say this topic addresses a different issue and your question is a bit outside the scope. Though, I would recommend you to open a separate discussion describing your problem in details and may be providing minimum code reproducible example. There might be various reasons: big batch size or creating duplicate variables.",
    "2594492": "Based on the [source code](https://github.com/cgohlke/tifffile/blob/master/tifffile/tifffile.py#L961) you are correct.\nIf you call imread with `aszarr=True`, you get a Zarr store object which can support lazy loading, but the default is to return an np array"
  },
  "source": "meta"
}