{
  "id": 228559,
  "title": "Memory usage",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/228559",
  "author_name": "",
  "post_date": "2021-03-25T08:49:01.271481800Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hello friends,<br>\nI would like to ask if there are some techniques/ solutions that reduce memory consumption every suggestion is welcomed !</p>",
  "messages": [
    {
      "id": "1251910",
      "postDate": "03/25/2021 08:49:01",
      "content": "<p>Hello friends,<br>\nI would like to ask if there are some techniques/ solutions that reduce memory consumption every suggestion is welcomed !</p>",
      "rawMarkdown": "Hello friends,\nI would like to ask if there are some techniques/ solutions that reduce memory consumption every suggestion is welcomed !",
      "votes": null
    },
    {
      "id": "1252395",
      "postDate": "03/25/2021 16:45:16",
      "content": "<p>In my case, only use gc.collect and del variable was not sufficient to reduce memory overflow in the previous dataset. Right now im not facing the same problem, but the solution was to read only the tile of the tiff image from disk (described here: <a href=\"https://www.kaggle.com/nheathpatterson/discussion\" target=\"_blank\">https://www.kaggle.com/nheathpatterson/discussion</a> ). There was two problems with this solution: 1) It decreases a lot the performance when you read tile by tile (which is obvious, of course, because HDs are much slower than ram), but this can be solved reading whole \"rows\" of tiles. And 2) It uses zarr library that isn't installed in kaggles notebook by default, so you need to install without internet during the commit and submission of your code. But I spent my time to do this so here is the library: <a href=\"https://www.kaggle.com/joaorodriguez/zarrlib\" target=\"_blank\">https://www.kaggle.com/joaorodriguez/zarrlib</a> and here is the code for installation:</p>\n<pre><code>import os\nos.system('pip install ../input/zarrlib/fasteners-0.16-py2.py3-none-any.whl')\nos.system('pip install ../input/zarrlib/asciitree-0.3.3-py2.py3-none-any.whl')\nos.system('pip install ../input/zarrlib/numcodecs-0.7.2-cp37-cp37m-manylinux2010_x86_64.whl')\nos.system('pip install ../input/zarrlib/zarr-2.6.1-py3-none-any.whl')\nimport zarr\nimport dask.array as da\n\ndef tifffile_to_dask(im_fp):\n    imdata = zarr.open(tifffile.imread(im_fp, aszarr=True))\n    if isinstance(imdata, zarr.hierarchy.Group):\n        # always return base layer if pyramidal\n        imdata = da.from_zarr(imdata[0])\n    else:\n        imdata = da.from_zarr(imdata)\n    return imdata\n\n\n#To read the tiff image: \n\nim = tifffile_to_dask(filename)\n</code></pre>\n<p>And another solution is to use an optimized rle converter (described here: <a href=\"https://www.kaggle.com/bguberfain/memory-aware-rle-encoding)\" target=\"_blank\">https://www.kaggle.com/bguberfain/memory-aware-rle-encoding)</a>.</p>",
      "rawMarkdown": "In my case, only use gc.collect and del variable was not sufficient to reduce memory overflow in the previous dataset. Right now im not facing the same problem, but the solution was to read only the tile of the tiff image from disk (described here: https://www.kaggle.com/nheathpatterson/discussion ). There was two problems with this solution: 1) It decreases a lot the performance when you read tile by tile (which is obvious, of course, because HDs are much slower than ram), but this can be solved reading whole \"rows\" of tiles. And 2) It uses zarr library that isn't installed in kaggles notebook by default, so you need to install without internet during the commit and submission of your code. But I spent my time to do this so here is the library: https://www.kaggle.com/joaorodriguez/zarrlib and here is the code for installation:\n\n```\nimport os\nos.system('pip install ../input/zarrlib/fasteners-0.16-py2.py3-none-any.whl')\nos.system('pip install ../input/zarrlib/asciitree-0.3.3-py2.py3-none-any.whl')\nos.system('pip install ../input/zarrlib/numcodecs-0.7.2-cp37-cp37m-manylinux2010_x86_64.whl')\nos.system('pip install ../input/zarrlib/zarr-2.6.1-py3-none-any.whl')\nimport zarr\nimport dask.array as da\n\ndef tifffile_to_dask(im_fp):\n    imdata = zarr.open(tifffile.imread(im_fp, aszarr=True))\n    if isinstance(imdata, zarr.hierarchy.Group):\n        # always return base layer if pyramidal\n        imdata = da.from_zarr(imdata[0])\n    else:\n        imdata = da.from_zarr(imdata)\n    return imdata\n\n\n#To read the tiff image: \n\nim = tifffile_to_dask(filename)\n\n```\n\nAnd another solution is to use an optimized rle converter (described here: https://www.kaggle.com/bguberfain/memory-aware-rle-encoding).",
      "votes": null
    },
    {
      "id": "1253200",
      "postDate": "03/26/2021 13:15:15",
      "content": "<p>Thanks for ur help it's really interesting, but there's one more question, the obtained image has the same dimension as the original one ?</p>",
      "rawMarkdown": "Thanks for ur help it's really interesting, but there's one more question, the obtained image has the same dimension as the original one ?",
      "votes": null
    },
    {
      "id": "1253207",
      "postDate": "03/26/2021 13:19:55",
      "content": "<p>Yes, and you manipulate as an array, for example im[0:512,0:512] extract from disk the first tile of 512x512px</p>",
      "rawMarkdown": "Yes, and you manipulate as an array, for example im[0:512,0:512] extract from disk the first tile of 512x512px",
      "votes": null
    },
    {
      "id": "1253226",
      "postDate": "03/26/2021 13:42:07",
      "content": "<p>That's OK Thanks very muuch for ur help 🙂</p>",
      "rawMarkdown": "That's OK Thanks very muuch for ur help 🙂",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1252395,
      "author_name": "joaorodriguez",
      "author_url": "",
      "post_date": "03/25/2021 16:45:16",
      "content": "<p>In my case, only use gc.collect and del variable was not sufficient to reduce memory overflow in the previous dataset. Right now im not facing the same problem, but the solution was to read only the tile of the tiff image from disk (described here: <a href=\"https://www.kaggle.com/nheathpatterson/discussion\" target=\"_blank\">https://www.kaggle.com/nheathpatterson/discussion</a> ). There was two problems with this solution: 1) It decreases a lot the performance when you read tile by tile (which is obvious, of course, because HDs are much slower than ram), but this can be solved reading whole \"rows\" of tiles. And 2) It uses zarr library that isn't installed in kaggles notebook by default, so you need to install without internet during the commit and submission of your code. But I spent my time to do this so here is the library: <a href=\"https://www.kaggle.com/joaorodriguez/zarrlib\" target=\"_blank\">https://www.kaggle.com/joaorodriguez/zarrlib</a> and here is the code for installation:</p>\n<pre><code>import os\nos.system('pip install ../input/zarrlib/fasteners-0.16-py2.py3-none-any.whl')\nos.system('pip install ../input/zarrlib/asciitree-0.3.3-py2.py3-none-any.whl')\nos.system('pip install ../input/zarrlib/numcodecs-0.7.2-cp37-cp37m-manylinux2010_x86_64.whl')\nos.system('pip install ../input/zarrlib/zarr-2.6.1-py3-none-any.whl')\nimport zarr\nimport dask.array as da\n\ndef tifffile_to_dask(im_fp):\n    imdata = zarr.open(tifffile.imread(im_fp, aszarr=True))\n    if isinstance(imdata, zarr.hierarchy.Group):\n        # always return base layer if pyramidal\n        imdata = da.from_zarr(imdata[0])\n    else:\n        imdata = da.from_zarr(imdata)\n    return imdata\n\n\n#To read the tiff image: \n\nim = tifffile_to_dask(filename)\n</code></pre>\n<p>And another solution is to use an optimized rle converter (described here: <a href=\"https://www.kaggle.com/bguberfain/memory-aware-rle-encoding)\" target=\"_blank\">https://www.kaggle.com/bguberfain/memory-aware-rle-encoding)</a>.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1253200,
          "author_name": "houssemayed",
          "author_url": "",
          "post_date": "03/26/2021 13:15:15",
          "content": "<p>Thanks for ur help it's really interesting, but there's one more question, the obtained image has the same dimension as the original one ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1253207,
          "author_name": "joaorodriguez",
          "author_url": "",
          "post_date": "03/26/2021 13:19:55",
          "content": "<p>Yes, and you manipulate as an array, for example im[0:512,0:512] extract from disk the first tile of 512x512px</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1253226,
          "author_name": "houssemayed",
          "author_url": "",
          "post_date": "03/26/2021 13:42:07",
          "content": "<p>That's OK Thanks very muuch for ur help 🙂</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1251910": "Hello friends,\nI would like to ask if there are some techniques/ solutions that reduce memory consumption every suggestion is welcomed !",
    "1252395": "In my case, only use gc.collect and del variable was not sufficient to reduce memory overflow in the previous dataset. Right now im not facing the same problem, but the solution was to read only the tile of the tiff image from disk (described here: https://www.kaggle.com/nheathpatterson/discussion ). There was two problems with this solution: 1) It decreases a lot the performance when you read tile by tile (which is obvious, of course, because HDs are much slower than ram), but this can be solved reading whole \"rows\" of tiles. And 2) It uses zarr library that isn't installed in kaggles notebook by default, so you need to install without internet during the commit and submission of your code. But I spent my time to do this so here is the library: https://www.kaggle.com/joaorodriguez/zarrlib and here is the code for installation:\n\n```\nimport os\nos.system('pip install ../input/zarrlib/fasteners-0.16-py2.py3-none-any.whl')\nos.system('pip install ../input/zarrlib/asciitree-0.3.3-py2.py3-none-any.whl')\nos.system('pip install ../input/zarrlib/numcodecs-0.7.2-cp37-cp37m-manylinux2010_x86_64.whl')\nos.system('pip install ../input/zarrlib/zarr-2.6.1-py3-none-any.whl')\nimport zarr\nimport dask.array as da\n\ndef tifffile_to_dask(im_fp):\n    imdata = zarr.open(tifffile.imread(im_fp, aszarr=True))\n    if isinstance(imdata, zarr.hierarchy.Group):\n        # always return base layer if pyramidal\n        imdata = da.from_zarr(imdata[0])\n    else:\n        imdata = da.from_zarr(imdata)\n    return imdata\n\n\n#To read the tiff image: \n\nim = tifffile_to_dask(filename)\n\n```\n\nAnd another solution is to use an optimized rle converter (described here: https://www.kaggle.com/bguberfain/memory-aware-rle-encoding).",
    "1253200": "Thanks for ur help it's really interesting, but there's one more question, the obtained image has the same dimension as the original one ?",
    "1253207": "Yes, and you manipulate as an array, for example im[0:512,0:512] extract from disk the first tile of 512x512px",
    "1253226": "That's OK Thanks very muuch for ur help 🙂"
  },
  "source": "meta"
}