{
  "id": 200145,
  "title": "Issue of Allocating more memory",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/200145",
  "author_name": "",
  "post_date": "2020-11-29T05:19:15.810992400Z",
  "votes": 10,
  "comment_count": 12,
  "views": 0,
  "content": "<p>I cant inference, my model, due to allocation of too much memory in RAM<br>\nHave anyone fixed this issue.<br>\n<a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> said it will raise an issue like this on Private dataset.<br>\nis there any possible ways to avoid<br>\nI think current top LB scorer may have fixed them<br>\n<a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> <a href=\"https://www.kaggle.com/vladimirgroza\" target=\"_blank\">@vladimirgroza</a> <a href=\"https://www.kaggle.com/tivfrvqhs5\" target=\"_blank\">@tivfrvqhs5</a> also addresses to host even</p>",
  "messages": [
    {
      "id": "1094929",
      "postDate": "11/29/2020 05:19:15",
      "content": "<p>I cant inference, my model, due to allocation of too much memory in RAM<br>\nHave anyone fixed this issue.<br>\n<a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> said it will raise an issue like this on Private dataset.<br>\nis there any possible ways to avoid<br>\nI think current top LB scorer may have fixed them<br>\n<a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> <a href=\"https://www.kaggle.com/vladimirgroza\" target=\"_blank\">@vladimirgroza</a> <a href=\"https://www.kaggle.com/tivfrvqhs5\" target=\"_blank\">@tivfrvqhs5</a> also addresses to host even</p>",
      "rawMarkdown": "I cant inference, my model, due to allocation of too much memory in RAM\nHave anyone fixed this issue.\n@iafoss said it will raise an issue like this on Private dataset.\nis there any possible ways to avoid\nI think current top LB scorer may have fixed them\n@haqishen @vladimirgroza @tivfrvqhs5 also addresses to host even",
      "votes": null
    },
    {
      "id": "1102320",
      "postDate": "12/04/2020 19:42:21",
      "content": "<p><a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a>:</p>\n<p>I solved this problem by using the hard disk as a working space. Basically I convert each image to TFREcord format(in tiles), write them to a file and then pass it as a tf.Data.Dataset to the model that will then read it serially. I was able to submit to private LB with no problems. Here is the notebook:<br>\n<a href=\"https://www.kaggle.com/marcosnovaes/hubmap-memory-efficient-submission-using-disk/\" target=\"_blank\">https://www.kaggle.com/marcosnovaes/hubmap-memory-efficient-submission-using-disk/</a></p>\n<p>Hope it helps! Cheers!</p>\n<p>--Marcos</p>",
      "rawMarkdown": "morizin:\n\nI solved this problem by using the hard disk as a working space. Basically I convert each image to TFREcord format(in tiles), write them to a file and then pass it as a tf.Data.Dataset to the model that will then read it serially. I was able to submit to private LB with no problems. Here is the notebook:\n[https://www.kaggle.com/marcosnovaes/hubmap-memory-efficient-submission-using-disk/](https://www.kaggle.com/marcosnovaes/hubmap-memory-efficient-submission-using-disk/)\n\nHope it helps! Cheers!\n\n--Marcos",
      "votes": null
    },
    {
      "id": "1102543",
      "postDate": "12/05/2020 03:30:06",
      "content": "<p><a href=\"https://www.kaggle.com/marcosnovaes\" target=\"_blank\">@marcosnovaes</a> <br>\nSorry , i lost to mention I am a pytorch user. please can you provide information about it</p>",
      "rawMarkdown": "marcosnovaes \nSorry , i lost to mention I am a pytorch user. please can you provide information about it",
      "votes": null
    },
    {
      "id": "1103077",
      "postDate": "12/05/2020 16:02:19",
      "content": "<p>You can maybe try Google Cloud AI Notebooks where you allocate larger RAM?</p>",
      "rawMarkdown": "You can maybe try Google Cloud AI Notebooks where you allocate larger RAM?",
      "votes": null
    },
    {
      "id": "1103093",
      "postDate": "12/05/2020 16:19:44",
      "content": "<p>is it possible? . because this is a Code competition and the submission must be done from Kaggle Kernel or else I was wrong.<br>\nHave you succeeded in inferencing by making that notebook a Google Cloud AI Notebook ? </p>",
      "rawMarkdown": "is it possible? . because this is a Code competition and the submission must be done from Kaggle Kernel or else I was wrong.\nHave you succeeded in inferencing by making that notebook a Google Cloud AI Notebook ?",
      "votes": null
    },
    {
      "id": "1103102",
      "postDate": "12/05/2020 16:31:14",
      "content": "<p>You are right sorry didn't think of that, haven't tried it yet. There's also information found here <a href=\"https://www.kaggle.com/product-feedback/159602\" target=\"_blank\">https://www.kaggle.com/product-feedback/159602</a>. Tensorflow Cloud is also available <a href=\"https://www.kaggle.com/rosebv/train-model-with-tensorflow-cloud\" target=\"_blank\">https://www.kaggle.com/rosebv/train-model-with-tensorflow-cloud</a>, but you're using pytorch I see.  Hopefully someone will comment who has tried it, it would like to know these things as well.</p>",
      "rawMarkdown": "You are right sorry didn't think of that, haven't tried it yet. There's also information found here https://www.kaggle.com/product-feedback/159602. Tensorflow Cloud is also available https://www.kaggle.com/rosebv/train-model-with-tensorflow-cloud, but you're using pytorch I see.  Hopefully someone will comment who has tried it, it would like to know these things as well.",
      "votes": null
    },
    {
      "id": "1104836",
      "postDate": "12/07/2020 09:10:32",
      "content": "<p>Having the same problem here ☹️ I think it is about time for Kaggle to increase the memory on the kernels.<br>\nI've spent more time trying to work around the limitations, by writing intermediate results to disk, using garbage collector, etc then actually solving the problem. Very frustrating.<br>\n<a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> Did you find a solution for the memory issue?</p>",
      "rawMarkdown": "Having the same problem here ☹️ I think it is about time for Kaggle to increase the memory on the kernels.\nI've spent more time trying to work around the limitations, by writing intermediate results to disk, using garbage collector, etc then actually solving the problem. Very frustrating.\n@morizin Did you find a solution for the memory issue?",
      "votes": null
    },
    {
      "id": "1104850",
      "postDate": "12/07/2020 09:32:19",
      "content": "<p>i was wondering if the organizer can expose the largest private test image file dimensions (e.g. width x height).<br>\nThen can I can work out the memory usage and developed locally.</p>\n<p>if not I will be wasting submission to solve out of memory issues.</p>\n<p>i think many solutions are possible, but we need an environment to test and debug our code.<br>\n(e.g. i google several python/C package to read large tiff partially, etc)</p>\n<hr>\n<p>I am exploring slideio</p>\n<p>e.g. <a href=\"https://towardsdatascience.com/slideio-a-new-python-library-for-reading-medical-images-11858a522059\" target=\"_blank\">https://towardsdatascience.com/slideio-a-new-python-library-for-reading-medical-images-11858a522059</a><br>\n<a href=\"https://pypi.org/project/slideio/\" target=\"_blank\">https://pypi.org/project/slideio/</a><br>\n<a href=\"http://slideio.com/\" target=\"_blank\">http://slideio.com/</a></p>\n<p>\"Normally it is not possible to read the whole image at the original scale because of the large size. In this case, the program can retrieve a region of the image, down-scale it to the acceptable size, or retrieve a down-scaled region.  </p>\n<pre><code>The code snippet below reads a rectangle region from the image and down-scales it to a 500 pixels width picture.\n\nimage = scene.read_block((5000, 5000, 5000, 5000), size=(500,0))\nplt.imshow(image)\n</code></pre>",
      "rawMarkdown": "i was wondering if the organizer can expose the largest private test image file dimensions (e.g. width x height).\nThen can I can work out the memory usage and developed locally.\n\nif not I will be wasting submission to solve out of memory issues.\n\ni think many solutions are possible, but we need an environment to test and debug our code.\n(e.g. i google several python/C package to read large tiff partially, etc)\n\n---\nI am exploring slideio\n\ne.g. https://towardsdatascience.com/slideio-a-new-python-library-for-reading-medical-images-11858a522059\nhttps://pypi.org/project/slideio/\nhttp://slideio.com/\n\n\"Normally it is not possible to read the whole image at the original scale because of the large size. In this case, the program can retrieve a region of the image, down-scale it to the acceptable size, or retrieve a down-scaled region.  \n```\n\nThe code snippet below reads a rectangle region from the image and down-scales it to a 500 pixels width picture.\n\nimage = scene.read_block((5000, 5000, 5000, 5000), size=(500,0))\nplt.imshow(image)\n```",
      "votes": null
    },
    {
      "id": "1104929",
      "postDate": "12/07/2020 11:28:08",
      "content": "<p>Still no                          </p>",
      "rawMarkdown": "Still no",
      "votes": null
    },
    {
      "id": "1109415",
      "postDate": "12/11/2020 16:45:48",
      "content": "<p>Is anyone has found yet a solution for this</p>",
      "rawMarkdown": "Is anyone has found yet a solution for this",
      "votes": null
    },
    {
      "id": "1109462",
      "postDate": "12/11/2020 17:54:55",
      "content": "<p>Another idea…</p>\n<p>You can use a combination of <code>tifffile&gt;=2020.10.1</code>, <code>zarr</code>, and <code>dask</code> to read sub-regions of the image without loading the entire images into memory.</p>\n<p><code>tifffile</code> can map the tiles within the images as a <code>zarr</code> store and and <code>dask</code> allows distributed computing from the store and <code>numpy</code>-like slicing.</p>\n<pre><code>import zarr\nimport dask.array as da\nfrom tifffile import imread\n\ndef tifffile_to_dask(im_fp):\n    imdata = zarr.open(imread(im_fp, aszarr=True))\n    if isinstance(imdata, zarr.hierarchy.Group):\n        # always return base layer if pyramidal\n        imdata = da.from_zarr(imdata[0])\n    else:\n        imdata = da.from_zarr(imdata)\n    return imdata\n\n# map data in tiff file to a dask.array\ndask_im = tifffile_to_dask(image_fp)\n\ndask_subregion = dask_im[0:1000,0:1000,:]\n</code></pre>\n<p>This process will only consume a small amount of memory.</p>",
      "rawMarkdown": "Another idea...\n\nYou can use a combination of `tifffile>=2020.10.1`, `zarr`, and `dask` to read sub-regions of the image without loading the entire images into memory.\n\n`tifffile` can map the tiles within the images as a `zarr` store and and `dask` allows distributed computing from the store and `numpy`-like slicing.\n\n```python\nimport zarr\nimport dask.array as da\nfrom tifffile import imread\n\ndef tifffile_to_dask(im_fp):\n    imdata = zarr.open(imread(im_fp, aszarr=True))\n    if isinstance(imdata, zarr.hierarchy.Group):\n        # always return base layer if pyramidal\n        imdata = da.from_zarr(imdata[0])\n    else:\n        imdata = da.from_zarr(imdata)\n    return imdata\n\n# map data in tiff file to a dask.array\ndask_im = tifffile_to_dask(image_fp)\n\ndask_subregion = dask_im[0:1000,0:1000,:]\n\n```\n\nThis process will only consume a small amount of memory.",
      "votes": null
    },
    {
      "id": "1110065",
      "postDate": "12/12/2020 11:51:22",
      "content": "<p>You can read in the TIF file, then map the image numpy array to disk using numpy.memmap(). A little slower than having the whole image in memory, but uses a lot less memory. You can find the code for mapping image to disk in <a href=\"https://www.kaggle.com/mistag/inference-hubmap-u-net-mobilenetv2-256x256\" target=\"_blank\">this notebook</a></p>",
      "rawMarkdown": "You can read in the TIF file, then map the image numpy array to disk using numpy.memmap(). A little slower than having the whole image in memory, but uses a lot less memory. You can find the code for mapping image to disk in [this notebook](https://www.kaggle.com/mistag/inference-hubmap-u-net-mobilenetv2-256x256)",
      "votes": null
    },
    {
      "id": "1110324",
      "postDate": "12/12/2020 16:12:08",
      "content": "<p>Thank you even if don't work. I will make a try with this.</p>",
      "rawMarkdown": "Thank you even if don't work. I will make a try with this.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1102320,
      "author_name": "marcosnovaes",
      "author_url": "",
      "post_date": "12/04/2020 19:42:21",
      "content": "<p><a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a>:</p>\n<p>I solved this problem by using the hard disk as a working space. Basically I convert each image to TFREcord format(in tiles), write them to a file and then pass it as a tf.Data.Dataset to the model that will then read it serially. I was able to submit to private LB with no problems. Here is the notebook:<br>\n<a href=\"https://www.kaggle.com/marcosnovaes/hubmap-memory-efficient-submission-using-disk/\" target=\"_blank\">https://www.kaggle.com/marcosnovaes/hubmap-memory-efficient-submission-using-disk/</a></p>\n<p>Hope it helps! Cheers!</p>\n<p>--Marcos</p>",
      "votes": null,
      "replies": [
        {
          "id": 1102543,
          "author_name": "morizin",
          "author_url": "",
          "post_date": "12/05/2020 03:30:06",
          "content": "<p><a href=\"https://www.kaggle.com/marcosnovaes\" target=\"_blank\">@marcosnovaes</a> <br>\nSorry , i lost to mention I am a pytorch user. please can you provide information about it</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1103077,
      "author_name": "bouweceunen",
      "author_url": "",
      "post_date": "12/05/2020 16:02:19",
      "content": "<p>You can maybe try Google Cloud AI Notebooks where you allocate larger RAM?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1103093,
          "author_name": "morizin",
          "author_url": "",
          "post_date": "12/05/2020 16:19:44",
          "content": "<p>is it possible? . because this is a Code competition and the submission must be done from Kaggle Kernel or else I was wrong.<br>\nHave you succeeded in inferencing by making that notebook a Google Cloud AI Notebook ? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1103102,
          "author_name": "bouweceunen",
          "author_url": "",
          "post_date": "12/05/2020 16:31:14",
          "content": "<p>You are right sorry didn't think of that, haven't tried it yet. There's also information found here <a href=\"https://www.kaggle.com/product-feedback/159602\" target=\"_blank\">https://www.kaggle.com/product-feedback/159602</a>. Tensorflow Cloud is also available <a href=\"https://www.kaggle.com/rosebv/train-model-with-tensorflow-cloud\" target=\"_blank\">https://www.kaggle.com/rosebv/train-model-with-tensorflow-cloud</a>, but you're using pytorch I see.  Hopefully someone will comment who has tried it, it would like to know these things as well.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1104836,
      "author_name": "ngcferreira",
      "author_url": "",
      "post_date": "12/07/2020 09:10:32",
      "content": "<p>Having the same problem here ☹️ I think it is about time for Kaggle to increase the memory on the kernels.<br>\nI've spent more time trying to work around the limitations, by writing intermediate results to disk, using garbage collector, etc then actually solving the problem. Very frustrating.<br>\n<a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> Did you find a solution for the memory issue?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1104929,
          "author_name": "morizin",
          "author_url": "",
          "post_date": "12/07/2020 11:28:08",
          "content": "<p>Still no                          </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1104850,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "12/07/2020 09:32:19",
      "content": "<p>i was wondering if the organizer can expose the largest private test image file dimensions (e.g. width x height).<br>\nThen can I can work out the memory usage and developed locally.</p>\n<p>if not I will be wasting submission to solve out of memory issues.</p>\n<p>i think many solutions are possible, but we need an environment to test and debug our code.<br>\n(e.g. i google several python/C package to read large tiff partially, etc)</p>\n<hr>\n<p>I am exploring slideio</p>\n<p>e.g. <a href=\"https://towardsdatascience.com/slideio-a-new-python-library-for-reading-medical-images-11858a522059\" target=\"_blank\">https://towardsdatascience.com/slideio-a-new-python-library-for-reading-medical-images-11858a522059</a><br>\n<a href=\"https://pypi.org/project/slideio/\" target=\"_blank\">https://pypi.org/project/slideio/</a><br>\n<a href=\"http://slideio.com/\" target=\"_blank\">http://slideio.com/</a></p>\n<p>\"Normally it is not possible to read the whole image at the original scale because of the large size. In this case, the program can retrieve a region of the image, down-scale it to the acceptable size, or retrieve a down-scaled region.  </p>\n<pre><code>The code snippet below reads a rectangle region from the image and down-scales it to a 500 pixels width picture.\n\nimage = scene.read_block((5000, 5000, 5000, 5000), size=(500,0))\nplt.imshow(image)\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1109462,
          "author_name": "nheathpatterson",
          "author_url": "",
          "post_date": "12/11/2020 17:54:55",
          "content": "<p>Another idea…</p>\n<p>You can use a combination of <code>tifffile&gt;=2020.10.1</code>, <code>zarr</code>, and <code>dask</code> to read sub-regions of the image without loading the entire images into memory.</p>\n<p><code>tifffile</code> can map the tiles within the images as a <code>zarr</code> store and and <code>dask</code> allows distributed computing from the store and <code>numpy</code>-like slicing.</p>\n<pre><code>import zarr\nimport dask.array as da\nfrom tifffile import imread\n\ndef tifffile_to_dask(im_fp):\n    imdata = zarr.open(imread(im_fp, aszarr=True))\n    if isinstance(imdata, zarr.hierarchy.Group):\n        # always return base layer if pyramidal\n        imdata = da.from_zarr(imdata[0])\n    else:\n        imdata = da.from_zarr(imdata)\n    return imdata\n\n# map data in tiff file to a dask.array\ndask_im = tifffile_to_dask(image_fp)\n\ndask_subregion = dask_im[0:1000,0:1000,:]\n</code></pre>\n<p>This process will only consume a small amount of memory.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1109415,
      "author_name": "morizin",
      "author_url": "",
      "post_date": "12/11/2020 16:45:48",
      "content": "<p>Is anyone has found yet a solution for this</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1110065,
      "author_name": "mistag",
      "author_url": "",
      "post_date": "12/12/2020 11:51:22",
      "content": "<p>You can read in the TIF file, then map the image numpy array to disk using numpy.memmap(). A little slower than having the whole image in memory, but uses a lot less memory. You can find the code for mapping image to disk in <a href=\"https://www.kaggle.com/mistag/inference-hubmap-u-net-mobilenetv2-256x256\" target=\"_blank\">this notebook</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1110324,
          "author_name": "morizin",
          "author_url": "",
          "post_date": "12/12/2020 16:12:08",
          "content": "<p>Thank you even if don't work. I will make a try with this.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1094929": "I cant inference, my model, due to allocation of too much memory in RAM\nHave anyone fixed this issue.\n@iafoss said it will raise an issue like this on Private dataset.\nis there any possible ways to avoid\nI think current top LB scorer may have fixed them\n@haqishen @vladimirgroza @tivfrvqhs5 also addresses to host even",
    "1102320": "morizin:\n\nI solved this problem by using the hard disk as a working space. Basically I convert each image to TFREcord format(in tiles), write them to a file and then pass it as a tf.Data.Dataset to the model that will then read it serially. I was able to submit to private LB with no problems. Here is the notebook:\n[https://www.kaggle.com/marcosnovaes/hubmap-memory-efficient-submission-using-disk/](https://www.kaggle.com/marcosnovaes/hubmap-memory-efficient-submission-using-disk/)\n\nHope it helps! Cheers!\n\n--Marcos",
    "1102543": "marcosnovaes \nSorry , i lost to mention I am a pytorch user. please can you provide information about it",
    "1103077": "You can maybe try Google Cloud AI Notebooks where you allocate larger RAM?",
    "1103093": "is it possible? . because this is a Code competition and the submission must be done from Kaggle Kernel or else I was wrong.\nHave you succeeded in inferencing by making that notebook a Google Cloud AI Notebook ?",
    "1103102": "You are right sorry didn't think of that, haven't tried it yet. There's also information found here https://www.kaggle.com/product-feedback/159602. Tensorflow Cloud is also available https://www.kaggle.com/rosebv/train-model-with-tensorflow-cloud, but you're using pytorch I see.  Hopefully someone will comment who has tried it, it would like to know these things as well.",
    "1104836": "Having the same problem here ☹️ I think it is about time for Kaggle to increase the memory on the kernels.\nI've spent more time trying to work around the limitations, by writing intermediate results to disk, using garbage collector, etc then actually solving the problem. Very frustrating.\n@morizin Did you find a solution for the memory issue?",
    "1104850": "i was wondering if the organizer can expose the largest private test image file dimensions (e.g. width x height).\nThen can I can work out the memory usage and developed locally.\n\nif not I will be wasting submission to solve out of memory issues.\n\ni think many solutions are possible, but we need an environment to test and debug our code.\n(e.g. i google several python/C package to read large tiff partially, etc)\n\n---\nI am exploring slideio\n\ne.g. https://towardsdatascience.com/slideio-a-new-python-library-for-reading-medical-images-11858a522059\nhttps://pypi.org/project/slideio/\nhttp://slideio.com/\n\n\"Normally it is not possible to read the whole image at the original scale because of the large size. In this case, the program can retrieve a region of the image, down-scale it to the acceptable size, or retrieve a down-scaled region.  \n```\n\nThe code snippet below reads a rectangle region from the image and down-scales it to a 500 pixels width picture.\n\nimage = scene.read_block((5000, 5000, 5000, 5000), size=(500,0))\nplt.imshow(image)\n```",
    "1104929": "Still no",
    "1109415": "Is anyone has found yet a solution for this",
    "1109462": "Another idea...\n\nYou can use a combination of `tifffile>=2020.10.1`, `zarr`, and `dask` to read sub-regions of the image without loading the entire images into memory.\n\n`tifffile` can map the tiles within the images as a `zarr` store and and `dask` allows distributed computing from the store and `numpy`-like slicing.\n\n```python\nimport zarr\nimport dask.array as da\nfrom tifffile import imread\n\ndef tifffile_to_dask(im_fp):\n    imdata = zarr.open(imread(im_fp, aszarr=True))\n    if isinstance(imdata, zarr.hierarchy.Group):\n        # always return base layer if pyramidal\n        imdata = da.from_zarr(imdata[0])\n    else:\n        imdata = da.from_zarr(imdata)\n    return imdata\n\n# map data in tiff file to a dask.array\ndask_im = tifffile_to_dask(image_fp)\n\ndask_subregion = dask_im[0:1000,0:1000,:]\n\n```\n\nThis process will only consume a small amount of memory.",
    "1110065": "You can read in the TIF file, then map the image numpy array to disk using numpy.memmap(). A little slower than having the whole image in memory, but uses a lot less memory. You can find the code for mapping image to disk in [this notebook](https://www.kaggle.com/mistag/inference-hubmap-u-net-mobilenetv2-256x256)",
    "1110324": "Thank you even if don't work. I will make a try with this."
  },
  "source": "meta"
}