{
  "id": 235605,
  "title": "OOM issues",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/235605",
  "author_name": "",
  "post_date": "2021-04-30T11:32:37.247112900Z",
  "votes": 5,
  "comment_count": 2,
  "views": 0,
  "content": "<p>After 22h of GPU time used to debug the inference code, I still don't manage to solve OOM issues …</p>\n<p>When simply committing the notebook, I check the RAM just looking at the values provided by the kaggle interface: the maximum RAM usage reported is below 10Gb. However, this value is presumably smoothed over a given time range and doesn't necessary show the max RAM usage ?</p>\n<p>Is there a way to track more precisely the RAM usage ?</p>\n<p>Does somebody have a tip on the way to deal with such issue ?</p>",
  "messages": [
    {
      "id": "1288829",
      "postDate": "04/30/2021 11:32:37",
      "content": "<p>After 22h of GPU time used to debug the inference code, I still don't manage to solve OOM issues …</p>\n<p>When simply committing the notebook, I check the RAM just looking at the values provided by the kaggle interface: the maximum RAM usage reported is below 10Gb. However, this value is presumably smoothed over a given time range and doesn't necessary show the max RAM usage ?</p>\n<p>Is there a way to track more precisely the RAM usage ?</p>\n<p>Does somebody have a tip on the way to deal with such issue ?</p>",
      "rawMarkdown": "After 22h of GPU time used to debug the inference code, I still don't manage to solve OOM issues ...\n\nWhen simply committing the notebook, I check the RAM just looking at the values provided by the kaggle interface: the maximum RAM usage reported is below 10Gb. However, this value is presumably smoothed over a given time range and doesn't necessary show the max RAM usage ?\n\nIs there a way to track more precisely the RAM usage ?\n\nDoes somebody have a tip on the way to deal with such issue ?",
      "votes": null
    },
    {
      "id": "1289227",
      "postDate": "04/30/2021 18:59:22",
      "content": "<p><a href=\"https://www.kaggle.com/fabiendaniel\" target=\"_blank\">@fabiendaniel</a> You may use <code>psutil</code> package (in the recent RiiD competition was quite helpful) e.g. <br>\n<code>print('VM (%)', psutil.virtual_memory().percent)</code>; <br>\n<code>if psutil.virtual_memory().percent &gt;90: do_something</code><br>\nor search for similar commands </p>",
      "rawMarkdown": "fabiendaniel You may use `psutil` package (in the recent RiiD competition was quite helpful) e.g. \n`print('VM (%)', psutil.virtual_memory().percent)`; \n`if psutil.virtual_memory().percent >90: do_something`\nor search for similar commands",
      "votes": null
    },
    {
      "id": "1289547",
      "postDate": "05/01/2021 06:08:32",
      "content": "<p>Not sure what your approach is, but these are some ideas from the forum, some are old but should still apply. Ideas like using zarr, writing to disk may help. Also if you are using multiples of TTA, several models ensembled, perhaps reduce to a minimum to see what works.  If you are using pytorch, try torch.cuda.empty_cache() after finishing with an image, model, where appropriate. And gc.collect()</p>\n<p><a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/211638\" target=\"_blank\">Simple trick if you need more RAM</a></p>\n<p><a href=\"https://www.kaggle.com/mistag/inference-hubmap-u-net-mobilenetv2-256x256\" target=\"_blank\">to save memory the images are mapped to disk using numpy.memmap()</a></p>\n<p><a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/200145\" target=\"_blank\">Issue of Allocating more memory</a></p>\n<p>For reference, <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/161500\" target=\"_blank\">this</a>  a discussion from Prostate Cancer competition, see also in it, Dustin (Kaggle staff) comment on scratch space available in /tmp. </p>",
      "rawMarkdown": "Not sure what your approach is, but these are some ideas from the forum, some are old but should still apply. Ideas like using zarr, writing to disk may help. Also if you are using multiples of TTA, several models ensembled, perhaps reduce to a minimum to see what works.  If you are using pytorch, try torch.cuda.empty_cache() after finishing with an image, model, where appropriate. And gc.collect()\n\n[Simple trick if you need more RAM](https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/211638)\n\n[to save memory the images are mapped to disk using numpy.memmap()](https://www.kaggle.com/mistag/inference-hubmap-u-net-mobilenetv2-256x256)\n\n[Issue of Allocating more memory](https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/200145)\n\nFor reference, [this](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/161500)  a discussion from Prostate Cancer competition, see also in it, Dustin (Kaggle staff) comment on scratch space available in /tmp.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1289227,
      "author_name": "imeintanis",
      "author_url": "",
      "post_date": "04/30/2021 18:59:22",
      "content": "<p><a href=\"https://www.kaggle.com/fabiendaniel\" target=\"_blank\">@fabiendaniel</a> You may use <code>psutil</code> package (in the recent RiiD competition was quite helpful) e.g. <br>\n<code>print('VM (%)', psutil.virtual_memory().percent)</code>; <br>\n<code>if psutil.virtual_memory().percent &gt;90: do_something</code><br>\nor search for similar commands </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1289547,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "05/01/2021 06:08:32",
      "content": "<p>Not sure what your approach is, but these are some ideas from the forum, some are old but should still apply. Ideas like using zarr, writing to disk may help. Also if you are using multiples of TTA, several models ensembled, perhaps reduce to a minimum to see what works.  If you are using pytorch, try torch.cuda.empty_cache() after finishing with an image, model, where appropriate. And gc.collect()</p>\n<p><a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/211638\" target=\"_blank\">Simple trick if you need more RAM</a></p>\n<p><a href=\"https://www.kaggle.com/mistag/inference-hubmap-u-net-mobilenetv2-256x256\" target=\"_blank\">to save memory the images are mapped to disk using numpy.memmap()</a></p>\n<p><a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/200145\" target=\"_blank\">Issue of Allocating more memory</a></p>\n<p>For reference, <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/161500\" target=\"_blank\">this</a>  a discussion from Prostate Cancer competition, see also in it, Dustin (Kaggle staff) comment on scratch space available in /tmp. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1288829": "After 22h of GPU time used to debug the inference code, I still don't manage to solve OOM issues ...\n\nWhen simply committing the notebook, I check the RAM just looking at the values provided by the kaggle interface: the maximum RAM usage reported is below 10Gb. However, this value is presumably smoothed over a given time range and doesn't necessary show the max RAM usage ?\n\nIs there a way to track more precisely the RAM usage ?\n\nDoes somebody have a tip on the way to deal with such issue ?",
    "1289227": "fabiendaniel You may use `psutil` package (in the recent RiiD competition was quite helpful) e.g. \n`print('VM (%)', psutil.virtual_memory().percent)`; \n`if psutil.virtual_memory().percent >90: do_something`\nor search for similar commands",
    "1289547": "Not sure what your approach is, but these are some ideas from the forum, some are old but should still apply. Ideas like using zarr, writing to disk may help. Also if you are using multiples of TTA, several models ensembled, perhaps reduce to a minimum to see what works.  If you are using pytorch, try torch.cuda.empty_cache() after finishing with an image, model, where appropriate. And gc.collect()\n\n[Simple trick if you need more RAM](https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/211638)\n\n[to save memory the images are mapped to disk using numpy.memmap()](https://www.kaggle.com/mistag/inference-hubmap-u-net-mobilenetv2-256x256)\n\n[Issue of Allocating more memory](https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/200145)\n\nFor reference, [this](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/161500)  a discussion from Prostate Cancer competition, see also in it, Dustin (Kaggle staff) comment on scratch space available in /tmp."
  },
  "source": "meta"
}