{
  "id": 477246,
  "title": "How do you clean up RAM after big data transformations",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/477246",
  "author_name": "",
  "post_date": "2024-02-15T10:17:54.620612600Z",
  "votes": 11,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Python is notorious for its memory management, but we still (have to) love it.</p>\n<p>So how do you free up RAM after some memory-heavy operations?</p>\n<p>My strategy is to Run scripts in my Kaggle notebook.</p>\n<pre><code>%%\n!python run_feature_engineering  train    _3\n</code></pre>\n<p>After this, my RAM usage always comes back to ~900MB.</p>",
  "messages": [
    {
      "id": "2653251",
      "postDate": "02/15/2024 10:17:54",
      "content": "<p>Python is notorious for its memory management, but we still (have to) love it.</p>\n<p>So how do you free up RAM after some memory-heavy operations?</p>\n<p>My strategy is to Run scripts in my Kaggle notebook.</p>\n<pre><code>%%\n!python run_feature_engineering  train    _3\n</code></pre>\n<p>After this, my RAM usage always comes back to ~900MB.</p>",
      "rawMarkdown": "Python is notorious for its memory management, but we still (have to) love it.\n\nSo how do you free up RAM after some memory-heavy operations?\n\nMy strategy is to Run scripts in my Kaggle notebook.\n\n```\n%%time\n!python run_feature_engineering.py --ds train --level 2 --shard 1_3\n```\nAfter this, my RAM usage always comes back to ~900MB.",
      "votes": null
    },
    {
      "id": "2653358",
      "postDate": "02/15/2024 11:28:17",
      "content": "<p>Generally speaking:</p>\n<ul>\n<li>handle data types properly, typically shrink_dtypes for polars, track how dtypes are handled trough the DE pipeline</li>\n<li>avoid to create intermediate values, specifically within functions</li>\n<li>check big object in memory with:</li>\n</ul>\n<pre><code> ():\n    ()\n    ()\n     var_name  ():\n          var_name.startswith()  sys.getsizeof((var_name)) &gt; max_size:\n            ()\n</code></pre>\n<ul>\n<li>delete (del) named object that stay in memory</li>\n</ul>",
      "rawMarkdown": "Generally speaking:\n- handle data types properly, typically shrink_dtypes for polars, track how dtypes are handled trough the DE pipeline\n- avoid to create intermediate values, specifically within functions\n- check big object in memory with:\n\n```python\ndef print_memory_usage(max_size=10_000):\n    print(f\"|{'Variable Name': >25}|{'Memory': >10}|\")\n    print(\" ------------------------------------ \")\n    for var_name in dir():\n        if not var_name.startswith(\"_\") and sys.getsizeof(eval(var_name)) > max_size:\n            print(f\"|{var_name: >25}|{sys.getsizeof(eval(var_name)): >10}|\")\n```\n- delete (del) named object that stay in memory",
      "votes": null
    },
    {
      "id": "2654060",
      "postDate": "02/15/2024 20:08:02",
      "content": "<p>I usually use <code>gc.collect</code> and <code>ctypes.cdll().malloc.trim(0)</code>. I also avoid assigning copies of data-frames unnecessarily and use polars.lazyframes whenever possible <a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a> </p>",
      "rawMarkdown": "I usually use `gc.collect` and `ctypes.cdll().malloc.trim(0)`. I also avoid assigning copies of data-frames unnecessarily and use polars.lazyframes whenever possible @narsil",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2653358,
      "author_name": "lucasmorin",
      "author_url": "",
      "post_date": "02/15/2024 11:28:17",
      "content": "<p>Generally speaking:</p>\n<ul>\n<li>handle data types properly, typically shrink_dtypes for polars, track how dtypes are handled trough the DE pipeline</li>\n<li>avoid to create intermediate values, specifically within functions</li>\n<li>check big object in memory with:</li>\n</ul>\n<pre><code> ():\n    ()\n    ()\n     var_name  ():\n          var_name.startswith()  sys.getsizeof((var_name)) &gt; max_size:\n            ()\n</code></pre>\n<ul>\n<li>delete (del) named object that stay in memory</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2654060,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "02/15/2024 20:08:02",
      "content": "<p>I usually use <code>gc.collect</code> and <code>ctypes.cdll().malloc.trim(0)</code>. I also avoid assigning copies of data-frames unnecessarily and use polars.lazyframes whenever possible <a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2653251": "Python is notorious for its memory management, but we still (have to) love it.\n\nSo how do you free up RAM after some memory-heavy operations?\n\nMy strategy is to Run scripts in my Kaggle notebook.\n\n```\n%%time\n!python run_feature_engineering.py --ds train --level 2 --shard 1_3\n```\nAfter this, my RAM usage always comes back to ~900MB.",
    "2653358": "Generally speaking:\n- handle data types properly, typically shrink_dtypes for polars, track how dtypes are handled trough the DE pipeline\n- avoid to create intermediate values, specifically within functions\n- check big object in memory with:\n\n```python\ndef print_memory_usage(max_size=10_000):\n    print(f\"|{'Variable Name': >25}|{'Memory': >10}|\")\n    print(\" ------------------------------------ \")\n    for var_name in dir():\n        if not var_name.startswith(\"_\") and sys.getsizeof(eval(var_name)) > max_size:\n            print(f\"|{var_name: >25}|{sys.getsizeof(eval(var_name)): >10}|\")\n```\n- delete (del) named object that stay in memory",
    "2654060": "I usually use `gc.collect` and `ctypes.cdll().malloc.trim(0)`. I also avoid assigning copies of data-frames unnecessarily and use polars.lazyframes whenever possible @narsil"
  },
  "source": "meta"
}