{
  "id": 203020,
  "title": "How to Debug Memory usage and time consuming process",
  "url": "/competitions/riiid-test-answer-prediction/discussion/203020",
  "author_name": "",
  "post_date": "2020-12-13T09:20:30.377904500Z",
  "votes": 107,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Hi folks,<br>\nI have been struggling with limited memory in kaggle notebooks for feature engineering and I guess some of you do!<br>\nI'd like to share what I'm using  for that here. Hope it helps!</p>\n<h2>How to use</h2>\n<p>You just need to add <code>with trace('title')</code> to your code as follows.<br>\nIn this case you're merging questions_df which <em>might</em> increase memory usage drastically.</p>\n<pre><code>with trace(\"merge questions\"):\n    train_df = train_df.merge(questions_df, how='left', left_on='content_id', right_on='question_id')\n    valid_df = valid_df.merge(questions_df, how='left', left_on='content_id', right_on='question_id')\n</code></pre>\n<h2>What you'll see</h2>\n<p>It magically tells you how much memory you're using!<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2199749%2F7af0dd24a84aa139bd73d7d5976207e9%2F2020-12-13%2018.11.46.png?generation=1607851068654566&amp;alt=media\" alt=\"\"> </p>\n<h2>The code</h2>\n<p>Put this in your code :)</p>\n<pre><code>import psutil\nimport os\nimport time\nimport sys\nimport math\nfrom contextlib import contextmanager\n\n@contextmanager\ndef trace(title):\n    t0 = time.time()\n    p = psutil.Process(os.getpid())\n    m0 = p.memory_info().rss / 2. ** 30\n    yield\n    m1 = p.memory_info().rss / 2. ** 30\n    delta = m1 - m0\n    sign = '+' if delta &gt;= 0 else '-'\n    delta = math.fabs(delta)\n    print(f\"[{m1:.1f}GB({sign}{delta:.1f}GB):{time.time() - t0:.1f}sec] {title} \", file=sys.stderr)\n</code></pre>\n<p>2/22/2021: Updated to use memory_info().rss.</p>",
  "messages": [
    {
      "id": "1110995",
      "postDate": "12/13/2020 09:20:30",
      "content": "<p>Hi folks,<br>\nI have been struggling with limited memory in kaggle notebooks for feature engineering and I guess some of you do!<br>\nI'd like to share what I'm using  for that here. Hope it helps!</p>\n<h2>How to use</h2>\n<p>You just need to add <code>with trace('title')</code> to your code as follows.<br>\nIn this case you're merging questions_df which <em>might</em> increase memory usage drastically.</p>\n<pre><code>with trace(\"merge questions\"):\n    train_df = train_df.merge(questions_df, how='left', left_on='content_id', right_on='question_id')\n    valid_df = valid_df.merge(questions_df, how='left', left_on='content_id', right_on='question_id')\n</code></pre>\n<h2>What you'll see</h2>\n<p>It magically tells you how much memory you're using!<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2199749%2F7af0dd24a84aa139bd73d7d5976207e9%2F2020-12-13%2018.11.46.png?generation=1607851068654566&amp;alt=media\" alt=\"\"> </p>\n<h2>The code</h2>\n<p>Put this in your code :)</p>\n<pre><code>import psutil\nimport os\nimport time\nimport sys\nimport math\nfrom contextlib import contextmanager\n\n@contextmanager\ndef trace(title):\n    t0 = time.time()\n    p = psutil.Process(os.getpid())\n    m0 = p.memory_info().rss / 2. ** 30\n    yield\n    m1 = p.memory_info().rss / 2. ** 30\n    delta = m1 - m0\n    sign = '+' if delta &gt;= 0 else '-'\n    delta = math.fabs(delta)\n    print(f\"[{m1:.1f}GB({sign}{delta:.1f}GB):{time.time() - t0:.1f}sec] {title} \", file=sys.stderr)\n</code></pre>\n<p>2/22/2021: Updated to use memory_info().rss.</p>",
      "rawMarkdown": "Hi folks,\nI have been struggling with limited memory in kaggle notebooks for feature engineering and I guess some of you do!\nI'd like to share what I'm using  for that here. Hope it helps!\n\n## How to use\nYou just need to add ``with trace('title')`` to your code as follows.\nIn this case you're merging questions_df which *might* increase memory usage drastically.\n```\nwith trace(\"merge questions\"):\n    train_df = train_df.merge(questions_df, how='left', left_on='content_id', right_on='question_id')\n    valid_df = valid_df.merge(questions_df, how='left', left_on='content_id', right_on='question_id')\n```\n\n## What you'll see\nIt magically tells you how much memory you're using!\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2199749%2F7af0dd24a84aa139bd73d7d5976207e9%2F2020-12-13%2018.11.46.png?generation=1607851068654566&alt=media) \n\n## The code\nPut this in your code :)\n```\nimport psutil\nimport os\nimport time\nimport sys\nimport math\nfrom contextlib import contextmanager\n\n@contextmanager\ndef trace(title):\n    t0 = time.time()\n    p = psutil.Process(os.getpid())\n    m0 = p.memory_info().rss / 2. ** 30\n    yield\n    m1 = p.memory_info().rss / 2. ** 30\n    delta = m1 - m0\n    sign = '+' if delta >= 0 else '-'\n    delta = math.fabs(delta)\n    print(f\"[{m1:.1f}GB({sign}{delta:.1f}GB):{time.time() - t0:.1f}sec] {title} \", file=sys.stderr)\n\n```\n\n 2/22/2021: Updated to use memory_info().rss.",
      "votes": null
    },
    {
      "id": "1111011",
      "postDate": "12/13/2020 09:45:43",
      "content": "<p>Great, I will use this function!</p>",
      "rawMarkdown": "Great, I will use this function!",
      "votes": null
    },
    {
      "id": "1111022",
      "postDate": "12/13/2020 09:59:31",
      "content": "<pre><code>class Colors:\n    \"\"\"Defining Color Codes to color the text displayed on terminal.\n    \"\"\"\n\n    blue = \"\\033[94m\"\n    green = \"\\033[92m\"\n    yellow = \"\\033[93m\"\n    red = \"\\033[91m\"\n    end = \"\\033[0m\"\n\n\ndef color(string: str, color: Colors = Colors.yellow) -&gt; str:\n    return f\"{color}{string}{Colors.end}\"\n\n\n@contextmanager\ndef timer(label: str) -&gt; None:\n    \"\"\"compute the time the code block takes to run.\n    \"\"\"\n    p = psutil.Process(os.getpid())\n    start = time()  # Setup - __enter__\n    m0 = p.memory_info()[0] / 2. ** 30\n    print(color(f\"{label}: Start at {start}; RAM USAGE AT START {m0}\"))\n    try:\n        yield  # yield to body of `with` statement\n    finally:  # Teardown - __exit__\n        m1 = p.memory_info()[0] / 2. ** 30\n        delta = m1 - m0\n        sign = '+' if delta &gt;= 0 else '-'\n        delta = math.fabs(delta)\n        end = time()\n        print(color(f\"{label}: End at {end} ({end - start} elapsed); RAM USAGE AT END {m1:.2f}GB ({sign}{delta:.2f}GB)\", color=Colors.red))\n</code></pre>\n<p>I use this for tracking time! Will add memory tracking as well! (It adds colors to the texts etc). <br>\nEdit -&gt; I added yours as well now!</p>",
      "rawMarkdown": "```\nclass Colors:\n    \"\"\"Defining Color Codes to color the text displayed on terminal.\n    \"\"\"\n\n    blue = \"\\033[94m\"\n    green = \"\\033[92m\"\n    yellow = \"\\033[93m\"\n    red = \"\\033[91m\"\n    end = \"\\033[0m\"\n\n\ndef color(string: str, color: Colors = Colors.yellow) -> str:\n    return f\"{color}{string}{Colors.end}\"\n\n\n@contextmanager\ndef timer(label: str) -> None:\n    \"\"\"compute the time the code block takes to run.\n    \"\"\"\n    p = psutil.Process(os.getpid())\n    start = time()  # Setup - __enter__\n    m0 = p.memory_info()[0] / 2. ** 30\n    print(color(f\"{label}: Start at {start}; RAM USAGE AT START {m0}\"))\n    try:\n        yield  # yield to body of `with` statement\n    finally:  # Teardown - __exit__\n        m1 = p.memory_info()[0] / 2. ** 30\n        delta = m1 - m0\n        sign = '+' if delta >= 0 else '-'\n        delta = math.fabs(delta)\n        end = time()\n        print(color(f\"{label}: End at {end} ({end - start} elapsed); RAM USAGE AT END {m1:.2f}GB ({sign}{delta:.2f}GB)\", color=Colors.red))\n```\n\nI use this for tracking time! Will add memory tracking as well! (It adds colors to the texts etc). \nEdit -> I added yours as well now!",
      "votes": null
    },
    {
      "id": "1111036",
      "postDate": "12/13/2020 10:27:27",
      "content": "<p>Thank you. Glad to know you like it.</p>",
      "rawMarkdown": "Thank you. Glad to know you like it.",
      "votes": null
    },
    {
      "id": "1111038",
      "postDate": "12/13/2020 10:28:22",
      "content": "<p>Thank you. I like your idea of showing color!</p>",
      "rawMarkdown": "Thank you. I like your idea of showing color!",
      "votes": null
    },
    {
      "id": "1111835",
      "postDate": "12/14/2020 04:31:16",
      "content": "<p>I had similar problems trying to read large parquet files for the Bengali AI competition. Here is my notebook showing techniques for memory profiling:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/jamesmcguigan/reading-parquet-files-ram-cpu-optimization\" target=\"_blank\">https://www.kaggle.com/jamesmcguigan/reading-parquet-files-ram-cpu-optimization</a></li>\n</ul>",
      "rawMarkdown": "I had similar problems trying to read large parquet files for the Bengali AI competition. Here is my notebook showing techniques for memory profiling:\n- https://www.kaggle.com/jamesmcguigan/reading-parquet-files-ram-cpu-optimization",
      "votes": null
    },
    {
      "id": "1113167",
      "postDate": "12/15/2020 08:51:40",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/higepon\" target=\"_blank\">@higepon</a> 👍</p>",
      "rawMarkdown": "Thanks for sharing @higepon 👍",
      "votes": null
    },
    {
      "id": "1114054",
      "postDate": "12/16/2020 00:09:54",
      "content": "<p>Thanks for sharing James. Great summary.</p>",
      "rawMarkdown": "Thanks for sharing James. Great summary.",
      "votes": null
    },
    {
      "id": "1115119",
      "postDate": "12/16/2020 02:08:18",
      "content": "<p>Bravo, very useful, thanks for sharing</p>",
      "rawMarkdown": "Bravo, very useful, thanks for sharing",
      "votes": null
    },
    {
      "id": "1116155",
      "postDate": "12/16/2020 22:57:56",
      "content": "<p>Cool, very useful :) Thaks for sharing!</p>",
      "rawMarkdown": "Cool, very useful :) Thaks for sharing!",
      "votes": null
    },
    {
      "id": "1116296",
      "postDate": "12/17/2020 03:32:15",
      "content": "<p>Great, this is very useful for code competitions. Thanks😃</p>",
      "rawMarkdown": "Great, this is very useful for code competitions. Thanks😃",
      "votes": null
    },
    {
      "id": "1140848",
      "postDate": "01/06/2021 10:10:55",
      "content": "<p>If my submission does not exceed 9h limit in the current public leaderboard, can it be sure that it will not exceed the 9h limit later in the private leaderboard? Thanks.</p>",
      "rawMarkdown": "If my submission does not exceed 9h limit in the current public leaderboard, can it be sure that it will not exceed the 9h limit later in the private leaderboard? Thanks.",
      "votes": null
    },
    {
      "id": "1140855",
      "postDate": "01/06/2021 10:19:27",
      "content": "<p>In my understanding Yes, when u submit you already run inference for private data, only the score of private data is not shown</p>",
      "rawMarkdown": "In my understanding Yes, when u submit you already run inference for private data, only the score of private data is not shown",
      "votes": null
    },
    {
      "id": "1140866",
      "postDate": "01/06/2021 10:31:46",
      "content": "<p>Oh really? I even suspected that we are only inferencing for 20% of the total data now. So it means now we already inference for all the data, not only 20%?</p>",
      "rawMarkdown": "Oh really? I even suspected that we are only inferencing for 20% of the total data now. So it means now we already inference for all the data, not only 20%?",
      "votes": null
    },
    {
      "id": "1142150",
      "postDate": "01/07/2021 07:38:14",
      "content": "<p>You can also add <code>%%time</code> to the top of a jupyter notebook cell</p>",
      "rawMarkdown": "You can also add `%%time` to the top of a jupyter notebook cell",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1111011,
      "author_name": "mamasinkgs",
      "author_url": "",
      "post_date": "12/13/2020 09:45:43",
      "content": "<p>Great, I will use this function!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1111036,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "12/13/2020 10:27:27",
          "content": "<p>Thank you. Glad to know you like it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1111022,
      "author_name": "adityaecdrid",
      "author_url": "",
      "post_date": "12/13/2020 09:59:31",
      "content": "<pre><code>class Colors:\n    \"\"\"Defining Color Codes to color the text displayed on terminal.\n    \"\"\"\n\n    blue = \"\\033[94m\"\n    green = \"\\033[92m\"\n    yellow = \"\\033[93m\"\n    red = \"\\033[91m\"\n    end = \"\\033[0m\"\n\n\ndef color(string: str, color: Colors = Colors.yellow) -&gt; str:\n    return f\"{color}{string}{Colors.end}\"\n\n\n@contextmanager\ndef timer(label: str) -&gt; None:\n    \"\"\"compute the time the code block takes to run.\n    \"\"\"\n    p = psutil.Process(os.getpid())\n    start = time()  # Setup - __enter__\n    m0 = p.memory_info()[0] / 2. ** 30\n    print(color(f\"{label}: Start at {start}; RAM USAGE AT START {m0}\"))\n    try:\n        yield  # yield to body of `with` statement\n    finally:  # Teardown - __exit__\n        m1 = p.memory_info()[0] / 2. ** 30\n        delta = m1 - m0\n        sign = '+' if delta &gt;= 0 else '-'\n        delta = math.fabs(delta)\n        end = time()\n        print(color(f\"{label}: End at {end} ({end - start} elapsed); RAM USAGE AT END {m1:.2f}GB ({sign}{delta:.2f}GB)\", color=Colors.red))\n</code></pre>\n<p>I use this for tracking time! Will add memory tracking as well! (It adds colors to the texts etc). <br>\nEdit -&gt; I added yours as well now!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1111038,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "12/13/2020 10:28:22",
          "content": "<p>Thank you. I like your idea of showing color!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1142150,
          "author_name": "jamesmcguigan",
          "author_url": "",
          "post_date": "01/07/2021 07:38:14",
          "content": "<p>You can also add <code>%%time</code> to the top of a jupyter notebook cell</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1111835,
      "author_name": "jamesmcguigan",
      "author_url": "",
      "post_date": "12/14/2020 04:31:16",
      "content": "<p>I had similar problems trying to read large parquet files for the Bengali AI competition. Here is my notebook showing techniques for memory profiling:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/jamesmcguigan/reading-parquet-files-ram-cpu-optimization\" target=\"_blank\">https://www.kaggle.com/jamesmcguigan/reading-parquet-files-ram-cpu-optimization</a></li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1114054,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "12/16/2020 00:09:54",
          "content": "<p>Thanks for sharing James. Great summary.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1113167,
      "author_name": "datajameson",
      "author_url": "",
      "post_date": "12/15/2020 08:51:40",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/higepon\" target=\"_blank\">@higepon</a> 👍</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1115119,
      "author_name": "superchenhao",
      "author_url": "",
      "post_date": "12/16/2020 02:08:18",
      "content": "<p>Bravo, very useful, thanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1116155,
      "author_name": "nyanpn",
      "author_url": "",
      "post_date": "12/16/2020 22:57:56",
      "content": "<p>Cool, very useful :) Thaks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1116296,
      "author_name": "ttahara",
      "author_url": "",
      "post_date": "12/17/2020 03:32:15",
      "content": "<p>Great, this is very useful for code competitions. Thanks😃</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1140848,
      "author_name": "",
      "author_url": "",
      "post_date": "01/06/2021 10:10:55",
      "content": "<p>If my submission does not exceed 9h limit in the current public leaderboard, can it be sure that it will not exceed the 9h limit later in the private leaderboard? Thanks.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1140855,
          "author_name": "superchenhao",
          "author_url": "",
          "post_date": "01/06/2021 10:19:27",
          "content": "<p>In my understanding Yes, when u submit you already run inference for private data, only the score of private data is not shown</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1140866,
      "author_name": "",
      "author_url": "",
      "post_date": "01/06/2021 10:31:46",
      "content": "<p>Oh really? I even suspected that we are only inferencing for 20% of the total data now. So it means now we already inference for all the data, not only 20%?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1110995": "Hi folks,\nI have been struggling with limited memory in kaggle notebooks for feature engineering and I guess some of you do!\nI'd like to share what I'm using  for that here. Hope it helps!\n\n## How to use\nYou just need to add ``with trace('title')`` to your code as follows.\nIn this case you're merging questions_df which *might* increase memory usage drastically.\n```\nwith trace(\"merge questions\"):\n    train_df = train_df.merge(questions_df, how='left', left_on='content_id', right_on='question_id')\n    valid_df = valid_df.merge(questions_df, how='left', left_on='content_id', right_on='question_id')\n```\n\n## What you'll see\nIt magically tells you how much memory you're using!\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2199749%2F7af0dd24a84aa139bd73d7d5976207e9%2F2020-12-13%2018.11.46.png?generation=1607851068654566&alt=media) \n\n## The code\nPut this in your code :)\n```\nimport psutil\nimport os\nimport time\nimport sys\nimport math\nfrom contextlib import contextmanager\n\n@contextmanager\ndef trace(title):\n    t0 = time.time()\n    p = psutil.Process(os.getpid())\n    m0 = p.memory_info().rss / 2. ** 30\n    yield\n    m1 = p.memory_info().rss / 2. ** 30\n    delta = m1 - m0\n    sign = '+' if delta >= 0 else '-'\n    delta = math.fabs(delta)\n    print(f\"[{m1:.1f}GB({sign}{delta:.1f}GB):{time.time() - t0:.1f}sec] {title} \", file=sys.stderr)\n\n```\n\n 2/22/2021: Updated to use memory_info().rss.",
    "1111011": "Great, I will use this function!",
    "1111022": "```\nclass Colors:\n    \"\"\"Defining Color Codes to color the text displayed on terminal.\n    \"\"\"\n\n    blue = \"\\033[94m\"\n    green = \"\\033[92m\"\n    yellow = \"\\033[93m\"\n    red = \"\\033[91m\"\n    end = \"\\033[0m\"\n\n\ndef color(string: str, color: Colors = Colors.yellow) -> str:\n    return f\"{color}{string}{Colors.end}\"\n\n\n@contextmanager\ndef timer(label: str) -> None:\n    \"\"\"compute the time the code block takes to run.\n    \"\"\"\n    p = psutil.Process(os.getpid())\n    start = time()  # Setup - __enter__\n    m0 = p.memory_info()[0] / 2. ** 30\n    print(color(f\"{label}: Start at {start}; RAM USAGE AT START {m0}\"))\n    try:\n        yield  # yield to body of `with` statement\n    finally:  # Teardown - __exit__\n        m1 = p.memory_info()[0] / 2. ** 30\n        delta = m1 - m0\n        sign = '+' if delta >= 0 else '-'\n        delta = math.fabs(delta)\n        end = time()\n        print(color(f\"{label}: End at {end} ({end - start} elapsed); RAM USAGE AT END {m1:.2f}GB ({sign}{delta:.2f}GB)\", color=Colors.red))\n```\n\nI use this for tracking time! Will add memory tracking as well! (It adds colors to the texts etc). \nEdit -> I added yours as well now!",
    "1111036": "Thank you. Glad to know you like it.",
    "1111038": "Thank you. I like your idea of showing color!",
    "1111835": "I had similar problems trying to read large parquet files for the Bengali AI competition. Here is my notebook showing techniques for memory profiling:\n- https://www.kaggle.com/jamesmcguigan/reading-parquet-files-ram-cpu-optimization",
    "1113167": "Thanks for sharing @higepon 👍",
    "1114054": "Thanks for sharing James. Great summary.",
    "1115119": "Bravo, very useful, thanks for sharing",
    "1116155": "Cool, very useful :) Thaks for sharing!",
    "1116296": "Great, this is very useful for code competitions. Thanks😃",
    "1140848": "If my submission does not exceed 9h limit in the current public leaderboard, can it be sure that it will not exceed the 9h limit later in the private leaderboard? Thanks.",
    "1140855": "In my understanding Yes, when u submit you already run inference for private data, only the score of private data is not shown",
    "1140866": "Oh really? I even suspected that we are only inferencing for 20% of the total data now. So it means now we already inference for all the data, not only 20%?",
    "1142150": "You can also add `%%time` to the top of a jupyter notebook cell"
  },
  "source": "meta"
}