{
  "id": 201444,
  "title": "Memory Profiling in Python | Reduce your RAM Overload",
  "url": "/competitions/riiid-test-answer-prediction/discussion/201444",
  "author_name": "",
  "post_date": "2020-12-05T04:23:35.063436200Z",
  "votes": 49,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi there,</p>\n<p>Are you facing Out of Memory issues while running Kernels? Are you able to debug them? Many people have pointed out that this competition is as much as about Model Building as of Software Engineering. Here are two things I use to solve memory issues for my Kernels. </p>\n<h2>Summary Tracker</h2>\n<p>Summary Tracker helps us to find out the objects with their memory usage. We can use sorting to find the top 10 objects using the code below.</p>\n<pre><code>from operator import itemgetter\nfrom pympler import tracker\n\nmem = tracker.SummaryTracker()\nprint(sorted(mem.create_summary(), reverse=True, key=itemgetter(2))[:10])\n</code></pre>\n<p>This way we know what is causing the issue. If you facing memory overload, try to do profiling step by step and understand the bottleneck.</p>\n<h2>Time vs Memory Tradeoff - Save Data to Save Memory</h2>\n<p>It's always about Time vs Memory Tradeoff. Since we have 9 hours (GPU usage also) to submit our predictions, it should be plenty of time to try techniques that reduce the memory overload unless you are doing model building and inference at the same time or using the test data to retrain your models periodically during the inference process. The following tips might help you:</p>\n<ul>\n<li>Save RAM by saving that data that you don't need to complete the current step which is causing OOM issues. Ex: If merging Train and Questions DataFrames is causing the issue, then remove the other objects (Ex: You don't need to do feature engineering for the train and valid data at the same time, consider saving train and valid to memory and reloading it after the bottleneck is over). But remember that only 19.6GB of disk space is available at a time.</li>\n<li>Use Garbage Collection. If it is not removing memory from your RAM, then you need to look at why some objects are not getting removed. Python garbage collection works by removing all the objects which don't have any references. Use <code>get_referrers</code> to find all the pointers to the object and try to reduce them.</li>\n</ul>\n<pre><code>import gc\ngc.get_referrers(object)\n</code></pre>\n<ul>\n<li>Understand how Pandas works. Try to do in place operations and reduce creating copies of DataFrames when not required (Understand the concept of view vs copy in Pandas. It is confusing at times but look out for the memory usage, sometimes it helps).</li>\n<li>Use %reset to reset ram in Jupyter notebooks. This frees up RAM (if -f is used). You can save the data to a hard disk (either to /kaggle/working/ or /kaggle/temp/) and reload the data again. Thanks, <a href=\"https://www.kaggle.com/erikbruin\" target=\"_blank\">@erikbruin</a> for suggesting this.</li>\n</ul>\n<p>Have other ideas which help, please let me know. </p>",
  "messages": [
    {
      "id": "1102568",
      "postDate": "12/05/2020 04:23:35",
      "content": "<p>Hi there,</p>\n<p>Are you facing Out of Memory issues while running Kernels? Are you able to debug them? Many people have pointed out that this competition is as much as about Model Building as of Software Engineering. Here are two things I use to solve memory issues for my Kernels. </p>\n<h2>Summary Tracker</h2>\n<p>Summary Tracker helps us to find out the objects with their memory usage. We can use sorting to find the top 10 objects using the code below.</p>\n<pre><code>from operator import itemgetter\nfrom pympler import tracker\n\nmem = tracker.SummaryTracker()\nprint(sorted(mem.create_summary(), reverse=True, key=itemgetter(2))[:10])\n</code></pre>\n<p>This way we know what is causing the issue. If you facing memory overload, try to do profiling step by step and understand the bottleneck.</p>\n<h2>Time vs Memory Tradeoff - Save Data to Save Memory</h2>\n<p>It's always about Time vs Memory Tradeoff. Since we have 9 hours (GPU usage also) to submit our predictions, it should be plenty of time to try techniques that reduce the memory overload unless you are doing model building and inference at the same time or using the test data to retrain your models periodically during the inference process. The following tips might help you:</p>\n<ul>\n<li>Save RAM by saving that data that you don't need to complete the current step which is causing OOM issues. Ex: If merging Train and Questions DataFrames is causing the issue, then remove the other objects (Ex: You don't need to do feature engineering for the train and valid data at the same time, consider saving train and valid to memory and reloading it after the bottleneck is over). But remember that only 19.6GB of disk space is available at a time.</li>\n<li>Use Garbage Collection. If it is not removing memory from your RAM, then you need to look at why some objects are not getting removed. Python garbage collection works by removing all the objects which don't have any references. Use <code>get_referrers</code> to find all the pointers to the object and try to reduce them.</li>\n</ul>\n<pre><code>import gc\ngc.get_referrers(object)\n</code></pre>\n<ul>\n<li>Understand how Pandas works. Try to do in place operations and reduce creating copies of DataFrames when not required (Understand the concept of view vs copy in Pandas. It is confusing at times but look out for the memory usage, sometimes it helps).</li>\n<li>Use %reset to reset ram in Jupyter notebooks. This frees up RAM (if -f is used). You can save the data to a hard disk (either to /kaggle/working/ or /kaggle/temp/) and reload the data again. Thanks, <a href=\"https://www.kaggle.com/erikbruin\" target=\"_blank\">@erikbruin</a> for suggesting this.</li>\n</ul>\n<p>Have other ideas which help, please let me know. </p>",
      "rawMarkdown": "Hi there,\n\nAre you facing Out of Memory issues while running Kernels? Are you able to debug them? Many people have pointed out that this competition is as much as about Model Building as of Software Engineering. Here are two things I use to solve memory issues for my Kernels. \n\n## Summary Tracker\nSummary Tracker helps us to find out the objects with their memory usage. We can use sorting to find the top 10 objects using the code below.\n\n```python\nfrom operator import itemgetter\nfrom pympler import tracker\n\nmem = tracker.SummaryTracker()\nprint(sorted(mem.create_summary(), reverse=True, key=itemgetter(2))[:10])\n```\n\nThis way we know what is causing the issue. If you facing memory overload, try to do profiling step by step and understand the bottleneck.\n\n## Time vs Memory Tradeoff - Save Data to Save Memory\nIt's always about Time vs Memory Tradeoff. Since we have 9 hours (GPU usage also) to submit our predictions, it should be plenty of time to try techniques that reduce the memory overload unless you are doing model building and inference at the same time or using the test data to retrain your models periodically during the inference process. The following tips might help you:\n- Save RAM by saving that data that you don't need to complete the current step which is causing OOM issues. Ex: If merging Train and Questions DataFrames is causing the issue, then remove the other objects (Ex: You don't need to do feature engineering for the train and valid data at the same time, consider saving train and valid to memory and reloading it after the bottleneck is over). But remember that only 19.6GB of disk space is available at a time.\n- Use Garbage Collection. If it is not removing memory from your RAM, then you need to look at why some objects are not getting removed. Python garbage collection works by removing all the objects which don't have any references. Use `get_referrers` to find all the pointers to the object and try to reduce them.\n``` \nimport gc\ngc.get_referrers(object)\n```\n- Understand how Pandas works. Try to do in place operations and reduce creating copies of DataFrames when not required (Understand the concept of view vs copy in Pandas. It is confusing at times but look out for the memory usage, sometimes it helps).\n- Use %reset to reset ram in Jupyter notebooks. This frees up RAM (if -f is used). You can save the data to a hard disk (either to /kaggle/working/ or /kaggle/temp/) and reload the data again. Thanks, @erikbruin for suggesting this.\n\nHave other ideas which help, please let me know.",
      "votes": null
    },
    {
      "id": "1102607",
      "postDate": "12/05/2020 05:37:45",
      "content": "<p>Thank you for the insights!</p>",
      "rawMarkdown": "Thank you for the insights!",
      "votes": null
    },
    {
      "id": "1103267",
      "postDate": "12/05/2020 19:17:14",
      "content": "<p>Great tips, <a href=\"https://www.kaggle.com/manikanthr5\" target=\"_blank\">@manikanthr5</a>! Many thanks 😃</p>",
      "rawMarkdown": "Great tips, @manikanthr5! Many thanks 😃",
      "votes": null
    },
    {
      "id": "1106978",
      "postDate": "12/09/2020 09:11:03",
      "content": "<p>There are also Jupyter notebook specific issues. I have had cases where the object is deleted from the namespace, but memory did not seem to be released back (even after specifically calling gc.collect). These things happen when you print things to your screen (such as a simple head()). According to the documentation, we can remove those by calling %reset out, but it seems as if this does not always clear all references to the object (RAM does not go down while monitoring the RAM usage while running the notebook).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1443335%2Fe9a911587440fcddcfde467b1196e0b0%2Freset.png?generation=1607504897856996&amp;alt=media\" alt=\"\"></p>\n<p>By the way, clearing everything in RAM with %reset -f always works (I have used that in my EDA).</p>",
      "rawMarkdown": "There are also Jupyter notebook specific issues. I have had cases where the object is deleted from the namespace, but memory did not seem to be released back (even after specifically calling gc.collect). These things happen when you print things to your screen (such as a simple head()). According to the documentation, we can remove those by calling %reset out, but it seems as if this does not always clear all references to the object (RAM does not go down while monitoring the RAM usage while running the notebook).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1443335%2Fe9a911587440fcddcfde467b1196e0b0%2Freset.png?generation=1607504897856996&alt=media)\n\nBy the way, clearing everything in RAM with %reset -f always works (I have used that in my EDA).",
      "votes": null
    },
    {
      "id": "1107019",
      "postDate": "12/09/2020 09:48:03",
      "content": "<p>Yes. After looking at your EDA notebook, I started to regularly use %reset command in my notebooks. This is really useful to know. </p>",
      "rawMarkdown": "Yes. After looking at your EDA notebook, I started to regularly use %reset command in my notebooks. This is really useful to know.",
      "votes": null
    },
    {
      "id": "1112240",
      "postDate": "12/14/2020 12:23:32",
      "content": "<p>I had similar problems trying to read large parquet files for the Bengali AI competition. Here is my notebook showing techniques for memory profiling:</p>\n<p><a href=\"https://www.kaggle.com/jamesmcguigan/reading-parquet-files-ram-cpu-optimization\" target=\"_blank\">https://www.kaggle.com/jamesmcguigan/reading-parquet-files-ram-cpu-optimization</a></p>",
      "rawMarkdown": "I had similar problems trying to read large parquet files for the Bengali AI competition. Here is my notebook showing techniques for memory profiling:\n\nhttps://www.kaggle.com/jamesmcguigan/reading-parquet-files-ram-cpu-optimization",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1102607,
      "author_name": "yolandakaggle",
      "author_url": "",
      "post_date": "12/05/2020 05:37:45",
      "content": "<p>Thank you for the insights!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1103267,
      "author_name": "pedrocouto39",
      "author_url": "",
      "post_date": "12/05/2020 19:17:14",
      "content": "<p>Great tips, <a href=\"https://www.kaggle.com/manikanthr5\" target=\"_blank\">@manikanthr5</a>! Many thanks 😃</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1106978,
      "author_name": "erikbruin",
      "author_url": "",
      "post_date": "12/09/2020 09:11:03",
      "content": "<p>There are also Jupyter notebook specific issues. I have had cases where the object is deleted from the namespace, but memory did not seem to be released back (even after specifically calling gc.collect). These things happen when you print things to your screen (such as a simple head()). According to the documentation, we can remove those by calling %reset out, but it seems as if this does not always clear all references to the object (RAM does not go down while monitoring the RAM usage while running the notebook).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1443335%2Fe9a911587440fcddcfde467b1196e0b0%2Freset.png?generation=1607504897856996&amp;alt=media\" alt=\"\"></p>\n<p>By the way, clearing everything in RAM with %reset -f always works (I have used that in my EDA).</p>",
      "votes": null,
      "replies": [
        {
          "id": 1107019,
          "author_name": "manikanthr5",
          "author_url": "",
          "post_date": "12/09/2020 09:48:03",
          "content": "<p>Yes. After looking at your EDA notebook, I started to regularly use %reset command in my notebooks. This is really useful to know. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1112240,
      "author_name": "jamesmcguigan",
      "author_url": "",
      "post_date": "12/14/2020 12:23:32",
      "content": "<p>I had similar problems trying to read large parquet files for the Bengali AI competition. Here is my notebook showing techniques for memory profiling:</p>\n<p><a href=\"https://www.kaggle.com/jamesmcguigan/reading-parquet-files-ram-cpu-optimization\" target=\"_blank\">https://www.kaggle.com/jamesmcguigan/reading-parquet-files-ram-cpu-optimization</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1102568": "Hi there,\n\nAre you facing Out of Memory issues while running Kernels? Are you able to debug them? Many people have pointed out that this competition is as much as about Model Building as of Software Engineering. Here are two things I use to solve memory issues for my Kernels. \n\n## Summary Tracker\nSummary Tracker helps us to find out the objects with their memory usage. We can use sorting to find the top 10 objects using the code below.\n\n```python\nfrom operator import itemgetter\nfrom pympler import tracker\n\nmem = tracker.SummaryTracker()\nprint(sorted(mem.create_summary(), reverse=True, key=itemgetter(2))[:10])\n```\n\nThis way we know what is causing the issue. If you facing memory overload, try to do profiling step by step and understand the bottleneck.\n\n## Time vs Memory Tradeoff - Save Data to Save Memory\nIt's always about Time vs Memory Tradeoff. Since we have 9 hours (GPU usage also) to submit our predictions, it should be plenty of time to try techniques that reduce the memory overload unless you are doing model building and inference at the same time or using the test data to retrain your models periodically during the inference process. The following tips might help you:\n- Save RAM by saving that data that you don't need to complete the current step which is causing OOM issues. Ex: If merging Train and Questions DataFrames is causing the issue, then remove the other objects (Ex: You don't need to do feature engineering for the train and valid data at the same time, consider saving train and valid to memory and reloading it after the bottleneck is over). But remember that only 19.6GB of disk space is available at a time.\n- Use Garbage Collection. If it is not removing memory from your RAM, then you need to look at why some objects are not getting removed. Python garbage collection works by removing all the objects which don't have any references. Use `get_referrers` to find all the pointers to the object and try to reduce them.\n``` \nimport gc\ngc.get_referrers(object)\n```\n- Understand how Pandas works. Try to do in place operations and reduce creating copies of DataFrames when not required (Understand the concept of view vs copy in Pandas. It is confusing at times but look out for the memory usage, sometimes it helps).\n- Use %reset to reset ram in Jupyter notebooks. This frees up RAM (if -f is used). You can save the data to a hard disk (either to /kaggle/working/ or /kaggle/temp/) and reload the data again. Thanks, @erikbruin for suggesting this.\n\nHave other ideas which help, please let me know.",
    "1102607": "Thank you for the insights!",
    "1103267": "Great tips, @manikanthr5! Many thanks 😃",
    "1106978": "There are also Jupyter notebook specific issues. I have had cases where the object is deleted from the namespace, but memory did not seem to be released back (even after specifically calling gc.collect). These things happen when you print things to your screen (such as a simple head()). According to the documentation, we can remove those by calling %reset out, but it seems as if this does not always clear all references to the object (RAM does not go down while monitoring the RAM usage while running the notebook).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1443335%2Fe9a911587440fcddcfde467b1196e0b0%2Freset.png?generation=1607504897856996&alt=media)\n\nBy the way, clearing everything in RAM with %reset -f always works (I have used that in my EDA).",
    "1107019": "Yes. After looking at your EDA notebook, I started to regularly use %reset command in my notebooks. This is really useful to know.",
    "1112240": "I had similar problems trying to read large parquet files for the Bengali AI competition. Here is my notebook showing techniques for memory profiling:\n\nhttps://www.kaggle.com/jamesmcguigan/reading-parquet-files-ram-cpu-optimization"
  },
  "source": "meta"
}