{
  "id": 314829,
  "title": "make the most of RAM/memory",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/314829",
  "author_name": "",
  "post_date": "2022-03-24T16:22:55.105098100Z",
  "votes": 17,
  "comment_count": 3,
  "views": 0,
  "content": "<p>In trying to avoid the dreaded <code>out of memory</code> error, it's helpful to delete any objects that are no longer needed.</p>\n<p>But I've been frustrated in situations where there are objects I don't need now, but will need later on in the notebook.</p>\n<p>For example, I'm loading the transactions csv file as a pandas/cudf dataset, preprocessing it, and then working with a subset of the data. <br>\nAt that point I no longer need the entire transactions dataset, and can delete it.<br>\nBut then later, I'll want the transactions dataset again, so that I can work with a different subset of the data.</p>\n<p>A good solution is to pickle the objects before deleting them, and then reload them as needed.</p>\n<p>I'm using the following two functions to manage the pickling/unpickling by variable string names.<br>\nHope you find them useful!</p>\n<pre><code>def pickle_variables(variable_names, delete_variables=False):\n    \"\"\"\n    saves to current directory\n    if delete_variables=True, will also delete variables\n\n    pickle variables with format &lt;variable_name&gt;.pkl\n    \"\"\"\n    for variable_name in variable_names:\n        with open(f\"{variable_name}.pkl\", \"wb\") as f:\n            pkl.dump(globals()[variable_name], f)\n        if delete_variables:\n            del globals()[variable_name]\n</code></pre>\n<pre><code>def unpickle_variables(variable_names, delete_files=False):\n    \"\"\"\n    load pickled objects from current directory\n    if delete_objs=True, will also delete pickled files\n\n    will look for pickled files with format &lt;variable_name&gt;.pkl\n    \"\"\"\n\n    for variable_name in variable_names:\n        file_name = f\"{variable_name}.pkl\"\n        with open(file_name, \"rb\") as f:\n            globals()[variable_name] = pkl.load(f)\n        if delete_files:\n            os.remove(file_name)\n</code></pre>",
  "messages": [
    {
      "id": "1733812",
      "postDate": "03/24/2022 16:22:55",
      "content": "<p>In trying to avoid the dreaded <code>out of memory</code> error, it's helpful to delete any objects that are no longer needed.</p>\n<p>But I've been frustrated in situations where there are objects I don't need now, but will need later on in the notebook.</p>\n<p>For example, I'm loading the transactions csv file as a pandas/cudf dataset, preprocessing it, and then working with a subset of the data. <br>\nAt that point I no longer need the entire transactions dataset, and can delete it.<br>\nBut then later, I'll want the transactions dataset again, so that I can work with a different subset of the data.</p>\n<p>A good solution is to pickle the objects before deleting them, and then reload them as needed.</p>\n<p>I'm using the following two functions to manage the pickling/unpickling by variable string names.<br>\nHope you find them useful!</p>\n<pre><code>def pickle_variables(variable_names, delete_variables=False):\n    \"\"\"\n    saves to current directory\n    if delete_variables=True, will also delete variables\n\n    pickle variables with format &lt;variable_name&gt;.pkl\n    \"\"\"\n    for variable_name in variable_names:\n        with open(f\"{variable_name}.pkl\", \"wb\") as f:\n            pkl.dump(globals()[variable_name], f)\n        if delete_variables:\n            del globals()[variable_name]\n</code></pre>\n<pre><code>def unpickle_variables(variable_names, delete_files=False):\n    \"\"\"\n    load pickled objects from current directory\n    if delete_objs=True, will also delete pickled files\n\n    will look for pickled files with format &lt;variable_name&gt;.pkl\n    \"\"\"\n\n    for variable_name in variable_names:\n        file_name = f\"{variable_name}.pkl\"\n        with open(file_name, \"rb\") as f:\n            globals()[variable_name] = pkl.load(f)\n        if delete_files:\n            os.remove(file_name)\n</code></pre>",
      "rawMarkdown": "In trying to avoid the dreaded `out of memory` error, it's helpful to delete any objects that are no longer needed.\n\nBut I've been frustrated in situations where there are objects I don't need now, but will need later on in the notebook.\n\nFor example, I'm loading the transactions csv file as a pandas/cudf dataset, preprocessing it, and then working with a subset of the data. \nAt that point I no longer need the entire transactions dataset, and can delete it.\nBut then later, I'll want the transactions dataset again, so that I can work with a different subset of the data.\n\nA good solution is to pickle the objects before deleting them, and then reload them as needed.\n\nI'm using the following two functions to manage the pickling/unpickling by variable string names.\nHope you find them useful!\n\n```\ndef pickle_variables(variable_names, delete_variables=False):\n    \"\"\"\n    saves to current directory\n    if delete_variables=True, will also delete variables\n\n    pickle variables with format <variable_name>.pkl\n    \"\"\"\n    for variable_name in variable_names:\n        with open(f\"{variable_name}.pkl\", \"wb\") as f:\n            pkl.dump(globals()[variable_name], f)\n        if delete_variables:\n            del globals()[variable_name]\n```\n\n```\ndef unpickle_variables(variable_names, delete_files=False):\n    \"\"\"\n    load pickled objects from current directory\n    if delete_objs=True, will also delete pickled files\n\n    will look for pickled files with format <variable_name>.pkl\n    \"\"\"\n\n    for variable_name in variable_names:\n        file_name = f\"{variable_name}.pkl\"\n        with open(file_name, \"rb\") as f:\n            globals()[variable_name] = pkl.load(f)\n        if delete_files:\n            os.remove(file_name)\n```",
      "votes": null
    },
    {
      "id": "1735506",
      "postDate": "03/26/2022 10:52:18",
      "content": "<p>Thanks for the techs👏</p>",
      "rawMarkdown": "Thanks for the techs👏",
      "votes": null
    },
    {
      "id": "1748651",
      "postDate": "04/07/2022 20:32:14",
      "content": "<p>Thanks, good technique.</p>\n<p>On a side note, I was stuck on struggling to find out why RAM was slowly filling up as I was running  optuna studies with many trials, until I realized the cell's growing text outputs were the culprit.   So now, I try to limit output/verbosity too.</p>",
      "rawMarkdown": "Thanks, good technique.\n\nOn a side note, I was stuck on struggling to find out why RAM was slowly filling up as I was running  optuna studies with many trials, until I realized the cell's growing text outputs were the culprit.   So now, I try to limit output/verbosity too.",
      "votes": null
    },
    {
      "id": "1749591",
      "postDate": "04/08/2022 18:25:18",
      "content": "<p>LGBM <code>model.predict()</code> seems to be a huge one - it converts everything to float32/64 - doesn't support any other types.<br>\nDoing prediction in batches as advised <a href=\"https://github.com/microsoft/LightGBM/issues/4033#issuecomment-787930426\" target=\"_blank\">here</a> can help.</p>",
      "rawMarkdown": "LGBM `model.predict()` seems to be a huge one - it converts everything to float32/64 - doesn't support any other types.\nDoing prediction in batches as advised [here](https://github.com/microsoft/LightGBM/issues/4033#issuecomment-787930426) can help.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1735506,
      "author_name": "alanhabrony",
      "author_url": "",
      "post_date": "03/26/2022 10:52:18",
      "content": "<p>Thanks for the techs👏</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1748651,
      "author_name": "josephramon",
      "author_url": "",
      "post_date": "04/07/2022 20:32:14",
      "content": "<p>Thanks, good technique.</p>\n<p>On a side note, I was stuck on struggling to find out why RAM was slowly filling up as I was running  optuna studies with many trials, until I realized the cell's growing text outputs were the culprit.   So now, I try to limit output/verbosity too.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1749591,
      "author_name": "jacob34",
      "author_url": "",
      "post_date": "04/08/2022 18:25:18",
      "content": "<p>LGBM <code>model.predict()</code> seems to be a huge one - it converts everything to float32/64 - doesn't support any other types.<br>\nDoing prediction in batches as advised <a href=\"https://github.com/microsoft/LightGBM/issues/4033#issuecomment-787930426\" target=\"_blank\">here</a> can help.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1733812": "In trying to avoid the dreaded `out of memory` error, it's helpful to delete any objects that are no longer needed.\n\nBut I've been frustrated in situations where there are objects I don't need now, but will need later on in the notebook.\n\nFor example, I'm loading the transactions csv file as a pandas/cudf dataset, preprocessing it, and then working with a subset of the data. \nAt that point I no longer need the entire transactions dataset, and can delete it.\nBut then later, I'll want the transactions dataset again, so that I can work with a different subset of the data.\n\nA good solution is to pickle the objects before deleting them, and then reload them as needed.\n\nI'm using the following two functions to manage the pickling/unpickling by variable string names.\nHope you find them useful!\n\n```\ndef pickle_variables(variable_names, delete_variables=False):\n    \"\"\"\n    saves to current directory\n    if delete_variables=True, will also delete variables\n\n    pickle variables with format <variable_name>.pkl\n    \"\"\"\n    for variable_name in variable_names:\n        with open(f\"{variable_name}.pkl\", \"wb\") as f:\n            pkl.dump(globals()[variable_name], f)\n        if delete_variables:\n            del globals()[variable_name]\n```\n\n```\ndef unpickle_variables(variable_names, delete_files=False):\n    \"\"\"\n    load pickled objects from current directory\n    if delete_objs=True, will also delete pickled files\n\n    will look for pickled files with format <variable_name>.pkl\n    \"\"\"\n\n    for variable_name in variable_names:\n        file_name = f\"{variable_name}.pkl\"\n        with open(file_name, \"rb\") as f:\n            globals()[variable_name] = pkl.load(f)\n        if delete_files:\n            os.remove(file_name)\n```",
    "1735506": "Thanks for the techs👏",
    "1748651": "Thanks, good technique.\n\nOn a side note, I was stuck on struggling to find out why RAM was slowly filling up as I was running  optuna studies with many trials, until I realized the cell's growing text outputs were the culprit.   So now, I try to limit output/verbosity too.",
    "1749591": "LGBM `model.predict()` seems to be a huge one - it converts everything to float32/64 - doesn't support any other types.\nDoing prediction in batches as advised [here](https://github.com/microsoft/LightGBM/issues/4033#issuecomment-787930426) can help."
  },
  "source": "meta"
}