{
  "id": 316193,
  "title": "RAM doesn't add up?",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/316193",
  "author_name": "",
  "post_date": "2022-03-31T18:21:28.448618300Z",
  "votes": 7,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I'm using <code>sys.getsizeof()</code> to keep track of how individual DataFrames are using memory.<br>\nBut I can't get it to add up!</p>\n<p>For a simple example, I read the <code>transactions_train.csv</code> file with:</p>\n<p><code>t = cudf.read_csv(\"../input/h-and-m-personalized-fashion-recommendations/transactions_train.csv\")</code></p>\n<p>The, using <code>sys.getsizeof()</code> gives 3.37 GB.</p>\n<p>But looking at the kaggle system before vs. after shows an increase of usage of 4.1 GB of GPU RAM, and 1.5 GB of CPU RAM. (See image below)</p>\n<p>This is a simple example, but I'm finding that it's often the case that things don't add up, and it's making it very hard to manage memory properly.</p>\n<p>Any help would be greatly appreciated!!</p>\n<p><a href=\"https://postimg.cc/BLPd7BQp\" target=\"_blank\"><img src=\"https://i.postimg.cc/C1cS1c33/memory-before-and-after.jpg\" alt=\"memory-before-and-after.jpg\"></a></p>",
  "messages": [
    {
      "id": "1741342",
      "postDate": "03/31/2022 18:21:28",
      "content": "<p>I'm using <code>sys.getsizeof()</code> to keep track of how individual DataFrames are using memory.<br>\nBut I can't get it to add up!</p>\n<p>For a simple example, I read the <code>transactions_train.csv</code> file with:</p>\n<p><code>t = cudf.read_csv(\"../input/h-and-m-personalized-fashion-recommendations/transactions_train.csv\")</code></p>\n<p>The, using <code>sys.getsizeof()</code> gives 3.37 GB.</p>\n<p>But looking at the kaggle system before vs. after shows an increase of usage of 4.1 GB of GPU RAM, and 1.5 GB of CPU RAM. (See image below)</p>\n<p>This is a simple example, but I'm finding that it's often the case that things don't add up, and it's making it very hard to manage memory properly.</p>\n<p>Any help would be greatly appreciated!!</p>\n<p><a href=\"https://postimg.cc/BLPd7BQp\" target=\"_blank\"><img src=\"https://i.postimg.cc/C1cS1c33/memory-before-and-after.jpg\" alt=\"memory-before-and-after.jpg\"></a></p>",
      "rawMarkdown": "I'm using `sys.getsizeof()` to keep track of how individual DataFrames are using memory.\nBut I can't get it to add up!\n\nFor a simple example, I read the `transactions_train.csv` file with:\n\n`t = cudf.read_csv(\"../input/h-and-m-personalized-fashion-recommendations/transactions_train.csv\")`\n\nThe, using `sys.getsizeof()` gives 3.37 GB.\n\nBut looking at the kaggle system before vs. after shows an increase of usage of 4.1 GB of GPU RAM, and 1.5 GB of CPU RAM. (See image below)\n\nThis is a simple example, but I'm finding that it's often the case that things don't add up, and it's making it very hard to manage memory properly.\n\nAny help would be greatly appreciated!!\n\n[![memory-before-and-after.jpg](https://i.postimg.cc/C1cS1c33/memory-before-and-after.jpg)](https://postimg.cc/BLPd7BQp)",
      "votes": null
    },
    {
      "id": "1741424",
      "postDate": "03/31/2022 19:50:33",
      "content": "<p>Its not easy to handle memory in script languages.</p>\n<p>I recommend to save memory converting your DataFrame variables to <code>int32</code> or <code>float32</code>. And when possible (low cardinality, binary features) use int16 and even int8 to save even more. <br>\neg.  <code>df['feature'] = df['feature'].astype('int32')</code></p>\n<p>Also call garbage collector often <code>gc.collect()</code> </p>",
      "rawMarkdown": "Its not easy to handle memory in script languages.\n\nI recommend to save memory converting your DataFrame variables to `int32` or `float32`. And when possible (low cardinality, binary features) use int16 and even int8 to save even more. \neg.  `df['feature'] = df['feature'].astype('int32')`\n\nAlso call garbage collector often `gc.collect()`",
      "votes": null
    },
    {
      "id": "1742118",
      "postDate": "04/01/2022 13:26:48",
      "content": "<p>Thank you, <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a>! <br>\nI'm familiar with your dtype conversion trick from cdeotte's post.<br>\ngc.collect() seems to be helping in many cases.</p>",
      "rawMarkdown": "Thank you, @titericz! \nI'm familiar with your dtype conversion trick from cdeotte's post.\ngc.collect() seems to be helping in many cases.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1741424,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "03/31/2022 19:50:33",
      "content": "<p>Its not easy to handle memory in script languages.</p>\n<p>I recommend to save memory converting your DataFrame variables to <code>int32</code> or <code>float32</code>. And when possible (low cardinality, binary features) use int16 and even int8 to save even more. <br>\neg.  <code>df['feature'] = df['feature'].astype('int32')</code></p>\n<p>Also call garbage collector often <code>gc.collect()</code> </p>",
      "votes": null,
      "replies": [
        {
          "id": 1742118,
          "author_name": "jacob34",
          "author_url": "",
          "post_date": "04/01/2022 13:26:48",
          "content": "<p>Thank you, <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a>! <br>\nI'm familiar with your dtype conversion trick from cdeotte's post.<br>\ngc.collect() seems to be helping in many cases.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1741342": "I'm using `sys.getsizeof()` to keep track of how individual DataFrames are using memory.\nBut I can't get it to add up!\n\nFor a simple example, I read the `transactions_train.csv` file with:\n\n`t = cudf.read_csv(\"../input/h-and-m-personalized-fashion-recommendations/transactions_train.csv\")`\n\nThe, using `sys.getsizeof()` gives 3.37 GB.\n\nBut looking at the kaggle system before vs. after shows an increase of usage of 4.1 GB of GPU RAM, and 1.5 GB of CPU RAM. (See image below)\n\nThis is a simple example, but I'm finding that it's often the case that things don't add up, and it's making it very hard to manage memory properly.\n\nAny help would be greatly appreciated!!\n\n[![memory-before-and-after.jpg](https://i.postimg.cc/C1cS1c33/memory-before-and-after.jpg)](https://postimg.cc/BLPd7BQp)",
    "1741424": "Its not easy to handle memory in script languages.\n\nI recommend to save memory converting your DataFrame variables to `int32` or `float32`. And when possible (low cardinality, binary features) use int16 and even int8 to save even more. \neg.  `df['feature'] = df['feature'].astype('int32')`\n\nAlso call garbage collector often `gc.collect()`",
    "1742118": "Thank you, @titericz! \nI'm familiar with your dtype conversion trick from cdeotte's post.\ngc.collect() seems to be helping in many cases."
  },
  "source": "meta"
}