{
  "id": 390741,
  "title": "RAM Utilization can't be explained",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/390741",
  "author_name": "",
  "post_date": "2023-02-27T02:54:34.726743800Z",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<p>When I start running the kernel, the RAM utilization show is 729.5MB,Then I read the train data by pd.read_csv(), and the RAM utilization is 5.6GB<br>\nand then by a reduce memory function turn the dtype to low memory use dtype,like turn int64 to int8,and the output show before modify train_data dataframe used memory is 1.96GB,after modify the train_data dataframe used memory is 0.76GB, now the RAM utilization show is 4.9GB, I don't understand what process or variable used so many RAM, and when continue run the  kernel, the RAM will crash and restart<br>\n!<br>\nwish your advice and thanks very much</p>\n<p>`</p>\n<p>def reduce_memory(data_df):</p>\n<pre><code>now_memory_used = data_df.memory_usage().sum() / 1024**3\nprint('before reduce memory modify, dataframe memory used is %.2fGB'%(now_memory_used))\nfor col in data_df.columns:\n    dtype = data_df[col].dtype.name\n    if re.search('datatime', dtype) or re.search('category', dtype):\n        continue\n    if not dtype == 'object':\n        #print(col,dtype,dtype == 'object')\n        min = data_df[col].min()\n        max = data_df[col].max()\n        if re.search('int', dtype):    \n            if min &gt;= np.iinfo(np.int8).min and max &lt;= np.iinfo(np.int8).max:\n                data_df[col] = data_df[col].astype(np.int8)\n            elif min &gt;= np.iinfo(np.int16).min and max &lt;= np.iinfo(np.int16).max:\n                data_df[col] = data_df[col].astype(np.int16)\n            elif min &gt;= np.iinfo(np.int32).min and max &lt;= np.iinfo(np.int32).max:\n                data_df[col] = data_df[col].astype(np.int32)\n            elif min &gt;= np.iinfo(np.int64).min and max &lt;= np.iinfo(np.int64).max:\n                data_df[col] = data_df[col].astype(np.int64)\n        elif re.search('float', dtype):\n            if min &gt;= np.finfo(np.float16).min and max &lt;= np.finfo(np.float16).max:\n                data_df[col] = data_df[col].astype(np.float16)\n            elif min &gt;= np.finfo(np.float32).min and max &lt;= np.finfo(np.float32).max:\n                data_df[col] = data_df[col].astype(np.float32)\n    else:\n        data_df[col] = data_df[col].astype('category')\nafter_modify_memory_used = data_df.memory_usage().sum() / 1024**3\nprint('after reduce memory modify, dataframe memory used is %.2fGB'%(after_modify_memory_used))\n\nreduced_memory = now_memory_used - after_modify_memory_used\nprint('by the reduce memroy operation, the reduced memory amount is %.2fGB'%(reduced_memory))\nreturn data_df\n</code></pre>\n<p>`</p>",
  "messages": [
    {
      "id": "2160831",
      "postDate": "02/27/2023 02:54:34",
      "content": "<p>When I start running the kernel, the RAM utilization show is 729.5MB,Then I read the train data by pd.read_csv(), and the RAM utilization is 5.6GB<br>\nand then by a reduce memory function turn the dtype to low memory use dtype,like turn int64 to int8,and the output show before modify train_data dataframe used memory is 1.96GB,after modify the train_data dataframe used memory is 0.76GB, now the RAM utilization show is 4.9GB, I don't understand what process or variable used so many RAM, and when continue run the  kernel, the RAM will crash and restart<br>\n!<br>\nwish your advice and thanks very much</p>\n<p>`</p>\n<p>def reduce_memory(data_df):</p>\n<pre><code>now_memory_used = data_df.memory_usage().sum() / 1024**3\nprint('before reduce memory modify, dataframe memory used is %.2fGB'%(now_memory_used))\nfor col in data_df.columns:\n    dtype = data_df[col].dtype.name\n    if re.search('datatime', dtype) or re.search('category', dtype):\n        continue\n    if not dtype == 'object':\n        #print(col,dtype,dtype == 'object')\n        min = data_df[col].min()\n        max = data_df[col].max()\n        if re.search('int', dtype):    \n            if min &gt;= np.iinfo(np.int8).min and max &lt;= np.iinfo(np.int8).max:\n                data_df[col] = data_df[col].astype(np.int8)\n            elif min &gt;= np.iinfo(np.int16).min and max &lt;= np.iinfo(np.int16).max:\n                data_df[col] = data_df[col].astype(np.int16)\n            elif min &gt;= np.iinfo(np.int32).min and max &lt;= np.iinfo(np.int32).max:\n                data_df[col] = data_df[col].astype(np.int32)\n            elif min &gt;= np.iinfo(np.int64).min and max &lt;= np.iinfo(np.int64).max:\n                data_df[col] = data_df[col].astype(np.int64)\n        elif re.search('float', dtype):\n            if min &gt;= np.finfo(np.float16).min and max &lt;= np.finfo(np.float16).max:\n                data_df[col] = data_df[col].astype(np.float16)\n            elif min &gt;= np.finfo(np.float32).min and max &lt;= np.finfo(np.float32).max:\n                data_df[col] = data_df[col].astype(np.float32)\n    else:\n        data_df[col] = data_df[col].astype('category')\nafter_modify_memory_used = data_df.memory_usage().sum() / 1024**3\nprint('after reduce memory modify, dataframe memory used is %.2fGB'%(after_modify_memory_used))\n\nreduced_memory = now_memory_used - after_modify_memory_used\nprint('by the reduce memroy operation, the reduced memory amount is %.2fGB'%(reduced_memory))\nreturn data_df\n</code></pre>\n<p>`</p>",
      "rawMarkdown": "When I start running the kernel, the RAM utilization show is 729.5MB,Then I read the train data by pd.read_csv(), and the RAM utilization is 5.6GB\nand then by a reduce memory function turn the dtype to low memory use dtype,like turn int64 to int8,and the output show before modify train_data dataframe used memory is 1.96GB,after modify the train_data dataframe used memory is 0.76GB, now the RAM utilization show is 4.9GB, I don't understand what process or variable used so many RAM, and when continue run the  kernel, the RAM will crash and restart\n!\nwish your advice and thanks very much\n\n`\n\ndef reduce_memory(data_df):\n\n    now_memory_used = data_df.memory_usage().sum() / 1024**3\n    print('before reduce memory modify, dataframe memory used is %.2fGB'%(now_memory_used))\n    for col in data_df.columns:\n        dtype = data_df[col].dtype.name\n        if re.search('datatime', dtype) or re.search('category', dtype):\n            continue\n        if not dtype == 'object':\n            #print(col,dtype,dtype == 'object')\n            min = data_df[col].min()\n            max = data_df[col].max()\n            if re.search('int', dtype):    \n                if min >= np.iinfo(np.int8).min and max <= np.iinfo(np.int8).max:\n                    data_df[col] = data_df[col].astype(np.int8)\n                elif min >= np.iinfo(np.int16).min and max <= np.iinfo(np.int16).max:\n                    data_df[col] = data_df[col].astype(np.int16)\n                elif min >= np.iinfo(np.int32).min and max <= np.iinfo(np.int32).max:\n                    data_df[col] = data_df[col].astype(np.int32)\n                elif min >= np.iinfo(np.int64).min and max <= np.iinfo(np.int64).max:\n                    data_df[col] = data_df[col].astype(np.int64)\n            elif re.search('float', dtype):\n                if min >= np.finfo(np.float16).min and max <= np.finfo(np.float16).max:\n                    data_df[col] = data_df[col].astype(np.float16)\n                elif min >= np.finfo(np.float32).min and max <= np.finfo(np.float32).max:\n                    data_df[col] = data_df[col].astype(np.float32)\n        else:\n            data_df[col] = data_df[col].astype('category')\n    after_modify_memory_used = data_df.memory_usage().sum() / 1024**3\n    print('after reduce memory modify, dataframe memory used is %.2fGB'%(after_modify_memory_used))\n    \n    reduced_memory = now_memory_used - after_modify_memory_used\n    print('by the reduce memroy operation, the reduced memory amount is %.2fGB'%(reduced_memory))\n    return data_df\n`",
      "votes": null
    },
    {
      "id": "2161893",
      "postDate": "02/27/2023 20:03:12",
      "content": "<p>The function <code>reduce_memory</code> uses default profiling settings:<br>\n<code>data_df.memory_usage().sum()</code></p>\n<p>This setup doesn't introspect the data deeply and hence doesn't cover full memory usage. </p>\n<p>Use this setting to see full memory usage for specific dataframe:<br>\n<code>data_df.memory_usage(index=True, deep=True).sum()</code></p>\n<p>Also note that <a href=\"https://www.kaggle.com/code/demche/student-performance-from-game-play-eda\" target=\"_blank\">three columns are missing</a> (fullscreen, hq, music) on both train and test. It's OK to drop them from the dataset and save some memory.</p>\n<p>Besides that, any further calculations consume memory, so it's better to avoid any heavy operations (as long as you're using pandas).</p>",
      "rawMarkdown": "The function `reduce_memory` uses default profiling settings:\n`data_df.memory_usage().sum()`\n\nThis setup doesn't introspect the data deeply and hence doesn't cover full memory usage. \n\nUse this setting to see full memory usage for specific dataframe:\n`data_df.memory_usage(index=True, deep=True).sum()`\n\nAlso note that [three columns are missing](https://www.kaggle.com/code/demche/student-performance-from-game-play-eda) (fullscreen, hq, music) on both train and test. It's OK to drop them from the dataset and save some memory.\n\nBesides that, any further calculations consume memory, so it's better to avoid any heavy operations (as long as you're using pandas).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2161893,
      "author_name": "demche",
      "author_url": "",
      "post_date": "02/27/2023 20:03:12",
      "content": "<p>The function <code>reduce_memory</code> uses default profiling settings:<br>\n<code>data_df.memory_usage().sum()</code></p>\n<p>This setup doesn't introspect the data deeply and hence doesn't cover full memory usage. </p>\n<p>Use this setting to see full memory usage for specific dataframe:<br>\n<code>data_df.memory_usage(index=True, deep=True).sum()</code></p>\n<p>Also note that <a href=\"https://www.kaggle.com/code/demche/student-performance-from-game-play-eda\" target=\"_blank\">three columns are missing</a> (fullscreen, hq, music) on both train and test. It's OK to drop them from the dataset and save some memory.</p>\n<p>Besides that, any further calculations consume memory, so it's better to avoid any heavy operations (as long as you're using pandas).</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2160831": "When I start running the kernel, the RAM utilization show is 729.5MB,Then I read the train data by pd.read_csv(), and the RAM utilization is 5.6GB\nand then by a reduce memory function turn the dtype to low memory use dtype,like turn int64 to int8,and the output show before modify train_data dataframe used memory is 1.96GB,after modify the train_data dataframe used memory is 0.76GB, now the RAM utilization show is 4.9GB, I don't understand what process or variable used so many RAM, and when continue run the  kernel, the RAM will crash and restart\n!\nwish your advice and thanks very much\n\n`\n\ndef reduce_memory(data_df):\n\n    now_memory_used = data_df.memory_usage().sum() / 1024**3\n    print('before reduce memory modify, dataframe memory used is %.2fGB'%(now_memory_used))\n    for col in data_df.columns:\n        dtype = data_df[col].dtype.name\n        if re.search('datatime', dtype) or re.search('category', dtype):\n            continue\n        if not dtype == 'object':\n            #print(col,dtype,dtype == 'object')\n            min = data_df[col].min()\n            max = data_df[col].max()\n            if re.search('int', dtype):    \n                if min >= np.iinfo(np.int8).min and max <= np.iinfo(np.int8).max:\n                    data_df[col] = data_df[col].astype(np.int8)\n                elif min >= np.iinfo(np.int16).min and max <= np.iinfo(np.int16).max:\n                    data_df[col] = data_df[col].astype(np.int16)\n                elif min >= np.iinfo(np.int32).min and max <= np.iinfo(np.int32).max:\n                    data_df[col] = data_df[col].astype(np.int32)\n                elif min >= np.iinfo(np.int64).min and max <= np.iinfo(np.int64).max:\n                    data_df[col] = data_df[col].astype(np.int64)\n            elif re.search('float', dtype):\n                if min >= np.finfo(np.float16).min and max <= np.finfo(np.float16).max:\n                    data_df[col] = data_df[col].astype(np.float16)\n                elif min >= np.finfo(np.float32).min and max <= np.finfo(np.float32).max:\n                    data_df[col] = data_df[col].astype(np.float32)\n        else:\n            data_df[col] = data_df[col].astype('category')\n    after_modify_memory_used = data_df.memory_usage().sum() / 1024**3\n    print('after reduce memory modify, dataframe memory used is %.2fGB'%(after_modify_memory_used))\n    \n    reduced_memory = now_memory_used - after_modify_memory_used\n    print('by the reduce memroy operation, the reduced memory amount is %.2fGB'%(reduced_memory))\n    return data_df\n`",
    "2161893": "The function `reduce_memory` uses default profiling settings:\n`data_df.memory_usage().sum()`\n\nThis setup doesn't introspect the data deeply and hence doesn't cover full memory usage. \n\nUse this setting to see full memory usage for specific dataframe:\n`data_df.memory_usage(index=True, deep=True).sum()`\n\nAlso note that [three columns are missing](https://www.kaggle.com/code/demche/student-performance-from-game-play-eda) (fullscreen, hq, music) on both train and test. It's OK to drop them from the dataset and save some memory.\n\nBesides that, any further calculations consume memory, so it's better to avoid any heavy operations (as long as you're using pandas)."
  },
  "source": "meta"
}