{
  "id": 371058,
  "title": "Saving GPU memory when processing features == (Speedup + Efficiency)",
  "url": "/competitions/otto-recommender-system/discussion/371058",
  "author_name": "",
  "post_date": "2022-12-07T19:16:58.394766400Z",
  "votes": 42,
  "comment_count": 4,
  "views": 0,
  "content": "<p>A new functionality of Cudf (since 22.08) is setting the default data type to be 32 or 64 bit. <br>\nUsing 32 bit can help save GPU memory when processing features.<br>\nThe code bellow sets 32 bit as the default:</p>\n<pre><code>import cudf\ncudf.set_option(\"default_integer_bitwidth\", 32)\ncudf.set_option(\"default_float_bitwidth\", 32)\n</code></pre>\n<p>Another possibility is converting all the DataFrame from 64bit to 32bit on the fly and using Garbage Collector at the end:</p>\n<pre><code>import gc\ndef freemem(df):\n    for col in df.columns:\n        if df[col].dtype == 'int64':\n            df[col] = df[col].astype('int32')\n        elif df[col].dtype == 'float64':\n            df[col] = df[col].astype('float32')\n    gc.collect()\n    return\n\nfreemem(mydataframe)\n</code></pre>",
  "messages": [
    {
      "id": "2058288",
      "postDate": "12/07/2022 19:16:58",
      "content": "<p>A new functionality of Cudf (since 22.08) is setting the default data type to be 32 or 64 bit. <br>\nUsing 32 bit can help save GPU memory when processing features.<br>\nThe code bellow sets 32 bit as the default:</p>\n<pre><code>import cudf\ncudf.set_option(\"default_integer_bitwidth\", 32)\ncudf.set_option(\"default_float_bitwidth\", 32)\n</code></pre>\n<p>Another possibility is converting all the DataFrame from 64bit to 32bit on the fly and using Garbage Collector at the end:</p>\n<pre><code>import gc\ndef freemem(df):\n    for col in df.columns:\n        if df[col].dtype == 'int64':\n            df[col] = df[col].astype('int32')\n        elif df[col].dtype == 'float64':\n            df[col] = df[col].astype('float32')\n    gc.collect()\n    return\n\nfreemem(mydataframe)\n</code></pre>",
      "rawMarkdown": "A new functionality of Cudf (since 22.08) is setting the default data type to be 32 or 64 bit. \nUsing 32 bit can help save GPU memory when processing features.\nThe code bellow sets 32 bit as the default:\n\n```\nimport cudf\ncudf.set_option(\"default_integer_bitwidth\", 32)\ncudf.set_option(\"default_float_bitwidth\", 32)\n```\n\n\nAnother possibility is converting all the DataFrame from 64bit to 32bit on the fly and using Garbage Collector at the end:\n\n```\nimport gc\ndef freemem(df):\n    for col in df.columns:\n        if df[col].dtype == 'int64':\n            df[col] = df[col].astype('int32')\n        elif df[col].dtype == 'float64':\n            df[col] = df[col].astype('float32')\n    gc.collect()\n    return\n\nfreemem(mydataframe)\n```",
      "votes": null
    },
    {
      "id": "2058295",
      "postDate": "12/07/2022 19:24:43",
      "content": "<p>Great suggestion. Memory management, accelerating code, and efficient disk usage is very important in this competition.</p>",
      "rawMarkdown": "Great suggestion. Memory management, accelerating code, and efficient disk usage is very important in this competition.",
      "votes": null
    },
    {
      "id": "2059026",
      "postDate": "12/08/2022 12:09:59",
      "content": "<p>I can't directly install cudf&gt;=22.08 by pip. It seems that they didn't update cudf on pip for a long time.</p>",
      "rawMarkdown": "I can't directly install cudf>=22.08 by pip. It seems that they didn't update cudf on pip for a long time.",
      "votes": null
    },
    {
      "id": "2059262",
      "postDate": "12/08/2022 16:22:13",
      "content": "<p>nice suggestion. Thanks for sharing </p>",
      "rawMarkdown": "nice suggestion. Thanks for sharing",
      "votes": null
    },
    {
      "id": "2090699",
      "postDate": "01/07/2023 15:45:44",
      "content": "<p>Thank you for this excelent hint. Reducing memory is always important, but it is especially beneficial for this competition.</p>",
      "rawMarkdown": "Thank you for this excelent hint. Reducing memory is always important, but it is especially beneficial for this competition.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2058295,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "12/07/2022 19:24:43",
      "content": "<p>Great suggestion. Memory management, accelerating code, and efficient disk usage is very important in this competition.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2059026,
      "author_name": "dylanliuofficial",
      "author_url": "",
      "post_date": "12/08/2022 12:09:59",
      "content": "<p>I can't directly install cudf&gt;=22.08 by pip. It seems that they didn't update cudf on pip for a long time.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2059262,
      "author_name": "pvtrmalli",
      "author_url": "",
      "post_date": "12/08/2022 16:22:13",
      "content": "<p>nice suggestion. Thanks for sharing </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2090699,
      "author_name": "gpreda",
      "author_url": "",
      "post_date": "01/07/2023 15:45:44",
      "content": "<p>Thank you for this excelent hint. Reducing memory is always important, but it is especially beneficial for this competition.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2058288": "A new functionality of Cudf (since 22.08) is setting the default data type to be 32 or 64 bit. \nUsing 32 bit can help save GPU memory when processing features.\nThe code bellow sets 32 bit as the default:\n\n```\nimport cudf\ncudf.set_option(\"default_integer_bitwidth\", 32)\ncudf.set_option(\"default_float_bitwidth\", 32)\n```\n\n\nAnother possibility is converting all the DataFrame from 64bit to 32bit on the fly and using Garbage Collector at the end:\n\n```\nimport gc\ndef freemem(df):\n    for col in df.columns:\n        if df[col].dtype == 'int64':\n            df[col] = df[col].astype('int32')\n        elif df[col].dtype == 'float64':\n            df[col] = df[col].astype('float32')\n    gc.collect()\n    return\n\nfreemem(mydataframe)\n```",
    "2058295": "Great suggestion. Memory management, accelerating code, and efficient disk usage is very important in this competition.",
    "2059026": "I can't directly install cudf>=22.08 by pip. It seems that they didn't update cudf on pip for a long time.",
    "2059262": "nice suggestion. Thanks for sharing",
    "2090699": "Thank you for this excelent hint. Reducing memory is always important, but it is especially beneficial for this competition."
  },
  "source": "meta"
}