{
  "id": 490931,
  "title": "Trick to reduce memory usage",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/490931",
  "author_name": "mavillan",
  "post_date": "2024-04-04T00:06:39.673000",
  "votes": 24,
  "comment_count": 5,
  "views": 0,
  "content": "<p>While I was running my preproc pipeline over <code>credit_bureau_a_1</code> I found a way to greatly reduce memory usage in polars, almost effortless. </p>\n<p>This was the log of my preproc in my local:</p>\n<pre><code>[info] processing credit_bureau_a_1 dataset\n[info] loading data/parquet_files/train/train_credit_bureau_a_1_0.parquet\n[info] loading data/parquet_files/train/train_credit_bureau_a_1_1.parquet\n[info] loading data/parquet_files/train/train_credit_bureau_a_1_3.parquet\n[info] loading data/parquet_files/train/train_credit_bureau_a_1_2.parquet\n[info] releasing memory\n[info] initial size: 8.55 GB\n[info] final size: 5.20 GB\n</code></pre>\n<p>And this was the same log in Kaggle:</p>\n<pre><code>[info] processing credit_bureau_a_1 dataset\n[info] loading ../input/home-credit-credit-risk-model-stability/parquet_files/train/train_credit_bureau_a_1_3.parquet\n[info] loading ../input/home-credit-credit-risk-model-stability/parquet_files/train/train_credit_bureau_a_1_2.parquet\n[info] loading ../input/home-credit-credit-risk-model-stability/parquet_files/train/train_credit_bureau_a_1_0.parquet\n[info] loading ../input/home-credit-credit-risk-model-stability/parquet_files/train/train_credit_bureau_a_1_1.parquet\n[info] releasing memory\n[info] initial size: 12.24 GB\n[info] final size: 8.88 GB\n</code></pre>\n<p>There's a huge difference in memory usage (before and after memory reduction). The difference: Kaggle uses <strong>polars-0.20.10</strong> while I was using <strong>polars-0.20.18</strong> in my local. So, just I upgraded polars in Kaggle and got exactly the same memory stats. </p>\n<p>Given that we are not allowed to install packages from internet in submission notebooks, here is how I did it:</p>\n<ol>\n<li>Add this dataset to your notebook: <a href=\"https://www.kaggle.com/datasets/mavillan/polars-0-20-18\" target=\"_blank\">https://www.kaggle.com/datasets/mavillan/polars-0-20-18</a> (are just the polars artifacts from pypi)</li>\n<li>Install it by placing this code at the top of your notebook:</li>\n</ol>\n<pre><code>!pip install /kaggle/input/polars-0-20-18/polars-0.20.18-cp38-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl\n</code></pre>\n<p>Hope this helps you deal with the huge <code>credit_bureau_a_1</code>, it worked for me :-)</p>",
  "messages": [
    {
      "id": 2734121,
      "postDate": "2024-04-04T00:06:39.673Z",
      "content": "<p>While I was running my preproc pipeline over <code>credit_bureau_a_1</code> I found a way to greatly reduce memory usage in polars, almost effortless. </p>\n<p>This was the log of my preproc in my local:</p>\n<pre><code>[info] processing credit_bureau_a_1 dataset\n[info] loading data/parquet_files/train/train_credit_bureau_a_1_0.parquet\n[info] loading data/parquet_files/train/train_credit_bureau_a_1_1.parquet\n[info] loading data/parquet_files/train/train_credit_bureau_a_1_3.parquet\n[info] loading data/parquet_files/train/train_credit_bureau_a_1_2.parquet\n[info] releasing memory\n[info] initial size: 8.55 GB\n[info] final size: 5.20 GB\n</code></pre>\n<p>And this was the same log in Kaggle:</p>\n<pre><code>[info] processing credit_bureau_a_1 dataset\n[info] loading ../input/home-credit-credit-risk-model-stability/parquet_files/train/train_credit_bureau_a_1_3.parquet\n[info] loading ../input/home-credit-credit-risk-model-stability/parquet_files/train/train_credit_bureau_a_1_2.parquet\n[info] loading ../input/home-credit-credit-risk-model-stability/parquet_files/train/train_credit_bureau_a_1_0.parquet\n[info] loading ../input/home-credit-credit-risk-model-stability/parquet_files/train/train_credit_bureau_a_1_1.parquet\n[info] releasing memory\n[info] initial size: 12.24 GB\n[info] final size: 8.88 GB\n</code></pre>\n<p>There's a huge difference in memory usage (before and after memory reduction). The difference: Kaggle uses <strong>polars-0.20.10</strong> while I was using <strong>polars-0.20.18</strong> in my local. So, just I upgraded polars in Kaggle and got exactly the same memory stats. </p>\n<p>Given that we are not allowed to install packages from internet in submission notebooks, here is how I did it:</p>\n<ol>\n<li>Add this dataset to your notebook: <a href=\"https://www.kaggle.com/datasets/mavillan/polars-0-20-18\" target=\"_blank\">https://www.kaggle.com/datasets/mavillan/polars-0-20-18</a> (are just the polars artifacts from pypi)</li>\n<li>Install it by placing this code at the top of your notebook:</li>\n</ol>\n<pre><code>!pip install /kaggle/input/polars-0-20-18/polars-0.20.18-cp38-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl\n</code></pre>\n<p>Hope this helps you deal with the huge <code>credit_bureau_a_1</code>, it worked for me :-)</p>",
      "rawMarkdown": "While I was running my preproc pipeline over `credit_bureau_a_1` I found a way to greatly reduce memory usage in polars, almost effortless. \n\nThis was the log of my preproc in my local:\n```bash\n[info] processing credit_bureau_a_1 dataset\n[info] loading data/parquet_files/train/train_credit_bureau_a_1_0.parquet\n[info] loading data/parquet_files/train/train_credit_bureau_a_1_1.parquet\n[info] loading data/parquet_files/train/train_credit_bureau_a_1_3.parquet\n[info] loading data/parquet_files/train/train_credit_bureau_a_1_2.parquet\n[info] releasing memory\n[info] initial size: 8.55 GB\n[info] final size: 5.20 GB\n```\n\nAnd this was the same log in Kaggle:\n```bash\n[info] processing credit_bureau_a_1 dataset\n[info] loading ../input/home-credit-credit-risk-model-stability/parquet_files/train/train_credit_bureau_a_1_3.parquet\n[info] loading ../input/home-credit-credit-risk-model-stability/parquet_files/train/train_credit_bureau_a_1_2.parquet\n[info] loading ../input/home-credit-credit-risk-model-stability/parquet_files/train/train_credit_bureau_a_1_0.parquet\n[info] loading ../input/home-credit-credit-risk-model-stability/parquet_files/train/train_credit_bureau_a_1_1.parquet\n[info] releasing memory\n[info] initial size: 12.24 GB\n[info] final size: 8.88 GB\n```\n\nThere's a huge difference in memory usage (before and after memory reduction). The difference: Kaggle uses **polars-0.20.10** while I was using **polars-0.20.18** in my local. So, just I upgraded polars in Kaggle and got exactly the same memory stats. \n\nGiven that we are not allowed to install packages from internet in submission notebooks, here is how I did it:\n1. Add this dataset to your notebook: https://www.kaggle.com/datasets/mavillan/polars-0-20-18 (are just the polars artifacts from pypi)\n2. Install it by placing this code at the top of your notebook:\n```bash\n!pip install /kaggle/input/polars-0-20-18/polars-0.20.18-cp38-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl\n```\n\nHope this helps you deal with the huge `credit_bureau_a_1`, it worked for me :-)",
      "votes": 24
    },
    {
      "id": 2741053,
      "postDate": "2024-04-08T05:32:10.440Z",
      "content": "<p>Same memory used after &amp; before polars lib updated version !! </p>\n<p>Please send some referance notebook or any code to test , how did you arrived this!!!  <a href=\"https://www.kaggle.com/mavillan\" target=\"_blank\">@mavillan</a> </p>",
      "rawMarkdown": "Same memory used after & before polars lib updated version !! \n\nPlease send some referance notebook or any code to test , how did you arrived this!!!  @mavillan \n\n"
    },
    {
      "id": 2736627,
      "postDate": "2024-04-05T09:57:01.510Z",
      "content": "<p>Convert your object dtype to category and profit from a great size decrease</p>",
      "rawMarkdown": "Convert your object dtype to category and profit from a great size decrease"
    },
    {
      "id": 2734193,
      "postDate": "2024-04-04T00:52:40.517Z",
      "content": "<p>perfect! something like this is what i need, thanks</p>",
      "rawMarkdown": "perfect! something like this is what i need, thanks"
    },
    {
      "id": 2737091,
      "postDate": "2024-04-05T15:19:55.610Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2814673,
      "postDate": "2024-05-15T13:38:16.810Z",
      "content": "<p>thank u for sharing!!!</p>",
      "rawMarkdown": "thank u for sharing!!!"
    }
  ],
  "comments": [
    {
      "id": 2741053,
      "author_name": "Anil Kumar Reddy",
      "author_url": "",
      "post_date": "2024-04-08T05:32:10.440000",
      "content": "<p>Same memory used after &amp; before polars lib updated version !! </p>\n<p>Please send some referance notebook or any code to test , how did you arrived this!!!  <a href=\"https://www.kaggle.com/mavillan\" target=\"_blank\">@mavillan</a> </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2736627,
      "author_name": "Danu A.",
      "author_url": "",
      "post_date": "2024-04-05T09:57:01.510000",
      "content": "<p>Convert your object dtype to category and profit from a great size decrease</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2734193,
      "author_name": "Shreyas Bhatt",
      "author_url": "",
      "post_date": "2024-04-04T00:52:40.517000",
      "content": "<p>perfect! something like this is what i need, thanks</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2737091,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-05T15:19:55.610000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2814673,
      "author_name": "Heather233",
      "author_url": "",
      "post_date": "2024-05-15T13:38:16.810000",
      "content": "<p>thank u for sharing!!!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2734121": "While I was running my preproc pipeline over `credit_bureau_a_1` I found a way to greatly reduce memory usage in polars, almost effortless. \n\nThis was the log of my preproc in my local:\n```bash\n[info] processing credit_bureau_a_1 dataset\n[info] loading data/parquet_files/train/train_credit_bureau_a_1_0.parquet\n[info] loading data/parquet_files/train/train_credit_bureau_a_1_1.parquet\n[info] loading data/parquet_files/train/train_credit_bureau_a_1_3.parquet\n[info] loading data/parquet_files/train/train_credit_bureau_a_1_2.parquet\n[info] releasing memory\n[info] initial size: 8.55 GB\n[info] final size: 5.20 GB\n```\n\nAnd this was the same log in Kaggle:\n```bash\n[info] processing credit_bureau_a_1 dataset\n[info] loading ../input/home-credit-credit-risk-model-stability/parquet_files/train/train_credit_bureau_a_1_3.parquet\n[info] loading ../input/home-credit-credit-risk-model-stability/parquet_files/train/train_credit_bureau_a_1_2.parquet\n[info] loading ../input/home-credit-credit-risk-model-stability/parquet_files/train/train_credit_bureau_a_1_0.parquet\n[info] loading ../input/home-credit-credit-risk-model-stability/parquet_files/train/train_credit_bureau_a_1_1.parquet\n[info] releasing memory\n[info] initial size: 12.24 GB\n[info] final size: 8.88 GB\n```\n\nThere's a huge difference in memory usage (before and after memory reduction). The difference: Kaggle uses **polars-0.20.10** while I was using **polars-0.20.18** in my local. So, just I upgraded polars in Kaggle and got exactly the same memory stats. \n\nGiven that we are not allowed to install packages from internet in submission notebooks, here is how I did it:\n1. Add this dataset to your notebook: https://www.kaggle.com/datasets/mavillan/polars-0-20-18 (are just the polars artifacts from pypi)\n2. Install it by placing this code at the top of your notebook:\n```bash\n!pip install /kaggle/input/polars-0-20-18/polars-0.20.18-cp38-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl\n```\n\nHope this helps you deal with the huge `credit_bureau_a_1`, it worked for me :-)",
    "2741053": "Same memory used after & before polars lib updated version !! \n\nPlease send some referance notebook or any code to test , how did you arrived this!!!  @mavillan \n\n",
    "2736627": "Convert your object dtype to category and profit from a great size decrease",
    "2734193": "perfect! something like this is what i need, thanks",
    "2737091": "",
    "2814673": "thank u for sharing!!!"
  }
}