{
  "id": 327333,
  "title": "How to reduce pandas memory while loading dataframe ?",
  "url": "/competitions/amex-default-prediction/discussion/327333",
  "author_name": "",
  "post_date": "2022-05-26T17:26:03.915743800Z",
  "votes": 17,
  "comment_count": 3,
  "views": 0,
  "content": "<blockquote>\n  <h2>How to reduce pandas memory while loading dataframe :</h2>\n  <p>1.Dropping columns<br>\n  2.Lower-range numerical dtypes<br>\n  3.Categoricals<br>\n  4.Sparse columns <br>\n  5.Reading in chunks</p>\n<pre><code>import pandas as pd\n</code></pre>\n  <p><strong>Technique 1 : Don’t load all the columns</strong></p>\n<pre><code>df = pd.read_csv(\"bigcsvfile.csv\", usecols=[\"col1\", \"col2\"])\n</code></pre>\n  <p><strong>Technique 2: Shrink numerical columns with smaller dtypes</strong></p>\n<pre><code>int8 can store integers from -128 to 127.\nint16 can store integers from -32768 to 32767.\nint64 can store integers from -9223372036854775808 to 9223372036854775807.\n</code></pre>\n  <p>When Pandas loads a CSV, it guesses the dtypes. If a column  is integer, by default it assigns that column int64 as the dtype.</p>\n  <p><em>In case if you know that the numbers in a particular column will never be higher than 32767, you can use an int16 and reduce the memory usage</em></p>\n<pre><code>df = pd.read_csv(\"bigcsvfile.csv\", dtype={\"numericalcolumn\": \"int8\"})\n</code></pre>\n  <p><strong>Technique 3: Shrink categorical data using Categorical dtypes</strong><br>\n  By default column for qualitative value parsed as a string.<br>\n  However a more compact representation for data with only a limited number of values is a custom dtype called Categorical, whose memory usage is tied to the number of different values.</p>\n<pre><code>df = pd.read_csv( \"bigcsvfile.csv\", dtype={\"categoricalcolumn\" : \"category\"})\n</code></pre>\n  <p><strong>Technique 4: Sparse series</strong></p>\n  <p>If you have a column with lots of empty values, usually represented as NaNs, you can save memory by using a sparse column representation.It won’t waste memory storing all those empty values.</p>\n<pre><code>df = pd.read_csv(\"bigcsvfile.csv\")\nseries = df[\"columnwithNA\"]\nsparse_series = series.astype(\"Sparse[str]\")\n</code></pre>\n  <p><strong>Technique 5: Reading in chunks</strong></p>\n<pre><code>chunksize = 10 ** 6\nwith pd.read_csv(\"bigcsvfile.csv\", chunksize=chunksize) as reader:\n   for chunk in reader:\n       process(chunk)\n</code></pre>\n</blockquote>",
  "messages": [
    {
      "id": "1802403",
      "postDate": "05/26/2022 17:26:03",
      "content": "<blockquote>\n  <h2>How to reduce pandas memory while loading dataframe :</h2>\n  <p>1.Dropping columns<br>\n  2.Lower-range numerical dtypes<br>\n  3.Categoricals<br>\n  4.Sparse columns <br>\n  5.Reading in chunks</p>\n<pre><code>import pandas as pd\n</code></pre>\n  <p><strong>Technique 1 : Don’t load all the columns</strong></p>\n<pre><code>df = pd.read_csv(\"bigcsvfile.csv\", usecols=[\"col1\", \"col2\"])\n</code></pre>\n  <p><strong>Technique 2: Shrink numerical columns with smaller dtypes</strong></p>\n<pre><code>int8 can store integers from -128 to 127.\nint16 can store integers from -32768 to 32767.\nint64 can store integers from -9223372036854775808 to 9223372036854775807.\n</code></pre>\n  <p>When Pandas loads a CSV, it guesses the dtypes. If a column  is integer, by default it assigns that column int64 as the dtype.</p>\n  <p><em>In case if you know that the numbers in a particular column will never be higher than 32767, you can use an int16 and reduce the memory usage</em></p>\n<pre><code>df = pd.read_csv(\"bigcsvfile.csv\", dtype={\"numericalcolumn\": \"int8\"})\n</code></pre>\n  <p><strong>Technique 3: Shrink categorical data using Categorical dtypes</strong><br>\n  By default column for qualitative value parsed as a string.<br>\n  However a more compact representation for data with only a limited number of values is a custom dtype called Categorical, whose memory usage is tied to the number of different values.</p>\n<pre><code>df = pd.read_csv( \"bigcsvfile.csv\", dtype={\"categoricalcolumn\" : \"category\"})\n</code></pre>\n  <p><strong>Technique 4: Sparse series</strong></p>\n  <p>If you have a column with lots of empty values, usually represented as NaNs, you can save memory by using a sparse column representation.It won’t waste memory storing all those empty values.</p>\n<pre><code>df = pd.read_csv(\"bigcsvfile.csv\")\nseries = df[\"columnwithNA\"]\nsparse_series = series.astype(\"Sparse[str]\")\n</code></pre>\n  <p><strong>Technique 5: Reading in chunks</strong></p>\n<pre><code>chunksize = 10 ** 6\nwith pd.read_csv(\"bigcsvfile.csv\", chunksize=chunksize) as reader:\n   for chunk in reader:\n       process(chunk)\n</code></pre>\n</blockquote>",
      "rawMarkdown": "> ## How to reduce pandas memory while loading dataframe :\n> 1.Dropping columns\n> 2.Lower-range numerical dtypes\n> 3.Categoricals\n> 4.Sparse columns \n> 5.Reading in chunks\n\n> ```\n>import pandas as pd\n>```\n> **Technique 1 : Don’t load all the columns**\n> \n> ```\n> df = pd.read_csv(\"bigcsvfile.csv\", usecols=[\"col1\", \"col2\"])\n> ```\n> \n> **Technique 2: Shrink numerical columns with smaller dtypes**\n> \n> ```\n> int8 can store integers from -128 to 127.\n> int16 can store integers from -32768 to 32767.\n> int64 can store integers from -9223372036854775808 to 9223372036854775807.\n> ```\n> \n> \n> When Pandas loads a CSV, it guesses the dtypes. If a column  is integer, by default it assigns that column int64 as the dtype.\n> \n> *In case if you know that the numbers in a particular column will never be higher than 32767, you can use an int16 and reduce the memory usage*\n> \n> ```\n> df = pd.read_csv(\"bigcsvfile.csv\", dtype={\"numericalcolumn\": \"int8\"})\n> ```\n> \n> **Technique 3: Shrink categorical data using Categorical dtypes**\n> By default column for qualitative value parsed as a string.\n> However a more compact representation for data with only a limited number of values is a custom dtype called Categorical, whose memory usage is tied to the number of different values.\n> \n> ```\n> df = pd.read_csv( \"bigcsvfile.csv\", dtype={\"categoricalcolumn\" : \"category\"})\n> ```\n> \n> \n> **Technique 4: Sparse series**\n> \n> If you have a column with lots of empty values, usually represented as NaNs, you can save memory by using a sparse column representation.It won’t waste memory storing all those empty values.\n> \n> ```\n> df = pd.read_csv(\"bigcsvfile.csv\")\n> series = df[\"columnwithNA\"]\n> sparse_series = series.astype(\"Sparse[str]\")\n> ```\n>**Technique 5: Reading in chunks**\n>```\n>chunksize = 10 ** 6\n>with pd.read_csv(\"bigcsvfile.csv\", chunksize=chunksize) as reader:\n>    for chunk in reader:\n>        process(chunk)\n>```",
      "votes": null
    },
    {
      "id": "2044758",
      "postDate": "11/26/2022 18:45:45",
      "content": "<p>Thanks for this, I started with Technique 5, but the others also seem promising. Do you have any other options in mind recently? </p>",
      "rawMarkdown": "Thanks for this, I started with Technique 5, but the others also seem promising. Do you have any other options in mind recently?",
      "votes": null
    },
    {
      "id": "2044761",
      "postDate": "11/26/2022 18:47:06",
      "content": "<p>FYR: <a href=\"https://www.kaggle.com/code/shravankumar147/1-data-exploration-amex-default-prediction\" target=\"_blank\">My Notebook</a></p>",
      "rawMarkdown": "FYR: [My Notebook](https://www.kaggle.com/code/shravankumar147/1-data-exploration-amex-default-prediction)",
      "votes": null
    },
    {
      "id": "2045343",
      "postDate": "11/27/2022 09:20:46",
      "content": "<p>very good techniques <a href=\"https://www.kaggle.com/vidyasagarbhargava\" target=\"_blank\">@vidyasagarbhargava</a> . i will definitely remember these techniques the next time i work with large datasets. thanks for sharing. keep it up. simple and objective.</p>\n<p>my newer notebooks handle this type of optimization in another way: sampling. five notebooks with different sampling techniques that help me to assemble a representative subset of data, optimizing the execution time of my exploratory analyzes and algorithms. if you could take a look and see if it makes sense to you, i would be very grateful.</p>",
      "rawMarkdown": "very good techniques @vidyasagarbhargava . i will definitely remember these techniques the next time i work with large datasets. thanks for sharing. keep it up. simple and objective.\n\nmy newer notebooks handle this type of optimization in another way: sampling. five notebooks with different sampling techniques that help me to assemble a representative subset of data, optimizing the execution time of my exploratory analyzes and algorithms. if you could take a look and see if it makes sense to you, i would be very grateful.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2044758,
      "author_name": "shravankumar147",
      "author_url": "",
      "post_date": "11/26/2022 18:45:45",
      "content": "<p>Thanks for this, I started with Technique 5, but the others also seem promising. Do you have any other options in mind recently? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2044761,
      "author_name": "shravankumar147",
      "author_url": "",
      "post_date": "11/26/2022 18:47:06",
      "content": "<p>FYR: <a href=\"https://www.kaggle.com/code/shravankumar147/1-data-exploration-amex-default-prediction\" target=\"_blank\">My Notebook</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2045343,
      "author_name": "jardelnascimento",
      "author_url": "",
      "post_date": "11/27/2022 09:20:46",
      "content": "<p>very good techniques <a href=\"https://www.kaggle.com/vidyasagarbhargava\" target=\"_blank\">@vidyasagarbhargava</a> . i will definitely remember these techniques the next time i work with large datasets. thanks for sharing. keep it up. simple and objective.</p>\n<p>my newer notebooks handle this type of optimization in another way: sampling. five notebooks with different sampling techniques that help me to assemble a representative subset of data, optimizing the execution time of my exploratory analyzes and algorithms. if you could take a look and see if it makes sense to you, i would be very grateful.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1802403": "> ## How to reduce pandas memory while loading dataframe :\n> 1.Dropping columns\n> 2.Lower-range numerical dtypes\n> 3.Categoricals\n> 4.Sparse columns \n> 5.Reading in chunks\n\n> ```\n>import pandas as pd\n>```\n> **Technique 1 : Don’t load all the columns**\n> \n> ```\n> df = pd.read_csv(\"bigcsvfile.csv\", usecols=[\"col1\", \"col2\"])\n> ```\n> \n> **Technique 2: Shrink numerical columns with smaller dtypes**\n> \n> ```\n> int8 can store integers from -128 to 127.\n> int16 can store integers from -32768 to 32767.\n> int64 can store integers from -9223372036854775808 to 9223372036854775807.\n> ```\n> \n> \n> When Pandas loads a CSV, it guesses the dtypes. If a column  is integer, by default it assigns that column int64 as the dtype.\n> \n> *In case if you know that the numbers in a particular column will never be higher than 32767, you can use an int16 and reduce the memory usage*\n> \n> ```\n> df = pd.read_csv(\"bigcsvfile.csv\", dtype={\"numericalcolumn\": \"int8\"})\n> ```\n> \n> **Technique 3: Shrink categorical data using Categorical dtypes**\n> By default column for qualitative value parsed as a string.\n> However a more compact representation for data with only a limited number of values is a custom dtype called Categorical, whose memory usage is tied to the number of different values.\n> \n> ```\n> df = pd.read_csv( \"bigcsvfile.csv\", dtype={\"categoricalcolumn\" : \"category\"})\n> ```\n> \n> \n> **Technique 4: Sparse series**\n> \n> If you have a column with lots of empty values, usually represented as NaNs, you can save memory by using a sparse column representation.It won’t waste memory storing all those empty values.\n> \n> ```\n> df = pd.read_csv(\"bigcsvfile.csv\")\n> series = df[\"columnwithNA\"]\n> sparse_series = series.astype(\"Sparse[str]\")\n> ```\n>**Technique 5: Reading in chunks**\n>```\n>chunksize = 10 ** 6\n>with pd.read_csv(\"bigcsvfile.csv\", chunksize=chunksize) as reader:\n>    for chunk in reader:\n>        process(chunk)\n>```",
    "2044758": "Thanks for this, I started with Technique 5, but the others also seem promising. Do you have any other options in mind recently?",
    "2044761": "FYR: [My Notebook](https://www.kaggle.com/code/shravankumar147/1-data-exploration-amex-default-prediction)",
    "2045343": "very good techniques @vidyasagarbhargava . i will definitely remember these techniques the next time i work with large datasets. thanks for sharing. keep it up. simple and objective.\n\nmy newer notebooks handle this type of optimization in another way: sampling. five notebooks with different sampling techniques that help me to assemble a representative subset of data, optimizing the execution time of my exploratory analyzes and algorithms. if you could take a look and see if it makes sense to you, i would be very grateful."
  },
  "source": "meta"
}