{
  "id": 327142,
  "title": "Understanding the Input data using Pandas",
  "url": "/competitions/amex-default-prediction/discussion/327142",
  "author_name": "",
  "post_date": "2022-05-25T19:46:59.448215400Z",
  "votes": 22,
  "comment_count": 1,
  "views": 0,
  "content": "<p>For this competition, the memory of the given dataset is enormous, <strong>train_data - 16.39 GB, test_data - 33.82 GB</strong>, and loading them directly into Kaggle Notebook, will throw a memory error.</p>\n<p>So, I tried the following methods to understand the dataset, by loading them in small chunks.</p>\n<ol>\n<li><p>First step, I imported the train data using <strong>pd.read_csv() with nrows option</strong><br>\n<code>df = pd.read_csv('../input/amex-default-prediction/train_data.csv', nrows=1000)</code></p></li>\n<li><p>With this small data frame, I found out the total no. of columns and data type of each column<br>\nSo, totally we have <strong>190 columns - float64 (185 cols), int64(1 cols), object(4 cols)</strong></p></li>\n<li><p>Using variable with datatype <strong>int64</strong>, I tried to import train and test data using below command to get total rows in train and test data<br>\n<code>test_data = pd.read_csv('../input/amex-default-prediction/test_data.csv',usecols=['B_31'])</code></p></li>\n<li><p>Shape of the datasets - <br>\n  <strong>In train dataset - 5531451 rows, 190 columns</strong><br>\n  <strong>In test dataset - 11363762 rows, 190 columns</strong></p></li>\n<li><p>We have 185 columns with float64 data type, which can be reduced to float32 or float16, to reduce memory size. To check this, I created a user function <code>get_min_max(df)</code> to get the min and max values and store them in a dictionary</p></li>\n<li><p>To get the min and max values, I imported data in chunks and kept updating the min-max values for each float column, as shown:</p></li>\n</ol>\n<pre><code>chunksize = 1000000 #rows subset from large dataframe\nfor chunk in pd.read_csv('../input/amex-default-prediction/train_data.csv',usecols=float_cols, chunksize=chunksize):\n    get_min_max(chunk)\n</code></pre>\n<p>With the extracted min-max values for all 185 columns, it is evident that the values are with very high precision. <strong>Lowest value - 1.463473564555784e-11 &amp; highest value - 5755.075986356201</strong></p>\n<p>Based on the above observation, I changed the data type from <strong>float64 to float16</strong>. By doing this, I will lose information in the data, as the precision is going to be reduced, but still, this is worth the try. With this, I was able to reduce memory to <strong>2.1 GB (train data)</strong></p>\n<p>Attached is <a href=\"https://www.kaggle.com/code/balabaskar/memory-reduction-using-pandas\" target=\"_blank\">my notebook</a> here for reference.</p>\n<p>What do you think about this approach? Can this be done in an efficient way? Please suggest some improvements.</p>",
  "messages": [
    {
      "id": "1801489",
      "postDate": "05/25/2022 19:46:59",
      "content": "<p>For this competition, the memory of the given dataset is enormous, <strong>train_data - 16.39 GB, test_data - 33.82 GB</strong>, and loading them directly into Kaggle Notebook, will throw a memory error.</p>\n<p>So, I tried the following methods to understand the dataset, by loading them in small chunks.</p>\n<ol>\n<li><p>First step, I imported the train data using <strong>pd.read_csv() with nrows option</strong><br>\n<code>df = pd.read_csv('../input/amex-default-prediction/train_data.csv', nrows=1000)</code></p></li>\n<li><p>With this small data frame, I found out the total no. of columns and data type of each column<br>\nSo, totally we have <strong>190 columns - float64 (185 cols), int64(1 cols), object(4 cols)</strong></p></li>\n<li><p>Using variable with datatype <strong>int64</strong>, I tried to import train and test data using below command to get total rows in train and test data<br>\n<code>test_data = pd.read_csv('../input/amex-default-prediction/test_data.csv',usecols=['B_31'])</code></p></li>\n<li><p>Shape of the datasets - <br>\n  <strong>In train dataset - 5531451 rows, 190 columns</strong><br>\n  <strong>In test dataset - 11363762 rows, 190 columns</strong></p></li>\n<li><p>We have 185 columns with float64 data type, which can be reduced to float32 or float16, to reduce memory size. To check this, I created a user function <code>get_min_max(df)</code> to get the min and max values and store them in a dictionary</p></li>\n<li><p>To get the min and max values, I imported data in chunks and kept updating the min-max values for each float column, as shown:</p></li>\n</ol>\n<pre><code>chunksize = 1000000 #rows subset from large dataframe\nfor chunk in pd.read_csv('../input/amex-default-prediction/train_data.csv',usecols=float_cols, chunksize=chunksize):\n    get_min_max(chunk)\n</code></pre>\n<p>With the extracted min-max values for all 185 columns, it is evident that the values are with very high precision. <strong>Lowest value - 1.463473564555784e-11 &amp; highest value - 5755.075986356201</strong></p>\n<p>Based on the above observation, I changed the data type from <strong>float64 to float16</strong>. By doing this, I will lose information in the data, as the precision is going to be reduced, but still, this is worth the try. With this, I was able to reduce memory to <strong>2.1 GB (train data)</strong></p>\n<p>Attached is <a href=\"https://www.kaggle.com/code/balabaskar/memory-reduction-using-pandas\" target=\"_blank\">my notebook</a> here for reference.</p>\n<p>What do you think about this approach? Can this be done in an efficient way? Please suggest some improvements.</p>",
      "rawMarkdown": "For this competition, the memory of the given dataset is enormous, **train_data - 16.39 GB, test_data - 33.82 GB**, and loading them directly into Kaggle Notebook, will throw a memory error.\n\nSo, I tried the following methods to understand the dataset, by loading them in small chunks.\n\n1. First step, I imported the train data using **pd.read_csv() with nrows option**\n```df = pd.read_csv('../input/amex-default-prediction/train_data.csv', nrows=1000)```\n2. With this small data frame, I found out the total no. of columns and data type of each column\n    So, totally we have **190 columns - float64 (185 cols), int64(1 cols), object(4 cols)**\n3. Using variable with datatype **int64**, I tried to import train and test data using below command to get total rows in train and test data\n```test_data = pd.read_csv('../input/amex-default-prediction/test_data.csv',usecols=['B_31'])```\n4. Shape of the datasets - \n      **In train dataset - 5531451 rows, 190 columns**\n      **In test dataset - 11363762 rows, 190 columns**\n\n5. We have 185 columns with float64 data type, which can be reduced to float32 or float16, to reduce memory size. To check this, I created a user function ```get_min_max(df)``` to get the min and max values and store them in a dictionary\n6. To get the min and max values, I imported data in chunks and kept updating the min-max values for each float column, as shown:\n```\nchunksize = 1000000 #rows subset from large dataframe\nfor chunk in pd.read_csv('../input/amex-default-prediction/train_data.csv',usecols=float_cols, chunksize=chunksize):\n    get_min_max(chunk)\n```\n\nWith the extracted min-max values for all 185 columns, it is evident that the values are with very high precision. **Lowest value - 1.463473564555784e-11 & highest value - 5755.075986356201**\n\nBased on the above observation, I changed the data type from **float64 to float16**. By doing this, I will lose information in the data, as the precision is going to be reduced, but still, this is worth the try. With this, I was able to reduce memory to **2.1 GB (train data)**\n\nAttached is [my notebook](https://www.kaggle.com/code/balabaskar/memory-reduction-using-pandas) here for reference.\n\nWhat do you think about this approach? Can this be done in an efficient way? Please suggest some improvements.",
      "votes": null
    },
    {
      "id": "1817982",
      "postDate": "06/12/2022 05:05:54",
      "content": "<p>I will approach this by performing the analysis on subset of the data and for training use incremental learning library like 'vaex'.</p>",
      "rawMarkdown": "I will approach this by performing the analysis on subset of the data and for training use incremental learning library like 'vaex'.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1817982,
      "author_name": "ashwinshetgaonkar",
      "author_url": "",
      "post_date": "06/12/2022 05:05:54",
      "content": "<p>I will approach this by performing the analysis on subset of the data and for training use incremental learning library like 'vaex'.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1801489": "For this competition, the memory of the given dataset is enormous, **train_data - 16.39 GB, test_data - 33.82 GB**, and loading them directly into Kaggle Notebook, will throw a memory error.\n\nSo, I tried the following methods to understand the dataset, by loading them in small chunks.\n\n1. First step, I imported the train data using **pd.read_csv() with nrows option**\n```df = pd.read_csv('../input/amex-default-prediction/train_data.csv', nrows=1000)```\n2. With this small data frame, I found out the total no. of columns and data type of each column\n    So, totally we have **190 columns - float64 (185 cols), int64(1 cols), object(4 cols)**\n3. Using variable with datatype **int64**, I tried to import train and test data using below command to get total rows in train and test data\n```test_data = pd.read_csv('../input/amex-default-prediction/test_data.csv',usecols=['B_31'])```\n4. Shape of the datasets - \n      **In train dataset - 5531451 rows, 190 columns**\n      **In test dataset - 11363762 rows, 190 columns**\n\n5. We have 185 columns with float64 data type, which can be reduced to float32 or float16, to reduce memory size. To check this, I created a user function ```get_min_max(df)``` to get the min and max values and store them in a dictionary\n6. To get the min and max values, I imported data in chunks and kept updating the min-max values for each float column, as shown:\n```\nchunksize = 1000000 #rows subset from large dataframe\nfor chunk in pd.read_csv('../input/amex-default-prediction/train_data.csv',usecols=float_cols, chunksize=chunksize):\n    get_min_max(chunk)\n```\n\nWith the extracted min-max values for all 185 columns, it is evident that the values are with very high precision. **Lowest value - 1.463473564555784e-11 & highest value - 5755.075986356201**\n\nBased on the above observation, I changed the data type from **float64 to float16**. By doing this, I will lose information in the data, as the precision is going to be reduced, but still, this is worth the try. With this, I was able to reduce memory to **2.1 GB (train data)**\n\nAttached is [my notebook](https://www.kaggle.com/code/balabaskar/memory-reduction-using-pandas) here for reference.\n\nWhat do you think about this approach? Can this be done in an efficient way? Please suggest some improvements.",
    "1817982": "I will approach this by performing the analysis on subset of the data and for training use incremental learning library like 'vaex'."
  },
  "source": "meta"
}