{
  "id": 327400,
  "title": "⚡[FAST LOADING] only 1.4GB Training Data using Feather",
  "url": "/competitions/amex-default-prediction/discussion/327400",
  "author_name": "",
  "post_date": "2022-05-27T04:25:32.461833300Z",
  "votes": 33,
  "comment_count": 1,
  "views": 0,
  "content": "<p>As you might have noticed, the tabular files for this competition are very HUGE, load these huge csv files may task too much time and RAM.</p>\n<p>Inspired by <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327143\" target=\"_blank\">9x Data Compression achieved with Feather</a>, we also optimized data storage format using <a href=\"https://arrow.apache.org/docs/python/feather.html\" target=\"_blank\">feather</a> and some preprocessing was done to achieve higher compression rate (9x -&gt; <strong>more than 11x</strong>).</p>\n<p>💳<strong>Compressed Dataset</strong><br>\nTo help you read the data faster, I've created feather versions of the csv files and create a dataset. <a href=\"https://www.kaggle.com/datasets/seefun/AMEX-Default-Prediction-Feather\" target=\"_blank\"><strong>DATASET HERE</strong></a></p>\n<table>\n<thead>\n<tr>\n<th>Data</th>\n<th>Size of csv file</th>\n<th>Size of feather file(Ruchi)</th>\n<th><strong>Size of feather file(Ours)</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>train</td>\n<td>16.4 GB</td>\n<td>1.8GB</td>\n<td><strong>1.4 GB</strong></td>\n</tr>\n<tr>\n<td>test</td>\n<td>33.8 GB</td>\n<td>3.6GB</td>\n<td><strong>2.9 GB</strong></td>\n</tr>\n</tbody>\n</table>\n<p>⚙️<strong>Process</strong><br>\nChanged the datatypes of the columns as follows:</p>\n<p>Numerical columns -&gt; <code>float16</code><br>\nCategorical columns -&gt; preprocess to <code>int8</code> (NAN-&gt;-1)</p>\n<p>zstd compression is also added</p>\n<p>you can also find the preprocess script <a href=\"https://www.kaggle.com/datasets/seefun/amex-default-prediction-feather\" target=\"_blank\">our dataset</a>.</p>\n<p>⚡<strong>Load</strong><br>\n<code>pd.read_feather('path/to/train.feather')</code></p>",
  "messages": [
    {
      "id": "1802711",
      "postDate": "05/27/2022 04:25:32",
      "content": "<p>As you might have noticed, the tabular files for this competition are very HUGE, load these huge csv files may task too much time and RAM.</p>\n<p>Inspired by <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327143\" target=\"_blank\">9x Data Compression achieved with Feather</a>, we also optimized data storage format using <a href=\"https://arrow.apache.org/docs/python/feather.html\" target=\"_blank\">feather</a> and some preprocessing was done to achieve higher compression rate (9x -&gt; <strong>more than 11x</strong>).</p>\n<p>💳<strong>Compressed Dataset</strong><br>\nTo help you read the data faster, I've created feather versions of the csv files and create a dataset. <a href=\"https://www.kaggle.com/datasets/seefun/AMEX-Default-Prediction-Feather\" target=\"_blank\"><strong>DATASET HERE</strong></a></p>\n<table>\n<thead>\n<tr>\n<th>Data</th>\n<th>Size of csv file</th>\n<th>Size of feather file(Ruchi)</th>\n<th><strong>Size of feather file(Ours)</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>train</td>\n<td>16.4 GB</td>\n<td>1.8GB</td>\n<td><strong>1.4 GB</strong></td>\n</tr>\n<tr>\n<td>test</td>\n<td>33.8 GB</td>\n<td>3.6GB</td>\n<td><strong>2.9 GB</strong></td>\n</tr>\n</tbody>\n</table>\n<p>⚙️<strong>Process</strong><br>\nChanged the datatypes of the columns as follows:</p>\n<p>Numerical columns -&gt; <code>float16</code><br>\nCategorical columns -&gt; preprocess to <code>int8</code> (NAN-&gt;-1)</p>\n<p>zstd compression is also added</p>\n<p>you can also find the preprocess script <a href=\"https://www.kaggle.com/datasets/seefun/amex-default-prediction-feather\" target=\"_blank\">our dataset</a>.</p>\n<p>⚡<strong>Load</strong><br>\n<code>pd.read_feather('path/to/train.feather')</code></p>",
      "rawMarkdown": "As you might have noticed, the tabular files for this competition are very HUGE, load these huge csv files may task too much time and RAM.\n\nInspired by [9x Data Compression achieved with Feather](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327143), we also optimized data storage format using [feather](https://arrow.apache.org/docs/python/feather.html) and some preprocessing was done to achieve higher compression rate (9x -> **more than 11x**).\n\n 💳**Compressed Dataset**\nTo help you read the data faster, I've created feather versions of the csv files and create a dataset. [**DATASET HERE**](https://www.kaggle.com/datasets/seefun/AMEX-Default-Prediction-Feather)\n\n|  Data  | Size of csv file | Size of feather file(Ruchi) | **Size of feather file(Ours)** |\n|  ----  | ---- | ---- | ---- |\n| train  | 16.4 GB | 1.8GB | **1.4 GB** |\n| test  | 33.8 GB | 3.6GB | **2.9 GB** |\n\n\n⚙️**Process**\nChanged the datatypes of the columns as follows:\n\nNumerical columns -> `float16`\nCategorical columns -> preprocess to `int8` (NAN->-1)\n\nzstd compression is also added\n\nyou can also find the preprocess script [our dataset](https://www.kaggle.com/datasets/seefun/amex-default-prediction-feather).\n\n⚡**Load**\n`pd.read_feather('path/to/train.feather')`",
      "votes": null
    },
    {
      "id": "1855384",
      "postDate": "07/14/2022 14:34:32",
      "content": "<p>can you please share the scripts for converting csv to feather ?</p>",
      "rawMarkdown": "can you please share the scripts for converting csv to feather ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1855384,
      "author_name": "ptrikp",
      "author_url": "",
      "post_date": "07/14/2022 14:34:32",
      "content": "<p>can you please share the scripts for converting csv to feather ?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1802711": "As you might have noticed, the tabular files for this competition are very HUGE, load these huge csv files may task too much time and RAM.\n\nInspired by [9x Data Compression achieved with Feather](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327143), we also optimized data storage format using [feather](https://arrow.apache.org/docs/python/feather.html) and some preprocessing was done to achieve higher compression rate (9x -> **more than 11x**).\n\n 💳**Compressed Dataset**\nTo help you read the data faster, I've created feather versions of the csv files and create a dataset. [**DATASET HERE**](https://www.kaggle.com/datasets/seefun/AMEX-Default-Prediction-Feather)\n\n|  Data  | Size of csv file | Size of feather file(Ruchi) | **Size of feather file(Ours)** |\n|  ----  | ---- | ---- | ---- |\n| train  | 16.4 GB | 1.8GB | **1.4 GB** |\n| test  | 33.8 GB | 3.6GB | **2.9 GB** |\n\n\n⚙️**Process**\nChanged the datatypes of the columns as follows:\n\nNumerical columns -> `float16`\nCategorical columns -> preprocess to `int8` (NAN->-1)\n\nzstd compression is also added\n\nyou can also find the preprocess script [our dataset](https://www.kaggle.com/datasets/seefun/amex-default-prediction-feather).\n\n⚡**Load**\n`pd.read_feather('path/to/train.feather')`",
    "1855384": "can you please share the scripts for converting csv to feather ?"
  },
  "source": "meta"
}