{
  "id": 327143,
  "title": "⚡ 9x Data Compression achieved with Feather🕊️",
  "url": "/competitions/amex-default-prediction/discussion/327143",
  "author_name": "Ruchi Bhatia",
  "post_date": "2022-05-25T19:57:39.245000",
  "votes": 157,
  "comment_count": 36,
  "views": 0,
  "content": "<p>As you might have noticed, the tabular files for this competition are HUGE 🤯</p>\n<h4>💡What are the Feather and Parquet file formats?</h4>\n<ul>\n<li>These file formats are <strong>designed to handle complex data in bulk</strong>.</li>\n<li>They are preferred due to their <strong>high efficiency</strong> in data compression and decompression.</li>\n</ul>\n<h4>💳 Compressed Dataset</h4>\n<p>To help you read the data faster, I've created <a href=\"https://www.kaggle.com/datasets/ruchi798/parquet-files-amexdefault-prediction\" target=\"_blank\">feather and parquet versions</a> of the csv files and curated a dataset.</p>\n<table>\n<thead>\n<tr>\n<th>Data</th>\n<th>Size of csv file</th>\n<th>Size of parquet file</th>\n<th><strong>Size of feather file</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>train_data</td>\n<td>16.4 GB</td>\n<td>6.7 GB</td>\n<td><strong>1.8 GB</strong></td>\n</tr>\n<tr>\n<td>test_data</td>\n<td>33.8 GB</td>\n<td>13.7 GB</td>\n<td><strong>3.6 GB</strong></td>\n</tr>\n</tbody>\n</table>\n<h4>⚙️ Process</h4>\n<p>Changed the datatypes of the columns as follows:</p>\n<ul>\n<li>Numerical columns -&gt; <strong><code>float16</code></strong> and</li>\n<li>Categorical columns -&gt; <strong><code>category</code></strong></li>\n</ul>",
  "messages": [
    {
      "id": 1801493,
      "postDate": "2022-05-25T19:57:39.247Z",
      "content": "<p>As you might have noticed, the tabular files for this competition are HUGE 🤯</p>\n<h4>💡What are the Feather and Parquet file formats?</h4>\n<ul>\n<li>These file formats are <strong>designed to handle complex data in bulk</strong>.</li>\n<li>They are preferred due to their <strong>high efficiency</strong> in data compression and decompression.</li>\n</ul>\n<h4>💳 Compressed Dataset</h4>\n<p>To help you read the data faster, I've created <a href=\"https://www.kaggle.com/datasets/ruchi798/parquet-files-amexdefault-prediction\" target=\"_blank\">feather and parquet versions</a> of the csv files and curated a dataset.</p>\n<table>\n<thead>\n<tr>\n<th>Data</th>\n<th>Size of csv file</th>\n<th>Size of parquet file</th>\n<th><strong>Size of feather file</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>train_data</td>\n<td>16.4 GB</td>\n<td>6.7 GB</td>\n<td><strong>1.8 GB</strong></td>\n</tr>\n<tr>\n<td>test_data</td>\n<td>33.8 GB</td>\n<td>13.7 GB</td>\n<td><strong>3.6 GB</strong></td>\n</tr>\n</tbody>\n</table>\n<h4>⚙️ Process</h4>\n<p>Changed the datatypes of the columns as follows:</p>\n<ul>\n<li>Numerical columns -&gt; <strong><code>float16</code></strong> and</li>\n<li>Categorical columns -&gt; <strong><code>category</code></strong></li>\n</ul>",
      "rawMarkdown": "As you might have noticed, the tabular files for this competition are HUGE 🤯\n\n#### 💡What are the Feather and Parquet file formats?\n- These file formats are **designed to handle complex data in bulk**.\n- They are preferred due to their **high efficiency** in data compression and decompression.\n\n#### 💳 Compressed Dataset\nTo help you read the data faster, I've created [feather and parquet versions](https://www.kaggle.com/datasets/ruchi798/parquet-files-amexdefault-prediction) of the csv files and curated a dataset.\n\n| Data       | Size of csv file | Size of parquet file  | **Size of feather file**  |\n|------------|----------|---------------|---------------|\n| train_data | 16.4 GB  | 6.7 GB        | **1.8 GB**        |\n| test_data  | 33.8 GB  | 13.7 GB       | **3.6 GB**        |\n\n#### ⚙️ Process\nChanged the datatypes of the columns as follows:\n- Numerical columns -> **`float16`** and\n- Categorical columns -> **`category`**",
      "votes": 155
    },
    {
      "id": 1801806,
      "postDate": "2022-05-26T06:57:46.837Z",
      "content": "<p>How did you achieve such compression? When I save it <code>df.to_parquet</code> or <code>df.to_feather</code>, it turns out to be about the same size.</p>",
      "rawMarkdown": "How did you achieve such compression? When I save it `df.to_parquet` or `df.to_feather`, it turns out to be about the same size.",
      "votes": 3,
      "replies": [
        {
          "id": 1801809,
          "postDate": "2022-05-26T07:00:23.503Z",
          "content": "<p>Hello,<br>\nI converted the datatypes of the columns:</p>\n<ul>\n<li>Numerical columns to <code>float16</code> and </li>\n<li>Categorical columns to <code>category</code></li>\n</ul>\n<p>This really helped me to <em>reduce the size of the dataset immensely</em>.<br>\nSince we have so many rows in both the train and test dataset, making even a <em>tiny</em> transformation like downcasting can play a huge role!</p>\n<p>We can <strong>process our data faster</strong> since the dataset takes significantly <strong>less memory</strong> ✨</p>",
          "rawMarkdown": "Hello,\nI converted the datatypes of the columns:\n- Numerical columns to `float16` and \n- Categorical columns to `category`\n\nThis really helped me to *reduce the size of the dataset immensely*.\nSince we have so many rows in both the train and test dataset, making even a *tiny* transformation like downcasting can play a huge role!\n\nWe can **process our data faster** since the dataset takes significantly **less memory** ✨",
          "votes": 17
        }
      ]
    },
    {
      "id": 1919361,
      "postDate": "2022-08-30T10:37:59.453Z",
      "content": "<p>Useful content <a href=\"https://www.kaggle.com/ruchi798\" target=\"_blank\">@ruchi798</a> </p>",
      "rawMarkdown": "Useful content @ruchi798 ",
      "votes": 1
    },
    {
      "id": 1855965,
      "postDate": "2022-07-15T03:28:35.617Z",
      "content": "<p>Thanks for sharing.you explain it in a very simple way. </p>",
      "rawMarkdown": "Thanks for sharing.you explain it in a very simple way. ",
      "votes": 1
    },
    {
      "id": 1826691,
      "postDate": "2022-06-20T14:37:22.687Z",
      "content": "<p>Very helpful</p>",
      "rawMarkdown": "Very helpful",
      "votes": 1
    },
    {
      "id": 1826648,
      "postDate": "2022-06-20T14:12:40.757Z",
      "content": "<p>Thaks for sharing <a href=\"https://www.kaggle.com/ruchi798\" target=\"_blank\">@ruchi798</a> </p>",
      "rawMarkdown": "Thaks for sharing @ruchi798 ",
      "votes": 1
    },
    {
      "id": 1826591,
      "postDate": "2022-06-20T13:21:34.193Z",
      "content": "<p>Thanks for the impressive work. it's really needed for this huge type of dataset.</p>",
      "rawMarkdown": "Thanks for the impressive work. it's really needed for this huge type of dataset.",
      "votes": 1
    },
    {
      "id": 1824582,
      "postDate": "2022-06-18T13:11:17.607Z",
      "content": "<p>Thanks for sharing this! Wow the compression is really quite impressive! I had no idea feather even existed and looks like it's worth exploring and learning.</p>",
      "rawMarkdown": "Thanks for sharing this! Wow the compression is really quite impressive! I had no idea feather even existed and looks like it's worth exploring and learning.",
      "votes": 1
    },
    {
      "id": 1816777,
      "postDate": "2022-06-10T14:54:43.487Z",
      "content": "<p>this is great!</p>",
      "rawMarkdown": "this is great!",
      "votes": 1
    },
    {
      "id": 1809046,
      "postDate": "2022-06-02T10:47:08.513Z",
      "content": "<p>Well done, thank you for this solution, <a href=\"https://www.kaggle.com/ruchi798\" target=\"_blank\">@ruchi798</a> 🤜<br>\nHappy Kaggling!</p>",
      "rawMarkdown": "Well done, thank you for this solution, @ruchi798 🤜\nHappy Kaggling!",
      "votes": 1
    },
    {
      "id": 1808912,
      "postDate": "2022-06-02T08:36:49.503Z",
      "content": "<p>Again job well done! <a href=\"https://www.kaggle.com/ruchi798\" target=\"_blank\">@ruchi798</a> 👍</p>",
      "rawMarkdown": "Again job well done! @ruchi798 👍",
      "votes": 1
    },
    {
      "id": 1808362,
      "postDate": "2022-06-01T18:27:54.087Z",
      "content": "<p>Great help <a href=\"https://www.kaggle.com/ruchi798\" target=\"_blank\">@ruchi798</a> </p>",
      "rawMarkdown": "Great help @ruchi798 ",
      "votes": 1
    },
    {
      "id": 1807663,
      "postDate": "2022-06-01T08:48:45.200Z",
      "content": "<p>For folks looking on how to convert to feather a good reference notebook : <a href=\"https://www.kaggle.com/code/kmader/convert-to-feather-for-use-in-other-kernels/notebook\" target=\"_blank\">https://www.kaggle.com/code/kmader/convert-to-feather-for-use-in-other-kernels/notebook</a></p>",
      "rawMarkdown": "For folks looking on how to convert to feather a good reference notebook : https://www.kaggle.com/code/kmader/convert-to-feather-for-use-in-other-kernels/notebook",
      "votes": 1
    },
    {
      "id": 1806119,
      "postDate": "2022-05-30T19:56:14.073Z",
      "content": "<p>Feather does really a good job at compression it seems. Eager to learn more about it.</p>",
      "rawMarkdown": "Feather does really a good job at compression it seems. Eager to learn more about it.",
      "votes": 1
    },
    {
      "id": 1805206,
      "postDate": "2022-05-29T23:20:25.277Z",
      "content": "<p>Hey! Feather seems really interesting! Any opensource projects when it has been widely used?</p>",
      "rawMarkdown": "Hey! Feather seems really interesting! Any opensource projects when it has been widely used?",
      "votes": 1
    },
    {
      "id": 1804365,
      "postDate": "2022-05-28T20:14:24.557Z",
      "content": "<p>Thank you for sharing this dataset with us. I'm already using it for my notebook. Honestly, I've never heard before about the feather format.</p>\n<p>Great job!</p>",
      "rawMarkdown": "Thank you for sharing this dataset with us. I'm already using it for my notebook. Honestly, I've never heard before about the feather format.\n\nGreat job!",
      "votes": 1
    },
    {
      "id": 1803911,
      "postDate": "2022-05-28T09:46:05.040Z",
      "content": "<p><strong>This is the ultimate compression</strong>!. Interesting how changing the data types makes a huge difference! Thanks again for sharing!  </p>",
      "rawMarkdown": "**This is the ultimate compression**!. Interesting how changing the data types makes a huge difference! Thanks again for sharing!  ",
      "votes": 1
    },
    {
      "id": 1803782,
      "postDate": "2022-05-28T07:21:00.543Z",
      "content": "<p>It's really helpful for smooth runtime, thanks for sharing 👍</p>",
      "rawMarkdown": "It's really helpful for smooth runtime, thanks for sharing 👍",
      "votes": 1
    },
    {
      "id": 1803705,
      "postDate": "2022-05-28T04:50:24.980Z",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/rohitgarud\" target=\"_blank\">@rohitgarud</a> and <a href=\"https://www.kaggle.com/cv13j0\" target=\"_blank\">@cv13j0</a>,<br>\nTo answer your questions, </p>\n<ul>\n<li><strong>float16</strong> has reduced precision compared to float64 but is <strong>much faster</strong> than the latter.</li>\n<li><strong>float64</strong> consumes <strong>much more memory</strong> than float16 due to its higher precision digits, resulting in the large dataset size we have.</li>\n</ul>",
      "rawMarkdown": "Hello @rohitgarud and @cv13j0,\nTo answer your questions, \n- **float16** has reduced precision compared to float64 but is **much faster** than the latter.\n- **float64** consumes **much more memory** than float16 due to its higher precision digits, resulting in the large dataset size we have.",
      "votes": 1,
      "replies": [
        {
          "id": 1808909,
          "postDate": "2022-06-02T08:36:12.310Z",
          "content": "<p><a href=\"https://www.kaggle.com/ruchi798\" target=\"_blank\">@ruchi798</a> thank you for the help.</p>",
          "rawMarkdown": "@ruchi798 thank you for the help."
        }
      ]
    },
    {
      "id": 1803480,
      "postDate": "2022-05-27T21:02:01.983Z",
      "content": "<p>Thanks for the dataset, <a href=\"https://www.kaggle.com/ruchi798\" target=\"_blank\">@ruchi798</a>; I will start using them today. Do you anticipate any precision loss or impacts to the ML model by decreasing the numeric columns to fload16?</p>",
      "rawMarkdown": "Thanks for the dataset, @ruchi798; I will start using them today. Do you anticipate any precision loss or impacts to the ML model by decreasing the numeric columns to fload16?",
      "votes": 1,
      "replies": [
        {
          "id": 1806114,
          "postDate": "2022-05-30T19:52:30.960Z",
          "content": "<p>I would expect the model trained with float16 instead of float64 to be on average ~0.0000003814 more \"off\", which by most means would be negligible. However, if you want to test it train two models, one of float64 data and one of float16 data and see if there's a difference in scores.</p>",
          "rawMarkdown": "I would expect the model trained with float16 instead of float64 to be on average ~0.0000003814 more \"off\", which by most means would be negligible. However, if you want to test it train two models, one of float64 data and one of float16 data and see if there's a difference in scores."
        }
      ]
    },
    {
      "id": 1802199,
      "postDate": "2022-05-26T14:42:16.997Z",
      "content": "<p>Great job! Thanks very much for sharing these two methods 😊 I have never used the feather file format before. Do you have any links that you could share about it? I have always just assumed that the parquet format most the most efficient.</p>",
      "rawMarkdown": "Great job! Thanks very much for sharing these two methods 😊 I have never used the feather file format before. Do you have any links that you could share about it? I have always just assumed that the parquet format most the most efficient.",
      "votes": 1
    },
    {
      "id": 1808771,
      "postDate": "2022-06-02T06:30:55.303Z",
      "content": "<p>if you consider using feather/parquet data format, you could also consider cleaned up dataset which has some columns as integers:</p>\n<p><a href=\"https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\" target=\"_blank\">https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format</a></p>\n<p>train data less than 1.7gb</p>",
      "rawMarkdown": "if you consider using feather/parquet data format, you could also consider cleaned up dataset which has some columns as integers:\n\nhttps://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\n\ntrain data less than 1.7gb\n",
      "votes": 2,
      "replies": [
        {
          "id": 1841792,
          "postDate": "2022-07-03T13:14:45.360Z",
          "content": "<p>could you please sharing the notebook for converting dataset?</p>",
          "rawMarkdown": "could you please sharing the notebook for converting dataset?"
        }
      ]
    },
    {
      "id": 1808128,
      "postDate": "2022-06-01T14:50:42.053Z",
      "content": "<p>The length of the training data in this data is 5531451 rows and the raw data (csv) of the competition is 458913 rows.<br>\nPlease let me know the reason if you don't mind.<br>\nThank you in advance.🙌</p>",
      "rawMarkdown": "The length of the training data in this data is 5531451 rows and the raw data (csv) of the competition is 458913 rows.\nPlease let me know the reason if you don't mind.\nThank you in advance.🙌",
      "votes": 2
    },
    {
      "id": 1802985,
      "postDate": "2022-05-27T11:01:18.350Z",
      "content": "<p>I was surprised to learn that the feather format can compress the full version down to less than 6 gigs!<br>\nThanks for sharing!👍</p>",
      "rawMarkdown": "I was surprised to learn that the feather format can compress the full version down to less than 6 gigs!\nThanks for sharing!👍",
      "votes": 2
    },
    {
      "id": 1801761,
      "postDate": "2022-05-26T05:56:04.117Z",
      "content": "<p>What are differences between parquet files and feather files, like if feather file is way smaller why use parquet and csv files anyway? </p>",
      "rawMarkdown": "What are differences between parquet files and feather files, like if feather file is way smaller why use parquet and csv files anyway? ",
      "votes": 2,
      "replies": [
        {
          "id": 1801767,
          "postDate": "2022-05-26T06:01:46.833Z",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/luna4444\" target=\"_blank\">@luna4444</a>,<br>\nIt depends on the use case of the application 😄<br>\nSometimes one file format is better than the other when it comes to reading/writing in terms of time!</p>",
          "rawMarkdown": "Hey @luna4444,\nIt depends on the use case of the application 😄\nSometimes one file format is better than the other when it comes to reading/writing in terms of time!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1865870,
      "postDate": "2022-07-22T06:31:19.523Z",
      "content": "<p>Did any one try the <strong>Terality</strong> python module? It should make data loading and preprocessing much faster than pandas😀</p>",
      "rawMarkdown": "Did any one try the **Terality** python module? It should make data loading and preprocessing much faster than pandas😀"
    },
    {
      "id": 1803354,
      "postDate": "2022-05-27T17:51:15.230Z",
      "content": "<p>Thank you very much for sharing.. Just one question.. does it reduce the precision of the data?</p>",
      "rawMarkdown": "Thank you very much for sharing.. Just one question.. does it reduce the precision of the data?",
      "isDeleted": true
    },
    {
      "id": 1809014,
      "postDate": "2022-06-02T10:19:41.447Z",
      "content": "<p>Thanks for sharing👍</p>",
      "rawMarkdown": "Thanks for sharing👍",
      "votes": 1
    },
    {
      "id": 1807197,
      "postDate": "2022-05-31T20:05:00.410Z",
      "content": "<p>It's really helpful for me, thanks</p>",
      "rawMarkdown": "It's really helpful for me, thanks",
      "votes": 1
    },
    {
      "id": 1802184,
      "postDate": "2022-05-26T14:32:41.257Z",
      "content": "<p>Really Thanks for sharing!👍</p>",
      "rawMarkdown": "Really Thanks for sharing!👍",
      "votes": 1
    },
    {
      "id": 1801834,
      "postDate": "2022-05-26T07:35:48.597Z",
      "content": "<p>good job! thanks for sharing!👍</p>",
      "rawMarkdown": "good job! thanks for sharing!👍",
      "votes": 1
    },
    {
      "id": 1801814,
      "postDate": "2022-05-26T07:04:17.470Z",
      "content": "<p>Thanks for sharing! It was helpful!</p>",
      "rawMarkdown": "Thanks for sharing! It was helpful!",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1801806,
      "author_name": "Pavel Orlov",
      "author_url": "",
      "post_date": "2022-05-26T06:57:46.837000",
      "content": "<p>How did you achieve such compression? When I save it <code>df.to_parquet</code> or <code>df.to_feather</code>, it turns out to be about the same size.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1801809,
          "author_name": "Ruchi Bhatia",
          "author_url": "",
          "post_date": "2022-05-26T07:00:23.503000",
          "content": "<p>Hello,<br>\nI converted the datatypes of the columns:</p>\n<ul>\n<li>Numerical columns to <code>float16</code> and </li>\n<li>Categorical columns to <code>category</code></li>\n</ul>\n<p>This really helped me to <em>reduce the size of the dataset immensely</em>.<br>\nSince we have so many rows in both the train and test dataset, making even a <em>tiny</em> transformation like downcasting can play a huge role!</p>\n<p>We can <strong>process our data faster</strong> since the dataset takes significantly <strong>less memory</strong> ✨</p>",
          "votes": 17,
          "replies": []
        }
      ]
    },
    {
      "id": 1919361,
      "author_name": "Prasad",
      "author_url": "",
      "post_date": "2022-08-30T10:37:59.453000",
      "content": "<p>Useful content <a href=\"https://www.kaggle.com/ruchi798\" target=\"_blank\">@ruchi798</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1855965,
      "author_name": "Hafiz Sayyed Ali Hamdani",
      "author_url": "",
      "post_date": "2022-07-15T03:28:35.617000",
      "content": "<p>Thanks for sharing.you explain it in a very simple way. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1826691,
      "author_name": "mayur1992",
      "author_url": "",
      "post_date": "2022-06-20T14:37:22.687000",
      "content": "<p>Very helpful</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1826648,
      "author_name": "Prathwish Mestha",
      "author_url": "",
      "post_date": "2022-06-20T14:12:40.757000",
      "content": "<p>Thaks for sharing <a href=\"https://www.kaggle.com/ruchi798\" target=\"_blank\">@ruchi798</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1826591,
      "author_name": "Md. Rasel Meya",
      "author_url": "",
      "post_date": "2022-06-20T13:21:34.193000",
      "content": "<p>Thanks for the impressive work. it's really needed for this huge type of dataset.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1824582,
      "author_name": "Kevin Kwan",
      "author_url": "",
      "post_date": "2022-06-18T13:11:17.607000",
      "content": "<p>Thanks for sharing this! Wow the compression is really quite impressive! I had no idea feather even existed and looks like it's worth exploring and learning.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1816777,
      "author_name": "Roopesh P",
      "author_url": "",
      "post_date": "2022-06-10T14:54:43.487000",
      "content": "<p>this is great!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1809046,
      "author_name": "Daniel Valyano",
      "author_url": "",
      "post_date": "2022-06-02T10:47:08.513000",
      "content": "<p>Well done, thank you for this solution, <a href=\"https://www.kaggle.com/ruchi798\" target=\"_blank\">@ruchi798</a> 🤜<br>\nHappy Kaggling!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1808912,
      "author_name": "Hassan Shehzad",
      "author_url": "",
      "post_date": "2022-06-02T08:36:49.503000",
      "content": "<p>Again job well done! <a href=\"https://www.kaggle.com/ruchi798\" target=\"_blank\">@ruchi798</a> 👍</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1808362,
      "author_name": "Karan Dora",
      "author_url": "",
      "post_date": "2022-06-01T18:27:54.087000",
      "content": "<p>Great help <a href=\"https://www.kaggle.com/ruchi798\" target=\"_blank\">@ruchi798</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1807663,
      "author_name": "@2L",
      "author_url": "",
      "post_date": "2022-06-01T08:48:45.200000",
      "content": "<p>For folks looking on how to convert to feather a good reference notebook : <a href=\"https://www.kaggle.com/code/kmader/convert-to-feather-for-use-in-other-kernels/notebook\" target=\"_blank\">https://www.kaggle.com/code/kmader/convert-to-feather-for-use-in-other-kernels/notebook</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1806119,
      "author_name": "Abdur Rakib Mollah",
      "author_url": "",
      "post_date": "2022-05-30T19:56:14.073000",
      "content": "<p>Feather does really a good job at compression it seems. Eager to learn more about it.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1805206,
      "author_name": "Anubhav Chhabra",
      "author_url": "",
      "post_date": "2022-05-29T23:20:25.277000",
      "content": "<p>Hey! Feather seems really interesting! Any opensource projects when it has been widely used?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1804365,
      "author_name": "Robert Kwiatkowski",
      "author_url": "",
      "post_date": "2022-05-28T20:14:24.557000",
      "content": "<p>Thank you for sharing this dataset with us. I'm already using it for my notebook. Honestly, I've never heard before about the feather format.</p>\n<p>Great job!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1803911,
      "author_name": "wuuthraad",
      "author_url": "",
      "post_date": "2022-05-28T09:46:05.040000",
      "content": "<p><strong>This is the ultimate compression</strong>!. Interesting how changing the data types makes a huge difference! Thanks again for sharing!  </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1803782,
      "author_name": "Mr.Dheeraj",
      "author_url": "",
      "post_date": "2022-05-28T07:21:00.543000",
      "content": "<p>It's really helpful for smooth runtime, thanks for sharing 👍</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1803705,
      "author_name": "Ruchi Bhatia",
      "author_url": "",
      "post_date": "2022-05-28T04:50:24.980000",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/rohitgarud\" target=\"_blank\">@rohitgarud</a> and <a href=\"https://www.kaggle.com/cv13j0\" target=\"_blank\">@cv13j0</a>,<br>\nTo answer your questions, </p>\n<ul>\n<li><strong>float16</strong> has reduced precision compared to float64 but is <strong>much faster</strong> than the latter.</li>\n<li><strong>float64</strong> consumes <strong>much more memory</strong> than float16 due to its higher precision digits, resulting in the large dataset size we have.</li>\n</ul>",
      "votes": 1,
      "replies": [
        {
          "id": 1808909,
          "author_name": "Hassan Shehzad",
          "author_url": "",
          "post_date": "2022-06-02T08:36:12.310000",
          "content": "<p><a href=\"https://www.kaggle.com/ruchi798\" target=\"_blank\">@ruchi798</a> thank you for the help.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1803480,
      "author_name": "C4rl05/V",
      "author_url": "",
      "post_date": "2022-05-27T21:02:01.983000",
      "content": "<p>Thanks for the dataset, <a href=\"https://www.kaggle.com/ruchi798\" target=\"_blank\">@ruchi798</a>; I will start using them today. Do you anticipate any precision loss or impacts to the ML model by decreasing the numeric columns to fload16?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1806114,
          "author_name": "Sola Sky",
          "author_url": "",
          "post_date": "2022-05-30T19:52:30.960000",
          "content": "<p>I would expect the model trained with float16 instead of float64 to be on average ~0.0000003814 more \"off\", which by most means would be negligible. However, if you want to test it train two models, one of float64 data and one of float16 data and see if there's a difference in scores.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1802199,
      "author_name": "James McNeill",
      "author_url": "",
      "post_date": "2022-05-26T14:42:16.997000",
      "content": "<p>Great job! Thanks very much for sharing these two methods 😊 I have never used the feather file format before. Do you have any links that you could share about it? I have always just assumed that the parquet format most the most efficient.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1808771,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "2022-06-02T06:30:55.303000",
      "content": "<p>if you consider using feather/parquet data format, you could also consider cleaned up dataset which has some columns as integers:</p>\n<p><a href=\"https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\" target=\"_blank\">https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format</a></p>\n<p>train data less than 1.7gb</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1841792,
          "author_name": "dragon zhang",
          "author_url": "",
          "post_date": "2022-07-03T13:14:45.360000",
          "content": "<p>could you please sharing the notebook for converting dataset?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1808128,
      "author_name": "chinchilla",
      "author_url": "",
      "post_date": "2022-06-01T14:50:42.053000",
      "content": "<p>The length of the training data in this data is 5531451 rows and the raw data (csv) of the competition is 458913 rows.<br>\nPlease let me know the reason if you don't mind.<br>\nThank you in advance.🙌</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1802985,
      "author_name": "chinchilla",
      "author_url": "",
      "post_date": "2022-05-27T11:01:18.350000",
      "content": "<p>I was surprised to learn that the feather format can compress the full version down to less than 6 gigs!<br>\nThanks for sharing!👍</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1801761,
      "author_name": "Luna_4444",
      "author_url": "",
      "post_date": "2022-05-26T05:56:04.117000",
      "content": "<p>What are differences between parquet files and feather files, like if feather file is way smaller why use parquet and csv files anyway? </p>",
      "votes": 2,
      "replies": [
        {
          "id": 1801767,
          "author_name": "Ruchi Bhatia",
          "author_url": "",
          "post_date": "2022-05-26T06:01:46.833000",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/luna4444\" target=\"_blank\">@luna4444</a>,<br>\nIt depends on the use case of the application 😄<br>\nSometimes one file format is better than the other when it comes to reading/writing in terms of time!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1865870,
      "author_name": "Sunnymoon Sultan",
      "author_url": "",
      "post_date": "2022-07-22T06:31:19.523000",
      "content": "<p>Did any one try the <strong>Terality</strong> python module? It should make data loading and preprocessing much faster than pandas😀</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1803354,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-27T17:51:15.230000",
      "content": "<p>Thank you very much for sharing.. Just one question.. does it reduce the precision of the data?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1809014,
      "author_name": "Paddy",
      "author_url": "",
      "post_date": "2022-06-02T10:19:41.447000",
      "content": "<p>Thanks for sharing👍</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1807197,
      "author_name": "MIkhail Donskoy",
      "author_url": "",
      "post_date": "2022-05-31T20:05:00.410000",
      "content": "<p>It's really helpful for me, thanks</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1802184,
      "author_name": "Mohit",
      "author_url": "",
      "post_date": "2022-05-26T14:32:41.257000",
      "content": "<p>Really Thanks for sharing!👍</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1801834,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-26T07:35:48.597000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1801814,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-26T07:04:17.470000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1801493": "As you might have noticed, the tabular files for this competition are HUGE 🤯\n\n#### 💡What are the Feather and Parquet file formats?\n- These file formats are **designed to handle complex data in bulk**.\n- They are preferred due to their **high efficiency** in data compression and decompression.\n\n#### 💳 Compressed Dataset\nTo help you read the data faster, I've created [feather and parquet versions](https://www.kaggle.com/datasets/ruchi798/parquet-files-amexdefault-prediction) of the csv files and curated a dataset.\n\n| Data       | Size of csv file | Size of parquet file  | **Size of feather file**  |\n|------------|----------|---------------|---------------|\n| train_data | 16.4 GB  | 6.7 GB        | **1.8 GB**        |\n| test_data  | 33.8 GB  | 13.7 GB       | **3.6 GB**        |\n\n#### ⚙️ Process\nChanged the datatypes of the columns as follows:\n- Numerical columns -> **`float16`** and\n- Categorical columns -> **`category`**",
    "1801806": "How did you achieve such compression? When I save it `df.to_parquet` or `df.to_feather`, it turns out to be about the same size.",
    "1919361": "Useful content @ruchi798 ",
    "1855965": "Thanks for sharing.you explain it in a very simple way. ",
    "1826691": "Very helpful",
    "1826648": "Thaks for sharing @ruchi798 ",
    "1826591": "Thanks for the impressive work. it's really needed for this huge type of dataset.",
    "1824582": "Thanks for sharing this! Wow the compression is really quite impressive! I had no idea feather even existed and looks like it's worth exploring and learning.",
    "1816777": "this is great!",
    "1809046": "Well done, thank you for this solution, @ruchi798 🤜\nHappy Kaggling!",
    "1808912": "Again job well done! @ruchi798 👍",
    "1808362": "Great help @ruchi798 ",
    "1807663": "For folks looking on how to convert to feather a good reference notebook : https://www.kaggle.com/code/kmader/convert-to-feather-for-use-in-other-kernels/notebook",
    "1806119": "Feather does really a good job at compression it seems. Eager to learn more about it.",
    "1805206": "Hey! Feather seems really interesting! Any opensource projects when it has been widely used?",
    "1804365": "Thank you for sharing this dataset with us. I'm already using it for my notebook. Honestly, I've never heard before about the feather format.\n\nGreat job!",
    "1803911": "**This is the ultimate compression**!. Interesting how changing the data types makes a huge difference! Thanks again for sharing!  ",
    "1803782": "It's really helpful for smooth runtime, thanks for sharing 👍",
    "1803705": "Hello @rohitgarud and @cv13j0,\nTo answer your questions, \n- **float16** has reduced precision compared to float64 but is **much faster** than the latter.\n- **float64** consumes **much more memory** than float16 due to its higher precision digits, resulting in the large dataset size we have.",
    "1803480": "Thanks for the dataset, @ruchi798; I will start using them today. Do you anticipate any precision loss or impacts to the ML model by decreasing the numeric columns to fload16?",
    "1802199": "Great job! Thanks very much for sharing these two methods 😊 I have never used the feather file format before. Do you have any links that you could share about it? I have always just assumed that the parquet format most the most efficient.",
    "1808771": "if you consider using feather/parquet data format, you could also consider cleaned up dataset which has some columns as integers:\n\nhttps://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\n\ntrain data less than 1.7gb\n",
    "1808128": "The length of the training data in this data is 5531451 rows and the raw data (csv) of the competition is 458913 rows.\nPlease let me know the reason if you don't mind.\nThank you in advance.🙌",
    "1802985": "I was surprised to learn that the feather format can compress the full version down to less than 6 gigs!\nThanks for sharing!👍",
    "1801761": "What are differences between parquet files and feather files, like if feather file is way smaller why use parquet and csv files anyway? ",
    "1865870": "Did any one try the **Terality** python module? It should make data loading and preprocessing much faster than pandas😀",
    "1803354": "Thank you very much for sharing.. Just one question.. does it reduce the precision of the data?",
    "1809014": "Thanks for sharing👍",
    "1807197": "It's really helpful for me, thanks",
    "1802184": "Really Thanks for sharing!👍",
    "1801834": "good job! thanks for sharing!👍",
    "1801814": "Thanks for sharing! It was helpful!"
  }
}