{
  "id": 327441,
  "title": "🗜️10x Compression of whole dataset using Feather",
  "url": "/competitions/amex-default-prediction/discussion/327441",
  "author_name": "",
  "post_date": "2022-05-27T09:27:25.556386800Z",
  "votes": 6,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hey everyone,<br>\nI have created feather format files from train_data, train_labels, and test_data CSV files.</p>\n<p>Total Dataset: 50.31 GB -&gt; <strong>5.1 GB</strong><br>\ntrain_data: 16.39 GB -&gt; <strong>1.67 GB</strong><br>\ntest_data: 33.82 GB -&gt; <strong>3.44 GB</strong><br>\ntrain_labels: 30.75 MB -&gt; <strong>2.37 MB</strong></p>\n<p>Compression - <br>\nConvert all float64 dtypes to float16<br>\nSet Categorical columns to category dtype <br>\nConvert all int64 dtypes to int8<br>\nConverted customer_ID to int32 dtype to save some memory usage. Credit to <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/308635\" target=\"_blank\">cdeotte</a><br>\nStoring the DateTime data (S_2) as 3 separate columns i.e year, month, and date (int8 dtype) to preserve more space.</p>\n<p>Link to Dataset - <a href=\"https://www.kaggle.com/datasets/beraamit/amex-default-prediction-10x-compression\" target=\"_blank\">https://www.kaggle.com/datasets/beraamit/amex-default-prediction-10x-compression</a></p>",
  "messages": [
    {
      "id": "1802892",
      "postDate": "05/27/2022 09:27:25",
      "content": "<p>Hey everyone,<br>\nI have created feather format files from train_data, train_labels, and test_data CSV files.</p>\n<p>Total Dataset: 50.31 GB -&gt; <strong>5.1 GB</strong><br>\ntrain_data: 16.39 GB -&gt; <strong>1.67 GB</strong><br>\ntest_data: 33.82 GB -&gt; <strong>3.44 GB</strong><br>\ntrain_labels: 30.75 MB -&gt; <strong>2.37 MB</strong></p>\n<p>Compression - <br>\nConvert all float64 dtypes to float16<br>\nSet Categorical columns to category dtype <br>\nConvert all int64 dtypes to int8<br>\nConverted customer_ID to int32 dtype to save some memory usage. Credit to <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/308635\" target=\"_blank\">cdeotte</a><br>\nStoring the DateTime data (S_2) as 3 separate columns i.e year, month, and date (int8 dtype) to preserve more space.</p>\n<p>Link to Dataset - <a href=\"https://www.kaggle.com/datasets/beraamit/amex-default-prediction-10x-compression\" target=\"_blank\">https://www.kaggle.com/datasets/beraamit/amex-default-prediction-10x-compression</a></p>",
      "rawMarkdown": "Hey everyone,\nI have created feather format files from train_data, train_labels, and test_data CSV files.\n\nTotal Dataset: 50.31 GB -> **5.1 GB**\ntrain_data: 16.39 GB -> **1.67 GB**\ntest_data: 33.82 GB -> **3.44 GB**\ntrain_labels: 30.75 MB -> **2.37 MB**\n\nCompression - \nConvert all float64 dtypes to float16\nSet Categorical columns to category dtype \nConvert all int64 dtypes to int8\nConverted customer_ID to int32 dtype to save some memory usage. Credit to [cdeotte](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/308635)\nStoring the DateTime data (S_2) as 3 separate columns i.e year, month, and date (int8 dtype) to preserve more space.\n\nLink to Dataset - https://www.kaggle.com/datasets/beraamit/amex-default-prediction-10x-compression",
      "votes": null
    },
    {
      "id": "1803628",
      "postDate": "05/28/2022 02:59:26",
      "content": "<p>Thanks for your great work! Are we allowed to use these files for the competition?</p>",
      "rawMarkdown": "Thanks for your great work! Are we allowed to use these files for the competition?",
      "votes": null
    },
    {
      "id": "1803640",
      "postDate": "05/28/2022 03:32:34",
      "content": "<p>yes, you can use this file, this doesnt break any rule</p>",
      "rawMarkdown": "yes, you can use this file, this doesnt break any rule",
      "votes": null
    },
    {
      "id": "1882654",
      "postDate": "08/03/2022 11:16:18",
      "content": "<p>Thanks. How did you make these file? can you please explain the whole procedure. Many of the competitors are using files that have size in MBs </p>",
      "rawMarkdown": "Thanks. How did you make these file? can you please explain the whole procedure. Many of the competitors are using files that have size in MBs",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1803628,
      "author_name": "hoale2908",
      "author_url": "",
      "post_date": "05/28/2022 02:59:26",
      "content": "<p>Thanks for your great work! Are we allowed to use these files for the competition?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1803640,
          "author_name": "beraamit",
          "author_url": "",
          "post_date": "05/28/2022 03:32:34",
          "content": "<p>yes, you can use this file, this doesnt break any rule</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1882654,
      "author_name": "kulkarniishwarinitin",
      "author_url": "",
      "post_date": "08/03/2022 11:16:18",
      "content": "<p>Thanks. How did you make these file? can you please explain the whole procedure. Many of the competitors are using files that have size in MBs </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1802892": "Hey everyone,\nI have created feather format files from train_data, train_labels, and test_data CSV files.\n\nTotal Dataset: 50.31 GB -> **5.1 GB**\ntrain_data: 16.39 GB -> **1.67 GB**\ntest_data: 33.82 GB -> **3.44 GB**\ntrain_labels: 30.75 MB -> **2.37 MB**\n\nCompression - \nConvert all float64 dtypes to float16\nSet Categorical columns to category dtype \nConvert all int64 dtypes to int8\nConverted customer_ID to int32 dtype to save some memory usage. Credit to [cdeotte](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/308635)\nStoring the DateTime data (S_2) as 3 separate columns i.e year, month, and date (int8 dtype) to preserve more space.\n\nLink to Dataset - https://www.kaggle.com/datasets/beraamit/amex-default-prediction-10x-compression",
    "1803628": "Thanks for your great work! Are we allowed to use these files for the competition?",
    "1803640": "yes, you can use this file, this doesnt break any rule",
    "1882654": "Thanks. How did you make these file? can you please explain the whole procedure. Many of the competitors are using files that have size in MBs"
  },
  "source": "meta"
}