{
  "id": 356582,
  "title": "9x data compression  Faster using Feather",
  "url": "/competitions/tabular-playground-series-oct-2022/discussion/356582",
  "author_name": "",
  "post_date": "2022-10-01T06:39:20.432242200Z",
  "votes": 5,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I create a dataset using feather which helps you to load data 8 times faster than CSV format data. </p>\n<table>\n<thead>\n<tr>\n<th>data</th>\n<th>Size of CSV</th>\n<th>size of Feature</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Train_0</td>\n<td>954 MB</td>\n<td>568 MB</td>\n</tr>\n<tr>\n<td>Test</td>\n<td>296 MB</td>\n<td>208 MB</td>\n</tr>\n</tbody>\n</table>\n<p>My Dataset : <a href=\"https://www.kaggle.com/datasets/gazu468/tpsoct22-feather-files\">this link</a><br>\nMy Notebook : <a href=\"https://www.kaggle.com/code/gazu468/feather-to-compress-your-data-6x-faster\"> check this</a><br>\nUsing XGboost Classifier Publish a notebook : <a href=\"https://www.kaggle.com/code/gazu468/tps-oct-22-simple-eda-and-modeling/notebook\"> Use that</a></p>",
  "messages": [
    {
      "id": "1965097",
      "postDate": "10/01/2022 06:39:20",
      "content": "<p>I create a dataset using feather which helps you to load data 8 times faster than CSV format data. </p>\n<table>\n<thead>\n<tr>\n<th>data</th>\n<th>Size of CSV</th>\n<th>size of Feature</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Train_0</td>\n<td>954 MB</td>\n<td>568 MB</td>\n</tr>\n<tr>\n<td>Test</td>\n<td>296 MB</td>\n<td>208 MB</td>\n</tr>\n</tbody>\n</table>\n<p>My Dataset : <a href=\"https://www.kaggle.com/datasets/gazu468/tpsoct22-feather-files\">this link</a><br>\nMy Notebook : <a href=\"https://www.kaggle.com/code/gazu468/feather-to-compress-your-data-6x-faster\"> check this</a><br>\nUsing XGboost Classifier Publish a notebook : <a href=\"https://www.kaggle.com/code/gazu468/tps-oct-22-simple-eda-and-modeling/notebook\"> Use that</a></p>",
      "rawMarkdown": "I create a dataset using feather which helps you to load data 8 times faster than CSV format data. \n| data | Size of CSV |size of Feature\n| --- | --- |\n| Train_0 |954 MB |568 MB\n|Test|296 MB |208 MB|\n\n\nMy Dataset : <a href=\"https://www.kaggle.com/datasets/gazu468/tpsoct22-feather-files\">this link</a>\nMy Notebook : <a href=\"https://www.kaggle.com/code/gazu468/feather-to-compress-your-data-6x-faster\"> check this</a>\nUsing XGboost Classifier Publish a notebook : <a href=\"https://www.kaggle.com/code/gazu468/tps-oct-22-simple-eda-and-modeling/notebook\"> Use that</a>",
      "votes": null
    },
    {
      "id": "1966140",
      "postDate": "10/01/2022 18:20:59",
      "content": "<p><a href=\"https://www.kaggle.com/gazu468\" target=\"_blank\">@gazu468</a> I am wondering why I never saw this library till now🙄</p>",
      "rawMarkdown": "gazu468 I am wondering why I never saw this library till now🙄",
      "votes": null
    },
    {
      "id": "1966258",
      "postDate": "10/01/2022 19:59:05",
      "content": "<p>There are more libraries like this one of those is <a href=\"https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.to_parquet.html\">Parquet</a></p>",
      "rawMarkdown": "There are more libraries like this one of those is <a href=\"https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.to_parquet.html\">Parquet</a>",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1966140,
      "author_name": "abrafey",
      "author_url": "",
      "post_date": "10/01/2022 18:20:59",
      "content": "<p><a href=\"https://www.kaggle.com/gazu468\" target=\"_blank\">@gazu468</a> I am wondering why I never saw this library till now🙄</p>",
      "votes": null,
      "replies": [
        {
          "id": 1966258,
          "author_name": "gazu468",
          "author_url": "",
          "post_date": "10/01/2022 19:59:05",
          "content": "<p>There are more libraries like this one of those is <a href=\"https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.to_parquet.html\">Parquet</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1965097": "I create a dataset using feather which helps you to load data 8 times faster than CSV format data. \n| data | Size of CSV |size of Feature\n| --- | --- |\n| Train_0 |954 MB |568 MB\n|Test|296 MB |208 MB|\n\n\nMy Dataset : <a href=\"https://www.kaggle.com/datasets/gazu468/tpsoct22-feather-files\">this link</a>\nMy Notebook : <a href=\"https://www.kaggle.com/code/gazu468/feather-to-compress-your-data-6x-faster\"> check this</a>\nUsing XGboost Classifier Publish a notebook : <a href=\"https://www.kaggle.com/code/gazu468/tps-oct-22-simple-eda-and-modeling/notebook\"> Use that</a>",
    "1966140": "gazu468 I am wondering why I never saw this library till now🙄",
    "1966258": "There are more libraries like this one of those is <a href=\"https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.to_parquet.html\">Parquet</a>"
  },
  "source": "meta"
}