{
  "id": 341166,
  "title": "How to convert \"csv-file\" to \"parquet-file\"",
  "url": "/competitions/amex-default-prediction/discussion/341166",
  "author_name": "",
  "post_date": "2022-08-01T15:06:01.475515600Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I can't convert \"train_data.csv\" to \"train_data.parquet\".<br>\nIn my way, it takes long time to load \"train_data.csv\" by pandas.</p>\n<p>I want to know how to convert \"csv-file\" to \"parquet-file\" smoothly.<br>\nIf someone give me advice, I am happy.</p>",
  "messages": [
    {
      "id": "1880255",
      "postDate": "08/01/2022 15:06:01",
      "content": "<p>I can't convert \"train_data.csv\" to \"train_data.parquet\".<br>\nIn my way, it takes long time to load \"train_data.csv\" by pandas.</p>\n<p>I want to know how to convert \"csv-file\" to \"parquet-file\" smoothly.<br>\nIf someone give me advice, I am happy.</p>",
      "rawMarkdown": "I can't convert \"train_data.csv\" to \"train_data.parquet\".\nIn my way, it takes long time to load \"train_data.csv\" by pandas.\n\nI want to know how to convert \"csv-file\" to \"parquet-file\" smoothly.\nIf someone give me advice, I am happy.",
      "votes": null
    },
    {
      "id": "1880444",
      "postDate": "08/01/2022 18:10:36",
      "content": "<p>The best suggestion is to use Raddar's parquet <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514\" target=\"_blank\">here</a>. When converting CSV to Parquet you must do more than <code>read_csv()</code> then <code>to_parquet()</code>. After reading CSV, you must convert each column into using less bytes per row. This refers to label encoding strings, removing unnecessary noise, and converting dtypes. The full process is described <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328054\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "The best suggestion is to use Raddar's parquet [here][1]. When converting CSV to Parquet you must do more than `read_csv()` then `to_parquet()`. After reading CSV, you must convert each column into using less bytes per row. This refers to label encoding strings, removing unnecessary noise, and converting dtypes. The full process is described [here][2]\n\n[1]: https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514\n[2]: https://www.kaggle.com/competitions/amex-default-prediction/discussion/328054",
      "votes": null
    },
    {
      "id": "1881474",
      "postDate": "08/02/2022 14:37:37",
      "content": "<p>Thanks to your kind advice, I have solved my problem ! <br>\nThank you ! </p>",
      "rawMarkdown": "Thanks to your kind advice, I have solved my problem ! \nThank you !",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1880444,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "08/01/2022 18:10:36",
      "content": "<p>The best suggestion is to use Raddar's parquet <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514\" target=\"_blank\">here</a>. When converting CSV to Parquet you must do more than <code>read_csv()</code> then <code>to_parquet()</code>. After reading CSV, you must convert each column into using less bytes per row. This refers to label encoding strings, removing unnecessary noise, and converting dtypes. The full process is described <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328054\" target=\"_blank\">here</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1881474,
      "author_name": "rabishibata",
      "author_url": "",
      "post_date": "08/02/2022 14:37:37",
      "content": "<p>Thanks to your kind advice, I have solved my problem ! <br>\nThank you ! </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1880255": "I can't convert \"train_data.csv\" to \"train_data.parquet\".\nIn my way, it takes long time to load \"train_data.csv\" by pandas.\n\nI want to know how to convert \"csv-file\" to \"parquet-file\" smoothly.\nIf someone give me advice, I am happy.",
    "1880444": "The best suggestion is to use Raddar's parquet [here][1]. When converting CSV to Parquet you must do more than `read_csv()` then `to_parquet()`. After reading CSV, you must convert each column into using less bytes per row. This refers to label encoding strings, removing unnecessary noise, and converting dtypes. The full process is described [here][2]\n\n[1]: https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514\n[2]: https://www.kaggle.com/competitions/amex-default-prediction/discussion/328054",
    "1881474": "Thanks to your kind advice, I have solved my problem ! \nThank you !"
  },
  "source": "meta"
}