{
  "id": 335362,
  "title": "Strange overflow error (df.write_parquet()",
  "url": "/competitions/amex-default-prediction/discussion/335362",
  "author_name": "Xavier R Nogueira",
  "post_date": "2022-07-05T21:18:00.194000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p>So I was looking to import the data from .csv into pandas, convert dtypes to how I want them (object -&gt; category, int -&gt; int64, 'S_2' -&gt; datetime[ns]), then save to .parquet which holds the dtypes so I can restart the notebook much quicker from parquet and not have to reformat columns.</p>\n<p>Seems easy enough, but once I get the dtypes set the to_parquet() function fails with a mysterious <a href=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8611170%2F0298ef13af34b16a0b2a1b0f74ebe6cd%2Foverflow.PNG?generation=1657055853494274&amp;alt=media\" target=\"_blank\">overflow error</a>. At first I thought it was because some int column contains values &gt; the system max, but upon inspection this is not the case as all int columns as int64 which should be able to store very large numbers.</p>\n<p>I am using the following command on test_data and train_data:<code>df.to_parquet(out_title, engine='fastparquet', compression='gzip')</code></p>\n<p>Anyone else run into a similar issue? Would love to be able to share the formatted .parquet data!</p>\n<p>Thanks!</p>",
  "messages": [
    {
      "id": 1844853,
      "postDate": "2022-07-05T21:18:00.193Z",
      "content": "<p>So I was looking to import the data from .csv into pandas, convert dtypes to how I want them (object -&gt; category, int -&gt; int64, 'S_2' -&gt; datetime[ns]), then save to .parquet which holds the dtypes so I can restart the notebook much quicker from parquet and not have to reformat columns.</p>\n<p>Seems easy enough, but once I get the dtypes set the to_parquet() function fails with a mysterious <a href=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8611170%2F0298ef13af34b16a0b2a1b0f74ebe6cd%2Foverflow.PNG?generation=1657055853494274&amp;alt=media\" target=\"_blank\">overflow error</a>. At first I thought it was because some int column contains values &gt; the system max, but upon inspection this is not the case as all int columns as int64 which should be able to store very large numbers.</p>\n<p>I am using the following command on test_data and train_data:<code>df.to_parquet(out_title, engine='fastparquet', compression='gzip')</code></p>\n<p>Anyone else run into a similar issue? Would love to be able to share the formatted .parquet data!</p>\n<p>Thanks!</p>",
      "rawMarkdown": "So I was looking to import the data from .csv into pandas, convert dtypes to how I want them (object -> category, int -> int64, 'S_2' -> datetime[ns]), then save to .parquet which holds the dtypes so I can restart the notebook much quicker from parquet and not have to reformat columns.\n\nSeems easy enough, but once I get the dtypes set the to_parquet() function fails with a mysterious [overflow error](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8611170%2F0298ef13af34b16a0b2a1b0f74ebe6cd%2Foverflow.PNG?generation=1657055853494274&alt=media). At first I thought it was because some int column contains values > the system max, but upon inspection this is not the case as all int columns as int64 which should be able to store very large numbers.\n\nI am using the following command on test_data and train_data:`df.to_parquet(out_title, engine='fastparquet', compression='gzip')`\n\nAnyone else run into a similar issue? Would love to be able to share the formatted .parquet data!\n\nThanks!\n\n\n\n"
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1844853": "So I was looking to import the data from .csv into pandas, convert dtypes to how I want them (object -> category, int -> int64, 'S_2' -> datetime[ns]), then save to .parquet which holds the dtypes so I can restart the notebook much quicker from parquet and not have to reformat columns.\n\nSeems easy enough, but once I get the dtypes set the to_parquet() function fails with a mysterious [overflow error](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8611170%2F0298ef13af34b16a0b2a1b0f74ebe6cd%2Foverflow.PNG?generation=1657055853494274&alt=media). At first I thought it was because some int column contains values > the system max, but upon inspection this is not the case as all int columns as int64 which should be able to store very large numbers.\n\nI am using the following command on test_data and train_data:`df.to_parquet(out_title, engine='fastparquet', compression='gzip')`\n\nAnyone else run into a similar issue? Would love to be able to share the formatted .parquet data!\n\nThanks!\n\n\n\n"
  }
}