{
  "id": 381652,
  "title": "Memory leak in reading parquet files",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/381652",
  "author_name": "Kaan Güven",
  "post_date": "2023-01-27T15:06:14.593000",
  "votes": 7,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi all! I want to share what I learned the hard way about reading parquet files.<br>\nIt was driving me crazy, and I hope this will help some of us.</p>\n<p>Say we read a parquet file with pandas:<br>\n<code>pd.read_parquet(\"/kaggle/input/icecube-neutrinos-in-deep-ice/train/batch_1.parquet\")</code><br>\nIn notebooks this also displays the file. If you run this code cell many times I run out of memory.<br>\nHowever, the following does not replicate the same behaviour:<br>\n<code>df = pd.read_parquet(\"/kaggle/input/icecube-neutrinos-in-deep-ice/train/batch_1.parquet\")</code></p>\n<p>Trying to track down the memory leak, and fix it I tried the <code>gc.collect()</code> method.<br>\nAlso I tried to manually delete the variables from locals().items().<br>\nNothing worked, but I found what causes this. It is pyarrow's memory allocator.<br>\nThere is a function that I used to check the memory allocated by pyarrow:<br>\n<code>pyarrow.total_allocated_bytes()</code></p>\n<p>So, I don't know how to get the memory back once it is allocated, but at least I now know that<br>\nI shouldn't display any files if I am not doing exploratory data analysis.<br>\nMaybe you know more about the topic, and want to enlighten us.<br>\nIf so please write your comments on this issue. :)</p>",
  "messages": [
    {
      "id": 2117808,
      "postDate": "2023-01-27T15:06:14.593Z",
      "content": "<p>Hi all! I want to share what I learned the hard way about reading parquet files.<br>\nIt was driving me crazy, and I hope this will help some of us.</p>\n<p>Say we read a parquet file with pandas:<br>\n<code>pd.read_parquet(\"/kaggle/input/icecube-neutrinos-in-deep-ice/train/batch_1.parquet\")</code><br>\nIn notebooks this also displays the file. If you run this code cell many times I run out of memory.<br>\nHowever, the following does not replicate the same behaviour:<br>\n<code>df = pd.read_parquet(\"/kaggle/input/icecube-neutrinos-in-deep-ice/train/batch_1.parquet\")</code></p>\n<p>Trying to track down the memory leak, and fix it I tried the <code>gc.collect()</code> method.<br>\nAlso I tried to manually delete the variables from locals().items().<br>\nNothing worked, but I found what causes this. It is pyarrow's memory allocator.<br>\nThere is a function that I used to check the memory allocated by pyarrow:<br>\n<code>pyarrow.total_allocated_bytes()</code></p>\n<p>So, I don't know how to get the memory back once it is allocated, but at least I now know that<br>\nI shouldn't display any files if I am not doing exploratory data analysis.<br>\nMaybe you know more about the topic, and want to enlighten us.<br>\nIf so please write your comments on this issue. :)</p>",
      "rawMarkdown": "Hi all! I want to share what I learned the hard way about reading parquet files.\nIt was driving me crazy, and I hope this will help some of us.\n\nSay we read a parquet file with pandas:\n`pd.read_parquet(\"/kaggle/input/icecube-neutrinos-in-deep-ice/train/batch_1.parquet\")`\nIn notebooks this also displays the file. If you run this code cell many times I run out of memory.\nHowever, the following does not replicate the same behaviour:\n`df = pd.read_parquet(\"/kaggle/input/icecube-neutrinos-in-deep-ice/train/batch_1.parquet\")`\n\nTrying to track down the memory leak, and fix it I tried the `gc.collect()` method.\nAlso I tried to manually delete the variables from locals().items().\nNothing worked, but I found what causes this. It is pyarrow's memory allocator.\nThere is a function that I used to check the memory allocated by pyarrow:\n`pyarrow.total_allocated_bytes()`\n\nSo, I don't know how to get the memory back once it is allocated, but at least I now know that\nI shouldn't display any files if I am not doing exploratory data analysis.\nMaybe you know more about the topic, and want to enlighten us.\nIf so please write your comments on this issue. :)\n",
      "votes": 7
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2117808": "Hi all! I want to share what I learned the hard way about reading parquet files.\nIt was driving me crazy, and I hope this will help some of us.\n\nSay we read a parquet file with pandas:\n`pd.read_parquet(\"/kaggle/input/icecube-neutrinos-in-deep-ice/train/batch_1.parquet\")`\nIn notebooks this also displays the file. If you run this code cell many times I run out of memory.\nHowever, the following does not replicate the same behaviour:\n`df = pd.read_parquet(\"/kaggle/input/icecube-neutrinos-in-deep-ice/train/batch_1.parquet\")`\n\nTrying to track down the memory leak, and fix it I tried the `gc.collect()` method.\nAlso I tried to manually delete the variables from locals().items().\nNothing worked, but I found what causes this. It is pyarrow's memory allocator.\nThere is a function that I used to check the memory allocated by pyarrow:\n`pyarrow.total_allocated_bytes()`\n\nSo, I don't know how to get the memory back once it is allocated, but at least I now know that\nI shouldn't display any files if I am not doing exploratory data analysis.\nMaybe you know more about the topic, and want to enlighten us.\nIf so please write your comments on this issue. :)\n"
  }
}