{
  "id": 135563,
  "title": "Notebook: Reading Parquet Files - RAM/CPU Optimization",
  "url": "/competitions/bengaliai-cv19/discussion/135563",
  "author_name": "James McGuigan",
  "post_date": "2020-03-14T15:46:23.547000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>One of the big problems I have been facing with the Bengali.AI competition is getting \"Failed. Exited with code 137\" (Out of Memory) errors on Commit, and even worse the dreaded \"Notebook Exceeded Allowed Compute\" when submitting the final submission.csv</p>\n\n<p>As such, I decided I needed to do a more indepth analysis of how to read Parquet files, and do some careful CPU/Memory profiling of how much data was being loaded into RAM and the true effect of using different dtypes (by default python casts <code>int</code> -&gt; <code>float64</code> which memory wise is 8x <code>uint8</code> and 4x <code>float16</code>).</p>\n\n<p>I also figured out how to write a generator function around <code>pandas.read_parquet()</code> that allows trading 50% RAM for 2x disk IO. I also discovered that <code>skimage.measure.block_reduce(output, (1,2,2,1), func=np.mean, cval=0)</code> can be used to downsample the raw image data (with 4x RAM saving).</p>\n\n<p><a href=\"https://www.kaggle.com/jamesmcguigan/reading-parquet-files-ram-cpu-optimization\">https://www.kaggle.com/jamesmcguigan/reading-parquet-files-ram-cpu-optimization</a></p>",
  "messages": [
    {
      "id": 771768,
      "postDate": "2020-03-14T15:46:23.547Z",
      "content": "<p>One of the big problems I have been facing with the Bengali.AI competition is getting \"Failed. Exited with code 137\" (Out of Memory) errors on Commit, and even worse the dreaded \"Notebook Exceeded Allowed Compute\" when submitting the final submission.csv</p>\n\n<p>As such, I decided I needed to do a more indepth analysis of how to read Parquet files, and do some careful CPU/Memory profiling of how much data was being loaded into RAM and the true effect of using different dtypes (by default python casts <code>int</code> -&gt; <code>float64</code> which memory wise is 8x <code>uint8</code> and 4x <code>float16</code>).</p>\n\n<p>I also figured out how to write a generator function around <code>pandas.read_parquet()</code> that allows trading 50% RAM for 2x disk IO. I also discovered that <code>skimage.measure.block_reduce(output, (1,2,2,1), func=np.mean, cval=0)</code> can be used to downsample the raw image data (with 4x RAM saving).</p>\n\n<p><a href=\"https://www.kaggle.com/jamesmcguigan/reading-parquet-files-ram-cpu-optimization\">https://www.kaggle.com/jamesmcguigan/reading-parquet-files-ram-cpu-optimization</a></p>",
      "rawMarkdown": "One of the big problems I have been facing with the Bengali.AI competition is getting \"Failed. Exited with code 137\" (Out of Memory) errors on Commit, and even worse the dreaded \"Notebook Exceeded Allowed Compute\" when submitting the final submission.csv\n\nAs such, I decided I needed to do a more indepth analysis of how to read Parquet files, and do some careful CPU/Memory profiling of how much data was being loaded into RAM and the true effect of using different dtypes (by default python casts `int` -&gt; `float64` which memory wise is 8x `uint8` and 4x `float16`).\n\nI also figured out how to write a generator function around `pandas.read_parquet()` that allows trading 50% RAM for 2x disk IO. I also discovered that `skimage.measure.block_reduce(output, (1,2,2,1), func=np.mean, cval=0)` can be used to downsample the raw image data (with 4x RAM saving).\n\nhttps://www.kaggle.com/jamesmcguigan/reading-parquet-files-ram-cpu-optimization",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "771768": "One of the big problems I have been facing with the Bengali.AI competition is getting \"Failed. Exited with code 137\" (Out of Memory) errors on Commit, and even worse the dreaded \"Notebook Exceeded Allowed Compute\" when submitting the final submission.csv\n\nAs such, I decided I needed to do a more indepth analysis of how to read Parquet files, and do some careful CPU/Memory profiling of how much data was being loaded into RAM and the true effect of using different dtypes (by default python casts `int` -&gt; `float64` which memory wise is 8x `uint8` and 4x `float16`).\n\nI also figured out how to write a generator function around `pandas.read_parquet()` that allows trading 50% RAM for 2x disk IO. I also discovered that `skimage.measure.block_reduce(output, (1,2,2,1), func=np.mean, cval=0)` can be used to downsample the raw image data (with 4x RAM saving).\n\nhttps://www.kaggle.com/jamesmcguigan/reading-parquet-files-ram-cpu-optimization"
  }
}