{
  "id": 381658,
  "title": "Reading Parquet Files - RAM/CPU Optimization",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/381658",
  "author_name": "",
  "post_date": "2023-01-27T16:03:51.618276Z",
  "votes": 4,
  "comment_count": 1,
  "views": 0,
  "content": "<p>The dataset for this competition is stored in parquet files (pandas + pyarrow).</p>\n<p>I wrote this notebook exploring how to performance optimize reading the parquet file format into memory, which was a infrastructure-level challenge for Bengali AI, which had oversized datasets.</p>\n<p>Hopefully this will be of use to some people in this competition.</p>\n<p><a href=\"https://www.kaggle.com/code/jamesmcguigan/reading-parquet-files-ram-cpu-optimization\" target=\"_blank\">https://www.kaggle.com/code/jamesmcguigan/reading-parquet-files-ram-cpu-optimization</a> </p>",
  "messages": [
    {
      "id": "2117886",
      "postDate": "01/27/2023 16:03:51",
      "content": "<p>The dataset for this competition is stored in parquet files (pandas + pyarrow).</p>\n<p>I wrote this notebook exploring how to performance optimize reading the parquet file format into memory, which was a infrastructure-level challenge for Bengali AI, which had oversized datasets.</p>\n<p>Hopefully this will be of use to some people in this competition.</p>\n<p><a href=\"https://www.kaggle.com/code/jamesmcguigan/reading-parquet-files-ram-cpu-optimization\" target=\"_blank\">https://www.kaggle.com/code/jamesmcguigan/reading-parquet-files-ram-cpu-optimization</a> </p>",
      "rawMarkdown": "The dataset for this competition is stored in parquet files (pandas + pyarrow).\n\nI wrote this notebook exploring how to performance optimize reading the parquet file format into memory, which was a infrastructure-level challenge for Bengali AI, which had oversized datasets.\n\nHopefully this will be of use to some people in this competition.\n\nhttps://www.kaggle.com/code/jamesmcguigan/reading-parquet-files-ram-cpu-optimization",
      "votes": null
    },
    {
      "id": "2119088",
      "postDate": "01/28/2023 13:45:55",
      "content": "<p>Thanks for your sharing..</p>",
      "rawMarkdown": "Thanks for your sharing..",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2119088,
      "author_name": "ibrahimkaratas",
      "author_url": "",
      "post_date": "01/28/2023 13:45:55",
      "content": "<p>Thanks for your sharing..</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2117886": "The dataset for this competition is stored in parquet files (pandas + pyarrow).\n\nI wrote this notebook exploring how to performance optimize reading the parquet file format into memory, which was a infrastructure-level challenge for Bengali AI, which had oversized datasets.\n\nHopefully this will be of use to some people in this competition.\n\nhttps://www.kaggle.com/code/jamesmcguigan/reading-parquet-files-ram-cpu-optimization",
    "2119088": "Thanks for your sharing.."
  },
  "source": "meta"
}