{
  "id": 538076,
  "title": "handle parquet files",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/538076",
  "author_name": "",
  "post_date": "2024-10-06T19:25:14.280672400Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello, <br>\nI can't handle parquet files, when i open them with pandas i get memory error. How do you it?</p>",
  "messages": [
    {
      "id": "3008589",
      "postDate": "10/06/2024 19:25:14",
      "content": "<p>Hello, <br>\nI can't handle parquet files, when i open them with pandas i get memory error. How do you it?</p>",
      "rawMarkdown": "Hello, \nI can't handle parquet files, when i open them with pandas i get memory error. How do you it?",
      "votes": null
    },
    {
      "id": "3008979",
      "postDate": "10/07/2024 11:11:45",
      "content": "<p>Hello. I have the same problem</p>",
      "rawMarkdown": "Hello. I have the same problem",
      "votes": null
    },
    {
      "id": "3009332",
      "postDate": "10/07/2024 18:23:20",
      "content": "<p>Hope this will Help you <a href=\"https://www.kaggle.com/elyoussfiazeddine\" target=\"_blank\">@elyoussfiazeddine</a> </p>\n<pre><code>def process:\n    df = pd.read)\n    df.drop('step', axis=, inplace=True)\n    return df.describe.values.reshape, filename.split()\n\ndef load -&gt; pd.DataFrame:\n    ids = os.listdir(dirname)\n\n       executor:\n        results = (tqdm(executor.map(lambda fname: process, ids), total=len(ids)))\n\n    stats, indexes = zip(*results)\n\n    df = pd.)])\n    df = indexes\n\n    return df\n</code></pre>\n<p>For More Detail you can see notebook also : <a href=\"https://www.kaggle.com/code/abdmental01/cmi-best-single-model\" target=\"_blank\">https://www.kaggle.com/code/abdmental01/cmi-best-single-model</a></p>",
      "rawMarkdown": "Hope this will Help you @elyoussfiazeddine \n\n```\ndef process_file(filename, dirname):\n    df = pd.read_parquet(os.path.join(dirname, filename, 'part-0.parquet'))\n    df.drop('step', axis=1, inplace=True)\n    return df.describe().values.reshape(-1), filename.split('=')[1]\n\ndef load_time_series(dirname) -> pd.DataFrame:\n    ids = os.listdir(dirname)\n    \n    with ThreadPoolExecutor() as executor:\n        results = list(tqdm(executor.map(lambda fname: process_file(fname, dirname), ids), total=len(ids)))\n    \n    stats, indexes = zip(*results)\n    \n    df = pd.DataFrame(stats, columns=[f\"Stat_{i}\" for i in range(len(stats[0]))])\n    df['id'] = indexes\n    \n    return df\n```\nFor More Detail you can see notebook also : https://www.kaggle.com/code/abdmental01/cmi-best-single-model",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3008979,
      "author_name": "oksanatemruk",
      "author_url": "",
      "post_date": "10/07/2024 11:11:45",
      "content": "<p>Hello. I have the same problem</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3009332,
      "author_name": "abdmental01",
      "author_url": "",
      "post_date": "10/07/2024 18:23:20",
      "content": "<p>Hope this will Help you <a href=\"https://www.kaggle.com/elyoussfiazeddine\" target=\"_blank\">@elyoussfiazeddine</a> </p>\n<pre><code>def process:\n    df = pd.read)\n    df.drop('step', axis=, inplace=True)\n    return df.describe.values.reshape, filename.split()\n\ndef load -&gt; pd.DataFrame:\n    ids = os.listdir(dirname)\n\n       executor:\n        results = (tqdm(executor.map(lambda fname: process, ids), total=len(ids)))\n\n    stats, indexes = zip(*results)\n\n    df = pd.)])\n    df = indexes\n\n    return df\n</code></pre>\n<p>For More Detail you can see notebook also : <a href=\"https://www.kaggle.com/code/abdmental01/cmi-best-single-model\" target=\"_blank\">https://www.kaggle.com/code/abdmental01/cmi-best-single-model</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3008589": "Hello, \nI can't handle parquet files, when i open them with pandas i get memory error. How do you it?",
    "3008979": "Hello. I have the same problem",
    "3009332": "Hope this will Help you @elyoussfiazeddine \n\n```\ndef process_file(filename, dirname):\n    df = pd.read_parquet(os.path.join(dirname, filename, 'part-0.parquet'))\n    df.drop('step', axis=1, inplace=True)\n    return df.describe().values.reshape(-1), filename.split('=')[1]\n\ndef load_time_series(dirname) -> pd.DataFrame:\n    ids = os.listdir(dirname)\n    \n    with ThreadPoolExecutor() as executor:\n        results = list(tqdm(executor.map(lambda fname: process_file(fname, dirname), ids), total=len(ids)))\n    \n    stats, indexes = zip(*results)\n    \n    df = pd.DataFrame(stats, columns=[f\"Stat_{i}\" for i in range(len(stats[0]))])\n    df['id'] = indexes\n    \n    return df\n```\nFor More Detail you can see notebook also : https://www.kaggle.com/code/abdmental01/cmi-best-single-model"
  },
  "source": "meta"
}