{
  "id": 545583,
  "title": "What is series_test.parquet?",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/545583",
  "author_name": "",
  "post_date": "2024-11-11T03:55:28.763170100Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hey this is my first competition on kaggle and figuring out my way here… It might be silly but what is \"parquet\" and what does it actually represent? I've got 2 folders one with train and other with test so… I am not sure what to do with them?</p>",
  "messages": [
    {
      "id": "3041980",
      "postDate": "11/11/2024 03:55:28",
      "content": "<p>Hey this is my first competition on kaggle and figuring out my way here… It might be silly but what is \"parquet\" and what does it actually represent? I've got 2 folders one with train and other with test so… I am not sure what to do with them?</p>",
      "rawMarkdown": "Hey this is my first competition on kaggle and figuring out my way here... It might be silly but what is \"parquet\" and what does it actually represent? I've got 2 folders one with train and other with test so... I am not sure what to do with them?",
      "votes": null
    },
    {
      "id": "3042367",
      "postDate": "11/11/2024 13:18:54",
      "content": "<p>Hi, I also beginner of kaggle and recently, I tried using parquet, so I share the method of using parquet. <br>\nParquet is a columnar, open-source file format designed for efficient storage and analysis of large datasets. It's particularly popular in big data environments like data warehouses and data lakes.</p>\n<p>As an example,the code below allows you to extract statistical information from a parquet dataset.</p>\n<pre><code> numpy  np \n pandas  pd \n os\n concurrent.futures  ThreadPoolExecutor\n tqdm  tqdm\n\n ():\n    df = pd.read_parquet(os.path.join(dirname, filename, )) \n\n    df.drop(, axis=, inplace=)  \n\n     df.describe().values.reshape(-), filename.split()[] \n\n\n () -&gt; pd.DataFrame:\n    ids = os.listdir(dirname)  \n\n     ThreadPoolExecutor()  executor:\n        results = (tqdm(executor.( fname: process_file(fname, dirname), ids), total=(ids)))\n\n    stats, indexes = (*results)  \n\n    df = pd.DataFrame(stats, columns=[  i  ((stats[]))])  \n\n    df[] = indexes  \n\n     df\n\ntrain_ts = load_time_series()\n</code></pre>",
      "rawMarkdown": "Hi, I also beginner of kaggle and recently, I tried using parquet, so I share the method of using parquet. \nParquet is a columnar, open-source file format designed for efficient storage and analysis of large datasets. It's particularly popular in big data environments like data warehouses and data lakes.\n\nAs an example,the code below allows you to extract statistical information from a parquet dataset.\n\n\n```\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport os\nfrom concurrent.futures import ThreadPoolExecutor\nfrom tqdm import tqdm\n\ndef process_file(filename, dirname):\n    df = pd.read_parquet(os.path.join(dirname, filename, 'part-0.parquet')) \n    \n    df.drop('step', axis=1, inplace=True)  \n    \n    return df.describe().values.reshape(-1), filename.split('=')[1] \n\n\ndef load_time_series(dirname) -> pd.DataFrame:\n    ids = os.listdir(dirname)  \n    \n    with ThreadPoolExecutor() as executor:\n        results = list(tqdm(executor.map(lambda fname: process_file(fname, dirname), ids), total=len(ids)))\n    \n    stats, indexes = zip(*results)  \n    \n    df = pd.DataFrame(stats, columns=[f\"stat_{i}\" for i in range(len(stats[0]))])  \n    \n    df['id'] = indexes  \n    \n    return df\n\ntrain_ts = load_time_series(\"/kaggle/input/child-mind-institute-problematic-internet-use/series_train.parquet\")\n\n```",
      "votes": null
    },
    {
      "id": "3044979",
      "postDate": "11/14/2024 03:39:41",
      "content": "<p>parquet is a format of file to store data, just like <code>csv</code>, there's many open available code to read parquet file, example</p>\n<pre><code>def (filename, dirname):\n    df = pd.(os.path.(dirname, filename, ))\n    df.(, axis=, inplace=True)\n    return df.().values.(-), filename.()[]\n\n\ndef (dirname) -&gt; pd.DataFrame:\n    ids = os.(dirname)\n\n    with () as executor:\n        results = ((executor.(lambda fname: (fname, dirname), ids), total=(ids)))\n\n    stats, indexes = (*results)\n\n    df = pd.(stats, columns=[f for i in ((stats[]))])\n    df[] = indexes\n    return df\n</code></pre>",
      "rawMarkdown": "parquet is a format of file to store data, just like `csv`, there's many open available code to read parquet file, example\n\n```\ndef process_file(filename, dirname):\n    df = pd.read_parquet(os.path.join(dirname, filename, 'part-0.parquet'))\n    df.drop('step', axis=1, inplace=True)\n    return df.describe().values.reshape(-1), filename.split('=')[1]\n\n\ndef load_time_series(dirname) -> pd.DataFrame:\n    ids = os.listdir(dirname)\n    \n    with ThreadPoolExecutor() as executor:\n        results = list(tqdm(executor.map(lambda fname: process_file(fname, dirname), ids), total=len(ids)))\n    \n    stats, indexes = zip(*results)\n    \n    df = pd.DataFrame(stats, columns=[f\"stat_{i}\" for i in range(len(stats[0]))])\n    df['id'] = indexes\n    return df\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3042367,
      "author_name": "wagakana",
      "author_url": "",
      "post_date": "11/11/2024 13:18:54",
      "content": "<p>Hi, I also beginner of kaggle and recently, I tried using parquet, so I share the method of using parquet. <br>\nParquet is a columnar, open-source file format designed for efficient storage and analysis of large datasets. It's particularly popular in big data environments like data warehouses and data lakes.</p>\n<p>As an example,the code below allows you to extract statistical information from a parquet dataset.</p>\n<pre><code> numpy  np \n pandas  pd \n os\n concurrent.futures  ThreadPoolExecutor\n tqdm  tqdm\n\n ():\n    df = pd.read_parquet(os.path.join(dirname, filename, )) \n\n    df.drop(, axis=, inplace=)  \n\n     df.describe().values.reshape(-), filename.split()[] \n\n\n () -&gt; pd.DataFrame:\n    ids = os.listdir(dirname)  \n\n     ThreadPoolExecutor()  executor:\n        results = (tqdm(executor.( fname: process_file(fname, dirname), ids), total=(ids)))\n\n    stats, indexes = (*results)  \n\n    df = pd.DataFrame(stats, columns=[  i  ((stats[]))])  \n\n    df[] = indexes  \n\n     df\n\ntrain_ts = load_time_series()\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3044979,
      "author_name": "zhangyue199",
      "author_url": "",
      "post_date": "11/14/2024 03:39:41",
      "content": "<p>parquet is a format of file to store data, just like <code>csv</code>, there's many open available code to read parquet file, example</p>\n<pre><code>def (filename, dirname):\n    df = pd.(os.path.(dirname, filename, ))\n    df.(, axis=, inplace=True)\n    return df.().values.(-), filename.()[]\n\n\ndef (dirname) -&gt; pd.DataFrame:\n    ids = os.(dirname)\n\n    with () as executor:\n        results = ((executor.(lambda fname: (fname, dirname), ids), total=(ids)))\n\n    stats, indexes = (*results)\n\n    df = pd.(stats, columns=[f for i in ((stats[]))])\n    df[] = indexes\n    return df\n</code></pre>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3041980": "Hey this is my first competition on kaggle and figuring out my way here... It might be silly but what is \"parquet\" and what does it actually represent? I've got 2 folders one with train and other with test so... I am not sure what to do with them?",
    "3042367": "Hi, I also beginner of kaggle and recently, I tried using parquet, so I share the method of using parquet. \nParquet is a columnar, open-source file format designed for efficient storage and analysis of large datasets. It's particularly popular in big data environments like data warehouses and data lakes.\n\nAs an example,the code below allows you to extract statistical information from a parquet dataset.\n\n\n```\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport os\nfrom concurrent.futures import ThreadPoolExecutor\nfrom tqdm import tqdm\n\ndef process_file(filename, dirname):\n    df = pd.read_parquet(os.path.join(dirname, filename, 'part-0.parquet')) \n    \n    df.drop('step', axis=1, inplace=True)  \n    \n    return df.describe().values.reshape(-1), filename.split('=')[1] \n\n\ndef load_time_series(dirname) -> pd.DataFrame:\n    ids = os.listdir(dirname)  \n    \n    with ThreadPoolExecutor() as executor:\n        results = list(tqdm(executor.map(lambda fname: process_file(fname, dirname), ids), total=len(ids)))\n    \n    stats, indexes = zip(*results)  \n    \n    df = pd.DataFrame(stats, columns=[f\"stat_{i}\" for i in range(len(stats[0]))])  \n    \n    df['id'] = indexes  \n    \n    return df\n\ntrain_ts = load_time_series(\"/kaggle/input/child-mind-institute-problematic-internet-use/series_train.parquet\")\n\n```",
    "3044979": "parquet is a format of file to store data, just like `csv`, there's many open available code to read parquet file, example\n\n```\ndef process_file(filename, dirname):\n    df = pd.read_parquet(os.path.join(dirname, filename, 'part-0.parquet'))\n    df.drop('step', axis=1, inplace=True)\n    return df.describe().values.reshape(-1), filename.split('=')[1]\n\n\ndef load_time_series(dirname) -> pd.DataFrame:\n    ids = os.listdir(dirname)\n    \n    with ThreadPoolExecutor() as executor:\n        results = list(tqdm(executor.map(lambda fname: process_file(fname, dirname), ids), total=len(ids)))\n    \n    stats, indexes = zip(*results)\n    \n    df = pd.DataFrame(stats, columns=[f\"stat_{i}\" for i in range(len(stats[0]))])\n    df['id'] = indexes\n    return df\n```"
  },
  "source": "meta"
}