{
  "id": 380161,
  "title": "tips & tricks to deal with parquet files effectively ",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/380161",
  "author_name": "Rashmi Margani",
  "post_date": "2023-01-22T11:42:27.132000",
  "votes": 10,
  "comment_count": 5,
  "views": 0,
  "content": "<p><a href=\"https://jpweytjens.be/read-multiple-files-with-pandas-fast/\" target=\"_blank\">Read multiple (parquet) files with pandas fast</a> : </p>\n<pre><code>from functools import partial\n\nimport pandas as pd\nimport pyarrow as pa\nfrom tqdm.auto import tqdm\nfrom tqdm.contrib.concurrent import process_map\n\n\ndef _read_parquet(filename, columns=None):\n    \"\"\"\n    Wrapper to pass to a ProcessPoolExecutor to read parquet files as fast as possible. The PyArrow engine (v4.0.0) is faster than the fastparquet engine (v0.7.0) as it can read columns in parallel. Explicitly enable multithreaded column reading with `use_threads == true`.\n\n    Parameters\n    ----------\n\n    filename : str\n        Path of the parquet file to read.\n    columns : list, default=None\n        List of columns to read from the parquet file. If None, reads all columns.\n\n    Returns\n    -------\n    pandas Dataframe\n    \"\"\"\n\n    return pd.read_parquet(\n        filename, columns=columns, engine=\"pyarrow\", use_threads=True\n    )\n\n\ndef read_parquet(\n    files,\n    columns=None,\n    parallel=True,\n    n_concurrent_files=8,\n    n_concurrent_columns=4,\n    show_progress=True,\n    ignore_index=True,\n    chunksize=None,\n):\n    \"\"\"\n    Read a single parquet file or a list of parquet files and return a pandas DataFrame.\n\n    If `parallel==True`, it's on average 50% faster than `pd.read_parquet(..., engine=\"fastparquet\")`. Limited benchmarks indicate that the default values for `n_concurrent_files` and `n_concurrent_columns` are the fastest combination on a 32 core CPU. `n_concurrent_files` * `n_concurrent_columns` &lt;= the number of available cores.\n\n    Parameters\n    ----------\n\n    files : list or str\n        String with path or list of strings with paths of the parqiaet file(s) to be read.\n    columns : list, default=None\n        List of columns to read from the parquet file(s). If None, reads all columns.\n    parallel : bool, default=True\n        If True, reads both files and columns in parallel. If False, read the files serially while still reading the columns in parallel.\n    n_concurrent_files : int, default=8\n        Number of files to read in parallel.\n    n_concurrent_columns : int, default=4\n        Number of columns to read in parallel.\n    show_progress : bool, default=True\n        If True, shows a tqdm progress bar with the number of files that have already been read.\n    ignore_index : bool, default=True\n        If True, do not use the index values along the concatenation axis. The resulting axis will be labeled 0, ..., n-1. This is useful if you are concatenating objects where the concatention axis does not have meaningful indexing information.\n    chunksize : int, default=None\n        Number of files to pass as a single task to a single process. Values greater than 1 can improve performance if each task is expected to take a similar amount of time to complete and `len(files) &gt; n_concurrent_files`. If None, chunksize is set to `len(files) / n_concurrent_files` if `len(files) &gt; n_concurrent_files` else it's set to 1.\n\n    Returns\n    ------\n    pandas DataFrame\n    \"\"\"\n\n    # ensure files is a list when reading a single file\n    if isinstance(files, str):\n        files = [files]\n\n    # no need for more cpu's then files\n    if len(files) &lt; n_concurrent_files:\n        n_concurrent_files = len(files)\n\n    # no need for more workers than columns\n    if columns:\n        if len(columns) &lt; n_concurrent_columns:\n            n_concurrent_columns = len(columns)\n\n    # set number of threads used for reading the columns of each parquet files\n    pa.set_cpu_count(n_concurrent_columns)\n\n    # try to optimize the chunksize based on\n    # https://stackoverflow.com/questions/53751050/python-multiprocessing-understanding-logic-behind-chunksize\n    # this assumes each task takes roughly the same amount of time to complete\n    # i.e. each dataset is roughly the same size if there are only a few files\n    # to be read, i.e. ´len(files) &lt; n_concurrent_files´, give each cpu a single file to read\n    # when there are more files than cpu's give chunks of multiple files to each cpu\n    # this is in an attempt to minimize the overhead of assigning files after every completed file read\n    if (chunksize is None) and (len(files) &gt; n_concurrent_files):\n        chunksize, remainder = divmod(len(files), n_concurrent_files)\n        if remainder:\n            chunksize += 1\n    else:\n        chunksize = 1\n\n    if parallel is True:\n        _read_parquet_map = partial(_read_parquet, columns=columns)\n        dfs = process_map(\n            _read_parquet_map,\n            files,\n            max_workers=n_concurrent_files,\n            chunksize=chunksize,\n            disabled=not show_progress,\n        )\n\n    else:\n        dfs = [_read_parquet(file) for file in tqdm(files, disabled=not show_progress)]\n\n    # reduce the list of dataframes to a single dataframe\n    df = pd.concat(dfs, ignore_index=ignore_index)\n\n    return df\n</code></pre>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328054\" target=\"_blank\">How to reduce size</a> by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>\n<pre><code>read_parquet = pd.read_parquet(\"/kaggle/input/icecube-neutrinos-in-deep-ice/train/batch_10.parquet\")\n\nread_parquet.info()\n\n&lt;class 'pandas.core.frame.DataFrame'&gt;\nInt64Index: 33243258 entries, 29296372 to 32567683\nData columns (total 4 columns):\n #   Column     Dtype  \n---  ------     -----  \n 0   sensor_id  int16  \n 1   time       int64  \n 2   charge     float64\n 3   auxiliary  bool   \ndtypes: bool(1), float64(1), int16(1), int64(1)\nmemory usage: 856.0 MB\n\nread_parquet['sensor_id'] = read_parquet['sensor_id'].astype('int16')\nread_parquet['time'] = read_parquet['time'].astype('int16')\nread_parquet['charge'] = read_parquet['charge'].astype('float32')\n\nread_parquet.info()\n&lt;class 'pandas.core.frame.DataFrame'&gt;\nInt64Index: 33243258 entries, 29296372 to 32567683\nData columns (total 4 columns):\n #   Column     Dtype  \n---  ------     -----  \n 0   sensor_id  int16  \n 1   time       int16  \n 2   charge     float32\n 3   auxiliary  bool   \ndtypes: bool(1), float32(1), int16(2)\nmemory usage: 539.0 MB\n</code></pre>",
  "messages": [
    {
      "id": 2110740,
      "postDate": "2023-01-22T11:42:27.133Z",
      "content": "<p><a href=\"https://jpweytjens.be/read-multiple-files-with-pandas-fast/\" target=\"_blank\">Read multiple (parquet) files with pandas fast</a> : </p>\n<pre><code>from functools import partial\n\nimport pandas as pd\nimport pyarrow as pa\nfrom tqdm.auto import tqdm\nfrom tqdm.contrib.concurrent import process_map\n\n\ndef _read_parquet(filename, columns=None):\n    \"\"\"\n    Wrapper to pass to a ProcessPoolExecutor to read parquet files as fast as possible. The PyArrow engine (v4.0.0) is faster than the fastparquet engine (v0.7.0) as it can read columns in parallel. Explicitly enable multithreaded column reading with `use_threads == true`.\n\n    Parameters\n    ----------\n\n    filename : str\n        Path of the parquet file to read.\n    columns : list, default=None\n        List of columns to read from the parquet file. If None, reads all columns.\n\n    Returns\n    -------\n    pandas Dataframe\n    \"\"\"\n\n    return pd.read_parquet(\n        filename, columns=columns, engine=\"pyarrow\", use_threads=True\n    )\n\n\ndef read_parquet(\n    files,\n    columns=None,\n    parallel=True,\n    n_concurrent_files=8,\n    n_concurrent_columns=4,\n    show_progress=True,\n    ignore_index=True,\n    chunksize=None,\n):\n    \"\"\"\n    Read a single parquet file or a list of parquet files and return a pandas DataFrame.\n\n    If `parallel==True`, it's on average 50% faster than `pd.read_parquet(..., engine=\"fastparquet\")`. Limited benchmarks indicate that the default values for `n_concurrent_files` and `n_concurrent_columns` are the fastest combination on a 32 core CPU. `n_concurrent_files` * `n_concurrent_columns` &lt;= the number of available cores.\n\n    Parameters\n    ----------\n\n    files : list or str\n        String with path or list of strings with paths of the parqiaet file(s) to be read.\n    columns : list, default=None\n        List of columns to read from the parquet file(s). If None, reads all columns.\n    parallel : bool, default=True\n        If True, reads both files and columns in parallel. If False, read the files serially while still reading the columns in parallel.\n    n_concurrent_files : int, default=8\n        Number of files to read in parallel.\n    n_concurrent_columns : int, default=4\n        Number of columns to read in parallel.\n    show_progress : bool, default=True\n        If True, shows a tqdm progress bar with the number of files that have already been read.\n    ignore_index : bool, default=True\n        If True, do not use the index values along the concatenation axis. The resulting axis will be labeled 0, ..., n-1. This is useful if you are concatenating objects where the concatention axis does not have meaningful indexing information.\n    chunksize : int, default=None\n        Number of files to pass as a single task to a single process. Values greater than 1 can improve performance if each task is expected to take a similar amount of time to complete and `len(files) &gt; n_concurrent_files`. If None, chunksize is set to `len(files) / n_concurrent_files` if `len(files) &gt; n_concurrent_files` else it's set to 1.\n\n    Returns\n    ------\n    pandas DataFrame\n    \"\"\"\n\n    # ensure files is a list when reading a single file\n    if isinstance(files, str):\n        files = [files]\n\n    # no need for more cpu's then files\n    if len(files) &lt; n_concurrent_files:\n        n_concurrent_files = len(files)\n\n    # no need for more workers than columns\n    if columns:\n        if len(columns) &lt; n_concurrent_columns:\n            n_concurrent_columns = len(columns)\n\n    # set number of threads used for reading the columns of each parquet files\n    pa.set_cpu_count(n_concurrent_columns)\n\n    # try to optimize the chunksize based on\n    # https://stackoverflow.com/questions/53751050/python-multiprocessing-understanding-logic-behind-chunksize\n    # this assumes each task takes roughly the same amount of time to complete\n    # i.e. each dataset is roughly the same size if there are only a few files\n    # to be read, i.e. ´len(files) &lt; n_concurrent_files´, give each cpu a single file to read\n    # when there are more files than cpu's give chunks of multiple files to each cpu\n    # this is in an attempt to minimize the overhead of assigning files after every completed file read\n    if (chunksize is None) and (len(files) &gt; n_concurrent_files):\n        chunksize, remainder = divmod(len(files), n_concurrent_files)\n        if remainder:\n            chunksize += 1\n    else:\n        chunksize = 1\n\n    if parallel is True:\n        _read_parquet_map = partial(_read_parquet, columns=columns)\n        dfs = process_map(\n            _read_parquet_map,\n            files,\n            max_workers=n_concurrent_files,\n            chunksize=chunksize,\n            disabled=not show_progress,\n        )\n\n    else:\n        dfs = [_read_parquet(file) for file in tqdm(files, disabled=not show_progress)]\n\n    # reduce the list of dataframes to a single dataframe\n    df = pd.concat(dfs, ignore_index=ignore_index)\n\n    return df\n</code></pre>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328054\" target=\"_blank\">How to reduce size</a> by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>\n<pre><code>read_parquet = pd.read_parquet(\"/kaggle/input/icecube-neutrinos-in-deep-ice/train/batch_10.parquet\")\n\nread_parquet.info()\n\n&lt;class 'pandas.core.frame.DataFrame'&gt;\nInt64Index: 33243258 entries, 29296372 to 32567683\nData columns (total 4 columns):\n #   Column     Dtype  \n---  ------     -----  \n 0   sensor_id  int16  \n 1   time       int64  \n 2   charge     float64\n 3   auxiliary  bool   \ndtypes: bool(1), float64(1), int16(1), int64(1)\nmemory usage: 856.0 MB\n\nread_parquet['sensor_id'] = read_parquet['sensor_id'].astype('int16')\nread_parquet['time'] = read_parquet['time'].astype('int16')\nread_parquet['charge'] = read_parquet['charge'].astype('float32')\n\nread_parquet.info()\n&lt;class 'pandas.core.frame.DataFrame'&gt;\nInt64Index: 33243258 entries, 29296372 to 32567683\nData columns (total 4 columns):\n #   Column     Dtype  \n---  ------     -----  \n 0   sensor_id  int16  \n 1   time       int16  \n 2   charge     float32\n 3   auxiliary  bool   \ndtypes: bool(1), float32(1), int16(2)\nmemory usage: 539.0 MB\n</code></pre>",
      "rawMarkdown": "[Read multiple (parquet) files with pandas fast](https://jpweytjens.be/read-multiple-files-with-pandas-fast/) : \n\n```\nfrom functools import partial\n\nimport pandas as pd\nimport pyarrow as pa\nfrom tqdm.auto import tqdm\nfrom tqdm.contrib.concurrent import process_map\n\n\ndef _read_parquet(filename, columns=None):\n    \"\"\"\n    Wrapper to pass to a ProcessPoolExecutor to read parquet files as fast as possible. The PyArrow engine (v4.0.0) is faster than the fastparquet engine (v0.7.0) as it can read columns in parallel. Explicitly enable multithreaded column reading with `use_threads == true`.\n\n    Parameters\n    ----------\n\n    filename : str\n        Path of the parquet file to read.\n    columns : list, default=None\n        List of columns to read from the parquet file. If None, reads all columns.\n\n    Returns\n    -------\n    pandas Dataframe\n    \"\"\"\n\n    return pd.read_parquet(\n        filename, columns=columns, engine=\"pyarrow\", use_threads=True\n    )\n\n\ndef read_parquet(\n    files,\n    columns=None,\n    parallel=True,\n    n_concurrent_files=8,\n    n_concurrent_columns=4,\n    show_progress=True,\n    ignore_index=True,\n    chunksize=None,\n):\n    \"\"\"\n    Read a single parquet file or a list of parquet files and return a pandas DataFrame.\n\n    If `parallel==True`, it's on average 50% faster than `pd.read_parquet(..., engine=\"fastparquet\")`. Limited benchmarks indicate that the default values for `n_concurrent_files` and `n_concurrent_columns` are the fastest combination on a 32 core CPU. `n_concurrent_files` * `n_concurrent_columns` <= the number of available cores.\n\n    Parameters\n    ----------\n\n    files : list or str\n        String with path or list of strings with paths of the parqiaet file(s) to be read.\n    columns : list, default=None\n        List of columns to read from the parquet file(s). If None, reads all columns.\n    parallel : bool, default=True\n        If True, reads both files and columns in parallel. If False, read the files serially while still reading the columns in parallel.\n    n_concurrent_files : int, default=8\n        Number of files to read in parallel.\n    n_concurrent_columns : int, default=4\n        Number of columns to read in parallel.\n    show_progress : bool, default=True\n        If True, shows a tqdm progress bar with the number of files that have already been read.\n    ignore_index : bool, default=True\n        If True, do not use the index values along the concatenation axis. The resulting axis will be labeled 0, ..., n-1. This is useful if you are concatenating objects where the concatention axis does not have meaningful indexing information.\n    chunksize : int, default=None\n        Number of files to pass as a single task to a single process. Values greater than 1 can improve performance if each task is expected to take a similar amount of time to complete and `len(files) > n_concurrent_files`. If None, chunksize is set to `len(files) / n_concurrent_files` if `len(files) > n_concurrent_files` else it's set to 1.\n\n    Returns\n    ------\n    pandas DataFrame\n    \"\"\"\n\n    # ensure files is a list when reading a single file\n    if isinstance(files, str):\n        files = [files]\n\n    # no need for more cpu's then files\n    if len(files) < n_concurrent_files:\n        n_concurrent_files = len(files)\n\n    # no need for more workers than columns\n    if columns:\n        if len(columns) < n_concurrent_columns:\n            n_concurrent_columns = len(columns)\n\n    # set number of threads used for reading the columns of each parquet files\n    pa.set_cpu_count(n_concurrent_columns)\n\n    # try to optimize the chunksize based on\n    # https://stackoverflow.com/questions/53751050/python-multiprocessing-understanding-logic-behind-chunksize\n    # this assumes each task takes roughly the same amount of time to complete\n    # i.e. each dataset is roughly the same size if there are only a few files\n    # to be read, i.e. ´len(files) < n_concurrent_files´, give each cpu a single file to read\n    # when there are more files than cpu's give chunks of multiple files to each cpu\n    # this is in an attempt to minimize the overhead of assigning files after every completed file read\n    if (chunksize is None) and (len(files) > n_concurrent_files):\n        chunksize, remainder = divmod(len(files), n_concurrent_files)\n        if remainder:\n            chunksize += 1\n    else:\n        chunksize = 1\n\n    if parallel is True:\n        _read_parquet_map = partial(_read_parquet, columns=columns)\n        dfs = process_map(\n            _read_parquet_map,\n            files,\n            max_workers=n_concurrent_files,\n            chunksize=chunksize,\n            disabled=not show_progress,\n        )\n\n    else:\n        dfs = [_read_parquet(file) for file in tqdm(files, disabled=not show_progress)]\n\n    # reduce the list of dataframes to a single dataframe\n    df = pd.concat(dfs, ignore_index=ignore_index)\n\n    return df\n\n```\n\n[How to reduce size](https://www.kaggle.com/competitions/amex-default-prediction/discussion/328054) by @cdeotte \n\n```\nread_parquet = pd.read_parquet(\"/kaggle/input/icecube-neutrinos-in-deep-ice/train/batch_10.parquet\")\n\nread_parquet.info()\n\n<class 'pandas.core.frame.DataFrame'>\nInt64Index: 33243258 entries, 29296372 to 32567683\nData columns (total 4 columns):\n #   Column     Dtype  \n---  ------     -----  \n 0   sensor_id  int16  \n 1   time       int64  \n 2   charge     float64\n 3   auxiliary  bool   \ndtypes: bool(1), float64(1), int16(1), int64(1)\nmemory usage: 856.0 MB\n\nread_parquet['sensor_id'] = read_parquet['sensor_id'].astype('int16')\nread_parquet['time'] = read_parquet['time'].astype('int16')\nread_parquet['charge'] = read_parquet['charge'].astype('float32')\n\nread_parquet.info()\n<class 'pandas.core.frame.DataFrame'>\nInt64Index: 33243258 entries, 29296372 to 32567683\nData columns (total 4 columns):\n #   Column     Dtype  \n---  ------     -----  \n 0   sensor_id  int16  \n 1   time       int16  \n 2   charge     float32\n 3   auxiliary  bool   \ndtypes: bool(1), float32(1), int16(2)\nmemory usage: 539.0 MB\n\n```",
      "votes": 10
    },
    {
      "id": 2113400,
      "postDate": "2023-01-24T09:11:31.167Z",
      "content": "<p>While polar bears are only found near the North pole, the <a href=\"https://pola-rs.github.io/polars-book/user-guide/introduction.html\" target=\"_blank\">Polars</a> package might also be a great fit for this South Pole dataset :)</p>",
      "rawMarkdown": "While polar bears are only found near the North pole, the [Polars](https://pola-rs.github.io/polars-book/user-guide/introduction.html) package might also be a great fit for this South Pole dataset :)",
      "votes": 1
    },
    {
      "id": 2112665,
      "postDate": "2023-01-23T18:36:09.553Z",
      "content": "<p>If you really want to read multiple files at a time I would encourage you to consider existing libraries like Dask: <a href=\"https://docs.dask.org/en/stable/dataframe-parquet.html\" target=\"_blank\">https://docs.dask.org/en/stable/dataframe-parquet.html</a></p>",
      "rawMarkdown": "If you really want to read multiple files at a time I would encourage you to consider existing libraries like Dask: https://docs.dask.org/en/stable/dataframe-parquet.html",
      "replies": [
        {
          "id": 2113304,
          "postDate": "2023-01-24T07:59:00.020Z",
          "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>, I am also looking into Dask as well. Thanks for the mention.</p>",
          "rawMarkdown": "@sohier, I am also looking into Dask as well. Thanks for the mention."
        }
      ]
    },
    {
      "id": 2111561,
      "postDate": "2023-01-23T01:30:15.630Z",
      "content": "<p>Maybe include them in a notebook, it will be much easier to understand this way.</p>\n<p>The Devastator.</p>",
      "rawMarkdown": "Maybe include them in a notebook, it will be much easier to understand this way.\n\nThe Devastator.\n",
      "replies": [
        {
          "id": 2111799,
          "postDate": "2023-01-23T07:34:08.820Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2113400,
      "author_name": "datasaurus",
      "author_url": "",
      "post_date": "2023-01-24T09:11:31.167000",
      "content": "<p>While polar bears are only found near the North pole, the <a href=\"https://pola-rs.github.io/polars-book/user-guide/introduction.html\" target=\"_blank\">Polars</a> package might also be a great fit for this South Pole dataset :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2112665,
      "author_name": "Sohier Dane",
      "author_url": "",
      "post_date": "2023-01-23T18:36:09.553000",
      "content": "<p>If you really want to read multiple files at a time I would encourage you to consider existing libraries like Dask: <a href=\"https://docs.dask.org/en/stable/dataframe-parquet.html\" target=\"_blank\">https://docs.dask.org/en/stable/dataframe-parquet.html</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 2113304,
          "author_name": "Rashmi Margani",
          "author_url": "",
          "post_date": "2023-01-24T07:59:00.020000",
          "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>, I am also looking into Dask as well. Thanks for the mention.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2111561,
      "author_name": "The Devastator",
      "author_url": "",
      "post_date": "2023-01-23T01:30:15.630000",
      "content": "<p>Maybe include them in a notebook, it will be much easier to understand this way.</p>\n<p>The Devastator.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2111799,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-01-23T07:34:08.820000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2110740": "[Read multiple (parquet) files with pandas fast](https://jpweytjens.be/read-multiple-files-with-pandas-fast/) : \n\n```\nfrom functools import partial\n\nimport pandas as pd\nimport pyarrow as pa\nfrom tqdm.auto import tqdm\nfrom tqdm.contrib.concurrent import process_map\n\n\ndef _read_parquet(filename, columns=None):\n    \"\"\"\n    Wrapper to pass to a ProcessPoolExecutor to read parquet files as fast as possible. The PyArrow engine (v4.0.0) is faster than the fastparquet engine (v0.7.0) as it can read columns in parallel. Explicitly enable multithreaded column reading with `use_threads == true`.\n\n    Parameters\n    ----------\n\n    filename : str\n        Path of the parquet file to read.\n    columns : list, default=None\n        List of columns to read from the parquet file. If None, reads all columns.\n\n    Returns\n    -------\n    pandas Dataframe\n    \"\"\"\n\n    return pd.read_parquet(\n        filename, columns=columns, engine=\"pyarrow\", use_threads=True\n    )\n\n\ndef read_parquet(\n    files,\n    columns=None,\n    parallel=True,\n    n_concurrent_files=8,\n    n_concurrent_columns=4,\n    show_progress=True,\n    ignore_index=True,\n    chunksize=None,\n):\n    \"\"\"\n    Read a single parquet file or a list of parquet files and return a pandas DataFrame.\n\n    If `parallel==True`, it's on average 50% faster than `pd.read_parquet(..., engine=\"fastparquet\")`. Limited benchmarks indicate that the default values for `n_concurrent_files` and `n_concurrent_columns` are the fastest combination on a 32 core CPU. `n_concurrent_files` * `n_concurrent_columns` <= the number of available cores.\n\n    Parameters\n    ----------\n\n    files : list or str\n        String with path or list of strings with paths of the parqiaet file(s) to be read.\n    columns : list, default=None\n        List of columns to read from the parquet file(s). If None, reads all columns.\n    parallel : bool, default=True\n        If True, reads both files and columns in parallel. If False, read the files serially while still reading the columns in parallel.\n    n_concurrent_files : int, default=8\n        Number of files to read in parallel.\n    n_concurrent_columns : int, default=4\n        Number of columns to read in parallel.\n    show_progress : bool, default=True\n        If True, shows a tqdm progress bar with the number of files that have already been read.\n    ignore_index : bool, default=True\n        If True, do not use the index values along the concatenation axis. The resulting axis will be labeled 0, ..., n-1. This is useful if you are concatenating objects where the concatention axis does not have meaningful indexing information.\n    chunksize : int, default=None\n        Number of files to pass as a single task to a single process. Values greater than 1 can improve performance if each task is expected to take a similar amount of time to complete and `len(files) > n_concurrent_files`. If None, chunksize is set to `len(files) / n_concurrent_files` if `len(files) > n_concurrent_files` else it's set to 1.\n\n    Returns\n    ------\n    pandas DataFrame\n    \"\"\"\n\n    # ensure files is a list when reading a single file\n    if isinstance(files, str):\n        files = [files]\n\n    # no need for more cpu's then files\n    if len(files) < n_concurrent_files:\n        n_concurrent_files = len(files)\n\n    # no need for more workers than columns\n    if columns:\n        if len(columns) < n_concurrent_columns:\n            n_concurrent_columns = len(columns)\n\n    # set number of threads used for reading the columns of each parquet files\n    pa.set_cpu_count(n_concurrent_columns)\n\n    # try to optimize the chunksize based on\n    # https://stackoverflow.com/questions/53751050/python-multiprocessing-understanding-logic-behind-chunksize\n    # this assumes each task takes roughly the same amount of time to complete\n    # i.e. each dataset is roughly the same size if there are only a few files\n    # to be read, i.e. ´len(files) < n_concurrent_files´, give each cpu a single file to read\n    # when there are more files than cpu's give chunks of multiple files to each cpu\n    # this is in an attempt to minimize the overhead of assigning files after every completed file read\n    if (chunksize is None) and (len(files) > n_concurrent_files):\n        chunksize, remainder = divmod(len(files), n_concurrent_files)\n        if remainder:\n            chunksize += 1\n    else:\n        chunksize = 1\n\n    if parallel is True:\n        _read_parquet_map = partial(_read_parquet, columns=columns)\n        dfs = process_map(\n            _read_parquet_map,\n            files,\n            max_workers=n_concurrent_files,\n            chunksize=chunksize,\n            disabled=not show_progress,\n        )\n\n    else:\n        dfs = [_read_parquet(file) for file in tqdm(files, disabled=not show_progress)]\n\n    # reduce the list of dataframes to a single dataframe\n    df = pd.concat(dfs, ignore_index=ignore_index)\n\n    return df\n\n```\n\n[How to reduce size](https://www.kaggle.com/competitions/amex-default-prediction/discussion/328054) by @cdeotte \n\n```\nread_parquet = pd.read_parquet(\"/kaggle/input/icecube-neutrinos-in-deep-ice/train/batch_10.parquet\")\n\nread_parquet.info()\n\n<class 'pandas.core.frame.DataFrame'>\nInt64Index: 33243258 entries, 29296372 to 32567683\nData columns (total 4 columns):\n #   Column     Dtype  \n---  ------     -----  \n 0   sensor_id  int16  \n 1   time       int64  \n 2   charge     float64\n 3   auxiliary  bool   \ndtypes: bool(1), float64(1), int16(1), int64(1)\nmemory usage: 856.0 MB\n\nread_parquet['sensor_id'] = read_parquet['sensor_id'].astype('int16')\nread_parquet['time'] = read_parquet['time'].astype('int16')\nread_parquet['charge'] = read_parquet['charge'].astype('float32')\n\nread_parquet.info()\n<class 'pandas.core.frame.DataFrame'>\nInt64Index: 33243258 entries, 29296372 to 32567683\nData columns (total 4 columns):\n #   Column     Dtype  \n---  ------     -----  \n 0   sensor_id  int16  \n 1   time       int16  \n 2   charge     float32\n 3   auxiliary  bool   \ndtypes: bool(1), float32(1), int16(2)\nmemory usage: 539.0 MB\n\n```",
    "2113400": "While polar bears are only found near the North pole, the [Polars](https://pola-rs.github.io/polars-book/user-guide/introduction.html) package might also be a great fit for this South Pole dataset :)",
    "2112665": "If you really want to read multiple files at a time I would encourage you to consider existing libraries like Dask: https://docs.dask.org/en/stable/dataframe-parquet.html",
    "2111561": "Maybe include them in a notebook, it will be much easier to understand this way.\n\nThe Devastator.\n"
  }
}