{
  "id": 388651,
  "title": "How to parallelize polars partition_by result dataframes processing?",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/388651",
  "author_name": "",
  "post_date": "2023-02-18T19:11:49.831196600Z",
  "votes": 4,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Is it possible to parallelize the processing of a list of dataframes received from a polars partition_by? Taking into account the fact that polars already parallelizes calculations, but as I understand it in this context, polars will not parallelize calculations between processing individual dataframes from the list. I would like to speed up dataframe processing.</p>\n<pre><code>def processing_df(df):\n    ...\n    ...\n    return result # np.array for example\n\ndf_gr = df_sourse.partition_by(groups=\"group_col\", maintain_order=True)\n\nresults = []\nfor df in df_gr:\n    res = processing_df(df)\n    results.append(res)\n</code></pre>\n<p>P.S. I tried use <a href=\"https://www.machinelearningplus.com/python/parallel-processing-python/\" target=\"_blank\">this approaches</a>, but processes were not running.</p>",
  "messages": [
    {
      "id": "2149895",
      "postDate": "02/18/2023 19:11:49",
      "content": "<p>Is it possible to parallelize the processing of a list of dataframes received from a polars partition_by? Taking into account the fact that polars already parallelizes calculations, but as I understand it in this context, polars will not parallelize calculations between processing individual dataframes from the list. I would like to speed up dataframe processing.</p>\n<pre><code>def processing_df(df):\n    ...\n    ...\n    return result # np.array for example\n\ndf_gr = df_sourse.partition_by(groups=\"group_col\", maintain_order=True)\n\nresults = []\nfor df in df_gr:\n    res = processing_df(df)\n    results.append(res)\n</code></pre>\n<p>P.S. I tried use <a href=\"https://www.machinelearningplus.com/python/parallel-processing-python/\" target=\"_blank\">this approaches</a>, but processes were not running.</p>",
      "rawMarkdown": "Is it possible to parallelize the processing of a list of dataframes received from a polars partition_by? Taking into account the fact that polars already parallelizes calculations, but as I understand it in this context, polars will not parallelize calculations between processing individual dataframes from the list. I would like to speed up dataframe processing.\n\n```\ndef processing_df(df):\n    ...\n    ...\n    return result # np.array for example\n\ndf_gr = df_sourse.partition_by(groups=\"group_col\", maintain_order=True)\n\nresults = []\nfor df in df_gr:\n    res = processing_df(df)\n    results.append(res)\n```\n\nP.S. I tried use [this approaches](https://www.machinelearningplus.com/python/parallel-processing-python/), but processes were not running.",
      "votes": null
    },
    {
      "id": "2149899",
      "postDate": "02/18/2023 19:14:24",
      "content": "<p>It is impossible to cope only with aggregation functions, since I use not only them</p>",
      "rawMarkdown": "It is impossible to cope only with aggregation functions, since I use not only them",
      "votes": null
    },
    {
      "id": "2149961",
      "postDate": "02/18/2023 20:26:52",
      "content": "<p>To parallelize the processing of multiple partitioned dataframes using Polars partition_by function, you can make use of Python's built-in multiprocessing module.</p>\n<p>Here's a high-level overview of the approach:</p>\n<p>Create a list of the partitioned dataframes using the partition_by function.<br>\nSplit the list of dataframes into smaller chunks, based on the number of CPU cores available or some other factor.<br>\nUse the multiprocessing.Pool class to create a pool of worker processes.<br>\nUse the Pool.map method to apply a function to each chunk of dataframes in parallel.<br>\nFinally, use the concat method to combine the results from the processed dataframes into a single dataframe.</p>\n<p>Here's some sample code that demonstrates this approach:</p>\n<pre><code> polars  pl\n multiprocessing  mp\n\n\ndf = pl.DataFrame({\n    : [, , , , , ],\n    : [, , , , , ]\n})\n\n\npartitions = df.partition_by()\n\n\n ():\n    \n    result = partitions.select(, ).groupby().()\n     result\n\n\nchunk_size = (partitions) // mp.cpu_count()\npartition_chunks = [partitions[i:i+chunk_size]  i  (, (partitions), chunk_size)]\n\n\npool = mp.Pool()\n\n\nresults = pool.(process_partitions, partition_chunks)\n\n\nfinal_result = pl.concat(results)\n</code></pre>\n<p>In this example, the process_partitions function is applied to each chunk of partitions in parallel using the Pool.map method. The resulting dataframes are concatenated into a single dataframe using the concat method.</p>",
      "rawMarkdown": "To parallelize the processing of multiple partitioned dataframes using Polars partition_by function, you can make use of Python's built-in multiprocessing module.\n\nHere's a high-level overview of the approach:\n\nCreate a list of the partitioned dataframes using the partition_by function.\nSplit the list of dataframes into smaller chunks, based on the number of CPU cores available or some other factor.\nUse the multiprocessing.Pool class to create a pool of worker processes.\nUse the Pool.map method to apply a function to each chunk of dataframes in parallel.\nFinally, use the concat method to combine the results from the processed dataframes into a single dataframe.\n\nHere's some sample code that demonstrates this approach:\n```python\n\nimport polars as pl\nimport multiprocessing as mp\n\n# create a DataFrame\ndf = pl.DataFrame({\n    'foo': ['A', 'A', 'B', 'B', 'C', 'C'],\n    'bar': [1, 2, 3, 4, 5, 6]\n})\n\n# partition the DataFrame by 'foo'\npartitions = df.partition_by('foo')\n\n# define a function to process a chunk of partitions\ndef process_partitions(partitions):\n    # do some processing on the partitions\n    result = partitions.select('foo', 'bar').groupby('foo').sum()\n    return result\n\n# split the list of partitions into chunks\nchunk_size = len(partitions) // mp.cpu_count()\npartition_chunks = [partitions[i:i+chunk_size] for i in range(0, len(partitions), chunk_size)]\n\n# create a pool of worker processes\npool = mp.Pool()\n\n# apply the function to each chunk of partitions in parallel\nresults = pool.map(process_partitions, partition_chunks)\n\n# concatenate the results into a single dataframe\nfinal_result = pl.concat(results)\n```\n\n\nIn this example, the process_partitions function is applied to each chunk of partitions in parallel using the Pool.map method. The resulting dataframes are concatenated into a single dataframe using the concat method.",
      "votes": null
    },
    {
      "id": "2150346",
      "postDate": "02/19/2023 07:09:55",
      "content": "<p>Thanks you very much for the answer! I tried to use this method, but the functions were not executed. I now tried the method with joblib Parallel, delayed - this way it worked</p>\n<pre><code>from joblib import Parallel, delayed\nimport polars as pl\n\ndf = pl.DataFrame({\n    'foo': ['A', 'A', 'B', 'B', 'C', 'C'],\n    'bar': [1, 2, 3, 4, 5, 6]\n})\n\n# partition the DataFrame by 'foo'\npartitions = df.partition_by('foo')\n\n# define a function to process a chunk of partitions\ndef process_partitions(partitions):\n    # do some processing on the partitions\n    result = [part.select('foo', 'bar').groupby('foo').sum() for part in partitions]\n    return result\n\n# split the list of partitions into chunks\nchunk_size = len(partitions) // mp.cpu_count()\npartition_chunks = [partitions[i:i+chunk_size] for i in range(0, len(partitions), chunk_size)]\n\nres = Parallel(\n    n_jobs=5\n)(\n    delayed(process_partitions)(chunk) for chink in partition_chunks\n)\n</code></pre>",
      "rawMarkdown": "Thanks you very much for the answer! I tried to use this method, but the functions were not executed. I now tried the method with joblib Parallel, delayed - this way it worked\n\n```\nfrom joblib import Parallel, delayed\nimport polars as pl\n\ndf = pl.DataFrame({\n    'foo': ['A', 'A', 'B', 'B', 'C', 'C'],\n    'bar': [1, 2, 3, 4, 5, 6]\n})\n\n# partition the DataFrame by 'foo'\npartitions = df.partition_by('foo')\n\n# define a function to process a chunk of partitions\ndef process_partitions(partitions):\n    # do some processing on the partitions\n    result = [part.select('foo', 'bar').groupby('foo').sum() for part in partitions]\n    return result\n\n# split the list of partitions into chunks\nchunk_size = len(partitions) // mp.cpu_count()\npartition_chunks = [partitions[i:i+chunk_size] for i in range(0, len(partitions), chunk_size)]\n\nres = Parallel(\n    n_jobs=5\n)(\n    delayed(process_partitions)(chunk) for chink in partition_chunks\n)\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2149899,
      "author_name": "dmdgik",
      "author_url": "",
      "post_date": "02/18/2023 19:14:24",
      "content": "<p>It is impossible to cope only with aggregation functions, since I use not only them</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2149961,
      "author_name": "soheiltehranipour",
      "author_url": "",
      "post_date": "02/18/2023 20:26:52",
      "content": "<p>To parallelize the processing of multiple partitioned dataframes using Polars partition_by function, you can make use of Python's built-in multiprocessing module.</p>\n<p>Here's a high-level overview of the approach:</p>\n<p>Create a list of the partitioned dataframes using the partition_by function.<br>\nSplit the list of dataframes into smaller chunks, based on the number of CPU cores available or some other factor.<br>\nUse the multiprocessing.Pool class to create a pool of worker processes.<br>\nUse the Pool.map method to apply a function to each chunk of dataframes in parallel.<br>\nFinally, use the concat method to combine the results from the processed dataframes into a single dataframe.</p>\n<p>Here's some sample code that demonstrates this approach:</p>\n<pre><code> polars  pl\n multiprocessing  mp\n\n\ndf = pl.DataFrame({\n    : [, , , , , ],\n    : [, , , , , ]\n})\n\n\npartitions = df.partition_by()\n\n\n ():\n    \n    result = partitions.select(, ).groupby().()\n     result\n\n\nchunk_size = (partitions) // mp.cpu_count()\npartition_chunks = [partitions[i:i+chunk_size]  i  (, (partitions), chunk_size)]\n\n\npool = mp.Pool()\n\n\nresults = pool.(process_partitions, partition_chunks)\n\n\nfinal_result = pl.concat(results)\n</code></pre>\n<p>In this example, the process_partitions function is applied to each chunk of partitions in parallel using the Pool.map method. The resulting dataframes are concatenated into a single dataframe using the concat method.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2150346,
          "author_name": "dmdgik",
          "author_url": "",
          "post_date": "02/19/2023 07:09:55",
          "content": "<p>Thanks you very much for the answer! I tried to use this method, but the functions were not executed. I now tried the method with joblib Parallel, delayed - this way it worked</p>\n<pre><code>from joblib import Parallel, delayed\nimport polars as pl\n\ndf = pl.DataFrame({\n    'foo': ['A', 'A', 'B', 'B', 'C', 'C'],\n    'bar': [1, 2, 3, 4, 5, 6]\n})\n\n# partition the DataFrame by 'foo'\npartitions = df.partition_by('foo')\n\n# define a function to process a chunk of partitions\ndef process_partitions(partitions):\n    # do some processing on the partitions\n    result = [part.select('foo', 'bar').groupby('foo').sum() for part in partitions]\n    return result\n\n# split the list of partitions into chunks\nchunk_size = len(partitions) // mp.cpu_count()\npartition_chunks = [partitions[i:i+chunk_size] for i in range(0, len(partitions), chunk_size)]\n\nres = Parallel(\n    n_jobs=5\n)(\n    delayed(process_partitions)(chunk) for chink in partition_chunks\n)\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2149895": "Is it possible to parallelize the processing of a list of dataframes received from a polars partition_by? Taking into account the fact that polars already parallelizes calculations, but as I understand it in this context, polars will not parallelize calculations between processing individual dataframes from the list. I would like to speed up dataframe processing.\n\n```\ndef processing_df(df):\n    ...\n    ...\n    return result # np.array for example\n\ndf_gr = df_sourse.partition_by(groups=\"group_col\", maintain_order=True)\n\nresults = []\nfor df in df_gr:\n    res = processing_df(df)\n    results.append(res)\n```\n\nP.S. I tried use [this approaches](https://www.machinelearningplus.com/python/parallel-processing-python/), but processes were not running.",
    "2149899": "It is impossible to cope only with aggregation functions, since I use not only them",
    "2149961": "To parallelize the processing of multiple partitioned dataframes using Polars partition_by function, you can make use of Python's built-in multiprocessing module.\n\nHere's a high-level overview of the approach:\n\nCreate a list of the partitioned dataframes using the partition_by function.\nSplit the list of dataframes into smaller chunks, based on the number of CPU cores available or some other factor.\nUse the multiprocessing.Pool class to create a pool of worker processes.\nUse the Pool.map method to apply a function to each chunk of dataframes in parallel.\nFinally, use the concat method to combine the results from the processed dataframes into a single dataframe.\n\nHere's some sample code that demonstrates this approach:\n```python\n\nimport polars as pl\nimport multiprocessing as mp\n\n# create a DataFrame\ndf = pl.DataFrame({\n    'foo': ['A', 'A', 'B', 'B', 'C', 'C'],\n    'bar': [1, 2, 3, 4, 5, 6]\n})\n\n# partition the DataFrame by 'foo'\npartitions = df.partition_by('foo')\n\n# define a function to process a chunk of partitions\ndef process_partitions(partitions):\n    # do some processing on the partitions\n    result = partitions.select('foo', 'bar').groupby('foo').sum()\n    return result\n\n# split the list of partitions into chunks\nchunk_size = len(partitions) // mp.cpu_count()\npartition_chunks = [partitions[i:i+chunk_size] for i in range(0, len(partitions), chunk_size)]\n\n# create a pool of worker processes\npool = mp.Pool()\n\n# apply the function to each chunk of partitions in parallel\nresults = pool.map(process_partitions, partition_chunks)\n\n# concatenate the results into a single dataframe\nfinal_result = pl.concat(results)\n```\n\n\nIn this example, the process_partitions function is applied to each chunk of partitions in parallel using the Pool.map method. The resulting dataframes are concatenated into a single dataframe using the concat method.",
    "2150346": "Thanks you very much for the answer! I tried to use this method, but the functions were not executed. I now tried the method with joblib Parallel, delayed - this way it worked\n\n```\nfrom joblib import Parallel, delayed\nimport polars as pl\n\ndf = pl.DataFrame({\n    'foo': ['A', 'A', 'B', 'B', 'C', 'C'],\n    'bar': [1, 2, 3, 4, 5, 6]\n})\n\n# partition the DataFrame by 'foo'\npartitions = df.partition_by('foo')\n\n# define a function to process a chunk of partitions\ndef process_partitions(partitions):\n    # do some processing on the partitions\n    result = [part.select('foo', 'bar').groupby('foo').sum() for part in partitions]\n    return result\n\n# split the list of partitions into chunks\nchunk_size = len(partitions) // mp.cpu_count()\npartition_chunks = [partitions[i:i+chunk_size] for i in range(0, len(partitions), chunk_size)]\n\nres = Parallel(\n    n_jobs=5\n)(\n    delayed(process_partitions)(chunk) for chink in partition_chunks\n)\n```"
  },
  "source": "meta"
}