{
  "id": 393311,
  "title": "fastparquet is faster than pandas",
  "url": "/competitions/asl-signs/discussion/393311",
  "author_name": "",
  "post_date": "2023-03-08T20:26:27.484057800Z",
  "votes": 2,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hello Kagglers  ✋,</p>\n<p>I'm sharing with you here a quick benchmark of execution time between pandas and fastparquet to read parquet files. A 25% improvement in the loading time 💥. This is very interesting especially with the large number of parquet files we have.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3775861%2Fd7086d1507ca65bc790fa630770a8677%2FCapture.PNG?generation=1678307032908608&amp;alt=media\" alt=\"\"></p>\n<p>Good luck!</p>",
  "messages": [
    {
      "id": "2174082",
      "postDate": "03/08/2023 20:26:27",
      "content": "<p>Hello Kagglers  ✋,</p>\n<p>I'm sharing with you here a quick benchmark of execution time between pandas and fastparquet to read parquet files. A 25% improvement in the loading time 💥. This is very interesting especially with the large number of parquet files we have.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3775861%2Fd7086d1507ca65bc790fa630770a8677%2FCapture.PNG?generation=1678307032908608&amp;alt=media\" alt=\"\"></p>\n<p>Good luck!</p>",
      "rawMarkdown": "Hello Kagglers  ✋,\n\nI'm sharing with you here a quick benchmark of execution time between pandas and fastparquet to read parquet files. A 25% improvement in the loading time 💥. This is very interesting especially with the large number of parquet files we have.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3775861%2Fd7086d1507ca65bc790fa630770a8677%2FCapture.PNG?generation=1678307032908608&alt=media)\n\nGood luck!",
      "votes": null
    },
    {
      "id": "2174097",
      "postDate": "03/08/2023 20:38:42",
      "content": "<p>Is it possible to give some statistics about the parquet file you're reading in? (namely the size, # records, # columns, compression if used, average sizes of each column etc)?<br>\nThat would be very interesting to me.</p>",
      "rawMarkdown": "Is it possible to give some statistics about the parquet file you're reading in? (namely the size, # records, # columns, compression if used, average sizes of each column etc)?\nThat would be very interesting to me.",
      "votes": null
    },
    {
      "id": "2174104",
      "postDate": "03/08/2023 20:50:32",
      "content": "<p>It has 12489 raws, 7 columns (3 float, two int and two string) and the size of the file is 350 KB</p>",
      "rawMarkdown": "It has 12489 raws, 7 columns (3 float, two int and two string) and the size of the file is 350 KB",
      "votes": null
    },
    {
      "id": "2174596",
      "postDate": "03/09/2023 09:05:59",
      "content": "<p>How does it compare when you actually try to do anything with the data?  I ask because, with minimal benchmarks like this, there's always a chance that one of them is doing lazy loading (fast to load, but then slow to access) and the other is doing eager loading (slow to load but then fast to access).</p>",
      "rawMarkdown": "How does it compare when you actually try to do anything with the data?  I ask because, with minimal benchmarks like this, there's always a chance that one of them is doing lazy loading (fast to load, but then slow to access) and the other is doing eager loading (slow to load but then fast to access).",
      "votes": null
    },
    {
      "id": "2174623",
      "postDate": "03/09/2023 09:33:25",
      "content": "<p>I was curious about your remark so I made a small test:<br>\nfastparquet v 2023.2.0</p>\n<pre><code>%%timeit\npa_df = ParquetFile(\"out.parquet\").to_pandas()\npa_df['target'] = pa_df.target + 1\n==&gt; 81.2 ms ± 652 µs per loop (mean ± std. dev. of 7 runs, 10 loops each)\n\n%%timeit\npd_df = pd.read_parquet(\"out.parquet\")\npd_df['target'] = pd_df.target + 1\n==&gt; 174 ms ± 4.77 ms per loop (mean ± std. dev. of 7 runs, 10 loops each)\n</code></pre>\n<p>So still faster.</p>\n<p>But then I wanted to check all the columns and I got:</p>\n<pre><code>pa_df['customer_ID'] = pa_df.customer_ID.str[-10:]\n</code></pre>\n<p>==&gt; 173 ms ± 2.3 ms per loop (mean ± std. dev. of 7 runs, 10 loops each)<br>\n==&gt; 258 ms ± 2.68 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)</p>\n<p>So I believe <a href=\"https://www.kaggle.com/andrewrrose\" target=\"_blank\">@andrewrrose</a> is accurate in his assumption that there are optimizations while reading.<br>\nParquet is a tabular format with column-based-optimizations, and it shows in my test here above.</p>",
      "rawMarkdown": "I was curious about your remark so I made a small test:\nfastparquet v 2023.2.0\n\n```\n%%timeit\npa_df = ParquetFile(\"out.parquet\").to_pandas()\npa_df['target'] = pa_df.target + 1\n==> 81.2 ms ± 652 µs per loop (mean ± std. dev. of 7 runs, 10 loops each)\n\n%%timeit\npd_df = pd.read_parquet(\"out.parquet\")\npd_df['target'] = pd_df.target + 1\n==> 174 ms ± 4.77 ms per loop (mean ± std. dev. of 7 runs, 10 loops each)\n```\nSo still faster.\n\nBut then I wanted to check all the columns and I got:\n```\npa_df['customer_ID'] = pa_df.customer_ID.str[-10:]\n```\n==> 173 ms ± 2.3 ms per loop (mean ± std. dev. of 7 runs, 10 loops each)\n==> 258 ms ± 2.68 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)\n\nSo I believe @andrewrrose is accurate in his assumption that there are optimizations while reading.\nParquet is a tabular format with column-based-optimizations, and it shows in my test here above.",
      "votes": null
    },
    {
      "id": "2174624",
      "postDate": "03/09/2023 09:33:45",
      "content": "<p>But in this case I'm loading using fastparquet then converting to pandas dataframe and it's faster than loading directly in pandas. So  Time (loading in fast parquet + converting to pandas) &lt; Time (loading in pandas). </p>",
      "rawMarkdown": "But in this case I'm loading using fastparquet then converting to pandas dataframe and it's faster than loading directly in pandas. So  Time (loading in fast parquet + converting to pandas) < Time (loading in pandas).",
      "votes": null
    },
    {
      "id": "2174634",
      "postDate": "03/09/2023 09:37:55",
      "content": "<p>Yes I agree with the assumptions of <a href=\"https://www.kaggle.com/andrewrrose\" target=\"_blank\">@andrewrrose</a> . But if you load in fastparquet then you do the processing in pandas. You still have an optimization in the loading time because Time (loading in fast parquet + converting to pandas) &lt; Time (loading in pandas). Then you'll be using pandas to do the other operations. In this competition it takes 30min to load all parquet files in pandas so it's still usefull to reduce this bottlneck.</p>",
      "rawMarkdown": "Yes I agree with the assumptions of @andrewrrose . But if you load in fastparquet then you do the processing in pandas. You still have an optimization in the loading time because Time (loading in fast parquet + converting to pandas) < Time (loading in pandas). Then you'll be using pandas to do the other operations. In this competition it takes 30min to load all parquet files in pandas so it's still usefull to reduce this bottlneck.",
      "votes": null
    },
    {
      "id": "2174663",
      "postDate": "03/09/2023 09:49:46",
      "content": "<p>It certainly is faster to get a dataframe back. However, once you need the actual data it'll be very similar in loading times.<br>\nThe main advantage is that if you do not need to use a column, it doesn't get loaded from disc. Only the types etc will be loaded from the metadata in the parquet file.</p>",
      "rawMarkdown": "It certainly is faster to get a dataframe back. However, once you need the actual data it'll be very similar in loading times.\nThe main advantage is that if you do not need to use a column, it doesn't get loaded from disc. Only the types etc will be loaded from the metadata in the parquet file.",
      "votes": null
    },
    {
      "id": "2174693",
      "postDate": "03/09/2023 09:56:27",
      "content": "<p>I agree with your point yes.</p>",
      "rawMarkdown": "I agree with your point yes.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2174097,
      "author_name": "svaningelgem",
      "author_url": "",
      "post_date": "03/08/2023 20:38:42",
      "content": "<p>Is it possible to give some statistics about the parquet file you're reading in? (namely the size, # records, # columns, compression if used, average sizes of each column etc)?<br>\nThat would be very interesting to me.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2174104,
          "author_name": "jarvisai7",
          "author_url": "",
          "post_date": "03/08/2023 20:50:32",
          "content": "<p>It has 12489 raws, 7 columns (3 float, two int and two string) and the size of the file is 350 KB</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2174596,
      "author_name": "andrewrrose",
      "author_url": "",
      "post_date": "03/09/2023 09:05:59",
      "content": "<p>How does it compare when you actually try to do anything with the data?  I ask because, with minimal benchmarks like this, there's always a chance that one of them is doing lazy loading (fast to load, but then slow to access) and the other is doing eager loading (slow to load but then fast to access).</p>",
      "votes": null,
      "replies": [
        {
          "id": 2174623,
          "author_name": "svaningelgem",
          "author_url": "",
          "post_date": "03/09/2023 09:33:25",
          "content": "<p>I was curious about your remark so I made a small test:<br>\nfastparquet v 2023.2.0</p>\n<pre><code>%%timeit\npa_df = ParquetFile(\"out.parquet\").to_pandas()\npa_df['target'] = pa_df.target + 1\n==&gt; 81.2 ms ± 652 µs per loop (mean ± std. dev. of 7 runs, 10 loops each)\n\n%%timeit\npd_df = pd.read_parquet(\"out.parquet\")\npd_df['target'] = pd_df.target + 1\n==&gt; 174 ms ± 4.77 ms per loop (mean ± std. dev. of 7 runs, 10 loops each)\n</code></pre>\n<p>So still faster.</p>\n<p>But then I wanted to check all the columns and I got:</p>\n<pre><code>pa_df['customer_ID'] = pa_df.customer_ID.str[-10:]\n</code></pre>\n<p>==&gt; 173 ms ± 2.3 ms per loop (mean ± std. dev. of 7 runs, 10 loops each)<br>\n==&gt; 258 ms ± 2.68 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)</p>\n<p>So I believe <a href=\"https://www.kaggle.com/andrewrrose\" target=\"_blank\">@andrewrrose</a> is accurate in his assumption that there are optimizations while reading.<br>\nParquet is a tabular format with column-based-optimizations, and it shows in my test here above.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2174634,
              "author_name": "jarvisai7",
              "author_url": "",
              "post_date": "03/09/2023 09:37:55",
              "content": "<p>Yes I agree with the assumptions of <a href=\"https://www.kaggle.com/andrewrrose\" target=\"_blank\">@andrewrrose</a> . But if you load in fastparquet then you do the processing in pandas. You still have an optimization in the loading time because Time (loading in fast parquet + converting to pandas) &lt; Time (loading in pandas). Then you'll be using pandas to do the other operations. In this competition it takes 30min to load all parquet files in pandas so it's still usefull to reduce this bottlneck.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2174663,
                  "author_name": "svaningelgem",
                  "author_url": "",
                  "post_date": "03/09/2023 09:49:46",
                  "content": "<p>It certainly is faster to get a dataframe back. However, once you need the actual data it'll be very similar in loading times.<br>\nThe main advantage is that if you do not need to use a column, it doesn't get loaded from disc. Only the types etc will be loaded from the metadata in the parquet file.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2174693,
                      "author_name": "jarvisai7",
                      "author_url": "",
                      "post_date": "03/09/2023 09:56:27",
                      "content": "<p>I agree with your point yes.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        },
        {
          "id": 2174624,
          "author_name": "jarvisai7",
          "author_url": "",
          "post_date": "03/09/2023 09:33:45",
          "content": "<p>But in this case I'm loading using fastparquet then converting to pandas dataframe and it's faster than loading directly in pandas. So  Time (loading in fast parquet + converting to pandas) &lt; Time (loading in pandas). </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2174082": "Hello Kagglers  ✋,\n\nI'm sharing with you here a quick benchmark of execution time between pandas and fastparquet to read parquet files. A 25% improvement in the loading time 💥. This is very interesting especially with the large number of parquet files we have.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3775861%2Fd7086d1507ca65bc790fa630770a8677%2FCapture.PNG?generation=1678307032908608&alt=media)\n\nGood luck!",
    "2174097": "Is it possible to give some statistics about the parquet file you're reading in? (namely the size, # records, # columns, compression if used, average sizes of each column etc)?\nThat would be very interesting to me.",
    "2174104": "It has 12489 raws, 7 columns (3 float, two int and two string) and the size of the file is 350 KB",
    "2174596": "How does it compare when you actually try to do anything with the data?  I ask because, with minimal benchmarks like this, there's always a chance that one of them is doing lazy loading (fast to load, but then slow to access) and the other is doing eager loading (slow to load but then fast to access).",
    "2174623": "I was curious about your remark so I made a small test:\nfastparquet v 2023.2.0\n\n```\n%%timeit\npa_df = ParquetFile(\"out.parquet\").to_pandas()\npa_df['target'] = pa_df.target + 1\n==> 81.2 ms ± 652 µs per loop (mean ± std. dev. of 7 runs, 10 loops each)\n\n%%timeit\npd_df = pd.read_parquet(\"out.parquet\")\npd_df['target'] = pd_df.target + 1\n==> 174 ms ± 4.77 ms per loop (mean ± std. dev. of 7 runs, 10 loops each)\n```\nSo still faster.\n\nBut then I wanted to check all the columns and I got:\n```\npa_df['customer_ID'] = pa_df.customer_ID.str[-10:]\n```\n==> 173 ms ± 2.3 ms per loop (mean ± std. dev. of 7 runs, 10 loops each)\n==> 258 ms ± 2.68 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)\n\nSo I believe @andrewrrose is accurate in his assumption that there are optimizations while reading.\nParquet is a tabular format with column-based-optimizations, and it shows in my test here above.",
    "2174624": "But in this case I'm loading using fastparquet then converting to pandas dataframe and it's faster than loading directly in pandas. So  Time (loading in fast parquet + converting to pandas) < Time (loading in pandas).",
    "2174634": "Yes I agree with the assumptions of @andrewrrose . But if you load in fastparquet then you do the processing in pandas. You still have an optimization in the loading time because Time (loading in fast parquet + converting to pandas) < Time (loading in pandas). Then you'll be using pandas to do the other operations. In this competition it takes 30min to load all parquet files in pandas so it's still usefull to reduce this bottlneck.",
    "2174663": "It certainly is faster to get a dataframe back. However, once you need the actual data it'll be very similar in loading times.\nThe main advantage is that if you do not need to use a column, it doesn't get loaded from disc. Only the types etc will be loaded from the metadata in the parquet file.",
    "2174693": "I agree with your point yes."
  },
  "source": "meta"
}