{
  "id": 475389,
  "title": "Faster ways to load competition's data",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/475389",
  "author_name": "",
  "post_date": "2024-02-08T09:10:58.988526800Z",
  "votes": 54,
  "comment_count": 10,
  "views": 0,
  "content": "<p>For this competition we've got quite a lot of data: dozens of files with the overall size of ~26 GB. Loading data is going to be an intrinsic part of every kernel, so it is obviously important to minimize reading times as much as possible. </p>\n<p>According to the official dataset description, all files can be found in both CSV and Parquet formats. So what is the fastest way to read them? In <a href=\"https://www.kaggle.com/code/kononenko/home-credit-faster-data-loading/\" target=\"_blank\">home-credit-faster-data-loading</a> notebook I did some benchmarks of a few popular packages: <a href=\"https://pandas.pydata.org/docs/\" target=\"_blank\">pandas</a>, <a href=\"https://pola.rs/\" target=\"_blank\">polars</a> and <a href=\"https://datatable.readthedocs.io/\" target=\"_blank\">datatable</a>. In this post I will briefly summarize the main findings.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2099265%2Fe314f1bee56e388ecb7ec1b090aec82d%2Fpandas_polars_datatable.png?generation=1707791559484413&amp;alt=media\"></p>\n<p><strong>TL;DR:</strong> use <code>polars</code> to read Parquet or <code>datatable</code> to load memory-mapped files almost instantly.</p>\n<h1>1. Loading data from CSV files</h1>\n<p>The table below presents time taken by packages to load all the competition's data from the CSV files.</p>\n<table>\n<thead>\n<tr>\n<th>Package</th>\n<th>Time to read all CSV files [s]</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>pandas</td>\n<td>1100.03</td>\n</tr>\n<tr>\n<td>polars</td>\n<td>142.73</td>\n</tr>\n<tr>\n<td>datatable</td>\n<td>118.37</td>\n</tr>\n</tbody>\n</table>\n<p><br></p>\n<p>We can also look at this benchmark file-wise for the first few files in the training dataset.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2099265%2Fb546572054224b09182d6e8be084afd5%2Fload_csv.png?generation=1707382589451428&amp;alt=media\"></p>\n<p>As we see,<code>pandas</code> is pretty slow when it comes to load larger files. <code>datatable</code> is doing a little bit better than <code>polars</code>, though both these packages are quite fast demonstrating similar performance file-wise.</p>\n<h1>2. Loading data from Parquet files</h1>\n<p>Loading data from Parquet is much faster, because that's a binary format. We only compare <code>pandas</code> and <code>polars</code>, because <code>datatable</code> doesn't have a dedicated function to load <code>.parquet</code> relying on the <code>arrow</code> library.</p>\n<table>\n<thead>\n<tr>\n<th>Package</th>\n<th>Time to read all Parquet files [s]</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>pandas</td>\n<td>129.83</td>\n</tr>\n<tr>\n<td>polars</td>\n<td>55.22</td>\n</tr>\n</tbody>\n</table>\n<p><br></p>\n<p>Similarly, we can also plot file-wise comparison.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2099265%2Fdd96fa932bf9aeba7d17c60f7fc6de90%2Fload_parquet.png?generation=1707382782914363&amp;alt=media\"></p>\n<p>As expected, both packages process Parquet much faster than CSV with <code>polars</code> demonstrating better performance. However, processing times are still non-negligible when it comes to loading all the data.</p>\n<h1>3. Loading data with <a href=\"https://en.wikipedia.org/wiki/Memory-mapped_file\" target=\"_blank\">memory mapping</a></h1>\n<p>Once data is loaded into a dataframe, we can actually export it in a binary format appropriate for memory mapping: <a href=\"https://docs.pola.rs/py-polars/html/reference/api/polars.DataFrame.write_ipc.html\" target=\"_blank\">ipc</a> for <code>polars</code> and <a href=\"https://datatable.readthedocs.io/en/latest/api/frame/to_jay.html\" target=\"_blank\">jay</a> for <code>datatable</code>. Reading data back should literally take no time, because files are mapped into the memory address space very fast. To confirm that, I benchmarked <code>.ipc</code> and <code>.jay</code> loading for five first files from the competition's data.</p>\n<table>\n<thead>\n<tr>\n<th>Package</th>\n<th>Time to memory map five files [s]</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>polars</td>\n<td>1.25</td>\n</tr>\n<tr>\n<td>datatable</td>\n<td>0.04</td>\n</tr>\n</tbody>\n</table>\n<p><br></p>\n<p>The measured timings are much smaller than even those for Parquet with <code>datatable</code> memory mapping being almost immediate. When comparing file-wise, one can see that <code>datatable</code>'s performance it almost file independent, while <code>polars</code> timings change quite a lot.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2099265%2F78f6b9f2c118f52878c4e01b83394008%2Fload_mmap.png?generation=1707382990401744&amp;alt=media\"></p>\n<h1>Conclusion</h1>\n<p>To sum up, CSV is the slowest way to get the data, reading Parquet is about twice faster and memory mapping could be almost instant. The final benchmarking results are also summarized in the table below. Note, we are not using actual numbers in this table, because those could vary on different systems from run to run.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>pandas</th>\n<th>polars</th>\n<th>datatable</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>CSV</td>\n<td>Slow</td>\n<td>Fast</td>\n<td>Fast</td>\n</tr>\n<tr>\n<td>Parquet</td>\n<td>OK</td>\n<td>Fast</td>\n<td>N/A</td>\n</tr>\n<tr>\n<td>Memory mapping</td>\n<td>N/A</td>\n<td>OK</td>\n<td>Fast</td>\n</tr>\n</tbody>\n</table>\n<p><br></p>\n<p>For more details, please refer to <a href=\"https://www.kaggle.com/code/kononenko/home-credit-faster-data-loading/\" target=\"_blank\">home-credit-faster-data-loading</a> notebook. Hope it helps and good luck with this awesome competition!</p>",
  "messages": [
    {
      "id": "2642588",
      "postDate": "02/08/2024 09:10:58",
      "content": "<p>For this competition we've got quite a lot of data: dozens of files with the overall size of ~26 GB. Loading data is going to be an intrinsic part of every kernel, so it is obviously important to minimize reading times as much as possible. </p>\n<p>According to the official dataset description, all files can be found in both CSV and Parquet formats. So what is the fastest way to read them? In <a href=\"https://www.kaggle.com/code/kononenko/home-credit-faster-data-loading/\" target=\"_blank\">home-credit-faster-data-loading</a> notebook I did some benchmarks of a few popular packages: <a href=\"https://pandas.pydata.org/docs/\" target=\"_blank\">pandas</a>, <a href=\"https://pola.rs/\" target=\"_blank\">polars</a> and <a href=\"https://datatable.readthedocs.io/\" target=\"_blank\">datatable</a>. In this post I will briefly summarize the main findings.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2099265%2Fe314f1bee56e388ecb7ec1b090aec82d%2Fpandas_polars_datatable.png?generation=1707791559484413&amp;alt=media\"></p>\n<p><strong>TL;DR:</strong> use <code>polars</code> to read Parquet or <code>datatable</code> to load memory-mapped files almost instantly.</p>\n<h1>1. Loading data from CSV files</h1>\n<p>The table below presents time taken by packages to load all the competition's data from the CSV files.</p>\n<table>\n<thead>\n<tr>\n<th>Package</th>\n<th>Time to read all CSV files [s]</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>pandas</td>\n<td>1100.03</td>\n</tr>\n<tr>\n<td>polars</td>\n<td>142.73</td>\n</tr>\n<tr>\n<td>datatable</td>\n<td>118.37</td>\n</tr>\n</tbody>\n</table>\n<p><br></p>\n<p>We can also look at this benchmark file-wise for the first few files in the training dataset.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2099265%2Fb546572054224b09182d6e8be084afd5%2Fload_csv.png?generation=1707382589451428&amp;alt=media\"></p>\n<p>As we see,<code>pandas</code> is pretty slow when it comes to load larger files. <code>datatable</code> is doing a little bit better than <code>polars</code>, though both these packages are quite fast demonstrating similar performance file-wise.</p>\n<h1>2. Loading data from Parquet files</h1>\n<p>Loading data from Parquet is much faster, because that's a binary format. We only compare <code>pandas</code> and <code>polars</code>, because <code>datatable</code> doesn't have a dedicated function to load <code>.parquet</code> relying on the <code>arrow</code> library.</p>\n<table>\n<thead>\n<tr>\n<th>Package</th>\n<th>Time to read all Parquet files [s]</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>pandas</td>\n<td>129.83</td>\n</tr>\n<tr>\n<td>polars</td>\n<td>55.22</td>\n</tr>\n</tbody>\n</table>\n<p><br></p>\n<p>Similarly, we can also plot file-wise comparison.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2099265%2Fdd96fa932bf9aeba7d17c60f7fc6de90%2Fload_parquet.png?generation=1707382782914363&amp;alt=media\"></p>\n<p>As expected, both packages process Parquet much faster than CSV with <code>polars</code> demonstrating better performance. However, processing times are still non-negligible when it comes to loading all the data.</p>\n<h1>3. Loading data with <a href=\"https://en.wikipedia.org/wiki/Memory-mapped_file\" target=\"_blank\">memory mapping</a></h1>\n<p>Once data is loaded into a dataframe, we can actually export it in a binary format appropriate for memory mapping: <a href=\"https://docs.pola.rs/py-polars/html/reference/api/polars.DataFrame.write_ipc.html\" target=\"_blank\">ipc</a> for <code>polars</code> and <a href=\"https://datatable.readthedocs.io/en/latest/api/frame/to_jay.html\" target=\"_blank\">jay</a> for <code>datatable</code>. Reading data back should literally take no time, because files are mapped into the memory address space very fast. To confirm that, I benchmarked <code>.ipc</code> and <code>.jay</code> loading for five first files from the competition's data.</p>\n<table>\n<thead>\n<tr>\n<th>Package</th>\n<th>Time to memory map five files [s]</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>polars</td>\n<td>1.25</td>\n</tr>\n<tr>\n<td>datatable</td>\n<td>0.04</td>\n</tr>\n</tbody>\n</table>\n<p><br></p>\n<p>The measured timings are much smaller than even those for Parquet with <code>datatable</code> memory mapping being almost immediate. When comparing file-wise, one can see that <code>datatable</code>'s performance it almost file independent, while <code>polars</code> timings change quite a lot.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2099265%2F78f6b9f2c118f52878c4e01b83394008%2Fload_mmap.png?generation=1707382990401744&amp;alt=media\"></p>\n<h1>Conclusion</h1>\n<p>To sum up, CSV is the slowest way to get the data, reading Parquet is about twice faster and memory mapping could be almost instant. The final benchmarking results are also summarized in the table below. Note, we are not using actual numbers in this table, because those could vary on different systems from run to run.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>pandas</th>\n<th>polars</th>\n<th>datatable</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>CSV</td>\n<td>Slow</td>\n<td>Fast</td>\n<td>Fast</td>\n</tr>\n<tr>\n<td>Parquet</td>\n<td>OK</td>\n<td>Fast</td>\n<td>N/A</td>\n</tr>\n<tr>\n<td>Memory mapping</td>\n<td>N/A</td>\n<td>OK</td>\n<td>Fast</td>\n</tr>\n</tbody>\n</table>\n<p><br></p>\n<p>For more details, please refer to <a href=\"https://www.kaggle.com/code/kononenko/home-credit-faster-data-loading/\" target=\"_blank\">home-credit-faster-data-loading</a> notebook. Hope it helps and good luck with this awesome competition!</p>",
      "rawMarkdown": "For this competition we've got quite a lot of data: dozens of files with the overall size of ~26 GB. Loading data is going to be an intrinsic part of every kernel, so it is obviously important to minimize reading times as much as possible. \n\nAccording to the official dataset description, all files can be found in both CSV and Parquet formats. So what is the fastest way to read them? In [home-credit-faster-data-loading](https://www.kaggle.com/code/kononenko/home-credit-faster-data-loading/) notebook I did some benchmarks of a few popular packages: [pandas](https://pandas.pydata.org/docs/), [polars](https://pola.rs/) and [datatable](https://datatable.readthedocs.io/). In this post I will briefly summarize the main findings.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2099265%2Fe314f1bee56e388ecb7ec1b090aec82d%2Fpandas_polars_datatable.png?generation=1707791559484413&alt=media)\n\n**TL;DR:** use `polars` to read Parquet or `datatable` to load memory-mapped files almost instantly.\n\n# 1. Loading data from CSV files\n\nThe table below presents time taken by packages to load all the competition's data from the CSV files.\n\n| Package | Time to read all CSV files [s] |\n|----------|----------|\n| pandas    | 1100.03 |\n| polars      |142.73 |\n| datatable | 118.37 |\n\n<br>\n\nWe can also look at this benchmark file-wise for the first few files in the training dataset.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2099265%2Fb546572054224b09182d6e8be084afd5%2Fload_csv.png?generation=1707382589451428&alt=media)\n\nAs we see,`pandas` is pretty slow when it comes to load larger files. `datatable` is doing a little bit better than `polars`, though both these packages are quite fast demonstrating similar performance file-wise.\n\n# 2. Loading data from Parquet files\n\nLoading data from Parquet is much faster, because that's a binary format. We only compare `pandas` and `polars`, because `datatable` doesn't have a dedicated function to load `.parquet` relying on the `arrow` library.\n\n| Package | Time to read all Parquet files [s] |\n|----------|----------|\n| pandas    | 129.83 |\n| polars      |55.22 |\n\n<br>\n\nSimilarly, we can also plot file-wise comparison.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2099265%2Fdd96fa932bf9aeba7d17c60f7fc6de90%2Fload_parquet.png?generation=1707382782914363&alt=media)\n\nAs expected, both packages process Parquet much faster than CSV with `polars` demonstrating better performance. However, processing times are still non-negligible when it comes to loading all the data.\n\n# 3. Loading data with [memory mapping](https://en.wikipedia.org/wiki/Memory-mapped_file)\n\nOnce data is loaded into a dataframe, we can actually export it in a binary format appropriate for memory mapping: [ipc](https://docs.pola.rs/py-polars/html/reference/api/polars.DataFrame.write_ipc.html) for `polars` and [jay](https://datatable.readthedocs.io/en/latest/api/frame/to_jay.html) for `datatable`. Reading data back should literally take no time, because files are mapped into the memory address space very fast. To confirm that, I benchmarked `.ipc` and `.jay` loading for five first files from the competition's data.\n\n| Package | Time to memory map five files [s] |\n|----------|----------|\n| polars    | 1.25 |\n| datatable      | 0.04 |\n\n<br>\n\nThe measured timings are much smaller than even those for Parquet with `datatable` memory mapping being almost immediate. When comparing file-wise, one can see that `datatable`'s performance it almost file independent, while `polars` timings change quite a lot.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2099265%2F78f6b9f2c118f52878c4e01b83394008%2Fload_mmap.png?generation=1707382990401744&alt=media)\n\n# Conclusion\n\nTo sum up, CSV is the slowest way to get the data, reading Parquet is about twice faster and memory mapping could be almost instant. The final benchmarking results are also summarized in the table below. Note, we are not using actual numbers in this table, because those could vary on different systems from run to run.\n\n|               | pandas | polars | datatable |\n|---------------|--------|--------|-----------|\n| CSV           | Slow   | Fast   | Fast      | \n| Parquet       | OK     | Fast   | N/A       |\n| Memory mapping| N/A    | OK     | Fast      |\n\n<br>\n\nFor more details, please refer to [home-credit-faster-data-loading](https://www.kaggle.com/code/kononenko/home-credit-faster-data-loading/) notebook. Hope it helps and good luck with this awesome competition!",
      "votes": null
    },
    {
      "id": "2642745",
      "postDate": "02/08/2024 11:31:23",
      "content": "<p>Impressive, thanks for the effort and comparison.</p>",
      "rawMarkdown": "Impressive, thanks for the effort and comparison.",
      "votes": null
    },
    {
      "id": "2646893",
      "postDate": "02/11/2024 09:11:58",
      "content": "<p>Hi, do you think loading data faster is crucial in this competion?</p>",
      "rawMarkdown": "Hi, do you think loading data faster is crucial in this competion?",
      "votes": null
    },
    {
      "id": "2646926",
      "postDate": "02/11/2024 09:42:03",
      "content": "<p>I think so, because if one makes an incorrect choice by sticking to <code>pandas</code> to read all the <code>.csv</code> files, it will take 1000+ seconds just to load the data, and this is something that needs to be done as often as we start a new session in our competition notebook… </p>",
      "rawMarkdown": "I think so, because if one makes an incorrect choice by sticking to `pandas` to read all the `.csv` files, it will take 1000+ seconds just to load the data, and this is something that needs to be done as often as we start a new session in our competition notebook...",
      "votes": null
    },
    {
      "id": "2646941",
      "postDate": "02/11/2024 09:46:40",
      "content": "<p>But pandas and parquet would be ok in your opinion?</p>",
      "rawMarkdown": "But pandas and parquet would be ok in your opinion?",
      "votes": null
    },
    {
      "id": "2646946",
      "postDate": "02/11/2024 09:51:58",
      "content": "<p>It is not that bad, see benchmarks <a href=\"https://www.kaggle.com/code/kononenko/home-credit-faster-data-loading/#2.-Loading-data-from-Parquet-files\" target=\"_blank\">here</a>, about 2 minutes to load all the data, <code>polars</code> would do it in 1 minute. So it is up to you to decide, if you need 50% speed-up or not.</p>",
      "rawMarkdown": "It is not that bad, see benchmarks [here](https://www.kaggle.com/code/kononenko/home-credit-faster-data-loading/#2.-Loading-data-from-Parquet-files), about 2 minutes to load all the data, `polars` would do it in 1 minute. So it is up to you to decide, if you need 50% speed-up or not.",
      "votes": null
    },
    {
      "id": "2647054",
      "postDate": "02/11/2024 10:44:24",
      "content": "<p>Thank You dear Oleksiy <a href=\"https://www.kaggle.com/kononenko\" target=\"_blank\">@kononenko</a> for Your comparison and getting me familiar with datatable! I will try to use them both!</p>",
      "rawMarkdown": "Thank You dear Oleksiy @kononenko for Your comparison and getting me familiar with datatable! I will try to use them both!",
      "votes": null
    },
    {
      "id": "2652002",
      "postDate": "02/14/2024 13:59:30",
      "content": "<p>Thanks for sharing </p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": null
    },
    {
      "id": "2671195",
      "postDate": "02/27/2024 11:30:43",
      "content": "<p>Thanks for the comparison, very interesting! </p>\n<p>From a kaggle newbie: Do you plan to then manipulate the data with the loading library, or do you load it with the fastest library and then convert it to your favorite one? (or the most capable one?). For example, I feel comfortable with pandas, and learning to do all the same things in polars is not very appealing.</p>",
      "rawMarkdown": "Thanks for the comparison, very interesting! \n\nFrom a kaggle newbie: Do you plan to then manipulate the data with the loading library, or do you load it with the fastest library and then convert it to your favorite one? (or the most capable one?). For example, I feel comfortable with pandas, and learning to do all the same things in polars is not very appealing.",
      "votes": null
    },
    {
      "id": "2671268",
      "postDate": "02/27/2024 12:21:32",
      "content": "<p>You can go with any of the options you mention, just additional benchmarks are needed as to how fast/memory hungry are conversions, <code>join</code> and other munging capabilities. </p>\n<p>One thing to keep in mind is that at the very end one needs to pass the resulting frames to modeling packages like LightGBM and those may not support all the possible frame types, so conversion to <code>pandas</code> or <code>arrow</code> may be necessary, that could take additional time.</p>",
      "rawMarkdown": "You can go with any of the options you mention, just additional benchmarks are needed as to how fast/memory hungry are conversions, `join` and other munging capabilities. \n\nOne thing to keep in mind is that at the very end one needs to pass the resulting frames to modeling packages like LightGBM and those may not support all the possible frame types, so conversion to `pandas` or `arrow` may be necessary, that could take additional time.",
      "votes": null
    },
    {
      "id": "2720508",
      "postDate": "03/28/2024 12:00:18",
      "content": "<p>thank you share,very good !</p>",
      "rawMarkdown": "thank you share,very good !",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2642745,
      "author_name": "jetakow",
      "author_url": "",
      "post_date": "02/08/2024 11:31:23",
      "content": "<p>Impressive, thanks for the effort and comparison.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2646893,
      "author_name": "lucamtb",
      "author_url": "",
      "post_date": "02/11/2024 09:11:58",
      "content": "<p>Hi, do you think loading data faster is crucial in this competion?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2646926,
          "author_name": "kononenko",
          "author_url": "",
          "post_date": "02/11/2024 09:42:03",
          "content": "<p>I think so, because if one makes an incorrect choice by sticking to <code>pandas</code> to read all the <code>.csv</code> files, it will take 1000+ seconds just to load the data, and this is something that needs to be done as often as we start a new session in our competition notebook… </p>",
          "votes": null,
          "replies": [
            {
              "id": 2646941,
              "author_name": "lucamtb",
              "author_url": "",
              "post_date": "02/11/2024 09:46:40",
              "content": "<p>But pandas and parquet would be ok in your opinion?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2646946,
                  "author_name": "kononenko",
                  "author_url": "",
                  "post_date": "02/11/2024 09:51:58",
                  "content": "<p>It is not that bad, see benchmarks <a href=\"https://www.kaggle.com/code/kononenko/home-credit-faster-data-loading/#2.-Loading-data-from-Parquet-files\" target=\"_blank\">here</a>, about 2 minutes to load all the data, <code>polars</code> would do it in 1 minute. So it is up to you to decide, if you need 50% speed-up or not.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2647054,
      "author_name": "kapturovalexander",
      "author_url": "",
      "post_date": "02/11/2024 10:44:24",
      "content": "<p>Thank You dear Oleksiy <a href=\"https://www.kaggle.com/kononenko\" target=\"_blank\">@kononenko</a> for Your comparison and getting me familiar with datatable! I will try to use them both!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2652002,
      "author_name": "",
      "author_url": "",
      "post_date": "02/14/2024 13:59:30",
      "content": "<p>Thanks for sharing </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2671195,
      "author_name": "rpicatoste",
      "author_url": "",
      "post_date": "02/27/2024 11:30:43",
      "content": "<p>Thanks for the comparison, very interesting! </p>\n<p>From a kaggle newbie: Do you plan to then manipulate the data with the loading library, or do you load it with the fastest library and then convert it to your favorite one? (or the most capable one?). For example, I feel comfortable with pandas, and learning to do all the same things in polars is not very appealing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2671268,
          "author_name": "kononenko",
          "author_url": "",
          "post_date": "02/27/2024 12:21:32",
          "content": "<p>You can go with any of the options you mention, just additional benchmarks are needed as to how fast/memory hungry are conversions, <code>join</code> and other munging capabilities. </p>\n<p>One thing to keep in mind is that at the very end one needs to pass the resulting frames to modeling packages like LightGBM and those may not support all the possible frame types, so conversion to <code>pandas</code> or <code>arrow</code> may be necessary, that could take additional time.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2720508,
      "author_name": "wo281954971",
      "author_url": "",
      "post_date": "03/28/2024 12:00:18",
      "content": "<p>thank you share,very good !</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2642588": "For this competition we've got quite a lot of data: dozens of files with the overall size of ~26 GB. Loading data is going to be an intrinsic part of every kernel, so it is obviously important to minimize reading times as much as possible. \n\nAccording to the official dataset description, all files can be found in both CSV and Parquet formats. So what is the fastest way to read them? In [home-credit-faster-data-loading](https://www.kaggle.com/code/kononenko/home-credit-faster-data-loading/) notebook I did some benchmarks of a few popular packages: [pandas](https://pandas.pydata.org/docs/), [polars](https://pola.rs/) and [datatable](https://datatable.readthedocs.io/). In this post I will briefly summarize the main findings.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2099265%2Fe314f1bee56e388ecb7ec1b090aec82d%2Fpandas_polars_datatable.png?generation=1707791559484413&alt=media)\n\n**TL;DR:** use `polars` to read Parquet or `datatable` to load memory-mapped files almost instantly.\n\n# 1. Loading data from CSV files\n\nThe table below presents time taken by packages to load all the competition's data from the CSV files.\n\n| Package | Time to read all CSV files [s] |\n|----------|----------|\n| pandas    | 1100.03 |\n| polars      |142.73 |\n| datatable | 118.37 |\n\n<br>\n\nWe can also look at this benchmark file-wise for the first few files in the training dataset.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2099265%2Fb546572054224b09182d6e8be084afd5%2Fload_csv.png?generation=1707382589451428&alt=media)\n\nAs we see,`pandas` is pretty slow when it comes to load larger files. `datatable` is doing a little bit better than `polars`, though both these packages are quite fast demonstrating similar performance file-wise.\n\n# 2. Loading data from Parquet files\n\nLoading data from Parquet is much faster, because that's a binary format. We only compare `pandas` and `polars`, because `datatable` doesn't have a dedicated function to load `.parquet` relying on the `arrow` library.\n\n| Package | Time to read all Parquet files [s] |\n|----------|----------|\n| pandas    | 129.83 |\n| polars      |55.22 |\n\n<br>\n\nSimilarly, we can also plot file-wise comparison.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2099265%2Fdd96fa932bf9aeba7d17c60f7fc6de90%2Fload_parquet.png?generation=1707382782914363&alt=media)\n\nAs expected, both packages process Parquet much faster than CSV with `polars` demonstrating better performance. However, processing times are still non-negligible when it comes to loading all the data.\n\n# 3. Loading data with [memory mapping](https://en.wikipedia.org/wiki/Memory-mapped_file)\n\nOnce data is loaded into a dataframe, we can actually export it in a binary format appropriate for memory mapping: [ipc](https://docs.pola.rs/py-polars/html/reference/api/polars.DataFrame.write_ipc.html) for `polars` and [jay](https://datatable.readthedocs.io/en/latest/api/frame/to_jay.html) for `datatable`. Reading data back should literally take no time, because files are mapped into the memory address space very fast. To confirm that, I benchmarked `.ipc` and `.jay` loading for five first files from the competition's data.\n\n| Package | Time to memory map five files [s] |\n|----------|----------|\n| polars    | 1.25 |\n| datatable      | 0.04 |\n\n<br>\n\nThe measured timings are much smaller than even those for Parquet with `datatable` memory mapping being almost immediate. When comparing file-wise, one can see that `datatable`'s performance it almost file independent, while `polars` timings change quite a lot.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2099265%2F78f6b9f2c118f52878c4e01b83394008%2Fload_mmap.png?generation=1707382990401744&alt=media)\n\n# Conclusion\n\nTo sum up, CSV is the slowest way to get the data, reading Parquet is about twice faster and memory mapping could be almost instant. The final benchmarking results are also summarized in the table below. Note, we are not using actual numbers in this table, because those could vary on different systems from run to run.\n\n|               | pandas | polars | datatable |\n|---------------|--------|--------|-----------|\n| CSV           | Slow   | Fast   | Fast      | \n| Parquet       | OK     | Fast   | N/A       |\n| Memory mapping| N/A    | OK     | Fast      |\n\n<br>\n\nFor more details, please refer to [home-credit-faster-data-loading](https://www.kaggle.com/code/kononenko/home-credit-faster-data-loading/) notebook. Hope it helps and good luck with this awesome competition!",
    "2642745": "Impressive, thanks for the effort and comparison.",
    "2646893": "Hi, do you think loading data faster is crucial in this competion?",
    "2646926": "I think so, because if one makes an incorrect choice by sticking to `pandas` to read all the `.csv` files, it will take 1000+ seconds just to load the data, and this is something that needs to be done as often as we start a new session in our competition notebook...",
    "2646941": "But pandas and parquet would be ok in your opinion?",
    "2646946": "It is not that bad, see benchmarks [here](https://www.kaggle.com/code/kononenko/home-credit-faster-data-loading/#2.-Loading-data-from-Parquet-files), about 2 minutes to load all the data, `polars` would do it in 1 minute. So it is up to you to decide, if you need 50% speed-up or not.",
    "2647054": "Thank You dear Oleksiy @kononenko for Your comparison and getting me familiar with datatable! I will try to use them both!",
    "2652002": "Thanks for sharing",
    "2671195": "Thanks for the comparison, very interesting! \n\nFrom a kaggle newbie: Do you plan to then manipulate the data with the loading library, or do you load it with the fastest library and then convert it to your favorite one? (or the most capable one?). For example, I feel comfortable with pandas, and learning to do all the same things in polars is not very appealing.",
    "2671268": "You can go with any of the options you mention, just additional benchmarks are needed as to how fast/memory hungry are conversions, `join` and other munging capabilities. \n\nOne thing to keep in mind is that at the very end one needs to pass the resulting frames to modeling packages like LightGBM and those may not support all the possible frame types, so conversion to `pandas` or `arrow` may be necessary, that could take additional time.",
    "2720508": "thank you share,very good !"
  },
  "source": "meta"
}