{
  "id": 543371,
  "title": "Why Pandas? ",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/543371",
  "author_name": "",
  "post_date": "2024-10-30T06:42:56.596316400Z",
  "votes": 2,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I had never heard of the polars library prior to this competition, so I got curious and found <a href=\"https://pola.rs/posts/benchmarks/\" target=\"_blank\">these benchmarks</a>. Question is, why is pandas so widespread if it's abysmally slower than other libraries? Are we just too accustomed to the library to make the swap? Does pandas have any inherent capabilities that are superior to polars?</p>",
  "messages": [
    {
      "id": "3031813",
      "postDate": "10/30/2024 06:42:56",
      "content": "<p>I had never heard of the polars library prior to this competition, so I got curious and found <a href=\"https://pola.rs/posts/benchmarks/\" target=\"_blank\">these benchmarks</a>. Question is, why is pandas so widespread if it's abysmally slower than other libraries? Are we just too accustomed to the library to make the swap? Does pandas have any inherent capabilities that are superior to polars?</p>",
      "rawMarkdown": "I had never heard of the polars library prior to this competition, so I got curious and found [these benchmarks](https://pola.rs/posts/benchmarks/). Question is, why is pandas so widespread if it's abysmally slower than other libraries? Are we just too accustomed to the library to make the swap? Does pandas have any inherent capabilities that are superior to polars?",
      "votes": null
    },
    {
      "id": "3031866",
      "postDate": "10/30/2024 08:13:02",
      "content": "<p>Benchmark gaming aside: pit single-threaded (pandas) vs multi-threaded (polars) on a 22-core machine. </p>\n<p>The sweet spot for both libraries is medium data.  <strong>Wes McKinney</strong>: Medium Data (n): <em>Not too big for a single machine, but too big to be dumb about.</em></p>\n<p>Pandas has been the standard in this space for a long time. I don't like lots about it, but It's pretty versatile, robust, reliable, and has the power of incumbency. Call it Stockholm syndrome, but I will not be switching over lots of existing code - which means maintaining a level of expertise in pandas.</p>\n<p>But for new projects?  That gets interesting. Polars has an appeal in the \"medium space\", but I'm interested in scaling seamlessly from small to big data. </p>\n<p>I'm leaning more towards <a href=\"https://ibis-project.org/why\" target=\"_blank\">Ibis</a>. Ibis provides a consistent API across different backends (including DuckDB, spark and Polars). We're not locked into a single-machine execution engine and can scale from local to production.</p>",
      "rawMarkdown": "Benchmark gaming aside: pit single-threaded (pandas) vs multi-threaded (polars) on a 22-core machine. \n\nThe sweet spot for both libraries is medium data.  **Wes McKinney**: Medium Data (n): *Not too big for a single machine, but too big to be dumb about.*\n\nPandas has been the standard in this space for a long time. I don't like lots about it, but It's pretty versatile, robust, reliable, and has the power of incumbency. Call it Stockholm syndrome, but I will not be switching over lots of existing code - which means maintaining a level of expertise in pandas.\n\nBut for new projects?  That gets interesting. Polars has an appeal in the \"medium space\", but I'm interested in scaling seamlessly from small to big data. \n\nI'm leaning more towards [Ibis](https://ibis-project.org/why). Ibis provides a consistent API across different backends (including DuckDB, spark and Polars). We're not locked into a single-machine execution engine and can scale from local to production.",
      "votes": null
    },
    {
      "id": "3031885",
      "postDate": "10/30/2024 08:47:09",
      "content": "<p>Pandas has been in the game for so long that any ML library is compatible with its data structures (DataFrame, Series), so it's very comfortable to use it with other libraries. While today most of the libraries also work well with polars, there are still some that aren't. There are some preprocessing libraries and methods that only work with the pandas DataFrame. Another reason is legacy code - for older projects it may be hard and not worth it to refactor and change everything into polars. But I think the primary reason is the people are just afraid to learn new library. The syntax is a bit different to pandas and as you said, people are too accustomed to it to make the swap.</p>\n<p>For me, switching to polars was great. I find the syntax to be very comfortable (reminds me of SQL, which I really like), so I rarely use pandas these days. I use pandas only for older projects and when I need to be able to use libraries that can't work with polars.</p>",
      "rawMarkdown": "Pandas has been in the game for so long that any ML library is compatible with its data structures (DataFrame, Series), so it's very comfortable to use it with other libraries. While today most of the libraries also work well with polars, there are still some that aren't. There are some preprocessing libraries and methods that only work with the pandas DataFrame. Another reason is legacy code - for older projects it may be hard and not worth it to refactor and change everything into polars. But I think the primary reason is the people are just afraid to learn new library. The syntax is a bit different to pandas and as you said, people are too accustomed to it to make the swap.\n\nFor me, switching to polars was great. I find the syntax to be very comfortable (reminds me of SQL, which I really like), so I rarely use pandas these days. I use pandas only for older projects and when I need to be able to use libraries that can't work with polars.",
      "votes": null
    },
    {
      "id": "3031901",
      "postDate": "10/30/2024 09:16:48",
      "content": "<p>Polars is relatively new… it lacks some capabilities related to pandas and its support is not widespread in existing libraries. But its getting traction.</p>",
      "rawMarkdown": "Polars is relatively new... it lacks some capabilities related to pandas and its support is not widespread in existing libraries. But its getting traction.",
      "votes": null
    },
    {
      "id": "3032619",
      "postDate": "10/31/2024 06:43:53",
      "content": "<p>this is my first time using polars too, i will name a few things that makes me not switch to polars:</p>\n<ul>\n<li>Learning curve and syntax differences though it's not difficult to switch but it has a Less extensive documentation and examples .</li>\n<li>Doesn't support fp16 and handles categorical types poorly.</li>\n<li>It requires conversion to use with some ML libraries.</li>\n<li>Lacks native plotting support, requiring you to convert to Pandas or use third-party libraries.</li>\n<li>Iterative and row-wise operations are less intuitive with polars.</li>\n</ul>",
      "rawMarkdown": "this is my first time using polars too, i will name a few things that makes me not switch to polars:\n- Learning curve and syntax differences though it's not difficult to switch but it has a Less extensive documentation and examples .\n- Doesn't support fp16 and handles categorical types poorly.\n- It requires conversion to use with some ML libraries.\n- Lacks native plotting support, requiring you to convert to Pandas or use third-party libraries.\n- Iterative and row-wise operations are less intuitive with polars.",
      "votes": null
    },
    {
      "id": "3033124",
      "postDate": "10/31/2024 19:07:12",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/paddykb\" target=\"_blank\">@paddykb</a> </p>\n<p>Many thanks for the heads-up regarding Ibis. </p>\n<p>I did notice whilst perusing the documentation that </p>\n<blockquote>\n  <p>…matplotlib has not implemented the dataframe interchange protocol so it is recommended to call to_pandas() on the Ibis table before plotting. <a href=\"https://ibis-project.org/how-to/visualization/matplotlib#using-matplotlib-with-ibis\" target=\"_blank\">[Ref]</a></p>\n</blockquote>\n<p>so it could be  a little clunky when it comes to performing extensive EDA.</p>\n<p>All the best,<br>\ncarl</p>",
      "rawMarkdown": "Dear @paddykb \n\nMany thanks for the heads-up regarding Ibis. \n\nI did notice whilst perusing the documentation that \n\n> ...matplotlib has not implemented the dataframe interchange protocol so it is recommended to call to_pandas() on the Ibis table before plotting. [[Ref]](https://ibis-project.org/how-to/visualization/matplotlib#using-matplotlib-with-ibis)\n\nso it could be  a little clunky when it comes to performing extensive EDA.\n\nAll the best,\ncarl",
      "votes": null
    },
    {
      "id": "3042071",
      "postDate": "11/11/2024 06:35:08",
      "content": "<p>Inplace assignments like this <code>df.loc[validation_mask, 'prediction'] = df.loc[validation_mask, 'prediction'].clip(lower=-5, upper=5)</code> are slower in polars and they return a copy of the original dataframe when you use <code>df.with_columns</code> so the memory usage spikes when execution reaches that point. I'm using polars for faster joins and feature engineering but then I cast it to pandas dataframe before my training loop.</p>",
      "rawMarkdown": "Inplace assignments like this `df.loc[validation_mask, 'prediction'] = df.loc[validation_mask, 'prediction'].clip(lower=-5, upper=5)` are slower in polars and they return a copy of the original dataframe when you use `df.with_columns` so the memory usage spikes when execution reaches that point. I'm using polars for faster joins and feature engineering but then I cast it to pandas dataframe before my training loop.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3031866,
      "author_name": "paddykb",
      "author_url": "",
      "post_date": "10/30/2024 08:13:02",
      "content": "<p>Benchmark gaming aside: pit single-threaded (pandas) vs multi-threaded (polars) on a 22-core machine. </p>\n<p>The sweet spot for both libraries is medium data.  <strong>Wes McKinney</strong>: Medium Data (n): <em>Not too big for a single machine, but too big to be dumb about.</em></p>\n<p>Pandas has been the standard in this space for a long time. I don't like lots about it, but It's pretty versatile, robust, reliable, and has the power of incumbency. Call it Stockholm syndrome, but I will not be switching over lots of existing code - which means maintaining a level of expertise in pandas.</p>\n<p>But for new projects?  That gets interesting. Polars has an appeal in the \"medium space\", but I'm interested in scaling seamlessly from small to big data. </p>\n<p>I'm leaning more towards <a href=\"https://ibis-project.org/why\" target=\"_blank\">Ibis</a>. Ibis provides a consistent API across different backends (including DuckDB, spark and Polars). We're not locked into a single-machine execution engine and can scale from local to production.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3033124,
          "author_name": "carlmcbrideellis",
          "author_url": "",
          "post_date": "10/31/2024 19:07:12",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/paddykb\" target=\"_blank\">@paddykb</a> </p>\n<p>Many thanks for the heads-up regarding Ibis. </p>\n<p>I did notice whilst perusing the documentation that </p>\n<blockquote>\n  <p>…matplotlib has not implemented the dataframe interchange protocol so it is recommended to call to_pandas() on the Ibis table before plotting. <a href=\"https://ibis-project.org/how-to/visualization/matplotlib#using-matplotlib-with-ibis\" target=\"_blank\">[Ref]</a></p>\n</blockquote>\n<p>so it could be  a little clunky when it comes to performing extensive EDA.</p>\n<p>All the best,<br>\ncarl</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3031885,
      "author_name": "dinezra11",
      "author_url": "",
      "post_date": "10/30/2024 08:47:09",
      "content": "<p>Pandas has been in the game for so long that any ML library is compatible with its data structures (DataFrame, Series), so it's very comfortable to use it with other libraries. While today most of the libraries also work well with polars, there are still some that aren't. There are some preprocessing libraries and methods that only work with the pandas DataFrame. Another reason is legacy code - for older projects it may be hard and not worth it to refactor and change everything into polars. But I think the primary reason is the people are just afraid to learn new library. The syntax is a bit different to pandas and as you said, people are too accustomed to it to make the swap.</p>\n<p>For me, switching to polars was great. I find the syntax to be very comfortable (reminds me of SQL, which I really like), so I rarely use pandas these days. I use pandas only for older projects and when I need to be able to use libraries that can't work with polars.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3031901,
      "author_name": "lucasmorin",
      "author_url": "",
      "post_date": "10/30/2024 09:16:48",
      "content": "<p>Polars is relatively new… it lacks some capabilities related to pandas and its support is not widespread in existing libraries. But its getting traction.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3032619,
      "author_name": "younesbenalia",
      "author_url": "",
      "post_date": "10/31/2024 06:43:53",
      "content": "<p>this is my first time using polars too, i will name a few things that makes me not switch to polars:</p>\n<ul>\n<li>Learning curve and syntax differences though it's not difficult to switch but it has a Less extensive documentation and examples .</li>\n<li>Doesn't support fp16 and handles categorical types poorly.</li>\n<li>It requires conversion to use with some ML libraries.</li>\n<li>Lacks native plotting support, requiring you to convert to Pandas or use third-party libraries.</li>\n<li>Iterative and row-wise operations are less intuitive with polars.</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3042071,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "11/11/2024 06:35:08",
      "content": "<p>Inplace assignments like this <code>df.loc[validation_mask, 'prediction'] = df.loc[validation_mask, 'prediction'].clip(lower=-5, upper=5)</code> are slower in polars and they return a copy of the original dataframe when you use <code>df.with_columns</code> so the memory usage spikes when execution reaches that point. I'm using polars for faster joins and feature engineering but then I cast it to pandas dataframe before my training loop.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3031813": "I had never heard of the polars library prior to this competition, so I got curious and found [these benchmarks](https://pola.rs/posts/benchmarks/). Question is, why is pandas so widespread if it's abysmally slower than other libraries? Are we just too accustomed to the library to make the swap? Does pandas have any inherent capabilities that are superior to polars?",
    "3031866": "Benchmark gaming aside: pit single-threaded (pandas) vs multi-threaded (polars) on a 22-core machine. \n\nThe sweet spot for both libraries is medium data.  **Wes McKinney**: Medium Data (n): *Not too big for a single machine, but too big to be dumb about.*\n\nPandas has been the standard in this space for a long time. I don't like lots about it, but It's pretty versatile, robust, reliable, and has the power of incumbency. Call it Stockholm syndrome, but I will not be switching over lots of existing code - which means maintaining a level of expertise in pandas.\n\nBut for new projects?  That gets interesting. Polars has an appeal in the \"medium space\", but I'm interested in scaling seamlessly from small to big data. \n\nI'm leaning more towards [Ibis](https://ibis-project.org/why). Ibis provides a consistent API across different backends (including DuckDB, spark and Polars). We're not locked into a single-machine execution engine and can scale from local to production.",
    "3031885": "Pandas has been in the game for so long that any ML library is compatible with its data structures (DataFrame, Series), so it's very comfortable to use it with other libraries. While today most of the libraries also work well with polars, there are still some that aren't. There are some preprocessing libraries and methods that only work with the pandas DataFrame. Another reason is legacy code - for older projects it may be hard and not worth it to refactor and change everything into polars. But I think the primary reason is the people are just afraid to learn new library. The syntax is a bit different to pandas and as you said, people are too accustomed to it to make the swap.\n\nFor me, switching to polars was great. I find the syntax to be very comfortable (reminds me of SQL, which I really like), so I rarely use pandas these days. I use pandas only for older projects and when I need to be able to use libraries that can't work with polars.",
    "3031901": "Polars is relatively new... it lacks some capabilities related to pandas and its support is not widespread in existing libraries. But its getting traction.",
    "3032619": "this is my first time using polars too, i will name a few things that makes me not switch to polars:\n- Learning curve and syntax differences though it's not difficult to switch but it has a Less extensive documentation and examples .\n- Doesn't support fp16 and handles categorical types poorly.\n- It requires conversion to use with some ML libraries.\n- Lacks native plotting support, requiring you to convert to Pandas or use third-party libraries.\n- Iterative and row-wise operations are less intuitive with polars.",
    "3033124": "Dear @paddykb \n\nMany thanks for the heads-up regarding Ibis. \n\nI did notice whilst perusing the documentation that \n\n> ...matplotlib has not implemented the dataframe interchange protocol so it is recommended to call to_pandas() on the Ibis table before plotting. [[Ref]](https://ibis-project.org/how-to/visualization/matplotlib#using-matplotlib-with-ibis)\n\nso it could be  a little clunky when it comes to performing extensive EDA.\n\nAll the best,\ncarl",
    "3042071": "Inplace assignments like this `df.loc[validation_mask, 'prediction'] = df.loc[validation_mask, 'prediction'].clip(lower=-5, upper=5)` are slower in polars and they return a copy of the original dataframe when you use `df.with_columns` so the memory usage spikes when execution reaches that point. I'm using polars for faster joins and feature engineering but then I cast it to pandas dataframe before my training loop."
  },
  "source": "meta"
}