{
  "id": 542997,
  "title": "Use Polars-SQL to improve your workflows",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/542997",
  "author_name": "",
  "post_date": "2024-10-28T06:53:09.775360600Z",
  "votes": 27,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Hello all,</p>\n<p>I am sure many of us are new to Polars and are finding it a little difficult to navigate though the basics. However, I am also certain that many of us have good SQL knowledge from our past learnings. Polars has an excellent SQL interface without any speed compromises and offers several mediums for a user to use SQL in expressions, context manager and globally too! <br>\nOne could use this to good effect to circumnavigate any syntactic issues one may face with native polars syntax. I have illustrated a few simple examples <a href=\"https://www.kaggle.com/code/ravi20076/janestreet2024-polarssqldemo\" target=\"_blank\">here</a> of various ways of using SQL in lazy-frames and eager-frames alike. </p>\n<p>Also, Polars released their latest version==1.1.12 yesterday (27Oct2024). Since this has no dependency issues, I am sure we could also collectively upgrade from Polars 1.9.0 in our Kaggle base environment to the latest version without issues. Wheel file is provided <a href=\"https://www.kaggle.com/code/ravi20076/janestreet2024-imports-v1\" target=\"_blank\">here</a> for ready usage. This has NVidia RAPIDS integration to use GPU while collecting lazy-frames to eager-frames. Just use the below code to speed-up your workflow- <br>\n<code>mylazyframe.collect(engine = \"gpu\")</code> and ensure you have your GPU on!</p>\n<p>SQL references from polars documentation are here below- <br>\n<a href=\"https://docs.pola.rs/py-polars/html/reference/sql/index.html\" target=\"_blank\">https://docs.pola.rs/py-polars/html/reference/sql/index.html</a></p>\n<p>NVIDIA integration with Polars is described here below-</p>\n<ul>\n<li><a href=\"https://pola.rs/posts/polars-on-gpu/\" target=\"_blank\">https://pola.rs/posts/polars-on-gpu/</a></li>\n<li><a href=\"https://developer.nvidia.com/blog/polars-gpu-engine-powered-by-rapids-cudf-now-available-in-open-beta/\" target=\"_blank\">https://developer.nvidia.com/blog/polars-gpu-engine-powered-by-rapids-cudf-now-available-in-open-beta/</a></li>\n<li><a href=\"https://rapids.ai/polars-gpu-engine/\" target=\"_blank\">https://rapids.ai/polars-gpu-engine/</a></li>\n<li><a href=\"https://www.datacamp.com/blog/polars-gpu-engine\" target=\"_blank\">https://www.datacamp.com/blog/polars-gpu-engine</a></li>\n<li><a href=\"https://docs.pola.rs/user-guide/lazy/gpu/\" target=\"_blank\">https://docs.pola.rs/user-guide/lazy/gpu/</a></li>\n</ul>\n<p>Best wishes and regards!</p>",
  "messages": [
    {
      "id": "3030094",
      "postDate": "10/28/2024 06:53:09",
      "content": "<p>Hello all,</p>\n<p>I am sure many of us are new to Polars and are finding it a little difficult to navigate though the basics. However, I am also certain that many of us have good SQL knowledge from our past learnings. Polars has an excellent SQL interface without any speed compromises and offers several mediums for a user to use SQL in expressions, context manager and globally too! <br>\nOne could use this to good effect to circumnavigate any syntactic issues one may face with native polars syntax. I have illustrated a few simple examples <a href=\"https://www.kaggle.com/code/ravi20076/janestreet2024-polarssqldemo\" target=\"_blank\">here</a> of various ways of using SQL in lazy-frames and eager-frames alike. </p>\n<p>Also, Polars released their latest version==1.1.12 yesterday (27Oct2024). Since this has no dependency issues, I am sure we could also collectively upgrade from Polars 1.9.0 in our Kaggle base environment to the latest version without issues. Wheel file is provided <a href=\"https://www.kaggle.com/code/ravi20076/janestreet2024-imports-v1\" target=\"_blank\">here</a> for ready usage. This has NVidia RAPIDS integration to use GPU while collecting lazy-frames to eager-frames. Just use the below code to speed-up your workflow- <br>\n<code>mylazyframe.collect(engine = \"gpu\")</code> and ensure you have your GPU on!</p>\n<p>SQL references from polars documentation are here below- <br>\n<a href=\"https://docs.pola.rs/py-polars/html/reference/sql/index.html\" target=\"_blank\">https://docs.pola.rs/py-polars/html/reference/sql/index.html</a></p>\n<p>NVIDIA integration with Polars is described here below-</p>\n<ul>\n<li><a href=\"https://pola.rs/posts/polars-on-gpu/\" target=\"_blank\">https://pola.rs/posts/polars-on-gpu/</a></li>\n<li><a href=\"https://developer.nvidia.com/blog/polars-gpu-engine-powered-by-rapids-cudf-now-available-in-open-beta/\" target=\"_blank\">https://developer.nvidia.com/blog/polars-gpu-engine-powered-by-rapids-cudf-now-available-in-open-beta/</a></li>\n<li><a href=\"https://rapids.ai/polars-gpu-engine/\" target=\"_blank\">https://rapids.ai/polars-gpu-engine/</a></li>\n<li><a href=\"https://www.datacamp.com/blog/polars-gpu-engine\" target=\"_blank\">https://www.datacamp.com/blog/polars-gpu-engine</a></li>\n<li><a href=\"https://docs.pola.rs/user-guide/lazy/gpu/\" target=\"_blank\">https://docs.pola.rs/user-guide/lazy/gpu/</a></li>\n</ul>\n<p>Best wishes and regards!</p>",
      "rawMarkdown": "Hello all,\n\nI am sure many of us are new to Polars and are finding it a little difficult to navigate though the basics. However, I am also certain that many of us have good SQL knowledge from our past learnings. Polars has an excellent SQL interface without any speed compromises and offers several mediums for a user to use SQL in expressions, context manager and globally too! \nOne could use this to good effect to circumnavigate any syntactic issues one may face with native polars syntax. I have illustrated a few simple examples [here](https://www.kaggle.com/code/ravi20076/janestreet2024-polarssqldemo) of various ways of using SQL in lazy-frames and eager-frames alike. \n\nAlso, Polars released their latest version==1.1.12 yesterday (27Oct2024). Since this has no dependency issues, I am sure we could also collectively upgrade from Polars 1.9.0 in our Kaggle base environment to the latest version without issues. Wheel file is provided [here](https://www.kaggle.com/code/ravi20076/janestreet2024-imports-v1) for ready usage. This has NVidia RAPIDS integration to use GPU while collecting lazy-frames to eager-frames. Just use the below code to speed-up your workflow- \n`mylazyframe.collect(engine = \"gpu\")` and ensure you have your GPU on!\n\nSQL references from polars documentation are here below- \nhttps://docs.pola.rs/py-polars/html/reference/sql/index.html\n\nNVIDIA integration with Polars is described here below-\n- https://pola.rs/posts/polars-on-gpu/\n- https://developer.nvidia.com/blog/polars-gpu-engine-powered-by-rapids-cudf-now-available-in-open-beta/\n- https://rapids.ai/polars-gpu-engine/\n- https://www.datacamp.com/blog/polars-gpu-engine\n- https://docs.pola.rs/user-guide/lazy/gpu/\n\nBest wishes and regards!",
      "votes": null
    },
    {
      "id": "3030860",
      "postDate": "10/29/2024 02:03:07",
      "content": "<p>interesting, but it makes sense. Do you have some experience how much improvement this makes for kaggle competitions?</p>",
      "rawMarkdown": "interesting, but it makes sense. Do you have some experience how much improvement this makes for kaggle competitions?",
      "votes": null
    },
    {
      "id": "3030880",
      "postDate": "10/29/2024 02:32:08",
      "content": "<p>This is an interesting tool. Thanks for sharing</p>",
      "rawMarkdown": "This is an interesting tool. Thanks for sharing",
      "votes": null
    },
    {
      "id": "3030967",
      "postDate": "10/29/2024 05:50:07",
      "content": "<p>In the early days of this competition, I tried to use pandas for predictions, and received timeout error. (It was basic FE)<br>\nReplacing it with polars really helped. I can't give you quantifiable measurement as to how much. But yeah it helped, since the data itself it 12GB.</p>",
      "rawMarkdown": "In the early days of this competition, I tried to use pandas for predictions, and received timeout error. (It was basic FE)\nReplacing it with polars really helped. I can't give you quantifiable measurement as to how much. But yeah it helped, since the data itself it 12GB.",
      "votes": null
    },
    {
      "id": "3031062",
      "postDate": "10/29/2024 09:33:21",
      "content": "<p>I use duckdb during EDA in a similar way.</p>\n<p>eg. create a view over the training data files:</p>\n<pre><code>duckdb.query(f)\n</code></pre>\n<p>And then continue with <strong>normal sql</strong>.  Very similar to <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a>'s examples - <em>but without \"limit\" guardrails</em> :)</p>\n<p>For example:</p>\n<pre><code>duckdb.query(f'''\n     \n        partition_id,\n        (*)  row_count,\n        ((date_id))  date_count, \n        (date_id)  min_date, \n        (date_id)  max_date,\n        ((symbol_id))  symbol_count\n     train_data\n      partition_id\n      partition_id\n    </code></pre>\n<table>\n<thead>\n<tr>\n<th>partition_id</th>\n<th>row_count</th>\n<th>date_count</th>\n<th>min_date</th>\n<th>max_date</th>\n<th>symbol_count</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>1944210</td>\n<td>170</td>\n<td>0</td>\n<td>169</td>\n<td>20</td>\n</tr>\n<tr>\n<td>1</td>\n<td>2804247</td>\n<td>170</td>\n<td>170</td>\n<td>339</td>\n<td>20</td>\n</tr>\n<tr>\n<td>2</td>\n<td>3036873</td>\n<td>170</td>\n<td>340</td>\n<td>509</td>\n<td>29</td>\n</tr>\n<tr>\n<td>3</td>\n<td>4016784</td>\n<td>170</td>\n<td>510</td>\n<td>679</td>\n<td>29</td>\n</tr>\n<tr>\n<td>4</td>\n<td>5022952</td>\n<td>170</td>\n<td>680</td>\n<td>849</td>\n<td>33</td>\n</tr>\n<tr>\n<td>5</td>\n<td>5348200</td>\n<td>170</td>\n<td>850</td>\n<td>1019</td>\n<td>36</td>\n</tr>\n<tr>\n<td>6</td>\n<td>6203912</td>\n<td>170</td>\n<td>1020</td>\n<td>1189</td>\n<td>39</td>\n</tr>\n<tr>\n<td>7</td>\n<td>6335560</td>\n<td>170</td>\n<td>1190</td>\n<td>1359</td>\n<td>39</td>\n</tr>\n<tr>\n<td>8</td>\n<td>6140024</td>\n<td>170</td>\n<td>1360</td>\n<td>1529</td>\n<td>39</td>\n</tr>\n<tr>\n<td>9</td>\n<td>6274576</td>\n<td>169</td>\n<td>1530</td>\n<td>1698</td>\n<td>39</td>\n</tr>\n</tbody>\n</table>\n<p>You can return the results as pandas (.df) or polars (.pl) data frames.</p>\n<p>Many queries execute quicker than you can load the parquet files using pandas.</p>\n<p>Polars / pandas can be quicker when you can trust the ordering of rows. A sql engine will have to sort.</p>\n<p>We don't need to worry (<em>too much</em>) about memory. Duckdb will not drag everything into RAM… unless you try to materialise a huge result set as a data frame :P </p>\n<p>Duckdb has a full SQL dialect with windowing, analytics, and some statistical functions. Polars SQL is a more limited subset of SQL.</p>\n<p>Easy things are easy and quick to test. <br>\neg Is weight is the same for all time_ids within a date_id?</p>\n<pre><code>duckdb.query(f)\n\n&gt; CPU times: user  s, sys:  ms, total:  s\n&gt; Wall time:  s\n&gt;  rows\n</code></pre>\n<p>I could write that in pandas, but it'd be a lot more faff. (<em>just to avoid memory problems</em>).</p>",
      "rawMarkdown": "I use duckdb during EDA in a similar way.\n\neg. create a view over the training data files:\n```\nduckdb.query(f'''\n    create or replace view train_data as\n    SELECT \n        *\n    FROM read_parquet('{train_path}', union_by_name=true) t\n    ''')\n```\n\nAnd then continue with **normal sql**.  Very similar to @ravi20076's examples - *but without \"limit\" guardrails* :)\n\nFor example:\n```\nduckdb.query(f'''\n    SELECT \n        partition_id,\n        count(*) as row_count,\n        count(distinct(date_id)) as date_count, \n        min(date_id) as min_date, \n        max(date_id) as max_date,\n        count(distinct(symbol_id)) as symbol_count\n    FROM train_data\n    group by partition_id\n    order by partition_id\n    ''').df()\n\n> CPU times: user 5.37 s, sys: 52.7 ms, total: 5.42 s\n> Wall time: 1.53 s\n```\n| partition_id | row_count | date_count | min_date | max_date | symbol_count |\n|---|---|---|---|---|---|\n| 0 | 1944210 | 170 | 0 | 169 | 20 |\n| 1 | 2804247 | 170 | 170 | 339 | 20 |\n| 2 | 3036873 | 170 | 340 | 509 | 29 |\n| 3 | 4016784 | 170 | 510 | 679 | 29 |\n| 4 | 5022952 | 170 | 680 | 849 | 33 |\n| 5 | 5348200 | 170 | 850 | 1019 | 36 |\n| 6 | 6203912 | 170 | 1020 | 1189 | 39 |\n| 7 | 6335560 | 170 | 1190 | 1359 | 39 |\n| 8 | 6140024 | 170 | 1360 | 1529 | 39 |\n| 9 | 6274576 | 169 | 1530 | 1698 | 39 |\n\nYou can return the results as pandas (.df) or polars (.pl) data frames.\n\nMany queries execute quicker than you can load the parquet files using pandas.\n\nPolars / pandas can be quicker when you can trust the ordering of rows. A sql engine will have to sort.\n\nWe don't need to worry (*too much*) about memory. Duckdb will not drag everything into RAM... unless you try to materialise a huge result set as a data frame :P \n\nDuckdb has a full SQL dialect with windowing, analytics, and some statistical functions. Polars SQL is a more limited subset of SQL.\n\nEasy things are easy and quick to test. \neg Is weight is the same for all time_ids within a date_id?\n\n```\nduckdb.query(f'''\n    SELECT \n        -- rows with any variation of weight within a date?\n        date_id, symbol_id\n    FROM train_data\n    GROUP BY date_id, symbol_id\n    HAVING stddev(weight) > 0\n    ''')\n\n> CPU times: user 3.56 s, sys: 62.8 ms, total: 3.62 s\n> Wall time: 1 s\n> 0 rows\n```\n\nI could write that in pandas, but it'd be a lot more faff. (*just to avoid memory problems*).",
      "votes": null
    },
    {
      "id": "3031249",
      "postDate": "10/29/2024 14:06:09",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/jayshrivastava\" target=\"_blank\">@jayshrivastava</a> , thanks for sharing. There is no doubt that Polars can process large datasets better than pandas. The question here is the difference between polars WITH GPU compared to polars WITHOUT GPU. </p>\n<p>If you have a ML algorithm implemented directly on polars, I believe there will be a measurable difference. If you have a streaming realtime use case and use polars for data engineering, I believe there will be a measurable difference. For a one-time processing of a large dataset: I don't know. That's why I am asking    </p>",
      "rawMarkdown": "Hi @jayshrivastava , thanks for sharing. There is no doubt that Polars can process large datasets better than pandas. The question here is the difference between polars WITH GPU compared to polars WITHOUT GPU. \n\nIf you have a ML algorithm implemented directly on polars, I believe there will be a measurable difference. If you have a streaming realtime use case and use polars for data engineering, I believe there will be a measurable difference. For a one-time processing of a large dataset: I don't know. That's why I am asking",
      "votes": null
    },
    {
      "id": "3031260",
      "postDate": "10/29/2024 14:13:37",
      "content": "<p>So interesting. Thanks for sharing this tool</p>",
      "rawMarkdown": "So interesting. Thanks for sharing this tool",
      "votes": null
    },
    {
      "id": "3031296",
      "postDate": "10/29/2024 14:47:26",
      "content": "<p>Agreed <a href=\"https://www.kaggle.com/paddykb\" target=\"_blank\">@paddykb</a> <br>\nI frequently use duckdb outside of Kaggle and shall use it here as well. SQL dialect in duckdb is way better than polars SQL </p>",
      "rawMarkdown": "Agreed @paddykb \nI frequently use duckdb outside of Kaggle and shall use it here as well. SQL dialect in duckdb is way better than polars SQL",
      "votes": null
    },
    {
      "id": "3031369",
      "postDate": "10/29/2024 16:13:26",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": null
    },
    {
      "id": "3031412",
      "postDate": "10/29/2024 17:00:22",
      "content": "<p>Since I started using polars, I never looked back. I almost don't use pandas at all now (unless I need to use another library that is not compatible with polars's DataFrame yet).<br>\nThanks for sharing! <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> </p>",
      "rawMarkdown": "Since I started using polars, I never looked back. I almost don't use pandas at all now (unless I need to use another library that is not compatible with polars's DataFrame yet).\nThanks for sharing! @ravi20076",
      "votes": null
    },
    {
      "id": "3031452",
      "postDate": "10/29/2024 17:39:40",
      "content": "<p>Thanks for sharing this useful information </p>",
      "rawMarkdown": "Thanks for sharing this useful information",
      "votes": null
    },
    {
      "id": "3031938",
      "postDate": "10/30/2024 10:25:13",
      "content": "<p>Thanks for sharing this useful information</p>",
      "rawMarkdown": "Thanks for sharing this useful information",
      "votes": null
    },
    {
      "id": "3033038",
      "postDate": "10/31/2024 17:41:42",
      "content": "<p>It's great to hear this</p>",
      "rawMarkdown": "It's great to hear this",
      "votes": null
    },
    {
      "id": "3033086",
      "postDate": "10/31/2024 18:27:23",
      "content": "<p>Thanks a lot for this! Just recently came across polars, definitely will keep this in mind</p>",
      "rawMarkdown": "Thanks a lot for this! Just recently came across polars, definitely will keep this in mind",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3030860,
      "author_name": "fredericnicholson",
      "author_url": "",
      "post_date": "10/29/2024 02:03:07",
      "content": "<p>interesting, but it makes sense. Do you have some experience how much improvement this makes for kaggle competitions?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3030967,
          "author_name": "jayshrivastava",
          "author_url": "",
          "post_date": "10/29/2024 05:50:07",
          "content": "<p>In the early days of this competition, I tried to use pandas for predictions, and received timeout error. (It was basic FE)<br>\nReplacing it with polars really helped. I can't give you quantifiable measurement as to how much. But yeah it helped, since the data itself it 12GB.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3031249,
              "author_name": "fredericnicholson",
              "author_url": "",
              "post_date": "10/29/2024 14:06:09",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/jayshrivastava\" target=\"_blank\">@jayshrivastava</a> , thanks for sharing. There is no doubt that Polars can process large datasets better than pandas. The question here is the difference between polars WITH GPU compared to polars WITHOUT GPU. </p>\n<p>If you have a ML algorithm implemented directly on polars, I believe there will be a measurable difference. If you have a streaming realtime use case and use polars for data engineering, I believe there will be a measurable difference. For a one-time processing of a large dataset: I don't know. That's why I am asking    </p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 3031062,
          "author_name": "paddykb",
          "author_url": "",
          "post_date": "10/29/2024 09:33:21",
          "content": "<p>I use duckdb during EDA in a similar way.</p>\n<p>eg. create a view over the training data files:</p>\n<pre><code>duckdb.query(f)\n</code></pre>\n<p>And then continue with <strong>normal sql</strong>.  Very similar to <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a>'s examples - <em>but without \"limit\" guardrails</em> :)</p>\n<p>For example:</p>\n<pre><code>duckdb.query(f'''\n     \n        partition_id,\n        (*)  row_count,\n        ((date_id))  date_count, \n        (date_id)  min_date, \n        (date_id)  max_date,\n        ((symbol_id))  symbol_count\n     train_data\n      partition_id\n      partition_id\n    </code></pre>\n<table>\n<thead>\n<tr>\n<th>partition_id</th>\n<th>row_count</th>\n<th>date_count</th>\n<th>min_date</th>\n<th>max_date</th>\n<th>symbol_count</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>1944210</td>\n<td>170</td>\n<td>0</td>\n<td>169</td>\n<td>20</td>\n</tr>\n<tr>\n<td>1</td>\n<td>2804247</td>\n<td>170</td>\n<td>170</td>\n<td>339</td>\n<td>20</td>\n</tr>\n<tr>\n<td>2</td>\n<td>3036873</td>\n<td>170</td>\n<td>340</td>\n<td>509</td>\n<td>29</td>\n</tr>\n<tr>\n<td>3</td>\n<td>4016784</td>\n<td>170</td>\n<td>510</td>\n<td>679</td>\n<td>29</td>\n</tr>\n<tr>\n<td>4</td>\n<td>5022952</td>\n<td>170</td>\n<td>680</td>\n<td>849</td>\n<td>33</td>\n</tr>\n<tr>\n<td>5</td>\n<td>5348200</td>\n<td>170</td>\n<td>850</td>\n<td>1019</td>\n<td>36</td>\n</tr>\n<tr>\n<td>6</td>\n<td>6203912</td>\n<td>170</td>\n<td>1020</td>\n<td>1189</td>\n<td>39</td>\n</tr>\n<tr>\n<td>7</td>\n<td>6335560</td>\n<td>170</td>\n<td>1190</td>\n<td>1359</td>\n<td>39</td>\n</tr>\n<tr>\n<td>8</td>\n<td>6140024</td>\n<td>170</td>\n<td>1360</td>\n<td>1529</td>\n<td>39</td>\n</tr>\n<tr>\n<td>9</td>\n<td>6274576</td>\n<td>169</td>\n<td>1530</td>\n<td>1698</td>\n<td>39</td>\n</tr>\n</tbody>\n</table>\n<p>You can return the results as pandas (.df) or polars (.pl) data frames.</p>\n<p>Many queries execute quicker than you can load the parquet files using pandas.</p>\n<p>Polars / pandas can be quicker when you can trust the ordering of rows. A sql engine will have to sort.</p>\n<p>We don't need to worry (<em>too much</em>) about memory. Duckdb will not drag everything into RAM… unless you try to materialise a huge result set as a data frame :P </p>\n<p>Duckdb has a full SQL dialect with windowing, analytics, and some statistical functions. Polars SQL is a more limited subset of SQL.</p>\n<p>Easy things are easy and quick to test. <br>\neg Is weight is the same for all time_ids within a date_id?</p>\n<pre><code>duckdb.query(f)\n\n&gt; CPU times: user  s, sys:  ms, total:  s\n&gt; Wall time:  s\n&gt;  rows\n</code></pre>\n<p>I could write that in pandas, but it'd be a lot more faff. (<em>just to avoid memory problems</em>).</p>",
          "votes": null,
          "replies": [
            {
              "id": 3031296,
              "author_name": "ravi20076",
              "author_url": "",
              "post_date": "10/29/2024 14:47:26",
              "content": "<p>Agreed <a href=\"https://www.kaggle.com/paddykb\" target=\"_blank\">@paddykb</a> <br>\nI frequently use duckdb outside of Kaggle and shall use it here as well. SQL dialect in duckdb is way better than polars SQL </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3030880,
      "author_name": "skylord",
      "author_url": "",
      "post_date": "10/29/2024 02:32:08",
      "content": "<p>This is an interesting tool. Thanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3031260,
      "author_name": "gabrielabilleira",
      "author_url": "",
      "post_date": "10/29/2024 14:13:37",
      "content": "<p>So interesting. Thanks for sharing this tool</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3031369,
      "author_name": "humayrakhanomrime",
      "author_url": "",
      "post_date": "10/29/2024 16:13:26",
      "content": "<p>Thanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3031412,
      "author_name": "dinezra11",
      "author_url": "",
      "post_date": "10/29/2024 17:00:22",
      "content": "<p>Since I started using polars, I never looked back. I almost don't use pandas at all now (unless I need to use another library that is not compatible with polars's DataFrame yet).<br>\nThanks for sharing! <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3031452,
      "author_name": "anitarostami",
      "author_url": "",
      "post_date": "10/29/2024 17:39:40",
      "content": "<p>Thanks for sharing this useful information </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3031938,
      "author_name": "mrsimple07",
      "author_url": "",
      "post_date": "10/30/2024 10:25:13",
      "content": "<p>Thanks for sharing this useful information</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3033038,
      "author_name": "farhankardan",
      "author_url": "",
      "post_date": "10/31/2024 17:41:42",
      "content": "<p>It's great to hear this</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3033086,
      "author_name": "advaithsanilkumar",
      "author_url": "",
      "post_date": "10/31/2024 18:27:23",
      "content": "<p>Thanks a lot for this! Just recently came across polars, definitely will keep this in mind</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3030094": "Hello all,\n\nI am sure many of us are new to Polars and are finding it a little difficult to navigate though the basics. However, I am also certain that many of us have good SQL knowledge from our past learnings. Polars has an excellent SQL interface without any speed compromises and offers several mediums for a user to use SQL in expressions, context manager and globally too! \nOne could use this to good effect to circumnavigate any syntactic issues one may face with native polars syntax. I have illustrated a few simple examples [here](https://www.kaggle.com/code/ravi20076/janestreet2024-polarssqldemo) of various ways of using SQL in lazy-frames and eager-frames alike. \n\nAlso, Polars released their latest version==1.1.12 yesterday (27Oct2024). Since this has no dependency issues, I am sure we could also collectively upgrade from Polars 1.9.0 in our Kaggle base environment to the latest version without issues. Wheel file is provided [here](https://www.kaggle.com/code/ravi20076/janestreet2024-imports-v1) for ready usage. This has NVidia RAPIDS integration to use GPU while collecting lazy-frames to eager-frames. Just use the below code to speed-up your workflow- \n`mylazyframe.collect(engine = \"gpu\")` and ensure you have your GPU on!\n\nSQL references from polars documentation are here below- \nhttps://docs.pola.rs/py-polars/html/reference/sql/index.html\n\nNVIDIA integration with Polars is described here below-\n- https://pola.rs/posts/polars-on-gpu/\n- https://developer.nvidia.com/blog/polars-gpu-engine-powered-by-rapids-cudf-now-available-in-open-beta/\n- https://rapids.ai/polars-gpu-engine/\n- https://www.datacamp.com/blog/polars-gpu-engine\n- https://docs.pola.rs/user-guide/lazy/gpu/\n\nBest wishes and regards!",
    "3030860": "interesting, but it makes sense. Do you have some experience how much improvement this makes for kaggle competitions?",
    "3030880": "This is an interesting tool. Thanks for sharing",
    "3030967": "In the early days of this competition, I tried to use pandas for predictions, and received timeout error. (It was basic FE)\nReplacing it with polars really helped. I can't give you quantifiable measurement as to how much. But yeah it helped, since the data itself it 12GB.",
    "3031062": "I use duckdb during EDA in a similar way.\n\neg. create a view over the training data files:\n```\nduckdb.query(f'''\n    create or replace view train_data as\n    SELECT \n        *\n    FROM read_parquet('{train_path}', union_by_name=true) t\n    ''')\n```\n\nAnd then continue with **normal sql**.  Very similar to @ravi20076's examples - *but without \"limit\" guardrails* :)\n\nFor example:\n```\nduckdb.query(f'''\n    SELECT \n        partition_id,\n        count(*) as row_count,\n        count(distinct(date_id)) as date_count, \n        min(date_id) as min_date, \n        max(date_id) as max_date,\n        count(distinct(symbol_id)) as symbol_count\n    FROM train_data\n    group by partition_id\n    order by partition_id\n    ''').df()\n\n> CPU times: user 5.37 s, sys: 52.7 ms, total: 5.42 s\n> Wall time: 1.53 s\n```\n| partition_id | row_count | date_count | min_date | max_date | symbol_count |\n|---|---|---|---|---|---|\n| 0 | 1944210 | 170 | 0 | 169 | 20 |\n| 1 | 2804247 | 170 | 170 | 339 | 20 |\n| 2 | 3036873 | 170 | 340 | 509 | 29 |\n| 3 | 4016784 | 170 | 510 | 679 | 29 |\n| 4 | 5022952 | 170 | 680 | 849 | 33 |\n| 5 | 5348200 | 170 | 850 | 1019 | 36 |\n| 6 | 6203912 | 170 | 1020 | 1189 | 39 |\n| 7 | 6335560 | 170 | 1190 | 1359 | 39 |\n| 8 | 6140024 | 170 | 1360 | 1529 | 39 |\n| 9 | 6274576 | 169 | 1530 | 1698 | 39 |\n\nYou can return the results as pandas (.df) or polars (.pl) data frames.\n\nMany queries execute quicker than you can load the parquet files using pandas.\n\nPolars / pandas can be quicker when you can trust the ordering of rows. A sql engine will have to sort.\n\nWe don't need to worry (*too much*) about memory. Duckdb will not drag everything into RAM... unless you try to materialise a huge result set as a data frame :P \n\nDuckdb has a full SQL dialect with windowing, analytics, and some statistical functions. Polars SQL is a more limited subset of SQL.\n\nEasy things are easy and quick to test. \neg Is weight is the same for all time_ids within a date_id?\n\n```\nduckdb.query(f'''\n    SELECT \n        -- rows with any variation of weight within a date?\n        date_id, symbol_id\n    FROM train_data\n    GROUP BY date_id, symbol_id\n    HAVING stddev(weight) > 0\n    ''')\n\n> CPU times: user 3.56 s, sys: 62.8 ms, total: 3.62 s\n> Wall time: 1 s\n> 0 rows\n```\n\nI could write that in pandas, but it'd be a lot more faff. (*just to avoid memory problems*).",
    "3031249": "Hi @jayshrivastava , thanks for sharing. There is no doubt that Polars can process large datasets better than pandas. The question here is the difference between polars WITH GPU compared to polars WITHOUT GPU. \n\nIf you have a ML algorithm implemented directly on polars, I believe there will be a measurable difference. If you have a streaming realtime use case and use polars for data engineering, I believe there will be a measurable difference. For a one-time processing of a large dataset: I don't know. That's why I am asking",
    "3031260": "So interesting. Thanks for sharing this tool",
    "3031296": "Agreed @paddykb \nI frequently use duckdb outside of Kaggle and shall use it here as well. SQL dialect in duckdb is way better than polars SQL",
    "3031369": "Thanks for sharing",
    "3031412": "Since I started using polars, I never looked back. I almost don't use pandas at all now (unless I need to use another library that is not compatible with polars's DataFrame yet).\nThanks for sharing! @ravi20076",
    "3031452": "Thanks for sharing this useful information",
    "3031938": "Thanks for sharing this useful information",
    "3033038": "It's great to hear this",
    "3033086": "Thanks a lot for this! Just recently came across polars, definitely will keep this in mind"
  },
  "source": "meta"
}