{
  "id": 540739,
  "title": "Consider using your GPU using NVidia RAPIDS + Polars framework for data wrangling",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/540739",
  "author_name": "Ravi Ramakrishnan",
  "post_date": "2024-10-15T21:50:56.103000",
  "votes": 4,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hello all,</p>\n<p>I presume we are all collectively using polars for feature engineering and data loads herewith. I am sure we are also facing a few problems with data loads as the data is quite huge here. </p>\n<p>Recently, polars integrated with NVidia RAPIDS framework to improve the speed of data wrangling using GPUs without explicit code changes as well. References for this are as below-</p>\n<ul>\n<li><a href=\"https://pola.rs/posts/polars-on-gpu/\" target=\"_blank\">https://pola.rs/posts/polars-on-gpu/</a></li>\n<li><a href=\"https://pola.rs/posts/gpu-engine-release/\" target=\"_blank\">https://pola.rs/posts/gpu-engine-release/</a></li>\n<li><a href=\"https://developer.nvidia.com/blog/polars-gpu-engine-powered-by-rapids-cudf-now-available-in-open-beta/\" target=\"_blank\">https://developer.nvidia.com/blog/polars-gpu-engine-powered-by-rapids-cudf-now-available-in-open-beta/</a></li>\n<li><a href=\"https://www.datacamp.com/blog/polars-gpu-engine\" target=\"_blank\">https://www.datacamp.com/blog/polars-gpu-engine</a></li>\n<li><a href=\"https://www.youtube.com/watch?v=SFdKFhOjVLQ\" target=\"_blank\">https://www.youtube.com/watch?v=SFdKFhOjVLQ</a></li>\n<li><a href=\"https://rapids.ai/polars-gpu-engine/\" target=\"_blank\">https://rapids.ai/polars-gpu-engine/</a></li>\n</ul>\n<p>One simply needs to install polars with GPU as below- <br>\n<code>!pip install --extra-index-url=https://pypi.nvidia.com polars[gpu]==1.9.0 -q -d /kaggle/working/polars</code></p>\n<p>While collecting from lazy-frame to eager-frames, one needs to specify the engine as GPU as below-<br>\n<code>df.collect(engine = \"gpu\")</code></p>\n<p>This has the potential to speed up your features tremendously! Are you going to try this in your pipelines?<br>\nUpdated wheel file is present here-<br>\n<a href=\"https://www.kaggle.com/code/ravi20076/janestreet2024-imports-v1\" target=\"_blank\">https://www.kaggle.com/code/ravi20076/janestreet2024-imports-v1</a></p>\n<p>Usage of FE with GPU is here-<br>\n<a href=\"https://www.kaggle.com/code/ravi20076/janestreet2024-baseline-train-v1\" target=\"_blank\">https://www.kaggle.com/code/ravi20076/janestreet2024-baseline-train-v1</a></p>\n<p>Best regards!</p>",
  "messages": [
    {
      "id": 3018595,
      "postDate": "2024-10-15T21:50:56.103Z",
      "content": "<p>Hello all,</p>\n<p>I presume we are all collectively using polars for feature engineering and data loads herewith. I am sure we are also facing a few problems with data loads as the data is quite huge here. </p>\n<p>Recently, polars integrated with NVidia RAPIDS framework to improve the speed of data wrangling using GPUs without explicit code changes as well. References for this are as below-</p>\n<ul>\n<li><a href=\"https://pola.rs/posts/polars-on-gpu/\" target=\"_blank\">https://pola.rs/posts/polars-on-gpu/</a></li>\n<li><a href=\"https://pola.rs/posts/gpu-engine-release/\" target=\"_blank\">https://pola.rs/posts/gpu-engine-release/</a></li>\n<li><a href=\"https://developer.nvidia.com/blog/polars-gpu-engine-powered-by-rapids-cudf-now-available-in-open-beta/\" target=\"_blank\">https://developer.nvidia.com/blog/polars-gpu-engine-powered-by-rapids-cudf-now-available-in-open-beta/</a></li>\n<li><a href=\"https://www.datacamp.com/blog/polars-gpu-engine\" target=\"_blank\">https://www.datacamp.com/blog/polars-gpu-engine</a></li>\n<li><a href=\"https://www.youtube.com/watch?v=SFdKFhOjVLQ\" target=\"_blank\">https://www.youtube.com/watch?v=SFdKFhOjVLQ</a></li>\n<li><a href=\"https://rapids.ai/polars-gpu-engine/\" target=\"_blank\">https://rapids.ai/polars-gpu-engine/</a></li>\n</ul>\n<p>One simply needs to install polars with GPU as below- <br>\n<code>!pip install --extra-index-url=https://pypi.nvidia.com polars[gpu]==1.9.0 -q -d /kaggle/working/polars</code></p>\n<p>While collecting from lazy-frame to eager-frames, one needs to specify the engine as GPU as below-<br>\n<code>df.collect(engine = \"gpu\")</code></p>\n<p>This has the potential to speed up your features tremendously! Are you going to try this in your pipelines?<br>\nUpdated wheel file is present here-<br>\n<a href=\"https://www.kaggle.com/code/ravi20076/janestreet2024-imports-v1\" target=\"_blank\">https://www.kaggle.com/code/ravi20076/janestreet2024-imports-v1</a></p>\n<p>Usage of FE with GPU is here-<br>\n<a href=\"https://www.kaggle.com/code/ravi20076/janestreet2024-baseline-train-v1\" target=\"_blank\">https://www.kaggle.com/code/ravi20076/janestreet2024-baseline-train-v1</a></p>\n<p>Best regards!</p>",
      "rawMarkdown": "Hello all,\n\nI presume we are all collectively using polars for feature engineering and data loads herewith. I am sure we are also facing a few problems with data loads as the data is quite huge here. \n\nRecently, polars integrated with NVidia RAPIDS framework to improve the speed of data wrangling using GPUs without explicit code changes as well. References for this are as below-\n- https://pola.rs/posts/polars-on-gpu/\n- https://pola.rs/posts/gpu-engine-release/\n- https://developer.nvidia.com/blog/polars-gpu-engine-powered-by-rapids-cudf-now-available-in-open-beta/\n- https://www.datacamp.com/blog/polars-gpu-engine\n- https://www.youtube.com/watch?v=SFdKFhOjVLQ\n- https://rapids.ai/polars-gpu-engine/\n\nOne simply needs to install polars with GPU as below- \n`!pip install --extra-index-url=https://pypi.nvidia.com polars[gpu]==1.9.0 -q -d /kaggle/working/polars`\n\nWhile collecting from lazy-frame to eager-frames, one needs to specify the engine as GPU as below-\n`df.collect(engine = \"gpu\")`\n\nThis has the potential to speed up your features tremendously! Are you going to try this in your pipelines?\nUpdated wheel file is present here-\nhttps://www.kaggle.com/code/ravi20076/janestreet2024-imports-v1\n\nUsage of FE with GPU is here-\nhttps://www.kaggle.com/code/ravi20076/janestreet2024-baseline-train-v1\n\nBest regards!",
      "votes": 5
    },
    {
      "id": 3020014,
      "postDate": "2024-10-17T05:24:20.240Z",
      "content": "<p>Is this comp viable to those without GPUs? It seems like volume of data is quite a constraint. On an M2 Macbook here.</p>",
      "rawMarkdown": "Is this comp viable to those without GPUs? It seems like volume of data is quite a constraint. On an M2 Macbook here.",
      "votes": 1,
      "replies": [
        {
          "id": 3020293,
          "postDate": "2024-10-17T11:18:32.010Z",
          "content": "<p><a href=\"https://www.kaggle.com/jcc310\" target=\"_blank\">@jcc310</a> you can avoid GPU usage as well, but you will have to be creative with your code!</p>",
          "rawMarkdown": "@jcc310 you can avoid GPU usage as well, but you will have to be creative with your code!"
        }
      ]
    },
    {
      "id": 3018626,
      "postDate": "2024-10-15T23:19:52.827Z",
      "content": "<p>Hey Ravi, I tried to use polars for \"used car prices regression competition\" as well. But i faced numerous problems to achieve same data engineering using pandas and numpy. Do you have any other resource focusing on polars and pandas syntax difference. (apart from documentation - Which is very clear and provides good comparison to pandas).</p>",
      "rawMarkdown": "Hey Ravi, I tried to use polars for \"used car prices regression competition\" as well. But i faced numerous problems to achieve same data engineering using pandas and numpy. Do you have any other resource focusing on polars and pandas syntax difference. (apart from documentation - Which is very clear and provides good comparison to pandas).",
      "votes": 2,
      "replies": [
        {
          "id": 3018989,
          "postDate": "2024-10-16T07:31:00.127Z",
          "rawMarkdown": "",
          "votes": -1,
          "isDeleted": true,
          "replies": [
            {
              "id": 3019381,
              "postDate": "2024-10-16T14:26:32.063Z",
              "content": "<p>Polars has very good user documentation<br>\n<a href=\"https://docs.pola.rs/\" target=\"_blank\">https://docs.pola.rs/</a></p>",
              "rawMarkdown": "Polars has very good user documentation\nhttps://docs.pola.rs/",
              "votes": 4
            }
          ]
        },
        {
          "id": 3031038,
          "postDate": "2024-10-29T08:28:35.977Z",
          "content": "<p>I would advise you to forget about thinking in a \"pandas\" way, and just treat it as a completely different way to structure data. I think it's too conceptually different to be able to translate directly, and it may be better to just think of it in abstract terms of what you want to achieve with it. That's my personal approach, at least.</p>",
          "rawMarkdown": "I would advise you to forget about thinking in a \"pandas\" way, and just treat it as a completely different way to structure data. I think it's too conceptually different to be able to translate directly, and it may be better to just think of it in abstract terms of what you want to achieve with it. That's my personal approach, at least.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3020014,
      "author_name": "JamesC",
      "author_url": "",
      "post_date": "2024-10-17T05:24:20.240000",
      "content": "<p>Is this comp viable to those without GPUs? It seems like volume of data is quite a constraint. On an M2 Macbook here.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3020293,
          "author_name": "Ravi Ramakrishnan",
          "author_url": "",
          "post_date": "2024-10-17T11:18:32.010000",
          "content": "<p><a href=\"https://www.kaggle.com/jcc310\" target=\"_blank\">@jcc310</a> you can avoid GPU usage as well, but you will have to be creative with your code!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3018626,
      "author_name": "Jay Shrivastava",
      "author_url": "",
      "post_date": "2024-10-15T23:19:52.827000",
      "content": "<p>Hey Ravi, I tried to use polars for \"used car prices regression competition\" as well. But i faced numerous problems to achieve same data engineering using pandas and numpy. Do you have any other resource focusing on polars and pandas syntax difference. (apart from documentation - Which is very clear and provides good comparison to pandas).</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3018989,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-10-16T07:31:00.127000",
          "content": "",
          "votes": -1,
          "replies": [
            {
              "id": 3019381,
              "author_name": "Simon Veitner",
              "author_url": "",
              "post_date": "2024-10-16T14:26:32.063000",
              "content": "<p>Polars has very good user documentation<br>\n<a href=\"https://docs.pola.rs/\" target=\"_blank\">https://docs.pola.rs/</a></p>",
              "votes": 4,
              "replies": []
            }
          ]
        },
        {
          "id": 3031038,
          "author_name": "Abstr Phil",
          "author_url": "",
          "post_date": "2024-10-29T08:28:35.977000",
          "content": "<p>I would advise you to forget about thinking in a \"pandas\" way, and just treat it as a completely different way to structure data. I think it's too conceptually different to be able to translate directly, and it may be better to just think of it in abstract terms of what you want to achieve with it. That's my personal approach, at least.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3018595": "Hello all,\n\nI presume we are all collectively using polars for feature engineering and data loads herewith. I am sure we are also facing a few problems with data loads as the data is quite huge here. \n\nRecently, polars integrated with NVidia RAPIDS framework to improve the speed of data wrangling using GPUs without explicit code changes as well. References for this are as below-\n- https://pola.rs/posts/polars-on-gpu/\n- https://pola.rs/posts/gpu-engine-release/\n- https://developer.nvidia.com/blog/polars-gpu-engine-powered-by-rapids-cudf-now-available-in-open-beta/\n- https://www.datacamp.com/blog/polars-gpu-engine\n- https://www.youtube.com/watch?v=SFdKFhOjVLQ\n- https://rapids.ai/polars-gpu-engine/\n\nOne simply needs to install polars with GPU as below- \n`!pip install --extra-index-url=https://pypi.nvidia.com polars[gpu]==1.9.0 -q -d /kaggle/working/polars`\n\nWhile collecting from lazy-frame to eager-frames, one needs to specify the engine as GPU as below-\n`df.collect(engine = \"gpu\")`\n\nThis has the potential to speed up your features tremendously! Are you going to try this in your pipelines?\nUpdated wheel file is present here-\nhttps://www.kaggle.com/code/ravi20076/janestreet2024-imports-v1\n\nUsage of FE with GPU is here-\nhttps://www.kaggle.com/code/ravi20076/janestreet2024-baseline-train-v1\n\nBest regards!",
    "3020014": "Is this comp viable to those without GPUs? It seems like volume of data is quite a constraint. On an M2 Macbook here.",
    "3018626": "Hey Ravi, I tried to use polars for \"used car prices regression competition\" as well. But i faced numerous problems to achieve same data engineering using pandas and numpy. Do you have any other resource focusing on polars and pandas syntax difference. (apart from documentation - Which is very clear and provides good comparison to pandas)."
  }
}