{
  "id": 384605,
  "title": "Will polars be a force in this competition?",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/384605",
  "author_name": "",
  "post_date": "2023-02-08T16:33:11.218651200Z",
  "votes": 2,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Polars was very helpful in processing large data sets for the OTTO competition. <br>\nHowever, polars may not have much advantage in the inference phase of this competition, since small data are output as mini batches.  </p>\n<p>Please let us know what you think!</p>",
  "messages": [
    {
      "id": "2135449",
      "postDate": "02/08/2023 16:33:11",
      "content": "<p>Polars was very helpful in processing large data sets for the OTTO competition. <br>\nHowever, polars may not have much advantage in the inference phase of this competition, since small data are output as mini batches.  </p>\n<p>Please let us know what you think!</p>",
      "rawMarkdown": "Polars was very helpful in processing large data sets for the OTTO competition. \nHowever, polars may not have much advantage in the inference phase of this competition, since small data are output as mini batches.  \n\nPlease let us know what you think!",
      "votes": null
    },
    {
      "id": "2135458",
      "postDate": "02/08/2023 16:38:55",
      "content": "<p>Might be helpful in training. But in inference internet is turned off. So you can't install it.</p>\n<p>And yes, test batches are small so there's no big utility boost.</p>",
      "rawMarkdown": "Might be helpful in training. But in inference internet is turned off. So you can't install it.\n\nAnd yes, test batches are small so there's no big utility boost.",
      "votes": null
    },
    {
      "id": "2135480",
      "postDate": "02/08/2023 16:56:28",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/mayukh18\" target=\"_blank\">@mayukh18</a> <br>\nThanks for the comment.<br>\nI think if we put the wheel in Datasets we can import polars without internet access, <br>\nbut I agree that inference will not have much of an effect</p>",
      "rawMarkdown": "Hi, @mayukh18 \nThanks for the comment.\nI think if we put the wheel in Datasets we can import polars without internet access, \nbut I agree that inference will not have much of an effect",
      "votes": null
    },
    {
      "id": "2135736",
      "postDate": "02/08/2023 21:11:36",
      "content": "<p>In the context of efficiency prize, it's hard to say. 35 seconds to install it, and then getting the test data in unknown sized batches already in pandas format. These work against it. But there's need for feature engineering and low CPU memory available, so it could still end up being faster?</p>\n<p>Because of getting presumably small batches in pandas format, I'll probably try pure numpy (+numba) before I'd try Polars.</p>",
      "rawMarkdown": "In the context of efficiency prize, it's hard to say. 35 seconds to install it, and then getting the test data in unknown sized batches already in pandas format. These work against it. But there's need for feature engineering and low CPU memory available, so it could still end up being faster?\n\nBecause of getting presumably small batches in pandas format, I'll probably try pure numpy (+numba) before I'd try Polars.",
      "votes": null
    },
    {
      "id": "2135858",
      "postDate": "02/08/2023 22:43:20",
      "content": "<p>Could I know if the speed difference between Numpy and Numba is significant enough to give it a try? I always go with numpy, but didn't try numba.</p>",
      "rawMarkdown": "Could I know if the speed difference between Numpy and Numba is significant enough to give it a try? I always go with numpy, but didn't try numba.",
      "votes": null
    },
    {
      "id": "2135911",
      "postDate": "02/09/2023 00:33:31",
      "content": "<p>Actually I haven't really tried either. Read about it recently when doing basic internet digging on performance optimizations. I got some specific error the one time I tried to use it, and didn't have a chance to come back to it. (Yet)</p>",
      "rawMarkdown": "Actually I haven't really tried either. Read about it recently when doing basic internet digging on performance optimizations. I got some specific error the one time I tried to use it, and didn't have a chance to come back to it. (Yet)",
      "votes": null
    },
    {
      "id": "2138194",
      "postDate": "02/10/2023 16:21:30",
      "content": "<p>can't we just upload polars as a dataset and then just reference it using <code>sys</code>? Maybe this can help avoid installing it</p>",
      "rawMarkdown": "can't we just upload polars as a dataset and then just reference it using `sys`? Maybe this can help avoid installing it",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2135458,
      "author_name": "mayukh18",
      "author_url": "",
      "post_date": "02/08/2023 16:38:55",
      "content": "<p>Might be helpful in training. But in inference internet is turned off. So you can't install it.</p>\n<p>And yes, test batches are small so there's no big utility boost.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2135480,
          "author_name": "t88take",
          "author_url": "",
          "post_date": "02/08/2023 16:56:28",
          "content": "<p>Hi, <a href=\"https://www.kaggle.com/mayukh18\" target=\"_blank\">@mayukh18</a> <br>\nThanks for the comment.<br>\nI think if we put the wheel in Datasets we can import polars without internet access, <br>\nbut I agree that inference will not have much of an effect</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2135736,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "02/08/2023 21:11:36",
      "content": "<p>In the context of efficiency prize, it's hard to say. 35 seconds to install it, and then getting the test data in unknown sized batches already in pandas format. These work against it. But there's need for feature engineering and low CPU memory available, so it could still end up being faster?</p>\n<p>Because of getting presumably small batches in pandas format, I'll probably try pure numpy (+numba) before I'd try Polars.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2135858,
          "author_name": "mohammad2012191",
          "author_url": "",
          "post_date": "02/08/2023 22:43:20",
          "content": "<p>Could I know if the speed difference between Numpy and Numba is significant enough to give it a try? I always go with numpy, but didn't try numba.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2135911,
              "author_name": "roberthatch",
              "author_url": "",
              "post_date": "02/09/2023 00:33:31",
              "content": "<p>Actually I haven't really tried either. Read about it recently when doing basic internet digging on performance optimizations. I got some specific error the one time I tried to use it, and didn't have a chance to come back to it. (Yet)</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 2138194,
          "author_name": "alejopaullier",
          "author_url": "",
          "post_date": "02/10/2023 16:21:30",
          "content": "<p>can't we just upload polars as a dataset and then just reference it using <code>sys</code>? Maybe this can help avoid installing it</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2135449": "Polars was very helpful in processing large data sets for the OTTO competition. \nHowever, polars may not have much advantage in the inference phase of this competition, since small data are output as mini batches.  \n\nPlease let us know what you think!",
    "2135458": "Might be helpful in training. But in inference internet is turned off. So you can't install it.\n\nAnd yes, test batches are small so there's no big utility boost.",
    "2135480": "Hi, @mayukh18 \nThanks for the comment.\nI think if we put the wheel in Datasets we can import polars without internet access, \nbut I agree that inference will not have much of an effect",
    "2135736": "In the context of efficiency prize, it's hard to say. 35 seconds to install it, and then getting the test data in unknown sized batches already in pandas format. These work against it. But there's need for feature engineering and low CPU memory available, so it could still end up being faster?\n\nBecause of getting presumably small batches in pandas format, I'll probably try pure numpy (+numba) before I'd try Polars.",
    "2135858": "Could I know if the speed difference between Numpy and Numba is significant enough to give it a try? I always go with numpy, but didn't try numba.",
    "2135911": "Actually I haven't really tried either. Read about it recently when doing basic internet digging on performance optimizations. I got some specific error the one time I tried to use it, and didn't have a chance to come back to it. (Yet)",
    "2138194": "can't we just upload polars as a dataset and then just reference it using `sys`? Maybe this can help avoid installing it"
  },
  "source": "meta"
}