{
  "id": 388804,
  "title": "[LB 0.679 / 19min / CV 0.680] Polars training and inference baseline here!",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/388804",
  "author_name": "",
  "post_date": "2023-02-19T15:53:17.317701Z",
  "votes": 39,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Someone mentioned <strong><a href=\"https://pola-rs.github.io/polars-book/user-guide/introduction.html\" target=\"_blank\">Polars</a></strong> in discussion before, but there is no polars-based public code for this competition. </p>\n<p>I noticed that <strong>the current Kaggle environment (2023-02-17) has <code>polars</code> supported!</strong> So I made a notebook series to implement Chris's idea using polars. And I also added a new <strong>\"elapse-time diff\"</strong> feature named \"time_past\" for you to better get used to polars column operations.</p>\n<p>The series has two parts:<br>\nPart1 - <strong>CPU CatBoost Baseline (Using Polars) - Train</strong> - <a href=\"https://www.kaggle.com/code/carnozhao/cpu-catboost-baseline-using-polars-train\" target=\"_blank\">Link</a><br>\nPart2 - <strong>CPU CatBoost Baseline (Using Polars) - Inference</strong> - <a href=\"https://www.kaggle.com/code/carnozhao/cpu-catboost-baseline-using-polars-inference\" target=\"_blank\">Link</a></p>\n<p>As shown in training notebook above:</p>\n<pre><code>columns = [\n    (\n        (pl.col(\"elapsed_time\") - pl.col(\"elapsed_time\").shift(1))\n         .fill_null(0)\n         .clip(0, 1e9)\n         .over([\"session_id\", \"level_group\"])\n         .alias(\"time_past\")\n    ),\n]\n\n\naggs = [\n    *[pl.col(c).drop_nulls().n_unique().alias(f\"{c}_unique\") for c in CATS],\n    *[pl.col(c).mean().alias(f\"{c}_mean\") for c in NUMS],\n    *[pl.col(c).std().alias(f\"{c}_std\") for c in NUMS],\n    *[(pl.col(\"event_name\") == c).sum().alias(f\"{c}_sum\") for c in EVENTS],\n]\n</code></pre>\n<p>You can write few more lines to add hundreds of feature using polars's fancy expression syntax.</p>\n<p>IMHO, this training pipeline might also have comparable speed like CuDF's version. And the inference pipeline can be much faster than pandas version without any tricky optimization.</p>\n<p>Hope it helps. Thanks you <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for such an amazing idea. And <a href=\"https://www.kaggle.com/shashwatraman\" target=\"_blank\">@shashwatraman</a> for the CuDF version.<br>\n😊 </p>",
  "messages": [
    {
      "id": "2150798",
      "postDate": "02/19/2023 15:53:17",
      "content": "<p>Someone mentioned <strong><a href=\"https://pola-rs.github.io/polars-book/user-guide/introduction.html\" target=\"_blank\">Polars</a></strong> in discussion before, but there is no polars-based public code for this competition. </p>\n<p>I noticed that <strong>the current Kaggle environment (2023-02-17) has <code>polars</code> supported!</strong> So I made a notebook series to implement Chris's idea using polars. And I also added a new <strong>\"elapse-time diff\"</strong> feature named \"time_past\" for you to better get used to polars column operations.</p>\n<p>The series has two parts:<br>\nPart1 - <strong>CPU CatBoost Baseline (Using Polars) - Train</strong> - <a href=\"https://www.kaggle.com/code/carnozhao/cpu-catboost-baseline-using-polars-train\" target=\"_blank\">Link</a><br>\nPart2 - <strong>CPU CatBoost Baseline (Using Polars) - Inference</strong> - <a href=\"https://www.kaggle.com/code/carnozhao/cpu-catboost-baseline-using-polars-inference\" target=\"_blank\">Link</a></p>\n<p>As shown in training notebook above:</p>\n<pre><code>columns = [\n    (\n        (pl.col(\"elapsed_time\") - pl.col(\"elapsed_time\").shift(1))\n         .fill_null(0)\n         .clip(0, 1e9)\n         .over([\"session_id\", \"level_group\"])\n         .alias(\"time_past\")\n    ),\n]\n\n\naggs = [\n    *[pl.col(c).drop_nulls().n_unique().alias(f\"{c}_unique\") for c in CATS],\n    *[pl.col(c).mean().alias(f\"{c}_mean\") for c in NUMS],\n    *[pl.col(c).std().alias(f\"{c}_std\") for c in NUMS],\n    *[(pl.col(\"event_name\") == c).sum().alias(f\"{c}_sum\") for c in EVENTS],\n]\n</code></pre>\n<p>You can write few more lines to add hundreds of feature using polars's fancy expression syntax.</p>\n<p>IMHO, this training pipeline might also have comparable speed like CuDF's version. And the inference pipeline can be much faster than pandas version without any tricky optimization.</p>\n<p>Hope it helps. Thanks you <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for such an amazing idea. And <a href=\"https://www.kaggle.com/shashwatraman\" target=\"_blank\">@shashwatraman</a> for the CuDF version.<br>\n😊 </p>",
      "rawMarkdown": "Someone mentioned **[Polars](https://pola-rs.github.io/polars-book/user-guide/introduction.html)** in discussion before, but there is no polars-based public code for this competition. \n\nI noticed that **the current Kaggle environment (2023-02-17) has `polars` supported!** So I made a notebook series to implement Chris's idea using polars. And I also added a new **\"elapse-time diff\"** feature named \"time_past\" for you to better get used to polars column operations.\n\nThe series has two parts:\nPart1 - **CPU CatBoost Baseline (Using Polars) - Train** - [Link](https://www.kaggle.com/code/carnozhao/cpu-catboost-baseline-using-polars-train)\nPart2 - **CPU CatBoost Baseline (Using Polars) - Inference** - [Link](https://www.kaggle.com/code/carnozhao/cpu-catboost-baseline-using-polars-inference)\n\nAs shown in training notebook above:\n\n```\ncolumns = [\n    (\n        (pl.col(\"elapsed_time\") - pl.col(\"elapsed_time\").shift(1))\n         .fill_null(0)\n         .clip(0, 1e9)\n         .over([\"session_id\", \"level_group\"])\n         .alias(\"time_past\")\n    ),\n]\n\n\naggs = [\n    *[pl.col(c).drop_nulls().n_unique().alias(f\"{c}_unique\") for c in CATS],\n    *[pl.col(c).mean().alias(f\"{c}_mean\") for c in NUMS],\n    *[pl.col(c).std().alias(f\"{c}_std\") for c in NUMS],\n    *[(pl.col(\"event_name\") == c).sum().alias(f\"{c}_sum\") for c in EVENTS],\n]\n```\n\nYou can write few more lines to add hundreds of feature using polars's fancy expression syntax.\n\nIMHO, this training pipeline might also have comparable speed like CuDF's version. And the inference pipeline can be much faster than pandas version without any tricky optimization.\n\nHope it helps. Thanks you @cdeotte for such an amazing idea. And @shashwatraman for the CuDF version.\n😊",
      "votes": null
    },
    {
      "id": "2169399",
      "postDate": "03/05/2023 04:34:14",
      "content": "<p>Hi boss ,  sorry that i found that the \"index\" is not consistent with the \"elapsed_time\", and i saw ur code is without sorting , so the data should follow the existing sequence , and then the negtive time would be calculated , right ?  (but the negtive amount will be coverted to 0.) </p>",
      "rawMarkdown": "Hi boss ,  sorry that i found that the \"index\" is not consistent with the \"elapsed_time\", and i saw ur code is without sorting , so the data should follow the existing sequence , and then the negtive time would be calculated , right ?  (but the negtive amount will be coverted to 0.)",
      "votes": null
    },
    {
      "id": "2195169",
      "postDate": "03/24/2023 13:28:34",
      "content": "<p>I compared polars and cudf in Otto competition, and there is no discussion: using cudf is way faster, see <a href=\"https://www.kaggle.com/code/cpmpml/matrix-factorization-with-gpu\" target=\"_blank\">https://www.kaggle.com/code/cpmpml/matrix-factorization-with-gpu</a></p>\n<p>It therefore makes sense to use GPU for training in this competition also. See <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> post on this topic for instance.</p>",
      "rawMarkdown": "I compared polars and cudf in Otto competition, and there is no discussion: using cudf is way faster, see https://www.kaggle.com/code/cpmpml/matrix-factorization-with-gpu\n\nIt therefore makes sense to use GPU for training in this competition also. See @cdeotte post on this topic for instance.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2169399,
      "author_name": "johnfjliu",
      "author_url": "",
      "post_date": "03/05/2023 04:34:14",
      "content": "<p>Hi boss ,  sorry that i found that the \"index\" is not consistent with the \"elapsed_time\", and i saw ur code is without sorting , so the data should follow the existing sequence , and then the negtive time would be calculated , right ?  (but the negtive amount will be coverted to 0.) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2195169,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "03/24/2023 13:28:34",
      "content": "<p>I compared polars and cudf in Otto competition, and there is no discussion: using cudf is way faster, see <a href=\"https://www.kaggle.com/code/cpmpml/matrix-factorization-with-gpu\" target=\"_blank\">https://www.kaggle.com/code/cpmpml/matrix-factorization-with-gpu</a></p>\n<p>It therefore makes sense to use GPU for training in this competition also. See <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> post on this topic for instance.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2150798": "Someone mentioned **[Polars](https://pola-rs.github.io/polars-book/user-guide/introduction.html)** in discussion before, but there is no polars-based public code for this competition. \n\nI noticed that **the current Kaggle environment (2023-02-17) has `polars` supported!** So I made a notebook series to implement Chris's idea using polars. And I also added a new **\"elapse-time diff\"** feature named \"time_past\" for you to better get used to polars column operations.\n\nThe series has two parts:\nPart1 - **CPU CatBoost Baseline (Using Polars) - Train** - [Link](https://www.kaggle.com/code/carnozhao/cpu-catboost-baseline-using-polars-train)\nPart2 - **CPU CatBoost Baseline (Using Polars) - Inference** - [Link](https://www.kaggle.com/code/carnozhao/cpu-catboost-baseline-using-polars-inference)\n\nAs shown in training notebook above:\n\n```\ncolumns = [\n    (\n        (pl.col(\"elapsed_time\") - pl.col(\"elapsed_time\").shift(1))\n         .fill_null(0)\n         .clip(0, 1e9)\n         .over([\"session_id\", \"level_group\"])\n         .alias(\"time_past\")\n    ),\n]\n\n\naggs = [\n    *[pl.col(c).drop_nulls().n_unique().alias(f\"{c}_unique\") for c in CATS],\n    *[pl.col(c).mean().alias(f\"{c}_mean\") for c in NUMS],\n    *[pl.col(c).std().alias(f\"{c}_std\") for c in NUMS],\n    *[(pl.col(\"event_name\") == c).sum().alias(f\"{c}_sum\") for c in EVENTS],\n]\n```\n\nYou can write few more lines to add hundreds of feature using polars's fancy expression syntax.\n\nIMHO, this training pipeline might also have comparable speed like CuDF's version. And the inference pipeline can be much faster than pandas version without any tricky optimization.\n\nHope it helps. Thanks you @cdeotte for such an amazing idea. And @shashwatraman for the CuDF version.\n😊",
    "2169399": "Hi boss ,  sorry that i found that the \"index\" is not consistent with the \"elapsed_time\", and i saw ur code is without sorting , so the data should follow the existing sequence , and then the negtive time would be calculated , right ?  (but the negtive amount will be coverted to 0.)",
    "2195169": "I compared polars and cudf in Otto competition, and there is no discussion: using cudf is way faster, see https://www.kaggle.com/code/cpmpml/matrix-factorization-with-gpu\n\nIt therefore makes sense to use GPU for training in this competition also. See @cdeotte post on this topic for instance."
  },
  "source": "meta"
}