{
  "id": 542697,
  "title": "How to 1.8 times [updated → 2.0 times] speed up of inference in LGBM !",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/542697",
  "author_name": "chumajin",
  "post_date": "2024-10-26T10:44:53.811000",
  "votes": 45,
  "comment_count": 17,
  "views": 0,
  "content": "<p>In this topic, I will explain how I reduced the inference time of my LightGBM single model from 32 minutes to 18 minutes (updated 16 min, please see <a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/542697#3029289\" target=\"_blank\">comment</a> ). </p>\n<p><strong>In short, convert your data to NumPy instead of using pandas, then train and perform inference with LightGBM.</strong></p>\n<p>In this competition, test data is provided using Polars, so converting it to pandas for inference slows down the process. However, by converting it directly to NumPy from the start and then performing training and inference, you can achieve faster inference.</p>\n<p>The following is example.</p>\n<p>□ training</p>\n<p>train : polars DataFrame, valid : polas DataFrame</p>\n<pre><code>lgb_train = lgb(train(features)(), train(target)())\nlgb_valid = lgb(valid(features)(), valid(target)())\n</code></pre>\n<p>□ inference (please see updated <a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/542697#3029289\" target=\"_blank\">comment</a> )</p>\n<pre><code>def (: .DataFrame, lags: .DataFrame | None) -&gt; .DataFrame | pd.DataFrame:\n\n    preds = model.(.select(features).to_numpy())\n\n    predictions = .select(\n        'row_id',\n        .Series(preds).alias('responder_6'),\n    )\n\n     predictions\n</code></pre>\n<p>Additionally, there was a discussion in this <a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/542022\" target=\"_blank\">topic</a> about the simulation time not matching the submission time. However, in my case, the <a href=\"https://www.kaggle.com/code/chumajin/janestreet-updated-simulator-for-time-series-api\" target=\"_blank\">simulation</a> using the Kaggle evaluation API aligns quite well with this submission method.</p>\n<p>Conventinal LGBM using pandas : simulation 27 min, submission 32 min<br>\nThis LGBM using numpy(above predict function) : simulation 18 min, submission 18 min</p>\n<p>Furthermore, if you are not using lags, removing the following code can make the process about 2 minutes faster.</p>\n<pre><code> lags_\n     lags   :\n        lags_ = lags\n</code></pre>\n<p>I hope this helps you.<br>\nEnjoy kaggling ~ .</p>",
  "messages": [
    {
      "id": 3028627,
      "postDate": "2024-10-26T10:44:53.810Z",
      "content": "<p>In this topic, I will explain how I reduced the inference time of my LightGBM single model from 32 minutes to 18 minutes (updated 16 min, please see <a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/542697#3029289\" target=\"_blank\">comment</a> ). </p>\n<p><strong>In short, convert your data to NumPy instead of using pandas, then train and perform inference with LightGBM.</strong></p>\n<p>In this competition, test data is provided using Polars, so converting it to pandas for inference slows down the process. However, by converting it directly to NumPy from the start and then performing training and inference, you can achieve faster inference.</p>\n<p>The following is example.</p>\n<p>□ training</p>\n<p>train : polars DataFrame, valid : polas DataFrame</p>\n<pre><code>lgb_train = lgb(train(features)(), train(target)())\nlgb_valid = lgb(valid(features)(), valid(target)())\n</code></pre>\n<p>□ inference (please see updated <a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/542697#3029289\" target=\"_blank\">comment</a> )</p>\n<pre><code>def (: .DataFrame, lags: .DataFrame | None) -&gt; .DataFrame | pd.DataFrame:\n\n    preds = model.(.select(features).to_numpy())\n\n    predictions = .select(\n        'row_id',\n        .Series(preds).alias('responder_6'),\n    )\n\n     predictions\n</code></pre>\n<p>Additionally, there was a discussion in this <a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/542022\" target=\"_blank\">topic</a> about the simulation time not matching the submission time. However, in my case, the <a href=\"https://www.kaggle.com/code/chumajin/janestreet-updated-simulator-for-time-series-api\" target=\"_blank\">simulation</a> using the Kaggle evaluation API aligns quite well with this submission method.</p>\n<p>Conventinal LGBM using pandas : simulation 27 min, submission 32 min<br>\nThis LGBM using numpy(above predict function) : simulation 18 min, submission 18 min</p>\n<p>Furthermore, if you are not using lags, removing the following code can make the process about 2 minutes faster.</p>\n<pre><code> lags_\n     lags   :\n        lags_ = lags\n</code></pre>\n<p>I hope this helps you.<br>\nEnjoy kaggling ~ .</p>",
      "rawMarkdown": "In this topic, I will explain how I reduced the inference time of my LightGBM single model from 32 minutes to 18 minutes (updated 16 min, please see [comment](https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/542697#3029289) ). \n\n**In short, convert your data to NumPy instead of using pandas, then train and perform inference with LightGBM.**\n\nIn this competition, test data is provided using Polars, so converting it to pandas for inference slows down the process. However, by converting it directly to NumPy from the start and then performing training and inference, you can achieve faster inference.\n\nThe following is example.\n\n□ training\n\ntrain : polars DataFrame, valid : polas DataFrame\n\n~~~\nlgb_train = lgb.Dataset(train.select(features).to_numpy(), train.select(target).to_numpy())\nlgb_valid = lgb.Dataset(valid.select(features).to_numpy(), valid.select(target).to_numpy())\n~~~\n\n□ inference (please see updated [comment](https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/542697#3029289) )\n\n~~~\ndef predict(test: pl.DataFrame, lags: pl.DataFrame | None) -> pl.DataFrame | pd.DataFrame:\n\n    preds = model.predict(test.select(features).to_numpy())\n\n    predictions = test.select(\n        'row_id',\n        pl.Series(preds).alias('responder_6'),\n    )\n\n    return predictions\n~~~\n\nAdditionally, there was a discussion in this [topic](https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/542022) about the simulation time not matching the submission time. However, in my case, the [simulation](https://www.kaggle.com/code/chumajin/janestreet-updated-simulator-for-time-series-api) using the Kaggle evaluation API aligns quite well with this submission method.\n\nConventinal LGBM using pandas : simulation 27 min, submission 32 min\nThis LGBM using numpy(above predict function) : simulation 18 min, submission 18 min\n\nFurthermore, if you are not using lags, removing the following code can make the process about 2 minutes faster.\n\n~~~\nglobal lags_\n    if lags is not None:\n        lags_ = lags\n~~~\n\nI hope this helps you.\nEnjoy kaggling ~ .",
      "votes": 45
    },
    {
      "id": 3028696,
      "postDate": "2024-10-26T12:49:17.543Z",
      "content": "<p>If you want to go even faster there is some ways to do GPU inference of LGBM</p>\n<p>~5x Faster Inference - <a href=\"https://opensource.microsoft.com/blog/2020/09/29/accelerate-machine-learning-models-gpu-onnx-runtime-hummingbird/\" target=\"_blank\">https://opensource.microsoft.com/blog/2020/09/29/accelerate-machine-learning-models-gpu-onnx-runtime-hummingbird/</a></p>\n<p>~10x Faster Inference - <a href=\"https://github.com/siboehm/lleaves\" target=\"_blank\">https://github.com/siboehm/lleaves</a></p>\n<p>I have used some of these Tree Inference packages before and can confirm they do work (don't neccesarily get the speeds they claim but usually at least 2x) - however I would only really use them if you have a very large tree.</p>",
      "rawMarkdown": "If you want to go even faster there is some ways to do GPU inference of LGBM\n\n~5x Faster Inference - https://opensource.microsoft.com/blog/2020/09/29/accelerate-machine-learning-models-gpu-onnx-runtime-hummingbird/\n\n~10x Faster Inference - https://github.com/siboehm/lleaves\n\nI have used some of these Tree Inference packages before and can confirm they do work (don't neccesarily get the speeds they claim but usually at least 2x) - however I would only really use them if you have a very large tree.\n\n",
      "votes": 6,
      "replies": [
        {
          "id": 3029129,
          "postDate": "2024-10-26T23:12:05.040Z",
          "content": "<p><a href=\"https://www.kaggle.com/julianmukaj\" target=\"_blank\">@julianmukaj</a> Thank you very much! These articles are very interesting. They seem very useful. I will try it !</p>",
          "rawMarkdown": "@julianmukaj Thank you very much! These articles are very interesting. They seem very useful. I will try it !",
          "votes": 2,
          "replies": [
            {
              "id": 3029728,
              "postDate": "2024-10-27T16:50:15.650Z",
              "content": "<p>Extremely helpful.You Guys are really awesome. Such a homy community.</p>",
              "rawMarkdown": "Extremely helpful.You Guys are really awesome. Such a homy community.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3030192,
      "postDate": "2024-10-28T09:27:55.460Z",
      "content": "<p>Do you have any recommendations on OOM issues? Can't seem to load and train the entire dataset in the Kaggle notebook environment.</p>",
      "rawMarkdown": "Do you have any recommendations on OOM issues? Can't seem to load and train the entire dataset in the Kaggle notebook environment.",
      "votes": 1,
      "replies": [
        {
          "id": 3030419,
          "postDate": "2024-10-28T14:12:00.490Z",
          "content": "<p><a href=\"https://www.kaggle.com/yipvictor\" target=\"_blank\">@yipvictor</a> Regarding the OOM (Out Of Memory) issue, it’s not possible to process all the training data in a Kaggle notebook. If you wish to do so, you should consider using a different environment like Google Colab or Google Cloud for training, and perform inference in the Kaggle notebook. However, it's uncertain whether you actually need all the training data.</p>",
          "rawMarkdown": "@yipvictor Regarding the OOM (Out Of Memory) issue, it’s not possible to process all the training data in a Kaggle notebook. If you wish to do so, you should consider using a different environment like Google Colab or Google Cloud for training, and perform inference in the Kaggle notebook. However, it's uncertain whether you actually need all the training data."
        },
        {
          "id": 3043298,
          "postDate": "2024-11-12T09:49:51.600Z",
          "content": "<p>You can do it, check this notebook - <a href=\"https://www.kaggle.com/code/daniilvolkov/js24-prepare-data-pkl-gz-output\" target=\"_blank\">https://www.kaggle.com/code/daniilvolkov/js24-prepare-data-pkl-gz-output</a></p>",
          "rawMarkdown": "You can do it, check this notebook - https://www.kaggle.com/code/daniilvolkov/js24-prepare-data-pkl-gz-output",
          "votes": 1
        }
      ]
    },
    {
      "id": 3029167,
      "postDate": "2024-10-27T00:59:36.833Z",
      "content": "<p>May I ask if I misunderstood the meaning? My experimental results are not ideal? runtime is 155 and 284.</p>\n<pre><code> polars  pl\n pandas  pd\n numpy  np\n\n  lightgbm  LGBMRegressor,LGBMClassifier,log_evaluation,early_stopping\n time\n\ntrain=pl.read_parquet()\ntrain=train.to_pandas()\nfeatures=[ i&lt;     i  ()]\ntest=train[-:][features+[]]\ntrain=train[:-][features+[]]\n((train),(test))\n\nstart_time=time.time()\nmodel=LGBMRegressor()\nmodel.fit(train[features],train[])\nmodel.predict(test[features])\nend_time=time.time()\n()\n\nstart_time=time.time()\nmodel=LGBMRegressor()\nmodel.fit(train[features].values,train[].values)\nmodel.predict(test[features].values)\nend_time=time.time()\n()\n</code></pre>",
      "rawMarkdown": "May I ask if I misunderstood the meaning? My experimental results are not ideal? runtime is 155 and 284.\n```python\nimport polars as pl#similar to pandas, but with better performance when dealing with large datasets.\nimport pandas as pd#read csv,parquet\nimport numpy as np#for scientific computation of matrices\n#models(lgb,xgb,cat)\nfrom  lightgbm import LGBMRegressor,LGBMClassifier,log_evaluation,early_stopping\nimport time\n\ntrain=pl.read_parquet(\"/kaggle/input/jane-street-real-time-market-data-forecasting/train.parquet/partition_id=9/part-0.parquet\")\ntrain=train.to_pandas()\nfeatures=[f'feature_0{i}'if i<10  else f'feature_{i}' for i in range(79)]\ntest=train[-2000000:][features+['responder_6']]\ntrain=train[:-2000000][features+['responder_6']]\nprint(len(train),len(test))\n\nstart_time=time.time()\nmodel=LGBMRegressor()\nmodel.fit(train[features],train['responder_6'])\nmodel.predict(test[features])\nend_time=time.time()\nprint(f\"run_time:{end_time-start_time}\")\n\nstart_time=time.time()\nmodel=LGBMRegressor()\nmodel.fit(train[features].values,train['responder_6'].values)\nmodel.predict(test[features].values)\nend_time=time.time()\nprint(f\"run_time:{end_time-start_time}\")\n```",
      "votes": 1,
      "replies": [
        {
          "id": 3029287,
          "postDate": "2024-10-27T05:25:03.410Z",
          "content": "<p><a href=\"https://www.kaggle.com/yunsuxiaozi\" target=\"_blank\">@yunsuxiaozi</a> Thank you for comment. This code is something wrong. When I check the following code, which changes the order of evaluation for pandas and numpy. I got 151 sec and 286 sec. Maybe initialization is required.</p>\n<pre><code>import polars as pl to pandas, but with better performance when dealing with large datasets.\nimport pandas as pd csv,parquet\nimport numpy as np scientific computation of matrices\n(lgb,xgb,cat)\nfrom  lightgbm import LGBMRegressor,LGBMClassifier,log_evaluation,early_stopping\nimport \n\ntrain=pl()\ntrain=train()\nfeatures=\ntest=train]\ntrain=train]\n,(test))\n\nstart_time=()\nmodel=()\nmodel(train,train.values)\nmodel(test.values)\nend_time=()\n\n\nstart_time=()\nmodel=()\nmodel(train,train)\nmodel(test)\nend_time=()\n\n</code></pre>",
          "rawMarkdown": "@yunsuxiaozi Thank you for comment. This code is something wrong. When I check the following code, which changes the order of evaluation for pandas and numpy. I got 151 sec and 286 sec. Maybe initialization is required.\n\n~~~\nimport polars as pl#similar to pandas, but with better performance when dealing with large datasets.\nimport pandas as pd#read csv,parquet\nimport numpy as np#for scientific computation of matrices\n#models(lgb,xgb,cat)\nfrom  lightgbm import LGBMRegressor,LGBMClassifier,log_evaluation,early_stopping\nimport time\n\ntrain=pl.read_parquet(\"/kaggle/input/jane-street-real-time-market-data-forecasting/train.parquet/partition_id=9/part-0.parquet\")\ntrain=train.to_pandas()\nfeatures=[f'feature_0{i}'if i<10  else f'feature_{i}' for i in range(79)]\ntest=train[-2000000:][features+['responder_6']]\ntrain=train[:-2000000][features+['responder_6']]\nprint(len(train),len(test))\n\nstart_time=time.time()\nmodel=LGBMRegressor()\nmodel.fit(train[features].values,train['responder_6'].values)\nmodel.predict(test[features].values)\nend_time=time.time()\nprint(f\"run_time:{end_time-start_time}\")\n\nstart_time=time.time()\nmodel=LGBMRegressor()\nmodel.fit(train[features],train['responder_6'])\nmodel.predict(test[features])\nend_time=time.time()\nprint(f\"run_time:{end_time-start_time}\")\n~~~"
        }
      ]
    },
    {
      "id": 3030925,
      "postDate": "2024-10-29T04:00:35.280Z",
      "content": "<p>Here are some optimization methods you can implement to speed up inference times even further for your LightGBM model:</p>\n<p>Quantize the Model: Use LightGBM's built-in model quantization (quantize) to reduce precision, which speeds up both memory usage and inference time without major loss in accuracy. <br>\nmodel = lgb.Booster(model_file='your_model.txt')<br>\nmodel.quantize()<br>\nReduce Data Transfer: By running inference directly on the GPU, data transfer between CPU and GPU is minimized, which helps in achieving faster inference. For this:</p>\n<p>Use predict with parameter pred_early_stop=True (if compatible with your model setup) to reduce computation by stopping predictions early for instances that are highly confident.<br>\nAvoid Row-by-Row Inference: Use batch prediction instead of processing row-by-row. Process multiple rows together to maximize hardware utilization:</p>\n<p>preds = model.predict(test.select(features).to_numpy(), num_iteration=model.best_iteration)<br>\nThreading and Parallelization: Increase LightGBM's parallel processing with n_jobs to leverage all CPU cores:</p>\n<p>model.predict(test.select(features).to_numpy(), num_iteration=model.best_iteration, n_jobs=-1)<br>\nUse dart for Smaller Trees: If your model’s configuration allows, experiment with the dart boosting type, which can sometimes reduce inference time, as it often produces smaller trees.</p>\n<p>Optimize Feature Selection: Drop unnecessary features and optimize feature engineering by selecting only essential features for your prediction, reducing the data dimensions handled during inference.</p>\n<p>These combined approaches can further reduce your inference times below 16 minutes if they align with your competition requirements and hardware.</p>",
      "rawMarkdown": "Here are some optimization methods you can implement to speed up inference times even further for your LightGBM model:\n\nQuantize the Model: Use LightGBM's built-in model quantization (quantize) to reduce precision, which speeds up both memory usage and inference time without major loss in accuracy. \nmodel = lgb.Booster(model_file='your_model.txt')\nmodel.quantize()\nReduce Data Transfer: By running inference directly on the GPU, data transfer between CPU and GPU is minimized, which helps in achieving faster inference. For this:\n\nUse predict with parameter pred_early_stop=True (if compatible with your model setup) to reduce computation by stopping predictions early for instances that are highly confident.\nAvoid Row-by-Row Inference: Use batch prediction instead of processing row-by-row. Process multiple rows together to maximize hardware utilization:\n\n\npreds = model.predict(test.select(features).to_numpy(), num_iteration=model.best_iteration)\nThreading and Parallelization: Increase LightGBM's parallel processing with n_jobs to leverage all CPU cores:\n\n\nmodel.predict(test.select(features).to_numpy(), num_iteration=model.best_iteration, n_jobs=-1)\nUse dart for Smaller Trees: If your model’s configuration allows, experiment with the dart boosting type, which can sometimes reduce inference time, as it often produces smaller trees.\n\nOptimize Feature Selection: Drop unnecessary features and optimize feature engineering by selecting only essential features for your prediction, reducing the data dimensions handled during inference.\n\nThese combined approaches can further reduce your inference times below 16 minutes if they align with your competition requirements and hardware.",
      "replies": [
        {
          "id": 3030938,
          "postDate": "2024-10-29T04:37:04.743Z",
          "content": "<p><a href=\"https://www.kaggle.com/bhaveshtf\" target=\"_blank\">@bhaveshtf</a> Thank you for all the insights. There was so much I didn’t know. I'll give it a try!</p>",
          "rawMarkdown": "@bhaveshtf Thank you for all the insights. There was so much I didn’t know. I'll give it a try!"
        },
        {
          "id": 3031058,
          "postDate": "2024-10-29T09:18:29.593Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 3031550,
          "postDate": "2024-10-29T20:08:41.847Z",
          "content": "<p>ChatGPT response? pred_early_stop ︎is only for classification problems 😆</p>\n<blockquote>\n  <p>Quantize the Model: Use LightGBM's built-in model quantization (quantize) to reduce precision, which speeds up both memory usage and inference time without major loss in accuracy.<br>\n  model = lgb.Booster(model_file='your_model.txt')<br>\n  model.quantize()</p>\n</blockquote>\n<p>No such method exists: <a href=\"https://lightgbm.readthedocs.io/en/latest/pythonapi/lightgbm.Booster.html\" target=\"_blank\">https://lightgbm.readthedocs.io/en/latest/pythonapi/lightgbm.Booster.html</a></p>",
          "rawMarkdown": "ChatGPT response? pred_early_stop ︎is only for classification problems 😆\n\n>Quantize the Model: Use LightGBM's built-in model quantization (quantize) to reduce precision, which speeds up both memory usage and inference time without major loss in accuracy.\nmodel = lgb.Booster(model_file='your_model.txt')\nmodel.quantize()\n\nNo such method exists: https://lightgbm.readthedocs.io/en/latest/pythonapi/lightgbm.Booster.html",
          "votes": 8
        }
      ]
    },
    {
      "id": 3029289,
      "postDate": "2024-10-27T05:28:54.070Z",
      "content": "<p>Updated : I can achieve 2 times faster (16 min) like this in inference code. The difference from above is that it involves converting to NumPy in advance before making predictions.</p>\n<pre><code>def (: .DataFrame, lags: .DataFrame | None) -&gt; .DataFrame | pd.DataFrame:\n\n    test_X = .select(features).to_numpy()\n\n    preds = model.(test_X)\n\n    predictions = .select(\n        'row_id',\n        .Series(preds).alias('responder_6'),\n    )\n\n     predictions\n</code></pre>",
      "rawMarkdown": "Updated : I can achieve 2 times faster (16 min) like this in inference code. The difference from above is that it involves converting to NumPy in advance before making predictions.\n\n~~~\ndef predict(test: pl.DataFrame, lags: pl.DataFrame | None) -> pl.DataFrame | pd.DataFrame:\n    \n    test_X = test.select(features).to_numpy()\n          \n    preds = model.predict(test_X)\n\n    predictions = test.select(\n        'row_id',\n        pl.Series(preds).alias('responder_6'),\n    )\n\n    return predictions\n~~~",
      "replies": [
        {
          "id": 3030191,
          "postDate": "2024-10-28T09:27:17.727Z",
          "content": "<p>Is this not identical to what you had before or does assigning it to a variable first somehow reduce compute</p>",
          "rawMarkdown": "Is this not identical to what you had before or does assigning it to a variable first somehow reduce compute",
          "votes": 1,
          "replies": [
            {
              "id": 3030426,
              "postDate": "2024-10-28T14:18:37.403Z",
              "content": "<p><a href=\"https://www.kaggle.com/yipvictor\" target=\"_blank\">@yipvictor</a> We might avoid unnecessary overhead and speed up inference, due to differences in internal processing like type checking and data conversion.</p>",
              "rawMarkdown": "@yipvictor We might avoid unnecessary overhead and speed up inference, due to differences in internal processing like type checking and data conversion."
            }
          ]
        }
      ]
    },
    {
      "id": 3029734,
      "postDate": "2024-10-27T16:57:20.870Z",
      "content": "<p>Thank you for such an incredible idea.</p>",
      "rawMarkdown": "Thank you for such an incredible idea.",
      "votes": 1
    },
    {
      "id": 3028779,
      "postDate": "2024-10-26T13:54:23.417Z",
      "content": "<p>Thanks such an amazing idea</p>",
      "rawMarkdown": "Thanks such an amazing idea",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 3028696,
      "author_name": "JM",
      "author_url": "",
      "post_date": "2024-10-26T12:49:17.543000",
      "content": "<p>If you want to go even faster there is some ways to do GPU inference of LGBM</p>\n<p>~5x Faster Inference - <a href=\"https://opensource.microsoft.com/blog/2020/09/29/accelerate-machine-learning-models-gpu-onnx-runtime-hummingbird/\" target=\"_blank\">https://opensource.microsoft.com/blog/2020/09/29/accelerate-machine-learning-models-gpu-onnx-runtime-hummingbird/</a></p>\n<p>~10x Faster Inference - <a href=\"https://github.com/siboehm/lleaves\" target=\"_blank\">https://github.com/siboehm/lleaves</a></p>\n<p>I have used some of these Tree Inference packages before and can confirm they do work (don't neccesarily get the speeds they claim but usually at least 2x) - however I would only really use them if you have a very large tree.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 3029129,
          "author_name": "chumajin",
          "author_url": "",
          "post_date": "2024-10-26T23:12:05.040000",
          "content": "<p><a href=\"https://www.kaggle.com/julianmukaj\" target=\"_blank\">@julianmukaj</a> Thank you very much! These articles are very interesting. They seem very useful. I will try it !</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3029728,
              "author_name": "FarhaanKhan ",
              "author_url": "",
              "post_date": "2024-10-27T16:50:15.650000",
              "content": "<p>Extremely helpful.You Guys are really awesome. Such a homy community.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3030192,
      "author_name": "Victor Yip",
      "author_url": "",
      "post_date": "2024-10-28T09:27:55.460000",
      "content": "<p>Do you have any recommendations on OOM issues? Can't seem to load and train the entire dataset in the Kaggle notebook environment.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3030419,
          "author_name": "chumajin",
          "author_url": "",
          "post_date": "2024-10-28T14:12:00.490000",
          "content": "<p><a href=\"https://www.kaggle.com/yipvictor\" target=\"_blank\">@yipvictor</a> Regarding the OOM (Out Of Memory) issue, it’s not possible to process all the training data in a Kaggle notebook. If you wish to do so, you should consider using a different environment like Google Colab or Google Cloud for training, and perform inference in the Kaggle notebook. However, it's uncertain whether you actually need all the training data.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3043298,
          "author_name": "Daniil Volkov",
          "author_url": "",
          "post_date": "2024-11-12T09:49:51.600000",
          "content": "<p>You can do it, check this notebook - <a href=\"https://www.kaggle.com/code/daniilvolkov/js24-prepare-data-pkl-gz-output\" target=\"_blank\">https://www.kaggle.com/code/daniilvolkov/js24-prepare-data-pkl-gz-output</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3029167,
      "author_name": "yunsuxiaozi",
      "author_url": "",
      "post_date": "2024-10-27T00:59:36.833000",
      "content": "<p>May I ask if I misunderstood the meaning? My experimental results are not ideal? runtime is 155 and 284.</p>\n<pre><code> polars  pl\n pandas  pd\n numpy  np\n\n  lightgbm  LGBMRegressor,LGBMClassifier,log_evaluation,early_stopping\n time\n\ntrain=pl.read_parquet()\ntrain=train.to_pandas()\nfeatures=[ i&lt;     i  ()]\ntest=train[-:][features+[]]\ntrain=train[:-][features+[]]\n((train),(test))\n\nstart_time=time.time()\nmodel=LGBMRegressor()\nmodel.fit(train[features],train[])\nmodel.predict(test[features])\nend_time=time.time()\n()\n\nstart_time=time.time()\nmodel=LGBMRegressor()\nmodel.fit(train[features].values,train[].values)\nmodel.predict(test[features].values)\nend_time=time.time()\n()\n</code></pre>",
      "votes": 1,
      "replies": [
        {
          "id": 3029287,
          "author_name": "chumajin",
          "author_url": "",
          "post_date": "2024-10-27T05:25:03.410000",
          "content": "<p><a href=\"https://www.kaggle.com/yunsuxiaozi\" target=\"_blank\">@yunsuxiaozi</a> Thank you for comment. This code is something wrong. When I check the following code, which changes the order of evaluation for pandas and numpy. I got 151 sec and 286 sec. Maybe initialization is required.</p>\n<pre><code>import polars as pl to pandas, but with better performance when dealing with large datasets.\nimport pandas as pd csv,parquet\nimport numpy as np scientific computation of matrices\n(lgb,xgb,cat)\nfrom  lightgbm import LGBMRegressor,LGBMClassifier,log_evaluation,early_stopping\nimport \n\ntrain=pl()\ntrain=train()\nfeatures=\ntest=train]\ntrain=train]\n,(test))\n\nstart_time=()\nmodel=()\nmodel(train,train.values)\nmodel(test.values)\nend_time=()\n\n\nstart_time=()\nmodel=()\nmodel(train,train)\nmodel(test)\nend_time=()\n\n</code></pre>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3030925,
      "author_name": "Bhavesh@tf",
      "author_url": "",
      "post_date": "2024-10-29T04:00:35.280000",
      "content": "<p>Here are some optimization methods you can implement to speed up inference times even further for your LightGBM model:</p>\n<p>Quantize the Model: Use LightGBM's built-in model quantization (quantize) to reduce precision, which speeds up both memory usage and inference time without major loss in accuracy. <br>\nmodel = lgb.Booster(model_file='your_model.txt')<br>\nmodel.quantize()<br>\nReduce Data Transfer: By running inference directly on the GPU, data transfer between CPU and GPU is minimized, which helps in achieving faster inference. For this:</p>\n<p>Use predict with parameter pred_early_stop=True (if compatible with your model setup) to reduce computation by stopping predictions early for instances that are highly confident.<br>\nAvoid Row-by-Row Inference: Use batch prediction instead of processing row-by-row. Process multiple rows together to maximize hardware utilization:</p>\n<p>preds = model.predict(test.select(features).to_numpy(), num_iteration=model.best_iteration)<br>\nThreading and Parallelization: Increase LightGBM's parallel processing with n_jobs to leverage all CPU cores:</p>\n<p>model.predict(test.select(features).to_numpy(), num_iteration=model.best_iteration, n_jobs=-1)<br>\nUse dart for Smaller Trees: If your model’s configuration allows, experiment with the dart boosting type, which can sometimes reduce inference time, as it often produces smaller trees.</p>\n<p>Optimize Feature Selection: Drop unnecessary features and optimize feature engineering by selecting only essential features for your prediction, reducing the data dimensions handled during inference.</p>\n<p>These combined approaches can further reduce your inference times below 16 minutes if they align with your competition requirements and hardware.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3030938,
          "author_name": "chumajin",
          "author_url": "",
          "post_date": "2024-10-29T04:37:04.743000",
          "content": "<p><a href=\"https://www.kaggle.com/bhaveshtf\" target=\"_blank\">@bhaveshtf</a> Thank you for all the insights. There was so much I didn’t know. I'll give it a try!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3031058,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-10-29T09:18:29.593000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3031550,
          "author_name": "JM",
          "author_url": "",
          "post_date": "2024-10-29T20:08:41.847000",
          "content": "<p>ChatGPT response? pred_early_stop ︎is only for classification problems 😆</p>\n<blockquote>\n  <p>Quantize the Model: Use LightGBM's built-in model quantization (quantize) to reduce precision, which speeds up both memory usage and inference time without major loss in accuracy.<br>\n  model = lgb.Booster(model_file='your_model.txt')<br>\n  model.quantize()</p>\n</blockquote>\n<p>No such method exists: <a href=\"https://lightgbm.readthedocs.io/en/latest/pythonapi/lightgbm.Booster.html\" target=\"_blank\">https://lightgbm.readthedocs.io/en/latest/pythonapi/lightgbm.Booster.html</a></p>",
          "votes": 8,
          "replies": []
        }
      ]
    },
    {
      "id": 3029289,
      "author_name": "chumajin",
      "author_url": "",
      "post_date": "2024-10-27T05:28:54.070000",
      "content": "<p>Updated : I can achieve 2 times faster (16 min) like this in inference code. The difference from above is that it involves converting to NumPy in advance before making predictions.</p>\n<pre><code>def (: .DataFrame, lags: .DataFrame | None) -&gt; .DataFrame | pd.DataFrame:\n\n    test_X = .select(features).to_numpy()\n\n    preds = model.(test_X)\n\n    predictions = .select(\n        'row_id',\n        .Series(preds).alias('responder_6'),\n    )\n\n     predictions\n</code></pre>",
      "votes": 0,
      "replies": [
        {
          "id": 3030191,
          "author_name": "Victor Yip",
          "author_url": "",
          "post_date": "2024-10-28T09:27:17.727000",
          "content": "<p>Is this not identical to what you had before or does assigning it to a variable first somehow reduce compute</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3030426,
              "author_name": "chumajin",
              "author_url": "",
              "post_date": "2024-10-28T14:18:37.403000",
              "content": "<p><a href=\"https://www.kaggle.com/yipvictor\" target=\"_blank\">@yipvictor</a> We might avoid unnecessary overhead and speed up inference, due to differences in internal processing like type checking and data conversion.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3029734,
      "author_name": "Humayra Khanom Rime",
      "author_url": "",
      "post_date": "2024-10-27T16:57:20.870000",
      "content": "<p>Thank you for such an incredible idea.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3028779,
      "author_name": "Farhan Kardan",
      "author_url": "",
      "post_date": "2024-10-26T13:54:23.417000",
      "content": "<p>Thanks such an amazing idea</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3028627": "In this topic, I will explain how I reduced the inference time of my LightGBM single model from 32 minutes to 18 minutes (updated 16 min, please see [comment](https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/542697#3029289) ). \n\n**In short, convert your data to NumPy instead of using pandas, then train and perform inference with LightGBM.**\n\nIn this competition, test data is provided using Polars, so converting it to pandas for inference slows down the process. However, by converting it directly to NumPy from the start and then performing training and inference, you can achieve faster inference.\n\nThe following is example.\n\n□ training\n\ntrain : polars DataFrame, valid : polas DataFrame\n\n~~~\nlgb_train = lgb.Dataset(train.select(features).to_numpy(), train.select(target).to_numpy())\nlgb_valid = lgb.Dataset(valid.select(features).to_numpy(), valid.select(target).to_numpy())\n~~~\n\n□ inference (please see updated [comment](https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/542697#3029289) )\n\n~~~\ndef predict(test: pl.DataFrame, lags: pl.DataFrame | None) -> pl.DataFrame | pd.DataFrame:\n\n    preds = model.predict(test.select(features).to_numpy())\n\n    predictions = test.select(\n        'row_id',\n        pl.Series(preds).alias('responder_6'),\n    )\n\n    return predictions\n~~~\n\nAdditionally, there was a discussion in this [topic](https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/542022) about the simulation time not matching the submission time. However, in my case, the [simulation](https://www.kaggle.com/code/chumajin/janestreet-updated-simulator-for-time-series-api) using the Kaggle evaluation API aligns quite well with this submission method.\n\nConventinal LGBM using pandas : simulation 27 min, submission 32 min\nThis LGBM using numpy(above predict function) : simulation 18 min, submission 18 min\n\nFurthermore, if you are not using lags, removing the following code can make the process about 2 minutes faster.\n\n~~~\nglobal lags_\n    if lags is not None:\n        lags_ = lags\n~~~\n\nI hope this helps you.\nEnjoy kaggling ~ .",
    "3028696": "If you want to go even faster there is some ways to do GPU inference of LGBM\n\n~5x Faster Inference - https://opensource.microsoft.com/blog/2020/09/29/accelerate-machine-learning-models-gpu-onnx-runtime-hummingbird/\n\n~10x Faster Inference - https://github.com/siboehm/lleaves\n\nI have used some of these Tree Inference packages before and can confirm they do work (don't neccesarily get the speeds they claim but usually at least 2x) - however I would only really use them if you have a very large tree.\n\n",
    "3030192": "Do you have any recommendations on OOM issues? Can't seem to load and train the entire dataset in the Kaggle notebook environment.",
    "3029167": "May I ask if I misunderstood the meaning? My experimental results are not ideal? runtime is 155 and 284.\n```python\nimport polars as pl#similar to pandas, but with better performance when dealing with large datasets.\nimport pandas as pd#read csv,parquet\nimport numpy as np#for scientific computation of matrices\n#models(lgb,xgb,cat)\nfrom  lightgbm import LGBMRegressor,LGBMClassifier,log_evaluation,early_stopping\nimport time\n\ntrain=pl.read_parquet(\"/kaggle/input/jane-street-real-time-market-data-forecasting/train.parquet/partition_id=9/part-0.parquet\")\ntrain=train.to_pandas()\nfeatures=[f'feature_0{i}'if i<10  else f'feature_{i}' for i in range(79)]\ntest=train[-2000000:][features+['responder_6']]\ntrain=train[:-2000000][features+['responder_6']]\nprint(len(train),len(test))\n\nstart_time=time.time()\nmodel=LGBMRegressor()\nmodel.fit(train[features],train['responder_6'])\nmodel.predict(test[features])\nend_time=time.time()\nprint(f\"run_time:{end_time-start_time}\")\n\nstart_time=time.time()\nmodel=LGBMRegressor()\nmodel.fit(train[features].values,train['responder_6'].values)\nmodel.predict(test[features].values)\nend_time=time.time()\nprint(f\"run_time:{end_time-start_time}\")\n```",
    "3030925": "Here are some optimization methods you can implement to speed up inference times even further for your LightGBM model:\n\nQuantize the Model: Use LightGBM's built-in model quantization (quantize) to reduce precision, which speeds up both memory usage and inference time without major loss in accuracy. \nmodel = lgb.Booster(model_file='your_model.txt')\nmodel.quantize()\nReduce Data Transfer: By running inference directly on the GPU, data transfer between CPU and GPU is minimized, which helps in achieving faster inference. For this:\n\nUse predict with parameter pred_early_stop=True (if compatible with your model setup) to reduce computation by stopping predictions early for instances that are highly confident.\nAvoid Row-by-Row Inference: Use batch prediction instead of processing row-by-row. Process multiple rows together to maximize hardware utilization:\n\n\npreds = model.predict(test.select(features).to_numpy(), num_iteration=model.best_iteration)\nThreading and Parallelization: Increase LightGBM's parallel processing with n_jobs to leverage all CPU cores:\n\n\nmodel.predict(test.select(features).to_numpy(), num_iteration=model.best_iteration, n_jobs=-1)\nUse dart for Smaller Trees: If your model’s configuration allows, experiment with the dart boosting type, which can sometimes reduce inference time, as it often produces smaller trees.\n\nOptimize Feature Selection: Drop unnecessary features and optimize feature engineering by selecting only essential features for your prediction, reducing the data dimensions handled during inference.\n\nThese combined approaches can further reduce your inference times below 16 minutes if they align with your competition requirements and hardware.",
    "3029289": "Updated : I can achieve 2 times faster (16 min) like this in inference code. The difference from above is that it involves converting to NumPy in advance before making predictions.\n\n~~~\ndef predict(test: pl.DataFrame, lags: pl.DataFrame | None) -> pl.DataFrame | pd.DataFrame:\n    \n    test_X = test.select(features).to_numpy()\n          \n    preds = model.predict(test_X)\n\n    predictions = test.select(\n        'row_id',\n        pl.Series(preds).alias('responder_6'),\n    )\n\n    return predictions\n~~~",
    "3029734": "Thank you for such an incredible idea.",
    "3028779": "Thanks such an amazing idea"
  }
}