{
  "id": 328606,
  "title": "Speed Up XGB, CatBoost, and LGBM by 20x",
  "url": "/competitions/amex-default-prediction/discussion/328606",
  "author_name": "Chris Deotte",
  "post_date": "2022-06-02T04:08:33.937000",
  "votes": 211,
  "comment_count": 43,
  "views": 0,
  "content": "<h1>CatBoost GPU</h1>\n<p>Note that the best public CatBoost notebook does not use GPU. It is running on CPU and takes 8 hours to train 5-Folds. To turn on GPU, first choose <code>accelerator = GPU</code> in the right sidebar of the Kaggle jupyter notebook. Next add <code>task_type = 'GPU'</code> to the creation of the classifier:</p>\n<pre><code>clf = CatBoostClassifier(iterations=5000, random_state=22, \n                         task_type = 'GPU')\n</code></pre>\n<p>That's it! Afterward, CatBoost will only take 25 minutes to train 5-folds instead of 8 hours! Wow! </p>\n<p>When running in a Kaggle notebook, we will need to optimize memory usage so that the notebook doesn't have memory errors. If you start with the most popular public notebook, add this to the end of each K-fold for-loop to avoid GPU memory error:</p>\n<pre><code>import gc\ndel clf, tr_x, val_x, tr_y, val_y, preds, preds_test\ngc.collect()\n</code></pre>\n<p>And if you run CatBoost offline, note that CatBoost will use all your available GPUs and train even faster!</p>\n<h1>XGB GPU</h1>\n<p>We can also train fast boosted trees with XGB on GPU. To use GPU with XGB, add the two following parameters to your XGB code:</p>\n<pre><code>xgb_parms = { \n    'tree_method':'gpu_hist',\n    'predictor':'gpu_predictor',\n}\nmodel = xgb.train(xgb_parms)\n</code></pre>\n<p>I posted a starter notebook <a href=\"https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\" target=\"_blank\">here</a>. It trains 5-folds in a total of 9 minutes! And it achieves CV 0.792 and LB 0.794! That's fast and accurate!</p>\n<h1>LGBM GPU</h1>\n<p>And lastly, we can install and use LGBM GPU version to accelerate LGBM. More info <a href=\"https://www.kaggle.com/competitions/santander-customer-transaction-prediction/discussion/89004\" target=\"_blank\">here</a>. </p>\n<p><strong>UPDATE 1</strong>: to use LGBM GPU we do not need to reinstall LGBM as the outdated link suggests. We can just add parameters <code>'device': 'gpu', 'gpu_platform_id': 0, and 'gpu_device_id': 0</code>. However when i try this in Kaggle notebooks, and watch GPU usage it doesn't seem to use much GPU. So perhaps LGBM doesn't give much speed up. We will need to investigate this. </p>\n<p><strong>UPDATE 2</strong>: Using LGBM with Kaggle's P100 GPU is 1.67x faster than Kaggle's 2-core CPU and 1.33x faster than Kaggle's 4-core CPU. See comments below for details.</p>",
  "messages": [
    {
      "id": 1808640,
      "postDate": "2022-06-02T04:08:33.937Z",
      "content": "<h1>CatBoost GPU</h1>\n<p>Note that the best public CatBoost notebook does not use GPU. It is running on CPU and takes 8 hours to train 5-Folds. To turn on GPU, first choose <code>accelerator = GPU</code> in the right sidebar of the Kaggle jupyter notebook. Next add <code>task_type = 'GPU'</code> to the creation of the classifier:</p>\n<pre><code>clf = CatBoostClassifier(iterations=5000, random_state=22, \n                         task_type = 'GPU')\n</code></pre>\n<p>That's it! Afterward, CatBoost will only take 25 minutes to train 5-folds instead of 8 hours! Wow! </p>\n<p>When running in a Kaggle notebook, we will need to optimize memory usage so that the notebook doesn't have memory errors. If you start with the most popular public notebook, add this to the end of each K-fold for-loop to avoid GPU memory error:</p>\n<pre><code>import gc\ndel clf, tr_x, val_x, tr_y, val_y, preds, preds_test\ngc.collect()\n</code></pre>\n<p>And if you run CatBoost offline, note that CatBoost will use all your available GPUs and train even faster!</p>\n<h1>XGB GPU</h1>\n<p>We can also train fast boosted trees with XGB on GPU. To use GPU with XGB, add the two following parameters to your XGB code:</p>\n<pre><code>xgb_parms = { \n    'tree_method':'gpu_hist',\n    'predictor':'gpu_predictor',\n}\nmodel = xgb.train(xgb_parms)\n</code></pre>\n<p>I posted a starter notebook <a href=\"https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\" target=\"_blank\">here</a>. It trains 5-folds in a total of 9 minutes! And it achieves CV 0.792 and LB 0.794! That's fast and accurate!</p>\n<h1>LGBM GPU</h1>\n<p>And lastly, we can install and use LGBM GPU version to accelerate LGBM. More info <a href=\"https://www.kaggle.com/competitions/santander-customer-transaction-prediction/discussion/89004\" target=\"_blank\">here</a>. </p>\n<p><strong>UPDATE 1</strong>: to use LGBM GPU we do not need to reinstall LGBM as the outdated link suggests. We can just add parameters <code>'device': 'gpu', 'gpu_platform_id': 0, and 'gpu_device_id': 0</code>. However when i try this in Kaggle notebooks, and watch GPU usage it doesn't seem to use much GPU. So perhaps LGBM doesn't give much speed up. We will need to investigate this. </p>\n<p><strong>UPDATE 2</strong>: Using LGBM with Kaggle's P100 GPU is 1.67x faster than Kaggle's 2-core CPU and 1.33x faster than Kaggle's 4-core CPU. See comments below for details.</p>",
      "rawMarkdown": "# CatBoost GPU\nNote that the best public CatBoost notebook does not use GPU. It is running on CPU and takes 8 hours to train 5-Folds. To turn on GPU, first choose `accelerator = GPU` in the right sidebar of the Kaggle jupyter notebook. Next add `task_type = 'GPU'` to the creation of the classifier:\n\n    clf = CatBoostClassifier(iterations=5000, random_state=22, \n                             task_type = 'GPU')\n\nThat's it! Afterward, CatBoost will only take 25 minutes to train 5-folds instead of 8 hours! Wow! \n\nWhen running in a Kaggle notebook, we will need to optimize memory usage so that the notebook doesn't have memory errors. If you start with the most popular public notebook, add this to the end of each K-fold for-loop to avoid GPU memory error:\n\n    import gc\n    del clf, tr_x, val_x, tr_y, val_y, preds, preds_test\n    gc.collect()\n\nAnd if you run CatBoost offline, note that CatBoost will use all your available GPUs and train even faster!\n\n# XGB GPU\nWe can also train fast boosted trees with XGB on GPU. To use GPU with XGB, add the two following parameters to your XGB code:\n\n    xgb_parms = { \n        'tree_method':'gpu_hist',\n        'predictor':'gpu_predictor',\n    }\n    model = xgb.train(xgb_parms)\n\n\nI posted a starter notebook [here][1]. It trains 5-folds in a total of 9 minutes! And it achieves CV 0.792 and LB 0.794! That's fast and accurate!\n\n# LGBM GPU\nAnd lastly, we can install and use LGBM GPU version to accelerate LGBM. More info [here][2]. \n\n**UPDATE 1**: to use LGBM GPU we do not need to reinstall LGBM as the outdated link suggests. We can just add parameters `'device': 'gpu', 'gpu_platform_id': 0, and 'gpu_device_id': 0`. However when i try this in Kaggle notebooks, and watch GPU usage it doesn't seem to use much GPU. So perhaps LGBM doesn't give much speed up. We will need to investigate this. \n\n**UPDATE 2**: Using LGBM with Kaggle's P100 GPU is 1.67x faster than Kaggle's 2-core CPU and 1.33x faster than Kaggle's 4-core CPU. See comments below for details.\n\n[1]: https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\n[2]: https://www.kaggle.com/competitions/santander-customer-transaction-prediction/discussion/89004\n[3]: https://www.kaggle.com/code/youheitomio/gpu-version-amex-lgbm-by-youhei-tomio\n[4]: https://www.kaggle.com/code/youheitomio/gpu-version-amex-lgbm-by-youhei-tomio?scriptVersionId=99773844\n",
      "votes": 211
    },
    {
      "id": 1808995,
      "postDate": "2022-06-02T09:59:25.590Z",
      "content": "<p>Didn't try lgbm with gpu recently by as far as I remember GPU speed was just slightly better in previous lgbm versions. </p>\n<p>--</p>\n<p>Catboost and XGB -&gt; GPU for sure (if have some resources)</p>",
      "rawMarkdown": "Didn't try lgbm with gpu recently by as far as I remember GPU speed was just slightly better in previous lgbm versions. \n\n--\n\nCatboost and XGB -> GPU for sure (if have some resources)",
      "votes": 10,
      "replies": [
        {
          "id": 1885538,
          "postDate": "2022-08-05T08:45:08.800Z",
          "content": "<p>Konstantin, greetings!<br>\nDo You use catboost on GPU with amex metric?<br>\nOn my PC it causes an error:<br>\n<em>CatBoostError: User defined loss functions, metrics and callbacks are not supported for GPU</em></p>",
          "rawMarkdown": "Konstantin, greetings!\nDo You use catboost on GPU with amex metric?\nOn my PC it causes an error:\n*CatBoostError: User defined loss functions, metrics and callbacks are not supported for GPU*",
          "votes": 1
        }
      ]
    },
    {
      "id": 1840983,
      "postDate": "2022-07-02T18:05:15.377Z",
      "content": "<p><strong>UPDATE 1</strong>: Youhei Tomio, Takuya Fushihara, Yuzuki Kitajima posted a LGBM GPU Kaggle notebook <a href=\"https://www.kaggle.com/code/youheitomio/gpu-version-amex-lgbm-by-youhei-tomio\" target=\"_blank\">here</a> where they reinstall a GPU version of LGBM. Check out notebook <a href=\"https://www.kaggle.com/code/youheitomio/gpu-version-amex-lgbm-by-youhei-tomio?scriptVersionId=99773844\" target=\"_blank\">version 5</a> where each fold takes 5 minutes. (In the newer versions they are adding DART which slows down both CPU and GPU LGBM). </p>\n<p><strong>UPDATE 2</strong>: I ran some experiments to determine LGBM runtimes. Here are the results. Using LGBM with GPU does not require re-installing LGBM GPU. If we use the LGBM that is already installed in Kaggle's notebook it is just as fast as reinstalling LGBM-GPU. Using LGBM with Kaggle's P100 GPU is 1.67x faster than Kaggle's 2-core CPU and 1.33x faster than Kaggle's 4-core CPU. See comments below for details.</p>",
      "rawMarkdown": "**UPDATE 1**: Youhei Tomio, Takuya Fushihara, Yuzuki Kitajima posted a LGBM GPU Kaggle notebook [here][3] where they reinstall a GPU version of LGBM. Check out notebook [version 5][2] where each fold takes 5 minutes. (In the newer versions they are adding DART which slows down both CPU and GPU LGBM). \n\n**UPDATE 2**: I ran some experiments to determine LGBM runtimes. Here are the results. Using LGBM with GPU does not require re-installing LGBM GPU. If we use the LGBM that is already installed in Kaggle's notebook it is just as fast as reinstalling LGBM-GPU. Using LGBM with Kaggle's P100 GPU is 1.67x faster than Kaggle's 2-core CPU and 1.33x faster than Kaggle's 4-core CPU. See comments below for details.\n\n[2]: https://www.kaggle.com/code/youheitomio/gpu-version-amex-lgbm-by-youhei-tomio?scriptVersionId=99773844\n[3]: https://www.kaggle.com/code/youheitomio/gpu-version-amex-lgbm-by-youhei-tomio",
      "votes": 5,
      "replies": [
        {
          "id": 1841030,
          "postDate": "2022-07-02T19:31:45.153Z",
          "content": "<p>setting a <code>gpu</code> param does not mean gpu is used. run it yourself. gpu for lightgbm is on windows machines only</p>\n<p>Update: reading docs it seems gpu for linux is possible. However kaggle notebooks does not it have it compiled to run?</p>",
          "rawMarkdown": "setting a `gpu` param does not mean gpu is used. run it yourself. gpu for lightgbm is on windows machines only\n\nUpdate: reading docs it seems gpu for linux is possible. However kaggle notebooks does not it have it compiled to run?",
          "votes": 4
        },
        {
          "id": 1841118,
          "postDate": "2022-07-02T20:37:54.047Z",
          "content": "<p>Thanks for the heads up, i'll check it out. It has been a while since I ran LGBM GPU in a Kaggle notebook. I remember doing it in Santander Comp 3 years ago following the instructions in this notebook <a href=\"https://www.kaggle.com/code/vinhnguyen/gpu-acceleration-for-lightgbm/notebook\" target=\"_blank\">here</a> (and I checked and it was using GPU). Back then you needed to re-install LGBM with GPU support. I'm not sure what is required today.</p>",
          "rawMarkdown": "Thanks for the heads up, i'll check it out. It has been a while since I ran LGBM GPU in a Kaggle notebook. I remember doing it in Santander Comp 3 years ago following the instructions in this notebook [here][1] (and I checked and it was using GPU). Back then you needed to re-install LGBM with GPU support. I'm not sure what is required today.\n\n[1]: https://www.kaggle.com/code/vinhnguyen/gpu-acceleration-for-lightgbm/notebook"
        },
        {
          "id": 1841135,
          "postDate": "2022-07-02T20:47:23.977Z",
          "content": "<p><a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> I checked it out. Version 5 of that notebook does indeed use the GPU in Kaggle notebook. Note that that Kaggle notebook reinstalls LGBM GPU version before running LGBM with <code>gpu</code> parameter.<br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Jun-2022/gpu.png\" alt=\"\"></p>",
          "rawMarkdown": "@raddar I checked it out. Version 5 of that notebook does indeed use the GPU in Kaggle notebook. Note that that Kaggle notebook reinstalls LGBM GPU version before running LGBM with `gpu` parameter.\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Jun-2022/gpu.png)",
          "votes": 2
        },
        {
          "id": 1841136,
          "postDate": "2022-07-02T20:49:09.243Z",
          "content": "<p>I'm not sure how efficiently LGBM uses the GPU. In the image above, it is only using 10% of the GPU. When i get some time, i can run some time tests to compare CPU vs. GPU speeds. But i have heard from others that LGBM doesn't achieve as big a speed up from GPU as XGB and CatBoost do.</p>",
          "rawMarkdown": "I'm not sure how efficiently LGBM uses the GPU. In the image above, it is only using 10% of the GPU. When i get some time, i can run some time tests to compare CPU vs. GPU speeds. But i have heard from others that LGBM doesn't achieve as big a speed up from GPU as XGB and CatBoost do."
        },
        {
          "id": 1841139,
          "postDate": "2022-07-02T20:51:59.690Z",
          "content": "<p>These GPU spikes are rare. i monitored how one fold trained. mostly its 0%. there were like 10 spikes of GPU usage in that time</p>",
          "rawMarkdown": "These GPU spikes are rare. i monitored how one fold trained. mostly its 0%. there were like 10 spikes of GPU usage in that time",
          "votes": 2
        },
        {
          "id": 1841155,
          "postDate": "2022-07-02T21:17:36.793Z",
          "content": "<p>There are 2 possible explanations:</p>\n<ol>\n<li>Amex custom metric doesn’t use multi threading and also could be cpu one - so model constantly switching between cpu and gpu</li>\n<li>Lgbm just doesn’t  like gpu)))</li>\n</ol>",
          "rawMarkdown": "There are 2 possible explanations:\n1. Amex custom metric doesn’t use multi threading and also could be cpu one - so model constantly switching between cpu and gpu\n2. Lgbm just doesn’t  like gpu)))",
          "votes": 1
        },
        {
          "id": 1841343,
          "postDate": "2022-07-03T03:02:40.420Z",
          "content": "<p><a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> <a href=\"https://www.kaggle.com/kyakovlev\" target=\"_blank\">@kyakovlev</a> I compared run times. I started with the notebook <a href=\"https://www.kaggle.com/code/youheitomio/gpu-version-amex-lgbm-by-youhei-tomio?scriptVersionId=99773844\" target=\"_blank\">here</a> version 5 without DART. For each experiment, I trained 3 folds for 1100 iterations. The train dataset has 458913 rows and  918 columns:</p>\n<ul>\n<li>15 minutes using GPU and Kaggle's LGBM</li>\n<li>15 minutes using GPU and re-installing LGBM-GPU</li>\n<li>25 minutes using CPU 2-core</li>\n<li>20 minutes using CPU 4-core</li>\n</ul>\n<p>So, using GPU was faster but not by much. Using Kaggle's installed LGBM with parameter <code>gpu</code> was just as fast as re-installing the LGBM with GPU support and then using parameter <code>gpu</code>. When using a Kaggle GPU notebook which has GPU Nvidia P100 and CPU 2-core, we see that GPU was 1.67x times faster than the corresponding CPU 2-core. When running LGBM in a Kaggle CPU notebook, then we have CPU 4-core and that compared to GPU notebook was only 1.33x times faster.</p>\n<p>I'm curious does the LGBM Windows GPU version achieve more speed up?</p>",
          "rawMarkdown": "@raddar @kyakovlev I compared run times. I started with the notebook [here][1] version 5 without DART. For each experiment, I trained 3 folds for 1100 iterations. The train dataset has 458913 rows and  918 columns:\n\n* 15 minutes using GPU and Kaggle's LGBM\n* 15 minutes using GPU and re-installing LGBM-GPU\n* 25 minutes using CPU 2-core\n* 20 minutes using CPU 4-core\n\nSo, using GPU was faster but not by much. Using Kaggle's installed LGBM with parameter `gpu` was just as fast as re-installing the LGBM with GPU support and then using parameter `gpu`. When using a Kaggle GPU notebook which has GPU Nvidia P100 and CPU 2-core, we see that GPU was 1.67x times faster than the corresponding CPU 2-core. When running LGBM in a Kaggle CPU notebook, then we have CPU 4-core and that compared to GPU notebook was only 1.33x times faster.\n\nI'm curious does the LGBM Windows GPU version achieve more speed up?\n\n[1]: https://www.kaggle.com/code/youheitomio/gpu-version-amex-lgbm-by-youhei-tomio?scriptVersionId=99773844",
          "votes": 3
        }
      ]
    },
    {
      "id": 1882195,
      "postDate": "2022-08-03T05:54:16.513Z",
      "content": "<p>I was looking this ! Thanks. I've saved my time</p>",
      "rawMarkdown": "I was looking this ! Thanks. I've saved my time",
      "votes": 1
    },
    {
      "id": 1842528,
      "postDate": "2022-07-04T05:16:51.397Z",
      "content": "<p>Thats is what I was looking for. <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thank you  very much for sharing this excellent information with us. It has been pretty useful for me indeed. </p>",
      "rawMarkdown": "Thats is what I was looking for. @cdeotte Thank you  very much for sharing this excellent information with us. It has been pretty useful for me indeed. ",
      "votes": 1
    },
    {
      "id": 1841316,
      "postDate": "2022-07-03T02:01:53.177Z",
      "content": "<p>Thanks for the tip! Can you please suggest if I can use gpu with stacking classifier or not? <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
      "rawMarkdown": "Thanks for the tip! Can you please suggest if I can use gpu with stacking classifier or not? @cdeotte ",
      "votes": 1,
      "replies": [
        {
          "id": 1841348,
          "postDate": "2022-07-03T03:13:50.850Z",
          "content": "<p>Yes you can use GPU with all GBT stacking classifiers (XGB, CatBoost, LGBM), all NN stacking classifiers (PyTorch, TensorFlow), and all ML stacking classifiers (RAPIDS cuML).</p>",
          "rawMarkdown": "Yes you can use GPU with all GBT stacking classifiers (XGB, CatBoost, LGBM), all NN stacking classifiers (PyTorch, TensorFlow), and all ML stacking classifiers (RAPIDS cuML).",
          "votes": 3
        }
      ]
    },
    {
      "id": 1816456,
      "postDate": "2022-06-10T08:06:28.353Z",
      "content": "<p>Wow, thanks a lot. BTW, How does lgbm call the GPU in kaggle notebook?</p>",
      "rawMarkdown": "Wow, thanks a lot. BTW, How does lgbm call the GPU in kaggle notebook?",
      "votes": 1
    },
    {
      "id": 1909310,
      "postDate": "2022-08-22T14:02:05.543Z",
      "content": "<p>Thanks for sharing! I also tried to run different models on GPUs, and the one I'm using is RTX 3090 which has 25.4 GB VRAM.  It was doing OK when running XGB model as it speeded up the whole parameter tuning process. But when I switched to Catboost, it just squeezed all the VRAM and the notebook restarted… I tried double the VRAM but the same thing happened. Can you tell us the specs of your GPUs? Thanks!</p>",
      "rawMarkdown": "Thanks for sharing! I also tried to run different models on GPUs, and the one I'm using is RTX 3090 which has 25.4 GB VRAM.  It was doing OK when running XGB model as it speeded up the whole parameter tuning process. But when I switched to Catboost, it just squeezed all the VRAM and the notebook restarted... I tried double the VRAM but the same thing happened. Can you tell us the specs of your GPUs? Thanks!",
      "votes": 2,
      "replies": [
        {
          "id": 1909444,
          "postDate": "2022-08-22T16:04:57.080Z",
          "content": "<p>If you are running out of GPU VRAM, i suggest using feature reduction to reduce the number of your features. There is a discussion <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/342284\" target=\"_blank\">here</a> about the number of features that teams are using.</p>",
          "rawMarkdown": "If you are running out of GPU VRAM, i suggest using feature reduction to reduce the number of your features. There is a discussion [here][1] about the number of features that teams are using.\n\n[1]: https://www.kaggle.com/competitions/amex-default-prediction/discussion/342284"
        }
      ]
    },
    {
      "id": 1814868,
      "postDate": "2022-06-08T12:24:59.367Z",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> thanks for sharing this. Will definitely explore these to speed up training workloads.</p>",
      "rawMarkdown": "@cdeotte thanks for sharing this. Will definitely explore these to speed up training workloads.",
      "votes": 1
    },
    {
      "id": 1813588,
      "postDate": "2022-06-07T03:09:34.953Z",
      "content": "<p>This sharing helps me explore more functions about ML, Thanks a lot!</p>",
      "rawMarkdown": "This sharing helps me explore more functions about ML, Thanks a lot!",
      "votes": 1,
      "replies": [
        {
          "id": 2527059,
          "postDate": "2023-11-16T08:51:32.113Z",
          "content": "<p><a href=\"https://www.kaggle.com/zyhchasel\" target=\"_blank\">@zyhchasel</a> when you raise your hand for a team, do reply somehow?!<br>\nin the context of this current competition <br>\n<a href=\"https://www.kaggle.com/competitions/optiver-trading-at-the-close/overview\" target=\"_blank\">https://www.kaggle.com/competitions/optiver-trading-at-the-close/overview</a></p>",
          "rawMarkdown": "@zyhchasel when you raise your hand for a team, do reply somehow?!\nin the context of this current competition \nhttps://www.kaggle.com/competitions/optiver-trading-at-the-close/overview",
          "replies": [
            {
              "id": 2527060,
              "postDate": "2023-11-16T08:51:52.547Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 1812597,
      "postDate": "2022-06-06T03:37:11.737Z",
      "content": "<p>thx for sharing</p>",
      "rawMarkdown": "thx for sharing",
      "votes": 1
    },
    {
      "id": 1811526,
      "postDate": "2022-06-04T19:16:30.370Z",
      "content": "<p>Really liked the garbage collection trick! We do not see this quite often in popular ML notebooks! </p>",
      "rawMarkdown": "Really liked the garbage collection trick! We do not see this quite often in popular ML notebooks! ",
      "votes": 1
    },
    {
      "id": 1809333,
      "postDate": "2022-06-02T16:07:53.187Z",
      "content": "<p>From my former experiments, there will be a performance drop if train CatBoost on GPU</p>",
      "rawMarkdown": "From my former experiments, there will be a performance drop if train CatBoost on GPU",
      "votes": 1,
      "replies": [
        {
          "id": 1809345,
          "postDate": "2022-06-02T16:18:27.443Z",
          "content": "<p>Interesting. Perhaps it depends on the features. If you train the most popular CatBoost public notebook on GPU vs. CPU then the CV improves to 0.792 instead of 0.790. So in this case, GPU improves the performance (maybe it helps remove dataset noise).</p>\n<p>If you observe performance drop, then you can run your experiments with GPU to allow you to test more features and then use CPU for your submission to LB.</p>",
          "rawMarkdown": "Interesting. Perhaps it depends on the features. If you train the most popular CatBoost public notebook on GPU vs. CPU then the CV improves to 0.792 instead of 0.790. So in this case, GPU improves the performance (maybe it helps remove dataset noise).\n\nIf you observe performance drop, then you can run your experiments with GPU to allow you to test more features and then use CPU for your submission to LB.",
          "votes": 5,
          "replies": [
            {
              "id": 2200942,
              "postDate": "2023-03-28T23:25:59.250Z",
              "content": "<p>I too noticed a difference in results in a recent playground competition using CatBoost GPU vs CPU, with the GPU being worse.  I finally discovered CatBoost uses different default parameters for GPU and CPU, and it appears impossible to force them to be identical as some of the parameters are unique to device type. Use the method get_all_params() to see the differences.</p>",
              "rawMarkdown": "I too noticed a difference in results in a recent playground competition using CatBoost GPU vs CPU, with the GPU being worse.  I finally discovered CatBoost uses different default parameters for GPU and CPU, and it appears impossible to force them to be identical as some of the parameters are unique to device type. Use the method get_all_params() to see the differences."
            }
          ]
        },
        {
          "id": 1809347,
          "postDate": "2022-06-02T16:20:33.880Z",
          "content": "<p>I think this is related with default bootstrap_type parameter which is 'MVS' for CPU and only can be used in CPU. For GPU it is 'Bernoulli'</p>",
          "rawMarkdown": "I think this is related with default bootstrap_type parameter which is 'MVS' for CPU and only can be used in CPU. For GPU it is 'Bernoulli'",
          "votes": 16
        }
      ]
    },
    {
      "id": 1851123,
      "postDate": "2022-07-11T03:24:18.417Z",
      "content": "<p>yes, I found that too.<br>\ngpu version of catboost and xgboost can speed up a lot! specially catboost can run very fast!<br>\nwhile lightgbm only speed up a little by gpu.<br>\neven 32 core cpus is not slow than gpu.</p>",
      "rawMarkdown": "yes, I found that too.\ngpu version of catboost and xgboost can speed up a lot! specially catboost can run very fast!\nwhile lightgbm only speed up a little by gpu.\neven 32 core cpus is not slow than gpu.",
      "votes": 2,
      "replies": [
        {
          "id": 1851366,
          "postDate": "2022-07-11T07:32:33.320Z",
          "content": "<p>Waovv 32 core yeah </p>",
          "rawMarkdown": "Waovv 32 core yeah ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1809498,
      "postDate": "2022-06-02T18:04:21.080Z",
      "content": "<p>Running catboost on 4 dual GPU systems using a huge swap file so that one of my machines can process the full data set.    With my current set of features the CPU ram hits around 150GB per System Monitor - runs at a decent speed (for sure would be faster if I had 256GB of hardware ram :) but does occasionally just hang during one of the folds.</p>\n<p>Realized with your tip of memory that I am not deleting any data frames - will add that ASAP.   Is there any other tricks you might have when using large swap files (on Ubuntu) ?</p>",
      "rawMarkdown": "Running catboost on 4 dual GPU systems using a huge swap file so that one of my machines can process the full data set.    With my current set of features the CPU ram hits around 150GB per System Monitor - runs at a decent speed (for sure would be faster if I had 256GB of hardware ram :) but does occasionally just hang during one of the folds.\n\nRealized with your tip of memory that I am not deleting any data frames - will add that ASAP.   Is there any other tricks you might have when using large swap files (on Ubuntu) ?",
      "votes": 2,
      "replies": [
        {
          "id": 1809521,
          "postDate": "2022-06-02T18:34:34.880Z",
          "content": "<p>The most important thing is to make the training data as small as possible (while still containing all the signal). This step is explained <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328054\" target=\"_blank\">here</a>.</p>\n<p>The next step is <strong>feature selection</strong>. Useless columns need to be located and removed. This will free much memory. For example, my XGB notebook <a href=\"https://tinyurl.com/28vk5xnn\" target=\"_blank\">here</a> uses all 918 features. However, if you look at the feature importance (dataframe is saved in output <a href=\"https://tinyurl.com/2p8me793\" target=\"_blank\">here</a>), only some are important. We need to remove the unimportant features before we continue with new experiments. </p>",
          "rawMarkdown": "The most important thing is to make the training data as small as possible (while still containing all the signal). This step is explained [here][1].\n\nThe next step is **feature selection**. Useless columns need to be located and removed. This will free much memory. For example, my XGB notebook [here][2] uses all 918 features. However, if you look at the feature importance (dataframe is saved in output [here][3]), only some are important. We need to remove the unimportant features before we continue with new experiments. \n\n[1]: https://www.kaggle.com/competitions/amex-default-prediction/discussion/328054\n[2]: https://tinyurl.com/28vk5xnn\n[3]: https://tinyurl.com/2p8me793",
          "votes": 6
        },
        {
          "id": 1809634,
          "postDate": "2022-06-02T21:59:21.400Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">Chris</a><br>\nGood answer but was not my question …</p>\n<blockquote>\n  <p>(while still containing all the signal)&gt; </p>\n</blockquote>\n<p>This the piece of interest - since I have 4 machines - three can be very happy running training data with as small a size as possible.  My interest is for the fourth machine which is very eager to run the full data set but it occasionally chokes for no apparent reason when 150GB - 200GB is the system monitor report of memory use (64GB hardware + the rest on an SSD swap file).   Was hoping you might have some general root cause / corrective actions for large swap files on Linux.</p>",
          "rawMarkdown": "[Chris](https://www.kaggle.com/cdeotte)\nGood answer but was not my question ...\n> (while still containing all the signal)> \n\nThis the piece of interest - since I have 4 machines - three can be very happy running training data with as small a size as possible.  My interest is for the fourth machine which is very eager to run the full data set but it occasionally chokes for no apparent reason when 150GB - 200GB is the system monitor report of memory use (64GB hardware + the rest on an SSD swap file).   Was hoping you might have some general root cause / corrective actions for large swap files on Linux.",
          "votes": 2
        },
        {
          "id": 1809635,
          "postDate": "2022-06-02T22:02:49.907Z",
          "content": "<p>Unfortunately, i don't know what is happening. Sounds frustrating.</p>",
          "rawMarkdown": "Unfortunately, i don't know what is happening. Sounds frustrating.",
          "votes": 1
        },
        {
          "id": 1809639,
          "postDate": "2022-06-02T22:11:20.127Z",
          "content": "<p>Thanks - Yep - I probably would be happier if it did not work all the time - success 80% of the time now leads me to keep trying.  If I find something that gets me to 100% success will put the fix here.  </p>",
          "rawMarkdown": "Thanks - Yep - I probably would be happier if it did not work all the time - success 80% of the time now leads me to keep trying.  If I find something that gets me to 100% success will put the fix here.  ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1882170,
      "postDate": "2022-08-03T05:37:29.097Z",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>  Wouldn't you consider finding your teammates to fight for the gold medal together? I see your ranking dropped to eighth. Looking forward to your gold medal.  👊   I see <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a>  is looking for teammates and looking forward to working with you. 😄</p>",
      "rawMarkdown": "@cdeotte  Wouldn't you consider finding your teammates to fight for the gold medal together? I see your ranking dropped to eighth. Looking forward to your gold medal.  👊   I see @raddar  is looking for teammates and looking forward to working with you. 😄"
    },
    {
      "id": 1831717,
      "postDate": "2022-06-24T11:27:50.270Z",
      "content": "<p>I often use classification algorithms. CatBoost and LightGBM doesn't support working with GPU for binary classification. Also, CatBoost developer (I think) shared: LOOOOOL 😂</p>\n<ul>\n<li>Hello!<br>\nCurrently we don't support RSM on GPU for non-pairwise losses, because we don't see need for it - rsm won't provide significant speedup, but it's implementation will require lot's of work, so for now - just don't set RSM while training on GPU.<br>\nThanks!</li>\n</ul>\n<p>CPU overall time about [01:43&lt;00:00,  4.81it/s] with 500 run iteration. But GPU added code spent [32:03&lt;00:00,  3.85s/it] 30 fold time. Really an enormous difference and bad performance. Performance metrics resulted same.</p>\n<p>With GPU:<br>\nTrain Accuracy score: 0.97<br>\nTrain ROC AUC score: 0.96<br>\nTest Accuracy score: 0.82<br>\nTest ROC-AUC score: 0.777</p>\n<p>Without GPU:<br>\nTrain Accuracy score: 0.97<br>\nTrain ROC AUC score: 0.96<br>\nTest Accuracy score: 0.82<br>\nTest ROC-AUC score: 0.777</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8699995%2Ffef474688c489a1a160d0a27b177027f%2FEkran%20Resmi%202022-06-24%2014.10.33.png?generation=1656069076329264&amp;alt=media\" alt=\"xgboost\"></p>\n<p>My device is Colab Pro GPU:</p>\n<p>+-----------------------------------------------------------------------------------------------+<br>\n| NVIDIA-SMI 460.32.03    Driver Version: 460.32.03    CUDA Version: 11.2     | 15109MiB            |<br>\n|-------------------------------+----------------------+----------------------------------------+</p>\n<p>I will trying with regression models.</p>",
      "rawMarkdown": "I often use classification algorithms. CatBoost and LightGBM doesn't support working with GPU for binary classification. Also, CatBoost developer (I think) shared: LOOOOOL 😂\n\n- Hello!\nCurrently we don't support RSM on GPU for non-pairwise losses, because we don't see need for it - rsm won't provide significant speedup, but it's implementation will require lot's of work, so for now - just don't set RSM while training on GPU.\nThanks!\n\nCPU overall time about [01:43<00:00,  4.81it/s] with 500 run iteration. But GPU added code spent [32:03<00:00,  3.85s/it] 30 fold time. Really an enormous difference and bad performance. Performance metrics resulted same.\n\nWith GPU:\nTrain Accuracy score: 0.97\nTrain ROC AUC score: 0.96\nTest Accuracy score: 0.82\nTest ROC-AUC score: 0.777\n\nWithout GPU:\nTrain Accuracy score: 0.97\nTrain ROC AUC score: 0.96\nTest Accuracy score: 0.82\nTest ROC-AUC score: 0.777\n\n![xgboost](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8699995%2Ffef474688c489a1a160d0a27b177027f%2FEkran%20Resmi%202022-06-24%2014.10.33.png?generation=1656069076329264&alt=media)\n\nMy device is Colab Pro GPU:\n\n+-----------------------------------------------------------------------------------------------+\n| NVIDIA-SMI 460.32.03    Driver Version: 460.32.03    CUDA Version: 11.2     | 15109MiB            |\n|-------------------------------+----------------------+----------------------------------------+\n\nI will trying with regression models."
    },
    {
      "id": 1809040,
      "postDate": "2022-06-02T10:35:04.093Z",
      "rawMarkdown": "",
      "votes": 3,
      "isDeleted": true
    },
    {
      "id": 1898123,
      "postDate": "2022-08-14T10:24:26.763Z",
      "content": "<p>Thanks for the tip! </p>",
      "rawMarkdown": "Thanks for the tip! ",
      "votes": 1
    },
    {
      "id": 1841194,
      "postDate": "2022-07-02T22:13:18.493Z",
      "content": "<p>Thank you for sharing!<br>\nUseful tips</p>",
      "rawMarkdown": "Thank you for sharing!\nUseful tips",
      "votes": 1
    },
    {
      "id": 1811889,
      "postDate": "2022-06-05T08:56:21.780Z",
      "content": "<p>Very useful! Thanks for sharing</p>",
      "rawMarkdown": "Very useful! Thanks for sharing",
      "votes": 1
    },
    {
      "id": 1811023,
      "postDate": "2022-06-04T07:23:08.700Z",
      "content": "<p>Thank you for sharing!</p>",
      "rawMarkdown": "Thank you for sharing!\n\n",
      "votes": 1
    },
    {
      "id": 1810611,
      "postDate": "2022-06-03T18:04:40.683Z",
      "content": "<p>Thank you for sharing!</p>",
      "rawMarkdown": "Thank you for sharing!",
      "votes": 1
    },
    {
      "id": 1809673,
      "postDate": "2022-06-03T00:07:34.723Z",
      "content": "<p>Thank you for sharing</p>",
      "rawMarkdown": "Thank you for sharing",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1808995,
      "author_name": "Konstantin Yakovlev",
      "author_url": "",
      "post_date": "2022-06-02T09:59:25.590000",
      "content": "<p>Didn't try lgbm with gpu recently by as far as I remember GPU speed was just slightly better in previous lgbm versions. </p>\n<p>--</p>\n<p>Catboost and XGB -&gt; GPU for sure (if have some resources)</p>",
      "votes": 10,
      "replies": [
        {
          "id": 1885538,
          "author_name": "Oleg Khudyakov",
          "author_url": "",
          "post_date": "2022-08-05T08:45:08.800000",
          "content": "<p>Konstantin, greetings!<br>\nDo You use catboost on GPU with amex metric?<br>\nOn my PC it causes an error:<br>\n<em>CatBoostError: User defined loss functions, metrics and callbacks are not supported for GPU</em></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1840983,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2022-07-02T18:05:15.377000",
      "content": "<p><strong>UPDATE 1</strong>: Youhei Tomio, Takuya Fushihara, Yuzuki Kitajima posted a LGBM GPU Kaggle notebook <a href=\"https://www.kaggle.com/code/youheitomio/gpu-version-amex-lgbm-by-youhei-tomio\" target=\"_blank\">here</a> where they reinstall a GPU version of LGBM. Check out notebook <a href=\"https://www.kaggle.com/code/youheitomio/gpu-version-amex-lgbm-by-youhei-tomio?scriptVersionId=99773844\" target=\"_blank\">version 5</a> where each fold takes 5 minutes. (In the newer versions they are adding DART which slows down both CPU and GPU LGBM). </p>\n<p><strong>UPDATE 2</strong>: I ran some experiments to determine LGBM runtimes. Here are the results. Using LGBM with GPU does not require re-installing LGBM GPU. If we use the LGBM that is already installed in Kaggle's notebook it is just as fast as reinstalling LGBM-GPU. Using LGBM with Kaggle's P100 GPU is 1.67x faster than Kaggle's 2-core CPU and 1.33x faster than Kaggle's 4-core CPU. See comments below for details.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1841030,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2022-07-02T19:31:45.153000",
          "content": "<p>setting a <code>gpu</code> param does not mean gpu is used. run it yourself. gpu for lightgbm is on windows machines only</p>\n<p>Update: reading docs it seems gpu for linux is possible. However kaggle notebooks does not it have it compiled to run?</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1841118,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-07-02T20:37:54.047000",
          "content": "<p>Thanks for the heads up, i'll check it out. It has been a while since I ran LGBM GPU in a Kaggle notebook. I remember doing it in Santander Comp 3 years ago following the instructions in this notebook <a href=\"https://www.kaggle.com/code/vinhnguyen/gpu-acceleration-for-lightgbm/notebook\" target=\"_blank\">here</a> (and I checked and it was using GPU). Back then you needed to re-install LGBM with GPU support. I'm not sure what is required today.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1841135,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-07-02T20:47:23.977000",
          "content": "<p><a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> I checked it out. Version 5 of that notebook does indeed use the GPU in Kaggle notebook. Note that that Kaggle notebook reinstalls LGBM GPU version before running LGBM with <code>gpu</code> parameter.<br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Jun-2022/gpu.png\" alt=\"\"></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1841136,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-07-02T20:49:09.243000",
          "content": "<p>I'm not sure how efficiently LGBM uses the GPU. In the image above, it is only using 10% of the GPU. When i get some time, i can run some time tests to compare CPU vs. GPU speeds. But i have heard from others that LGBM doesn't achieve as big a speed up from GPU as XGB and CatBoost do.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1841139,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2022-07-02T20:51:59.690000",
          "content": "<p>These GPU spikes are rare. i monitored how one fold trained. mostly its 0%. there were like 10 spikes of GPU usage in that time</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1841155,
          "author_name": "Konstantin Yakovlev",
          "author_url": "",
          "post_date": "2022-07-02T21:17:36.793000",
          "content": "<p>There are 2 possible explanations:</p>\n<ol>\n<li>Amex custom metric doesn’t use multi threading and also could be cpu one - so model constantly switching between cpu and gpu</li>\n<li>Lgbm just doesn’t  like gpu)))</li>\n</ol>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1841343,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-07-03T03:02:40.420000",
          "content": "<p><a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> <a href=\"https://www.kaggle.com/kyakovlev\" target=\"_blank\">@kyakovlev</a> I compared run times. I started with the notebook <a href=\"https://www.kaggle.com/code/youheitomio/gpu-version-amex-lgbm-by-youhei-tomio?scriptVersionId=99773844\" target=\"_blank\">here</a> version 5 without DART. For each experiment, I trained 3 folds for 1100 iterations. The train dataset has 458913 rows and  918 columns:</p>\n<ul>\n<li>15 minutes using GPU and Kaggle's LGBM</li>\n<li>15 minutes using GPU and re-installing LGBM-GPU</li>\n<li>25 minutes using CPU 2-core</li>\n<li>20 minutes using CPU 4-core</li>\n</ul>\n<p>So, using GPU was faster but not by much. Using Kaggle's installed LGBM with parameter <code>gpu</code> was just as fast as re-installing the LGBM with GPU support and then using parameter <code>gpu</code>. When using a Kaggle GPU notebook which has GPU Nvidia P100 and CPU 2-core, we see that GPU was 1.67x times faster than the corresponding CPU 2-core. When running LGBM in a Kaggle CPU notebook, then we have CPU 4-core and that compared to GPU notebook was only 1.33x times faster.</p>\n<p>I'm curious does the LGBM Windows GPU version achieve more speed up?</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1882195,
      "author_name": "Agustin222",
      "author_url": "",
      "post_date": "2022-08-03T05:54:16.513000",
      "content": "<p>I was looking this ! Thanks. I've saved my time</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1842528,
      "author_name": "Mustanger",
      "author_url": "",
      "post_date": "2022-07-04T05:16:51.397000",
      "content": "<p>Thats is what I was looking for. <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thank you  very much for sharing this excellent information with us. It has been pretty useful for me indeed. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1841316,
      "author_name": "Towhidul.Tonmoy",
      "author_url": "",
      "post_date": "2022-07-03T02:01:53.177000",
      "content": "<p>Thanks for the tip! Can you please suggest if I can use gpu with stacking classifier or not? <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1841348,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-07-03T03:13:50.850000",
          "content": "<p>Yes you can use GPU with all GBT stacking classifiers (XGB, CatBoost, LGBM), all NN stacking classifiers (PyTorch, TensorFlow), and all ML stacking classifiers (RAPIDS cuML).</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1816456,
      "author_name": "SgangX",
      "author_url": "",
      "post_date": "2022-06-10T08:06:28.353000",
      "content": "<p>Wow, thanks a lot. BTW, How does lgbm call the GPU in kaggle notebook?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1909310,
      "author_name": "LYCHEE",
      "author_url": "",
      "post_date": "2022-08-22T14:02:05.543000",
      "content": "<p>Thanks for sharing! I also tried to run different models on GPUs, and the one I'm using is RTX 3090 which has 25.4 GB VRAM.  It was doing OK when running XGB model as it speeded up the whole parameter tuning process. But when I switched to Catboost, it just squeezed all the VRAM and the notebook restarted… I tried double the VRAM but the same thing happened. Can you tell us the specs of your GPUs? Thanks!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1909444,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-08-22T16:04:57.080000",
          "content": "<p>If you are running out of GPU VRAM, i suggest using feature reduction to reduce the number of your features. There is a discussion <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/342284\" target=\"_blank\">here</a> about the number of features that teams are using.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1814868,
      "author_name": "Ankit Pal",
      "author_url": "",
      "post_date": "2022-06-08T12:24:59.367000",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> thanks for sharing this. Will definitely explore these to speed up training workloads.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1813588,
      "author_name": "NODE",
      "author_url": "",
      "post_date": "2022-06-07T03:09:34.953000",
      "content": "<p>This sharing helps me explore more functions about ML, Thanks a lot!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2527059,
          "author_name": "MAT",
          "author_url": "",
          "post_date": "2023-11-16T08:51:32.113000",
          "content": "<p><a href=\"https://www.kaggle.com/zyhchasel\" target=\"_blank\">@zyhchasel</a> when you raise your hand for a team, do reply somehow?!<br>\nin the context of this current competition <br>\n<a href=\"https://www.kaggle.com/competitions/optiver-trading-at-the-close/overview\" target=\"_blank\">https://www.kaggle.com/competitions/optiver-trading-at-the-close/overview</a></p>",
          "votes": 0,
          "replies": [
            {
              "id": 2527060,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-11-16T08:51:52.547000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 1812597,
      "author_name": "dragon zhang",
      "author_url": "",
      "post_date": "2022-06-06T03:37:11.737000",
      "content": "<p>thx for sharing</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1811526,
      "author_name": "Anubhav Chhabra",
      "author_url": "",
      "post_date": "2022-06-04T19:16:30.370000",
      "content": "<p>Really liked the garbage collection trick! We do not see this quite often in popular ML notebooks! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1809333,
      "author_name": "Waipaang",
      "author_url": "",
      "post_date": "2022-06-02T16:07:53.187000",
      "content": "<p>From my former experiments, there will be a performance drop if train CatBoost on GPU</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1809345,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-06-02T16:18:27.443000",
          "content": "<p>Interesting. Perhaps it depends on the features. If you train the most popular CatBoost public notebook on GPU vs. CPU then the CV improves to 0.792 instead of 0.790. So in this case, GPU improves the performance (maybe it helps remove dataset noise).</p>\n<p>If you observe performance drop, then you can run your experiments with GPU to allow you to test more features and then use CPU for your submission to LB.</p>",
          "votes": 5,
          "replies": [
            {
              "id": 2200942,
              "author_name": "Random Draw",
              "author_url": "",
              "post_date": "2023-03-28T23:25:59.250000",
              "content": "<p>I too noticed a difference in results in a recent playground competition using CatBoost GPU vs CPU, with the GPU being worse.  I finally discovered CatBoost uses different default parameters for GPU and CPU, and it appears impossible to force them to be identical as some of the parameters are unique to device type. Use the method get_all_params() to see the differences.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 1809347,
          "author_name": "huseyincotel",
          "author_url": "",
          "post_date": "2022-06-02T16:20:33.880000",
          "content": "<p>I think this is related with default bootstrap_type parameter which is 'MVS' for CPU and only can be used in CPU. For GPU it is 'Bernoulli'</p>",
          "votes": 16,
          "replies": []
        }
      ]
    },
    {
      "id": 1851123,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-11T03:24:18.417000",
      "content": "<p>yes, I found that too.<br>\ngpu version of catboost and xgboost can speed up a lot! specially catboost can run very fast!<br>\nwhile lightgbm only speed up a little by gpu.<br>\neven 32 core cpus is not slow than gpu.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1851366,
          "author_name": "Izzet Turkalp Akbasli",
          "author_url": "",
          "post_date": "2022-07-11T07:32:33.320000",
          "content": "<p>Waovv 32 core yeah </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1809498,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2022-06-02T18:04:21.080000",
      "content": "<p>Running catboost on 4 dual GPU systems using a huge swap file so that one of my machines can process the full data set.    With my current set of features the CPU ram hits around 150GB per System Monitor - runs at a decent speed (for sure would be faster if I had 256GB of hardware ram :) but does occasionally just hang during one of the folds.</p>\n<p>Realized with your tip of memory that I am not deleting any data frames - will add that ASAP.   Is there any other tricks you might have when using large swap files (on Ubuntu) ?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1809521,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-06-02T18:34:34.880000",
          "content": "<p>The most important thing is to make the training data as small as possible (while still containing all the signal). This step is explained <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328054\" target=\"_blank\">here</a>.</p>\n<p>The next step is <strong>feature selection</strong>. Useless columns need to be located and removed. This will free much memory. For example, my XGB notebook <a href=\"https://tinyurl.com/28vk5xnn\" target=\"_blank\">here</a> uses all 918 features. However, if you look at the feature importance (dataframe is saved in output <a href=\"https://tinyurl.com/2p8me793\" target=\"_blank\">here</a>), only some are important. We need to remove the unimportant features before we continue with new experiments. </p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1809634,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2022-06-02T21:59:21.400000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">Chris</a><br>\nGood answer but was not my question …</p>\n<blockquote>\n  <p>(while still containing all the signal)&gt; </p>\n</blockquote>\n<p>This the piece of interest - since I have 4 machines - three can be very happy running training data with as small a size as possible.  My interest is for the fourth machine which is very eager to run the full data set but it occasionally chokes for no apparent reason when 150GB - 200GB is the system monitor report of memory use (64GB hardware + the rest on an SSD swap file).   Was hoping you might have some general root cause / corrective actions for large swap files on Linux.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1809635,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-06-02T22:02:49.907000",
          "content": "<p>Unfortunately, i don't know what is happening. Sounds frustrating.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1809639,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2022-06-02T22:11:20.127000",
          "content": "<p>Thanks - Yep - I probably would be happier if it did not work all the time - success 80% of the time now leads me to keep trying.  If I find something that gets me to 100% success will put the fix here.  </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1882170,
      "author_name": "kgxiao",
      "author_url": "",
      "post_date": "2022-08-03T05:37:29.097000",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>  Wouldn't you consider finding your teammates to fight for the gold medal together? I see your ranking dropped to eighth. Looking forward to your gold medal.  👊   I see <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a>  is looking for teammates and looking forward to working with you. 😄</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1831717,
      "author_name": "Izzet Turkalp Akbasli",
      "author_url": "",
      "post_date": "2022-06-24T11:27:50.270000",
      "content": "<p>I often use classification algorithms. CatBoost and LightGBM doesn't support working with GPU for binary classification. Also, CatBoost developer (I think) shared: LOOOOOL 😂</p>\n<ul>\n<li>Hello!<br>\nCurrently we don't support RSM on GPU for non-pairwise losses, because we don't see need for it - rsm won't provide significant speedup, but it's implementation will require lot's of work, so for now - just don't set RSM while training on GPU.<br>\nThanks!</li>\n</ul>\n<p>CPU overall time about [01:43&lt;00:00,  4.81it/s] with 500 run iteration. But GPU added code spent [32:03&lt;00:00,  3.85s/it] 30 fold time. Really an enormous difference and bad performance. Performance metrics resulted same.</p>\n<p>With GPU:<br>\nTrain Accuracy score: 0.97<br>\nTrain ROC AUC score: 0.96<br>\nTest Accuracy score: 0.82<br>\nTest ROC-AUC score: 0.777</p>\n<p>Without GPU:<br>\nTrain Accuracy score: 0.97<br>\nTrain ROC AUC score: 0.96<br>\nTest Accuracy score: 0.82<br>\nTest ROC-AUC score: 0.777</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8699995%2Ffef474688c489a1a160d0a27b177027f%2FEkran%20Resmi%202022-06-24%2014.10.33.png?generation=1656069076329264&amp;alt=media\" alt=\"xgboost\"></p>\n<p>My device is Colab Pro GPU:</p>\n<p>+-----------------------------------------------------------------------------------------------+<br>\n| NVIDIA-SMI 460.32.03    Driver Version: 460.32.03    CUDA Version: 11.2     | 15109MiB            |<br>\n|-------------------------------+----------------------+----------------------------------------+</p>\n<p>I will trying with regression models.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1809040,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-02T10:35:04.093000",
      "content": "",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1898123,
      "author_name": "bertges",
      "author_url": "",
      "post_date": "2022-08-14T10:24:26.763000",
      "content": "<p>Thanks for the tip! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1841194,
      "author_name": "Saman",
      "author_url": "",
      "post_date": "2022-07-02T22:13:18.493000",
      "content": "<p>Thank you for sharing!<br>\nUseful tips</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1811889,
      "author_name": "Dhamu",
      "author_url": "",
      "post_date": "2022-06-05T08:56:21.780000",
      "content": "<p>Very useful! Thanks for sharing</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1811023,
      "author_name": "IMvision12",
      "author_url": "",
      "post_date": "2022-06-04T07:23:08.700000",
      "content": "<p>Thank you for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1810611,
      "author_name": "Mark Danovich",
      "author_url": "",
      "post_date": "2022-06-03T18:04:40.683000",
      "content": "<p>Thank you for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1809673,
      "author_name": "Making TARS",
      "author_url": "",
      "post_date": "2022-06-03T00:07:34.723000",
      "content": "<p>Thank you for sharing</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1808640": "# CatBoost GPU\nNote that the best public CatBoost notebook does not use GPU. It is running on CPU and takes 8 hours to train 5-Folds. To turn on GPU, first choose `accelerator = GPU` in the right sidebar of the Kaggle jupyter notebook. Next add `task_type = 'GPU'` to the creation of the classifier:\n\n    clf = CatBoostClassifier(iterations=5000, random_state=22, \n                             task_type = 'GPU')\n\nThat's it! Afterward, CatBoost will only take 25 minutes to train 5-folds instead of 8 hours! Wow! \n\nWhen running in a Kaggle notebook, we will need to optimize memory usage so that the notebook doesn't have memory errors. If you start with the most popular public notebook, add this to the end of each K-fold for-loop to avoid GPU memory error:\n\n    import gc\n    del clf, tr_x, val_x, tr_y, val_y, preds, preds_test\n    gc.collect()\n\nAnd if you run CatBoost offline, note that CatBoost will use all your available GPUs and train even faster!\n\n# XGB GPU\nWe can also train fast boosted trees with XGB on GPU. To use GPU with XGB, add the two following parameters to your XGB code:\n\n    xgb_parms = { \n        'tree_method':'gpu_hist',\n        'predictor':'gpu_predictor',\n    }\n    model = xgb.train(xgb_parms)\n\n\nI posted a starter notebook [here][1]. It trains 5-folds in a total of 9 minutes! And it achieves CV 0.792 and LB 0.794! That's fast and accurate!\n\n# LGBM GPU\nAnd lastly, we can install and use LGBM GPU version to accelerate LGBM. More info [here][2]. \n\n**UPDATE 1**: to use LGBM GPU we do not need to reinstall LGBM as the outdated link suggests. We can just add parameters `'device': 'gpu', 'gpu_platform_id': 0, and 'gpu_device_id': 0`. However when i try this in Kaggle notebooks, and watch GPU usage it doesn't seem to use much GPU. So perhaps LGBM doesn't give much speed up. We will need to investigate this. \n\n**UPDATE 2**: Using LGBM with Kaggle's P100 GPU is 1.67x faster than Kaggle's 2-core CPU and 1.33x faster than Kaggle's 4-core CPU. See comments below for details.\n\n[1]: https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\n[2]: https://www.kaggle.com/competitions/santander-customer-transaction-prediction/discussion/89004\n[3]: https://www.kaggle.com/code/youheitomio/gpu-version-amex-lgbm-by-youhei-tomio\n[4]: https://www.kaggle.com/code/youheitomio/gpu-version-amex-lgbm-by-youhei-tomio?scriptVersionId=99773844\n",
    "1808995": "Didn't try lgbm with gpu recently by as far as I remember GPU speed was just slightly better in previous lgbm versions. \n\n--\n\nCatboost and XGB -> GPU for sure (if have some resources)",
    "1840983": "**UPDATE 1**: Youhei Tomio, Takuya Fushihara, Yuzuki Kitajima posted a LGBM GPU Kaggle notebook [here][3] where they reinstall a GPU version of LGBM. Check out notebook [version 5][2] where each fold takes 5 minutes. (In the newer versions they are adding DART which slows down both CPU and GPU LGBM). \n\n**UPDATE 2**: I ran some experiments to determine LGBM runtimes. Here are the results. Using LGBM with GPU does not require re-installing LGBM GPU. If we use the LGBM that is already installed in Kaggle's notebook it is just as fast as reinstalling LGBM-GPU. Using LGBM with Kaggle's P100 GPU is 1.67x faster than Kaggle's 2-core CPU and 1.33x faster than Kaggle's 4-core CPU. See comments below for details.\n\n[2]: https://www.kaggle.com/code/youheitomio/gpu-version-amex-lgbm-by-youhei-tomio?scriptVersionId=99773844\n[3]: https://www.kaggle.com/code/youheitomio/gpu-version-amex-lgbm-by-youhei-tomio",
    "1882195": "I was looking this ! Thanks. I've saved my time",
    "1842528": "Thats is what I was looking for. @cdeotte Thank you  very much for sharing this excellent information with us. It has been pretty useful for me indeed. ",
    "1841316": "Thanks for the tip! Can you please suggest if I can use gpu with stacking classifier or not? @cdeotte ",
    "1816456": "Wow, thanks a lot. BTW, How does lgbm call the GPU in kaggle notebook?",
    "1909310": "Thanks for sharing! I also tried to run different models on GPUs, and the one I'm using is RTX 3090 which has 25.4 GB VRAM.  It was doing OK when running XGB model as it speeded up the whole parameter tuning process. But when I switched to Catboost, it just squeezed all the VRAM and the notebook restarted... I tried double the VRAM but the same thing happened. Can you tell us the specs of your GPUs? Thanks!",
    "1814868": "@cdeotte thanks for sharing this. Will definitely explore these to speed up training workloads.",
    "1813588": "This sharing helps me explore more functions about ML, Thanks a lot!",
    "1812597": "thx for sharing",
    "1811526": "Really liked the garbage collection trick! We do not see this quite often in popular ML notebooks! ",
    "1809333": "From my former experiments, there will be a performance drop if train CatBoost on GPU",
    "1851123": "yes, I found that too.\ngpu version of catboost and xgboost can speed up a lot! specially catboost can run very fast!\nwhile lightgbm only speed up a little by gpu.\neven 32 core cpus is not slow than gpu.",
    "1809498": "Running catboost on 4 dual GPU systems using a huge swap file so that one of my machines can process the full data set.    With my current set of features the CPU ram hits around 150GB per System Monitor - runs at a decent speed (for sure would be faster if I had 256GB of hardware ram :) but does occasionally just hang during one of the folds.\n\nRealized with your tip of memory that I am not deleting any data frames - will add that ASAP.   Is there any other tricks you might have when using large swap files (on Ubuntu) ?",
    "1882170": "@cdeotte  Wouldn't you consider finding your teammates to fight for the gold medal together? I see your ranking dropped to eighth. Looking forward to your gold medal.  👊   I see @raddar  is looking for teammates and looking forward to working with you. 😄",
    "1831717": "I often use classification algorithms. CatBoost and LightGBM doesn't support working with GPU for binary classification. Also, CatBoost developer (I think) shared: LOOOOOL 😂\n\n- Hello!\nCurrently we don't support RSM on GPU for non-pairwise losses, because we don't see need for it - rsm won't provide significant speedup, but it's implementation will require lot's of work, so for now - just don't set RSM while training on GPU.\nThanks!\n\nCPU overall time about [01:43<00:00,  4.81it/s] with 500 run iteration. But GPU added code spent [32:03<00:00,  3.85s/it] 30 fold time. Really an enormous difference and bad performance. Performance metrics resulted same.\n\nWith GPU:\nTrain Accuracy score: 0.97\nTrain ROC AUC score: 0.96\nTest Accuracy score: 0.82\nTest ROC-AUC score: 0.777\n\nWithout GPU:\nTrain Accuracy score: 0.97\nTrain ROC AUC score: 0.96\nTest Accuracy score: 0.82\nTest ROC-AUC score: 0.777\n\n![xgboost](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8699995%2Ffef474688c489a1a160d0a27b177027f%2FEkran%20Resmi%202022-06-24%2014.10.33.png?generation=1656069076329264&alt=media)\n\nMy device is Colab Pro GPU:\n\n+-----------------------------------------------------------------------------------------------+\n| NVIDIA-SMI 460.32.03    Driver Version: 460.32.03    CUDA Version: 11.2     | 15109MiB            |\n|-------------------------------+----------------------+----------------------------------------+\n\nI will trying with regression models.",
    "1809040": "",
    "1898123": "Thanks for the tip! ",
    "1841194": "Thank you for sharing!\nUseful tips",
    "1811889": "Very useful! Thanks for sharing",
    "1811023": "Thank you for sharing!\n\n",
    "1810611": "Thank you for sharing!",
    "1809673": "Thank you for sharing"
  }
}