{
  "id": 333507,
  "title": "Tips needed for getting good performance in Collab with GPU for cudf,cupy,XGB and LGMB ",
  "url": "/competitions/amex-default-prediction/discussion/333507",
  "author_name": "",
  "post_date": "2022-06-27T00:31:59.653293500Z",
  "votes": 9,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi just started with the competitions and was looking at some advice on improving performance of training and experiments on collab . So far I have observed my training and inference suprisingly is faster on Kaggle compared to collab . I guess it might be due to custom steps we need to install these libraries on the collab container .</p>\n<p>So far this is what I have done :-</p>\n<ul>\n<li><strong>XGB with cudf cupy</strong></li>\n</ul>\n<p>Installing Cudf and rapids utils like below but still the training is slow compared to in Kaggle . Cuda version in kaggle though is <strong>21.10.01</strong> compared to on collab  <strong>21.12.02</strong>.On top of this the below installation is very very slow need to see if Nvidia folks here have a better alternative.</p>\n<pre><code>!git clone https://github.com/rapidsai/rapidsai-csp-utils.git\n!python rapidsai-csp-utils/colab/env-check.py\n\n!bash rapidsai-csp-utils/colab/update_gcc.sh\nimport os\nos._exit(00)\n\nimport condacolab\ncondacolab.install()\n\nimport condacolab\ncondacolab.check()\n\n!python rapidsai-csp-utils/colab/install_rapids.py stable\nimport os\nos.environ['NUMBAPRO_NVVM'] = '/usr/local/cuda/nvvm/lib64/libnvvm.so'\nos.environ['NUMBAPRO_LIBDEVICE'] = '/usr/local/cuda/nvvm/libdevice/'\nos.environ['CONDA_PREFIX'] = '/usr/local'\n</code></pre>\n<ul>\n<li><strong>LGBM</strong></li>\n</ul>\n<p>Did below steps to recompile lgbm for GPU but seems as I think observed in kaggle somehow LGBM doesn't use GPU that efficiently like XGboost.</p>\n<pre><code>! git clone --recursive https://github.com/Microsoft/LightGBM\n! cd LightGBM &amp;&amp; rm -rf build &amp;&amp; mkdir build &amp;&amp; cd build &amp;&amp; cmake -DUSE_GPU=1 ../../LightGBM &amp;&amp; make -j4 &amp;&amp; cd ../python-package &amp;&amp; python3 setup.py install --precompile --gpu;\n</code></pre>\n<ul>\n<li><strong>CATBOOST</strong></li>\n</ul>\n<p>Not related to this but I have seen CATBOOST working pretty well and fast on GPUs on collab and making most of the available GPUS .</p>\n<p>If anyone has any tips please do let me know, I am definitely missing something 😊 Thanks .</p>",
  "messages": [
    {
      "id": "1834454",
      "postDate": "06/27/2022 00:31:59",
      "content": "<p>Hi just started with the competitions and was looking at some advice on improving performance of training and experiments on collab . So far I have observed my training and inference suprisingly is faster on Kaggle compared to collab . I guess it might be due to custom steps we need to install these libraries on the collab container .</p>\n<p>So far this is what I have done :-</p>\n<ul>\n<li><strong>XGB with cudf cupy</strong></li>\n</ul>\n<p>Installing Cudf and rapids utils like below but still the training is slow compared to in Kaggle . Cuda version in kaggle though is <strong>21.10.01</strong> compared to on collab  <strong>21.12.02</strong>.On top of this the below installation is very very slow need to see if Nvidia folks here have a better alternative.</p>\n<pre><code>!git clone https://github.com/rapidsai/rapidsai-csp-utils.git\n!python rapidsai-csp-utils/colab/env-check.py\n\n!bash rapidsai-csp-utils/colab/update_gcc.sh\nimport os\nos._exit(00)\n\nimport condacolab\ncondacolab.install()\n\nimport condacolab\ncondacolab.check()\n\n!python rapidsai-csp-utils/colab/install_rapids.py stable\nimport os\nos.environ['NUMBAPRO_NVVM'] = '/usr/local/cuda/nvvm/lib64/libnvvm.so'\nos.environ['NUMBAPRO_LIBDEVICE'] = '/usr/local/cuda/nvvm/libdevice/'\nos.environ['CONDA_PREFIX'] = '/usr/local'\n</code></pre>\n<ul>\n<li><strong>LGBM</strong></li>\n</ul>\n<p>Did below steps to recompile lgbm for GPU but seems as I think observed in kaggle somehow LGBM doesn't use GPU that efficiently like XGboost.</p>\n<pre><code>! git clone --recursive https://github.com/Microsoft/LightGBM\n! cd LightGBM &amp;&amp; rm -rf build &amp;&amp; mkdir build &amp;&amp; cd build &amp;&amp; cmake -DUSE_GPU=1 ../../LightGBM &amp;&amp; make -j4 &amp;&amp; cd ../python-package &amp;&amp; python3 setup.py install --precompile --gpu;\n</code></pre>\n<ul>\n<li><strong>CATBOOST</strong></li>\n</ul>\n<p>Not related to this but I have seen CATBOOST working pretty well and fast on GPUs on collab and making most of the available GPUS .</p>\n<p>If anyone has any tips please do let me know, I am definitely missing something 😊 Thanks .</p>",
      "rawMarkdown": "Hi just started with the competitions and was looking at some advice on improving performance of training and experiments on collab . So far I have observed my training and inference suprisingly is faster on Kaggle compared to collab . I guess it might be due to custom steps we need to install these libraries on the collab container .\n\nSo far this is what I have done :-\n\n- **XGB with cudf cupy**\n\nInstalling Cudf and rapids utils like below but still the training is slow compared to in Kaggle . Cuda version in kaggle though is **21.10.01** compared to on collab  **21.12.02**.On top of this the below installation is very very slow need to see if Nvidia folks here have a better alternative.\n\n```python\n!git clone https://github.com/rapidsai/rapidsai-csp-utils.git\n!python rapidsai-csp-utils/colab/env-check.py\n\n!bash rapidsai-csp-utils/colab/update_gcc.sh\nimport os\nos._exit(00)\n\nimport condacolab\ncondacolab.install()\n\nimport condacolab\ncondacolab.check()\n\n!python rapidsai-csp-utils/colab/install_rapids.py stable\nimport os\nos.environ['NUMBAPRO_NVVM'] = '/usr/local/cuda/nvvm/lib64/libnvvm.so'\nos.environ['NUMBAPRO_LIBDEVICE'] = '/usr/local/cuda/nvvm/libdevice/'\nos.environ['CONDA_PREFIX'] = '/usr/local'\n```\n- **LGBM**\n\nDid below steps to recompile lgbm for GPU but seems as I think observed in kaggle somehow LGBM doesn't use GPU that efficiently like XGboost.\n\n```python\n! git clone --recursive https://github.com/Microsoft/LightGBM\n! cd LightGBM && rm -rf build && mkdir build && cd build && cmake -DUSE_GPU=1 ../../LightGBM && make -j4 && cd ../python-package && python3 setup.py install --precompile --gpu;\n```\n- **CATBOOST**\n\nNot related to this but I have seen CATBOOST working pretty well and fast on GPUs on collab and making most of the available GPUS .\n\nIf anyone has any tips please do let me know, I am definitely missing something 😊 Thanks .",
      "votes": null
    },
    {
      "id": "1835094",
      "postDate": "06/27/2022 13:11:34",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/gauravbrills\" target=\"_blank\">@gauravbrills</a>, thanks for sharing, it really takes some time to find out how to enable GPU tools on Colab</p>\n<p>Couple of my personal hints there:<br>\n1) For LGBM with GPU you can use simple pip install right AFTER rapids.ai installation script <code>(!python rapidsai-csp-utils/colab/install_rapids.py stable)</code>:</p>\n<pre><code>import os\n%pip install -U lightgbm --install-option=--gpu\nos._exit(00)\n</code></pre>\n<p>That will install LGBM <strong>with GPU support</strong></p>\n<p>2) To speed up RAPIDS.ai installation script you can pass only component to install, like cudf:<br>\n<code>!python rapidsai-csp-utils/colab/install_rapids.py stable cudf</code></p>\n<p>As by default RAPIDS.ai install all the components</p>",
      "rawMarkdown": "Hi, @gauravbrills, thanks for sharing, it really takes some time to find out how to enable GPU tools on Colab\n\nCouple of my personal hints there:\n1) For LGBM with GPU you can use simple pip install right AFTER rapids.ai installation script `(!python rapidsai-csp-utils/colab/install_rapids.py stable)`:\n```\nimport os\n%pip install -U lightgbm --install-option=--gpu\nos._exit(00)\n```\nThat will install LGBM **with GPU support**\n\n2) To speed up RAPIDS.ai installation script you can pass only component to install, like cudf:\n` !python rapidsai-csp-utils/colab/install_rapids.py stable cudf`\n\nAs by default RAPIDS.ai install all the components",
      "votes": null
    },
    {
      "id": "1835136",
      "postDate": "06/27/2022 13:59:27",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/romaupgini\" target=\"_blank\">@romaupgini</a> these make sense will try and revert . But from prior experience with LGBM have observed it doesnt use GPU that efficiently and still need to check how it rolls with cudf </p>",
      "rawMarkdown": "Thanks @romaupgini these make sense will try and revert . But from prior experience with LGBM have observed it doesnt use GPU that efficiently and still need to check how it rolls with cudf",
      "votes": null
    },
    {
      "id": "1836016",
      "postDate": "06/28/2022 09:31:05",
      "content": "<p>As far a I know LGBM doesnt support cudf directly (in contrast to XGB), and after several trials I found out that the most memory efficient way to pass cudf to LGBM Dataset would be following:<br>\n<code>dtrain = lgb.Dataset(train.loc[train_idx, FEATURES].fillna(NAN_VALUE).as_gpu_matrix(),\n                            train.loc[train_idx, 'target'].values.get())</code><br>\nAs as_gpu_matrix() doesn't create copy of dataset. BTW you'll get warning on that and proposal to use to_cupy() what creates A COPY. Seems like a lot of inefficiency under the hood in cudf/cupy.</p>",
      "rawMarkdown": "As far a I know LGBM doesnt support cudf directly (in contrast to XGB), and after several trials I found out that the most memory efficient way to pass cudf to LGBM Dataset would be following:\n`dtrain = lgb.Dataset(train.loc[train_idx, FEATURES].fillna(NAN_VALUE).as_gpu_matrix(),\n                            train.loc[train_idx, 'target'].values.get())`\nAs as_gpu_matrix() doesn't create copy of dataset. BTW you'll get warning on that and proposal to use to_cupy() what creates A COPY. Seems like a lot of inefficiency under the hood in cudf/cupy.",
      "votes": null
    },
    {
      "id": "1836323",
      "postDate": "06/28/2022 14:31:34",
      "content": "<p>yup I am doing the same right now but not much performance gains so sticking to CPU . </p>",
      "rawMarkdown": "yup I am doing the same right now but not much performance gains so sticking to CPU .",
      "votes": null
    },
    {
      "id": "1836593",
      "postDate": "06/28/2022 20:27:33",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/gauravbrills\" target=\"_blank\">@gauravbrills</a> This approach somehow stopped working on Colab for me. <br>\nStarting with missing module 'pynvml' and ending in install_rapids.py completely failing due to Conda failed with initial frozen solve.</p>\n<p>Did you encounter any problems or can you share some insights on your current Colab setup?</p>",
      "rawMarkdown": "Hi, @gauravbrills This approach somehow stopped working on Colab for me. \nStarting with missing module 'pynvml' and ending in install_rapids.py completely failing due to Conda failed with initial frozen solve.\n\nDid you encounter any problems or can you share some insights on your current Colab setup?",
      "votes": null
    },
    {
      "id": "1836743",
      "postDate": "06/29/2022 03:02:22",
      "content": "<p>Which setup specifically . I was able to run XGB fine with the above code , You might need to run the above snippet in separate cells as they do restart in between the COllab container .</p>",
      "rawMarkdown": "Which setup specifically . I was able to run XGB fine with the above code , You might need to run the above snippet in separate cells as they do restart in between the COllab container .",
      "votes": null
    },
    {
      "id": "1836915",
      "postDate": "06/29/2022 08:02:35",
      "content": "<p><a href=\"https://www.kaggle.com/matthiasanderer\" target=\"_blank\">@matthiasanderer</a>  hope this helps <a href=\"https://colab.research.google.com/drive/1kS61kCHnbZOw5dO82Um0W2FO7UWOvfpJ?usp=sharing\" target=\"_blank\">https://colab.research.google.com/drive/1kS61kCHnbZOw5dO82Um0W2FO7UWOvfpJ?usp=sharing</a></p>",
      "rawMarkdown": "matthiasanderer  hope this helps https://colab.research.google.com/drive/1kS61kCHnbZOw5dO82Um0W2FO7UWOvfpJ?usp=sharing",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1835094,
      "author_name": "romaupgini",
      "author_url": "",
      "post_date": "06/27/2022 13:11:34",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/gauravbrills\" target=\"_blank\">@gauravbrills</a>, thanks for sharing, it really takes some time to find out how to enable GPU tools on Colab</p>\n<p>Couple of my personal hints there:<br>\n1) For LGBM with GPU you can use simple pip install right AFTER rapids.ai installation script <code>(!python rapidsai-csp-utils/colab/install_rapids.py stable)</code>:</p>\n<pre><code>import os\n%pip install -U lightgbm --install-option=--gpu\nos._exit(00)\n</code></pre>\n<p>That will install LGBM <strong>with GPU support</strong></p>\n<p>2) To speed up RAPIDS.ai installation script you can pass only component to install, like cudf:<br>\n<code>!python rapidsai-csp-utils/colab/install_rapids.py stable cudf</code></p>\n<p>As by default RAPIDS.ai install all the components</p>",
      "votes": null,
      "replies": [
        {
          "id": 1835136,
          "author_name": "gauravbrills",
          "author_url": "",
          "post_date": "06/27/2022 13:59:27",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/romaupgini\" target=\"_blank\">@romaupgini</a> these make sense will try and revert . But from prior experience with LGBM have observed it doesnt use GPU that efficiently and still need to check how it rolls with cudf </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1836016,
          "author_name": "romaupgini",
          "author_url": "",
          "post_date": "06/28/2022 09:31:05",
          "content": "<p>As far a I know LGBM doesnt support cudf directly (in contrast to XGB), and after several trials I found out that the most memory efficient way to pass cudf to LGBM Dataset would be following:<br>\n<code>dtrain = lgb.Dataset(train.loc[train_idx, FEATURES].fillna(NAN_VALUE).as_gpu_matrix(),\n                            train.loc[train_idx, 'target'].values.get())</code><br>\nAs as_gpu_matrix() doesn't create copy of dataset. BTW you'll get warning on that and proposal to use to_cupy() what creates A COPY. Seems like a lot of inefficiency under the hood in cudf/cupy.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1836323,
          "author_name": "gauravbrills",
          "author_url": "",
          "post_date": "06/28/2022 14:31:34",
          "content": "<p>yup I am doing the same right now but not much performance gains so sticking to CPU . </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1836593,
      "author_name": "matthiasanderer",
      "author_url": "",
      "post_date": "06/28/2022 20:27:33",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/gauravbrills\" target=\"_blank\">@gauravbrills</a> This approach somehow stopped working on Colab for me. <br>\nStarting with missing module 'pynvml' and ending in install_rapids.py completely failing due to Conda failed with initial frozen solve.</p>\n<p>Did you encounter any problems or can you share some insights on your current Colab setup?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1836743,
          "author_name": "gauravbrills",
          "author_url": "",
          "post_date": "06/29/2022 03:02:22",
          "content": "<p>Which setup specifically . I was able to run XGB fine with the above code , You might need to run the above snippet in separate cells as they do restart in between the COllab container .</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1836915,
          "author_name": "romaupgini",
          "author_url": "",
          "post_date": "06/29/2022 08:02:35",
          "content": "<p><a href=\"https://www.kaggle.com/matthiasanderer\" target=\"_blank\">@matthiasanderer</a>  hope this helps <a href=\"https://colab.research.google.com/drive/1kS61kCHnbZOw5dO82Um0W2FO7UWOvfpJ?usp=sharing\" target=\"_blank\">https://colab.research.google.com/drive/1kS61kCHnbZOw5dO82Um0W2FO7UWOvfpJ?usp=sharing</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1834454": "Hi just started with the competitions and was looking at some advice on improving performance of training and experiments on collab . So far I have observed my training and inference suprisingly is faster on Kaggle compared to collab . I guess it might be due to custom steps we need to install these libraries on the collab container .\n\nSo far this is what I have done :-\n\n- **XGB with cudf cupy**\n\nInstalling Cudf and rapids utils like below but still the training is slow compared to in Kaggle . Cuda version in kaggle though is **21.10.01** compared to on collab  **21.12.02**.On top of this the below installation is very very slow need to see if Nvidia folks here have a better alternative.\n\n```python\n!git clone https://github.com/rapidsai/rapidsai-csp-utils.git\n!python rapidsai-csp-utils/colab/env-check.py\n\n!bash rapidsai-csp-utils/colab/update_gcc.sh\nimport os\nos._exit(00)\n\nimport condacolab\ncondacolab.install()\n\nimport condacolab\ncondacolab.check()\n\n!python rapidsai-csp-utils/colab/install_rapids.py stable\nimport os\nos.environ['NUMBAPRO_NVVM'] = '/usr/local/cuda/nvvm/lib64/libnvvm.so'\nos.environ['NUMBAPRO_LIBDEVICE'] = '/usr/local/cuda/nvvm/libdevice/'\nos.environ['CONDA_PREFIX'] = '/usr/local'\n```\n- **LGBM**\n\nDid below steps to recompile lgbm for GPU but seems as I think observed in kaggle somehow LGBM doesn't use GPU that efficiently like XGboost.\n\n```python\n! git clone --recursive https://github.com/Microsoft/LightGBM\n! cd LightGBM && rm -rf build && mkdir build && cd build && cmake -DUSE_GPU=1 ../../LightGBM && make -j4 && cd ../python-package && python3 setup.py install --precompile --gpu;\n```\n- **CATBOOST**\n\nNot related to this but I have seen CATBOOST working pretty well and fast on GPUs on collab and making most of the available GPUS .\n\nIf anyone has any tips please do let me know, I am definitely missing something 😊 Thanks .",
    "1835094": "Hi, @gauravbrills, thanks for sharing, it really takes some time to find out how to enable GPU tools on Colab\n\nCouple of my personal hints there:\n1) For LGBM with GPU you can use simple pip install right AFTER rapids.ai installation script `(!python rapidsai-csp-utils/colab/install_rapids.py stable)`:\n```\nimport os\n%pip install -U lightgbm --install-option=--gpu\nos._exit(00)\n```\nThat will install LGBM **with GPU support**\n\n2) To speed up RAPIDS.ai installation script you can pass only component to install, like cudf:\n` !python rapidsai-csp-utils/colab/install_rapids.py stable cudf`\n\nAs by default RAPIDS.ai install all the components",
    "1835136": "Thanks @romaupgini these make sense will try and revert . But from prior experience with LGBM have observed it doesnt use GPU that efficiently and still need to check how it rolls with cudf",
    "1836016": "As far a I know LGBM doesnt support cudf directly (in contrast to XGB), and after several trials I found out that the most memory efficient way to pass cudf to LGBM Dataset would be following:\n`dtrain = lgb.Dataset(train.loc[train_idx, FEATURES].fillna(NAN_VALUE).as_gpu_matrix(),\n                            train.loc[train_idx, 'target'].values.get())`\nAs as_gpu_matrix() doesn't create copy of dataset. BTW you'll get warning on that and proposal to use to_cupy() what creates A COPY. Seems like a lot of inefficiency under the hood in cudf/cupy.",
    "1836323": "yup I am doing the same right now but not much performance gains so sticking to CPU .",
    "1836593": "Hi, @gauravbrills This approach somehow stopped working on Colab for me. \nStarting with missing module 'pynvml' and ending in install_rapids.py completely failing due to Conda failed with initial frozen solve.\n\nDid you encounter any problems or can you share some insights on your current Colab setup?",
    "1836743": "Which setup specifically . I was able to run XGB fine with the above code , You might need to run the above snippet in separate cells as they do restart in between the COllab container .",
    "1836915": "matthiasanderer  hope this helps https://colab.research.google.com/drive/1kS61kCHnbZOw5dO82Um0W2FO7UWOvfpJ?usp=sharing"
  },
  "source": "meta"
}