{
  "id": 205045,
  "title": "Has anyone tried GPU training with LightGBM for this competition?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/205045",
  "author_name": "",
  "post_date": "2020-12-18T08:25:53.563542300Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Looks LGBM is the most popular model in this competition, and training/inference time and memory consumption is critical here. I<code>m wondering what</code>s your training/inference time improvement if you are using LGBM with GPU, as I heard GPU acceleration for LGBM is not significant so far, is that true?</p>",
  "messages": [
    {
      "id": "1117586",
      "postDate": "12/18/2020 08:25:53",
      "content": "<p>Looks LGBM is the most popular model in this competition, and training/inference time and memory consumption is critical here. I<code>m wondering what</code>s your training/inference time improvement if you are using LGBM with GPU, as I heard GPU acceleration for LGBM is not significant so far, is that true?</p>",
      "rawMarkdown": "Looks LGBM is the most popular model in this competition, and training/inference time and memory consumption is critical here. I`m wondering what`s your training/inference time improvement if you are using LGBM with GPU, as I heard GPU acceleration for LGBM is not significant so far, is that true?",
      "votes": null
    },
    {
      "id": "1117820",
      "postDate": "12/18/2020 13:35:02",
      "content": "<p>GPU is fast when the main workload of training is matrix-vector, tensor-matrix multiplication. I'm not even sure the LightGBM installed in the current Kaggle docker is the GPU-accelerated one (<a href=\"https://lightgbm.readthedocs.io/en/latest/GPU-Performance.html)\" target=\"_blank\">https://lightgbm.readthedocs.io/en/latest/GPU-Performance.html)</a>. Even if it is, there is significant overhead of transferring the data between because still a considerable amount of work on CPU for any tree-based method. The reason is that you have to check if a condition is met for a sample, and this recursive style thing is slow on GPU (unless all your features are categorical).</p>\n<p>TL;DR: stick to CPU for LGB.</p>",
      "rawMarkdown": "GPU is fast when the main workload of training is matrix-vector, tensor-matrix multiplication. I'm not even sure the LightGBM installed in the current Kaggle docker is the GPU-accelerated one (https://lightgbm.readthedocs.io/en/latest/GPU-Performance.html). Even if it is, there is significant overhead of transferring the data between because still a considerable amount of work on CPU for any tree-based method. The reason is that you have to check if a condition is met for a sample, and this recursive style thing is slow on GPU (unless all your features are categorical).\n\nTL;DR: stick to CPU for LGB.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1117820,
      "author_name": "scaomath",
      "author_url": "",
      "post_date": "12/18/2020 13:35:02",
      "content": "<p>GPU is fast when the main workload of training is matrix-vector, tensor-matrix multiplication. I'm not even sure the LightGBM installed in the current Kaggle docker is the GPU-accelerated one (<a href=\"https://lightgbm.readthedocs.io/en/latest/GPU-Performance.html)\" target=\"_blank\">https://lightgbm.readthedocs.io/en/latest/GPU-Performance.html)</a>. Even if it is, there is significant overhead of transferring the data between because still a considerable amount of work on CPU for any tree-based method. The reason is that you have to check if a condition is met for a sample, and this recursive style thing is slow on GPU (unless all your features are categorical).</p>\n<p>TL;DR: stick to CPU for LGB.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1117586": "Looks LGBM is the most popular model in this competition, and training/inference time and memory consumption is critical here. I`m wondering what`s your training/inference time improvement if you are using LGBM with GPU, as I heard GPU acceleration for LGBM is not significant so far, is that true?",
    "1117820": "GPU is fast when the main workload of training is matrix-vector, tensor-matrix multiplication. I'm not even sure the LightGBM installed in the current Kaggle docker is the GPU-accelerated one (https://lightgbm.readthedocs.io/en/latest/GPU-Performance.html). Even if it is, there is significant overhead of transferring the data between because still a considerable amount of work on CPU for any tree-based method. The reason is that you have to check if a condition is met for a sample, and this recursive style thing is slow on GPU (unless all your features are categorical).\n\nTL;DR: stick to CPU for LGB."
  },
  "source": "meta"
}