{
  "id": 89909,
  "title": "Good summary of XGBoost vs CatBoost vs LightGBM",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/89909",
  "author_name": "cyberia",
  "post_date": "2019-04-18T15:35:13.764000",
  "votes": 100,
  "comment_count": 20,
  "views": 0,
  "content": "<p>Most of the participants are using one or a combination of these top three libraries. Recently I found a good overview of these libraries in one place.\nFull article <a href=\"https://towardsdatascience.com/catboost-vs-light-gbm-vs-xgboost-5f93620723db\">https://towardsdatascience.com/catboost-vs-light-gbm-vs-xgboost-5f93620723db</a>\n</p>\n\n<p>There is a paper on similar topic for the curious ones <a href=\"https://arxiv.org/pdf/1809.04559.pdf\">https://arxiv.org/pdf/1809.04559.pdf</a>\n<strong>Benchmarking and Optimization of Gradient Boosting Decision Tree Algorithms</strong></p>\n\n<p><img src=\"https://ai2-s2-public.s3.amazonaws.com/figures/2017-08-08/afe6fd762fea88797740fdc7c2d28505f2a316d9/6-Figure2-1.png\" alt=\"\"></p>\n\n<p><strong>Conclusion (<em>short version</em>)</strong> - for a fixed set of hyper-parameters, XGBoost provides the largest reduction in training time when using a GPU relative to a CPU. Moreover, we observe that one is often able to utilize this speedup to converge to a good set of hyper-parameters in a short time. However, there are tasks for which LightGBM, albeit slower, can converge to a solution that generalizes better. Furthermore, for datasets with a large number of features, XGBoost cannot run due to memory limitations, and Catboost converges to a good solution in the shortest time.</p>",
  "messages": [
    {
      "id": 519241,
      "postDate": "2019-04-18T15:35:13.763Z",
      "content": "<p>Most of the participants are using one or a combination of these top three libraries. Recently I found a good overview of these libraries in one place.\nFull article <a href=\"https://towardsdatascience.com/catboost-vs-light-gbm-vs-xgboost-5f93620723db\">https://towardsdatascience.com/catboost-vs-light-gbm-vs-xgboost-5f93620723db</a>\n</p>\n\n<p>There is a paper on similar topic for the curious ones <a href=\"https://arxiv.org/pdf/1809.04559.pdf\">https://arxiv.org/pdf/1809.04559.pdf</a>\n<strong>Benchmarking and Optimization of Gradient Boosting Decision Tree Algorithms</strong></p>\n\n<p><img src=\"https://ai2-s2-public.s3.amazonaws.com/figures/2017-08-08/afe6fd762fea88797740fdc7c2d28505f2a316d9/6-Figure2-1.png\" alt=\"\"></p>\n\n<p><strong>Conclusion (<em>short version</em>)</strong> - for a fixed set of hyper-parameters, XGBoost provides the largest reduction in training time when using a GPU relative to a CPU. Moreover, we observe that one is often able to utilize this speedup to converge to a good set of hyper-parameters in a short time. However, there are tasks for which LightGBM, albeit slower, can converge to a solution that generalizes better. Furthermore, for datasets with a large number of features, XGBoost cannot run due to memory limitations, and Catboost converges to a good solution in the shortest time.</p>",
      "rawMarkdown": "Most of the participants are using one or a combination of these top three libraries. Recently I found a good overview of these libraries in one place.\nFull article https://towardsdatascience.com/catboost-vs-light-gbm-vs-xgboost-5f93620723db\n![](https://cdn-images-1.medium.com/max/2600/1*A0b_ahXOrrijazzJengwYw.png)\n\nThere is a paper on similar topic for the curious ones https://arxiv.org/pdf/1809.04559.pdf\n**Benchmarking and Optimization of Gradient Boosting Decision Tree Algorithms**\n\n![](https://ai2-s2-public.s3.amazonaws.com/figures/2017-08-08/afe6fd762fea88797740fdc7c2d28505f2a316d9/6-Figure2-1.png)\n\n**Conclusion (*short version*)** - for a fixed set of hyper-parameters, XGBoost provides the largest reduction in training time when using a GPU relative to a CPU. Moreover, we observe that one is often able to utilize this speedup to converge to a good set of hyper-parameters in a short time. However, there are tasks for which LightGBM, albeit slower, can converge to a solution that generalizes better. Furthermore, for datasets with a large number of features, XGBoost cannot run due to memory limitations, and Catboost converges to a good solution in the shortest time.",
      "votes": 100
    },
    {
      "id": 751986,
      "postDate": "2020-02-20T17:38:40.840Z",
      "content": "<p>Great</p>",
      "rawMarkdown": "Great",
      "votes": 14
    },
    {
      "id": 914461,
      "postDate": "2020-07-03T23:02:29.880Z",
      "content": "<p>Great!!!</p>",
      "rawMarkdown": "Great!!!",
      "votes": 9
    },
    {
      "id": 519582,
      "postDate": "2019-04-19T08:26:42.623Z",
      "content": "<p>Interesting stuff, I think for me, those results at the end are typical for Kaggle problems, where the best algorithm depends on the problem. The only way of knowing which is the best is to test for yourself.</p>",
      "rawMarkdown": "Interesting stuff, I think for me, those results at the end are typical for Kaggle problems, where the best algorithm depends on the problem. The only way of knowing which is the best is to test for yourself.",
      "votes": 2
    },
    {
      "id": 1164407,
      "postDate": "2021-01-22T10:59:56.773Z",
      "content": "<p>nice summary</p>",
      "rawMarkdown": "nice summary"
    },
    {
      "id": 1010345,
      "postDate": "2020-09-14T17:22:58.937Z",
      "content": "<p>Very Nice Summary! Thanks for your effort.</p>",
      "rawMarkdown": "Very Nice Summary! Thanks for your effort."
    },
    {
      "id": 975810,
      "postDate": "2020-08-18T13:29:21.460Z",
      "content": "<p>Thank you for sharing. Simple, clear and all in one place :)</p>",
      "rawMarkdown": "Thank you for sharing. Simple, clear and all in one place :)"
    },
    {
      "id": 751538,
      "postDate": "2020-02-20T09:04:44.253Z",
      "content": "<p>Good!</p>",
      "rawMarkdown": "Good!"
    },
    {
      "id": 749516,
      "postDate": "2020-02-18T18:41:19.837Z",
      "content": "<p><a href=\"/cyberia\">@cyberia</a> Thank for sharing, I was looking for something like this. </p>",
      "rawMarkdown": "@cyberia Thank for sharing, I was looking for something like this. "
    },
    {
      "id": 520304,
      "postDate": "2019-04-20T16:52:23.787Z",
      "content": "<p>Good knowledge sharing. Different projects could end up using the same methodology, to really get a project well done,  the selected methodology might be just one of many considerations.</p>",
      "rawMarkdown": "Good knowledge sharing. Different projects could end up using the same methodology, to really get a project well done,  the selected methodology might be just one of many considerations."
    },
    {
      "id": 804666,
      "postDate": "2020-04-11T19:31:18.260Z",
      "rawMarkdown": "",
      "votes": 4,
      "isDeleted": true
    },
    {
      "id": 752059,
      "postDate": "2020-02-20T18:15:03.460Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 941986,
      "postDate": "2020-07-23T14:29:10.040Z",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": 1
    },
    {
      "id": 814226,
      "postDate": "2020-04-20T13:31:19.320Z",
      "content": "<p>Thank you, useful comparison.</p>",
      "rawMarkdown": "Thank you, useful comparison.",
      "votes": 1
    },
    {
      "id": 792748,
      "postDate": "2020-03-31T14:24:47.723Z",
      "content": "<p>Thanks, I used it</p>",
      "rawMarkdown": "Thanks, I used it",
      "votes": 1
    },
    {
      "id": 3031682,
      "postDate": "2024-10-30T02:10:58.223Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!"
    },
    {
      "id": 2034008,
      "postDate": "2022-11-17T18:16:31.197Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!"
    },
    {
      "id": 1711160,
      "postDate": "2022-03-03T17:09:21.970Z",
      "content": "<p>Thanks for sharing, nice summary.</p>",
      "rawMarkdown": "Thanks for sharing, nice summary."
    },
    {
      "id": 1240100,
      "postDate": "2021-03-16T08:23:11.623Z",
      "content": "<p>thank you very much</p>",
      "rawMarkdown": "thank you very much"
    },
    {
      "id": 1180490,
      "postDate": "2021-02-01T09:40:17.147Z",
      "content": "<p>cool thanks for sharing </p>",
      "rawMarkdown": "cool thanks for sharing "
    },
    {
      "id": 749499,
      "postDate": "2020-02-18T18:28:54.600Z",
      "content": "<p>Hmm very cool! Thanks for sharing.</p>",
      "rawMarkdown": "Hmm very cool! Thanks for sharing."
    }
  ],
  "comments": [
    {
      "id": 751986,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-20T17:38:40.840000",
      "content": "<p>Great</p>",
      "votes": 14,
      "replies": []
    },
    {
      "id": 914461,
      "author_name": "Felipe Oliveira ",
      "author_url": "",
      "post_date": "2020-07-03T23:02:29.880000",
      "content": "<p>Great!!!</p>",
      "votes": 9,
      "replies": []
    },
    {
      "id": 519582,
      "author_name": "anonemaus",
      "author_url": "",
      "post_date": "2019-04-19T08:26:42.623000",
      "content": "<p>Interesting stuff, I think for me, those results at the end are typical for Kaggle problems, where the best algorithm depends on the problem. The only way of knowing which is the best is to test for yourself.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1164407,
      "author_name": "Akash kumar",
      "author_url": "",
      "post_date": "2021-01-22T10:59:56.773000",
      "content": "<p>nice summary</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1010345,
      "author_name": "konwarsky",
      "author_url": "",
      "post_date": "2020-09-14T17:22:58.937000",
      "content": "<p>Very Nice Summary! Thanks for your effort.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 975810,
      "author_name": "Oleksandr Pustovoi",
      "author_url": "",
      "post_date": "2020-08-18T13:29:21.460000",
      "content": "<p>Thank you for sharing. Simple, clear and all in one place :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 751538,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-20T09:04:44.253000",
      "content": "<p>Good!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 749516,
      "author_name": "Sonia Sharma",
      "author_url": "",
      "post_date": "2020-02-18T18:41:19.837000",
      "content": "<p><a href=\"/cyberia\">@cyberia</a> Thank for sharing, I was looking for something like this. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 520304,
      "author_name": "AC_datascientist",
      "author_url": "",
      "post_date": "2019-04-20T16:52:23.787000",
      "content": "<p>Good knowledge sharing. Different projects could end up using the same methodology, to really get a project well done,  the selected methodology might be just one of many considerations.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 804666,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-11T19:31:18.260000",
      "content": "",
      "votes": 4,
      "replies": []
    },
    {
      "id": 752059,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-20T18:15:03.460000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 941986,
      "author_name": "Stanley Hua",
      "author_url": "",
      "post_date": "2020-07-23T14:29:10.040000",
      "content": "<p>Thanks!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 814226,
      "author_name": "Caesar Balona",
      "author_url": "",
      "post_date": "2020-04-20T13:31:19.320000",
      "content": "<p>Thank you, useful comparison.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 792748,
      "author_name": "Mohma L",
      "author_url": "",
      "post_date": "2020-03-31T14:24:47.723000",
      "content": "<p>Thanks, I used it</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3031682,
      "author_name": "V W.",
      "author_url": "",
      "post_date": "2024-10-30T02:10:58.223000",
      "content": "<p>Thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2034008,
      "author_name": "JMVJ",
      "author_url": "",
      "post_date": "2022-11-17T18:16:31.197000",
      "content": "<p>Thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1711160,
      "author_name": "Erik Dzul",
      "author_url": "",
      "post_date": "2022-03-03T17:09:21.970000",
      "content": "<p>Thanks for sharing, nice summary.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1240100,
      "author_name": "smith.zhang",
      "author_url": "",
      "post_date": "2021-03-16T08:23:11.623000",
      "content": "<p>thank you very much</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1180490,
      "author_name": "Maciej Gronczynski",
      "author_url": "",
      "post_date": "2021-02-01T09:40:17.147000",
      "content": "<p>cool thanks for sharing </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 749499,
      "author_name": "saltysnowball",
      "author_url": "",
      "post_date": "2020-02-18T18:28:54.600000",
      "content": "<p>Hmm very cool! Thanks for sharing.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "519241": "Most of the participants are using one or a combination of these top three libraries. Recently I found a good overview of these libraries in one place.\nFull article https://towardsdatascience.com/catboost-vs-light-gbm-vs-xgboost-5f93620723db\n![](https://cdn-images-1.medium.com/max/2600/1*A0b_ahXOrrijazzJengwYw.png)\n\nThere is a paper on similar topic for the curious ones https://arxiv.org/pdf/1809.04559.pdf\n**Benchmarking and Optimization of Gradient Boosting Decision Tree Algorithms**\n\n![](https://ai2-s2-public.s3.amazonaws.com/figures/2017-08-08/afe6fd762fea88797740fdc7c2d28505f2a316d9/6-Figure2-1.png)\n\n**Conclusion (*short version*)** - for a fixed set of hyper-parameters, XGBoost provides the largest reduction in training time when using a GPU relative to a CPU. Moreover, we observe that one is often able to utilize this speedup to converge to a good set of hyper-parameters in a short time. However, there are tasks for which LightGBM, albeit slower, can converge to a solution that generalizes better. Furthermore, for datasets with a large number of features, XGBoost cannot run due to memory limitations, and Catboost converges to a good solution in the shortest time.",
    "751986": "Great",
    "914461": "Great!!!",
    "519582": "Interesting stuff, I think for me, those results at the end are typical for Kaggle problems, where the best algorithm depends on the problem. The only way of knowing which is the best is to test for yourself.",
    "1164407": "nice summary",
    "1010345": "Very Nice Summary! Thanks for your effort.",
    "975810": "Thank you for sharing. Simple, clear and all in one place :)",
    "751538": "Good!",
    "749516": "@cyberia Thank for sharing, I was looking for something like this. ",
    "520304": "Good knowledge sharing. Different projects could end up using the same methodology, to really get a project well done,  the selected methodology might be just one of many considerations.",
    "804666": "",
    "752059": "",
    "941986": "Thanks!",
    "814226": "Thank you, useful comparison.",
    "792748": "Thanks, I used it",
    "3031682": "Thanks for sharing!",
    "2034008": "Thanks for sharing!",
    "1711160": "Thanks for sharing, nice summary.",
    "1240100": "thank you very much",
    "1180490": "cool thanks for sharing ",
    "749499": "Hmm very cool! Thanks for sharing."
  }
}