{
  "id": 498280,
  "title": "Train CatBoost faster using two gpus",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/498280",
  "author_name": "",
  "post_date": "2024-04-27T16:12:14.864903400Z",
  "votes": 21,
  "comment_count": 8,
  "views": 0,
  "content": "<p>To this point, it's clear that GBDT algorithms, namely CatBoost and LightGBM, are the best models so far, with both models supporting GPU training. Unlike LightGBM's VRAM usage (400 to 800 MB), CatBoost uses all available VRAM (15.1/16 GB), for 1 gpu. For CatBoost, using two gpus (T4 x 2) - as per the CatBoost documentation, <a href=\"https://catboost.ai/en/docs/features/training-on-gpu\" target=\"_blank\">here</a> - we can use 14/15 GB from each gpu and thus, have faster training times - meaning wasting less of the valuable gpu quota:</p>\n<blockquote>\n  <p>devices='0:1'</p>\n</blockquote>\n<p>This is the code I use to access both T4 gpus. </p>\n<p>For me, CatBoost takes almost 3 hours to train on P100, which is a lot, comparing to the 30 minutes LightGBM needs. With the T4 x 2 setup, I will need no more than 2 hours. </p>\n<p>Hope you find this information useful if you are using CatBoost!</p>",
  "messages": [
    {
      "id": "2779343",
      "postDate": "04/27/2024 16:12:14",
      "content": "<p>To this point, it's clear that GBDT algorithms, namely CatBoost and LightGBM, are the best models so far, with both models supporting GPU training. Unlike LightGBM's VRAM usage (400 to 800 MB), CatBoost uses all available VRAM (15.1/16 GB), for 1 gpu. For CatBoost, using two gpus (T4 x 2) - as per the CatBoost documentation, <a href=\"https://catboost.ai/en/docs/features/training-on-gpu\" target=\"_blank\">here</a> - we can use 14/15 GB from each gpu and thus, have faster training times - meaning wasting less of the valuable gpu quota:</p>\n<blockquote>\n  <p>devices='0:1'</p>\n</blockquote>\n<p>This is the code I use to access both T4 gpus. </p>\n<p>For me, CatBoost takes almost 3 hours to train on P100, which is a lot, comparing to the 30 minutes LightGBM needs. With the T4 x 2 setup, I will need no more than 2 hours. </p>\n<p>Hope you find this information useful if you are using CatBoost!</p>",
      "rawMarkdown": "To this point, it's clear that GBDT algorithms, namely CatBoost and LightGBM, are the best models so far, with both models supporting GPU training. Unlike LightGBM's VRAM usage (400 to 800 MB), CatBoost uses all available VRAM (15.1/16 GB), for 1 gpu. For CatBoost, using two gpus (T4 x 2) - as per the CatBoost documentation, [here](https://catboost.ai/en/docs/features/training-on-gpu) - we can use 14/15 GB from each gpu and thus, have faster training times - meaning wasting less of the valuable gpu quota:\n\n>devices='0:1'\n\nThis is the code I use to access both T4 gpus. \n\nFor me, CatBoost takes almost 3 hours to train on P100, which is a lot, comparing to the 30 minutes LightGBM needs. With the T4 x 2 setup, I will need no more than 2 hours. \n\nHope you find this information useful if you are using CatBoost!",
      "votes": null
    },
    {
      "id": "2780402",
      "postDate": "04/28/2024 07:26:05",
      "content": "<p>2 hours vs a 3 hour speedup isn't much when lgbm can be trained in 30 mins</p>\n<p>but thanks for the insight, that extra can be handy</p>",
      "rawMarkdown": "2 hours vs a 3 hour speedup isn't much when lgbm can be trained in 30 mins\n\nbut thanks for the insight, that extra can be handy",
      "votes": null
    },
    {
      "id": "2798768",
      "postDate": "05/07/2024 12:08:02",
      "content": "<p>Helpful tip. Thanks a lot!</p>",
      "rawMarkdown": "Helpful tip. Thanks a lot!",
      "votes": null
    },
    {
      "id": "2804719",
      "postDate": "05/10/2024 07:14:53",
      "content": "<p>thanks for posting</p>",
      "rawMarkdown": "thanks for posting",
      "votes": null
    },
    {
      "id": "2804800",
      "postDate": "05/10/2024 08:07:35",
      "content": "<p>It's not mandatory to indicate devices='0:1' as it uses all the gpus anyway, 'task_type': 'GPU' in parameters is suffiecient. But the performance of GPUs is not only about their RAM, more important is the number of processors, the power is in parallel processing. A bunch of less powerfull processors can be better than a less number of more powerful, that is the case of 2 x T4 vs 1 x P100 for CatBoost, but in the case of LGBM or XGB the P100 is a better choice as they use just 1 GPU and it's more powerful than 1 x T4</p>",
      "rawMarkdown": "It's not mandatory to indicate devices='0:1' as it uses all the gpus anyway, 'task_type': 'GPU' in parameters is suffiecient. But the performance of GPUs is not only about their RAM, more important is the number of processors, the power is in parallel processing. A bunch of less powerfull processors can be better than a less number of more powerful, that is the case of 2 x T4 vs 1 x P100 for CatBoost, but in the case of LGBM or XGB the P100 is a better choice as they use just 1 GPU and it's more powerful than 1 x T4",
      "votes": null
    },
    {
      "id": "2805012",
      "postDate": "05/10/2024 10:12:15",
      "content": "<p>For CatBoost, I have to explicitly define </p>\n<blockquote>\n  <p>devices='0:1'</p>\n</blockquote>\n<p>to use both T4 gpus. For the LGB models however - which needs around 700MB (depending on num_leaves, number of features etc.) - indeed, a P100 is faster due to the HBM2 memory (vs the GDDR6 T4s use). CTB though can use more VRAM for training thus more VRAM &gt; better processing cores</p>\n<p>Ironically, T4 gpus are newer than P100 (Tesla vs Pascal architecture, released in 2018 and 2016 respectively).</p>",
      "rawMarkdown": "For CatBoost, I have to explicitly define \n\n> devices='0:1'\n\nto use both T4 gpus. For the LGB models however - which needs around 700MB (depending on num_leaves, number of features etc.) - indeed, a P100 is faster due to the HBM2 memory (vs the GDDR6 T4s use). CTB though can use more VRAM for training thus more VRAM > better processing cores\n\nIronically, T4 gpus are newer than P100 (Tesla vs Pascal architecture, released in 2018 and 2016 respectively).",
      "votes": null
    },
    {
      "id": "2805023",
      "postDate": "05/10/2024 10:19:07",
      "content": "<p>If you don't indicate devices at all your GPU utilization is just on one of the two T4? In my case both are working without any devices indicated.</p>",
      "rawMarkdown": "If you don't indicate devices at all your GPU utilization is just on one of the two T4? In my case both are working without any devices indicated.",
      "votes": null
    },
    {
      "id": "2805031",
      "postDate": "05/10/2024 10:24:22",
      "content": "<p>It has been almost two weeks since I have trained a CatBoost model. Tomorrow when my gpu quota is restored, I will try again, without the \"devices='0:1' argument and see the outcome.</p>\n<blockquote>\n  <p>If you don't indicate devices at all your GPU utilization is just on one of the two T4?</p>\n</blockquote>\n<p>If I recall correctly, yes, only one T4 device was used.</p>",
      "rawMarkdown": "It has been almost two weeks since I have trained a CatBoost model. Tomorrow when my gpu quota is restored, I will try again, without the \"devices='0:1' argument and see the outcome.\n\n> If you don't indicate devices at all your GPU utilization is just on one of the two T4?\n\nIf I recall correctly, yes, only one T4 device was used.",
      "votes": null
    },
    {
      "id": "2806460",
      "postDate": "05/11/2024 05:32:26",
      "content": "<p><a href=\"https://www.kaggle.com/eu1234\" target=\"_blank\">@eu1234</a> You were right: without setting </p>\n<blockquote>\n  <p>devices='0:1'</p>\n</blockquote>\n<p>I can still use the two T4 devices - regardless, I've made the original post since this param exist on the CatBoost documentation and I thought that this is the best way to do it…</p>",
      "rawMarkdown": "eu1234 You were right: without setting \n\n> devices='0:1'\n\nI can still use the two T4 devices - regardless, I've made the original post since this param exist on the CatBoost documentation and I thought that this is the best way to do it...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2780402,
      "author_name": "luciferisback",
      "author_url": "",
      "post_date": "04/28/2024 07:26:05",
      "content": "<p>2 hours vs a 3 hour speedup isn't much when lgbm can be trained in 30 mins</p>\n<p>but thanks for the insight, that extra can be handy</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2798768,
      "author_name": "faithk7u",
      "author_url": "",
      "post_date": "05/07/2024 12:08:02",
      "content": "<p>Helpful tip. Thanks a lot!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2804719,
      "author_name": "financialtimes",
      "author_url": "",
      "post_date": "05/10/2024 07:14:53",
      "content": "<p>thanks for posting</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2804800,
      "author_name": "eu1234",
      "author_url": "",
      "post_date": "05/10/2024 08:07:35",
      "content": "<p>It's not mandatory to indicate devices='0:1' as it uses all the gpus anyway, 'task_type': 'GPU' in parameters is suffiecient. But the performance of GPUs is not only about their RAM, more important is the number of processors, the power is in parallel processing. A bunch of less powerfull processors can be better than a less number of more powerful, that is the case of 2 x T4 vs 1 x P100 for CatBoost, but in the case of LGBM or XGB the P100 is a better choice as they use just 1 GPU and it's more powerful than 1 x T4</p>",
      "votes": null,
      "replies": [
        {
          "id": 2805012,
          "author_name": "andreasbis",
          "author_url": "",
          "post_date": "05/10/2024 10:12:15",
          "content": "<p>For CatBoost, I have to explicitly define </p>\n<blockquote>\n  <p>devices='0:1'</p>\n</blockquote>\n<p>to use both T4 gpus. For the LGB models however - which needs around 700MB (depending on num_leaves, number of features etc.) - indeed, a P100 is faster due to the HBM2 memory (vs the GDDR6 T4s use). CTB though can use more VRAM for training thus more VRAM &gt; better processing cores</p>\n<p>Ironically, T4 gpus are newer than P100 (Tesla vs Pascal architecture, released in 2018 and 2016 respectively).</p>",
          "votes": null,
          "replies": [
            {
              "id": 2805023,
              "author_name": "eu1234",
              "author_url": "",
              "post_date": "05/10/2024 10:19:07",
              "content": "<p>If you don't indicate devices at all your GPU utilization is just on one of the two T4? In my case both are working without any devices indicated.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2805031,
                  "author_name": "andreasbis",
                  "author_url": "",
                  "post_date": "05/10/2024 10:24:22",
                  "content": "<p>It has been almost two weeks since I have trained a CatBoost model. Tomorrow when my gpu quota is restored, I will try again, without the \"devices='0:1' argument and see the outcome.</p>\n<blockquote>\n  <p>If you don't indicate devices at all your GPU utilization is just on one of the two T4?</p>\n</blockquote>\n<p>If I recall correctly, yes, only one T4 device was used.</p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 2806460,
                  "author_name": "andreasbis",
                  "author_url": "",
                  "post_date": "05/11/2024 05:32:26",
                  "content": "<p><a href=\"https://www.kaggle.com/eu1234\" target=\"_blank\">@eu1234</a> You were right: without setting </p>\n<blockquote>\n  <p>devices='0:1'</p>\n</blockquote>\n<p>I can still use the two T4 devices - regardless, I've made the original post since this param exist on the CatBoost documentation and I thought that this is the best way to do it…</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2779343": "To this point, it's clear that GBDT algorithms, namely CatBoost and LightGBM, are the best models so far, with both models supporting GPU training. Unlike LightGBM's VRAM usage (400 to 800 MB), CatBoost uses all available VRAM (15.1/16 GB), for 1 gpu. For CatBoost, using two gpus (T4 x 2) - as per the CatBoost documentation, [here](https://catboost.ai/en/docs/features/training-on-gpu) - we can use 14/15 GB from each gpu and thus, have faster training times - meaning wasting less of the valuable gpu quota:\n\n>devices='0:1'\n\nThis is the code I use to access both T4 gpus. \n\nFor me, CatBoost takes almost 3 hours to train on P100, which is a lot, comparing to the 30 minutes LightGBM needs. With the T4 x 2 setup, I will need no more than 2 hours. \n\nHope you find this information useful if you are using CatBoost!",
    "2780402": "2 hours vs a 3 hour speedup isn't much when lgbm can be trained in 30 mins\n\nbut thanks for the insight, that extra can be handy",
    "2798768": "Helpful tip. Thanks a lot!",
    "2804719": "thanks for posting",
    "2804800": "It's not mandatory to indicate devices='0:1' as it uses all the gpus anyway, 'task_type': 'GPU' in parameters is suffiecient. But the performance of GPUs is not only about their RAM, more important is the number of processors, the power is in parallel processing. A bunch of less powerfull processors can be better than a less number of more powerful, that is the case of 2 x T4 vs 1 x P100 for CatBoost, but in the case of LGBM or XGB the P100 is a better choice as they use just 1 GPU and it's more powerful than 1 x T4",
    "2805012": "For CatBoost, I have to explicitly define \n\n> devices='0:1'\n\nto use both T4 gpus. For the LGB models however - which needs around 700MB (depending on num_leaves, number of features etc.) - indeed, a P100 is faster due to the HBM2 memory (vs the GDDR6 T4s use). CTB though can use more VRAM for training thus more VRAM > better processing cores\n\nIronically, T4 gpus are newer than P100 (Tesla vs Pascal architecture, released in 2018 and 2016 respectively).",
    "2805023": "If you don't indicate devices at all your GPU utilization is just on one of the two T4? In my case both are working without any devices indicated.",
    "2805031": "It has been almost two weeks since I have trained a CatBoost model. Tomorrow when my gpu quota is restored, I will try again, without the \"devices='0:1' argument and see the outcome.\n\n> If you don't indicate devices at all your GPU utilization is just on one of the two T4?\n\nIf I recall correctly, yes, only one T4 device was used.",
    "2806460": "eu1234 You were right: without setting \n\n> devices='0:1'\n\nI can still use the two T4 devices - regardless, I've made the original post since this param exist on the CatBoost documentation and I thought that this is the best way to do it..."
  },
  "source": "meta"
}