{
  "id": 665764,
  "title": "Question about GPU Support for Training Deep Learning Models",
  "url": "/competitions/vesuvius-challenge-surface-detection/discussion/665764",
  "author_name": "hamzah",
  "post_date": "2026-01-03T15:13:24.633000",
  "votes": 5,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Hi Everyone I'm new to these type of Deep Learning Competition on Kaggle, I want to know that Kaggle GPU Quota is enough to train Deep Learning Models or I need external GPU's for Training Deep Learning Models for these Type of Competitions!</p>\n<p>Thanks!</p>",
  "messages": [
    {
      "id": 3385583,
      "postDate": "2026-01-03T15:13:24.633Z",
      "content": "<p>Hi Everyone I'm new to these type of Deep Learning Competition on Kaggle, I want to know that Kaggle GPU Quota is enough to train Deep Learning Models or I need external GPU's for Training Deep Learning Models for these Type of Competitions!</p>\n<p>Thanks!</p>",
      "rawMarkdown": "Hi Everyone I'm new to these type of Deep Learning Competition on Kaggle, I want to know that Kaggle GPU Quota is enough to train Deep Learning Models or I need external GPU's for Training Deep Learning Models for these Type of Competitions!\n\nThanks!",
      "votes": 5
    },
    {
      "id": 3385762,
      "postDate": "2026-01-03T22:33:06.337Z",
      "content": "<p>I trained a model (just a seg model) with a 0.597 lb score on an rtx 3090 over 22 hrs,  which can be rented on many cloud providers for around 50 cents/hr , you don’t need anything super expensive. </p>",
      "rawMarkdown": "I trained a model (just a seg model) with a 0.597 lb score on an rtx 3090 over 22 hrs,  which can be rented on many cloud providers for around 50 cents/hr , you don’t need anything super expensive. ",
      "votes": 4,
      "replies": [
        {
          "id": 3385779,
          "postDate": "2026-01-04T00:06:02.147Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/seanjohnsonsp\" target=\"_blank\">@seanjohnsonsp</a>, thanks for hosting such an awesome competition, it’s been really exciting to participate so far.</p>\n<p>Just to clarify regarding your comment: is the metric you’re reporting computed on the public leaderboard using the same training data split that participants have access to? The score you mentioned seems quite high compared to what most competitors are currently achieving. From what I’ve seen, many teams are using nnUNet approaches too, yet none appear to be reaching results in that range, which made me curious if it's the exact same evaluation setup</p>",
          "rawMarkdown": "Hi @seanjohnsonsp, thanks for hosting such an awesome competition, it’s been really exciting to participate so far.\n\nJust to clarify regarding your comment: is the metric you’re reporting computed on the public leaderboard using the same training data split that participants have access to? The score you mentioned seems quite high compared to what most competitors are currently achieving. From what I’ve seen, many teams are using nnUNet approaches too, yet none appear to be reaching results in that range, which made me curious if it's the exact same evaluation setup",
          "votes": 5,
          "replies": [
            {
              "id": 3385805,
              "postDate": "2026-01-04T03:52:28.317Z",
              "content": "<p>Hey! Happy to hear you are enjoying it! The metric is indeed the exact same metric and the test split is the same.</p>\n<p>The primary difference I probably have is I trained a single model on the <em>entire train set</em> and evaluated afterwards , which is a luxury competitors without external data would not have. </p>\n<p>I used a fork of the skeleton recall loss dkfz repository and switched the 3d skeletonized to a 2d computed over z slices, with a bit more heavy augmentation. I can post about it a bit more when I am back home (at a wedding so a bit predisposed for next day or so) </p>",
              "rawMarkdown": "Hey! Happy to hear you are enjoying it! The metric is indeed the exact same metric and the test split is the same.\n\nThe primary difference I probably have is I trained a single model on the *entire train set* and evaluated afterwards , which is a luxury competitors without external data would not have. \n\nI used a fork of the skeleton recall loss dkfz repository and switched the 3d skeletonized to a 2d computed over z slices, with a bit more heavy augmentation. I can post about it a bit more when I am back home (at a wedding so a bit predisposed for next day or so) ",
              "votes": 6
            },
            {
              "id": 3385829,
              "postDate": "2026-01-04T05:23:50.503Z",
              "content": "<p><a href=\"https://www.kaggle.com/seanjohnsonsp\" target=\"_blank\">@seanjohnsonsp</a> \nThank you for sharing this information. I explored your official codebase and noticed two additional data augmentation techniques:</p>\n<ol>\n<li>BlankRectangleTransform</li>\n<li>InhomogeneousSliceIlluminationTransform</li>\n</ol>\n<p>I attempted to incorporate these into our current baseline model, but unfortunately, I didn't observe any significant difference in performance.\nAdditionally, I tested the following trainers provided in the code:</p>\n<ol>\n<li>nnUNetTrainerSkeletonRecall</li>\n<li>nnUNetTrainerSkeletonRecallNoTube</li>\n</ol>\n<p>However, these didn't seem to produce better results either.</p>\n<p>I am unsure if the augmentation and loss methods you previously mentioned correspond to the specific code I tested above. It would be greatly appreciated if you could clarify this for me.</p>",
              "rawMarkdown": "@seanjohnsonsp \nThank you for sharing this information. I explored your official codebase and noticed two additional data augmentation techniques:\n1. BlankRectangleTransform\n2. InhomogeneousSliceIlluminationTransform\n\nI attempted to incorporate these into our current baseline model, but unfortunately, I didn't observe any significant difference in performance.\nAdditionally, I tested the following trainers provided in the code:\n1. nnUNetTrainerSkeletonRecall\n2. nnUNetTrainerSkeletonRecallNoTube\n\nHowever, these didn't seem to produce better results either.\n\nI am unsure if the augmentation and loss methods you previously mentioned correspond to the specific code I tested above. It would be greatly appreciated if you could clarify this for me.",
              "votes": 5
            },
            {
              "id": 3386250,
              "postDate": "2026-01-04T22:56:21.287Z",
              "content": "<p>The trainer used for this one is nnUNetTrainerMedialSurfaceRecall , the skeleton trainers in that fork still use the 3d skeletonization, i think the rest of that repo version is accurate to what i have locally, i will check tonight. </p>",
              "rawMarkdown": "The trainer used for this one is nnUNetTrainerMedialSurfaceRecall , the skeleton trainers in that fork still use the 3d skeletonization, i think the rest of that repo version is accurate to what i have locally, i will check tonight. ",
              "votes": 2
            },
            {
              "id": 3386308,
              "postDate": "2026-01-05T03:25:29.270Z",
              "content": "<p>Thanks. I'll give it a try.</p>",
              "rawMarkdown": "Thanks. I'll give it a try."
            },
            {
              "id": 3387101,
              "postDate": "2026-01-06T11:39:17.717Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/seanjohnsonsp\" target=\"_blank\">@seanjohnsonsp</a> </p>\n<blockquote>\n  <p>The primary difference I probably have is I trained a single model on the entire train set and evaluated afterwards , which is a luxury competitors without external data would not have.</p>\n</blockquote>\n<p>What is the meaning of <code>entire train set</code>?</p>",
              "rawMarkdown": "Hi @seanjohnsonsp \n> The primary difference I probably have is I trained a single model on the entire train set and evaluated afterwards , which is a luxury competitors without external data would not have.\n\nWhat is the meaning of `entire train set`?"
            },
            {
              "id": 3387106,
              "postDate": "2026-01-06T11:45:53.173Z",
              "content": "<p>He means that he did not use any split for cross validation. He used the full available set for training. However, we still have to score this model through a sandbox submission on Kaggle. The reported score was computed using a local version of the metrics, but I am not entirely sure that the exact scoring logic implemented on Kaggle is being used. We are going to figure this out soon, and possibly share this baseline in a standalone discussion post.</p>",
              "rawMarkdown": "He means that he did not use any split for cross validation. He used the full available set for training. However, we still have to score this model through a sandbox submission on Kaggle. The reported score was computed using a local version of the metrics, but I am not entirely sure that the exact scoring logic implemented on Kaggle is being used. We are going to figure this out soon, and possibly share this baseline in a standalone discussion post.",
              "votes": 3
            }
          ]
        },
        {
          "id": 3388836,
          "postDate": "2026-01-09T17:08:55.823Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 3386067,
      "postDate": "2026-01-04T14:27:22.757Z",
      "content": "<p>Trained model on Kaggle's 2xT4. Tho, you might need to train using 20 or more hours, so essentially that makes you test on the leaderboard once a week. </p>",
      "rawMarkdown": "Trained model on Kaggle's 2xT4. Tho, you might need to train using 20 or more hours, so essentially that makes you test on the leaderboard once a week. ",
      "votes": 1
    },
    {
      "id": 3385741,
      "postDate": "2026-01-03T20:07:34.380Z",
      "content": "<p>You probably could use kaggle gpus. But to realistically iterate fast enough and train models of this size you probably need cloud compute elsewhere. </p>",
      "rawMarkdown": "You probably could use kaggle gpus. But to realistically iterate fast enough and train models of this size you probably need cloud compute elsewhere. ",
      "votes": 1
    },
    {
      "id": 3385623,
      "postDate": "2026-01-03T16:31:17.627Z",
      "content": "<p>External GPUs are absolutely needed for these competitions. Kaggle GPUs may be used only for submissions <a href=\"https://www.kaggle.com/mhamza0810\" target=\"_blank\">@mhamza0810</a> </p>",
      "rawMarkdown": "External GPUs are absolutely needed for these competitions. Kaggle GPUs may be used only for submissions @mhamza0810 ",
      "votes": 1,
      "replies": [
        {
          "id": 3399309,
          "postDate": "2026-01-30T16:02:09.780Z",
          "content": "<p>Hi, but don't we specifically need to submit the notebook as well in this competition? We wait for the time it's computed on Kaggle's end regardless of how good local GPUs are, as far as I understand</p>",
          "rawMarkdown": "Hi, but don't we specifically need to submit the notebook as well in this competition? We wait for the time it's computed on Kaggle's end regardless of how good local GPUs are, as far as I understand"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3385762,
      "author_name": "Sean Johnson_SP",
      "author_url": "",
      "post_date": "2026-01-03T22:33:06.337000",
      "content": "<p>I trained a model (just a seg model) with a 0.597 lb score on an rtx 3090 over 22 hrs,  which can be rented on many cloud providers for around 50 cents/hr , you don’t need anything super expensive. </p>",
      "votes": 4,
      "replies": [
        {
          "id": 3385779,
          "author_name": "Sergio Alvarez",
          "author_url": "",
          "post_date": "2026-01-04T00:06:02.147000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/seanjohnsonsp\" target=\"_blank\">@seanjohnsonsp</a>, thanks for hosting such an awesome competition, it’s been really exciting to participate so far.</p>\n<p>Just to clarify regarding your comment: is the metric you’re reporting computed on the public leaderboard using the same training data split that participants have access to? The score you mentioned seems quite high compared to what most competitors are currently achieving. From what I’ve seen, many teams are using nnUNet approaches too, yet none appear to be reaching results in that range, which made me curious if it's the exact same evaluation setup</p>",
          "votes": 5,
          "replies": [
            {
              "id": 3385805,
              "author_name": "Sean Johnson_SP",
              "author_url": "",
              "post_date": "2026-01-04T03:52:28.317000",
              "content": "<p>Hey! Happy to hear you are enjoying it! The metric is indeed the exact same metric and the test split is the same.</p>\n<p>The primary difference I probably have is I trained a single model on the <em>entire train set</em> and evaluated afterwards , which is a luxury competitors without external data would not have. </p>\n<p>I used a fork of the skeleton recall loss dkfz repository and switched the 3d skeletonized to a 2d computed over z slices, with a bit more heavy augmentation. I can post about it a bit more when I am back home (at a wedding so a bit predisposed for next day or so) </p>",
              "votes": 6,
              "replies": []
            },
            {
              "id": 3385829,
              "author_name": "tingyi",
              "author_url": "",
              "post_date": "2026-01-04T05:23:50.503000",
              "content": "<p><a href=\"https://www.kaggle.com/seanjohnsonsp\" target=\"_blank\">@seanjohnsonsp</a> \nThank you for sharing this information. I explored your official codebase and noticed two additional data augmentation techniques:</p>\n<ol>\n<li>BlankRectangleTransform</li>\n<li>InhomogeneousSliceIlluminationTransform</li>\n</ol>\n<p>I attempted to incorporate these into our current baseline model, but unfortunately, I didn't observe any significant difference in performance.\nAdditionally, I tested the following trainers provided in the code:</p>\n<ol>\n<li>nnUNetTrainerSkeletonRecall</li>\n<li>nnUNetTrainerSkeletonRecallNoTube</li>\n</ol>\n<p>However, these didn't seem to produce better results either.</p>\n<p>I am unsure if the augmentation and loss methods you previously mentioned correspond to the specific code I tested above. It would be greatly appreciated if you could clarify this for me.</p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 3386250,
              "author_name": "Sean Johnson_SP",
              "author_url": "",
              "post_date": "2026-01-04T22:56:21.287000",
              "content": "<p>The trainer used for this one is nnUNetTrainerMedialSurfaceRecall , the skeleton trainers in that fork still use the 3d skeletonization, i think the rest of that repo version is accurate to what i have locally, i will check tonight. </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3386308,
              "author_name": "tingyi",
              "author_url": "",
              "post_date": "2026-01-05T03:25:29.270000",
              "content": "<p>Thanks. I'll give it a try.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3387101,
              "author_name": "ForcewithMe",
              "author_url": "",
              "post_date": "2026-01-06T11:39:17.717000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/seanjohnsonsp\" target=\"_blank\">@seanjohnsonsp</a> </p>\n<blockquote>\n  <p>The primary difference I probably have is I trained a single model on the entire train set and evaluated afterwards , which is a luxury competitors without external data would not have.</p>\n</blockquote>\n<p>What is the meaning of <code>entire train set</code>?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3387106,
              "author_name": "Giorgio Angelotti",
              "author_url": "",
              "post_date": "2026-01-06T11:45:53.173000",
              "content": "<p>He means that he did not use any split for cross validation. He used the full available set for training. However, we still have to score this model through a sandbox submission on Kaggle. The reported score was computed using a local version of the metrics, but I am not entirely sure that the exact scoring logic implemented on Kaggle is being used. We are going to figure this out soon, and possibly share this baseline in a standalone discussion post.</p>",
              "votes": 3,
              "replies": []
            }
          ]
        },
        {
          "id": 3388836,
          "author_name": "",
          "author_url": "",
          "post_date": "2026-01-09T17:08:55.823000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3386067,
      "author_name": "Ahmed Samir",
      "author_url": "",
      "post_date": "2026-01-04T14:27:22.757000",
      "content": "<p>Trained model on Kaggle's 2xT4. Tho, you might need to train using 20 or more hours, so essentially that makes you test on the leaderboard once a week. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3385741,
      "author_name": "Cody_Null",
      "author_url": "",
      "post_date": "2026-01-03T20:07:34.380000",
      "content": "<p>You probably could use kaggle gpus. But to realistically iterate fast enough and train models of this size you probably need cloud compute elsewhere. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3385623,
      "author_name": "Ravi Ramakrishnan",
      "author_url": "",
      "post_date": "2026-01-03T16:31:17.627000",
      "content": "<p>External GPUs are absolutely needed for these competitions. Kaggle GPUs may be used only for submissions <a href=\"https://www.kaggle.com/mhamza0810\" target=\"_blank\">@mhamza0810</a> </p>",
      "votes": 1,
      "replies": [
        {
          "id": 3399309,
          "author_name": "Garage",
          "author_url": "",
          "post_date": "2026-01-30T16:02:09.780000",
          "content": "<p>Hi, but don't we specifically need to submit the notebook as well in this competition? We wait for the time it's computed on Kaggle's end regardless of how good local GPUs are, as far as I understand</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3385583": "Hi Everyone I'm new to these type of Deep Learning Competition on Kaggle, I want to know that Kaggle GPU Quota is enough to train Deep Learning Models or I need external GPU's for Training Deep Learning Models for these Type of Competitions!\n\nThanks!",
    "3385762": "I trained a model (just a seg model) with a 0.597 lb score on an rtx 3090 over 22 hrs,  which can be rented on many cloud providers for around 50 cents/hr , you don’t need anything super expensive. ",
    "3386067": "Trained model on Kaggle's 2xT4. Tho, you might need to train using 20 or more hours, so essentially that makes you test on the leaderboard once a week. ",
    "3385741": "You probably could use kaggle gpus. But to realistically iterate fast enough and train models of this size you probably need cloud compute elsewhere. ",
    "3385623": "External GPUs are absolutely needed for these competitions. Kaggle GPUs may be used only for submissions @mhamza0810 "
  }
}