{
  "id": 642922,
  "title": "Training large models on Kaggle without your own GPU",
  "url": "/competitions/vesuvius-challenge-surface-detection/discussion/642922",
  "author_name": "",
  "post_date": "2025-11-28T13:15:52.255508500Z",
  "votes": 5,
  "comment_count": 12,
  "views": 0,
  "content": "<p>I’ve been experimenting with ways to handle longer training runs within Kaggle’s resource limits. One approach that worked well for me is kernel checkpointing:</p>\n<ul>\n<li>Train until you hit the kernel run limit</li>\n<li>Offload checkpoints to a Kaggle Dataset or Model</li>\n<li>Resume training later by attaching the saved artifact back to the kernel</li>\n</ul>\n<p>This makes it possible to continue training across multiple sessions, and it also doubles as a neat inference setup if you train models elsewhere (e.g. Google Colab) and want to run them on Kaggle.</p>\n<p>Curious to hear how others manage long training jobs or combine Kaggle with external resources—always looking for new tricks to make workflows smoother.</p>",
  "messages": [
    {
      "id": "3351505",
      "postDate": "11/28/2025 13:15:52",
      "content": "<p>I’ve been experimenting with ways to handle longer training runs within Kaggle’s resource limits. One approach that worked well for me is kernel checkpointing:</p>\n<ul>\n<li>Train until you hit the kernel run limit</li>\n<li>Offload checkpoints to a Kaggle Dataset or Model</li>\n<li>Resume training later by attaching the saved artifact back to the kernel</li>\n</ul>\n<p>This makes it possible to continue training across multiple sessions, and it also doubles as a neat inference setup if you train models elsewhere (e.g. Google Colab) and want to run them on Kaggle.</p>\n<p>Curious to hear how others manage long training jobs or combine Kaggle with external resources—always looking for new tricks to make workflows smoother.</p>",
      "rawMarkdown": "I’ve been experimenting with ways to handle longer training runs within Kaggle’s resource limits. One approach that worked well for me is kernel checkpointing:\n\n- Train until you hit the kernel run limit\n- Offload checkpoints to a Kaggle Dataset or Model\n- Resume training later by attaching the saved artifact back to the kernel\n\nThis makes it possible to continue training across multiple sessions, and it also doubles as a neat inference setup if you train models elsewhere (e.g. Google Colab) and want to run them on Kaggle.\n\nCurious to hear how others manage long training jobs or combine Kaggle with external resources—always looking for new tricks to make workflows smoother.",
      "votes": null
    },
    {
      "id": "3352481",
      "postDate": "11/29/2025 08:02:08",
      "content": "<p>many people have their own GPUs or using cloud  gpu/tpu.</p>",
      "rawMarkdown": "many people have their own GPUs or using cloud  gpu/tpu.",
      "votes": null
    },
    {
      "id": "3352543",
      "postDate": "11/29/2025 08:55:01",
      "content": "<p><a href=\"https://www.kaggle.com/jirkaborovec\" target=\"_blank\">@jirkaborovec</a> \nThis suggestion is temporary and might be not recommended. </p>\n<p>The <a href=\"https://www.kaggle.com/competitions/ai-mathematical-olympiad-progress-prize-3\" target=\"_blank\">AIMO</a> competition gives H100 quota. This GPU is only for AIMO competition. But it can be used for your cause.</p>\n<ul>\n<li>Create notebook in AIMO competition.</li>\n<li>Attach your dataset (i.e. npy format of vesuvius challenge or any).</li>\n<li>Attach H100</li>\n<li>Boom.</li>\n</ul>\n<p>Only for training. Though I tried it one or two times. The speed is blazing fast.</p>",
      "rawMarkdown": "jirkaborovec \nThis suggestion is temporary and might be not recommended. \n\nThe [AIMO](https://www.kaggle.com/competitions/ai-mathematical-olympiad-progress-prize-3) competition gives H100 quota. This GPU is only for AIMO competition. But it can be used for your cause.\n\n- Create notebook in AIMO competition.\n- Attach your dataset (i.e. npy format of vesuvius challenge or any).\n- Attach H100\n- Boom.\n\nOnly for training. Though I tried it one or two times. The speed is blazing fast.",
      "votes": null
    },
    {
      "id": "3352676",
      "postDate": "11/29/2025 10:53:51",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F23060793%2F273658e0cc95233aa559727be46b2f2a%2FScreenshot%202025-11-29%20175312.png?generation=1764413624534990&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F23060793%2F273658e0cc95233aa559727be46b2f2a%2FScreenshot%202025-11-29%20175312.png?generation=1764413624534990&alt=media)",
      "votes": null
    },
    {
      "id": "3352712",
      "postDate": "11/29/2025 11:19:44",
      "content": "<p>That's why I mentioned, 'might be not recommended.'</p>",
      "rawMarkdown": "That's why I mentioned, 'might be not recommended.'",
      "votes": null
    },
    {
      "id": "3356454",
      "postDate": "12/01/2025 00:22:27",
      "content": "<p>These are what I did 4 years ago when I was just a 3rd-grade bachelor starter, I ran all the training and experiments purely on kaggle. And I just felt desperate when seeing kaggle gpu quota is empty. </p>",
      "rawMarkdown": "These are what I did 4 years ago when I was just a 3rd-grade bachelor starter, I ran all the training and experiments purely on kaggle. And I just felt desperate when seeing kaggle gpu quota is empty.",
      "votes": null
    },
    {
      "id": "3356756",
      "postDate": "12/01/2025 03:41:21",
      "content": "<p>For the most part, if you want to run outside of the kaggle env due to runtime caps or other reasons, i'd recommend just renting something like a 3090/4090/5090 on one of the many available cloud providers. i've personally used runpod quite a bit (no affiliation with them, just the one i ended up using), primarily due to their persistent \"volumes\". You shouldnt need something like an h100 to train a model which could win this competition, i'd imagine a 24/32gb vram would be enough, though it would be smaller batch sizes over longer runs. They're also available for very cheap from many providers, i've seen well under 0.50/hr frequently for 3090s </p>",
      "rawMarkdown": "For the most part, if you want to run outside of the kaggle env due to runtime caps or other reasons, i'd recommend just renting something like a 3090/4090/5090 on one of the many available cloud providers. i've personally used runpod quite a bit (no affiliation with them, just the one i ended up using), primarily due to their persistent \"volumes\". You shouldnt need something like an h100 to train a model which could win this competition, i'd imagine a 24/32gb vram would be enough, though it would be smaller batch sizes over longer runs. They're also available for very cheap from many providers, i've seen well under 0.50/hr frequently for 3090s",
      "votes": null
    },
    {
      "id": "3364286",
      "postDate": "12/06/2025 11:46:12",
      "content": "<p>many, but most students participating don't have that luxury.</p>",
      "rawMarkdown": "many, but most students participating don't have that luxury.",
      "votes": null
    },
    {
      "id": "3364346",
      "postDate": "12/06/2025 12:28:53",
      "content": "<p>If anyone attempts to use TPU, you can train the model with less than 1 minute.</p>\n<pre><code>Keras 3:\n  - jax backend\n  - tf.data API (MUST)\n</code></pre>\n<p>I was able to train a <strong>70M</strong> param 3D model on TPU with epoch less than 1 mintute. Input shape <strong>(128,128,128,1)</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2Fff668b460599327c11f05b09076d37ad%2FScreenshot%202025-12-06%20182402.png?generation=1765023866467460&amp;alt=media\" alt=\"\"></p>\n<p>If you're solely torch user, still you can possibly leverage this if you use pytorch-lightning. Looks like TPU setup is comparatively easy here than pure torch. Though I didn't tested it, pytorch-lightning isn't available in my region, so couldn't check the api guide :(</p>",
      "rawMarkdown": "If anyone attempts to use TPU, you can train the model with less than 1 minute.\n```bash\nKeras 3:\n  - jax backend\n  - tf.data API (MUST)\n```\nI was able to train a **70M** param 3D model on TPU with epoch less than 1 mintute. Input shape **(128,128,128,1)**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2Fff668b460599327c11f05b09076d37ad%2FScreenshot%202025-12-06%20182402.png?generation=1765023866467460&alt=media)\n\nIf you're solely torch user, still you can possibly leverage this if you use pytorch-lightning. Looks like TPU setup is comparatively easy here than pure torch. Though I didn't tested it, pytorch-lightning isn't available in my region, so couldn't check the api guide :(",
      "votes": null
    },
    {
      "id": "3364352",
      "postDate": "12/06/2025 12:32:18",
      "content": "<p>What is the model size and any augmentation? Otherwise with all the extra tasks on GPU is about 3 minuts even for larger volume 160<em>160</em>160</p>",
      "rawMarkdown": "What is the model size and any augmentation? Otherwise with all the extra tasks on GPU is about 3 minuts even for larger volume 160*160*160",
      "votes": null
    },
    {
      "id": "3364371",
      "postDate": "12/06/2025 12:44:25",
      "content": "<p>The model is <a href=\"https://github.com/innat/medic-ai/blob/main/medicai/models/transunet/README.md\" target=\"_blank\">TransUNet</a> with SE-ResNet50 (param 70M). Augmentaitons are same I tried to use <a href=\"https://www.kaggle.com/code/ipythonx/train-vesuvius-surface-3d-detection-in-lightning\" target=\"_blank\">here, cell no. 10</a>.</p>\n<p>Offloading augmentation to the GPU is generally beneficial, but for 3D models with their high memory footprint and limited GPU resources, this approach may lead to performance or memory issues. But for 2D approach, this is safe; I wrote a quick <a href=\"https://stackoverflow.com/a/79327724/9215780\" target=\"_blank\">benchmark</a> regarding this while back.</p>",
      "rawMarkdown": "The model is [TransUNet](https://github.com/innat/medic-ai/blob/main/medicai/models/transunet/README.md) with SE-ResNet50 (param 70M). Augmentaitons are same I tried to use [here, cell no. 10](https://www.kaggle.com/code/ipythonx/train-vesuvius-surface-3d-detection-in-lightning).\n\nOffloading augmentation to the GPU is generally beneficial, but for 3D models with their high memory footprint and limited GPU resources, this approach may lead to performance or memory issues. But for 2D approach, this is safe; I wrote a quick [benchmark](https://stackoverflow.com/a/79327724/9215780) regarding this while back.",
      "votes": null
    },
    {
      "id": "3364386",
      "postDate": "12/06/2025 13:00:41",
      "content": "<p>I have used recently B200 card for about $4 with Lightning.ai</p>",
      "rawMarkdown": "I have used recently B200 card for about $4 with Lightning.ai",
      "votes": null
    },
    {
      "id": "3364470",
      "postDate": "12/06/2025 14:11:24",
      "content": "<p>not 'might be not recommended.'\nit is prohibited</p>",
      "rawMarkdown": "not 'might be not recommended.'\nit is prohibited",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3352481,
      "author_name": "dragonzhang",
      "author_url": "",
      "post_date": "11/29/2025 08:02:08",
      "content": "<p>many people have their own GPUs or using cloud  gpu/tpu.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3364286,
          "author_name": "choudharymanas",
          "author_url": "",
          "post_date": "12/06/2025 11:46:12",
          "content": "<p>many, but most students participating don't have that luxury.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3352543,
      "author_name": "ipythonx",
      "author_url": "",
      "post_date": "11/29/2025 08:55:01",
      "content": "<p><a href=\"https://www.kaggle.com/jirkaborovec\" target=\"_blank\">@jirkaborovec</a> \nThis suggestion is temporary and might be not recommended. </p>\n<p>The <a href=\"https://www.kaggle.com/competitions/ai-mathematical-olympiad-progress-prize-3\" target=\"_blank\">AIMO</a> competition gives H100 quota. This GPU is only for AIMO competition. But it can be used for your cause.</p>\n<ul>\n<li>Create notebook in AIMO competition.</li>\n<li>Attach your dataset (i.e. npy format of vesuvius challenge or any).</li>\n<li>Attach H100</li>\n<li>Boom.</li>\n</ul>\n<p>Only for training. Though I tried it one or two times. The speed is blazing fast.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3352676,
          "author_name": "mahoganybuttstrings",
          "author_url": "",
          "post_date": "11/29/2025 10:53:51",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F23060793%2F273658e0cc95233aa559727be46b2f2a%2FScreenshot%202025-11-29%20175312.png?generation=1764413624534990&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": [
            {
              "id": 3352712,
              "author_name": "ipythonx",
              "author_url": "",
              "post_date": "11/29/2025 11:19:44",
              "content": "<p>That's why I mentioned, 'might be not recommended.'</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3364470,
                  "author_name": "hengck23",
                  "author_url": "",
                  "post_date": "12/06/2025 14:11:24",
                  "content": "<p>not 'might be not recommended.'\nit is prohibited</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3356454,
      "author_name": "tom99763",
      "author_url": "",
      "post_date": "12/01/2025 00:22:27",
      "content": "<p>These are what I did 4 years ago when I was just a 3rd-grade bachelor starter, I ran all the training and experiments purely on kaggle. And I just felt desperate when seeing kaggle gpu quota is empty. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3356756,
      "author_name": "seanjohnsonsp",
      "author_url": "",
      "post_date": "12/01/2025 03:41:21",
      "content": "<p>For the most part, if you want to run outside of the kaggle env due to runtime caps or other reasons, i'd recommend just renting something like a 3090/4090/5090 on one of the many available cloud providers. i've personally used runpod quite a bit (no affiliation with them, just the one i ended up using), primarily due to their persistent \"volumes\". You shouldnt need something like an h100 to train a model which could win this competition, i'd imagine a 24/32gb vram would be enough, though it would be smaller batch sizes over longer runs. They're also available for very cheap from many providers, i've seen well under 0.50/hr frequently for 3090s </p>",
      "votes": null,
      "replies": [
        {
          "id": 3364386,
          "author_name": "jirkaborovec",
          "author_url": "",
          "post_date": "12/06/2025 13:00:41",
          "content": "<p>I have used recently B200 card for about $4 with Lightning.ai</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3364346,
      "author_name": "ipythonx",
      "author_url": "",
      "post_date": "12/06/2025 12:28:53",
      "content": "<p>If anyone attempts to use TPU, you can train the model with less than 1 minute.</p>\n<pre><code>Keras 3:\n  - jax backend\n  - tf.data API (MUST)\n</code></pre>\n<p>I was able to train a <strong>70M</strong> param 3D model on TPU with epoch less than 1 mintute. Input shape <strong>(128,128,128,1)</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2Fff668b460599327c11f05b09076d37ad%2FScreenshot%202025-12-06%20182402.png?generation=1765023866467460&amp;alt=media\" alt=\"\"></p>\n<p>If you're solely torch user, still you can possibly leverage this if you use pytorch-lightning. Looks like TPU setup is comparatively easy here than pure torch. Though I didn't tested it, pytorch-lightning isn't available in my region, so couldn't check the api guide :(</p>",
      "votes": null,
      "replies": [
        {
          "id": 3364352,
          "author_name": "jirkaborovec",
          "author_url": "",
          "post_date": "12/06/2025 12:32:18",
          "content": "<p>What is the model size and any augmentation? Otherwise with all the extra tasks on GPU is about 3 minuts even for larger volume 160<em>160</em>160</p>",
          "votes": null,
          "replies": [
            {
              "id": 3364371,
              "author_name": "ipythonx",
              "author_url": "",
              "post_date": "12/06/2025 12:44:25",
              "content": "<p>The model is <a href=\"https://github.com/innat/medic-ai/blob/main/medicai/models/transunet/README.md\" target=\"_blank\">TransUNet</a> with SE-ResNet50 (param 70M). Augmentaitons are same I tried to use <a href=\"https://www.kaggle.com/code/ipythonx/train-vesuvius-surface-3d-detection-in-lightning\" target=\"_blank\">here, cell no. 10</a>.</p>\n<p>Offloading augmentation to the GPU is generally beneficial, but for 3D models with their high memory footprint and limited GPU resources, this approach may lead to performance or memory issues. But for 2D approach, this is safe; I wrote a quick <a href=\"https://stackoverflow.com/a/79327724/9215780\" target=\"_blank\">benchmark</a> regarding this while back.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3351505": "I’ve been experimenting with ways to handle longer training runs within Kaggle’s resource limits. One approach that worked well for me is kernel checkpointing:\n\n- Train until you hit the kernel run limit\n- Offload checkpoints to a Kaggle Dataset or Model\n- Resume training later by attaching the saved artifact back to the kernel\n\nThis makes it possible to continue training across multiple sessions, and it also doubles as a neat inference setup if you train models elsewhere (e.g. Google Colab) and want to run them on Kaggle.\n\nCurious to hear how others manage long training jobs or combine Kaggle with external resources—always looking for new tricks to make workflows smoother.",
    "3352481": "many people have their own GPUs or using cloud  gpu/tpu.",
    "3352543": "jirkaborovec \nThis suggestion is temporary and might be not recommended. \n\nThe [AIMO](https://www.kaggle.com/competitions/ai-mathematical-olympiad-progress-prize-3) competition gives H100 quota. This GPU is only for AIMO competition. But it can be used for your cause.\n\n- Create notebook in AIMO competition.\n- Attach your dataset (i.e. npy format of vesuvius challenge or any).\n- Attach H100\n- Boom.\n\nOnly for training. Though I tried it one or two times. The speed is blazing fast.",
    "3352676": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F23060793%2F273658e0cc95233aa559727be46b2f2a%2FScreenshot%202025-11-29%20175312.png?generation=1764413624534990&alt=media)",
    "3352712": "That's why I mentioned, 'might be not recommended.'",
    "3356454": "These are what I did 4 years ago when I was just a 3rd-grade bachelor starter, I ran all the training and experiments purely on kaggle. And I just felt desperate when seeing kaggle gpu quota is empty.",
    "3356756": "For the most part, if you want to run outside of the kaggle env due to runtime caps or other reasons, i'd recommend just renting something like a 3090/4090/5090 on one of the many available cloud providers. i've personally used runpod quite a bit (no affiliation with them, just the one i ended up using), primarily due to their persistent \"volumes\". You shouldnt need something like an h100 to train a model which could win this competition, i'd imagine a 24/32gb vram would be enough, though it would be smaller batch sizes over longer runs. They're also available for very cheap from many providers, i've seen well under 0.50/hr frequently for 3090s",
    "3364286": "many, but most students participating don't have that luxury.",
    "3364346": "If anyone attempts to use TPU, you can train the model with less than 1 minute.\n```bash\nKeras 3:\n  - jax backend\n  - tf.data API (MUST)\n```\nI was able to train a **70M** param 3D model on TPU with epoch less than 1 mintute. Input shape **(128,128,128,1)**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2Fff668b460599327c11f05b09076d37ad%2FScreenshot%202025-12-06%20182402.png?generation=1765023866467460&alt=media)\n\nIf you're solely torch user, still you can possibly leverage this if you use pytorch-lightning. Looks like TPU setup is comparatively easy here than pure torch. Though I didn't tested it, pytorch-lightning isn't available in my region, so couldn't check the api guide :(",
    "3364352": "What is the model size and any augmentation? Otherwise with all the extra tasks on GPU is about 3 minuts even for larger volume 160*160*160",
    "3364371": "The model is [TransUNet](https://github.com/innat/medic-ai/blob/main/medicai/models/transunet/README.md) with SE-ResNet50 (param 70M). Augmentaitons are same I tried to use [here, cell no. 10](https://www.kaggle.com/code/ipythonx/train-vesuvius-surface-3d-detection-in-lightning).\n\nOffloading augmentation to the GPU is generally beneficial, but for 3D models with their high memory footprint and limited GPU resources, this approach may lead to performance or memory issues. But for 2D approach, this is safe; I wrote a quick [benchmark](https://stackoverflow.com/a/79327724/9215780) regarding this while back.",
    "3364386": "I have used recently B200 card for about $4 with Lightning.ai",
    "3364470": "not 'might be not recommended.'\nit is prohibited"
  },
  "source": "meta"
}