{
  "id": 475564,
  "title": "What CPU/GPU/MEM workstation did you use for training?",
  "url": "/competitions/blood-vessel-segmentation/discussion/475564",
  "author_name": "velangovan",
  "post_date": "2024-02-08T22:22:27.890000",
  "votes": 3,
  "comment_count": 13,
  "views": 0,
  "content": "<p>One of the main challenges I faced when working on this competition is the 30 hours per week GPU quota limit in Kaggle. I experimented with using GoogleColab, it was not usable to me because of the lack of background session capability, lack of input data persistence between sessions, google drive mounting/copying too slow, CPU memory lower than Kaggle on the free version etc. I also tried the “Upgrade to Google Cloud AI notebooks” and it turns out the $300 free credits are not usable for GPU instances, and the GPU instances equivalent to Kaggle runtime were priced at around $500 per month. I saw that <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> shared in a notebook that they were using the following which is in a whole different league than Kaggle resources.</p>\n<p>Z8-G4 Data Science Workstation</p>\n<ul>\n<li>GPU: 2x Nvidia Quadro RTX 8000, each with VRAM 48 GB</li>\n<li>CPU: Intel® Xeon(R) Gold 6240 CPU @ 2.60GHz, 72 cores</li>\n<li>Memory: 376 GB RAM</li>\n</ul>\n<p>Now that I read the 1st place solution by <a href=\"https://www.kaggle.com/clevert\" target=\"_blank\">@clevert</a>, I realized that there is no way that all kidneys training at that resolution will fit in a Kaggle runtime. Some other 3D training solutions posted would also not fit in Kaggle runtime.</p>\n<p>I am curious to know how many teams here, especially medal winning ones, used Kaggle CPU/GPU exclusively for training, cross validation and inference? Please raise your hand by replying to this post. Special Kudos to those fellow teams!</p>\n<p>Also, to those who used another more powerful workstation, would you be willing to share the specs of here?</p>",
  "messages": [
    {
      "id": 2643652,
      "postDate": "2024-02-09T01:59:01.440Z",
      "content": "<p>wow, everyone has high gpu computation and memory!</p>\n<p>How a newbie self-study datascience plan is changing:</p>\n<p>year 2012:</p>\n<ul>\n<li>read paper </li>\n<li>participate in kaggle</li>\n</ul>\n<p>year 2014:</p>\n<ul>\n<li>read paper </li>\n<li>read github </li>\n<li>learn online course </li>\n<li>win a bronze in kaggle</li>\n</ul>\n<p>year 2018:</p>\n<ul>\n<li>read paper </li>\n<li>read github </li>\n<li>learn online course</li>\n<li>setup various autotune tools, train framework tools   </li>\n<li>win a silver in kaggle</li>\n</ul>\n<p>year 2022:</p>\n<ul>\n<li>read paper </li>\n<li>read github </li>\n<li>learn online course</li>\n<li>setup various autotune tools, train framework tools     + accelerate/distribution training framework </li>\n<li>setup cloud gpu for scaleup training</li>\n<li>earn money or find sponsor to buy gpu machine</li>\n<li>learn from discord, etc</li>\n<li>win a gold in kaggle</li>\n</ul>\n<p>year 2024:</p>\n<ul>\n<li>read paper </li>\n<li>read github </li>\n<li>learn online course</li>\n<li>setup various autotune tools, train framework tools</li>\n<li>setup cloud gpu for scaleup training</li>\n<li>earn money or find sponsor to buy gpu machine</li>\n<li>learn from discord, etc</li>\n<li>how to ask chatgpt for help</li>\n<li>foundation models to label more data?</li>\n<li>top 3 in kaggle</li>\n</ul>",
      "rawMarkdown": "wow, everyone has high gpu computation and memory!\n\nHow a newbie self-study datascience plan is changing:\n\nyear 2012:\n- read paper \n- participate in kaggle\n\nyear 2014:\n- read paper \n- read github \n- learn online course \n- win a bronze in kaggle\n\nyear 2018:\n- read paper \n- read github \n- learn online course\n- setup various autotune tools, train framework tools   \n- win a silver in kaggle\n\nyear 2022:\n- read paper \n- read github \n- learn online course\n- setup various autotune tools, train framework tools     + accelerate/distribution training framework \n- setup cloud gpu for scaleup training\n- earn money or find sponsor to buy gpu machine\n- learn from discord, etc\n- win a gold in kaggle\n\n\nyear 2024:\n- read paper \n- read github \n- learn online course\n- setup various autotune tools, train framework tools\n- setup cloud gpu for scaleup training\n- earn money or find sponsor to buy gpu machine\n- learn from discord, etc\n- how to ask chatgpt for help\n- foundation models to label more data?\n- top 3 in kaggle\n",
      "votes": 3,
      "replies": [
        {
          "id": 2643656,
          "postDate": "2024-02-09T02:06:29.830Z",
          "rawMarkdown": "",
          "isDeleted": true,
          "replies": [
            {
              "id": 2643664,
              "postDate": "2024-02-09T02:27:51.317Z",
              "content": "<p>though uncommon, but not rare.<br>\ni do receive a couple of emails, expressing interest in sponsoring GPU, tools etc … from time to time</p>\n<p>if you check website, you can also find invitation (for being ambassador ) via write in for a couple of products.</p>\n<p>i would suggest why don't you just write in and try. but before that, give good proposal (why should people sponsor you a gpu or cloud gpu) … it is good if you have large followers in youtube, tiktok, kaggler, blogs, linkedin, face-to-face/virtual meetup etc becuase these are good for marketing, etc</p>",
              "rawMarkdown": "though uncommon, but not rare.\ni do receive a couple of emails, expressing interest in sponsoring GPU, tools etc ... from time to time\n\nif you check website, you can also find invitation (for being ambassador ) via write in for a couple of products.\n\ni would suggest why don't you just write in and try. but before that, give good proposal (why should people sponsor you a gpu or cloud gpu) ... it is good if you have large followers in youtube, tiktok, kaggler, blogs, linkedin, face-to-face/virtual meetup etc becuase these are good for marketing, etc",
              "votes": 1
            },
            {
              "id": 2643670,
              "postDate": "2024-02-09T02:35:00.280Z",
              "content": "<p>just take this competition as an example.</p>\n<ol>\n<li>start a discord group and ask people to signup</li>\n<li>what you want to do is to create a study group so that each of the members can modify their code to get same results as the top. </li>\n<li>in the study group we share code, experiment results, discussion … do post submisisons, more ablation study, presnetation, etc …</li>\n<li>if you do this over time you can get followers … and then ….</li>\n</ol>\n<hr>\n<p>in short, if you want to to be very sucessful in kaggle, you need powerful resources, particularly gpu. you have to find them. finding resources is part of the competition</p>",
              "rawMarkdown": "just take this competition as an example.\n\n1. start a discord group and ask people to signup\n2. what you want to do is to create a study group so that each of the members can modify their code to get same results as the top. \n3. in the study group we share code, experiment results, discussion ... do post submisisons, more ablation study, presnetation, etc ...\n4. if you do this over time you can get followers ... and then ....\n\n\n---\n\nin short, if you want to to be very sucessful in kaggle, you need powerful resources, particularly gpu. you have to find them. finding resources is part of the competition\n "
            },
            {
              "id": 2643672,
              "postDate": "2024-02-09T02:35:09.597Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 2643543,
      "postDate": "2024-02-08T22:22:27.890Z",
      "content": "<p>One of the main challenges I faced when working on this competition is the 30 hours per week GPU quota limit in Kaggle. I experimented with using GoogleColab, it was not usable to me because of the lack of background session capability, lack of input data persistence between sessions, google drive mounting/copying too slow, CPU memory lower than Kaggle on the free version etc. I also tried the “Upgrade to Google Cloud AI notebooks” and it turns out the $300 free credits are not usable for GPU instances, and the GPU instances equivalent to Kaggle runtime were priced at around $500 per month. I saw that <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> shared in a notebook that they were using the following which is in a whole different league than Kaggle resources.</p>\n<p>Z8-G4 Data Science Workstation</p>\n<ul>\n<li>GPU: 2x Nvidia Quadro RTX 8000, each with VRAM 48 GB</li>\n<li>CPU: Intel® Xeon(R) Gold 6240 CPU @ 2.60GHz, 72 cores</li>\n<li>Memory: 376 GB RAM</li>\n</ul>\n<p>Now that I read the 1st place solution by <a href=\"https://www.kaggle.com/clevert\" target=\"_blank\">@clevert</a>, I realized that there is no way that all kidneys training at that resolution will fit in a Kaggle runtime. Some other 3D training solutions posted would also not fit in Kaggle runtime.</p>\n<p>I am curious to know how many teams here, especially medal winning ones, used Kaggle CPU/GPU exclusively for training, cross validation and inference? Please raise your hand by replying to this post. Special Kudos to those fellow teams!</p>\n<p>Also, to those who used another more powerful workstation, would you be willing to share the specs of here?</p>",
      "rawMarkdown": "One of the main challenges I faced when working on this competition is the 30 hours per week GPU quota limit in Kaggle. I experimented with using GoogleColab, it was not usable to me because of the lack of background session capability, lack of input data persistence between sessions, google drive mounting/copying too slow, CPU memory lower than Kaggle on the free version etc. I also tried the “Upgrade to Google Cloud AI notebooks” and it turns out the $300 free credits are not usable for GPU instances, and the GPU instances equivalent to Kaggle runtime were priced at around $500 per month. I saw that @hengck23 shared in a notebook that they were using the following which is in a whole different league than Kaggle resources.\n\nZ8-G4 Data Science Workstation\n* GPU: 2x Nvidia Quadro RTX 8000, each with VRAM 48 GB\n* CPU: Intel® Xeon(R) Gold 6240 CPU @ 2.60GHz, 72 cores\n* Memory: 376 GB RAM\n\nNow that I read the 1st place solution by @clevert, I realized that there is no way that all kidneys training at that resolution will fit in a Kaggle runtime. Some other 3D training solutions posted would also not fit in Kaggle runtime.\n\nI am curious to know how many teams here, especially medal winning ones, used Kaggle CPU/GPU exclusively for training, cross validation and inference? Please raise your hand by replying to this post. Special Kudos to those fellow teams!\n\nAlso, to those who used another more powerful workstation, would you be willing to share the specs of here?",
      "votes": 3
    },
    {
      "id": 2643649,
      "postDate": "2024-02-09T01:43:30.847Z",
      "content": "<p>I was actually able to get away with training on 40GB of gpu this time around but I only trained with 2 kidney types. </p>",
      "rawMarkdown": "I was actually able to get away with training on 40GB of gpu this time around but I only trained with 2 kidney types. ",
      "votes": 1
    },
    {
      "id": 2643612,
      "postDate": "2024-02-09T00:37:00.737Z",
      "content": "<p>I used an rtx 3060 (12 gb vram) for most of this competition. I could implement most ideas, just needed to keep batch size small or number of filters low.</p>",
      "rawMarkdown": "I used an rtx 3060 (12 gb vram) for most of this competition. I could implement most ideas, just needed to keep batch size small or number of filters low.",
      "votes": 1
    },
    {
      "id": 2643631,
      "postDate": "2024-02-09T01:20:30.890Z",
      "content": "<p>1x4090, 80GB CPU RAM.  I also used laptop 4060 for code writing and code testing short runs.  </p>",
      "rawMarkdown": "1x4090, 80GB CPU RAM.  I also used laptop 4060 for code writing and code testing short runs.  ",
      "votes": 2
    },
    {
      "id": 2643597,
      "postDate": "2024-02-09T00:08:27.803Z",
      "content": "<p>\"Now that I read the 1st place solution by <a href=\"https://www.kaggle.com/clevert\" target=\"_blank\">@clevert</a>, I realized that there is no way that all kidneys training at that resolution will fit in a Kaggle runtime.\"</p>\n<p>actaully if your model did not have batch norm (e.g. repalce with layer norm or group norm), you can use gradient accumation:</p>\n<pre><code>  batch  train data\n     gradient\n      image   batch:\n              loss = net( image)\n              compute gradient ( accumulate )  backward()  \n     backprop \n</code></pre>\n<p>this is how i train my 3d for crop size 256x256x256</p>",
      "rawMarkdown": "\"Now that I read the 1st place solution by @clevert, I realized that there is no way that all kidneys training at that resolution will fit in a Kaggle runtime.\"\n\n\nactaully if your model did not have batch norm (e.g. repalce with layer norm or group norm), you can use gradient accumation:\n\n```\nfor each batch in train data\n    zero gradient\n    for each image  in batch:\n              loss = net(one image)\n              compute gradient (and accumulate ) using backward()  # (loss/batch_size).backward()\n    do backprop # optimizer.step()     \n\n```\n\nthis is how i train my 3d for crop size 256x256x256",
      "votes": 2,
      "replies": [
        {
          "id": 2643611,
          "postDate": "2024-02-09T00:35:40.083Z",
          "content": "<p>The tip about later norm or group norm is a good one. I want to learn more about those, as well as gradient accumulation, next time. Another thing relevant to OP's post is mixed precision training, which is also on my list.</p>",
          "rawMarkdown": "The tip about later norm or group norm is a good one. I want to learn more about those, as well as gradient accumulation, next time. Another thing relevant to OP's post is mixed precision training, which is also on my list.",
          "replies": [
            {
              "id": 2643627,
              "postDate": "2024-02-09T01:05:03.950Z",
              "content": "<p>the gradient accumation is not enough, you still gradient checkpointing<br>\nnow kaggle gpu has about 16 GB, with  gradient checkpointing, you will have virtually x2 (32 GB) of train vram, but slower training iteration.</p>\n<p><a href=\"https://github.com/prigoyal/pytorch_memonger/blob/master/tutorial/Checkpointing_for_PyTorch_models.ipynb\" target=\"_blank\">https://github.com/prigoyal/pytorch_memonger/blob/master/tutorial/Checkpointing_for_PyTorch_models.ipynb</a><br>\n<a href=\"https://huggingface.co/docs/transformers/v4.20.1/en/perf_train_gpu_one\" target=\"_blank\">https://huggingface.co/docs/transformers/v4.20.1/en/perf_train_gpu_one</a><br>\n<a href=\"https://huggingface.co/docs/transformers/v4.18.0/en/performance\" target=\"_blank\">https://huggingface.co/docs/transformers/v4.18.0/en/performance</a></p>\n<p><a href=\"https://medium.com/geekculture/training-larger-models-over-your-average-gpu-with-gradient-checkpointing-in-pytorch-571b4b5c2068\" target=\"_blank\">https://medium.com/geekculture/training-larger-models-over-your-average-gpu-with-gradient-checkpointing-in-pytorch-571b4b5c2068</a></p>",
              "rawMarkdown": "the gradient accumation is not enough, you still gradient checkpointing\nnow kaggle gpu has about 16 GB, with  gradient checkpointing, you will have virtually x2 (32 GB) of train vram, but slower training iteration.\n\nhttps://github.com/prigoyal/pytorch_memonger/blob/master/tutorial/Checkpointing_for_PyTorch_models.ipynb\nhttps://huggingface.co/docs/transformers/v4.20.1/en/perf_train_gpu_one\nhttps://huggingface.co/docs/transformers/v4.18.0/en/performance\n\nhttps://medium.com/geekculture/training-larger-models-over-your-average-gpu-with-gradient-checkpointing-in-pytorch-571b4b5c2068"
            },
            {
              "id": 2643724,
              "postDate": "2024-02-09T03:53:53.397Z",
              "content": "<p>Thanks for the references <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> ! I will read up on this :)</p>",
              "rawMarkdown": "Thanks for the references @hengck23 ! I will read up on this :)"
            }
          ]
        }
      ]
    },
    {
      "id": 2643549,
      "postDate": "2024-02-08T22:26:54.273Z",
      "content": "<p>I used Kaggle GPUs for trainings only in my 1st competition on Kaggle. Back then you could get up to 6 instances that were running for (I think) 9 hours straight. </p>\n<p>Quickly after that the quotas were reduced, and I had to move to my personal hardware. It was always lacking in power, but in the past year managed to build a devbox with 2xA6000 Ada. So 96GB of VRAM and (almost) speed of RTX 4090. Finally, I could implement most of ideas in competitions. </p>",
      "rawMarkdown": "I used Kaggle GPUs for trainings only in my 1st competition on Kaggle. Back then you could get up to 6 instances that were running for (I think) 9 hours straight. \n\nQuickly after that the quotas were reduced, and I had to move to my personal hardware. It was always lacking in power, but in the past year managed to build a devbox with 2xA6000 Ada. So 96GB of VRAM and (almost) speed of RTX 4090. Finally, I could implement most of ideas in competitions. ",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 2643652,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-02-09T01:59:01.440000",
      "content": "<p>wow, everyone has high gpu computation and memory!</p>\n<p>How a newbie self-study datascience plan is changing:</p>\n<p>year 2012:</p>\n<ul>\n<li>read paper </li>\n<li>participate in kaggle</li>\n</ul>\n<p>year 2014:</p>\n<ul>\n<li>read paper </li>\n<li>read github </li>\n<li>learn online course </li>\n<li>win a bronze in kaggle</li>\n</ul>\n<p>year 2018:</p>\n<ul>\n<li>read paper </li>\n<li>read github </li>\n<li>learn online course</li>\n<li>setup various autotune tools, train framework tools   </li>\n<li>win a silver in kaggle</li>\n</ul>\n<p>year 2022:</p>\n<ul>\n<li>read paper </li>\n<li>read github </li>\n<li>learn online course</li>\n<li>setup various autotune tools, train framework tools     + accelerate/distribution training framework </li>\n<li>setup cloud gpu for scaleup training</li>\n<li>earn money or find sponsor to buy gpu machine</li>\n<li>learn from discord, etc</li>\n<li>win a gold in kaggle</li>\n</ul>\n<p>year 2024:</p>\n<ul>\n<li>read paper </li>\n<li>read github </li>\n<li>learn online course</li>\n<li>setup various autotune tools, train framework tools</li>\n<li>setup cloud gpu for scaleup training</li>\n<li>earn money or find sponsor to buy gpu machine</li>\n<li>learn from discord, etc</li>\n<li>how to ask chatgpt for help</li>\n<li>foundation models to label more data?</li>\n<li>top 3 in kaggle</li>\n</ul>",
      "votes": 3,
      "replies": [
        {
          "id": 2643656,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-02-09T02:06:29.830000",
          "content": "",
          "votes": 0,
          "replies": [
            {
              "id": 2643664,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-02-09T02:27:51.317000",
              "content": "<p>though uncommon, but not rare.<br>\ni do receive a couple of emails, expressing interest in sponsoring GPU, tools etc … from time to time</p>\n<p>if you check website, you can also find invitation (for being ambassador ) via write in for a couple of products.</p>\n<p>i would suggest why don't you just write in and try. but before that, give good proposal (why should people sponsor you a gpu or cloud gpu) … it is good if you have large followers in youtube, tiktok, kaggler, blogs, linkedin, face-to-face/virtual meetup etc becuase these are good for marketing, etc</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2643670,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-02-09T02:35:00.280000",
              "content": "<p>just take this competition as an example.</p>\n<ol>\n<li>start a discord group and ask people to signup</li>\n<li>what you want to do is to create a study group so that each of the members can modify their code to get same results as the top. </li>\n<li>in the study group we share code, experiment results, discussion … do post submisisons, more ablation study, presnetation, etc …</li>\n<li>if you do this over time you can get followers … and then ….</li>\n</ol>\n<hr>\n<p>in short, if you want to to be very sucessful in kaggle, you need powerful resources, particularly gpu. you have to find them. finding resources is part of the competition</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2643672,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-02-09T02:35:09.597000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2643649,
      "author_name": "Cody_Null",
      "author_url": "",
      "post_date": "2024-02-09T01:43:30.847000",
      "content": "<p>I was actually able to get away with training on 40GB of gpu this time around but I only trained with 2 kidney types. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2643612,
      "author_name": "chemdatafarmer",
      "author_url": "",
      "post_date": "2024-02-09T00:37:00.737000",
      "content": "<p>I used an rtx 3060 (12 gb vram) for most of this competition. I could implement most ideas, just needed to keep batch size small or number of filters low.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2643631,
      "author_name": "joejeo1",
      "author_url": "",
      "post_date": "2024-02-09T01:20:30.890000",
      "content": "<p>1x4090, 80GB CPU RAM.  I also used laptop 4060 for code writing and code testing short runs.  </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2643597,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-02-09T00:08:27.803000",
      "content": "<p>\"Now that I read the 1st place solution by <a href=\"https://www.kaggle.com/clevert\" target=\"_blank\">@clevert</a>, I realized that there is no way that all kidneys training at that resolution will fit in a Kaggle runtime.\"</p>\n<p>actaully if your model did not have batch norm (e.g. repalce with layer norm or group norm), you can use gradient accumation:</p>\n<pre><code>  batch  train data\n     gradient\n      image   batch:\n              loss = net( image)\n              compute gradient ( accumulate )  backward()  \n     backprop \n</code></pre>\n<p>this is how i train my 3d for crop size 256x256x256</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2643611,
          "author_name": "chemdatafarmer",
          "author_url": "",
          "post_date": "2024-02-09T00:35:40.083000",
          "content": "<p>The tip about later norm or group norm is a good one. I want to learn more about those, as well as gradient accumulation, next time. Another thing relevant to OP's post is mixed precision training, which is also on my list.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2643627,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-02-09T01:05:03.950000",
              "content": "<p>the gradient accumation is not enough, you still gradient checkpointing<br>\nnow kaggle gpu has about 16 GB, with  gradient checkpointing, you will have virtually x2 (32 GB) of train vram, but slower training iteration.</p>\n<p><a href=\"https://github.com/prigoyal/pytorch_memonger/blob/master/tutorial/Checkpointing_for_PyTorch_models.ipynb\" target=\"_blank\">https://github.com/prigoyal/pytorch_memonger/blob/master/tutorial/Checkpointing_for_PyTorch_models.ipynb</a><br>\n<a href=\"https://huggingface.co/docs/transformers/v4.20.1/en/perf_train_gpu_one\" target=\"_blank\">https://huggingface.co/docs/transformers/v4.20.1/en/perf_train_gpu_one</a><br>\n<a href=\"https://huggingface.co/docs/transformers/v4.18.0/en/performance\" target=\"_blank\">https://huggingface.co/docs/transformers/v4.18.0/en/performance</a></p>\n<p><a href=\"https://medium.com/geekculture/training-larger-models-over-your-average-gpu-with-gradient-checkpointing-in-pytorch-571b4b5c2068\" target=\"_blank\">https://medium.com/geekculture/training-larger-models-over-your-average-gpu-with-gradient-checkpointing-in-pytorch-571b4b5c2068</a></p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2643724,
              "author_name": "chemdatafarmer",
              "author_url": "",
              "post_date": "2024-02-09T03:53:53.397000",
              "content": "<p>Thanks for the references <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> ! I will read up on this :)</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2643549,
      "author_name": "Ivan Panshin",
      "author_url": "",
      "post_date": "2024-02-08T22:26:54.273000",
      "content": "<p>I used Kaggle GPUs for trainings only in my 1st competition on Kaggle. Back then you could get up to 6 instances that were running for (I think) 9 hours straight. </p>\n<p>Quickly after that the quotas were reduced, and I had to move to my personal hardware. It was always lacking in power, but in the past year managed to build a devbox with 2xA6000 Ada. So 96GB of VRAM and (almost) speed of RTX 4090. Finally, I could implement most of ideas in competitions. </p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2643652": "wow, everyone has high gpu computation and memory!\n\nHow a newbie self-study datascience plan is changing:\n\nyear 2012:\n- read paper \n- participate in kaggle\n\nyear 2014:\n- read paper \n- read github \n- learn online course \n- win a bronze in kaggle\n\nyear 2018:\n- read paper \n- read github \n- learn online course\n- setup various autotune tools, train framework tools   \n- win a silver in kaggle\n\nyear 2022:\n- read paper \n- read github \n- learn online course\n- setup various autotune tools, train framework tools     + accelerate/distribution training framework \n- setup cloud gpu for scaleup training\n- earn money or find sponsor to buy gpu machine\n- learn from discord, etc\n- win a gold in kaggle\n\n\nyear 2024:\n- read paper \n- read github \n- learn online course\n- setup various autotune tools, train framework tools\n- setup cloud gpu for scaleup training\n- earn money or find sponsor to buy gpu machine\n- learn from discord, etc\n- how to ask chatgpt for help\n- foundation models to label more data?\n- top 3 in kaggle\n",
    "2643543": "One of the main challenges I faced when working on this competition is the 30 hours per week GPU quota limit in Kaggle. I experimented with using GoogleColab, it was not usable to me because of the lack of background session capability, lack of input data persistence between sessions, google drive mounting/copying too slow, CPU memory lower than Kaggle on the free version etc. I also tried the “Upgrade to Google Cloud AI notebooks” and it turns out the $300 free credits are not usable for GPU instances, and the GPU instances equivalent to Kaggle runtime were priced at around $500 per month. I saw that @hengck23 shared in a notebook that they were using the following which is in a whole different league than Kaggle resources.\n\nZ8-G4 Data Science Workstation\n* GPU: 2x Nvidia Quadro RTX 8000, each with VRAM 48 GB\n* CPU: Intel® Xeon(R) Gold 6240 CPU @ 2.60GHz, 72 cores\n* Memory: 376 GB RAM\n\nNow that I read the 1st place solution by @clevert, I realized that there is no way that all kidneys training at that resolution will fit in a Kaggle runtime. Some other 3D training solutions posted would also not fit in Kaggle runtime.\n\nI am curious to know how many teams here, especially medal winning ones, used Kaggle CPU/GPU exclusively for training, cross validation and inference? Please raise your hand by replying to this post. Special Kudos to those fellow teams!\n\nAlso, to those who used another more powerful workstation, would you be willing to share the specs of here?",
    "2643649": "I was actually able to get away with training on 40GB of gpu this time around but I only trained with 2 kidney types. ",
    "2643612": "I used an rtx 3060 (12 gb vram) for most of this competition. I could implement most ideas, just needed to keep batch size small or number of filters low.",
    "2643631": "1x4090, 80GB CPU RAM.  I also used laptop 4060 for code writing and code testing short runs.  ",
    "2643597": "\"Now that I read the 1st place solution by @clevert, I realized that there is no way that all kidneys training at that resolution will fit in a Kaggle runtime.\"\n\n\nactaully if your model did not have batch norm (e.g. repalce with layer norm or group norm), you can use gradient accumation:\n\n```\nfor each batch in train data\n    zero gradient\n    for each image  in batch:\n              loss = net(one image)\n              compute gradient (and accumulate ) using backward()  # (loss/batch_size).backward()\n    do backprop # optimizer.step()     \n\n```\n\nthis is how i train my 3d for crop size 256x256x256",
    "2643549": "I used Kaggle GPUs for trainings only in my 1st competition on Kaggle. Back then you could get up to 6 instances that were running for (I think) 9 hours straight. \n\nQuickly after that the quotas were reduced, and I had to move to my personal hardware. It was always lacking in power, but in the past year managed to build a devbox with 2xA6000 Ada. So 96GB of VRAM and (almost) speed of RTX 4090. Finally, I could implement most of ideas in competitions. "
  }
}