{
  "id": 313506,
  "title": "How to train without HUGE GPU?",
  "url": "/competitions/happy-whale-and-dolphin/discussion/313506",
  "author_name": "",
  "post_date": "2022-03-17T13:48:50.239299600Z",
  "votes": 10,
  "comment_count": 19,
  "views": 0,
  "content": "<p>I personally own a RTX 2080 GPU and I trained Arc Face basic NN (embeddings to find new individuals not used yet) within my 8GB GPU RAM. Used detic cropped dataset, with image resized to 56x56, barely got to a score of 0.1 . Then I tried vast ai GPU providers to rent out 4x3090s to train on 400x400 images, same procedure, score now jumped to 0.4. But it did cost me a lot, like 50 dollars. It is going to cost a lot more if I have to experiment with various architectures, methodologies etc. <br>\nSo here's my question, are some kaggle competitions, inherently this costly? or am I missing something here? Thanks.</p>",
  "messages": [
    {
      "id": "1725884",
      "postDate": "03/17/2022 13:48:50",
      "content": "<p>I personally own a RTX 2080 GPU and I trained Arc Face basic NN (embeddings to find new individuals not used yet) within my 8GB GPU RAM. Used detic cropped dataset, with image resized to 56x56, barely got to a score of 0.1 . Then I tried vast ai GPU providers to rent out 4x3090s to train on 400x400 images, same procedure, score now jumped to 0.4. But it did cost me a lot, like 50 dollars. It is going to cost a lot more if I have to experiment with various architectures, methodologies etc. <br>\nSo here's my question, are some kaggle competitions, inherently this costly? or am I missing something here? Thanks.</p>",
      "rawMarkdown": "I personally own a RTX 2080 GPU and I trained Arc Face basic NN (embeddings to find new individuals not used yet) within my 8GB GPU RAM. Used detic cropped dataset, with image resized to 56x56, barely got to a score of 0.1 . Then I tried vast ai GPU providers to rent out 4x3090s to train on 400x400 images, same procedure, score now jumped to 0.4. But it did cost me a lot, like 50 dollars. It is going to cost a lot more if I have to experiment with various architectures, methodologies etc. \nSo here's my question, are some kaggle competitions, inherently this costly? or am I missing something here? Thanks.",
      "votes": null
    },
    {
      "id": "1725899",
      "postDate": "03/17/2022 13:57:04",
      "content": "<p>It is unspoken truth that Kagglers with great hardware have significant advantage.<br>\nYou can find heated discussions about this if you dig deep in topics.<br>\nNevertheless, keep up the good work!</p>",
      "rawMarkdown": "It is unspoken truth that Kagglers with great hardware have significant advantage.\nYou can find heated discussions about this if you dig deep in topics.\nNevertheless, keep up the good work!",
      "votes": null
    },
    {
      "id": "1725906",
      "postDate": "03/17/2022 13:59:21",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/vivovinco\" target=\"_blank\">@vivovinco</a> </p>",
      "rawMarkdown": "Thanks @vivovinco",
      "votes": null
    },
    {
      "id": "1725933",
      "postDate": "03/17/2022 14:23:07",
      "content": "<p>then  just learn the models and pipelines. </p>",
      "rawMarkdown": "then  just learn the models and pipelines.",
      "votes": null
    },
    {
      "id": "1725946",
      "postDate": "03/17/2022 14:31:47",
      "content": "<ol>\n<li>Use the gradient accumulation. You can find a really brief tutorial <a href=\"https://kozodoi.me/python/deep%20learning/pytorch/tutorial/2021/02/19/gradient-accumulation.html\" target=\"_blank\">here</a>.</li>\n<li>Reduce memory usage by computing in FP16. If you use PyTorch as a base framework, here is the <a href=\"https://pytorch.org/blog/accelerating-training-on-nvidia-gpus-with-pytorch-automatic-mixed-precision/\" target=\"_blank\">documentation</a>. It is also important here to choose architecture, which is very beneficial in FP16 training</li>\n</ol>\n<p>I believe you can get my current score using 2080, far from the top but something. I use one GPU with 12Gb of memory</p>",
      "rawMarkdown": "1. Use the gradient accumulation. You can find a really brief tutorial [here](https://kozodoi.me/python/deep%20learning/pytorch/tutorial/2021/02/19/gradient-accumulation.html).\n2. Reduce memory usage by computing in FP16. If you use PyTorch as a base framework, here is the [documentation](https://pytorch.org/blog/accelerating-training-on-nvidia-gpus-with-pytorch-automatic-mixed-precision/). It is also important here to choose architecture, which is very beneficial in FP16 training\n\nI believe you can get my current score using 2080, far from the top but something. I use one GPU with 12Gb of memory",
      "votes": null
    },
    {
      "id": "1725972",
      "postDate": "03/17/2022 14:55:41",
      "content": "<p>Please refer or point to something, that will be much more helpful, rather than throw around generic terms.</p>",
      "rawMarkdown": "Please refer or point to something, that will be much more helpful, rather than throw around generic terms.",
      "votes": null
    },
    {
      "id": "1725987",
      "postDate": "03/17/2022 15:05:38",
      "content": "<p>Thanks that was helpful, while I was already aware of them, did not think they would give enough results, like 4x24 GB (64 batch size) cannot be down-scaled to 8GB suddenly, even with these. Mind telling about your batch size and image res on 12 GB GPU? Thanks.<br>\nOne more thing is we cannot directly 'choose' architectures right? We gotta experiment to conclude on them concretely.</p>",
      "rawMarkdown": "Thanks that was helpful, while I was already aware of them, did not think they would give enough results, like 4x24 GB (64 batch size) cannot be down-scaled to 8GB suddenly, even with these. Mind telling about your batch size and image res on 12 GB GPU? Thanks.\nOne more thing is we cannot directly 'choose' architectures right? We gotta experiment to conclude on them concretely.",
      "votes": null
    },
    {
      "id": "1725992",
      "postDate": "03/17/2022 15:11:21",
      "content": "<p>I use colab TPU. It can be a remedy for you.<br>\ncolab pro is just $10</p>",
      "rawMarkdown": "I use colab TPU. It can be a remedy for you.\ncolab pro is just $10",
      "votes": null
    },
    {
      "id": "1725999",
      "postDate": "03/17/2022 15:15:48",
      "content": "<p>Can I ask small ques here that when FP16 help the model train faster, is FP16 worse than FP32 so much? or do FP16 just reduce some small accuracy compare to FP32?</p>",
      "rawMarkdown": "Can I ask small ques here that when FP16 help the model train faster, is FP16 worse than FP32 so much? or do FP16 just reduce some small accuracy compare to FP32?",
      "votes": null
    },
    {
      "id": "1726000",
      "postDate": "03/17/2022 15:15:57",
      "content": "<p>Mind telling most ram, time of training and gpu memorys you have extracted from pro? Thanks</p>",
      "rawMarkdown": "Mind telling most ram, time of training and gpu memorys you have extracted from pro? Thanks",
      "votes": null
    },
    {
      "id": "1726009",
      "postDate": "03/17/2022 15:22:21",
      "content": "<p><a href=\"https://www.kaggle.com/dqhdqmcttdqx\" target=\"_blank\">@dqhdqmcttdqx</a> For most recent RTX GPUs, you should see a nice speedup and minimal effect on convergence. I had written a blog on how this works, if you're interested, <a href=\"https://hackernoon.com/rtx-2080ti-vs-gtx-1080ti-fastai-mixed-precision-training-comparisons-on-cifar-100-761d8f615d7f\" target=\"_blank\">here's</a> the link</p>",
      "rawMarkdown": "dqhdqmcttdqx For most recent RTX GPUs, you should see a nice speedup and minimal effect on convergence. I had written a blog on how this works, if you're interested, [here's](https://hackernoon.com/rtx-2080ti-vs-gtx-1080ti-fastai-mixed-precision-training-comparisons-on-cifar-100-761d8f615d7f) the link",
      "votes": null
    },
    {
      "id": "1726105",
      "postDate": "03/17/2022 16:49:53",
      "content": "<p>I've been using only kaggle TPU. You just need to be very cautious with the amount of time used with TPU. I have been training EffNetB5 with 224 img size which means I can do over 20 experiments per week. Then upgrade to a better model and img size when you are ready but beware that each model will take much longer to train (2x, 3x, maybe longer). I modified this <a href=\"https://www.kaggle.com/code/dragonzhang/happywhale-effnet-b7-fork-with-detic-crop\" target=\"_blank\">script for training and inference</a>. You can definitely get a silver medal and maybe a gold medal with only kaggle TPU. </p>",
      "rawMarkdown": "I've been using only kaggle TPU. You just need to be very cautious with the amount of time used with TPU. I have been training EffNetB5 with 224 img size which means I can do over 20 experiments per week. Then upgrade to a better model and img size when you are ready but beware that each model will take much longer to train (2x, 3x, maybe longer). I modified this [script for training and inference](https://www.kaggle.com/code/dragonzhang/happywhale-effnet-b7-fork-with-detic-crop). You can definitely get a silver medal and maybe a gold medal with only kaggle TPU.",
      "votes": null
    },
    {
      "id": "1727941",
      "postDate": "03/18/2022 13:04:35",
      "content": "<p>Just tried fp16 pytorch, was able to come from batchsize 2 to 4 for 600x600 image res.</p>",
      "rawMarkdown": "Just tried fp16 pytorch, was able to come from batchsize 2 to 4 for 600x600 image res.",
      "votes": null
    },
    {
      "id": "1728140",
      "postDate": "03/18/2022 16:01:05",
      "content": "<p><a href=\"https://www.kaggle.com/init27\" target=\"_blank\">@init27</a> Thanks for sharing 💯</p>",
      "rawMarkdown": "init27 Thanks for sharing 💯",
      "votes": null
    },
    {
      "id": "1728195",
      "postDate": "03/18/2022 16:43:47",
      "content": "<p>In my case, i when i'm using tpu, you can train efficient net b6 with image size 512 and have up to batch size of 80 (10 images * 8 core TPUs), and time for each epochs about 9 minutes. <br>\nfor specs:</p>\n<ul>\n<li>ram: 13gb available. (you can change to high ram-&gt;35gb) </li>\n<li>disk: 256gb available (when you use gpu it's change to 167gb)</li>\n</ul>",
      "rawMarkdown": "In my case, i when i'm using tpu, you can train efficient net b6 with image size 512 and have up to batch size of 80 (10 images * 8 core TPUs), and time for each epochs about 9 minutes. \nfor specs:\n- ram: 13gb available. (you can change to high ram->35gb) \n- disk: 256gb available (when you use gpu it's change to 167gb)",
      "votes": null
    },
    {
      "id": "1728233",
      "postDate": "03/18/2022 17:40:16",
      "content": "<p>No question that vision competitions need more compute than tabular.  And generally higher resolution images model better than low.  More compute can mean more money - or more training time.  More compute will also mean your skill set needs to be better.</p>\n<p>As noted by several, you can have success using kaggle TPU/GPU.</p>\n<ul>\n<li><p>On local PC's I have never seen significant loss in performance by using fp16 - pretty much allows for doubling of batch size which reduces training time.</p></li>\n<li><p>Image sizes of around 224 have been good for experimentation  (larger image sizes I use the kaggle TPU).  64x64 has almost always been a waste of time.</p></li>\n<li><p>Training times of around 16 hours or less have been good for experimentation.</p></li>\n<li><p>I create a large swap file on Ubuntu to avoid any cpu memory issues (200GB on an ssd works nicely).</p></li>\n<li><p>I have dual GPU's on all 4 of my machines, but GPU prices these days are crazy. </p></li>\n<li><p>Don't waste TPU quota doing cpu level work - do pipeline work in separate cpu/gpu kernels to create data sets.</p></li>\n<li><p>Don't get caught in the bs trap that folks with big hardware have a huge advantage - of course they do - turn on the news - the world is not fair.</p></li>\n</ul>\n<p>Have fun and learn</p>",
      "rawMarkdown": "No question that vision competitions need more compute than tabular.  And generally higher resolution images model better than low.  More compute can mean more money - or more training time.  More compute will also mean your skill set needs to be better.\n\nAs noted by several, you can have success using kaggle TPU/GPU.\n\n- On local PC's I have never seen significant loss in performance by using fp16 - pretty much allows for doubling of batch size which reduces training time.\n\n- Image sizes of around 224 have been good for experimentation  (larger image sizes I use the kaggle TPU).  64x64 has almost always been a waste of time.\n\n- Training times of around 16 hours or less have been good for experimentation.\n\n- I create a large swap file on Ubuntu to avoid any cpu memory issues (200GB on an ssd works nicely).\n\n- I have dual GPU's on all 4 of my machines, but GPU prices these days are crazy. \n\n- Don't waste TPU quota doing cpu level work - do pipeline work in separate cpu/gpu kernels to create data sets.\n\n- Don't get caught in the bs trap that folks with big hardware have a huge advantage - of course they do - turn on the news - the world is not fair.\n\nHave fun and learn",
      "votes": null
    },
    {
      "id": "1728401",
      "postDate": "03/18/2022 21:25:33",
      "content": "<p>for the more in-depth guide refer to:<br>\n<a href=\"https://docs.nvidia.com/deeplearning/performance/mixed-precision-training/index.html\" target=\"_blank\">https://docs.nvidia.com/deeplearning/performance/mixed-precision-training/index.html</a></p>\n<p>you'll get the idea on why you need such thing as GradScaler</p>",
      "rawMarkdown": "for the more in-depth guide refer to:\nhttps://docs.nvidia.com/deeplearning/performance/mixed-precision-training/index.html\n\nyou'll get the idea on why you need such thing as GradScaler",
      "votes": null
    },
    {
      "id": "1730105",
      "postDate": "03/20/2022 23:52:23",
      "content": "<p><a href=\"https://www.kaggle.com/achilles38\" target=\"_blank\">@achilles38</a> can you tell how much time does training of your model take?</p>",
      "rawMarkdown": "achilles38 can you tell how much time does training of your model take?",
      "votes": null
    },
    {
      "id": "1730247",
      "postDate": "03/21/2022 04:03:16",
      "content": "<p>one model takes 3 hours to train</p>",
      "rawMarkdown": "one model takes 3 hours to train",
      "votes": null
    },
    {
      "id": "1740537",
      "postDate": "03/31/2022 03:23:17",
      "content": "<p><a href=\"https://www.kaggle.com/deepkim\" target=\"_blank\">@deepkim</a> Great score considering you used only Colab TPU.</p>",
      "rawMarkdown": "deepkim Great score considering you used only Colab TPU.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1725899,
      "author_name": "vivovinco",
      "author_url": "",
      "post_date": "03/17/2022 13:57:04",
      "content": "<p>It is unspoken truth that Kagglers with great hardware have significant advantage.<br>\nYou can find heated discussions about this if you dig deep in topics.<br>\nNevertheless, keep up the good work!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1725906,
          "author_name": "aaftaabv",
          "author_url": "",
          "post_date": "03/17/2022 13:59:21",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/vivovinco\" target=\"_blank\">@vivovinco</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1725933,
      "author_name": "dragonzhang",
      "author_url": "",
      "post_date": "03/17/2022 14:23:07",
      "content": "<p>then  just learn the models and pipelines. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1725972,
          "author_name": "aaftaabv",
          "author_url": "",
          "post_date": "03/17/2022 14:55:41",
          "content": "<p>Please refer or point to something, that will be much more helpful, rather than throw around generic terms.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1725946,
      "author_name": "achilles38",
      "author_url": "",
      "post_date": "03/17/2022 14:31:47",
      "content": "<ol>\n<li>Use the gradient accumulation. You can find a really brief tutorial <a href=\"https://kozodoi.me/python/deep%20learning/pytorch/tutorial/2021/02/19/gradient-accumulation.html\" target=\"_blank\">here</a>.</li>\n<li>Reduce memory usage by computing in FP16. If you use PyTorch as a base framework, here is the <a href=\"https://pytorch.org/blog/accelerating-training-on-nvidia-gpus-with-pytorch-automatic-mixed-precision/\" target=\"_blank\">documentation</a>. It is also important here to choose architecture, which is very beneficial in FP16 training</li>\n</ol>\n<p>I believe you can get my current score using 2080, far from the top but something. I use one GPU with 12Gb of memory</p>",
      "votes": null,
      "replies": [
        {
          "id": 1725987,
          "author_name": "aaftaabv",
          "author_url": "",
          "post_date": "03/17/2022 15:05:38",
          "content": "<p>Thanks that was helpful, while I was already aware of them, did not think they would give enough results, like 4x24 GB (64 batch size) cannot be down-scaled to 8GB suddenly, even with these. Mind telling about your batch size and image res on 12 GB GPU? Thanks.<br>\nOne more thing is we cannot directly 'choose' architectures right? We gotta experiment to conclude on them concretely.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1725999,
          "author_name": "dqhdqmcttdqx",
          "author_url": "",
          "post_date": "03/17/2022 15:15:48",
          "content": "<p>Can I ask small ques here that when FP16 help the model train faster, is FP16 worse than FP32 so much? or do FP16 just reduce some small accuracy compare to FP32?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1726009,
          "author_name": "init27",
          "author_url": "",
          "post_date": "03/17/2022 15:22:21",
          "content": "<p><a href=\"https://www.kaggle.com/dqhdqmcttdqx\" target=\"_blank\">@dqhdqmcttdqx</a> For most recent RTX GPUs, you should see a nice speedup and minimal effect on convergence. I had written a blog on how this works, if you're interested, <a href=\"https://hackernoon.com/rtx-2080ti-vs-gtx-1080ti-fastai-mixed-precision-training-comparisons-on-cifar-100-761d8f615d7f\" target=\"_blank\">here's</a> the link</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1727941,
          "author_name": "aaftaabv",
          "author_url": "",
          "post_date": "03/18/2022 13:04:35",
          "content": "<p>Just tried fp16 pytorch, was able to come from batchsize 2 to 4 for 600x600 image res.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1728140,
          "author_name": "dqhdqmcttdqx",
          "author_url": "",
          "post_date": "03/18/2022 16:01:05",
          "content": "<p><a href=\"https://www.kaggle.com/init27\" target=\"_blank\">@init27</a> Thanks for sharing 💯</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1728401,
          "author_name": "martynoveduard",
          "author_url": "",
          "post_date": "03/18/2022 21:25:33",
          "content": "<p>for the more in-depth guide refer to:<br>\n<a href=\"https://docs.nvidia.com/deeplearning/performance/mixed-precision-training/index.html\" target=\"_blank\">https://docs.nvidia.com/deeplearning/performance/mixed-precision-training/index.html</a></p>\n<p>you'll get the idea on why you need such thing as GradScaler</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1730105,
          "author_name": "paawel",
          "author_url": "",
          "post_date": "03/20/2022 23:52:23",
          "content": "<p><a href=\"https://www.kaggle.com/achilles38\" target=\"_blank\">@achilles38</a> can you tell how much time does training of your model take?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1730247,
          "author_name": "achilles38",
          "author_url": "",
          "post_date": "03/21/2022 04:03:16",
          "content": "<p>one model takes 3 hours to train</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1725992,
      "author_name": "deepkim",
      "author_url": "",
      "post_date": "03/17/2022 15:11:21",
      "content": "<p>I use colab TPU. It can be a remedy for you.<br>\ncolab pro is just $10</p>",
      "votes": null,
      "replies": [
        {
          "id": 1726000,
          "author_name": "aaftaabv",
          "author_url": "",
          "post_date": "03/17/2022 15:15:57",
          "content": "<p>Mind telling most ram, time of training and gpu memorys you have extracted from pro? Thanks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1728195,
          "author_name": "locbaop",
          "author_url": "",
          "post_date": "03/18/2022 16:43:47",
          "content": "<p>In my case, i when i'm using tpu, you can train efficient net b6 with image size 512 and have up to batch size of 80 (10 images * 8 core TPUs), and time for each epochs about 9 minutes. <br>\nfor specs:</p>\n<ul>\n<li>ram: 13gb available. (you can change to high ram-&gt;35gb) </li>\n<li>disk: 256gb available (when you use gpu it's change to 167gb)</li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1740537,
          "author_name": "devanshchowdhury",
          "author_url": "",
          "post_date": "03/31/2022 03:23:17",
          "content": "<p><a href=\"https://www.kaggle.com/deepkim\" target=\"_blank\">@deepkim</a> Great score considering you used only Colab TPU.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1726105,
      "author_name": "vexxingbanana",
      "author_url": "",
      "post_date": "03/17/2022 16:49:53",
      "content": "<p>I've been using only kaggle TPU. You just need to be very cautious with the amount of time used with TPU. I have been training EffNetB5 with 224 img size which means I can do over 20 experiments per week. Then upgrade to a better model and img size when you are ready but beware that each model will take much longer to train (2x, 3x, maybe longer). I modified this <a href=\"https://www.kaggle.com/code/dragonzhang/happywhale-effnet-b7-fork-with-detic-crop\" target=\"_blank\">script for training and inference</a>. You can definitely get a silver medal and maybe a gold medal with only kaggle TPU. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1728233,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "03/18/2022 17:40:16",
      "content": "<p>No question that vision competitions need more compute than tabular.  And generally higher resolution images model better than low.  More compute can mean more money - or more training time.  More compute will also mean your skill set needs to be better.</p>\n<p>As noted by several, you can have success using kaggle TPU/GPU.</p>\n<ul>\n<li><p>On local PC's I have never seen significant loss in performance by using fp16 - pretty much allows for doubling of batch size which reduces training time.</p></li>\n<li><p>Image sizes of around 224 have been good for experimentation  (larger image sizes I use the kaggle TPU).  64x64 has almost always been a waste of time.</p></li>\n<li><p>Training times of around 16 hours or less have been good for experimentation.</p></li>\n<li><p>I create a large swap file on Ubuntu to avoid any cpu memory issues (200GB on an ssd works nicely).</p></li>\n<li><p>I have dual GPU's on all 4 of my machines, but GPU prices these days are crazy. </p></li>\n<li><p>Don't waste TPU quota doing cpu level work - do pipeline work in separate cpu/gpu kernels to create data sets.</p></li>\n<li><p>Don't get caught in the bs trap that folks with big hardware have a huge advantage - of course they do - turn on the news - the world is not fair.</p></li>\n</ul>\n<p>Have fun and learn</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1725884": "I personally own a RTX 2080 GPU and I trained Arc Face basic NN (embeddings to find new individuals not used yet) within my 8GB GPU RAM. Used detic cropped dataset, with image resized to 56x56, barely got to a score of 0.1 . Then I tried vast ai GPU providers to rent out 4x3090s to train on 400x400 images, same procedure, score now jumped to 0.4. But it did cost me a lot, like 50 dollars. It is going to cost a lot more if I have to experiment with various architectures, methodologies etc. \nSo here's my question, are some kaggle competitions, inherently this costly? or am I missing something here? Thanks.",
    "1725899": "It is unspoken truth that Kagglers with great hardware have significant advantage.\nYou can find heated discussions about this if you dig deep in topics.\nNevertheless, keep up the good work!",
    "1725906": "Thanks @vivovinco",
    "1725933": "then  just learn the models and pipelines.",
    "1725946": "1. Use the gradient accumulation. You can find a really brief tutorial [here](https://kozodoi.me/python/deep%20learning/pytorch/tutorial/2021/02/19/gradient-accumulation.html).\n2. Reduce memory usage by computing in FP16. If you use PyTorch as a base framework, here is the [documentation](https://pytorch.org/blog/accelerating-training-on-nvidia-gpus-with-pytorch-automatic-mixed-precision/). It is also important here to choose architecture, which is very beneficial in FP16 training\n\nI believe you can get my current score using 2080, far from the top but something. I use one GPU with 12Gb of memory",
    "1725972": "Please refer or point to something, that will be much more helpful, rather than throw around generic terms.",
    "1725987": "Thanks that was helpful, while I was already aware of them, did not think they would give enough results, like 4x24 GB (64 batch size) cannot be down-scaled to 8GB suddenly, even with these. Mind telling about your batch size and image res on 12 GB GPU? Thanks.\nOne more thing is we cannot directly 'choose' architectures right? We gotta experiment to conclude on them concretely.",
    "1725992": "I use colab TPU. It can be a remedy for you.\ncolab pro is just $10",
    "1725999": "Can I ask small ques here that when FP16 help the model train faster, is FP16 worse than FP32 so much? or do FP16 just reduce some small accuracy compare to FP32?",
    "1726000": "Mind telling most ram, time of training and gpu memorys you have extracted from pro? Thanks",
    "1726009": "dqhdqmcttdqx For most recent RTX GPUs, you should see a nice speedup and minimal effect on convergence. I had written a blog on how this works, if you're interested, [here's](https://hackernoon.com/rtx-2080ti-vs-gtx-1080ti-fastai-mixed-precision-training-comparisons-on-cifar-100-761d8f615d7f) the link",
    "1726105": "I've been using only kaggle TPU. You just need to be very cautious with the amount of time used with TPU. I have been training EffNetB5 with 224 img size which means I can do over 20 experiments per week. Then upgrade to a better model and img size when you are ready but beware that each model will take much longer to train (2x, 3x, maybe longer). I modified this [script for training and inference](https://www.kaggle.com/code/dragonzhang/happywhale-effnet-b7-fork-with-detic-crop). You can definitely get a silver medal and maybe a gold medal with only kaggle TPU.",
    "1727941": "Just tried fp16 pytorch, was able to come from batchsize 2 to 4 for 600x600 image res.",
    "1728140": "init27 Thanks for sharing 💯",
    "1728195": "In my case, i when i'm using tpu, you can train efficient net b6 with image size 512 and have up to batch size of 80 (10 images * 8 core TPUs), and time for each epochs about 9 minutes. \nfor specs:\n- ram: 13gb available. (you can change to high ram->35gb) \n- disk: 256gb available (when you use gpu it's change to 167gb)",
    "1728233": "No question that vision competitions need more compute than tabular.  And generally higher resolution images model better than low.  More compute can mean more money - or more training time.  More compute will also mean your skill set needs to be better.\n\nAs noted by several, you can have success using kaggle TPU/GPU.\n\n- On local PC's I have never seen significant loss in performance by using fp16 - pretty much allows for doubling of batch size which reduces training time.\n\n- Image sizes of around 224 have been good for experimentation  (larger image sizes I use the kaggle TPU).  64x64 has almost always been a waste of time.\n\n- Training times of around 16 hours or less have been good for experimentation.\n\n- I create a large swap file on Ubuntu to avoid any cpu memory issues (200GB on an ssd works nicely).\n\n- I have dual GPU's on all 4 of my machines, but GPU prices these days are crazy. \n\n- Don't waste TPU quota doing cpu level work - do pipeline work in separate cpu/gpu kernels to create data sets.\n\n- Don't get caught in the bs trap that folks with big hardware have a huge advantage - of course they do - turn on the news - the world is not fair.\n\nHave fun and learn",
    "1728401": "for the more in-depth guide refer to:\nhttps://docs.nvidia.com/deeplearning/performance/mixed-precision-training/index.html\n\nyou'll get the idea on why you need such thing as GradScaler",
    "1730105": "achilles38 can you tell how much time does training of your model take?",
    "1730247": "one model takes 3 hours to train",
    "1740537": "deepkim Great score considering you used only Colab TPU."
  },
  "source": "meta"
}