{
  "id": 224198,
  "title": "Multi-GPU training issue",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/224198",
  "author_name": "",
  "post_date": "2021-03-07T10:12:35.253653800Z",
  "votes": 3,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Has anyone had this issue while using multiple GPUs to train a pytorch model?<br>\nI trained a pytorch model by using single GPU with batchsize=8, which took about 11G GPU memory. Then I added lines below to my code to train model on multi-GPUs, WITHOUT changing the batchsize, two GPUs were used but both of them took almost 11G memory as just use a single GPU. </p>\n<p><strong>device = torch.device(\"cuda\")</strong></p>\n<p><strong>model =Mymodel()</strong></p>\n<p><strong>model = nn.DataParallel(model)</strong></p>\n<p><strong>model.to(device)</strong></p>\n<p>And I did check the batchsize was split to 4 on each. This is strange, I expected the memory usage of both GPU would be half of using single one.</p>\n<p>Any clues of how to solve this issue?<br>\nI am using Pytorch 1.6.0, Nvidia Cuda: 10.1</p>",
  "messages": [
    {
      "id": "1229399",
      "postDate": "03/07/2021 10:12:35",
      "content": "<p>Has anyone had this issue while using multiple GPUs to train a pytorch model?<br>\nI trained a pytorch model by using single GPU with batchsize=8, which took about 11G GPU memory. Then I added lines below to my code to train model on multi-GPUs, WITHOUT changing the batchsize, two GPUs were used but both of them took almost 11G memory as just use a single GPU. </p>\n<p><strong>device = torch.device(\"cuda\")</strong></p>\n<p><strong>model =Mymodel()</strong></p>\n<p><strong>model = nn.DataParallel(model)</strong></p>\n<p><strong>model.to(device)</strong></p>\n<p>And I did check the batchsize was split to 4 on each. This is strange, I expected the memory usage of both GPU would be half of using single one.</p>\n<p>Any clues of how to solve this issue?<br>\nI am using Pytorch 1.6.0, Nvidia Cuda: 10.1</p>",
      "rawMarkdown": "Has anyone had this issue while using multiple GPUs to train a pytorch model?\nI trained a pytorch model by using single GPU with batchsize=8, which took about 11G GPU memory. Then I added lines below to my code to train model on multi-GPUs, WITHOUT changing the batchsize, two GPUs were used but both of them took almost 11G memory as just use a single GPU. \n\n**device = torch.device(\"cuda\")**\n\n**model =Mymodel()**\n\n**model = nn.DataParallel(model)**\n\n**model.to(device)**\n\nAnd I did check the batchsize was split to 4 on each. This is strange, I expected the memory usage of both GPU would be half of using single one.\n\nAny clues of how to solve this issue?\nI am using Pytorch 1.6.0, Nvidia Cuda: 10.1",
      "votes": null
    },
    {
      "id": "1229453",
      "postDate": "03/07/2021 11:10:54",
      "content": "<p>Probably, it is not autocasted. </p>\n<p>The following code could fix the issue.</p>\n<pre><code>MyModel(nn.Module):\n    ...\n    @autocast()\n    def forward(self, input):\n       ...\n</code></pre>\n<p>from<br>\n<a href=\"https://pytorch.org/docs/stable/notes/amp_examples.html\" target=\"_blank\">https://pytorch.org/docs/stable/notes/amp_examples.html</a></p>",
      "rawMarkdown": "Probably, it is not autocasted. \n\nThe following code could fix the issue.\n```\nMyModel(nn.Module):\n    ...\n    @autocast()\n    def forward(self, input):\n       ...\n\n```\nfrom\nhttps://pytorch.org/docs/stable/notes/amp_examples.html",
      "votes": null
    },
    {
      "id": "1229465",
      "postDate": "03/07/2021 11:20:42",
      "content": "<p>Thanks, do you mean add this line **@autocast() **<br>\nto Mymodel?</p>",
      "rawMarkdown": "Thanks, do you mean add this line **@autocast() **\nto Mymodel?",
      "votes": null
    },
    {
      "id": "1229518",
      "postDate": "03/07/2021 12:01:40",
      "content": "<p>Yes.<br>\n<code>The issues described here only affect autocast. GradScaler’s usage is unchanged.</code> <br>\nfrom pytorch official.</p>",
      "rawMarkdown": "Yes.\n`The issues described here only affect autocast. GradScaler’s usage is unchanged.` \nfrom pytorch official.",
      "votes": null
    },
    {
      "id": "1229921",
      "postDate": "03/07/2021 17:34:30",
      "content": "<p>Supposedly, PyTorch Lightning makes this sort of thing (incl. TPU use) easier. Have never tried it with multiple GPUs though.</p>",
      "rawMarkdown": "Supposedly, PyTorch Lightning makes this sort of thing (incl. TPU use) easier. Have never tried it with multiple GPUs though.",
      "votes": null
    },
    {
      "id": "1230302",
      "postDate": "03/08/2021 02:08:13",
      "content": "<p>what he said is to use mix-precision, it will save gpu ram, and also you can try distributed trianing, it will save gpu ram and accelerate training speed. <a href=\"https://www.kaggle.com/lftuwujie\" target=\"_blank\">@lftuwujie</a> </p>",
      "rawMarkdown": "what he said is to use mix-precision, it will save gpu ram, and also you can try distributed trianing, it will save gpu ram and accelerate training speed. @lftuwujie",
      "votes": null
    },
    {
      "id": "1230303",
      "postDate": "03/08/2021 02:10:06",
      "content": "<p>dataparallel,容易造成显存不均衡（dataparallel 很多时候并行bs &lt; 单卡训练bs*卡数），加速效果也有限</p>",
      "rawMarkdown": "dataparallel,容易造成显存不均衡（dataparallel 很多时候并行bs < 单卡训练bs*卡数），加速效果也有限",
      "votes": null
    },
    {
      "id": "1230877",
      "postDate": "03/08/2021 14:26:07",
      "content": "<p>Yes, that's true. But my problem is indeed what <a href=\"https://www.kaggle.com/tmhrkt\" target=\"_blank\">@tmhrkt</a> said.</p>",
      "rawMarkdown": "Yes, that's true. But my problem is indeed what @tmhrkt said.",
      "votes": null
    },
    {
      "id": "1231228",
      "postDate": "03/08/2021 19:00:45",
      "content": "<p>I recall when using nn.DataParallel, your model first need to be on cuda:<br>\n<code>model = nn.DataParallel(model.cuda())</code> in order for DP to work correct.<br>\nAlso, consider using Distributed Data Parallel, which is faster than DP and loads multiple GPUs more evenly.</p>",
      "rawMarkdown": "I recall when using nn.DataParallel, your model first need to be on cuda:\n`model = nn.DataParallel(model.cuda())` in order for DP to work correct.\nAlso, consider using Distributed Data Parallel, which is faster than DP and loads multiple GPUs more evenly.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1229453,
      "author_name": "tmhrkt",
      "author_url": "",
      "post_date": "03/07/2021 11:10:54",
      "content": "<p>Probably, it is not autocasted. </p>\n<p>The following code could fix the issue.</p>\n<pre><code>MyModel(nn.Module):\n    ...\n    @autocast()\n    def forward(self, input):\n       ...\n</code></pre>\n<p>from<br>\n<a href=\"https://pytorch.org/docs/stable/notes/amp_examples.html\" target=\"_blank\">https://pytorch.org/docs/stable/notes/amp_examples.html</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1229465,
          "author_name": "lftuwujie",
          "author_url": "",
          "post_date": "03/07/2021 11:20:42",
          "content": "<p>Thanks, do you mean add this line **@autocast() **<br>\nto Mymodel?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1229518,
          "author_name": "tmhrkt",
          "author_url": "",
          "post_date": "03/07/2021 12:01:40",
          "content": "<p>Yes.<br>\n<code>The issues described here only affect autocast. GradScaler’s usage is unchanged.</code> <br>\nfrom pytorch official.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1230302,
          "author_name": "cswwp347724",
          "author_url": "",
          "post_date": "03/08/2021 02:08:13",
          "content": "<p>what he said is to use mix-precision, it will save gpu ram, and also you can try distributed trianing, it will save gpu ram and accelerate training speed. <a href=\"https://www.kaggle.com/lftuwujie\" target=\"_blank\">@lftuwujie</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1229921,
      "author_name": "bjoernholzhauer",
      "author_url": "",
      "post_date": "03/07/2021 17:34:30",
      "content": "<p>Supposedly, PyTorch Lightning makes this sort of thing (incl. TPU use) easier. Have never tried it with multiple GPUs though.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1230303,
      "author_name": "cswwp347724",
      "author_url": "",
      "post_date": "03/08/2021 02:10:06",
      "content": "<p>dataparallel,容易造成显存不均衡（dataparallel 很多时候并行bs &lt; 单卡训练bs*卡数），加速效果也有限</p>",
      "votes": null,
      "replies": [
        {
          "id": 1230877,
          "author_name": "lftuwujie",
          "author_url": "",
          "post_date": "03/08/2021 14:26:07",
          "content": "<p>Yes, that's true. But my problem is indeed what <a href=\"https://www.kaggle.com/tmhrkt\" target=\"_blank\">@tmhrkt</a> said.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1231228,
      "author_name": "bloodaxe",
      "author_url": "",
      "post_date": "03/08/2021 19:00:45",
      "content": "<p>I recall when using nn.DataParallel, your model first need to be on cuda:<br>\n<code>model = nn.DataParallel(model.cuda())</code> in order for DP to work correct.<br>\nAlso, consider using Distributed Data Parallel, which is faster than DP and loads multiple GPUs more evenly.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1229399": "Has anyone had this issue while using multiple GPUs to train a pytorch model?\nI trained a pytorch model by using single GPU with batchsize=8, which took about 11G GPU memory. Then I added lines below to my code to train model on multi-GPUs, WITHOUT changing the batchsize, two GPUs were used but both of them took almost 11G memory as just use a single GPU. \n\n**device = torch.device(\"cuda\")**\n\n**model =Mymodel()**\n\n**model = nn.DataParallel(model)**\n\n**model.to(device)**\n\nAnd I did check the batchsize was split to 4 on each. This is strange, I expected the memory usage of both GPU would be half of using single one.\n\nAny clues of how to solve this issue?\nI am using Pytorch 1.6.0, Nvidia Cuda: 10.1",
    "1229453": "Probably, it is not autocasted. \n\nThe following code could fix the issue.\n```\nMyModel(nn.Module):\n    ...\n    @autocast()\n    def forward(self, input):\n       ...\n\n```\nfrom\nhttps://pytorch.org/docs/stable/notes/amp_examples.html",
    "1229465": "Thanks, do you mean add this line **@autocast() **\nto Mymodel?",
    "1229518": "Yes.\n`The issues described here only affect autocast. GradScaler’s usage is unchanged.` \nfrom pytorch official.",
    "1229921": "Supposedly, PyTorch Lightning makes this sort of thing (incl. TPU use) easier. Have never tried it with multiple GPUs though.",
    "1230302": "what he said is to use mix-precision, it will save gpu ram, and also you can try distributed trianing, it will save gpu ram and accelerate training speed. @lftuwujie",
    "1230303": "dataparallel,容易造成显存不均衡（dataparallel 很多时候并行bs < 单卡训练bs*卡数），加速效果也有限",
    "1230877": "Yes, that's true. But my problem is indeed what @tmhrkt said.",
    "1231228": "I recall when using nn.DataParallel, your model first need to be on cuda:\n`model = nn.DataParallel(model.cuda())` in order for DP to work correct.\nAlso, consider using Distributed Data Parallel, which is faster than DP and loads multiple GPUs more evenly."
  },
  "source": "meta"
}