{
  "id": 40922,
  "title": "Multi-gpu error:the accuracy of one gpu is much lower than two gpu",
  "url": "/competitions/cdiscount-image-classification-challenge/discussion/40922",
  "author_name": "",
  "post_date": "2017-10-10T08:55:23.335893Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>When training with 2 GPUs, it always says segmentation fault (core dumped), and the accuracy of train is much lower than one GPU. But when training with one GPU , no error, and the accuracy is higher.</p>\n\n<p>This is more details about my error:\n<a href=\"https://discuss.pytorch.org/t/multi-gpu-error-the-accuracy-of-one-gpu-is-much-lower-than-two-gpus/8492\">https://discuss.pytorch.org/t/multi-gpu-error-the-accuracy-of-one-gpu-is-much-lower-than-two-gpus/8492</a></p>\n\n<p>Hoping anyone to give some advice.</p>\n\n<p>Thanks!</p>",
  "messages": [
    {
      "id": "229724",
      "postDate": "10/10/2017 08:55:23",
      "content": "<p>When training with 2 GPUs, it always says segmentation fault (core dumped), and the accuracy of train is much lower than one GPU. But when training with one GPU , no error, and the accuracy is higher.</p>\n\n<p>This is more details about my error:\n<a href=\"https://discuss.pytorch.org/t/multi-gpu-error-the-accuracy-of-one-gpu-is-much-lower-than-two-gpus/8492\">https://discuss.pytorch.org/t/multi-gpu-error-the-accuracy-of-one-gpu-is-much-lower-than-two-gpus/8492</a></p>\n\n<p>Hoping anyone to give some advice.</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "When training with 2 GPUs, it always says segmentation fault (core dumped), and the accuracy of train is much lower than one GPU. But when training with one GPU , no error, and the accuracy is higher.\n\nThis is more details about my error:\nhttps://discuss.pytorch.org/t/multi-gpu-error-the-accuracy-of-one-gpu-is-much-lower-than-two-gpus/8492\n\nHoping anyone to give some advice.\n\nThanks!",
      "votes": null
    },
    {
      "id": "231985",
      "postDate": "10/16/2017 15:06:55",
      "content": "<p>this is the code i used for multi-gpu training in pytorch</p>",
      "rawMarkdown": "this is the code i used for multi-gpu training in pytorch",
      "votes": null
    },
    {
      "id": "232005",
      "postDate": "10/16/2017 15:34:23",
      "content": "<p>Heng, why do you apply nn.DataParallel at every iteration? Doesn't it make more sense to use it once at the beginning? <code>net = nn.DataParallel(net)</code></p>",
      "rawMarkdown": "Heng, why do you apply nn.DataParallel at every iteration? Doesn't it make more sense to use it once at the beginning? `net = nn.DataParallel(net)`",
      "votes": null
    },
    {
      "id": "234090",
      "postDate": "10/22/2017 09:16:49",
      "content": "<p>you are correct.</p>\n\n<p>you can put \"net = nn.DataParallel(net)\" at the beginning and use once. It is much faster. But do take care of the saving:</p>\n\n<pre><code>        if i in iter_save: \n            #https://discuss.pytorch.org/t/dataparallel-optim-and-saving-correctness/4054 \n            if NUM_CUDA_DEVICES != 1: \n                torch.save(net.module.state_dict(),out_dir +'/checkpoint/%08d_model.pth'%(i))\n            else:\n                torch.save(net.state_dict(),out_dir +'/checkpoint/%08d_model.pth'%(i))\n\n            torch.save({\n                'optimizer': optimizer.state_dict(),\n                'iter'     : i,\n                'epoch'    : epoch,\n            }, out_dir +'/checkpoint/%08d_optimizer.pth'%(i))\n</code></pre>",
      "rawMarkdown": "you are correct.\n\nyou can put \"net = nn.DataParallel(net)\" at the beginning and use once. It is much faster. But do take care of the saving:\n\n           \n            if i in iter_save: \n\t\t\t\t#https://discuss.pytorch.org/t/dataparallel-optim-and-saving-correctness/4054 \n\t\t\t\tif NUM_CUDA_DEVICES != 1: \n                \ttorch.save(net.module.state_dict(),out_dir +'/checkpoint/%08d_model.pth'%(i))\n\t\t\t\telse:\n                \ttorch.save(net.state_dict(),out_dir +'/checkpoint/%08d_model.pth'%(i))\n\n                torch.save({\n                    'optimizer': optimizer.state_dict(),\n                    'iter'     : i,\n                    'epoch'    : epoch,\n                }, out_dir +'/checkpoint/%08d_optimizer.pth'%(i))",
      "votes": null
    },
    {
      "id": "234100",
      "postDate": "10/22/2017 09:58:06",
      "content": "<p>Thanks you for the tip on saving!</p>",
      "rawMarkdown": "Thanks you for the tip on saving!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 231985,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "10/16/2017 15:06:55",
      "content": "<p>this is the code i used for multi-gpu training in pytorch</p>",
      "votes": null,
      "replies": [
        {
          "id": 232005,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "10/16/2017 15:34:23",
          "content": "<p>Heng, why do you apply nn.DataParallel at every iteration? Doesn't it make more sense to use it once at the beginning? <code>net = nn.DataParallel(net)</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 234090,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/22/2017 09:16:49",
          "content": "<p>you are correct.</p>\n\n<p>you can put \"net = nn.DataParallel(net)\" at the beginning and use once. It is much faster. But do take care of the saving:</p>\n\n<pre><code>        if i in iter_save: \n            #https://discuss.pytorch.org/t/dataparallel-optim-and-saving-correctness/4054 \n            if NUM_CUDA_DEVICES != 1: \n                torch.save(net.module.state_dict(),out_dir +'/checkpoint/%08d_model.pth'%(i))\n            else:\n                torch.save(net.state_dict(),out_dir +'/checkpoint/%08d_model.pth'%(i))\n\n            torch.save({\n                'optimizer': optimizer.state_dict(),\n                'iter'     : i,\n                'epoch'    : epoch,\n            }, out_dir +'/checkpoint/%08d_optimizer.pth'%(i))\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 234100,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "10/22/2017 09:58:06",
          "content": "<p>Thanks you for the tip on saving!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "229724": "When training with 2 GPUs, it always says segmentation fault (core dumped), and the accuracy of train is much lower than one GPU. But when training with one GPU , no error, and the accuracy is higher.\n\nThis is more details about my error:\nhttps://discuss.pytorch.org/t/multi-gpu-error-the-accuracy-of-one-gpu-is-much-lower-than-two-gpus/8492\n\nHoping anyone to give some advice.\n\nThanks!",
    "231985": "this is the code i used for multi-gpu training in pytorch",
    "232005": "Heng, why do you apply nn.DataParallel at every iteration? Doesn't it make more sense to use it once at the beginning? `net = nn.DataParallel(net)`",
    "234090": "you are correct.\n\nyou can put \"net = nn.DataParallel(net)\" at the beginning and use once. It is much faster. But do take care of the saving:\n\n           \n            if i in iter_save: \n\t\t\t\t#https://discuss.pytorch.org/t/dataparallel-optim-and-saving-correctness/4054 \n\t\t\t\tif NUM_CUDA_DEVICES != 1: \n                \ttorch.save(net.module.state_dict(),out_dir +'/checkpoint/%08d_model.pth'%(i))\n\t\t\t\telse:\n                \ttorch.save(net.state_dict(),out_dir +'/checkpoint/%08d_model.pth'%(i))\n\n                torch.save({\n                    'optimizer': optimizer.state_dict(),\n                    'iter'     : i,\n                    'epoch'    : epoch,\n                }, out_dir +'/checkpoint/%08d_optimizer.pth'%(i))",
    "234100": "Thanks you for the tip on saving!"
  },
  "source": "meta"
}