{
  "id": 241383,
  "title": "Multi-GPU (Model parallelism) for EfficientNet-B0",
  "url": "/competitions/seti-breakthrough-listen/discussion/241383",
  "author_name": "John Clarke",
  "post_date": "2021-05-24T10:23:12.215000",
  "votes": 0,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I am trying to partition a NN based on EfficientNet B0 across multiple NVidia K80s. In the code below I am using 2 K80's, but still running out of GPU memory on GPU 0. The machine has 8 K80's attached so perhaps even finer parallelism would work.</p>\n<p>I'm still learning the nitty gritty details of pytorch and EfficientNet so any help or pointers would be really appreciated. (I have also reached out to Luke M-L, the efficientnet-pytorch author, so hopefully he will respond.)</p>\n<p>Below is my best effort so far. I think I need to drill down into efficientnet itself to use more GPUs, but where to start?</p>\n<p>Thanks!</p>\n<blockquote>\n  <p>baseline_name = 'efficientnet-b0'<br>\n     pretrained_model = {<br>\n         'efficientnet-b0': '../input/efficientnet-pytorch/efficientnet-b0-08094119.pth',<br>\n         'efficientnet-b1': '../input/efficientnet-pytorch/efficientnet-b1-dbc7070a.pth',<br>\n         'efficientnet-b2': '../input/efficientnet-pytorch/efficientnet-b2-27687264.pth',<br>\n         'efficientnet-b3': '../input/efficientnet-pytorch/efficientnet-b3-c8376fa2.pth',<br>\n         'efficientnet-b4': '../input/efficientnet-pytorch/efficientnet-b4-e116e8b3.pth',<br>\n         'efficientnet-b5': '../input/efficientnet-pytorch/efficientnet-b5-586e6cc6.pth',<br>\n         'efficientnet-b6': '../input/efficientnet-pytorch/efficientnet-b6-c76e70fd.pth',<br>\n         'efficientnet-b7': '../input/efficientnet-pytorch/efficientnet-b7-dcc49843.pth'<br>\n     } # pretrained_model</p>\n  <p>class enetv2(nn.Module):<br>\n         def <strong>init</strong>(self, backbone, out_dim):<br>\n             super(enetv2, self).<strong>init</strong>()<br>\n             self.enet = enet.EfficientNet.from_name(backbone)<br>\n             # self.enet.load_state_dict(torch.load(pretrained_model[backbone]))<br>\n             self.enet.load_state_dict(torch.load(pretrained_model[backbone], map_location=\"cpu\"))<br>\n             self.myfc = nn.Linear(self.enet._fc.in_features, out_dim).to('cuda:0')<br>\n             self.enet._fc = nn.Identity()<br>\n             self.conv1 = nn.Conv2d(1, 3, kernel_size=3, stride=1, padding=3, bias=False).to('cuda:1') # in_channels, out_channels, kernel_size</p>\n</blockquote>\n<pre><code>def extract(self, x):\n    return self.enet(x)\n\ndef forward(self, x):\n    #x = self.conv1(x)\n    #x = self.extract(x)\n    #x = self.myfc(x)\n    x = self.extract(self.conv1(x.to('cuda:1')))\n    return self.myfc(x.to('cuda:0'))\n</code></pre>",
  "messages": [
    {
      "id": 1323742,
      "postDate": "2021-05-26T12:24:26.250Z",
      "content": "<p>Using pytorch Lightning or Fastai frameworks along with <code>nn.DataParallel(model, device_ids=[0,1,2,...])</code> works well (based on my experience) </p>",
      "rawMarkdown": "Using pytorch Lightning or Fastai frameworks along with `nn.DataParallel(model, device_ids=[0,1,2,...])` works well (based on my experience) ",
      "votes": 1
    },
    {
      "id": 1332115,
      "postDate": "2021-06-01T23:13:24.663Z",
      "content": "<p>Hi Everyone,</p>\n<p>Just though I'd close the loop on this. I took the Pytorch Lightning recommendations and spent several sessions working through it. The model now runs on multi GPU. Luckily I did not have to go model parallel, instead reducing the batch size was enough to get it to fit. So thanks so much for your guidance.</p>\n<p>Sadly the denoising algorithm I'm trying doesn't seem to help on the larger images and so is not competitive. I'll keep trying a bit longer because I think denoising will be useful in its own right. But at some point I'll give up and move on to data augmentation.</p>\n<p>For those who are interested here's the model code I arrived at. I'm still not 100% certain that <code>test_epoch_end()</code> and <code>validation_epoch_end()</code> are quite correct. They run on every GPU so is it possible there are race conditions? I.e. could the <code>self.all_gather()</code> functions pull in data from other GPUs before it's ready?</p>\n<p>`baseline_name = 'efficientnet-b0'<br>\npretrained_model = {<br>\n    'efficientnet-b0': '../input/efficientnet-pytorch/efficientnet-b0-08094119.pth',<br>\n    'efficientnet-b1': '../input/efficientnet-pytorch/efficientnet-b1-dbc7070a.pth',<br>\n    'efficientnet-b2': '../input/efficientnet-pytorch/efficientnet-b2-27687264.pth',<br>\n    'efficientnet-b3': '../input/efficientnet-pytorch/efficientnet-b3-c8376fa2.pth',<br>\n    'efficientnet-b4': '../input/efficientnet-pytorch/efficientnet-b4-e116e8b3.pth',<br>\n    'efficientnet-b5': '../input/efficientnet-pytorch/efficientnet-b5-586e6cc6.pth',<br>\n    'efficientnet-b6': '../input/efficientnet-pytorch/efficientnet-b6-c76e70fd.pth',<br>\n    'efficientnet-b7': '../input/efficientnet-pytorch/efficientnet-b7-dcc49843.pth'<br>\n} # pretrained_model</p>\n<p>class lit_enetv2(pl.LightningModule):<br>\n    def <strong>init</strong>(self, backbone):<br>\n        super(lit_enetv2, self).<strong>init</strong>()<br>\n        self.enet = enet.EfficientNet.from_name(backbone)<br>\n        self.enet.load_state_dict(torch.load(pretrained_model[backbone]))<br>\n        self.myfc = nn.Linear(self.enet._fc.in_features, 1) # out_dim = 1<br>\n        self.enet._fc = nn.Identity()<br>\n        self.conv1 = nn.Conv2d(1, 3, kernel_size=3, stride=1, padding=3, bias=False) # in_channels, out_channels, kernel_size</p>\n<pre><code>    self.classifier = nn.Sequential(self.conv1, self.enet, self.myfc)\n\ndef forward(self, x):\n    return torch.squeeze(self.classifier(x), 1)\n\ndef step(self, step_type, batch, batch_idx):\n    x = batch['image']\n    y = batch['targets']\n    y_hat = self(x) # self.classifier(x)\n    loss = F.binary_cross_entropy_with_logits(y_hat, y) # training criterion\n    self.log(step_type + '_loss', loss, on_step=False, on_epoch=True, prog_bar=True, logger=True, sync_dist=(step_type != 'train'))\n    return loss, y_hat\n\ndef training_step(self, batch, batch_idx):\n    loss, _ = self.step('train', batch, batch_idx)\n    return loss\n\ndef validation_step(self, batch, batch_idx):\n    _, y_hat = self.step('xval', batch, batch_idx)\n    return {\n        'y_hat': y_hat,\n        'y': batch['targets'],\n    } # return\n\ndef validation_epoch_end(self, validation_step_outputs):\n    validation_step_outputs = self.all_gather(validation_step_outputs)\n\n    y_hats = []\n    ys = []\n    for vso in validation_step_outputs:\n        yh = vso['y_hat'].reshape(-1,)\n        y_hats.append(yh)\n        y = vso['y'].reshape(-1,)\n        ys.append(y)\n\n    y_hats = torch.cat(y_hats)\n    ys  = torch.cat(ys)\n    auroc = metrics.roc_auc_score(ys.cpu(), y_hats.cpu())\n    self.log('xval_auroc', auroc, on_step=False, on_epoch=True, prog_bar=True, logger=True, sync_dist=True)\n\ndef test_step(self, batch, batch_idx):\n    _, y_hat = self.step('xval', batch, batch_idx)\n    return {\n        'y_hat': y_hat,\n        'items': batch['item'],\n    } # return\n\ndef test_epoch_end(self, outputs):\n    outputs = self.all_gather(outputs)\n\n    y_hats = []\n    items = []\n    for op in outputs:\n        yh = op['y_hat'].reshape(-1,)\n        y_hats.append(yh)\n        im = op['items'].reshape(-1,)\n        items.append(im)\n\n    y_hats = torch.cat(y_hats)\n    items  = torch.cat(items)\n\n    y_hats_sorted = y_hats.clone().detach()\n    for i, ii in enumerate(items):\n        y_hats_sorted[ii] = y_hats[i]\n    self.test_output = y_hats_sorted\n\ndef configure_optimizers(self):\n    optimizer = torch.optim.Adam(self.parameters(), lr=CFG.learning_rate)\n    return optimizer\n</code></pre>\n<p>`</p>",
      "rawMarkdown": "Hi Everyone,\n\nJust though I'd close the loop on this. I took the Pytorch Lightning recommendations and spent several sessions working through it. The model now runs on multi GPU. Luckily I did not have to go model parallel, instead reducing the batch size was enough to get it to fit. So thanks so much for your guidance.\n\nSadly the denoising algorithm I'm trying doesn't seem to help on the larger images and so is not competitive. I'll keep trying a bit longer because I think denoising will be useful in its own right. But at some point I'll give up and move on to data augmentation.\n\nFor those who are interested here's the model code I arrived at. I'm still not 100% certain that `test_epoch_end()` and `validation_epoch_end()` are quite correct. They run on every GPU so is it possible there are race conditions? I.e. could the `self.all_gather()` functions pull in data from other GPUs before it's ready?\n\n`baseline_name = 'efficientnet-b0'\npretrained_model = {\n    'efficientnet-b0': '../input/efficientnet-pytorch/efficientnet-b0-08094119.pth',\n    'efficientnet-b1': '../input/efficientnet-pytorch/efficientnet-b1-dbc7070a.pth',\n    'efficientnet-b2': '../input/efficientnet-pytorch/efficientnet-b2-27687264.pth',\n    'efficientnet-b3': '../input/efficientnet-pytorch/efficientnet-b3-c8376fa2.pth',\n    'efficientnet-b4': '../input/efficientnet-pytorch/efficientnet-b4-e116e8b3.pth',\n    'efficientnet-b5': '../input/efficientnet-pytorch/efficientnet-b5-586e6cc6.pth',\n    'efficientnet-b6': '../input/efficientnet-pytorch/efficientnet-b6-c76e70fd.pth',\n    'efficientnet-b7': '../input/efficientnet-pytorch/efficientnet-b7-dcc49843.pth'\n} # pretrained_model\n\nclass lit_enetv2(pl.LightningModule):\n    def __init__(self, backbone):\n        super(lit_enetv2, self).__init__()\n        self.enet = enet.EfficientNet.from_name(backbone)\n        self.enet.load_state_dict(torch.load(pretrained_model[backbone]))\n        self.myfc = nn.Linear(self.enet._fc.in_features, 1) # out_dim = 1\n        self.enet._fc = nn.Identity()\n        self.conv1 = nn.Conv2d(1, 3, kernel_size=3, stride=1, padding=3, bias=False) # in_channels, out_channels, kernel_size\n\n        self.classifier = nn.Sequential(self.conv1, self.enet, self.myfc)\n\n    def forward(self, x):\n        return torch.squeeze(self.classifier(x), 1)\n\n    def step(self, step_type, batch, batch_idx):\n        x = batch['image']\n        y = batch['targets']\n        y_hat = self(x) # self.classifier(x)\n        loss = F.binary_cross_entropy_with_logits(y_hat, y) # training criterion\n        self.log(step_type + '_loss', loss, on_step=False, on_epoch=True, prog_bar=True, logger=True, sync_dist=(step_type != 'train'))\n        return loss, y_hat\n\n    def training_step(self, batch, batch_idx):\n        loss, _ = self.step('train', batch, batch_idx)\n        return loss\n\n    def validation_step(self, batch, batch_idx):\n        _, y_hat = self.step('xval', batch, batch_idx)\n        return {\n            'y_hat': y_hat,\n            'y': batch['targets'],\n        } # return\n\n    def validation_epoch_end(self, validation_step_outputs):\n        validation_step_outputs = self.all_gather(validation_step_outputs)\n\n        y_hats = []\n        ys = []\n        for vso in validation_step_outputs:\n            yh = vso['y_hat'].reshape(-1,)\n            y_hats.append(yh)\n            y = vso['y'].reshape(-1,)\n            ys.append(y)\n\n        y_hats = torch.cat(y_hats)\n        ys  = torch.cat(ys)\n        auroc = metrics.roc_auc_score(ys.cpu(), y_hats.cpu())\n        self.log('xval_auroc', auroc, on_step=False, on_epoch=True, prog_bar=True, logger=True, sync_dist=True)\n\n    def test_step(self, batch, batch_idx):\n        _, y_hat = self.step('xval', batch, batch_idx)\n        return {\n            'y_hat': y_hat,\n            'items': batch['item'],\n        } # return\n\n    def test_epoch_end(self, outputs):\n        outputs = self.all_gather(outputs)\n\n        y_hats = []\n        items = []\n        for op in outputs:\n            yh = op['y_hat'].reshape(-1,)\n            y_hats.append(yh)\n            im = op['items'].reshape(-1,)\n            items.append(im)\n\n        y_hats = torch.cat(y_hats)\n        items  = torch.cat(items)\n        \n        y_hats_sorted = y_hats.clone().detach()\n        for i, ii in enumerate(items):\n            y_hats_sorted[ii] = y_hats[i]\n        self.test_output = y_hats_sorted\n\n    def configure_optimizers(self):\n        optimizer = torch.optim.Adam(self.parameters(), lr=CFG.learning_rate)\n        return optimizer\n`"
    },
    {
      "id": 1320805,
      "postDate": "2021-05-24T10:35:09.093Z",
      "content": "<p>Edit: sorry I realise I misunderstood your intention. You don't wish to use DataParallel, but instead put different stages of 1 model across multiple GPUs?</p>",
      "rawMarkdown": "Edit: sorry I realise I misunderstood your intention. You don't wish to use DataParallel, but instead put different stages of 1 model across multiple GPUs?",
      "replies": [
        {
          "id": 1320917,
          "postDate": "2021-05-24T12:16:54.587Z",
          "content": "<p>Exactly! </p>\n<p>That means drilling down into how EfficientNet is implemented - which is still a bit beyond my level of knowledge.</p>\n<p>Do you know if there's an easy way to display the structure of the NN model? I'm looking for the # channels, sizes, etc. at each level. I don't understand why the memory requirements have blown out so suddenly.</p>",
          "rawMarkdown": "Exactly! \n\nThat means drilling down into how EfficientNet is implemented - which is still a bit beyond my level of knowledge.\n\nDo you know if there's an easy way to display the structure of the NN model? I'm looking for the # channels, sizes, etc. at each level. I don't understand why the memory requirements have blown out so suddenly.\n"
        },
        {
          "id": 1321091,
          "postDate": "2021-05-24T14:11:21.643Z",
          "content": "<p>What you're doing is quite unusual. I advise you to take a look at the pytorch lightning library, it's very easy to use multiple gpus training. Your code will much more structured, concise and efficient.<br>\nUnless you do that for learning or whatever, the problem you face has been faced many times before by many people and efficient solutions have been found. Don't waste to much time and effort on such things imo.</p>",
          "rawMarkdown": "What you're doing is quite unusual. I advise you to take a look at the pytorch lightning library, it's very easy to use multiple gpus training. Your code will much more structured, concise and efficient.\nUnless you do that for learning or whatever, the problem you face has been faced many times before by many people and efficient solutions have been found. Don't waste to much time and effort on such things imo.",
          "votes": 2
        },
        {
          "id": 1321351,
          "postDate": "2021-05-24T16:32:32.933Z",
          "content": "<p>Agree with <a href=\"https://www.kaggle.com/rodolphelampe\" target=\"_blank\">Rodolphe</a> - been using dual GPU's on 4 different machines over the last two years.  Getting both GPU running successfully in Pytorch or Keras is often a bitch - setting up the right nvidia libraries, etc all make it a real chore.  Pytorch and tensorflow changes in recent times have made this a bit easier but still find myself doing lots of system refreshes when a major update to Pytorch or Tensorflow happens until I get out of Nvidia Hell.   </p>\n<p>Trying to learn a new way of using GPU's - that's going to be a huge bundle of pain I don't think you really want.</p>\n<p>Out of memory - use nvidia-smi in terminal to see the amount of memory use on both GPU's.  Decent chance that only one GPU loaded when you think two are in use.  Batch size being too big is 99% of the reasons I get memory issues on my dual GPU's, most often because I have left more than one Jupyter script running.</p>",
          "rawMarkdown": "Agree with [Rodolphe](https://www.kaggle.com/rodolphelampe) - been using dual GPU's on 4 different machines over the last two years.  Getting both GPU running successfully in Pytorch or Keras is often a bitch - setting up the right nvidia libraries, etc all make it a real chore.  Pytorch and tensorflow changes in recent times have made this a bit easier but still find myself doing lots of system refreshes when a major update to Pytorch or Tensorflow happens until I get out of Nvidia Hell.   \n\nTrying to learn a new way of using GPU's - that's going to be a huge bundle of pain I don't think you really want.\n\nOut of memory - use nvidia-smi in terminal to see the amount of memory use on both GPU's.  Decent chance that only one GPU loaded when you think two are in use.  Batch size being too big is 99% of the reasons I get memory issues on my dual GPU's, most often because I have left more than one Jupyter script running.",
          "votes": 2
        },
        {
          "id": 1321358,
          "postDate": "2021-05-24T16:40:13.453Z",
          "content": "<p>Thanks for that suggestion. I'll give pytorch-lightning a try right now. Will let you know how it goes.</p>",
          "rawMarkdown": "Thanks for that suggestion. I'll give pytorch-lightning a try right now. Will let you know how it goes."
        }
      ]
    },
    {
      "id": 1320794,
      "postDate": "2021-05-24T10:23:12.217Z",
      "content": "<p>I am trying to partition a NN based on EfficientNet B0 across multiple NVidia K80s. In the code below I am using 2 K80's, but still running out of GPU memory on GPU 0. The machine has 8 K80's attached so perhaps even finer parallelism would work.</p>\n<p>I'm still learning the nitty gritty details of pytorch and EfficientNet so any help or pointers would be really appreciated. (I have also reached out to Luke M-L, the efficientnet-pytorch author, so hopefully he will respond.)</p>\n<p>Below is my best effort so far. I think I need to drill down into efficientnet itself to use more GPUs, but where to start?</p>\n<p>Thanks!</p>\n<blockquote>\n  <p>baseline_name = 'efficientnet-b0'<br>\n     pretrained_model = {<br>\n         'efficientnet-b0': '../input/efficientnet-pytorch/efficientnet-b0-08094119.pth',<br>\n         'efficientnet-b1': '../input/efficientnet-pytorch/efficientnet-b1-dbc7070a.pth',<br>\n         'efficientnet-b2': '../input/efficientnet-pytorch/efficientnet-b2-27687264.pth',<br>\n         'efficientnet-b3': '../input/efficientnet-pytorch/efficientnet-b3-c8376fa2.pth',<br>\n         'efficientnet-b4': '../input/efficientnet-pytorch/efficientnet-b4-e116e8b3.pth',<br>\n         'efficientnet-b5': '../input/efficientnet-pytorch/efficientnet-b5-586e6cc6.pth',<br>\n         'efficientnet-b6': '../input/efficientnet-pytorch/efficientnet-b6-c76e70fd.pth',<br>\n         'efficientnet-b7': '../input/efficientnet-pytorch/efficientnet-b7-dcc49843.pth'<br>\n     } # pretrained_model</p>\n  <p>class enetv2(nn.Module):<br>\n         def <strong>init</strong>(self, backbone, out_dim):<br>\n             super(enetv2, self).<strong>init</strong>()<br>\n             self.enet = enet.EfficientNet.from_name(backbone)<br>\n             # self.enet.load_state_dict(torch.load(pretrained_model[backbone]))<br>\n             self.enet.load_state_dict(torch.load(pretrained_model[backbone], map_location=\"cpu\"))<br>\n             self.myfc = nn.Linear(self.enet._fc.in_features, out_dim).to('cuda:0')<br>\n             self.enet._fc = nn.Identity()<br>\n             self.conv1 = nn.Conv2d(1, 3, kernel_size=3, stride=1, padding=3, bias=False).to('cuda:1') # in_channels, out_channels, kernel_size</p>\n</blockquote>\n<pre><code>def extract(self, x):\n    return self.enet(x)\n\ndef forward(self, x):\n    #x = self.conv1(x)\n    #x = self.extract(x)\n    #x = self.myfc(x)\n    x = self.extract(self.conv1(x.to('cuda:1')))\n    return self.myfc(x.to('cuda:0'))\n</code></pre>",
      "rawMarkdown": "I am trying to partition a NN based on EfficientNet B0 across multiple NVidia K80s. In the code below I am using 2 K80's, but still running out of GPU memory on GPU 0. The machine has 8 K80's attached so perhaps even finer parallelism would work.\n\nI'm still learning the nitty gritty details of pytorch and EfficientNet so any help or pointers would be really appreciated. (I have also reached out to Luke M-L, the efficientnet-pytorch author, so hopefully he will respond.)\n\nBelow is my best effort so far. I think I need to drill down into efficientnet itself to use more GPUs, but where to start?\n\nThanks!\n\n\n>    baseline_name = 'efficientnet-b0'\n   pretrained_model = {\n       'efficientnet-b0': '../input/efficientnet-pytorch/efficientnet-b0-08094119.pth',\n       'efficientnet-b1': '../input/efficientnet-pytorch/efficientnet-b1-dbc7070a.pth',\n       'efficientnet-b2': '../input/efficientnet-pytorch/efficientnet-b2-27687264.pth',\n       'efficientnet-b3': '../input/efficientnet-pytorch/efficientnet-b3-c8376fa2.pth',\n       'efficientnet-b4': '../input/efficientnet-pytorch/efficientnet-b4-e116e8b3.pth',\n       'efficientnet-b5': '../input/efficientnet-pytorch/efficientnet-b5-586e6cc6.pth',\n       'efficientnet-b6': '../input/efficientnet-pytorch/efficientnet-b6-c76e70fd.pth',\n       'efficientnet-b7': '../input/efficientnet-pytorch/efficientnet-b7-dcc49843.pth'\n   } # pretrained_model\n\n   > class enetv2(nn.Module):\n       def __init__(self, backbone, out_dim):\n           super(enetv2, self).__init__()\n           self.enet = enet.EfficientNet.from_name(backbone)\n           # self.enet.load_state_dict(torch.load(pretrained_model[backbone]))\n           self.enet.load_state_dict(torch.load(pretrained_model[backbone], map_location=\"cpu\"))\n           self.myfc = nn.Linear(self.enet._fc.in_features, out_dim).to('cuda:0')\n           self.enet._fc = nn.Identity()\n           self.conv1 = nn.Conv2d(1, 3, kernel_size=3, stride=1, padding=3, bias=False).to('cuda:1') # in_channels, out_channels, kernel_size\n\n    def extract(self, x):\n        return self.enet(x)\n\n    def forward(self, x):\n        #x = self.conv1(x)\n        #x = self.extract(x)\n        #x = self.myfc(x)\n        x = self.extract(self.conv1(x.to('cuda:1')))\n        return self.myfc(x.to('cuda:0'))\n"
    }
  ],
  "comments": [
    {
      "id": 1323742,
      "author_name": "Ioannis M",
      "author_url": "",
      "post_date": "2021-05-26T12:24:26.250000",
      "content": "<p>Using pytorch Lightning or Fastai frameworks along with <code>nn.DataParallel(model, device_ids=[0,1,2,...])</code> works well (based on my experience) </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1332115,
      "author_name": "John Clarke",
      "author_url": "",
      "post_date": "2021-06-01T23:13:24.663000",
      "content": "<p>Hi Everyone,</p>\n<p>Just though I'd close the loop on this. I took the Pytorch Lightning recommendations and spent several sessions working through it. The model now runs on multi GPU. Luckily I did not have to go model parallel, instead reducing the batch size was enough to get it to fit. So thanks so much for your guidance.</p>\n<p>Sadly the denoising algorithm I'm trying doesn't seem to help on the larger images and so is not competitive. I'll keep trying a bit longer because I think denoising will be useful in its own right. But at some point I'll give up and move on to data augmentation.</p>\n<p>For those who are interested here's the model code I arrived at. I'm still not 100% certain that <code>test_epoch_end()</code> and <code>validation_epoch_end()</code> are quite correct. They run on every GPU so is it possible there are race conditions? I.e. could the <code>self.all_gather()</code> functions pull in data from other GPUs before it's ready?</p>\n<p>`baseline_name = 'efficientnet-b0'<br>\npretrained_model = {<br>\n    'efficientnet-b0': '../input/efficientnet-pytorch/efficientnet-b0-08094119.pth',<br>\n    'efficientnet-b1': '../input/efficientnet-pytorch/efficientnet-b1-dbc7070a.pth',<br>\n    'efficientnet-b2': '../input/efficientnet-pytorch/efficientnet-b2-27687264.pth',<br>\n    'efficientnet-b3': '../input/efficientnet-pytorch/efficientnet-b3-c8376fa2.pth',<br>\n    'efficientnet-b4': '../input/efficientnet-pytorch/efficientnet-b4-e116e8b3.pth',<br>\n    'efficientnet-b5': '../input/efficientnet-pytorch/efficientnet-b5-586e6cc6.pth',<br>\n    'efficientnet-b6': '../input/efficientnet-pytorch/efficientnet-b6-c76e70fd.pth',<br>\n    'efficientnet-b7': '../input/efficientnet-pytorch/efficientnet-b7-dcc49843.pth'<br>\n} # pretrained_model</p>\n<p>class lit_enetv2(pl.LightningModule):<br>\n    def <strong>init</strong>(self, backbone):<br>\n        super(lit_enetv2, self).<strong>init</strong>()<br>\n        self.enet = enet.EfficientNet.from_name(backbone)<br>\n        self.enet.load_state_dict(torch.load(pretrained_model[backbone]))<br>\n        self.myfc = nn.Linear(self.enet._fc.in_features, 1) # out_dim = 1<br>\n        self.enet._fc = nn.Identity()<br>\n        self.conv1 = nn.Conv2d(1, 3, kernel_size=3, stride=1, padding=3, bias=False) # in_channels, out_channels, kernel_size</p>\n<pre><code>    self.classifier = nn.Sequential(self.conv1, self.enet, self.myfc)\n\ndef forward(self, x):\n    return torch.squeeze(self.classifier(x), 1)\n\ndef step(self, step_type, batch, batch_idx):\n    x = batch['image']\n    y = batch['targets']\n    y_hat = self(x) # self.classifier(x)\n    loss = F.binary_cross_entropy_with_logits(y_hat, y) # training criterion\n    self.log(step_type + '_loss', loss, on_step=False, on_epoch=True, prog_bar=True, logger=True, sync_dist=(step_type != 'train'))\n    return loss, y_hat\n\ndef training_step(self, batch, batch_idx):\n    loss, _ = self.step('train', batch, batch_idx)\n    return loss\n\ndef validation_step(self, batch, batch_idx):\n    _, y_hat = self.step('xval', batch, batch_idx)\n    return {\n        'y_hat': y_hat,\n        'y': batch['targets'],\n    } # return\n\ndef validation_epoch_end(self, validation_step_outputs):\n    validation_step_outputs = self.all_gather(validation_step_outputs)\n\n    y_hats = []\n    ys = []\n    for vso in validation_step_outputs:\n        yh = vso['y_hat'].reshape(-1,)\n        y_hats.append(yh)\n        y = vso['y'].reshape(-1,)\n        ys.append(y)\n\n    y_hats = torch.cat(y_hats)\n    ys  = torch.cat(ys)\n    auroc = metrics.roc_auc_score(ys.cpu(), y_hats.cpu())\n    self.log('xval_auroc', auroc, on_step=False, on_epoch=True, prog_bar=True, logger=True, sync_dist=True)\n\ndef test_step(self, batch, batch_idx):\n    _, y_hat = self.step('xval', batch, batch_idx)\n    return {\n        'y_hat': y_hat,\n        'items': batch['item'],\n    } # return\n\ndef test_epoch_end(self, outputs):\n    outputs = self.all_gather(outputs)\n\n    y_hats = []\n    items = []\n    for op in outputs:\n        yh = op['y_hat'].reshape(-1,)\n        y_hats.append(yh)\n        im = op['items'].reshape(-1,)\n        items.append(im)\n\n    y_hats = torch.cat(y_hats)\n    items  = torch.cat(items)\n\n    y_hats_sorted = y_hats.clone().detach()\n    for i, ii in enumerate(items):\n        y_hats_sorted[ii] = y_hats[i]\n    self.test_output = y_hats_sorted\n\ndef configure_optimizers(self):\n    optimizer = torch.optim.Adam(self.parameters(), lr=CFG.learning_rate)\n    return optimizer\n</code></pre>\n<p>`</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1320805,
      "author_name": "James Howard",
      "author_url": "",
      "post_date": "2021-05-24T10:35:09.093000",
      "content": "<p>Edit: sorry I realise I misunderstood your intention. You don't wish to use DataParallel, but instead put different stages of 1 model across multiple GPUs?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1320917,
          "author_name": "John Clarke",
          "author_url": "",
          "post_date": "2021-05-24T12:16:54.587000",
          "content": "<p>Exactly! </p>\n<p>That means drilling down into how EfficientNet is implemented - which is still a bit beyond my level of knowledge.</p>\n<p>Do you know if there's an easy way to display the structure of the NN model? I'm looking for the # channels, sizes, etc. at each level. I don't understand why the memory requirements have blown out so suddenly.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1321091,
          "author_name": "Rodolphe Lampe",
          "author_url": "",
          "post_date": "2021-05-24T14:11:21.643000",
          "content": "<p>What you're doing is quite unusual. I advise you to take a look at the pytorch lightning library, it's very easy to use multiple gpus training. Your code will much more structured, concise and efficient.<br>\nUnless you do that for learning or whatever, the problem you face has been faced many times before by many people and efficient solutions have been found. Don't waste to much time and effort on such things imo.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1321351,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2021-05-24T16:32:32.933000",
          "content": "<p>Agree with <a href=\"https://www.kaggle.com/rodolphelampe\" target=\"_blank\">Rodolphe</a> - been using dual GPU's on 4 different machines over the last two years.  Getting both GPU running successfully in Pytorch or Keras is often a bitch - setting up the right nvidia libraries, etc all make it a real chore.  Pytorch and tensorflow changes in recent times have made this a bit easier but still find myself doing lots of system refreshes when a major update to Pytorch or Tensorflow happens until I get out of Nvidia Hell.   </p>\n<p>Trying to learn a new way of using GPU's - that's going to be a huge bundle of pain I don't think you really want.</p>\n<p>Out of memory - use nvidia-smi in terminal to see the amount of memory use on both GPU's.  Decent chance that only one GPU loaded when you think two are in use.  Batch size being too big is 99% of the reasons I get memory issues on my dual GPU's, most often because I have left more than one Jupyter script running.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1321358,
          "author_name": "John Clarke",
          "author_url": "",
          "post_date": "2021-05-24T16:40:13.453000",
          "content": "<p>Thanks for that suggestion. I'll give pytorch-lightning a try right now. Will let you know how it goes.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1323742": "Using pytorch Lightning or Fastai frameworks along with `nn.DataParallel(model, device_ids=[0,1,2,...])` works well (based on my experience) ",
    "1332115": "Hi Everyone,\n\nJust though I'd close the loop on this. I took the Pytorch Lightning recommendations and spent several sessions working through it. The model now runs on multi GPU. Luckily I did not have to go model parallel, instead reducing the batch size was enough to get it to fit. So thanks so much for your guidance.\n\nSadly the denoising algorithm I'm trying doesn't seem to help on the larger images and so is not competitive. I'll keep trying a bit longer because I think denoising will be useful in its own right. But at some point I'll give up and move on to data augmentation.\n\nFor those who are interested here's the model code I arrived at. I'm still not 100% certain that `test_epoch_end()` and `validation_epoch_end()` are quite correct. They run on every GPU so is it possible there are race conditions? I.e. could the `self.all_gather()` functions pull in data from other GPUs before it's ready?\n\n`baseline_name = 'efficientnet-b0'\npretrained_model = {\n    'efficientnet-b0': '../input/efficientnet-pytorch/efficientnet-b0-08094119.pth',\n    'efficientnet-b1': '../input/efficientnet-pytorch/efficientnet-b1-dbc7070a.pth',\n    'efficientnet-b2': '../input/efficientnet-pytorch/efficientnet-b2-27687264.pth',\n    'efficientnet-b3': '../input/efficientnet-pytorch/efficientnet-b3-c8376fa2.pth',\n    'efficientnet-b4': '../input/efficientnet-pytorch/efficientnet-b4-e116e8b3.pth',\n    'efficientnet-b5': '../input/efficientnet-pytorch/efficientnet-b5-586e6cc6.pth',\n    'efficientnet-b6': '../input/efficientnet-pytorch/efficientnet-b6-c76e70fd.pth',\n    'efficientnet-b7': '../input/efficientnet-pytorch/efficientnet-b7-dcc49843.pth'\n} # pretrained_model\n\nclass lit_enetv2(pl.LightningModule):\n    def __init__(self, backbone):\n        super(lit_enetv2, self).__init__()\n        self.enet = enet.EfficientNet.from_name(backbone)\n        self.enet.load_state_dict(torch.load(pretrained_model[backbone]))\n        self.myfc = nn.Linear(self.enet._fc.in_features, 1) # out_dim = 1\n        self.enet._fc = nn.Identity()\n        self.conv1 = nn.Conv2d(1, 3, kernel_size=3, stride=1, padding=3, bias=False) # in_channels, out_channels, kernel_size\n\n        self.classifier = nn.Sequential(self.conv1, self.enet, self.myfc)\n\n    def forward(self, x):\n        return torch.squeeze(self.classifier(x), 1)\n\n    def step(self, step_type, batch, batch_idx):\n        x = batch['image']\n        y = batch['targets']\n        y_hat = self(x) # self.classifier(x)\n        loss = F.binary_cross_entropy_with_logits(y_hat, y) # training criterion\n        self.log(step_type + '_loss', loss, on_step=False, on_epoch=True, prog_bar=True, logger=True, sync_dist=(step_type != 'train'))\n        return loss, y_hat\n\n    def training_step(self, batch, batch_idx):\n        loss, _ = self.step('train', batch, batch_idx)\n        return loss\n\n    def validation_step(self, batch, batch_idx):\n        _, y_hat = self.step('xval', batch, batch_idx)\n        return {\n            'y_hat': y_hat,\n            'y': batch['targets'],\n        } # return\n\n    def validation_epoch_end(self, validation_step_outputs):\n        validation_step_outputs = self.all_gather(validation_step_outputs)\n\n        y_hats = []\n        ys = []\n        for vso in validation_step_outputs:\n            yh = vso['y_hat'].reshape(-1,)\n            y_hats.append(yh)\n            y = vso['y'].reshape(-1,)\n            ys.append(y)\n\n        y_hats = torch.cat(y_hats)\n        ys  = torch.cat(ys)\n        auroc = metrics.roc_auc_score(ys.cpu(), y_hats.cpu())\n        self.log('xval_auroc', auroc, on_step=False, on_epoch=True, prog_bar=True, logger=True, sync_dist=True)\n\n    def test_step(self, batch, batch_idx):\n        _, y_hat = self.step('xval', batch, batch_idx)\n        return {\n            'y_hat': y_hat,\n            'items': batch['item'],\n        } # return\n\n    def test_epoch_end(self, outputs):\n        outputs = self.all_gather(outputs)\n\n        y_hats = []\n        items = []\n        for op in outputs:\n            yh = op['y_hat'].reshape(-1,)\n            y_hats.append(yh)\n            im = op['items'].reshape(-1,)\n            items.append(im)\n\n        y_hats = torch.cat(y_hats)\n        items  = torch.cat(items)\n        \n        y_hats_sorted = y_hats.clone().detach()\n        for i, ii in enumerate(items):\n            y_hats_sorted[ii] = y_hats[i]\n        self.test_output = y_hats_sorted\n\n    def configure_optimizers(self):\n        optimizer = torch.optim.Adam(self.parameters(), lr=CFG.learning_rate)\n        return optimizer\n`",
    "1320805": "Edit: sorry I realise I misunderstood your intention. You don't wish to use DataParallel, but instead put different stages of 1 model across multiple GPUs?",
    "1320794": "I am trying to partition a NN based on EfficientNet B0 across multiple NVidia K80s. In the code below I am using 2 K80's, but still running out of GPU memory on GPU 0. The machine has 8 K80's attached so perhaps even finer parallelism would work.\n\nI'm still learning the nitty gritty details of pytorch and EfficientNet so any help or pointers would be really appreciated. (I have also reached out to Luke M-L, the efficientnet-pytorch author, so hopefully he will respond.)\n\nBelow is my best effort so far. I think I need to drill down into efficientnet itself to use more GPUs, but where to start?\n\nThanks!\n\n\n>    baseline_name = 'efficientnet-b0'\n   pretrained_model = {\n       'efficientnet-b0': '../input/efficientnet-pytorch/efficientnet-b0-08094119.pth',\n       'efficientnet-b1': '../input/efficientnet-pytorch/efficientnet-b1-dbc7070a.pth',\n       'efficientnet-b2': '../input/efficientnet-pytorch/efficientnet-b2-27687264.pth',\n       'efficientnet-b3': '../input/efficientnet-pytorch/efficientnet-b3-c8376fa2.pth',\n       'efficientnet-b4': '../input/efficientnet-pytorch/efficientnet-b4-e116e8b3.pth',\n       'efficientnet-b5': '../input/efficientnet-pytorch/efficientnet-b5-586e6cc6.pth',\n       'efficientnet-b6': '../input/efficientnet-pytorch/efficientnet-b6-c76e70fd.pth',\n       'efficientnet-b7': '../input/efficientnet-pytorch/efficientnet-b7-dcc49843.pth'\n   } # pretrained_model\n\n   > class enetv2(nn.Module):\n       def __init__(self, backbone, out_dim):\n           super(enetv2, self).__init__()\n           self.enet = enet.EfficientNet.from_name(backbone)\n           # self.enet.load_state_dict(torch.load(pretrained_model[backbone]))\n           self.enet.load_state_dict(torch.load(pretrained_model[backbone], map_location=\"cpu\"))\n           self.myfc = nn.Linear(self.enet._fc.in_features, out_dim).to('cuda:0')\n           self.enet._fc = nn.Identity()\n           self.conv1 = nn.Conv2d(1, 3, kernel_size=3, stride=1, padding=3, bias=False).to('cuda:1') # in_channels, out_channels, kernel_size\n\n    def extract(self, x):\n        return self.enet(x)\n\n    def forward(self, x):\n        #x = self.conv1(x)\n        #x = self.extract(x)\n        #x = self.myfc(x)\n        x = self.extract(self.conv1(x.to('cuda:1')))\n        return self.myfc(x.to('cuda:0'))\n"
  }
}