{
  "id": 390849,
  "title": "Train GraphNeT on multiple GPUs",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/390849",
  "author_name": "",
  "post_date": "2023-02-27T14:39:42.367667500Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Is where a way to train GraphNeT model on multiple GPUs? I tried to use <code>gpu = [0, 1]</code> and <code>distribution_strategy = 'dp'</code> in <code>model.fit()</code>, but got a following error: </p>\n<p><code>Expected all tensors to be on the same device, but found at least two devices, cuda:0 and cpu! (when checking argument for argument index in method wrapper_scatter_add_)</code></p>",
  "messages": [
    {
      "id": "2161482",
      "postDate": "02/27/2023 14:39:42",
      "content": "<p>Is where a way to train GraphNeT model on multiple GPUs? I tried to use <code>gpu = [0, 1]</code> and <code>distribution_strategy = 'dp'</code> in <code>model.fit()</code>, but got a following error: </p>\n<p><code>Expected all tensors to be on the same device, but found at least two devices, cuda:0 and cpu! (when checking argument for argument index in method wrapper_scatter_add_)</code></p>",
      "rawMarkdown": "Is where a way to train GraphNeT model on multiple GPUs? I tried to use `gpu = [0, 1]` and `distribution_strategy = 'dp'` in `model.fit()`, but got a following error: \n\n`Expected all tensors to be on the same device, but found at least two devices, cuda:0 and cpu! (when checking argument for argument index in method wrapper_scatter_add_)`",
      "votes": null
    },
    {
      "id": "2161539",
      "postDate": "02/27/2023 15:15:49",
      "content": "<p>Hi,<br>\nI had similar problem. <br>\nTry to run <code>distribution_strategy = 'ddp'</code>. And it's better to run pure python script instead of jupyter notebook.<br>\n<code>trainer = Trainer(strategy=\"ddp\", accelerator=\"gpu\", devices=2)</code></p>\n<p>From lightning <a href=\"https://pytorch-lightning.readthedocs.io/en/stable/accelerators/gpu_intermediate.html#data-parallel\" target=\"_blank\">doc</a>:</p>\n<blockquote>\n  <p>DP use is discouraged by PyTorch and Lightning. State is not maintained on the replicas created by</p>\n</blockquote>",
      "rawMarkdown": "Hi,\nI had similar problem. \nTry to run `distribution_strategy = 'ddp'`. And it's better to run pure python script instead of jupyter notebook.\n`trainer = Trainer(strategy=\"ddp\", accelerator=\"gpu\", devices=2)`\n\nFrom lightning [doc](https://pytorch-lightning.readthedocs.io/en/stable/accelerators/gpu_intermediate.html#data-parallel):\n>DP use is discouraged by PyTorch and Lightning. State is not maintained on the replicas created by",
      "votes": null
    },
    {
      "id": "2161623",
      "postDate": "02/27/2023 16:05:06",
      "content": "<p>Hi Ruslan, have you tried to run inference on two gpus (T4x2)?</p>\n<p>you can find more details here about my question.<br>\n<a href=\"https://www.kaggle.com/code/rasmusrse/graphnet-baseline-submission/comments#2159385\" target=\"_blank\">https://www.kaggle.com/code/rasmusrse/graphnet-baseline-submission/comments#2159385</a></p>",
      "rawMarkdown": "Hi Ruslan, have you tried to run inference on two gpus (T4x2)?\n\nyou can find more details here about my question.\nhttps://www.kaggle.com/code/rasmusrse/graphnet-baseline-submission/comments#2159385",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2161539,
      "author_name": "rusg77",
      "author_url": "",
      "post_date": "02/27/2023 15:15:49",
      "content": "<p>Hi,<br>\nI had similar problem. <br>\nTry to run <code>distribution_strategy = 'ddp'</code>. And it's better to run pure python script instead of jupyter notebook.<br>\n<code>trainer = Trainer(strategy=\"ddp\", accelerator=\"gpu\", devices=2)</code></p>\n<p>From lightning <a href=\"https://pytorch-lightning.readthedocs.io/en/stable/accelerators/gpu_intermediate.html#data-parallel\" target=\"_blank\">doc</a>:</p>\n<blockquote>\n  <p>DP use is discouraged by PyTorch and Lightning. State is not maintained on the replicas created by</p>\n</blockquote>",
      "votes": null,
      "replies": [
        {
          "id": 2161623,
          "author_name": "mohammadrahmati",
          "author_url": "",
          "post_date": "02/27/2023 16:05:06",
          "content": "<p>Hi Ruslan, have you tried to run inference on two gpus (T4x2)?</p>\n<p>you can find more details here about my question.<br>\n<a href=\"https://www.kaggle.com/code/rasmusrse/graphnet-baseline-submission/comments#2159385\" target=\"_blank\">https://www.kaggle.com/code/rasmusrse/graphnet-baseline-submission/comments#2159385</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2161482": "Is where a way to train GraphNeT model on multiple GPUs? I tried to use `gpu = [0, 1]` and `distribution_strategy = 'dp'` in `model.fit()`, but got a following error: \n\n`Expected all tensors to be on the same device, but found at least two devices, cuda:0 and cpu! (when checking argument for argument index in method wrapper_scatter_add_)`",
    "2161539": "Hi,\nI had similar problem. \nTry to run `distribution_strategy = 'ddp'`. And it's better to run pure python script instead of jupyter notebook.\n`trainer = Trainer(strategy=\"ddp\", accelerator=\"gpu\", devices=2)`\n\nFrom lightning [doc](https://pytorch-lightning.readthedocs.io/en/stable/accelerators/gpu_intermediate.html#data-parallel):\n>DP use is discouraged by PyTorch and Lightning. State is not maintained on the replicas created by",
    "2161623": "Hi Ruslan, have you tried to run inference on two gpus (T4x2)?\n\nyou can find more details here about my question.\nhttps://www.kaggle.com/code/rasmusrse/graphnet-baseline-submission/comments#2159385"
  },
  "source": "meta"
}