{
  "id": 228511,
  "title": "How to use multi-gpus properly?",
  "url": "/competitions/bms-molecular-translation/discussion/228511",
  "author_name": "",
  "post_date": "2021-03-25T03:10:46.066580500Z",
  "votes": 6,
  "comment_count": 4,
  "views": 0,
  "content": "<p>When i using <br>\n<code>decoder = torch.nn.DataParallel(decoder, device_ids=device_ids)</code>, <br>\nit get the bug <br>\n<code>RuntimeError: Input tensor at index 1 has invalid shape [64, 171, 193], but expected [64, 151, 193]</code>. </p>\n<p>Does anyone know how to solve it? Thank you first</p>",
  "messages": [
    {
      "id": "1251689",
      "postDate": "03/25/2021 03:10:46",
      "content": "<p>When i using <br>\n<code>decoder = torch.nn.DataParallel(decoder, device_ids=device_ids)</code>, <br>\nit get the bug <br>\n<code>RuntimeError: Input tensor at index 1 has invalid shape [64, 171, 193], but expected [64, 151, 193]</code>. </p>\n<p>Does anyone know how to solve it? Thank you first</p>",
      "rawMarkdown": "When i using \n`decoder = torch.nn.DataParallel(decoder, device_ids=device_ids)`, \nit get the bug \n`RuntimeError: Input tensor at index 1 has invalid shape [64, 171, 193], but expected [64, 151, 193]`. \n\nDoes anyone know how to solve it? Thank you first",
      "votes": null
    },
    {
      "id": "1252097",
      "postDate": "03/25/2021 12:23:01",
      "content": "<p>You can solve this problem 2 ways:</p>\n<p>1) You have to pad every sequence that generate using your data loader to the same dimension.<br>\n2) Use <code>DistributedDataParallel</code> training. If you don't want to spend time implementing I would recommend switching training loop to Pytorch Lightning (pretty easy to do). Once you have your <code>trainer</code> you can just have to modify <code>distributed_backend = dpp</code></p>\n<p>example: <br>\n<code>trainer = Trainer(experiment=exp, gpus=[0, 1], max_nb_epochs=20, distributed_backend='dpp')</code></p>",
      "rawMarkdown": "You can solve this problem 2 ways:\n\n1) You have to pad every sequence that generate using your data loader to the same dimension.\n2) Use `DistributedDataParallel` training. If you don't want to spend time implementing I would recommend switching training loop to Pytorch Lightning (pretty easy to do). Once you have your `trainer` you can just have to modify `distributed_backend = dpp`\n\nexample: \n`trainer = Trainer(experiment=exp, gpus=[0, 1], max_nb_epochs=20, distributed_backend='dpp')`",
      "votes": null
    },
    {
      "id": "1252123",
      "postDate": "03/25/2021 12:45:11",
      "content": "<p>can i ask you about something <a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> when i use tpu for my code which already run successfully on gpu it gives the error <br>\nRuntimeError: Caught RuntimeError in DataLoader worker process 0.<br>\nOriginal Traceback (most recent call last):<br>\n  File \"/opt/conda/lib/python3.7/site-packages/torch/utils/data/_utils/worker.py\", line 202, in _worker_loop<br>\n    data = fetcher.fetch(index)<br>\n  File \"/opt/conda/lib/python3.7/site-packages/torch/utils/data/_utils/fetch.py\", line 47, in fetch<br>\n    return self.collate_fn(data)<br>\n  File \"/opt/conda/lib/python3.7/site-packages/torch/utils/data/_utils/collate.py\", line 84, in default_collate<br>\n    return [default_collate(samples) for samples in transposed]<br>\n  File \"/opt/conda/lib/python3.7/site-packages/torch/utils/data/_utils/collate.py\", line 84, in <br>\n    return [default_collate(samples) for samples in transposed]<br>\n  File \"/opt/conda/lib/python3.7/site-packages/torch/utils/data/_utils/collate.py\", line 56, in default_collate<br>\n    return torch.stack(batch, 0, out=out)<br>\nRuntimeError: stack expects each tensor to be equal size, but got [101] at entry 0 and [75] at entry 1</p>",
      "rawMarkdown": "can i ask you about something @drhabib when i use tpu for my code which already run successfully on gpu it gives the error \nRuntimeError: Caught RuntimeError in DataLoader worker process 0.\nOriginal Traceback (most recent call last):\n  File \"/opt/conda/lib/python3.7/site-packages/torch/utils/data/_utils/worker.py\", line 202, in _worker_loop\n    data = fetcher.fetch(index)\n  File \"/opt/conda/lib/python3.7/site-packages/torch/utils/data/_utils/fetch.py\", line 47, in fetch\n    return self.collate_fn(data)\n  File \"/opt/conda/lib/python3.7/site-packages/torch/utils/data/_utils/collate.py\", line 84, in default_collate\n    return [default_collate(samples) for samples in transposed]\n  File \"/opt/conda/lib/python3.7/site-packages/torch/utils/data/_utils/collate.py\", line 84, in <listcomp>\n    return [default_collate(samples) for samples in transposed]\n  File \"/opt/conda/lib/python3.7/site-packages/torch/utils/data/_utils/collate.py\", line 56, in default_collate\n    return torch.stack(batch, 0, out=out)\nRuntimeError: stack expects each tensor to be equal size, but got [101] at entry 0 and [75] at entry 1",
      "votes": null
    },
    {
      "id": "1252175",
      "postDate": "03/25/2021 13:25:55",
      "content": "<p>See this thread as well (same question and discussion):</p>\n<p><a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/225856#1245187\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/225856#1245187</a></p>",
      "rawMarkdown": "See this thread as well (same question and discussion):\n\nhttps://www.kaggle.com/c/bms-molecular-translation/discussion/225856#1245187",
      "votes": null
    },
    {
      "id": "1252189",
      "postDate": "03/25/2021 13:40:45",
      "content": "<p>Unfortunately I am not familiar with tpu.</p>",
      "rawMarkdown": "Unfortunately I am not familiar with tpu.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1252097,
      "author_name": "drhabib",
      "author_url": "",
      "post_date": "03/25/2021 12:23:01",
      "content": "<p>You can solve this problem 2 ways:</p>\n<p>1) You have to pad every sequence that generate using your data loader to the same dimension.<br>\n2) Use <code>DistributedDataParallel</code> training. If you don't want to spend time implementing I would recommend switching training loop to Pytorch Lightning (pretty easy to do). Once you have your <code>trainer</code> you can just have to modify <code>distributed_backend = dpp</code></p>\n<p>example: <br>\n<code>trainer = Trainer(experiment=exp, gpus=[0, 1], max_nb_epochs=20, distributed_backend='dpp')</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 1252123,
          "author_name": "hussensehs",
          "author_url": "",
          "post_date": "03/25/2021 12:45:11",
          "content": "<p>can i ask you about something <a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> when i use tpu for my code which already run successfully on gpu it gives the error <br>\nRuntimeError: Caught RuntimeError in DataLoader worker process 0.<br>\nOriginal Traceback (most recent call last):<br>\n  File \"/opt/conda/lib/python3.7/site-packages/torch/utils/data/_utils/worker.py\", line 202, in _worker_loop<br>\n    data = fetcher.fetch(index)<br>\n  File \"/opt/conda/lib/python3.7/site-packages/torch/utils/data/_utils/fetch.py\", line 47, in fetch<br>\n    return self.collate_fn(data)<br>\n  File \"/opt/conda/lib/python3.7/site-packages/torch/utils/data/_utils/collate.py\", line 84, in default_collate<br>\n    return [default_collate(samples) for samples in transposed]<br>\n  File \"/opt/conda/lib/python3.7/site-packages/torch/utils/data/_utils/collate.py\", line 84, in <br>\n    return [default_collate(samples) for samples in transposed]<br>\n  File \"/opt/conda/lib/python3.7/site-packages/torch/utils/data/_utils/collate.py\", line 56, in default_collate<br>\n    return torch.stack(batch, 0, out=out)<br>\nRuntimeError: stack expects each tensor to be equal size, but got [101] at entry 0 and [75] at entry 1</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1252189,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "03/25/2021 13:40:45",
          "content": "<p>Unfortunately I am not familiar with tpu.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1252175,
      "author_name": "marketneutral",
      "author_url": "",
      "post_date": "03/25/2021 13:25:55",
      "content": "<p>See this thread as well (same question and discussion):</p>\n<p><a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/225856#1245187\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/225856#1245187</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1251689": "When i using \n`decoder = torch.nn.DataParallel(decoder, device_ids=device_ids)`, \nit get the bug \n`RuntimeError: Input tensor at index 1 has invalid shape [64, 171, 193], but expected [64, 151, 193]`. \n\nDoes anyone know how to solve it? Thank you first",
    "1252097": "You can solve this problem 2 ways:\n\n1) You have to pad every sequence that generate using your data loader to the same dimension.\n2) Use `DistributedDataParallel` training. If you don't want to spend time implementing I would recommend switching training loop to Pytorch Lightning (pretty easy to do). Once you have your `trainer` you can just have to modify `distributed_backend = dpp`\n\nexample: \n`trainer = Trainer(experiment=exp, gpus=[0, 1], max_nb_epochs=20, distributed_backend='dpp')`",
    "1252123": "can i ask you about something @drhabib when i use tpu for my code which already run successfully on gpu it gives the error \nRuntimeError: Caught RuntimeError in DataLoader worker process 0.\nOriginal Traceback (most recent call last):\n  File \"/opt/conda/lib/python3.7/site-packages/torch/utils/data/_utils/worker.py\", line 202, in _worker_loop\n    data = fetcher.fetch(index)\n  File \"/opt/conda/lib/python3.7/site-packages/torch/utils/data/_utils/fetch.py\", line 47, in fetch\n    return self.collate_fn(data)\n  File \"/opt/conda/lib/python3.7/site-packages/torch/utils/data/_utils/collate.py\", line 84, in default_collate\n    return [default_collate(samples) for samples in transposed]\n  File \"/opt/conda/lib/python3.7/site-packages/torch/utils/data/_utils/collate.py\", line 84, in <listcomp>\n    return [default_collate(samples) for samples in transposed]\n  File \"/opt/conda/lib/python3.7/site-packages/torch/utils/data/_utils/collate.py\", line 56, in default_collate\n    return torch.stack(batch, 0, out=out)\nRuntimeError: stack expects each tensor to be equal size, but got [101] at entry 0 and [75] at entry 1",
    "1252175": "See this thread as well (same question and discussion):\n\nhttps://www.kaggle.com/c/bms-molecular-translation/discussion/225856#1245187",
    "1252189": "Unfortunately I am not familiar with tpu."
  },
  "source": "meta"
}