{
  "id": 553886,
  "title": "If you are familiar with ddp training in pytorch lightning, then help me.",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/553886",
  "author_name": "",
  "post_date": "2024-12-29T04:14:51.707487600Z",
  "votes": 3,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Notebook - <a href=\"https://www.kaggle.com/code/vigneshwar472/baseline-residualunetse3d\" target=\"_blank\">https://www.kaggle.com/code/vigneshwar472/baseline-residualunetse3d</a><br>\nGithub repo (ResidualUNetSE3D implementation) - <a href=\"https://github.com/wolny/pytorch-3dunet/tree/master\" target=\"_blank\">https://github.com/wolny/pytorch-3dunet/tree/master</a></p>\n<p>Issue - Training starts on 1x P100 GPU but it does not start on 2x T4 GPU</p>\n<h3>1x P100 GPU</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15856017%2Feb0516914782c23ca6f71ce26a4a0f0a%2FScreenshot%202024-12-29%20093405.png?generation=1735445558456464&amp;alt=media\" alt=\"\"></p>\n<h3>2x T4 GPU (Training does not start and GPUs were not in use)</h3>\n<pre><code>n = len\n\nfor i in range:\n    print\n    model = ResidualUNetSE3D\n    lm = CZIILightningModule\n    logger = CSVLogger\n    trainer = Trainer\n    trainer.fit, \n                val_dataloaders=DataLoader)\n    del model, lm, logger, trainer\n    print\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15856017%2F9a609bc04602f6749d34bf7be03f903a%2FScreenshot%202024-12-29%20092937.png?generation=1735445583781330&amp;alt=media\" alt=\"\"></p>\n<p>I have no idea \"why it's not working\".</p>",
  "messages": [
    {
      "id": "3083128",
      "postDate": "12/29/2024 04:14:51",
      "content": "<p>Notebook - <a href=\"https://www.kaggle.com/code/vigneshwar472/baseline-residualunetse3d\" target=\"_blank\">https://www.kaggle.com/code/vigneshwar472/baseline-residualunetse3d</a><br>\nGithub repo (ResidualUNetSE3D implementation) - <a href=\"https://github.com/wolny/pytorch-3dunet/tree/master\" target=\"_blank\">https://github.com/wolny/pytorch-3dunet/tree/master</a></p>\n<p>Issue - Training starts on 1x P100 GPU but it does not start on 2x T4 GPU</p>\n<h3>1x P100 GPU</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15856017%2Feb0516914782c23ca6f71ce26a4a0f0a%2FScreenshot%202024-12-29%20093405.png?generation=1735445558456464&amp;alt=media\" alt=\"\"></p>\n<h3>2x T4 GPU (Training does not start and GPUs were not in use)</h3>\n<pre><code>n = len\n\nfor i in range:\n    print\n    model = ResidualUNetSE3D\n    lm = CZIILightningModule\n    logger = CSVLogger\n    trainer = Trainer\n    trainer.fit, \n                val_dataloaders=DataLoader)\n    del model, lm, logger, trainer\n    print\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15856017%2F9a609bc04602f6749d34bf7be03f903a%2FScreenshot%202024-12-29%20092937.png?generation=1735445583781330&amp;alt=media\" alt=\"\"></p>\n<p>I have no idea \"why it's not working\".</p>",
      "rawMarkdown": "Notebook - https://www.kaggle.com/code/vigneshwar472/baseline-residualunetse3d\nGithub repo (ResidualUNetSE3D implementation) - https://github.com/wolny/pytorch-3dunet/tree/master\n\nIssue - Training starts on 1x P100 GPU but it does not start on 2x T4 GPU\n\n### 1x P100 GPU\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15856017%2Feb0516914782c23ca6f71ce26a4a0f0a%2FScreenshot%202024-12-29%20093405.png?generation=1735445558456464&alt=media)\n\n### 2x T4 GPU (Training does not start and GPUs were not in use)\n```\nn = len(folds)\n\nfor i in range(n):\n    print(f'fold {i} started....')\n    model = ResidualUNetSE3D(in_channels=1, out_channels=6)\n    lm = CZIILightningModule(model=model)\n    logger = CSVLogger(save_dir='/kaggle/working/training_results', name=f'fold_{i}')\n    trainer = Trainer(accelerator='gpu',\n                     strategy='ddp_notebook',\n                     devices=2,\n                     precision='32',\n                     gradient_clip_val=None, \n                     logger=logger,\n                     max_epochs=15,\n                     enable_checkpointing=True,\n                     enable_progress_bar=True,\n                     enable_model_summary=False,\n                     inference_mode=True,\n                     default_root_dir='/kaggle/working/training_results',\n                     num_sanity_val_steps=0)\n    trainer.fit(model=lm, \n                train_dataloaders=DataLoader(folds[i][0], batch_size=1, num_workers=4, shuffle=True), \n                val_dataloaders=DataLoader(folds[i][1], batch_size=1, num_workers=4, shuffle=False))\n    del model, lm, logger, trainer\n    print(f'fold {i} completed....')\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15856017%2F9a609bc04602f6749d34bf7be03f903a%2FScreenshot%202024-12-29%20092937.png?generation=1735445583781330&alt=media)\n\nI have no idea \"why it's not working\".",
      "votes": null
    },
    {
      "id": "3083277",
      "postDate": "12/29/2024 08:29:22",
      "content": "<p>Hi. No familiar with lightning but I'm curious… No error crash, just a never ending pause?</p>",
      "rawMarkdown": "Hi. No familiar with lightning but I'm curious... No error crash, just a never ending pause?",
      "votes": null
    },
    {
      "id": "3083750",
      "postDate": "12/30/2024 01:51:42",
      "content": "<p><a href=\"https://www.kaggle.com/vigneshwar472\" target=\"_blank\">@vigneshwar472</a> it used to be a pain to train notebooks with ddp in pytorch lightning. You might try <a href=\"https://lightning.ai/docs/pytorch/stable/accelerators/gpu_intermediate.html#distributed-data-parallel-in-notebooks\" target=\"_blank\">strategy=\"ddp_notebook\"</a>.</p>",
      "rawMarkdown": "vigneshwar472 it used to be a pain to train notebooks with ddp in pytorch lightning. You might try [strategy=\"ddp_notebook\"](https://lightning.ai/docs/pytorch/stable/accelerators/gpu_intermediate.html#distributed-data-parallel-in-notebooks).",
      "votes": null
    },
    {
      "id": "3083789",
      "postDate": "12/30/2024 03:11:39",
      "content": "<p>I am using strategy=\"ddp_notebook\"</p>",
      "rawMarkdown": "I am using strategy=\"ddp_notebook\"",
      "votes": null
    },
    {
      "id": "3083792",
      "postDate": "12/30/2024 03:12:43",
      "content": "<p>I also opened an issue on GitHub. <br>\n<a href=\"https://github.com/Lightning-AI/pytorch-lightning/issues/20523\" target=\"_blank\">https://github.com/Lightning-AI/pytorch-lightning/issues/20523</a></p>",
      "rawMarkdown": "I also opened an issue on GitHub. \nhttps://github.com/Lightning-AI/pytorch-lightning/issues/20523",
      "votes": null
    },
    {
      "id": "3083987",
      "postDate": "12/30/2024 08:52:33",
      "content": "<p>I tried the notebook you shared, and no errors occurred. That’s quite strange.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5483160%2F583f4ffce7e8616e814e84401f5d9657%2F_405e8469-f854-43a5-8b90-4efbd0b341ed.png?generation=1735548724169279&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I tried the notebook you shared, and no errors occurred. That’s quite strange.![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5483160%2F583f4ffce7e8616e814e84401f5d9657%2F_405e8469-f854-43a5-8b90-4efbd0b341ed.png?generation=1735548724169279&alt=media)",
      "votes": null
    },
    {
      "id": "3084004",
      "postDate": "12/30/2024 09:06:50",
      "content": "<ol>\n<li>Turn on 2x T4 GPUs</li>\n<li>Uncomment the <strong>\"strategy\"</strong> hook in the trainer </li>\n<li>Set <strong>devices = 2</strong> </li>\n</ol>\n<p>Now run it. You will encounter the situation.</p>",
      "rawMarkdown": "1. Turn on 2x T4 GPUs\n2. Uncomment the **\"strategy\"** hook in the trainer \n3. Set **devices = 2** \n\nNow run it. You will encounter the situation.",
      "votes": null
    },
    {
      "id": "3084367",
      "postDate": "12/30/2024 18:15:00",
      "content": "<p>The times I used T4 x2 I had to add a line like <code>model = nn.DataParallel(model)</code> in my code.</p>\n<p>More explicitely:</p>\n<p><code>model = torch.nn.DataParallel(HMSmodel([EEGmodel,SPECmodel,customSPECmodel]), device_ids = [0,1]).to(device)</code></p>",
      "rawMarkdown": "The times I used T4 x2 I had to add a line like `model = nn.DataParallel(model)` in my code.\n\nMore explicitely:\n\n`model = torch.nn.DataParallel(HMSmodel([EEGmodel,SPECmodel,customSPECmodel]), device_ids = [0,1]).to(device)`",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3083277,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "12/29/2024 08:29:22",
      "content": "<p>Hi. No familiar with lightning but I'm curious… No error crash, just a never ending pause?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3083750,
      "author_name": "sergiosaharovskiy",
      "author_url": "",
      "post_date": "12/30/2024 01:51:42",
      "content": "<p><a href=\"https://www.kaggle.com/vigneshwar472\" target=\"_blank\">@vigneshwar472</a> it used to be a pain to train notebooks with ddp in pytorch lightning. You might try <a href=\"https://lightning.ai/docs/pytorch/stable/accelerators/gpu_intermediate.html#distributed-data-parallel-in-notebooks\" target=\"_blank\">strategy=\"ddp_notebook\"</a>.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3083789,
          "author_name": "vigneshwar472",
          "author_url": "",
          "post_date": "12/30/2024 03:11:39",
          "content": "<p>I am using strategy=\"ddp_notebook\"</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3083792,
      "author_name": "vigneshwar472",
      "author_url": "",
      "post_date": "12/30/2024 03:12:43",
      "content": "<p>I also opened an issue on GitHub. <br>\n<a href=\"https://github.com/Lightning-AI/pytorch-lightning/issues/20523\" target=\"_blank\">https://github.com/Lightning-AI/pytorch-lightning/issues/20523</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3083987,
      "author_name": "sjtuwangshuo",
      "author_url": "",
      "post_date": "12/30/2024 08:52:33",
      "content": "<p>I tried the notebook you shared, and no errors occurred. That’s quite strange.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5483160%2F583f4ffce7e8616e814e84401f5d9657%2F_405e8469-f854-43a5-8b90-4efbd0b341ed.png?generation=1735548724169279&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 3084004,
          "author_name": "vigneshwar472",
          "author_url": "",
          "post_date": "12/30/2024 09:06:50",
          "content": "<ol>\n<li>Turn on 2x T4 GPUs</li>\n<li>Uncomment the <strong>\"strategy\"</strong> hook in the trainer </li>\n<li>Set <strong>devices = 2</strong> </li>\n</ol>\n<p>Now run it. You will encounter the situation.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3084367,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "12/30/2024 18:15:00",
      "content": "<p>The times I used T4 x2 I had to add a line like <code>model = nn.DataParallel(model)</code> in my code.</p>\n<p>More explicitely:</p>\n<p><code>model = torch.nn.DataParallel(HMSmodel([EEGmodel,SPECmodel,customSPECmodel]), device_ids = [0,1]).to(device)</code></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3083128": "Notebook - https://www.kaggle.com/code/vigneshwar472/baseline-residualunetse3d\nGithub repo (ResidualUNetSE3D implementation) - https://github.com/wolny/pytorch-3dunet/tree/master\n\nIssue - Training starts on 1x P100 GPU but it does not start on 2x T4 GPU\n\n### 1x P100 GPU\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15856017%2Feb0516914782c23ca6f71ce26a4a0f0a%2FScreenshot%202024-12-29%20093405.png?generation=1735445558456464&alt=media)\n\n### 2x T4 GPU (Training does not start and GPUs were not in use)\n```\nn = len(folds)\n\nfor i in range(n):\n    print(f'fold {i} started....')\n    model = ResidualUNetSE3D(in_channels=1, out_channels=6)\n    lm = CZIILightningModule(model=model)\n    logger = CSVLogger(save_dir='/kaggle/working/training_results', name=f'fold_{i}')\n    trainer = Trainer(accelerator='gpu',\n                     strategy='ddp_notebook',\n                     devices=2,\n                     precision='32',\n                     gradient_clip_val=None, \n                     logger=logger,\n                     max_epochs=15,\n                     enable_checkpointing=True,\n                     enable_progress_bar=True,\n                     enable_model_summary=False,\n                     inference_mode=True,\n                     default_root_dir='/kaggle/working/training_results',\n                     num_sanity_val_steps=0)\n    trainer.fit(model=lm, \n                train_dataloaders=DataLoader(folds[i][0], batch_size=1, num_workers=4, shuffle=True), \n                val_dataloaders=DataLoader(folds[i][1], batch_size=1, num_workers=4, shuffle=False))\n    del model, lm, logger, trainer\n    print(f'fold {i} completed....')\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15856017%2F9a609bc04602f6749d34bf7be03f903a%2FScreenshot%202024-12-29%20092937.png?generation=1735445583781330&alt=media)\n\nI have no idea \"why it's not working\".",
    "3083277": "Hi. No familiar with lightning but I'm curious... No error crash, just a never ending pause?",
    "3083750": "vigneshwar472 it used to be a pain to train notebooks with ddp in pytorch lightning. You might try [strategy=\"ddp_notebook\"](https://lightning.ai/docs/pytorch/stable/accelerators/gpu_intermediate.html#distributed-data-parallel-in-notebooks).",
    "3083789": "I am using strategy=\"ddp_notebook\"",
    "3083792": "I also opened an issue on GitHub. \nhttps://github.com/Lightning-AI/pytorch-lightning/issues/20523",
    "3083987": "I tried the notebook you shared, and no errors occurred. That’s quite strange.![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5483160%2F583f4ffce7e8616e814e84401f5d9657%2F_405e8469-f854-43a5-8b90-4efbd0b341ed.png?generation=1735548724169279&alt=media)",
    "3084004": "1. Turn on 2x T4 GPUs\n2. Uncomment the **\"strategy\"** hook in the trainer \n3. Set **devices = 2** \n\nNow run it. You will encounter the situation.",
    "3084367": "The times I used T4 x2 I had to add a line like `model = nn.DataParallel(model)` in my code.\n\nMore explicitely:\n\n`model = torch.nn.DataParallel(HMSmodel([EEGmodel,SPECmodel,customSPECmodel]), device_ids = [0,1]).to(device)`"
  },
  "source": "meta"
}