{
  "id": 586109,
  "title": "A question about nn.DataParallel",
  "url": "/competitions/waveform-inversion/discussion/586109",
  "author_name": "",
  "post_date": "2025-06-25T07:21:57.130902100Z",
  "votes": 5,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi everyone, recently I  tried to modify the code:  <a href=\"https://www.kaggle.com/code/muhammadqasimshabbir/gwi-unet-with-float16-dataset22\" target=\"_blank\">https://www.kaggle.com/code/muhammadqasimshabbir/gwi-unet-with-float16-dataset22</a> to run on two T4 GPUs from Kaggle using nn.DataParallel and also changed the batchsize to 128(2x compared to original). However, it seems like the running time have no improvement, has anyone encountered the same issue before?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14642567%2F1d24e8059a59d81559474ea6c94dc86d%2Fori.png?generation=1750836042600934&amp;alt=media\" alt=\"original\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14642567%2Fb3dcdd0491fc74de30eea6cd7c93f7d9%2Fmy.png?generation=1750836082321816&amp;alt=media\" alt=\"mine\"></p>",
  "messages": [
    {
      "id": "3231961",
      "postDate": "06/25/2025 07:21:57",
      "content": "<p>Hi everyone, recently I  tried to modify the code:  <a href=\"https://www.kaggle.com/code/muhammadqasimshabbir/gwi-unet-with-float16-dataset22\" target=\"_blank\">https://www.kaggle.com/code/muhammadqasimshabbir/gwi-unet-with-float16-dataset22</a> to run on two T4 GPUs from Kaggle using nn.DataParallel and also changed the batchsize to 128(2x compared to original). However, it seems like the running time have no improvement, has anyone encountered the same issue before?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14642567%2F1d24e8059a59d81559474ea6c94dc86d%2Fori.png?generation=1750836042600934&amp;alt=media\" alt=\"original\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14642567%2Fb3dcdd0491fc74de30eea6cd7c93f7d9%2Fmy.png?generation=1750836082321816&amp;alt=media\" alt=\"mine\"></p>",
      "rawMarkdown": "Hi everyone, recently I  tried to modify the code:  https://www.kaggle.com/code/muhammadqasimshabbir/gwi-unet-with-float16-dataset22 to run on two T4 GPUs from Kaggle using nn.DataParallel and also changed the batchsize to 128(2x compared to original). However, it seems like the running time have no improvement, has anyone encountered the same issue before?\n![original](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14642567%2F1d24e8059a59d81559474ea6c94dc86d%2Fori.png?generation=1750836042600934&alt=media)\n\n![mine](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14642567%2Fb3dcdd0491fc74de30eea6cd7c93f7d9%2Fmy.png?generation=1750836082321816&alt=media)",
      "votes": null
    },
    {
      "id": "3232011",
      "postDate": "06/25/2025 09:07:28",
      "content": "<p>DDP (DistributedDataParallel) is more efficient. DP should never be used. You can find an example of DDP in the notebook by <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> <a href=\"https://www.kaggle.com/code/brendanartley/caformer-full-resolution-improved\" target=\"_blank\">https://www.kaggle.com/code/brendanartley/caformer-full-resolution-improved</a></p>",
      "rawMarkdown": "DDP (DistributedDataParallel) is more efficient. DP should never be used. You can find an example of DDP in the notebook by @brendanartley https://www.kaggle.com/code/brendanartley/caformer-full-resolution-improved",
      "votes": null
    },
    {
      "id": "3232229",
      "postDate": "06/25/2025 14:12:14",
      "content": "<p>i tried both nn.DataParallel and DDP (DistributedDataParallel) .</p>\n<p>My suggestion is that don't use nn.DataParallel  anymore.<br>\nnn.DataParallel  is much slower than DDP.</p>\n<p>this is becuase on how the grad and other computation are sync from one gpu to another.<br>\n(you can ask chatgpt on why and how the gpu computation are sync for more infor)</p>\n<hr>\n<p>to test if nn.DataParallel is correctly setup,</p>\n<ul>\n<li>make sure you can now process 2x data as before</li>\n<li>there is GPU utlization on BOTH gpu</li>\n<li>you only get speedup if GPU utlization is high enough on BOTH GPU</li>\n</ul>\n<p>e.g. if i use one gpu, utlization is 100%.<br>\nbut if i use both, they become 33% each. then maybe there is inefficiency to split and transfer data to 2 gpu,etc  or inefficiecy to sunc gpu computation among gpus, etc</p>",
      "rawMarkdown": "i tried both nn.DataParallel and DDP (DistributedDataParallel) .\n\nMy suggestion is that don't use nn.DataParallel  anymore.\nnn.DataParallel  is much slower than DDP.\n\nthis is becuase on how the grad and other computation are sync from one gpu to another.\n(you can ask chatgpt on why and how the gpu computation are sync for more infor)\n\n---\n\nto test if nn.DataParallel is correctly setup,\n- make sure you can now process 2x data as before\n- there is GPU utlization on BOTH gpu\n- you only get speedup if GPU utlization is high enough on BOTH GPU\n\ne.g. if i use one gpu, utlization is 100%.\nbut if i use both, they become 33% each. then maybe there is inefficiency to split and transfer data to 2 gpu,etc  or inefficiecy to sunc gpu computation among gpus, etc",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3232011,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/25/2025 09:07:28",
      "content": "<p>DDP (DistributedDataParallel) is more efficient. DP should never be used. You can find an example of DDP in the notebook by <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> <a href=\"https://www.kaggle.com/code/brendanartley/caformer-full-resolution-improved\" target=\"_blank\">https://www.kaggle.com/code/brendanartley/caformer-full-resolution-improved</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3232229,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/25/2025 14:12:14",
      "content": "<p>i tried both nn.DataParallel and DDP (DistributedDataParallel) .</p>\n<p>My suggestion is that don't use nn.DataParallel  anymore.<br>\nnn.DataParallel  is much slower than DDP.</p>\n<p>this is becuase on how the grad and other computation are sync from one gpu to another.<br>\n(you can ask chatgpt on why and how the gpu computation are sync for more infor)</p>\n<hr>\n<p>to test if nn.DataParallel is correctly setup,</p>\n<ul>\n<li>make sure you can now process 2x data as before</li>\n<li>there is GPU utlization on BOTH gpu</li>\n<li>you only get speedup if GPU utlization is high enough on BOTH GPU</li>\n</ul>\n<p>e.g. if i use one gpu, utlization is 100%.<br>\nbut if i use both, they become 33% each. then maybe there is inefficiency to split and transfer data to 2 gpu,etc  or inefficiecy to sunc gpu computation among gpus, etc</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3231961": "Hi everyone, recently I  tried to modify the code:  https://www.kaggle.com/code/muhammadqasimshabbir/gwi-unet-with-float16-dataset22 to run on two T4 GPUs from Kaggle using nn.DataParallel and also changed the batchsize to 128(2x compared to original). However, it seems like the running time have no improvement, has anyone encountered the same issue before?\n![original](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14642567%2F1d24e8059a59d81559474ea6c94dc86d%2Fori.png?generation=1750836042600934&alt=media)\n\n![mine](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14642567%2Fb3dcdd0491fc74de30eea6cd7c93f7d9%2Fmy.png?generation=1750836082321816&alt=media)",
    "3232011": "DDP (DistributedDataParallel) is more efficient. DP should never be used. You can find an example of DDP in the notebook by @brendanartley https://www.kaggle.com/code/brendanartley/caformer-full-resolution-improved",
    "3232229": "i tried both nn.DataParallel and DDP (DistributedDataParallel) .\n\nMy suggestion is that don't use nn.DataParallel  anymore.\nnn.DataParallel  is much slower than DDP.\n\nthis is becuase on how the grad and other computation are sync from one gpu to another.\n(you can ask chatgpt on why and how the gpu computation are sync for more infor)\n\n---\n\nto test if nn.DataParallel is correctly setup,\n- make sure you can now process 2x data as before\n- there is GPU utlization on BOTH gpu\n- you only get speedup if GPU utlization is high enough on BOTH GPU\n\ne.g. if i use one gpu, utlization is 100%.\nbut if i use both, they become 33% each. then maybe there is inefficiency to split and transfer data to 2 gpu,etc  or inefficiecy to sunc gpu computation among gpus, etc"
  },
  "source": "meta"
}