{
  "id": 461892,
  "title": "Sharing Multi GPU Inference Script (3D with 2D masking)",
  "url": "/competitions/blood-vessel-segmentation/discussion/461892",
  "author_name": "E/S Pronk",
  "post_date": "2023-12-17T03:38:26.421000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Inference takes a LOT of time on only one GPU, and a lot of RAM if using multiple GPUs. Kidney 2 segmentation takes well over 2 hours on a single NVIDIA A10 for me, so I decided to make a distributed inference script using DistributedDataParallel. This script completes loading, segmenting and masking kidney 2 in 32 minutes using 4 GPUs for me.</p>\n<p>The predictions are stored only once, in (cpu) RAM, on rank 0. The other GPUs' predictions are gathered after each iteration. There are 2 models, one that segments the entire kidney from 2d slices, calculated combined over XY,XZ and YZ views of the volume. The other one takes 3d blocks of the volume to predict 3d segmentation masks. The prediction by the second model is multiplied by the mask predicted by the first model.</p>\n<p>Both models are stubbed, so there is some DIY involved in getting this running.</p>\n<p><a href=\"https://www.kaggle.com/limitz/multi-gpu-inference-2d-3d\" target=\"_blank\">https://www.kaggle.com/limitz/multi-gpu-inference-2d-3d</a></p>\n<p>NOTE: copy to something like <code>inference.py</code> and run it in the terminal. Also include <code>inline.py</code> and <code>meters.py</code> found in the utility scripts section of the notebook for saving images and progress display resp.</p>",
  "messages": [
    {
      "id": 2564343,
      "postDate": "2023-12-17T03:38:26.420Z",
      "content": "<p>Inference takes a LOT of time on only one GPU, and a lot of RAM if using multiple GPUs. Kidney 2 segmentation takes well over 2 hours on a single NVIDIA A10 for me, so I decided to make a distributed inference script using DistributedDataParallel. This script completes loading, segmenting and masking kidney 2 in 32 minutes using 4 GPUs for me.</p>\n<p>The predictions are stored only once, in (cpu) RAM, on rank 0. The other GPUs' predictions are gathered after each iteration. There are 2 models, one that segments the entire kidney from 2d slices, calculated combined over XY,XZ and YZ views of the volume. The other one takes 3d blocks of the volume to predict 3d segmentation masks. The prediction by the second model is multiplied by the mask predicted by the first model.</p>\n<p>Both models are stubbed, so there is some DIY involved in getting this running.</p>\n<p><a href=\"https://www.kaggle.com/limitz/multi-gpu-inference-2d-3d\" target=\"_blank\">https://www.kaggle.com/limitz/multi-gpu-inference-2d-3d</a></p>\n<p>NOTE: copy to something like <code>inference.py</code> and run it in the terminal. Also include <code>inline.py</code> and <code>meters.py</code> found in the utility scripts section of the notebook for saving images and progress display resp.</p>",
      "rawMarkdown": "Inference takes a LOT of time on only one GPU, and a lot of RAM if using multiple GPUs. Kidney 2 segmentation takes well over 2 hours on a single NVIDIA A10 for me, so I decided to make a distributed inference script using DistributedDataParallel. This script completes loading, segmenting and masking kidney 2 in 32 minutes using 4 GPUs for me.\n\nThe predictions are stored only once, in (cpu) RAM, on rank 0. The other GPUs' predictions are gathered after each iteration. There are 2 models, one that segments the entire kidney from 2d slices, calculated combined over XY,XZ and YZ views of the volume. The other one takes 3d blocks of the volume to predict 3d segmentation masks. The prediction by the second model is multiplied by the mask predicted by the first model.\n\nBoth models are stubbed, so there is some DIY involved in getting this running.\n\nhttps://www.kaggle.com/limitz/multi-gpu-inference-2d-3d\n\nNOTE: copy to something like `inference.py` and run it in the terminal. Also include `inline.py` and `meters.py` found in the utility scripts section of the notebook for saving images and progress display resp.",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2564343": "Inference takes a LOT of time on only one GPU, and a lot of RAM if using multiple GPUs. Kidney 2 segmentation takes well over 2 hours on a single NVIDIA A10 for me, so I decided to make a distributed inference script using DistributedDataParallel. This script completes loading, segmenting and masking kidney 2 in 32 minutes using 4 GPUs for me.\n\nThe predictions are stored only once, in (cpu) RAM, on rank 0. The other GPUs' predictions are gathered after each iteration. There are 2 models, one that segments the entire kidney from 2d slices, calculated combined over XY,XZ and YZ views of the volume. The other one takes 3d blocks of the volume to predict 3d segmentation masks. The prediction by the second model is multiplied by the mask predicted by the first model.\n\nBoth models are stubbed, so there is some DIY involved in getting this running.\n\nhttps://www.kaggle.com/limitz/multi-gpu-inference-2d-3d\n\nNOTE: copy to something like `inference.py` and run it in the terminal. Also include `inline.py` and `meters.py` found in the utility scripts section of the notebook for saving images and progress display resp."
  }
}