{
  "id": 679231,
  "title": "88th solution with Topology-Preserving 3D U-Net  and 5-Loss Stack  HybridCon",
  "url": "/competitions/vesuvius-challenge-surface-detection/discussion/679231",
  "author_name": "Manish Swami",
  "post_date": "2026-02-28T02:47:08.399000",
  "votes": 5,
  "comment_count": 0,
  "views": 0,
  "content": "<h2>1. Problem Understanding</h2>\n<p>The Vesuvius 2025 Surface Detection task requires segmenting thin papyrus surfaces from 3D CT volumes of ancient scrolls. The competition metric is:</p>\n<p><strong>LB = 0.30 × TopoScore + 0.35 × SurfaceDice + 0.35 × VOI</strong></p>\n<p>This means <strong>topology matters as much as pixel overlap</strong>. A model that gets good Dice but merges or splits surface sheets will score poorly on Betti matching and VOI. This insight drove every design decision — from the loss function to the post-processing.</p>\n<hr>\n<h2>2. Architecture: TopoPreservingUNet3D (~10M params)</h2>\n<p>Custom 3D U-Net with 6 encoder/decoder stages designed specifically for thin surface detection in CT volumes.</p>\n<p><strong>Feature channels:</strong> <code>[32, 64, 128, 256, 320, 320]</code><br>\n<strong>Residual blocks per stage:</strong> <code>[1, 2, 3, 4, 6, 6]</code></p>\n<h3>Key Design Choices</h3>\n<p><strong>HybridConv3d</strong> — Instead of standard <code>3×3×3</code> convolutions, I decouple XY and Z processing: <code>Conv3d(kernel=(1,3,3))</code> for in-plane features and <code>Conv3d(kernel=(3,1,1))</code> for cross-slice features, then concatenate. CT data is anisotropic — in-plane resolution typically differs from slice spacing — so decoupled processing respects the data geometry.</p>\n<p><strong>MultiScaleResBlock</strong> — Res2Net-style blocks where the channel dimension is split into groups that process hierarchically. Each group receives input from the previous group, building multi-scale receptive fields within a single residual block. Applied at every encoder and decoder stage.</p>\n<p><strong>AttentionBlock (stages 2–5)</strong> — Combined channel attention (squeeze-excitation via <code>AdaptiveAvgPool3d</code>) and spatial attention (<code>7×7</code> conv on concatenated mean+max pooled features). Helps the model focus on surface regions while suppressing irrelevant background.</p>\n<p><strong>SurfaceRefinementBlock (decoder stage 0)</strong> — This was critical for performance. At the highest resolution decoder stage, an edge convolution branch computes <code>|conv(x)|</code> to detect edges, concatenates with original features, and refines through two conv-norm-activation layers. Surfaces in the Vesuvius data can be <strong>1–2 voxels thick</strong>, so having dedicated edge-aware processing at full resolution made a noticeable difference.</p>\n<p><strong>Other details:</strong></p>\n<ul>\n<li><code>GroupNorm</code> throughout (stable with small batch size 4, unlike <code>BatchNorm</code>)</li>\n<li><code>LeakyReLU(0.01)</code> activation</li>\n<li>Strided convolution downsampling (learnable, instead of max-pool)</li>\n<li><code>ConvTranspose3d</code> upsampling</li>\n<li>Deep supervision heads at 3 decoder levels during training</li>\n</ul>\n<hr>\n<h2>3. Loss Function: 5-Component Topology-Aware Stack</h2>\n<p>The competition metric penalizes topological errors (Betti matching, VOI), so I designed the loss function with topology preservation as a first-class objective. All 5 losses are active from epoch 0.</p>\n<table>\n<thead>\n<tr>\n<th>Loss</th>\n<th>Weight</th>\n<th>Purpose</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Dice</strong></td>\n<td>0.25</td>\n<td>Core overlap signal for segmentation</td>\n</tr>\n<tr>\n<td><strong>BCE</strong></td>\n<td>0.10</td>\n<td>Per-voxel calibration, prevents over-confident predictions</td>\n</tr>\n<tr>\n<td><strong>clDice</strong></td>\n<td>0.30</td>\n<td>Centerline Dice — penalizes breaks in thin surface sheets</td>\n</tr>\n<tr>\n<td><strong>Surface</strong></td>\n<td>0.15</td>\n<td>Distance-weighted boundary loss — forces sharper edges</td>\n</tr>\n<tr>\n<td><strong>Topology</strong></td>\n<td>0.20</td>\n<td>Laplacian-based loss — preserves connected components</td>\n</tr>\n</tbody>\n</table>\n<h3>How each loss works</h3>\n<p><strong>clDice (centerline Dice)</strong> — Computes soft skeletonization via iterative min-pooling at half resolution (for speed), then evaluates Dice on the skeleton. A break in a 1-voxel-thick surface sheet destroys the skeleton connectivity, so clDice directly optimizes the topological continuity that TopoScore measures. This loss received the highest weight (0.30) because it was the most effective at improving topology.</p>\n<p><strong>Surface Loss</strong> — Computes GPU-approximate signed distance maps via iterative morphological dilation (5 iterations), then weights prediction errors by distance to the ground truth boundary. This forces the model to get surface boundaries right rather than just filling interiors.</p>\n<p><strong>Topology Loss</strong> — Applies a discrete 3D Laplacian kernel to both prediction and target, then penalizes the difference. The Laplacian highlights topological features (holes, tunnels, component boundaries) — exactly the features that Betti matching and VOI measure. Uses exponential weighting to focus on boundary regions.</p>\n<p><strong>Deep Supervision</strong> — During training, auxiliary Dice loss heads at 3 intermediate decoder resolutions (weights: 0.5, 0.25, 0.125) help gradient flow to deeper layers.</p>\n<hr>\n<h2>4. Training Details</h2>\n<table>\n<thead>\n<tr>\n<th>Setting</th>\n<th>Value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Folds</strong></td>\n<td>3-fold StratifiedKFold (stratified on <code>scroll_id</code>, seed=42)</td>\n</tr>\n<tr>\n<td><strong>Data</strong></td>\n<td>786 valid volumes across 6 scrolls</td>\n</tr>\n<tr>\n<td><strong>Patch size</strong></td>\n<td>192 × 192 × 192</td>\n</tr>\n<tr>\n<td><strong>Batch size</strong></td>\n<td>4</td>\n</tr>\n<tr>\n<td><strong>Optimizer</strong></td>\n<td>AdamW (lr=3e-4, weight_decay=1e-2)</td>\n</tr>\n<tr>\n<td><strong>LR warmup</strong></td>\n<td>5 epochs linear warmup</td>\n</tr>\n<tr>\n<td><strong>LR schedule</strong></td>\n<td>CosineAnnealingLR (T_max=795, eta_min=1e-6)</td>\n</tr>\n<tr>\n<td><strong>Precision</strong></td>\n<td>bfloat16</td>\n</tr>\n<tr>\n<td><strong>Grad clipping</strong></td>\n<td>max_norm=1.0</td>\n</tr>\n<tr>\n<td><strong>Epochs</strong></td>\n<td>800 per fold</td>\n</tr>\n<tr>\n<td><strong>Validation</strong></td>\n<td>Every 5 epochs, patch-based Dice</td>\n</tr>\n</tbody>\n</table>\n<h3>Data Preprocessing</h3>\n<ul>\n<li><strong>Robust Z-score normalization</strong>: Percentile clipping (0.5th – 99.5th percentile) followed by Z-score. Standard medical imaging approach (used by nnU-Net, MONAI). Handles CT artifacts and outlier voxels before computing mean/std.</li>\n<li><strong>Foreground oversampling</strong>: 60% of patches centered on foreground voxels — important because surfaces are sparse (often &lt;10% of volume).</li>\n</ul>\n<h3>GPU Augmentations (all on GPU, no CPU bottleneck)</h3>\n<p>All augmentations run on GPU in bfloat16 using pure PyTorch operations:</p>\n<ul>\n<li><strong>Random 3D flips</strong> (all axes) + <strong>90° rotations</strong> in HW plane</li>\n<li><strong>Elastic deformation</strong>: Low-res random displacement upsampled with trilinear interpolation (σ=2.0 — intentionally mild to avoid breaking thin surfaces)</li>\n<li><strong>Affine scaling</strong>: Uniform(0.9, 1.1)</li>\n<li><strong>Gaussian noise</strong>: σ=0.05</li>\n<li><strong>Contrast/brightness jitter</strong></li>\n<li><strong>Random cuboid occlusion</strong>: Up to 3 small cubes zeroed out per sample</li>\n</ul>\n<p>All spatial augmentations use <code>F.grid_sample</code> for batched GPU processing. Labels use nearest-neighbor interpolation to preserve discrete values.</p>\n<blockquote>\n  <p><strong>Key insight:</strong> Mild augmentation was critical. Heavy elastic deformation or affine transforms break 1–2 voxel thick surfaces, destroying the topology the model needs to learn.</p>\n</blockquote>\n<hr>\n<h2>5. Inference Pipeline</h2>\n<h3>Single Model Inference</h3>\n<p>Used the fold 0 best checkpoint (epoch 249, best Dice 0.5711) for submission.</p>\n<h3>Sliding Window Inference (SWI)</h3>\n<ol>\n<li>Load test volume as float32 TIFF</li>\n<li>Normalize with identical percentile clipping (0.5–99.5%) + Z-score as training</li>\n<li>Pad volume if smaller than patch size (reflect padding)</li>\n<li>Generate overlapping 192×192×192 patch positions with <strong>50% overlap</strong></li>\n<li>Create 3D Gaussian weight kernel (σ = 0.125 × patch_size) for smooth blending at patch boundaries</li>\n<li>Process patches in batches of 2 (one per T4 GPU)</li>\n<li>Accumulate weighted predictions and normalize by weight sum</li>\n</ol>\n<h3>Test-Time Augmentation (TTA)</h3>\n<ul>\n<li><strong>Flip TTA</strong>: Original + flip along Z, Y, X axes = <strong>4 forward passes</strong></li>\n<li>Predictions averaged in probability space before thresholding</li>\n</ul>\n<h3>Multi-GPU Setup</h3>\n<ul>\n<li><code>nn.DataParallel</code> across 2× Tesla T4 (15.6 GB each)</li>\n<li>Batch size 2 (1 patch per GPU) for 192³ patches</li>\n<li>float16 inference for speed</li>\n</ul>\n<hr>\n<h2>6. Post-Processing (Topology-Safe)</h2>\n<p>The key insight: every morphological operation is wrapped in a <strong>topology-safety check</strong>. Before and after each operation, we count 3D connected components (26-connectivity). If the operation would <strong>reduce</strong> the component count (i.e., merge separate surface sheets), it is <strong>reverted</strong>. This is cheap to compute and prevents post-processing from destroying the topology the model learned.</p>\n<h3>Pipeline</h3>\n<ol>\n<li><strong>Fixed threshold at 0.5</strong> — standard sigmoid midpoint</li>\n<li><strong>Remove small components</strong> (&lt;50 voxels) — noise cleanup using 26-connectivity labeling</li>\n<li><strong>2D slicewise closing</strong> (4-connectivity, 1 iteration, topology-safe) — fills small gaps within each slice</li>\n<li><strong>2D slicewise hole filling</strong> (all 3 axes, topology-safe) — fills enclosed holes per-slice</li>\n<li><strong>2D slicewise opening</strong> (4-connectivity, 1 iteration, topology-safe) — removes small protrusions</li>\n<li><strong>Final small component removal</strong> — second cleanup pass</li>\n</ol>\n<blockquote>\n  <p><strong>Critical design choice: 2D slicewise, not 3D.</strong> All morphology is done slice-by-slice independently. 3D morphological operations with 3D structuring elements are too aggressive and destroy 1-voxel-thick surfaces. 2D slicewise processing is much gentler and preserves thin structures.</p>\n</blockquote>\n<hr>\n<h2>7. What Worked ✅</h2>\n<ol>\n<li><strong>Topology-aware loss stack</strong> — clDice(0.30) + Topology(0.20) directly optimize what the competition metric measures. Without them, the model learns decent Dice but poor Betti matching and VOI scores.</li>\n<li><strong>HybridConv3d</strong> — Decoupled XY/Z convolutions respect CT anisotropy. More effective than isotropic 3×3×3 convolutions when in-plane and cross-slice resolutions differ.</li>\n<li><strong>SurfaceRefinementBlock</strong> — Edge-aware processing at full decoder resolution is essential for 1–2 voxel thick surfaces. The <code>|conv(x)|</code> edge detection branch gives the model an explicit edge signal to refine.</li>\n<li><strong>2D slicewise post-processing</strong> — Much gentler than 3D morphology. Thin surface structures survive.</li>\n<li><strong>Topology-safe operation guard</strong> — Simple component-count check before/after each morphological operation prevents accidental merging of separate sheets.</li>\n<li><strong>GPU augmentations</strong> — All augmentations on GPU with <code>F.grid_sample</code> eliminates CPU bottleneck. Training speed limited only by forward/backward pass.</li>\n<li><strong>bfloat16 training</strong> — Same quality as float32, 2× memory savings, no GradScaler complexity.</li>\n<li><strong>Foreground oversampling (60%)</strong> — Critical when surfaces are &lt;10% of volume. Without it, the model sees mostly background patches.</li>\n</ol>\n<hr>\n<h2>8. What Didn't Work ❌</h2>\n<ol>\n<li><strong>Frangi vesselness filter</strong> in post-processing — Designed for tubular structures, not sheets. Reduced mean prediction confidence from 0.118 to 0.089, hurting overall score.</li>\n<li><strong>Adaptive thresholding</strong> (e.g., 0.30) — Created too many false positive regions. The fixed 0.5 threshold was consistently better.</li>\n<li><strong>3D morphological operations</strong> — Too aggressive for 1-voxel-thick surfaces. Switching to 2D slicewise processing was a critical improvement.</li>\n<li><strong>Heavy augmentation</strong> — Strong elastic deformation (σ &gt; 5) or large affine transforms broke thin surface topology during training, making the model worse at preserving connectivity.</li>\n</ol>\n<hr>\n<h2>9. Hardware &amp; Runtime</h2>\n<table>\n<thead>\n<tr>\n<th>Stage</th>\n<th>Hardware</th>\n<th>Time</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Training (per fold)</td>\n<td>NVIDIA H100 80GB</td>\n<td>800 epochs, ~10–15 min/epoch</td>\n</tr>\n<tr>\n<td>Inference (per volume)</td>\n<td>2× Tesla T4 (DataParallel)</td>\n<td>~2–3 min with flip TTA</td>\n</tr>\n</tbody>\n</table>\n<p>Training was done on Vast.ai with H100 GPU. Inference runs on 2× T4 within the competition runtime limits.</p>\n<hr>\n<h2>10. Code</h2>\n<p>Both notebooks are fully self-contained single-file implementations. No external repositories or custom packages required.</p>\n<table>\n<thead>\n<tr>\n<th>Notebook</th>\n<th>What it does</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Training</strong></td>\n<td>Full <code>TopoPreservingUNet3D</code> architecture + 5-loss stack + GPU augmentations + 3-fold stratified CV + checkpoint management</td>\n</tr>\n<tr>\n<td><strong>Inference</strong></td>\n<td>Model loading + sliding window inference with Gaussian blending + 3-axis flip TTA + topology-safe 2D post-processing + TIFF submission</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Dependencies:</strong> PyTorch, NumPy, SciPy, scikit-image, tifffile, pandas</p>\n<hr>\n<p><em>Thanks for reading!</em></p>",
  "messages": [
    {
      "id": 3414979,
      "postDate": "2026-02-28T02:47:08.400Z",
      "content": "<h2>1. Problem Understanding</h2>\n<p>The Vesuvius 2025 Surface Detection task requires segmenting thin papyrus surfaces from 3D CT volumes of ancient scrolls. The competition metric is:</p>\n<p><strong>LB = 0.30 × TopoScore + 0.35 × SurfaceDice + 0.35 × VOI</strong></p>\n<p>This means <strong>topology matters as much as pixel overlap</strong>. A model that gets good Dice but merges or splits surface sheets will score poorly on Betti matching and VOI. This insight drove every design decision — from the loss function to the post-processing.</p>\n<hr>\n<h2>2. Architecture: TopoPreservingUNet3D (~10M params)</h2>\n<p>Custom 3D U-Net with 6 encoder/decoder stages designed specifically for thin surface detection in CT volumes.</p>\n<p><strong>Feature channels:</strong> <code>[32, 64, 128, 256, 320, 320]</code><br>\n<strong>Residual blocks per stage:</strong> <code>[1, 2, 3, 4, 6, 6]</code></p>\n<h3>Key Design Choices</h3>\n<p><strong>HybridConv3d</strong> — Instead of standard <code>3×3×3</code> convolutions, I decouple XY and Z processing: <code>Conv3d(kernel=(1,3,3))</code> for in-plane features and <code>Conv3d(kernel=(3,1,1))</code> for cross-slice features, then concatenate. CT data is anisotropic — in-plane resolution typically differs from slice spacing — so decoupled processing respects the data geometry.</p>\n<p><strong>MultiScaleResBlock</strong> — Res2Net-style blocks where the channel dimension is split into groups that process hierarchically. Each group receives input from the previous group, building multi-scale receptive fields within a single residual block. Applied at every encoder and decoder stage.</p>\n<p><strong>AttentionBlock (stages 2–5)</strong> — Combined channel attention (squeeze-excitation via <code>AdaptiveAvgPool3d</code>) and spatial attention (<code>7×7</code> conv on concatenated mean+max pooled features). Helps the model focus on surface regions while suppressing irrelevant background.</p>\n<p><strong>SurfaceRefinementBlock (decoder stage 0)</strong> — This was critical for performance. At the highest resolution decoder stage, an edge convolution branch computes <code>|conv(x)|</code> to detect edges, concatenates with original features, and refines through two conv-norm-activation layers. Surfaces in the Vesuvius data can be <strong>1–2 voxels thick</strong>, so having dedicated edge-aware processing at full resolution made a noticeable difference.</p>\n<p><strong>Other details:</strong></p>\n<ul>\n<li><code>GroupNorm</code> throughout (stable with small batch size 4, unlike <code>BatchNorm</code>)</li>\n<li><code>LeakyReLU(0.01)</code> activation</li>\n<li>Strided convolution downsampling (learnable, instead of max-pool)</li>\n<li><code>ConvTranspose3d</code> upsampling</li>\n<li>Deep supervision heads at 3 decoder levels during training</li>\n</ul>\n<hr>\n<h2>3. Loss Function: 5-Component Topology-Aware Stack</h2>\n<p>The competition metric penalizes topological errors (Betti matching, VOI), so I designed the loss function with topology preservation as a first-class objective. All 5 losses are active from epoch 0.</p>\n<table>\n<thead>\n<tr>\n<th>Loss</th>\n<th>Weight</th>\n<th>Purpose</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Dice</strong></td>\n<td>0.25</td>\n<td>Core overlap signal for segmentation</td>\n</tr>\n<tr>\n<td><strong>BCE</strong></td>\n<td>0.10</td>\n<td>Per-voxel calibration, prevents over-confident predictions</td>\n</tr>\n<tr>\n<td><strong>clDice</strong></td>\n<td>0.30</td>\n<td>Centerline Dice — penalizes breaks in thin surface sheets</td>\n</tr>\n<tr>\n<td><strong>Surface</strong></td>\n<td>0.15</td>\n<td>Distance-weighted boundary loss — forces sharper edges</td>\n</tr>\n<tr>\n<td><strong>Topology</strong></td>\n<td>0.20</td>\n<td>Laplacian-based loss — preserves connected components</td>\n</tr>\n</tbody>\n</table>\n<h3>How each loss works</h3>\n<p><strong>clDice (centerline Dice)</strong> — Computes soft skeletonization via iterative min-pooling at half resolution (for speed), then evaluates Dice on the skeleton. A break in a 1-voxel-thick surface sheet destroys the skeleton connectivity, so clDice directly optimizes the topological continuity that TopoScore measures. This loss received the highest weight (0.30) because it was the most effective at improving topology.</p>\n<p><strong>Surface Loss</strong> — Computes GPU-approximate signed distance maps via iterative morphological dilation (5 iterations), then weights prediction errors by distance to the ground truth boundary. This forces the model to get surface boundaries right rather than just filling interiors.</p>\n<p><strong>Topology Loss</strong> — Applies a discrete 3D Laplacian kernel to both prediction and target, then penalizes the difference. The Laplacian highlights topological features (holes, tunnels, component boundaries) — exactly the features that Betti matching and VOI measure. Uses exponential weighting to focus on boundary regions.</p>\n<p><strong>Deep Supervision</strong> — During training, auxiliary Dice loss heads at 3 intermediate decoder resolutions (weights: 0.5, 0.25, 0.125) help gradient flow to deeper layers.</p>\n<hr>\n<h2>4. Training Details</h2>\n<table>\n<thead>\n<tr>\n<th>Setting</th>\n<th>Value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Folds</strong></td>\n<td>3-fold StratifiedKFold (stratified on <code>scroll_id</code>, seed=42)</td>\n</tr>\n<tr>\n<td><strong>Data</strong></td>\n<td>786 valid volumes across 6 scrolls</td>\n</tr>\n<tr>\n<td><strong>Patch size</strong></td>\n<td>192 × 192 × 192</td>\n</tr>\n<tr>\n<td><strong>Batch size</strong></td>\n<td>4</td>\n</tr>\n<tr>\n<td><strong>Optimizer</strong></td>\n<td>AdamW (lr=3e-4, weight_decay=1e-2)</td>\n</tr>\n<tr>\n<td><strong>LR warmup</strong></td>\n<td>5 epochs linear warmup</td>\n</tr>\n<tr>\n<td><strong>LR schedule</strong></td>\n<td>CosineAnnealingLR (T_max=795, eta_min=1e-6)</td>\n</tr>\n<tr>\n<td><strong>Precision</strong></td>\n<td>bfloat16</td>\n</tr>\n<tr>\n<td><strong>Grad clipping</strong></td>\n<td>max_norm=1.0</td>\n</tr>\n<tr>\n<td><strong>Epochs</strong></td>\n<td>800 per fold</td>\n</tr>\n<tr>\n<td><strong>Validation</strong></td>\n<td>Every 5 epochs, patch-based Dice</td>\n</tr>\n</tbody>\n</table>\n<h3>Data Preprocessing</h3>\n<ul>\n<li><strong>Robust Z-score normalization</strong>: Percentile clipping (0.5th – 99.5th percentile) followed by Z-score. Standard medical imaging approach (used by nnU-Net, MONAI). Handles CT artifacts and outlier voxels before computing mean/std.</li>\n<li><strong>Foreground oversampling</strong>: 60% of patches centered on foreground voxels — important because surfaces are sparse (often &lt;10% of volume).</li>\n</ul>\n<h3>GPU Augmentations (all on GPU, no CPU bottleneck)</h3>\n<p>All augmentations run on GPU in bfloat16 using pure PyTorch operations:</p>\n<ul>\n<li><strong>Random 3D flips</strong> (all axes) + <strong>90° rotations</strong> in HW plane</li>\n<li><strong>Elastic deformation</strong>: Low-res random displacement upsampled with trilinear interpolation (σ=2.0 — intentionally mild to avoid breaking thin surfaces)</li>\n<li><strong>Affine scaling</strong>: Uniform(0.9, 1.1)</li>\n<li><strong>Gaussian noise</strong>: σ=0.05</li>\n<li><strong>Contrast/brightness jitter</strong></li>\n<li><strong>Random cuboid occlusion</strong>: Up to 3 small cubes zeroed out per sample</li>\n</ul>\n<p>All spatial augmentations use <code>F.grid_sample</code> for batched GPU processing. Labels use nearest-neighbor interpolation to preserve discrete values.</p>\n<blockquote>\n  <p><strong>Key insight:</strong> Mild augmentation was critical. Heavy elastic deformation or affine transforms break 1–2 voxel thick surfaces, destroying the topology the model needs to learn.</p>\n</blockquote>\n<hr>\n<h2>5. Inference Pipeline</h2>\n<h3>Single Model Inference</h3>\n<p>Used the fold 0 best checkpoint (epoch 249, best Dice 0.5711) for submission.</p>\n<h3>Sliding Window Inference (SWI)</h3>\n<ol>\n<li>Load test volume as float32 TIFF</li>\n<li>Normalize with identical percentile clipping (0.5–99.5%) + Z-score as training</li>\n<li>Pad volume if smaller than patch size (reflect padding)</li>\n<li>Generate overlapping 192×192×192 patch positions with <strong>50% overlap</strong></li>\n<li>Create 3D Gaussian weight kernel (σ = 0.125 × patch_size) for smooth blending at patch boundaries</li>\n<li>Process patches in batches of 2 (one per T4 GPU)</li>\n<li>Accumulate weighted predictions and normalize by weight sum</li>\n</ol>\n<h3>Test-Time Augmentation (TTA)</h3>\n<ul>\n<li><strong>Flip TTA</strong>: Original + flip along Z, Y, X axes = <strong>4 forward passes</strong></li>\n<li>Predictions averaged in probability space before thresholding</li>\n</ul>\n<h3>Multi-GPU Setup</h3>\n<ul>\n<li><code>nn.DataParallel</code> across 2× Tesla T4 (15.6 GB each)</li>\n<li>Batch size 2 (1 patch per GPU) for 192³ patches</li>\n<li>float16 inference for speed</li>\n</ul>\n<hr>\n<h2>6. Post-Processing (Topology-Safe)</h2>\n<p>The key insight: every morphological operation is wrapped in a <strong>topology-safety check</strong>. Before and after each operation, we count 3D connected components (26-connectivity). If the operation would <strong>reduce</strong> the component count (i.e., merge separate surface sheets), it is <strong>reverted</strong>. This is cheap to compute and prevents post-processing from destroying the topology the model learned.</p>\n<h3>Pipeline</h3>\n<ol>\n<li><strong>Fixed threshold at 0.5</strong> — standard sigmoid midpoint</li>\n<li><strong>Remove small components</strong> (&lt;50 voxels) — noise cleanup using 26-connectivity labeling</li>\n<li><strong>2D slicewise closing</strong> (4-connectivity, 1 iteration, topology-safe) — fills small gaps within each slice</li>\n<li><strong>2D slicewise hole filling</strong> (all 3 axes, topology-safe) — fills enclosed holes per-slice</li>\n<li><strong>2D slicewise opening</strong> (4-connectivity, 1 iteration, topology-safe) — removes small protrusions</li>\n<li><strong>Final small component removal</strong> — second cleanup pass</li>\n</ol>\n<blockquote>\n  <p><strong>Critical design choice: 2D slicewise, not 3D.</strong> All morphology is done slice-by-slice independently. 3D morphological operations with 3D structuring elements are too aggressive and destroy 1-voxel-thick surfaces. 2D slicewise processing is much gentler and preserves thin structures.</p>\n</blockquote>\n<hr>\n<h2>7. What Worked ✅</h2>\n<ol>\n<li><strong>Topology-aware loss stack</strong> — clDice(0.30) + Topology(0.20) directly optimize what the competition metric measures. Without them, the model learns decent Dice but poor Betti matching and VOI scores.</li>\n<li><strong>HybridConv3d</strong> — Decoupled XY/Z convolutions respect CT anisotropy. More effective than isotropic 3×3×3 convolutions when in-plane and cross-slice resolutions differ.</li>\n<li><strong>SurfaceRefinementBlock</strong> — Edge-aware processing at full decoder resolution is essential for 1–2 voxel thick surfaces. The <code>|conv(x)|</code> edge detection branch gives the model an explicit edge signal to refine.</li>\n<li><strong>2D slicewise post-processing</strong> — Much gentler than 3D morphology. Thin surface structures survive.</li>\n<li><strong>Topology-safe operation guard</strong> — Simple component-count check before/after each morphological operation prevents accidental merging of separate sheets.</li>\n<li><strong>GPU augmentations</strong> — All augmentations on GPU with <code>F.grid_sample</code> eliminates CPU bottleneck. Training speed limited only by forward/backward pass.</li>\n<li><strong>bfloat16 training</strong> — Same quality as float32, 2× memory savings, no GradScaler complexity.</li>\n<li><strong>Foreground oversampling (60%)</strong> — Critical when surfaces are &lt;10% of volume. Without it, the model sees mostly background patches.</li>\n</ol>\n<hr>\n<h2>8. What Didn't Work ❌</h2>\n<ol>\n<li><strong>Frangi vesselness filter</strong> in post-processing — Designed for tubular structures, not sheets. Reduced mean prediction confidence from 0.118 to 0.089, hurting overall score.</li>\n<li><strong>Adaptive thresholding</strong> (e.g., 0.30) — Created too many false positive regions. The fixed 0.5 threshold was consistently better.</li>\n<li><strong>3D morphological operations</strong> — Too aggressive for 1-voxel-thick surfaces. Switching to 2D slicewise processing was a critical improvement.</li>\n<li><strong>Heavy augmentation</strong> — Strong elastic deformation (σ &gt; 5) or large affine transforms broke thin surface topology during training, making the model worse at preserving connectivity.</li>\n</ol>\n<hr>\n<h2>9. Hardware &amp; Runtime</h2>\n<table>\n<thead>\n<tr>\n<th>Stage</th>\n<th>Hardware</th>\n<th>Time</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Training (per fold)</td>\n<td>NVIDIA H100 80GB</td>\n<td>800 epochs, ~10–15 min/epoch</td>\n</tr>\n<tr>\n<td>Inference (per volume)</td>\n<td>2× Tesla T4 (DataParallel)</td>\n<td>~2–3 min with flip TTA</td>\n</tr>\n</tbody>\n</table>\n<p>Training was done on Vast.ai with H100 GPU. Inference runs on 2× T4 within the competition runtime limits.</p>\n<hr>\n<h2>10. Code</h2>\n<p>Both notebooks are fully self-contained single-file implementations. No external repositories or custom packages required.</p>\n<table>\n<thead>\n<tr>\n<th>Notebook</th>\n<th>What it does</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Training</strong></td>\n<td>Full <code>TopoPreservingUNet3D</code> architecture + 5-loss stack + GPU augmentations + 3-fold stratified CV + checkpoint management</td>\n</tr>\n<tr>\n<td><strong>Inference</strong></td>\n<td>Model loading + sliding window inference with Gaussian blending + 3-axis flip TTA + topology-safe 2D post-processing + TIFF submission</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Dependencies:</strong> PyTorch, NumPy, SciPy, scikit-image, tifffile, pandas</p>\n<hr>\n<p><em>Thanks for reading!</em></p>",
      "rawMarkdown": "## 1. Problem Understanding\n\nThe Vesuvius 2025 Surface Detection task requires segmenting thin papyrus surfaces from 3D CT volumes of ancient scrolls. The competition metric is:\n\n**LB = 0.30 × TopoScore + 0.35 × SurfaceDice + 0.35 × VOI**\n\nThis means **topology matters as much as pixel overlap**. A model that gets good Dice but merges or splits surface sheets will score poorly on Betti matching and VOI. This insight drove every design decision — from the loss function to the post-processing.\n\n---\n\n## 2. Architecture: TopoPreservingUNet3D (~10M params)\n\nCustom 3D U-Net with 6 encoder/decoder stages designed specifically for thin surface detection in CT volumes.\n\n**Feature channels:** `[32, 64, 128, 256, 320, 320]`  \n**Residual blocks per stage:** `[1, 2, 3, 4, 6, 6]`\n\n### Key Design Choices\n\n**HybridConv3d** — Instead of standard `3×3×3` convolutions, I decouple XY and Z processing: `Conv3d(kernel=(1,3,3))` for in-plane features and `Conv3d(kernel=(3,1,1))` for cross-slice features, then concatenate. CT data is anisotropic — in-plane resolution typically differs from slice spacing — so decoupled processing respects the data geometry.\n\n**MultiScaleResBlock** — Res2Net-style blocks where the channel dimension is split into groups that process hierarchically. Each group receives input from the previous group, building multi-scale receptive fields within a single residual block. Applied at every encoder and decoder stage.\n\n**AttentionBlock (stages 2–5)** — Combined channel attention (squeeze-excitation via `AdaptiveAvgPool3d`) and spatial attention (`7×7` conv on concatenated mean+max pooled features). Helps the model focus on surface regions while suppressing irrelevant background.\n\n**SurfaceRefinementBlock (decoder stage 0)** — This was critical for performance. At the highest resolution decoder stage, an edge convolution branch computes `|conv(x)|` to detect edges, concatenates with original features, and refines through two conv-norm-activation layers. Surfaces in the Vesuvius data can be **1–2 voxels thick**, so having dedicated edge-aware processing at full resolution made a noticeable difference.\n\n**Other details:**\n- `GroupNorm` throughout (stable with small batch size 4, unlike `BatchNorm`)\n- `LeakyReLU(0.01)` activation\n- Strided convolution downsampling (learnable, instead of max-pool)\n- `ConvTranspose3d` upsampling\n- Deep supervision heads at 3 decoder levels during training\n\n---\n\n## 3. Loss Function: 5-Component Topology-Aware Stack\n\nThe competition metric penalizes topological errors (Betti matching, VOI), so I designed the loss function with topology preservation as a first-class objective. All 5 losses are active from epoch 0.\n\n| Loss | Weight | Purpose |\n|:-----|:------:|:--------|\n| **Dice** | 0.25 | Core overlap signal for segmentation |\n| **BCE** | 0.10 | Per-voxel calibration, prevents over-confident predictions |\n| **clDice** | 0.30 | Centerline Dice — penalizes breaks in thin surface sheets |\n| **Surface** | 0.15 | Distance-weighted boundary loss — forces sharper edges |\n| **Topology** | 0.20 | Laplacian-based loss — preserves connected components |\n\n### How each loss works\n\n**clDice (centerline Dice)** — Computes soft skeletonization via iterative min-pooling at half resolution (for speed), then evaluates Dice on the skeleton. A break in a 1-voxel-thick surface sheet destroys the skeleton connectivity, so clDice directly optimizes the topological continuity that TopoScore measures. This loss received the highest weight (0.30) because it was the most effective at improving topology.\n\n**Surface Loss** — Computes GPU-approximate signed distance maps via iterative morphological dilation (5 iterations), then weights prediction errors by distance to the ground truth boundary. This forces the model to get surface boundaries right rather than just filling interiors.\n\n**Topology Loss** — Applies a discrete 3D Laplacian kernel to both prediction and target, then penalizes the difference. The Laplacian highlights topological features (holes, tunnels, component boundaries) — exactly the features that Betti matching and VOI measure. Uses exponential weighting to focus on boundary regions.\n\n**Deep Supervision** — During training, auxiliary Dice loss heads at 3 intermediate decoder resolutions (weights: 0.5, 0.25, 0.125) help gradient flow to deeper layers.\n\n---\n\n## 4. Training Details\n\n| Setting | Value |\n|:--------|:------|\n| **Folds** | 3-fold StratifiedKFold (stratified on `scroll_id`, seed=42) |\n| **Data** | 786 valid volumes across 6 scrolls |\n| **Patch size** | 192 × 192 × 192 |\n| **Batch size** | 4 |\n| **Optimizer** | AdamW (lr=3e-4, weight\\_decay=1e-2) |\n| **LR warmup** | 5 epochs linear warmup |\n| **LR schedule** | CosineAnnealingLR (T\\_max=795, eta\\_min=1e-6) |\n| **Precision** | bfloat16 |\n| **Grad clipping** | max\\_norm=1.0 |\n| **Epochs** | 800 per fold |\n| **Validation** | Every 5 epochs, patch-based Dice |\n\n### Data Preprocessing\n\n- **Robust Z-score normalization**: Percentile clipping (0.5th – 99.5th percentile) followed by Z-score. Standard medical imaging approach (used by nnU-Net, MONAI). Handles CT artifacts and outlier voxels before computing mean/std.\n- **Foreground oversampling**: 60% of patches centered on foreground voxels — important because surfaces are sparse (often <10% of volume).\n\n### GPU Augmentations (all on GPU, no CPU bottleneck)\n\nAll augmentations run on GPU in bfloat16 using pure PyTorch operations:\n\n- **Random 3D flips** (all axes) + **90° rotations** in HW plane\n- **Elastic deformation**: Low-res random displacement upsampled with trilinear interpolation (σ=2.0 — intentionally mild to avoid breaking thin surfaces)\n- **Affine scaling**: Uniform(0.9, 1.1)\n- **Gaussian noise**: σ=0.05\n- **Contrast/brightness jitter**\n- **Random cuboid occlusion**: Up to 3 small cubes zeroed out per sample\n\nAll spatial augmentations use `F.grid_sample` for batched GPU processing. Labels use nearest-neighbor interpolation to preserve discrete values.\n\n> **Key insight:** Mild augmentation was critical. Heavy elastic deformation or affine transforms break 1–2 voxel thick surfaces, destroying the topology the model needs to learn.\n\n---\n\n## 5. Inference Pipeline\n\n### Single Model Inference\n\nUsed the fold 0 best checkpoint (epoch 249, best Dice 0.5711) for submission.\n\n### Sliding Window Inference (SWI)\n\n1. Load test volume as float32 TIFF\n2. Normalize with identical percentile clipping (0.5–99.5%) + Z-score as training\n3. Pad volume if smaller than patch size (reflect padding)\n4. Generate overlapping 192×192×192 patch positions with **50% overlap**\n5. Create 3D Gaussian weight kernel (σ = 0.125 × patch\\_size) for smooth blending at patch boundaries\n6. Process patches in batches of 2 (one per T4 GPU)\n7. Accumulate weighted predictions and normalize by weight sum\n\n### Test-Time Augmentation (TTA)\n\n- **Flip TTA**: Original + flip along Z, Y, X axes = **4 forward passes**\n- Predictions averaged in probability space before thresholding\n\n### Multi-GPU Setup\n\n- `nn.DataParallel` across 2× Tesla T4 (15.6 GB each)\n- Batch size 2 (1 patch per GPU) for 192³ patches\n- float16 inference for speed\n\n---\n\n## 6. Post-Processing (Topology-Safe)\n\nThe key insight: every morphological operation is wrapped in a **topology-safety check**. Before and after each operation, we count 3D connected components (26-connectivity). If the operation would **reduce** the component count (i.e., merge separate surface sheets), it is **reverted**. This is cheap to compute and prevents post-processing from destroying the topology the model learned.\n\n### Pipeline\n\n1. **Fixed threshold at 0.5** — standard sigmoid midpoint\n2. **Remove small components** (<50 voxels) — noise cleanup using 26-connectivity labeling\n3. **2D slicewise closing** (4-connectivity, 1 iteration, topology-safe) — fills small gaps within each slice\n4. **2D slicewise hole filling** (all 3 axes, topology-safe) — fills enclosed holes per-slice\n5. **2D slicewise opening** (4-connectivity, 1 iteration, topology-safe) — removes small protrusions\n6. **Final small component removal** — second cleanup pass\n\n> **Critical design choice: 2D slicewise, not 3D.** All morphology is done slice-by-slice independently. 3D morphological operations with 3D structuring elements are too aggressive and destroy 1-voxel-thick surfaces. 2D slicewise processing is much gentler and preserves thin structures.\n\n---\n\n## 7. What Worked ✅\n\n1. **Topology-aware loss stack** — clDice(0.30) + Topology(0.20) directly optimize what the competition metric measures. Without them, the model learns decent Dice but poor Betti matching and VOI scores.\n2. **HybridConv3d** — Decoupled XY/Z convolutions respect CT anisotropy. More effective than isotropic 3×3×3 convolutions when in-plane and cross-slice resolutions differ.\n3. **SurfaceRefinementBlock** — Edge-aware processing at full decoder resolution is essential for 1–2 voxel thick surfaces. The `|conv(x)|` edge detection branch gives the model an explicit edge signal to refine.\n4. **2D slicewise post-processing** — Much gentler than 3D morphology. Thin surface structures survive.\n5. **Topology-safe operation guard** — Simple component-count check before/after each morphological operation prevents accidental merging of separate sheets.\n6. **GPU augmentations** — All augmentations on GPU with `F.grid_sample` eliminates CPU bottleneck. Training speed limited only by forward/backward pass.\n7. **bfloat16 training** — Same quality as float32, 2× memory savings, no GradScaler complexity.\n8. **Foreground oversampling (60%)** — Critical when surfaces are <10% of volume. Without it, the model sees mostly background patches.\n\n---\n\n## 8. What Didn't Work ❌\n\n1. **Frangi vesselness filter** in post-processing — Designed for tubular structures, not sheets. Reduced mean prediction confidence from 0.118 to 0.089, hurting overall score.\n2. **Adaptive thresholding** (e.g., 0.30) — Created too many false positive regions. The fixed 0.5 threshold was consistently better.\n3. **3D morphological operations** — Too aggressive for 1-voxel-thick surfaces. Switching to 2D slicewise processing was a critical improvement.\n4. **Heavy augmentation** — Strong elastic deformation (σ > 5) or large affine transforms broke thin surface topology during training, making the model worse at preserving connectivity.\n\n---\n\n## 9. Hardware & Runtime\n\n| Stage | Hardware | Time |\n|:------|:---------|:-----|\n| Training (per fold) | NVIDIA H100 80GB | 800 epochs, ~10–15 min/epoch |\n| Inference (per volume) | 2× Tesla T4 (DataParallel) | ~2–3 min with flip TTA |\n\nTraining was done on Vast.ai with H100 GPU. Inference runs on 2× T4 within the competition runtime limits.\n\n---\n\n## 10. Code\n\nBoth notebooks are fully self-contained single-file implementations. No external repositories or custom packages required.\n\n| Notebook | What it does |\n|:---------|:-------------|\n| **Training** | Full `TopoPreservingUNet3D` architecture + 5-loss stack + GPU augmentations + 3-fold stratified CV + checkpoint management |\n| **Inference** | Model loading + sliding window inference with Gaussian blending + 3-axis flip TTA + topology-safe 2D post-processing + TIFF submission |\n\n**Dependencies:** PyTorch, NumPy, SciPy, scikit-image, tifffile, pandas\n\n---\n\n*Thanks for reading!*\n",
      "votes": 5
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3414979": "## 1. Problem Understanding\n\nThe Vesuvius 2025 Surface Detection task requires segmenting thin papyrus surfaces from 3D CT volumes of ancient scrolls. The competition metric is:\n\n**LB = 0.30 × TopoScore + 0.35 × SurfaceDice + 0.35 × VOI**\n\nThis means **topology matters as much as pixel overlap**. A model that gets good Dice but merges or splits surface sheets will score poorly on Betti matching and VOI. This insight drove every design decision — from the loss function to the post-processing.\n\n---\n\n## 2. Architecture: TopoPreservingUNet3D (~10M params)\n\nCustom 3D U-Net with 6 encoder/decoder stages designed specifically for thin surface detection in CT volumes.\n\n**Feature channels:** `[32, 64, 128, 256, 320, 320]`  \n**Residual blocks per stage:** `[1, 2, 3, 4, 6, 6]`\n\n### Key Design Choices\n\n**HybridConv3d** — Instead of standard `3×3×3` convolutions, I decouple XY and Z processing: `Conv3d(kernel=(1,3,3))` for in-plane features and `Conv3d(kernel=(3,1,1))` for cross-slice features, then concatenate. CT data is anisotropic — in-plane resolution typically differs from slice spacing — so decoupled processing respects the data geometry.\n\n**MultiScaleResBlock** — Res2Net-style blocks where the channel dimension is split into groups that process hierarchically. Each group receives input from the previous group, building multi-scale receptive fields within a single residual block. Applied at every encoder and decoder stage.\n\n**AttentionBlock (stages 2–5)** — Combined channel attention (squeeze-excitation via `AdaptiveAvgPool3d`) and spatial attention (`7×7` conv on concatenated mean+max pooled features). Helps the model focus on surface regions while suppressing irrelevant background.\n\n**SurfaceRefinementBlock (decoder stage 0)** — This was critical for performance. At the highest resolution decoder stage, an edge convolution branch computes `|conv(x)|` to detect edges, concatenates with original features, and refines through two conv-norm-activation layers. Surfaces in the Vesuvius data can be **1–2 voxels thick**, so having dedicated edge-aware processing at full resolution made a noticeable difference.\n\n**Other details:**\n- `GroupNorm` throughout (stable with small batch size 4, unlike `BatchNorm`)\n- `LeakyReLU(0.01)` activation\n- Strided convolution downsampling (learnable, instead of max-pool)\n- `ConvTranspose3d` upsampling\n- Deep supervision heads at 3 decoder levels during training\n\n---\n\n## 3. Loss Function: 5-Component Topology-Aware Stack\n\nThe competition metric penalizes topological errors (Betti matching, VOI), so I designed the loss function with topology preservation as a first-class objective. All 5 losses are active from epoch 0.\n\n| Loss | Weight | Purpose |\n|:-----|:------:|:--------|\n| **Dice** | 0.25 | Core overlap signal for segmentation |\n| **BCE** | 0.10 | Per-voxel calibration, prevents over-confident predictions |\n| **clDice** | 0.30 | Centerline Dice — penalizes breaks in thin surface sheets |\n| **Surface** | 0.15 | Distance-weighted boundary loss — forces sharper edges |\n| **Topology** | 0.20 | Laplacian-based loss — preserves connected components |\n\n### How each loss works\n\n**clDice (centerline Dice)** — Computes soft skeletonization via iterative min-pooling at half resolution (for speed), then evaluates Dice on the skeleton. A break in a 1-voxel-thick surface sheet destroys the skeleton connectivity, so clDice directly optimizes the topological continuity that TopoScore measures. This loss received the highest weight (0.30) because it was the most effective at improving topology.\n\n**Surface Loss** — Computes GPU-approximate signed distance maps via iterative morphological dilation (5 iterations), then weights prediction errors by distance to the ground truth boundary. This forces the model to get surface boundaries right rather than just filling interiors.\n\n**Topology Loss** — Applies a discrete 3D Laplacian kernel to both prediction and target, then penalizes the difference. The Laplacian highlights topological features (holes, tunnels, component boundaries) — exactly the features that Betti matching and VOI measure. Uses exponential weighting to focus on boundary regions.\n\n**Deep Supervision** — During training, auxiliary Dice loss heads at 3 intermediate decoder resolutions (weights: 0.5, 0.25, 0.125) help gradient flow to deeper layers.\n\n---\n\n## 4. Training Details\n\n| Setting | Value |\n|:--------|:------|\n| **Folds** | 3-fold StratifiedKFold (stratified on `scroll_id`, seed=42) |\n| **Data** | 786 valid volumes across 6 scrolls |\n| **Patch size** | 192 × 192 × 192 |\n| **Batch size** | 4 |\n| **Optimizer** | AdamW (lr=3e-4, weight\\_decay=1e-2) |\n| **LR warmup** | 5 epochs linear warmup |\n| **LR schedule** | CosineAnnealingLR (T\\_max=795, eta\\_min=1e-6) |\n| **Precision** | bfloat16 |\n| **Grad clipping** | max\\_norm=1.0 |\n| **Epochs** | 800 per fold |\n| **Validation** | Every 5 epochs, patch-based Dice |\n\n### Data Preprocessing\n\n- **Robust Z-score normalization**: Percentile clipping (0.5th – 99.5th percentile) followed by Z-score. Standard medical imaging approach (used by nnU-Net, MONAI). Handles CT artifacts and outlier voxels before computing mean/std.\n- **Foreground oversampling**: 60% of patches centered on foreground voxels — important because surfaces are sparse (often <10% of volume).\n\n### GPU Augmentations (all on GPU, no CPU bottleneck)\n\nAll augmentations run on GPU in bfloat16 using pure PyTorch operations:\n\n- **Random 3D flips** (all axes) + **90° rotations** in HW plane\n- **Elastic deformation**: Low-res random displacement upsampled with trilinear interpolation (σ=2.0 — intentionally mild to avoid breaking thin surfaces)\n- **Affine scaling**: Uniform(0.9, 1.1)\n- **Gaussian noise**: σ=0.05\n- **Contrast/brightness jitter**\n- **Random cuboid occlusion**: Up to 3 small cubes zeroed out per sample\n\nAll spatial augmentations use `F.grid_sample` for batched GPU processing. Labels use nearest-neighbor interpolation to preserve discrete values.\n\n> **Key insight:** Mild augmentation was critical. Heavy elastic deformation or affine transforms break 1–2 voxel thick surfaces, destroying the topology the model needs to learn.\n\n---\n\n## 5. Inference Pipeline\n\n### Single Model Inference\n\nUsed the fold 0 best checkpoint (epoch 249, best Dice 0.5711) for submission.\n\n### Sliding Window Inference (SWI)\n\n1. Load test volume as float32 TIFF\n2. Normalize with identical percentile clipping (0.5–99.5%) + Z-score as training\n3. Pad volume if smaller than patch size (reflect padding)\n4. Generate overlapping 192×192×192 patch positions with **50% overlap**\n5. Create 3D Gaussian weight kernel (σ = 0.125 × patch\\_size) for smooth blending at patch boundaries\n6. Process patches in batches of 2 (one per T4 GPU)\n7. Accumulate weighted predictions and normalize by weight sum\n\n### Test-Time Augmentation (TTA)\n\n- **Flip TTA**: Original + flip along Z, Y, X axes = **4 forward passes**\n- Predictions averaged in probability space before thresholding\n\n### Multi-GPU Setup\n\n- `nn.DataParallel` across 2× Tesla T4 (15.6 GB each)\n- Batch size 2 (1 patch per GPU) for 192³ patches\n- float16 inference for speed\n\n---\n\n## 6. Post-Processing (Topology-Safe)\n\nThe key insight: every morphological operation is wrapped in a **topology-safety check**. Before and after each operation, we count 3D connected components (26-connectivity). If the operation would **reduce** the component count (i.e., merge separate surface sheets), it is **reverted**. This is cheap to compute and prevents post-processing from destroying the topology the model learned.\n\n### Pipeline\n\n1. **Fixed threshold at 0.5** — standard sigmoid midpoint\n2. **Remove small components** (<50 voxels) — noise cleanup using 26-connectivity labeling\n3. **2D slicewise closing** (4-connectivity, 1 iteration, topology-safe) — fills small gaps within each slice\n4. **2D slicewise hole filling** (all 3 axes, topology-safe) — fills enclosed holes per-slice\n5. **2D slicewise opening** (4-connectivity, 1 iteration, topology-safe) — removes small protrusions\n6. **Final small component removal** — second cleanup pass\n\n> **Critical design choice: 2D slicewise, not 3D.** All morphology is done slice-by-slice independently. 3D morphological operations with 3D structuring elements are too aggressive and destroy 1-voxel-thick surfaces. 2D slicewise processing is much gentler and preserves thin structures.\n\n---\n\n## 7. What Worked ✅\n\n1. **Topology-aware loss stack** — clDice(0.30) + Topology(0.20) directly optimize what the competition metric measures. Without them, the model learns decent Dice but poor Betti matching and VOI scores.\n2. **HybridConv3d** — Decoupled XY/Z convolutions respect CT anisotropy. More effective than isotropic 3×3×3 convolutions when in-plane and cross-slice resolutions differ.\n3. **SurfaceRefinementBlock** — Edge-aware processing at full decoder resolution is essential for 1–2 voxel thick surfaces. The `|conv(x)|` edge detection branch gives the model an explicit edge signal to refine.\n4. **2D slicewise post-processing** — Much gentler than 3D morphology. Thin surface structures survive.\n5. **Topology-safe operation guard** — Simple component-count check before/after each morphological operation prevents accidental merging of separate sheets.\n6. **GPU augmentations** — All augmentations on GPU with `F.grid_sample` eliminates CPU bottleneck. Training speed limited only by forward/backward pass.\n7. **bfloat16 training** — Same quality as float32, 2× memory savings, no GradScaler complexity.\n8. **Foreground oversampling (60%)** — Critical when surfaces are <10% of volume. Without it, the model sees mostly background patches.\n\n---\n\n## 8. What Didn't Work ❌\n\n1. **Frangi vesselness filter** in post-processing — Designed for tubular structures, not sheets. Reduced mean prediction confidence from 0.118 to 0.089, hurting overall score.\n2. **Adaptive thresholding** (e.g., 0.30) — Created too many false positive regions. The fixed 0.5 threshold was consistently better.\n3. **3D morphological operations** — Too aggressive for 1-voxel-thick surfaces. Switching to 2D slicewise processing was a critical improvement.\n4. **Heavy augmentation** — Strong elastic deformation (σ > 5) or large affine transforms broke thin surface topology during training, making the model worse at preserving connectivity.\n\n---\n\n## 9. Hardware & Runtime\n\n| Stage | Hardware | Time |\n|:------|:---------|:-----|\n| Training (per fold) | NVIDIA H100 80GB | 800 epochs, ~10–15 min/epoch |\n| Inference (per volume) | 2× Tesla T4 (DataParallel) | ~2–3 min with flip TTA |\n\nTraining was done on Vast.ai with H100 GPU. Inference runs on 2× T4 within the competition runtime limits.\n\n---\n\n## 10. Code\n\nBoth notebooks are fully self-contained single-file implementations. No external repositories or custom packages required.\n\n| Notebook | What it does |\n|:---------|:-------------|\n| **Training** | Full `TopoPreservingUNet3D` architecture + 5-loss stack + GPU augmentations + 3-fold stratified CV + checkpoint management |\n| **Inference** | Model loading + sliding window inference with Gaussian blending + 3-axis flip TTA + topology-safe 2D post-processing + TIFF submission |\n\n**Dependencies:** PyTorch, NumPy, SciPy, scikit-image, tifffile, pandas\n\n---\n\n*Thanks for reading!*\n"
  }
}