{
  "id": 679360,
  "title": "5th Place Solution",
  "url": "/competitions/vesuvius-challenge-surface-detection/discussion/679360",
  "author_name": "Dieter",
  "post_date": "2026-02-28T22:30:40.699000",
  "votes": 51,
  "comment_count": 21,
  "views": 0,
  "content": "<p>First, let me thank the Vesuvius Challenge team and Kaggle for organizing this fascinating competition. </p>\n<h2>TLDR</h2>\n<p>My solution is an ensemble of UNets mainy based on a custom SEResNeXt152 encoder with an Attention UNet decoder. All models are trained to regress a signed distance field (SDF) rather than a binary mask. I am using a novel(?) SDF L1 + SDF mass loss which is the sdf equivalent of BCE + Dice. SDF predictions are averaged across checkpoints and up to 8 TTA flips. At the end all is binarized at an SDF threshold of 0.3, and then refined with an iterative persistence-homology-based tunnel-filling post-processing pipeline. Although training on 160x160x160 crops, I inferred on full 320x320x320 to prevent any artifacts from a sliding window approach. </p>\n<h2>Cross-Validation</h2>\n<p>I used 4-fold cross-validation. During training I tracked three metrics locally - Surface Dice, VOI accelerated by GPU on original size, and the topology score on 10x downsampled size on CPU as the bettimatching algorithm is sequential. Neither local CV nor public LB correlated well with private LB, hence my selection of final submission went poorly. </p>\n<h2>Training Routine</h2>\n<p>All models are trained on 160^3 ROI patches randomly cropped from the npy saved 320^3 training volumes. SDF targets are calculated on the fly. Training uses the Adam optimizer with learning rate 1e-3, cosine annealing schedule. Mixed-precision training with bfloat16 is used throughout, and gradient checkpointing is enabled for the larger SEResNeXt152 models to fit within GPU memory. Weight clamping to the fp16-safe range is applied at checkpointing time to ensure clean float16 inference later. Batch size is 16.</p>\n<p>Data augmentation consists of random 3D flips along each spatial axis (p=0.5 each) and random 90-degree rotations in the (axis-1, axis-2) plane (p=0.5). CutMix (beta=1.0) was used in earlier model families but disabled for the final SEResNeXt152 runs, where it did not improve scores.</p>\n<h2>Loss Function: SDF L1 and \"mass\"</h2>\n<p>The key design decision was to train models to regress a signed distance field rather than a binary segmentation mask. The SDF target is computed from the binary label via Euclidean distance transforms on the fly on GPU:</p>\n<pre><code>fg_edt = distance_transform_edt(foreground_mask)\nbg_edt = distance_transform_edt(~foreground_mask)\nsdf = bg_edt - fg_edt          # negative inside, positive outside, zero at boundary\nsdf = clamp(sdf, -100, 5)\n</code></pre>\n<p>The loss has two components with equal weight:</p>\n<p><strong>Weighted SDF L1 loss (weight 0.5)</strong>: A voxel-wise L1 loss between predicted and target SDF, weighted by a Gaussian bell-curve matrix that concentrates supervision near the surface boundary:</p>\n<pre><code>W(sdf) = 1 + w0 * exp(-((sdf + core_radius)^2) / (2 * sigma^2))\n</code></pre>\n<p>with <code>w0 = 8</code>, <code>sigma = 4</code>, <code>core_radius = 5</code>. This gives a peak weight of 9.0 at <code>sdf = -5</code> (inside the surface at depth 5 voxels), plateauing at 9.0 for deeper interior, and decaying to ~1.0 far from the surface. The effect is that the model receives a strong gradient signal near the surface while distant background voxels contribute minimally.</p>\n<p><strong>SDF-based Dice loss (weight 0.5)</strong>: Instead of thresholding to binary, a continuous \"mass\" is derived via <code>ReLU(-sdf)</code> for both prediction and target, and a soft Dice coefficient is computed from these masses. This provides a global volumetric overlap signal that complements the voxel-wise L1.</p>\n<p>The following figure illustrates the three quantities at a central slice (axis_0 = 160) of a training volume:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1424766%2Fd1f6c2cf94b96afe18f924803ac4b03e%2Fsdf_target_visualization.png?generation=1772316847712852&amp;alt=media\" alt=\"\"></p>\n<p><em>Left: original binary target (0 = background, 1 = foreground, 2 = ignore). Center: SDF target (negative inside surface, positive outside; black contour at SDF = 0). Right: Gaussian weight matrix (higher weights near and inside the surface boundary).</em></p>\n<p>This loss combined with gaussian weighting solves a lot of the key challenges of this competition in an elegant and efficient way: Sheet separation, sheet continuity, boundary sharpness, precise skeletons etc. </p>\n<h2>Model Architectures</h2>\n<p>I explored a wide range of 3D architectures. Instead of relying on frameworks like nnUnet or medicai I asked an agent to recode relevant model families from scratch. This gives much more control. For the keras based SEResnext model from medcai it was also 60% (!) faster for training. The final ensemble draws from three architecture lineages:</p>\n<p><strong>SEResNeXt152 + Attention UNet</strong> -- the main workhorse and strongest single model. A custom pure-PyTorch 3D SE-ResNeXt152 encoder (layers <code>[3, 8, 36, 3]</code>, groups=32, width_per_group=8, SE reduction=16, DropPath rate 0.3) paired with an Attention UNet decoder (channels <code>[256, 128, 64, 32, 16]</code>). Attention gates on skip connections let the decoder focus on relevant encoder features. Gradient checkpointing is essential here due to the 36-block third stage. One variant using <code>seresnext152_48x8d</code> (48 groups, width 8) for additional capacity.</p>\n<p><strong>ResNet152 + skip maxpool</strong> -- Expensive UNet with a ResNet152 encoder, custom decoder channel configuration, and deep supervision during training (auxiliary losses at intermediate decoder stages). I skip one early maxpool so resolution through the model is higher. At inference, only the final head is used.</p>\n<p><strong>ResNet152 + UNet with SDF loss heavy augs and external data</strong> -- earlier ResNet152-based UNets trained with the SDF + Dice loss. One model extends the training set with additional scrolls. I used an additional binary input channel to give the model pixelwise information that this data is external (and hence labels are \"bad\") These models provide useful diversity to the ensemble despite being individually weaker than SEResNeXt152. </p>\n<p>I also experimented with SwinUNETR, UNETR++, MedNeXt, SegFormer3D, ConvNeXt V2, U-Mamba, and SEResNeXt200, but none surpassed the SEResNeXt152 + AttUNet architecture on local validation. Some topology-aware loss variants (persistence diagram matching loss, Euler characteristic loss) were explored. One PD-matching variant is included in the final ensemble but its finetuned only for a few epochs from another model because training was very slow.</p>\n<h2>Ensemble</h2>\n<p>Predictions are combined by simple averaging of raw SDF logits across all checkpoints and TTA augmentations per sample. For the larger SEResNeXt152 models, TTA is reduced to 2 augmentations (identity + triple-flip) to stay within the 9-hour Kaggle kernel runtime; smaller models use the full 8-flip TTA (3 single-axis flips + 3 dual-axis + 1 triple-axis + identity).</p>\n<p>NaN safety is critical: the SEResNeXt152 models are trained in bfloat16 but inference runs in float16. A safety conversion routine clamps all weights to <code>[-65504, 65504]</code> and registers forward hooks that replace any NaN/Inf activations.</p>\n<h2>Inference</h2>\n<p>All inference is done on full 320x320x320 image! No strided slices, as those create artifacts that are very harmful to topology.\nTwo GPUs are used in parallel via data sharding: even-indexed test samples go to GPU 0, odd-indexed to GPU 1. Each GPU processes all models sequentially for its shard and offload predictions to disk to prevent OOM. </p>\n<h2>Post-Processing</h2>\n<p>Post-processing proved essential for the topology score, which heavily penalizes tunnels (H1 topological features) in the predicted surface. My pipeline:</p>\n<p><strong>Step 1 -- Binarization</strong>: Threshold the averaged SDF logits at 0.3 (voxels with SDF &lt; 0.3 become foreground).</p>\n<p><strong>Step 2 -- Dust removal</strong>: Remove small connected components (&lt; 50,000 voxels, 6-connectivity), preserving components that touch at least 3 volume boundaries regardless of size.</p>\n<p><strong>Step 3 -- Iterative H1 tunnel filling</strong> (13 iterations):</p>\n<ol>\n<li><strong>SDF filtration</strong>: Overwrite foreground voxels in the raw SDF to <code>threshold - 1.0</code>, creating a cubical filtration where the surface sits at the threshold.</li>\n<li><strong>Persistence barcode computation</strong>: Split the volume into 2x2x2 = 8 octant tiles, compute persistence homology on each tile using a custom C++ module (<code>barcode3d_fast_v3</code>) which was derived from betti matching library. This identifies H1 features (tunnels) with birth/death values and coordinates.</li>\n<li><strong>Straddling filter</strong>: Keep only H1 features where <code>birth &lt; sdf_threshold &lt; death</code>, i.e., tunnels that cross the binarization surface.</li>\n<li><strong>Adaptive radius</strong>: For each tunnel's death coordinate, estimate the tunnel width from the background EDT and set the fill radius to <code>clip(edt_width + 1.0, 3.5, 7.0)</code>.</li>\n<li><strong>Bridge detection</strong>: Simulate filling a ball at each death coordinate and check whether it would merge separate connected components. Skip coordinates that would create bridges.</li>\n<li><strong>Ball filling</strong>: Fill spherical balls at the remaining death coordinates. On the final iteration, also apply morphological hole filling.</li>\n<li><strong>Component count guard</strong>: If filling reduced the number of connected components (accidental merge), revert to the pre-fill state.</li>\n</ol>\n<h2>Key Insights</h2>\n<ul>\n<li><strong>SDF regression &gt; binary classification</strong> for this metric suite. The SDF naturally encodes distance-to-surface information, and thresholding the SDF at different values during validation gives a smooth trade-off between Surface Dice and topology metrics. Binary cross-entropy models consistently scored lower on the topology component.</li>\n<li><strong>Gaussian weighting</strong> focuses the loss on the skeleton region, but still extents to sheet boundary. Background is heavily downweighted. </li>\n<li><strong>Iterative tunnel filling</strong> with bridge detection is the single most impactful post-processing step. Going from 0 to 13 iterations improved the topology score by ~0.08 on local validation, with diminishing returns beyond 11-13 iterations.</li>\n</ul>",
  "messages": [
    {
      "id": 3415426,
      "postDate": "2026-02-28T22:30:40.700Z",
      "content": "<p>First, let me thank the Vesuvius Challenge team and Kaggle for organizing this fascinating competition. </p>\n<h2>TLDR</h2>\n<p>My solution is an ensemble of UNets mainy based on a custom SEResNeXt152 encoder with an Attention UNet decoder. All models are trained to regress a signed distance field (SDF) rather than a binary mask. I am using a novel(?) SDF L1 + SDF mass loss which is the sdf equivalent of BCE + Dice. SDF predictions are averaged across checkpoints and up to 8 TTA flips. At the end all is binarized at an SDF threshold of 0.3, and then refined with an iterative persistence-homology-based tunnel-filling post-processing pipeline. Although training on 160x160x160 crops, I inferred on full 320x320x320 to prevent any artifacts from a sliding window approach. </p>\n<h2>Cross-Validation</h2>\n<p>I used 4-fold cross-validation. During training I tracked three metrics locally - Surface Dice, VOI accelerated by GPU on original size, and the topology score on 10x downsampled size on CPU as the bettimatching algorithm is sequential. Neither local CV nor public LB correlated well with private LB, hence my selection of final submission went poorly. </p>\n<h2>Training Routine</h2>\n<p>All models are trained on 160^3 ROI patches randomly cropped from the npy saved 320^3 training volumes. SDF targets are calculated on the fly. Training uses the Adam optimizer with learning rate 1e-3, cosine annealing schedule. Mixed-precision training with bfloat16 is used throughout, and gradient checkpointing is enabled for the larger SEResNeXt152 models to fit within GPU memory. Weight clamping to the fp16-safe range is applied at checkpointing time to ensure clean float16 inference later. Batch size is 16.</p>\n<p>Data augmentation consists of random 3D flips along each spatial axis (p=0.5 each) and random 90-degree rotations in the (axis-1, axis-2) plane (p=0.5). CutMix (beta=1.0) was used in earlier model families but disabled for the final SEResNeXt152 runs, where it did not improve scores.</p>\n<h2>Loss Function: SDF L1 and \"mass\"</h2>\n<p>The key design decision was to train models to regress a signed distance field rather than a binary segmentation mask. The SDF target is computed from the binary label via Euclidean distance transforms on the fly on GPU:</p>\n<pre><code>fg_edt = distance_transform_edt(foreground_mask)\nbg_edt = distance_transform_edt(~foreground_mask)\nsdf = bg_edt - fg_edt          # negative inside, positive outside, zero at boundary\nsdf = clamp(sdf, -100, 5)\n</code></pre>\n<p>The loss has two components with equal weight:</p>\n<p><strong>Weighted SDF L1 loss (weight 0.5)</strong>: A voxel-wise L1 loss between predicted and target SDF, weighted by a Gaussian bell-curve matrix that concentrates supervision near the surface boundary:</p>\n<pre><code>W(sdf) = 1 + w0 * exp(-((sdf + core_radius)^2) / (2 * sigma^2))\n</code></pre>\n<p>with <code>w0 = 8</code>, <code>sigma = 4</code>, <code>core_radius = 5</code>. This gives a peak weight of 9.0 at <code>sdf = -5</code> (inside the surface at depth 5 voxels), plateauing at 9.0 for deeper interior, and decaying to ~1.0 far from the surface. The effect is that the model receives a strong gradient signal near the surface while distant background voxels contribute minimally.</p>\n<p><strong>SDF-based Dice loss (weight 0.5)</strong>: Instead of thresholding to binary, a continuous \"mass\" is derived via <code>ReLU(-sdf)</code> for both prediction and target, and a soft Dice coefficient is computed from these masses. This provides a global volumetric overlap signal that complements the voxel-wise L1.</p>\n<p>The following figure illustrates the three quantities at a central slice (axis_0 = 160) of a training volume:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1424766%2Fd1f6c2cf94b96afe18f924803ac4b03e%2Fsdf_target_visualization.png?generation=1772316847712852&amp;alt=media\" alt=\"\"></p>\n<p><em>Left: original binary target (0 = background, 1 = foreground, 2 = ignore). Center: SDF target (negative inside surface, positive outside; black contour at SDF = 0). Right: Gaussian weight matrix (higher weights near and inside the surface boundary).</em></p>\n<p>This loss combined with gaussian weighting solves a lot of the key challenges of this competition in an elegant and efficient way: Sheet separation, sheet continuity, boundary sharpness, precise skeletons etc. </p>\n<h2>Model Architectures</h2>\n<p>I explored a wide range of 3D architectures. Instead of relying on frameworks like nnUnet or medicai I asked an agent to recode relevant model families from scratch. This gives much more control. For the keras based SEResnext model from medcai it was also 60% (!) faster for training. The final ensemble draws from three architecture lineages:</p>\n<p><strong>SEResNeXt152 + Attention UNet</strong> -- the main workhorse and strongest single model. A custom pure-PyTorch 3D SE-ResNeXt152 encoder (layers <code>[3, 8, 36, 3]</code>, groups=32, width_per_group=8, SE reduction=16, DropPath rate 0.3) paired with an Attention UNet decoder (channels <code>[256, 128, 64, 32, 16]</code>). Attention gates on skip connections let the decoder focus on relevant encoder features. Gradient checkpointing is essential here due to the 36-block third stage. One variant using <code>seresnext152_48x8d</code> (48 groups, width 8) for additional capacity.</p>\n<p><strong>ResNet152 + skip maxpool</strong> -- Expensive UNet with a ResNet152 encoder, custom decoder channel configuration, and deep supervision during training (auxiliary losses at intermediate decoder stages). I skip one early maxpool so resolution through the model is higher. At inference, only the final head is used.</p>\n<p><strong>ResNet152 + UNet with SDF loss heavy augs and external data</strong> -- earlier ResNet152-based UNets trained with the SDF + Dice loss. One model extends the training set with additional scrolls. I used an additional binary input channel to give the model pixelwise information that this data is external (and hence labels are \"bad\") These models provide useful diversity to the ensemble despite being individually weaker than SEResNeXt152. </p>\n<p>I also experimented with SwinUNETR, UNETR++, MedNeXt, SegFormer3D, ConvNeXt V2, U-Mamba, and SEResNeXt200, but none surpassed the SEResNeXt152 + AttUNet architecture on local validation. Some topology-aware loss variants (persistence diagram matching loss, Euler characteristic loss) were explored. One PD-matching variant is included in the final ensemble but its finetuned only for a few epochs from another model because training was very slow.</p>\n<h2>Ensemble</h2>\n<p>Predictions are combined by simple averaging of raw SDF logits across all checkpoints and TTA augmentations per sample. For the larger SEResNeXt152 models, TTA is reduced to 2 augmentations (identity + triple-flip) to stay within the 9-hour Kaggle kernel runtime; smaller models use the full 8-flip TTA (3 single-axis flips + 3 dual-axis + 1 triple-axis + identity).</p>\n<p>NaN safety is critical: the SEResNeXt152 models are trained in bfloat16 but inference runs in float16. A safety conversion routine clamps all weights to <code>[-65504, 65504]</code> and registers forward hooks that replace any NaN/Inf activations.</p>\n<h2>Inference</h2>\n<p>All inference is done on full 320x320x320 image! No strided slices, as those create artifacts that are very harmful to topology.\nTwo GPUs are used in parallel via data sharding: even-indexed test samples go to GPU 0, odd-indexed to GPU 1. Each GPU processes all models sequentially for its shard and offload predictions to disk to prevent OOM. </p>\n<h2>Post-Processing</h2>\n<p>Post-processing proved essential for the topology score, which heavily penalizes tunnels (H1 topological features) in the predicted surface. My pipeline:</p>\n<p><strong>Step 1 -- Binarization</strong>: Threshold the averaged SDF logits at 0.3 (voxels with SDF &lt; 0.3 become foreground).</p>\n<p><strong>Step 2 -- Dust removal</strong>: Remove small connected components (&lt; 50,000 voxels, 6-connectivity), preserving components that touch at least 3 volume boundaries regardless of size.</p>\n<p><strong>Step 3 -- Iterative H1 tunnel filling</strong> (13 iterations):</p>\n<ol>\n<li><strong>SDF filtration</strong>: Overwrite foreground voxels in the raw SDF to <code>threshold - 1.0</code>, creating a cubical filtration where the surface sits at the threshold.</li>\n<li><strong>Persistence barcode computation</strong>: Split the volume into 2x2x2 = 8 octant tiles, compute persistence homology on each tile using a custom C++ module (<code>barcode3d_fast_v3</code>) which was derived from betti matching library. This identifies H1 features (tunnels) with birth/death values and coordinates.</li>\n<li><strong>Straddling filter</strong>: Keep only H1 features where <code>birth &lt; sdf_threshold &lt; death</code>, i.e., tunnels that cross the binarization surface.</li>\n<li><strong>Adaptive radius</strong>: For each tunnel's death coordinate, estimate the tunnel width from the background EDT and set the fill radius to <code>clip(edt_width + 1.0, 3.5, 7.0)</code>.</li>\n<li><strong>Bridge detection</strong>: Simulate filling a ball at each death coordinate and check whether it would merge separate connected components. Skip coordinates that would create bridges.</li>\n<li><strong>Ball filling</strong>: Fill spherical balls at the remaining death coordinates. On the final iteration, also apply morphological hole filling.</li>\n<li><strong>Component count guard</strong>: If filling reduced the number of connected components (accidental merge), revert to the pre-fill state.</li>\n</ol>\n<h2>Key Insights</h2>\n<ul>\n<li><strong>SDF regression &gt; binary classification</strong> for this metric suite. The SDF naturally encodes distance-to-surface information, and thresholding the SDF at different values during validation gives a smooth trade-off between Surface Dice and topology metrics. Binary cross-entropy models consistently scored lower on the topology component.</li>\n<li><strong>Gaussian weighting</strong> focuses the loss on the skeleton region, but still extents to sheet boundary. Background is heavily downweighted. </li>\n<li><strong>Iterative tunnel filling</strong> with bridge detection is the single most impactful post-processing step. Going from 0 to 13 iterations improved the topology score by ~0.08 on local validation, with diminishing returns beyond 11-13 iterations.</li>\n</ul>",
      "rawMarkdown": "First, let me thank the Vesuvius Challenge team and Kaggle for organizing this fascinating competition. \n\n\n## TLDR\n\n\nMy solution is an ensemble of UNets mainy based on a custom SEResNeXt152 encoder with an Attention UNet decoder. All models are trained to regress a signed distance field (SDF) rather than a binary mask. I am using a novel(?) SDF L1 + SDF mass loss which is the sdf equivalent of BCE + Dice. SDF predictions are averaged across checkpoints and up to 8 TTA flips. At the end all is binarized at an SDF threshold of 0.3, and then refined with an iterative persistence-homology-based tunnel-filling post-processing pipeline. Although training on 160x160x160 crops, I inferred on full 320x320x320 to prevent any artifacts from a sliding window approach. \n\n\n## Cross-Validation\n\n\nI used 4-fold cross-validation. During training I tracked three metrics locally - Surface Dice, VOI accelerated by GPU on original size, and the topology score on 10x downsampled size on CPU as the bettimatching algorithm is sequential. Neither local CV nor public LB correlated well with private LB, hence my selection of final submission went poorly. \n\n\n## Training Routine\n\n\nAll models are trained on 160^3 ROI patches randomly cropped from the npy saved 320^3 training volumes. SDF targets are calculated on the fly. Training uses the Adam optimizer with learning rate 1e-3, cosine annealing schedule. Mixed-precision training with bfloat16 is used throughout, and gradient checkpointing is enabled for the larger SEResNeXt152 models to fit within GPU memory. Weight clamping to the fp16-safe range is applied at checkpointing time to ensure clean float16 inference later. Batch size is 16.\n\n\nData augmentation consists of random 3D flips along each spatial axis (p=0.5 each) and random 90-degree rotations in the (axis-1, axis-2) plane (p=0.5). CutMix (beta=1.0) was used in earlier model families but disabled for the final SEResNeXt152 runs, where it did not improve scores.\n\n\n## Loss Function: SDF L1 and \"mass\"\n\n\nThe key design decision was to train models to regress a signed distance field rather than a binary segmentation mask. The SDF target is computed from the binary label via Euclidean distance transforms on the fly on GPU:\n\n\n```\nfg_edt = distance_transform_edt(foreground_mask)\nbg_edt = distance_transform_edt(~foreground_mask)\nsdf = bg_edt - fg_edt          # negative inside, positive outside, zero at boundary\nsdf = clamp(sdf, -100, 5)\n```\n\n\nThe loss has two components with equal weight:\n\n\n**Weighted SDF L1 loss (weight 0.5)**: A voxel-wise L1 loss between predicted and target SDF, weighted by a Gaussian bell-curve matrix that concentrates supervision near the surface boundary:\n\n\n```\nW(sdf) = 1 + w0 * exp(-((sdf + core_radius)^2) / (2 * sigma^2))\n```\n\n\nwith `w0 = 8`, `sigma = 4`, `core_radius = 5`. This gives a peak weight of 9.0 at `sdf = -5` (inside the surface at depth 5 voxels), plateauing at 9.0 for deeper interior, and decaying to ~1.0 far from the surface. The effect is that the model receives a strong gradient signal near the surface while distant background voxels contribute minimally.\n\n\n**SDF-based Dice loss (weight 0.5)**: Instead of thresholding to binary, a continuous \"mass\" is derived via `ReLU(-sdf)` for both prediction and target, and a soft Dice coefficient is computed from these masses. This provides a global volumetric overlap signal that complements the voxel-wise L1.\n\n\nThe following figure illustrates the three quantities at a central slice (axis_0 = 160) of a training volume:\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1424766%2Fd1f6c2cf94b96afe18f924803ac4b03e%2Fsdf_target_visualization.png?generation=1772316847712852&alt=media)\n\n\n*Left: original binary target (0 = background, 1 = foreground, 2 = ignore). Center: SDF target (negative inside surface, positive outside; black contour at SDF = 0). Right: Gaussian weight matrix (higher weights near and inside the surface boundary).*\n\nThis loss combined with gaussian weighting solves a lot of the key challenges of this competition in an elegant and efficient way: Sheet separation, sheet continuity, boundary sharpness, precise skeletons etc. \n\n\n## Model Architectures\n\n\nI explored a wide range of 3D architectures. Instead of relying on frameworks like nnUnet or medicai I asked an agent to recode relevant model families from scratch. This gives much more control. For the keras based SEResnext model from medcai it was also 60% (!) faster for training. The final ensemble draws from three architecture lineages:\n\n\n**SEResNeXt152 + Attention UNet** -- the main workhorse and strongest single model. A custom pure-PyTorch 3D SE-ResNeXt152 encoder (layers `[3, 8, 36, 3]`, groups=32, width_per_group=8, SE reduction=16, DropPath rate 0.3) paired with an Attention UNet decoder (channels `[256, 128, 64, 32, 16]`). Attention gates on skip connections let the decoder focus on relevant encoder features. Gradient checkpointing is essential here due to the 36-block third stage. One variant using `seresnext152_48x8d` (48 groups, width 8) for additional capacity.\n\n\n**ResNet152 + skip maxpool** -- Expensive UNet with a ResNet152 encoder, custom decoder channel configuration, and deep supervision during training (auxiliary losses at intermediate decoder stages). I skip one early maxpool so resolution through the model is higher. At inference, only the final head is used.\n\n\n**ResNet152 + UNet with SDF loss heavy augs and external data** -- earlier ResNet152-based UNets trained with the SDF + Dice loss. One model extends the training set with additional scrolls. I used an additional binary input channel to give the model pixelwise information that this data is external (and hence labels are \"bad\") These models provide useful diversity to the ensemble despite being individually weaker than SEResNeXt152. \n\n\nI also experimented with SwinUNETR, UNETR++, MedNeXt, SegFormer3D, ConvNeXt V2, U-Mamba, and SEResNeXt200, but none surpassed the SEResNeXt152 + AttUNet architecture on local validation. Some topology-aware loss variants (persistence diagram matching loss, Euler characteristic loss) were explored. One PD-matching variant is included in the final ensemble but its finetuned only for a few epochs from another model because training was very slow.\n\n\n## Ensemble\n\n\nPredictions are combined by simple averaging of raw SDF logits across all checkpoints and TTA augmentations per sample. For the larger SEResNeXt152 models, TTA is reduced to 2 augmentations (identity + triple-flip) to stay within the 9-hour Kaggle kernel runtime; smaller models use the full 8-flip TTA (3 single-axis flips + 3 dual-axis + 1 triple-axis + identity).\n\n\nNaN safety is critical: the SEResNeXt152 models are trained in bfloat16 but inference runs in float16. A safety conversion routine clamps all weights to `[-65504, 65504]` and registers forward hooks that replace any NaN/Inf activations.\n\n\n## Inference\n\n\nAll inference is done on full 320x320x320 image! No strided slices, as those create artifacts that are very harmful to topology.\nTwo GPUs are used in parallel via data sharding: even-indexed test samples go to GPU 0, odd-indexed to GPU 1. Each GPU processes all models sequentially for its shard and offload predictions to disk to prevent OOM. \n\n\n## Post-Processing\n\n\nPost-processing proved essential for the topology score, which heavily penalizes tunnels (H1 topological features) in the predicted surface. My pipeline:\n\n\n**Step 1 -- Binarization**: Threshold the averaged SDF logits at 0.3 (voxels with SDF < 0.3 become foreground).\n\n\n**Step 2 -- Dust removal**: Remove small connected components (< 50,000 voxels, 6-connectivity), preserving components that touch at least 3 volume boundaries regardless of size.\n\n\n**Step 3 -- Iterative H1 tunnel filling** (13 iterations):\n\n\n1. **SDF filtration**: Overwrite foreground voxels in the raw SDF to `threshold - 1.0`, creating a cubical filtration where the surface sits at the threshold.\n2. **Persistence barcode computation**: Split the volume into 2x2x2 = 8 octant tiles, compute persistence homology on each tile using a custom C++ module (`barcode3d_fast_v3`) which was derived from betti matching library. This identifies H1 features (tunnels) with birth/death values and coordinates.\n3. **Straddling filter**: Keep only H1 features where `birth < sdf_threshold < death`, i.e., tunnels that cross the binarization surface.\n4. **Adaptive radius**: For each tunnel's death coordinate, estimate the tunnel width from the background EDT and set the fill radius to `clip(edt_width + 1.0, 3.5, 7.0)`.\n5. **Bridge detection**: Simulate filling a ball at each death coordinate and check whether it would merge separate connected components. Skip coordinates that would create bridges.\n6. **Ball filling**: Fill spherical balls at the remaining death coordinates. On the final iteration, also apply morphological hole filling.\n7. **Component count guard**: If filling reduced the number of connected components (accidental merge), revert to the pre-fill state.\n\n\n## Key Insights\n\n\n- **SDF regression > binary classification** for this metric suite. The SDF naturally encodes distance-to-surface information, and thresholding the SDF at different values during validation gives a smooth trade-off between Surface Dice and topology metrics. Binary cross-entropy models consistently scored lower on the topology component.\n- **Gaussian weighting** focuses the loss on the skeleton region, but still extents to sheet boundary. Background is heavily downweighted. \n- **Iterative tunnel filling** with bridge detection is the single most impactful post-processing step. Going from 0 to 13 iterations improved the topology score by ~0.08 on local validation, with diminishing returns beyond 11-13 iterations.\n\n\n\n\n\n\n\n\n\n",
      "votes": 51
    },
    {
      "id": 3415538,
      "postDate": "2026-03-01T01:29:16.250Z",
      "content": "<p>Thank you for sharing and congrats! the sdf regression in particular is interesting. i've never necessarily loved semantic segmentation for this task as the model has a tendency to just over \"blur\" tough regions, where this can provide a bit more signal to the model in terms of \"interpolating\" areas it has trouble separating. </p>",
      "rawMarkdown": "Thank you for sharing and congrats! the sdf regression in particular is interesting. i've never necessarily loved semantic segmentation for this task as the model has a tendency to just over \"blur\" tough regions, where this can provide a bit more signal to the model in terms of \"interpolating\" areas it has trouble separating. ",
      "votes": 4
    },
    {
      "id": 3416025,
      "postDate": "2026-03-01T21:11:34.050Z",
      "content": "<p>Glad to see another gold with an interesting approach. I originally wanted to try out SEResNeXt152 type models but we had so much success with the nnunet structure and some of the custom models that we started with that I never even tried it. Very jealous of being able to infer on full 320,320,320 with 8xTTA. As our approach had so many steps by the time we got to post processing we only had about an hour for inference time. I would be curious to see how much impact that inferencing resolution had on your cv/lb? Another question I had was our model preformed worse when averaging on raw logits instead of the probabilities, despite the logits being the more logical approach, so I would be curious if you tried that. Love the post processing ideas as well. </p>\n<p>I really thought after seeing everyones solutions that we would have a solid pipeline and just missed with the post processing and that slid us down to 10th but so far all the post processing from other top competitiors actually just lowers our score. </p>\n<p>Congrats on your placement!</p>",
      "rawMarkdown": "Glad to see another gold with an interesting approach. I originally wanted to try out SEResNeXt152 type models but we had so much success with the nnunet structure and some of the custom models that we started with that I never even tried it. Very jealous of being able to infer on full 320,320,320 with 8xTTA. As our approach had so many steps by the time we got to post processing we only had about an hour for inference time. I would be curious to see how much impact that inferencing resolution had on your cv/lb? Another question I had was our model preformed worse when averaging on raw logits instead of the probabilities, despite the logits being the more logical approach, so I would be curious if you tried that. Love the post processing ideas as well. \n\nI really thought after seeing everyones solutions that we would have a solid pipeline and just missed with the post processing and that slid us down to 10th but so far all the post processing from other top competitiors actually just lowers our score. \n\nCongrats on your placement!",
      "votes": 2,
      "replies": [
        {
          "id": 3416166,
          "postDate": "2026-03-02T07:57:14.437Z",
          "content": "<blockquote>\n  <p>Very jealous of being able to infer on full 320,320,320 with 8xTTA</p>\n</blockquote>\n<p>Reading your inference kernel it seems you used an overlap of (0.5,0.5,0.5) with a 160³ window. Thats 27 patches for a 320³ image and 3x the amount of voxels put through the model. So even with TTA8 my inference is 1.5x cheaper without creating border artifacts.</p>\n<p>My model simply has no probabilities. It predicts sdf values (ranging from -10 to 5) and then binary mask, by thresholding at 0.3. As <a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> explained you can create artificial probabilities  by introducing a temperature and do sigmoid(-sdf/T), but T is another hyperparameter and probs will be on a slightly different scale than those from BCE models. You can (probably) ensemble my model with other models by a voxel rank transform like we did in our <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/writeups/daddies-1st-place-solution-segmentation-with-partl\" target=\"_blank\">winning solution to the  CryoET competition</a> section Ensembling </p>",
          "rawMarkdown": ">Very jealous of being able to infer on full 320,320,320 with 8xTTA\n\nReading your inference kernel it seems you used an overlap of (0.5,0.5,0.5) with a 160³ window. Thats 27 patches for a 320³ image and 3x the amount of voxels put through the model. So even with TTA8 my inference is 1.5x cheaper without creating border artifacts.\n\nMy model simply has no probabilities. It predicts sdf values (ranging from -10 to 5) and then binary mask, by thresholding at 0.3. As @giorgioangelotti explained you can create artificial probabilities  by introducing a temperature and do sigmoid(-sdf/T), but T is another hyperparameter and probs will be on a slightly different scale than those from BCE models. You can (probably) ensemble my model with other models by a voxel rank transform like we did in our [winning solution to the  CryoET competition](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/writeups/daddies-1st-place-solution-segmentation-with-partl) section Ensembling ",
          "votes": 1,
          "replies": [
            {
              "id": 3416402,
              "postDate": "2026-03-02T19:02:17.843Z",
              "content": "<p>Congrats on 50 golds <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>, great achievement!</p>\n<p>Do you design your architecture/training pipeline any differently when the inference patch size differs from the training patch size?</p>\n<p>I often struggle to match the performance achieved when using the same patch size during training and inference (in this competition and past 3D competitions).</p>",
              "rawMarkdown": "Congrats on 50 golds @christofhenkel, great achievement!\n\nDo you design your architecture/training pipeline any differently when the inference patch size differs from the training patch size?\n\nI often struggle to match the performance achieved when using the same patch size during training and inference (in this competition and past 3D competitions).\n",
              "votes": 1
            },
            {
              "id": 3416418,
              "postDate": "2026-03-02T19:51:34.867Z",
              "content": "<p>Its the first time I used much larger patch size for inference. Learned the trick from <a href=\"https://www.kaggle.com/bloodaxe\" target=\"_blank\">@bloodaxe</a> in CryoET comp. I think the metric in this competition favors not doing slided window inference (SWI) a bit because border artifacts and inhomogenous amount of voxel predictions (think about that with SWI some voxels are predicted more often than others) hurt the topo score. And what I saw here is that SWI hurt the score more than inconsistencies created by the patch size mismatch. E.g. I normalized each patch by mean/ std and thats inconsistent for different patch sizes. </p>",
              "rawMarkdown": "Its the first time I used much larger patch size for inference. Learned the trick from @bloodaxe in CryoET comp. I think the metric in this competition favors not doing slided window inference (SWI) a bit because border artifacts and inhomogenous amount of voxel predictions (think about that with SWI some voxels are predicted more often than others) hurt the topo score. And what I saw here is that SWI hurt the score more than inconsistencies created by the patch size mismatch. E.g. I normalized each patch by mean/ std and thats inconsistent for different patch sizes. ",
              "votes": 2
            },
            {
              "id": 3416434,
              "postDate": "2026-03-02T21:40:03.887Z",
              "content": "<p>Oh yeah we had a whole lot going on because we missed the median filter blurring idea. With it we get first, without it we use the full 9 hours and barely get 10th! All a learning experience :). I saw increased score for larger inference patch for BYU but just didnt have the time to even think about it for his comp</p>",
              "rawMarkdown": "Oh yeah we had a whole lot going on because we missed the median filter blurring idea. With it we get first, without it we use the full 9 hours and barely get 10th! All a learning experience :). I saw increased score for larger inference patch for BYU but just didnt have the time to even think about it for his comp\n",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3415750,
      "postDate": "2026-03-01T09:03:17.683Z",
      "content": "<p>This solution is really brilliant and I like seeing architectures coming out of the nnUNet framework!\nI find the formulation as an SDF regression really elegant. When we launched this challenge we thought that some solutions could be inspirational also for another task we are working on: ink detection in the scrolls. I can see this solution very easily readapted for ink detection as well!</p>\n<p>I have two questions</p>\n<ol>\n<li><p>The part of the loss with mass from the SDF and the soft Dice made me think of <a href=\"https://arxiv.org/abs/1911.02278\" target=\"_blank\">a paper a read</a> a while ago where the authors said that in tasks with uncertainty in the labels, optimization of the soft dice can lead to biased estimates of the volume of the region to segment, while the cross entropy leads ofc to worse Dice score but unbiased volume estimates. Maybe an overestimation of the volume can lead to unwanted mergers, irregardless of the better Dice? You say that* Binary cross-entropy models consistently scored lower on the topology component*, but I am curious to know whether you tested also a BCE variant of the loss on a probability derived from the SDF, e.g. sigmoid(-sdf/T) (with T some temperature)?</p></li>\n<li><p>Within the architectures that you are using, the ResNe(X)ts have a lot of layers. Do you feel that this higher capacity is really needed for this task? I am also interested in the skip of the first maxpool to try to preserve finer structure. I think this can be even more important for ink!</p></li>\n</ol>\n<p>Thank you!</p>",
      "rawMarkdown": "This solution is really brilliant and I like seeing architectures coming out of the nnUNet framework!\nI find the formulation as an SDF regression really elegant. When we launched this challenge we thought that some solutions could be inspirational also for another task we are working on: ink detection in the scrolls. I can see this solution very easily readapted for ink detection as well!\n\nI have two questions\n1. The part of the loss with mass from the SDF and the soft Dice made me think of [a paper a read](https://arxiv.org/abs/1911.02278) a while ago where the authors said that in tasks with uncertainty in the labels, optimization of the soft dice can lead to biased estimates of the volume of the region to segment, while the cross entropy leads ofc to worse Dice score but unbiased volume estimates. Maybe an overestimation of the volume can lead to unwanted mergers, irregardless of the better Dice? You say that* Binary cross-entropy models consistently scored lower on the topology component*, but I am curious to know whether you tested also a BCE variant of the loss on a probability derived from the SDF, e.g. sigmoid(-sdf/T) (with T some temperature)?\n\n2. Within the architectures that you are using, the ResNe(X)ts have a lot of layers. Do you feel that this higher capacity is really needed for this task? I am also interested in the skip of the first maxpool to try to preserve finer structure. I think this can be even more important for ink!\n\nThank you!",
      "votes": 2,
      "replies": [
        {
          "id": 3416156,
          "postDate": "2026-03-02T07:31:18.530Z",
          "content": "<ol>\n<li><p>The paper author claim \"systematic under or overestimation of the predicted volume\". I guess this can be true. However note that I am using a binarization threshold of 0.3 which systematically predicts a slightly \"larger\" volume than the model predicted (higher sdf means further away from sheet center). This threshold is optimized via grid-search on CV and incoorporate trade-off of the 3 metrics used. In short, you can control merges with the threshold. I also tried more targeted prevention of merges via adding an sdf weight at voronoi lines, but it had no effect, because model was already really good at separation. I think reason is that sdf-approach results in really good skeletons. I tested probablity  derived from SDF, especially since this then can be used for frangii filter etc, but it performed worse and its tricky to tune T.</p></li>\n<li><p>Yes higher capacity has significant impact. I tend to start with very simple and fast models to optimize the number of experiments/ time for a competition and only increase capacity if really needed. </p></li>\n</ol>",
          "rawMarkdown": "1. The paper author claim \"systematic under or overestimation of the predicted volume\". I guess this can be true. However note that I am using a binarization threshold of 0.3 which systematically predicts a slightly \"larger\" volume than the model predicted (higher sdf means further away from sheet center). This threshold is optimized via grid-search on CV and incoorporate trade-off of the 3 metrics used. In short, you can control merges with the threshold. I also tried more targeted prevention of merges via adding an sdf weight at voronoi lines, but it had no effect, because model was already really good at separation. I think reason is that sdf-approach results in really good skeletons. I tested probablity  derived from SDF, especially since this then can be used for frangii filter etc, but it performed worse and its tricky to tune T.\n\n2. Yes higher capacity has significant impact. I tend to start with very simple and fast models to optimize the number of experiments/ time for a competition and only increase capacity if really needed. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 3415640,
      "postDate": "2026-03-01T04:23:40.267Z",
      "content": "<p><a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> <strong>Sir Huge Congrats on achieving your 50th Competition Gold Medal. 🥳🎉</strong>\nThis is a remarkable milestone and truly a half-century of excellence.</p>\n<p>Your approach is amazing sir, and this is finally a solution that feels genuinely unique, with many things to learn from it\nMost other write-ups I read were either not as useful or did not share this much detail.</p>",
      "rawMarkdown": "@christofhenkel **Sir Huge Congrats on achieving your 50th Competition Gold Medal. 🥳🎉**\nThis is a remarkable milestone and truly a half-century of excellence.\n\nYour approach is amazing sir, and this is finally a solution that feels genuinely unique, with many things to learn from it\nMost other write-ups I read were either not as useful or did not share this much detail.",
      "votes": 2
    },
    {
      "id": 3417535,
      "postDate": "2026-03-05T16:00:50.927Z",
      "content": "<p>Brilliant combination of geometric learning, large-scale engineering, and topology-aware post-processing. \nTurning surface detection into an SDF regression problem and closing the loop with persistence homology is genuinely insightful. A masterclass in metric-aligned modeling.</p>",
      "rawMarkdown": "Brilliant combination of geometric learning, large-scale engineering, and topology-aware post-processing. \nTurning surface detection into an SDF regression problem and closing the loop with persistence homology is genuinely insightful. A masterclass in metric-aligned modeling.",
      "replies": [
        {
          "id": 3417695,
          "postDate": "2026-03-06T01:04:32.417Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 3416206,
      "postDate": "2026-03-02T09:59:31.100Z",
      "content": "<p>Really inspiring solution, initially I was expecting the final solutions to be made up of varying  architectures, methods like predicting sdf, vector fields, etc. but nnUNet somehow overpowered and people stuck to it so what I thought initially didn't happen.\nSticking to something unconventional is itself really cool, winning is a cherry on top. Congratss</p>",
      "rawMarkdown": "Really inspiring solution, initially I was expecting the final solutions to be made up of varying  architectures, methods like predicting sdf, vector fields, etc. but nnUNet somehow overpowered and people stuck to it so what I thought initially didn't happen.\nSticking to something unconventional is itself really cool, winning is a cherry on top. Congratss"
    },
    {
      "id": 3415873,
      "postDate": "2026-03-01T14:30:13.613Z",
      "content": "<p>Thanks for share your non nnUNet/TransUNet top solution. I've been working with custom 3D Unet too but with resnet50 encoder and no attention decoder (in my experiments I didn't fount evidence of improvement). Also I've trained with a plane dice + focal loss combo on the provided hard masks. I've inferred too in full volumes whatever they were. I did 4 folds based on the 4 most available scrolls ids. With all I couldn't go higher than .461/.462 with that.</p>\n<p>At first read I think loss and post-processing have played an important role. Thanks again for sharing a solution to actually learn something from.</p>\n<p>EDIT:</p>\n<blockquote>\n  <p>Data augmentation consists of random 3D flips along each spatial axis (p=0.5 each) and random 90-degree rotations in the (axis-1, axis-2) plane (p=0.5). CutMix (beta=1.0) was used in earlier model families but disabled for the final SEResNeXt152 runs, where it did not improve scores.</p>\n</blockquote>\n<p>I skipped cutmix and restricted flips to slices planes. So I did similar but modest augmentations. I've been surprised with some high scoring public notebooks with totally free rotations, those were terrible in my pipeline.</p>",
      "rawMarkdown": "Thanks for share your non nnUNet/TransUNet top solution. I've been working with custom 3D Unet too but with resnet50 encoder and no attention decoder (in my experiments I didn't fount evidence of improvement). Also I've trained with a plane dice + focal loss combo on the provided hard masks. I've inferred too in full volumes whatever they were. I did 4 folds based on the 4 most available scrolls ids. With all I couldn't go higher than .461/.462 with that.\n\nAt first read I think loss and post-processing have played an important role. Thanks again for sharing a solution to actually learn something from.\n\nEDIT:\n\n>Data augmentation consists of random 3D flips along each spatial axis (p=0.5 each) and random 90-degree rotations in the (axis-1, axis-2) plane (p=0.5). CutMix (beta=1.0) was used in earlier model families but disabled for the final SEResNeXt152 runs, where it did not improve scores.\n\nI skipped cutmix and restricted flips to slices planes. So I did similar but modest augmentations. I've been surprised with some high scoring public notebooks with totally free rotations, those were terrible in my pipeline."
    },
    {
      "id": 3415675,
      "postDate": "2026-03-01T05:43:40.037Z",
      "content": "<p><a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> \nAmazing experiment. I really enjoyed reading your write-up. This approach is quite different from what others have explored. Congratulations on the solo win!</p>\n<p>Have you open-sourced your training and inference code? I would love to explore it for learning purposes.</p>\n<blockquote>\n  <p>For the keras based SEResnext model from medcai it was also 60% (!) faster for training. </p>\n</blockquote>\n<p>That is a huge improvement. Could you please elaborate on this in more detail? Did you run controlled benchmarks or compare training times against publicly available notebooks? It would really help me understand whether there is still room to optimize the medicai implementation.</p>",
      "rawMarkdown": "@christofhenkel \nAmazing experiment. I really enjoyed reading your write-up. This approach is quite different from what others have explored. Congratulations on the solo win!\n\nHave you open-sourced your training and inference code? I would love to explore it for learning purposes.\n\n> For the keras based SEResnext model from medcai it was also 60% (!) faster for training. \n\nThat is a huge improvement. Could you please elaborate on this in more detail? Did you run controlled benchmarks or compare training times against publicly available notebooks? It would really help me understand whether there is still room to optimize the medicai implementation.",
      "replies": [
        {
          "id": 3415689,
          "postDate": "2026-03-01T06:29:14.963Z",
          "content": "<p>I did not run a controlled benchmark, but I can share the before and after. I am really happy that medicai exist, it has a perfect repository structure and doumentation. Only problem is keras. Its important to note that I used keras with pytroch backend, and I guess that just does not work well. I got a memory leak with that model and used jit trace to fix it, so first I used this</p>\n<pre><code>import os\nos.environ[\"KERAS_BACKEND\"] = 'torch'\nfrom medicai.models import AttentionUNet\n\nmodel = AttentionUNet(encoder_name='seresnext50', input_shape=(160,160,160,1), classifier_activation=None, num_classes=1\ndummy_input = torch.randn(size=[1] + cfg.roi_size + [1])\ntraced_model = torch.jit.trace(model, dummy_input, strict=False, check_trace=False)\nself.backbone = traced_model\n</code></pre>\n<p>The reimplementation is attached. </p>",
          "rawMarkdown": "I did not run a controlled benchmark, but I can share the before and after. I am really happy that medicai exist, it has a perfect repository structure and doumentation. Only problem is keras. Its important to note that I used keras with pytroch backend, and I guess that just does not work well. I got a memory leak with that model and used jit trace to fix it, so first I used this\n\n```python\n\nimport os\nos.environ[\"KERAS_BACKEND\"] = 'torch'\nfrom medicai.models import AttentionUNet\n\nmodel = AttentionUNet(encoder_name='seresnext50', input_shape=(160,160,160,1), classifier_activation=None, num_classes=1\ndummy_input = torch.randn(size=[1] + cfg.roi_size + [1])\ntraced_model = torch.jit.trace(model, dummy_input, strict=False, check_trace=False)\nself.backbone = traced_model\n```\n\nThe reimplementation is attached. ",
          "votes": 1,
          "replies": [
            {
              "id": 3415992,
              "postDate": "2026-03-01T18:55:32.737Z",
              "content": "<p><a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> Thanks for the feedback and details response. </p>\n<p>About training speed, </p>\n<ul>\n<li>If you use the Keras built-in training API, you would call <code>model.compile()</code>, where <code>jit_compile=\"auto\"</code> is the default. According to the Keras documentation, this runs the model in eager mode. Setting <code>jit_compile=True</code> enables <code>torch.compile</code> with the Inductor backend. I haven’t tested this in this competition setting yet, but I suspect it could make a noticeable difference.</li>\n<li>If you use a custom training loop (like in the <a href=\"https://www.kaggle.com/code/ipythonx/train-vesuvius-surface-3d-detection-in-pytorch\" target=\"_blank\">code1</a> and <a href=\"https://www.kaggle.com/code/cdeotte/train-bronze-medal-uunet-by-chatgpt\" target=\"_blank\">code2</a> examples), you’re less likely to encounter memory leak issues. That said, I can’t be completely certain without benchmarking both approaches under the same conditions.</li>\n</ul>\n<p>About inference,</p>\n<p>You attempted to use <code>torch.jit.trace</code> to mitigate the memory leak when running with the <code>torch</code> backend. However, since the same model isn’t being optimized or exercised under other backends, there’s a possibility that some subtle or unexpected bugs are still present in the implementation. I’ll take a closer look into this.</p>",
              "rawMarkdown": "@christofhenkel Thanks for the feedback and details response. \n\nAbout training speed, \n\n- If you use the Keras built-in training API, you would call `model.compile()`, where `jit_compile=\"auto\"` is the default. According to the Keras documentation, this runs the model in eager mode. Setting `jit_compile=True` enables `torch.compile` with the Inductor backend. I haven’t tested this in this competition setting yet, but I suspect it could make a noticeable difference.\n- If you use a custom training loop (like in the [code1](https://www.kaggle.com/code/ipythonx/train-vesuvius-surface-3d-detection-in-pytorch) and [code2](https://www.kaggle.com/code/cdeotte/train-bronze-medal-uunet-by-chatgpt) examples), you’re less likely to encounter memory leak issues. That said, I can’t be completely certain without benchmarking both approaches under the same conditions.\n\nAbout inference,\n\nYou attempted to use `torch.jit.trace` to mitigate the memory leak when running with the `torch` backend. However, since the same model isn’t being optimized or exercised under other backends, there’s a possibility that some subtle or unexpected bugs are still present in the implementation. I’ll take a closer look into this."
            }
          ]
        }
      ]
    },
    {
      "id": 3415551,
      "postDate": "2026-03-01T01:52:14.280Z",
      "content": "<p>Congratulations! An elegant solution.</p>\n<p>SDF regression instead of binary classification is pretty interesting. Also are the birth and death coordinate based tunnel and bridge detection.</p>",
      "rawMarkdown": "Congratulations! An elegant solution.\n\nSDF regression instead of binary classification is pretty interesting. Also are the birth and death coordinate based tunnel and bridge detection."
    },
    {
      "id": 3417689,
      "postDate": "2026-03-06T00:56:16.563Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 3415597,
      "postDate": "2026-03-01T03:09:02.003Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3416754,
      "postDate": "2026-03-03T16:44:37.280Z",
      "content": "<p>Thank you so much for sharing!</p>",
      "rawMarkdown": "Thank you so much for sharing!"
    },
    {
      "id": 3416084,
      "postDate": "2026-03-02T02:48:43.773Z",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing."
    }
  ],
  "comments": [
    {
      "id": 3415538,
      "author_name": "Sean Johnson_SP",
      "author_url": "",
      "post_date": "2026-03-01T01:29:16.250000",
      "content": "<p>Thank you for sharing and congrats! the sdf regression in particular is interesting. i've never necessarily loved semantic segmentation for this task as the model has a tendency to just over \"blur\" tough regions, where this can provide a bit more signal to the model in terms of \"interpolating\" areas it has trouble separating. </p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 3416025,
      "author_name": "Cody_Null",
      "author_url": "",
      "post_date": "2026-03-01T21:11:34.050000",
      "content": "<p>Glad to see another gold with an interesting approach. I originally wanted to try out SEResNeXt152 type models but we had so much success with the nnunet structure and some of the custom models that we started with that I never even tried it. Very jealous of being able to infer on full 320,320,320 with 8xTTA. As our approach had so many steps by the time we got to post processing we only had about an hour for inference time. I would be curious to see how much impact that inferencing resolution had on your cv/lb? Another question I had was our model preformed worse when averaging on raw logits instead of the probabilities, despite the logits being the more logical approach, so I would be curious if you tried that. Love the post processing ideas as well. </p>\n<p>I really thought after seeing everyones solutions that we would have a solid pipeline and just missed with the post processing and that slid us down to 10th but so far all the post processing from other top competitiors actually just lowers our score. </p>\n<p>Congrats on your placement!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3416166,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2026-03-02T07:57:14.437000",
          "content": "<blockquote>\n  <p>Very jealous of being able to infer on full 320,320,320 with 8xTTA</p>\n</blockquote>\n<p>Reading your inference kernel it seems you used an overlap of (0.5,0.5,0.5) with a 160³ window. Thats 27 patches for a 320³ image and 3x the amount of voxels put through the model. So even with TTA8 my inference is 1.5x cheaper without creating border artifacts.</p>\n<p>My model simply has no probabilities. It predicts sdf values (ranging from -10 to 5) and then binary mask, by thresholding at 0.3. As <a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> explained you can create artificial probabilities  by introducing a temperature and do sigmoid(-sdf/T), but T is another hyperparameter and probs will be on a slightly different scale than those from BCE models. You can (probably) ensemble my model with other models by a voxel rank transform like we did in our <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/writeups/daddies-1st-place-solution-segmentation-with-partl\" target=\"_blank\">winning solution to the  CryoET competition</a> section Ensembling </p>",
          "votes": 1,
          "replies": [
            {
              "id": 3416402,
              "author_name": "Bartley",
              "author_url": "",
              "post_date": "2026-03-02T19:02:17.843000",
              "content": "<p>Congrats on 50 golds <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>, great achievement!</p>\n<p>Do you design your architecture/training pipeline any differently when the inference patch size differs from the training patch size?</p>\n<p>I often struggle to match the performance achieved when using the same patch size during training and inference (in this competition and past 3D competitions).</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3416418,
              "author_name": "Dieter",
              "author_url": "",
              "post_date": "2026-03-02T19:51:34.867000",
              "content": "<p>Its the first time I used much larger patch size for inference. Learned the trick from <a href=\"https://www.kaggle.com/bloodaxe\" target=\"_blank\">@bloodaxe</a> in CryoET comp. I think the metric in this competition favors not doing slided window inference (SWI) a bit because border artifacts and inhomogenous amount of voxel predictions (think about that with SWI some voxels are predicted more often than others) hurt the topo score. And what I saw here is that SWI hurt the score more than inconsistencies created by the patch size mismatch. E.g. I normalized each patch by mean/ std and thats inconsistent for different patch sizes. </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3416434,
              "author_name": "Cody_Null",
              "author_url": "",
              "post_date": "2026-03-02T21:40:03.887000",
              "content": "<p>Oh yeah we had a whole lot going on because we missed the median filter blurring idea. With it we get first, without it we use the full 9 hours and barely get 10th! All a learning experience :). I saw increased score for larger inference patch for BYU but just didnt have the time to even think about it for his comp</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3415750,
      "author_name": "Giorgio Angelotti",
      "author_url": "",
      "post_date": "2026-03-01T09:03:17.683000",
      "content": "<p>This solution is really brilliant and I like seeing architectures coming out of the nnUNet framework!\nI find the formulation as an SDF regression really elegant. When we launched this challenge we thought that some solutions could be inspirational also for another task we are working on: ink detection in the scrolls. I can see this solution very easily readapted for ink detection as well!</p>\n<p>I have two questions</p>\n<ol>\n<li><p>The part of the loss with mass from the SDF and the soft Dice made me think of <a href=\"https://arxiv.org/abs/1911.02278\" target=\"_blank\">a paper a read</a> a while ago where the authors said that in tasks with uncertainty in the labels, optimization of the soft dice can lead to biased estimates of the volume of the region to segment, while the cross entropy leads ofc to worse Dice score but unbiased volume estimates. Maybe an overestimation of the volume can lead to unwanted mergers, irregardless of the better Dice? You say that* Binary cross-entropy models consistently scored lower on the topology component*, but I am curious to know whether you tested also a BCE variant of the loss on a probability derived from the SDF, e.g. sigmoid(-sdf/T) (with T some temperature)?</p></li>\n<li><p>Within the architectures that you are using, the ResNe(X)ts have a lot of layers. Do you feel that this higher capacity is really needed for this task? I am also interested in the skip of the first maxpool to try to preserve finer structure. I think this can be even more important for ink!</p></li>\n</ol>\n<p>Thank you!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3416156,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2026-03-02T07:31:18.530000",
          "content": "<ol>\n<li><p>The paper author claim \"systematic under or overestimation of the predicted volume\". I guess this can be true. However note that I am using a binarization threshold of 0.3 which systematically predicts a slightly \"larger\" volume than the model predicted (higher sdf means further away from sheet center). This threshold is optimized via grid-search on CV and incoorporate trade-off of the 3 metrics used. In short, you can control merges with the threshold. I also tried more targeted prevention of merges via adding an sdf weight at voronoi lines, but it had no effect, because model was already really good at separation. I think reason is that sdf-approach results in really good skeletons. I tested probablity  derived from SDF, especially since this then can be used for frangii filter etc, but it performed worse and its tricky to tune T.</p></li>\n<li><p>Yes higher capacity has significant impact. I tend to start with very simple and fast models to optimize the number of experiments/ time for a competition and only increase capacity if really needed. </p></li>\n</ol>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3415640,
      "author_name": "Komil Parmar",
      "author_url": "",
      "post_date": "2026-03-01T04:23:40.267000",
      "content": "<p><a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> <strong>Sir Huge Congrats on achieving your 50th Competition Gold Medal. 🥳🎉</strong>\nThis is a remarkable milestone and truly a half-century of excellence.</p>\n<p>Your approach is amazing sir, and this is finally a solution that feels genuinely unique, with many things to learn from it\nMost other write-ups I read were either not as useful or did not share this much detail.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3417535,
      "author_name": "Durga Kumari",
      "author_url": "",
      "post_date": "2026-03-05T16:00:50.927000",
      "content": "<p>Brilliant combination of geometric learning, large-scale engineering, and topology-aware post-processing. \nTurning surface detection into an SDF regression problem and closing the loop with persistence homology is genuinely insightful. A masterclass in metric-aligned modeling.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3417695,
          "author_name": "",
          "author_url": "",
          "post_date": "2026-03-06T01:04:32.417000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3416206,
      "author_name": "Manas Choudhary",
      "author_url": "",
      "post_date": "2026-03-02T09:59:31.100000",
      "content": "<p>Really inspiring solution, initially I was expecting the final solutions to be made up of varying  architectures, methods like predicting sdf, vector fields, etc. but nnUNet somehow overpowered and people stuck to it so what I thought initially didn't happen.\nSticking to something unconventional is itself really cool, winning is a cherry on top. Congratss</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3415873,
      "author_name": "Ángel Jacinto Sánchez Ruiz",
      "author_url": "",
      "post_date": "2026-03-01T14:30:13.613000",
      "content": "<p>Thanks for share your non nnUNet/TransUNet top solution. I've been working with custom 3D Unet too but with resnet50 encoder and no attention decoder (in my experiments I didn't fount evidence of improvement). Also I've trained with a plane dice + focal loss combo on the provided hard masks. I've inferred too in full volumes whatever they were. I did 4 folds based on the 4 most available scrolls ids. With all I couldn't go higher than .461/.462 with that.</p>\n<p>At first read I think loss and post-processing have played an important role. Thanks again for sharing a solution to actually learn something from.</p>\n<p>EDIT:</p>\n<blockquote>\n  <p>Data augmentation consists of random 3D flips along each spatial axis (p=0.5 each) and random 90-degree rotations in the (axis-1, axis-2) plane (p=0.5). CutMix (beta=1.0) was used in earlier model families but disabled for the final SEResNeXt152 runs, where it did not improve scores.</p>\n</blockquote>\n<p>I skipped cutmix and restricted flips to slices planes. So I did similar but modest augmentations. I've been surprised with some high scoring public notebooks with totally free rotations, those were terrible in my pipeline.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3415675,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2026-03-01T05:43:40.037000",
      "content": "<p><a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> \nAmazing experiment. I really enjoyed reading your write-up. This approach is quite different from what others have explored. Congratulations on the solo win!</p>\n<p>Have you open-sourced your training and inference code? I would love to explore it for learning purposes.</p>\n<blockquote>\n  <p>For the keras based SEResnext model from medcai it was also 60% (!) faster for training. </p>\n</blockquote>\n<p>That is a huge improvement. Could you please elaborate on this in more detail? Did you run controlled benchmarks or compare training times against publicly available notebooks? It would really help me understand whether there is still room to optimize the medicai implementation.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3415689,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2026-03-01T06:29:14.963000",
          "content": "<p>I did not run a controlled benchmark, but I can share the before and after. I am really happy that medicai exist, it has a perfect repository structure and doumentation. Only problem is keras. Its important to note that I used keras with pytroch backend, and I guess that just does not work well. I got a memory leak with that model and used jit trace to fix it, so first I used this</p>\n<pre><code>import os\nos.environ[\"KERAS_BACKEND\"] = 'torch'\nfrom medicai.models import AttentionUNet\n\nmodel = AttentionUNet(encoder_name='seresnext50', input_shape=(160,160,160,1), classifier_activation=None, num_classes=1\ndummy_input = torch.randn(size=[1] + cfg.roi_size + [1])\ntraced_model = torch.jit.trace(model, dummy_input, strict=False, check_trace=False)\nself.backbone = traced_model\n</code></pre>\n<p>The reimplementation is attached. </p>",
          "votes": 1,
          "replies": [
            {
              "id": 3415992,
              "author_name": "Innat",
              "author_url": "",
              "post_date": "2026-03-01T18:55:32.737000",
              "content": "<p><a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> Thanks for the feedback and details response. </p>\n<p>About training speed, </p>\n<ul>\n<li>If you use the Keras built-in training API, you would call <code>model.compile()</code>, where <code>jit_compile=\"auto\"</code> is the default. According to the Keras documentation, this runs the model in eager mode. Setting <code>jit_compile=True</code> enables <code>torch.compile</code> with the Inductor backend. I haven’t tested this in this competition setting yet, but I suspect it could make a noticeable difference.</li>\n<li>If you use a custom training loop (like in the <a href=\"https://www.kaggle.com/code/ipythonx/train-vesuvius-surface-3d-detection-in-pytorch\" target=\"_blank\">code1</a> and <a href=\"https://www.kaggle.com/code/cdeotte/train-bronze-medal-uunet-by-chatgpt\" target=\"_blank\">code2</a> examples), you’re less likely to encounter memory leak issues. That said, I can’t be completely certain without benchmarking both approaches under the same conditions.</li>\n</ul>\n<p>About inference,</p>\n<p>You attempted to use <code>torch.jit.trace</code> to mitigate the memory leak when running with the <code>torch</code> backend. However, since the same model isn’t being optimized or exercised under other backends, there’s a possibility that some subtle or unexpected bugs are still present in the implementation. I’ll take a closer look into this.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3415551,
      "author_name": "Vineet K Reddy",
      "author_url": "",
      "post_date": "2026-03-01T01:52:14.280000",
      "content": "<p>Congratulations! An elegant solution.</p>\n<p>SDF regression instead of binary classification is pretty interesting. Also are the birth and death coordinate based tunnel and bridge detection.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3417689,
      "author_name": "",
      "author_url": "",
      "post_date": "2026-03-06T00:56:16.563000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3415597,
      "author_name": "",
      "author_url": "",
      "post_date": "2026-03-01T03:09:02.003000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3416754,
      "author_name": "Bhawesh Sinha",
      "author_url": "",
      "post_date": "2026-03-03T16:44:37.280000",
      "content": "<p>Thank you so much for sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3416084,
      "author_name": "dragon zhang",
      "author_url": "",
      "post_date": "2026-03-02T02:48:43.773000",
      "content": "<p>Thanks for sharing.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3415426": "First, let me thank the Vesuvius Challenge team and Kaggle for organizing this fascinating competition. \n\n\n## TLDR\n\n\nMy solution is an ensemble of UNets mainy based on a custom SEResNeXt152 encoder with an Attention UNet decoder. All models are trained to regress a signed distance field (SDF) rather than a binary mask. I am using a novel(?) SDF L1 + SDF mass loss which is the sdf equivalent of BCE + Dice. SDF predictions are averaged across checkpoints and up to 8 TTA flips. At the end all is binarized at an SDF threshold of 0.3, and then refined with an iterative persistence-homology-based tunnel-filling post-processing pipeline. Although training on 160x160x160 crops, I inferred on full 320x320x320 to prevent any artifacts from a sliding window approach. \n\n\n## Cross-Validation\n\n\nI used 4-fold cross-validation. During training I tracked three metrics locally - Surface Dice, VOI accelerated by GPU on original size, and the topology score on 10x downsampled size on CPU as the bettimatching algorithm is sequential. Neither local CV nor public LB correlated well with private LB, hence my selection of final submission went poorly. \n\n\n## Training Routine\n\n\nAll models are trained on 160^3 ROI patches randomly cropped from the npy saved 320^3 training volumes. SDF targets are calculated on the fly. Training uses the Adam optimizer with learning rate 1e-3, cosine annealing schedule. Mixed-precision training with bfloat16 is used throughout, and gradient checkpointing is enabled for the larger SEResNeXt152 models to fit within GPU memory. Weight clamping to the fp16-safe range is applied at checkpointing time to ensure clean float16 inference later. Batch size is 16.\n\n\nData augmentation consists of random 3D flips along each spatial axis (p=0.5 each) and random 90-degree rotations in the (axis-1, axis-2) plane (p=0.5). CutMix (beta=1.0) was used in earlier model families but disabled for the final SEResNeXt152 runs, where it did not improve scores.\n\n\n## Loss Function: SDF L1 and \"mass\"\n\n\nThe key design decision was to train models to regress a signed distance field rather than a binary segmentation mask. The SDF target is computed from the binary label via Euclidean distance transforms on the fly on GPU:\n\n\n```\nfg_edt = distance_transform_edt(foreground_mask)\nbg_edt = distance_transform_edt(~foreground_mask)\nsdf = bg_edt - fg_edt          # negative inside, positive outside, zero at boundary\nsdf = clamp(sdf, -100, 5)\n```\n\n\nThe loss has two components with equal weight:\n\n\n**Weighted SDF L1 loss (weight 0.5)**: A voxel-wise L1 loss between predicted and target SDF, weighted by a Gaussian bell-curve matrix that concentrates supervision near the surface boundary:\n\n\n```\nW(sdf) = 1 + w0 * exp(-((sdf + core_radius)^2) / (2 * sigma^2))\n```\n\n\nwith `w0 = 8`, `sigma = 4`, `core_radius = 5`. This gives a peak weight of 9.0 at `sdf = -5` (inside the surface at depth 5 voxels), plateauing at 9.0 for deeper interior, and decaying to ~1.0 far from the surface. The effect is that the model receives a strong gradient signal near the surface while distant background voxels contribute minimally.\n\n\n**SDF-based Dice loss (weight 0.5)**: Instead of thresholding to binary, a continuous \"mass\" is derived via `ReLU(-sdf)` for both prediction and target, and a soft Dice coefficient is computed from these masses. This provides a global volumetric overlap signal that complements the voxel-wise L1.\n\n\nThe following figure illustrates the three quantities at a central slice (axis_0 = 160) of a training volume:\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1424766%2Fd1f6c2cf94b96afe18f924803ac4b03e%2Fsdf_target_visualization.png?generation=1772316847712852&alt=media)\n\n\n*Left: original binary target (0 = background, 1 = foreground, 2 = ignore). Center: SDF target (negative inside surface, positive outside; black contour at SDF = 0). Right: Gaussian weight matrix (higher weights near and inside the surface boundary).*\n\nThis loss combined with gaussian weighting solves a lot of the key challenges of this competition in an elegant and efficient way: Sheet separation, sheet continuity, boundary sharpness, precise skeletons etc. \n\n\n## Model Architectures\n\n\nI explored a wide range of 3D architectures. Instead of relying on frameworks like nnUnet or medicai I asked an agent to recode relevant model families from scratch. This gives much more control. For the keras based SEResnext model from medcai it was also 60% (!) faster for training. The final ensemble draws from three architecture lineages:\n\n\n**SEResNeXt152 + Attention UNet** -- the main workhorse and strongest single model. A custom pure-PyTorch 3D SE-ResNeXt152 encoder (layers `[3, 8, 36, 3]`, groups=32, width_per_group=8, SE reduction=16, DropPath rate 0.3) paired with an Attention UNet decoder (channels `[256, 128, 64, 32, 16]`). Attention gates on skip connections let the decoder focus on relevant encoder features. Gradient checkpointing is essential here due to the 36-block third stage. One variant using `seresnext152_48x8d` (48 groups, width 8) for additional capacity.\n\n\n**ResNet152 + skip maxpool** -- Expensive UNet with a ResNet152 encoder, custom decoder channel configuration, and deep supervision during training (auxiliary losses at intermediate decoder stages). I skip one early maxpool so resolution through the model is higher. At inference, only the final head is used.\n\n\n**ResNet152 + UNet with SDF loss heavy augs and external data** -- earlier ResNet152-based UNets trained with the SDF + Dice loss. One model extends the training set with additional scrolls. I used an additional binary input channel to give the model pixelwise information that this data is external (and hence labels are \"bad\") These models provide useful diversity to the ensemble despite being individually weaker than SEResNeXt152. \n\n\nI also experimented with SwinUNETR, UNETR++, MedNeXt, SegFormer3D, ConvNeXt V2, U-Mamba, and SEResNeXt200, but none surpassed the SEResNeXt152 + AttUNet architecture on local validation. Some topology-aware loss variants (persistence diagram matching loss, Euler characteristic loss) were explored. One PD-matching variant is included in the final ensemble but its finetuned only for a few epochs from another model because training was very slow.\n\n\n## Ensemble\n\n\nPredictions are combined by simple averaging of raw SDF logits across all checkpoints and TTA augmentations per sample. For the larger SEResNeXt152 models, TTA is reduced to 2 augmentations (identity + triple-flip) to stay within the 9-hour Kaggle kernel runtime; smaller models use the full 8-flip TTA (3 single-axis flips + 3 dual-axis + 1 triple-axis + identity).\n\n\nNaN safety is critical: the SEResNeXt152 models are trained in bfloat16 but inference runs in float16. A safety conversion routine clamps all weights to `[-65504, 65504]` and registers forward hooks that replace any NaN/Inf activations.\n\n\n## Inference\n\n\nAll inference is done on full 320x320x320 image! No strided slices, as those create artifacts that are very harmful to topology.\nTwo GPUs are used in parallel via data sharding: even-indexed test samples go to GPU 0, odd-indexed to GPU 1. Each GPU processes all models sequentially for its shard and offload predictions to disk to prevent OOM. \n\n\n## Post-Processing\n\n\nPost-processing proved essential for the topology score, which heavily penalizes tunnels (H1 topological features) in the predicted surface. My pipeline:\n\n\n**Step 1 -- Binarization**: Threshold the averaged SDF logits at 0.3 (voxels with SDF < 0.3 become foreground).\n\n\n**Step 2 -- Dust removal**: Remove small connected components (< 50,000 voxels, 6-connectivity), preserving components that touch at least 3 volume boundaries regardless of size.\n\n\n**Step 3 -- Iterative H1 tunnel filling** (13 iterations):\n\n\n1. **SDF filtration**: Overwrite foreground voxels in the raw SDF to `threshold - 1.0`, creating a cubical filtration where the surface sits at the threshold.\n2. **Persistence barcode computation**: Split the volume into 2x2x2 = 8 octant tiles, compute persistence homology on each tile using a custom C++ module (`barcode3d_fast_v3`) which was derived from betti matching library. This identifies H1 features (tunnels) with birth/death values and coordinates.\n3. **Straddling filter**: Keep only H1 features where `birth < sdf_threshold < death`, i.e., tunnels that cross the binarization surface.\n4. **Adaptive radius**: For each tunnel's death coordinate, estimate the tunnel width from the background EDT and set the fill radius to `clip(edt_width + 1.0, 3.5, 7.0)`.\n5. **Bridge detection**: Simulate filling a ball at each death coordinate and check whether it would merge separate connected components. Skip coordinates that would create bridges.\n6. **Ball filling**: Fill spherical balls at the remaining death coordinates. On the final iteration, also apply morphological hole filling.\n7. **Component count guard**: If filling reduced the number of connected components (accidental merge), revert to the pre-fill state.\n\n\n## Key Insights\n\n\n- **SDF regression > binary classification** for this metric suite. The SDF naturally encodes distance-to-surface information, and thresholding the SDF at different values during validation gives a smooth trade-off between Surface Dice and topology metrics. Binary cross-entropy models consistently scored lower on the topology component.\n- **Gaussian weighting** focuses the loss on the skeleton region, but still extents to sheet boundary. Background is heavily downweighted. \n- **Iterative tunnel filling** with bridge detection is the single most impactful post-processing step. Going from 0 to 13 iterations improved the topology score by ~0.08 on local validation, with diminishing returns beyond 11-13 iterations.\n\n\n\n\n\n\n\n\n\n",
    "3415538": "Thank you for sharing and congrats! the sdf regression in particular is interesting. i've never necessarily loved semantic segmentation for this task as the model has a tendency to just over \"blur\" tough regions, where this can provide a bit more signal to the model in terms of \"interpolating\" areas it has trouble separating. ",
    "3416025": "Glad to see another gold with an interesting approach. I originally wanted to try out SEResNeXt152 type models but we had so much success with the nnunet structure and some of the custom models that we started with that I never even tried it. Very jealous of being able to infer on full 320,320,320 with 8xTTA. As our approach had so many steps by the time we got to post processing we only had about an hour for inference time. I would be curious to see how much impact that inferencing resolution had on your cv/lb? Another question I had was our model preformed worse when averaging on raw logits instead of the probabilities, despite the logits being the more logical approach, so I would be curious if you tried that. Love the post processing ideas as well. \n\nI really thought after seeing everyones solutions that we would have a solid pipeline and just missed with the post processing and that slid us down to 10th but so far all the post processing from other top competitiors actually just lowers our score. \n\nCongrats on your placement!",
    "3415750": "This solution is really brilliant and I like seeing architectures coming out of the nnUNet framework!\nI find the formulation as an SDF regression really elegant. When we launched this challenge we thought that some solutions could be inspirational also for another task we are working on: ink detection in the scrolls. I can see this solution very easily readapted for ink detection as well!\n\nI have two questions\n1. The part of the loss with mass from the SDF and the soft Dice made me think of [a paper a read](https://arxiv.org/abs/1911.02278) a while ago where the authors said that in tasks with uncertainty in the labels, optimization of the soft dice can lead to biased estimates of the volume of the region to segment, while the cross entropy leads ofc to worse Dice score but unbiased volume estimates. Maybe an overestimation of the volume can lead to unwanted mergers, irregardless of the better Dice? You say that* Binary cross-entropy models consistently scored lower on the topology component*, but I am curious to know whether you tested also a BCE variant of the loss on a probability derived from the SDF, e.g. sigmoid(-sdf/T) (with T some temperature)?\n\n2. Within the architectures that you are using, the ResNe(X)ts have a lot of layers. Do you feel that this higher capacity is really needed for this task? I am also interested in the skip of the first maxpool to try to preserve finer structure. I think this can be even more important for ink!\n\nThank you!",
    "3415640": "@christofhenkel **Sir Huge Congrats on achieving your 50th Competition Gold Medal. 🥳🎉**\nThis is a remarkable milestone and truly a half-century of excellence.\n\nYour approach is amazing sir, and this is finally a solution that feels genuinely unique, with many things to learn from it\nMost other write-ups I read were either not as useful or did not share this much detail.",
    "3417535": "Brilliant combination of geometric learning, large-scale engineering, and topology-aware post-processing. \nTurning surface detection into an SDF regression problem and closing the loop with persistence homology is genuinely insightful. A masterclass in metric-aligned modeling.",
    "3416206": "Really inspiring solution, initially I was expecting the final solutions to be made up of varying  architectures, methods like predicting sdf, vector fields, etc. but nnUNet somehow overpowered and people stuck to it so what I thought initially didn't happen.\nSticking to something unconventional is itself really cool, winning is a cherry on top. Congratss",
    "3415873": "Thanks for share your non nnUNet/TransUNet top solution. I've been working with custom 3D Unet too but with resnet50 encoder and no attention decoder (in my experiments I didn't fount evidence of improvement). Also I've trained with a plane dice + focal loss combo on the provided hard masks. I've inferred too in full volumes whatever they were. I did 4 folds based on the 4 most available scrolls ids. With all I couldn't go higher than .461/.462 with that.\n\nAt first read I think loss and post-processing have played an important role. Thanks again for sharing a solution to actually learn something from.\n\nEDIT:\n\n>Data augmentation consists of random 3D flips along each spatial axis (p=0.5 each) and random 90-degree rotations in the (axis-1, axis-2) plane (p=0.5). CutMix (beta=1.0) was used in earlier model families but disabled for the final SEResNeXt152 runs, where it did not improve scores.\n\nI skipped cutmix and restricted flips to slices planes. So I did similar but modest augmentations. I've been surprised with some high scoring public notebooks with totally free rotations, those were terrible in my pipeline.",
    "3415675": "@christofhenkel \nAmazing experiment. I really enjoyed reading your write-up. This approach is quite different from what others have explored. Congratulations on the solo win!\n\nHave you open-sourced your training and inference code? I would love to explore it for learning purposes.\n\n> For the keras based SEResnext model from medcai it was also 60% (!) faster for training. \n\nThat is a huge improvement. Could you please elaborate on this in more detail? Did you run controlled benchmarks or compare training times against publicly available notebooks? It would really help me understand whether there is still room to optimize the medicai implementation.",
    "3415551": "Congratulations! An elegant solution.\n\nSDF regression instead of binary classification is pretty interesting. Also are the birth and death coordinate based tunnel and bridge detection.",
    "3417689": "",
    "3415597": "",
    "3416754": "Thank you so much for sharing!",
    "3416084": "Thanks for sharing."
  }
}