{
  "id": 679227,
  "title": "10th place solution ",
  "url": "/competitions/vesuvius-challenge-surface-detection/discussion/679227",
  "author_name": "Tom",
  "post_date": "2026-02-28T02:02:19.265000",
  "votes": 27,
  "comment_count": 30,
  "views": 0,
  "content": "<p>We thank Kaggle and the Vesuvius Challenge organizers for a fascinating competition.</p>\n<p>Although majority of training pipeline was implemented outside the original nnUNet framework, most architectures are derived from work produced by MIC-DKFZ. We are grateful to them for advancing the 3D segmentation domain.</p>\n<p>Below is an overview of our approach.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4310004%2Ff83bcfcfbbbbd3d8bad0af47c331cae8%2Ff45.png?generation=1773033543025607&amp;alt=media\" alt=\"\"></p>\n<h2>1st Stage — Initial Segmentation</h2>\n<p>We trained three independent models:</p>\n<ul>\n<li><strong>ResEnc-L UNet</strong> (4 × TTA): The standard nnUNet-style residual encoder UNet with channels (32, 64, 128, 256, 320, 320) and <code>n_blocks=(1,3,4,6,6,6)</code>.</li>\n<li><strong>Primus-B</strong>: A transformer based segmentation model from dynamic-network-architectures library of MIC-DKFZ.</li>\n<li><strong>Primus-B V2</strong>: An improved variant of Primus-B.</li>\n</ul>\n<p>Models were first pretrained on approximately 1,600 annotated images from other scrolls: <a href=\"https://dl.ash2txt.org/datasets/seg-derived-recto-surfaces/\" target=\"_blank\">https://dl.ash2txt.org/datasets/seg-derived-recto-surfaces/</a>. Labels were slightly different but helped fine-tuning afterward to converge much faster. For example, Primus was taking 700 epochs or more to fully converge, and with pretraining it took 400.</p>\n<p>All models were trained at patch size 160³, with AdamW (<code>lr=5e-5</code>, <code>wd=1e-4</code>) and mixed-precision training for approximately 400 epochs.</p>\n<p>For losses, we used a combination of Dice Loss, Cross-Entropy Loss, Skeleton Recall Loss, and Surface Dice Loss (a modified version of clDice for 2D manifolds in 3D environments; implementation was for CryoET membranes). <em>(Note <a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a>: I had the opportunity to watch Lorenz Lamm's presentation at the CZII competition workshop, very nice to find a use almost a year later in a totally different task.)</em></p>\n<p>From local evaluation, the best models were  Primus-B V2, ResEnc-L and Primus-B respectively. We will add scores later. The good thing is that the models are very complementary to each other.</p>\n<h2>2nd Stage — Ensemble</h2>\n<p>The idea here was to let a model learn the complementarity between predictions. The model used a 4-channel input: [Image, ResEnc-L binary mask, Primus binary mask, PrimusV2 binary mask].</p>\n<ul>\n<li><strong>ResEnc-L UNet</strong> ensembles initial stage predictions.</li>\n</ul>\n<p>We trained with a random threshold on the fly ranging from 0.1 to 0.7, but used a threshold of 0.3 at inference.</p>\n<h2>3rd Stage — Refinement</h2>\n<p><em>(Note <a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a>: I really liked the discussions about algorithms to fill holes, make the sheets more continuous, and so on. But I'm not clever enough to do this with an algorithm. So I did what I could: train a model to do that.)</em></p>\n<p>We used the same random threshold strategy on the fly. If the threshold is low, we hope the model learns to correct merging sheets, if it is high, we hope the model learns to correct broken sheets.\nThis behavior can be seen in the image below:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2221915%2F6351f874a7cd7b5169eb2f874ebf62e7%2FCaptura%20de%20tela%202026-03-09%20165943.png?generation=1773086517474158&amp;alt=media\" alt=\"\"></p>\n<p>We also added a strong <code>randCoarseDropout</code> to mask the input to help model learn continuity and fix holes.</p>\n<p>Input is [Mask, 2nd stage binary].</p>\n<ul>\n<li>ResEnc-L UNet refinement.</li>\n</ul>\n<h2>4th Stage — Refinement</h2>\n<p>Just another refinement network. Same idea of stage 3, help model learn to correct silly errors,</p>\n<p>Input is [Mask, 3rd stage binary].</p>\n<ul>\n<li>ResEnc-L UNet — refinement.</li>\n</ul>\n<h2>5th Stage — Diffeomorphic Stage (biggest boost)</h2>\n<p>The diffeomorphic network takes the role of shape calibratin and thickness modification. This approach refers Tom’s paper (<strong>Perceptual Contrastive Generative Adversarial Network based on image warping for unsupervised image-to-image translation</strong>) and the <strong>FlowNet2</strong>. Unlike classic way like previous stage which just reproduces a better segmentation prediction, this network predicts the stationary velocity field which can manipulate the input mask previous stage it recieves. Intuitively, it is a vector field that determine how the mesh changes in 3d space, and we decide how many step it should move. </p>\n<h3>Core Designs and Intuition</h3>\n<p>This model is mainly based on Tom’s paper. It performs warping first and then refines the content afterward.</p>\n<p>Diffeomorphic Step:</p>\n<p>Using a lower threshold to generate the hard mask and applying uniform blur are two crucial steps for helping the model learn the SVF logits. Since the OOF mask is already strong, you need to intentionally introduce “errors” so the model can learn from them and be rewarded for correcting them.</p>\n<ul>\n<li>Using a small threshold reveals more suppressed predictions. Some of these predictions contribute to thicker mask regions, which allows the model to identify and fix such issues.</li>\n<li>A hard mask alone cannot provide useful gradients to the model. We also found that warping a blurred mask is very beneficial for improving topology and boosting VOI performance.</li>\n</ul>\n<p>Blurring Example\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4310004%2F6ca13dd6aa7ec55274b64e0193386444%2Ff1.png?generation=1772451933189014&amp;alt=media\" alt=\"\"></p>\n<pre><code># ----------------------------------------\n# 1. Blur the input mask\n# ----------------------------------------\nHard_Mask = Soft_Mask&gt;0.3\n\nBlurred_mask = Gaussian_Blur(Hard_Mask, kernel_size=3, sigma=5)\n# (Produces soft mask M in [0,1])\n\n# ----------------------------------------\n# 2. Predict stationary velocity field\n# ----------------------------------------\n\ngamma = Diffeomorphic_Prediction(Blurred_mask)\n# gamma ∈ R^{3×D×H×W}  (SVF logits)\n\nv = tanh(gamma) * max_v\n# v : Ω → R³\n# bounded stationary velocity field\n\n# ----------------------------------------\n# 3. Scaling &amp; Squaring (Lie exponential)\n# ----------------------------------------\n\n# Initialize small deformation\nphi_0 = v / (2^N)\n\nphi = phi_0\n\nRepeat N times:\n\n    # Compose deformation with itself\n    # (φ ∘ φ)(x) = φ(x + φ(x))\n\n    phi = phi + Warp(phi, phi)\n\n# After N steps:\n# phi ≈ exp(v)\n\n# ----------------------------------------\n# 4. Warp mask with diffeomorphic transform\n# ----------------------------------------\n\nWarped_mask(x) = Blurred_mask(x + phi(x))\n</code></pre>\n<p>Signed Distance Topology Shift Prediction:</p>\n<p>Optical flow–based methods usually suffer from a “folding” issue. This occurs when the model makes abrupt deformations, which significantly damage the topology. Tom struggled with this problem for quite some time. He eventually addressed it by introducing an additional output channel that predicts the “shift” in the SDF space of the warped mask.</p>\n<pre><code># ----------------------------------------\n# Topology-aware correction\n# ----------------------------------------\n\n# Convert to soft signed distance\nSDF = log(Warped_mask + ε) - log(1 - Warped_mask + ε)\n\n# Learn correction field from 4th channel r_t\nt      = sigmoid(r_t)                 # gating map\ndelta  = max_offset * tanh(r_t)       # signed bounded offset\n\nSDF_corrected = SDF + t * delta\n\nFinal_mask = sigmoid(SDF_corrected)\n</code></pre>\n<p>Loss functions for warping</p>\n<ul>\n<li>Minimizing SVF Smoothing to avoid folding by surpressing the magnitutes of velocities</li>\n</ul>\n<pre><code>def svf_smoothness(self, v):\n    dz = (v[:, :, 1:] - v[:, :, :-1]).pow(2).mean()\n    dy = (v[:, :, :, 1:] - v[:, :, :, :-1]).pow(2).pow(1).mean()\n    dx = (v[:, :, :, :, 1:] - v[:, :, :, :, :-1]).pow(2).mean()\n\n    return (dz + dy + dx) / 3.0\n</code></pre>\n<ul>\n<li>Minimizing jacobian log barrier to keep flow active and preventing a compressible elastic material that strongly resists collapsing or flipping</li>\n</ul>\n<pre><code>def jacobian_determinant(flow):\n    \"\"\"\n    flow: (B, 3, D, H, W) displacement field u(x)\n    returns: (B, D, H, W) jacobian determinant of φ(x)=x+u(x)\n    \"\"\"\n    B, C, D, H, W = flow.shape\n    assert C == 3\n\n    # gradients wrt spatial axes (z = depth, y = height, x = width)\n    du_dx = torch.gradient(flow, dim=4)[0]  # width axis\n    du_dy = torch.gradient(flow, dim=3)[0]  # height axis\n    du_dz = torch.gradient(flow, dim=2)[0]  # depth axis\n\n    # components\n    ux_x = du_dx[:,0]; ux_y = du_dy[:,0]; ux_z = du_dz[:,0]\n    uy_x = du_dx[:,1]; uy_y = du_dy[:,1]; uy_z = du_dz[:,1]\n    uz_x = du_dx[:,2]; uz_y = du_dy[:,2]; uz_z = du_dz[:,2]\n\n    # deformation gradient J = I + ∇u\n    j11 = 1 + ux_x; j12 =     ux_y; j13 =     ux_z\n    j21 =     uy_x; j22 = 1 + uy_y; j23 =     uy_z\n    j31 =     uz_x; j32 =     uz_y; j33 = 1 + uz_z\n\n    det = (\n        j11 * (j22 * j33 - j23 * j32)\n        - j12 * (j21 * j33 - j23 * j31)\n        + j13 * (j21 * j32 - j22 * j31)\n    )\n    return det\n\ndef jacobian_log_barrier(flow, eps=1e-6):\n    det = jacobian_determinant(flow)\n    det_clamped = torch.clamp(det, min=eps)\n    loss = -torch.log(det_clamped).mean()\n    return loss\n</code></pre>\n<p>Loss function for SDF topology fixing </p>\n<p>The challenge <a href=\"https://www.kaggle.com/Tom\" target=\"_blank\">@Tom</a> faces is avoiding the forth channel cheating and dominating the prediction, which letting the diffeomorphic step useless. We use three loss functions to solve this issue.</p>\n<ul>\n<li>Sparisty loss: Keep sparse topology correction, discouraged from editing everywhere and only activates correction when necessary</li>\n</ul>\n<pre><code>#suppose t = gating map in topoFix\ndef topo_sparsity(self, t):\n        return t.mean()\n</code></pre>\n<ul>\n<li>Total variation loss: Making topology edits to be region-based, not noisy voxel-wise toggles</li>\n</ul>\n<pre><code>def topo_tv(self, t):\n    dz = (t[:, :, 1:] - t[:, :, :-1]).abs().mean()\n    dy = (t[:, :, :, 1:] - t[:, :, :, :-1]).abs().mean()\n    dx = (t[:, :, :, :, 1:] - t[:, :, :, :, :-1]).abs().mean()\n\n    return (dz + dy + dx) / 3.0\n</code></pre>\n<ul>\n<li>Boundary loss: Encourage edits near zero-level set and only modify topology near the object surface</li>\n</ul>\n<pre><code>def topo_boundary(self, t, sdf):\n    boundary = torch.exp(-sdf.abs())\n    return (t * (1.0 - boundary)).mean()\n</code></pre>\n<h3>Variations of Implementations</h3>\n<p>We have two implementation of Diffeomorphic Network, one is in <a href=\"https://www.kaggle.com/Tom\" target=\"_blank\">@Tom</a> repositery and another is <a href=\"https://www.kaggle.com/Sergio\" target=\"_blank\">@Sergio</a> repositery, they have different augmentation, processing and loss functions. All the network architecture are all identical, using the nnUNet-style residual encoder UNet but outputs with 4 channels. </p>\n<ul>\n<li>Loss differneces</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Version</th>\n<th>CLDice</th>\n<th>SoftSDFLoss</th>\n<th>Skeleton Recall</th>\n<th>Dice + CE</th>\n<th>Surface Dice Loss</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Tom</td>\n<td>o</td>\n<td>o</td>\n<td>o</td>\n<td>o</td>\n<td>x</td>\n</tr>\n<tr>\n<td>Sergio</td>\n<td>x</td>\n<td>x</td>\n<td>o</td>\n<td>o</td>\n<td>o</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>Augmentation differences</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Version</th>\n<th>Online thresholding</th>\n<th>Mask augmentation</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Tom</strong></td>\n<td><code>threshold = 0.3 + torch.randn(1, device=prob_mask_oof.device) * 0.01</code><br><code>threshold = torch.clamp(threshold, 0.1, 0.5)</code></td>\n<td>• Random Affine (3° rotation)<br>• RandCoarseDropout (12 holes, spatial size = 10)</td>\n</tr>\n<tr>\n<td><strong>Sergio</strong></td>\n<td><code>threshold = 0.3 + np.random.uniform(-0.2, 0.5)</code></td>\n<td><code>RandCoarseDropoutdWithRanges(keys=oof_keys, prob=1, shared_holes_range=(10,15), independent_holes_range=(15,15), spatial_size_range=(10,30), fill_value=0.0)</code><br><br><code>RandCoarseDropoutdWithRanges(keys=oof_keys, prob=0.8, shared_holes_range=(20,40), independent_holes_range=(30,60), spatial_size_range=(1,3), fill_value=0.0)</code></td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li><p>Prediciton results: Sergio's vs Tom's\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4310004%2F7d3f271456fe68aa47f7a89350dae37c%2Ff44.png?generation=1772452069836953&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4310004%2F0b7202034b6a6d1bd159c1793043ef74%2Ff2.png?generation=1772451950172111&amp;alt=media\" alt=\"\"></p></li>\n<li><p>Learned components: Sergio's vs Tom's\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4310004%2F2512060ffeeebfed8d8326580e73cdac%2Ff45.png?generation=1772452213966109&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4310004%2Fb250011f987349d85a1191d0c3501027%2Ff3.png?generation=1772452235538826&amp;alt=media\" alt=\"\"></p></li>\n<li><p>SVF results: Sergio's vs Tom's\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4310004%2Fcd5ed636e60ebaac44005da3ea7508a7%2Ffig1.png?generation=1772452012248337&amp;alt=media\" alt=\"\"></p></li>\n</ul>\n<h3>Comparing different diffeomorphic setup</h3>\n<p>Before the deadline, we conducted simulation trials to approximate the public subset (40 samples per subset, 200 trials) and compared the diffeomorphic networks from the repo of <a href=\"https://www.kaggle.com/Tom\" target=\"_blank\">@Tom</a> and <a href=\"https://www.kaggle.com/Sergio\" target=\"_blank\">@Sergio</a>.</p>\n<p>As shown in the plot, there is a clear difference in metric preference between the two approaches. <a href=\"https://www.kaggle.com/Tom\" target=\"_blank\">@Tom</a>’s model outperforms <a href=\"https://www.kaggle.com/Sergio\" target=\"_blank\">@Sergio</a>’s on Surface Dice, while <a href=\"https://www.kaggle.com/Sergio\" target=\"_blank\">@Sergio</a>’s setup achieves better results on Topology and VOI scores. In most cases, <a href=\"https://www.kaggle.com/Sergio\" target=\"_blank\">@Sergio</a>’s configuration achieves the best overall performance.</p>\n<p>However, we also observed that in a few subsets, <a href=\"https://www.kaggle.com/Tom\" target=\"_blank\">@Tom</a>’s overall score is higher, which is consistent with its LB. This suggests that the public subset may resemble those specific subsets where <a href=\"https://www.kaggle.com/Tom\" target=\"_blank\">@Tom</a>’s model performs better, and a small number of samples could have disproportionately boosted the public score, essentially a favorable sampling effect.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4310004%2F142aa44e7586adc8190b77670022c9ad%2F112.png?generation=1772418585149240&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Validation results</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>CV</th>\n<th>LB</th>\n<th>PB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Tom</td>\n<td>0.619</td>\n<td>0.595</td>\n<td>0.612</td>\n</tr>\n<tr>\n<td>Tom + Sergio</td>\n<td>0.621</td>\n<td>0.588</td>\n<td>0.615</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>Hypothesis testing (H0: <a href=\"https://www.kaggle.com/Tom\" target=\"_blank\">@Tom</a> = <a href=\"https://www.kaggle.com/Sergio\" target=\"_blank\">@Sergio</a>, H1: <a href=\"https://www.kaggle.com/Tom\" target=\"_blank\">@Tom</a> &gt; <a href=\"https://www.kaggle.com/Sergio\" target=\"_blank\">@Sergio</a> )</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Metric</th>\n<th>Mean Difference</th>\n<th>T-Statistic</th>\n<th>P-Value</th>\n<th>Reject?</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Surface_dice</td>\n<td>+0.0086</td>\n<td>74.97</td>\n<td>&lt;0.0001</td>\n<td>Y</td>\n</tr>\n<tr>\n<td>Score</td>\n<td>-0.0017</td>\n<td>-7.64</td>\n<td>1</td>\n<td>N</td>\n</tr>\n<tr>\n<td>Topo</td>\n<td>-0.0124</td>\n<td>-20.11</td>\n<td>1</td>\n<td>N</td>\n</tr>\n<tr>\n<td>Voi</td>\n<td>-0.004</td>\n<td>-117.68</td>\n<td>1</td>\n<td>N</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>Based on the hypothesis testing, these two methods may not have a significant difference in topology or VOI score on the public subset. In that case, the method with the higher surface Dice score dominates, resulting in a 0.007 boost on the leaderboard. The other one we hoped improving lb experiences a substantial drop, scoring 0.588, which does bad in surface dice.</li>\n</ul>\n<h2>Post-Processing</h2>\n<table>\n<thead>\n<tr>\n<th>Post Processing</th>\n<th>LB</th>\n<th>PB</th>\n<th>Comments</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>remove cc &lt; 3000</td>\n<td>0.592</td>\n<td>0.614</td>\n<td></td>\n</tr>\n<tr>\n<td>x7 median filter</td>\n<td>0.596</td>\n<td>0.623</td>\n<td></td>\n</tr>\n<tr>\n<td>x6 median filter</td>\n<td>0.597</td>\n<td>0.623</td>\n<td></td>\n</tr>\n<tr>\n<td>x8 median filter</td>\n<td>0.596</td>\n<td>0.624</td>\n<td></td>\n</tr>\n<tr>\n<td>x9 median filter</td>\n<td>0.596</td>\n<td>0.624</td>\n<td></td>\n</tr>\n<tr>\n<td>x10 median filter</td>\n<td>0.597</td>\n<td>0.624</td>\n<td></td>\n</tr>\n<tr>\n<td>binary closing</td>\n<td>0.588</td>\n<td>0.615</td>\n<td></td>\n</tr>\n<tr>\n<td>closing(7)+hole patching+cavity fill</td>\n<td>0.583</td>\n<td>0.604</td>\n<td></td>\n</tr>\n<tr>\n<td>Gaussian smooth + rethreshold + close/open + fill</td>\n<td>0.585</td>\n<td>0.603</td>\n<td></td>\n</tr>\n<tr>\n<td>Remove internal cavities</td>\n<td>0.589</td>\n<td>0.614</td>\n<td></td>\n</tr>\n<tr>\n<td>Erosion / Dilation</td>\n<td>0.592</td>\n<td>0.613</td>\n<td></td>\n</tr>\n<tr>\n<td>Enforced minimum sheet thickness</td>\n<td>0.463</td>\n<td>0.454</td>\n<td>Topo destroyed; Fails small sample</td>\n</tr>\n<tr>\n<td>Z-consistent smoothing</td>\n<td>0.592</td>\n<td>0.615</td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h2>Code</h2>\n<p><a href=\"https://github.com/Sersasj/vesuvius-challenge-10th-solution\" target=\"_blank\">https://github.com/Sersasj/vesuvius-challenge-10th-solution</a></p>\n<h2>References</h2>\n<ul>\n<li>Primus: Enforcing Attention Usage for 3D Medical Image Segmentation — <a href=\"https://openreview.net/forum?id=YWwGmmObri\" target=\"_blank\">https://openreview.net/forum?id=YWwGmmObri</a></li>\n<li>MemBrain v2: An end-to-end tool for the analysis of membranes in cryo-electron tomography (Surface Dice Loss) — <a href=\"https://www.biorxiv.org/content/10.1101/2024.01.05.574336v1.full.pdf\" target=\"_blank\">https://www.biorxiv.org/content/10.1101/2024.01.05.574336v1.full.pdf</a></li>\n<li>Skeleton Recall Loss for Connectivity Conserving and Resource Efficient Segmentation of Thin Tubular Structures — <a href=\"https://www.ecva.net/papers/eccv_2024/papers_ECCV/papers/09904.pdf\" target=\"_blank\">https://www.ecva.net/papers/eccv_2024/papers_ECCV/papers/09904.pdf</a></li>\n<li>dynamic-network-architectures library — <a href=\"https://github.com/MIC-DKFZ/dynamic-network-architectures\" target=\"_blank\">https://github.com/MIC-DKFZ/dynamic-network-architectures</a></li>\n<li>PrimusV2 code (for some reason it's not merged, we got lucky to find this fork) — <a href=\"https://github.com/TaWald/dynamic-network-architectures/blob/main/dynamic_network_architectures/architectures/primus.py\" target=\"_blank\">https://github.com/TaWald/dynamic-network-architectures/blob/main/dynamic_network_architectures/architectures/primus.py</a></li>\n<li><strong>Perceptual Contrastive Generative Adversarial Network based on image warping for unsupervised image-to-image translation(</strong><a href=\"https://www.sciencedirect.com/science/article/abs/pii/S0893608023003684**\" target=\"_blank\">https://www.sciencedirect.com/science/article/abs/pii/S0893608023003684**</a>)**</li>\n<li><strong>FlowNet 2.0: Evolution of Optical Flow Estimation with Deep Networks(</strong><a href=\"https://arxiv.org/abs/1612.01925**\" target=\"_blank\">https://arxiv.org/abs/1612.01925**</a>)**</li>\n</ul>",
  "messages": [
    {
      "id": 3414960,
      "postDate": "2026-02-28T02:02:19.267Z",
      "content": "<p>We thank Kaggle and the Vesuvius Challenge organizers for a fascinating competition.</p>\n<p>Although majority of training pipeline was implemented outside the original nnUNet framework, most architectures are derived from work produced by MIC-DKFZ. We are grateful to them for advancing the 3D segmentation domain.</p>\n<p>Below is an overview of our approach.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4310004%2Ff83bcfcfbbbbd3d8bad0af47c331cae8%2Ff45.png?generation=1773033543025607&amp;alt=media\" alt=\"\"></p>\n<h2>1st Stage — Initial Segmentation</h2>\n<p>We trained three independent models:</p>\n<ul>\n<li><strong>ResEnc-L UNet</strong> (4 × TTA): The standard nnUNet-style residual encoder UNet with channels (32, 64, 128, 256, 320, 320) and <code>n_blocks=(1,3,4,6,6,6)</code>.</li>\n<li><strong>Primus-B</strong>: A transformer based segmentation model from dynamic-network-architectures library of MIC-DKFZ.</li>\n<li><strong>Primus-B V2</strong>: An improved variant of Primus-B.</li>\n</ul>\n<p>Models were first pretrained on approximately 1,600 annotated images from other scrolls: <a href=\"https://dl.ash2txt.org/datasets/seg-derived-recto-surfaces/\" target=\"_blank\">https://dl.ash2txt.org/datasets/seg-derived-recto-surfaces/</a>. Labels were slightly different but helped fine-tuning afterward to converge much faster. For example, Primus was taking 700 epochs or more to fully converge, and with pretraining it took 400.</p>\n<p>All models were trained at patch size 160³, with AdamW (<code>lr=5e-5</code>, <code>wd=1e-4</code>) and mixed-precision training for approximately 400 epochs.</p>\n<p>For losses, we used a combination of Dice Loss, Cross-Entropy Loss, Skeleton Recall Loss, and Surface Dice Loss (a modified version of clDice for 2D manifolds in 3D environments; implementation was for CryoET membranes). <em>(Note <a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a>: I had the opportunity to watch Lorenz Lamm's presentation at the CZII competition workshop, very nice to find a use almost a year later in a totally different task.)</em></p>\n<p>From local evaluation, the best models were  Primus-B V2, ResEnc-L and Primus-B respectively. We will add scores later. The good thing is that the models are very complementary to each other.</p>\n<h2>2nd Stage — Ensemble</h2>\n<p>The idea here was to let a model learn the complementarity between predictions. The model used a 4-channel input: [Image, ResEnc-L binary mask, Primus binary mask, PrimusV2 binary mask].</p>\n<ul>\n<li><strong>ResEnc-L UNet</strong> ensembles initial stage predictions.</li>\n</ul>\n<p>We trained with a random threshold on the fly ranging from 0.1 to 0.7, but used a threshold of 0.3 at inference.</p>\n<h2>3rd Stage — Refinement</h2>\n<p><em>(Note <a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a>: I really liked the discussions about algorithms to fill holes, make the sheets more continuous, and so on. But I'm not clever enough to do this with an algorithm. So I did what I could: train a model to do that.)</em></p>\n<p>We used the same random threshold strategy on the fly. If the threshold is low, we hope the model learns to correct merging sheets, if it is high, we hope the model learns to correct broken sheets.\nThis behavior can be seen in the image below:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2221915%2F6351f874a7cd7b5169eb2f874ebf62e7%2FCaptura%20de%20tela%202026-03-09%20165943.png?generation=1773086517474158&amp;alt=media\" alt=\"\"></p>\n<p>We also added a strong <code>randCoarseDropout</code> to mask the input to help model learn continuity and fix holes.</p>\n<p>Input is [Mask, 2nd stage binary].</p>\n<ul>\n<li>ResEnc-L UNet refinement.</li>\n</ul>\n<h2>4th Stage — Refinement</h2>\n<p>Just another refinement network. Same idea of stage 3, help model learn to correct silly errors,</p>\n<p>Input is [Mask, 3rd stage binary].</p>\n<ul>\n<li>ResEnc-L UNet — refinement.</li>\n</ul>\n<h2>5th Stage — Diffeomorphic Stage (biggest boost)</h2>\n<p>The diffeomorphic network takes the role of shape calibratin and thickness modification. This approach refers Tom’s paper (<strong>Perceptual Contrastive Generative Adversarial Network based on image warping for unsupervised image-to-image translation</strong>) and the <strong>FlowNet2</strong>. Unlike classic way like previous stage which just reproduces a better segmentation prediction, this network predicts the stationary velocity field which can manipulate the input mask previous stage it recieves. Intuitively, it is a vector field that determine how the mesh changes in 3d space, and we decide how many step it should move. </p>\n<h3>Core Designs and Intuition</h3>\n<p>This model is mainly based on Tom’s paper. It performs warping first and then refines the content afterward.</p>\n<p>Diffeomorphic Step:</p>\n<p>Using a lower threshold to generate the hard mask and applying uniform blur are two crucial steps for helping the model learn the SVF logits. Since the OOF mask is already strong, you need to intentionally introduce “errors” so the model can learn from them and be rewarded for correcting them.</p>\n<ul>\n<li>Using a small threshold reveals more suppressed predictions. Some of these predictions contribute to thicker mask regions, which allows the model to identify and fix such issues.</li>\n<li>A hard mask alone cannot provide useful gradients to the model. We also found that warping a blurred mask is very beneficial for improving topology and boosting VOI performance.</li>\n</ul>\n<p>Blurring Example\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4310004%2F6ca13dd6aa7ec55274b64e0193386444%2Ff1.png?generation=1772451933189014&amp;alt=media\" alt=\"\"></p>\n<pre><code># ----------------------------------------\n# 1. Blur the input mask\n# ----------------------------------------\nHard_Mask = Soft_Mask&gt;0.3\n\nBlurred_mask = Gaussian_Blur(Hard_Mask, kernel_size=3, sigma=5)\n# (Produces soft mask M in [0,1])\n\n# ----------------------------------------\n# 2. Predict stationary velocity field\n# ----------------------------------------\n\ngamma = Diffeomorphic_Prediction(Blurred_mask)\n# gamma ∈ R^{3×D×H×W}  (SVF logits)\n\nv = tanh(gamma) * max_v\n# v : Ω → R³\n# bounded stationary velocity field\n\n# ----------------------------------------\n# 3. Scaling &amp; Squaring (Lie exponential)\n# ----------------------------------------\n\n# Initialize small deformation\nphi_0 = v / (2^N)\n\nphi = phi_0\n\nRepeat N times:\n\n    # Compose deformation with itself\n    # (φ ∘ φ)(x) = φ(x + φ(x))\n\n    phi = phi + Warp(phi, phi)\n\n# After N steps:\n# phi ≈ exp(v)\n\n# ----------------------------------------\n# 4. Warp mask with diffeomorphic transform\n# ----------------------------------------\n\nWarped_mask(x) = Blurred_mask(x + phi(x))\n</code></pre>\n<p>Signed Distance Topology Shift Prediction:</p>\n<p>Optical flow–based methods usually suffer from a “folding” issue. This occurs when the model makes abrupt deformations, which significantly damage the topology. Tom struggled with this problem for quite some time. He eventually addressed it by introducing an additional output channel that predicts the “shift” in the SDF space of the warped mask.</p>\n<pre><code># ----------------------------------------\n# Topology-aware correction\n# ----------------------------------------\n\n# Convert to soft signed distance\nSDF = log(Warped_mask + ε) - log(1 - Warped_mask + ε)\n\n# Learn correction field from 4th channel r_t\nt      = sigmoid(r_t)                 # gating map\ndelta  = max_offset * tanh(r_t)       # signed bounded offset\n\nSDF_corrected = SDF + t * delta\n\nFinal_mask = sigmoid(SDF_corrected)\n</code></pre>\n<p>Loss functions for warping</p>\n<ul>\n<li>Minimizing SVF Smoothing to avoid folding by surpressing the magnitutes of velocities</li>\n</ul>\n<pre><code>def svf_smoothness(self, v):\n    dz = (v[:, :, 1:] - v[:, :, :-1]).pow(2).mean()\n    dy = (v[:, :, :, 1:] - v[:, :, :, :-1]).pow(2).pow(1).mean()\n    dx = (v[:, :, :, :, 1:] - v[:, :, :, :, :-1]).pow(2).mean()\n\n    return (dz + dy + dx) / 3.0\n</code></pre>\n<ul>\n<li>Minimizing jacobian log barrier to keep flow active and preventing a compressible elastic material that strongly resists collapsing or flipping</li>\n</ul>\n<pre><code>def jacobian_determinant(flow):\n    \"\"\"\n    flow: (B, 3, D, H, W) displacement field u(x)\n    returns: (B, D, H, W) jacobian determinant of φ(x)=x+u(x)\n    \"\"\"\n    B, C, D, H, W = flow.shape\n    assert C == 3\n\n    # gradients wrt spatial axes (z = depth, y = height, x = width)\n    du_dx = torch.gradient(flow, dim=4)[0]  # width axis\n    du_dy = torch.gradient(flow, dim=3)[0]  # height axis\n    du_dz = torch.gradient(flow, dim=2)[0]  # depth axis\n\n    # components\n    ux_x = du_dx[:,0]; ux_y = du_dy[:,0]; ux_z = du_dz[:,0]\n    uy_x = du_dx[:,1]; uy_y = du_dy[:,1]; uy_z = du_dz[:,1]\n    uz_x = du_dx[:,2]; uz_y = du_dy[:,2]; uz_z = du_dz[:,2]\n\n    # deformation gradient J = I + ∇u\n    j11 = 1 + ux_x; j12 =     ux_y; j13 =     ux_z\n    j21 =     uy_x; j22 = 1 + uy_y; j23 =     uy_z\n    j31 =     uz_x; j32 =     uz_y; j33 = 1 + uz_z\n\n    det = (\n        j11 * (j22 * j33 - j23 * j32)\n        - j12 * (j21 * j33 - j23 * j31)\n        + j13 * (j21 * j32 - j22 * j31)\n    )\n    return det\n\ndef jacobian_log_barrier(flow, eps=1e-6):\n    det = jacobian_determinant(flow)\n    det_clamped = torch.clamp(det, min=eps)\n    loss = -torch.log(det_clamped).mean()\n    return loss\n</code></pre>\n<p>Loss function for SDF topology fixing </p>\n<p>The challenge <a href=\"https://www.kaggle.com/Tom\" target=\"_blank\">@Tom</a> faces is avoiding the forth channel cheating and dominating the prediction, which letting the diffeomorphic step useless. We use three loss functions to solve this issue.</p>\n<ul>\n<li>Sparisty loss: Keep sparse topology correction, discouraged from editing everywhere and only activates correction when necessary</li>\n</ul>\n<pre><code>#suppose t = gating map in topoFix\ndef topo_sparsity(self, t):\n        return t.mean()\n</code></pre>\n<ul>\n<li>Total variation loss: Making topology edits to be region-based, not noisy voxel-wise toggles</li>\n</ul>\n<pre><code>def topo_tv(self, t):\n    dz = (t[:, :, 1:] - t[:, :, :-1]).abs().mean()\n    dy = (t[:, :, :, 1:] - t[:, :, :, :-1]).abs().mean()\n    dx = (t[:, :, :, :, 1:] - t[:, :, :, :, :-1]).abs().mean()\n\n    return (dz + dy + dx) / 3.0\n</code></pre>\n<ul>\n<li>Boundary loss: Encourage edits near zero-level set and only modify topology near the object surface</li>\n</ul>\n<pre><code>def topo_boundary(self, t, sdf):\n    boundary = torch.exp(-sdf.abs())\n    return (t * (1.0 - boundary)).mean()\n</code></pre>\n<h3>Variations of Implementations</h3>\n<p>We have two implementation of Diffeomorphic Network, one is in <a href=\"https://www.kaggle.com/Tom\" target=\"_blank\">@Tom</a> repositery and another is <a href=\"https://www.kaggle.com/Sergio\" target=\"_blank\">@Sergio</a> repositery, they have different augmentation, processing and loss functions. All the network architecture are all identical, using the nnUNet-style residual encoder UNet but outputs with 4 channels. </p>\n<ul>\n<li>Loss differneces</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Version</th>\n<th>CLDice</th>\n<th>SoftSDFLoss</th>\n<th>Skeleton Recall</th>\n<th>Dice + CE</th>\n<th>Surface Dice Loss</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Tom</td>\n<td>o</td>\n<td>o</td>\n<td>o</td>\n<td>o</td>\n<td>x</td>\n</tr>\n<tr>\n<td>Sergio</td>\n<td>x</td>\n<td>x</td>\n<td>o</td>\n<td>o</td>\n<td>o</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>Augmentation differences</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Version</th>\n<th>Online thresholding</th>\n<th>Mask augmentation</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Tom</strong></td>\n<td><code>threshold = 0.3 + torch.randn(1, device=prob_mask_oof.device) * 0.01</code><br><code>threshold = torch.clamp(threshold, 0.1, 0.5)</code></td>\n<td>• Random Affine (3° rotation)<br>• RandCoarseDropout (12 holes, spatial size = 10)</td>\n</tr>\n<tr>\n<td><strong>Sergio</strong></td>\n<td><code>threshold = 0.3 + np.random.uniform(-0.2, 0.5)</code></td>\n<td><code>RandCoarseDropoutdWithRanges(keys=oof_keys, prob=1, shared_holes_range=(10,15), independent_holes_range=(15,15), spatial_size_range=(10,30), fill_value=0.0)</code><br><br><code>RandCoarseDropoutdWithRanges(keys=oof_keys, prob=0.8, shared_holes_range=(20,40), independent_holes_range=(30,60), spatial_size_range=(1,3), fill_value=0.0)</code></td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li><p>Prediciton results: Sergio's vs Tom's\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4310004%2F7d3f271456fe68aa47f7a89350dae37c%2Ff44.png?generation=1772452069836953&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4310004%2F0b7202034b6a6d1bd159c1793043ef74%2Ff2.png?generation=1772451950172111&amp;alt=media\" alt=\"\"></p></li>\n<li><p>Learned components: Sergio's vs Tom's\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4310004%2F2512060ffeeebfed8d8326580e73cdac%2Ff45.png?generation=1772452213966109&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4310004%2Fb250011f987349d85a1191d0c3501027%2Ff3.png?generation=1772452235538826&amp;alt=media\" alt=\"\"></p></li>\n<li><p>SVF results: Sergio's vs Tom's\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4310004%2Fcd5ed636e60ebaac44005da3ea7508a7%2Ffig1.png?generation=1772452012248337&amp;alt=media\" alt=\"\"></p></li>\n</ul>\n<h3>Comparing different diffeomorphic setup</h3>\n<p>Before the deadline, we conducted simulation trials to approximate the public subset (40 samples per subset, 200 trials) and compared the diffeomorphic networks from the repo of <a href=\"https://www.kaggle.com/Tom\" target=\"_blank\">@Tom</a> and <a href=\"https://www.kaggle.com/Sergio\" target=\"_blank\">@Sergio</a>.</p>\n<p>As shown in the plot, there is a clear difference in metric preference between the two approaches. <a href=\"https://www.kaggle.com/Tom\" target=\"_blank\">@Tom</a>’s model outperforms <a href=\"https://www.kaggle.com/Sergio\" target=\"_blank\">@Sergio</a>’s on Surface Dice, while <a href=\"https://www.kaggle.com/Sergio\" target=\"_blank\">@Sergio</a>’s setup achieves better results on Topology and VOI scores. In most cases, <a href=\"https://www.kaggle.com/Sergio\" target=\"_blank\">@Sergio</a>’s configuration achieves the best overall performance.</p>\n<p>However, we also observed that in a few subsets, <a href=\"https://www.kaggle.com/Tom\" target=\"_blank\">@Tom</a>’s overall score is higher, which is consistent with its LB. This suggests that the public subset may resemble those specific subsets where <a href=\"https://www.kaggle.com/Tom\" target=\"_blank\">@Tom</a>’s model performs better, and a small number of samples could have disproportionately boosted the public score, essentially a favorable sampling effect.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4310004%2F142aa44e7586adc8190b77670022c9ad%2F112.png?generation=1772418585149240&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Validation results</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>CV</th>\n<th>LB</th>\n<th>PB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Tom</td>\n<td>0.619</td>\n<td>0.595</td>\n<td>0.612</td>\n</tr>\n<tr>\n<td>Tom + Sergio</td>\n<td>0.621</td>\n<td>0.588</td>\n<td>0.615</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>Hypothesis testing (H0: <a href=\"https://www.kaggle.com/Tom\" target=\"_blank\">@Tom</a> = <a href=\"https://www.kaggle.com/Sergio\" target=\"_blank\">@Sergio</a>, H1: <a href=\"https://www.kaggle.com/Tom\" target=\"_blank\">@Tom</a> &gt; <a href=\"https://www.kaggle.com/Sergio\" target=\"_blank\">@Sergio</a> )</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Metric</th>\n<th>Mean Difference</th>\n<th>T-Statistic</th>\n<th>P-Value</th>\n<th>Reject?</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Surface_dice</td>\n<td>+0.0086</td>\n<td>74.97</td>\n<td>&lt;0.0001</td>\n<td>Y</td>\n</tr>\n<tr>\n<td>Score</td>\n<td>-0.0017</td>\n<td>-7.64</td>\n<td>1</td>\n<td>N</td>\n</tr>\n<tr>\n<td>Topo</td>\n<td>-0.0124</td>\n<td>-20.11</td>\n<td>1</td>\n<td>N</td>\n</tr>\n<tr>\n<td>Voi</td>\n<td>-0.004</td>\n<td>-117.68</td>\n<td>1</td>\n<td>N</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>Based on the hypothesis testing, these two methods may not have a significant difference in topology or VOI score on the public subset. In that case, the method with the higher surface Dice score dominates, resulting in a 0.007 boost on the leaderboard. The other one we hoped improving lb experiences a substantial drop, scoring 0.588, which does bad in surface dice.</li>\n</ul>\n<h2>Post-Processing</h2>\n<table>\n<thead>\n<tr>\n<th>Post Processing</th>\n<th>LB</th>\n<th>PB</th>\n<th>Comments</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>remove cc &lt; 3000</td>\n<td>0.592</td>\n<td>0.614</td>\n<td></td>\n</tr>\n<tr>\n<td>x7 median filter</td>\n<td>0.596</td>\n<td>0.623</td>\n<td></td>\n</tr>\n<tr>\n<td>x6 median filter</td>\n<td>0.597</td>\n<td>0.623</td>\n<td></td>\n</tr>\n<tr>\n<td>x8 median filter</td>\n<td>0.596</td>\n<td>0.624</td>\n<td></td>\n</tr>\n<tr>\n<td>x9 median filter</td>\n<td>0.596</td>\n<td>0.624</td>\n<td></td>\n</tr>\n<tr>\n<td>x10 median filter</td>\n<td>0.597</td>\n<td>0.624</td>\n<td></td>\n</tr>\n<tr>\n<td>binary closing</td>\n<td>0.588</td>\n<td>0.615</td>\n<td></td>\n</tr>\n<tr>\n<td>closing(7)+hole patching+cavity fill</td>\n<td>0.583</td>\n<td>0.604</td>\n<td></td>\n</tr>\n<tr>\n<td>Gaussian smooth + rethreshold + close/open + fill</td>\n<td>0.585</td>\n<td>0.603</td>\n<td></td>\n</tr>\n<tr>\n<td>Remove internal cavities</td>\n<td>0.589</td>\n<td>0.614</td>\n<td></td>\n</tr>\n<tr>\n<td>Erosion / Dilation</td>\n<td>0.592</td>\n<td>0.613</td>\n<td></td>\n</tr>\n<tr>\n<td>Enforced minimum sheet thickness</td>\n<td>0.463</td>\n<td>0.454</td>\n<td>Topo destroyed; Fails small sample</td>\n</tr>\n<tr>\n<td>Z-consistent smoothing</td>\n<td>0.592</td>\n<td>0.615</td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h2>Code</h2>\n<p><a href=\"https://github.com/Sersasj/vesuvius-challenge-10th-solution\" target=\"_blank\">https://github.com/Sersasj/vesuvius-challenge-10th-solution</a></p>\n<h2>References</h2>\n<ul>\n<li>Primus: Enforcing Attention Usage for 3D Medical Image Segmentation — <a href=\"https://openreview.net/forum?id=YWwGmmObri\" target=\"_blank\">https://openreview.net/forum?id=YWwGmmObri</a></li>\n<li>MemBrain v2: An end-to-end tool for the analysis of membranes in cryo-electron tomography (Surface Dice Loss) — <a href=\"https://www.biorxiv.org/content/10.1101/2024.01.05.574336v1.full.pdf\" target=\"_blank\">https://www.biorxiv.org/content/10.1101/2024.01.05.574336v1.full.pdf</a></li>\n<li>Skeleton Recall Loss for Connectivity Conserving and Resource Efficient Segmentation of Thin Tubular Structures — <a href=\"https://www.ecva.net/papers/eccv_2024/papers_ECCV/papers/09904.pdf\" target=\"_blank\">https://www.ecva.net/papers/eccv_2024/papers_ECCV/papers/09904.pdf</a></li>\n<li>dynamic-network-architectures library — <a href=\"https://github.com/MIC-DKFZ/dynamic-network-architectures\" target=\"_blank\">https://github.com/MIC-DKFZ/dynamic-network-architectures</a></li>\n<li>PrimusV2 code (for some reason it's not merged, we got lucky to find this fork) — <a href=\"https://github.com/TaWald/dynamic-network-architectures/blob/main/dynamic_network_architectures/architectures/primus.py\" target=\"_blank\">https://github.com/TaWald/dynamic-network-architectures/blob/main/dynamic_network_architectures/architectures/primus.py</a></li>\n<li><strong>Perceptual Contrastive Generative Adversarial Network based on image warping for unsupervised image-to-image translation(</strong><a href=\"https://www.sciencedirect.com/science/article/abs/pii/S0893608023003684**\" target=\"_blank\">https://www.sciencedirect.com/science/article/abs/pii/S0893608023003684**</a>)**</li>\n<li><strong>FlowNet 2.0: Evolution of Optical Flow Estimation with Deep Networks(</strong><a href=\"https://arxiv.org/abs/1612.01925**\" target=\"_blank\">https://arxiv.org/abs/1612.01925**</a>)**</li>\n</ul>",
      "rawMarkdown": "We thank Kaggle and the Vesuvius Challenge organizers for a fascinating competition.\n\nAlthough majority of training pipeline was implemented outside the original nnUNet framework, most architectures are derived from work produced by MIC-DKFZ. We are grateful to them for advancing the 3D segmentation domain.\n\nBelow is an overview of our approach.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4310004%2Ff83bcfcfbbbbd3d8bad0af47c331cae8%2Ff45.png?generation=1773033543025607&alt=media)\n\n## 1st Stage — Initial Segmentation\n\nWe trained three independent models:\n\n- **ResEnc-L UNet** (4 × TTA): The standard nnUNet-style residual encoder UNet with channels (32, 64, 128, 256, 320, 320) and `n_blocks=(1,3,4,6,6,6)`.\n- **Primus-B**: A transformer based segmentation model from dynamic-network-architectures library of MIC-DKFZ.\n- **Primus-B V2**: An improved variant of Primus-B.\n\nModels were first pretrained on approximately 1,600 annotated images from other scrolls: https://dl.ash2txt.org/datasets/seg-derived-recto-surfaces/. Labels were slightly different but helped fine-tuning afterward to converge much faster. For example, Primus was taking 700 epochs or more to fully converge, and with pretraining it took 400.\n\nAll models were trained at patch size 160³, with AdamW (`lr=5e-5`, `wd=1e-4`) and mixed-precision training for approximately 400 epochs.\n\nFor losses, we used a combination of Dice Loss, Cross-Entropy Loss, Skeleton Recall Loss, and Surface Dice Loss (a modified version of clDice for 2D manifolds in 3D environments; implementation was for CryoET membranes). *(Note @sersasj: I had the opportunity to watch Lorenz Lamm's presentation at the CZII competition workshop, very nice to find a use almost a year later in a totally different task.)*\n\nFrom local evaluation, the best models were  Primus-B V2, ResEnc-L and Primus-B respectively. We will add scores later. The good thing is that the models are very complementary to each other.\n\n## 2nd Stage — Ensemble\n\nThe idea here was to let a model learn the complementarity between predictions. The model used a 4-channel input: [Image, ResEnc-L binary mask, Primus binary mask, PrimusV2 binary mask].\n\n- **ResEnc-L UNet** ensembles initial stage predictions.\n\nWe trained with a random threshold on the fly ranging from 0.1 to 0.7, but used a threshold of 0.3 at inference.\n\n## 3rd Stage — Refinement\n\n*(Note @sersasj: I really liked the discussions about algorithms to fill holes, make the sheets more continuous, and so on. But I'm not clever enough to do this with an algorithm. So I did what I could: train a model to do that.)*\n\nWe used the same random threshold strategy on the fly. If the threshold is low, we hope the model learns to correct merging sheets, if it is high, we hope the model learns to correct broken sheets.\nThis behavior can be seen in the image below:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2221915%2F6351f874a7cd7b5169eb2f874ebf62e7%2FCaptura%20de%20tela%202026-03-09%20165943.png?generation=1773086517474158&alt=media)\n\nWe also added a strong `randCoarseDropout` to mask the input to help model learn continuity and fix holes.\n\nInput is [Mask, 2nd stage binary].\n\n- ResEnc-L UNet refinement.\n\n## 4th Stage — Refinement\n\nJust another refinement network. Same idea of stage 3, help model learn to correct silly errors,\n\nInput is [Mask, 3rd stage binary].\n\n- ResEnc-L UNet — refinement.\n\n## 5th Stage — Diffeomorphic Stage (biggest boost)\n\nThe diffeomorphic network takes the role of shape calibratin and thickness modification. This approach refers Tom’s paper (**Perceptual Contrastive Generative Adversarial Network based on image warping for unsupervised image-to-image translation**) and the **FlowNet2**. Unlike classic way like previous stage which just reproduces a better segmentation prediction, this network predicts the stationary velocity field which can manipulate the input mask previous stage it recieves. Intuitively, it is a vector field that determine how the mesh changes in 3d space, and we decide how many step it should move. \n\n### Core Designs and Intuition\n\nThis model is mainly based on Tom’s paper. It performs warping first and then refines the content afterward.\n\nDiffeomorphic Step:\n\nUsing a lower threshold to generate the hard mask and applying uniform blur are two crucial steps for helping the model learn the SVF logits. Since the OOF mask is already strong, you need to intentionally introduce “errors” so the model can learn from them and be rewarded for correcting them.\n\n- Using a small threshold reveals more suppressed predictions. Some of these predictions contribute to thicker mask regions, which allows the model to identify and fix such issues.\n- A hard mask alone cannot provide useful gradients to the model. We also found that warping a blurred mask is very beneficial for improving topology and boosting VOI performance.\n\nBlurring Example\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4310004%2F6ca13dd6aa7ec55274b64e0193386444%2Ff1.png?generation=1772451933189014&alt=media)\n\n\n```\n# ----------------------------------------\n# 1. Blur the input mask\n# ----------------------------------------\nHard_Mask = Soft_Mask>0.3\n\nBlurred_mask = Gaussian_Blur(Hard_Mask, kernel_size=3, sigma=5)\n# (Produces soft mask M in [0,1])\n\n# ----------------------------------------\n# 2. Predict stationary velocity field\n# ----------------------------------------\n\ngamma = Diffeomorphic_Prediction(Blurred_mask)\n# gamma ∈ R^{3×D×H×W}  (SVF logits)\n\nv = tanh(gamma) * max_v\n# v : Ω → R³\n# bounded stationary velocity field\n\n# ----------------------------------------\n# 3. Scaling & Squaring (Lie exponential)\n# ----------------------------------------\n\n# Initialize small deformation\nphi_0 = v / (2^N)\n\nphi = phi_0\n\nRepeat N times:\n\n    # Compose deformation with itself\n    # (φ ∘ φ)(x) = φ(x + φ(x))\n    \n    phi = phi + Warp(phi, phi)\n\n# After N steps:\n# phi ≈ exp(v)\n\n# ----------------------------------------\n# 4. Warp mask with diffeomorphic transform\n# ----------------------------------------\n\nWarped_mask(x) = Blurred_mask(x + phi(x))\n\n```\n\n\nSigned Distance Topology Shift Prediction:\n\nOptical flow–based methods usually suffer from a “folding” issue. This occurs when the model makes abrupt deformations, which significantly damage the topology. Tom struggled with this problem for quite some time. He eventually addressed it by introducing an additional output channel that predicts the “shift” in the SDF space of the warped mask.\n\n```\n# ----------------------------------------\n# Topology-aware correction\n# ----------------------------------------\n\n# Convert to soft signed distance\nSDF = log(Warped_mask + ε) - log(1 - Warped_mask + ε)\n\n# Learn correction field from 4th channel r_t\nt      = sigmoid(r_t)                 # gating map\ndelta  = max_offset * tanh(r_t)       # signed bounded offset\n\nSDF_corrected = SDF + t * delta\n\nFinal_mask = sigmoid(SDF_corrected)\n```\n\nLoss functions for warping\n\n- Minimizing SVF Smoothing to avoid folding by surpressing the magnitutes of velocities\n\n```python\ndef svf_smoothness(self, v):\n    dz = (v[:, :, 1:] - v[:, :, :-1]).pow(2).mean()\n    dy = (v[:, :, :, 1:] - v[:, :, :, :-1]).pow(2).pow(1).mean()\n    dx = (v[:, :, :, :, 1:] - v[:, :, :, :, :-1]).pow(2).mean()\n\n    return (dz + dy + dx) / 3.0\n```\n\n- Minimizing jacobian log barrier to keep flow active and preventing a compressible elastic material that strongly resists collapsing or flipping\n\n```python\ndef jacobian_determinant(flow):\n    \"\"\"\n    flow: (B, 3, D, H, W) displacement field u(x)\n    returns: (B, D, H, W) jacobian determinant of φ(x)=x+u(x)\n    \"\"\"\n    B, C, D, H, W = flow.shape\n    assert C == 3\n\n    # gradients wrt spatial axes (z = depth, y = height, x = width)\n    du_dx = torch.gradient(flow, dim=4)[0]  # width axis\n    du_dy = torch.gradient(flow, dim=3)[0]  # height axis\n    du_dz = torch.gradient(flow, dim=2)[0]  # depth axis\n\n    # components\n    ux_x = du_dx[:,0]; ux_y = du_dy[:,0]; ux_z = du_dz[:,0]\n    uy_x = du_dx[:,1]; uy_y = du_dy[:,1]; uy_z = du_dz[:,1]\n    uz_x = du_dx[:,2]; uz_y = du_dy[:,2]; uz_z = du_dz[:,2]\n\n    # deformation gradient J = I + ∇u\n    j11 = 1 + ux_x; j12 =     ux_y; j13 =     ux_z\n    j21 =     uy_x; j22 = 1 + uy_y; j23 =     uy_z\n    j31 =     uz_x; j32 =     uz_y; j33 = 1 + uz_z\n\n    det = (\n        j11 * (j22 * j33 - j23 * j32)\n        - j12 * (j21 * j33 - j23 * j31)\n        + j13 * (j21 * j32 - j22 * j31)\n    )\n    return det\n    \ndef jacobian_log_barrier(flow, eps=1e-6):\n    det = jacobian_determinant(flow)\n    det_clamped = torch.clamp(det, min=eps)\n    loss = -torch.log(det_clamped).mean()\n    return loss\n```\n\nLoss function for SDF topology fixing \n\nThe challenge @Tom faces is avoiding the forth channel cheating and dominating the prediction, which letting the diffeomorphic step useless. We use three loss functions to solve this issue.\n\n- Sparisty loss: Keep sparse topology correction, discouraged from editing everywhere and only activates correction when necessary\n\n```python\n#suppose t = gating map in topoFix\ndef topo_sparsity(self, t):\n\t\treturn t.mean()\n```\n\n- Total variation loss: Making topology edits to be region-based, not noisy voxel-wise toggles\n\n```python\ndef topo_tv(self, t):\n    dz = (t[:, :, 1:] - t[:, :, :-1]).abs().mean()\n    dy = (t[:, :, :, 1:] - t[:, :, :, :-1]).abs().mean()\n    dx = (t[:, :, :, :, 1:] - t[:, :, :, :, :-1]).abs().mean()\n\n    return (dz + dy + dx) / 3.0\n```\n\n- Boundary loss: Encourage edits near zero-level set and only modify topology near the object surface\n```python\ndef topo_boundary(self, t, sdf):\n    boundary = torch.exp(-sdf.abs())\n    return (t * (1.0 - boundary)).mean()\n```\n\n\n### Variations of Implementations\n\nWe have two implementation of Diffeomorphic Network, one is in @Tom repositery and another is @Sergio repositery, they have different augmentation, processing and loss functions. All the network architecture are all identical, using the nnUNet-style residual encoder UNet but outputs with 4 channels. \n\n* Loss differneces\n\n| Version | CLDice | SoftSDFLoss | Skeleton Recall | Dice + CE | Surface Dice Loss |\n| --- | --- | --- | --- | --- | --- |\n| Tom | o | o | o | o | x |\n| Sergio | x | x | o | o | o |\n\n\n* Augmentation differences\n\n| Version  | Online thresholding | Mask augmentation |\n|--------|--------------------|------------------|\n| **Tom** | `threshold = 0.3 + torch.randn(1, device=prob_mask_oof.device) * 0.01`<br>`threshold = torch.clamp(threshold, 0.1, 0.5)` | • Random Affine (3° rotation)<br>• RandCoarseDropout (12 holes, spatial size = 10) |\n| **Sergio** | `threshold = 0.3 + np.random.uniform(-0.2, 0.5)` | `RandCoarseDropoutdWithRanges(keys=oof_keys, prob=1, shared_holes_range=(10,15), independent_holes_range=(15,15), spatial_size_range=(10,30), fill_value=0.0)`<br><br>`RandCoarseDropoutdWithRanges(keys=oof_keys, prob=0.8, shared_holes_range=(20,40), independent_holes_range=(30,60), spatial_size_range=(1,3), fill_value=0.0)` |\n\n* Prediciton results: Sergio's vs Tom's\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4310004%2F7d3f271456fe68aa47f7a89350dae37c%2Ff44.png?generation=1772452069836953&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4310004%2F0b7202034b6a6d1bd159c1793043ef74%2Ff2.png?generation=1772451950172111&alt=media)\n\n\n* Learned components: Sergio's vs Tom's\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4310004%2F2512060ffeeebfed8d8326580e73cdac%2Ff45.png?generation=1772452213966109&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4310004%2Fb250011f987349d85a1191d0c3501027%2Ff3.png?generation=1772452235538826&alt=media)\n\n* SVF results: Sergio's vs Tom's\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4310004%2Fcd5ed636e60ebaac44005da3ea7508a7%2Ffig1.png?generation=1772452012248337&alt=media)\n\n\n### Comparing different diffeomorphic setup\n\nBefore the deadline, we conducted simulation trials to approximate the public subset (40 samples per subset, 200 trials) and compared the diffeomorphic networks from the repo of @Tom and @Sergio.\n\nAs shown in the plot, there is a clear difference in metric preference between the two approaches. @Tom’s model outperforms @Sergio’s on Surface Dice, while @Sergio’s setup achieves better results on Topology and VOI scores. In most cases, @Sergio’s configuration achieves the best overall performance.\n\nHowever, we also observed that in a few subsets, @Tom’s overall score is higher, which is consistent with its LB. This suggests that the public subset may resemble those specific subsets where @Tom’s model performs better, and a small number of samples could have disproportionately boosted the public score, essentially a favorable sampling effect.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4310004%2F142aa44e7586adc8190b77670022c9ad%2F112.png?generation=1772418585149240&alt=media)\n\n* Validation results\n\n|  | CV | LB | PB |\n| --- | --- | --- | --- |\n| Tom | 0.619 | 0.595 | 0.612 |\n| Tom + Sergio | 0.621 | 0.588 | 0.615 |\n\n* Hypothesis testing (H0: @Tom = @Sergio, H1: @Tom > @Sergio )\n\n| Metric | Mean Difference | T-Statistic | P-Value | Reject? |\n| --- | --- | --- | --- | --- |\n| Surface_dice | +0.0086 | 74.97 | <0.0001 | Y |\n| Score | -0.0017 | -7.64 | 1 | N |\n| Topo | -0.0124 | -20.11 | 1 | N |\n| Voi | -0.004 | -117.68 | 1 | N |\n\n* Based on the hypothesis testing, these two methods may not have a significant difference in topology or VOI score on the public subset. In that case, the method with the higher surface Dice score dominates, resulting in a 0.007 boost on the leaderboard. The other one we hoped improving lb experiences a substantial drop, scoring 0.588, which does bad in surface dice.\n\n## Post-Processing\n\n| Post Processing | LB | PB | Comments |\n| --- | --- | --- | --- |\n| remove cc < 3000 | 0.592 | 0.614 | |\n| x7 median filter | 0.596 | 0.623 | |\n| x6 median filter | 0.597 | 0.623 | |\n| x8 median filter | 0.596 | 0.624 | |\n| x9 median filter | 0.596 | 0.624 | |\n| x10 median filter | 0.597 | 0.624 | |\n| binary closing | 0.588 | 0.615 | |\n| closing(7)+hole patching+cavity fill | 0.583 | 0.604 | |\n| Gaussian smooth + rethreshold + close/open + fill | 0.585 | 0.603 | |\n| Remove internal cavities | 0.589 | 0.614 | |\n| Erosion / Dilation | 0.592 | 0.613 | |\n| Enforced minimum sheet thickness | 0.463 | 0.454 | Topo destroyed; Fails small sample |\n| Z-consistent smoothing | 0.592 | 0.615 | |\n\n---\n\n## Code \n\nhttps://github.com/Sersasj/vesuvius-challenge-10th-solution\n\n## References\n\n- Primus: Enforcing Attention Usage for 3D Medical Image Segmentation — https://openreview.net/forum?id=YWwGmmObri\n- MemBrain v2: An end-to-end tool for the analysis of membranes in cryo-electron tomography (Surface Dice Loss) — https://www.biorxiv.org/content/10.1101/2024.01.05.574336v1.full.pdf\n- Skeleton Recall Loss for Connectivity Conserving and Resource Efficient Segmentation of Thin Tubular Structures — https://www.ecva.net/papers/eccv_2024/papers_ECCV/papers/09904.pdf\n- dynamic-network-architectures library — https://github.com/MIC-DKFZ/dynamic-network-architectures\n- PrimusV2 code (for some reason it's not merged, we got lucky to find this fork) — https://github.com/TaWald/dynamic-network-architectures/blob/main/dynamic_network_architectures/architectures/primus.py\n- **Perceptual Contrastive Generative Adversarial Network based on image warping for unsupervised image-to-image translation(**https://www.sciencedirect.com/science/article/abs/pii/S0893608023003684**)**\n- **FlowNet 2.0: Evolution of Optical Flow Estimation with Deep Networks(**https://arxiv.org/abs/1612.01925**)**",
      "votes": 27
    },
    {
      "id": 3414971,
      "postDate": "2026-02-28T02:37:54.573Z",
      "content": "<p>This is the second time I have served as the leader of this team. This competition has given me far more than technical experience. It has provided valuable lessons beyond the contest itself. As a team leader, I had to learn how to respond when my teammates felt discouraged after being overtaken or when we could not come up with a solution. I needed to find ways to bring everyone back into the competition and refocus our efforts. I also had to manage my own emotions and stay composed so that we could continue moving forward even when frustration arose.</p>\n<p>I am truly grateful to my teammates for staying with me throughout this journey, never giving up, and continuing to experiment and develop workable solutions. I hope to keep working with this group and strive together for even greater achievements and more success in future competitions.</p>\n<p>After seeing the diffeomorphic network approach applied in this competition, I believe it will likely be adopted by many teams in future related contests. Personally, I hope more researchers will explore this method further and develop even more innovative and distinctive approaches based on it. </p>",
      "rawMarkdown": "This is the second time I have served as the leader of this team. This competition has given me far more than technical experience. It has provided valuable lessons beyond the contest itself. As a team leader, I had to learn how to respond when my teammates felt discouraged after being overtaken or when we could not come up with a solution. I needed to find ways to bring everyone back into the competition and refocus our efforts. I also had to manage my own emotions and stay composed so that we could continue moving forward even when frustration arose.\n\nI am truly grateful to my teammates for staying with me throughout this journey, never giving up, and continuing to experiment and develop workable solutions. I hope to keep working with this group and strive together for even greater achievements and more success in future competitions.\n\nAfter seeing the diffeomorphic network approach applied in this competition, I believe it will likely be adopted by many teams in future related contests. Personally, I hope more researchers will explore this method further and develop even more innovative and distinctive approaches based on it. ",
      "votes": 8,
      "replies": [
        {
          "id": 3414977,
          "postDate": "2026-02-28T02:45:15.037Z",
          "content": "<p>Congratulations Tom and team achieving 10th place Cash Gold finish! Having a team leader is important. Teams can easily become \"everyone working and fighting alone\" and then just blend separate models. With a good leader, all the code can be combined (i.e. ideas from one person can be incorporated into another person's pipeline) to maximum performance and everyone can feel connected and encouraged and work stronger! Congrats team!</p>",
          "rawMarkdown": "Congratulations Tom and team achieving 10th place Cash Gold finish! Having a team leader is important. Teams can easily become \"everyone working and fighting alone\" and then just blend separate models. With a good leader, all the code can be combined (i.e. ideas from one person can be incorporated into another person's pipeline) to maximum performance and everyone can feel connected and encouraged and work stronger! Congrats team!",
          "votes": 2,
          "replies": [
            {
              "id": 3415403,
              "postDate": "2026-02-28T21:11:17.320Z",
              "content": "<p>Always happy to compete with the best! I personally might be ready for a bit of a break and do this march madness comp after a stressful 2 golds with Vesuvius and the Santa comp back to back haha</p>",
              "rawMarkdown": "Always happy to compete with the best! I personally might be ready for a bit of a break and do this march madness comp after a stressful 2 golds with Vesuvius and the Santa comp back to back haha",
              "votes": 1
            }
          ]
        },
        {
          "id": 3415336,
          "postDate": "2026-02-28T17:44:02.997Z",
          "rawMarkdown": "",
          "votes": -2,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 3416467,
      "postDate": "2026-03-02T23:53:08.750Z",
      "content": "<h3>Post Deadline Experiments</h3>\n<p>Since we had not focused much on post-processing, for post deadline experiment, we incorporated the simple iterative median filter used in the <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/writeups/18th-median-filter-x-7-post-processing-is-very-s\" target=\"_blank\">18th-place solution</a> into <a href=\"https://www.kaggle.com/code/tom99763/writeup-exp-with-median-filter-5-stages\" target=\"_blank\">our pipeline</a>. This addition substantially improved our results, yielding a significant boost in the topological score. Our five-stage setup with only <a href=\"https://www.kaggle.com/tom99763\" target=\"_blank\">@tom99763</a> diffeomorphic network benefited greatly from this refinement.</p>\n<table>\n<thead>\n<tr>\n<th>Configuration</th>\n<th>LB</th>\n<th>PB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>5-stage + remove small object (baseline)</td>\n<td>0.592</td>\n<td>0.614</td>\n</tr>\n<tr>\n<td>6-stage + remove small object (baseline)</td>\n<td>0.588</td>\n<td>0.615</td>\n</tr>\n<tr>\n<td>6-stage + 7x median filter + remove small object</td>\n<td>0.593</td>\n<td>0.618</td>\n</tr>\n<tr>\n<td>5-stage + 6x median filter + remove small object</td>\n<td>0.597</td>\n<td>0.623</td>\n</tr>\n<tr>\n<td>5-stage + 7x median filter + remove small object</td>\n<td>0.596</td>\n<td>0.623</td>\n</tr>\n<tr>\n<td>5-stage + 8x median filter + remove small object</td>\n<td>0.596</td>\n<td>0.624</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>CV analysis (clarify: 1-stage = baseline, 2-stage = baseline+median filter)\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4310004%2Fc40896deaeaf40f0747196ced2e11617%2Fimage.png?generation=1772495587219649&amp;alt=media\" alt=\"\"></li>\n</ul>",
      "rawMarkdown": "### Post Deadline Experiments\n\nSince we had not focused much on post-processing, for post deadline experiment, we incorporated the simple iterative median filter used in the [18th-place solution](https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/writeups/18th-median-filter-x-7-post-processing-is-very-s) into [our pipeline](https://www.kaggle.com/code/tom99763/writeup-exp-with-median-filter-5-stages). This addition substantially improved our results, yielding a significant boost in the topological score. Our five-stage setup with only @tom99763 diffeomorphic network benefited greatly from this refinement.\n\n| Configuration | LB | PB |\n| --- | --- | --- |\n| 5-stage + remove small object (baseline) | 0.592 | 0.614 |\n| 6-stage + remove small object (baseline) | 0.588 | 0.615 |\n| 6-stage + 7x median filter + remove small object | 0.593 | 0.618 |\n| 5-stage + 6x median filter + remove small object | 0.597 | 0.623 |\n| 5-stage + 7x median filter + remove small object | 0.596 | 0.623 |\n| 5-stage + 8x median filter + remove small object | 0.596 | 0.624 |\n\n\n* CV analysis (clarify: 1-stage = baseline, 2-stage = baseline+median filter)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4310004%2Fc40896deaeaf40f0747196ced2e11617%2Fimage.png?generation=1772495587219649&alt=media)",
      "votes": 3,
      "replies": [
        {
          "id": 3416493,
          "postDate": "2026-03-03T02:40:52.030Z",
          "content": "<p>Wow, impressive. Post process makes a big effect in this comp. It's interesting how it affects private LB more than public LB. Does it also improve CV as much as private LB? (probably so since private and CV are large datasets and more correlated than public test small data).</p>",
          "rawMarkdown": "Wow, impressive. Post process makes a big effect in this comp. It's interesting how it affects private LB more than public LB. Does it also improve CV as much as private LB? (probably so since private and CV are large datasets and more correlated than public test small data).",
          "votes": 2,
          "replies": [
            {
              "id": 3416504,
              "postDate": "2026-03-03T02:59:09.660Z",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>  The plot I show below it's CV improvement results, about 0.6214-&gt;0.6288(+0.0074) boosts which is really close to pb boost 0.614-&gt;0.624 (+0.01).  Also the public lb obtains consistent boosting: 0.588-&gt;0.596 (+0.008)</p>\n<p>The biggest benefit and interesting thing is we choose our best surface dice model to do that, then  it directly boosts topo score cv without hurting other metric, which is really insane. \nA classic example of how just a few lines of code can take you straight to the top.</p>",
              "rawMarkdown": "@cdeotte  The plot I show below it's CV improvement results, about 0.6214->0.6288(+0.0074) boosts which is really close to pb boost 0.614->0.624 (+0.01).  Also the public lb obtains consistent boosting: 0.588->0.596 (+0.008)\n\nThe biggest benefit and interesting thing is we choose our best surface dice model to do that, then  it directly boosts topo score cv without hurting other metric, which is really insane. \nA classic example of how just a few lines of code can take you straight to the top.",
              "votes": 1
            },
            {
              "id": 3416510,
              "postDate": "2026-03-03T03:42:20.887Z",
              "content": "<p>Median filer seems differentiable after some hacks. I wonder if it used as nonlinear diffusion filter in training, would it be better?</p>",
              "rawMarkdown": "Median filer seems differentiable after some hacks. I wonder if it used as nonlinear diffusion filter in training, would it be better?",
              "votes": 2
            },
            {
              "id": 3416512,
              "postDate": "2026-03-03T03:51:50.767Z",
              "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>  When I examined the approach used in the 18th-place solution, I started thinking about the median filtering operation from a more structural perspective. A 3×3×3 median filter can essentially be interpreted as repeatedly performing a local ranking over a binary mask and selecting the middle-ranked element. In this sense, the output at each voxel is primarily determined by the density (i.e., the number of active neighbors) within the 3×3×3 neighborhood.</p>\n<p>From that viewpoint, the median filter is effectively implementing a deterministic, density-based voting mechanism. This suggests that the core behavior could be approximated by a very lightweight learnable model that aggregates local neighborhood statistics and iteratively selects or refines voxel states. Such a model could mimic the “neighbor selection” dynamics of median filtering while offering greater flexibility, e.g., learning anisotropic weights, adapting to local structures, or conditioning on additional features without being constrained to a fixed rank-selection rule.</p>",
              "rawMarkdown": "@hengck23  When I examined the approach used in the 18th-place solution, I started thinking about the median filtering operation from a more structural perspective. A 3×3×3 median filter can essentially be interpreted as repeatedly performing a local ranking over a binary mask and selecting the middle-ranked element. In this sense, the output at each voxel is primarily determined by the density (i.e., the number of active neighbors) within the 3×3×3 neighborhood.\n\nFrom that viewpoint, the median filter is effectively implementing a deterministic, density-based voting mechanism. This suggests that the core behavior could be approximated by a very lightweight learnable model that aggregates local neighborhood statistics and iteratively selects or refines voxel states. Such a model could mimic the “neighbor selection” dynamics of median filtering while offering greater flexibility, e.g., learning anisotropic weights, adapting to local structures, or conditioning on additional features without being constrained to a fixed rank-selection rule.",
              "votes": 2
            },
            {
              "id": 3416514,
              "postDate": "2026-03-03T03:59:50.010Z",
              "content": "<p>I tried conv diffusion filter to diffuse from seed from cc3d on differentiable voi loss.  Results are mixed</p>",
              "rawMarkdown": "I tried conv diffusion filter to diffuse from seed from cc3d on differentiable voi loss.  Results are mixed",
              "votes": 1
            },
            {
              "id": 3416515,
              "postDate": "2026-03-03T04:06:04.920Z",
              "content": "<p>How about applying the median filter in two sequential steps and then leveraging the intermediate responses to guide a diffusion process? By decomposing the operation this way, we could explicitly capture the neighborhood dynamics that are implicitly propagated across iterations such as effective padding directions, local influence patterns, or propagation vectors.</p>",
              "rawMarkdown": "How about applying the median filter in two sequential steps and then leveraging the intermediate responses to guide a diffusion process? By decomposing the operation this way, we could explicitly capture the neighborhood dynamics that are implicitly propagated across iterations such as effective padding directions, local influence patterns, or propagation vectors.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3415767,
      "postDate": "2026-03-01T09:32:23.433Z",
      "content": "<p>Congrats on the great submission.\nI am looking forward to read also the Diffeomorphic Stage part in more details 👀</p>",
      "rawMarkdown": "Congrats on the great submission.\nI am looking forward to read also the Diffeomorphic Stage part in more details 👀\n",
      "votes": 3,
      "replies": [
        {
          "id": 3415863,
          "postDate": "2026-03-01T13:59:39.383Z",
          "content": "<p><a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> I've added some descriptions in the diffeomorphic part. It might takes several days to complete because this method is complicated and contains a lot of details.</p>",
          "rawMarkdown": "@giorgioangelotti I've added some descriptions in the diffeomorphic part. It might takes several days to complete because this method is complicated and contains a lot of details.",
          "votes": 1,
          "replies": [
            {
              "id": 3416573,
              "postDate": "2026-03-03T08:31:28.383Z",
              "content": "<p>Thank you!\nThis is very cool. And I really like the idea that the vector field is learnt rather than being an heuristics. Did you notice a different in number of connected components after applying this stage (i.e. could it split some mergers)?</p>",
              "rawMarkdown": "Thank you!\nThis is very cool. And I really like the idea that the vector field is learnt rather than being an heuristics. Did you notice a different in number of connected components after applying this stage (i.e. could it split some mergers)?\n",
              "votes": 1
            },
            {
              "id": 3416643,
              "postDate": "2026-03-03T12:24:12.680Z",
              "content": "<p><a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a>  It does not explicitly guarantee merge–split correction. In practice, we handle holes and merging effects through the earlier stacked ResEncoder stages. This aspect is not currently described in the write-up. Empirically, we observe that under the same training setup, different out-of-fold (OOF) splits can lead ResUNet to either learn hole-filling behavior or sheet unmerging behavior.</p>\n<p>The primary goal of the diffeomorphic stage is to fine-tune the shape of the predicted structure. Most of the time, the overall topology is already preserved by the preceding network. However, the predicted sheet may still exhibit incorrect geometry. For example, being slightly too thick or having locally distorted shapes even if it appears visually acceptable and does not suffer from merging artifacts.</p>\n<p>At this stage, diffeomorphic refinement adjusts the geometry to better align with the ground-truth annotation. We observe a significant improvement in surface-based and topological metrics during this refinement phase.</p>\n<p>We are currently experimenting with other team's post processing since we didn't do much on that. I think diffeomorphic + good post processing is a very good combo. Like what I post above, just adding median filter repeated with 7 times, it can make huge boost on cv, lb and pb at the same time.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4310004%2Faf772ca25e8c407790ce2af261401f7c%2Fimage.png?generation=1772540347031710&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "@giorgioangelotti  It does not explicitly guarantee merge–split correction. In practice, we handle holes and merging effects through the earlier stacked ResEncoder stages. This aspect is not currently described in the write-up. Empirically, we observe that under the same training setup, different out-of-fold (OOF) splits can lead ResUNet to either learn hole-filling behavior or sheet unmerging behavior.\n\nThe primary goal of the diffeomorphic stage is to fine-tune the shape of the predicted structure. Most of the time, the overall topology is already preserved by the preceding network. However, the predicted sheet may still exhibit incorrect geometry. For example, being slightly too thick or having locally distorted shapes even if it appears visually acceptable and does not suffer from merging artifacts.\n\nAt this stage, diffeomorphic refinement adjusts the geometry to better align with the ground-truth annotation. We observe a significant improvement in surface-based and topological metrics during this refinement phase.\n\nWe are currently experimenting with other team's post processing since we didn't do much on that. I think diffeomorphic + good post processing is a very good combo. Like what I post above, just adding median filter repeated with 7 times, it can make huge boost on cv, lb and pb at the same time.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4310004%2Faf772ca25e8c407790ce2af261401f7c%2Fimage.png?generation=1772540347031710&alt=media)",
              "votes": 1
            },
            {
              "id": 3416649,
              "postDate": "2026-03-03T12:33:04.567Z",
              "content": "<p>Here we present our preliminary experiments comparing improvements over the baseline nnUNet by incorporating a twice-stacked ResEncoder (RefineNet) and an additional diffeomorphic stage (DeformNet). As shown in the results, introducing the diffeomorphic refinement leads to a significant improvement in Surface Dice.</p>\n<p>While this refinement slightly degrades the VOI score, combining the twice-stacked ResEncoder with the diffeomorphic stage ultimately yields strong topological improvements. Overall, this combination provides a good balance, with notable gains in topology-related metrics.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4310004%2Fdf7a775d33c000325b5ecf4750c8ecc8%2FAOI_d_80ZEOQHIyQ-tJdqBPLMm3Y-Oq-A5GbEnH8sGHhp0n0cB5IQLM4IBwvEVgrzoKT-VNjxBqpu2tySU8H1NjrrPb72gxhgsdm3MGysjNlCLacf3G5ppKsL-_bcfTBDvl5AsjyXAgYAJ0sQYmmq4QBQ5hKbufQSOBMCE-BW0_bdeafgCDTtws1600.png?generation=1772541033292991&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "Here we present our preliminary experiments comparing improvements over the baseline nnUNet by incorporating a twice-stacked ResEncoder (RefineNet) and an additional diffeomorphic stage (DeformNet). As shown in the results, introducing the diffeomorphic refinement leads to a significant improvement in Surface Dice.\n\nWhile this refinement slightly degrades the VOI score, combining the twice-stacked ResEncoder with the diffeomorphic stage ultimately yields strong topological improvements. Overall, this combination provides a good balance, with notable gains in topology-related metrics.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4310004%2Fdf7a775d33c000325b5ecf4750c8ecc8%2FAOI_d_80ZEOQHIyQ-tJdqBPLMm3Y-Oq-A5GbEnH8sGHhp0n0cB5IQLM4IBwvEVgrzoKT-VNjxBqpu2tySU8H1NjrrPb72gxhgsdm3MGysjNlCLacf3G5ppKsL-_bcfTBDvl5AsjyXAgYAJ0sQYmmq4QBQ5hKbufQSOBMCE-BW0_bdeafgCDTtws1600.png?generation=1772541033292991&alt=media)",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3416180,
      "postDate": "2026-03-02T08:26:47.657Z",
      "content": "<p>Congrats on the soultion!</p>",
      "rawMarkdown": "Congrats on the soultion!",
      "votes": 1
    },
    {
      "id": 3415047,
      "postDate": "2026-02-28T05:14:36.583Z",
      "content": "<p><a href=\"https://www.kaggle.com/tom99763\" target=\"_blank\">@tom99763</a> \nin the last week it read about this which you may be interested\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa26bf3fac75ced4087d346952bf54d78%2FSelection_2458.png?generation=1772255662756779&amp;alt=media\" alt=\"\"></p>\n<p>curveformer++: <a href=\"https://arxiv.org/html/2402.06423v2\" target=\"_blank\">https://arxiv.org/html/2402.06423v2</a></p>",
      "rawMarkdown": "@tom99763 \nin the last week it read about this which you may be interested\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa26bf3fac75ced4087d346952bf54d78%2FSelection_2458.png?generation=1772255662756779&alt=media)\n\ncurveformer++: https://arxiv.org/html/2402.06423v2",
      "votes": 1,
      "replies": [
        {
          "id": 3415057,
          "postDate": "2026-02-28T05:47:26.827Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> This is the approach I miss! I naively replace resenc backbone with primus in diffeonet, but result is terrible. Now I see this one and really make sense! </p>",
          "rawMarkdown": "@hengck23 This is the approach I miss! I naively replace resenc backbone with primus in diffeonet, but result is terrible. Now I see this one and really make sense! ",
          "replies": [
            {
              "id": 3415065,
              "postDate": "2026-02-28T05:53:34.477Z",
              "content": "<p>the trick in this paper is that he initializes the query curve with one that guarantees to intersect with the ground truth.\nHence this gives a valid sampling on the curve coords</p>",
              "rawMarkdown": "the trick in this paper is that he initializes the query curve with one that guarantees to intersect with the ground truth.\nHence this gives a valid sampling on the curve coords",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 3414996,
      "postDate": "2026-02-28T03:24:56.073Z",
      "content": "<p>I had assumed your solution would incorporate unlabeled data, but it appears not to, which is somewhat unexpected.\nIt appears that semi-supervised methods are not ultimately required for the competition. Having spent a month writing semi-supervised training code, the final results still fell short of those achieved through supervised training.</p>",
      "rawMarkdown": "I had assumed your solution would incorporate unlabeled data, but it appears not to, which is somewhat unexpected.\nIt appears that semi-supervised methods are not ultimately required for the competition. Having spent a month writing semi-supervised training code, the final results still fell short of those achieved through supervised training.",
      "votes": 1,
      "replies": [
        {
          "id": 3414999,
          "postDate": "2026-02-28T03:33:35.510Z",
          "content": "<p>We did actually experiment with semi-supervised methods. We tried Cross Pseudo Supervision, which seemed to help, but it doubled training time so we eventually dropped it.\nWe also did pseudo-label iterations: training on labeled data, running inference on unlabeled data, pretraining on pseudo-labels, and fine-tuning again, repeating this twice. It helped with convergence in the early stages. However, <a href=\"https://www.kaggle.com/p4rallax\" target=\"_blank\">@p4rallax</a> found additional labeled so we used that to pretrain</p>",
          "rawMarkdown": "We did actually experiment with semi-supervised methods. We tried Cross Pseudo Supervision, which seemed to help, but it doubled training time so we eventually dropped it.\nWe also did pseudo-label iterations: training on labeled data, running inference on unlabeled data, pretraining on pseudo-labels, and fine-tuning again, repeating this twice. It helped with convergence in the early stages. However, @p4rallax found additional labeled so we used that to pretrain",
          "votes": 1,
          "replies": [
            {
              "id": 3415001,
              "postDate": "2026-02-28T03:37:01.330Z",
              "content": "<p>I had thought semi-supervised learning might be the ultimate solution, and I devoted 80% of my time to this direction.🤣</p>",
              "rawMarkdown": "I had thought semi-supervised learning might be the ultimate solution, and I devoted 80% of my time to this direction.🤣"
            },
            {
              "id": 3415010,
              "postDate": "2026-02-28T04:12:15.007Z",
              "content": "<p><a href=\"https://www.kaggle.com/tanyong666\" target=\"_blank\">@tanyong666</a> do you develop a learning strategy for semi-sup? My initial thought is to maximize the entropy of unlabeled area first to let more sheet being discovered, then improving the topology and voi next</p>",
              "rawMarkdown": "@tanyong666 do you develop a learning strategy for semi-sup? My initial thought is to maximize the entropy of unlabeled area first to let more sheet being discovered, then improving the topology and voi next"
            },
            {
              "id": 3415011,
              "postDate": "2026-02-28T04:15:12.327Z",
              "content": "<p>I generated pseudo labels, reconstructed the semi-supervised training method and the corresponding loss function.</p>",
              "rawMarkdown": "I generated pseudo labels, reconstructed the semi-supervised training method and the corresponding loss function."
            },
            {
              "id": 3415014,
              "postDate": "2026-02-28T04:20:42.710Z",
              "content": "<p>I also thought of reconstruction, especially with GAN! Unfortunately we don't have enough time haha.</p>",
              "rawMarkdown": "I also thought of reconstruction, especially with GAN! Unfortunately we don't have enough time haha."
            }
          ]
        },
        {
          "id": 3416583,
          "postDate": "2026-03-03T09:20:45.243Z",
          "content": "<p><a href=\"https://www.kaggle.com/tanyong666\" target=\"_blank\">@tanyong666</a> we would be very interested in reading your write-up, if you don't mind sharing it! We are exploring semi-supervised training for ink detection, and we think it's a promising route!</p>",
          "rawMarkdown": "@tanyong666 we would be very interested in reading your write-up, if you don't mind sharing it! We are exploring semi-supervised training for ink detection, and we think it's a promising route!",
          "replies": [
            {
              "id": 3417293,
              "postDate": "2026-03-05T02:45:23.907Z",
              "content": "<p>Thanks for your interest! My approach was fairly straightforward without complex architectures:\nPseudo-labeling: Kept the best supervised model frozen, generated pseudo-labels for new scroll fragments, adding ~3,866 extra training samples.\nNo teacher model: Didn't train a separate teacher or design elaborate semi-supervised frameworks.\nTraining: Loaded the supervised checkpoint, applied different losses for labeled vs pseudo-labeled data, optimized against the official validation metric.\nDue to the large data volume and some data loss after the competition, details might vary slightly, but this was the core idea. Good luck with your exploration!</p>",
              "rawMarkdown": "Thanks for your interest! My approach was fairly straightforward without complex architectures:\nPseudo-labeling: Kept the best supervised model frozen, generated pseudo-labels for new scroll fragments, adding ~3,866 extra training samples.\nNo teacher model: Didn't train a separate teacher or design elaborate semi-supervised frameworks.\nTraining: Loaded the supervised checkpoint, applied different losses for labeled vs pseudo-labeled data, optimized against the official validation metric.\nDue to the large data volume and some data loss after the competition, details might vary slightly, but this was the core idea. Good luck with your exploration!"
            }
          ]
        }
      ]
    },
    {
      "id": 3415390,
      "postDate": "2026-02-28T20:30:38.967Z",
      "content": "<p>I like your use of multiple stages and inputting multiple models into a new refinement model. This is a common technique in tabular data stacks. It's cool to see it work in deep learning image segmentation problems too.</p>",
      "rawMarkdown": "I like your use of multiple stages and inputting multiple models into a new refinement model. This is a common technique in tabular data stacks. It's cool to see it work in deep learning image segmentation problems too.",
      "votes": 2,
      "replies": [
        {
          "id": 3415406,
          "postDate": "2026-02-28T21:15:18.017Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Our goal is to enable the stacked models to learn how to use OOF predictions as references to correct broken regions. We don't use probability map as the input to next stage to let it not takes reference of full discrmination from previous one. Interestingly, under same setup, we found that the characteristics of the OOF strongly influence the task the model implicitly learns. When provided with thicker OOF predictions, the model tends to learn erosion behavior (resulting in better Surface Dice). In contrast, when trained with thinner OOF predictions, it learns dilation behavior (leading to better topology and VOI metrics).</p>\n<p>So the pipeline is basically:</p>\n<pre><code>Init segmentation =&gt; Stack(iterative hole fix) =&gt; shape fix in the final of the stack\n</code></pre>\n<p>We'll release the details of oof experiments soon.</p>",
          "rawMarkdown": "@cdeotte Our goal is to enable the stacked models to learn how to use OOF predictions as references to correct broken regions. We don't use probability map as the input to next stage to let it not takes reference of full discrmination from previous one. Interestingly, under same setup, we found that the characteristics of the OOF strongly influence the task the model implicitly learns. When provided with thicker OOF predictions, the model tends to learn erosion behavior (resulting in better Surface Dice). In contrast, when trained with thinner OOF predictions, it learns dilation behavior (leading to better topology and VOI metrics).\n\nSo the pipeline is basically:\n\n```\nInit segmentation => Stack(iterative hole fix) => shape fix in the final of the stack\n```\nWe'll release the details of oof experiments soon.\n",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3414971,
      "author_name": "Tom",
      "author_url": "",
      "post_date": "2026-02-28T02:37:54.573000",
      "content": "<p>This is the second time I have served as the leader of this team. This competition has given me far more than technical experience. It has provided valuable lessons beyond the contest itself. As a team leader, I had to learn how to respond when my teammates felt discouraged after being overtaken or when we could not come up with a solution. I needed to find ways to bring everyone back into the competition and refocus our efforts. I also had to manage my own emotions and stay composed so that we could continue moving forward even when frustration arose.</p>\n<p>I am truly grateful to my teammates for staying with me throughout this journey, never giving up, and continuing to experiment and develop workable solutions. I hope to keep working with this group and strive together for even greater achievements and more success in future competitions.</p>\n<p>After seeing the diffeomorphic network approach applied in this competition, I believe it will likely be adopted by many teams in future related contests. Personally, I hope more researchers will explore this method further and develop even more innovative and distinctive approaches based on it. </p>",
      "votes": 8,
      "replies": [
        {
          "id": 3414977,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2026-02-28T02:45:15.037000",
          "content": "<p>Congratulations Tom and team achieving 10th place Cash Gold finish! Having a team leader is important. Teams can easily become \"everyone working and fighting alone\" and then just blend separate models. With a good leader, all the code can be combined (i.e. ideas from one person can be incorporated into another person's pipeline) to maximum performance and everyone can feel connected and encouraged and work stronger! Congrats team!</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3415403,
              "author_name": "Cody_Null",
              "author_url": "",
              "post_date": "2026-02-28T21:11:17.320000",
              "content": "<p>Always happy to compete with the best! I personally might be ready for a bit of a break and do this march madness comp after a stressful 2 golds with Vesuvius and the Santa comp back to back haha</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 3415336,
          "author_name": "",
          "author_url": "",
          "post_date": "2026-02-28T17:44:02.997000",
          "content": "",
          "votes": -2,
          "replies": []
        }
      ]
    },
    {
      "id": 3416467,
      "author_name": "Tom",
      "author_url": "",
      "post_date": "2026-03-02T23:53:08.750000",
      "content": "<h3>Post Deadline Experiments</h3>\n<p>Since we had not focused much on post-processing, for post deadline experiment, we incorporated the simple iterative median filter used in the <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/writeups/18th-median-filter-x-7-post-processing-is-very-s\" target=\"_blank\">18th-place solution</a> into <a href=\"https://www.kaggle.com/code/tom99763/writeup-exp-with-median-filter-5-stages\" target=\"_blank\">our pipeline</a>. This addition substantially improved our results, yielding a significant boost in the topological score. Our five-stage setup with only <a href=\"https://www.kaggle.com/tom99763\" target=\"_blank\">@tom99763</a> diffeomorphic network benefited greatly from this refinement.</p>\n<table>\n<thead>\n<tr>\n<th>Configuration</th>\n<th>LB</th>\n<th>PB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>5-stage + remove small object (baseline)</td>\n<td>0.592</td>\n<td>0.614</td>\n</tr>\n<tr>\n<td>6-stage + remove small object (baseline)</td>\n<td>0.588</td>\n<td>0.615</td>\n</tr>\n<tr>\n<td>6-stage + 7x median filter + remove small object</td>\n<td>0.593</td>\n<td>0.618</td>\n</tr>\n<tr>\n<td>5-stage + 6x median filter + remove small object</td>\n<td>0.597</td>\n<td>0.623</td>\n</tr>\n<tr>\n<td>5-stage + 7x median filter + remove small object</td>\n<td>0.596</td>\n<td>0.623</td>\n</tr>\n<tr>\n<td>5-stage + 8x median filter + remove small object</td>\n<td>0.596</td>\n<td>0.624</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>CV analysis (clarify: 1-stage = baseline, 2-stage = baseline+median filter)\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4310004%2Fc40896deaeaf40f0747196ced2e11617%2Fimage.png?generation=1772495587219649&amp;alt=media\" alt=\"\"></li>\n</ul>",
      "votes": 3,
      "replies": [
        {
          "id": 3416493,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2026-03-03T02:40:52.030000",
          "content": "<p>Wow, impressive. Post process makes a big effect in this comp. It's interesting how it affects private LB more than public LB. Does it also improve CV as much as private LB? (probably so since private and CV are large datasets and more correlated than public test small data).</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3416504,
              "author_name": "Tom",
              "author_url": "",
              "post_date": "2026-03-03T02:59:09.660000",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>  The plot I show below it's CV improvement results, about 0.6214-&gt;0.6288(+0.0074) boosts which is really close to pb boost 0.614-&gt;0.624 (+0.01).  Also the public lb obtains consistent boosting: 0.588-&gt;0.596 (+0.008)</p>\n<p>The biggest benefit and interesting thing is we choose our best surface dice model to do that, then  it directly boosts topo score cv without hurting other metric, which is really insane. \nA classic example of how just a few lines of code can take you straight to the top.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3416510,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2026-03-03T03:42:20.887000",
              "content": "<p>Median filer seems differentiable after some hacks. I wonder if it used as nonlinear diffusion filter in training, would it be better?</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3416512,
              "author_name": "Tom",
              "author_url": "",
              "post_date": "2026-03-03T03:51:50.767000",
              "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>  When I examined the approach used in the 18th-place solution, I started thinking about the median filtering operation from a more structural perspective. A 3×3×3 median filter can essentially be interpreted as repeatedly performing a local ranking over a binary mask and selecting the middle-ranked element. In this sense, the output at each voxel is primarily determined by the density (i.e., the number of active neighbors) within the 3×3×3 neighborhood.</p>\n<p>From that viewpoint, the median filter is effectively implementing a deterministic, density-based voting mechanism. This suggests that the core behavior could be approximated by a very lightweight learnable model that aggregates local neighborhood statistics and iteratively selects or refines voxel states. Such a model could mimic the “neighbor selection” dynamics of median filtering while offering greater flexibility, e.g., learning anisotropic weights, adapting to local structures, or conditioning on additional features without being constrained to a fixed rank-selection rule.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3416514,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2026-03-03T03:59:50.010000",
              "content": "<p>I tried conv diffusion filter to diffuse from seed from cc3d on differentiable voi loss.  Results are mixed</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3416515,
              "author_name": "Tom",
              "author_url": "",
              "post_date": "2026-03-03T04:06:04.920000",
              "content": "<p>How about applying the median filter in two sequential steps and then leveraging the intermediate responses to guide a diffusion process? By decomposing the operation this way, we could explicitly capture the neighborhood dynamics that are implicitly propagated across iterations such as effective padding directions, local influence patterns, or propagation vectors.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3415767,
      "author_name": "Giorgio Angelotti",
      "author_url": "",
      "post_date": "2026-03-01T09:32:23.433000",
      "content": "<p>Congrats on the great submission.\nI am looking forward to read also the Diffeomorphic Stage part in more details 👀</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3415863,
          "author_name": "Tom",
          "author_url": "",
          "post_date": "2026-03-01T13:59:39.383000",
          "content": "<p><a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> I've added some descriptions in the diffeomorphic part. It might takes several days to complete because this method is complicated and contains a lot of details.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3416573,
              "author_name": "Giorgio Angelotti",
              "author_url": "",
              "post_date": "2026-03-03T08:31:28.383000",
              "content": "<p>Thank you!\nThis is very cool. And I really like the idea that the vector field is learnt rather than being an heuristics. Did you notice a different in number of connected components after applying this stage (i.e. could it split some mergers)?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3416643,
              "author_name": "Tom",
              "author_url": "",
              "post_date": "2026-03-03T12:24:12.680000",
              "content": "<p><a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a>  It does not explicitly guarantee merge–split correction. In practice, we handle holes and merging effects through the earlier stacked ResEncoder stages. This aspect is not currently described in the write-up. Empirically, we observe that under the same training setup, different out-of-fold (OOF) splits can lead ResUNet to either learn hole-filling behavior or sheet unmerging behavior.</p>\n<p>The primary goal of the diffeomorphic stage is to fine-tune the shape of the predicted structure. Most of the time, the overall topology is already preserved by the preceding network. However, the predicted sheet may still exhibit incorrect geometry. For example, being slightly too thick or having locally distorted shapes even if it appears visually acceptable and does not suffer from merging artifacts.</p>\n<p>At this stage, diffeomorphic refinement adjusts the geometry to better align with the ground-truth annotation. We observe a significant improvement in surface-based and topological metrics during this refinement phase.</p>\n<p>We are currently experimenting with other team's post processing since we didn't do much on that. I think diffeomorphic + good post processing is a very good combo. Like what I post above, just adding median filter repeated with 7 times, it can make huge boost on cv, lb and pb at the same time.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4310004%2Faf772ca25e8c407790ce2af261401f7c%2Fimage.png?generation=1772540347031710&amp;alt=media\" alt=\"\"></p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3416649,
              "author_name": "Tom",
              "author_url": "",
              "post_date": "2026-03-03T12:33:04.567000",
              "content": "<p>Here we present our preliminary experiments comparing improvements over the baseline nnUNet by incorporating a twice-stacked ResEncoder (RefineNet) and an additional diffeomorphic stage (DeformNet). As shown in the results, introducing the diffeomorphic refinement leads to a significant improvement in Surface Dice.</p>\n<p>While this refinement slightly degrades the VOI score, combining the twice-stacked ResEncoder with the diffeomorphic stage ultimately yields strong topological improvements. Overall, this combination provides a good balance, with notable gains in topology-related metrics.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4310004%2Fdf7a775d33c000325b5ecf4750c8ecc8%2FAOI_d_80ZEOQHIyQ-tJdqBPLMm3Y-Oq-A5GbEnH8sGHhp0n0cB5IQLM4IBwvEVgrzoKT-VNjxBqpu2tySU8H1NjrrPb72gxhgsdm3MGysjNlCLacf3G5ppKsL-_bcfTBDvl5AsjyXAgYAJ0sQYmmq4QBQ5hKbufQSOBMCE-BW0_bdeafgCDTtws1600.png?generation=1772541033292991&amp;alt=media\" alt=\"\"></p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3416180,
      "author_name": "Navneet",
      "author_url": "",
      "post_date": "2026-03-02T08:26:47.657000",
      "content": "<p>Congrats on the soultion!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3415047,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2026-02-28T05:14:36.583000",
      "content": "<p><a href=\"https://www.kaggle.com/tom99763\" target=\"_blank\">@tom99763</a> \nin the last week it read about this which you may be interested\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa26bf3fac75ced4087d346952bf54d78%2FSelection_2458.png?generation=1772255662756779&amp;alt=media\" alt=\"\"></p>\n<p>curveformer++: <a href=\"https://arxiv.org/html/2402.06423v2\" target=\"_blank\">https://arxiv.org/html/2402.06423v2</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 3415057,
          "author_name": "Tom",
          "author_url": "",
          "post_date": "2026-02-28T05:47:26.827000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> This is the approach I miss! I naively replace resenc backbone with primus in diffeonet, but result is terrible. Now I see this one and really make sense! </p>",
          "votes": 0,
          "replies": [
            {
              "id": 3415065,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2026-02-28T05:53:34.477000",
              "content": "<p>the trick in this paper is that he initializes the query curve with one that guarantees to intersect with the ground truth.\nHence this gives a valid sampling on the curve coords</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3414996,
      "author_name": "huoxu",
      "author_url": "",
      "post_date": "2026-02-28T03:24:56.073000",
      "content": "<p>I had assumed your solution would incorporate unlabeled data, but it appears not to, which is somewhat unexpected.\nIt appears that semi-supervised methods are not ultimately required for the competition. Having spent a month writing semi-supervised training code, the final results still fell short of those achieved through supervised training.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3414999,
          "author_name": "Sergio Alvarez",
          "author_url": "",
          "post_date": "2026-02-28T03:33:35.510000",
          "content": "<p>We did actually experiment with semi-supervised methods. We tried Cross Pseudo Supervision, which seemed to help, but it doubled training time so we eventually dropped it.\nWe also did pseudo-label iterations: training on labeled data, running inference on unlabeled data, pretraining on pseudo-labels, and fine-tuning again, repeating this twice. It helped with convergence in the early stages. However, <a href=\"https://www.kaggle.com/p4rallax\" target=\"_blank\">@p4rallax</a> found additional labeled so we used that to pretrain</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3415001,
              "author_name": "huoxu",
              "author_url": "",
              "post_date": "2026-02-28T03:37:01.330000",
              "content": "<p>I had thought semi-supervised learning might be the ultimate solution, and I devoted 80% of my time to this direction.🤣</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3415010,
              "author_name": "Tom",
              "author_url": "",
              "post_date": "2026-02-28T04:12:15.007000",
              "content": "<p><a href=\"https://www.kaggle.com/tanyong666\" target=\"_blank\">@tanyong666</a> do you develop a learning strategy for semi-sup? My initial thought is to maximize the entropy of unlabeled area first to let more sheet being discovered, then improving the topology and voi next</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3415011,
              "author_name": "huoxu",
              "author_url": "",
              "post_date": "2026-02-28T04:15:12.327000",
              "content": "<p>I generated pseudo labels, reconstructed the semi-supervised training method and the corresponding loss function.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3415014,
              "author_name": "Tom",
              "author_url": "",
              "post_date": "2026-02-28T04:20:42.710000",
              "content": "<p>I also thought of reconstruction, especially with GAN! Unfortunately we don't have enough time haha.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3416583,
          "author_name": "Giorgio Angelotti",
          "author_url": "",
          "post_date": "2026-03-03T09:20:45.243000",
          "content": "<p><a href=\"https://www.kaggle.com/tanyong666\" target=\"_blank\">@tanyong666</a> we would be very interested in reading your write-up, if you don't mind sharing it! We are exploring semi-supervised training for ink detection, and we think it's a promising route!</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3417293,
              "author_name": "huoxu",
              "author_url": "",
              "post_date": "2026-03-05T02:45:23.907000",
              "content": "<p>Thanks for your interest! My approach was fairly straightforward without complex architectures:\nPseudo-labeling: Kept the best supervised model frozen, generated pseudo-labels for new scroll fragments, adding ~3,866 extra training samples.\nNo teacher model: Didn't train a separate teacher or design elaborate semi-supervised frameworks.\nTraining: Loaded the supervised checkpoint, applied different losses for labeled vs pseudo-labeled data, optimized against the official validation metric.\nDue to the large data volume and some data loss after the competition, details might vary slightly, but this was the core idea. Good luck with your exploration!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3415390,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2026-02-28T20:30:38.967000",
      "content": "<p>I like your use of multiple stages and inputting multiple models into a new refinement model. This is a common technique in tabular data stacks. It's cool to see it work in deep learning image segmentation problems too.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3415406,
          "author_name": "Tom",
          "author_url": "",
          "post_date": "2026-02-28T21:15:18.017000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Our goal is to enable the stacked models to learn how to use OOF predictions as references to correct broken regions. We don't use probability map as the input to next stage to let it not takes reference of full discrmination from previous one. Interestingly, under same setup, we found that the characteristics of the OOF strongly influence the task the model implicitly learns. When provided with thicker OOF predictions, the model tends to learn erosion behavior (resulting in better Surface Dice). In contrast, when trained with thinner OOF predictions, it learns dilation behavior (leading to better topology and VOI metrics).</p>\n<p>So the pipeline is basically:</p>\n<pre><code>Init segmentation =&gt; Stack(iterative hole fix) =&gt; shape fix in the final of the stack\n</code></pre>\n<p>We'll release the details of oof experiments soon.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3414960": "We thank Kaggle and the Vesuvius Challenge organizers for a fascinating competition.\n\nAlthough majority of training pipeline was implemented outside the original nnUNet framework, most architectures are derived from work produced by MIC-DKFZ. We are grateful to them for advancing the 3D segmentation domain.\n\nBelow is an overview of our approach.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4310004%2Ff83bcfcfbbbbd3d8bad0af47c331cae8%2Ff45.png?generation=1773033543025607&alt=media)\n\n## 1st Stage — Initial Segmentation\n\nWe trained three independent models:\n\n- **ResEnc-L UNet** (4 × TTA): The standard nnUNet-style residual encoder UNet with channels (32, 64, 128, 256, 320, 320) and `n_blocks=(1,3,4,6,6,6)`.\n- **Primus-B**: A transformer based segmentation model from dynamic-network-architectures library of MIC-DKFZ.\n- **Primus-B V2**: An improved variant of Primus-B.\n\nModels were first pretrained on approximately 1,600 annotated images from other scrolls: https://dl.ash2txt.org/datasets/seg-derived-recto-surfaces/. Labels were slightly different but helped fine-tuning afterward to converge much faster. For example, Primus was taking 700 epochs or more to fully converge, and with pretraining it took 400.\n\nAll models were trained at patch size 160³, with AdamW (`lr=5e-5`, `wd=1e-4`) and mixed-precision training for approximately 400 epochs.\n\nFor losses, we used a combination of Dice Loss, Cross-Entropy Loss, Skeleton Recall Loss, and Surface Dice Loss (a modified version of clDice for 2D manifolds in 3D environments; implementation was for CryoET membranes). *(Note @sersasj: I had the opportunity to watch Lorenz Lamm's presentation at the CZII competition workshop, very nice to find a use almost a year later in a totally different task.)*\n\nFrom local evaluation, the best models were  Primus-B V2, ResEnc-L and Primus-B respectively. We will add scores later. The good thing is that the models are very complementary to each other.\n\n## 2nd Stage — Ensemble\n\nThe idea here was to let a model learn the complementarity between predictions. The model used a 4-channel input: [Image, ResEnc-L binary mask, Primus binary mask, PrimusV2 binary mask].\n\n- **ResEnc-L UNet** ensembles initial stage predictions.\n\nWe trained with a random threshold on the fly ranging from 0.1 to 0.7, but used a threshold of 0.3 at inference.\n\n## 3rd Stage — Refinement\n\n*(Note @sersasj: I really liked the discussions about algorithms to fill holes, make the sheets more continuous, and so on. But I'm not clever enough to do this with an algorithm. So I did what I could: train a model to do that.)*\n\nWe used the same random threshold strategy on the fly. If the threshold is low, we hope the model learns to correct merging sheets, if it is high, we hope the model learns to correct broken sheets.\nThis behavior can be seen in the image below:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2221915%2F6351f874a7cd7b5169eb2f874ebf62e7%2FCaptura%20de%20tela%202026-03-09%20165943.png?generation=1773086517474158&alt=media)\n\nWe also added a strong `randCoarseDropout` to mask the input to help model learn continuity and fix holes.\n\nInput is [Mask, 2nd stage binary].\n\n- ResEnc-L UNet refinement.\n\n## 4th Stage — Refinement\n\nJust another refinement network. Same idea of stage 3, help model learn to correct silly errors,\n\nInput is [Mask, 3rd stage binary].\n\n- ResEnc-L UNet — refinement.\n\n## 5th Stage — Diffeomorphic Stage (biggest boost)\n\nThe diffeomorphic network takes the role of shape calibratin and thickness modification. This approach refers Tom’s paper (**Perceptual Contrastive Generative Adversarial Network based on image warping for unsupervised image-to-image translation**) and the **FlowNet2**. Unlike classic way like previous stage which just reproduces a better segmentation prediction, this network predicts the stationary velocity field which can manipulate the input mask previous stage it recieves. Intuitively, it is a vector field that determine how the mesh changes in 3d space, and we decide how many step it should move. \n\n### Core Designs and Intuition\n\nThis model is mainly based on Tom’s paper. It performs warping first and then refines the content afterward.\n\nDiffeomorphic Step:\n\nUsing a lower threshold to generate the hard mask and applying uniform blur are two crucial steps for helping the model learn the SVF logits. Since the OOF mask is already strong, you need to intentionally introduce “errors” so the model can learn from them and be rewarded for correcting them.\n\n- Using a small threshold reveals more suppressed predictions. Some of these predictions contribute to thicker mask regions, which allows the model to identify and fix such issues.\n- A hard mask alone cannot provide useful gradients to the model. We also found that warping a blurred mask is very beneficial for improving topology and boosting VOI performance.\n\nBlurring Example\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4310004%2F6ca13dd6aa7ec55274b64e0193386444%2Ff1.png?generation=1772451933189014&alt=media)\n\n\n```\n# ----------------------------------------\n# 1. Blur the input mask\n# ----------------------------------------\nHard_Mask = Soft_Mask>0.3\n\nBlurred_mask = Gaussian_Blur(Hard_Mask, kernel_size=3, sigma=5)\n# (Produces soft mask M in [0,1])\n\n# ----------------------------------------\n# 2. Predict stationary velocity field\n# ----------------------------------------\n\ngamma = Diffeomorphic_Prediction(Blurred_mask)\n# gamma ∈ R^{3×D×H×W}  (SVF logits)\n\nv = tanh(gamma) * max_v\n# v : Ω → R³\n# bounded stationary velocity field\n\n# ----------------------------------------\n# 3. Scaling & Squaring (Lie exponential)\n# ----------------------------------------\n\n# Initialize small deformation\nphi_0 = v / (2^N)\n\nphi = phi_0\n\nRepeat N times:\n\n    # Compose deformation with itself\n    # (φ ∘ φ)(x) = φ(x + φ(x))\n    \n    phi = phi + Warp(phi, phi)\n\n# After N steps:\n# phi ≈ exp(v)\n\n# ----------------------------------------\n# 4. Warp mask with diffeomorphic transform\n# ----------------------------------------\n\nWarped_mask(x) = Blurred_mask(x + phi(x))\n\n```\n\n\nSigned Distance Topology Shift Prediction:\n\nOptical flow–based methods usually suffer from a “folding” issue. This occurs when the model makes abrupt deformations, which significantly damage the topology. Tom struggled with this problem for quite some time. He eventually addressed it by introducing an additional output channel that predicts the “shift” in the SDF space of the warped mask.\n\n```\n# ----------------------------------------\n# Topology-aware correction\n# ----------------------------------------\n\n# Convert to soft signed distance\nSDF = log(Warped_mask + ε) - log(1 - Warped_mask + ε)\n\n# Learn correction field from 4th channel r_t\nt      = sigmoid(r_t)                 # gating map\ndelta  = max_offset * tanh(r_t)       # signed bounded offset\n\nSDF_corrected = SDF + t * delta\n\nFinal_mask = sigmoid(SDF_corrected)\n```\n\nLoss functions for warping\n\n- Minimizing SVF Smoothing to avoid folding by surpressing the magnitutes of velocities\n\n```python\ndef svf_smoothness(self, v):\n    dz = (v[:, :, 1:] - v[:, :, :-1]).pow(2).mean()\n    dy = (v[:, :, :, 1:] - v[:, :, :, :-1]).pow(2).pow(1).mean()\n    dx = (v[:, :, :, :, 1:] - v[:, :, :, :, :-1]).pow(2).mean()\n\n    return (dz + dy + dx) / 3.0\n```\n\n- Minimizing jacobian log barrier to keep flow active and preventing a compressible elastic material that strongly resists collapsing or flipping\n\n```python\ndef jacobian_determinant(flow):\n    \"\"\"\n    flow: (B, 3, D, H, W) displacement field u(x)\n    returns: (B, D, H, W) jacobian determinant of φ(x)=x+u(x)\n    \"\"\"\n    B, C, D, H, W = flow.shape\n    assert C == 3\n\n    # gradients wrt spatial axes (z = depth, y = height, x = width)\n    du_dx = torch.gradient(flow, dim=4)[0]  # width axis\n    du_dy = torch.gradient(flow, dim=3)[0]  # height axis\n    du_dz = torch.gradient(flow, dim=2)[0]  # depth axis\n\n    # components\n    ux_x = du_dx[:,0]; ux_y = du_dy[:,0]; ux_z = du_dz[:,0]\n    uy_x = du_dx[:,1]; uy_y = du_dy[:,1]; uy_z = du_dz[:,1]\n    uz_x = du_dx[:,2]; uz_y = du_dy[:,2]; uz_z = du_dz[:,2]\n\n    # deformation gradient J = I + ∇u\n    j11 = 1 + ux_x; j12 =     ux_y; j13 =     ux_z\n    j21 =     uy_x; j22 = 1 + uy_y; j23 =     uy_z\n    j31 =     uz_x; j32 =     uz_y; j33 = 1 + uz_z\n\n    det = (\n        j11 * (j22 * j33 - j23 * j32)\n        - j12 * (j21 * j33 - j23 * j31)\n        + j13 * (j21 * j32 - j22 * j31)\n    )\n    return det\n    \ndef jacobian_log_barrier(flow, eps=1e-6):\n    det = jacobian_determinant(flow)\n    det_clamped = torch.clamp(det, min=eps)\n    loss = -torch.log(det_clamped).mean()\n    return loss\n```\n\nLoss function for SDF topology fixing \n\nThe challenge @Tom faces is avoiding the forth channel cheating and dominating the prediction, which letting the diffeomorphic step useless. We use three loss functions to solve this issue.\n\n- Sparisty loss: Keep sparse topology correction, discouraged from editing everywhere and only activates correction when necessary\n\n```python\n#suppose t = gating map in topoFix\ndef topo_sparsity(self, t):\n\t\treturn t.mean()\n```\n\n- Total variation loss: Making topology edits to be region-based, not noisy voxel-wise toggles\n\n```python\ndef topo_tv(self, t):\n    dz = (t[:, :, 1:] - t[:, :, :-1]).abs().mean()\n    dy = (t[:, :, :, 1:] - t[:, :, :, :-1]).abs().mean()\n    dx = (t[:, :, :, :, 1:] - t[:, :, :, :, :-1]).abs().mean()\n\n    return (dz + dy + dx) / 3.0\n```\n\n- Boundary loss: Encourage edits near zero-level set and only modify topology near the object surface\n```python\ndef topo_boundary(self, t, sdf):\n    boundary = torch.exp(-sdf.abs())\n    return (t * (1.0 - boundary)).mean()\n```\n\n\n### Variations of Implementations\n\nWe have two implementation of Diffeomorphic Network, one is in @Tom repositery and another is @Sergio repositery, they have different augmentation, processing and loss functions. All the network architecture are all identical, using the nnUNet-style residual encoder UNet but outputs with 4 channels. \n\n* Loss differneces\n\n| Version | CLDice | SoftSDFLoss | Skeleton Recall | Dice + CE | Surface Dice Loss |\n| --- | --- | --- | --- | --- | --- |\n| Tom | o | o | o | o | x |\n| Sergio | x | x | o | o | o |\n\n\n* Augmentation differences\n\n| Version  | Online thresholding | Mask augmentation |\n|--------|--------------------|------------------|\n| **Tom** | `threshold = 0.3 + torch.randn(1, device=prob_mask_oof.device) * 0.01`<br>`threshold = torch.clamp(threshold, 0.1, 0.5)` | • Random Affine (3° rotation)<br>• RandCoarseDropout (12 holes, spatial size = 10) |\n| **Sergio** | `threshold = 0.3 + np.random.uniform(-0.2, 0.5)` | `RandCoarseDropoutdWithRanges(keys=oof_keys, prob=1, shared_holes_range=(10,15), independent_holes_range=(15,15), spatial_size_range=(10,30), fill_value=0.0)`<br><br>`RandCoarseDropoutdWithRanges(keys=oof_keys, prob=0.8, shared_holes_range=(20,40), independent_holes_range=(30,60), spatial_size_range=(1,3), fill_value=0.0)` |\n\n* Prediciton results: Sergio's vs Tom's\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4310004%2F7d3f271456fe68aa47f7a89350dae37c%2Ff44.png?generation=1772452069836953&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4310004%2F0b7202034b6a6d1bd159c1793043ef74%2Ff2.png?generation=1772451950172111&alt=media)\n\n\n* Learned components: Sergio's vs Tom's\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4310004%2F2512060ffeeebfed8d8326580e73cdac%2Ff45.png?generation=1772452213966109&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4310004%2Fb250011f987349d85a1191d0c3501027%2Ff3.png?generation=1772452235538826&alt=media)\n\n* SVF results: Sergio's vs Tom's\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4310004%2Fcd5ed636e60ebaac44005da3ea7508a7%2Ffig1.png?generation=1772452012248337&alt=media)\n\n\n### Comparing different diffeomorphic setup\n\nBefore the deadline, we conducted simulation trials to approximate the public subset (40 samples per subset, 200 trials) and compared the diffeomorphic networks from the repo of @Tom and @Sergio.\n\nAs shown in the plot, there is a clear difference in metric preference between the two approaches. @Tom’s model outperforms @Sergio’s on Surface Dice, while @Sergio’s setup achieves better results on Topology and VOI scores. In most cases, @Sergio’s configuration achieves the best overall performance.\n\nHowever, we also observed that in a few subsets, @Tom’s overall score is higher, which is consistent with its LB. This suggests that the public subset may resemble those specific subsets where @Tom’s model performs better, and a small number of samples could have disproportionately boosted the public score, essentially a favorable sampling effect.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4310004%2F142aa44e7586adc8190b77670022c9ad%2F112.png?generation=1772418585149240&alt=media)\n\n* Validation results\n\n|  | CV | LB | PB |\n| --- | --- | --- | --- |\n| Tom | 0.619 | 0.595 | 0.612 |\n| Tom + Sergio | 0.621 | 0.588 | 0.615 |\n\n* Hypothesis testing (H0: @Tom = @Sergio, H1: @Tom > @Sergio )\n\n| Metric | Mean Difference | T-Statistic | P-Value | Reject? |\n| --- | --- | --- | --- | --- |\n| Surface_dice | +0.0086 | 74.97 | <0.0001 | Y |\n| Score | -0.0017 | -7.64 | 1 | N |\n| Topo | -0.0124 | -20.11 | 1 | N |\n| Voi | -0.004 | -117.68 | 1 | N |\n\n* Based on the hypothesis testing, these two methods may not have a significant difference in topology or VOI score on the public subset. In that case, the method with the higher surface Dice score dominates, resulting in a 0.007 boost on the leaderboard. The other one we hoped improving lb experiences a substantial drop, scoring 0.588, which does bad in surface dice.\n\n## Post-Processing\n\n| Post Processing | LB | PB | Comments |\n| --- | --- | --- | --- |\n| remove cc < 3000 | 0.592 | 0.614 | |\n| x7 median filter | 0.596 | 0.623 | |\n| x6 median filter | 0.597 | 0.623 | |\n| x8 median filter | 0.596 | 0.624 | |\n| x9 median filter | 0.596 | 0.624 | |\n| x10 median filter | 0.597 | 0.624 | |\n| binary closing | 0.588 | 0.615 | |\n| closing(7)+hole patching+cavity fill | 0.583 | 0.604 | |\n| Gaussian smooth + rethreshold + close/open + fill | 0.585 | 0.603 | |\n| Remove internal cavities | 0.589 | 0.614 | |\n| Erosion / Dilation | 0.592 | 0.613 | |\n| Enforced minimum sheet thickness | 0.463 | 0.454 | Topo destroyed; Fails small sample |\n| Z-consistent smoothing | 0.592 | 0.615 | |\n\n---\n\n## Code \n\nhttps://github.com/Sersasj/vesuvius-challenge-10th-solution\n\n## References\n\n- Primus: Enforcing Attention Usage for 3D Medical Image Segmentation — https://openreview.net/forum?id=YWwGmmObri\n- MemBrain v2: An end-to-end tool for the analysis of membranes in cryo-electron tomography (Surface Dice Loss) — https://www.biorxiv.org/content/10.1101/2024.01.05.574336v1.full.pdf\n- Skeleton Recall Loss for Connectivity Conserving and Resource Efficient Segmentation of Thin Tubular Structures — https://www.ecva.net/papers/eccv_2024/papers_ECCV/papers/09904.pdf\n- dynamic-network-architectures library — https://github.com/MIC-DKFZ/dynamic-network-architectures\n- PrimusV2 code (for some reason it's not merged, we got lucky to find this fork) — https://github.com/TaWald/dynamic-network-architectures/blob/main/dynamic_network_architectures/architectures/primus.py\n- **Perceptual Contrastive Generative Adversarial Network based on image warping for unsupervised image-to-image translation(**https://www.sciencedirect.com/science/article/abs/pii/S0893608023003684**)**\n- **FlowNet 2.0: Evolution of Optical Flow Estimation with Deep Networks(**https://arxiv.org/abs/1612.01925**)**",
    "3414971": "This is the second time I have served as the leader of this team. This competition has given me far more than technical experience. It has provided valuable lessons beyond the contest itself. As a team leader, I had to learn how to respond when my teammates felt discouraged after being overtaken or when we could not come up with a solution. I needed to find ways to bring everyone back into the competition and refocus our efforts. I also had to manage my own emotions and stay composed so that we could continue moving forward even when frustration arose.\n\nI am truly grateful to my teammates for staying with me throughout this journey, never giving up, and continuing to experiment and develop workable solutions. I hope to keep working with this group and strive together for even greater achievements and more success in future competitions.\n\nAfter seeing the diffeomorphic network approach applied in this competition, I believe it will likely be adopted by many teams in future related contests. Personally, I hope more researchers will explore this method further and develop even more innovative and distinctive approaches based on it. ",
    "3416467": "### Post Deadline Experiments\n\nSince we had not focused much on post-processing, for post deadline experiment, we incorporated the simple iterative median filter used in the [18th-place solution](https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/writeups/18th-median-filter-x-7-post-processing-is-very-s) into [our pipeline](https://www.kaggle.com/code/tom99763/writeup-exp-with-median-filter-5-stages). This addition substantially improved our results, yielding a significant boost in the topological score. Our five-stage setup with only @tom99763 diffeomorphic network benefited greatly from this refinement.\n\n| Configuration | LB | PB |\n| --- | --- | --- |\n| 5-stage + remove small object (baseline) | 0.592 | 0.614 |\n| 6-stage + remove small object (baseline) | 0.588 | 0.615 |\n| 6-stage + 7x median filter + remove small object | 0.593 | 0.618 |\n| 5-stage + 6x median filter + remove small object | 0.597 | 0.623 |\n| 5-stage + 7x median filter + remove small object | 0.596 | 0.623 |\n| 5-stage + 8x median filter + remove small object | 0.596 | 0.624 |\n\n\n* CV analysis (clarify: 1-stage = baseline, 2-stage = baseline+median filter)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4310004%2Fc40896deaeaf40f0747196ced2e11617%2Fimage.png?generation=1772495587219649&alt=media)",
    "3415767": "Congrats on the great submission.\nI am looking forward to read also the Diffeomorphic Stage part in more details 👀\n",
    "3416180": "Congrats on the soultion!",
    "3415047": "@tom99763 \nin the last week it read about this which you may be interested\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa26bf3fac75ced4087d346952bf54d78%2FSelection_2458.png?generation=1772255662756779&alt=media)\n\ncurveformer++: https://arxiv.org/html/2402.06423v2",
    "3414996": "I had assumed your solution would incorporate unlabeled data, but it appears not to, which is somewhat unexpected.\nIt appears that semi-supervised methods are not ultimately required for the competition. Having spent a month writing semi-supervised training code, the final results still fell short of those achieved through supervised training.",
    "3415390": "I like your use of multiple stages and inputting multiple models into a new refinement model. This is a common technique in tabular data stacks. It's cool to see it work in deep learning image segmentation problems too."
  }
}