{
  "id": 679907,
  "title": "49th Place Solution - Vesuvius Surface Detection",
  "url": "/competitions/vesuvius-challenge-surface-detection/discussion/679907",
  "author_name": "TWEAK",
  "post_date": "2026-03-04T16:12:58.632000",
  "votes": 10,
  "comment_count": 0,
  "views": 0,
  "content": "<h1>Vesuvius Challenge - Surface Detection: 49th Place Solution</h1>\n<h3><em>Three skeleton-aware nnU-Nets and the power of loss function diversity</em></h3>\n<h2>Acknowledgements</h2>\n<p>Thank you to the Kaggle team and the Vesuvius Challenge team for all of their hard work and dedication to this challenge. It was a roller coaster -- multiple rescores and dataset updates kept everyone on their toes throughout the competition. The effort that went into curating the data, fixing ground truth labels, and maintaining the evaluation infrastructure was enormous, and we're grateful for the opportunity to contribute to this incredible mission.</p>\n<p>A special thank you to my teammate <strong>Sunny</strong> -- this solution wouldn't have been possible without your incredible work and dedication throughout the competition. From brainstorming ideas to grinding through late-night experiments, your contributions were invaluable. It was a privilege to tackle this challenge together.</p>\n<hr>\n<h2>Challenge Overview</h2>\n<p>The Vesuvius Challenge - Surface Detection competition tasked participants with detecting the thin papyrus surfaces within 3D micro-CT volumes of ancient Herculaneum scroll fragments. Each volume is approximately 320x320x320 voxels at isotropic 1.0mm spacing. The surfaces appear as faint, sheet-like structures embedded in the scan, and the goal is to produce a binary segmentation mask identifying these surfaces.</p>\n<p>The competition scoring combined three metrics:</p>\n<pre><code>Score = 0.30 x TopoScore + 0.35 x SurfaceDice@2.0 + 0.35 x VOI_score\n</code></pre>\n<p>This scoring formula heavily rewarded topologically clean predictions -- not just voxel-level accuracy, but the structural integrity of the predicted surfaces.</p>\n<p><strong>Final Result: Score 0.598, 49th place on the private leaderboard.</strong></p>\n<hr>\n<h2>Approach: Three Skeleton-Aware nnU-Nets</h2>\n<p>Our final submission was a <strong>3-model ensemble</strong> of nnU-Net v2 Residual Encoder U-Nets, all sharing the same architecture but trained with different skeleton-aware loss configurations. The ensemble combined models with complementary strengths -- a conservative model that avoided false positives, an aggressive model that maximized surface recall, and a longer-trained model that balanced both.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Trainer</th>\n<th>Skeleton Recall Weight</th>\n<th>FP Penalty Weight</th>\n<th>Epochs</th>\n<th>EMA Pseudo Dice</th>\n<th>Public LB (solo)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>skel_all</strong></td>\n<td>SkeletonBased</td>\n<td>0.5</td>\n<td>0.3</td>\n<td>1000</td>\n<td>0.5708</td>\n<td>0.560</td>\n</tr>\n<tr>\n<td><strong>aggressive_all</strong></td>\n<td>SkeletonAggressive</td>\n<td>0.7</td>\n<td>0.2</td>\n<td>1000</td>\n<td>0.5527</td>\n<td>0.563</td>\n</tr>\n<tr>\n<td><strong>basic1500</strong></td>\n<td>SkeletonBased1500</td>\n<td>0.5</td>\n<td>0.3</td>\n<td>1500</td>\n<td>0.6089</td>\n<td>0.587</td>\n</tr>\n</tbody>\n</table>\n<p>All three models were trained on all 737 volumes (<code>fold_all</code>) from Dataset504_VesuviusNewData3Fold. The ensemble logits were averaged before post-processing, and the combined prediction scored <strong>0.598 on the private leaderboard</strong> -- outperforming any individual model.</p>\n<h3>Why the Ensemble Worked</h3>\n<p>On the public leaderboard, basic1500 solo (0.587) appeared to beat every 2-model ensemble we tried (~0.582). This led us to believe ensembling was hurting. But the private leaderboard told a different story -- the 3-model ensemble generalized better to unseen data.</p>\n<p>The key was <strong>loss function diversity</strong>. Each model learned a slightly different trade-off between surface recall and precision:</p>\n<ul>\n<li><strong>skel_all</strong> (moderate, 1000 epochs): conservative baseline, fewer false positives</li>\n<li><strong>aggressive_all</strong> (aggressive, 1000 epochs): pushed harder on recall, caught surfaces the others missed</li>\n<li><strong>basic1500</strong> (moderate, 1500 epochs): the anchor -- longer training gave it the best individual performance</li>\n</ul>\n<p>Averaging their logits smoothed out each model's failure modes. Where aggressive_all over-predicted, skel_all and basic1500 pulled it back. Where the conservative models missed faint surfaces, aggressive_all filled in the gaps.</p>\n<hr>\n<h2>Model Architecture</h2>\n<p>All three models shared the same <strong>ResEncUNetM</strong> (Residual Encoder U-Net Medium) architecture from nnU-Net v2, planned by <code>nnUNetPlannerResEncM</code>:</p>\n<table>\n<thead>\n<tr>\n<th>Parameter</th>\n<th>Value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Network</td>\n<td><code>ResidualEncoderUNet</code> (fully 3D)</td>\n</tr>\n<tr>\n<td>Stages</td>\n<td>6</td>\n</tr>\n<tr>\n<td>Features per stage</td>\n<td>[32, 64, 128, 256, 320, 320]</td>\n</tr>\n<tr>\n<td>Convolution</td>\n<td><code>Conv3d</code>, 3x3x3 kernels throughout</td>\n</tr>\n<tr>\n<td>Strides</td>\n<td>[1,1,1] then [2,2,2] for stages 2-6</td>\n</tr>\n<tr>\n<td>Encoder blocks per stage</td>\n<td>[1, 3, 4, 6, 6, 6]</td>\n</tr>\n<tr>\n<td>Normalization</td>\n<td>InstanceNorm3d</td>\n</tr>\n<tr>\n<td>Activation</td>\n<td>LeakyReLU (inplace)</td>\n</tr>\n<tr>\n<td>Deep supervision</td>\n<td>5 scales, weights [0.533, 0.267, 0.133, 0.067, 0.0]</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h2>Custom Loss Function: SkeletonRecallPlusDiceLoss</h2>\n<p>Standard Dice + CE loss struggles with thin surface structures -- the extreme class imbalance (surfaces are just a few voxels thick in a 320^3 volume) means the model can achieve high Dice by simply being conservative. We designed a custom three-component loss to address this:</p>\n<pre><code>Total Loss = w_dice * (Dice + CE) + w_skel * SkeletonRecall + w_fp * FalsePositivePenalty\n</code></pre>\n<p><strong>Component 1: Base Dice + Cross-Entropy (weight 1.0)</strong>\nStandard nnU-Net segmentation loss on the original mask. Provides the foundation for voxel-accurate prediction.</p>\n<p><strong>Component 2: Skeleton Recall Loss (weight 0.5 or 0.7)</strong>\nOn-the-fly, we compute the morphological skeleton of the ground truth surface mask during training via a custom <code>SkeletonTransform</code>. The skeleton recall loss then penalizes the model for missing these skeleton voxels:</p>\n<pre><code>skeleton_recall = sum(pred * skeleton) / sum(skeleton)\nloss = 1.0 - skeleton_recall\n</code></pre>\n<p>This forces the model to maintain the thin medial axis of each surface, preventing it from \"eroding away\" thin structures even when that would improve the Dice score.</p>\n<p><strong>Component 3: False Positive Penalty (weight 0.3 or 0.2)</strong>\nPenalizes predictions in confirmed background regions by computing the mean foreground probability where the ground truth is background:</p>\n<pre><code>fp_loss = sum(pred_fg * background_mask) / sum(background_mask)\n</code></pre>\n<p>This prevents the model from over-predicting to boost skeleton recall, keeping surfaces clean and preventing merges between adjacent sheets.</p>\n<h3>The Two Loss Configurations</h3>\n<table>\n<thead>\n<tr>\n<th>Config</th>\n<th>Skeleton Recall</th>\n<th>FP Penalty</th>\n<th>Character</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Moderate</strong> (skel_all, basic1500)</td>\n<td>0.5</td>\n<td>0.3</td>\n<td>Conservative -- cleaner surfaces, fewer false positives</td>\n</tr>\n<tr>\n<td><strong>Aggressive</strong> (aggressive_all)</td>\n<td>0.7</td>\n<td>0.2</td>\n<td>Recall-focused -- catches more surfaces, noisier output</td>\n</tr>\n</tbody>\n</table>\n<p>The moderate configuration produced individually stronger models, but the aggressive variant contributed valuable diversity to the ensemble.</p>\n<hr>\n<h2>Training Data &amp; Augmentation</h2>\n<p><strong>Dataset</strong>: Dataset504_VesuviusNewData3Fold — 737 3D micro-CT volumes of ancient Herculaneum scroll fragments with binary surface labels (0=background, 1=surface, 2=ignore). The training ground truth was known to contain topology defects (holes, tunnels).</p>\n<h3>Custom Augmentation: Occlusion3DTransform</h3>\n<p>Beyond standard nnU-Net 3D augmentations (rotation, scaling, mirroring, gamma, noise), we added a custom <strong>Occlusion3DTransform</strong>:</p>\n<ul>\n<li>Randomly places 1-6 cuboid occlusions (size 2-8 voxels per dimension) over the input volume</li>\n<li>Applied with 30% probability per sample</li>\n<li>Occlusion regions are zeroed out, forcing the model to learn to predict surfaces even with missing data</li>\n<li>This improved robustness to the noisy, artifact-prone micro-CT inputs</li>\n</ul>\n<h3>Training Configuration</h3>\n<p>All models shared these settings:</p>\n<table>\n<thead>\n<tr>\n<th>Parameter</th>\n<th>Value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Dataset</td>\n<td>Dataset504_VesuviusNewData3Fold</td>\n</tr>\n<tr>\n<td>Training volumes</td>\n<td>737</td>\n</tr>\n<tr>\n<td>Fold</td>\n<td><code>fold_all</code> (all data, no holdout)</td>\n</tr>\n<tr>\n<td>Initial learning rate</td>\n<td>0.01</td>\n</tr>\n<tr>\n<td>LR schedule</td>\n<td>Polynomial decay</td>\n</tr>\n<tr>\n<td>Patch size</td>\n<td>128 x 128 x 128</td>\n</tr>\n<tr>\n<td>Batch size</td>\n<td>2</td>\n</tr>\n<tr>\n<td>torch.compile</td>\n<td>Enabled</td>\n</tr>\n<tr>\n<td>Labels</td>\n<td>0=background, 1=surface, 2=ignore</td>\n</tr>\n<tr>\n<td>Plans</td>\n<td>nnUNetResEncUNetMPlans</td>\n</tr>\n<tr>\n<td>Training hardware</td>\n<td>Single GPU (local, WSL2)</td>\n</tr>\n</tbody>\n</table>\n<p>The only differences were epoch count (1000 for skel_all and aggressive_all, 1500 for basic1500) and loss weights.</p>\n<p>Training on all 737 volumes (<code>fold_all</code>) was a deliberate choice -- with limited data and a high-capacity 3D model, we maximized the data available for learning.</p>\n<h3>Training Commands</h3>\n<pre><code># skel_all (moderate, 1000 epochs)\nnnUNetv2_train 504 3d_fullres fold_all \\\n    -tr nnUNetTrainerSkeletonBased \\\n    -p nnUNetResEncUNetMPlans\n\n# aggressive_all (aggressive, 1000 epochs)\nnnUNetv2_train 504 3d_fullres fold_all \\\n    -tr nnUNetTrainerSkeletonAggressive \\\n    -p nnUNetResEncUNetMPlans\n\n# basic1500 (moderate, 1500 epochs)\nnnUNetv2_train 504 3d_fullres fold_all \\\n    -tr nnUNetTrainerSkeletonBased1500 \\\n    -p nnUNetResEncUNetMPlans\n</code></pre>\n<hr>\n<h2>Inference Pipeline</h2>\n<p>Inference was run on <strong>Kaggle notebooks with 2x Tesla T4 GPUs</strong> within the 9-hour time limit.</p>\n<h3>Step 1: Ensemble Inference with TTA</h3>\n<p>Each of the three models ran nnU-Net's built-in sliding window predictor with <strong>Test Time Augmentation (TTA)</strong> -- 8 mirrored forward passes per volume (all axis-aligned flips). TTA was critical, providing a consistent +0.01-0.02 score improvement per model.</p>\n<p>Models were distributed across both T4 GPUs for parallel inference, achieving ~1.5x speedup (~274s vs 411s per volume on a single GPU).</p>\n<h3>Step 2: Weighted Logit Averaging</h3>\n<p>The softmax logits from all three models were combined using <strong>weighted averaging</strong> before thresholding:</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Weight</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>skel_all</td>\n<td>0.5</td>\n</tr>\n<tr>\n<td>aggressive_all</td>\n<td>0.5</td>\n</tr>\n<tr>\n<td>basic1500</td>\n<td>1.0</td>\n</tr>\n</tbody>\n</table>\n<p>basic1500 received double the weight of the other two models, reflecting its significantly stronger individual performance (0.587 vs 0.560/0.563 solo). The two supporting models each contributed half-weight, adding diversity without diluting the anchor. This is where the ensemble's balance paid off -- the aggressive model's higher recall was tempered by basic1500's dominance in the average, while skel_all and aggressive_all together still contributed enough signal to improve generalization on the private test set.</p>\n<h3>Step 3: Watershed Instance Separation</h3>\n<p>The key post-processing step was watershed-based instance separation, which prevents adjacent papyrus sheets from merging in the prediction:</p>\n<table>\n<thead>\n<tr>\n<th>Parameter</th>\n<th>Value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>High threshold (seed cores)</td>\n<td>0.55</td>\n</tr>\n<tr>\n<td>Low threshold (growth mask)</td>\n<td>0.25</td>\n</tr>\n<tr>\n<td>Boundary gap</td>\n<td>1</td>\n</tr>\n<tr>\n<td>Connectivity</td>\n<td>26</td>\n</tr>\n<tr>\n<td>Seed minimum size</td>\n<td>500</td>\n</tr>\n<tr>\n<td>Smoothing sigma</td>\n<td>0.6</td>\n</tr>\n<tr>\n<td>2D hole filling</td>\n<td>Enabled (max area 32)</td>\n</tr>\n<tr>\n<td>Min component size</td>\n<td>200</td>\n</tr>\n</tbody>\n</table>\n<p><strong>How it works:</strong></p>\n<ol>\n<li>Voxels above the high threshold (0.55) become seed cores for each surface instance</li>\n<li>Seeds smaller than 500 voxels are removed to prevent over-splitting</li>\n<li>The probability map is lightly smoothed (sigma=0.6) to reduce micro-basins</li>\n<li>Watershed flooding grows each seed into the low threshold mask (0.25)</li>\n<li>Boundaries between watershed regions create 1-voxel gaps separating instances</li>\n<li>Small holes within instances are filled (2D, max area 32)</li>\n<li>Components smaller than 200 voxels are removed</li>\n</ol>\n<h3>Step 4: Skeletonization and Re-dilation</h3>\n<p>After watershed separation, we applied 2D slice-wise skeletonization to reduce each surface to its medial axis, then re-dilated by 2 iterations as a safety buffer. This ensured predictions were thin and centered on the true surface location while maintaining sufficient coverage.</p>\n<h3>Step 5: Output</h3>\n<p>Final output: 320x320x320 uint8 binary TIF masks, packaged as <code>submission.zip</code>.</p>\n<hr>\n<h2>What Didn't Work (But We Tried)</h2>\n<p>A significant portion of competition time was spent on approaches that ultimately did not improve over our final ensemble:</p>\n<table>\n<thead>\n<tr>\n<th>Approach</th>\n<th>Result</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Otsu-thresholded inputs</strong> (Dataset511)</td>\n<td>Preprocessing inputs with Otsu thresholding dropped the score from 0.587 to 0.562. Critical intensity information was lost.</td>\n</tr>\n<tr>\n<td><strong>Cleaned ground truth</strong> (Dataset512)</td>\n<td>Cleaning the training GT by filling holes and closing gaps (3D binary closing with cube structuring element) produced similar training metrics but did not clearly improve the hidden test score.</td>\n</tr>\n<tr>\n<td><strong>Frangi surface filter</strong></td>\n<td>3D/2D Frangi filters (multiply, add, blend, repair modes) -- all disabled in the best submission.</td>\n</tr>\n<tr>\n<td><strong>CRF post-processing</strong></td>\n<td>Dense CRF smoothing added no benefit.</td>\n</tr>\n<tr>\n<td><strong>GraphCut</strong></td>\n<td>Memory intensive and no improvement.</td>\n</tr>\n<tr>\n<td><strong>PCA bridge-cut refinement</strong></td>\n<td>Attempted 0 cuts on test data -- no applicable cases found.</td>\n</tr>\n<tr>\n<td><strong>Binary morphological closing</strong></td>\n<td>CLOSE_ITERS=0 in best config.</td>\n</tr>\n<tr>\n<td><strong>Line pattern filter</strong></td>\n<td>Bit-pattern convolution for line detection -- disabled.</td>\n</tr>\n<tr>\n<td><strong>Marching Ants recursive splitting</strong></td>\n<td>Complex ray-casting + Dijkstra approach for splitting merged surfaces -- overkill for this data.</td>\n</tr>\n<tr>\n<td><strong>Spatial shift correction</strong></td>\n<td>Shifting predictions by +/-1 voxel to compensate for potential GT alignment issues -- no help.</td>\n</tr>\n<tr>\n<td><strong>2D medial surface approach</strong> (host_all)</td>\n<td>XY-plane medial surface detection scored 0.556 solo, well below the 3D approach.</td>\n</tr>\n<tr>\n<td><strong>Model soups / weight averaging</strong></td>\n<td>Averaging weights across checkpoints or models did not improve predictions.</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h2>The Rescore Effect</h2>\n<p>A pivotal moment in the competition was the mid-competition ground truth rescore. The competition hosts fixed topology defects in the test set ground truth -- filling cavities, closing tunnels, and performing manual topology corrections. This dramatically reshuffled the leaderboard.</p>\n<p>Models that had learned to reproduce the noisy original GT were penalized, while models producing clean, topologically sound surfaces were rewarded. Our skeleton-based loss naturally produced cleaner surfaces -- the skeleton recall component encouraged the model to preserve surface topology rather than match noisy voxel labels exactly.</p>\n<hr>\n<h2>Public vs Private Leaderboard</h2>\n<p>An interesting dynamic played out between the public and private leaderboards:</p>\n<table>\n<thead>\n<tr>\n<th>Configuration</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>basic1500 solo</td>\n<td>0.587</td>\n<td>0.596</td>\n</tr>\n<tr>\n<td>skel_all + aggressive_all</td>\n<td>0.587</td>\n<td>0.601</td>\n</tr>\n<tr>\n<td><strong>skel_all + aggressive_all + basic1500</strong> (selected)</td>\n<td>--</td>\n<td><strong>0.598</strong></td>\n</tr>\n</tbody>\n</table>\n<p>On the public leaderboard, basic1500 solo and the 2-model ensemble (skel_all + aggressive_all) were tied at 0.587. With no way to distinguish them publicly, we selected the 3-model ensemble as our final submission. In hindsight, the 2-model ensemble of skel_all + aggressive_all would have scored 0.601 on the private leaderboard -- our best result -- but we had no way of knowing that at the time.</p>\n<p>The lesson: <strong>public leaderboard ties hide real differences.</strong> Two submissions that look identical on a small public test set can diverge meaningfully on the full private set. When in doubt, diversity wins.</p>\n<hr>\n<h2>Key Takeaways</h2>\n<ol>\n<li><p><strong>Loss function design matters more than architecture.</strong> The custom skeleton-aware loss was the core innovation -- it directly addressed the thin-surface detection challenge in a way that standard Dice + CE could not.</p></li>\n<li><p><strong>Ensemble diversity through loss variation.</strong> Rather than ensembling different architectures, we ensembled the same architecture trained with different loss weightings. The moderate and aggressive configurations made complementary errors that cancelled out when averaged.</p></li>\n<li><p><strong>Don't trust the public leaderboard.</strong> basic1500 solo looked unbeatable on the public LB, but the 3-model ensemble proved its worth on the private test set. Diversity matters for generalization.</p></li>\n<li><p><strong>TTA is free performance.</strong> 8-way mirror augmentation at test time consistently added +0.01-0.02 with no architecture changes.</p></li>\n<li><p><strong>Post-processing should be surgical.</strong> Watershed separation for preventing sheet merges was valuable; everything else (Frangi, CRF, GraphCut, etc.) added complexity without improving scores.</p></li>\n<li><p><strong>Don't throw away information.</strong> Otsu-thresholding the inputs seemed like a reasonable preprocessing step but destroyed subtle intensity gradients the model relied on.</p></li>\n<li><p><strong>Train longer on all data.</strong> basic1500 (1500 epochs, fold_all) was the strongest individual model. More epochs and more data -- no magic, just patience.</p></li>\n</ol>\n<hr>\n<h2>Final Summary</h2>\n<table>\n<thead>\n<tr>\n<th></th>\n<th></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Models</strong></td>\n<td>3x nnU-Net v2 ResEncUNetM</td>\n</tr>\n<tr>\n<td><strong>Ensemble</strong></td>\n<td>skel_all + aggressive_all + basic1500</td>\n</tr>\n<tr>\n<td><strong>Loss</strong></td>\n<td>SkeletonRecallPlusDiceLoss (varied weights)</td>\n</tr>\n<tr>\n<td><strong>Data</strong></td>\n<td>737 volumes, fold_all</td>\n</tr>\n<tr>\n<td><strong>Epochs</strong></td>\n<td>1000 / 1000 / 1500</td>\n</tr>\n<tr>\n<td><strong>Post-processing</strong></td>\n<td>Logit averaging + Watershed + Skeletonize + Dilate(2)</td>\n</tr>\n<tr>\n<td><strong>Inference</strong></td>\n<td>2x T4 GPU, TTA (8 mirrors)</td>\n</tr>\n<tr>\n<td><strong>Private LB Score</strong></td>\n<td><strong>0.598</strong></td>\n</tr>\n<tr>\n<td><strong>Final Placement</strong></td>\n<td><strong>49th</strong></td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": 3417101,
      "postDate": "2026-03-04T16:12:58.633Z",
      "content": "<h1>Vesuvius Challenge - Surface Detection: 49th Place Solution</h1>\n<h3><em>Three skeleton-aware nnU-Nets and the power of loss function diversity</em></h3>\n<h2>Acknowledgements</h2>\n<p>Thank you to the Kaggle team and the Vesuvius Challenge team for all of their hard work and dedication to this challenge. It was a roller coaster -- multiple rescores and dataset updates kept everyone on their toes throughout the competition. The effort that went into curating the data, fixing ground truth labels, and maintaining the evaluation infrastructure was enormous, and we're grateful for the opportunity to contribute to this incredible mission.</p>\n<p>A special thank you to my teammate <strong>Sunny</strong> -- this solution wouldn't have been possible without your incredible work and dedication throughout the competition. From brainstorming ideas to grinding through late-night experiments, your contributions were invaluable. It was a privilege to tackle this challenge together.</p>\n<hr>\n<h2>Challenge Overview</h2>\n<p>The Vesuvius Challenge - Surface Detection competition tasked participants with detecting the thin papyrus surfaces within 3D micro-CT volumes of ancient Herculaneum scroll fragments. Each volume is approximately 320x320x320 voxels at isotropic 1.0mm spacing. The surfaces appear as faint, sheet-like structures embedded in the scan, and the goal is to produce a binary segmentation mask identifying these surfaces.</p>\n<p>The competition scoring combined three metrics:</p>\n<pre><code>Score = 0.30 x TopoScore + 0.35 x SurfaceDice@2.0 + 0.35 x VOI_score\n</code></pre>\n<p>This scoring formula heavily rewarded topologically clean predictions -- not just voxel-level accuracy, but the structural integrity of the predicted surfaces.</p>\n<p><strong>Final Result: Score 0.598, 49th place on the private leaderboard.</strong></p>\n<hr>\n<h2>Approach: Three Skeleton-Aware nnU-Nets</h2>\n<p>Our final submission was a <strong>3-model ensemble</strong> of nnU-Net v2 Residual Encoder U-Nets, all sharing the same architecture but trained with different skeleton-aware loss configurations. The ensemble combined models with complementary strengths -- a conservative model that avoided false positives, an aggressive model that maximized surface recall, and a longer-trained model that balanced both.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Trainer</th>\n<th>Skeleton Recall Weight</th>\n<th>FP Penalty Weight</th>\n<th>Epochs</th>\n<th>EMA Pseudo Dice</th>\n<th>Public LB (solo)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>skel_all</strong></td>\n<td>SkeletonBased</td>\n<td>0.5</td>\n<td>0.3</td>\n<td>1000</td>\n<td>0.5708</td>\n<td>0.560</td>\n</tr>\n<tr>\n<td><strong>aggressive_all</strong></td>\n<td>SkeletonAggressive</td>\n<td>0.7</td>\n<td>0.2</td>\n<td>1000</td>\n<td>0.5527</td>\n<td>0.563</td>\n</tr>\n<tr>\n<td><strong>basic1500</strong></td>\n<td>SkeletonBased1500</td>\n<td>0.5</td>\n<td>0.3</td>\n<td>1500</td>\n<td>0.6089</td>\n<td>0.587</td>\n</tr>\n</tbody>\n</table>\n<p>All three models were trained on all 737 volumes (<code>fold_all</code>) from Dataset504_VesuviusNewData3Fold. The ensemble logits were averaged before post-processing, and the combined prediction scored <strong>0.598 on the private leaderboard</strong> -- outperforming any individual model.</p>\n<h3>Why the Ensemble Worked</h3>\n<p>On the public leaderboard, basic1500 solo (0.587) appeared to beat every 2-model ensemble we tried (~0.582). This led us to believe ensembling was hurting. But the private leaderboard told a different story -- the 3-model ensemble generalized better to unseen data.</p>\n<p>The key was <strong>loss function diversity</strong>. Each model learned a slightly different trade-off between surface recall and precision:</p>\n<ul>\n<li><strong>skel_all</strong> (moderate, 1000 epochs): conservative baseline, fewer false positives</li>\n<li><strong>aggressive_all</strong> (aggressive, 1000 epochs): pushed harder on recall, caught surfaces the others missed</li>\n<li><strong>basic1500</strong> (moderate, 1500 epochs): the anchor -- longer training gave it the best individual performance</li>\n</ul>\n<p>Averaging their logits smoothed out each model's failure modes. Where aggressive_all over-predicted, skel_all and basic1500 pulled it back. Where the conservative models missed faint surfaces, aggressive_all filled in the gaps.</p>\n<hr>\n<h2>Model Architecture</h2>\n<p>All three models shared the same <strong>ResEncUNetM</strong> (Residual Encoder U-Net Medium) architecture from nnU-Net v2, planned by <code>nnUNetPlannerResEncM</code>:</p>\n<table>\n<thead>\n<tr>\n<th>Parameter</th>\n<th>Value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Network</td>\n<td><code>ResidualEncoderUNet</code> (fully 3D)</td>\n</tr>\n<tr>\n<td>Stages</td>\n<td>6</td>\n</tr>\n<tr>\n<td>Features per stage</td>\n<td>[32, 64, 128, 256, 320, 320]</td>\n</tr>\n<tr>\n<td>Convolution</td>\n<td><code>Conv3d</code>, 3x3x3 kernels throughout</td>\n</tr>\n<tr>\n<td>Strides</td>\n<td>[1,1,1] then [2,2,2] for stages 2-6</td>\n</tr>\n<tr>\n<td>Encoder blocks per stage</td>\n<td>[1, 3, 4, 6, 6, 6]</td>\n</tr>\n<tr>\n<td>Normalization</td>\n<td>InstanceNorm3d</td>\n</tr>\n<tr>\n<td>Activation</td>\n<td>LeakyReLU (inplace)</td>\n</tr>\n<tr>\n<td>Deep supervision</td>\n<td>5 scales, weights [0.533, 0.267, 0.133, 0.067, 0.0]</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h2>Custom Loss Function: SkeletonRecallPlusDiceLoss</h2>\n<p>Standard Dice + CE loss struggles with thin surface structures -- the extreme class imbalance (surfaces are just a few voxels thick in a 320^3 volume) means the model can achieve high Dice by simply being conservative. We designed a custom three-component loss to address this:</p>\n<pre><code>Total Loss = w_dice * (Dice + CE) + w_skel * SkeletonRecall + w_fp * FalsePositivePenalty\n</code></pre>\n<p><strong>Component 1: Base Dice + Cross-Entropy (weight 1.0)</strong>\nStandard nnU-Net segmentation loss on the original mask. Provides the foundation for voxel-accurate prediction.</p>\n<p><strong>Component 2: Skeleton Recall Loss (weight 0.5 or 0.7)</strong>\nOn-the-fly, we compute the morphological skeleton of the ground truth surface mask during training via a custom <code>SkeletonTransform</code>. The skeleton recall loss then penalizes the model for missing these skeleton voxels:</p>\n<pre><code>skeleton_recall = sum(pred * skeleton) / sum(skeleton)\nloss = 1.0 - skeleton_recall\n</code></pre>\n<p>This forces the model to maintain the thin medial axis of each surface, preventing it from \"eroding away\" thin structures even when that would improve the Dice score.</p>\n<p><strong>Component 3: False Positive Penalty (weight 0.3 or 0.2)</strong>\nPenalizes predictions in confirmed background regions by computing the mean foreground probability where the ground truth is background:</p>\n<pre><code>fp_loss = sum(pred_fg * background_mask) / sum(background_mask)\n</code></pre>\n<p>This prevents the model from over-predicting to boost skeleton recall, keeping surfaces clean and preventing merges between adjacent sheets.</p>\n<h3>The Two Loss Configurations</h3>\n<table>\n<thead>\n<tr>\n<th>Config</th>\n<th>Skeleton Recall</th>\n<th>FP Penalty</th>\n<th>Character</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Moderate</strong> (skel_all, basic1500)</td>\n<td>0.5</td>\n<td>0.3</td>\n<td>Conservative -- cleaner surfaces, fewer false positives</td>\n</tr>\n<tr>\n<td><strong>Aggressive</strong> (aggressive_all)</td>\n<td>0.7</td>\n<td>0.2</td>\n<td>Recall-focused -- catches more surfaces, noisier output</td>\n</tr>\n</tbody>\n</table>\n<p>The moderate configuration produced individually stronger models, but the aggressive variant contributed valuable diversity to the ensemble.</p>\n<hr>\n<h2>Training Data &amp; Augmentation</h2>\n<p><strong>Dataset</strong>: Dataset504_VesuviusNewData3Fold — 737 3D micro-CT volumes of ancient Herculaneum scroll fragments with binary surface labels (0=background, 1=surface, 2=ignore). The training ground truth was known to contain topology defects (holes, tunnels).</p>\n<h3>Custom Augmentation: Occlusion3DTransform</h3>\n<p>Beyond standard nnU-Net 3D augmentations (rotation, scaling, mirroring, gamma, noise), we added a custom <strong>Occlusion3DTransform</strong>:</p>\n<ul>\n<li>Randomly places 1-6 cuboid occlusions (size 2-8 voxels per dimension) over the input volume</li>\n<li>Applied with 30% probability per sample</li>\n<li>Occlusion regions are zeroed out, forcing the model to learn to predict surfaces even with missing data</li>\n<li>This improved robustness to the noisy, artifact-prone micro-CT inputs</li>\n</ul>\n<h3>Training Configuration</h3>\n<p>All models shared these settings:</p>\n<table>\n<thead>\n<tr>\n<th>Parameter</th>\n<th>Value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Dataset</td>\n<td>Dataset504_VesuviusNewData3Fold</td>\n</tr>\n<tr>\n<td>Training volumes</td>\n<td>737</td>\n</tr>\n<tr>\n<td>Fold</td>\n<td><code>fold_all</code> (all data, no holdout)</td>\n</tr>\n<tr>\n<td>Initial learning rate</td>\n<td>0.01</td>\n</tr>\n<tr>\n<td>LR schedule</td>\n<td>Polynomial decay</td>\n</tr>\n<tr>\n<td>Patch size</td>\n<td>128 x 128 x 128</td>\n</tr>\n<tr>\n<td>Batch size</td>\n<td>2</td>\n</tr>\n<tr>\n<td>torch.compile</td>\n<td>Enabled</td>\n</tr>\n<tr>\n<td>Labels</td>\n<td>0=background, 1=surface, 2=ignore</td>\n</tr>\n<tr>\n<td>Plans</td>\n<td>nnUNetResEncUNetMPlans</td>\n</tr>\n<tr>\n<td>Training hardware</td>\n<td>Single GPU (local, WSL2)</td>\n</tr>\n</tbody>\n</table>\n<p>The only differences were epoch count (1000 for skel_all and aggressive_all, 1500 for basic1500) and loss weights.</p>\n<p>Training on all 737 volumes (<code>fold_all</code>) was a deliberate choice -- with limited data and a high-capacity 3D model, we maximized the data available for learning.</p>\n<h3>Training Commands</h3>\n<pre><code># skel_all (moderate, 1000 epochs)\nnnUNetv2_train 504 3d_fullres fold_all \\\n    -tr nnUNetTrainerSkeletonBased \\\n    -p nnUNetResEncUNetMPlans\n\n# aggressive_all (aggressive, 1000 epochs)\nnnUNetv2_train 504 3d_fullres fold_all \\\n    -tr nnUNetTrainerSkeletonAggressive \\\n    -p nnUNetResEncUNetMPlans\n\n# basic1500 (moderate, 1500 epochs)\nnnUNetv2_train 504 3d_fullres fold_all \\\n    -tr nnUNetTrainerSkeletonBased1500 \\\n    -p nnUNetResEncUNetMPlans\n</code></pre>\n<hr>\n<h2>Inference Pipeline</h2>\n<p>Inference was run on <strong>Kaggle notebooks with 2x Tesla T4 GPUs</strong> within the 9-hour time limit.</p>\n<h3>Step 1: Ensemble Inference with TTA</h3>\n<p>Each of the three models ran nnU-Net's built-in sliding window predictor with <strong>Test Time Augmentation (TTA)</strong> -- 8 mirrored forward passes per volume (all axis-aligned flips). TTA was critical, providing a consistent +0.01-0.02 score improvement per model.</p>\n<p>Models were distributed across both T4 GPUs for parallel inference, achieving ~1.5x speedup (~274s vs 411s per volume on a single GPU).</p>\n<h3>Step 2: Weighted Logit Averaging</h3>\n<p>The softmax logits from all three models were combined using <strong>weighted averaging</strong> before thresholding:</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Weight</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>skel_all</td>\n<td>0.5</td>\n</tr>\n<tr>\n<td>aggressive_all</td>\n<td>0.5</td>\n</tr>\n<tr>\n<td>basic1500</td>\n<td>1.0</td>\n</tr>\n</tbody>\n</table>\n<p>basic1500 received double the weight of the other two models, reflecting its significantly stronger individual performance (0.587 vs 0.560/0.563 solo). The two supporting models each contributed half-weight, adding diversity without diluting the anchor. This is where the ensemble's balance paid off -- the aggressive model's higher recall was tempered by basic1500's dominance in the average, while skel_all and aggressive_all together still contributed enough signal to improve generalization on the private test set.</p>\n<h3>Step 3: Watershed Instance Separation</h3>\n<p>The key post-processing step was watershed-based instance separation, which prevents adjacent papyrus sheets from merging in the prediction:</p>\n<table>\n<thead>\n<tr>\n<th>Parameter</th>\n<th>Value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>High threshold (seed cores)</td>\n<td>0.55</td>\n</tr>\n<tr>\n<td>Low threshold (growth mask)</td>\n<td>0.25</td>\n</tr>\n<tr>\n<td>Boundary gap</td>\n<td>1</td>\n</tr>\n<tr>\n<td>Connectivity</td>\n<td>26</td>\n</tr>\n<tr>\n<td>Seed minimum size</td>\n<td>500</td>\n</tr>\n<tr>\n<td>Smoothing sigma</td>\n<td>0.6</td>\n</tr>\n<tr>\n<td>2D hole filling</td>\n<td>Enabled (max area 32)</td>\n</tr>\n<tr>\n<td>Min component size</td>\n<td>200</td>\n</tr>\n</tbody>\n</table>\n<p><strong>How it works:</strong></p>\n<ol>\n<li>Voxels above the high threshold (0.55) become seed cores for each surface instance</li>\n<li>Seeds smaller than 500 voxels are removed to prevent over-splitting</li>\n<li>The probability map is lightly smoothed (sigma=0.6) to reduce micro-basins</li>\n<li>Watershed flooding grows each seed into the low threshold mask (0.25)</li>\n<li>Boundaries between watershed regions create 1-voxel gaps separating instances</li>\n<li>Small holes within instances are filled (2D, max area 32)</li>\n<li>Components smaller than 200 voxels are removed</li>\n</ol>\n<h3>Step 4: Skeletonization and Re-dilation</h3>\n<p>After watershed separation, we applied 2D slice-wise skeletonization to reduce each surface to its medial axis, then re-dilated by 2 iterations as a safety buffer. This ensured predictions were thin and centered on the true surface location while maintaining sufficient coverage.</p>\n<h3>Step 5: Output</h3>\n<p>Final output: 320x320x320 uint8 binary TIF masks, packaged as <code>submission.zip</code>.</p>\n<hr>\n<h2>What Didn't Work (But We Tried)</h2>\n<p>A significant portion of competition time was spent on approaches that ultimately did not improve over our final ensemble:</p>\n<table>\n<thead>\n<tr>\n<th>Approach</th>\n<th>Result</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Otsu-thresholded inputs</strong> (Dataset511)</td>\n<td>Preprocessing inputs with Otsu thresholding dropped the score from 0.587 to 0.562. Critical intensity information was lost.</td>\n</tr>\n<tr>\n<td><strong>Cleaned ground truth</strong> (Dataset512)</td>\n<td>Cleaning the training GT by filling holes and closing gaps (3D binary closing with cube structuring element) produced similar training metrics but did not clearly improve the hidden test score.</td>\n</tr>\n<tr>\n<td><strong>Frangi surface filter</strong></td>\n<td>3D/2D Frangi filters (multiply, add, blend, repair modes) -- all disabled in the best submission.</td>\n</tr>\n<tr>\n<td><strong>CRF post-processing</strong></td>\n<td>Dense CRF smoothing added no benefit.</td>\n</tr>\n<tr>\n<td><strong>GraphCut</strong></td>\n<td>Memory intensive and no improvement.</td>\n</tr>\n<tr>\n<td><strong>PCA bridge-cut refinement</strong></td>\n<td>Attempted 0 cuts on test data -- no applicable cases found.</td>\n</tr>\n<tr>\n<td><strong>Binary morphological closing</strong></td>\n<td>CLOSE_ITERS=0 in best config.</td>\n</tr>\n<tr>\n<td><strong>Line pattern filter</strong></td>\n<td>Bit-pattern convolution for line detection -- disabled.</td>\n</tr>\n<tr>\n<td><strong>Marching Ants recursive splitting</strong></td>\n<td>Complex ray-casting + Dijkstra approach for splitting merged surfaces -- overkill for this data.</td>\n</tr>\n<tr>\n<td><strong>Spatial shift correction</strong></td>\n<td>Shifting predictions by +/-1 voxel to compensate for potential GT alignment issues -- no help.</td>\n</tr>\n<tr>\n<td><strong>2D medial surface approach</strong> (host_all)</td>\n<td>XY-plane medial surface detection scored 0.556 solo, well below the 3D approach.</td>\n</tr>\n<tr>\n<td><strong>Model soups / weight averaging</strong></td>\n<td>Averaging weights across checkpoints or models did not improve predictions.</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h2>The Rescore Effect</h2>\n<p>A pivotal moment in the competition was the mid-competition ground truth rescore. The competition hosts fixed topology defects in the test set ground truth -- filling cavities, closing tunnels, and performing manual topology corrections. This dramatically reshuffled the leaderboard.</p>\n<p>Models that had learned to reproduce the noisy original GT were penalized, while models producing clean, topologically sound surfaces were rewarded. Our skeleton-based loss naturally produced cleaner surfaces -- the skeleton recall component encouraged the model to preserve surface topology rather than match noisy voxel labels exactly.</p>\n<hr>\n<h2>Public vs Private Leaderboard</h2>\n<p>An interesting dynamic played out between the public and private leaderboards:</p>\n<table>\n<thead>\n<tr>\n<th>Configuration</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>basic1500 solo</td>\n<td>0.587</td>\n<td>0.596</td>\n</tr>\n<tr>\n<td>skel_all + aggressive_all</td>\n<td>0.587</td>\n<td>0.601</td>\n</tr>\n<tr>\n<td><strong>skel_all + aggressive_all + basic1500</strong> (selected)</td>\n<td>--</td>\n<td><strong>0.598</strong></td>\n</tr>\n</tbody>\n</table>\n<p>On the public leaderboard, basic1500 solo and the 2-model ensemble (skel_all + aggressive_all) were tied at 0.587. With no way to distinguish them publicly, we selected the 3-model ensemble as our final submission. In hindsight, the 2-model ensemble of skel_all + aggressive_all would have scored 0.601 on the private leaderboard -- our best result -- but we had no way of knowing that at the time.</p>\n<p>The lesson: <strong>public leaderboard ties hide real differences.</strong> Two submissions that look identical on a small public test set can diverge meaningfully on the full private set. When in doubt, diversity wins.</p>\n<hr>\n<h2>Key Takeaways</h2>\n<ol>\n<li><p><strong>Loss function design matters more than architecture.</strong> The custom skeleton-aware loss was the core innovation -- it directly addressed the thin-surface detection challenge in a way that standard Dice + CE could not.</p></li>\n<li><p><strong>Ensemble diversity through loss variation.</strong> Rather than ensembling different architectures, we ensembled the same architecture trained with different loss weightings. The moderate and aggressive configurations made complementary errors that cancelled out when averaged.</p></li>\n<li><p><strong>Don't trust the public leaderboard.</strong> basic1500 solo looked unbeatable on the public LB, but the 3-model ensemble proved its worth on the private test set. Diversity matters for generalization.</p></li>\n<li><p><strong>TTA is free performance.</strong> 8-way mirror augmentation at test time consistently added +0.01-0.02 with no architecture changes.</p></li>\n<li><p><strong>Post-processing should be surgical.</strong> Watershed separation for preventing sheet merges was valuable; everything else (Frangi, CRF, GraphCut, etc.) added complexity without improving scores.</p></li>\n<li><p><strong>Don't throw away information.</strong> Otsu-thresholding the inputs seemed like a reasonable preprocessing step but destroyed subtle intensity gradients the model relied on.</p></li>\n<li><p><strong>Train longer on all data.</strong> basic1500 (1500 epochs, fold_all) was the strongest individual model. More epochs and more data -- no magic, just patience.</p></li>\n</ol>\n<hr>\n<h2>Final Summary</h2>\n<table>\n<thead>\n<tr>\n<th></th>\n<th></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Models</strong></td>\n<td>3x nnU-Net v2 ResEncUNetM</td>\n</tr>\n<tr>\n<td><strong>Ensemble</strong></td>\n<td>skel_all + aggressive_all + basic1500</td>\n</tr>\n<tr>\n<td><strong>Loss</strong></td>\n<td>SkeletonRecallPlusDiceLoss (varied weights)</td>\n</tr>\n<tr>\n<td><strong>Data</strong></td>\n<td>737 volumes, fold_all</td>\n</tr>\n<tr>\n<td><strong>Epochs</strong></td>\n<td>1000 / 1000 / 1500</td>\n</tr>\n<tr>\n<td><strong>Post-processing</strong></td>\n<td>Logit averaging + Watershed + Skeletonize + Dilate(2)</td>\n</tr>\n<tr>\n<td><strong>Inference</strong></td>\n<td>2x T4 GPU, TTA (8 mirrors)</td>\n</tr>\n<tr>\n<td><strong>Private LB Score</strong></td>\n<td><strong>0.598</strong></td>\n</tr>\n<tr>\n<td><strong>Final Placement</strong></td>\n<td><strong>49th</strong></td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "# Vesuvius Challenge - Surface Detection: 49th Place Solution\n### *Three skeleton-aware nnU-Nets and the power of loss function diversity*\n\n## Acknowledgements\n\nThank you to the Kaggle team and the Vesuvius Challenge team for all of their hard work and dedication to this challenge. It was a roller coaster -- multiple rescores and dataset updates kept everyone on their toes throughout the competition. The effort that went into curating the data, fixing ground truth labels, and maintaining the evaluation infrastructure was enormous, and we're grateful for the opportunity to contribute to this incredible mission.\n\nA special thank you to my teammate **Sunny** -- this solution wouldn't have been possible without your incredible work and dedication throughout the competition. From brainstorming ideas to grinding through late-night experiments, your contributions were invaluable. It was a privilege to tackle this challenge together.\n\n---\n\n## Challenge Overview\n\nThe Vesuvius Challenge - Surface Detection competition tasked participants with detecting the thin papyrus surfaces within 3D micro-CT volumes of ancient Herculaneum scroll fragments. Each volume is approximately 320x320x320 voxels at isotropic 1.0mm spacing. The surfaces appear as faint, sheet-like structures embedded in the scan, and the goal is to produce a binary segmentation mask identifying these surfaces.\n\nThe competition scoring combined three metrics:\n\n```\nScore = 0.30 x TopoScore + 0.35 x SurfaceDice@2.0 + 0.35 x VOI_score\n```\n\nThis scoring formula heavily rewarded topologically clean predictions -- not just voxel-level accuracy, but the structural integrity of the predicted surfaces.\n\n**Final Result: Score 0.598, 49th place on the private leaderboard.**\n\n---\n\n## Approach: Three Skeleton-Aware nnU-Nets\n\nOur final submission was a **3-model ensemble** of nnU-Net v2 Residual Encoder U-Nets, all sharing the same architecture but trained with different skeleton-aware loss configurations. The ensemble combined models with complementary strengths -- a conservative model that avoided false positives, an aggressive model that maximized surface recall, and a longer-trained model that balanced both.\n\n| Model | Trainer | Skeleton Recall Weight | FP Penalty Weight | Epochs | EMA Pseudo Dice | Public LB (solo) |\n|---|---|---|---|---|---|---|\n| **skel_all** | SkeletonBased | 0.5 | 0.3 | 1000 | 0.5708 | 0.560 |\n| **aggressive_all** | SkeletonAggressive | 0.7 | 0.2 | 1000 | 0.5527 | 0.563 |\n| **basic1500** | SkeletonBased1500 | 0.5 | 0.3 | 1500 | 0.6089 | 0.587 |\n\nAll three models were trained on all 737 volumes (`fold_all`) from Dataset504_VesuviusNewData3Fold. The ensemble logits were averaged before post-processing, and the combined prediction scored **0.598 on the private leaderboard** -- outperforming any individual model.\n\n### Why the Ensemble Worked\n\nOn the public leaderboard, basic1500 solo (0.587) appeared to beat every 2-model ensemble we tried (~0.582). This led us to believe ensembling was hurting. But the private leaderboard told a different story -- the 3-model ensemble generalized better to unseen data.\n\nThe key was **loss function diversity**. Each model learned a slightly different trade-off between surface recall and precision:\n- **skel_all** (moderate, 1000 epochs): conservative baseline, fewer false positives\n- **aggressive_all** (aggressive, 1000 epochs): pushed harder on recall, caught surfaces the others missed\n- **basic1500** (moderate, 1500 epochs): the anchor -- longer training gave it the best individual performance\n\nAveraging their logits smoothed out each model's failure modes. Where aggressive_all over-predicted, skel_all and basic1500 pulled it back. Where the conservative models missed faint surfaces, aggressive_all filled in the gaps.\n\n---\n\n## Model Architecture\n\nAll three models shared the same **ResEncUNetM** (Residual Encoder U-Net Medium) architecture from nnU-Net v2, planned by `nnUNetPlannerResEncM`:\n\n| Parameter | Value |\n|---|---|\n| Network | `ResidualEncoderUNet` (fully 3D) |\n| Stages | 6 |\n| Features per stage | [32, 64, 128, 256, 320, 320] |\n| Convolution | `Conv3d`, 3x3x3 kernels throughout |\n| Strides | [1,1,1] then [2,2,2] for stages 2-6 |\n| Encoder blocks per stage | [1, 3, 4, 6, 6, 6] |\n| Normalization | InstanceNorm3d |\n| Activation | LeakyReLU (inplace) |\n| Deep supervision | 5 scales, weights [0.533, 0.267, 0.133, 0.067, 0.0] |\n\n---\n\n## Custom Loss Function: SkeletonRecallPlusDiceLoss\n\nStandard Dice + CE loss struggles with thin surface structures -- the extreme class imbalance (surfaces are just a few voxels thick in a 320^3 volume) means the model can achieve high Dice by simply being conservative. We designed a custom three-component loss to address this:\n\n```\nTotal Loss = w_dice * (Dice + CE) + w_skel * SkeletonRecall + w_fp * FalsePositivePenalty\n```\n\n**Component 1: Base Dice + Cross-Entropy (weight 1.0)**\nStandard nnU-Net segmentation loss on the original mask. Provides the foundation for voxel-accurate prediction.\n\n**Component 2: Skeleton Recall Loss (weight 0.5 or 0.7)**\nOn-the-fly, we compute the morphological skeleton of the ground truth surface mask during training via a custom `SkeletonTransform`. The skeleton recall loss then penalizes the model for missing these skeleton voxels:\n\n```python\nskeleton_recall = sum(pred * skeleton) / sum(skeleton)\nloss = 1.0 - skeleton_recall\n```\n\nThis forces the model to maintain the thin medial axis of each surface, preventing it from \"eroding away\" thin structures even when that would improve the Dice score.\n\n**Component 3: False Positive Penalty (weight 0.3 or 0.2)**\nPenalizes predictions in confirmed background regions by computing the mean foreground probability where the ground truth is background:\n\n```python\nfp_loss = sum(pred_fg * background_mask) / sum(background_mask)\n```\n\nThis prevents the model from over-predicting to boost skeleton recall, keeping surfaces clean and preventing merges between adjacent sheets.\n\n### The Two Loss Configurations\n\n| Config | Skeleton Recall | FP Penalty | Character |\n|---|---|---|---|\n| **Moderate** (skel_all, basic1500) | 0.5 | 0.3 | Conservative -- cleaner surfaces, fewer false positives |\n| **Aggressive** (aggressive_all) | 0.7 | 0.2 | Recall-focused -- catches more surfaces, noisier output |\n\nThe moderate configuration produced individually stronger models, but the aggressive variant contributed valuable diversity to the ensemble.\n\n---\n\n## Training Data & Augmentation\n\n**Dataset**: Dataset504_VesuviusNewData3Fold — 737 3D micro-CT volumes of ancient Herculaneum scroll fragments with binary surface labels (0=background, 1=surface, 2=ignore). The training ground truth was known to contain topology defects (holes, tunnels).\n\n### Custom Augmentation: Occlusion3DTransform\n\nBeyond standard nnU-Net 3D augmentations (rotation, scaling, mirroring, gamma, noise), we added a custom **Occlusion3DTransform**:\n\n- Randomly places 1-6 cuboid occlusions (size 2-8 voxels per dimension) over the input volume\n- Applied with 30% probability per sample\n- Occlusion regions are zeroed out, forcing the model to learn to predict surfaces even with missing data\n- This improved robustness to the noisy, artifact-prone micro-CT inputs\n\n### Training Configuration\n\nAll models shared these settings:\n\n| Parameter | Value |\n|---|---|\n| Dataset | Dataset504_VesuviusNewData3Fold |\n| Training volumes | 737 |\n| Fold | `fold_all` (all data, no holdout) |\n| Initial learning rate | 0.01 |\n| LR schedule | Polynomial decay |\n| Patch size | 128 x 128 x 128 |\n| Batch size | 2 |\n| torch.compile | Enabled |\n| Labels | 0=background, 1=surface, 2=ignore |\n| Plans | nnUNetResEncUNetMPlans |\n| Training hardware | Single GPU (local, WSL2) |\n\nThe only differences were epoch count (1000 for skel_all and aggressive_all, 1500 for basic1500) and loss weights.\n\nTraining on all 737 volumes (`fold_all`) was a deliberate choice -- with limited data and a high-capacity 3D model, we maximized the data available for learning.\n\n### Training Commands\n\n```bash\n# skel_all (moderate, 1000 epochs)\nnnUNetv2_train 504 3d_fullres fold_all \\\n    -tr nnUNetTrainerSkeletonBased \\\n    -p nnUNetResEncUNetMPlans\n\n# aggressive_all (aggressive, 1000 epochs)\nnnUNetv2_train 504 3d_fullres fold_all \\\n    -tr nnUNetTrainerSkeletonAggressive \\\n    -p nnUNetResEncUNetMPlans\n\n# basic1500 (moderate, 1500 epochs)\nnnUNetv2_train 504 3d_fullres fold_all \\\n    -tr nnUNetTrainerSkeletonBased1500 \\\n    -p nnUNetResEncUNetMPlans\n```\n\n---\n\n## Inference Pipeline\n\nInference was run on **Kaggle notebooks with 2x Tesla T4 GPUs** within the 9-hour time limit.\n\n### Step 1: Ensemble Inference with TTA\n\nEach of the three models ran nnU-Net's built-in sliding window predictor with **Test Time Augmentation (TTA)** -- 8 mirrored forward passes per volume (all axis-aligned flips). TTA was critical, providing a consistent +0.01-0.02 score improvement per model.\n\nModels were distributed across both T4 GPUs for parallel inference, achieving ~1.5x speedup (~274s vs 411s per volume on a single GPU).\n\n### Step 2: Weighted Logit Averaging\n\nThe softmax logits from all three models were combined using **weighted averaging** before thresholding:\n\n| Model | Weight |\n|---|---|\n| skel_all | 0.5 |\n| aggressive_all | 0.5 |\n| basic1500 | 1.0 |\n\nbasic1500 received double the weight of the other two models, reflecting its significantly stronger individual performance (0.587 vs 0.560/0.563 solo). The two supporting models each contributed half-weight, adding diversity without diluting the anchor. This is where the ensemble's balance paid off -- the aggressive model's higher recall was tempered by basic1500's dominance in the average, while skel_all and aggressive_all together still contributed enough signal to improve generalization on the private test set.\n\n### Step 3: Watershed Instance Separation\n\nThe key post-processing step was watershed-based instance separation, which prevents adjacent papyrus sheets from merging in the prediction:\n\n| Parameter | Value |\n|---|---|\n| High threshold (seed cores) | 0.55 |\n| Low threshold (growth mask) | 0.25 |\n| Boundary gap | 1 |\n| Connectivity | 26 |\n| Seed minimum size | 500 |\n| Smoothing sigma | 0.6 |\n| 2D hole filling | Enabled (max area 32) |\n| Min component size | 200 |\n\n**How it works:**\n1. Voxels above the high threshold (0.55) become seed cores for each surface instance\n2. Seeds smaller than 500 voxels are removed to prevent over-splitting\n3. The probability map is lightly smoothed (sigma=0.6) to reduce micro-basins\n4. Watershed flooding grows each seed into the low threshold mask (0.25)\n5. Boundaries between watershed regions create 1-voxel gaps separating instances\n6. Small holes within instances are filled (2D, max area 32)\n7. Components smaller than 200 voxels are removed\n\n### Step 4: Skeletonization and Re-dilation\n\nAfter watershed separation, we applied 2D slice-wise skeletonization to reduce each surface to its medial axis, then re-dilated by 2 iterations as a safety buffer. This ensured predictions were thin and centered on the true surface location while maintaining sufficient coverage.\n\n### Step 5: Output\n\nFinal output: 320x320x320 uint8 binary TIF masks, packaged as `submission.zip`.\n\n---\n\n## What Didn't Work (But We Tried)\n\nA significant portion of competition time was spent on approaches that ultimately did not improve over our final ensemble:\n\n| Approach | Result |\n|---|---|\n| **Otsu-thresholded inputs** (Dataset511) | Preprocessing inputs with Otsu thresholding dropped the score from 0.587 to 0.562. Critical intensity information was lost. |\n| **Cleaned ground truth** (Dataset512) | Cleaning the training GT by filling holes and closing gaps (3D binary closing with cube structuring element) produced similar training metrics but did not clearly improve the hidden test score. |\n| **Frangi surface filter** | 3D/2D Frangi filters (multiply, add, blend, repair modes) -- all disabled in the best submission. |\n| **CRF post-processing** | Dense CRF smoothing added no benefit. |\n| **GraphCut** | Memory intensive and no improvement. |\n| **PCA bridge-cut refinement** | Attempted 0 cuts on test data -- no applicable cases found. |\n| **Binary morphological closing** | CLOSE_ITERS=0 in best config. |\n| **Line pattern filter** | Bit-pattern convolution for line detection -- disabled. |\n| **Marching Ants recursive splitting** | Complex ray-casting + Dijkstra approach for splitting merged surfaces -- overkill for this data. |\n| **Spatial shift correction** | Shifting predictions by +/-1 voxel to compensate for potential GT alignment issues -- no help. |\n| **2D medial surface approach** (host_all) | XY-plane medial surface detection scored 0.556 solo, well below the 3D approach. |\n| **Model soups / weight averaging** | Averaging weights across checkpoints or models did not improve predictions. |\n\n---\n\n## The Rescore Effect\n\nA pivotal moment in the competition was the mid-competition ground truth rescore. The competition hosts fixed topology defects in the test set ground truth -- filling cavities, closing tunnels, and performing manual topology corrections. This dramatically reshuffled the leaderboard.\n\nModels that had learned to reproduce the noisy original GT were penalized, while models producing clean, topologically sound surfaces were rewarded. Our skeleton-based loss naturally produced cleaner surfaces -- the skeleton recall component encouraged the model to preserve surface topology rather than match noisy voxel labels exactly.\n\n---\n\n## Public vs Private Leaderboard\n\nAn interesting dynamic played out between the public and private leaderboards:\n\n| Configuration | Public LB | Private LB |\n|---|---|---|\n| basic1500 solo | 0.587 | 0.596 |\n| skel_all + aggressive_all | 0.587 | 0.601 |\n| **skel_all + aggressive_all + basic1500** (selected) | -- | **0.598** |\n\nOn the public leaderboard, basic1500 solo and the 2-model ensemble (skel_all + aggressive_all) were tied at 0.587. With no way to distinguish them publicly, we selected the 3-model ensemble as our final submission. In hindsight, the 2-model ensemble of skel_all + aggressive_all would have scored 0.601 on the private leaderboard -- our best result -- but we had no way of knowing that at the time.\n\nThe lesson: **public leaderboard ties hide real differences.** Two submissions that look identical on a small public test set can diverge meaningfully on the full private set. When in doubt, diversity wins.\n\n---\n\n## Key Takeaways\n\n1. **Loss function design matters more than architecture.** The custom skeleton-aware loss was the core innovation -- it directly addressed the thin-surface detection challenge in a way that standard Dice + CE could not.\n\n2. **Ensemble diversity through loss variation.** Rather than ensembling different architectures, we ensembled the same architecture trained with different loss weightings. The moderate and aggressive configurations made complementary errors that cancelled out when averaged.\n\n3. **Don't trust the public leaderboard.** basic1500 solo looked unbeatable on the public LB, but the 3-model ensemble proved its worth on the private test set. Diversity matters for generalization.\n\n4. **TTA is free performance.** 8-way mirror augmentation at test time consistently added +0.01-0.02 with no architecture changes.\n\n5. **Post-processing should be surgical.** Watershed separation for preventing sheet merges was valuable; everything else (Frangi, CRF, GraphCut, etc.) added complexity without improving scores.\n\n6. **Don't throw away information.** Otsu-thresholding the inputs seemed like a reasonable preprocessing step but destroyed subtle intensity gradients the model relied on.\n\n7. **Train longer on all data.** basic1500 (1500 epochs, fold_all) was the strongest individual model. More epochs and more data -- no magic, just patience.\n\n---\n\n## Final Summary\n\n| | |\n|---|---|\n| **Models** | 3x nnU-Net v2 ResEncUNetM |\n| **Ensemble** | skel_all + aggressive_all + basic1500 |\n| **Loss** | SkeletonRecallPlusDiceLoss (varied weights) |\n| **Data** | 737 volumes, fold_all |\n| **Epochs** | 1000 / 1000 / 1500 |\n| **Post-processing** | Logit averaging + Watershed + Skeletonize + Dilate(2) |\n| **Inference** | 2x T4 GPU, TTA (8 mirrors) |\n| **Private LB Score** | **0.598** |\n| **Final Placement** | **49th** |\n",
      "votes": 10
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3417101": "# Vesuvius Challenge - Surface Detection: 49th Place Solution\n### *Three skeleton-aware nnU-Nets and the power of loss function diversity*\n\n## Acknowledgements\n\nThank you to the Kaggle team and the Vesuvius Challenge team for all of their hard work and dedication to this challenge. It was a roller coaster -- multiple rescores and dataset updates kept everyone on their toes throughout the competition. The effort that went into curating the data, fixing ground truth labels, and maintaining the evaluation infrastructure was enormous, and we're grateful for the opportunity to contribute to this incredible mission.\n\nA special thank you to my teammate **Sunny** -- this solution wouldn't have been possible without your incredible work and dedication throughout the competition. From brainstorming ideas to grinding through late-night experiments, your contributions were invaluable. It was a privilege to tackle this challenge together.\n\n---\n\n## Challenge Overview\n\nThe Vesuvius Challenge - Surface Detection competition tasked participants with detecting the thin papyrus surfaces within 3D micro-CT volumes of ancient Herculaneum scroll fragments. Each volume is approximately 320x320x320 voxels at isotropic 1.0mm spacing. The surfaces appear as faint, sheet-like structures embedded in the scan, and the goal is to produce a binary segmentation mask identifying these surfaces.\n\nThe competition scoring combined three metrics:\n\n```\nScore = 0.30 x TopoScore + 0.35 x SurfaceDice@2.0 + 0.35 x VOI_score\n```\n\nThis scoring formula heavily rewarded topologically clean predictions -- not just voxel-level accuracy, but the structural integrity of the predicted surfaces.\n\n**Final Result: Score 0.598, 49th place on the private leaderboard.**\n\n---\n\n## Approach: Three Skeleton-Aware nnU-Nets\n\nOur final submission was a **3-model ensemble** of nnU-Net v2 Residual Encoder U-Nets, all sharing the same architecture but trained with different skeleton-aware loss configurations. The ensemble combined models with complementary strengths -- a conservative model that avoided false positives, an aggressive model that maximized surface recall, and a longer-trained model that balanced both.\n\n| Model | Trainer | Skeleton Recall Weight | FP Penalty Weight | Epochs | EMA Pseudo Dice | Public LB (solo) |\n|---|---|---|---|---|---|---|\n| **skel_all** | SkeletonBased | 0.5 | 0.3 | 1000 | 0.5708 | 0.560 |\n| **aggressive_all** | SkeletonAggressive | 0.7 | 0.2 | 1000 | 0.5527 | 0.563 |\n| **basic1500** | SkeletonBased1500 | 0.5 | 0.3 | 1500 | 0.6089 | 0.587 |\n\nAll three models were trained on all 737 volumes (`fold_all`) from Dataset504_VesuviusNewData3Fold. The ensemble logits were averaged before post-processing, and the combined prediction scored **0.598 on the private leaderboard** -- outperforming any individual model.\n\n### Why the Ensemble Worked\n\nOn the public leaderboard, basic1500 solo (0.587) appeared to beat every 2-model ensemble we tried (~0.582). This led us to believe ensembling was hurting. But the private leaderboard told a different story -- the 3-model ensemble generalized better to unseen data.\n\nThe key was **loss function diversity**. Each model learned a slightly different trade-off between surface recall and precision:\n- **skel_all** (moderate, 1000 epochs): conservative baseline, fewer false positives\n- **aggressive_all** (aggressive, 1000 epochs): pushed harder on recall, caught surfaces the others missed\n- **basic1500** (moderate, 1500 epochs): the anchor -- longer training gave it the best individual performance\n\nAveraging their logits smoothed out each model's failure modes. Where aggressive_all over-predicted, skel_all and basic1500 pulled it back. Where the conservative models missed faint surfaces, aggressive_all filled in the gaps.\n\n---\n\n## Model Architecture\n\nAll three models shared the same **ResEncUNetM** (Residual Encoder U-Net Medium) architecture from nnU-Net v2, planned by `nnUNetPlannerResEncM`:\n\n| Parameter | Value |\n|---|---|\n| Network | `ResidualEncoderUNet` (fully 3D) |\n| Stages | 6 |\n| Features per stage | [32, 64, 128, 256, 320, 320] |\n| Convolution | `Conv3d`, 3x3x3 kernels throughout |\n| Strides | [1,1,1] then [2,2,2] for stages 2-6 |\n| Encoder blocks per stage | [1, 3, 4, 6, 6, 6] |\n| Normalization | InstanceNorm3d |\n| Activation | LeakyReLU (inplace) |\n| Deep supervision | 5 scales, weights [0.533, 0.267, 0.133, 0.067, 0.0] |\n\n---\n\n## Custom Loss Function: SkeletonRecallPlusDiceLoss\n\nStandard Dice + CE loss struggles with thin surface structures -- the extreme class imbalance (surfaces are just a few voxels thick in a 320^3 volume) means the model can achieve high Dice by simply being conservative. We designed a custom three-component loss to address this:\n\n```\nTotal Loss = w_dice * (Dice + CE) + w_skel * SkeletonRecall + w_fp * FalsePositivePenalty\n```\n\n**Component 1: Base Dice + Cross-Entropy (weight 1.0)**\nStandard nnU-Net segmentation loss on the original mask. Provides the foundation for voxel-accurate prediction.\n\n**Component 2: Skeleton Recall Loss (weight 0.5 or 0.7)**\nOn-the-fly, we compute the morphological skeleton of the ground truth surface mask during training via a custom `SkeletonTransform`. The skeleton recall loss then penalizes the model for missing these skeleton voxels:\n\n```python\nskeleton_recall = sum(pred * skeleton) / sum(skeleton)\nloss = 1.0 - skeleton_recall\n```\n\nThis forces the model to maintain the thin medial axis of each surface, preventing it from \"eroding away\" thin structures even when that would improve the Dice score.\n\n**Component 3: False Positive Penalty (weight 0.3 or 0.2)**\nPenalizes predictions in confirmed background regions by computing the mean foreground probability where the ground truth is background:\n\n```python\nfp_loss = sum(pred_fg * background_mask) / sum(background_mask)\n```\n\nThis prevents the model from over-predicting to boost skeleton recall, keeping surfaces clean and preventing merges between adjacent sheets.\n\n### The Two Loss Configurations\n\n| Config | Skeleton Recall | FP Penalty | Character |\n|---|---|---|---|\n| **Moderate** (skel_all, basic1500) | 0.5 | 0.3 | Conservative -- cleaner surfaces, fewer false positives |\n| **Aggressive** (aggressive_all) | 0.7 | 0.2 | Recall-focused -- catches more surfaces, noisier output |\n\nThe moderate configuration produced individually stronger models, but the aggressive variant contributed valuable diversity to the ensemble.\n\n---\n\n## Training Data & Augmentation\n\n**Dataset**: Dataset504_VesuviusNewData3Fold — 737 3D micro-CT volumes of ancient Herculaneum scroll fragments with binary surface labels (0=background, 1=surface, 2=ignore). The training ground truth was known to contain topology defects (holes, tunnels).\n\n### Custom Augmentation: Occlusion3DTransform\n\nBeyond standard nnU-Net 3D augmentations (rotation, scaling, mirroring, gamma, noise), we added a custom **Occlusion3DTransform**:\n\n- Randomly places 1-6 cuboid occlusions (size 2-8 voxels per dimension) over the input volume\n- Applied with 30% probability per sample\n- Occlusion regions are zeroed out, forcing the model to learn to predict surfaces even with missing data\n- This improved robustness to the noisy, artifact-prone micro-CT inputs\n\n### Training Configuration\n\nAll models shared these settings:\n\n| Parameter | Value |\n|---|---|\n| Dataset | Dataset504_VesuviusNewData3Fold |\n| Training volumes | 737 |\n| Fold | `fold_all` (all data, no holdout) |\n| Initial learning rate | 0.01 |\n| LR schedule | Polynomial decay |\n| Patch size | 128 x 128 x 128 |\n| Batch size | 2 |\n| torch.compile | Enabled |\n| Labels | 0=background, 1=surface, 2=ignore |\n| Plans | nnUNetResEncUNetMPlans |\n| Training hardware | Single GPU (local, WSL2) |\n\nThe only differences were epoch count (1000 for skel_all and aggressive_all, 1500 for basic1500) and loss weights.\n\nTraining on all 737 volumes (`fold_all`) was a deliberate choice -- with limited data and a high-capacity 3D model, we maximized the data available for learning.\n\n### Training Commands\n\n```bash\n# skel_all (moderate, 1000 epochs)\nnnUNetv2_train 504 3d_fullres fold_all \\\n    -tr nnUNetTrainerSkeletonBased \\\n    -p nnUNetResEncUNetMPlans\n\n# aggressive_all (aggressive, 1000 epochs)\nnnUNetv2_train 504 3d_fullres fold_all \\\n    -tr nnUNetTrainerSkeletonAggressive \\\n    -p nnUNetResEncUNetMPlans\n\n# basic1500 (moderate, 1500 epochs)\nnnUNetv2_train 504 3d_fullres fold_all \\\n    -tr nnUNetTrainerSkeletonBased1500 \\\n    -p nnUNetResEncUNetMPlans\n```\n\n---\n\n## Inference Pipeline\n\nInference was run on **Kaggle notebooks with 2x Tesla T4 GPUs** within the 9-hour time limit.\n\n### Step 1: Ensemble Inference with TTA\n\nEach of the three models ran nnU-Net's built-in sliding window predictor with **Test Time Augmentation (TTA)** -- 8 mirrored forward passes per volume (all axis-aligned flips). TTA was critical, providing a consistent +0.01-0.02 score improvement per model.\n\nModels were distributed across both T4 GPUs for parallel inference, achieving ~1.5x speedup (~274s vs 411s per volume on a single GPU).\n\n### Step 2: Weighted Logit Averaging\n\nThe softmax logits from all three models were combined using **weighted averaging** before thresholding:\n\n| Model | Weight |\n|---|---|\n| skel_all | 0.5 |\n| aggressive_all | 0.5 |\n| basic1500 | 1.0 |\n\nbasic1500 received double the weight of the other two models, reflecting its significantly stronger individual performance (0.587 vs 0.560/0.563 solo). The two supporting models each contributed half-weight, adding diversity without diluting the anchor. This is where the ensemble's balance paid off -- the aggressive model's higher recall was tempered by basic1500's dominance in the average, while skel_all and aggressive_all together still contributed enough signal to improve generalization on the private test set.\n\n### Step 3: Watershed Instance Separation\n\nThe key post-processing step was watershed-based instance separation, which prevents adjacent papyrus sheets from merging in the prediction:\n\n| Parameter | Value |\n|---|---|\n| High threshold (seed cores) | 0.55 |\n| Low threshold (growth mask) | 0.25 |\n| Boundary gap | 1 |\n| Connectivity | 26 |\n| Seed minimum size | 500 |\n| Smoothing sigma | 0.6 |\n| 2D hole filling | Enabled (max area 32) |\n| Min component size | 200 |\n\n**How it works:**\n1. Voxels above the high threshold (0.55) become seed cores for each surface instance\n2. Seeds smaller than 500 voxels are removed to prevent over-splitting\n3. The probability map is lightly smoothed (sigma=0.6) to reduce micro-basins\n4. Watershed flooding grows each seed into the low threshold mask (0.25)\n5. Boundaries between watershed regions create 1-voxel gaps separating instances\n6. Small holes within instances are filled (2D, max area 32)\n7. Components smaller than 200 voxels are removed\n\n### Step 4: Skeletonization and Re-dilation\n\nAfter watershed separation, we applied 2D slice-wise skeletonization to reduce each surface to its medial axis, then re-dilated by 2 iterations as a safety buffer. This ensured predictions were thin and centered on the true surface location while maintaining sufficient coverage.\n\n### Step 5: Output\n\nFinal output: 320x320x320 uint8 binary TIF masks, packaged as `submission.zip`.\n\n---\n\n## What Didn't Work (But We Tried)\n\nA significant portion of competition time was spent on approaches that ultimately did not improve over our final ensemble:\n\n| Approach | Result |\n|---|---|\n| **Otsu-thresholded inputs** (Dataset511) | Preprocessing inputs with Otsu thresholding dropped the score from 0.587 to 0.562. Critical intensity information was lost. |\n| **Cleaned ground truth** (Dataset512) | Cleaning the training GT by filling holes and closing gaps (3D binary closing with cube structuring element) produced similar training metrics but did not clearly improve the hidden test score. |\n| **Frangi surface filter** | 3D/2D Frangi filters (multiply, add, blend, repair modes) -- all disabled in the best submission. |\n| **CRF post-processing** | Dense CRF smoothing added no benefit. |\n| **GraphCut** | Memory intensive and no improvement. |\n| **PCA bridge-cut refinement** | Attempted 0 cuts on test data -- no applicable cases found. |\n| **Binary morphological closing** | CLOSE_ITERS=0 in best config. |\n| **Line pattern filter** | Bit-pattern convolution for line detection -- disabled. |\n| **Marching Ants recursive splitting** | Complex ray-casting + Dijkstra approach for splitting merged surfaces -- overkill for this data. |\n| **Spatial shift correction** | Shifting predictions by +/-1 voxel to compensate for potential GT alignment issues -- no help. |\n| **2D medial surface approach** (host_all) | XY-plane medial surface detection scored 0.556 solo, well below the 3D approach. |\n| **Model soups / weight averaging** | Averaging weights across checkpoints or models did not improve predictions. |\n\n---\n\n## The Rescore Effect\n\nA pivotal moment in the competition was the mid-competition ground truth rescore. The competition hosts fixed topology defects in the test set ground truth -- filling cavities, closing tunnels, and performing manual topology corrections. This dramatically reshuffled the leaderboard.\n\nModels that had learned to reproduce the noisy original GT were penalized, while models producing clean, topologically sound surfaces were rewarded. Our skeleton-based loss naturally produced cleaner surfaces -- the skeleton recall component encouraged the model to preserve surface topology rather than match noisy voxel labels exactly.\n\n---\n\n## Public vs Private Leaderboard\n\nAn interesting dynamic played out between the public and private leaderboards:\n\n| Configuration | Public LB | Private LB |\n|---|---|---|\n| basic1500 solo | 0.587 | 0.596 |\n| skel_all + aggressive_all | 0.587 | 0.601 |\n| **skel_all + aggressive_all + basic1500** (selected) | -- | **0.598** |\n\nOn the public leaderboard, basic1500 solo and the 2-model ensemble (skel_all + aggressive_all) were tied at 0.587. With no way to distinguish them publicly, we selected the 3-model ensemble as our final submission. In hindsight, the 2-model ensemble of skel_all + aggressive_all would have scored 0.601 on the private leaderboard -- our best result -- but we had no way of knowing that at the time.\n\nThe lesson: **public leaderboard ties hide real differences.** Two submissions that look identical on a small public test set can diverge meaningfully on the full private set. When in doubt, diversity wins.\n\n---\n\n## Key Takeaways\n\n1. **Loss function design matters more than architecture.** The custom skeleton-aware loss was the core innovation -- it directly addressed the thin-surface detection challenge in a way that standard Dice + CE could not.\n\n2. **Ensemble diversity through loss variation.** Rather than ensembling different architectures, we ensembled the same architecture trained with different loss weightings. The moderate and aggressive configurations made complementary errors that cancelled out when averaged.\n\n3. **Don't trust the public leaderboard.** basic1500 solo looked unbeatable on the public LB, but the 3-model ensemble proved its worth on the private test set. Diversity matters for generalization.\n\n4. **TTA is free performance.** 8-way mirror augmentation at test time consistently added +0.01-0.02 with no architecture changes.\n\n5. **Post-processing should be surgical.** Watershed separation for preventing sheet merges was valuable; everything else (Frangi, CRF, GraphCut, etc.) added complexity without improving scores.\n\n6. **Don't throw away information.** Otsu-thresholding the inputs seemed like a reasonable preprocessing step but destroyed subtle intensity gradients the model relied on.\n\n7. **Train longer on all data.** basic1500 (1500 epochs, fold_all) was the strongest individual model. More epochs and more data -- no magic, just patience.\n\n---\n\n## Final Summary\n\n| | |\n|---|---|\n| **Models** | 3x nnU-Net v2 ResEncUNetM |\n| **Ensemble** | skel_all + aggressive_all + basic1500 |\n| **Loss** | SkeletonRecallPlusDiceLoss (varied weights) |\n| **Data** | 737 volumes, fold_all |\n| **Epochs** | 1000 / 1000 / 1500 |\n| **Post-processing** | Logit averaging + Watershed + Skeletonize + Dilate(2) |\n| **Inference** | 2x T4 GPU, TTA (8 mirrors) |\n| **Private LB Score** | **0.598** |\n| **Final Placement** | **49th** |\n"
  }
}