{
  "id": 679575,
  "title": "From Cell Segmentation to Ancient Scrolls: WNet3D + nnUNet (LB 113, 0.582)",
  "url": "/competitions/vesuvius-challenge-surface-detection/discussion/679575",
  "author_name": "pr4deepr",
  "post_date": "2026-03-02T10:51:44.721000",
  "votes": 2,
  "comment_count": 0,
  "views": 0,
  "content": "<p>My day job is bioimage analysis — building segmentation pipelines for microscopy data. When I saw this competition's task (detecting papyrus surfaces in 3D CT scans), it looked a lot like what I do at work: find structures in noisy volumetric data. So I borrowed a tool from that world. </p>\n<p>Best submission: <strong>0.582 private LB, 113th place</strong></p>\n<h2>Pipeline</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4210346%2Fa083b64e1f0a236f4847a75decccc4c4%2FBiological%20Context%20Feature-2026-07-30-155424.png?generation=1785426902397067&amp;alt=media\" alt=\"\"></p>\n<p><a href=\"https://kaggle.com/code/pr4deepr/wnet3d-nnunet-ensemble-125th-place\" target=\"_blank\"><strong><em>&gt;NOTEBOOK LINK</em></strong></a> </p>\n<p>Postprocessing was deliberately simple: hysteresis thresholding (T_low 0.30, T_high 0.90), z-closing (radius 1), dust removal (min_size 100) — applied to the averaged probability maps. It beat every tuned alternative I tried, including an Optuna-optimised version and the winning team's pipeline.</p>\n<h2>The WNet Idea</h2>\n<ul>\n<li>Raw CT intensity alone is a limited input — subtle contrast between sheet and background, multiple overlapping layers</li>\n<li><a href=\"https://elifesciences.org/articles/99848\" target=\"_blank\">WNet3D</a> is a dual 3D U-Net trained without labels via a self-supervised reconstruction objective</li>\n<li>Originally designed for cell segmentation in cleared neural tissue (light-sheet microscopy)</li>\n<li>We repurposed it as a learned feature extractor for scroll CT. WNet3D was configured to predict 3 output classes for larger structures with the hope that it would capture ink, background and a third class focusing on edges, ink or other structural features.</li>\n<li>One channel focused on ink, whereas other channels focused on background, specifically big and small gaps between sheets.</li>\n<li>All 3 feature maps were concatenated with raw CT → 4-channel input for nnUNet</li>\n<li>Observed faster nnUNet training convergence with WNet input vs raw CT (preliminary tests only — no controlled LB ablation)</li>\n</ul>\n<p>In figures above, you can see that WNet3D (100 epoch) predicts:</p>\n<ul>\n<li>1st channel: small gaps between sheets</li>\n<li>2nd channel: Sheets with bright spots which I believe are ink?</li>\n<li>3rd channel: Large gaps between sheet</li>\n</ul>\n<h2>What Worked</h2>\n<ul>\n<li><strong>WNet preprocessing</strong> — faster convergence during nnUNet training. All submitted models used WNet-enriched input.</li>\n<li><strong>GT cleaning</strong> — manually corrected misalignments/noise in ground truth (Dataset325). Solo: 0.575. But critical ingredient in best ensemble (0.582).</li>\n<li><strong>Sheet separation</strong> — Sheet merging didn't seem to be a big problem (Visual QC), possibly due to WNet3D channels focusing on big and small gaps.</li>\n<li><strong>Half the data is enough</strong> — nnUNet models trained on 50% of training data scored 0.564–0.566. Full data: 0.579. Ensembling two half-data models: 0.576. The training set has substantial redundancy.</li>\n<li><strong>Simple postprocessing</strong> — hysteresis thresholding (T_low=0.30, T_high=0.90), z-closing (radius=1), dust removal (min_size=100). Beat every \"optimised\" alternative.</li>\n<li><strong>Ensemble diversity via data</strong> — best result (0.582) came from ensembling models with different preprocessor checkpoints (WNet20 vs WNet100) AND different GT versions. Half-data ensembles (321+322) also showed clear gains (+0.010 over best single). Models sharing the same preprocessor but differing only in GT gave less benefit.</li>\n</ul>\n<h2>What Didn't Work</h2>\n<ul>\n<li><strong>Optuna postprocessing</strong> — Used Optuna to find best postprocessing parameters, but overfit to local validation. 0.577 vs 0.579 for hand-picked params on same model.</li>\n<li><strong>Fragmentation problem</strong> Fragmentation appeared to be the main issue and postprocessing needs to be optimized for this.</li>\n<li><strong><a href=\"https://www.kaggle.com/models/ipythonx/vsd-model/Keras/transunet/4\" target=\"_blank\">TransUNet</a> ensemble</strong> — Tried TransUNet model by <a href=\"https://www.kaggle.com/ipythonx\" target=\"_blank\">@ipythonx</a> at 25% weight with nnUNet 323: 0.579, identical to 323 solo. Correlated errors.</li>\n<li><strong>Winning solution postprocessing on our model</strong> — After the competition, I applied the winning postprocessing as-is and scored 0.582, same as our simple postprocessing. Their pipeline was co-optimised with their 4-model ensemble. Postprocessing is model-specific.</li>\n<li><strong>Public LB for model selection</strong> — best public (0.563) and best private (0.582) came from different submissions. Would have picked the wrong model.</li>\n<li><strong>Local metrics</strong> — best local topometric scores consistently corresponded to worse LB results. Train/test distribution mismatch.</li>\n</ul>\n<h2>Key Takeaways</h2>\n<ol>\n<li><strong>Self-supervised preprocessing is underexplored</strong> — WNet from microscopy transferred to scroll CT. No other team tried learned preprocessing.</li>\n<li><strong>Data quality &gt; model complexity</strong> — GT cleaning gave comparable gains to architectural changes.</li>\n<li><strong>You don't need all the data</strong> — 50% data used with nnUNet got within 0.013 of full-data model. (Caveat: WNet was trained on all data.)</li>\n<li><strong>Simple postprocessing beats tuned parameters</strong> — train/test distribution gap made local optimisation unreliable.</li>\n<li><strong>Postprocessing is model-specific</strong> — winning team's PP didn't help our model.</li>\n</ol>\n<h2>Future Directions</h2>\n<p>In this competition I tuned WNet to find sheet-like and larger structures, but with more classes and careful tuning of the SoftNCuts loss parameters (intensity sigma, spatial sigma, radius), I believe WNet could learn to separate ink from sheet from background. The ink channel could then be fed into nnUNet as an additional input, giving the segmentation model explicit ink-awareness — potentially helping in regions where ink and sheet are hard to distinguish.</p>\n<p>Other ideas I didn't have time to explore:</p>\n<ul>\n<li><strong>Controlled WNet ablation</strong> - train identical nnUNet models with and without WNet features to properly quantify the preprocessing contribution on the leaderboard, not just training convergence</li>\n<li><strong>Multi-scale WNet features</strong> - the winning solution gained a lot from multiple patch sizes. Extracting WNet features at different scales and stacking them could capture both fine ink texture and coarse sheet structure</li>\n<li><strong>WNet-guided postprocessing</strong> —-use WNet's class maps to inform topology repair, e.g. only fill holes in regions WNet classifies as sheet, or preserve gaps where WNet detects ink. It seems to detect large and small gaps between sheets</li>\n<li><strong>More WNet output classes</strong> — going beyond 3 classes to see if finer structural categories emerge (sheet interior, sheet boundary, ink, voids, noise)</li>\n</ul>\n<p>If anyone is interested in exploring the WNet preprocessing direction, happy to chat. <a href=\"https://www.kaggle.com/seanjohnsonsp\" target=\"_blank\">@seanjohnsonsp</a> <a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> </p>\n<h2>All Results</h2>\n<p>Here's every submission I made during the competition, roughly in order. The story is one of diminishing returns from model tweaks and hard-won lessons about what actually moves the needle.</p>\n<table>\n<thead>\n<tr>\n<th>Submission</th>\n<th>Description</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Dataset321</td>\n<td>nnUNet on first half of training data, WNet20 preprocessor</td>\n<td>0.553</td>\n<td>0.566</td>\n</tr>\n<tr>\n<td>Dataset322</td>\n<td>nnUNet on second half of training data, WNet20 preprocessor</td>\n<td>0.549</td>\n<td>0.564</td>\n</tr>\n<tr>\n<td>321+322 ensemble</td>\n<td>50/50 average of the two half-data models</td>\n<td>0.556</td>\n<td>0.576</td>\n</tr>\n<tr>\n<td>Dataset323</td>\n<td>nnUNet on full data, WNet100 preprocessor, DC+CE loss</td>\n<td>0.563</td>\n<td>0.579</td>\n</tr>\n<tr>\n<td>Dataset324</td>\n<td>nnUNet on full data, WNet20 preprocessor, DC+CE loss</td>\n<td>—</td>\n<td>0.579</td>\n</tr>\n<tr>\n<td>Dataset325</td>\n<td>nnUNet on full data, WNet100 preprocessor, BoundaryDoULoss, cleaned GT</td>\n<td>0.553</td>\n<td>0.575</td>\n</tr>\n<tr>\n<td>323 + TransUNet</td>\n<td>75/25 nnUNet/TransUNet ensemble</td>\n<td>0.558</td>\n<td>0.579</td>\n</tr>\n<tr>\n<td>323 + Optuna PP</td>\n<td>Optuna-optimised postprocessing on best solo model</td>\n<td>—</td>\n<td>0.577</td>\n</tr>\n<tr>\n<td><strong>324+325 ensemble</strong></td>\n<td><strong>0.5/0.5 ensemble, different WNet checkpoints + different GT</strong></td>\n<td><strong>0.557</strong></td>\n<td><strong>0.582</strong></td>\n</tr>\n<tr>\n<td>325 + winning PP</td>\n<td>Dataset325 with 1st place postprocessing (post-competition test)</td>\n<td>—</td>\n<td>0.582</td>\n</tr>\n</tbody>\n</table>\n<p>A few things stand out from this table. The half-data models (321, 322) are surprisingly competitive, only ~0.013 behind the full-data model (323), which suggests the training set has a lot of redundancy. The best ensemble (324+325) only edges out the best solo by 0.003, but it required genuine diversity: different WNet checkpoints (20 vs 100 epochs) and different ground truth versions. Models that were too similar (e.g. 323+TransUNet) just averaged their shared mistakes.</p>\n<p>The public/private LB gap is also notable. The best public score (0.563, Dataset323) and best private score (0.582, 324+325 ensemble) came from different submissions entirely. If I'd been selecting on public LB alone, I'd have picked the wrong model.</p>\n<h2>Comparison to Winning Solution (0.627)</h2>\n<p>It's worth looking at where the gap came from:</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>Ours (113th)</th>\n<th>1st Place</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Backbone</td>\n<td>nnUNet ResEncUNetXL</td>\n<td>nnUNet (4 models)</td>\n</tr>\n<tr>\n<td>Patch sizes</td>\n<td>Single</td>\n<td>128, 160, 192, 224, 256</td>\n</tr>\n<tr>\n<td>Preprocessing</td>\n<td>WNet3D (self-supervised)</td>\n<td>Raw CT</td>\n</tr>\n<tr>\n<td>Postprocessing</td>\n<td>Simple hysteresis</td>\n<td>Height-map patching, LUT hole repair, oriented closing</td>\n</tr>\n<tr>\n<td>GT handling</td>\n<td>Manual cleaning</td>\n<td>Standard</td>\n</tr>\n<tr>\n<td>Private LB</td>\n<td>0.582</td>\n<td>0.627</td>\n</tr>\n</tbody>\n</table>\n<p>The ~0.045 gap is substantial, and it came primarily from their topology-aware postprocessing and multi-scale ensemble diversity (five different patch sizes). Both solutions validated nnUNet as the right backbone we invested in preprocessing, they invested in postprocessing and model diversity. Crucially, when I applied their postprocessing to our model post-competition, it scored 0.582, which is identical to our simple approach. Their postprocessing was co-optimised with their specific ensemble's error profile and didn't transfer.</p>\n<h2>Technical Details</h2>\n<p><strong>nnUNet configuration:</strong> ResEncUNetXL backbone, 3d_fullres, trained with the standard nnUNet pipeline. Dataset324 used the default Dice + Cross-Entropy loss. Dataset325 used <a href=\"https://arxiv.org/pdf/2308.00220.pdf\" target=\"_blank\">BoundaryDoULoss3D</a>, adapted from <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/igor-krashenyi-4th-place-solution-boundary-dou-los\" target=\"_blank\">Igor Krashenyi's 4th place solution</a> in the SenNet+HOA blood vessel segmentation competition (<a href=\"https://github.com/sunfan-bvb/BoundaryDoULoss\" target=\"_blank\">original repo</a>). This boundary-aware loss was designed to improve segmentation at tissue boundaries — a natural fit for sheet detection.</p>\n<p><strong>WNet3D:</strong> From <a href=\"https://elifesciences.org/articles/99848\" target=\"_blank\">CellSeg3D</a> (Achard et al., eLife 2025). Trained two checkpoints — 20 epochs (patch 64) and 100 epochs (patch 128) — on all raw CT training data in a fully self-supervised way (no labels required). The 20 epoch was my first attempt, and then when I got access to more compute, I trained one with larger patch size for more epochs.  The encoder's intermediate feature maps (3 channels) were concatenated with the raw CT to form a 4-channel input for nnUNet. Using two different WNet checkpoints for the two nnUNet models was key to ensemble diversity.</p>\n<p><strong>Inference:</strong> Sliding window with TTA (mirror axes 0, 1), tile_step_size=0.5. Initially used 1xP100 GPU, but I've learned to run both models run in parallel on Kaggle's 2×T4 GPUs. </p>\n<p><strong>Postprocessing:</strong> Hysteresis thresholding (T_low=0.30, T_high=0.90) to get connected high-confidence regions and grow them into lower-confidence voxels. Morphological z-closing (radius=1) to fill single-slice gaps. Dust removal (min connected component size=100) to clean up small fragments. Simple, conservative, and it generalised better than anything tuned.</p>\n<h2>Reflections</h2>\n<p>This was my first Kaggle competition and I learned a huge deal. My background in bioimage analysis turned out to be more transferable than I expected, i.e., volumetric segmentation is volumetric segmentation whether it's cells or scrolls, and tools like nnUNet and WNet3D crossed domains surprisingly well. At the same time, I learned just as much from reading other people's discussions and solutions. The community here is incredibly generous with sharing ideas, and I picked up techniques I'd never have found on my own.</p>\n<p>Next time I'd love to work in a team. It seems like a great way to connect with people, learn faster, and push further than you can solo. I'd also aim to contribute more to the discussion forums during the competition rather than just lurking, it would be a great way to pay forward what others shared with me + can get input from many.</p>\n<p>Coming from biological and microscopy data, I'm especially appreciative of the effort that went into curating this dataset. Good ground truth is expensive and time-consuming to produce, and it's what makes competitions like this possible. Thanks to the organisers for putting it all together.</p>",
  "messages": [
    {
      "id": 3416229,
      "postDate": "2026-03-02T10:51:44.720Z",
      "content": "<p>My day job is bioimage analysis — building segmentation pipelines for microscopy data. When I saw this competition's task (detecting papyrus surfaces in 3D CT scans), it looked a lot like what I do at work: find structures in noisy volumetric data. So I borrowed a tool from that world. </p>\n<p>Best submission: <strong>0.582 private LB, 113th place</strong></p>\n<h2>Pipeline</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4210346%2Fa083b64e1f0a236f4847a75decccc4c4%2FBiological%20Context%20Feature-2026-07-30-155424.png?generation=1785426902397067&amp;alt=media\" alt=\"\"></p>\n<p><a href=\"https://kaggle.com/code/pr4deepr/wnet3d-nnunet-ensemble-125th-place\" target=\"_blank\"><strong><em>&gt;NOTEBOOK LINK</em></strong></a> </p>\n<p>Postprocessing was deliberately simple: hysteresis thresholding (T_low 0.30, T_high 0.90), z-closing (radius 1), dust removal (min_size 100) — applied to the averaged probability maps. It beat every tuned alternative I tried, including an Optuna-optimised version and the winning team's pipeline.</p>\n<h2>The WNet Idea</h2>\n<ul>\n<li>Raw CT intensity alone is a limited input — subtle contrast between sheet and background, multiple overlapping layers</li>\n<li><a href=\"https://elifesciences.org/articles/99848\" target=\"_blank\">WNet3D</a> is a dual 3D U-Net trained without labels via a self-supervised reconstruction objective</li>\n<li>Originally designed for cell segmentation in cleared neural tissue (light-sheet microscopy)</li>\n<li>We repurposed it as a learned feature extractor for scroll CT. WNet3D was configured to predict 3 output classes for larger structures with the hope that it would capture ink, background and a third class focusing on edges, ink or other structural features.</li>\n<li>One channel focused on ink, whereas other channels focused on background, specifically big and small gaps between sheets.</li>\n<li>All 3 feature maps were concatenated with raw CT → 4-channel input for nnUNet</li>\n<li>Observed faster nnUNet training convergence with WNet input vs raw CT (preliminary tests only — no controlled LB ablation)</li>\n</ul>\n<p>In figures above, you can see that WNet3D (100 epoch) predicts:</p>\n<ul>\n<li>1st channel: small gaps between sheets</li>\n<li>2nd channel: Sheets with bright spots which I believe are ink?</li>\n<li>3rd channel: Large gaps between sheet</li>\n</ul>\n<h2>What Worked</h2>\n<ul>\n<li><strong>WNet preprocessing</strong> — faster convergence during nnUNet training. All submitted models used WNet-enriched input.</li>\n<li><strong>GT cleaning</strong> — manually corrected misalignments/noise in ground truth (Dataset325). Solo: 0.575. But critical ingredient in best ensemble (0.582).</li>\n<li><strong>Sheet separation</strong> — Sheet merging didn't seem to be a big problem (Visual QC), possibly due to WNet3D channels focusing on big and small gaps.</li>\n<li><strong>Half the data is enough</strong> — nnUNet models trained on 50% of training data scored 0.564–0.566. Full data: 0.579. Ensembling two half-data models: 0.576. The training set has substantial redundancy.</li>\n<li><strong>Simple postprocessing</strong> — hysteresis thresholding (T_low=0.30, T_high=0.90), z-closing (radius=1), dust removal (min_size=100). Beat every \"optimised\" alternative.</li>\n<li><strong>Ensemble diversity via data</strong> — best result (0.582) came from ensembling models with different preprocessor checkpoints (WNet20 vs WNet100) AND different GT versions. Half-data ensembles (321+322) also showed clear gains (+0.010 over best single). Models sharing the same preprocessor but differing only in GT gave less benefit.</li>\n</ul>\n<h2>What Didn't Work</h2>\n<ul>\n<li><strong>Optuna postprocessing</strong> — Used Optuna to find best postprocessing parameters, but overfit to local validation. 0.577 vs 0.579 for hand-picked params on same model.</li>\n<li><strong>Fragmentation problem</strong> Fragmentation appeared to be the main issue and postprocessing needs to be optimized for this.</li>\n<li><strong><a href=\"https://www.kaggle.com/models/ipythonx/vsd-model/Keras/transunet/4\" target=\"_blank\">TransUNet</a> ensemble</strong> — Tried TransUNet model by <a href=\"https://www.kaggle.com/ipythonx\" target=\"_blank\">@ipythonx</a> at 25% weight with nnUNet 323: 0.579, identical to 323 solo. Correlated errors.</li>\n<li><strong>Winning solution postprocessing on our model</strong> — After the competition, I applied the winning postprocessing as-is and scored 0.582, same as our simple postprocessing. Their pipeline was co-optimised with their 4-model ensemble. Postprocessing is model-specific.</li>\n<li><strong>Public LB for model selection</strong> — best public (0.563) and best private (0.582) came from different submissions. Would have picked the wrong model.</li>\n<li><strong>Local metrics</strong> — best local topometric scores consistently corresponded to worse LB results. Train/test distribution mismatch.</li>\n</ul>\n<h2>Key Takeaways</h2>\n<ol>\n<li><strong>Self-supervised preprocessing is underexplored</strong> — WNet from microscopy transferred to scroll CT. No other team tried learned preprocessing.</li>\n<li><strong>Data quality &gt; model complexity</strong> — GT cleaning gave comparable gains to architectural changes.</li>\n<li><strong>You don't need all the data</strong> — 50% data used with nnUNet got within 0.013 of full-data model. (Caveat: WNet was trained on all data.)</li>\n<li><strong>Simple postprocessing beats tuned parameters</strong> — train/test distribution gap made local optimisation unreliable.</li>\n<li><strong>Postprocessing is model-specific</strong> — winning team's PP didn't help our model.</li>\n</ol>\n<h2>Future Directions</h2>\n<p>In this competition I tuned WNet to find sheet-like and larger structures, but with more classes and careful tuning of the SoftNCuts loss parameters (intensity sigma, spatial sigma, radius), I believe WNet could learn to separate ink from sheet from background. The ink channel could then be fed into nnUNet as an additional input, giving the segmentation model explicit ink-awareness — potentially helping in regions where ink and sheet are hard to distinguish.</p>\n<p>Other ideas I didn't have time to explore:</p>\n<ul>\n<li><strong>Controlled WNet ablation</strong> - train identical nnUNet models with and without WNet features to properly quantify the preprocessing contribution on the leaderboard, not just training convergence</li>\n<li><strong>Multi-scale WNet features</strong> - the winning solution gained a lot from multiple patch sizes. Extracting WNet features at different scales and stacking them could capture both fine ink texture and coarse sheet structure</li>\n<li><strong>WNet-guided postprocessing</strong> —-use WNet's class maps to inform topology repair, e.g. only fill holes in regions WNet classifies as sheet, or preserve gaps where WNet detects ink. It seems to detect large and small gaps between sheets</li>\n<li><strong>More WNet output classes</strong> — going beyond 3 classes to see if finer structural categories emerge (sheet interior, sheet boundary, ink, voids, noise)</li>\n</ul>\n<p>If anyone is interested in exploring the WNet preprocessing direction, happy to chat. <a href=\"https://www.kaggle.com/seanjohnsonsp\" target=\"_blank\">@seanjohnsonsp</a> <a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> </p>\n<h2>All Results</h2>\n<p>Here's every submission I made during the competition, roughly in order. The story is one of diminishing returns from model tweaks and hard-won lessons about what actually moves the needle.</p>\n<table>\n<thead>\n<tr>\n<th>Submission</th>\n<th>Description</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Dataset321</td>\n<td>nnUNet on first half of training data, WNet20 preprocessor</td>\n<td>0.553</td>\n<td>0.566</td>\n</tr>\n<tr>\n<td>Dataset322</td>\n<td>nnUNet on second half of training data, WNet20 preprocessor</td>\n<td>0.549</td>\n<td>0.564</td>\n</tr>\n<tr>\n<td>321+322 ensemble</td>\n<td>50/50 average of the two half-data models</td>\n<td>0.556</td>\n<td>0.576</td>\n</tr>\n<tr>\n<td>Dataset323</td>\n<td>nnUNet on full data, WNet100 preprocessor, DC+CE loss</td>\n<td>0.563</td>\n<td>0.579</td>\n</tr>\n<tr>\n<td>Dataset324</td>\n<td>nnUNet on full data, WNet20 preprocessor, DC+CE loss</td>\n<td>—</td>\n<td>0.579</td>\n</tr>\n<tr>\n<td>Dataset325</td>\n<td>nnUNet on full data, WNet100 preprocessor, BoundaryDoULoss, cleaned GT</td>\n<td>0.553</td>\n<td>0.575</td>\n</tr>\n<tr>\n<td>323 + TransUNet</td>\n<td>75/25 nnUNet/TransUNet ensemble</td>\n<td>0.558</td>\n<td>0.579</td>\n</tr>\n<tr>\n<td>323 + Optuna PP</td>\n<td>Optuna-optimised postprocessing on best solo model</td>\n<td>—</td>\n<td>0.577</td>\n</tr>\n<tr>\n<td><strong>324+325 ensemble</strong></td>\n<td><strong>0.5/0.5 ensemble, different WNet checkpoints + different GT</strong></td>\n<td><strong>0.557</strong></td>\n<td><strong>0.582</strong></td>\n</tr>\n<tr>\n<td>325 + winning PP</td>\n<td>Dataset325 with 1st place postprocessing (post-competition test)</td>\n<td>—</td>\n<td>0.582</td>\n</tr>\n</tbody>\n</table>\n<p>A few things stand out from this table. The half-data models (321, 322) are surprisingly competitive, only ~0.013 behind the full-data model (323), which suggests the training set has a lot of redundancy. The best ensemble (324+325) only edges out the best solo by 0.003, but it required genuine diversity: different WNet checkpoints (20 vs 100 epochs) and different ground truth versions. Models that were too similar (e.g. 323+TransUNet) just averaged their shared mistakes.</p>\n<p>The public/private LB gap is also notable. The best public score (0.563, Dataset323) and best private score (0.582, 324+325 ensemble) came from different submissions entirely. If I'd been selecting on public LB alone, I'd have picked the wrong model.</p>\n<h2>Comparison to Winning Solution (0.627)</h2>\n<p>It's worth looking at where the gap came from:</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>Ours (113th)</th>\n<th>1st Place</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Backbone</td>\n<td>nnUNet ResEncUNetXL</td>\n<td>nnUNet (4 models)</td>\n</tr>\n<tr>\n<td>Patch sizes</td>\n<td>Single</td>\n<td>128, 160, 192, 224, 256</td>\n</tr>\n<tr>\n<td>Preprocessing</td>\n<td>WNet3D (self-supervised)</td>\n<td>Raw CT</td>\n</tr>\n<tr>\n<td>Postprocessing</td>\n<td>Simple hysteresis</td>\n<td>Height-map patching, LUT hole repair, oriented closing</td>\n</tr>\n<tr>\n<td>GT handling</td>\n<td>Manual cleaning</td>\n<td>Standard</td>\n</tr>\n<tr>\n<td>Private LB</td>\n<td>0.582</td>\n<td>0.627</td>\n</tr>\n</tbody>\n</table>\n<p>The ~0.045 gap is substantial, and it came primarily from their topology-aware postprocessing and multi-scale ensemble diversity (five different patch sizes). Both solutions validated nnUNet as the right backbone we invested in preprocessing, they invested in postprocessing and model diversity. Crucially, when I applied their postprocessing to our model post-competition, it scored 0.582, which is identical to our simple approach. Their postprocessing was co-optimised with their specific ensemble's error profile and didn't transfer.</p>\n<h2>Technical Details</h2>\n<p><strong>nnUNet configuration:</strong> ResEncUNetXL backbone, 3d_fullres, trained with the standard nnUNet pipeline. Dataset324 used the default Dice + Cross-Entropy loss. Dataset325 used <a href=\"https://arxiv.org/pdf/2308.00220.pdf\" target=\"_blank\">BoundaryDoULoss3D</a>, adapted from <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/igor-krashenyi-4th-place-solution-boundary-dou-los\" target=\"_blank\">Igor Krashenyi's 4th place solution</a> in the SenNet+HOA blood vessel segmentation competition (<a href=\"https://github.com/sunfan-bvb/BoundaryDoULoss\" target=\"_blank\">original repo</a>). This boundary-aware loss was designed to improve segmentation at tissue boundaries — a natural fit for sheet detection.</p>\n<p><strong>WNet3D:</strong> From <a href=\"https://elifesciences.org/articles/99848\" target=\"_blank\">CellSeg3D</a> (Achard et al., eLife 2025). Trained two checkpoints — 20 epochs (patch 64) and 100 epochs (patch 128) — on all raw CT training data in a fully self-supervised way (no labels required). The 20 epoch was my first attempt, and then when I got access to more compute, I trained one with larger patch size for more epochs.  The encoder's intermediate feature maps (3 channels) were concatenated with the raw CT to form a 4-channel input for nnUNet. Using two different WNet checkpoints for the two nnUNet models was key to ensemble diversity.</p>\n<p><strong>Inference:</strong> Sliding window with TTA (mirror axes 0, 1), tile_step_size=0.5. Initially used 1xP100 GPU, but I've learned to run both models run in parallel on Kaggle's 2×T4 GPUs. </p>\n<p><strong>Postprocessing:</strong> Hysteresis thresholding (T_low=0.30, T_high=0.90) to get connected high-confidence regions and grow them into lower-confidence voxels. Morphological z-closing (radius=1) to fill single-slice gaps. Dust removal (min connected component size=100) to clean up small fragments. Simple, conservative, and it generalised better than anything tuned.</p>\n<h2>Reflections</h2>\n<p>This was my first Kaggle competition and I learned a huge deal. My background in bioimage analysis turned out to be more transferable than I expected, i.e., volumetric segmentation is volumetric segmentation whether it's cells or scrolls, and tools like nnUNet and WNet3D crossed domains surprisingly well. At the same time, I learned just as much from reading other people's discussions and solutions. The community here is incredibly generous with sharing ideas, and I picked up techniques I'd never have found on my own.</p>\n<p>Next time I'd love to work in a team. It seems like a great way to connect with people, learn faster, and push further than you can solo. I'd also aim to contribute more to the discussion forums during the competition rather than just lurking, it would be a great way to pay forward what others shared with me + can get input from many.</p>\n<p>Coming from biological and microscopy data, I'm especially appreciative of the effort that went into curating this dataset. Good ground truth is expensive and time-consuming to produce, and it's what makes competitions like this possible. Thanks to the organisers for putting it all together.</p>",
      "rawMarkdown": "My day job is bioimage analysis — building segmentation pipelines for microscopy data. When I saw this competition's task (detecting papyrus surfaces in 3D CT scans), it looked a lot like what I do at work: find structures in noisy volumetric data. So I borrowed a tool from that world. \n\nBest submission: **0.582 private LB, 113th place**\n\n## Pipeline\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4210346%2Fa083b64e1f0a236f4847a75decccc4c4%2FBiological%20Context%20Feature-2026-07-30-155424.png?generation=1785426902397067&alt=media)\n\n[***>NOTEBOOK LINK***](https://kaggle.com/code/pr4deepr/wnet3d-nnunet-ensemble-125th-place) \n\nPostprocessing was deliberately simple: hysteresis thresholding (T_low 0.30, T_high 0.90), z-closing (radius 1), dust removal (min_size 100) — applied to the averaged probability maps. It beat every tuned alternative I tried, including an Optuna-optimised version and the winning team's pipeline.\n\n## The WNet Idea\n\n- Raw CT intensity alone is a limited input — subtle contrast between sheet and background, multiple overlapping layers\n- [WNet3D](https://elifesciences.org/articles/99848) is a dual 3D U-Net trained without labels via a self-supervised reconstruction objective\n- Originally designed for cell segmentation in cleared neural tissue (light-sheet microscopy)\n- We repurposed it as a learned feature extractor for scroll CT. WNet3D was configured to predict 3 output classes for larger structures with the hope that it would capture ink, background and a third class focusing on edges, ink or other structural features.\n- One channel focused on ink, whereas other channels focused on background, specifically big and small gaps between sheets.\n- All 3 feature maps were concatenated with raw CT → 4-channel input for nnUNet\n- Observed faster nnUNet training convergence with WNet input vs raw CT (preliminary tests only — no controlled LB ablation)\n\nIn figures above, you can see that WNet3D (100 epoch) predicts:\n- 1st channel: small gaps between sheets\n- 2nd channel: Sheets with bright spots which I believe are ink?\n- 3rd channel: Large gaps between sheet\n\n## What Worked\n\n- **WNet preprocessing** — faster convergence during nnUNet training. All submitted models used WNet-enriched input.\n- **GT cleaning** — manually corrected misalignments/noise in ground truth (Dataset325). Solo: 0.575. But critical ingredient in best ensemble (0.582).\n- **Sheet separation** — Sheet merging didn't seem to be a big problem (Visual QC), possibly due to WNet3D channels focusing on big and small gaps.\n- **Half the data is enough** — nnUNet models trained on 50% of training data scored 0.564–0.566. Full data: 0.579. Ensembling two half-data models: 0.576. The training set has substantial redundancy.\n- **Simple postprocessing** — hysteresis thresholding (T_low=0.30, T_high=0.90), z-closing (radius=1), dust removal (min_size=100). Beat every \"optimised\" alternative.\n- **Ensemble diversity via data** — best result (0.582) came from ensembling models with different preprocessor checkpoints (WNet20 vs WNet100) AND different GT versions. Half-data ensembles (321+322) also showed clear gains (+0.010 over best single). Models sharing the same preprocessor but differing only in GT gave less benefit.\n\n\n## What Didn't Work\n\n- **Optuna postprocessing** — Used Optuna to find best postprocessing parameters, but overfit to local validation. 0.577 vs 0.579 for hand-picked params on same model.\n- **Fragmentation problem** Fragmentation appeared to be the main issue and postprocessing needs to be optimized for this.\n- **[TransUNet](https://www.kaggle.com/models/ipythonx/vsd-model/Keras/transunet/4) ensemble** — Tried TransUNet model by @ipythonx at 25% weight with nnUNet 323: 0.579, identical to 323 solo. Correlated errors.\n- **Winning solution postprocessing on our model** — After the competition, I applied the winning postprocessing as-is and scored 0.582, same as our simple postprocessing. Their pipeline was co-optimised with their 4-model ensemble. Postprocessing is model-specific.\n- **Public LB for model selection** — best public (0.563) and best private (0.582) came from different submissions. Would have picked the wrong model.\n- **Local metrics** — best local topometric scores consistently corresponded to worse LB results. Train/test distribution mismatch.\n\n## Key Takeaways\n\n1. **Self-supervised preprocessing is underexplored** — WNet from microscopy transferred to scroll CT. No other team tried learned preprocessing.\n2. **Data quality > model complexity** — GT cleaning gave comparable gains to architectural changes.\n3. **You don't need all the data** — 50% data used with nnUNet got within 0.013 of full-data model. (Caveat: WNet was trained on all data.)\n4. **Simple postprocessing beats tuned parameters** — train/test distribution gap made local optimisation unreliable.\n5. **Postprocessing is model-specific** — winning team's PP didn't help our model.\n\n\n## Future Directions\n\n In this competition I tuned WNet to find sheet-like and larger structures, but with more classes and careful tuning of the SoftNCuts loss parameters (intensity sigma, spatial sigma, radius), I believe WNet could learn to separate ink from sheet from background. The ink channel could then be fed into nnUNet as an additional input, giving the segmentation model explicit ink-awareness — potentially helping in regions where ink and sheet are hard to distinguish.\n\nOther ideas I didn't have time to explore:\n\n- **Controlled WNet ablation** - train identical nnUNet models with and without WNet features to properly quantify the preprocessing contribution on the leaderboard, not just training convergence\n- **Multi-scale WNet features** - the winning solution gained a lot from multiple patch sizes. Extracting WNet features at different scales and stacking them could capture both fine ink texture and coarse sheet structure\n- **WNet-guided postprocessing** —-use WNet's class maps to inform topology repair, e.g. only fill holes in regions WNet classifies as sheet, or preserve gaps where WNet detects ink. It seems to detect large and small gaps between sheets\n- **More WNet output classes** — going beyond 3 classes to see if finer structural categories emerge (sheet interior, sheet boundary, ink, voids, noise)\n\nIf anyone is interested in exploring the WNet preprocessing direction, happy to chat. @seanjohnsonsp @giorgioangelotti \n\n\n## All Results\n\nHere's every submission I made during the competition, roughly in order. The story is one of diminishing returns from model tweaks and hard-won lessons about what actually moves the needle.\n\n| Submission | Description | Public LB | Private LB |\n|---|---|---|---|\n| Dataset321 | nnUNet on first half of training data, WNet20 preprocessor | 0.553 | 0.566 |\n| Dataset322 | nnUNet on second half of training data, WNet20 preprocessor | 0.549 | 0.564 |\n| 321+322 ensemble | 50/50 average of the two half-data models | 0.556 | 0.576 |\n| Dataset323 | nnUNet on full data, WNet100 preprocessor, DC+CE loss | 0.563 | 0.579 |\n| Dataset324 | nnUNet on full data, WNet20 preprocessor, DC+CE loss | — | 0.579 |\n| Dataset325 | nnUNet on full data, WNet100 preprocessor, BoundaryDoULoss, cleaned GT | 0.553 | 0.575 |\n| 323 + TransUNet | 75/25 nnUNet/TransUNet ensemble | 0.558 | 0.579 |\n| 323 + Optuna PP | Optuna-optimised postprocessing on best solo model | — | 0.577 |\n| **324+325 ensemble** | **0.5/0.5 ensemble, different WNet checkpoints + different GT** | **0.557** | **0.582** |\n| 325 + winning PP | Dataset325 with 1st place postprocessing (post-competition test) | — | 0.582 |\n\nA few things stand out from this table. The half-data models (321, 322) are surprisingly competitive, only ~0.013 behind the full-data model (323), which suggests the training set has a lot of redundancy. The best ensemble (324+325) only edges out the best solo by 0.003, but it required genuine diversity: different WNet checkpoints (20 vs 100 epochs) and different ground truth versions. Models that were too similar (e.g. 323+TransUNet) just averaged their shared mistakes.\n\nThe public/private LB gap is also notable. The best public score (0.563, Dataset323) and best private score (0.582, 324+325 ensemble) came from different submissions entirely. If I'd been selecting on public LB alone, I'd have picked the wrong model.\n\n## Comparison to Winning Solution (0.627)\n\nIt's worth looking at where the gap came from:\n\n| | Ours (113th) | 1st Place |\n|---|---|---|\n| Backbone | nnUNet ResEncUNetXL | nnUNet (4 models) |\n| Patch sizes | Single | 128, 160, 192, 224, 256 |\n| Preprocessing | WNet3D (self-supervised) | Raw CT |\n| Postprocessing | Simple hysteresis | Height-map patching, LUT hole repair, oriented closing |\n| GT handling | Manual cleaning | Standard |\n| Private LB | 0.582 | 0.627 |\n\nThe ~0.045 gap is substantial, and it came primarily from their topology-aware postprocessing and multi-scale ensemble diversity (five different patch sizes). Both solutions validated nnUNet as the right backbone we invested in preprocessing, they invested in postprocessing and model diversity. Crucially, when I applied their postprocessing to our model post-competition, it scored 0.582, which is identical to our simple approach. Their postprocessing was co-optimised with their specific ensemble's error profile and didn't transfer.\n\n## Technical Details\n\n**nnUNet configuration:** ResEncUNetXL backbone, 3d_fullres, trained with the standard nnUNet pipeline. Dataset324 used the default Dice + Cross-Entropy loss. Dataset325 used [BoundaryDoULoss3D](https://arxiv.org/pdf/2308.00220.pdf), adapted from [Igor Krashenyi's 4th place solution](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/igor-krashenyi-4th-place-solution-boundary-dou-los) in the SenNet+HOA blood vessel segmentation competition ([original repo](https://github.com/sunfan-bvb/BoundaryDoULoss)). This boundary-aware loss was designed to improve segmentation at tissue boundaries — a natural fit for sheet detection.\n\n**WNet3D:** From [CellSeg3D](https://elifesciences.org/articles/99848) (Achard et al., eLife 2025). Trained two checkpoints — 20 epochs (patch 64) and 100 epochs (patch 128) — on all raw CT training data in a fully self-supervised way (no labels required). The 20 epoch was my first attempt, and then when I got access to more compute, I trained one with larger patch size for more epochs.  The encoder's intermediate feature maps (3 channels) were concatenated with the raw CT to form a 4-channel input for nnUNet. Using two different WNet checkpoints for the two nnUNet models was key to ensemble diversity.\n\n**Inference:** Sliding window with TTA (mirror axes 0, 1), tile_step_size=0.5. Initially used 1xP100 GPU, but I've learned to run both models run in parallel on Kaggle's 2×T4 GPUs. \n\n**Postprocessing:** Hysteresis thresholding (T_low=0.30, T_high=0.90) to get connected high-confidence regions and grow them into lower-confidence voxels. Morphological z-closing (radius=1) to fill single-slice gaps. Dust removal (min connected component size=100) to clean up small fragments. Simple, conservative, and it generalised better than anything tuned.\n\n\n## Reflections\n\nThis was my first Kaggle competition and I learned a huge deal. My background in bioimage analysis turned out to be more transferable than I expected, i.e., volumetric segmentation is volumetric segmentation whether it's cells or scrolls, and tools like nnUNet and WNet3D crossed domains surprisingly well. At the same time, I learned just as much from reading other people's discussions and solutions. The community here is incredibly generous with sharing ideas, and I picked up techniques I'd never have found on my own.\n\nNext time I'd love to work in a team. It seems like a great way to connect with people, learn faster, and push further than you can solo. I'd also aim to contribute more to the discussion forums during the competition rather than just lurking, it would be a great way to pay forward what others shared with me + can get input from many.\n\nComing from biological and microscopy data, I'm especially appreciative of the effort that went into curating this dataset. Good ground truth is expensive and time-consuming to produce, and it's what makes competitions like this possible. Thanks to the organisers for putting it all together.\n",
      "votes": 2
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3416229": "My day job is bioimage analysis — building segmentation pipelines for microscopy data. When I saw this competition's task (detecting papyrus surfaces in 3D CT scans), it looked a lot like what I do at work: find structures in noisy volumetric data. So I borrowed a tool from that world. \n\nBest submission: **0.582 private LB, 113th place**\n\n## Pipeline\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4210346%2Fa083b64e1f0a236f4847a75decccc4c4%2FBiological%20Context%20Feature-2026-07-30-155424.png?generation=1785426902397067&alt=media)\n\n[***>NOTEBOOK LINK***](https://kaggle.com/code/pr4deepr/wnet3d-nnunet-ensemble-125th-place) \n\nPostprocessing was deliberately simple: hysteresis thresholding (T_low 0.30, T_high 0.90), z-closing (radius 1), dust removal (min_size 100) — applied to the averaged probability maps. It beat every tuned alternative I tried, including an Optuna-optimised version and the winning team's pipeline.\n\n## The WNet Idea\n\n- Raw CT intensity alone is a limited input — subtle contrast between sheet and background, multiple overlapping layers\n- [WNet3D](https://elifesciences.org/articles/99848) is a dual 3D U-Net trained without labels via a self-supervised reconstruction objective\n- Originally designed for cell segmentation in cleared neural tissue (light-sheet microscopy)\n- We repurposed it as a learned feature extractor for scroll CT. WNet3D was configured to predict 3 output classes for larger structures with the hope that it would capture ink, background and a third class focusing on edges, ink or other structural features.\n- One channel focused on ink, whereas other channels focused on background, specifically big and small gaps between sheets.\n- All 3 feature maps were concatenated with raw CT → 4-channel input for nnUNet\n- Observed faster nnUNet training convergence with WNet input vs raw CT (preliminary tests only — no controlled LB ablation)\n\nIn figures above, you can see that WNet3D (100 epoch) predicts:\n- 1st channel: small gaps between sheets\n- 2nd channel: Sheets with bright spots which I believe are ink?\n- 3rd channel: Large gaps between sheet\n\n## What Worked\n\n- **WNet preprocessing** — faster convergence during nnUNet training. All submitted models used WNet-enriched input.\n- **GT cleaning** — manually corrected misalignments/noise in ground truth (Dataset325). Solo: 0.575. But critical ingredient in best ensemble (0.582).\n- **Sheet separation** — Sheet merging didn't seem to be a big problem (Visual QC), possibly due to WNet3D channels focusing on big and small gaps.\n- **Half the data is enough** — nnUNet models trained on 50% of training data scored 0.564–0.566. Full data: 0.579. Ensembling two half-data models: 0.576. The training set has substantial redundancy.\n- **Simple postprocessing** — hysteresis thresholding (T_low=0.30, T_high=0.90), z-closing (radius=1), dust removal (min_size=100). Beat every \"optimised\" alternative.\n- **Ensemble diversity via data** — best result (0.582) came from ensembling models with different preprocessor checkpoints (WNet20 vs WNet100) AND different GT versions. Half-data ensembles (321+322) also showed clear gains (+0.010 over best single). Models sharing the same preprocessor but differing only in GT gave less benefit.\n\n\n## What Didn't Work\n\n- **Optuna postprocessing** — Used Optuna to find best postprocessing parameters, but overfit to local validation. 0.577 vs 0.579 for hand-picked params on same model.\n- **Fragmentation problem** Fragmentation appeared to be the main issue and postprocessing needs to be optimized for this.\n- **[TransUNet](https://www.kaggle.com/models/ipythonx/vsd-model/Keras/transunet/4) ensemble** — Tried TransUNet model by @ipythonx at 25% weight with nnUNet 323: 0.579, identical to 323 solo. Correlated errors.\n- **Winning solution postprocessing on our model** — After the competition, I applied the winning postprocessing as-is and scored 0.582, same as our simple postprocessing. Their pipeline was co-optimised with their 4-model ensemble. Postprocessing is model-specific.\n- **Public LB for model selection** — best public (0.563) and best private (0.582) came from different submissions. Would have picked the wrong model.\n- **Local metrics** — best local topometric scores consistently corresponded to worse LB results. Train/test distribution mismatch.\n\n## Key Takeaways\n\n1. **Self-supervised preprocessing is underexplored** — WNet from microscopy transferred to scroll CT. No other team tried learned preprocessing.\n2. **Data quality > model complexity** — GT cleaning gave comparable gains to architectural changes.\n3. **You don't need all the data** — 50% data used with nnUNet got within 0.013 of full-data model. (Caveat: WNet was trained on all data.)\n4. **Simple postprocessing beats tuned parameters** — train/test distribution gap made local optimisation unreliable.\n5. **Postprocessing is model-specific** — winning team's PP didn't help our model.\n\n\n## Future Directions\n\n In this competition I tuned WNet to find sheet-like and larger structures, but with more classes and careful tuning of the SoftNCuts loss parameters (intensity sigma, spatial sigma, radius), I believe WNet could learn to separate ink from sheet from background. The ink channel could then be fed into nnUNet as an additional input, giving the segmentation model explicit ink-awareness — potentially helping in regions where ink and sheet are hard to distinguish.\n\nOther ideas I didn't have time to explore:\n\n- **Controlled WNet ablation** - train identical nnUNet models with and without WNet features to properly quantify the preprocessing contribution on the leaderboard, not just training convergence\n- **Multi-scale WNet features** - the winning solution gained a lot from multiple patch sizes. Extracting WNet features at different scales and stacking them could capture both fine ink texture and coarse sheet structure\n- **WNet-guided postprocessing** —-use WNet's class maps to inform topology repair, e.g. only fill holes in regions WNet classifies as sheet, or preserve gaps where WNet detects ink. It seems to detect large and small gaps between sheets\n- **More WNet output classes** — going beyond 3 classes to see if finer structural categories emerge (sheet interior, sheet boundary, ink, voids, noise)\n\nIf anyone is interested in exploring the WNet preprocessing direction, happy to chat. @seanjohnsonsp @giorgioangelotti \n\n\n## All Results\n\nHere's every submission I made during the competition, roughly in order. The story is one of diminishing returns from model tweaks and hard-won lessons about what actually moves the needle.\n\n| Submission | Description | Public LB | Private LB |\n|---|---|---|---|\n| Dataset321 | nnUNet on first half of training data, WNet20 preprocessor | 0.553 | 0.566 |\n| Dataset322 | nnUNet on second half of training data, WNet20 preprocessor | 0.549 | 0.564 |\n| 321+322 ensemble | 50/50 average of the two half-data models | 0.556 | 0.576 |\n| Dataset323 | nnUNet on full data, WNet100 preprocessor, DC+CE loss | 0.563 | 0.579 |\n| Dataset324 | nnUNet on full data, WNet20 preprocessor, DC+CE loss | — | 0.579 |\n| Dataset325 | nnUNet on full data, WNet100 preprocessor, BoundaryDoULoss, cleaned GT | 0.553 | 0.575 |\n| 323 + TransUNet | 75/25 nnUNet/TransUNet ensemble | 0.558 | 0.579 |\n| 323 + Optuna PP | Optuna-optimised postprocessing on best solo model | — | 0.577 |\n| **324+325 ensemble** | **0.5/0.5 ensemble, different WNet checkpoints + different GT** | **0.557** | **0.582** |\n| 325 + winning PP | Dataset325 with 1st place postprocessing (post-competition test) | — | 0.582 |\n\nA few things stand out from this table. The half-data models (321, 322) are surprisingly competitive, only ~0.013 behind the full-data model (323), which suggests the training set has a lot of redundancy. The best ensemble (324+325) only edges out the best solo by 0.003, but it required genuine diversity: different WNet checkpoints (20 vs 100 epochs) and different ground truth versions. Models that were too similar (e.g. 323+TransUNet) just averaged their shared mistakes.\n\nThe public/private LB gap is also notable. The best public score (0.563, Dataset323) and best private score (0.582, 324+325 ensemble) came from different submissions entirely. If I'd been selecting on public LB alone, I'd have picked the wrong model.\n\n## Comparison to Winning Solution (0.627)\n\nIt's worth looking at where the gap came from:\n\n| | Ours (113th) | 1st Place |\n|---|---|---|\n| Backbone | nnUNet ResEncUNetXL | nnUNet (4 models) |\n| Patch sizes | Single | 128, 160, 192, 224, 256 |\n| Preprocessing | WNet3D (self-supervised) | Raw CT |\n| Postprocessing | Simple hysteresis | Height-map patching, LUT hole repair, oriented closing |\n| GT handling | Manual cleaning | Standard |\n| Private LB | 0.582 | 0.627 |\n\nThe ~0.045 gap is substantial, and it came primarily from their topology-aware postprocessing and multi-scale ensemble diversity (five different patch sizes). Both solutions validated nnUNet as the right backbone we invested in preprocessing, they invested in postprocessing and model diversity. Crucially, when I applied their postprocessing to our model post-competition, it scored 0.582, which is identical to our simple approach. Their postprocessing was co-optimised with their specific ensemble's error profile and didn't transfer.\n\n## Technical Details\n\n**nnUNet configuration:** ResEncUNetXL backbone, 3d_fullres, trained with the standard nnUNet pipeline. Dataset324 used the default Dice + Cross-Entropy loss. Dataset325 used [BoundaryDoULoss3D](https://arxiv.org/pdf/2308.00220.pdf), adapted from [Igor Krashenyi's 4th place solution](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/igor-krashenyi-4th-place-solution-boundary-dou-los) in the SenNet+HOA blood vessel segmentation competition ([original repo](https://github.com/sunfan-bvb/BoundaryDoULoss)). This boundary-aware loss was designed to improve segmentation at tissue boundaries — a natural fit for sheet detection.\n\n**WNet3D:** From [CellSeg3D](https://elifesciences.org/articles/99848) (Achard et al., eLife 2025). Trained two checkpoints — 20 epochs (patch 64) and 100 epochs (patch 128) — on all raw CT training data in a fully self-supervised way (no labels required). The 20 epoch was my first attempt, and then when I got access to more compute, I trained one with larger patch size for more epochs.  The encoder's intermediate feature maps (3 channels) were concatenated with the raw CT to form a 4-channel input for nnUNet. Using two different WNet checkpoints for the two nnUNet models was key to ensemble diversity.\n\n**Inference:** Sliding window with TTA (mirror axes 0, 1), tile_step_size=0.5. Initially used 1xP100 GPU, but I've learned to run both models run in parallel on Kaggle's 2×T4 GPUs. \n\n**Postprocessing:** Hysteresis thresholding (T_low=0.30, T_high=0.90) to get connected high-confidence regions and grow them into lower-confidence voxels. Morphological z-closing (radius=1) to fill single-slice gaps. Dust removal (min connected component size=100) to clean up small fragments. Simple, conservative, and it generalised better than anything tuned.\n\n\n## Reflections\n\nThis was my first Kaggle competition and I learned a huge deal. My background in bioimage analysis turned out to be more transferable than I expected, i.e., volumetric segmentation is volumetric segmentation whether it's cells or scrolls, and tools like nnUNet and WNet3D crossed domains surprisingly well. At the same time, I learned just as much from reading other people's discussions and solutions. The community here is incredibly generous with sharing ideas, and I picked up techniques I'd never have found on my own.\n\nNext time I'd love to work in a team. It seems like a great way to connect with people, learn faster, and push further than you can solo. I'd also aim to contribute more to the discussion forums during the competition rather than just lurking, it would be a great way to pay forward what others shared with me + can get input from many.\n\nComing from biological and microscopy data, I'm especially appreciative of the effort that went into curating this dataset. Good ground truth is expensive and time-consuming to produce, and it's what makes competitions like this possible. Thanks to the organisers for putting it all together.\n"
  }
}