{
  "id": 679251,
  "title": "7th Place Solution for the Vesuvius Challenge",
  "url": "/competitions/vesuvius-challenge-surface-detection/writeups/7st-place-solution-for-the-vesuvius-challenge",
  "author_name": "",
  "post_date": "2026-02-28T06:03:43.340Z",
  "votes": 16,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Thanks to the organizers for putting together such a cool competition — learned a ton and had a lot of fun digging into volumetric topology.</p>\n<p>Congrats to all the winners and participants! Here's a quick summary of what I did.</p>\n<h2>Overview</h2>\n<p>7th place, Public LB 0.618 / Private LB 0.591. Like many others, nnU-Net carried the day. Simple pipeline: two nnU-Net models ensembled + TTA + post-processing.</p>\n<h2>Model Training</h2>\n<p>Both models are based on <strong>nnUNet</strong> with <code>nnUNetPlannerResEncM</code>, <code>3d_lowres</code> configuration.</p>\n<ul>\n<li>Patch size: <strong>128³</strong></li>\n<li>Epochs: <strong>2000</strong></li>\n<li>Optimizer: <strong>SGD</strong> (nnUNet default: lr=0.01, momentum=0.99, nesterov=True, wd=3e-5)</li>\n<li>Loss: <strong>CE Loss + Dice Loss + Skeleton Recall Loss</strong>, all weighted equally (1.0 each)</li>\n</ul>\n<p>The key difference between my two models:</p>\n<p><strong>Model A — Preprocessed training labels (key insight):</strong> I noticed that the organizers' GitHub repo has a specific processing pipeline for the <em>test set</em> data. I applied the <strong>exact same processing to the training labels</strong> before training. This was probably the single most impactful thing I did — the raw training labels don't quite match the test-time processing, and this mismatch hurts generalization. Aligning train/test label processing gave a clear and consistent boost.</p>\n<p><strong>Model B — Raw training labels:</strong> Standard training without any label preprocessing. This gives the ensemble some diversity.</p>\n<h2>Ensemble &amp; Inference</h2>\n<ul>\n<li>Simple ensemble of Model A + Model B</li>\n<li><strong>TTA</strong> (nnUNet's built-in test-time augmentation with mirroring)</li>\n</ul>\n<p>Nothing fancy here, just let nnUNet do its thing.</p>\n<h2>Post-Processing</h2>\n<p>A four-stage post-processing pipeline applied on binary predictions from nnU-Net to improve topological correctness, surface accuracy, and instance consistency.</p>\n<p><strong>Ridge Detection (Hessian-based Frangi Filter):</strong> The 3D Frangi sheetness filter is applied directly on the binarized prediction volume. It computes the Hessian matrix at each voxel via second-order Gaussian derivatives, then analyzes the sorted eigenvalues (|λ1| ≤ |λ2| ≤ |λ3|) to compute a sheetness response. Voxels forming planar/sheet-like structures satisfy |λ1| ≈ 0, |λ2| ≈ 0, |λ3| &gt;&gt; 0, which corresponds to the papyrus surface geometry. The response combines a background suppression term, a planarity term, and a blob rejection term. Only bright-on-dark structures (λ3 &lt; 0) are retained. The Frangi output is thresholded to produce a cleaned binary mask, effectively enhancing continuous surface regions while suppressing isolated noise points and non-planar false positives.</p>\n<p><strong>Coherence Enhancing Diffusion (CED):</strong> After ridge detection, Coherence Enhancing Diffusion is applied slice-by-slice (2D) on the Frangi-filtered volume to enhance spatial continuity along the papyrus surface. CED is a PDE-based anisotropic diffusion method governed by ∂u/∂t = div(D(Jρ)·∇u), where the diffusion tensor D is derived from the structure tensor Jρ. The structure tensor is computed using Pavel Holoborodko derivative kernels with Gaussian smoothing. The diffusion tensor diffuses strongly along the dominant structure orientation (tangent to the papyrus surface) while preserving edges in the perpendicular direction, implemented with GPU acceleration via PyTorch. The result is re-binarized. This step fills small holes within surfaces, reconnects fragmented segments, and smooths boundary noise while maintaining edge sharpness.</p>\n<p><strong>Anisotropic Morphological Closing:</strong> A 3D binary closing operation (dilation followed by erosion) is applied with an anisotropic ellipsoidal structuring element. The Z-radius is set larger than the XY-radius, accounting for the fact that papyrus layers may have small gaps along the depth axis due to scanning artifacts. This fills remaining small holes and bridges tiny discontinuities that survived the CED step, without over-connecting adjacent layers in the XY plane.</p>\n<p><strong>Dust Removal (Small Object Pruning):</strong> Small connected components below a voxel-count threshold are removed from the final binary mask. This eliminates residual noise fragments and transient false positives that do not form meaningful papyrus structures, reducing spurious connected components.</p>\n<h3>Pipeline Summary</h3>\n<pre><code>Binary Prediction (nnU-Net) + TTA\n    → Frangi Sheetness Filter (3D)\n    → CED Anisotropic Diffusion (2D slice-wise, GPU)\n    → Anisotropic Closing\n    → Dust Removal\n    → Final Prediction\n</code></pre>\n<h2>Note on Muon Optimizer</h2>\n<p>One thing worth mentioning — I experimented with the <strong>Muon optimizer</strong> as a replacement for SGD. Models trained with Muon consistently performed better on both Public and Private LB compared to SGD.</p>\n<p>However, I didn't have enough compute to fully train all my final models with Muon, so my submitted ensemble still uses SGD-trained models. Training speed with SGD is only about <strong>1.25x faster</strong> than Muon, so it's not a huge tradeoff.</p>\n<p>I've open-sourced my optimized Muon implementation: <strong><a href=\"https://github.com/Decem-Y/FastMuon\" target=\"_blank\">FastMuon</a></strong> — drop-in Turbo-Muon with Triton acceleration. If you have the hardware, worth trying.</p>\n<h2>What Didn't Work / Didn't Try</h2>\n<ul>\n<li>More aggressive data augmentation — tried cranking up augmentation strengths / probabilities but it consistently hurt performance. Ended up using lower-than-default nnUNet augmentation probabilities</li>\n<li>clDice Loss — didn't see improvement over the CE + Dice + Skeleton Recall combo</li>\n<li>Adding SE (Squeeze-and-Excitation) modules into the UNet — no meaningful gain, just extra VRAM usage</li>\n<li>Pseudo labels — tried using model predictions on unlabeled data as pseudo labels for semi-supervised training, but the noise was too much. Ended up degrading performance rather than helping</li>\n<li>More aggressive topological post-processing (PCA hole filling, Betti matching, etc.)</li>\n</ul>\n<hr>\n<p>DECEM</p>",
  "messages": [
    {
      "id": "3415073",
      "postDate": "02/28/2026 06:03:31",
      "content": "<p>Thanks to the organizers for putting together such a cool competition — learned a ton and had a lot of fun digging into volumetric topology.</p>\n<p>Congrats to all the winners and participants! Here's a quick summary of what I did.</p>\n<h2>Overview</h2>\n<p>7th place, Public LB 0.618 / Private LB 0.591. Like many others, nnU-Net carried the day. Simple pipeline: two nnU-Net models ensembled + TTA + post-processing.</p>\n<h2>Model Training</h2>\n<p>Both models are based on <strong>nnUNet</strong> with <code>nnUNetPlannerResEncM</code>, <code>3d_lowres</code> configuration.</p>\n<ul>\n<li>Patch size: <strong>128³</strong></li>\n<li>Epochs: <strong>2000</strong></li>\n<li>Optimizer: <strong>SGD</strong> (nnUNet default: lr=0.01, momentum=0.99, nesterov=True, wd=3e-5)</li>\n<li>Loss: <strong>CE Loss + Dice Loss + Skeleton Recall Loss</strong>, all weighted equally (1.0 each)</li>\n</ul>\n<p>The key difference between my two models:</p>\n<p><strong>Model A — Preprocessed training labels (key insight):</strong> I noticed that the organizers' GitHub repo has a specific processing pipeline for the <em>test set</em> data. I applied the <strong>exact same processing to the training labels</strong> before training. This was probably the single most impactful thing I did — the raw training labels don't quite match the test-time processing, and this mismatch hurts generalization. Aligning train/test label processing gave a clear and consistent boost.</p>\n<p><strong>Model B — Raw training labels:</strong> Standard training without any label preprocessing. This gives the ensemble some diversity.</p>\n<h2>Ensemble &amp; Inference</h2>\n<ul>\n<li>Simple ensemble of Model A + Model B</li>\n<li><strong>TTA</strong> (nnUNet's built-in test-time augmentation with mirroring)</li>\n</ul>\n<p>Nothing fancy here, just let nnUNet do its thing.</p>\n<h2>Post-Processing</h2>\n<p>A four-stage post-processing pipeline applied on binary predictions from nnU-Net to improve topological correctness, surface accuracy, and instance consistency.</p>\n<p><strong>Ridge Detection (Hessian-based Frangi Filter):</strong> The 3D Frangi sheetness filter is applied directly on the binarized prediction volume. It computes the Hessian matrix at each voxel via second-order Gaussian derivatives, then analyzes the sorted eigenvalues (|λ1| ≤ |λ2| ≤ |λ3|) to compute a sheetness response. Voxels forming planar/sheet-like structures satisfy |λ1| ≈ 0, |λ2| ≈ 0, |λ3| &gt;&gt; 0, which corresponds to the papyrus surface geometry. The response combines a background suppression term, a planarity term, and a blob rejection term. Only bright-on-dark structures (λ3 &lt; 0) are retained. The Frangi output is thresholded to produce a cleaned binary mask, effectively enhancing continuous surface regions while suppressing isolated noise points and non-planar false positives.</p>\n<p><strong>Coherence Enhancing Diffusion (CED):</strong> After ridge detection, Coherence Enhancing Diffusion is applied slice-by-slice (2D) on the Frangi-filtered volume to enhance spatial continuity along the papyrus surface. CED is a PDE-based anisotropic diffusion method governed by ∂u/∂t = div(D(Jρ)·∇u), where the diffusion tensor D is derived from the structure tensor Jρ. The structure tensor is computed using Pavel Holoborodko derivative kernels with Gaussian smoothing. The diffusion tensor diffuses strongly along the dominant structure orientation (tangent to the papyrus surface) while preserving edges in the perpendicular direction, implemented with GPU acceleration via PyTorch. The result is re-binarized. This step fills small holes within surfaces, reconnects fragmented segments, and smooths boundary noise while maintaining edge sharpness.</p>\n<p><strong>Anisotropic Morphological Closing:</strong> A 3D binary closing operation (dilation followed by erosion) is applied with an anisotropic ellipsoidal structuring element. The Z-radius is set larger than the XY-radius, accounting for the fact that papyrus layers may have small gaps along the depth axis due to scanning artifacts. This fills remaining small holes and bridges tiny discontinuities that survived the CED step, without over-connecting adjacent layers in the XY plane.</p>\n<p><strong>Dust Removal (Small Object Pruning):</strong> Small connected components below a voxel-count threshold are removed from the final binary mask. This eliminates residual noise fragments and transient false positives that do not form meaningful papyrus structures, reducing spurious connected components.</p>\n<h3>Pipeline Summary</h3>\n<pre><code>Binary Prediction (nnU-Net) + TTA\n    → Frangi Sheetness Filter (3D)\n    → CED Anisotropic Diffusion (2D slice-wise, GPU)\n    → Anisotropic Closing\n    → Dust Removal\n    → Final Prediction\n</code></pre>\n<h2>Note on Muon Optimizer</h2>\n<p>One thing worth mentioning — I experimented with the <strong>Muon optimizer</strong> as a replacement for SGD. Models trained with Muon consistently performed better on both Public and Private LB compared to SGD.</p>\n<p>However, I didn't have enough compute to fully train all my final models with Muon, so my submitted ensemble still uses SGD-trained models. Training speed with SGD is only about <strong>1.25x faster</strong> than Muon, so it's not a huge tradeoff.</p>\n<p>I've open-sourced my optimized Muon implementation: <strong><a href=\"https://github.com/Decem-Y/FastMuon\" target=\"_blank\">FastMuon</a></strong> — drop-in Turbo-Muon with Triton acceleration. If you have the hardware, worth trying.</p>\n<h2>What Didn't Work / Didn't Try</h2>\n<ul>\n<li>More aggressive data augmentation — tried cranking up augmentation strengths / probabilities but it consistently hurt performance. Ended up using lower-than-default nnUNet augmentation probabilities</li>\n<li>clDice Loss — didn't see improvement over the CE + Dice + Skeleton Recall combo</li>\n<li>Adding SE (Squeeze-and-Excitation) modules into the UNet — no meaningful gain, just extra VRAM usage</li>\n<li>Pseudo labels — tried using model predictions on unlabeled data as pseudo labels for semi-supervised training, but the noise was too much. Ended up degrading performance rather than helping</li>\n<li>More aggressive topological post-processing (PCA hole filling, Betti matching, etc.)</li>\n</ul>\n<hr>\n<p>DECEM</p>",
      "rawMarkdown": "Thanks to the organizers for putting together such a cool competition — learned a ton and had a lot of fun digging into volumetric topology.\n\nCongrats to all the winners and participants! Here's a quick summary of what I did.\n\n## Overview\n\n7th place, Public LB 0.618 / Private LB 0.591. Like many others, nnU-Net carried the day. Simple pipeline: two nnU-Net models ensembled + TTA + post-processing.\n\n## Model Training\n\nBoth models are based on **nnUNet** with `nnUNetPlannerResEncM`, `3d_lowres` configuration.\n\n- Patch size: **128³**\n- Epochs: **2000**\n- Optimizer: **SGD** (nnUNet default: lr=0.01, momentum=0.99, nesterov=True, wd=3e-5)\n- Loss: **CE Loss + Dice Loss + Skeleton Recall Loss**, all weighted equally (1.0 each)\n\nThe key difference between my two models:\n\n**Model A — Preprocessed training labels (key insight):** I noticed that the organizers' GitHub repo has a specific processing pipeline for the *test set* data. I applied the **exact same processing to the training labels** before training. This was probably the single most impactful thing I did — the raw training labels don't quite match the test-time processing, and this mismatch hurts generalization. Aligning train/test label processing gave a clear and consistent boost.\n\n**Model B — Raw training labels:** Standard training without any label preprocessing. This gives the ensemble some diversity.\n\n## Ensemble & Inference\n\n- Simple ensemble of Model A + Model B\n- **TTA** (nnUNet's built-in test-time augmentation with mirroring)\n\nNothing fancy here, just let nnUNet do its thing.\n\n## Post-Processing\n\nA four-stage post-processing pipeline applied on binary predictions from nnU-Net to improve topological correctness, surface accuracy, and instance consistency.\n\n**Ridge Detection (Hessian-based Frangi Filter):** The 3D Frangi sheetness filter is applied directly on the binarized prediction volume. It computes the Hessian matrix at each voxel via second-order Gaussian derivatives, then analyzes the sorted eigenvalues (|λ1| ≤ |λ2| ≤ |λ3|) to compute a sheetness response. Voxels forming planar/sheet-like structures satisfy |λ1| ≈ 0, |λ2| ≈ 0, |λ3| >> 0, which corresponds to the papyrus surface geometry. The response combines a background suppression term, a planarity term, and a blob rejection term. Only bright-on-dark structures (λ3 < 0) are retained. The Frangi output is thresholded to produce a cleaned binary mask, effectively enhancing continuous surface regions while suppressing isolated noise points and non-planar false positives.\n\n**Coherence Enhancing Diffusion (CED):** After ridge detection, Coherence Enhancing Diffusion is applied slice-by-slice (2D) on the Frangi-filtered volume to enhance spatial continuity along the papyrus surface. CED is a PDE-based anisotropic diffusion method governed by ∂u/∂t = div(D(Jρ)·∇u), where the diffusion tensor D is derived from the structure tensor Jρ. The structure tensor is computed using Pavel Holoborodko derivative kernels with Gaussian smoothing. The diffusion tensor diffuses strongly along the dominant structure orientation (tangent to the papyrus surface) while preserving edges in the perpendicular direction, implemented with GPU acceleration via PyTorch. The result is re-binarized. This step fills small holes within surfaces, reconnects fragmented segments, and smooths boundary noise while maintaining edge sharpness.\n\n**Anisotropic Morphological Closing:** A 3D binary closing operation (dilation followed by erosion) is applied with an anisotropic ellipsoidal structuring element. The Z-radius is set larger than the XY-radius, accounting for the fact that papyrus layers may have small gaps along the depth axis due to scanning artifacts. This fills remaining small holes and bridges tiny discontinuities that survived the CED step, without over-connecting adjacent layers in the XY plane.\n\n**Dust Removal (Small Object Pruning):** Small connected components below a voxel-count threshold are removed from the final binary mask. This eliminates residual noise fragments and transient false positives that do not form meaningful papyrus structures, reducing spurious connected components.\n\n### Pipeline Summary\n\n```\nBinary Prediction (nnU-Net) + TTA\n    → Frangi Sheetness Filter (3D)\n    → CED Anisotropic Diffusion (2D slice-wise, GPU)\n    → Anisotropic Closing\n    → Dust Removal\n    → Final Prediction\n```\n\n## Note on Muon Optimizer\n\nOne thing worth mentioning — I experimented with the **Muon optimizer** as a replacement for SGD. Models trained with Muon consistently performed better on both Public and Private LB compared to SGD.\n\nHowever, I didn't have enough compute to fully train all my final models with Muon, so my submitted ensemble still uses SGD-trained models. Training speed with SGD is only about **1.25x faster** than Muon, so it's not a huge tradeoff.\n\nI've open-sourced my optimized Muon implementation: **[FastMuon](https://github.com/Decem-Y/FastMuon)** — drop-in Turbo-Muon with Triton acceleration. If you have the hardware, worth trying.\n\n## What Didn't Work / Didn't Try\n\n- More aggressive data augmentation — tried cranking up augmentation strengths / probabilities but it consistently hurt performance. Ended up using lower-than-default nnUNet augmentation probabilities\n- clDice Loss — didn't see improvement over the CE + Dice + Skeleton Recall combo\n- Adding SE (Squeeze-and-Excitation) modules into the UNet — no meaningful gain, just extra VRAM usage\n- Pseudo labels — tried using model predictions on unlabeled data as pseudo labels for semi-supervised training, but the noise was too much. Ended up degrading performance rather than helping\n- More aggressive topological post-processing (PCA hole filling, Betti matching, etc.)\n\n---\n\nDECEM",
      "votes": null
    },
    {
      "id": "3415127",
      "postDate": "02/28/2026 08:01:35",
      "content": "<p>Nice solution! I am also interested about your finding on the Muon optimizer. This is something we have been testing lately also for ink detection. Could you please share more?</p>",
      "rawMarkdown": "Nice solution! I am also interested about your finding on the Muon optimizer. This is something we have been testing lately also for ink detection. Could you please share more?",
      "votes": null
    },
    {
      "id": "3415134",
      "postDate": "02/28/2026 08:25:02",
      "content": "<p>Thanks! Here are my findings. I trained with the same config (nnUNetPlannerResEncM, 3d_lowres, 128³ patch, single fold, same post-processing) for 1000 epochs, only swapping the optimizer:</p>\n<table>\n<thead>\n<tr>\n<th>Optimizer</th>\n<th>Implementation</th>\n<th>Key Params</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>SGD (nnUNet default)</td>\n<td><code>torch.optim.SGD</code></td>\n<td>lr=0.01, momentum=0.99, nesterov=True, wd=3e-5</td>\n<td>0.590</td>\n</tr>\n<tr>\n<td>AdamW</td>\n<td><code>torch.optim.AdamW</code></td>\n<td>lr=3e-4, betas=(0.9, 0.999), eps=1e-8, wd=3e-5</td>\n<td>0.594</td>\n</tr>\n<tr>\n<td>Muon+AdamW</td>\n<td><code>SingleDeviceMuonWithAuxAdam</code></td>\n<td>Muon (hidden weights, ndim&gt;=2): lr=3e-4, momentum=0.95, variant=turbo; AdamW (bias/norm/seg_head): lr=3e-4, betas=(0.9, 0.95), eps=1e-10; wd=0.01</td>\n<td>0.595</td>\n</tr>\n</tbody>\n</table>\n<p>The Muon variant uses <code>SingleDeviceMuonWithAuxAdam</code> from <a href=\"https://github.com/Decem-Y/FastMuon\" target=\"_blank\">FastMuon</a> — Turbo-mode (AOL preconditioning + 4 Newton-Schulz iterations) is applied to conv/linear weight matrices (ndim&gt;=2, excluding embed/head/seg_layers), while the built-in AdamW handles bias, normalization, and segmentation head parameters. All use PolyLR scheduling.</p>\n<p>SGD training is only ~1.25x faster than Muon, so the speed tradeoff is small.</p>\n<p>Note: I didn't have enough compute to fully explore Muon's hyperparameter space (e.g. higher lr like 0.02 as originally recommended). With proper tuning the gap could be larger.</p>",
      "rawMarkdown": "Thanks! Here are my findings. I trained with the same config (nnUNetPlannerResEncM, 3d_lowres, 128³ patch, single fold, same post-processing) for 1000 epochs, only swapping the optimizer:\n\n| Optimizer | Implementation | Key Params | Private LB |\n|---|---|---|---|\n| SGD (nnUNet default) | `torch.optim.SGD` | lr=0.01, momentum=0.99, nesterov=True, wd=3e-5 | 0.590 |\n| AdamW | `torch.optim.AdamW` | lr=3e-4, betas=(0.9, 0.999), eps=1e-8, wd=3e-5 | 0.594 |\n| Muon+AdamW | `SingleDeviceMuonWithAuxAdam` | Muon (hidden weights, ndim>=2): lr=3e-4, momentum=0.95, variant=turbo; AdamW (bias/norm/seg_head): lr=3e-4, betas=(0.9, 0.95), eps=1e-10; wd=0.01 | 0.595 |\n\nThe Muon variant uses `SingleDeviceMuonWithAuxAdam` from [FastMuon](https://github.com/Decem-Y/FastMuon) — Turbo-mode (AOL preconditioning + 4 Newton-Schulz iterations) is applied to conv/linear weight matrices (ndim>=2, excluding embed/head/seg_layers), while the built-in AdamW handles bias, normalization, and segmentation head parameters. All use PolyLR scheduling.\n\nSGD training is only ~1.25x faster than Muon, so the speed tradeoff is small.\n\nNote: I didn't have enough compute to fully explore Muon's hyperparameter space (e.g. higher lr like 0.02 as originally recommended). With proper tuning the gap could be larger.",
      "votes": null
    },
    {
      "id": "3415718",
      "postDate": "03/01/2026 07:35:00",
      "content": "<p>The post-processing code is available in my public notebook linked below the writeup. Feel free to take a look and let me know if you have any questions — happy to discuss further!\n<a href=\"https://www.kaggle.com/code/zy1343930734/nnunet-notebook\" target=\"_blank\">notebook</a></p>",
      "rawMarkdown": "The post-processing code is available in my public notebook linked below the writeup. Feel free to take a look and let me know if you have any questions — happy to discuss further!\n[notebook](https://www.kaggle.com/code/zy1343930734/nnunet-notebook)",
      "votes": null
    },
    {
      "id": "3416414",
      "postDate": "03/02/2026 19:22:47",
      "content": "<p><a href=\"https://www.kaggle.com/zy1343930734\" target=\"_blank\">@zy1343930734</a> did you try out the implementation from timm <a href=\"https://github.com/huggingface/pytorch-image-models/blob/main/timm/optim/muon.py\" target=\"_blank\">here</a>? </p>\n<p>I tried reducing ns_steps from 5-&gt;3 and it seems like training speed decreases without any noticeable performance loss.</p>",
      "rawMarkdown": "zy1343930734 did you try out the implementation from timm [here](https://github.com/huggingface/pytorch-image-models/blob/main/timm/optim/muon.py)? \n\nI tried reducing ns_steps from 5->3 and it seems like training speed decreases without any noticeable performance loss.",
      "votes": null
    },
    {
      "id": "3417145",
      "postDate": "03/04/2026 17:21:49",
      "content": "<p>Interesting comments about using Muon. Did you compare it with AdamW on your pipeline?</p>\n<p>Nevermind, I see you discuss AdamW in another comment.</p>",
      "rawMarkdown": "Interesting comments about using Muon. Did you compare it with AdamW on your pipeline?\n\nNevermind, I see you discuss AdamW in another comment.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3415127,
      "author_name": "giorgioangelotti",
      "author_url": "",
      "post_date": "02/28/2026 08:01:35",
      "content": "<p>Nice solution! I am also interested about your finding on the Muon optimizer. This is something we have been testing lately also for ink detection. Could you please share more?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3415134,
          "author_name": "zy1343930734",
          "author_url": "",
          "post_date": "02/28/2026 08:25:02",
          "content": "<p>Thanks! Here are my findings. I trained with the same config (nnUNetPlannerResEncM, 3d_lowres, 128³ patch, single fold, same post-processing) for 1000 epochs, only swapping the optimizer:</p>\n<table>\n<thead>\n<tr>\n<th>Optimizer</th>\n<th>Implementation</th>\n<th>Key Params</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>SGD (nnUNet default)</td>\n<td><code>torch.optim.SGD</code></td>\n<td>lr=0.01, momentum=0.99, nesterov=True, wd=3e-5</td>\n<td>0.590</td>\n</tr>\n<tr>\n<td>AdamW</td>\n<td><code>torch.optim.AdamW</code></td>\n<td>lr=3e-4, betas=(0.9, 0.999), eps=1e-8, wd=3e-5</td>\n<td>0.594</td>\n</tr>\n<tr>\n<td>Muon+AdamW</td>\n<td><code>SingleDeviceMuonWithAuxAdam</code></td>\n<td>Muon (hidden weights, ndim&gt;=2): lr=3e-4, momentum=0.95, variant=turbo; AdamW (bias/norm/seg_head): lr=3e-4, betas=(0.9, 0.95), eps=1e-10; wd=0.01</td>\n<td>0.595</td>\n</tr>\n</tbody>\n</table>\n<p>The Muon variant uses <code>SingleDeviceMuonWithAuxAdam</code> from <a href=\"https://github.com/Decem-Y/FastMuon\" target=\"_blank\">FastMuon</a> — Turbo-mode (AOL preconditioning + 4 Newton-Schulz iterations) is applied to conv/linear weight matrices (ndim&gt;=2, excluding embed/head/seg_layers), while the built-in AdamW handles bias, normalization, and segmentation head parameters. All use PolyLR scheduling.</p>\n<p>SGD training is only ~1.25x faster than Muon, so the speed tradeoff is small.</p>\n<p>Note: I didn't have enough compute to fully explore Muon's hyperparameter space (e.g. higher lr like 0.02 as originally recommended). With proper tuning the gap could be larger.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3416414,
          "author_name": "brendanartley",
          "author_url": "",
          "post_date": "03/02/2026 19:22:47",
          "content": "<p><a href=\"https://www.kaggle.com/zy1343930734\" target=\"_blank\">@zy1343930734</a> did you try out the implementation from timm <a href=\"https://github.com/huggingface/pytorch-image-models/blob/main/timm/optim/muon.py\" target=\"_blank\">here</a>? </p>\n<p>I tried reducing ns_steps from 5-&gt;3 and it seems like training speed decreases without any noticeable performance loss.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3415718,
      "author_name": "zy1343930734",
      "author_url": "",
      "post_date": "03/01/2026 07:35:00",
      "content": "<p>The post-processing code is available in my public notebook linked below the writeup. Feel free to take a look and let me know if you have any questions — happy to discuss further!\n<a href=\"https://www.kaggle.com/code/zy1343930734/nnunet-notebook\" target=\"_blank\">notebook</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3417145,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "03/04/2026 17:21:49",
      "content": "<p>Interesting comments about using Muon. Did you compare it with AdamW on your pipeline?</p>\n<p>Nevermind, I see you discuss AdamW in another comment.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3415073": "Thanks to the organizers for putting together such a cool competition — learned a ton and had a lot of fun digging into volumetric topology.\n\nCongrats to all the winners and participants! Here's a quick summary of what I did.\n\n## Overview\n\n7th place, Public LB 0.618 / Private LB 0.591. Like many others, nnU-Net carried the day. Simple pipeline: two nnU-Net models ensembled + TTA + post-processing.\n\n## Model Training\n\nBoth models are based on **nnUNet** with `nnUNetPlannerResEncM`, `3d_lowres` configuration.\n\n- Patch size: **128³**\n- Epochs: **2000**\n- Optimizer: **SGD** (nnUNet default: lr=0.01, momentum=0.99, nesterov=True, wd=3e-5)\n- Loss: **CE Loss + Dice Loss + Skeleton Recall Loss**, all weighted equally (1.0 each)\n\nThe key difference between my two models:\n\n**Model A — Preprocessed training labels (key insight):** I noticed that the organizers' GitHub repo has a specific processing pipeline for the *test set* data. I applied the **exact same processing to the training labels** before training. This was probably the single most impactful thing I did — the raw training labels don't quite match the test-time processing, and this mismatch hurts generalization. Aligning train/test label processing gave a clear and consistent boost.\n\n**Model B — Raw training labels:** Standard training without any label preprocessing. This gives the ensemble some diversity.\n\n## Ensemble & Inference\n\n- Simple ensemble of Model A + Model B\n- **TTA** (nnUNet's built-in test-time augmentation with mirroring)\n\nNothing fancy here, just let nnUNet do its thing.\n\n## Post-Processing\n\nA four-stage post-processing pipeline applied on binary predictions from nnU-Net to improve topological correctness, surface accuracy, and instance consistency.\n\n**Ridge Detection (Hessian-based Frangi Filter):** The 3D Frangi sheetness filter is applied directly on the binarized prediction volume. It computes the Hessian matrix at each voxel via second-order Gaussian derivatives, then analyzes the sorted eigenvalues (|λ1| ≤ |λ2| ≤ |λ3|) to compute a sheetness response. Voxels forming planar/sheet-like structures satisfy |λ1| ≈ 0, |λ2| ≈ 0, |λ3| >> 0, which corresponds to the papyrus surface geometry. The response combines a background suppression term, a planarity term, and a blob rejection term. Only bright-on-dark structures (λ3 < 0) are retained. The Frangi output is thresholded to produce a cleaned binary mask, effectively enhancing continuous surface regions while suppressing isolated noise points and non-planar false positives.\n\n**Coherence Enhancing Diffusion (CED):** After ridge detection, Coherence Enhancing Diffusion is applied slice-by-slice (2D) on the Frangi-filtered volume to enhance spatial continuity along the papyrus surface. CED is a PDE-based anisotropic diffusion method governed by ∂u/∂t = div(D(Jρ)·∇u), where the diffusion tensor D is derived from the structure tensor Jρ. The structure tensor is computed using Pavel Holoborodko derivative kernels with Gaussian smoothing. The diffusion tensor diffuses strongly along the dominant structure orientation (tangent to the papyrus surface) while preserving edges in the perpendicular direction, implemented with GPU acceleration via PyTorch. The result is re-binarized. This step fills small holes within surfaces, reconnects fragmented segments, and smooths boundary noise while maintaining edge sharpness.\n\n**Anisotropic Morphological Closing:** A 3D binary closing operation (dilation followed by erosion) is applied with an anisotropic ellipsoidal structuring element. The Z-radius is set larger than the XY-radius, accounting for the fact that papyrus layers may have small gaps along the depth axis due to scanning artifacts. This fills remaining small holes and bridges tiny discontinuities that survived the CED step, without over-connecting adjacent layers in the XY plane.\n\n**Dust Removal (Small Object Pruning):** Small connected components below a voxel-count threshold are removed from the final binary mask. This eliminates residual noise fragments and transient false positives that do not form meaningful papyrus structures, reducing spurious connected components.\n\n### Pipeline Summary\n\n```\nBinary Prediction (nnU-Net) + TTA\n    → Frangi Sheetness Filter (3D)\n    → CED Anisotropic Diffusion (2D slice-wise, GPU)\n    → Anisotropic Closing\n    → Dust Removal\n    → Final Prediction\n```\n\n## Note on Muon Optimizer\n\nOne thing worth mentioning — I experimented with the **Muon optimizer** as a replacement for SGD. Models trained with Muon consistently performed better on both Public and Private LB compared to SGD.\n\nHowever, I didn't have enough compute to fully train all my final models with Muon, so my submitted ensemble still uses SGD-trained models. Training speed with SGD is only about **1.25x faster** than Muon, so it's not a huge tradeoff.\n\nI've open-sourced my optimized Muon implementation: **[FastMuon](https://github.com/Decem-Y/FastMuon)** — drop-in Turbo-Muon with Triton acceleration. If you have the hardware, worth trying.\n\n## What Didn't Work / Didn't Try\n\n- More aggressive data augmentation — tried cranking up augmentation strengths / probabilities but it consistently hurt performance. Ended up using lower-than-default nnUNet augmentation probabilities\n- clDice Loss — didn't see improvement over the CE + Dice + Skeleton Recall combo\n- Adding SE (Squeeze-and-Excitation) modules into the UNet — no meaningful gain, just extra VRAM usage\n- Pseudo labels — tried using model predictions on unlabeled data as pseudo labels for semi-supervised training, but the noise was too much. Ended up degrading performance rather than helping\n- More aggressive topological post-processing (PCA hole filling, Betti matching, etc.)\n\n---\n\nDECEM",
    "3415127": "Nice solution! I am also interested about your finding on the Muon optimizer. This is something we have been testing lately also for ink detection. Could you please share more?",
    "3415134": "Thanks! Here are my findings. I trained with the same config (nnUNetPlannerResEncM, 3d_lowres, 128³ patch, single fold, same post-processing) for 1000 epochs, only swapping the optimizer:\n\n| Optimizer | Implementation | Key Params | Private LB |\n|---|---|---|---|\n| SGD (nnUNet default) | `torch.optim.SGD` | lr=0.01, momentum=0.99, nesterov=True, wd=3e-5 | 0.590 |\n| AdamW | `torch.optim.AdamW` | lr=3e-4, betas=(0.9, 0.999), eps=1e-8, wd=3e-5 | 0.594 |\n| Muon+AdamW | `SingleDeviceMuonWithAuxAdam` | Muon (hidden weights, ndim>=2): lr=3e-4, momentum=0.95, variant=turbo; AdamW (bias/norm/seg_head): lr=3e-4, betas=(0.9, 0.95), eps=1e-10; wd=0.01 | 0.595 |\n\nThe Muon variant uses `SingleDeviceMuonWithAuxAdam` from [FastMuon](https://github.com/Decem-Y/FastMuon) — Turbo-mode (AOL preconditioning + 4 Newton-Schulz iterations) is applied to conv/linear weight matrices (ndim>=2, excluding embed/head/seg_layers), while the built-in AdamW handles bias, normalization, and segmentation head parameters. All use PolyLR scheduling.\n\nSGD training is only ~1.25x faster than Muon, so the speed tradeoff is small.\n\nNote: I didn't have enough compute to fully explore Muon's hyperparameter space (e.g. higher lr like 0.02 as originally recommended). With proper tuning the gap could be larger.",
    "3415718": "The post-processing code is available in my public notebook linked below the writeup. Feel free to take a look and let me know if you have any questions — happy to discuss further!\n[notebook](https://www.kaggle.com/code/zy1343930734/nnunet-notebook)",
    "3416414": "zy1343930734 did you try out the implementation from timm [here](https://github.com/huggingface/pytorch-image-models/blob/main/timm/optim/muon.py)? \n\nI tried reducing ns_steps from 5->3 and it seems like training speed decreases without any noticeable performance loss.",
    "3417145": "Interesting comments about using Muon. Did you compare it with AdamW on your pipeline?\n\nNevermind, I see you discuss AdamW in another comment."
  },
  "source": "meta"
}