{
  "id": 679236,
  "title": " 3rd place solution",
  "url": "/competitions/vesuvius-challenge-surface-detection/writeups/quick-preview-of-the-3rd-place",
  "author_name": "",
  "post_date": "2026-02-28T02:58:44.420Z",
  "votes": 24,
  "comment_count": 12,
  "views": 0,
  "content": "<p>First and foremost, we would like to express our gratitude to the organizers for hosting this fantastic competition! We are honored to share our solution here. Our approach is primarily based on a highly customized <strong>nnU-Net v2</strong> framework. The core highlights include: an asymmetric patch size strategy for training and inference, two-stage cascade prediction (Cascade 3D), customized Test Time Augmentation (TTA), and a post-processing pipeline based on 3D Hessian matrix features for Ridge Detection.</p>\n<h2>1. Overall Architecture Pipeline</h2>\n<p>We built an efficient and robust two-stage cascade 3D image segmentation pipeline:</p>\n<ul>\n<li><strong>Stage 1 (3D Fullres)</strong>: Uses a full-resolution 3D model for initial prediction and performs a weighted ensemble of probability maps from multiple models.</li>\n<li><strong>Stage 2 (3D Cascade Fullres)</strong>: Takes the predictions from Stage 1 as spatial prior information (Previous Stage Predictions) and feeds them into the cascade model for refined prediction.</li>\n<li><strong>Post-processing</strong>: Combines Distance Transform and 3D Hessian matrix to perform fine-grained topological optimization and subtle structure extraction on the model's output probability maps.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F14279047%2F88cdc604565d1e046b23f663219d3a8b%2Fpipline.png?generation=1773204411449897&amp;alt=media\" alt=\"\"></p>\n<h2>2. Dataset Strategy &amp; Processing</h2>\n<ul>\n<li><strong>Full Dataset Training</strong>: To maximize the model's ability to fit the true data distribution, we abandoned the traditional K-Fold cross-validation (i.e., holding out a validation set) approach. The final submitted models were trained directly on the <strong>entire dataset (Train on all data)</strong> to achieve more robust generalization performance.</li>\n<li><strong>Pseudo-labeling Experiments</strong>: During the competition, we attempted to use external/unlabeled data to generate pseudo-labels to expand the training set. However, experimental results showed that this strategy did not bring significant performance improvements. To keep the solution concise and efficient, we did not adopt the pseudo-labeling strategy in our final version.</li>\n<li><strong>Data Augmentation</strong>: Building upon nnU-Net's default augmentation strategies, and tailored to the topological characteristics of the target, we forcefully enabled flipping across all axes (Mirror Axes: 0, 1, 2) and set a relatively high probability for rotation augmentation (Rotation Probability: 0.4-0.8).</li>\n</ul>\n<h2>3. Model Architecture &amp; Training Strategy</h2>\n<p>We deeply customized the default nnU-Net Trainer. The core improvements are as follows:</p>\n<ul>\n<li><strong>Asymmetric Patch Size Strategy</strong>:\nWe utilized the <strong>M</strong> size network architecture. Notably, we introduced a \"size decoupling\" strategy that was proven highly effective in the CZII competition: <strong>the Patch Size during training was set to 128, while the Patch Size during inference was increased to 192</strong>. This allows the model to maintain a higher Batch Size and iteration efficiency during training, while obtaining a larger receptive field during inference, effectively improving the global consistency of the segmentation results.</li>\n<li><strong>Topology-Aware Loss Function</strong>:\nIn tasks that heavily focus on structural connectivity, the classic Dice + CE combination often struggles to perfectly preserve complex elongated or mesh-like structures. Therefore, we introduced the <strong>clDice Loss</strong> on top of the default loss. clDice is specifically optimized at the skeleton level for the topological connectivity of tubular/mesh-like structures.</li>\n<li><strong>Optimizer Configuration</strong>:\nWe abandoned traditional learning rate schedulers and switched to the <strong><code>RAdamScheduleFree</code></strong> optimizer. Practice has proven that this optimizer not only makes model convergence smoother but also significantly improves training efficiency.</li>\n<li><strong>Cascade Model Training Path</strong>:\nTo build the Stage 2 Cascade model, we strictly followed a \"coarse-to-fine\" training paradigm: first training the Lowres (low-resolution) model, and subsequently training the Cascade model based on it. This ensures that the network correctly learns and utilizes the coarse predictions from the previous stage as reliable spatial priors.</li>\n</ul>\n<h2>4. Inference &amp; Model Ensemble (TTA)</h2>\n<p>During the inference stage, to balance prediction accuracy and inference speed, we rewrote the prediction logic in <code>predict_from_raw_data.py</code>:</p>\n<ul>\n<li><strong>Asymmetric Inference and Stage Fusion</strong>: As mentioned earlier, we used a large Patch Size of 192 for inference. In the two-stage pipeline, we directly used the output of the <strong>Fullres model</strong> as the feature prior input for the <strong>Cascade model</strong> to make the final prediction. Empirical evidence shows that this combination can capture the richest image details.</li>\n<li><strong>Customized TTA (Test Time Augmentation)</strong>:\nBesides conventional mirror flipping, we introduced <strong>in-plane rotation TTA (±15°)</strong> via Affine Grid Sample. Ultimately, the prediction for each patch fuses the outputs of the regular, flipped, and ±15° rotated versions. <strong>Regarding the choice of rotation angle:</strong> Since a random rotation augmentation of ±30° was used during training, our experimental evaluations showed that setting the TTA rotation angle to ±15° was the most optimal. We also tried introducing larger rotation augmentations during training, but experiments showed that it actually led to performance degradation.</li>\n<li><strong>Concurrent Extraction and Multi-Model Ensemble</strong>:\nWe wrote dedicated parallel inference scripts to bind multiple processes to specific GPUs, extracting <code>.npz</code> probability maps in parallel in a multi-GPU environment. Subsequently, we performed equal-weight or weighted fusion on Checkpoints from different iteration steps or specific settings.</li>\n</ul>\n<h2>5. Topology-Level Post-processing</h2>\n<p>This is a crucial part of our solution that achieved significant score improvements. Considering the strong spatial connectivity of the target, we designed the following powerful post-processing pipeline:</p>\n<ol>\n<li><strong>Inverse EDT Dilation</strong>: Utilizing the inverse mapping operation based on the Euclidean Distance Transform (EDT) to perform non-linear dilation on the binarized results.</li>\n<li><strong>3D Hessian Ridge Detection</strong>: We view the targets as \"ridges\" in space. After applying Gaussian smoothing to the predicted volume, we calculate the 3D Hessian matrix and use the matrix's eigenvalues to construct a feature space, thereby extracting structures with tubular features and drastically filtering out unstructured background noise.</li>\n<li><strong>Topological Pruning and Consistency Pruning</strong>:<ul>\n<li><strong>Z-axis Consistency Pruning</strong>: If a voxel is positive on the current Z-axis slice but negative on both its upper and lower adjacent slices, it is judged as a false positive and removed. This step played a decisive role in eliminating inter-slice outlier noise.</li>\n<li>We applied 2D morphological closing operations to the slices to smooth the boundaries, and performed anisotropic 3D closing on the overall volume in the XY plane to repair fractured parts.</li>\n<li>Finally, we removed 3D discrete noise with excessively small volumes.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F14279047%2F78d3564d6539e217e9c1264fcb1e33e7%2Fimage.png?generation=1772962223361424&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F14279047%2F56c87797c9e1cbde01f7c2d0ed5640ae%2Fpost-image.png?generation=1772962237802993&amp;alt=media\" alt=\"\"></li></ul></li>\n</ol>\n<blockquote>\n  <p><strong>Special Thanks:</strong> It is worth mentioning that many inspirations and core code logic in our post-processing pipeline were referenced and benefited from the <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/651532#3365613\" target=\"_blank\">wonderful discussions and code</a> shared by the organizers in the discussion forum. We would like to express our sincere gratitude to the organizers and the community for their selfless sharing!</p>\n</blockquote>\n<h2>6. Acknowledgments</h2>\n<p>Thanks again to the organizers for providing this challenging platform, and also thanks to my teammate <a href=\"https://www.kaggle.com/arunodhayan\" target=\"_blank\">@arunodhayan</a> for their hard work during the model training and exploration process!</p>\n<blockquote>\n  <p>Our Best Private LB Submission:\n  It is worth noting that our highest-scoring submission on the Private Leaderboard was achieved using a heavily trained Cascade model configuration. Specifically, this best-performing model was built by first training the <code>lowres</code> network for 8000 epochs, and subsequently training the <code>cascade</code> network on top of it for an additional 8000 epochs.<a href=\"https://www.kaggle.com/code/arunodhayan/nnunet-1-rot-tta-ensemble-cascade-no-fillin-0522f9?scriptVersionId=299432647\" target=\"_blank\">https://www.kaggle.com/code/arunodhayan/nnunet-1-rot-tta-ensemble-cascade-no-fillin-0522f9?scriptVersionId=299432647</a></p>\n</blockquote>",
  "messages": [
    {
      "id": "3414987",
      "postDate": "02/28/2026 02:57:46",
      "content": "<p>First and foremost, we would like to express our gratitude to the organizers for hosting this fantastic competition! We are honored to share our solution here. Our approach is primarily based on a highly customized <strong>nnU-Net v2</strong> framework. The core highlights include: an asymmetric patch size strategy for training and inference, two-stage cascade prediction (Cascade 3D), customized Test Time Augmentation (TTA), and a post-processing pipeline based on 3D Hessian matrix features for Ridge Detection.</p>\n<h2>1. Overall Architecture Pipeline</h2>\n<p>We built an efficient and robust two-stage cascade 3D image segmentation pipeline:</p>\n<ul>\n<li><strong>Stage 1 (3D Fullres)</strong>: Uses a full-resolution 3D model for initial prediction and performs a weighted ensemble of probability maps from multiple models.</li>\n<li><strong>Stage 2 (3D Cascade Fullres)</strong>: Takes the predictions from Stage 1 as spatial prior information (Previous Stage Predictions) and feeds them into the cascade model for refined prediction.</li>\n<li><strong>Post-processing</strong>: Combines Distance Transform and 3D Hessian matrix to perform fine-grained topological optimization and subtle structure extraction on the model's output probability maps.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F14279047%2F88cdc604565d1e046b23f663219d3a8b%2Fpipline.png?generation=1773204411449897&amp;alt=media\" alt=\"\"></p>\n<h2>2. Dataset Strategy &amp; Processing</h2>\n<ul>\n<li><strong>Full Dataset Training</strong>: To maximize the model's ability to fit the true data distribution, we abandoned the traditional K-Fold cross-validation (i.e., holding out a validation set) approach. The final submitted models were trained directly on the <strong>entire dataset (Train on all data)</strong> to achieve more robust generalization performance.</li>\n<li><strong>Pseudo-labeling Experiments</strong>: During the competition, we attempted to use external/unlabeled data to generate pseudo-labels to expand the training set. However, experimental results showed that this strategy did not bring significant performance improvements. To keep the solution concise and efficient, we did not adopt the pseudo-labeling strategy in our final version.</li>\n<li><strong>Data Augmentation</strong>: Building upon nnU-Net's default augmentation strategies, and tailored to the topological characteristics of the target, we forcefully enabled flipping across all axes (Mirror Axes: 0, 1, 2) and set a relatively high probability for rotation augmentation (Rotation Probability: 0.4-0.8).</li>\n</ul>\n<h2>3. Model Architecture &amp; Training Strategy</h2>\n<p>We deeply customized the default nnU-Net Trainer. The core improvements are as follows:</p>\n<ul>\n<li><strong>Asymmetric Patch Size Strategy</strong>:\nWe utilized the <strong>M</strong> size network architecture. Notably, we introduced a \"size decoupling\" strategy that was proven highly effective in the CZII competition: <strong>the Patch Size during training was set to 128, while the Patch Size during inference was increased to 192</strong>. This allows the model to maintain a higher Batch Size and iteration efficiency during training, while obtaining a larger receptive field during inference, effectively improving the global consistency of the segmentation results.</li>\n<li><strong>Topology-Aware Loss Function</strong>:\nIn tasks that heavily focus on structural connectivity, the classic Dice + CE combination often struggles to perfectly preserve complex elongated or mesh-like structures. Therefore, we introduced the <strong>clDice Loss</strong> on top of the default loss. clDice is specifically optimized at the skeleton level for the topological connectivity of tubular/mesh-like structures.</li>\n<li><strong>Optimizer Configuration</strong>:\nWe abandoned traditional learning rate schedulers and switched to the <strong><code>RAdamScheduleFree</code></strong> optimizer. Practice has proven that this optimizer not only makes model convergence smoother but also significantly improves training efficiency.</li>\n<li><strong>Cascade Model Training Path</strong>:\nTo build the Stage 2 Cascade model, we strictly followed a \"coarse-to-fine\" training paradigm: first training the Lowres (low-resolution) model, and subsequently training the Cascade model based on it. This ensures that the network correctly learns and utilizes the coarse predictions from the previous stage as reliable spatial priors.</li>\n</ul>\n<h2>4. Inference &amp; Model Ensemble (TTA)</h2>\n<p>During the inference stage, to balance prediction accuracy and inference speed, we rewrote the prediction logic in <code>predict_from_raw_data.py</code>:</p>\n<ul>\n<li><strong>Asymmetric Inference and Stage Fusion</strong>: As mentioned earlier, we used a large Patch Size of 192 for inference. In the two-stage pipeline, we directly used the output of the <strong>Fullres model</strong> as the feature prior input for the <strong>Cascade model</strong> to make the final prediction. Empirical evidence shows that this combination can capture the richest image details.</li>\n<li><strong>Customized TTA (Test Time Augmentation)</strong>:\nBesides conventional mirror flipping, we introduced <strong>in-plane rotation TTA (±15°)</strong> via Affine Grid Sample. Ultimately, the prediction for each patch fuses the outputs of the regular, flipped, and ±15° rotated versions. <strong>Regarding the choice of rotation angle:</strong> Since a random rotation augmentation of ±30° was used during training, our experimental evaluations showed that setting the TTA rotation angle to ±15° was the most optimal. We also tried introducing larger rotation augmentations during training, but experiments showed that it actually led to performance degradation.</li>\n<li><strong>Concurrent Extraction and Multi-Model Ensemble</strong>:\nWe wrote dedicated parallel inference scripts to bind multiple processes to specific GPUs, extracting <code>.npz</code> probability maps in parallel in a multi-GPU environment. Subsequently, we performed equal-weight or weighted fusion on Checkpoints from different iteration steps or specific settings.</li>\n</ul>\n<h2>5. Topology-Level Post-processing</h2>\n<p>This is a crucial part of our solution that achieved significant score improvements. Considering the strong spatial connectivity of the target, we designed the following powerful post-processing pipeline:</p>\n<ol>\n<li><strong>Inverse EDT Dilation</strong>: Utilizing the inverse mapping operation based on the Euclidean Distance Transform (EDT) to perform non-linear dilation on the binarized results.</li>\n<li><strong>3D Hessian Ridge Detection</strong>: We view the targets as \"ridges\" in space. After applying Gaussian smoothing to the predicted volume, we calculate the 3D Hessian matrix and use the matrix's eigenvalues to construct a feature space, thereby extracting structures with tubular features and drastically filtering out unstructured background noise.</li>\n<li><strong>Topological Pruning and Consistency Pruning</strong>:<ul>\n<li><strong>Z-axis Consistency Pruning</strong>: If a voxel is positive on the current Z-axis slice but negative on both its upper and lower adjacent slices, it is judged as a false positive and removed. This step played a decisive role in eliminating inter-slice outlier noise.</li>\n<li>We applied 2D morphological closing operations to the slices to smooth the boundaries, and performed anisotropic 3D closing on the overall volume in the XY plane to repair fractured parts.</li>\n<li>Finally, we removed 3D discrete noise with excessively small volumes.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F14279047%2F78d3564d6539e217e9c1264fcb1e33e7%2Fimage.png?generation=1772962223361424&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F14279047%2F56c87797c9e1cbde01f7c2d0ed5640ae%2Fpost-image.png?generation=1772962237802993&amp;alt=media\" alt=\"\"></li></ul></li>\n</ol>\n<blockquote>\n  <p><strong>Special Thanks:</strong> It is worth mentioning that many inspirations and core code logic in our post-processing pipeline were referenced and benefited from the <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/651532#3365613\" target=\"_blank\">wonderful discussions and code</a> shared by the organizers in the discussion forum. We would like to express our sincere gratitude to the organizers and the community for their selfless sharing!</p>\n</blockquote>\n<h2>6. Acknowledgments</h2>\n<p>Thanks again to the organizers for providing this challenging platform, and also thanks to my teammate <a href=\"https://www.kaggle.com/arunodhayan\" target=\"_blank\">@arunodhayan</a> for their hard work during the model training and exploration process!</p>\n<blockquote>\n  <p>Our Best Private LB Submission:\n  It is worth noting that our highest-scoring submission on the Private Leaderboard was achieved using a heavily trained Cascade model configuration. Specifically, this best-performing model was built by first training the <code>lowres</code> network for 8000 epochs, and subsequently training the <code>cascade</code> network on top of it for an additional 8000 epochs.<a href=\"https://www.kaggle.com/code/arunodhayan/nnunet-1-rot-tta-ensemble-cascade-no-fillin-0522f9?scriptVersionId=299432647\" target=\"_blank\">https://www.kaggle.com/code/arunodhayan/nnunet-1-rot-tta-ensemble-cascade-no-fillin-0522f9?scriptVersionId=299432647</a></p>\n</blockquote>",
      "rawMarkdown": "First and foremost, we would like to express our gratitude to the organizers for hosting this fantastic competition! We are honored to share our solution here. Our approach is primarily based on a highly customized **nnU-Net v2** framework. The core highlights include: an asymmetric patch size strategy for training and inference, two-stage cascade prediction (Cascade 3D), customized Test Time Augmentation (TTA), and a post-processing pipeline based on 3D Hessian matrix features for Ridge Detection.\n\n## 1. Overall Architecture Pipeline\nWe built an efficient and robust two-stage cascade 3D image segmentation pipeline:\n* **Stage 1 (3D Fullres)**: Uses a full-resolution 3D model for initial prediction and performs a weighted ensemble of probability maps from multiple models.\n* **Stage 2 (3D Cascade Fullres)**: Takes the predictions from Stage 1 as spatial prior information (Previous Stage Predictions) and feeds them into the cascade model for refined prediction.\n* **Post-processing**: Combines Distance Transform and 3D Hessian matrix to perform fine-grained topological optimization and subtle structure extraction on the model's output probability maps.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F14279047%2F88cdc604565d1e046b23f663219d3a8b%2Fpipline.png?generation=1773204411449897&alt=media)\n\n## 2. Dataset Strategy & Processing\n* **Full Dataset Training**: To maximize the model's ability to fit the true data distribution, we abandoned the traditional K-Fold cross-validation (i.e., holding out a validation set) approach. The final submitted models were trained directly on the **entire dataset (Train on all data)** to achieve more robust generalization performance.\n* **Pseudo-labeling Experiments**: During the competition, we attempted to use external/unlabeled data to generate pseudo-labels to expand the training set. However, experimental results showed that this strategy did not bring significant performance improvements. To keep the solution concise and efficient, we did not adopt the pseudo-labeling strategy in our final version.\n* **Data Augmentation**: Building upon nnU-Net's default augmentation strategies, and tailored to the topological characteristics of the target, we forcefully enabled flipping across all axes (Mirror Axes: 0, 1, 2) and set a relatively high probability for rotation augmentation (Rotation Probability: 0.4-0.8).\n\n## 3. Model Architecture & Training Strategy\nWe deeply customized the default nnU-Net Trainer. The core improvements are as follows:\n\n* **Asymmetric Patch Size Strategy**:\n  We utilized the **M** size network architecture. Notably, we introduced a \"size decoupling\" strategy that was proven highly effective in the CZII competition: **the Patch Size during training was set to 128, while the Patch Size during inference was increased to 192**. This allows the model to maintain a higher Batch Size and iteration efficiency during training, while obtaining a larger receptive field during inference, effectively improving the global consistency of the segmentation results.\n* **Topology-Aware Loss Function**:\n  In tasks that heavily focus on structural connectivity, the classic Dice + CE combination often struggles to perfectly preserve complex elongated or mesh-like structures. Therefore, we introduced the **clDice Loss** on top of the default loss. clDice is specifically optimized at the skeleton level for the topological connectivity of tubular/mesh-like structures.\n* **Optimizer Configuration**:\n  We abandoned traditional learning rate schedulers and switched to the **`RAdamScheduleFree`** optimizer. Practice has proven that this optimizer not only makes model convergence smoother but also significantly improves training efficiency.\n* **Cascade Model Training Path**:\n  To build the Stage 2 Cascade model, we strictly followed a \"coarse-to-fine\" training paradigm: first training the Lowres (low-resolution) model, and subsequently training the Cascade model based on it. This ensures that the network correctly learns and utilizes the coarse predictions from the previous stage as reliable spatial priors.\n\n## 4. Inference & Model Ensemble (TTA)\nDuring the inference stage, to balance prediction accuracy and inference speed, we rewrote the prediction logic in `predict_from_raw_data.py`:\n\n* **Asymmetric Inference and Stage Fusion**: As mentioned earlier, we used a large Patch Size of 192 for inference. In the two-stage pipeline, we directly used the output of the **Fullres model** as the feature prior input for the **Cascade model** to make the final prediction. Empirical evidence shows that this combination can capture the richest image details.\n* **Customized TTA (Test Time Augmentation)**:\n  Besides conventional mirror flipping, we introduced **in-plane rotation TTA (±15°)** via Affine Grid Sample. Ultimately, the prediction for each patch fuses the outputs of the regular, flipped, and ±15° rotated versions. **Regarding the choice of rotation angle:** Since a random rotation augmentation of ±30° was used during training, our experimental evaluations showed that setting the TTA rotation angle to ±15° was the most optimal. We also tried introducing larger rotation augmentations during training, but experiments showed that it actually led to performance degradation.\n* **Concurrent Extraction and Multi-Model Ensemble**:\n  We wrote dedicated parallel inference scripts to bind multiple processes to specific GPUs, extracting `.npz` probability maps in parallel in a multi-GPU environment. Subsequently, we performed equal-weight or weighted fusion on Checkpoints from different iteration steps or specific settings.\n\n## 5. Topology-Level Post-processing\nThis is a crucial part of our solution that achieved significant score improvements. Considering the strong spatial connectivity of the target, we designed the following powerful post-processing pipeline:\n\n\n\n1. **Inverse EDT Dilation**: Utilizing the inverse mapping operation based on the Euclidean Distance Transform (EDT) to perform non-linear dilation on the binarized results.\n2. **3D Hessian Ridge Detection**: We view the targets as \"ridges\" in space. After applying Gaussian smoothing to the predicted volume, we calculate the 3D Hessian matrix and use the matrix's eigenvalues to construct a feature space, thereby extracting structures with tubular features and drastically filtering out unstructured background noise.\n3. **Topological Pruning and Consistency Pruning**:\n   * **Z-axis Consistency Pruning**: If a voxel is positive on the current Z-axis slice but negative on both its upper and lower adjacent slices, it is judged as a false positive and removed. This step played a decisive role in eliminating inter-slice outlier noise.\n   * We applied 2D morphological closing operations to the slices to smooth the boundaries, and performed anisotropic 3D closing on the overall volume in the XY plane to repair fractured parts.\n   * Finally, we removed 3D discrete noise with excessively small volumes.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F14279047%2F78d3564d6539e217e9c1264fcb1e33e7%2Fimage.png?generation=1772962223361424&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F14279047%2F56c87797c9e1cbde01f7c2d0ed5640ae%2Fpost-image.png?generation=1772962237802993&alt=media)\n> **Special Thanks:** It is worth mentioning that many inspirations and core code logic in our post-processing pipeline were referenced and benefited from the [wonderful discussions and code](https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/651532#3365613) shared by the organizers in the discussion forum. We would like to express our sincere gratitude to the organizers and the community for their selfless sharing!\n\n## 6. Acknowledgments\nThanks again to the organizers for providing this challenging platform, and also thanks to my teammate @arunodhayan for their hard work during the model training and exploration process!\n\n> Our Best Private LB Submission:\nIt is worth noting that our highest-scoring submission on the Private Leaderboard was achieved using a heavily trained Cascade model configuration. Specifically, this best-performing model was built by first training the `lowres` network for 8000 epochs, and subsequently training the `cascade` network on top of it for an additional 8000 epochs.https://www.kaggle.com/code/arunodhayan/nnunet-1-rot-tta-ensemble-cascade-no-fillin-0522f9?scriptVersionId=299432647",
      "votes": null
    },
    {
      "id": "3414990",
      "postDate": "02/28/2026 03:07:20",
      "content": "<p>Hi could you share how much of an improvement your postprocessing did to the public and private score?</p>",
      "rawMarkdown": "Hi could you share how much of an improvement your postprocessing did to the public and private score?",
      "votes": null
    },
    {
      "id": "3414995",
      "postDate": "02/28/2026 03:24:37",
      "content": "<blockquote>\n  <p>Hi could you share how much of an improvement your postprocessing did to the public and private score?\n  I believe around 0.02 to 0.03</p>\n</blockquote>",
      "rawMarkdown": "> Hi could you share how much of an improvement your postprocessing did to the public and private score?\nI believe around 0.02 to 0.03",
      "votes": null
    },
    {
      "id": "3415003",
      "postDate": "02/28/2026 03:48:08",
      "content": "<p>Looking forward to the full writeup! Out of curiousity , is the ridge detection and edt dilation from the script in the Vesuvius python library? It sounds exactly like a label \"cleanup\" filter i run occasionally   </p>",
      "rawMarkdown": "Looking forward to the full writeup! Out of curiousity , is the ridge detection and edt dilation from the script in the Vesuvius python library? It sounds exactly like a label \"cleanup\" filter i run occasionally",
      "votes": null
    },
    {
      "id": "3415021",
      "postDate": "02/28/2026 04:30:49",
      "content": "<p>I'm curious about the runtime of your post processing. Hessian-based approach is one of the post proc I've tried, which takes very long.</p>",
      "rawMarkdown": "I'm curious about the runtime of your post processing. Hessian-based approach is one of the post proc I've tried, which takes very long.",
      "votes": null
    },
    {
      "id": "3415083",
      "postDate": "02/28/2026 06:31:34",
      "content": "<p>i'm curious on this, what was the training speed per batch iteration (256 patch? / 8x A100 run ?) to have the belief to run for 8000 epochs / like you looked on training loss alone for progression or topo score improvement every 10/20 validation when it comes to such scale? \nIn other words like the research mentality / thought process behind the solution.</p>",
      "rawMarkdown": "i'm curious on this, what was the training speed per batch iteration (256 patch? / 8x A100 run ?) to have the belief to run for 8000 epochs / like you looked on training loss alone for progression or topo score improvement every 10/20 validation when it comes to such scale? \nIn other words like the research mentality / thought process behind the solution.",
      "votes": null
    },
    {
      "id": "3415103",
      "postDate": "02/28/2026 07:10:04",
      "content": "<p>Our final post-processing code mostly comes from <a href=\"https://github.com/ScrollPrize/villa/blob/main/vesuvius/src/vesuvius/image_proc/run/edt_frangi_label.py\" target=\"_blank\">here</a>, and we made some simple modifications based on it. As for the processing time, I estimate it to be around an hour.</p>",
      "rawMarkdown": "Our final post-processing code mostly comes from [here](https://github.com/ScrollPrize/villa/blob/main/vesuvius/src/vesuvius/image_proc/run/edt_frangi_label.py), and we made some simple modifications based on it. As for the processing time, I estimate it to be around an hour.",
      "votes": null
    },
    {
      "id": "3415104",
      "postDate": "02/28/2026 07:11:06",
      "content": "<p>Yes, our main post-processing comes from the methods you <a href=\"https://github.com/ScrollPrize/villa/blob/main/vesuvius/src/vesuvius/image_proc/run/edt_frangi_label.py\" target=\"_blank\">shared</a></p>",
      "rawMarkdown": "Yes, our main post-processing comes from the methods you [shared](https://github.com/ScrollPrize/villa/blob/main/vesuvius/src/vesuvius/image_proc/run/edt_frangi_label.py)",
      "votes": null
    },
    {
      "id": "3415161",
      "postDate": "02/28/2026 10:13:21",
      "content": "<p><a href=\"https://www.kaggle.com/rajeshthevar\" target=\"_blank\">@rajeshthevar</a>  we used 2xH100 :  to train 3d low res it took 14hrs  and for 3d cascade it took 24hrs to completely train 8000 epochs</p>",
      "rawMarkdown": "rajeshthevar  we used 2xH100 :  to train 3d low res it took 14hrs  and for 3d cascade it took 24hrs to completely train 8000 epochs",
      "votes": null
    },
    {
      "id": "3415222",
      "postDate": "02/28/2026 12:53:16",
      "content": "<p>Great result.</p>\n<p>What implementation did you use fo Hessian based ridge detection? I spent a bit of time doing like you, extracting eigenvalues from it, and train models on it, but didn't manage to make it work. I wonder what I did wrong.</p>\n<p>Nevermind, I see you used <a href=\"https://github.com/ScrollPrize/villa/blob/main/vesuvius/src/vesuvius/image_proc/run/edt_frangi_label.py\" target=\"_blank\">https://github.com/ScrollPrize/villa/blob/main/vesuvius/src/vesuvius/image_proc/run/edt_frangi_label.py</a></p>\n<p>It was on my todo list to try it, should have prioritized it.</p>",
      "rawMarkdown": "Great result.\n\nWhat implementation did you use fo Hessian based ridge detection? I spent a bit of time doing like you, extracting eigenvalues from it, and train models on it, but didn't manage to make it work. I wonder what I did wrong.\n\nNevermind, I see you used https://github.com/ScrollPrize/villa/blob/main/vesuvius/src/vesuvius/image_proc/run/edt_frangi_label.py\n\nIt was on my todo list to try it, should have prioritized it.",
      "votes": null
    },
    {
      "id": "3416200",
      "postDate": "03/02/2026 09:42:00",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> Thankyou. Please find attached the notebook containing our implementation below.   <a href=\"https://www.kaggle.com/code/arunodhayan/nnunet-1-rot-tta-ensemble-cascade-no-fillin-0522f9?scriptVersionId=299432647\" target=\"_blank\">https://www.kaggle.com/code/arunodhayan/nnunet-1-rot-tta-ensemble-cascade-no-fillin-0522f9?scriptVersionId=299432647</a></p>",
      "rawMarkdown": "cpmpml Thankyou. Please find attached the notebook containing our implementation below.   https://www.kaggle.com/code/arunodhayan/nnunet-1-rot-tta-ensemble-cascade-no-fillin-0522f9?scriptVersionId=299432647",
      "votes": null
    },
    {
      "id": "3416345",
      "postDate": "03/02/2026 16:58:04",
      "content": "<p>Thanks, will look at it. I hope I'll see what i did wrong.</p>",
      "rawMarkdown": "Thanks, will look at it. I hope I'll see what i did wrong.",
      "votes": null
    },
    {
      "id": "3416376",
      "postDate": "03/02/2026 17:53:22",
      "content": "<p>Excellent work , \nthe asymmetric patch size strategy combined with Cascade 3D refinement is particularly interesting. Using larger inference patches to increase effective receptive field while maintaining training efficiency is a very elegant trade-off.</p>",
      "rawMarkdown": "Excellent work , \nthe asymmetric patch size strategy combined with Cascade 3D refinement is particularly interesting. Using larger inference patches to increase effective receptive field while maintaining training efficiency is a very elegant trade-off.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3414990,
      "author_name": "arjunashokbhandary",
      "author_url": "",
      "post_date": "02/28/2026 03:07:20",
      "content": "<p>Hi could you share how much of an improvement your postprocessing did to the public and private score?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3414995,
          "author_name": "arunodhayan",
          "author_url": "",
          "post_date": "02/28/2026 03:24:37",
          "content": "<blockquote>\n  <p>Hi could you share how much of an improvement your postprocessing did to the public and private score?\n  I believe around 0.02 to 0.03</p>\n</blockquote>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3415003,
      "author_name": "seanjohnsonsp",
      "author_url": "",
      "post_date": "02/28/2026 03:48:08",
      "content": "<p>Looking forward to the full writeup! Out of curiousity , is the ridge detection and edt dilation from the script in the Vesuvius python library? It sounds exactly like a label \"cleanup\" filter i run occasionally   </p>",
      "votes": null,
      "replies": [
        {
          "id": 3415104,
          "author_name": "peilwang",
          "author_url": "",
          "post_date": "02/28/2026 07:11:06",
          "content": "<p>Yes, our main post-processing comes from the methods you <a href=\"https://github.com/ScrollPrize/villa/blob/main/vesuvius/src/vesuvius/image_proc/run/edt_frangi_label.py\" target=\"_blank\">shared</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3415021,
      "author_name": "tom99763",
      "author_url": "",
      "post_date": "02/28/2026 04:30:49",
      "content": "<p>I'm curious about the runtime of your post processing. Hessian-based approach is one of the post proc I've tried, which takes very long.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3415103,
          "author_name": "peilwang",
          "author_url": "",
          "post_date": "02/28/2026 07:10:04",
          "content": "<p>Our final post-processing code mostly comes from <a href=\"https://github.com/ScrollPrize/villa/blob/main/vesuvius/src/vesuvius/image_proc/run/edt_frangi_label.py\" target=\"_blank\">here</a>, and we made some simple modifications based on it. As for the processing time, I estimate it to be around an hour.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3415083,
      "author_name": "rajeshthevar",
      "author_url": "",
      "post_date": "02/28/2026 06:31:34",
      "content": "<p>i'm curious on this, what was the training speed per batch iteration (256 patch? / 8x A100 run ?) to have the belief to run for 8000 epochs / like you looked on training loss alone for progression or topo score improvement every 10/20 validation when it comes to such scale? \nIn other words like the research mentality / thought process behind the solution.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3415161,
          "author_name": "arunodhayan",
          "author_url": "",
          "post_date": "02/28/2026 10:13:21",
          "content": "<p><a href=\"https://www.kaggle.com/rajeshthevar\" target=\"_blank\">@rajeshthevar</a>  we used 2xH100 :  to train 3d low res it took 14hrs  and for 3d cascade it took 24hrs to completely train 8000 epochs</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3415222,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "02/28/2026 12:53:16",
      "content": "<p>Great result.</p>\n<p>What implementation did you use fo Hessian based ridge detection? I spent a bit of time doing like you, extracting eigenvalues from it, and train models on it, but didn't manage to make it work. I wonder what I did wrong.</p>\n<p>Nevermind, I see you used <a href=\"https://github.com/ScrollPrize/villa/blob/main/vesuvius/src/vesuvius/image_proc/run/edt_frangi_label.py\" target=\"_blank\">https://github.com/ScrollPrize/villa/blob/main/vesuvius/src/vesuvius/image_proc/run/edt_frangi_label.py</a></p>\n<p>It was on my todo list to try it, should have prioritized it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3416200,
          "author_name": "arunodhayan",
          "author_url": "",
          "post_date": "03/02/2026 09:42:00",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> Thankyou. Please find attached the notebook containing our implementation below.   <a href=\"https://www.kaggle.com/code/arunodhayan/nnunet-1-rot-tta-ensemble-cascade-no-fillin-0522f9?scriptVersionId=299432647\" target=\"_blank\">https://www.kaggle.com/code/arunodhayan/nnunet-1-rot-tta-ensemble-cascade-no-fillin-0522f9?scriptVersionId=299432647</a></p>",
          "votes": null,
          "replies": [
            {
              "id": 3416345,
              "author_name": "cpmpml",
              "author_url": "",
              "post_date": "03/02/2026 16:58:04",
              "content": "<p>Thanks, will look at it. I hope I'll see what i did wrong.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3416376,
      "author_name": "durgakumari12",
      "author_url": "",
      "post_date": "03/02/2026 17:53:22",
      "content": "<p>Excellent work , \nthe asymmetric patch size strategy combined with Cascade 3D refinement is particularly interesting. Using larger inference patches to increase effective receptive field while maintaining training efficiency is a very elegant trade-off.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3414987": "First and foremost, we would like to express our gratitude to the organizers for hosting this fantastic competition! We are honored to share our solution here. Our approach is primarily based on a highly customized **nnU-Net v2** framework. The core highlights include: an asymmetric patch size strategy for training and inference, two-stage cascade prediction (Cascade 3D), customized Test Time Augmentation (TTA), and a post-processing pipeline based on 3D Hessian matrix features for Ridge Detection.\n\n## 1. Overall Architecture Pipeline\nWe built an efficient and robust two-stage cascade 3D image segmentation pipeline:\n* **Stage 1 (3D Fullres)**: Uses a full-resolution 3D model for initial prediction and performs a weighted ensemble of probability maps from multiple models.\n* **Stage 2 (3D Cascade Fullres)**: Takes the predictions from Stage 1 as spatial prior information (Previous Stage Predictions) and feeds them into the cascade model for refined prediction.\n* **Post-processing**: Combines Distance Transform and 3D Hessian matrix to perform fine-grained topological optimization and subtle structure extraction on the model's output probability maps.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F14279047%2F88cdc604565d1e046b23f663219d3a8b%2Fpipline.png?generation=1773204411449897&alt=media)\n\n## 2. Dataset Strategy & Processing\n* **Full Dataset Training**: To maximize the model's ability to fit the true data distribution, we abandoned the traditional K-Fold cross-validation (i.e., holding out a validation set) approach. The final submitted models were trained directly on the **entire dataset (Train on all data)** to achieve more robust generalization performance.\n* **Pseudo-labeling Experiments**: During the competition, we attempted to use external/unlabeled data to generate pseudo-labels to expand the training set. However, experimental results showed that this strategy did not bring significant performance improvements. To keep the solution concise and efficient, we did not adopt the pseudo-labeling strategy in our final version.\n* **Data Augmentation**: Building upon nnU-Net's default augmentation strategies, and tailored to the topological characteristics of the target, we forcefully enabled flipping across all axes (Mirror Axes: 0, 1, 2) and set a relatively high probability for rotation augmentation (Rotation Probability: 0.4-0.8).\n\n## 3. Model Architecture & Training Strategy\nWe deeply customized the default nnU-Net Trainer. The core improvements are as follows:\n\n* **Asymmetric Patch Size Strategy**:\n  We utilized the **M** size network architecture. Notably, we introduced a \"size decoupling\" strategy that was proven highly effective in the CZII competition: **the Patch Size during training was set to 128, while the Patch Size during inference was increased to 192**. This allows the model to maintain a higher Batch Size and iteration efficiency during training, while obtaining a larger receptive field during inference, effectively improving the global consistency of the segmentation results.\n* **Topology-Aware Loss Function**:\n  In tasks that heavily focus on structural connectivity, the classic Dice + CE combination often struggles to perfectly preserve complex elongated or mesh-like structures. Therefore, we introduced the **clDice Loss** on top of the default loss. clDice is specifically optimized at the skeleton level for the topological connectivity of tubular/mesh-like structures.\n* **Optimizer Configuration**:\n  We abandoned traditional learning rate schedulers and switched to the **`RAdamScheduleFree`** optimizer. Practice has proven that this optimizer not only makes model convergence smoother but also significantly improves training efficiency.\n* **Cascade Model Training Path**:\n  To build the Stage 2 Cascade model, we strictly followed a \"coarse-to-fine\" training paradigm: first training the Lowres (low-resolution) model, and subsequently training the Cascade model based on it. This ensures that the network correctly learns and utilizes the coarse predictions from the previous stage as reliable spatial priors.\n\n## 4. Inference & Model Ensemble (TTA)\nDuring the inference stage, to balance prediction accuracy and inference speed, we rewrote the prediction logic in `predict_from_raw_data.py`:\n\n* **Asymmetric Inference and Stage Fusion**: As mentioned earlier, we used a large Patch Size of 192 for inference. In the two-stage pipeline, we directly used the output of the **Fullres model** as the feature prior input for the **Cascade model** to make the final prediction. Empirical evidence shows that this combination can capture the richest image details.\n* **Customized TTA (Test Time Augmentation)**:\n  Besides conventional mirror flipping, we introduced **in-plane rotation TTA (±15°)** via Affine Grid Sample. Ultimately, the prediction for each patch fuses the outputs of the regular, flipped, and ±15° rotated versions. **Regarding the choice of rotation angle:** Since a random rotation augmentation of ±30° was used during training, our experimental evaluations showed that setting the TTA rotation angle to ±15° was the most optimal. We also tried introducing larger rotation augmentations during training, but experiments showed that it actually led to performance degradation.\n* **Concurrent Extraction and Multi-Model Ensemble**:\n  We wrote dedicated parallel inference scripts to bind multiple processes to specific GPUs, extracting `.npz` probability maps in parallel in a multi-GPU environment. Subsequently, we performed equal-weight or weighted fusion on Checkpoints from different iteration steps or specific settings.\n\n## 5. Topology-Level Post-processing\nThis is a crucial part of our solution that achieved significant score improvements. Considering the strong spatial connectivity of the target, we designed the following powerful post-processing pipeline:\n\n\n\n1. **Inverse EDT Dilation**: Utilizing the inverse mapping operation based on the Euclidean Distance Transform (EDT) to perform non-linear dilation on the binarized results.\n2. **3D Hessian Ridge Detection**: We view the targets as \"ridges\" in space. After applying Gaussian smoothing to the predicted volume, we calculate the 3D Hessian matrix and use the matrix's eigenvalues to construct a feature space, thereby extracting structures with tubular features and drastically filtering out unstructured background noise.\n3. **Topological Pruning and Consistency Pruning**:\n   * **Z-axis Consistency Pruning**: If a voxel is positive on the current Z-axis slice but negative on both its upper and lower adjacent slices, it is judged as a false positive and removed. This step played a decisive role in eliminating inter-slice outlier noise.\n   * We applied 2D morphological closing operations to the slices to smooth the boundaries, and performed anisotropic 3D closing on the overall volume in the XY plane to repair fractured parts.\n   * Finally, we removed 3D discrete noise with excessively small volumes.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F14279047%2F78d3564d6539e217e9c1264fcb1e33e7%2Fimage.png?generation=1772962223361424&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F14279047%2F56c87797c9e1cbde01f7c2d0ed5640ae%2Fpost-image.png?generation=1772962237802993&alt=media)\n> **Special Thanks:** It is worth mentioning that many inspirations and core code logic in our post-processing pipeline were referenced and benefited from the [wonderful discussions and code](https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/651532#3365613) shared by the organizers in the discussion forum. We would like to express our sincere gratitude to the organizers and the community for their selfless sharing!\n\n## 6. Acknowledgments\nThanks again to the organizers for providing this challenging platform, and also thanks to my teammate @arunodhayan for their hard work during the model training and exploration process!\n\n> Our Best Private LB Submission:\nIt is worth noting that our highest-scoring submission on the Private Leaderboard was achieved using a heavily trained Cascade model configuration. Specifically, this best-performing model was built by first training the `lowres` network for 8000 epochs, and subsequently training the `cascade` network on top of it for an additional 8000 epochs.https://www.kaggle.com/code/arunodhayan/nnunet-1-rot-tta-ensemble-cascade-no-fillin-0522f9?scriptVersionId=299432647",
    "3414990": "Hi could you share how much of an improvement your postprocessing did to the public and private score?",
    "3414995": "> Hi could you share how much of an improvement your postprocessing did to the public and private score?\nI believe around 0.02 to 0.03",
    "3415003": "Looking forward to the full writeup! Out of curiousity , is the ridge detection and edt dilation from the script in the Vesuvius python library? It sounds exactly like a label \"cleanup\" filter i run occasionally",
    "3415021": "I'm curious about the runtime of your post processing. Hessian-based approach is one of the post proc I've tried, which takes very long.",
    "3415083": "i'm curious on this, what was the training speed per batch iteration (256 patch? / 8x A100 run ?) to have the belief to run for 8000 epochs / like you looked on training loss alone for progression or topo score improvement every 10/20 validation when it comes to such scale? \nIn other words like the research mentality / thought process behind the solution.",
    "3415103": "Our final post-processing code mostly comes from [here](https://github.com/ScrollPrize/villa/blob/main/vesuvius/src/vesuvius/image_proc/run/edt_frangi_label.py), and we made some simple modifications based on it. As for the processing time, I estimate it to be around an hour.",
    "3415104": "Yes, our main post-processing comes from the methods you [shared](https://github.com/ScrollPrize/villa/blob/main/vesuvius/src/vesuvius/image_proc/run/edt_frangi_label.py)",
    "3415161": "rajeshthevar  we used 2xH100 :  to train 3d low res it took 14hrs  and for 3d cascade it took 24hrs to completely train 8000 epochs",
    "3415222": "Great result.\n\nWhat implementation did you use fo Hessian based ridge detection? I spent a bit of time doing like you, extracting eigenvalues from it, and train models on it, but didn't manage to make it work. I wonder what I did wrong.\n\nNevermind, I see you used https://github.com/ScrollPrize/villa/blob/main/vesuvius/src/vesuvius/image_proc/run/edt_frangi_label.py\n\nIt was on my todo list to try it, should have prioritized it.",
    "3416200": "cpmpml Thankyou. Please find attached the notebook containing our implementation below.   https://www.kaggle.com/code/arunodhayan/nnunet-1-rot-tta-ensemble-cascade-no-fillin-0522f9?scriptVersionId=299432647",
    "3416345": "Thanks, will look at it. I hope I'll see what i did wrong.",
    "3416376": "Excellent work , \nthe asymmetric patch size strategy combined with Cascade 3D refinement is particularly interesting. Using larger inference patches to increase effective receptive field while maintaining training efficiency is a very elegant trade-off."
  },
  "source": "meta"
}