{
  "id": 679255,
  "title": "43rd Place Solution - Two-stage nnU-Net for Hole Filling",
  "url": "/competitions/vesuvius-challenge-surface-detection/discussion/679255",
  "author_name": "Boredom",
  "post_date": "2026-02-28T07:00:17.514000",
  "votes": 10,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I would like to thank Kaggle and the sponsors for hosting such a fantastic competition. The contest was incredibly fierce, and while our final ranking wasn't quite as ideal as we had hoped, the entire journey has been an invaluable learning experience for me. Now, please allow me to walk you through our methodology:</p>\n<hr>\n<h1><strong>Stage 1: Model Ensemble</strong></h1>\n<h3><strong>Model Architecture</strong></h3>\n<p><strong>nnU-Net</strong> is an exceptionally powerful and robust framework that I frequently utilize in both my professional work and various competitions, including this one. For our Stage 1 pipeline, we relied on an ensemble of five models to ensure a stable and high-quality initial segmentation. \nThe ensemble consisted of the following components:</p>\n<ul>\n<li><strong>2x SegMambaV2 (3d_fullres)</strong>: We trained two versions of this architecture—one using the original official dataset and another incorporating external data to improve the model's generalization capabilities.</li>\n<li><strong>1x nnU-Net Residual Encoder (ResEncL_fullres)</strong></li>\n<li><strong>1x nnU-Net Base Model (3d_lowres)</strong>: This was included to provide a larger receptive field.</li>\n<li><strong>1x SMP-3D MIT-B5</strong>: A custom 3D Unet model contributed by my teammate <strong>lhwcv</strong>, which provided critical architectural diversity to our final Stage 1 predictions.</li>\n</ul>\n<h3><strong>Loss Function and Training Strategy</strong></h3>\n<p>We implemented a <strong>tri-phase loss training schedule</strong> to progressively refine the model's performance:</p>\n<ul>\n<li><strong>Phase 1: Foundation Training</strong><ul>\n<li>We began with the standard <strong>nnU-Net CE + Dice loss</strong>.</li>\n<li>To better align with the competition's evaluation objectives, we integrated a <strong>custom Surface Dice loss</strong> (inspired by the 1st place solution from the SenNet competition) and <strong>BoundaryDOULoss</strong>. </li>\n<li>Both of these losses were heavily modified to adapt to nnU-Net's $0-1$ dual-channel output while ensuring that <code>ignore</code> regions were correctly handled.</li></ul></li>\n<li><strong>Phase 2: Precision Enhancement</strong><ul>\n<li>Once training reached the midway point, we introduced <strong>Tversky Loss</strong> into the training objective.</li>\n<li>The primary goal here was to specifically suppress <strong>False Positives (FP)</strong>, improving the model's reliability in distinguishing subtle signals from noise.</li></ul></li>\n<li><strong>Phase 3: Topological Optimization</strong><ul>\n<li>In the final stage, we combined our custom loss functions with <strong>Topoloss</strong> (from the <em>SHAPR_torch</em> repository).</li>\n<li>This enabled the joint optimization of metrics highly correlated with the competition's scoring system. By focusing on topological features such as the <strong>number of connected components</strong> and the presence of <strong>holes</strong>, we were able to significantly enhance the structural integrity of our segmentations.</li></ul></li>\n</ul>\n<h3><strong>Augmentation</strong></h3>\n<ul>\n<li>Random <strong>flips across the X, Y, and Z axes</strong> and <strong>3D rotations</strong>.</li>\n<li><strong>Gaussian Noise</strong> and <strong>Gaussian Blur (Filtering)</strong> </li>\n<li><strong>Low-resolution Simulation</strong> and <strong>3D Cutout</strong></li>\n</ul>\n<hr>\n<h1><strong>Stage 2</strong></h1>\n<p>We used the best Stage 1 ensemble to generate out-of-fold (OOF) predictions for the full training set. These results were combined with the original images as a multi-channel input to train our Stage 2 models. The second stage integrated an ensemble of two models: <strong>nnU-Net_ResEncL</strong> and <strong>SegMambaV2</strong>.</p>\n<p>In Stage 2 training, we introduced custom augmentations specifically targeting the Stage 1 prediction mask (Channel 1) to simulate adhesions, holes, and fractures that typically impact the final score. These included:</p>\n<ol>\n<li><strong>Opening and Closing operations (0.2 probability)</strong>: Used to simulate imprecise surfaces and optimize the Surface Dice metric.</li>\n<li><strong>Simulated Holes (0.5 probability)</strong>: Applied only to non-ignore regions of the prediction mask. We generated 0 to $n$ simulated holes for each of the 1 to $n$ scroll instances in a case.</li>\n<li><strong>Sawtooth Fractures and Random Adhesions (0.5 probability)</strong>: Targeted non-ignore regions, where random adhesions were simulated between the two nearest instances.</li>\n</ol>\n<h1><strong>Post-processing</strong></h1>\n<p>We utilized only open-source post-processing methods. After seeing the top solutions, the gap between our post-processing strategies and theirs became clear. Although I attempted some of those advanced techniques, none were successful. This remains perhaps the only regret for me in this competition.</p>\n<h1><strong>Comparison of Inference Results</strong></h1>\n<p>The following shows the comparison between Stage 1 and Stage 2 inference results for cases not included in the training set:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F21284842%2F647fa23506160b9e4af2e364cd5f883e%2F20260228-150147.jpg?generation=1772262128657618&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 3415098,
      "postDate": "2026-02-28T07:00:17.513Z",
      "content": "<p>I would like to thank Kaggle and the sponsors for hosting such a fantastic competition. The contest was incredibly fierce, and while our final ranking wasn't quite as ideal as we had hoped, the entire journey has been an invaluable learning experience for me. Now, please allow me to walk you through our methodology:</p>\n<hr>\n<h1><strong>Stage 1: Model Ensemble</strong></h1>\n<h3><strong>Model Architecture</strong></h3>\n<p><strong>nnU-Net</strong> is an exceptionally powerful and robust framework that I frequently utilize in both my professional work and various competitions, including this one. For our Stage 1 pipeline, we relied on an ensemble of five models to ensure a stable and high-quality initial segmentation. \nThe ensemble consisted of the following components:</p>\n<ul>\n<li><strong>2x SegMambaV2 (3d_fullres)</strong>: We trained two versions of this architecture—one using the original official dataset and another incorporating external data to improve the model's generalization capabilities.</li>\n<li><strong>1x nnU-Net Residual Encoder (ResEncL_fullres)</strong></li>\n<li><strong>1x nnU-Net Base Model (3d_lowres)</strong>: This was included to provide a larger receptive field.</li>\n<li><strong>1x SMP-3D MIT-B5</strong>: A custom 3D Unet model contributed by my teammate <strong>lhwcv</strong>, which provided critical architectural diversity to our final Stage 1 predictions.</li>\n</ul>\n<h3><strong>Loss Function and Training Strategy</strong></h3>\n<p>We implemented a <strong>tri-phase loss training schedule</strong> to progressively refine the model's performance:</p>\n<ul>\n<li><strong>Phase 1: Foundation Training</strong><ul>\n<li>We began with the standard <strong>nnU-Net CE + Dice loss</strong>.</li>\n<li>To better align with the competition's evaluation objectives, we integrated a <strong>custom Surface Dice loss</strong> (inspired by the 1st place solution from the SenNet competition) and <strong>BoundaryDOULoss</strong>. </li>\n<li>Both of these losses were heavily modified to adapt to nnU-Net's $0-1$ dual-channel output while ensuring that <code>ignore</code> regions were correctly handled.</li></ul></li>\n<li><strong>Phase 2: Precision Enhancement</strong><ul>\n<li>Once training reached the midway point, we introduced <strong>Tversky Loss</strong> into the training objective.</li>\n<li>The primary goal here was to specifically suppress <strong>False Positives (FP)</strong>, improving the model's reliability in distinguishing subtle signals from noise.</li></ul></li>\n<li><strong>Phase 3: Topological Optimization</strong><ul>\n<li>In the final stage, we combined our custom loss functions with <strong>Topoloss</strong> (from the <em>SHAPR_torch</em> repository).</li>\n<li>This enabled the joint optimization of metrics highly correlated with the competition's scoring system. By focusing on topological features such as the <strong>number of connected components</strong> and the presence of <strong>holes</strong>, we were able to significantly enhance the structural integrity of our segmentations.</li></ul></li>\n</ul>\n<h3><strong>Augmentation</strong></h3>\n<ul>\n<li>Random <strong>flips across the X, Y, and Z axes</strong> and <strong>3D rotations</strong>.</li>\n<li><strong>Gaussian Noise</strong> and <strong>Gaussian Blur (Filtering)</strong> </li>\n<li><strong>Low-resolution Simulation</strong> and <strong>3D Cutout</strong></li>\n</ul>\n<hr>\n<h1><strong>Stage 2</strong></h1>\n<p>We used the best Stage 1 ensemble to generate out-of-fold (OOF) predictions for the full training set. These results were combined with the original images as a multi-channel input to train our Stage 2 models. The second stage integrated an ensemble of two models: <strong>nnU-Net_ResEncL</strong> and <strong>SegMambaV2</strong>.</p>\n<p>In Stage 2 training, we introduced custom augmentations specifically targeting the Stage 1 prediction mask (Channel 1) to simulate adhesions, holes, and fractures that typically impact the final score. These included:</p>\n<ol>\n<li><strong>Opening and Closing operations (0.2 probability)</strong>: Used to simulate imprecise surfaces and optimize the Surface Dice metric.</li>\n<li><strong>Simulated Holes (0.5 probability)</strong>: Applied only to non-ignore regions of the prediction mask. We generated 0 to $n$ simulated holes for each of the 1 to $n$ scroll instances in a case.</li>\n<li><strong>Sawtooth Fractures and Random Adhesions (0.5 probability)</strong>: Targeted non-ignore regions, where random adhesions were simulated between the two nearest instances.</li>\n</ol>\n<h1><strong>Post-processing</strong></h1>\n<p>We utilized only open-source post-processing methods. After seeing the top solutions, the gap between our post-processing strategies and theirs became clear. Although I attempted some of those advanced techniques, none were successful. This remains perhaps the only regret for me in this competition.</p>\n<h1><strong>Comparison of Inference Results</strong></h1>\n<p>The following shows the comparison between Stage 1 and Stage 2 inference results for cases not included in the training set:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F21284842%2F647fa23506160b9e4af2e364cd5f883e%2F20260228-150147.jpg?generation=1772262128657618&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I would like to thank Kaggle and the sponsors for hosting such a fantastic competition. The contest was incredibly fierce, and while our final ranking wasn't quite as ideal as we had hoped, the entire journey has been an invaluable learning experience for me. Now, please allow me to walk you through our methodology:\n\n---\n\n# **Stage 1: Model Ensemble**\n### **Model Architecture**\n**nnU-Net** is an exceptionally powerful and robust framework that I frequently utilize in both my professional work and various competitions, including this one. For our Stage 1 pipeline, we relied on an ensemble of five models to ensure a stable and high-quality initial segmentation. \nThe ensemble consisted of the following components:\n\n* **2x SegMambaV2 (3d_fullres)**: We trained two versions of this architecture—one using the original official dataset and another incorporating external data to improve the model's generalization capabilities.\n* **1x nnU-Net Residual Encoder (ResEncL_fullres)**\n* **1x nnU-Net Base Model (3d_lowres)**: This was included to provide a larger receptive field.\n* **1x SMP-3D MIT-B5**: A custom 3D Unet model contributed by my teammate **lhwcv**, which provided critical architectural diversity to our final Stage 1 predictions.\n\n\n### **Loss Function and Training Strategy**\nWe implemented a **tri-phase loss training schedule** to progressively refine the model's performance:\n* **Phase 1: Foundation Training**\n    * We began with the standard **nnU-Net CE + Dice loss**.\n    * To better align with the competition's evaluation objectives, we integrated a **custom Surface Dice loss** (inspired by the 1st place solution from the SenNet competition) and **BoundaryDOULoss**. \n    * Both of these losses were heavily modified to adapt to nnU-Net's $0-1$ dual-channel output while ensuring that `ignore` regions were correctly handled.\n* **Phase 2: Precision Enhancement**\n    * Once training reached the midway point, we introduced **Tversky Loss** into the training objective.\n    * The primary goal here was to specifically suppress **False Positives (FP)**, improving the model's reliability in distinguishing subtle signals from noise.\n* **Phase 3: Topological Optimization**\n    * In the final stage, we combined our custom loss functions with **Topoloss** (from the *SHAPR_torch* repository).\n    * This enabled the joint optimization of metrics highly correlated with the competition's scoring system. By focusing on topological features such as the **number of connected components** and the presence of **holes**, we were able to significantly enhance the structural integrity of our segmentations.\n\n### **Augmentation**\n* Random **flips across the X, Y, and Z axes** and **3D rotations**.\n* **Gaussian Noise** and **Gaussian Blur (Filtering)** \n* **Low-resolution Simulation** and **3D Cutout**\n\n---\n\n# **Stage 2**\nWe used the best Stage 1 ensemble to generate out-of-fold (OOF) predictions for the full training set. These results were combined with the original images as a multi-channel input to train our Stage 2 models. The second stage integrated an ensemble of two models: **nnU-Net_ResEncL** and **SegMambaV2**.\n\nIn Stage 2 training, we introduced custom augmentations specifically targeting the Stage 1 prediction mask (Channel 1) to simulate adhesions, holes, and fractures that typically impact the final score. These included:\n1. **Opening and Closing operations (0.2 probability)**: Used to simulate imprecise surfaces and optimize the Surface Dice metric.\n2. **Simulated Holes (0.5 probability)**: Applied only to non-ignore regions of the prediction mask. We generated 0 to $n$ simulated holes for each of the 1 to $n$ scroll instances in a case.\n3. **Sawtooth Fractures and Random Adhesions (0.5 probability)**: Targeted non-ignore regions, where random adhesions were simulated between the two nearest instances.\n\n# **Post-processing**\nWe utilized only open-source post-processing methods. After seeing the top solutions, the gap between our post-processing strategies and theirs became clear. Although I attempted some of those advanced techniques, none were successful. This remains perhaps the only regret for me in this competition.\n\n# **Comparison of Inference Results**\nThe following shows the comparison between Stage 1 and Stage 2 inference results for cases not included in the training set:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F21284842%2F647fa23506160b9e4af2e364cd5f883e%2F20260228-150147.jpg?generation=1772262128657618&alt=media)",
      "votes": 10
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3415098": "I would like to thank Kaggle and the sponsors for hosting such a fantastic competition. The contest was incredibly fierce, and while our final ranking wasn't quite as ideal as we had hoped, the entire journey has been an invaluable learning experience for me. Now, please allow me to walk you through our methodology:\n\n---\n\n# **Stage 1: Model Ensemble**\n### **Model Architecture**\n**nnU-Net** is an exceptionally powerful and robust framework that I frequently utilize in both my professional work and various competitions, including this one. For our Stage 1 pipeline, we relied on an ensemble of five models to ensure a stable and high-quality initial segmentation. \nThe ensemble consisted of the following components:\n\n* **2x SegMambaV2 (3d_fullres)**: We trained two versions of this architecture—one using the original official dataset and another incorporating external data to improve the model's generalization capabilities.\n* **1x nnU-Net Residual Encoder (ResEncL_fullres)**\n* **1x nnU-Net Base Model (3d_lowres)**: This was included to provide a larger receptive field.\n* **1x SMP-3D MIT-B5**: A custom 3D Unet model contributed by my teammate **lhwcv**, which provided critical architectural diversity to our final Stage 1 predictions.\n\n\n### **Loss Function and Training Strategy**\nWe implemented a **tri-phase loss training schedule** to progressively refine the model's performance:\n* **Phase 1: Foundation Training**\n    * We began with the standard **nnU-Net CE + Dice loss**.\n    * To better align with the competition's evaluation objectives, we integrated a **custom Surface Dice loss** (inspired by the 1st place solution from the SenNet competition) and **BoundaryDOULoss**. \n    * Both of these losses were heavily modified to adapt to nnU-Net's $0-1$ dual-channel output while ensuring that `ignore` regions were correctly handled.\n* **Phase 2: Precision Enhancement**\n    * Once training reached the midway point, we introduced **Tversky Loss** into the training objective.\n    * The primary goal here was to specifically suppress **False Positives (FP)**, improving the model's reliability in distinguishing subtle signals from noise.\n* **Phase 3: Topological Optimization**\n    * In the final stage, we combined our custom loss functions with **Topoloss** (from the *SHAPR_torch* repository).\n    * This enabled the joint optimization of metrics highly correlated with the competition's scoring system. By focusing on topological features such as the **number of connected components** and the presence of **holes**, we were able to significantly enhance the structural integrity of our segmentations.\n\n### **Augmentation**\n* Random **flips across the X, Y, and Z axes** and **3D rotations**.\n* **Gaussian Noise** and **Gaussian Blur (Filtering)** \n* **Low-resolution Simulation** and **3D Cutout**\n\n---\n\n# **Stage 2**\nWe used the best Stage 1 ensemble to generate out-of-fold (OOF) predictions for the full training set. These results were combined with the original images as a multi-channel input to train our Stage 2 models. The second stage integrated an ensemble of two models: **nnU-Net_ResEncL** and **SegMambaV2**.\n\nIn Stage 2 training, we introduced custom augmentations specifically targeting the Stage 1 prediction mask (Channel 1) to simulate adhesions, holes, and fractures that typically impact the final score. These included:\n1. **Opening and Closing operations (0.2 probability)**: Used to simulate imprecise surfaces and optimize the Surface Dice metric.\n2. **Simulated Holes (0.5 probability)**: Applied only to non-ignore regions of the prediction mask. We generated 0 to $n$ simulated holes for each of the 1 to $n$ scroll instances in a case.\n3. **Sawtooth Fractures and Random Adhesions (0.5 probability)**: Targeted non-ignore regions, where random adhesions were simulated between the two nearest instances.\n\n# **Post-processing**\nWe utilized only open-source post-processing methods. After seeing the top solutions, the gap between our post-processing strategies and theirs became clear. Although I attempted some of those advanced techniques, none were successful. This remains perhaps the only regret for me in this competition.\n\n# **Comparison of Inference Results**\nThe following shows the comparison between Stage 1 and Stage 2 inference results for cases not included in the training set:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F21284842%2F647fa23506160b9e4af2e364cd5f883e%2F20260228-150147.jpg?generation=1772262128657618&alt=media)"
  }
}