{
  "id": 679581,
  "title": "35th SilverSolution🥈High-Epoch nnU-Net Cascade with Optimized Patch ",
  "url": "/competitions/vesuvius-challenge-surface-detection/discussion/679581",
  "author_name": "Suzuki Taichi",
  "post_date": "2026-03-02T11:57:36.407000",
  "votes": 2,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Here is the updated summarized write-up including the cascade threshold/epoch table in a clean and integrated way:</p>\n<hr>\n<h1>35th Place Silver Solution</h1>\n<h2>High-Epoch nnU-Net Cascade with Patch and Threshold Optimization</h2>\n<h2>Overview</h2>\n<p>This solution is based on a carefully optimized 2-stage nnU-Net cascade framework. \nFinal LB score reached * <strong>0.604 Public LB</strong>, * <strong>0.587 Private LB</strong> .\n Instead of introducing architectural complexity, performance gains were achieved through:</p>\n<ul>\n<li>Long convergence training</li>\n<li>Careful cascade refinement</li>\n<li>Systematic threshold tuning</li>\n<li>Optimized inference crop geometry</li>\n<li>Lightweight ensembling</li>\n</ul>\n<p>The approach prioritizes stability and controlled context expansion.</p>\n<hr>\n<h1>Stage 1: High-Epoch Lower-Resolution Model</h1>\n<p>A 3D nnU-Net lower-resolution model was trained for <strong>1500 epochs</strong> (patch size 120³).</p>\n<p>Long training was essential for producing stable coarse surface predictions that serve as reliable inputs to the cascade stage.</p>\n<p>Performance improved steadily with training length:</p>\n<ul>\n<li>250 epochs → 0.492 LB</li>\n<li>750 epochs → 0.508 LB</li>\n<li>1200 epochs → 0.540 LB</li>\n<li>1500 epochs → 0.542 LB</li>\n</ul>\n<p>Increasing inference crop size from 160 to 320 (without retraining) further improved LB from <strong>0.542 → 0.552</strong>, highlighting the importance of inference geometry.</p>\n<hr>\n<h1>Stage 2: Cascade Full-Resolution Model</h1>\n<p>The cascade model refines predictions using:</p>\n<ul>\n<li>Full-resolution image</li>\n<li>Upsampled Stage 1 prediction as an additional channel</li>\n</ul>\n<p>Base cascade (200 epochs, argmax threshold):</p>\n<ul>\n<li>CV: 0.697</li>\n<li>Public LB: 0.573</li>\n</ul>\n<hr>\n<h1>Cascade Epoch &amp; Threshold Study</h1>\n<p>We systematically evaluated threshold and epoch effects.</p>\n<h3>200 Epoch Model</h3>\n<table>\n<thead>\n<tr>\n<th>Threshold</th>\n<th>CV</th>\n<th>Public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Argmax</td>\n<td>0.6973</td>\n<td>0.573</td>\n</tr>\n<tr>\n<td>0.5</td>\n<td>—</td>\n<td>0.560</td>\n</tr>\n<tr>\n<td>0.3</td>\n<td>—</td>\n<td>0.573</td>\n</tr>\n<tr>\n<td><strong>0.2</strong></td>\n<td>—</td>\n<td><strong>0.583</strong></td>\n</tr>\n<tr>\n<td>0.1</td>\n<td>—</td>\n<td>0.576</td>\n</tr>\n</tbody>\n</table>\n<p>Key finding:\nThreshold <strong>0.2 significantly outperformed argmax and default 0.5</strong>.</p>\n<hr>\n<h3>Extended Training</h3>\n<table>\n<thead>\n<tr>\n<th>Epoch</th>\n<th>Threshold</th>\n<th>CV</th>\n<th>Public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>228</td>\n<td>Argmax</td>\n<td>0.6984</td>\n<td>0.579</td>\n</tr>\n<tr>\n<td>228</td>\n<td>0.2</td>\n<td>0.6984</td>\n<td>0.578</td>\n</tr>\n<tr>\n<td>989</td>\n<td>Argmax</td>\n<td>0.7035</td>\n<td>0.579</td>\n</tr>\n</tbody>\n</table>\n<p>Observations:</p>\n<ul>\n<li>Extending training from 200 → 228 epochs improved LB.</li>\n<li>Very long training (989 epochs) improved CV but did not improve LB further.</li>\n<li>Performance plateaued despite increasing CV.</li>\n</ul>\n<p>Conclusion:\nModerate training extension (200–300 epochs) is effective.\nExtremely long cascade training yields diminishing leaderboard returns.</p>\n<hr>\n<h1>Inference Crop Size Study</h1>\n<p>We evaluated sliding window inference crop sizes:</p>\n<table>\n<thead>\n<tr>\n<th>Inference Crop</th>\n<th>Public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>360</td>\n<td>0.552</td>\n</tr>\n<tr>\n<td>256</td>\n<td>0.551</td>\n</tr>\n<tr>\n<td>180</td>\n<td>0.508</td>\n</tr>\n</tbody>\n</table>\n<p>Key insights:</p>\n<ul>\n<li>Large XY crops (256–360) perform similarly.</li>\n<li>Small crops (180) severely degrade structural continuity.</li>\n<li>Larger XY inference improves spatial coherence without retraining.</li>\n</ul>\n<p>Final inference used:</p>\n<p>(z, y, x) = (128, 320, 320)\nThreshold = 0.15 (final optimized value)</p>\n<hr>\n<h1>Ensembling</h1>\n<p>A lightweight ensemble of two cascade variants (standard + patch variant) achieved:</p>\n<ul>\n<li><strong>0.604 Public LB</strong></li>\n<li><strong>0.587 Private LB</strong></li>\n</ul>\n<p>The ensemble improved robustness while keeping the system simple.</p>\n<hr>\n<h1>Overall Insights</h1>\n<ol>\n<li>Long Stage 1 convergence stabilizes the cascade.</li>\n<li>Threshold tuning provides one of the largest performance gains.</li>\n<li>Cascade performance plateaus beyond moderate epoch counts.</li>\n<li>Inference patch size strongly affects spatial continuity.</li>\n<li>Controlled patch variation improves robustness.</li>\n<li>Simple ensembling is sufficient.</li>\n</ol>\n<hr>\n<p>Here is the updated summarized write-up including the cascade threshold/epoch table in a clean and integrated way:</p>\n<hr>\n<h1>35th Place Silver Solution</h1>\n<h2>High-Epoch nnU-Net Cascade with Patch and Threshold Optimization</h2>\n<h2>Overview</h2>\n<p>This solution is based on a carefully optimized 2-stage nnU-Net cascade framework. Instead of introducing architectural complexity, performance gains were achieved through:</p>\n<ul>\n<li>Long convergence training</li>\n<li>Careful cascade refinement</li>\n<li>Systematic threshold tuning</li>\n<li>Optimized inference crop geometry</li>\n<li>Lightweight ensembling</li>\n</ul>\n<p>The approach prioritizes stability and controlled context expansion.</p>\n<hr>\n<h1>Stage 1: High-Epoch Lower-Resolution Model</h1>\n<p>A 3D nnU-Net lower-resolution model was trained for <strong>1500 epochs</strong> (patch size 120³).</p>\n<p>Long training was essential for producing stable coarse surface predictions that serve as reliable inputs to the cascade stage.</p>\n<p>Performance improved steadily with training length:</p>\n<ul>\n<li>250 epochs → 0.492 LB</li>\n<li>750 epochs → 0.508 LB</li>\n<li>1200 epochs → 0.540 LB</li>\n<li>1500 epochs → 0.542 LB</li>\n</ul>\n<p>Increasing inference crop size from 160 to 320 (without retraining) further improved LB from <strong>0.542 → 0.552</strong>, highlighting the importance of inference geometry.</p>\n<hr>\n<h1>Stage 2: Cascade Full-Resolution Model</h1>\n<p>The cascade model refines predictions using:</p>\n<ul>\n<li>Full-resolution image</li>\n<li>Upsampled Stage 1 prediction as an additional channel</li>\n</ul>\n<p>Base cascade (200 epochs, argmax threshold):</p>\n<ul>\n<li>CV: 0.697</li>\n<li>Public LB: 0.573</li>\n</ul>\n<hr>\n<h1>Cascade Epoch &amp; Threshold Study</h1>\n<p>We systematically evaluated threshold and epoch effects.</p>\n<h3>200 Epoch Model</h3>\n<table>\n<thead>\n<tr>\n<th>Threshold</th>\n<th>CV</th>\n<th>Public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Argmax</td>\n<td>0.6973</td>\n<td>0.573</td>\n</tr>\n<tr>\n<td>0.5</td>\n<td>—</td>\n<td>0.560</td>\n</tr>\n<tr>\n<td>0.3</td>\n<td>—</td>\n<td>0.573</td>\n</tr>\n<tr>\n<td><strong>0.2</strong></td>\n<td>—</td>\n<td><strong>0.583</strong></td>\n</tr>\n<tr>\n<td>0.1</td>\n<td>—</td>\n<td>0.576</td>\n</tr>\n</tbody>\n</table>\n<p>Key finding:\nThreshold <strong>0.2 significantly outperformed argmax and default 0.5</strong>.</p>\n<hr>\n<h3>Extended Training</h3>\n<table>\n<thead>\n<tr>\n<th>Epoch</th>\n<th>Threshold</th>\n<th>CV</th>\n<th>Public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>228</td>\n<td>Argmax</td>\n<td>0.6984</td>\n<td>0.579</td>\n</tr>\n<tr>\n<td>228</td>\n<td>0.2</td>\n<td>0.6984</td>\n<td>0.578</td>\n</tr>\n<tr>\n<td>989</td>\n<td>Argmax</td>\n<td>0.7035</td>\n<td>0.579</td>\n</tr>\n</tbody>\n</table>\n<p>Observations:</p>\n<ul>\n<li>Extending training from 200 → 228 epochs improved LB.</li>\n<li>Very long training (989 epochs) improved CV but did not improve LB further.</li>\n<li>Performance plateaued despite increasing CV.</li>\n</ul>\n<p>Conclusion:\nModerate training extension (200–300 epochs) is effective.\nExtremely long cascade training yields diminishing leaderboard returns.</p>\n<hr>\n<h1>Inference Crop Size Study</h1>\n<p>We evaluated sliding window inference crop sizes:</p>\n<table>\n<thead>\n<tr>\n<th>Inference Crop</th>\n<th>Public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>360</td>\n<td>0.552</td>\n</tr>\n<tr>\n<td>256</td>\n<td>0.551</td>\n</tr>\n<tr>\n<td>180</td>\n<td>0.508</td>\n</tr>\n</tbody>\n</table>\n<p>Key insights:</p>\n<ul>\n<li>Large XY crops (256–360) perform similarly.</li>\n<li>Small crops (180) severely degrade structural continuity.</li>\n<li>Larger XY inference improves spatial coherence without retraining.</li>\n</ul>\n<p>Final inference used:</p>\n<p>(z, y, x) = (128, 320, 320)\nThreshold = 0.15 (final optimized value)</p>\n<hr>\n<h1>Ensembling</h1>\n<p>A lightweight ensemble of two cascade variants (standard + patch variant) achieved:</p>\n<ul>\n<li><strong>0.604 Public LB</strong></li>\n<li><strong>0.587 Private LB</strong></li>\n</ul>\n<p>The ensemble improved robustness while keeping the system simple.</p>\n<hr>\n<h1>Overall Insights</h1>\n<ol>\n<li>Long Stage 1 convergence stabilizes the cascade.</li>\n<li>Threshold tuning provides one of the largest performance gains.</li>\n<li>Cascade performance plateaus beyond moderate epoch counts.</li>\n<li>Inference patch size strongly affects spatial continuity.</li>\n<li>Controlled patch variation improves robustness.</li>\n<li>Simple ensembling is sufficient.</li>\n</ol>\n<hr>",
  "messages": [
    {
      "id": 3416245,
      "postDate": "2026-03-02T11:57:36.407Z",
      "content": "<p>Here is the updated summarized write-up including the cascade threshold/epoch table in a clean and integrated way:</p>\n<hr>\n<h1>35th Place Silver Solution</h1>\n<h2>High-Epoch nnU-Net Cascade with Patch and Threshold Optimization</h2>\n<h2>Overview</h2>\n<p>This solution is based on a carefully optimized 2-stage nnU-Net cascade framework. \nFinal LB score reached * <strong>0.604 Public LB</strong>, * <strong>0.587 Private LB</strong> .\n Instead of introducing architectural complexity, performance gains were achieved through:</p>\n<ul>\n<li>Long convergence training</li>\n<li>Careful cascade refinement</li>\n<li>Systematic threshold tuning</li>\n<li>Optimized inference crop geometry</li>\n<li>Lightweight ensembling</li>\n</ul>\n<p>The approach prioritizes stability and controlled context expansion.</p>\n<hr>\n<h1>Stage 1: High-Epoch Lower-Resolution Model</h1>\n<p>A 3D nnU-Net lower-resolution model was trained for <strong>1500 epochs</strong> (patch size 120³).</p>\n<p>Long training was essential for producing stable coarse surface predictions that serve as reliable inputs to the cascade stage.</p>\n<p>Performance improved steadily with training length:</p>\n<ul>\n<li>250 epochs → 0.492 LB</li>\n<li>750 epochs → 0.508 LB</li>\n<li>1200 epochs → 0.540 LB</li>\n<li>1500 epochs → 0.542 LB</li>\n</ul>\n<p>Increasing inference crop size from 160 to 320 (without retraining) further improved LB from <strong>0.542 → 0.552</strong>, highlighting the importance of inference geometry.</p>\n<hr>\n<h1>Stage 2: Cascade Full-Resolution Model</h1>\n<p>The cascade model refines predictions using:</p>\n<ul>\n<li>Full-resolution image</li>\n<li>Upsampled Stage 1 prediction as an additional channel</li>\n</ul>\n<p>Base cascade (200 epochs, argmax threshold):</p>\n<ul>\n<li>CV: 0.697</li>\n<li>Public LB: 0.573</li>\n</ul>\n<hr>\n<h1>Cascade Epoch &amp; Threshold Study</h1>\n<p>We systematically evaluated threshold and epoch effects.</p>\n<h3>200 Epoch Model</h3>\n<table>\n<thead>\n<tr>\n<th>Threshold</th>\n<th>CV</th>\n<th>Public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Argmax</td>\n<td>0.6973</td>\n<td>0.573</td>\n</tr>\n<tr>\n<td>0.5</td>\n<td>—</td>\n<td>0.560</td>\n</tr>\n<tr>\n<td>0.3</td>\n<td>—</td>\n<td>0.573</td>\n</tr>\n<tr>\n<td><strong>0.2</strong></td>\n<td>—</td>\n<td><strong>0.583</strong></td>\n</tr>\n<tr>\n<td>0.1</td>\n<td>—</td>\n<td>0.576</td>\n</tr>\n</tbody>\n</table>\n<p>Key finding:\nThreshold <strong>0.2 significantly outperformed argmax and default 0.5</strong>.</p>\n<hr>\n<h3>Extended Training</h3>\n<table>\n<thead>\n<tr>\n<th>Epoch</th>\n<th>Threshold</th>\n<th>CV</th>\n<th>Public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>228</td>\n<td>Argmax</td>\n<td>0.6984</td>\n<td>0.579</td>\n</tr>\n<tr>\n<td>228</td>\n<td>0.2</td>\n<td>0.6984</td>\n<td>0.578</td>\n</tr>\n<tr>\n<td>989</td>\n<td>Argmax</td>\n<td>0.7035</td>\n<td>0.579</td>\n</tr>\n</tbody>\n</table>\n<p>Observations:</p>\n<ul>\n<li>Extending training from 200 → 228 epochs improved LB.</li>\n<li>Very long training (989 epochs) improved CV but did not improve LB further.</li>\n<li>Performance plateaued despite increasing CV.</li>\n</ul>\n<p>Conclusion:\nModerate training extension (200–300 epochs) is effective.\nExtremely long cascade training yields diminishing leaderboard returns.</p>\n<hr>\n<h1>Inference Crop Size Study</h1>\n<p>We evaluated sliding window inference crop sizes:</p>\n<table>\n<thead>\n<tr>\n<th>Inference Crop</th>\n<th>Public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>360</td>\n<td>0.552</td>\n</tr>\n<tr>\n<td>256</td>\n<td>0.551</td>\n</tr>\n<tr>\n<td>180</td>\n<td>0.508</td>\n</tr>\n</tbody>\n</table>\n<p>Key insights:</p>\n<ul>\n<li>Large XY crops (256–360) perform similarly.</li>\n<li>Small crops (180) severely degrade structural continuity.</li>\n<li>Larger XY inference improves spatial coherence without retraining.</li>\n</ul>\n<p>Final inference used:</p>\n<p>(z, y, x) = (128, 320, 320)\nThreshold = 0.15 (final optimized value)</p>\n<hr>\n<h1>Ensembling</h1>\n<p>A lightweight ensemble of two cascade variants (standard + patch variant) achieved:</p>\n<ul>\n<li><strong>0.604 Public LB</strong></li>\n<li><strong>0.587 Private LB</strong></li>\n</ul>\n<p>The ensemble improved robustness while keeping the system simple.</p>\n<hr>\n<h1>Overall Insights</h1>\n<ol>\n<li>Long Stage 1 convergence stabilizes the cascade.</li>\n<li>Threshold tuning provides one of the largest performance gains.</li>\n<li>Cascade performance plateaus beyond moderate epoch counts.</li>\n<li>Inference patch size strongly affects spatial continuity.</li>\n<li>Controlled patch variation improves robustness.</li>\n<li>Simple ensembling is sufficient.</li>\n</ol>\n<hr>\n<p>Here is the updated summarized write-up including the cascade threshold/epoch table in a clean and integrated way:</p>\n<hr>\n<h1>35th Place Silver Solution</h1>\n<h2>High-Epoch nnU-Net Cascade with Patch and Threshold Optimization</h2>\n<h2>Overview</h2>\n<p>This solution is based on a carefully optimized 2-stage nnU-Net cascade framework. Instead of introducing architectural complexity, performance gains were achieved through:</p>\n<ul>\n<li>Long convergence training</li>\n<li>Careful cascade refinement</li>\n<li>Systematic threshold tuning</li>\n<li>Optimized inference crop geometry</li>\n<li>Lightweight ensembling</li>\n</ul>\n<p>The approach prioritizes stability and controlled context expansion.</p>\n<hr>\n<h1>Stage 1: High-Epoch Lower-Resolution Model</h1>\n<p>A 3D nnU-Net lower-resolution model was trained for <strong>1500 epochs</strong> (patch size 120³).</p>\n<p>Long training was essential for producing stable coarse surface predictions that serve as reliable inputs to the cascade stage.</p>\n<p>Performance improved steadily with training length:</p>\n<ul>\n<li>250 epochs → 0.492 LB</li>\n<li>750 epochs → 0.508 LB</li>\n<li>1200 epochs → 0.540 LB</li>\n<li>1500 epochs → 0.542 LB</li>\n</ul>\n<p>Increasing inference crop size from 160 to 320 (without retraining) further improved LB from <strong>0.542 → 0.552</strong>, highlighting the importance of inference geometry.</p>\n<hr>\n<h1>Stage 2: Cascade Full-Resolution Model</h1>\n<p>The cascade model refines predictions using:</p>\n<ul>\n<li>Full-resolution image</li>\n<li>Upsampled Stage 1 prediction as an additional channel</li>\n</ul>\n<p>Base cascade (200 epochs, argmax threshold):</p>\n<ul>\n<li>CV: 0.697</li>\n<li>Public LB: 0.573</li>\n</ul>\n<hr>\n<h1>Cascade Epoch &amp; Threshold Study</h1>\n<p>We systematically evaluated threshold and epoch effects.</p>\n<h3>200 Epoch Model</h3>\n<table>\n<thead>\n<tr>\n<th>Threshold</th>\n<th>CV</th>\n<th>Public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Argmax</td>\n<td>0.6973</td>\n<td>0.573</td>\n</tr>\n<tr>\n<td>0.5</td>\n<td>—</td>\n<td>0.560</td>\n</tr>\n<tr>\n<td>0.3</td>\n<td>—</td>\n<td>0.573</td>\n</tr>\n<tr>\n<td><strong>0.2</strong></td>\n<td>—</td>\n<td><strong>0.583</strong></td>\n</tr>\n<tr>\n<td>0.1</td>\n<td>—</td>\n<td>0.576</td>\n</tr>\n</tbody>\n</table>\n<p>Key finding:\nThreshold <strong>0.2 significantly outperformed argmax and default 0.5</strong>.</p>\n<hr>\n<h3>Extended Training</h3>\n<table>\n<thead>\n<tr>\n<th>Epoch</th>\n<th>Threshold</th>\n<th>CV</th>\n<th>Public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>228</td>\n<td>Argmax</td>\n<td>0.6984</td>\n<td>0.579</td>\n</tr>\n<tr>\n<td>228</td>\n<td>0.2</td>\n<td>0.6984</td>\n<td>0.578</td>\n</tr>\n<tr>\n<td>989</td>\n<td>Argmax</td>\n<td>0.7035</td>\n<td>0.579</td>\n</tr>\n</tbody>\n</table>\n<p>Observations:</p>\n<ul>\n<li>Extending training from 200 → 228 epochs improved LB.</li>\n<li>Very long training (989 epochs) improved CV but did not improve LB further.</li>\n<li>Performance plateaued despite increasing CV.</li>\n</ul>\n<p>Conclusion:\nModerate training extension (200–300 epochs) is effective.\nExtremely long cascade training yields diminishing leaderboard returns.</p>\n<hr>\n<h1>Inference Crop Size Study</h1>\n<p>We evaluated sliding window inference crop sizes:</p>\n<table>\n<thead>\n<tr>\n<th>Inference Crop</th>\n<th>Public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>360</td>\n<td>0.552</td>\n</tr>\n<tr>\n<td>256</td>\n<td>0.551</td>\n</tr>\n<tr>\n<td>180</td>\n<td>0.508</td>\n</tr>\n</tbody>\n</table>\n<p>Key insights:</p>\n<ul>\n<li>Large XY crops (256–360) perform similarly.</li>\n<li>Small crops (180) severely degrade structural continuity.</li>\n<li>Larger XY inference improves spatial coherence without retraining.</li>\n</ul>\n<p>Final inference used:</p>\n<p>(z, y, x) = (128, 320, 320)\nThreshold = 0.15 (final optimized value)</p>\n<hr>\n<h1>Ensembling</h1>\n<p>A lightweight ensemble of two cascade variants (standard + patch variant) achieved:</p>\n<ul>\n<li><strong>0.604 Public LB</strong></li>\n<li><strong>0.587 Private LB</strong></li>\n</ul>\n<p>The ensemble improved robustness while keeping the system simple.</p>\n<hr>\n<h1>Overall Insights</h1>\n<ol>\n<li>Long Stage 1 convergence stabilizes the cascade.</li>\n<li>Threshold tuning provides one of the largest performance gains.</li>\n<li>Cascade performance plateaus beyond moderate epoch counts.</li>\n<li>Inference patch size strongly affects spatial continuity.</li>\n<li>Controlled patch variation improves robustness.</li>\n<li>Simple ensembling is sufficient.</li>\n</ol>\n<hr>",
      "rawMarkdown": "Here is the updated summarized write-up including the cascade threshold/epoch table in a clean and integrated way:\n\n---\n\n# 35th Place Silver Solution\n\n## High-Epoch nnU-Net Cascade with Patch and Threshold Optimization\n\n## Overview\n\nThis solution is based on a carefully optimized 2-stage nnU-Net cascade framework. \nFinal LB score reached * **0.604 Public LB**, * **0.587 Private LB** .\n Instead of introducing architectural complexity, performance gains were achieved through:\n\n* Long convergence training\n* Careful cascade refinement\n* Systematic threshold tuning\n* Optimized inference crop geometry\n* Lightweight ensembling\n\nThe approach prioritizes stability and controlled context expansion.\n\n---\n\n# Stage 1: High-Epoch Lower-Resolution Model\n\nA 3D nnU-Net lower-resolution model was trained for **1500 epochs** (patch size 120³).\n\nLong training was essential for producing stable coarse surface predictions that serve as reliable inputs to the cascade stage.\n\nPerformance improved steadily with training length:\n\n* 250 epochs → 0.492 LB\n* 750 epochs → 0.508 LB\n* 1200 epochs → 0.540 LB\n* 1500 epochs → 0.542 LB\n\nIncreasing inference crop size from 160 to 320 (without retraining) further improved LB from **0.542 → 0.552**, highlighting the importance of inference geometry.\n\n---\n\n# Stage 2: Cascade Full-Resolution Model\n\nThe cascade model refines predictions using:\n\n* Full-resolution image\n* Upsampled Stage 1 prediction as an additional channel\n\nBase cascade (200 epochs, argmax threshold):\n\n* CV: 0.697\n* Public LB: 0.573\n\n---\n\n# Cascade Epoch & Threshold Study\n\nWe systematically evaluated threshold and epoch effects.\n\n### 200 Epoch Model\n\n| Threshold | CV     | Public LB |\n| --------- | ------ | --------- |\n| Argmax    | 0.6973 | 0.573     |\n| 0.5       | —      | 0.560     |\n| 0.3       | —      | 0.573     |\n| **0.2**   | —      | **0.583** |\n| 0.1       | —      | 0.576     |\n\nKey finding:\nThreshold **0.2 significantly outperformed argmax and default 0.5**.\n\n---\n\n### Extended Training\n\n| Epoch | Threshold | CV     | Public LB |\n| ----- | --------- | ------ | --------- |\n| 228   | Argmax    | 0.6984 | 0.579     |\n| 228   | 0.2       | 0.6984 | 0.578     |\n| 989   | Argmax    | 0.7035 | 0.579     |\n\nObservations:\n\n* Extending training from 200 → 228 epochs improved LB.\n* Very long training (989 epochs) improved CV but did not improve LB further.\n* Performance plateaued despite increasing CV.\n\nConclusion:\nModerate training extension (200–300 epochs) is effective.\nExtremely long cascade training yields diminishing leaderboard returns.\n\n---\n\n# Inference Crop Size Study\n\nWe evaluated sliding window inference crop sizes:\n\n| Inference Crop | Public LB |\n| -------------- | --------- |\n| 360            | 0.552     |\n| 256            | 0.551     |\n| 180            | 0.508     |\n\nKey insights:\n\n* Large XY crops (256–360) perform similarly.\n* Small crops (180) severely degrade structural continuity.\n* Larger XY inference improves spatial coherence without retraining.\n\nFinal inference used:\n\n(z, y, x) = (128, 320, 320)\nThreshold = 0.15 (final optimized value)\n\n---\n\n# Ensembling\n\nA lightweight ensemble of two cascade variants (standard + patch variant) achieved:\n\n* **0.604 Public LB**\n* **0.587 Private LB**\n\nThe ensemble improved robustness while keeping the system simple.\n\n---\n\n# Overall Insights\n\n1. Long Stage 1 convergence stabilizes the cascade.\n2. Threshold tuning provides one of the largest performance gains.\n3. Cascade performance plateaus beyond moderate epoch counts.\n4. Inference patch size strongly affects spatial continuity.\n5. Controlled patch variation improves robustness.\n6. Simple ensembling is sufficient.\n\n---\nHere is the updated summarized write-up including the cascade threshold/epoch table in a clean and integrated way:\n\n---\n\n# 35th Place Silver Solution\n\n## High-Epoch nnU-Net Cascade with Patch and Threshold Optimization\n\n## Overview\n\nThis solution is based on a carefully optimized 2-stage nnU-Net cascade framework. Instead of introducing architectural complexity, performance gains were achieved through:\n\n* Long convergence training\n* Careful cascade refinement\n* Systematic threshold tuning\n* Optimized inference crop geometry\n* Lightweight ensembling\n\nThe approach prioritizes stability and controlled context expansion.\n\n---\n\n# Stage 1: High-Epoch Lower-Resolution Model\n\nA 3D nnU-Net lower-resolution model was trained for **1500 epochs** (patch size 120³).\n\nLong training was essential for producing stable coarse surface predictions that serve as reliable inputs to the cascade stage.\n\nPerformance improved steadily with training length:\n\n* 250 epochs → 0.492 LB\n* 750 epochs → 0.508 LB\n* 1200 epochs → 0.540 LB\n* 1500 epochs → 0.542 LB\n\nIncreasing inference crop size from 160 to 320 (without retraining) further improved LB from **0.542 → 0.552**, highlighting the importance of inference geometry.\n\n---\n\n# Stage 2: Cascade Full-Resolution Model\n\nThe cascade model refines predictions using:\n\n* Full-resolution image\n* Upsampled Stage 1 prediction as an additional channel\n\nBase cascade (200 epochs, argmax threshold):\n\n* CV: 0.697\n* Public LB: 0.573\n\n---\n\n# Cascade Epoch & Threshold Study\n\nWe systematically evaluated threshold and epoch effects.\n\n### 200 Epoch Model\n\n| Threshold | CV     | Public LB |\n| --------- | ------ | --------- |\n| Argmax    | 0.6973 | 0.573     |\n| 0.5       | —      | 0.560     |\n| 0.3       | —      | 0.573     |\n| **0.2**   | —      | **0.583** |\n| 0.1       | —      | 0.576     |\n\nKey finding:\nThreshold **0.2 significantly outperformed argmax and default 0.5**.\n\n---\n\n### Extended Training\n\n| Epoch | Threshold | CV     | Public LB |\n| ----- | --------- | ------ | --------- |\n| 228   | Argmax    | 0.6984 | 0.579     |\n| 228   | 0.2       | 0.6984 | 0.578     |\n| 989   | Argmax    | 0.7035 | 0.579     |\n\nObservations:\n\n* Extending training from 200 → 228 epochs improved LB.\n* Very long training (989 epochs) improved CV but did not improve LB further.\n* Performance plateaued despite increasing CV.\n\nConclusion:\nModerate training extension (200–300 epochs) is effective.\nExtremely long cascade training yields diminishing leaderboard returns.\n\n---\n\n# Inference Crop Size Study\n\nWe evaluated sliding window inference crop sizes:\n\n| Inference Crop | Public LB |\n| -------------- | --------- |\n| 360            | 0.552     |\n| 256            | 0.551     |\n| 180            | 0.508     |\n\nKey insights:\n\n* Large XY crops (256–360) perform similarly.\n* Small crops (180) severely degrade structural continuity.\n* Larger XY inference improves spatial coherence without retraining.\n\nFinal inference used:\n\n(z, y, x) = (128, 320, 320)\nThreshold = 0.15 (final optimized value)\n\n---\n\n# Ensembling\n\nA lightweight ensemble of two cascade variants (standard + patch variant) achieved:\n\n* **0.604 Public LB**\n* **0.587 Private LB**\n\nThe ensemble improved robustness while keeping the system simple.\n\n---\n\n# Overall Insights\n\n1. Long Stage 1 convergence stabilizes the cascade.\n2. Threshold tuning provides one of the largest performance gains.\n3. Cascade performance plateaus beyond moderate epoch counts.\n4. Inference patch size strongly affects spatial continuity.\n5. Controlled patch variation improves robustness.\n6. Simple ensembling is sufficient.\n\n---\n",
      "votes": 2
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3416245": "Here is the updated summarized write-up including the cascade threshold/epoch table in a clean and integrated way:\n\n---\n\n# 35th Place Silver Solution\n\n## High-Epoch nnU-Net Cascade with Patch and Threshold Optimization\n\n## Overview\n\nThis solution is based on a carefully optimized 2-stage nnU-Net cascade framework. \nFinal LB score reached * **0.604 Public LB**, * **0.587 Private LB** .\n Instead of introducing architectural complexity, performance gains were achieved through:\n\n* Long convergence training\n* Careful cascade refinement\n* Systematic threshold tuning\n* Optimized inference crop geometry\n* Lightweight ensembling\n\nThe approach prioritizes stability and controlled context expansion.\n\n---\n\n# Stage 1: High-Epoch Lower-Resolution Model\n\nA 3D nnU-Net lower-resolution model was trained for **1500 epochs** (patch size 120³).\n\nLong training was essential for producing stable coarse surface predictions that serve as reliable inputs to the cascade stage.\n\nPerformance improved steadily with training length:\n\n* 250 epochs → 0.492 LB\n* 750 epochs → 0.508 LB\n* 1200 epochs → 0.540 LB\n* 1500 epochs → 0.542 LB\n\nIncreasing inference crop size from 160 to 320 (without retraining) further improved LB from **0.542 → 0.552**, highlighting the importance of inference geometry.\n\n---\n\n# Stage 2: Cascade Full-Resolution Model\n\nThe cascade model refines predictions using:\n\n* Full-resolution image\n* Upsampled Stage 1 prediction as an additional channel\n\nBase cascade (200 epochs, argmax threshold):\n\n* CV: 0.697\n* Public LB: 0.573\n\n---\n\n# Cascade Epoch & Threshold Study\n\nWe systematically evaluated threshold and epoch effects.\n\n### 200 Epoch Model\n\n| Threshold | CV     | Public LB |\n| --------- | ------ | --------- |\n| Argmax    | 0.6973 | 0.573     |\n| 0.5       | —      | 0.560     |\n| 0.3       | —      | 0.573     |\n| **0.2**   | —      | **0.583** |\n| 0.1       | —      | 0.576     |\n\nKey finding:\nThreshold **0.2 significantly outperformed argmax and default 0.5**.\n\n---\n\n### Extended Training\n\n| Epoch | Threshold | CV     | Public LB |\n| ----- | --------- | ------ | --------- |\n| 228   | Argmax    | 0.6984 | 0.579     |\n| 228   | 0.2       | 0.6984 | 0.578     |\n| 989   | Argmax    | 0.7035 | 0.579     |\n\nObservations:\n\n* Extending training from 200 → 228 epochs improved LB.\n* Very long training (989 epochs) improved CV but did not improve LB further.\n* Performance plateaued despite increasing CV.\n\nConclusion:\nModerate training extension (200–300 epochs) is effective.\nExtremely long cascade training yields diminishing leaderboard returns.\n\n---\n\n# Inference Crop Size Study\n\nWe evaluated sliding window inference crop sizes:\n\n| Inference Crop | Public LB |\n| -------------- | --------- |\n| 360            | 0.552     |\n| 256            | 0.551     |\n| 180            | 0.508     |\n\nKey insights:\n\n* Large XY crops (256–360) perform similarly.\n* Small crops (180) severely degrade structural continuity.\n* Larger XY inference improves spatial coherence without retraining.\n\nFinal inference used:\n\n(z, y, x) = (128, 320, 320)\nThreshold = 0.15 (final optimized value)\n\n---\n\n# Ensembling\n\nA lightweight ensemble of two cascade variants (standard + patch variant) achieved:\n\n* **0.604 Public LB**\n* **0.587 Private LB**\n\nThe ensemble improved robustness while keeping the system simple.\n\n---\n\n# Overall Insights\n\n1. Long Stage 1 convergence stabilizes the cascade.\n2. Threshold tuning provides one of the largest performance gains.\n3. Cascade performance plateaus beyond moderate epoch counts.\n4. Inference patch size strongly affects spatial continuity.\n5. Controlled patch variation improves robustness.\n6. Simple ensembling is sufficient.\n\n---\nHere is the updated summarized write-up including the cascade threshold/epoch table in a clean and integrated way:\n\n---\n\n# 35th Place Silver Solution\n\n## High-Epoch nnU-Net Cascade with Patch and Threshold Optimization\n\n## Overview\n\nThis solution is based on a carefully optimized 2-stage nnU-Net cascade framework. Instead of introducing architectural complexity, performance gains were achieved through:\n\n* Long convergence training\n* Careful cascade refinement\n* Systematic threshold tuning\n* Optimized inference crop geometry\n* Lightweight ensembling\n\nThe approach prioritizes stability and controlled context expansion.\n\n---\n\n# Stage 1: High-Epoch Lower-Resolution Model\n\nA 3D nnU-Net lower-resolution model was trained for **1500 epochs** (patch size 120³).\n\nLong training was essential for producing stable coarse surface predictions that serve as reliable inputs to the cascade stage.\n\nPerformance improved steadily with training length:\n\n* 250 epochs → 0.492 LB\n* 750 epochs → 0.508 LB\n* 1200 epochs → 0.540 LB\n* 1500 epochs → 0.542 LB\n\nIncreasing inference crop size from 160 to 320 (without retraining) further improved LB from **0.542 → 0.552**, highlighting the importance of inference geometry.\n\n---\n\n# Stage 2: Cascade Full-Resolution Model\n\nThe cascade model refines predictions using:\n\n* Full-resolution image\n* Upsampled Stage 1 prediction as an additional channel\n\nBase cascade (200 epochs, argmax threshold):\n\n* CV: 0.697\n* Public LB: 0.573\n\n---\n\n# Cascade Epoch & Threshold Study\n\nWe systematically evaluated threshold and epoch effects.\n\n### 200 Epoch Model\n\n| Threshold | CV     | Public LB |\n| --------- | ------ | --------- |\n| Argmax    | 0.6973 | 0.573     |\n| 0.5       | —      | 0.560     |\n| 0.3       | —      | 0.573     |\n| **0.2**   | —      | **0.583** |\n| 0.1       | —      | 0.576     |\n\nKey finding:\nThreshold **0.2 significantly outperformed argmax and default 0.5**.\n\n---\n\n### Extended Training\n\n| Epoch | Threshold | CV     | Public LB |\n| ----- | --------- | ------ | --------- |\n| 228   | Argmax    | 0.6984 | 0.579     |\n| 228   | 0.2       | 0.6984 | 0.578     |\n| 989   | Argmax    | 0.7035 | 0.579     |\n\nObservations:\n\n* Extending training from 200 → 228 epochs improved LB.\n* Very long training (989 epochs) improved CV but did not improve LB further.\n* Performance plateaued despite increasing CV.\n\nConclusion:\nModerate training extension (200–300 epochs) is effective.\nExtremely long cascade training yields diminishing leaderboard returns.\n\n---\n\n# Inference Crop Size Study\n\nWe evaluated sliding window inference crop sizes:\n\n| Inference Crop | Public LB |\n| -------------- | --------- |\n| 360            | 0.552     |\n| 256            | 0.551     |\n| 180            | 0.508     |\n\nKey insights:\n\n* Large XY crops (256–360) perform similarly.\n* Small crops (180) severely degrade structural continuity.\n* Larger XY inference improves spatial coherence without retraining.\n\nFinal inference used:\n\n(z, y, x) = (128, 320, 320)\nThreshold = 0.15 (final optimized value)\n\n---\n\n# Ensembling\n\nA lightweight ensemble of two cascade variants (standard + patch variant) achieved:\n\n* **0.604 Public LB**\n* **0.587 Private LB**\n\nThe ensemble improved robustness while keeping the system simple.\n\n---\n\n# Overall Insights\n\n1. Long Stage 1 convergence stabilizes the cascade.\n2. Threshold tuning provides one of the largest performance gains.\n3. Cascade performance plateaus beyond moderate epoch counts.\n4. Inference patch size strongly affects spatial continuity.\n5. Controlled patch variation improves robustness.\n6. Simple ensembling is sufficient.\n\n---\n"
  }
}