{
  "id": 679371,
  "title": "Why are long epochs needed? I'd like to discuss this.",
  "url": "/competitions/vesuvius-challenge-surface-detection/discussion/679371",
  "author_name": "",
  "post_date": "2026-03-01T00:34:22.481386900Z",
  "votes": 8,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I ran a few experiments based on my hypothesis, and I'd like to discuss the results.\nMy experimental setup may have some flaws or odd choices, but I hope this can spark a useful discussion.</p>\n<h2>Background</h2>\n<p>Surveying the top solutions of Vesuvius Challenge 2, we found that all top teams employed <strong>2,000-8,000 epochs</strong> of training (1st: 4,000ep, 3rd: 8,000ep, 12th: 2,000-4,000ep). We wanted to understand why such long training was effective.</p>\n<p>In our experiments, after training on all data (786 samples), we ran inference on the training set itself and found that predictions were significantly worse on low-quality images (high brightness, blur, low contrast). This led us to investigate the data distribution more closely.</p>\n<ul>\n<li><strong>Normal</strong>: High contrast, high sharpness (545 samples, 69%)</li>\n<li><strong>A:Bright</strong>: High intensity, low contrast (132 samples, 17%)</li>\n<li><strong>B:Blur</strong>: Blurry, low sharpness (39 samples, 5%)</li>\n<li><strong>C:LowContrast</strong>: Low contrast (58 samples, 7%)</li>\n<li><strong>B-:MildBlur</strong>: Mildly blurry (12 samples, 2%)</li>\n</ul>\n<p><strong>Hypothesis</strong>: Difficult data patterns (Bright, Blur, etc.) converge more slowly, requiring many more epochs to learn adequately.</p>\n<h2>Preliminary Analysis: t-SNE Data Visualization</h2>\n<p>To investigate the source of the performance gap, we visualized the feature space of all 786 samples. From each 3D volume, we extracted the central 5 slices along XY, XZ, and YZ planes, fed them through an ImageNet-pretrained ConvNeXt-Tiny (timm) to obtain feature vectors, and averaged across the 3 views (768 dimensions). We then applied t-SNE (perplexity=30, PCA initialization) and clustered with KMeans (k=8).</p>\n<p>The result showed clear separation by image quality pattern:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2Fc3849585c995f9ec2e0e597b80cfcfd2%2F00_tsne_patterns.png?generation=1772324957059654&amp;alt=media\" alt=\"t-SNE Patterns\"></p>\n<p>High-brightness samples (A:Bright) cluster in the bottom-right, blurry/low-contrast samples (B:Blur, C:LowContrast) in the bottom-left, and Normal data spreads across the upper-center region. This separation aligns strongly with image quality metrics (sharpness, contrast, mean intensity), and samples with lower model performance tend to belong to the peripheral clusters.</p>\n<h2>Experimental Setup</h2>\n<p>To test this hypothesis, we fine-tuned a pre-trained model (Phase 1, 200 epochs) on the full Phase 2 dataset (786 samples) for only 100 epochs, recording per-sample loss at every epoch. The 100-epoch run was designed as a measurement-oriented short run to observe early-stage training dynamics (production training used 500-4,000 epochs).</p>\n<h3>Pre-training</h3>\n<p>We used a checkpoint pre-trained for 200 epochs on Phase 1 data (clean labels, no ignore regions). Pre-training accelerates initial convergence.</p>\n<table>\n<thead>\n<tr>\n<th>Setting</th>\n<th>Value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Phase 1 Data</td>\n<td>Dataset120_Phase1</td>\n</tr>\n<tr>\n<td>Phase 1 Epochs</td>\n<td>200</td>\n</tr>\n<tr>\n<td>Phase 1 Trainer</td>\n<td>nnUNetTrainerSkelRecall_200ep</td>\n</tr>\n<tr>\n<td>Phase 1 Loss</td>\n<td>CE + Dice + SkelRecall</td>\n</tr>\n<tr>\n<td>Phase 1 LR</td>\n<td>0.005 (cosine annealing)</td>\n</tr>\n</tbody>\n</table>\n<h3>Main Experiment (Phase 2 Fine-tune + Loss Tracking)</h3>\n<table>\n<thead>\n<tr>\n<th>Setting</th>\n<th>Value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Base Model</td>\n<td>nnUNetTrainerSkelRecall (CE + Dice + SkelRecall)</td>\n</tr>\n<tr>\n<td>Pre-trained Weights</td>\n<td>Phase 1 200ep checkpoint_final.pth</td>\n</tr>\n<tr>\n<td>Epochs</td>\n<td>100 (measurement run; production: 500-4,000ep)</td>\n</tr>\n<tr>\n<td>Fold</td>\n<td>all (all 786 samples used for training)</td>\n</tr>\n<tr>\n<td>Patch Size</td>\n<td>128 x 160 x 160</td>\n</tr>\n<tr>\n<td>Batch Size</td>\n<td>5</td>\n</tr>\n<tr>\n<td>LR</td>\n<td>0.005 (cosine annealing)</td>\n</tr>\n<tr>\n<td>Loss Logging</td>\n<td>Per-sample loss recorded every epoch</td>\n</tr>\n<tr>\n<td>Clustering</td>\n<td>t-SNE + KMeans (k=8)</td>\n</tr>\n</tbody>\n</table>\n<h3>Data Augmentation</h3>\n<p>We used nnU-Net defaults plus domain-specific custom augmentations:</p>\n<p><strong>Spatial transforms:</strong></p>\n<ul>\n<li>Elastic deformation (p=0.3, magnitude 10-50 voxels)</li>\n<li>Rotation (p=0.5, +/-15 degrees per axis)</li>\n<li>Scaling (p=0.2, 0.7x-1.4x)</li>\n<li>Mirror (all 3 axes)</li>\n</ul>\n<p><strong>Intensity transforms:</strong></p>\n<ul>\n<li>Gaussian Noise (p=0.15, variance 0-0.15)</li>\n<li>Gaussian Blur (p=0.2, sigma 0.5-1.5)</li>\n<li>Brightness (p=0.15, 0.5x-1.5x)</li>\n<li>Contrast (p=0.15, 0.5x-1.5x)</li>\n<li>Low Resolution Simulation (p=0.25, scale 0.25-1.0)</li>\n<li>Gamma (p=0.3, gamma 0.7-1.5)</li>\n</ul>\n<p><strong>Custom transforms (papyrus-specific):</strong></p>\n<ul>\n<li>BlankRectangle (p=0.4): Blank out rectangular regions to simulate damage</li>\n<li>Smear (p=0.2): Z-axis ink bleeding simulation</li>\n<li>InhomogeneousIllumination (p=0.25): Uneven lighting simulation</li>\n</ul>\n<h2>Results</h2>\n<h3>Per-Pattern Loss Curves</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2F7588172f89a612044ba9115c0731d4e2%2F01_loss_by_pattern.png?generation=1772325014361578&amp;alt=media\" alt=\"\"></p>\n<table>\n<thead>\n<tr>\n<th>Pattern</th>\n<th>N</th>\n<th>Initial Loss</th>\n<th>Final Loss</th>\n<th>Improvement</th>\n<th>90% Convergence Epoch</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Normal</td>\n<td>545</td>\n<td>-0.883</td>\n<td>-1.111</td>\n<td>0.228</td>\n<td>60</td>\n</tr>\n<tr>\n<td>A:Bright</td>\n<td>132</td>\n<td>-0.407</td>\n<td>-0.602</td>\n<td>0.195</td>\n<td>33</td>\n</tr>\n<tr>\n<td>B:Blur</td>\n<td>39</td>\n<td>-0.662</td>\n<td>-0.815</td>\n<td>0.153</td>\n<td>54</td>\n</tr>\n<tr>\n<td>C:LowContrast</td>\n<td>58</td>\n<td>-0.605</td>\n<td>-0.769</td>\n<td>0.164</td>\n<td>96</td>\n</tr>\n<tr>\n<td>B-:MildBlur</td>\n<td>12</td>\n<td>-0.833</td>\n<td>-1.163</td>\n<td>0.330</td>\n<td>25</td>\n</tr>\n</tbody>\n</table>\n<h3>Per-Cluster Loss Curves</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2F4238e7c0c0f3c354cb489d420a7f30e2%2F02_loss_by_cluster.png?generation=1772325035027860&amp;alt=media\" alt=\"\"></p>\n<table>\n<thead>\n<tr>\n<th>Cluster</th>\n<th>N</th>\n<th>Initial Loss</th>\n<th>Final Loss</th>\n<th>Improvement</th>\n<th>Dominant Pattern</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>C0 (Normal)</td>\n<td>101</td>\n<td>-0.947</td>\n<td>-1.118</td>\n<td>0.171</td>\n<td>Normal:101</td>\n</tr>\n<tr>\n<td>C1 (Normal)</td>\n<td>87</td>\n<td>-0.793</td>\n<td>-1.013</td>\n<td>0.220</td>\n<td>Normal:86</td>\n</tr>\n<tr>\n<td>C2 (Normal)</td>\n<td>85</td>\n<td>-0.958</td>\n<td>-1.254</td>\n<td>0.296</td>\n<td>Normal:70, LowContrast:15</td>\n</tr>\n<tr>\n<td>C3 (Normal)</td>\n<td>105</td>\n<td>-0.814</td>\n<td>-1.024</td>\n<td>0.210</td>\n<td>Normal:93, Blur:8</td>\n</tr>\n<tr>\n<td>C4 (Normal)</td>\n<td>97</td>\n<td>-0.864</td>\n<td>-1.105</td>\n<td>0.241</td>\n<td>Normal:96</td>\n</tr>\n<tr>\n<td>C5 (Mixed)</td>\n<td>100</td>\n<td>-0.657</td>\n<td>-0.881</td>\n<td>0.224</td>\n<td>Normal:58, Bright:34</td>\n</tr>\n<tr>\n<td>C6 (Bright)</td>\n<td>101</td>\n<td>-0.382</td>\n<td>-0.586</td>\n<td>0.204</td>\n<td><strong>A:Bright:97</strong></td>\n</tr>\n<tr>\n<td>C7 (Mixed)</td>\n<td>110</td>\n<td>-0.783</td>\n<td>-0.961</td>\n<td>0.178</td>\n<td>LowContrast:38, Normal:37, Blur:31</td>\n</tr>\n</tbody>\n</table>\n<h3>Overall Summary</h3>\n<table>\n<thead>\n<tr>\n<th>Metric</th>\n<th>Value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Overall Initial Loss</td>\n<td>-0.771</td>\n</tr>\n<tr>\n<td>Overall Final Loss</td>\n<td>-0.986</td>\n</tr>\n<tr>\n<td>Overall Improvement</td>\n<td>0.216</td>\n</tr>\n<tr>\n<td>Worst Cluster (C6) Final Loss</td>\n<td>-0.586</td>\n</tr>\n<tr>\n<td>Best Cluster (C2) Final Loss</td>\n<td>-1.254</td>\n</tr>\n<tr>\n<td>Worst-Best Gap</td>\n<td><strong>0.669</strong></td>\n</tr>\n</tbody>\n</table>\n<h2>Detailed Analysis</h2>\n<h3>Final Loss Distribution by Cluster</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2F5608232a866f857d032a51e79f13bfe4%2F04_final_loss_boxplot.png?generation=1772325058699998&amp;alt=media\" alt=\"\"></p>\n<h3>Improvement Distribution by Cluster</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2F57a7d9a9a05a370203254b9b36baddeb%2F05_improvement_boxplot.png?generation=1772325075898632&amp;alt=media\" alt=\"\"></p>\n<h3>Top 30 Hardest Samples</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2F11e9b1a508337154ca886287bc134931%2F06_hardest_samples.png?generation=1772325090638670&amp;alt=media\" alt=\"\"></p>\n<h3>Loss Distribution Change on t-SNE Map</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2F580300c261c928dd6aaeac840cb32df0%2F08_tsne_loss_progression.png?generation=1772325111687075&amp;alt=media\" alt=\"\"></p>\n<p>The A:Bright cluster (bottom-right) remains red (high loss) even after training. The Normal clusters (upper region) transition to green (low loss) early in training.</p>\n<h2>Key Findings</h2>\n<h3>1. A:Bright Samples Show Severe Convergence Lag</h3>\n<p>The final loss of A:Bright after 100 epochs (<strong>-0.602</strong>) is <strong>0.281 worse</strong> than Normal's <em>initial</em> loss (<strong>-0.883</strong>). In other words, after 100 epochs of training, Bright samples haven't even reached the level Normal data starts at.</p>\n<p>This means learning on Bright data is fundamentally harder, and reaching Normal-level performance requires <strong>many more epochs</strong>.</p>\n<h3>2. Convergence Speed Varies Dramatically Across Patterns</h3>\n<p>Epoch at which 90% of total improvement is reached:</p>\n<ul>\n<li>B-:MildBlur: <strong>25ep</strong> (fastest; small sample count, high variance)</li>\n<li>A:Bright: <strong>33ep</strong> (quick plateau at a poor absolute level = 100ep is not enough)</li>\n<li>B:Blur: <strong>54ep</strong></li>\n<li>Normal: <strong>60ep</strong></li>\n<li>C:LowContrast: <strong>96ep</strong> (slowest; barely converging within 100ep)</li>\n</ul>\n<p>The early 90% convergence of A:Bright is misleading --- it reflects a \"plateau at a low level\" rather than fast learning. The total improvement is small, so the 90% threshold is reached early, but the absolute loss remains high.</p>\n<h3>3. Cluster C6 Is an Outlier in Difficulty</h3>\n<p>C6 (A:Bright 97/101) has a final loss of -0.586, with a 0.3-0.7 gap from all other clusters. This cluster acts as the bottleneck for overall training efficiency.</p>\n<h3>4. LowContrast Has the Slowest Convergence</h3>\n<p>C:LowContrast requires <strong>96 epochs</strong> to reach 90% convergence. At 100 epochs, it has barely made it. Significant further improvement can be expected with 500+ epochs.</p>\n<h2>Why Long Training Works (Discussion)</h2>\n<ol>\n<li><p><strong>Data diversity</strong>: 17% (Bright) + 7% (LowContrast) + 5% (Blur) of the 786 samples are difficult. Roughly 30% of the data has not reached Normal-level convergence at 100 epochs.</p></li>\n<li><p><strong>Batch size constraint</strong>: nnU-Net's batch size is 5 (GPU memory limited). At 250 iterations per epoch = 1,250 patches, only a fraction of the 786 samples are seen each epoch. Difficult samples require many epochs to be sampled sufficiently.</p></li>\n<li><p><strong>Gradient direction conflicts</strong>: Normal and Bright data may require different optimization directions. Short training biases toward the majority Normal pattern, leaving minority difficult patterns under-trained.</p></li>\n<li><p><strong>2,000-8,000 epochs is justified</strong>: A:Bright's loss at 100 epochs (-0.602) hasn't even reached Normal's initial value (-0.883). Closing this gap requires training on the order of thousands of epochs.</p></li>\n</ol>\n<h2>Conclusion</h2>\n<p>This analysis quantitatively confirmed that <strong>training convergence speed varies significantly across image quality patterns</strong>. A:Bright (high-brightness) data did not reach Normal's initial loss even after 100 epochs, supporting the rationale behind the 2,000-8,000 epoch training adopted by top teams.</p>\n<p>Based on this hypothesis, we attempted a <strong>specialist model</strong> approach: separating Blur/Bright/LowContrast samples via rule-based classification and fine-tuning dedicated models for each pattern, with quality-based routing at inference time. However, rule-based domain separation had fundamental limitations --- boundaries between patterns were ambiguous, and borderline samples were difficult to handle correctly. </p>\n<p>Looking back at the top teams' approaches, the solution to difficult samples was not model separation but <strong>sufficient training epochs to lift overall performance</strong>. This insight is valuable for training strategies in 3D segmentation tasks with high data diversity.</p>",
  "messages": [
    {
      "id": "3415494",
      "postDate": "03/01/2026 00:34:22",
      "content": "<p>I ran a few experiments based on my hypothesis, and I'd like to discuss the results.\nMy experimental setup may have some flaws or odd choices, but I hope this can spark a useful discussion.</p>\n<h2>Background</h2>\n<p>Surveying the top solutions of Vesuvius Challenge 2, we found that all top teams employed <strong>2,000-8,000 epochs</strong> of training (1st: 4,000ep, 3rd: 8,000ep, 12th: 2,000-4,000ep). We wanted to understand why such long training was effective.</p>\n<p>In our experiments, after training on all data (786 samples), we ran inference on the training set itself and found that predictions were significantly worse on low-quality images (high brightness, blur, low contrast). This led us to investigate the data distribution more closely.</p>\n<ul>\n<li><strong>Normal</strong>: High contrast, high sharpness (545 samples, 69%)</li>\n<li><strong>A:Bright</strong>: High intensity, low contrast (132 samples, 17%)</li>\n<li><strong>B:Blur</strong>: Blurry, low sharpness (39 samples, 5%)</li>\n<li><strong>C:LowContrast</strong>: Low contrast (58 samples, 7%)</li>\n<li><strong>B-:MildBlur</strong>: Mildly blurry (12 samples, 2%)</li>\n</ul>\n<p><strong>Hypothesis</strong>: Difficult data patterns (Bright, Blur, etc.) converge more slowly, requiring many more epochs to learn adequately.</p>\n<h2>Preliminary Analysis: t-SNE Data Visualization</h2>\n<p>To investigate the source of the performance gap, we visualized the feature space of all 786 samples. From each 3D volume, we extracted the central 5 slices along XY, XZ, and YZ planes, fed them through an ImageNet-pretrained ConvNeXt-Tiny (timm) to obtain feature vectors, and averaged across the 3 views (768 dimensions). We then applied t-SNE (perplexity=30, PCA initialization) and clustered with KMeans (k=8).</p>\n<p>The result showed clear separation by image quality pattern:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2Fc3849585c995f9ec2e0e597b80cfcfd2%2F00_tsne_patterns.png?generation=1772324957059654&amp;alt=media\" alt=\"t-SNE Patterns\"></p>\n<p>High-brightness samples (A:Bright) cluster in the bottom-right, blurry/low-contrast samples (B:Blur, C:LowContrast) in the bottom-left, and Normal data spreads across the upper-center region. This separation aligns strongly with image quality metrics (sharpness, contrast, mean intensity), and samples with lower model performance tend to belong to the peripheral clusters.</p>\n<h2>Experimental Setup</h2>\n<p>To test this hypothesis, we fine-tuned a pre-trained model (Phase 1, 200 epochs) on the full Phase 2 dataset (786 samples) for only 100 epochs, recording per-sample loss at every epoch. The 100-epoch run was designed as a measurement-oriented short run to observe early-stage training dynamics (production training used 500-4,000 epochs).</p>\n<h3>Pre-training</h3>\n<p>We used a checkpoint pre-trained for 200 epochs on Phase 1 data (clean labels, no ignore regions). Pre-training accelerates initial convergence.</p>\n<table>\n<thead>\n<tr>\n<th>Setting</th>\n<th>Value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Phase 1 Data</td>\n<td>Dataset120_Phase1</td>\n</tr>\n<tr>\n<td>Phase 1 Epochs</td>\n<td>200</td>\n</tr>\n<tr>\n<td>Phase 1 Trainer</td>\n<td>nnUNetTrainerSkelRecall_200ep</td>\n</tr>\n<tr>\n<td>Phase 1 Loss</td>\n<td>CE + Dice + SkelRecall</td>\n</tr>\n<tr>\n<td>Phase 1 LR</td>\n<td>0.005 (cosine annealing)</td>\n</tr>\n</tbody>\n</table>\n<h3>Main Experiment (Phase 2 Fine-tune + Loss Tracking)</h3>\n<table>\n<thead>\n<tr>\n<th>Setting</th>\n<th>Value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Base Model</td>\n<td>nnUNetTrainerSkelRecall (CE + Dice + SkelRecall)</td>\n</tr>\n<tr>\n<td>Pre-trained Weights</td>\n<td>Phase 1 200ep checkpoint_final.pth</td>\n</tr>\n<tr>\n<td>Epochs</td>\n<td>100 (measurement run; production: 500-4,000ep)</td>\n</tr>\n<tr>\n<td>Fold</td>\n<td>all (all 786 samples used for training)</td>\n</tr>\n<tr>\n<td>Patch Size</td>\n<td>128 x 160 x 160</td>\n</tr>\n<tr>\n<td>Batch Size</td>\n<td>5</td>\n</tr>\n<tr>\n<td>LR</td>\n<td>0.005 (cosine annealing)</td>\n</tr>\n<tr>\n<td>Loss Logging</td>\n<td>Per-sample loss recorded every epoch</td>\n</tr>\n<tr>\n<td>Clustering</td>\n<td>t-SNE + KMeans (k=8)</td>\n</tr>\n</tbody>\n</table>\n<h3>Data Augmentation</h3>\n<p>We used nnU-Net defaults plus domain-specific custom augmentations:</p>\n<p><strong>Spatial transforms:</strong></p>\n<ul>\n<li>Elastic deformation (p=0.3, magnitude 10-50 voxels)</li>\n<li>Rotation (p=0.5, +/-15 degrees per axis)</li>\n<li>Scaling (p=0.2, 0.7x-1.4x)</li>\n<li>Mirror (all 3 axes)</li>\n</ul>\n<p><strong>Intensity transforms:</strong></p>\n<ul>\n<li>Gaussian Noise (p=0.15, variance 0-0.15)</li>\n<li>Gaussian Blur (p=0.2, sigma 0.5-1.5)</li>\n<li>Brightness (p=0.15, 0.5x-1.5x)</li>\n<li>Contrast (p=0.15, 0.5x-1.5x)</li>\n<li>Low Resolution Simulation (p=0.25, scale 0.25-1.0)</li>\n<li>Gamma (p=0.3, gamma 0.7-1.5)</li>\n</ul>\n<p><strong>Custom transforms (papyrus-specific):</strong></p>\n<ul>\n<li>BlankRectangle (p=0.4): Blank out rectangular regions to simulate damage</li>\n<li>Smear (p=0.2): Z-axis ink bleeding simulation</li>\n<li>InhomogeneousIllumination (p=0.25): Uneven lighting simulation</li>\n</ul>\n<h2>Results</h2>\n<h3>Per-Pattern Loss Curves</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2F7588172f89a612044ba9115c0731d4e2%2F01_loss_by_pattern.png?generation=1772325014361578&amp;alt=media\" alt=\"\"></p>\n<table>\n<thead>\n<tr>\n<th>Pattern</th>\n<th>N</th>\n<th>Initial Loss</th>\n<th>Final Loss</th>\n<th>Improvement</th>\n<th>90% Convergence Epoch</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Normal</td>\n<td>545</td>\n<td>-0.883</td>\n<td>-1.111</td>\n<td>0.228</td>\n<td>60</td>\n</tr>\n<tr>\n<td>A:Bright</td>\n<td>132</td>\n<td>-0.407</td>\n<td>-0.602</td>\n<td>0.195</td>\n<td>33</td>\n</tr>\n<tr>\n<td>B:Blur</td>\n<td>39</td>\n<td>-0.662</td>\n<td>-0.815</td>\n<td>0.153</td>\n<td>54</td>\n</tr>\n<tr>\n<td>C:LowContrast</td>\n<td>58</td>\n<td>-0.605</td>\n<td>-0.769</td>\n<td>0.164</td>\n<td>96</td>\n</tr>\n<tr>\n<td>B-:MildBlur</td>\n<td>12</td>\n<td>-0.833</td>\n<td>-1.163</td>\n<td>0.330</td>\n<td>25</td>\n</tr>\n</tbody>\n</table>\n<h3>Per-Cluster Loss Curves</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2F4238e7c0c0f3c354cb489d420a7f30e2%2F02_loss_by_cluster.png?generation=1772325035027860&amp;alt=media\" alt=\"\"></p>\n<table>\n<thead>\n<tr>\n<th>Cluster</th>\n<th>N</th>\n<th>Initial Loss</th>\n<th>Final Loss</th>\n<th>Improvement</th>\n<th>Dominant Pattern</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>C0 (Normal)</td>\n<td>101</td>\n<td>-0.947</td>\n<td>-1.118</td>\n<td>0.171</td>\n<td>Normal:101</td>\n</tr>\n<tr>\n<td>C1 (Normal)</td>\n<td>87</td>\n<td>-0.793</td>\n<td>-1.013</td>\n<td>0.220</td>\n<td>Normal:86</td>\n</tr>\n<tr>\n<td>C2 (Normal)</td>\n<td>85</td>\n<td>-0.958</td>\n<td>-1.254</td>\n<td>0.296</td>\n<td>Normal:70, LowContrast:15</td>\n</tr>\n<tr>\n<td>C3 (Normal)</td>\n<td>105</td>\n<td>-0.814</td>\n<td>-1.024</td>\n<td>0.210</td>\n<td>Normal:93, Blur:8</td>\n</tr>\n<tr>\n<td>C4 (Normal)</td>\n<td>97</td>\n<td>-0.864</td>\n<td>-1.105</td>\n<td>0.241</td>\n<td>Normal:96</td>\n</tr>\n<tr>\n<td>C5 (Mixed)</td>\n<td>100</td>\n<td>-0.657</td>\n<td>-0.881</td>\n<td>0.224</td>\n<td>Normal:58, Bright:34</td>\n</tr>\n<tr>\n<td>C6 (Bright)</td>\n<td>101</td>\n<td>-0.382</td>\n<td>-0.586</td>\n<td>0.204</td>\n<td><strong>A:Bright:97</strong></td>\n</tr>\n<tr>\n<td>C7 (Mixed)</td>\n<td>110</td>\n<td>-0.783</td>\n<td>-0.961</td>\n<td>0.178</td>\n<td>LowContrast:38, Normal:37, Blur:31</td>\n</tr>\n</tbody>\n</table>\n<h3>Overall Summary</h3>\n<table>\n<thead>\n<tr>\n<th>Metric</th>\n<th>Value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Overall Initial Loss</td>\n<td>-0.771</td>\n</tr>\n<tr>\n<td>Overall Final Loss</td>\n<td>-0.986</td>\n</tr>\n<tr>\n<td>Overall Improvement</td>\n<td>0.216</td>\n</tr>\n<tr>\n<td>Worst Cluster (C6) Final Loss</td>\n<td>-0.586</td>\n</tr>\n<tr>\n<td>Best Cluster (C2) Final Loss</td>\n<td>-1.254</td>\n</tr>\n<tr>\n<td>Worst-Best Gap</td>\n<td><strong>0.669</strong></td>\n</tr>\n</tbody>\n</table>\n<h2>Detailed Analysis</h2>\n<h3>Final Loss Distribution by Cluster</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2F5608232a866f857d032a51e79f13bfe4%2F04_final_loss_boxplot.png?generation=1772325058699998&amp;alt=media\" alt=\"\"></p>\n<h3>Improvement Distribution by Cluster</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2F57a7d9a9a05a370203254b9b36baddeb%2F05_improvement_boxplot.png?generation=1772325075898632&amp;alt=media\" alt=\"\"></p>\n<h3>Top 30 Hardest Samples</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2F11e9b1a508337154ca886287bc134931%2F06_hardest_samples.png?generation=1772325090638670&amp;alt=media\" alt=\"\"></p>\n<h3>Loss Distribution Change on t-SNE Map</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2F580300c261c928dd6aaeac840cb32df0%2F08_tsne_loss_progression.png?generation=1772325111687075&amp;alt=media\" alt=\"\"></p>\n<p>The A:Bright cluster (bottom-right) remains red (high loss) even after training. The Normal clusters (upper region) transition to green (low loss) early in training.</p>\n<h2>Key Findings</h2>\n<h3>1. A:Bright Samples Show Severe Convergence Lag</h3>\n<p>The final loss of A:Bright after 100 epochs (<strong>-0.602</strong>) is <strong>0.281 worse</strong> than Normal's <em>initial</em> loss (<strong>-0.883</strong>). In other words, after 100 epochs of training, Bright samples haven't even reached the level Normal data starts at.</p>\n<p>This means learning on Bright data is fundamentally harder, and reaching Normal-level performance requires <strong>many more epochs</strong>.</p>\n<h3>2. Convergence Speed Varies Dramatically Across Patterns</h3>\n<p>Epoch at which 90% of total improvement is reached:</p>\n<ul>\n<li>B-:MildBlur: <strong>25ep</strong> (fastest; small sample count, high variance)</li>\n<li>A:Bright: <strong>33ep</strong> (quick plateau at a poor absolute level = 100ep is not enough)</li>\n<li>B:Blur: <strong>54ep</strong></li>\n<li>Normal: <strong>60ep</strong></li>\n<li>C:LowContrast: <strong>96ep</strong> (slowest; barely converging within 100ep)</li>\n</ul>\n<p>The early 90% convergence of A:Bright is misleading --- it reflects a \"plateau at a low level\" rather than fast learning. The total improvement is small, so the 90% threshold is reached early, but the absolute loss remains high.</p>\n<h3>3. Cluster C6 Is an Outlier in Difficulty</h3>\n<p>C6 (A:Bright 97/101) has a final loss of -0.586, with a 0.3-0.7 gap from all other clusters. This cluster acts as the bottleneck for overall training efficiency.</p>\n<h3>4. LowContrast Has the Slowest Convergence</h3>\n<p>C:LowContrast requires <strong>96 epochs</strong> to reach 90% convergence. At 100 epochs, it has barely made it. Significant further improvement can be expected with 500+ epochs.</p>\n<h2>Why Long Training Works (Discussion)</h2>\n<ol>\n<li><p><strong>Data diversity</strong>: 17% (Bright) + 7% (LowContrast) + 5% (Blur) of the 786 samples are difficult. Roughly 30% of the data has not reached Normal-level convergence at 100 epochs.</p></li>\n<li><p><strong>Batch size constraint</strong>: nnU-Net's batch size is 5 (GPU memory limited). At 250 iterations per epoch = 1,250 patches, only a fraction of the 786 samples are seen each epoch. Difficult samples require many epochs to be sampled sufficiently.</p></li>\n<li><p><strong>Gradient direction conflicts</strong>: Normal and Bright data may require different optimization directions. Short training biases toward the majority Normal pattern, leaving minority difficult patterns under-trained.</p></li>\n<li><p><strong>2,000-8,000 epochs is justified</strong>: A:Bright's loss at 100 epochs (-0.602) hasn't even reached Normal's initial value (-0.883). Closing this gap requires training on the order of thousands of epochs.</p></li>\n</ol>\n<h2>Conclusion</h2>\n<p>This analysis quantitatively confirmed that <strong>training convergence speed varies significantly across image quality patterns</strong>. A:Bright (high-brightness) data did not reach Normal's initial loss even after 100 epochs, supporting the rationale behind the 2,000-8,000 epoch training adopted by top teams.</p>\n<p>Based on this hypothesis, we attempted a <strong>specialist model</strong> approach: separating Blur/Bright/LowContrast samples via rule-based classification and fine-tuning dedicated models for each pattern, with quality-based routing at inference time. However, rule-based domain separation had fundamental limitations --- boundaries between patterns were ambiguous, and borderline samples were difficult to handle correctly. </p>\n<p>Looking back at the top teams' approaches, the solution to difficult samples was not model separation but <strong>sufficient training epochs to lift overall performance</strong>. This insight is valuable for training strategies in 3D segmentation tasks with high data diversity.</p>",
      "rawMarkdown": "I ran a few experiments based on my hypothesis, and I'd like to discuss the results.\nMy experimental setup may have some flaws or odd choices, but I hope this can spark a useful discussion.\n\n\n## Background\n\nSurveying the top solutions of Vesuvius Challenge 2, we found that all top teams employed **2,000-8,000 epochs** of training (1st: 4,000ep, 3rd: 8,000ep, 12th: 2,000-4,000ep). We wanted to understand why such long training was effective.\n\nIn our experiments, after training on all data (786 samples), we ran inference on the training set itself and found that predictions were significantly worse on low-quality images (high brightness, blur, low contrast). This led us to investigate the data distribution more closely.\n\n- **Normal**: High contrast, high sharpness (545 samples, 69%)\n- **A:Bright**: High intensity, low contrast (132 samples, 17%)\n- **B:Blur**: Blurry, low sharpness (39 samples, 5%)\n- **C:LowContrast**: Low contrast (58 samples, 7%)\n- **B-:MildBlur**: Mildly blurry (12 samples, 2%)\n\n**Hypothesis**: Difficult data patterns (Bright, Blur, etc.) converge more slowly, requiring many more epochs to learn adequately.\n\n## Preliminary Analysis: t-SNE Data Visualization\n\nTo investigate the source of the performance gap, we visualized the feature space of all 786 samples. From each 3D volume, we extracted the central 5 slices along XY, XZ, and YZ planes, fed them through an ImageNet-pretrained ConvNeXt-Tiny (timm) to obtain feature vectors, and averaged across the 3 views (768 dimensions). We then applied t-SNE (perplexity=30, PCA initialization) and clustered with KMeans (k=8).\n\nThe result showed clear separation by image quality pattern:\n\n![t-SNE Patterns](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2Fc3849585c995f9ec2e0e597b80cfcfd2%2F00_tsne_patterns.png?generation=1772324957059654&alt=media)\n\nHigh-brightness samples (A:Bright) cluster in the bottom-right, blurry/low-contrast samples (B:Blur, C:LowContrast) in the bottom-left, and Normal data spreads across the upper-center region. This separation aligns strongly with image quality metrics (sharpness, contrast, mean intensity), and samples with lower model performance tend to belong to the peripheral clusters.\n\n## Experimental Setup\n\nTo test this hypothesis, we fine-tuned a pre-trained model (Phase 1, 200 epochs) on the full Phase 2 dataset (786 samples) for only 100 epochs, recording per-sample loss at every epoch. The 100-epoch run was designed as a measurement-oriented short run to observe early-stage training dynamics (production training used 500-4,000 epochs).\n\n### Pre-training\n\nWe used a checkpoint pre-trained for 200 epochs on Phase 1 data (clean labels, no ignore regions). Pre-training accelerates initial convergence.\n\n| Setting | Value |\n|---------|-------|\n| Phase 1 Data | Dataset120_Phase1 |\n| Phase 1 Epochs | 200 |\n| Phase 1 Trainer | nnUNetTrainerSkelRecall_200ep |\n| Phase 1 Loss | CE + Dice + SkelRecall |\n| Phase 1 LR | 0.005 (cosine annealing) |\n\n### Main Experiment (Phase 2 Fine-tune + Loss Tracking)\n\n| Setting | Value |\n|---------|-------|\n| Base Model | nnUNetTrainerSkelRecall (CE + Dice + SkelRecall) |\n| Pre-trained Weights | Phase 1 200ep checkpoint_final.pth |\n| Epochs | 100 (measurement run; production: 500-4,000ep) |\n| Fold | all (all 786 samples used for training) |\n| Patch Size | 128 x 160 x 160 |\n| Batch Size | 5 |\n| LR | 0.005 (cosine annealing) |\n| Loss Logging | Per-sample loss recorded every epoch |\n| Clustering | t-SNE + KMeans (k=8) |\n\n### Data Augmentation\n\nWe used nnU-Net defaults plus domain-specific custom augmentations:\n\n**Spatial transforms:**\n- Elastic deformation (p=0.3, magnitude 10-50 voxels)\n- Rotation (p=0.5, +/-15 degrees per axis)\n- Scaling (p=0.2, 0.7x-1.4x)\n- Mirror (all 3 axes)\n\n**Intensity transforms:**\n- Gaussian Noise (p=0.15, variance 0-0.15)\n- Gaussian Blur (p=0.2, sigma 0.5-1.5)\n- Brightness (p=0.15, 0.5x-1.5x)\n- Contrast (p=0.15, 0.5x-1.5x)\n- Low Resolution Simulation (p=0.25, scale 0.25-1.0)\n- Gamma (p=0.3, gamma 0.7-1.5)\n\n**Custom transforms (papyrus-specific):**\n- BlankRectangle (p=0.4): Blank out rectangular regions to simulate damage\n- Smear (p=0.2): Z-axis ink bleeding simulation\n- InhomogeneousIllumination (p=0.25): Uneven lighting simulation\n\n## Results\n\n### Per-Pattern Loss Curves\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2F7588172f89a612044ba9115c0731d4e2%2F01_loss_by_pattern.png?generation=1772325014361578&alt=media)\n\n\n| Pattern | N | Initial Loss | Final Loss | Improvement | 90% Convergence Epoch |\n|---------|---:|-----------:|-----------:|----------:|---------------------:|\n| Normal | 545 | -0.883 | -1.111 | 0.228 | 60 |\n| A:Bright | 132 | -0.407 | -0.602 | 0.195 | 33 |\n| B:Blur | 39 | -0.662 | -0.815 | 0.153 | 54 |\n| C:LowContrast | 58 | -0.605 | -0.769 | 0.164 | 96 |\n| B-:MildBlur | 12 | -0.833 | -1.163 | 0.330 | 25 |\n\n### Per-Cluster Loss Curves\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2F4238e7c0c0f3c354cb489d420a7f30e2%2F02_loss_by_cluster.png?generation=1772325035027860&alt=media)\n\n\n| Cluster | N | Initial Loss | Final Loss | Improvement | Dominant Pattern |\n|---------|---:|-----------:|-----------:|----------:|:----------------|\n| C0 (Normal) | 101 | -0.947 | -1.118 | 0.171 | Normal:101 |\n| C1 (Normal) | 87 | -0.793 | -1.013 | 0.220 | Normal:86 |\n| C2 (Normal) | 85 | -0.958 | -1.254 | 0.296 | Normal:70, LowContrast:15 |\n| C3 (Normal) | 105 | -0.814 | -1.024 | 0.210 | Normal:93, Blur:8 |\n| C4 (Normal) | 97 | -0.864 | -1.105 | 0.241 | Normal:96 |\n| C5 (Mixed) | 100 | -0.657 | -0.881 | 0.224 | Normal:58, Bright:34 |\n| C6 (Bright) | 101 | -0.382 | -0.586 | 0.204 | **A:Bright:97** |\n| C7 (Mixed) | 110 | -0.783 | -0.961 | 0.178 | LowContrast:38, Normal:37, Blur:31 |\n\n### Overall Summary\n\n| Metric | Value |\n|--------|------:|\n| Overall Initial Loss | -0.771 |\n| Overall Final Loss | -0.986 |\n| Overall Improvement | 0.216 |\n| Worst Cluster (C6) Final Loss | -0.586 |\n| Best Cluster (C2) Final Loss | -1.254 |\n| Worst-Best Gap | **0.669** |\n\n## Detailed Analysis\n\n### Final Loss Distribution by Cluster\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2F5608232a866f857d032a51e79f13bfe4%2F04_final_loss_boxplot.png?generation=1772325058699998&alt=media)\n\n\n### Improvement Distribution by Cluster\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2F57a7d9a9a05a370203254b9b36baddeb%2F05_improvement_boxplot.png?generation=1772325075898632&alt=media)\n\n\n### Top 30 Hardest Samples\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2F11e9b1a508337154ca886287bc134931%2F06_hardest_samples.png?generation=1772325090638670&alt=media)\n\n\n### Loss Distribution Change on t-SNE Map\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2F580300c261c928dd6aaeac840cb32df0%2F08_tsne_loss_progression.png?generation=1772325111687075&alt=media)\n\n\nThe A:Bright cluster (bottom-right) remains red (high loss) even after training. The Normal clusters (upper region) transition to green (low loss) early in training.\n\n## Key Findings\n\n### 1. A:Bright Samples Show Severe Convergence Lag\n\nThe final loss of A:Bright after 100 epochs (**-0.602**) is **0.281 worse** than Normal's *initial* loss (**-0.883**). In other words, after 100 epochs of training, Bright samples haven't even reached the level Normal data starts at.\n\nThis means learning on Bright data is fundamentally harder, and reaching Normal-level performance requires **many more epochs**.\n\n### 2. Convergence Speed Varies Dramatically Across Patterns\n\nEpoch at which 90% of total improvement is reached:\n- B-:MildBlur: **25ep** (fastest; small sample count, high variance)\n- A:Bright: **33ep** (quick plateau at a poor absolute level = 100ep is not enough)\n- B:Blur: **54ep**\n- Normal: **60ep**\n- C:LowContrast: **96ep** (slowest; barely converging within 100ep)\n\nThe early 90% convergence of A:Bright is misleading --- it reflects a \"plateau at a low level\" rather than fast learning. The total improvement is small, so the 90% threshold is reached early, but the absolute loss remains high.\n\n### 3. Cluster C6 Is an Outlier in Difficulty\n\nC6 (A:Bright 97/101) has a final loss of -0.586, with a 0.3-0.7 gap from all other clusters. This cluster acts as the bottleneck for overall training efficiency.\n\n### 4. LowContrast Has the Slowest Convergence\n\nC:LowContrast requires **96 epochs** to reach 90% convergence. At 100 epochs, it has barely made it. Significant further improvement can be expected with 500+ epochs.\n\n## Why Long Training Works (Discussion)\n\n1. **Data diversity**: 17% (Bright) + 7% (LowContrast) + 5% (Blur) of the 786 samples are difficult. Roughly 30% of the data has not reached Normal-level convergence at 100 epochs.\n\n2. **Batch size constraint**: nnU-Net's batch size is 5 (GPU memory limited). At 250 iterations per epoch = 1,250 patches, only a fraction of the 786 samples are seen each epoch. Difficult samples require many epochs to be sampled sufficiently.\n\n3. **Gradient direction conflicts**: Normal and Bright data may require different optimization directions. Short training biases toward the majority Normal pattern, leaving minority difficult patterns under-trained.\n\n4. **2,000-8,000 epochs is justified**: A:Bright's loss at 100 epochs (-0.602) hasn't even reached Normal's initial value (-0.883). Closing this gap requires training on the order of thousands of epochs.\n\n## Conclusion\n\nThis analysis quantitatively confirmed that **training convergence speed varies significantly across image quality patterns**. A:Bright (high-brightness) data did not reach Normal's initial loss even after 100 epochs, supporting the rationale behind the 2,000-8,000 epoch training adopted by top teams.\n\nBased on this hypothesis, we attempted a **specialist model** approach: separating Blur/Bright/LowContrast samples via rule-based classification and fine-tuning dedicated models for each pattern, with quality-based routing at inference time. However, rule-based domain separation had fundamental limitations --- boundaries between patterns were ambiguous, and borderline samples were difficult to handle correctly. \n\nLooking back at the top teams' approaches, the solution to difficult samples was not model separation but **sufficient training epochs to lift overall performance**. This insight is valuable for training strategies in 3D segmentation tasks with high data diversity.",
      "votes": null
    },
    {
      "id": "3415507",
      "postDate": "03/01/2026 00:58:53",
      "content": "<p>Thank you for the excellent writeup! This is something that i've always had questions about. no matter how many epochs i've trained these models for they never seem to quite stop learning, and they have a strange habit of plateauing for hundreds of epochs and suddenly learning again. Maybe this explains it! </p>\n<p>As a quick note -- with the exception of the smear transform, which is just something i had chatgpt write up, the custom transforms mentioned seem like they may be from the augmentations used in the villa monorepo (or the vesuvius subpackage), these were themself lifted from a MIC-DKFZ repository here: <a href=\"https://github.com/MIC-DKFZ/MurineAirwaySegmentation\" target=\"_blank\">https://github.com/MIC-DKFZ/MurineAirwaySegmentation</a> , so i want to make sure if you've grabbed them from there the credit goes to the right people! </p>",
      "rawMarkdown": "Thank you for the excellent writeup! This is something that i've always had questions about. no matter how many epochs i've trained these models for they never seem to quite stop learning, and they have a strange habit of plateauing for hundreds of epochs and suddenly learning again. Maybe this explains it! \n\nAs a quick note -- with the exception of the smear transform, which is just something i had chatgpt write up, the custom transforms mentioned seem like they may be from the augmentations used in the villa monorepo (or the vesuvius subpackage), these were themself lifted from a MIC-DKFZ repository here: https://github.com/MIC-DKFZ/MurineAirwaySegmentation , so i want to make sure if you've grabbed them from there the credit goes to the right people!",
      "votes": null
    },
    {
      "id": "3415540",
      "postDate": "03/01/2026 01:32:24",
      "content": "<h2>Additional comment</h2>\n<p>Ideally, the loss should be similar for both Bright and Normal. In practice, however, Bright tends to start with a higher loss and takes longer to converge.</p>\n<p>With strong augmentation, those same aggressive transforms are also applied to Bright samples, which can destabilize training and potentially lead to failure. In an ideal world, augmentation could be tailored on a per-sample basis, but that’s difficult in practice—so longer training (more epochs) may be necessary.</p>\n<p>Normal samples converge faster and appear more often, so they dominate the optimization, and the Bright cases get drowned out.</p>",
      "rawMarkdown": "## Additional comment\nIdeally, the loss should be similar for both Bright and Normal. In practice, however, Bright tends to start with a higher loss and takes longer to converge.\n\nWith strong augmentation, those same aggressive transforms are also applied to Bright samples, which can destabilize training and potentially lead to failure. In an ideal world, augmentation could be tailored on a per-sample basis, but that’s difficult in practice—so longer training (more epochs) may be necessary.\n\nNormal samples converge faster and appear more often, so they dominate the optimization, and the Bright cases get drowned out.",
      "votes": null
    },
    {
      "id": "3415669",
      "postDate": "03/01/2026 05:16:15",
      "content": "<p>I'm also curious about this. My final solution was trained for 1000 epochs. But I am currently retraining my final solution on 4000 epochs. I want to see how much the CV and LB improves. </p>",
      "rawMarkdown": "I'm also curious about this. My final solution was trained for 1000 epochs. But I am currently retraining my final solution on 4000 epochs. I want to see how much the CV and LB improves.",
      "votes": null
    },
    {
      "id": "3415674",
      "postDate": "03/01/2026 05:42:39",
      "content": "<p>it would be helpful if you share your thoughts after it ;)</p>",
      "rawMarkdown": "it would be helpful if you share your thoughts after it ;)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3415507,
      "author_name": "seanjohnsonsp",
      "author_url": "",
      "post_date": "03/01/2026 00:58:53",
      "content": "<p>Thank you for the excellent writeup! This is something that i've always had questions about. no matter how many epochs i've trained these models for they never seem to quite stop learning, and they have a strange habit of plateauing for hundreds of epochs and suddenly learning again. Maybe this explains it! </p>\n<p>As a quick note -- with the exception of the smear transform, which is just something i had chatgpt write up, the custom transforms mentioned seem like they may be from the augmentations used in the villa monorepo (or the vesuvius subpackage), these were themself lifted from a MIC-DKFZ repository here: <a href=\"https://github.com/MIC-DKFZ/MurineAirwaySegmentation\" target=\"_blank\">https://github.com/MIC-DKFZ/MurineAirwaySegmentation</a> , so i want to make sure if you've grabbed them from there the credit goes to the right people! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3415540,
      "author_name": "sugupoko",
      "author_url": "",
      "post_date": "03/01/2026 01:32:24",
      "content": "<h2>Additional comment</h2>\n<p>Ideally, the loss should be similar for both Bright and Normal. In practice, however, Bright tends to start with a higher loss and takes longer to converge.</p>\n<p>With strong augmentation, those same aggressive transforms are also applied to Bright samples, which can destabilize training and potentially lead to failure. In an ideal world, augmentation could be tailored on a per-sample basis, but that’s difficult in practice—so longer training (more epochs) may be necessary.</p>\n<p>Normal samples converge faster and appear more often, so they dominate the optimization, and the Bright cases get drowned out.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3415669,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "03/01/2026 05:16:15",
      "content": "<p>I'm also curious about this. My final solution was trained for 1000 epochs. But I am currently retraining my final solution on 4000 epochs. I want to see how much the CV and LB improves. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3415674,
          "author_name": "ppilania1985",
          "author_url": "",
          "post_date": "03/01/2026 05:42:39",
          "content": "<p>it would be helpful if you share your thoughts after it ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3415494": "I ran a few experiments based on my hypothesis, and I'd like to discuss the results.\nMy experimental setup may have some flaws or odd choices, but I hope this can spark a useful discussion.\n\n\n## Background\n\nSurveying the top solutions of Vesuvius Challenge 2, we found that all top teams employed **2,000-8,000 epochs** of training (1st: 4,000ep, 3rd: 8,000ep, 12th: 2,000-4,000ep). We wanted to understand why such long training was effective.\n\nIn our experiments, after training on all data (786 samples), we ran inference on the training set itself and found that predictions were significantly worse on low-quality images (high brightness, blur, low contrast). This led us to investigate the data distribution more closely.\n\n- **Normal**: High contrast, high sharpness (545 samples, 69%)\n- **A:Bright**: High intensity, low contrast (132 samples, 17%)\n- **B:Blur**: Blurry, low sharpness (39 samples, 5%)\n- **C:LowContrast**: Low contrast (58 samples, 7%)\n- **B-:MildBlur**: Mildly blurry (12 samples, 2%)\n\n**Hypothesis**: Difficult data patterns (Bright, Blur, etc.) converge more slowly, requiring many more epochs to learn adequately.\n\n## Preliminary Analysis: t-SNE Data Visualization\n\nTo investigate the source of the performance gap, we visualized the feature space of all 786 samples. From each 3D volume, we extracted the central 5 slices along XY, XZ, and YZ planes, fed them through an ImageNet-pretrained ConvNeXt-Tiny (timm) to obtain feature vectors, and averaged across the 3 views (768 dimensions). We then applied t-SNE (perplexity=30, PCA initialization) and clustered with KMeans (k=8).\n\nThe result showed clear separation by image quality pattern:\n\n![t-SNE Patterns](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2Fc3849585c995f9ec2e0e597b80cfcfd2%2F00_tsne_patterns.png?generation=1772324957059654&alt=media)\n\nHigh-brightness samples (A:Bright) cluster in the bottom-right, blurry/low-contrast samples (B:Blur, C:LowContrast) in the bottom-left, and Normal data spreads across the upper-center region. This separation aligns strongly with image quality metrics (sharpness, contrast, mean intensity), and samples with lower model performance tend to belong to the peripheral clusters.\n\n## Experimental Setup\n\nTo test this hypothesis, we fine-tuned a pre-trained model (Phase 1, 200 epochs) on the full Phase 2 dataset (786 samples) for only 100 epochs, recording per-sample loss at every epoch. The 100-epoch run was designed as a measurement-oriented short run to observe early-stage training dynamics (production training used 500-4,000 epochs).\n\n### Pre-training\n\nWe used a checkpoint pre-trained for 200 epochs on Phase 1 data (clean labels, no ignore regions). Pre-training accelerates initial convergence.\n\n| Setting | Value |\n|---------|-------|\n| Phase 1 Data | Dataset120_Phase1 |\n| Phase 1 Epochs | 200 |\n| Phase 1 Trainer | nnUNetTrainerSkelRecall_200ep |\n| Phase 1 Loss | CE + Dice + SkelRecall |\n| Phase 1 LR | 0.005 (cosine annealing) |\n\n### Main Experiment (Phase 2 Fine-tune + Loss Tracking)\n\n| Setting | Value |\n|---------|-------|\n| Base Model | nnUNetTrainerSkelRecall (CE + Dice + SkelRecall) |\n| Pre-trained Weights | Phase 1 200ep checkpoint_final.pth |\n| Epochs | 100 (measurement run; production: 500-4,000ep) |\n| Fold | all (all 786 samples used for training) |\n| Patch Size | 128 x 160 x 160 |\n| Batch Size | 5 |\n| LR | 0.005 (cosine annealing) |\n| Loss Logging | Per-sample loss recorded every epoch |\n| Clustering | t-SNE + KMeans (k=8) |\n\n### Data Augmentation\n\nWe used nnU-Net defaults plus domain-specific custom augmentations:\n\n**Spatial transforms:**\n- Elastic deformation (p=0.3, magnitude 10-50 voxels)\n- Rotation (p=0.5, +/-15 degrees per axis)\n- Scaling (p=0.2, 0.7x-1.4x)\n- Mirror (all 3 axes)\n\n**Intensity transforms:**\n- Gaussian Noise (p=0.15, variance 0-0.15)\n- Gaussian Blur (p=0.2, sigma 0.5-1.5)\n- Brightness (p=0.15, 0.5x-1.5x)\n- Contrast (p=0.15, 0.5x-1.5x)\n- Low Resolution Simulation (p=0.25, scale 0.25-1.0)\n- Gamma (p=0.3, gamma 0.7-1.5)\n\n**Custom transforms (papyrus-specific):**\n- BlankRectangle (p=0.4): Blank out rectangular regions to simulate damage\n- Smear (p=0.2): Z-axis ink bleeding simulation\n- InhomogeneousIllumination (p=0.25): Uneven lighting simulation\n\n## Results\n\n### Per-Pattern Loss Curves\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2F7588172f89a612044ba9115c0731d4e2%2F01_loss_by_pattern.png?generation=1772325014361578&alt=media)\n\n\n| Pattern | N | Initial Loss | Final Loss | Improvement | 90% Convergence Epoch |\n|---------|---:|-----------:|-----------:|----------:|---------------------:|\n| Normal | 545 | -0.883 | -1.111 | 0.228 | 60 |\n| A:Bright | 132 | -0.407 | -0.602 | 0.195 | 33 |\n| B:Blur | 39 | -0.662 | -0.815 | 0.153 | 54 |\n| C:LowContrast | 58 | -0.605 | -0.769 | 0.164 | 96 |\n| B-:MildBlur | 12 | -0.833 | -1.163 | 0.330 | 25 |\n\n### Per-Cluster Loss Curves\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2F4238e7c0c0f3c354cb489d420a7f30e2%2F02_loss_by_cluster.png?generation=1772325035027860&alt=media)\n\n\n| Cluster | N | Initial Loss | Final Loss | Improvement | Dominant Pattern |\n|---------|---:|-----------:|-----------:|----------:|:----------------|\n| C0 (Normal) | 101 | -0.947 | -1.118 | 0.171 | Normal:101 |\n| C1 (Normal) | 87 | -0.793 | -1.013 | 0.220 | Normal:86 |\n| C2 (Normal) | 85 | -0.958 | -1.254 | 0.296 | Normal:70, LowContrast:15 |\n| C3 (Normal) | 105 | -0.814 | -1.024 | 0.210 | Normal:93, Blur:8 |\n| C4 (Normal) | 97 | -0.864 | -1.105 | 0.241 | Normal:96 |\n| C5 (Mixed) | 100 | -0.657 | -0.881 | 0.224 | Normal:58, Bright:34 |\n| C6 (Bright) | 101 | -0.382 | -0.586 | 0.204 | **A:Bright:97** |\n| C7 (Mixed) | 110 | -0.783 | -0.961 | 0.178 | LowContrast:38, Normal:37, Blur:31 |\n\n### Overall Summary\n\n| Metric | Value |\n|--------|------:|\n| Overall Initial Loss | -0.771 |\n| Overall Final Loss | -0.986 |\n| Overall Improvement | 0.216 |\n| Worst Cluster (C6) Final Loss | -0.586 |\n| Best Cluster (C2) Final Loss | -1.254 |\n| Worst-Best Gap | **0.669** |\n\n## Detailed Analysis\n\n### Final Loss Distribution by Cluster\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2F5608232a866f857d032a51e79f13bfe4%2F04_final_loss_boxplot.png?generation=1772325058699998&alt=media)\n\n\n### Improvement Distribution by Cluster\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2F57a7d9a9a05a370203254b9b36baddeb%2F05_improvement_boxplot.png?generation=1772325075898632&alt=media)\n\n\n### Top 30 Hardest Samples\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2F11e9b1a508337154ca886287bc134931%2F06_hardest_samples.png?generation=1772325090638670&alt=media)\n\n\n### Loss Distribution Change on t-SNE Map\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2F580300c261c928dd6aaeac840cb32df0%2F08_tsne_loss_progression.png?generation=1772325111687075&alt=media)\n\n\nThe A:Bright cluster (bottom-right) remains red (high loss) even after training. The Normal clusters (upper region) transition to green (low loss) early in training.\n\n## Key Findings\n\n### 1. A:Bright Samples Show Severe Convergence Lag\n\nThe final loss of A:Bright after 100 epochs (**-0.602**) is **0.281 worse** than Normal's *initial* loss (**-0.883**). In other words, after 100 epochs of training, Bright samples haven't even reached the level Normal data starts at.\n\nThis means learning on Bright data is fundamentally harder, and reaching Normal-level performance requires **many more epochs**.\n\n### 2. Convergence Speed Varies Dramatically Across Patterns\n\nEpoch at which 90% of total improvement is reached:\n- B-:MildBlur: **25ep** (fastest; small sample count, high variance)\n- A:Bright: **33ep** (quick plateau at a poor absolute level = 100ep is not enough)\n- B:Blur: **54ep**\n- Normal: **60ep**\n- C:LowContrast: **96ep** (slowest; barely converging within 100ep)\n\nThe early 90% convergence of A:Bright is misleading --- it reflects a \"plateau at a low level\" rather than fast learning. The total improvement is small, so the 90% threshold is reached early, but the absolute loss remains high.\n\n### 3. Cluster C6 Is an Outlier in Difficulty\n\nC6 (A:Bright 97/101) has a final loss of -0.586, with a 0.3-0.7 gap from all other clusters. This cluster acts as the bottleneck for overall training efficiency.\n\n### 4. LowContrast Has the Slowest Convergence\n\nC:LowContrast requires **96 epochs** to reach 90% convergence. At 100 epochs, it has barely made it. Significant further improvement can be expected with 500+ epochs.\n\n## Why Long Training Works (Discussion)\n\n1. **Data diversity**: 17% (Bright) + 7% (LowContrast) + 5% (Blur) of the 786 samples are difficult. Roughly 30% of the data has not reached Normal-level convergence at 100 epochs.\n\n2. **Batch size constraint**: nnU-Net's batch size is 5 (GPU memory limited). At 250 iterations per epoch = 1,250 patches, only a fraction of the 786 samples are seen each epoch. Difficult samples require many epochs to be sampled sufficiently.\n\n3. **Gradient direction conflicts**: Normal and Bright data may require different optimization directions. Short training biases toward the majority Normal pattern, leaving minority difficult patterns under-trained.\n\n4. **2,000-8,000 epochs is justified**: A:Bright's loss at 100 epochs (-0.602) hasn't even reached Normal's initial value (-0.883). Closing this gap requires training on the order of thousands of epochs.\n\n## Conclusion\n\nThis analysis quantitatively confirmed that **training convergence speed varies significantly across image quality patterns**. A:Bright (high-brightness) data did not reach Normal's initial loss even after 100 epochs, supporting the rationale behind the 2,000-8,000 epoch training adopted by top teams.\n\nBased on this hypothesis, we attempted a **specialist model** approach: separating Blur/Bright/LowContrast samples via rule-based classification and fine-tuning dedicated models for each pattern, with quality-based routing at inference time. However, rule-based domain separation had fundamental limitations --- boundaries between patterns were ambiguous, and borderline samples were difficult to handle correctly. \n\nLooking back at the top teams' approaches, the solution to difficult samples was not model separation but **sufficient training epochs to lift overall performance**. This insight is valuable for training strategies in 3D segmentation tasks with high data diversity.",
    "3415507": "Thank you for the excellent writeup! This is something that i've always had questions about. no matter how many epochs i've trained these models for they never seem to quite stop learning, and they have a strange habit of plateauing for hundreds of epochs and suddenly learning again. Maybe this explains it! \n\nAs a quick note -- with the exception of the smear transform, which is just something i had chatgpt write up, the custom transforms mentioned seem like they may be from the augmentations used in the villa monorepo (or the vesuvius subpackage), these were themself lifted from a MIC-DKFZ repository here: https://github.com/MIC-DKFZ/MurineAirwaySegmentation , so i want to make sure if you've grabbed them from there the credit goes to the right people!",
    "3415540": "## Additional comment\nIdeally, the loss should be similar for both Bright and Normal. In practice, however, Bright tends to start with a higher loss and takes longer to converge.\n\nWith strong augmentation, those same aggressive transforms are also applied to Bright samples, which can destabilize training and potentially lead to failure. In an ideal world, augmentation could be tailored on a per-sample basis, but that’s difficult in practice—so longer training (more epochs) may be necessary.\n\nNormal samples converge faster and appear more often, so they dominate the optimization, and the Bright cases get drowned out.",
    "3415669": "I'm also curious about this. My final solution was trained for 1000 epochs. But I am currently retraining my final solution on 4000 epochs. I want to see how much the CV and LB improves.",
    "3415674": "it would be helpful if you share your thoughts after it ;)"
  },
  "source": "meta"
}