{
  "id": 679276,
  "title": "Bronze Medal - 29 Hours of TPU Continued Training（TransUNet）",
  "url": "/competitions/vesuvius-challenge-surface-detection/writeups/bronze-medal-29-hours-of-tpu-continued-training",
  "author_name": "",
  "post_date": "2026-02-28T11:19:42.137Z",
  "votes": 5,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Thank you for the organizers and Kaggle team for this great competition! I also thank all open-source contributors for sharing high-quality code and models.</p>\n<h2>1. Model and Training Foundation</h2>\n<p>My model for continued training is based on Innat's pre-trained TransUNet model (<a href=\"https://www.kaggle.com/models/ipythonx/vsd-model/Keras/transunet/3?select=transunet.seresnext50.160px.comboloss.weights.h5)\" target=\"_blank\">https://www.kaggle.com/models/ipythonx/vsd-model/Keras/transunet/3?select=transunet.seresnext50.160px.comboloss.weights.h5)</a>, and the training code is adapted from Innat's notebook (<a href=\"https://www.kaggle.com/code/ipythonx/train-vesuvius-surface-3d-detection-on-tpu?scriptVersionId=294535277)\" target=\"_blank\">https://www.kaggle.com/code/ipythonx/train-vesuvius-surface-3d-detection-on-tpu?scriptVersionId=294535277)</a>.</p>\n<h2>2. Key Modifications</h2>\n<h3>2.1 Model Architecture</h3>\n<ul>\n<li>Replaced SegFormer (mit_b0) with TransUNet (seresnext50)</li>\n<li>Increased input size from 128³ to 160³</li>\n</ul>\n<h3>2.2 TPU Optimization &amp; Distributed Computing</h3>\n<ul>\n<li>Added JAX/TPU memory environment variables to enhance stability across 8 TPU cores</li>\n<li>Adapted distribution strategy for compatibility with different Keras versions</li>\n<li>Separated training and inference (TPU for training, CPU for inference) to avoid memory overflow</li>\n</ul>\n<h3>2.3 Sliding Window Inference Callback</h3>\n<ul>\n<li>Implemented a custom <code>UltimateStableSWICallback</code> to replace the original solution:<ul>\n<li>Automatic switching between CPU/TPU environments (inference failures do not interrupt training)</li>\n<li>Proactive garbage collection to free up memory</li></ul></li>\n</ul>\n<h3>2.4 Weight Management</h3>\n<ul>\n<li>Supported loading pre-trained weights for continued training (with <code>skip_mismatch=True</code>)</li>\n<li>Implemented a triple weight-saving mechanism: temporary weights, best model (by validation dice), and final model weights</li>\n</ul>\n<h3>2.5 Training Parameters</h3>\n<ul>\n<li>Expanded validation set from 1 to 2 files to improve robustness</li>\n<li>Set <code>warmup_steps=0</code> for continued training</li>\n<li>Increased shuffle buffer size from 100 to 400</li>\n</ul>\n<h2>3. Training Process &amp; Results</h2>\n<p>Modified training notebooks (only epoch number adjusted):</p>\n<ul>\n<li>50 epochs: <a href=\"https://www.kaggle.com/code/tonyai007/train-vesuvius-seresnext50-comboloss\" target=\"_blank\">https://www.kaggle.com/code/tonyai007/train-vesuvius-seresnext50-comboloss</a></li>\n<li>250 epochs: <a href=\"https://www.kaggle.com/code/tonyai007/train-vesuvius-seresnext50-comboloss-2\" target=\"_blank\">https://www.kaggle.com/code/tonyai007/train-vesuvius-seresnext50-comboloss-2</a></li>\n<li>450 epochs: <a href=\"https://www.kaggle.com/code/tonyai007/train-vesuvius-seresnext50-comboloss-3\" target=\"_blank\">https://www.kaggle.com/code/tonyai007/train-vesuvius-seresnext50-comboloss-3</a></li>\n<li>600 epochs: <a href=\"https://www.kaggle.com/code/tonyai007/train-vesuvius-seresnext50-comboloss-4\" target=\"_blank\">https://www.kaggle.com/code/tonyai007/train-vesuvius-seresnext50-comboloss-4</a></li>\n<li>830 epochs: <a href=\"https://www.kaggle.com/code/tonyai007/train-vesuvius-seresnext50-comboloss-5\" target=\"_blank\">https://www.kaggle.com/code/tonyai007/train-vesuvius-seresnext50-comboloss-5</a></li>\n</ul>\n<p>Training was halted at 830 epochs due to the competition deadline. Validation results by phase:</p>\n<table>\n<thead>\n<tr>\n<th>Epochs</th>\n<th>Validation Dice</th>\n<th>Validation Loss</th>\n<th>Best Epoch</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0-50</td>\n<td>0.8075</td>\n<td>0.9753</td>\n<td>48</td>\n</tr>\n<tr>\n<td>50-250</td>\n<td>0.8140</td>\n<td>0.9396</td>\n<td>249</td>\n</tr>\n<tr>\n<td>250-450</td>\n<td>0.8176</td>\n<td>0.9256</td>\n<td>444</td>\n</tr>\n<tr>\n<td>450-600</td>\n<td>0.8193</td>\n<td>0.9200</td>\n<td>598</td>\n</tr>\n<tr>\n<td>600-830</td>\n<td>0.8223</td>\n<td>0.8937</td>\n<td>816</td>\n</tr>\n</tbody>\n</table>\n<h2>4. Inference &amp; Results</h2>\n<h3>4.1 Single-Model Inference</h3>\n<p>The inference code is based on Tony Li · PaulG · Yiheng Wang's notebook (<a href=\"https://www.kaggle.com/code/tonylica/vesuvius-0-552)\" target=\"_blank\">https://www.kaggle.com/code/tonylica/vesuvius-0-552)</a>, with 3-level overlap sliding window and 7-fold TTA.</p>\n<p><strong>Best Submission (Code 1):</strong> <a href=\"https://www.kaggle.com/code/tonyai007/vesuvius-seresnext50-comboloss-830ft\" target=\"_blank\">https://www.kaggle.com/code/tonyai007/vesuvius-seresnext50-comboloss-830ft</a></p>\n<ul>\n<li>Parameters: overlap_public=0.42, overlap_base=0.48, overlap_hi=0.6, T_low=0.5, T_high=0.9, z_radius=0, xy_radius=2, dust_min_size=200</li>\n<li>Scores: Public LB 0.565, Private LB 0.582</li>\n</ul>\n<p><strong>Ablation on Inference Params:</strong></p>\n<ul>\n<li>Adjusting T_low (0.35) and T_high (0.85) led to lower <a href=\"https://www.kaggle.com/code/tonyai007/inference-vesuvius-surface-3d-detection-test?scriptVersionId=300377344\" target=\"_blank\">scores</a>: Public LB 0.558, Private LB 0.570</li>\n<li>z_radius impact: 0 &gt; 1 &gt; 2 &gt; 3 &gt; 4 (z_radius=0 achieved the best performance)</li>\n</ul>\n<p><strong>Ablation on Training Epochs (Code 1 params):</strong></p>\n<ul>\n<li>50 epochs: Public LB 0.558, Private LB 0.571 (z_radius=1)</li>\n<li>250 epochs: Public LB 0.563, Private LB 0.573 (z_radius=0)</li>\n<li>450 epochs: Public LB 0.556, Private LB 0.577 (z_radius=0)</li>\n<li>600 epochs: Public LB 0.562, Private LB 0.574 (z_radius=0)</li>\n<li>830 epochs: Public LB 0.565, Private LB 0.582 (z_radius=0)</li>\n</ul>\n<h3>4.2 Multi-Model Inference</h3>\n<p>Submission (Code 2): <a href=\"https://www.kaggle.com/code/tonyai007/inference-vesuvius-surface-3d-detection-epoch-x5?scriptVersionId=300146063\" target=\"_blank\">https://www.kaggle.com/code/tonyai007/inference-vesuvius-surface-3d-detection-epoch-x5?scriptVersionId=300146063</a></p>\n<ul>\n<li>Scores: Public LB 0.566, Private LB 0.577\nKey settings:\nCombination of 3 epoch + 7TTA: Adjustments to <strong>overlap</strong>, <strong>T_low/T_high</strong>, and <strong>z_radius</strong> affected public LB scores, but private LB remained stable at ≈0.577\nCombination of 5 epoch + 3/4TTA: Changes to the same hyperparameters impacted public LB, with private LB consistently around 0.571\nNote: Multi-model fusion improved public LB score but underperformed compared to the single 830-epoch model on private LB (0.577 vs. 0.582), which may have led to misinterpretation.</li>\n</ul>\n<h3>4.3 Baseline Comparisons</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/ipythonx/train-vesuvius-surface-3d-detection-on-tpu\" target=\"_blank\">SegFormer</a> (128³, 1200 epochs on TPU, ≈25 hours, 1overlap + 7TTA): Public/Private LB 0.524</li>\n<li><a href=\"https://www.kaggle.com/code/jirkaborovec/surface-nnunet-training-inference-with-2xt4\" target=\"_blank\">nnUNet</a> (3d_lowres/fold_all, 500 epochs on T4 GPU, ≈40 hours): Public/Private LB 0.53 (200ep: 0.513 / 300ep: 0.521 / 400ep: 0.523)</li>\n</ul>\n<h2>Acknowledgements</h2>\n<p>I thank all upvoters for their support, which helped me earn two Notebook Gold Medals.</p>",
  "messages": [
    {
      "id": "3415184",
      "postDate": "02/28/2026 11:15:54",
      "content": "<p>Thank you for the organizers and Kaggle team for this great competition! I also thank all open-source contributors for sharing high-quality code and models.</p>\n<h2>1. Model and Training Foundation</h2>\n<p>My model for continued training is based on Innat's pre-trained TransUNet model (<a href=\"https://www.kaggle.com/models/ipythonx/vsd-model/Keras/transunet/3?select=transunet.seresnext50.160px.comboloss.weights.h5)\" target=\"_blank\">https://www.kaggle.com/models/ipythonx/vsd-model/Keras/transunet/3?select=transunet.seresnext50.160px.comboloss.weights.h5)</a>, and the training code is adapted from Innat's notebook (<a href=\"https://www.kaggle.com/code/ipythonx/train-vesuvius-surface-3d-detection-on-tpu?scriptVersionId=294535277)\" target=\"_blank\">https://www.kaggle.com/code/ipythonx/train-vesuvius-surface-3d-detection-on-tpu?scriptVersionId=294535277)</a>.</p>\n<h2>2. Key Modifications</h2>\n<h3>2.1 Model Architecture</h3>\n<ul>\n<li>Replaced SegFormer (mit_b0) with TransUNet (seresnext50)</li>\n<li>Increased input size from 128³ to 160³</li>\n</ul>\n<h3>2.2 TPU Optimization &amp; Distributed Computing</h3>\n<ul>\n<li>Added JAX/TPU memory environment variables to enhance stability across 8 TPU cores</li>\n<li>Adapted distribution strategy for compatibility with different Keras versions</li>\n<li>Separated training and inference (TPU for training, CPU for inference) to avoid memory overflow</li>\n</ul>\n<h3>2.3 Sliding Window Inference Callback</h3>\n<ul>\n<li>Implemented a custom <code>UltimateStableSWICallback</code> to replace the original solution:<ul>\n<li>Automatic switching between CPU/TPU environments (inference failures do not interrupt training)</li>\n<li>Proactive garbage collection to free up memory</li></ul></li>\n</ul>\n<h3>2.4 Weight Management</h3>\n<ul>\n<li>Supported loading pre-trained weights for continued training (with <code>skip_mismatch=True</code>)</li>\n<li>Implemented a triple weight-saving mechanism: temporary weights, best model (by validation dice), and final model weights</li>\n</ul>\n<h3>2.5 Training Parameters</h3>\n<ul>\n<li>Expanded validation set from 1 to 2 files to improve robustness</li>\n<li>Set <code>warmup_steps=0</code> for continued training</li>\n<li>Increased shuffle buffer size from 100 to 400</li>\n</ul>\n<h2>3. Training Process &amp; Results</h2>\n<p>Modified training notebooks (only epoch number adjusted):</p>\n<ul>\n<li>50 epochs: <a href=\"https://www.kaggle.com/code/tonyai007/train-vesuvius-seresnext50-comboloss\" target=\"_blank\">https://www.kaggle.com/code/tonyai007/train-vesuvius-seresnext50-comboloss</a></li>\n<li>250 epochs: <a href=\"https://www.kaggle.com/code/tonyai007/train-vesuvius-seresnext50-comboloss-2\" target=\"_blank\">https://www.kaggle.com/code/tonyai007/train-vesuvius-seresnext50-comboloss-2</a></li>\n<li>450 epochs: <a href=\"https://www.kaggle.com/code/tonyai007/train-vesuvius-seresnext50-comboloss-3\" target=\"_blank\">https://www.kaggle.com/code/tonyai007/train-vesuvius-seresnext50-comboloss-3</a></li>\n<li>600 epochs: <a href=\"https://www.kaggle.com/code/tonyai007/train-vesuvius-seresnext50-comboloss-4\" target=\"_blank\">https://www.kaggle.com/code/tonyai007/train-vesuvius-seresnext50-comboloss-4</a></li>\n<li>830 epochs: <a href=\"https://www.kaggle.com/code/tonyai007/train-vesuvius-seresnext50-comboloss-5\" target=\"_blank\">https://www.kaggle.com/code/tonyai007/train-vesuvius-seresnext50-comboloss-5</a></li>\n</ul>\n<p>Training was halted at 830 epochs due to the competition deadline. Validation results by phase:</p>\n<table>\n<thead>\n<tr>\n<th>Epochs</th>\n<th>Validation Dice</th>\n<th>Validation Loss</th>\n<th>Best Epoch</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0-50</td>\n<td>0.8075</td>\n<td>0.9753</td>\n<td>48</td>\n</tr>\n<tr>\n<td>50-250</td>\n<td>0.8140</td>\n<td>0.9396</td>\n<td>249</td>\n</tr>\n<tr>\n<td>250-450</td>\n<td>0.8176</td>\n<td>0.9256</td>\n<td>444</td>\n</tr>\n<tr>\n<td>450-600</td>\n<td>0.8193</td>\n<td>0.9200</td>\n<td>598</td>\n</tr>\n<tr>\n<td>600-830</td>\n<td>0.8223</td>\n<td>0.8937</td>\n<td>816</td>\n</tr>\n</tbody>\n</table>\n<h2>4. Inference &amp; Results</h2>\n<h3>4.1 Single-Model Inference</h3>\n<p>The inference code is based on Tony Li · PaulG · Yiheng Wang's notebook (<a href=\"https://www.kaggle.com/code/tonylica/vesuvius-0-552)\" target=\"_blank\">https://www.kaggle.com/code/tonylica/vesuvius-0-552)</a>, with 3-level overlap sliding window and 7-fold TTA.</p>\n<p><strong>Best Submission (Code 1):</strong> <a href=\"https://www.kaggle.com/code/tonyai007/vesuvius-seresnext50-comboloss-830ft\" target=\"_blank\">https://www.kaggle.com/code/tonyai007/vesuvius-seresnext50-comboloss-830ft</a></p>\n<ul>\n<li>Parameters: overlap_public=0.42, overlap_base=0.48, overlap_hi=0.6, T_low=0.5, T_high=0.9, z_radius=0, xy_radius=2, dust_min_size=200</li>\n<li>Scores: Public LB 0.565, Private LB 0.582</li>\n</ul>\n<p><strong>Ablation on Inference Params:</strong></p>\n<ul>\n<li>Adjusting T_low (0.35) and T_high (0.85) led to lower <a href=\"https://www.kaggle.com/code/tonyai007/inference-vesuvius-surface-3d-detection-test?scriptVersionId=300377344\" target=\"_blank\">scores</a>: Public LB 0.558, Private LB 0.570</li>\n<li>z_radius impact: 0 &gt; 1 &gt; 2 &gt; 3 &gt; 4 (z_radius=0 achieved the best performance)</li>\n</ul>\n<p><strong>Ablation on Training Epochs (Code 1 params):</strong></p>\n<ul>\n<li>50 epochs: Public LB 0.558, Private LB 0.571 (z_radius=1)</li>\n<li>250 epochs: Public LB 0.563, Private LB 0.573 (z_radius=0)</li>\n<li>450 epochs: Public LB 0.556, Private LB 0.577 (z_radius=0)</li>\n<li>600 epochs: Public LB 0.562, Private LB 0.574 (z_radius=0)</li>\n<li>830 epochs: Public LB 0.565, Private LB 0.582 (z_radius=0)</li>\n</ul>\n<h3>4.2 Multi-Model Inference</h3>\n<p>Submission (Code 2): <a href=\"https://www.kaggle.com/code/tonyai007/inference-vesuvius-surface-3d-detection-epoch-x5?scriptVersionId=300146063\" target=\"_blank\">https://www.kaggle.com/code/tonyai007/inference-vesuvius-surface-3d-detection-epoch-x5?scriptVersionId=300146063</a></p>\n<ul>\n<li>Scores: Public LB 0.566, Private LB 0.577\nKey settings:\nCombination of 3 epoch + 7TTA: Adjustments to <strong>overlap</strong>, <strong>T_low/T_high</strong>, and <strong>z_radius</strong> affected public LB scores, but private LB remained stable at ≈0.577\nCombination of 5 epoch + 3/4TTA: Changes to the same hyperparameters impacted public LB, with private LB consistently around 0.571\nNote: Multi-model fusion improved public LB score but underperformed compared to the single 830-epoch model on private LB (0.577 vs. 0.582), which may have led to misinterpretation.</li>\n</ul>\n<h3>4.3 Baseline Comparisons</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/ipythonx/train-vesuvius-surface-3d-detection-on-tpu\" target=\"_blank\">SegFormer</a> (128³, 1200 epochs on TPU, ≈25 hours, 1overlap + 7TTA): Public/Private LB 0.524</li>\n<li><a href=\"https://www.kaggle.com/code/jirkaborovec/surface-nnunet-training-inference-with-2xt4\" target=\"_blank\">nnUNet</a> (3d_lowres/fold_all, 500 epochs on T4 GPU, ≈40 hours): Public/Private LB 0.53 (200ep: 0.513 / 300ep: 0.521 / 400ep: 0.523)</li>\n</ul>\n<h2>Acknowledgements</h2>\n<p>I thank all upvoters for their support, which helped me earn two Notebook Gold Medals.</p>",
      "rawMarkdown": "Thank you for the organizers and Kaggle team for this great competition! I also thank all open-source contributors for sharing high-quality code and models.\n\n## 1. Model and Training Foundation\nMy model for continued training is based on Innat's pre-trained TransUNet model (https://www.kaggle.com/models/ipythonx/vsd-model/Keras/transunet/3?select=transunet.seresnext50.160px.comboloss.weights.h5), and the training code is adapted from Innat's notebook (https://www.kaggle.com/code/ipythonx/train-vesuvius-surface-3d-detection-on-tpu?scriptVersionId=294535277).\n\n## 2. Key Modifications\n### 2.1 Model Architecture\n- Replaced SegFormer (mit_b0) with TransUNet (seresnext50)\n- Increased input size from 128³ to 160³\n### 2.2 TPU Optimization & Distributed Computing\n- Added JAX/TPU memory environment variables to enhance stability across 8 TPU cores\n- Adapted distribution strategy for compatibility with different Keras versions\n- Separated training and inference (TPU for training, CPU for inference) to avoid memory overflow\n### 2.3 Sliding Window Inference Callback\n- Implemented a custom `UltimateStableSWICallback` to replace the original solution:\n  - Automatic switching between CPU/TPU environments (inference failures do not interrupt training)\n  - Proactive garbage collection to free up memory\n### 2.4 Weight Management\n- Supported loading pre-trained weights for continued training (with `skip_mismatch=True`)\n- Implemented a triple weight-saving mechanism: temporary weights, best model (by validation dice), and final model weights\n### 2.5 Training Parameters\n- Expanded validation set from 1 to 2 files to improve robustness\n- Set `warmup_steps=0` for continued training\n- Increased shuffle buffer size from 100 to 400\n\n## 3. Training Process & Results\nModified training notebooks (only epoch number adjusted):\n- 50 epochs: https://www.kaggle.com/code/tonyai007/train-vesuvius-seresnext50-comboloss\n- 250 epochs: https://www.kaggle.com/code/tonyai007/train-vesuvius-seresnext50-comboloss-2\n- 450 epochs: https://www.kaggle.com/code/tonyai007/train-vesuvius-seresnext50-comboloss-3\n- 600 epochs: https://www.kaggle.com/code/tonyai007/train-vesuvius-seresnext50-comboloss-4\n- 830 epochs: https://www.kaggle.com/code/tonyai007/train-vesuvius-seresnext50-comboloss-5\n\nTraining was halted at 830 epochs due to the competition deadline. Validation results by phase:\n\n| Epochs | Validation Dice | Validation Loss | Best Epoch |\n|--------|-----------------|-----------------|------------|\n| 0-50   | 0.8075          | 0.9753          | 48         |\n| 50-250 | 0.8140          | 0.9396          | 249        |\n| 250-450| 0.8176          | 0.9256          | 444        |\n| 450-600| 0.8193          | 0.9200          | 598        |\n| 600-830| 0.8223          | 0.8937          | 816        |\n\n## 4. Inference & Results\n### 4.1 Single-Model Inference\nThe inference code is based on Tony Li · PaulG · Yiheng Wang's notebook (https://www.kaggle.com/code/tonylica/vesuvius-0-552), with 3-level overlap sliding window and 7-fold TTA.\n\n**Best Submission (Code 1):** https://www.kaggle.com/code/tonyai007/vesuvius-seresnext50-comboloss-830ft\n- Parameters: overlap_public=0.42, overlap_base=0.48, overlap_hi=0.6, T_low=0.5, T_high=0.9, z_radius=0, xy_radius=2, dust_min_size=200\n- Scores: Public LB 0.565, Private LB 0.582\n\n**Ablation on Inference Params:**\n- Adjusting T_low (0.35) and T_high (0.85) led to lower [scores](https://www.kaggle.com/code/tonyai007/inference-vesuvius-surface-3d-detection-test?scriptVersionId=300377344): Public LB 0.558, Private LB 0.570\n- z_radius impact: 0 > 1 > 2 > 3 > 4 (z_radius=0 achieved the best performance)\n\n**Ablation on Training Epochs (Code 1 params):**\n- 50 epochs: Public LB 0.558, Private LB 0.571 (z_radius=1)\n- 250 epochs: Public LB 0.563, Private LB 0.573 (z_radius=0)\n- 450 epochs: Public LB 0.556, Private LB 0.577 (z_radius=0)\n- 600 epochs: Public LB 0.562, Private LB 0.574 (z_radius=0)\n- 830 epochs: Public LB 0.565, Private LB 0.582 (z_radius=0)\n\n### 4.2 Multi-Model Inference\nSubmission (Code 2): https://www.kaggle.com/code/tonyai007/inference-vesuvius-surface-3d-detection-epoch-x5?scriptVersionId=300146063\n- Scores: Public LB 0.566, Private LB 0.577\nKey settings:\nCombination of 3 epoch + 7TTA: Adjustments to **overlap**, **T_low/T_high**, and **z_radius** affected public LB scores, but private LB remained stable at ≈0.577\nCombination of 5 epoch + 3/4TTA: Changes to the same hyperparameters impacted public LB, with private LB consistently around 0.571\nNote: Multi-model fusion improved public LB score but underperformed compared to the single 830-epoch model on private LB (0.577 vs. 0.582), which may have led to misinterpretation.\n\n### 4.3 Baseline Comparisons\n- [SegFormer](https://www.kaggle.com/code/ipythonx/train-vesuvius-surface-3d-detection-on-tpu) (128³, 1200 epochs on TPU, ≈25 hours, 1overlap + 7TTA): Public/Private LB 0.524\n- [nnUNet](https://www.kaggle.com/code/jirkaborovec/surface-nnunet-training-inference-with-2xt4) (3d_lowres/fold_all, 500 epochs on T4 GPU, ≈40 hours): Public/Private LB 0.53 (200ep: 0.513 / 300ep: 0.521 / 400ep: 0.523)\n\n## Acknowledgements\nI thank all upvoters for their support, which helped me earn two Notebook Gold Medals.",
      "votes": null
    },
    {
      "id": "3415191",
      "postDate": "02/28/2026 11:59:07",
      "content": "<p><a href=\"https://www.kaggle.com/tonyai007\" target=\"_blank\">@tonyai007</a> Congrats. Did you perform all your experiment in kaggle only?</p>",
      "rawMarkdown": "tonyai007 Congrats. Did you perform all your experiment in kaggle only?",
      "votes": null
    },
    {
      "id": "3415225",
      "postDate": "02/28/2026 13:00:12",
      "content": "<p><a href=\"https://www.kaggle.com/Innat\" target=\"_blank\">@Innat</a> Yes, I only used Kaggle's graphics card. Thank you for providing the training and high score prediction code. I have learned a lot</p>",
      "rawMarkdown": "Innat Yes, I only used Kaggle's graphics card. Thank you for providing the training and high score prediction code. I have learned a lot",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3415191,
      "author_name": "ipythonx",
      "author_url": "",
      "post_date": "02/28/2026 11:59:07",
      "content": "<p><a href=\"https://www.kaggle.com/tonyai007\" target=\"_blank\">@tonyai007</a> Congrats. Did you perform all your experiment in kaggle only?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3415225,
          "author_name": "tonyai007",
          "author_url": "",
          "post_date": "02/28/2026 13:00:12",
          "content": "<p><a href=\"https://www.kaggle.com/Innat\" target=\"_blank\">@Innat</a> Yes, I only used Kaggle's graphics card. Thank you for providing the training and high score prediction code. I have learned a lot</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3415184": "Thank you for the organizers and Kaggle team for this great competition! I also thank all open-source contributors for sharing high-quality code and models.\n\n## 1. Model and Training Foundation\nMy model for continued training is based on Innat's pre-trained TransUNet model (https://www.kaggle.com/models/ipythonx/vsd-model/Keras/transunet/3?select=transunet.seresnext50.160px.comboloss.weights.h5), and the training code is adapted from Innat's notebook (https://www.kaggle.com/code/ipythonx/train-vesuvius-surface-3d-detection-on-tpu?scriptVersionId=294535277).\n\n## 2. Key Modifications\n### 2.1 Model Architecture\n- Replaced SegFormer (mit_b0) with TransUNet (seresnext50)\n- Increased input size from 128³ to 160³\n### 2.2 TPU Optimization & Distributed Computing\n- Added JAX/TPU memory environment variables to enhance stability across 8 TPU cores\n- Adapted distribution strategy for compatibility with different Keras versions\n- Separated training and inference (TPU for training, CPU for inference) to avoid memory overflow\n### 2.3 Sliding Window Inference Callback\n- Implemented a custom `UltimateStableSWICallback` to replace the original solution:\n  - Automatic switching between CPU/TPU environments (inference failures do not interrupt training)\n  - Proactive garbage collection to free up memory\n### 2.4 Weight Management\n- Supported loading pre-trained weights for continued training (with `skip_mismatch=True`)\n- Implemented a triple weight-saving mechanism: temporary weights, best model (by validation dice), and final model weights\n### 2.5 Training Parameters\n- Expanded validation set from 1 to 2 files to improve robustness\n- Set `warmup_steps=0` for continued training\n- Increased shuffle buffer size from 100 to 400\n\n## 3. Training Process & Results\nModified training notebooks (only epoch number adjusted):\n- 50 epochs: https://www.kaggle.com/code/tonyai007/train-vesuvius-seresnext50-comboloss\n- 250 epochs: https://www.kaggle.com/code/tonyai007/train-vesuvius-seresnext50-comboloss-2\n- 450 epochs: https://www.kaggle.com/code/tonyai007/train-vesuvius-seresnext50-comboloss-3\n- 600 epochs: https://www.kaggle.com/code/tonyai007/train-vesuvius-seresnext50-comboloss-4\n- 830 epochs: https://www.kaggle.com/code/tonyai007/train-vesuvius-seresnext50-comboloss-5\n\nTraining was halted at 830 epochs due to the competition deadline. Validation results by phase:\n\n| Epochs | Validation Dice | Validation Loss | Best Epoch |\n|--------|-----------------|-----------------|------------|\n| 0-50   | 0.8075          | 0.9753          | 48         |\n| 50-250 | 0.8140          | 0.9396          | 249        |\n| 250-450| 0.8176          | 0.9256          | 444        |\n| 450-600| 0.8193          | 0.9200          | 598        |\n| 600-830| 0.8223          | 0.8937          | 816        |\n\n## 4. Inference & Results\n### 4.1 Single-Model Inference\nThe inference code is based on Tony Li · PaulG · Yiheng Wang's notebook (https://www.kaggle.com/code/tonylica/vesuvius-0-552), with 3-level overlap sliding window and 7-fold TTA.\n\n**Best Submission (Code 1):** https://www.kaggle.com/code/tonyai007/vesuvius-seresnext50-comboloss-830ft\n- Parameters: overlap_public=0.42, overlap_base=0.48, overlap_hi=0.6, T_low=0.5, T_high=0.9, z_radius=0, xy_radius=2, dust_min_size=200\n- Scores: Public LB 0.565, Private LB 0.582\n\n**Ablation on Inference Params:**\n- Adjusting T_low (0.35) and T_high (0.85) led to lower [scores](https://www.kaggle.com/code/tonyai007/inference-vesuvius-surface-3d-detection-test?scriptVersionId=300377344): Public LB 0.558, Private LB 0.570\n- z_radius impact: 0 > 1 > 2 > 3 > 4 (z_radius=0 achieved the best performance)\n\n**Ablation on Training Epochs (Code 1 params):**\n- 50 epochs: Public LB 0.558, Private LB 0.571 (z_radius=1)\n- 250 epochs: Public LB 0.563, Private LB 0.573 (z_radius=0)\n- 450 epochs: Public LB 0.556, Private LB 0.577 (z_radius=0)\n- 600 epochs: Public LB 0.562, Private LB 0.574 (z_radius=0)\n- 830 epochs: Public LB 0.565, Private LB 0.582 (z_radius=0)\n\n### 4.2 Multi-Model Inference\nSubmission (Code 2): https://www.kaggle.com/code/tonyai007/inference-vesuvius-surface-3d-detection-epoch-x5?scriptVersionId=300146063\n- Scores: Public LB 0.566, Private LB 0.577\nKey settings:\nCombination of 3 epoch + 7TTA: Adjustments to **overlap**, **T_low/T_high**, and **z_radius** affected public LB scores, but private LB remained stable at ≈0.577\nCombination of 5 epoch + 3/4TTA: Changes to the same hyperparameters impacted public LB, with private LB consistently around 0.571\nNote: Multi-model fusion improved public LB score but underperformed compared to the single 830-epoch model on private LB (0.577 vs. 0.582), which may have led to misinterpretation.\n\n### 4.3 Baseline Comparisons\n- [SegFormer](https://www.kaggle.com/code/ipythonx/train-vesuvius-surface-3d-detection-on-tpu) (128³, 1200 epochs on TPU, ≈25 hours, 1overlap + 7TTA): Public/Private LB 0.524\n- [nnUNet](https://www.kaggle.com/code/jirkaborovec/surface-nnunet-training-inference-with-2xt4) (3d_lowres/fold_all, 500 epochs on T4 GPU, ≈40 hours): Public/Private LB 0.53 (200ep: 0.513 / 300ep: 0.521 / 400ep: 0.523)\n\n## Acknowledgements\nI thank all upvoters for their support, which helped me earn two Notebook Gold Medals.",
    "3415191": "tonyai007 Congrats. Did you perform all your experiment in kaggle only?",
    "3415225": "Innat Yes, I only used Kaggle's graphics card. Thank you for providing the training and high score prediction code. I have learned a lot"
  },
  "source": "meta"
}