{
  "id": 681341,
  "title": "Amazing results after 4,000 epochs: PB=0.635",
  "url": "/competitions/vesuvius-challenge-surface-detection/discussion/681341",
  "author_name": "",
  "post_date": "2026-03-14T10:23:21.906294100Z",
  "votes": 9,
  "comment_count": 4,
  "views": 0,
  "content": "<p>After reading through the top-ranked solution, I came across the 4,000-epoch log file shared by <a href=\"https://www.kaggle.com/pgeiger\" target=\"_blank\">@pgeiger</a> .\nThe upward trend in the later stages of training caught my attention, so I decided to try running it for 4,000 epochs.</p>\n<table>\n<thead>\n<tr>\n<th>Patch/Epoch</th>\n<th>Mean Validation Dice</th>\n<th>Final Score</th>\n<th>Surf Dice</th>\n<th>VOI Score</th>\n<th>Topo Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>192/1000</strong></td>\n<td>0.6083</td>\n<td>0.6395</td>\n<td>0.8744</td>\n<td>0.5673</td>\n<td>0.4496</td>\n</tr>\n<tr>\n<td><strong>160/4000</strong></td>\n<td>0.6173</td>\n<td>0.6446</td>\n<td>0.8867</td>\n<td>0.5678</td>\n<td>0.4516</td>\n</tr>\n</tbody>\n</table>\n<p>Note: 192/1000 differs from the final CV in our 9th solution because we made minor adjustments.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15255201%2F2de6dec1eff1f4d8cedd49690bb4c192%2FScreenshot%202026-03-14%20181950.png?generation=1773483617375516&amp;alt=media\" alt=\"\">\n160/4000</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15255201%2F33ae42bbecad89cf34e8a170cd2b9e2f%2FScreenshot%202026-03-14%20182033.png?generation=1773483650704097&amp;alt=media\" alt=\"\">\n192/1000</p>\n<h1>By the way, if you use ReLU during training 160, the GPU memory usage is less than 16 GB, so it fits on a T4.</h1>",
  "messages": [
    {
      "id": "3420964",
      "postDate": "03/14/2026 10:23:21",
      "content": "<p>After reading through the top-ranked solution, I came across the 4,000-epoch log file shared by <a href=\"https://www.kaggle.com/pgeiger\" target=\"_blank\">@pgeiger</a> .\nThe upward trend in the later stages of training caught my attention, so I decided to try running it for 4,000 epochs.</p>\n<table>\n<thead>\n<tr>\n<th>Patch/Epoch</th>\n<th>Mean Validation Dice</th>\n<th>Final Score</th>\n<th>Surf Dice</th>\n<th>VOI Score</th>\n<th>Topo Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>192/1000</strong></td>\n<td>0.6083</td>\n<td>0.6395</td>\n<td>0.8744</td>\n<td>0.5673</td>\n<td>0.4496</td>\n</tr>\n<tr>\n<td><strong>160/4000</strong></td>\n<td>0.6173</td>\n<td>0.6446</td>\n<td>0.8867</td>\n<td>0.5678</td>\n<td>0.4516</td>\n</tr>\n</tbody>\n</table>\n<p>Note: 192/1000 differs from the final CV in our 9th solution because we made minor adjustments.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15255201%2F2de6dec1eff1f4d8cedd49690bb4c192%2FScreenshot%202026-03-14%20181950.png?generation=1773483617375516&amp;alt=media\" alt=\"\">\n160/4000</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15255201%2F33ae42bbecad89cf34e8a170cd2b9e2f%2FScreenshot%202026-03-14%20182033.png?generation=1773483650704097&amp;alt=media\" alt=\"\">\n192/1000</p>\n<h1>By the way, if you use ReLU during training 160, the GPU memory usage is less than 16 GB, so it fits on a T4.</h1>",
      "rawMarkdown": "After reading through the top-ranked solution, I came across the 4,000-epoch log file shared by @pgeiger .\nThe upward trend in the later stages of training caught my attention, so I decided to try running it for 4,000 epochs.\n\n| Patch/Epoch | Mean Validation Dice | Final Score | Surf Dice | VOI Score | Topo Score |\n|---|---|---|---|---|---|\n| **192/1000** | 0.6083 | 0.6395 | 0.8744 | 0.5673 | 0.4496 |\n| **160/4000** | 0.6173 | 0.6446 | 0.8867 | 0.5678 | 0.4516 |\n\nNote: 192/1000 differs from the final CV in our 9th solution because we made minor adjustments.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15255201%2F2de6dec1eff1f4d8cedd49690bb4c192%2FScreenshot%202026-03-14%20181950.png?generation=1773483617375516&alt=media)\n160/4000\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15255201%2F33ae42bbecad89cf34e8a170cd2b9e2f%2FScreenshot%202026-03-14%20182033.png?generation=1773483650704097&alt=media)\n192/1000\n\n#By the way, if you use ReLU during training 160, the GPU memory usage is less than 16 GB, so it fits on a T4.",
      "votes": null
    },
    {
      "id": "3420984",
      "postDate": "03/14/2026 11:12:50",
      "content": "<p>Nice result <a href=\"https://www.kaggle.com/ggayoayogg\" target=\"_blank\">@ggayoayogg</a>!</p>\n<p>Could you share whether the validation losses increased during this process? I tried continuing training from one of mine checkpoint and noticed that the Dice loss continued to go down while CE loss started to rise. However, I'm not using nnU-Net training pipeline, so the behavior might be different.</p>\n<p>I also wonder if this improvement is not related to this problem:\n <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/679224\" target=\"_blank\">https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/679224</a></p>\n<blockquote>\n  <p>78 out of 110 test samples (≈71%) come from scroll IDs also present in the training data, meaning a large portion of the test set shares scroll IDs with the training set. This suggests that models may benefit significantly from in-distribution generalization rather than purely out-of-distribution generalization.</p>\n</blockquote>\n<p>Meaning that by training more epochs we could simply be overfitting to some scrolls IDs and this would not be useful to improve result in others out-of-distribution scrolls.</p>",
      "rawMarkdown": "Nice result @ggayoayogg!\n\nCould you share whether the validation losses increased during this process? I tried continuing training from one of mine checkpoint and noticed that the Dice loss continued to go down while CE loss started to rise. However, I'm not using nnU-Net training pipeline, so the behavior might be different.\n\nI also wonder if this improvement is not related to this problem:\n https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/679224\n>78 out of 110 test samples (≈71%) come from scroll IDs also present in the training data, meaning a large portion of the test set shares scroll IDs with the training set. This suggests that models may benefit significantly from in-distribution generalization rather than purely out-of-distribution generalization.\n \n\nMeaning that by training more epochs we could simply be overfitting to some scrolls IDs and this would not be useful to improve result in others out-of-distribution scrolls.",
      "votes": null
    },
    {
      "id": "3420999",
      "postDate": "03/14/2026 11:59:29",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15255201%2F0f941ad5b5f00a73ccfc9423f045dba3%2Fprogress.png?generation=1773489345120790&amp;alt=media\" alt=\"\"></p>\n<p>I feel validation losses hasn't changed.\nI did not monitor the individual curves for DC_SkelREC_and_CE_loss; I only monitored the combined curve.</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15255201%2F0f941ad5b5f00a73ccfc9423f045dba3%2Fprogress.png?generation=1773489345120790&alt=media)\n\nI feel validation losses hasn't changed.\nI did not monitor the individual curves for DC_SkelREC_and_CE_loss; I only monitored the combined curve.",
      "votes": null
    },
    {
      "id": "3421045",
      "postDate": "03/14/2026 13:55:43",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/ggayoayogg\" target=\"_blank\">@ggayoayogg</a>!\nOur loss curves are very different. Maybe because of optmizer and weight decay of nnUnet vs ours. I still have to explore this. For comparison:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2F2132a90d0c480ec1f408ad5c8f36bb76%2Fcurves.png?generation=1773496420505700&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Thanks for sharing @ggayoayogg!\nOur loss curves are very different. Maybe because of optmizer and weight decay of nnUnet vs ours. I still have to explore this. For comparison:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2F2132a90d0c480ec1f408ad5c8f36bb76%2Fcurves.png?generation=1773496420505700&alt=media)",
      "votes": null
    },
    {
      "id": "3421057",
      "postDate": "03/14/2026 15:00:56",
      "content": "<p><a href=\"https://github.com/ultralytics/ultralytics/blob/main/ultralytics/optim/muon.py\" target=\"_blank\">https://github.com/ultralytics/ultralytics/blob/main/ultralytics/optim/muon.py</a></p>",
      "rawMarkdown": "https://github.com/ultralytics/ultralytics/blob/main/ultralytics/optim/muon.py",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3420984,
      "author_name": "sersasj",
      "author_url": "",
      "post_date": "03/14/2026 11:12:50",
      "content": "<p>Nice result <a href=\"https://www.kaggle.com/ggayoayogg\" target=\"_blank\">@ggayoayogg</a>!</p>\n<p>Could you share whether the validation losses increased during this process? I tried continuing training from one of mine checkpoint and noticed that the Dice loss continued to go down while CE loss started to rise. However, I'm not using nnU-Net training pipeline, so the behavior might be different.</p>\n<p>I also wonder if this improvement is not related to this problem:\n <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/679224\" target=\"_blank\">https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/679224</a></p>\n<blockquote>\n  <p>78 out of 110 test samples (≈71%) come from scroll IDs also present in the training data, meaning a large portion of the test set shares scroll IDs with the training set. This suggests that models may benefit significantly from in-distribution generalization rather than purely out-of-distribution generalization.</p>\n</blockquote>\n<p>Meaning that by training more epochs we could simply be overfitting to some scrolls IDs and this would not be useful to improve result in others out-of-distribution scrolls.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3420999,
          "author_name": "ggayoayogg",
          "author_url": "",
          "post_date": "03/14/2026 11:59:29",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15255201%2F0f941ad5b5f00a73ccfc9423f045dba3%2Fprogress.png?generation=1773489345120790&amp;alt=media\" alt=\"\"></p>\n<p>I feel validation losses hasn't changed.\nI did not monitor the individual curves for DC_SkelREC_and_CE_loss; I only monitored the combined curve.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3421045,
              "author_name": "sersasj",
              "author_url": "",
              "post_date": "03/14/2026 13:55:43",
              "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/ggayoayogg\" target=\"_blank\">@ggayoayogg</a>!\nOur loss curves are very different. Maybe because of optmizer and weight decay of nnUnet vs ours. I still have to explore this. For comparison:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2F2132a90d0c480ec1f408ad5c8f36bb76%2Fcurves.png?generation=1773496420505700&amp;alt=media\" alt=\"\"></p>",
              "votes": null,
              "replies": [
                {
                  "id": 3421057,
                  "author_name": "ggayoayogg",
                  "author_url": "",
                  "post_date": "03/14/2026 15:00:56",
                  "content": "<p><a href=\"https://github.com/ultralytics/ultralytics/blob/main/ultralytics/optim/muon.py\" target=\"_blank\">https://github.com/ultralytics/ultralytics/blob/main/ultralytics/optim/muon.py</a></p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3420964": "After reading through the top-ranked solution, I came across the 4,000-epoch log file shared by @pgeiger .\nThe upward trend in the later stages of training caught my attention, so I decided to try running it for 4,000 epochs.\n\n| Patch/Epoch | Mean Validation Dice | Final Score | Surf Dice | VOI Score | Topo Score |\n|---|---|---|---|---|---|\n| **192/1000** | 0.6083 | 0.6395 | 0.8744 | 0.5673 | 0.4496 |\n| **160/4000** | 0.6173 | 0.6446 | 0.8867 | 0.5678 | 0.4516 |\n\nNote: 192/1000 differs from the final CV in our 9th solution because we made minor adjustments.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15255201%2F2de6dec1eff1f4d8cedd49690bb4c192%2FScreenshot%202026-03-14%20181950.png?generation=1773483617375516&alt=media)\n160/4000\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15255201%2F33ae42bbecad89cf34e8a170cd2b9e2f%2FScreenshot%202026-03-14%20182033.png?generation=1773483650704097&alt=media)\n192/1000\n\n#By the way, if you use ReLU during training 160, the GPU memory usage is less than 16 GB, so it fits on a T4.",
    "3420984": "Nice result @ggayoayogg!\n\nCould you share whether the validation losses increased during this process? I tried continuing training from one of mine checkpoint and noticed that the Dice loss continued to go down while CE loss started to rise. However, I'm not using nnU-Net training pipeline, so the behavior might be different.\n\nI also wonder if this improvement is not related to this problem:\n https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/679224\n>78 out of 110 test samples (≈71%) come from scroll IDs also present in the training data, meaning a large portion of the test set shares scroll IDs with the training set. This suggests that models may benefit significantly from in-distribution generalization rather than purely out-of-distribution generalization.\n \n\nMeaning that by training more epochs we could simply be overfitting to some scrolls IDs and this would not be useful to improve result in others out-of-distribution scrolls.",
    "3420999": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15255201%2F0f941ad5b5f00a73ccfc9423f045dba3%2Fprogress.png?generation=1773489345120790&alt=media)\n\nI feel validation losses hasn't changed.\nI did not monitor the individual curves for DC_SkelREC_and_CE_loss; I only monitored the combined curve.",
    "3421045": "Thanks for sharing @ggayoayogg!\nOur loss curves are very different. Maybe because of optmizer and weight decay of nnUnet vs ours. I still have to explore this. For comparison:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2F2132a90d0c480ec1f408ad5c8f36bb76%2Fcurves.png?generation=1773496420505700&alt=media)",
    "3421057": "https://github.com/ultralytics/ultralytics/blob/main/ultralytics/optim/muon.py"
  },
  "source": "meta"
}