{
  "id": 672634,
  "title": "Update: Test set fix, submission rescore, and deadline extension",
  "url": "/competitions/vesuvius-challenge-surface-detection/discussion/672634",
  "author_name": "Giorgio Angelotti",
  "post_date": "2026-02-09T16:21:11.652000",
  "votes": 48,
  "comment_count": 93,
  "views": 0,
  "content": "<p>After identifying a <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/671160\" target=\"_blank\">critical issue</a> in the data and completing a joint analysis with the Kaggle Support team, we will <strong>update the test set</strong> and <strong>rescore submissions</strong>.</p>\n<ul>\n<li><strong>Rescore timing:</strong> The rescore will start later today. Due to the metric’s long runtime, it will likely take <strong>at least a day</strong>, and the public leaderboard may shift during this period.</li>\n<li><strong>Deadline extension:</strong> To ensure everyone has time to calibrate to the updated public leaderboard, we are <strong>extending the competition deadline by 2 weeks</strong> (the competition page will reflect the new date).</li>\n<li><strong>Training data:</strong> After considering community feedback and the work teams have already done to revise training pipelines, we believe updating the training set now could raise fairness concerns. <strong>The training set will not be updated.</strong></li>\n</ul>\n<p>Thanks for your understanding, and let’s keep pushing toward the best possible models to help unwrap the scrolls.</p>",
  "messages": [
    {
      "id": 3403956,
      "postDate": "2026-02-09T16:21:11.653Z",
      "content": "<p>After identifying a <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/671160\" target=\"_blank\">critical issue</a> in the data and completing a joint analysis with the Kaggle Support team, we will <strong>update the test set</strong> and <strong>rescore submissions</strong>.</p>\n<ul>\n<li><strong>Rescore timing:</strong> The rescore will start later today. Due to the metric’s long runtime, it will likely take <strong>at least a day</strong>, and the public leaderboard may shift during this period.</li>\n<li><strong>Deadline extension:</strong> To ensure everyone has time to calibrate to the updated public leaderboard, we are <strong>extending the competition deadline by 2 weeks</strong> (the competition page will reflect the new date).</li>\n<li><strong>Training data:</strong> After considering community feedback and the work teams have already done to revise training pipelines, we believe updating the training set now could raise fairness concerns. <strong>The training set will not be updated.</strong></li>\n</ul>\n<p>Thanks for your understanding, and let’s keep pushing toward the best possible models to help unwrap the scrolls.</p>",
      "rawMarkdown": "After identifying a [critical issue](https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/671160) in the data and completing a joint analysis with the Kaggle Support team, we will **update the test set** and **rescore submissions**.\n\n* **Rescore timing:** The rescore will start later today. Due to the metric’s long runtime, it will likely take **at least a day**, and the public leaderboard may shift during this period.\n* **Deadline extension:** To ensure everyone has time to calibrate to the updated public leaderboard, we are **extending the competition deadline by 2 weeks** (the competition page will reflect the new date).\n* **Training data:** After considering community feedback and the work teams have already done to revise training pipelines, we believe updating the training set now could raise fairness concerns. **The training set will not be updated.**\n\nThanks for your understanding, and let’s keep pushing toward the best possible models to help unwrap the scrolls.",
      "votes": 48
    },
    {
      "id": 3404453,
      "postDate": "2026-02-10T14:29:35.837Z",
      "content": "<p>I understand the frustration, believe me, as we were tied first place going into this change. However, I think maybe we should take it easier on the hosts as they have been participating a ton, answering questions on the weekends, and trying to help out. They really just want the best solution they can get as they are paying for it after all. I can't say I blame them. Obviously, it can be inconvenient for us and the timeframe is not ideal, but I think they are doing what they can.</p>",
      "rawMarkdown": "I understand the frustration, believe me, as we were tied first place going into this change. However, I think maybe we should take it easier on the hosts as they have been participating a ton, answering questions on the weekends, and trying to help out. They really just want the best solution they can get as they are paying for it after all. I can't say I blame them. Obviously, it can be inconvenient for us and the timeframe is not ideal, but I think they are doing what they can.",
      "votes": 29,
      "replies": [
        {
          "id": 3404469,
          "postDate": "2026-02-10T15:01:40.700Z",
          "content": "<p>I would like to echo this sentiment. After all, the end goal is reading the scrolls.</p>",
          "rawMarkdown": "I would like to echo this sentiment. After all, the end goal is reading the scrolls.",
          "votes": 11
        },
        {
          "id": 3404471,
          "postDate": "2026-02-10T15:05:55.740Z",
          "content": "<p>I think people should show respect to the host. The host team wants the best possible outcome, which is why they had to make this difficult decision. I understand that many people are upset about it, especially Chinese Kagglers, for whom the next few weeks will be very inconvenient.</p>\n<p>That said, the host has done almost everything they could and has tried their best to fix the issues and answer people’s questions. They have also taken responsibility for the situation. I hope that in the last two weeks, we can focus on finishing the competition instead of continuing to complain about what has already happened. </p>",
          "rawMarkdown": "I think people should show respect to the host. The host team wants the best possible outcome, which is why they had to make this difficult decision. I understand that many people are upset about it, especially Chinese Kagglers, for whom the next few weeks will be very inconvenient.\n\nThat said, the host has done almost everything they could and has tried their best to fix the issues and answer people’s questions. They have also taken responsibility for the situation. I hope that in the last two weeks, we can focus on finishing the competition instead of continuing to complain about what has already happened. ",
          "votes": 15
        }
      ]
    },
    {
      "id": 3404284,
      "postDate": "2026-02-10T07:32:24.553Z",
      "content": "<p>Not surprisingly, this is going to demotivate those who worked hard till now. Some comments show it already.</p>\n<p>What I don't get is why the training data isn't fixed either? Now you added an extra complexity with a distribution shift between training and testing data.</p>",
      "rawMarkdown": "Not surprisingly, this is going to demotivate those who worked hard till now. Some comments show it already.\n\nWhat I don't get is why the training data isn't fixed either? Now you added an extra complexity with a distribution shift between training and testing data.",
      "votes": 25,
      "replies": [
        {
          "id": 3404305,
          "postDate": "2026-02-10T08:55:41.413Z",
          "rawMarkdown": "",
          "votes": -1,
          "isDeleted": true,
          "replies": [
            {
              "id": 3404383,
              "postDate": "2026-02-10T11:48:36.443Z",
              "content": "<p>Yes, he said that, and I don't understand why they didn't update train as well. Probably because it is time consuming. Which makes our life more difficult if we want to \"fix\" train as well.</p>",
              "rawMarkdown": "Yes, he said that, and I don't understand why they didn't update train as well. Probably because it is time consuming. Which makes our life more difficult if we want to \"fix\" train as well.",
              "votes": 1
            }
          ]
        },
        {
          "id": 3404324,
          "postDate": "2026-02-10T10:08:57.400Z",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> are you willing to share the \"other\" issue you discovered with training data? i do think its unfair they don't clearly state how the test labels have been updated…</p>",
          "rawMarkdown": "@cpmpml are you willing to share the \"other\" issue you discovered with training data? i do think its unfair they don't clearly state how the test labels have been updated...",
          "votes": 2,
          "replies": [
            {
              "id": 3404355,
              "postDate": "2026-02-10T11:05:03.307Z",
              "content": "<p>will do if they confirm the deadline extension.</p>",
              "rawMarkdown": "will do if they confirm the deadline extension.\n\n",
              "votes": 1
            },
            {
              "id": 3404368,
              "postDate": "2026-02-10T11:18:35.383Z",
              "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> The deadline has been extended on the overview page.</p>",
              "rawMarkdown": "@cpmpml The deadline has been extended on the overview page.\n"
            },
            {
              "id": 3404382,
              "postDate": "2026-02-10T11:46:29.497Z",
              "content": "<p>I know, but who knows what they'll decide given the reactions of people here. I prefer to wait for the dust to settle. Will share within 24 hours if there are no changes.</p>",
              "rawMarkdown": "I know, but who knows what they'll decide given the reactions of people here. I prefer to wait for the dust to settle. Will share within 24 hours if there are no changes.\n",
              "votes": 2
            },
            {
              "id": 3404390,
              "postDate": "2026-02-10T12:10:12.467Z",
              "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> i think they will wait to see how much shake up from results of rescore…</p>",
              "rawMarkdown": "@cpmpml i think they will wait to see how much shake up from results of rescore..."
            },
            {
              "id": 3404574,
              "postDate": "2026-02-10T18:26:20.660Z",
              "content": "<p>Rescore being in progress, I thinkt he change is committed to. I'll share tomorrow.</p>",
              "rawMarkdown": "Rescore being in progress, I thinkt he change is committed to. I'll share tomorrow.",
              "votes": 2
            },
            {
              "id": 3405212,
              "postDate": "2026-02-12T12:07:38.103Z",
              "content": "<p>The issue is in my code.</p>\n<p>I have been training with corrupted labels for weeks now…</p>\n<p>Sorry for the false alarm.</p>",
              "rawMarkdown": "The issue is in my code.\n\nI have been training with corrupted labels for weeks now...\n\nSorry for the false alarm.",
              "votes": 1
            },
            {
              "id": 3405218,
              "postDate": "2026-02-12T12:16:20.677Z",
              "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>  I’m sure many people are waiting for you to share it. They must be very disappointed right now.</p>",
              "rawMarkdown": "@cpmpml  I’m sure many people are waiting for you to share it. They must be very disappointed right now.",
              "votes": 1
            },
            {
              "id": 3405231,
              "postDate": "2026-02-12T13:05:24.297Z",
              "content": "<p>I have nothing to share mas explained above.</p>\n<p>I was resizing images and labels  and this introduced spurious positive voxels on the mask boundary. Using the original images and labels removed the issue.</p>\n<p>I found it while working on sharing the issue here. I found the issues, it was in my code (resizing labels).</p>\n<p>I have to redo everything I did. </p>",
              "rawMarkdown": "I have nothing to share mas explained above.\n\nI was resizing images and labels  and this introduced spurious positive voxels on the mask boundary. Using the original images and labels removed the issue.\n\n I found it while working on sharing the issue here. I found the issues, it was in my code (resizing labels).\n\nI have to redo everything I did. ",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 3404005,
      "postDate": "2026-02-09T17:47:31.393Z",
      "content": "<p>Thanks a lot for the effort in handling this update and for keeping the competition aligned with its ultimate goal, helping to read the scrolls, rather than competing for the sake of competition. We really appreciate the transparency and the work from your team.</p>\n<p>Could you please clarify what the fix is exactly?</p>\n<p>More specifically, should participants interpret the rescore as applying the approach mentioned by <a href=\"https://www.kaggle.com/dankrstev\" target=\"_blank\">@dankrstev</a>? That is, keeping the data unchanged but adjusting the metric so it is computed as if no holes exist? Or was any additional processing applied, such as a binary closing or another morphological operation??\n<a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/671160#3403175\" target=\"_blank\">https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/671160#3403175</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2F3fa5a1afd5feab38662b5c03600d68a7%2Faa.png?generation=1770659017538322&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Thanks a lot for the effort in handling this update and for keeping the competition aligned with its ultimate goal, helping to read the scrolls, rather than competing for the sake of competition. We really appreciate the transparency and the work from your team.\n\nCould you please clarify what the fix is exactly?\n\nMore specifically, should participants interpret the rescore as applying the approach mentioned by @dankrstev? That is, keeping the data unchanged but adjusting the metric so it is computed as if no holes exist? Or was any additional processing applied, such as a binary closing or another morphological operation??\nhttps://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/671160#3403175\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2F3fa5a1afd5feab38662b5c03600d68a7%2Faa.png?generation=1770659017538322&alt=media)\n",
      "votes": 18,
      "replies": [
        {
          "id": 3404162,
          "postDate": "2026-02-10T00:36:27.527Z",
          "content": "<p>I need we need to have the patch released in the earliest time for is to react.<br>\ne.g.<br>\n1) is the patch in metric computation?<br>\nthen relese the new code  </p>\n<p>2) is the patch in test data?<br>\nthen release new train labels that is patched similarly as the test or<br>\ncode to transform the data.  (eg filling micro holes)</p>",
          "rawMarkdown": "I need we need to have the patch released in the earliest time for is to react.  \ne.g.   \n1) is the patch in metric computation?  \nthen relese the new code  \n\n2) is the patch in test data?  \nthen release new train labels that is patched similarly as the test or  \ncode to transform the data.  (eg filling micro holes)",
          "votes": 13
        }
      ]
    },
    {
      "id": 3404334,
      "postDate": "2026-02-10T10:26:48Z",
      "content": "<p>Could you please disclose fix points?\nWe can't fix local pipeline…</p>\n<p>if this points are not important, we don't need extend long deadline… </p>",
      "rawMarkdown": "Could you please disclose fix points?\nWe can't fix local pipeline...\n\nif this points are not important, we don't need extend long deadline... ",
      "votes": 20,
      "replies": [
        {
          "id": 3404337,
          "postDate": "2026-02-10T10:32:45.347Z",
          "content": "<blockquote>\n  <p>if this points are not important, we don't need extend long deadline…</p>\n</blockquote>\n<p>Exactly, totally agree! </p>",
          "rawMarkdown": "> if this points are not important, we don't need extend long deadline…\n\nExactly, totally agree! ",
          "votes": 6
        }
      ]
    },
    {
      "id": 3404247,
      "postDate": "2026-02-10T05:25:53.033Z",
      "content": "<p>Haha, I am so angry I can't make any constructive comment. Not sure I will continue this comp. </p>",
      "rawMarkdown": "Haha, I am so angry I can't make any constructive comment. Not sure I will continue this comp. ",
      "votes": 21,
      "replies": [
        {
          "id": 3404251,
          "postDate": "2026-02-10T05:37:58.043Z",
          "content": "<p>Looks like I quit at the right time 😉, but I think one week extension was the most that should have been given, simply because only the leaderboard score has changed and not the data given to participants.</p>",
          "rawMarkdown": "Looks like I quit at the right time 😉, but I think one week extension was the most that should have been given, simply because only the leaderboard score has changed and not the data given to participants.",
          "votes": 2
        },
        {
          "id": 3404260,
          "postDate": "2026-02-10T05:54:29.237Z",
          "content": "<p>At the beginning, I chose to join this competition because it was scheduled to end just one day before the Chinese New Year, which was perfect timing for me and most Chinese people. Now, without any prior notice (although updates were expected, no one ever indicated it would be extended), the competition has been extended by two weeks, completely overlapping with the Spring Festival. I feel both angry and powerless🤐.</p>",
          "rawMarkdown": "At the beginning, I chose to join this competition because it was scheduled to end just one day before the Chinese New Year, which was perfect timing for me and most Chinese people. Now, without any prior notice (although updates were expected, no one ever indicated it would be extended), the competition has been extended by two weeks, completely overlapping with the Spring Festival. I feel both angry and powerless🤐.",
          "votes": 12,
          "replies": [
            {
              "id": 3404266,
              "postDate": "2026-02-10T06:27:07.850Z",
              "content": "<p>Yes, people plan vacations etc around competition end. Same for me. What makes me angry is the proportionality. Impacting 1600 participants time plans, just to weight topo score a bit higher. </p>",
              "rawMarkdown": "Yes, people plan vacations etc around competition end. Same for me. What makes me angry is the proportionality. Impacting 1600 participants time plans, just to weight topo score a bit higher. ",
              "votes": 11
            },
            {
              "id": 3404856,
              "postDate": "2026-02-11T11:19:17.417Z",
              "content": "<p>Yes, totally agree! Extending it by two weeks would completely mess up my vacation.</p>",
              "rawMarkdown": "Yes, totally agree! Extending it by two weeks would completely mess up my vacation."
            }
          ]
        }
      ]
    },
    {
      "id": 3404281,
      "postDate": "2026-02-10T07:23:37.637Z",
      "content": "<p>Honestly, I’m quite disappointed about the deadline extension. One big reason I joined this competition was because it was supposed to end before the Lunar New Year. Many of us planned our time around that. Extending it into the holiday really affects those plans and makes it very hard to balance family time and the competition.\nI understand that fixing the data is important, but changing the timeline this late — especially into a major holiday — feels really tough for a lot of Chinese participants.</p>",
      "rawMarkdown": "Honestly, I’m quite disappointed about the deadline extension. One big reason I joined this competition was because it was supposed to end before the Lunar New Year. Many of us planned our time around that. Extending it into the holiday really affects those plans and makes it very hard to balance family time and the competition.\nI understand that fixing the data is important, but changing the timeline this late — especially into a major holiday — feels really tough for a lot of Chinese participants.",
      "votes": 16
    },
    {
      "id": 3404373,
      "postDate": "2026-02-10T11:24:45.403Z",
      "content": "<p>I believe that revising the test set is not directly related to postponing the competition, since the training set will not change and most teams have already  completed their models. I hope the organizers can take participants’ feelings into consideration—everyone has put a great deal of effort into this competition, whether in terms of time or the cost of renting GPUs. A short extension would be understandable, but a two-week delay is really too long and quite stressful for us.</p>",
      "rawMarkdown": "I believe that revising the test set is not directly related to postponing the competition, since the training set will not change and most teams have already  completed their models. I hope the organizers can take participants’ feelings into consideration—everyone has put a great deal of effort into this competition, whether in terms of time or the cost of renting GPUs. A short extension would be understandable, but a two-week delay is really too long and quite stressful for us.",
      "votes": 13,
      "replies": [
        {
          "id": 3404404,
          "postDate": "2026-02-10T12:34:11.480Z",
          "content": "<p>Especially for friends in the East who need to celebrate the Chinese/Lunar New Year</p>",
          "rawMarkdown": "Especially for friends in the East who need to celebrate the Chinese/Lunar New Year",
          "votes": 6
        }
      ]
    },
    {
      "id": 3404026,
      "postDate": "2026-02-09T18:36:33.917Z",
      "content": "<p>Let’s be real: this extension is a double-edged sword for Chinese participants. We are entering the Spring Festival—the most important time for family reunions in our culture. While the extra time is appreciated, it effectively means choosing between quality time with family and squeezing out that last bit of model performance. </p>\n<p>It’s going to be a tough grind during the holidays! </p>",
      "rawMarkdown": "Let’s be real: this extension is a double-edged sword for Chinese participants. We are entering the Spring Festival—the most important time for family reunions in our culture. While the extra time is appreciated, it effectively means choosing between quality time with family and squeezing out that last bit of model performance. \n\nIt’s going to be a tough grind during the holidays! ",
      "votes": 14,
      "replies": [
        {
          "id": 3404229,
          "postDate": "2026-02-10T04:33:31.410Z",
          "content": "<p>That’s so struggling. For someone working in the Internet Company, it may be the most relaxing time in 1 year. Now it may become the most exhausting time.</p>",
          "rawMarkdown": "That’s so struggling. For someone working in the Internet Company, it may be the most relaxing time in 1 year. Now it may become the most exhausting time.",
          "votes": 4
        },
        {
          "id": 3404270,
          "postDate": "2026-02-10T06:38:04.353Z",
          "content": "<p>抱抱 过年期间还要打比赛却是压力太大了</p>",
          "rawMarkdown": "抱抱 过年期间还要打比赛却是压力太大了",
          "votes": 5
        }
      ]
    },
    {
      "id": 3403961,
      "postDate": "2026-02-09T16:32:53.523Z",
      "content": "<p><a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> \nThanks. I’m curious, after fixing the test set and evaluation metrics, what score did your baseline nnU-Net achieve without any post-processing?</p>",
      "rawMarkdown": "@giorgioangelotti \nThanks. I’m curious, after fixing the test set and evaluation metrics, what score did your baseline nnU-Net achieve without any post-processing?",
      "votes": 12
    },
    {
      "id": 3404439,
      "postDate": "2026-02-10T13:54:14Z",
      "content": "<p>Has anyone received their re-graded scores yet? I noticed that my new uploads still seem to use the old grading method.</p>",
      "rawMarkdown": "Has anyone received their re-graded scores yet? I noticed that my new uploads still seem to use the old grading method.",
      "votes": 9,
      "replies": [
        {
          "id": 3404458,
          "postDate": "2026-02-10T14:42:37.893Z",
          "content": "<p>My score also have not changed. And I face some strange issues: 1. one submission show \"kaggle error\", but after few minutes late it turn to success and get a score. 2. I save and run a notebook version, it should be end in about 6 minutes, but it has runned for more than 10h yet. And I can't stop it.</p>\n<p>It is so confusing.</p>",
          "rawMarkdown": "My score also have not changed. And I face some strange issues: 1. one submission show \"kaggle error\", but after few minutes late it turn to success and get a score. 2. I save and run a notebook version, it should be end in about 6 minutes, but it has runned for more than 10h yet. And I can't stop it.\n\n It is so confusing.",
          "votes": 2,
          "replies": [
            {
              "id": 3404461,
              "postDate": "2026-02-10T14:47:09.653Z",
              "content": "<p>Local testing shows that the CV without holes provides roughly a 1% boost, but current results are identical. Some players on the leaderboard appear to have received a buff.</p>",
              "rawMarkdown": "Local testing shows that the CV without holes provides roughly a 1% boost, but current results are identical. Some players on the leaderboard appear to have received a buff.",
              "votes": 2
            },
            {
              "id": 3404466,
              "postDate": "2026-02-10T14:55:33.887Z",
              "content": "<ol>\n<li><p>Will hole filled train samples changed training results? Likely no. It is unlikely that your model model learned to predict holes in your old training.</p></li>\n<li><p>Will hole filled test samples affected results? Yes if you have very accurate prediction. No if your prediction is not accurate. </p></li>\n</ol>\n<p>Here is a test: try perturbation ( small and big shift etc) on hole and filled gt. Then measure the topo score (eg score(filled, shift filled)). My experiments showed that if your local lb is in 0.60 range, it is more likely to be affected.</p>",
              "rawMarkdown": "1. Will hole filled train samples changed training results? Likely no. It is unlikely that your model model learned to predict holes in your old training.\n\n2. Will hole filled test samples affected results? Yes if you have very accurate prediction. No if your prediction is not accurate. \n\nHere is a test: try perturbation ( small and big shift etc) on hole and filled gt. Then measure the topo score (eg score(filled, shift filled)). My experiments showed that if your local lb is in 0.60 range, it is more likely to be affected.",
              "votes": 6
            }
          ]
        },
        {
          "id": 3404477,
          "postDate": "2026-02-10T15:27:21Z",
          "content": "<p>How can you know if they use the old scoring method or the new one?</p>",
          "rawMarkdown": "How can you know if they use the old scoring method or the new one?",
          "replies": [
            {
              "id": 3404478,
              "postDate": "2026-02-10T15:28:15.573Z",
              "content": "<p>Control group。\nHave you received your new scores? Looks like you've made significant progress.</p>",
              "rawMarkdown": "Control group。\nHave you received your new scores? Looks like you've made significant progress."
            },
            {
              "id": 3404481,
              "postDate": "2026-02-10T15:33:18.437Z",
              "content": "<p>I think it’s the score after updating.</p>",
              "rawMarkdown": "I think it’s the score after updating.",
              "votes": 1
            },
            {
              "id": 3404482,
              "postDate": "2026-02-10T15:33:18.477Z",
              "content": "<p>I'm actually not sure if it's the new or old method. I'm just judging whether they used the new scoring system based on the fact that the scores haven't changed.</p>",
              "rawMarkdown": "I'm actually not sure if it's the new or old method. I'm just judging whether they used the new scoring system based on the fact that the scores haven't changed.",
              "votes": 2
            }
          ]
        },
        {
          "id": 3404489,
          "postDate": "2026-02-10T15:50:38.433Z",
          "content": "<p>Same here. None of my old submissions seem to have been rescored and new ones are pretty much in the range of what I would expect with the old scoring process.\nConsidering that it takes ~3-4 hours to score the predictions and the thousands of submissions to process, I would not be surprised if it takes another few days before LB stabilises.</p>",
          "rawMarkdown": "Same here. None of my old submissions seem to have been rescored and new ones are pretty much in the range of what I would expect with the old scoring process.\nConsidering that it takes ~3-4 hours to score the predictions and the thousands of submissions to process, I would not be surprised if it takes another few days before LB stabilises.",
          "votes": 3,
          "replies": [
            {
              "id": 3404521,
              "postDate": "2026-02-10T17:03:22.010Z",
              "content": "<p>Same here. Our recent submission on the leaderboard was sent during the rescore, but none of my other submissions have changed. I’m not sure if this is the new scoring yet; I was using a brand-new model, so there’s no correlation to my old submissions. Because of that, I can't tell if the 0.005 bump was from the new scoring or the model itself.</p>",
              "rawMarkdown": "Same here. Our recent submission on the leaderboard was sent during the rescore, but none of my other submissions have changed. I’m not sure if this is the new scoring yet; I was using a brand-new model, so there’s no correlation to my old submissions. Because of that, I can't tell if the 0.005 bump was from the new scoring or the model itself.",
              "votes": 3
            }
          ]
        },
        {
          "id": 3404540,
          "postDate": "2026-02-10T17:35:23.807Z",
          "content": "<p>The rescore is in progress. We currently have about 8% of submissions rescored. Because this metric is fairly computationally intensive, it's likely going to be a couple days before everything is complete. The new scores are posted as soon as they're ready, so the leaderboard will show a mix of old and new until it's done. New submissions should be using the updated labels though.</p>",
          "rawMarkdown": "The rescore is in progress. We currently have about 8% of submissions rescored. Because this metric is fairly computationally intensive, it's likely going to be a couple days before everything is complete. The new scores are posted as soon as they're ready, so the leaderboard will show a mix of old and new until it's done. New submissions should be using the updated labels though.",
          "votes": 11,
          "replies": [
            {
              "id": 3404543,
              "postDate": "2026-02-10T17:41:24.067Z",
              "rawMarkdown": "",
              "votes": -2,
              "isDeleted": true
            },
            {
              "id": 3404586,
              "postDate": "2026-02-10T18:43:31.250Z",
              "content": "<p>Are new submissions scored with the new metric during this transition time?</p>",
              "rawMarkdown": "Are new submissions scored with the new metric during this transition time?"
            },
            {
              "id": 3404594,
              "postDate": "2026-02-10T19:01:56.353Z",
              "content": "<p>Yes. The new test labels are live on the site.</p>",
              "rawMarkdown": "Yes. The new test labels are live on the site.",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 3403985,
      "postDate": "2026-02-09T17:09:55.030Z",
      "content": "<p>Spent tons of time during the past days and tired, now  it would take more time because of extension deadline even in Spring festival 😱. Tired and tired.. But anyway, this kind of update I think it is the best way for fairness of the competetion.</p>",
      "rawMarkdown": "Spent tons of time during the past days and tired, now  it would take more time because of extension deadline even in Spring festival 😱. Tired and tired.. But anyway, this kind of update I think it is the best way for fairness of the competetion.",
      "votes": 8
    },
    {
      "id": 3404838,
      "postDate": "2026-02-11T10:50:24.883Z",
      "content": "<p>Since only the test set was updated, I believe this has introduced an additional challenge of dealing with a potential domain shift. Could you please share a bit more detail on what was changed?</p>\n<p>For example:</p>\n<ul>\n<li>Applied some algorithmic post-processing to fill small holes in the labels.</li>\n<li>Removed highly porous samples from the test set, or otherwise filtered/replaced them.</li>\n</ul>",
      "rawMarkdown": "Since only the test set was updated, I believe this has introduced an additional challenge of dealing with a potential domain shift. Could you please share a bit more detail on what was changed?\n\nFor example:\n- Applied some algorithmic post-processing to fill small holes in the labels.\n- Removed highly porous samples from the test set, or otherwise filtered/replaced them.",
      "votes": 8
    },
    {
      "id": 3404318,
      "postDate": "2026-02-10T09:46:37.773Z",
      "content": "<p>I didn't understand the reason for the extension, since only the test data was updated.</p>",
      "rawMarkdown": "I didn't understand the reason for the extension, since only the test data was updated.",
      "votes": 8
    },
    {
      "id": 3404406,
      "postDate": "2026-02-10T12:37:18.337Z",
      "content": "<p>I also think that the organisers are totally ignoring that <strong>humans</strong> are taking part in this competition. People have worked really hard, made sacrifices, and organised their personal life in such a way that they would be able to work extra hard until Friday. I was personally really looking forward to Friday.</p>\n<p>Moreover, changing the test set and not the training set is problematic. 2 weeks is not that much time to adapt our strategy, especially when the rescoring is taking time and the distribution shift between training and test sets means we now rely on LB even more. It does not leave much time to retrain and fine-tune models considering that the models take days to train with a good machine.</p>\n<p>While I love this project and I still have plenty of ideas to improve my solution, I am still unsure whether I will continue the competition as it's been an absolute chaos and new potential bugs have already been mentioned. Yes, the metric is \"a bit\" broken and it is clearly an issue from a data science point of view, but I think moving the goal posts at the last minute is not the solution.</p>",
      "rawMarkdown": "I also think that the organisers are totally ignoring that **humans** are taking part in this competition. People have worked really hard, made sacrifices, and organised their personal life in such a way that they would be able to work extra hard until Friday. I was personally really looking forward to Friday.\n\nMoreover, changing the test set and not the training set is problematic. 2 weeks is not that much time to adapt our strategy, especially when the rescoring is taking time and the distribution shift between training and test sets means we now rely on LB even more. It does not leave much time to retrain and fine-tune models considering that the models take days to train with a good machine.\n\nWhile I love this project and I still have plenty of ideas to improve my solution, I am still unsure whether I will continue the competition as it's been an absolute chaos and new potential bugs have already been mentioned. Yes, the metric is \"a bit\" broken and it is clearly an issue from a data science point of view, but I think moving the goal posts at the last minute is not the solution.",
      "votes": 9
    },
    {
      "id": 3404899,
      "postDate": "2026-02-11T14:16:17.323Z",
      "content": "<p>The rescore is now complete. Please let me know if you have any questions or concerns about the rescore.</p>",
      "rawMarkdown": "The rescore is now complete. Please let me know if you have any questions or concerns about the rescore.",
      "votes": 7,
      "isPinned": true,
      "replies": [
        {
          "id": 3404905,
          "postDate": "2026-02-11T14:22:20.563Z",
          "content": "<p>So what was fixed in the test set?</p>",
          "rawMarkdown": "So what was fixed in the test set?",
          "votes": 3
        },
        {
          "id": 3404914,
          "postDate": "2026-02-11T15:00:36.260Z",
          "content": "<p>From the comments in the discussion it sounds like the metric change was a patch to g_k = 0. I am currently trying to compute cv based off a monkey patch based on this but if we could be more concrete about exactly what changed so we can work on solving the problems with the newly alloted time that would be helpful. Below is an example of how I am doing it as well as some gpt comments (I dont have certainty this is correct)\n`</p>\n<h1>============================================================</h1>\n<h1>SCORE submission.zip ON TRAIN IDS (Official TopoMetric) + pred_holes diagnostic</h1>\n<h1>============================================================</h1>\n<p>import os, zipfile, subprocess, sys, importlib, shutil\nimport numpy as np\nimport pandas as pd\nfrom pathlib import Path\nfrom PIL import Image, ImageSequence\nfrom collections import deque</p>\n<h1>-----------------------------</h1>\n<h1>Config</h1>\n<h1>-----------------------------</h1>\n<p>SUBMISSION_ZIP = Path(\"submission.zip\")\nEXTRACT_DIR    = Path(\"./_sub_eval_extract\")\nLABEL_DIR      = Path(\"/kaggle/input/vesuvius-challenge-surface-detection/train_labels\")</p>\n<p>EVAL_IDS = [\n    \"2257172177\", \"663106834\", \"2555675774\",\n    \"1294570892\", \"2550619349\", \"3854101708\",\n    \"406944815\", \"2059872339\", \"360339268\",\n    \"1693721638\", \"508375143\", \"687559918\",\n    \"3834183540\", \"2689046967\", \"2454492741\",\n    \"370627915\", \"1837890067\", \"398977019\",\n    \"315189226\", \"239904888\", \"725727411\"\n]\nSURFACE_TOLERANCE = 2.0</p>\n<h1>============================================================</h1>\n<h1>HELPERS</h1>\n<h1>============================================================</h1>\n<p>def load_volume_tif(path: Path) -&gt; np.ndarray:\n    \"\"\"Load TIFF stack -&gt; (D,H,W)\"\"\"\n    im = Image.open(path)\n    slices = [np.array(frame) for frame in ImageSequence.Iterator(im)]\n    return np.stack(slices, axis=0)</p>\n<p>def maybe_transpose_gt(gt: np.ndarray, pred_shape: tuple) -&gt; np.ndarray:\n    \"\"\"\n    Fix common GT orientation mismatch (GT stored as (H,W,D)).\n    If GT last dim matches pred depth, transpose.\n    \"\"\"\n    if gt.shape == pred_shape:\n        return gt\n    if gt.ndim == 3 and gt.shape[-1] == pred_shape[0] and gt.shape[:2] == pred_shape[1:]:\n        return np.transpose(gt, (2, 0, 1))\n    return gt</p>\n<h1>============================================================</h1>\n<h1>PREDICTED HOLES (ENCLOSED VOIDS) DIAGNOSTIC</h1>\n<h1>============================================================</h1>\n<p>def count_enclosed_voids(binary_vol: np.ndarray) -&gt; int:\n    \"\"\"\n    Approximate 'holes' as enclosed 0-regions in the complement of predicted foreground.\n    binary_vol: (D,H,W) uint8 {0,1}\n    Returns number of enclosed void components.\n    \"\"\"\n    vol = (binary_vol &gt; 0)\n    D, H, W = vol.shape</p>\n<pre><code># Complement background\nbg = ~vol\nseen = np.zeros(bg.shape, dtype=bool)\n\nq = deque()\n\ndef push(z, y, x):\n    if 0 &lt;= z &lt; D and 0 &lt;= y &lt; H and 0 &lt;= x &lt; W and bg[z, y, x] and not seen[z, y, x]:\n        seen[z, y, x] = True\n        q.append((z, y, x))\n\n# Flood fill from boundary background voxels -&gt; marks \"outside\"\nfor z in (0, D - 1):\n    for y in range(H):\n        for x in range(W):\n            push(z, y, x)\nfor z in range(D):\n    for y in (0, H - 1):\n        for x in range(W):\n            push(z, y, x)\nfor z in range(D):\n    for y in range(H):\n        for x in (0, W - 1):\n            push(z, y, x)\n\n# 6-neighborhood\nnbrs = [(1, 0, 0), (-1, 0, 0),\n        (0, 1, 0), (0, -1, 0),\n        (0, 0, 1), (0, 0, -1)]\n\nwhile q:\n    z, y, x = q.popleft()\n    for dz, dy, dx in nbrs:\n        push(z + dz, y + dy, x + dx)\n\n# Background not reached from boundary = enclosed cavities\ninside = bg &amp; (~seen)\n\n# Count connected components in inside (6-neighborhood)\nholes = 0\nseen2 = np.zeros_like(inside, dtype=bool)\nfor z in range(D):\n    for y in range(H):\n        for x in range(W):\n            if inside[z, y, x] and not seen2[z, y, x]:\n                holes += 1\n                qq = deque([(z, y, x)])\n                seen2[z, y, x] = True\n                while qq:\n                    zz, yy, xx = qq.popleft()\n                    for dz, dy, dx in nbrs:\n                        nz, ny, nx = zz + dz, yy + dy, xx + dx\n                        if 0 &lt;= nz &lt; D and 0 &lt;= ny &lt; H and 0 &lt;= nx &lt; W:\n                            if inside[nz, ny, nx] and not seen2[nz, ny, nx]:\n                                seen2[nz, ny, nx] = True\n                                qq.append((nz, ny, nx))\nreturn holes\n</code></pre>\n<h1>============================================================</h1>\n<h1>TOPO METRICS INSTALL</h1>\n<h1>============================================================</h1>\n<p>_TOPO_OK = False</p>\n<p>def install_topometrics():\n    global _TOPO_OK\n    if _TOPO_OK:\n        return\n    try:\n        import topometrics.leaderboard\n        _TOPO_OK = True\n        return\n    except:\n        pass</p>\n<pre><code>resources = \"/kaggle/input/vesuvius-metric-resources\"\nworkdir   = \"/kaggle/working/topological-metrics-kaggle\"\n\nsubprocess.run(\n    f\"cd {resources} &amp;&amp; uv pip install --no-index --find-links=wheels \"\n    f\"-r topological-metrics-kaggle/requirements.txt\",\n    shell=True, check=True,\n)\n\nsubprocess.run(\n    f\"cp -r {resources}/topological-metrics-kaggle /kaggle/working/\",\n    shell=True, check=True\n)\n\nsubprocess.run(\n    f\"cd {workdir} &amp;&amp; chmod +x scripts/setup_submodules.sh scripts/build_betti.sh &amp;&amp; make build-betti\",\n    shell=True, check=True\n)\n\nsubprocess.run(\n    f\"cd {workdir} &amp;&amp; uv pip install -e . --no-deps --no-index --no-build-isolation\",\n    shell=True, check=True\n)\n\nsys.path.append(f\"{workdir}/src\")\nimportlib.invalidate_caches()\n\nimport topometrics.leaderboard\n_TOPO_OK = True\n</code></pre>\n<h1>============================================================</h1>\n<h1>OFFICIAL SCORING CALL</h1>\n<h1>============================================================</h1>\n<p>def score_official(gt_u8, pr_u8):\n    \"\"\"\n    gt_u8: (D,H,W) uint8 with ignore label=2 preserved\n    pr_u8: (D,H,W) uint8 {0,1}\n    \"\"\"\n    install_topometrics()\n    import topometrics.leaderboard as lb</p>\n<pre><code>r = lb.compute_leaderboard_score(\n    predictions=pr_u8.astype(np.uint8),\n    labels=gt_u8.astype(np.uint8),\n    dims=(0, 1, 2),\n    spacing=(1.0, 1.0, 1.0),\n    surface_tolerance=SURFACE_TOLERANCE,\n    voi_connectivity=26,\n    voi_transform=\"one_over_one_plus\",\n    voi_alpha=0.3,\n    combine_weights=(0.3, 0.35, 0.35),\n    ignore_label=2,\n)\n\n# VOI attribute differs in some versions\nvoi_score = r.voi.voi_score if hasattr(r.voi, \"voi_score\") else r.voi.oi_score\n\nreturn dict(\n    score=float(r.score),\n    topo=float(r.topo.toposcore),\n    surface_dice=float(r.surface_dice),\n    voi=float(voi_score),\n)\n</code></pre>\n<h1>============================================================</h1>\n<h1>UNZIP SUBMISSION</h1>\n<h1>============================================================</h1>\n<p>if EXTRACT_DIR.exists():\n    shutil.rmtree(EXTRACT_DIR)\nEXTRACT_DIR.mkdir(parents=True, exist_ok=True)</p>\n<p>with zipfile.ZipFile(SUBMISSION_ZIP, \"r\") as z:\n    z.extractall(EXTRACT_DIR)</p>\n<p>print(f\"✓ Extracted {SUBMISSION_ZIP} → {EXTRACT_DIR}\")</p>\n<h1>============================================================</h1>\n<h1>SCORE EACH ID</h1>\n<h1>============================================================</h1>\n<p>rows = []</p>\n<p>for sid in EVAL_IDS:\n    pred_path = EXTRACT_DIR / f\"{sid}.tif\"\n    gt_path   = LABEL_DIR / f\"{sid}.tif\"</p>\n<pre><code>if not pred_path.exists():\n    print(f\"❌ Missing prediction for {sid}\")\n    continue\nif not gt_path.exists():\n    print(f\"❌ Missing GT for {sid}\")\n    continue\n\npred = load_volume_tif(pred_path)\ngt   = load_volume_tif(gt_path)\n\n# Prediction is binary {0,1}\npr_u8 = (pred &gt; 0).astype(np.uint8)\n\n# GT must keep ignore label=2\n# Map: 2 → 2, (0 → 0), (nonzero but not 2 → 1)\ngt_raw = gt.astype(np.uint8)\ngt_u8 = np.where(gt_raw == 2, 2, (gt_raw &gt; 0).astype(np.uint8))\n\n# Fix possible (H,W,D) → (D,H,W)\ngt_u8 = maybe_transpose_gt(gt_u8, pr_u8.shape)\n\nif gt_u8.shape != pr_u8.shape:\n    raise ValueError(f\"Shape mismatch for {sid}: pred={pr_u8.shape}, gt={gt_u8.shape}\")\n\n# Diagnostic: predicted enclosed voids (\"holes\")\npred_holes = count_enclosed_voids(pr_u8)\n\nprint(f\"[{sid}] pred uniques={np.unique(pr_u8)}  gt uniques={np.unique(gt_u8)}  pred_holes={pred_holes}\")\n\nm = score_official(gt_u8, pr_u8)\nm[\"id\"] = sid\nm[\"pred_holes\"] = int(pred_holes)\nrows.append(m)\n\nprint(f\"✓ {sid}: score={m['score']:.4f} | topo={m['topo']:.4f} | sd={m['surface_dice']:.4f} | voi={m['voi']:.4f} | holes={m['pred_holes']}\")\n</code></pre>\n<h1>============================================================</h1>\n<h1>SUMMARY TABLE</h1>\n<h1>============================================================</h1>\n<p>df = pd.DataFrame(rows)</p>\n<p>print(\"\\n=== Per-volume ===\")\ncols = [\"id\", \"score\", \"topo\", \"surface_dice\", \"voi\", \"pred_holes\"]\nprint(df[cols].to_string(index=False))</p>\n<p>agg_cols = [\"score\", \"topo\", \"surface_dice\", \"voi\", \"pred_holes\"]\nmean = df[agg_cols].mean().add_prefix(\"mean_\")\nstd  = df[agg_cols].std(ddof=1).add_prefix(\"std_\")</p>\n<h1>CV (coefficient of variation) not very meaningful for pred_holes if mean is small,</h1>\n<h1>but we compute anyway for consistency.</h1>\n<p>cv   = (df[agg_cols].std(ddof=1) / df[agg_cols].mean().replace(0, np.nan)).add_prefix(\"cv_\")</p>\n<p>summary = pd.concat([mean, std, cv], axis=0).to_frame().T</p>\n<p>print(\"\\n=== Summary ===\")\nprint(summary.to_string(index=False))</p>\n<h1>Optional: save outputs</h1>\n<p>df.to_csv(\"cv_scores_with_holes.csv\", index=False)\nsummary.to_csv(\"cv_summary_with_holes.csv\", index=False)</p>\n<p>summary\n`</p>",
          "rawMarkdown": "From the comments in the discussion it sounds like the metric change was a patch to g_k = 0. I am currently trying to compute cv based off a monkey patch based on this but if we could be more concrete about exactly what changed so we can work on solving the problems with the newly alloted time that would be helpful. Below is an example of how I am doing it as well as some gpt comments (I dont have certainty this is correct)\n`\n# ============================================================\n# SCORE submission.zip ON TRAIN IDS (Official TopoMetric) + pred_holes diagnostic\n# ============================================================\nimport os, zipfile, subprocess, sys, importlib, shutil\nimport numpy as np\nimport pandas as pd\nfrom pathlib import Path\nfrom PIL import Image, ImageSequence\nfrom collections import deque\n\n# -----------------------------\n# Config\n# -----------------------------\nSUBMISSION_ZIP = Path(\"submission.zip\")\nEXTRACT_DIR    = Path(\"./_sub_eval_extract\")\nLABEL_DIR      = Path(\"/kaggle/input/vesuvius-challenge-surface-detection/train_labels\")\n\nEVAL_IDS = [\n    \"2257172177\", \"663106834\", \"2555675774\",\n    \"1294570892\", \"2550619349\", \"3854101708\",\n    \"406944815\", \"2059872339\", \"360339268\",\n    \"1693721638\", \"508375143\", \"687559918\",\n    \"3834183540\", \"2689046967\", \"2454492741\",\n    \"370627915\", \"1837890067\", \"398977019\",\n    \"315189226\", \"239904888\", \"725727411\"\n]\nSURFACE_TOLERANCE = 2.0\n\n\n# ============================================================\n# HELPERS\n# ============================================================\ndef load_volume_tif(path: Path) -> np.ndarray:\n    \"\"\"Load TIFF stack -> (D,H,W)\"\"\"\n    im = Image.open(path)\n    slices = [np.array(frame) for frame in ImageSequence.Iterator(im)]\n    return np.stack(slices, axis=0)\n\n\ndef maybe_transpose_gt(gt: np.ndarray, pred_shape: tuple) -> np.ndarray:\n    \"\"\"\n    Fix common GT orientation mismatch (GT stored as (H,W,D)).\n    If GT last dim matches pred depth, transpose.\n    \"\"\"\n    if gt.shape == pred_shape:\n        return gt\n    if gt.ndim == 3 and gt.shape[-1] == pred_shape[0] and gt.shape[:2] == pred_shape[1:]:\n        return np.transpose(gt, (2, 0, 1))\n    return gt\n\n\n# ============================================================\n# PREDICTED HOLES (ENCLOSED VOIDS) DIAGNOSTIC\n# ============================================================\ndef count_enclosed_voids(binary_vol: np.ndarray) -> int:\n    \"\"\"\n    Approximate 'holes' as enclosed 0-regions in the complement of predicted foreground.\n    binary_vol: (D,H,W) uint8 {0,1}\n    Returns number of enclosed void components.\n    \"\"\"\n    vol = (binary_vol > 0)\n    D, H, W = vol.shape\n\n    # Complement background\n    bg = ~vol\n    seen = np.zeros(bg.shape, dtype=bool)\n\n    q = deque()\n\n    def push(z, y, x):\n        if 0 <= z < D and 0 <= y < H and 0 <= x < W and bg[z, y, x] and not seen[z, y, x]:\n            seen[z, y, x] = True\n            q.append((z, y, x))\n\n    # Flood fill from boundary background voxels -> marks \"outside\"\n    for z in (0, D - 1):\n        for y in range(H):\n            for x in range(W):\n                push(z, y, x)\n    for z in range(D):\n        for y in (0, H - 1):\n            for x in range(W):\n                push(z, y, x)\n    for z in range(D):\n        for y in range(H):\n            for x in (0, W - 1):\n                push(z, y, x)\n\n    # 6-neighborhood\n    nbrs = [(1, 0, 0), (-1, 0, 0),\n            (0, 1, 0), (0, -1, 0),\n            (0, 0, 1), (0, 0, -1)]\n\n    while q:\n        z, y, x = q.popleft()\n        for dz, dy, dx in nbrs:\n            push(z + dz, y + dy, x + dx)\n\n    # Background not reached from boundary = enclosed cavities\n    inside = bg & (~seen)\n\n    # Count connected components in inside (6-neighborhood)\n    holes = 0\n    seen2 = np.zeros_like(inside, dtype=bool)\n    for z in range(D):\n        for y in range(H):\n            for x in range(W):\n                if inside[z, y, x] and not seen2[z, y, x]:\n                    holes += 1\n                    qq = deque([(z, y, x)])\n                    seen2[z, y, x] = True\n                    while qq:\n                        zz, yy, xx = qq.popleft()\n                        for dz, dy, dx in nbrs:\n                            nz, ny, nx = zz + dz, yy + dy, xx + dx\n                            if 0 <= nz < D and 0 <= ny < H and 0 <= nx < W:\n                                if inside[nz, ny, nx] and not seen2[nz, ny, nx]:\n                                    seen2[nz, ny, nx] = True\n                                    qq.append((nz, ny, nx))\n    return holes\n\n\n# ============================================================\n# TOPO METRICS INSTALL\n# ============================================================\n_TOPO_OK = False\n\ndef install_topometrics():\n    global _TOPO_OK\n    if _TOPO_OK:\n        return\n    try:\n        import topometrics.leaderboard\n        _TOPO_OK = True\n        return\n    except:\n        pass\n\n    resources = \"/kaggle/input/vesuvius-metric-resources\"\n    workdir   = \"/kaggle/working/topological-metrics-kaggle\"\n\n    subprocess.run(\n        f\"cd {resources} && uv pip install --no-index --find-links=wheels \"\n        f\"-r topological-metrics-kaggle/requirements.txt\",\n        shell=True, check=True,\n    )\n\n    subprocess.run(\n        f\"cp -r {resources}/topological-metrics-kaggle /kaggle/working/\",\n        shell=True, check=True\n    )\n\n    subprocess.run(\n        f\"cd {workdir} && chmod +x scripts/setup_submodules.sh scripts/build_betti.sh && make build-betti\",\n        shell=True, check=True\n    )\n\n    subprocess.run(\n        f\"cd {workdir} && uv pip install -e . --no-deps --no-index --no-build-isolation\",\n        shell=True, check=True\n    )\n\n    sys.path.append(f\"{workdir}/src\")\n    importlib.invalidate_caches()\n\n    import topometrics.leaderboard\n    _TOPO_OK = True\n\n\n# ============================================================\n# OFFICIAL SCORING CALL\n# ============================================================\ndef score_official(gt_u8, pr_u8):\n    \"\"\"\n    gt_u8: (D,H,W) uint8 with ignore label=2 preserved\n    pr_u8: (D,H,W) uint8 {0,1}\n    \"\"\"\n    install_topometrics()\n    import topometrics.leaderboard as lb\n\n    r = lb.compute_leaderboard_score(\n        predictions=pr_u8.astype(np.uint8),\n        labels=gt_u8.astype(np.uint8),\n        dims=(0, 1, 2),\n        spacing=(1.0, 1.0, 1.0),\n        surface_tolerance=SURFACE_TOLERANCE,\n        voi_connectivity=26,\n        voi_transform=\"one_over_one_plus\",\n        voi_alpha=0.3,\n        combine_weights=(0.3, 0.35, 0.35),\n        ignore_label=2,\n    )\n\n    # VOI attribute differs in some versions\n    voi_score = r.voi.voi_score if hasattr(r.voi, \"voi_score\") else r.voi.oi_score\n\n    return dict(\n        score=float(r.score),\n        topo=float(r.topo.toposcore),\n        surface_dice=float(r.surface_dice),\n        voi=float(voi_score),\n    )\n\n\n# ============================================================\n# UNZIP SUBMISSION\n# ============================================================\nif EXTRACT_DIR.exists():\n    shutil.rmtree(EXTRACT_DIR)\nEXTRACT_DIR.mkdir(parents=True, exist_ok=True)\n\nwith zipfile.ZipFile(SUBMISSION_ZIP, \"r\") as z:\n    z.extractall(EXTRACT_DIR)\n\nprint(f\"✓ Extracted {SUBMISSION_ZIP} → {EXTRACT_DIR}\")\n\n\n# ============================================================\n# SCORE EACH ID\n# ============================================================\nrows = []\n\nfor sid in EVAL_IDS:\n    pred_path = EXTRACT_DIR / f\"{sid}.tif\"\n    gt_path   = LABEL_DIR / f\"{sid}.tif\"\n\n    if not pred_path.exists():\n        print(f\"❌ Missing prediction for {sid}\")\n        continue\n    if not gt_path.exists():\n        print(f\"❌ Missing GT for {sid}\")\n        continue\n\n    pred = load_volume_tif(pred_path)\n    gt   = load_volume_tif(gt_path)\n\n    # Prediction is binary {0,1}\n    pr_u8 = (pred > 0).astype(np.uint8)\n\n    # GT must keep ignore label=2\n    # Map: 2 → 2, (0 → 0), (nonzero but not 2 → 1)\n    gt_raw = gt.astype(np.uint8)\n    gt_u8 = np.where(gt_raw == 2, 2, (gt_raw > 0).astype(np.uint8))\n\n    # Fix possible (H,W,D) → (D,H,W)\n    gt_u8 = maybe_transpose_gt(gt_u8, pr_u8.shape)\n\n    if gt_u8.shape != pr_u8.shape:\n        raise ValueError(f\"Shape mismatch for {sid}: pred={pr_u8.shape}, gt={gt_u8.shape}\")\n\n    # Diagnostic: predicted enclosed voids (\"holes\")\n    pred_holes = count_enclosed_voids(pr_u8)\n\n    print(f\"[{sid}] pred uniques={np.unique(pr_u8)}  gt uniques={np.unique(gt_u8)}  pred_holes={pred_holes}\")\n\n    m = score_official(gt_u8, pr_u8)\n    m[\"id\"] = sid\n    m[\"pred_holes\"] = int(pred_holes)\n    rows.append(m)\n\n    print(f\"✓ {sid}: score={m['score']:.4f} | topo={m['topo']:.4f} | sd={m['surface_dice']:.4f} | voi={m['voi']:.4f} | holes={m['pred_holes']}\")\n\n\n# ============================================================\n# SUMMARY TABLE\n# ============================================================\ndf = pd.DataFrame(rows)\n\nprint(\"\\n=== Per-volume ===\")\ncols = [\"id\", \"score\", \"topo\", \"surface_dice\", \"voi\", \"pred_holes\"]\nprint(df[cols].to_string(index=False))\n\nagg_cols = [\"score\", \"topo\", \"surface_dice\", \"voi\", \"pred_holes\"]\nmean = df[agg_cols].mean().add_prefix(\"mean_\")\nstd  = df[agg_cols].std(ddof=1).add_prefix(\"std_\")\n\n# CV (coefficient of variation) not very meaningful for pred_holes if mean is small,\n# but we compute anyway for consistency.\ncv   = (df[agg_cols].std(ddof=1) / df[agg_cols].mean().replace(0, np.nan)).add_prefix(\"cv_\")\n\nsummary = pd.concat([mean, std, cv], axis=0).to_frame().T\n\nprint(\"\\n=== Summary ===\")\nprint(summary.to_string(index=False))\n\n# Optional: save outputs\ndf.to_csv(\"cv_scores_with_holes.csv\", index=False)\nsummary.to_csv(\"cv_summary_with_holes.csv\", index=False)\n\nsummary\n`",
          "votes": 5
        }
      ]
    },
    {
      "id": 3404011,
      "postDate": "2026-02-09T18:10:00.410Z",
      "content": "<p>I have extremely mixed feelings about this. On one hand, I am happy to solve the real problem better. But on the other hand, we have already spent months working on the current solution and we couldn’t possibly retest all of the ideas on cross validation that we once had in two weeks. I am certain that my wife will not be thrilled about me expecting two more weeks on this competition so any advice people may have on that would be fabulous. 😂😂 I just hope that we remain a strong competitor and that we can help give the solution you guys need! </p>",
      "rawMarkdown": "I have extremely mixed feelings about this. On one hand, I am happy to solve the real problem better. But on the other hand, we have already spent months working on the current solution and we couldn’t possibly retest all of the ideas on cross validation that we once had in two weeks. I am certain that my wife will not be thrilled about me expecting two more weeks on this competition so any advice people may have on that would be fabulous. 😂😂 I just hope that we remain a strong competitor and that we can help give the solution you guys need! ",
      "votes": 6,
      "replies": [
        {
          "id": 3404206,
          "postDate": "2026-02-10T03:35:19.667Z",
          "content": "<p>You need to do something on Feb 14 to continue competing on Kaggle.</p>\n<p>Lesson learned: GPU only is not all you need.</p>",
          "rawMarkdown": "You need to do something on Feb 14 to continue competing on Kaggle.\n\nLesson learned: GPU only is not all you need.",
          "votes": 7
        },
        {
          "id": 3404220,
          "postDate": "2026-02-10T04:15:56.307Z",
          "content": "<p>I planned to have a rest on the holiday, and back to focus on the Deep Past Challenge 😅 Now I decide to skip that one, and wait for a new challenge.</p>",
          "rawMarkdown": "I planned to have a rest on the holiday, and back to focus on the Deep Past Challenge 😅 Now I decide to skip that one, and wait for a new challenge.",
          "votes": 7
        }
      ]
    },
    {
      "id": 3404237,
      "postDate": "2026-02-10T04:45:59.777Z",
      "content": "<p>I had left working on this challenge for half a month because of some other hackathons, and now some internship, I assumed it's ending soon so i'll be getting a silver atleast. Turns out team had other plans, and now it overlaps with my exams also 😭😭.</p>\n<p>ps--&gt; I still think, if only the test set is changed, then nothing fundamental has changed much, only if you were trying to iteratively improve leaderboard, then It becomes problematic for you</p>",
      "rawMarkdown": "I had left working on this challenge for half a month because of some other hackathons, and now some internship, I assumed it's ending soon so i'll be getting a silver atleast. Turns out team had other plans, and now it overlaps with my exams also 😭😭.\n\nps--> I still think, if only the test set is changed, then nothing fundamental has changed much, only if you were trying to iteratively improve leaderboard, then It becomes problematic for you",
      "votes": 5,
      "replies": [
        {
          "id": 3404277,
          "postDate": "2026-02-10T07:10:18.190Z",
          "content": "<p>I too have a lot of other commitments starting earlier this month, which was my main motivation to team up.\nLot of my plans were centered around the comp ending mid Feb. This decision is very very frustrating.</p>",
          "rawMarkdown": "I too have a lot of other commitments starting earlier this month, which was my main motivation to team up.\nLot of my plans were centered around the comp ending mid Feb. This decision is very very frustrating.",
          "votes": 3
        }
      ]
    },
    {
      "id": 3404181,
      "postDate": "2026-02-10T02:08:44.930Z",
      "content": "<p><a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> I also highly suggest while fixing the data and rescoring to also patch one of the longest known issues in this competition - the faulty logic of the validation mask cutting the predictions and spawning multiple components/holes if the mask intersects our predictions discussed <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/653482\" target=\"_blank\">here</a>:</p>\n<p>Right now our prediction is zeroed out with the unlabelled mask.</p>\n<pre><code> # Apply ignore: neutralize by setting both PR and GT to background at ignored voxels\n # (Done on float/int volumes before any binarization or topology)\n\n    pr_eval = np.where(ign, 0, pr_raw)\n    gt_eval = np.where(ign, 0, gt_raw)\n</code></pre>\n<p>Instead of zeroing the models predictions and label we should switch to a more elegant solution. Instead of tampering with the prediction and causing unwanted and unregulated side effects (we don't know how much components/holes we will spawn if the prediction intersects the mask) we simply <strong>ignore the predicted components if their birth coordinates lie within the ignore mask</strong>. </p>\n<p>We simply do two steps:</p>\n<ol>\n<li>We calculate the topology for the <strong>full prediction volume</strong>.</li>\n<li>Discard the <strong>unmatched components from the list</strong> if they lie within the <strong>ignore region</strong>. (there could be NO matched components in the ignore region as gt contains nothing there, so we just filter the unmatched list)</li>\n<li>Our topology score now contains matched, predicted, ground truth counts <strong>only from the valid regions</strong>. And we didn't spawn any extra holes/components in doing so.</li>\n</ol>\n<p>This is the original function. Right now we return the unmatched birth coordinates without any filtration.</p>\n<pre><code>    # @classmethod\n    # def _counts_for_dim(cls, result: Any, k: int) -&gt; Tuple[int, int, int]:\n    #     m_k = cls._count_points(cls._safe_index(getattr(result, \"input1_matched_birth_coordinates\", []), k))\n    #     p_un_k = cls._count_points(cls._safe_index(getattr(result, \"input1_unmatched_birth_coordinates\", []), k))\n    #     g_un_k = cls._count_points(cls._safe_index(getattr(result, \"input2_unmatched_birth_coordinates\", []), k))\n    #     p_k = m_k + p_un_k\n    #     g_k = m_k + g_un_k\n    #     return m_k, p_k, g_k\n</code></pre>\n<p>And this is the solution with few lines of changed code - basic if else checks. We calculate the topology for the whole prediction and we don't tamper it or zero it out causing a mess, instead we simply check if the [z,y,x] birth coordinates lie in the ignore region mask! If they do, then discard them - they should not contribute to the total metric, as they should simply be ignored.</p>\n<pre><code>@classmethod\n    def _counts_for_dim(\n        cls, \n        result: Any, \n        k: int, \n        ignore_mask: Optional[np.ndarray]\n    ) -&gt; Tuple[int, int, int]:\n\n        # 1. Fetch raw lists\n        # Note: These are lists of \"birth coordinates\"\n        matched = cls._safe_index(getattr(result, \"input1_matched_birth_coordinates\", []), k)\n        p_unmatched = cls._safe_index(getattr(result, \"input1_unmatched_birth_coordinates\", []), k)\n        g_unmatched = cls._safe_index(getattr(result, \"input2_unmatched_birth_coordinates\", []), k)\n\n        # 2.Filter Helper\n        def count_valid(coords, mask):\n            \"\"\"Returns count of coords NOT in the ignore mask.\"\"\"\n            if coords is None: return 0\n            n_total = cls._count_points(coords)\n            if mask is None or n_total == 0:\n                return n_total\n\n            arr = np.asarray(coords)\n            if arr.ndim != 2 or arr.shape[1] != 3:\n                return n_total \n\n            valid_count = 0\n            for i in range(arr.shape[0]):\n                c0, c1, c2 = int(arr[i, 0]), int(arr[i, 1]), int(arr[i, 2])\n\n                if (0 &lt;= c0 &lt; mask.shape[0] and \n                    0 &lt;= c1 &lt; mask.shape[1] and \n                    0 &lt;= c2 &lt; mask.shape[2]):\n\n                    # If mask is True, it is IGNORED -&gt; Do not count\n                    if not mask[c0, c1, c2]:\n                        valid_count += 1\n                else:\n                    valid_count += 1\n            return valid_count\n\n        # 3. Calculate Counts\n        m_k_count = cls._count_points(matched)\n        g_un_k_count = count_valid(g_unmatched, ignore_mask) \n        p_un_k_count = count_valid(p_unmatched, ignore_mask)\n\n        p_k = m_k_count + p_un_k_count\n        g_k = m_k_count + g_un_k_count\n\n        return m_k_count, p_k, g_k\n</code></pre>\n<p>We just pass the ignore mask down to this function and use it as is.(Note: the mask should be sliced as well, because we are using (2,2,2) splits for the pred/gt)</p>\n<p>In hindsight, the solution to this was simple, but maybe we dimissed the possibility that something can be done to fix it.</p>\n<p>If there are no edge-cases or bugs in the fix (I encourage competitors to think on it as well, so we can find mistakes if there are any), I think this will also solve a long known issue that surfaced over two months ago, making the metric waterproof - and taking chance and lucky (unlucky) mask hitting out of the equation.</p>",
      "rawMarkdown": "@giorgioangelotti I also highly suggest while fixing the data and rescoring to also patch one of the longest known issues in this competition - the faulty logic of the validation mask cutting the predictions and spawning multiple components/holes if the mask intersects our predictions discussed [here](https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/653482):\n\nRight now our prediction is zeroed out with the unlabelled mask.\n```   \n # Apply ignore: neutralize by setting both PR and GT to background at ignored voxels\n # (Done on float/int volumes before any binarization or topology)\n\n    pr_eval = np.where(ign, 0, pr_raw)\n    gt_eval = np.where(ign, 0, gt_raw)\n```\n\nInstead of zeroing the models predictions and label we should switch to a more elegant solution. Instead of tampering with the prediction and causing unwanted and unregulated side effects (we don't know how much components/holes we will spawn if the prediction intersects the mask) we simply **ignore the predicted components if their birth coordinates lie within the ignore mask**. \n\nWe simply do two steps:\n 1. We calculate the topology for the **full prediction volume**.\n 2. Discard the **unmatched components from the list** if they lie within the **ignore region**. (there could be NO matched components in the ignore region as gt contains nothing there, so we just filter the unmatched list)\n 3. Our topology score now contains matched, predicted, ground truth counts **only from the valid regions**. And we didn't spawn any extra holes/components in doing so.\n\nThis is the original function. Right now we return the unmatched birth coordinates without any filtration.\n```\n    # @classmethod\n    # def _counts_for_dim(cls, result: Any, k: int) -> Tuple[int, int, int]:\n    #     m_k = cls._count_points(cls._safe_index(getattr(result, \"input1_matched_birth_coordinates\", []), k))\n    #     p_un_k = cls._count_points(cls._safe_index(getattr(result, \"input1_unmatched_birth_coordinates\", []), k))\n    #     g_un_k = cls._count_points(cls._safe_index(getattr(result, \"input2_unmatched_birth_coordinates\", []), k))\n    #     p_k = m_k + p_un_k\n    #     g_k = m_k + g_un_k\n    #     return m_k, p_k, g_k\n```\n\nAnd this is the solution with few lines of changed code - basic if else checks. We calculate the topology for the whole prediction and we don't tamper it or zero it out causing a mess, instead we simply check if the [z,y,x] birth coordinates lie in the ignore region mask! If they do, then discard them - they should not contribute to the total metric, as they should simply be ignored.\n\n```\n@classmethod\n    def _counts_for_dim(\n        cls, \n        result: Any, \n        k: int, \n        ignore_mask: Optional[np.ndarray]\n    ) -> Tuple[int, int, int]:\n        \n        # 1. Fetch raw lists\n        # Note: These are lists of \"birth coordinates\"\n        matched = cls._safe_index(getattr(result, \"input1_matched_birth_coordinates\", []), k)\n        p_unmatched = cls._safe_index(getattr(result, \"input1_unmatched_birth_coordinates\", []), k)\n        g_unmatched = cls._safe_index(getattr(result, \"input2_unmatched_birth_coordinates\", []), k)\n\n        # 2.Filter Helper\n        def count_valid(coords, mask):\n            \"\"\"Returns count of coords NOT in the ignore mask.\"\"\"\n            if coords is None: return 0\n            n_total = cls._count_points(coords)\n            if mask is None or n_total == 0:\n                return n_total\n\n            arr = np.asarray(coords)\n            if arr.ndim != 2 or arr.shape[1] != 3:\n                return n_total \n\n            valid_count = 0\n            for i in range(arr.shape[0]):\n                c0, c1, c2 = int(arr[i, 0]), int(arr[i, 1]), int(arr[i, 2])\n                \n                if (0 <= c0 < mask.shape[0] and \n                    0 <= c1 < mask.shape[1] and \n                    0 <= c2 < mask.shape[2]):\n                    \n                    # If mask is True, it is IGNORED -> Do not count\n                    if not mask[c0, c1, c2]:\n                        valid_count += 1\n                else:\n                    valid_count += 1\n            return valid_count\n\n        # 3. Calculate Counts\n        m_k_count = cls._count_points(matched)\n        g_un_k_count = count_valid(g_unmatched, ignore_mask) \n        p_un_k_count = count_valid(p_unmatched, ignore_mask)\n\n        p_k = m_k_count + p_un_k_count\n        g_k = m_k_count + g_un_k_count\n        \n        return m_k_count, p_k, g_k\n```\n\nWe just pass the ignore mask down to this function and use it as is.(Note: the mask should be sliced as well, because we are using (2,2,2) splits for the pred/gt)\n\nIn hindsight, the solution to this was simple, but maybe we dimissed the possibility that something can be done to fix it.\n\nIf there are no edge-cases or bugs in the fix (I encourage competitors to think on it as well, so we can find mistakes if there are any), I think this will also solve a long known issue that surfaced over two months ago, making the metric waterproof - and taking chance and lucky (unlucky) mask hitting out of the equation.",
      "votes": 4,
      "replies": [
        {
          "id": 3404184,
          "postDate": "2026-02-10T02:17:15.290Z",
          "content": "<p>totally agree.</p>",
          "rawMarkdown": "totally agree.",
          "votes": 1
        },
        {
          "id": 3404262,
          "postDate": "2026-02-10T06:07:27.647Z",
          "content": "<p>I havent checked how the birth coordinates are calculated, but your suggestion would remove a correct segmentation if it were to \"birth\" in label 2. Which might happen alot because of the 3 voxel labeled 2 at each edge/face of the volume?</p>",
          "rawMarkdown": "I havent checked how the birth coordinates are calculated, but your suggestion would remove a correct segmentation if it were to \"birth\" in label 2. Which might happen alot because of the 3 voxel labeled 2 at each edge/face of the volume?"
        }
      ]
    },
    {
      "id": 3405061,
      "postDate": "2026-02-12T01:17:37.953Z",
      "content": "<p><a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> Hi, Could you please clarify what exactly was fixed in the test set and how the evaluation metric was updated? For example, were holes in the labels filled algorithmically?\nThis information is important for fairness, especially since the training set remains unchanged while the test set was modified. Without knowing the specific fixes, it’s difficult to properly adapt and validate our models locally.</p>\n<p>Thank you.</p>",
      "rawMarkdown": "@giorgioangelotti @ryanholbrook Hi, Could you please clarify what exactly was fixed in the test set and how the evaluation metric was updated? For example, were holes in the labels filled algorithmically?\nThis information is important for fairness, especially since the training set remains unchanged while the test set was modified. Without knowing the specific fixes, it’s difficult to properly adapt and validate our models locally.\n\nThank you.",
      "votes": 3
    },
    {
      "id": 3404272,
      "postDate": "2026-02-10T06:45:42.023Z",
      "content": "<p>haha, totally not cool with the fact that the new timeline now extends into the lunar new yr holiday. </p>",
      "rawMarkdown": "haha, totally not cool with the fact that the new timeline now extends into the lunar new yr holiday. ",
      "votes": 4
    },
    {
      "id": 3404763,
      "postDate": "2026-02-11T07:13:54.773Z",
      "content": "<p>Seeing everyone’s scores getting so high, I just realized I’m the one swimming naked. 😀</p>",
      "rawMarkdown": "Seeing everyone’s scores getting so high, I just realized I’m the one swimming naked. 😀",
      "votes": 4,
      "replies": [
        {
          "id": 3404859,
          "postDate": "2026-02-11T11:33:59.313Z",
          "content": "<p>What I'm 100% about is that, the final rankings on private will be far more random than this overhaul.</p>",
          "rawMarkdown": "What I'm 100% about is that, the final rankings on private will be far more random than this overhaul.",
          "votes": 1
        }
      ]
    },
    {
      "id": 3404468,
      "postDate": "2026-02-10T15:00:46.253Z",
      "content": "<p>Has anyone gotten their updated scores after the re-evaluation? It looks like my recent submissions are still being assessed using the previous grading criteria.</p>",
      "rawMarkdown": "Has anyone gotten their updated scores after the re-evaluation? It looks like my recent submissions are still being assessed using the previous grading criteria.\n",
      "votes": 1
    },
    {
      "id": 3404103,
      "postDate": "2026-02-09T21:55:33.293Z",
      "content": "<p>You do not change the train set but if one would independently retrain on fixed labels, that would give a big advantage on the test set.</p>",
      "rawMarkdown": "You do not change the train set but if one would independently retrain on fixed labels, that would give a big advantage on the test set.",
      "votes": 1
    },
    {
      "id": 3404261,
      "postDate": "2026-02-10T05:59:15.483Z",
      "content": "<p>I have a small question - do both Betti number 1 and Betti number 2 get corrected, or only Betti number 2?</p>",
      "rawMarkdown": "I have a small question - do both Betti number 1 and Betti number 2 get corrected, or only Betti number 2?",
      "votes": 2
    },
    {
      "id": 3404222,
      "postDate": "2026-02-10T04:17:25.973Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> <a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> , What will happen to the submissions that was submitted before the update, but complete after the update. Is the score old or new?</p>",
      "rawMarkdown": "Hi @ryanholbrook @giorgioangelotti , What will happen to the submissions that was submitted before the update, but complete after the update. Is the score old or new?",
      "votes": 2,
      "replies": [
        {
          "id": 3404443,
          "postDate": "2026-02-10T14:03:21.937Z",
          "content": "<p>They would be scored with the updated labels only. But since we're rescoring all submissions with the updated labels, there won't be any difference in the end.</p>",
          "rawMarkdown": "They would be scored with the updated labels only. But since we're rescoring all submissions with the updated labels, there won't be any difference in the end.",
          "votes": 1,
          "replies": [
            {
              "id": 3404450,
              "postDate": "2026-02-10T14:22:34.020Z",
              "content": "<p>Ahh, I am just now understanding this. If we submit now we get new score, otherwise we are waiting on all submissions to be updated? </p>",
              "rawMarkdown": "Ahh, I am just now understanding this. If we submit now we get new score, otherwise we are waiting on all submissions to be updated? "
            },
            {
              "id": 3404452,
              "postDate": "2026-02-10T14:26:05.177Z",
              "content": "<p>However, the newly submitted version from our team doesn't seem to have any difference in score.</p>",
              "rawMarkdown": "However, the newly submitted version from our team doesn't seem to have any difference in score."
            }
          ]
        }
      ]
    },
    {
      "id": 3403993,
      "postDate": "2026-02-09T17:19:37.503Z",
      "content": "<p>Hi Giorgio — thanks again for handling the rescore, and for all the work running the competition.</p>\n<p>Quick question for clarity: for this rescore, will you re-evaluate all submissions, or only a subset (e.g., each team’s best / final-selected submissions)? If it’s a subset, could you share how submissions are chosen?</p>\n<p>Thanks a lot!</p>",
      "rawMarkdown": "Hi Giorgio — thanks again for handling the rescore, and for all the work running the competition.\n\nQuick question for clarity: for this rescore, will you re-evaluate all submissions, or only a subset (e.g., each team’s best / final-selected submissions)? If it’s a subset, could you share how submissions are chosen?\n\nThanks a lot!",
      "votes": 1,
      "replies": [
        {
          "id": 3404072,
          "postDate": "2026-02-09T20:49:40.150Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/tonylica\" target=\"_blank\">@tonylica</a>, We will rescore all submissions (that is, run the metric over the saved predictions, not rerun the entire model).</p>",
          "rawMarkdown": "Hi @tonylica, We will rescore all submissions (that is, run the metric over the saved predictions, not rerun the entire model).",
          "votes": 5,
          "replies": [
            {
              "id": 3404074,
              "postDate": "2026-02-09T21:00:37.600Z",
              "content": "<p>Hi Ryan — thank you for the clarification. That makes perfect sense. Appreciate the quick response and all the work on this.</p>",
              "rawMarkdown": "Hi Ryan — thank you for the clarification. That makes perfect sense. Appreciate the quick response and all the work on this.",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 3405130,
      "postDate": "2026-02-12T07:19:49.410Z",
      "content": "<p>Thank you for the update <a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> </p>",
      "rawMarkdown": "Thank you for the update @giorgioangelotti \n"
    },
    {
      "id": 3404354,
      "postDate": "2026-02-10T11:01:03.057Z",
      "content": "<p>The worst orgs ever.</p>",
      "rawMarkdown": "The worst orgs ever.",
      "votes": -16
    },
    {
      "id": 3407003,
      "postDate": "2026-02-17T11:02:37.020Z",
      "content": "<p>its just some ever-some </p>",
      "rawMarkdown": "its just some ever-some ",
      "votes": -2
    },
    {
      "id": 3406677,
      "postDate": "2026-02-16T10:23:16.917Z",
      "content": "<p>\"Great update, thanks for keeping the competition fair for everyone. Let's keep pushing!\"</p>",
      "rawMarkdown": "\"Great update, thanks for keeping the competition fair for everyone. Let's keep pushing!\"",
      "votes": -2
    },
    {
      "id": 3406035,
      "postDate": "2026-02-14T12:01:43.687Z",
      "content": "<p>im a beginner how can i learn fast and efficient</p>",
      "rawMarkdown": "im a beginner how can i learn fast and efficient",
      "votes": -3,
      "replies": [
        {
          "id": 3406543,
          "postDate": "2026-02-16T00:49:35.397Z",
          "content": "<p>There's nothing in this world you can learn fast and become expert overnight, over a week or over a month. Amount of hours directly translate to expertise.</p>",
          "rawMarkdown": "There's nothing in this world you can learn fast and become expert overnight, over a week or over a month. Amount of hours directly translate to expertise.",
          "votes": 4
        },
        {
          "id": 3406612,
          "postDate": "2026-02-16T07:27:53.550Z",
          "content": "<p>Enter kaggle competitions. Start with playground competitions. Starting with a featured competition like vesuvius is extremely challenging.</p>",
          "rawMarkdown": "Enter kaggle competitions. Start with playground competitions. Starting with a featured competition like vesuvius is extremely challenging.",
          "votes": 4
        }
      ]
    },
    {
      "id": 3404812,
      "postDate": "2026-02-11T09:58:43.403Z",
      "content": "<p>Have the scores been updated?</p>",
      "rawMarkdown": "Have the scores been updated?\n",
      "votes": -1
    },
    {
      "id": 3408579,
      "postDate": "2026-02-20T22:13:10.250Z",
      "content": "<p>Thanks for the transparency and quick action on this.\nReally appreciate the coordination with Kaggle Support and the deadline extension that definitely helps teams recalibrate fairly.\nLooking forward to seeing how the leaderboard evolves after the rescore!</p>",
      "rawMarkdown": "Thanks for the transparency and quick action on this.\nReally appreciate the coordination with Kaggle Support and the deadline extension that definitely helps teams recalibrate fairly.\nLooking forward to seeing how the leaderboard evolves after the rescore!"
    },
    {
      "id": 3404170,
      "postDate": "2026-02-10T01:25:25.623Z",
      "content": "<p>My model performs better locally but worse on the leaderboard—it's frustrating. Thankfully, the Host ultimately opted to re-evaluate the scores.</p>",
      "rawMarkdown": "My model performs better locally but worse on the leaderboard—it's frustrating. Thankfully, the Host ultimately opted to re-evaluate the scores.",
      "replies": [
        {
          "id": 3404391,
          "postDate": "2026-02-10T12:13:52.790Z",
          "content": "<p>That means you did it right.</p>",
          "rawMarkdown": "That means you did it right.",
          "replies": [
            {
              "id": 3404402,
              "postDate": "2026-02-10T12:33:12.723Z",
              "content": "<p>The new plan I've been working on these past few days is not working, leaving me feeling a bit overwhelmed.</p>",
              "rawMarkdown": "The new plan I've been working on these past few days is not working, leaving me feeling a bit overwhelmed."
            }
          ]
        }
      ]
    },
    {
      "id": 3403995,
      "postDate": "2026-02-09T17:26:25.430Z",
      "content": "<p>so all submission will be rescored again which we submitted </p>",
      "rawMarkdown": "so all submission will be rescored again which we submitted "
    },
    {
      "id": 3404412,
      "postDate": "2026-02-10T12:45:18.110Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    },
    {
      "id": 3404223,
      "postDate": "2026-02-10T04:17:56.827Z",
      "rawMarkdown": "",
      "votes": -3,
      "isDeleted": true
    },
    {
      "id": 3404155,
      "postDate": "2026-02-10T00:06:04.290Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3404154,
      "postDate": "2026-02-10T00:02:56.320Z",
      "rawMarkdown": "",
      "votes": -4,
      "isDeleted": true
    },
    {
      "id": 3403998,
      "postDate": "2026-02-09T17:36:13.310Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true,
      "replies": [
        {
          "id": 3404002,
          "postDate": "2026-02-09T17:42:40.723Z",
          "content": "<p>Thanks — that makes sense if that’s the case.</p>\n<p>So Kaggle caches the submission files for the full submission history. Now I understand why the AI Math competition (using the Eval API) only allows one submission.</p>",
          "rawMarkdown": "Thanks — that makes sense if that’s the case.\n\nSo Kaggle caches the submission files for the full submission history. Now I understand why the AI Math competition (using the Eval API) only allows one submission.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3404453,
      "author_name": "Cody_Null",
      "author_url": "",
      "post_date": "2026-02-10T14:29:35.837000",
      "content": "<p>I understand the frustration, believe me, as we were tied first place going into this change. However, I think maybe we should take it easier on the hosts as they have been participating a ton, answering questions on the weekends, and trying to help out. They really just want the best solution they can get as they are paying for it after all. I can't say I blame them. Obviously, it can be inconvenient for us and the timeframe is not ideal, but I think they are doing what they can.</p>",
      "votes": 29,
      "replies": [
        {
          "id": 3404469,
          "author_name": "Temple of Dionysus",
          "author_url": "",
          "post_date": "2026-02-10T15:01:40.700000",
          "content": "<p>I would like to echo this sentiment. After all, the end goal is reading the scrolls.</p>",
          "votes": 11,
          "replies": []
        },
        {
          "id": 3404471,
          "author_name": "Tom",
          "author_url": "",
          "post_date": "2026-02-10T15:05:55.740000",
          "content": "<p>I think people should show respect to the host. The host team wants the best possible outcome, which is why they had to make this difficult decision. I understand that many people are upset about it, especially Chinese Kagglers, for whom the next few weeks will be very inconvenient.</p>\n<p>That said, the host has done almost everything they could and has tried their best to fix the issues and answer people’s questions. They have also taken responsibility for the situation. I hope that in the last two weeks, we can focus on finishing the competition instead of continuing to complain about what has already happened. </p>",
          "votes": 15,
          "replies": []
        }
      ]
    },
    {
      "id": 3404284,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2026-02-10T07:32:24.553000",
      "content": "<p>Not surprisingly, this is going to demotivate those who worked hard till now. Some comments show it already.</p>\n<p>What I don't get is why the training data isn't fixed either? Now you added an extra complexity with a distribution shift between training and testing data.</p>",
      "votes": 25,
      "replies": [
        {
          "id": 3404305,
          "author_name": "",
          "author_url": "",
          "post_date": "2026-02-10T08:55:41.413000",
          "content": "",
          "votes": -1,
          "replies": [
            {
              "id": 3404383,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2026-02-10T11:48:36.443000",
              "content": "<p>Yes, he said that, and I don't understand why they didn't update train as well. Probably because it is time consuming. Which makes our life more difficult if we want to \"fix\" train as well.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 3404324,
          "author_name": "cm391",
          "author_url": "",
          "post_date": "2026-02-10T10:08:57.400000",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> are you willing to share the \"other\" issue you discovered with training data? i do think its unfair they don't clearly state how the test labels have been updated…</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3404355,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2026-02-10T11:05:03.307000",
              "content": "<p>will do if they confirm the deadline extension.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3404368,
              "author_name": "Muhammad Ibrahim",
              "author_url": "",
              "post_date": "2026-02-10T11:18:35.383000",
              "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> The deadline has been extended on the overview page.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3404382,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2026-02-10T11:46:29.497000",
              "content": "<p>I know, but who knows what they'll decide given the reactions of people here. I prefer to wait for the dust to settle. Will share within 24 hours if there are no changes.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3404390,
              "author_name": "cm391",
              "author_url": "",
              "post_date": "2026-02-10T12:10:12.467000",
              "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> i think they will wait to see how much shake up from results of rescore…</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3404574,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2026-02-10T18:26:20.660000",
              "content": "<p>Rescore being in progress, I thinkt he change is committed to. I'll share tomorrow.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3405212,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2026-02-12T12:07:38.103000",
              "content": "<p>The issue is in my code.</p>\n<p>I have been training with corrupted labels for weeks now…</p>\n<p>Sorry for the false alarm.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3405218,
              "author_name": "Tom",
              "author_url": "",
              "post_date": "2026-02-12T12:16:20.677000",
              "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>  I’m sure many people are waiting for you to share it. They must be very disappointed right now.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3405231,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2026-02-12T13:05:24.297000",
              "content": "<p>I have nothing to share mas explained above.</p>\n<p>I was resizing images and labels  and this introduced spurious positive voxels on the mask boundary. Using the original images and labels removed the issue.</p>\n<p>I found it while working on sharing the issue here. I found the issues, it was in my code (resizing labels).</p>\n<p>I have to redo everything I did. </p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3404005,
      "author_name": "Sergio Alvarez",
      "author_url": "",
      "post_date": "2026-02-09T17:47:31.393000",
      "content": "<p>Thanks a lot for the effort in handling this update and for keeping the competition aligned with its ultimate goal, helping to read the scrolls, rather than competing for the sake of competition. We really appreciate the transparency and the work from your team.</p>\n<p>Could you please clarify what the fix is exactly?</p>\n<p>More specifically, should participants interpret the rescore as applying the approach mentioned by <a href=\"https://www.kaggle.com/dankrstev\" target=\"_blank\">@dankrstev</a>? That is, keeping the data unchanged but adjusting the metric so it is computed as if no holes exist? Or was any additional processing applied, such as a binary closing or another morphological operation??\n<a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/671160#3403175\" target=\"_blank\">https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/671160#3403175</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2F3fa5a1afd5feab38662b5c03600d68a7%2Faa.png?generation=1770659017538322&amp;alt=media\" alt=\"\"></p>",
      "votes": 18,
      "replies": [
        {
          "id": 3404162,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2026-02-10T00:36:27.527000",
          "content": "<p>I need we need to have the patch released in the earliest time for is to react.<br>\ne.g.<br>\n1) is the patch in metric computation?<br>\nthen relese the new code  </p>\n<p>2) is the patch in test data?<br>\nthen release new train labels that is patched similarly as the test or<br>\ncode to transform the data.  (eg filling micro holes)</p>",
          "votes": 13,
          "replies": []
        }
      ]
    },
    {
      "id": 3404334,
      "author_name": "tereka",
      "author_url": "",
      "post_date": "2026-02-10T10:26:48",
      "content": "<p>Could you please disclose fix points?\nWe can't fix local pipeline…</p>\n<p>if this points are not important, we don't need extend long deadline… </p>",
      "votes": 20,
      "replies": [
        {
          "id": 3404337,
          "author_name": "ForcewithMe",
          "author_url": "",
          "post_date": "2026-02-10T10:32:45.347000",
          "content": "<blockquote>\n  <p>if this points are not important, we don't need extend long deadline…</p>\n</blockquote>\n<p>Exactly, totally agree! </p>",
          "votes": 6,
          "replies": []
        }
      ]
    },
    {
      "id": 3404247,
      "author_name": "Dieter",
      "author_url": "",
      "post_date": "2026-02-10T05:25:53.033000",
      "content": "<p>Haha, I am so angry I can't make any constructive comment. Not sure I will continue this comp. </p>",
      "votes": 21,
      "replies": [
        {
          "id": 3404251,
          "author_name": "Manas Choudhary",
          "author_url": "",
          "post_date": "2026-02-10T05:37:58.043000",
          "content": "<p>Looks like I quit at the right time 😉, but I think one week extension was the most that should have been given, simply because only the leaderboard score has changed and not the data given to participants.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 3404260,
          "author_name": "ForcewithMe",
          "author_url": "",
          "post_date": "2026-02-10T05:54:29.237000",
          "content": "<p>At the beginning, I chose to join this competition because it was scheduled to end just one day before the Chinese New Year, which was perfect timing for me and most Chinese people. Now, without any prior notice (although updates were expected, no one ever indicated it would be extended), the competition has been extended by two weeks, completely overlapping with the Spring Festival. I feel both angry and powerless🤐.</p>",
          "votes": 12,
          "replies": [
            {
              "id": 3404266,
              "author_name": "Dieter",
              "author_url": "",
              "post_date": "2026-02-10T06:27:07.850000",
              "content": "<p>Yes, people plan vacations etc around competition end. Same for me. What makes me angry is the proportionality. Impacting 1600 participants time plans, just to weight topo score a bit higher. </p>",
              "votes": 11,
              "replies": []
            },
            {
              "id": 3404856,
              "author_name": "Roc",
              "author_url": "",
              "post_date": "2026-02-11T11:19:17.417000",
              "content": "<p>Yes, totally agree! Extending it by two weeks would completely mess up my vacation.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3404281,
      "author_name": "lingyundev",
      "author_url": "",
      "post_date": "2026-02-10T07:23:37.637000",
      "content": "<p>Honestly, I’m quite disappointed about the deadline extension. One big reason I joined this competition was because it was supposed to end before the Lunar New Year. Many of us planned our time around that. Extending it into the holiday really affects those plans and makes it very hard to balance family time and the competition.\nI understand that fixing the data is important, but changing the timeline this late — especially into a major holiday — feels really tough for a lot of Chinese participants.</p>",
      "votes": 16,
      "replies": []
    },
    {
      "id": 3404373,
      "author_name": "Yone",
      "author_url": "",
      "post_date": "2026-02-10T11:24:45.403000",
      "content": "<p>I believe that revising the test set is not directly related to postponing the competition, since the training set will not change and most teams have already  completed their models. I hope the organizers can take participants’ feelings into consideration—everyone has put a great deal of effort into this competition, whether in terms of time or the cost of renting GPUs. A short extension would be understandable, but a two-week delay is really too long and quite stressful for us.</p>",
      "votes": 13,
      "replies": [
        {
          "id": 3404404,
          "author_name": "Wang Zhiyao (王致尧)",
          "author_url": "",
          "post_date": "2026-02-10T12:34:11.480000",
          "content": "<p>Especially for friends in the East who need to celebrate the Chinese/Lunar New Year</p>",
          "votes": 6,
          "replies": []
        }
      ]
    },
    {
      "id": 3404026,
      "author_name": "DECEM",
      "author_url": "",
      "post_date": "2026-02-09T18:36:33.917000",
      "content": "<p>Let’s be real: this extension is a double-edged sword for Chinese participants. We are entering the Spring Festival—the most important time for family reunions in our culture. While the extra time is appreciated, it effectively means choosing between quality time with family and squeezing out that last bit of model performance. </p>\n<p>It’s going to be a tough grind during the holidays! </p>",
      "votes": 14,
      "replies": [
        {
          "id": 3404229,
          "author_name": "ForcewithMe",
          "author_url": "",
          "post_date": "2026-02-10T04:33:31.410000",
          "content": "<p>That’s so struggling. For someone working in the Internet Company, it may be the most relaxing time in 1 year. Now it may become the most exhausting time.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 3404270,
          "author_name": "hsiaosuan",
          "author_url": "",
          "post_date": "2026-02-10T06:38:04.353000",
          "content": "<p>抱抱 过年期间还要打比赛却是压力太大了</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 3403961,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2026-02-09T16:32:53.523000",
      "content": "<p><a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> \nThanks. I’m curious, after fixing the test set and evaluation metrics, what score did your baseline nnU-Net achieve without any post-processing?</p>",
      "votes": 12,
      "replies": []
    },
    {
      "id": 3404439,
      "author_name": "tingyi",
      "author_url": "",
      "post_date": "2026-02-10T13:54:14",
      "content": "<p>Has anyone received their re-graded scores yet? I noticed that my new uploads still seem to use the old grading method.</p>",
      "votes": 9,
      "replies": [
        {
          "id": 3404458,
          "author_name": "Starry",
          "author_url": "",
          "post_date": "2026-02-10T14:42:37.893000",
          "content": "<p>My score also have not changed. And I face some strange issues: 1. one submission show \"kaggle error\", but after few minutes late it turn to success and get a score. 2. I save and run a notebook version, it should be end in about 6 minutes, but it has runned for more than 10h yet. And I can't stop it.</p>\n<p>It is so confusing.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3404461,
              "author_name": "GG Ayo (AyoGG)",
              "author_url": "",
              "post_date": "2026-02-10T14:47:09.653000",
              "content": "<p>Local testing shows that the CV without holes provides roughly a 1% boost, but current results are identical. Some players on the leaderboard appear to have received a buff.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3404466,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2026-02-10T14:55:33.887000",
              "content": "<ol>\n<li><p>Will hole filled train samples changed training results? Likely no. It is unlikely that your model model learned to predict holes in your old training.</p></li>\n<li><p>Will hole filled test samples affected results? Yes if you have very accurate prediction. No if your prediction is not accurate. </p></li>\n</ol>\n<p>Here is a test: try perturbation ( small and big shift etc) on hole and filled gt. Then measure the topo score (eg score(filled, shift filled)). My experiments showed that if your local lb is in 0.60 range, it is more likely to be affected.</p>",
              "votes": 6,
              "replies": []
            }
          ]
        },
        {
          "id": 3404477,
          "author_name": "ForcewithMe",
          "author_url": "",
          "post_date": "2026-02-10T15:27:21",
          "content": "<p>How can you know if they use the old scoring method or the new one?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3404478,
              "author_name": "GG Ayo (AyoGG)",
              "author_url": "",
              "post_date": "2026-02-10T15:28:15.573000",
              "content": "<p>Control group。\nHave you received your new scores? Looks like you've made significant progress.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3404481,
              "author_name": "ForcewithMe",
              "author_url": "",
              "post_date": "2026-02-10T15:33:18.437000",
              "content": "<p>I think it’s the score after updating.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3404482,
              "author_name": "tingyi",
              "author_url": "",
              "post_date": "2026-02-10T15:33:18.477000",
              "content": "<p>I'm actually not sure if it's the new or old method. I'm just judging whether they used the new scoring system based on the fact that the scores haven't changed.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        },
        {
          "id": 3404489,
          "author_name": "Francois Lemarchand",
          "author_url": "",
          "post_date": "2026-02-10T15:50:38.433000",
          "content": "<p>Same here. None of my old submissions seem to have been rescored and new ones are pretty much in the range of what I would expect with the old scoring process.\nConsidering that it takes ~3-4 hours to score the predictions and the thousands of submissions to process, I would not be surprised if it takes another few days before LB stabilises.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 3404521,
              "author_name": "TWEAK",
              "author_url": "",
              "post_date": "2026-02-10T17:03:22.010000",
              "content": "<p>Same here. Our recent submission on the leaderboard was sent during the rescore, but none of my other submissions have changed. I’m not sure if this is the new scoring yet; I was using a brand-new model, so there’s no correlation to my old submissions. Because of that, I can't tell if the 0.005 bump was from the new scoring or the model itself.</p>",
              "votes": 3,
              "replies": []
            }
          ]
        },
        {
          "id": 3404540,
          "author_name": "Ryan Holbrook",
          "author_url": "",
          "post_date": "2026-02-10T17:35:23.807000",
          "content": "<p>The rescore is in progress. We currently have about 8% of submissions rescored. Because this metric is fairly computationally intensive, it's likely going to be a couple days before everything is complete. The new scores are posted as soon as they're ready, so the leaderboard will show a mix of old and new until it's done. New submissions should be using the updated labels though.</p>",
          "votes": 11,
          "replies": [
            {
              "id": 3404543,
              "author_name": "",
              "author_url": "",
              "post_date": "2026-02-10T17:41:24.067000",
              "content": "",
              "votes": -2,
              "replies": []
            },
            {
              "id": 3404586,
              "author_name": "Temple of Dionysus",
              "author_url": "",
              "post_date": "2026-02-10T18:43:31.250000",
              "content": "<p>Are new submissions scored with the new metric during this transition time?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3404594,
              "author_name": "Ryan Holbrook",
              "author_url": "",
              "post_date": "2026-02-10T19:01:56.353000",
              "content": "<p>Yes. The new test labels are live on the site.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3403985,
      "author_name": "Starry",
      "author_url": "",
      "post_date": "2026-02-09T17:09:55.030000",
      "content": "<p>Spent tons of time during the past days and tired, now  it would take more time because of extension deadline even in Spring festival 😱. Tired and tired.. But anyway, this kind of update I think it is the best way for fairness of the competetion.</p>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 3404838,
      "author_name": "sqrt4kaido",
      "author_url": "",
      "post_date": "2026-02-11T10:50:24.883000",
      "content": "<p>Since only the test set was updated, I believe this has introduced an additional challenge of dealing with a potential domain shift. Could you please share a bit more detail on what was changed?</p>\n<p>For example:</p>\n<ul>\n<li>Applied some algorithmic post-processing to fill small holes in the labels.</li>\n<li>Removed highly porous samples from the test set, or otherwise filtered/replaced them.</li>\n</ul>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 3404318,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2026-02-10T09:46:37.773000",
      "content": "<p>I didn't understand the reason for the extension, since only the test data was updated.</p>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 3404406,
      "author_name": "Francois Lemarchand",
      "author_url": "",
      "post_date": "2026-02-10T12:37:18.337000",
      "content": "<p>I also think that the organisers are totally ignoring that <strong>humans</strong> are taking part in this competition. People have worked really hard, made sacrifices, and organised their personal life in such a way that they would be able to work extra hard until Friday. I was personally really looking forward to Friday.</p>\n<p>Moreover, changing the test set and not the training set is problematic. 2 weeks is not that much time to adapt our strategy, especially when the rescoring is taking time and the distribution shift between training and test sets means we now rely on LB even more. It does not leave much time to retrain and fine-tune models considering that the models take days to train with a good machine.</p>\n<p>While I love this project and I still have plenty of ideas to improve my solution, I am still unsure whether I will continue the competition as it's been an absolute chaos and new potential bugs have already been mentioned. Yes, the metric is \"a bit\" broken and it is clearly an issue from a data science point of view, but I think moving the goal posts at the last minute is not the solution.</p>",
      "votes": 9,
      "replies": []
    },
    {
      "id": 3404899,
      "author_name": "Ryan Holbrook",
      "author_url": "",
      "post_date": "2026-02-11T14:16:17.323000",
      "content": "<p>The rescore is now complete. Please let me know if you have any questions or concerns about the rescore.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 3404905,
          "author_name": "GG Ayo (AyoGG)",
          "author_url": "",
          "post_date": "2026-02-11T14:22:20.563000",
          "content": "<p>So what was fixed in the test set?</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 3404914,
          "author_name": "Cody_Null",
          "author_url": "",
          "post_date": "2026-02-11T15:00:36.260000",
          "content": "<p>From the comments in the discussion it sounds like the metric change was a patch to g_k = 0. I am currently trying to compute cv based off a monkey patch based on this but if we could be more concrete about exactly what changed so we can work on solving the problems with the newly alloted time that would be helpful. Below is an example of how I am doing it as well as some gpt comments (I dont have certainty this is correct)\n`</p>\n<h1>============================================================</h1>\n<h1>SCORE submission.zip ON TRAIN IDS (Official TopoMetric) + pred_holes diagnostic</h1>\n<h1>============================================================</h1>\n<p>import os, zipfile, subprocess, sys, importlib, shutil\nimport numpy as np\nimport pandas as pd\nfrom pathlib import Path\nfrom PIL import Image, ImageSequence\nfrom collections import deque</p>\n<h1>-----------------------------</h1>\n<h1>Config</h1>\n<h1>-----------------------------</h1>\n<p>SUBMISSION_ZIP = Path(\"submission.zip\")\nEXTRACT_DIR    = Path(\"./_sub_eval_extract\")\nLABEL_DIR      = Path(\"/kaggle/input/vesuvius-challenge-surface-detection/train_labels\")</p>\n<p>EVAL_IDS = [\n    \"2257172177\", \"663106834\", \"2555675774\",\n    \"1294570892\", \"2550619349\", \"3854101708\",\n    \"406944815\", \"2059872339\", \"360339268\",\n    \"1693721638\", \"508375143\", \"687559918\",\n    \"3834183540\", \"2689046967\", \"2454492741\",\n    \"370627915\", \"1837890067\", \"398977019\",\n    \"315189226\", \"239904888\", \"725727411\"\n]\nSURFACE_TOLERANCE = 2.0</p>\n<h1>============================================================</h1>\n<h1>HELPERS</h1>\n<h1>============================================================</h1>\n<p>def load_volume_tif(path: Path) -&gt; np.ndarray:\n    \"\"\"Load TIFF stack -&gt; (D,H,W)\"\"\"\n    im = Image.open(path)\n    slices = [np.array(frame) for frame in ImageSequence.Iterator(im)]\n    return np.stack(slices, axis=0)</p>\n<p>def maybe_transpose_gt(gt: np.ndarray, pred_shape: tuple) -&gt; np.ndarray:\n    \"\"\"\n    Fix common GT orientation mismatch (GT stored as (H,W,D)).\n    If GT last dim matches pred depth, transpose.\n    \"\"\"\n    if gt.shape == pred_shape:\n        return gt\n    if gt.ndim == 3 and gt.shape[-1] == pred_shape[0] and gt.shape[:2] == pred_shape[1:]:\n        return np.transpose(gt, (2, 0, 1))\n    return gt</p>\n<h1>============================================================</h1>\n<h1>PREDICTED HOLES (ENCLOSED VOIDS) DIAGNOSTIC</h1>\n<h1>============================================================</h1>\n<p>def count_enclosed_voids(binary_vol: np.ndarray) -&gt; int:\n    \"\"\"\n    Approximate 'holes' as enclosed 0-regions in the complement of predicted foreground.\n    binary_vol: (D,H,W) uint8 {0,1}\n    Returns number of enclosed void components.\n    \"\"\"\n    vol = (binary_vol &gt; 0)\n    D, H, W = vol.shape</p>\n<pre><code># Complement background\nbg = ~vol\nseen = np.zeros(bg.shape, dtype=bool)\n\nq = deque()\n\ndef push(z, y, x):\n    if 0 &lt;= z &lt; D and 0 &lt;= y &lt; H and 0 &lt;= x &lt; W and bg[z, y, x] and not seen[z, y, x]:\n        seen[z, y, x] = True\n        q.append((z, y, x))\n\n# Flood fill from boundary background voxels -&gt; marks \"outside\"\nfor z in (0, D - 1):\n    for y in range(H):\n        for x in range(W):\n            push(z, y, x)\nfor z in range(D):\n    for y in (0, H - 1):\n        for x in range(W):\n            push(z, y, x)\nfor z in range(D):\n    for y in range(H):\n        for x in (0, W - 1):\n            push(z, y, x)\n\n# 6-neighborhood\nnbrs = [(1, 0, 0), (-1, 0, 0),\n        (0, 1, 0), (0, -1, 0),\n        (0, 0, 1), (0, 0, -1)]\n\nwhile q:\n    z, y, x = q.popleft()\n    for dz, dy, dx in nbrs:\n        push(z + dz, y + dy, x + dx)\n\n# Background not reached from boundary = enclosed cavities\ninside = bg &amp; (~seen)\n\n# Count connected components in inside (6-neighborhood)\nholes = 0\nseen2 = np.zeros_like(inside, dtype=bool)\nfor z in range(D):\n    for y in range(H):\n        for x in range(W):\n            if inside[z, y, x] and not seen2[z, y, x]:\n                holes += 1\n                qq = deque([(z, y, x)])\n                seen2[z, y, x] = True\n                while qq:\n                    zz, yy, xx = qq.popleft()\n                    for dz, dy, dx in nbrs:\n                        nz, ny, nx = zz + dz, yy + dy, xx + dx\n                        if 0 &lt;= nz &lt; D and 0 &lt;= ny &lt; H and 0 &lt;= nx &lt; W:\n                            if inside[nz, ny, nx] and not seen2[nz, ny, nx]:\n                                seen2[nz, ny, nx] = True\n                                qq.append((nz, ny, nx))\nreturn holes\n</code></pre>\n<h1>============================================================</h1>\n<h1>TOPO METRICS INSTALL</h1>\n<h1>============================================================</h1>\n<p>_TOPO_OK = False</p>\n<p>def install_topometrics():\n    global _TOPO_OK\n    if _TOPO_OK:\n        return\n    try:\n        import topometrics.leaderboard\n        _TOPO_OK = True\n        return\n    except:\n        pass</p>\n<pre><code>resources = \"/kaggle/input/vesuvius-metric-resources\"\nworkdir   = \"/kaggle/working/topological-metrics-kaggle\"\n\nsubprocess.run(\n    f\"cd {resources} &amp;&amp; uv pip install --no-index --find-links=wheels \"\n    f\"-r topological-metrics-kaggle/requirements.txt\",\n    shell=True, check=True,\n)\n\nsubprocess.run(\n    f\"cp -r {resources}/topological-metrics-kaggle /kaggle/working/\",\n    shell=True, check=True\n)\n\nsubprocess.run(\n    f\"cd {workdir} &amp;&amp; chmod +x scripts/setup_submodules.sh scripts/build_betti.sh &amp;&amp; make build-betti\",\n    shell=True, check=True\n)\n\nsubprocess.run(\n    f\"cd {workdir} &amp;&amp; uv pip install -e . --no-deps --no-index --no-build-isolation\",\n    shell=True, check=True\n)\n\nsys.path.append(f\"{workdir}/src\")\nimportlib.invalidate_caches()\n\nimport topometrics.leaderboard\n_TOPO_OK = True\n</code></pre>\n<h1>============================================================</h1>\n<h1>OFFICIAL SCORING CALL</h1>\n<h1>============================================================</h1>\n<p>def score_official(gt_u8, pr_u8):\n    \"\"\"\n    gt_u8: (D,H,W) uint8 with ignore label=2 preserved\n    pr_u8: (D,H,W) uint8 {0,1}\n    \"\"\"\n    install_topometrics()\n    import topometrics.leaderboard as lb</p>\n<pre><code>r = lb.compute_leaderboard_score(\n    predictions=pr_u8.astype(np.uint8),\n    labels=gt_u8.astype(np.uint8),\n    dims=(0, 1, 2),\n    spacing=(1.0, 1.0, 1.0),\n    surface_tolerance=SURFACE_TOLERANCE,\n    voi_connectivity=26,\n    voi_transform=\"one_over_one_plus\",\n    voi_alpha=0.3,\n    combine_weights=(0.3, 0.35, 0.35),\n    ignore_label=2,\n)\n\n# VOI attribute differs in some versions\nvoi_score = r.voi.voi_score if hasattr(r.voi, \"voi_score\") else r.voi.oi_score\n\nreturn dict(\n    score=float(r.score),\n    topo=float(r.topo.toposcore),\n    surface_dice=float(r.surface_dice),\n    voi=float(voi_score),\n)\n</code></pre>\n<h1>============================================================</h1>\n<h1>UNZIP SUBMISSION</h1>\n<h1>============================================================</h1>\n<p>if EXTRACT_DIR.exists():\n    shutil.rmtree(EXTRACT_DIR)\nEXTRACT_DIR.mkdir(parents=True, exist_ok=True)</p>\n<p>with zipfile.ZipFile(SUBMISSION_ZIP, \"r\") as z:\n    z.extractall(EXTRACT_DIR)</p>\n<p>print(f\"✓ Extracted {SUBMISSION_ZIP} → {EXTRACT_DIR}\")</p>\n<h1>============================================================</h1>\n<h1>SCORE EACH ID</h1>\n<h1>============================================================</h1>\n<p>rows = []</p>\n<p>for sid in EVAL_IDS:\n    pred_path = EXTRACT_DIR / f\"{sid}.tif\"\n    gt_path   = LABEL_DIR / f\"{sid}.tif\"</p>\n<pre><code>if not pred_path.exists():\n    print(f\"❌ Missing prediction for {sid}\")\n    continue\nif not gt_path.exists():\n    print(f\"❌ Missing GT for {sid}\")\n    continue\n\npred = load_volume_tif(pred_path)\ngt   = load_volume_tif(gt_path)\n\n# Prediction is binary {0,1}\npr_u8 = (pred &gt; 0).astype(np.uint8)\n\n# GT must keep ignore label=2\n# Map: 2 → 2, (0 → 0), (nonzero but not 2 → 1)\ngt_raw = gt.astype(np.uint8)\ngt_u8 = np.where(gt_raw == 2, 2, (gt_raw &gt; 0).astype(np.uint8))\n\n# Fix possible (H,W,D) → (D,H,W)\ngt_u8 = maybe_transpose_gt(gt_u8, pr_u8.shape)\n\nif gt_u8.shape != pr_u8.shape:\n    raise ValueError(f\"Shape mismatch for {sid}: pred={pr_u8.shape}, gt={gt_u8.shape}\")\n\n# Diagnostic: predicted enclosed voids (\"holes\")\npred_holes = count_enclosed_voids(pr_u8)\n\nprint(f\"[{sid}] pred uniques={np.unique(pr_u8)}  gt uniques={np.unique(gt_u8)}  pred_holes={pred_holes}\")\n\nm = score_official(gt_u8, pr_u8)\nm[\"id\"] = sid\nm[\"pred_holes\"] = int(pred_holes)\nrows.append(m)\n\nprint(f\"✓ {sid}: score={m['score']:.4f} | topo={m['topo']:.4f} | sd={m['surface_dice']:.4f} | voi={m['voi']:.4f} | holes={m['pred_holes']}\")\n</code></pre>\n<h1>============================================================</h1>\n<h1>SUMMARY TABLE</h1>\n<h1>============================================================</h1>\n<p>df = pd.DataFrame(rows)</p>\n<p>print(\"\\n=== Per-volume ===\")\ncols = [\"id\", \"score\", \"topo\", \"surface_dice\", \"voi\", \"pred_holes\"]\nprint(df[cols].to_string(index=False))</p>\n<p>agg_cols = [\"score\", \"topo\", \"surface_dice\", \"voi\", \"pred_holes\"]\nmean = df[agg_cols].mean().add_prefix(\"mean_\")\nstd  = df[agg_cols].std(ddof=1).add_prefix(\"std_\")</p>\n<h1>CV (coefficient of variation) not very meaningful for pred_holes if mean is small,</h1>\n<h1>but we compute anyway for consistency.</h1>\n<p>cv   = (df[agg_cols].std(ddof=1) / df[agg_cols].mean().replace(0, np.nan)).add_prefix(\"cv_\")</p>\n<p>summary = pd.concat([mean, std, cv], axis=0).to_frame().T</p>\n<p>print(\"\\n=== Summary ===\")\nprint(summary.to_string(index=False))</p>\n<h1>Optional: save outputs</h1>\n<p>df.to_csv(\"cv_scores_with_holes.csv\", index=False)\nsummary.to_csv(\"cv_summary_with_holes.csv\", index=False)</p>\n<p>summary\n`</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 3404011,
      "author_name": "Cody_Null",
      "author_url": "",
      "post_date": "2026-02-09T18:10:00.410000",
      "content": "<p>I have extremely mixed feelings about this. On one hand, I am happy to solve the real problem better. But on the other hand, we have already spent months working on the current solution and we couldn’t possibly retest all of the ideas on cross validation that we once had in two weeks. I am certain that my wife will not be thrilled about me expecting two more weeks on this competition so any advice people may have on that would be fabulous. 😂😂 I just hope that we remain a strong competitor and that we can help give the solution you guys need! </p>",
      "votes": 6,
      "replies": [
        {
          "id": 3404206,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2026-02-10T03:35:19.667000",
          "content": "<p>You need to do something on Feb 14 to continue competing on Kaggle.</p>\n<p>Lesson learned: GPU only is not all you need.</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 3404220,
          "author_name": "Yiheng Wang",
          "author_url": "",
          "post_date": "2026-02-10T04:15:56.307000",
          "content": "<p>I planned to have a rest on the holiday, and back to focus on the Deep Past Challenge 😅 Now I decide to skip that one, and wait for a new challenge.</p>",
          "votes": 7,
          "replies": []
        }
      ]
    },
    {
      "id": 3404237,
      "author_name": "Manas Choudhary",
      "author_url": "",
      "post_date": "2026-02-10T04:45:59.777000",
      "content": "<p>I had left working on this challenge for half a month because of some other hackathons, and now some internship, I assumed it's ending soon so i'll be getting a silver atleast. Turns out team had other plans, and now it overlaps with my exams also 😭😭.</p>\n<p>ps--&gt; I still think, if only the test set is changed, then nothing fundamental has changed much, only if you were trying to iteratively improve leaderboard, then It becomes problematic for you</p>",
      "votes": 5,
      "replies": [
        {
          "id": 3404277,
          "author_name": "ArjunB",
          "author_url": "",
          "post_date": "2026-02-10T07:10:18.190000",
          "content": "<p>I too have a lot of other commitments starting earlier this month, which was my main motivation to team up.\nLot of my plans were centered around the comp ending mid Feb. This decision is very very frustrating.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 3404181,
      "author_name": "dan4o",
      "author_url": "",
      "post_date": "2026-02-10T02:08:44.930000",
      "content": "<p><a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> I also highly suggest while fixing the data and rescoring to also patch one of the longest known issues in this competition - the faulty logic of the validation mask cutting the predictions and spawning multiple components/holes if the mask intersects our predictions discussed <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/653482\" target=\"_blank\">here</a>:</p>\n<p>Right now our prediction is zeroed out with the unlabelled mask.</p>\n<pre><code> # Apply ignore: neutralize by setting both PR and GT to background at ignored voxels\n # (Done on float/int volumes before any binarization or topology)\n\n    pr_eval = np.where(ign, 0, pr_raw)\n    gt_eval = np.where(ign, 0, gt_raw)\n</code></pre>\n<p>Instead of zeroing the models predictions and label we should switch to a more elegant solution. Instead of tampering with the prediction and causing unwanted and unregulated side effects (we don't know how much components/holes we will spawn if the prediction intersects the mask) we simply <strong>ignore the predicted components if their birth coordinates lie within the ignore mask</strong>. </p>\n<p>We simply do two steps:</p>\n<ol>\n<li>We calculate the topology for the <strong>full prediction volume</strong>.</li>\n<li>Discard the <strong>unmatched components from the list</strong> if they lie within the <strong>ignore region</strong>. (there could be NO matched components in the ignore region as gt contains nothing there, so we just filter the unmatched list)</li>\n<li>Our topology score now contains matched, predicted, ground truth counts <strong>only from the valid regions</strong>. And we didn't spawn any extra holes/components in doing so.</li>\n</ol>\n<p>This is the original function. Right now we return the unmatched birth coordinates without any filtration.</p>\n<pre><code>    # @classmethod\n    # def _counts_for_dim(cls, result: Any, k: int) -&gt; Tuple[int, int, int]:\n    #     m_k = cls._count_points(cls._safe_index(getattr(result, \"input1_matched_birth_coordinates\", []), k))\n    #     p_un_k = cls._count_points(cls._safe_index(getattr(result, \"input1_unmatched_birth_coordinates\", []), k))\n    #     g_un_k = cls._count_points(cls._safe_index(getattr(result, \"input2_unmatched_birth_coordinates\", []), k))\n    #     p_k = m_k + p_un_k\n    #     g_k = m_k + g_un_k\n    #     return m_k, p_k, g_k\n</code></pre>\n<p>And this is the solution with few lines of changed code - basic if else checks. We calculate the topology for the whole prediction and we don't tamper it or zero it out causing a mess, instead we simply check if the [z,y,x] birth coordinates lie in the ignore region mask! If they do, then discard them - they should not contribute to the total metric, as they should simply be ignored.</p>\n<pre><code>@classmethod\n    def _counts_for_dim(\n        cls, \n        result: Any, \n        k: int, \n        ignore_mask: Optional[np.ndarray]\n    ) -&gt; Tuple[int, int, int]:\n\n        # 1. Fetch raw lists\n        # Note: These are lists of \"birth coordinates\"\n        matched = cls._safe_index(getattr(result, \"input1_matched_birth_coordinates\", []), k)\n        p_unmatched = cls._safe_index(getattr(result, \"input1_unmatched_birth_coordinates\", []), k)\n        g_unmatched = cls._safe_index(getattr(result, \"input2_unmatched_birth_coordinates\", []), k)\n\n        # 2.Filter Helper\n        def count_valid(coords, mask):\n            \"\"\"Returns count of coords NOT in the ignore mask.\"\"\"\n            if coords is None: return 0\n            n_total = cls._count_points(coords)\n            if mask is None or n_total == 0:\n                return n_total\n\n            arr = np.asarray(coords)\n            if arr.ndim != 2 or arr.shape[1] != 3:\n                return n_total \n\n            valid_count = 0\n            for i in range(arr.shape[0]):\n                c0, c1, c2 = int(arr[i, 0]), int(arr[i, 1]), int(arr[i, 2])\n\n                if (0 &lt;= c0 &lt; mask.shape[0] and \n                    0 &lt;= c1 &lt; mask.shape[1] and \n                    0 &lt;= c2 &lt; mask.shape[2]):\n\n                    # If mask is True, it is IGNORED -&gt; Do not count\n                    if not mask[c0, c1, c2]:\n                        valid_count += 1\n                else:\n                    valid_count += 1\n            return valid_count\n\n        # 3. Calculate Counts\n        m_k_count = cls._count_points(matched)\n        g_un_k_count = count_valid(g_unmatched, ignore_mask) \n        p_un_k_count = count_valid(p_unmatched, ignore_mask)\n\n        p_k = m_k_count + p_un_k_count\n        g_k = m_k_count + g_un_k_count\n\n        return m_k_count, p_k, g_k\n</code></pre>\n<p>We just pass the ignore mask down to this function and use it as is.(Note: the mask should be sliced as well, because we are using (2,2,2) splits for the pred/gt)</p>\n<p>In hindsight, the solution to this was simple, but maybe we dimissed the possibility that something can be done to fix it.</p>\n<p>If there are no edge-cases or bugs in the fix (I encourage competitors to think on it as well, so we can find mistakes if there are any), I think this will also solve a long known issue that surfaced over two months ago, making the metric waterproof - and taking chance and lucky (unlucky) mask hitting out of the equation.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 3404184,
          "author_name": "Starry",
          "author_url": "",
          "post_date": "2026-02-10T02:17:15.290000",
          "content": "<p>totally agree.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 3404262,
          "author_name": "Marius Heuser",
          "author_url": "",
          "post_date": "2026-02-10T06:07:27.647000",
          "content": "<p>I havent checked how the birth coordinates are calculated, but your suggestion would remove a correct segmentation if it were to \"birth\" in label 2. Which might happen alot because of the 3 voxel labeled 2 at each edge/face of the volume?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3405061,
      "author_name": "lingyundev",
      "author_url": "",
      "post_date": "2026-02-12T01:17:37.953000",
      "content": "<p><a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> Hi, Could you please clarify what exactly was fixed in the test set and how the evaluation metric was updated? For example, were holes in the labels filled algorithmically?\nThis information is important for fairness, especially since the training set remains unchanged while the test set was modified. Without knowing the specific fixes, it’s difficult to properly adapt and validate our models locally.</p>\n<p>Thank you.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 3404272,
      "author_name": "hsiaosuan",
      "author_url": "",
      "post_date": "2026-02-10T06:45:42.023000",
      "content": "<p>haha, totally not cool with the fact that the new timeline now extends into the lunar new yr holiday. </p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 3404763,
      "author_name": "tingyi",
      "author_url": "",
      "post_date": "2026-02-11T07:13:54.773000",
      "content": "<p>Seeing everyone’s scores getting so high, I just realized I’m the one swimming naked. 😀</p>",
      "votes": 4,
      "replies": [
        {
          "id": 3404859,
          "author_name": "Manas Choudhary",
          "author_url": "",
          "post_date": "2026-02-11T11:33:59.313000",
          "content": "<p>What I'm 100% about is that, the final rankings on private will be far more random than this overhaul.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3404468,
      "author_name": "Santhosh20050206",
      "author_url": "",
      "post_date": "2026-02-10T15:00:46.253000",
      "content": "<p>Has anyone gotten their updated scores after the re-evaluation? It looks like my recent submissions are still being assessed using the previous grading criteria.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3404103,
      "author_name": "Roy Segalz",
      "author_url": "",
      "post_date": "2026-02-09T21:55:33.293000",
      "content": "<p>You do not change the train set but if one would independently retrain on fixed labels, that would give a big advantage on the test set.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3404261,
      "author_name": "tingyi",
      "author_url": "",
      "post_date": "2026-02-10T05:59:15.483000",
      "content": "<p>I have a small question - do both Betti number 1 and Betti number 2 get corrected, or only Betti number 2?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3404222,
      "author_name": "ForcewithMe",
      "author_url": "",
      "post_date": "2026-02-10T04:17:25.973000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> <a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> , What will happen to the submissions that was submitted before the update, but complete after the update. Is the score old or new?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3404443,
          "author_name": "Ryan Holbrook",
          "author_url": "",
          "post_date": "2026-02-10T14:03:21.937000",
          "content": "<p>They would be scored with the updated labels only. But since we're rescoring all submissions with the updated labels, there won't be any difference in the end.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3404450,
              "author_name": "Cody_Null",
              "author_url": "",
              "post_date": "2026-02-10T14:22:34.020000",
              "content": "<p>Ahh, I am just now understanding this. If we submit now we get new score, otherwise we are waiting on all submissions to be updated? </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3404452,
              "author_name": "GG Ayo (AyoGG)",
              "author_url": "",
              "post_date": "2026-02-10T14:26:05.177000",
              "content": "<p>However, the newly submitted version from our team doesn't seem to have any difference in score.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3403993,
      "author_name": "Tony Li",
      "author_url": "",
      "post_date": "2026-02-09T17:19:37.503000",
      "content": "<p>Hi Giorgio — thanks again for handling the rescore, and for all the work running the competition.</p>\n<p>Quick question for clarity: for this rescore, will you re-evaluate all submissions, or only a subset (e.g., each team’s best / final-selected submissions)? If it’s a subset, could you share how submissions are chosen?</p>\n<p>Thanks a lot!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3404072,
          "author_name": "Ryan Holbrook",
          "author_url": "",
          "post_date": "2026-02-09T20:49:40.150000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/tonylica\" target=\"_blank\">@tonylica</a>, We will rescore all submissions (that is, run the metric over the saved predictions, not rerun the entire model).</p>",
          "votes": 5,
          "replies": [
            {
              "id": 3404074,
              "author_name": "Tony Li",
              "author_url": "",
              "post_date": "2026-02-09T21:00:37.600000",
              "content": "<p>Hi Ryan — thank you for the clarification. That makes perfect sense. Appreciate the quick response and all the work on this.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3405130,
      "author_name": "Navneet",
      "author_url": "",
      "post_date": "2026-02-12T07:19:49.410000",
      "content": "<p>Thank you for the update <a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3404354,
      "author_name": "Sasha Turutin",
      "author_url": "",
      "post_date": "2026-02-10T11:01:03.057000",
      "content": "<p>The worst orgs ever.</p>",
      "votes": -16,
      "replies": []
    },
    {
      "id": 3407003,
      "author_name": "Victoria",
      "author_url": "",
      "post_date": "2026-02-17T11:02:37.020000",
      "content": "<p>its just some ever-some </p>",
      "votes": -2,
      "replies": []
    },
    {
      "id": 3406677,
      "author_name": "",
      "author_url": "",
      "post_date": "2026-02-16T10:23:16.917000",
      "content": "",
      "votes": -2,
      "replies": []
    },
    {
      "id": 3406035,
      "author_name": "",
      "author_url": "",
      "post_date": "2026-02-14T12:01:43.687000",
      "content": "",
      "votes": -3,
      "replies": [
        {
          "id": 3406543,
          "author_name": "",
          "author_url": "",
          "post_date": "2026-02-16T00:49:35.397000",
          "content": "",
          "votes": 4,
          "replies": []
        },
        {
          "id": 3406612,
          "author_name": "",
          "author_url": "",
          "post_date": "2026-02-16T07:27:53.550000",
          "content": "",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 3404812,
      "author_name": "",
      "author_url": "",
      "post_date": "2026-02-11T09:58:43.403000",
      "content": "",
      "votes": -1,
      "replies": []
    },
    {
      "id": 3408579,
      "author_name": "",
      "author_url": "",
      "post_date": "2026-02-20T22:13:10.250000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3404170,
      "author_name": "",
      "author_url": "",
      "post_date": "2026-02-10T01:25:25.623000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 3404391,
          "author_name": "",
          "author_url": "",
          "post_date": "2026-02-10T12:13:52.790000",
          "content": "",
          "votes": 0,
          "replies": [
            {
              "id": 3404402,
              "author_name": "",
              "author_url": "",
              "post_date": "2026-02-10T12:33:12.723000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3403995,
      "author_name": "",
      "author_url": "",
      "post_date": "2026-02-09T17:26:25.430000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3404412,
      "author_name": "",
      "author_url": "",
      "post_date": "2026-02-10T12:45:18.110000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3404223,
      "author_name": "",
      "author_url": "",
      "post_date": "2026-02-10T04:17:56.827000",
      "content": "",
      "votes": -3,
      "replies": []
    },
    {
      "id": 3404155,
      "author_name": "",
      "author_url": "",
      "post_date": "2026-02-10T00:06:04.290000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3404154,
      "author_name": "",
      "author_url": "",
      "post_date": "2026-02-10T00:02:56.320000",
      "content": "",
      "votes": -4,
      "replies": []
    },
    {
      "id": 3403998,
      "author_name": "",
      "author_url": "",
      "post_date": "2026-02-09T17:36:13.310000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 3404002,
          "author_name": "",
          "author_url": "",
          "post_date": "2026-02-09T17:42:40.723000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3403956": "After identifying a [critical issue](https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/671160) in the data and completing a joint analysis with the Kaggle Support team, we will **update the test set** and **rescore submissions**.\n\n* **Rescore timing:** The rescore will start later today. Due to the metric’s long runtime, it will likely take **at least a day**, and the public leaderboard may shift during this period.\n* **Deadline extension:** To ensure everyone has time to calibrate to the updated public leaderboard, we are **extending the competition deadline by 2 weeks** (the competition page will reflect the new date).\n* **Training data:** After considering community feedback and the work teams have already done to revise training pipelines, we believe updating the training set now could raise fairness concerns. **The training set will not be updated.**\n\nThanks for your understanding, and let’s keep pushing toward the best possible models to help unwrap the scrolls.",
    "3404453": "I understand the frustration, believe me, as we were tied first place going into this change. However, I think maybe we should take it easier on the hosts as they have been participating a ton, answering questions on the weekends, and trying to help out. They really just want the best solution they can get as they are paying for it after all. I can't say I blame them. Obviously, it can be inconvenient for us and the timeframe is not ideal, but I think they are doing what they can.",
    "3404284": "Not surprisingly, this is going to demotivate those who worked hard till now. Some comments show it already.\n\nWhat I don't get is why the training data isn't fixed either? Now you added an extra complexity with a distribution shift between training and testing data.",
    "3404005": "Thanks a lot for the effort in handling this update and for keeping the competition aligned with its ultimate goal, helping to read the scrolls, rather than competing for the sake of competition. We really appreciate the transparency and the work from your team.\n\nCould you please clarify what the fix is exactly?\n\nMore specifically, should participants interpret the rescore as applying the approach mentioned by @dankrstev? That is, keeping the data unchanged but adjusting the metric so it is computed as if no holes exist? Or was any additional processing applied, such as a binary closing or another morphological operation??\nhttps://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/671160#3403175\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2F3fa5a1afd5feab38662b5c03600d68a7%2Faa.png?generation=1770659017538322&alt=media)\n",
    "3404334": "Could you please disclose fix points?\nWe can't fix local pipeline...\n\nif this points are not important, we don't need extend long deadline... ",
    "3404247": "Haha, I am so angry I can't make any constructive comment. Not sure I will continue this comp. ",
    "3404281": "Honestly, I’m quite disappointed about the deadline extension. One big reason I joined this competition was because it was supposed to end before the Lunar New Year. Many of us planned our time around that. Extending it into the holiday really affects those plans and makes it very hard to balance family time and the competition.\nI understand that fixing the data is important, but changing the timeline this late — especially into a major holiday — feels really tough for a lot of Chinese participants.",
    "3404373": "I believe that revising the test set is not directly related to postponing the competition, since the training set will not change and most teams have already  completed their models. I hope the organizers can take participants’ feelings into consideration—everyone has put a great deal of effort into this competition, whether in terms of time or the cost of renting GPUs. A short extension would be understandable, but a two-week delay is really too long and quite stressful for us.",
    "3404026": "Let’s be real: this extension is a double-edged sword for Chinese participants. We are entering the Spring Festival—the most important time for family reunions in our culture. While the extra time is appreciated, it effectively means choosing between quality time with family and squeezing out that last bit of model performance. \n\nIt’s going to be a tough grind during the holidays! ",
    "3403961": "@giorgioangelotti \nThanks. I’m curious, after fixing the test set and evaluation metrics, what score did your baseline nnU-Net achieve without any post-processing?",
    "3404439": "Has anyone received their re-graded scores yet? I noticed that my new uploads still seem to use the old grading method.",
    "3403985": "Spent tons of time during the past days and tired, now  it would take more time because of extension deadline even in Spring festival 😱. Tired and tired.. But anyway, this kind of update I think it is the best way for fairness of the competetion.",
    "3404838": "Since only the test set was updated, I believe this has introduced an additional challenge of dealing with a potential domain shift. Could you please share a bit more detail on what was changed?\n\nFor example:\n- Applied some algorithmic post-processing to fill small holes in the labels.\n- Removed highly porous samples from the test set, or otherwise filtered/replaced them.",
    "3404318": "I didn't understand the reason for the extension, since only the test data was updated.",
    "3404406": "I also think that the organisers are totally ignoring that **humans** are taking part in this competition. People have worked really hard, made sacrifices, and organised their personal life in such a way that they would be able to work extra hard until Friday. I was personally really looking forward to Friday.\n\nMoreover, changing the test set and not the training set is problematic. 2 weeks is not that much time to adapt our strategy, especially when the rescoring is taking time and the distribution shift between training and test sets means we now rely on LB even more. It does not leave much time to retrain and fine-tune models considering that the models take days to train with a good machine.\n\nWhile I love this project and I still have plenty of ideas to improve my solution, I am still unsure whether I will continue the competition as it's been an absolute chaos and new potential bugs have already been mentioned. Yes, the metric is \"a bit\" broken and it is clearly an issue from a data science point of view, but I think moving the goal posts at the last minute is not the solution.",
    "3404899": "The rescore is now complete. Please let me know if you have any questions or concerns about the rescore.",
    "3404011": "I have extremely mixed feelings about this. On one hand, I am happy to solve the real problem better. But on the other hand, we have already spent months working on the current solution and we couldn’t possibly retest all of the ideas on cross validation that we once had in two weeks. I am certain that my wife will not be thrilled about me expecting two more weeks on this competition so any advice people may have on that would be fabulous. 😂😂 I just hope that we remain a strong competitor and that we can help give the solution you guys need! ",
    "3404237": "I had left working on this challenge for half a month because of some other hackathons, and now some internship, I assumed it's ending soon so i'll be getting a silver atleast. Turns out team had other plans, and now it overlaps with my exams also 😭😭.\n\nps--> I still think, if only the test set is changed, then nothing fundamental has changed much, only if you were trying to iteratively improve leaderboard, then It becomes problematic for you",
    "3404181": "@giorgioangelotti I also highly suggest while fixing the data and rescoring to also patch one of the longest known issues in this competition - the faulty logic of the validation mask cutting the predictions and spawning multiple components/holes if the mask intersects our predictions discussed [here](https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/653482):\n\nRight now our prediction is zeroed out with the unlabelled mask.\n```   \n # Apply ignore: neutralize by setting both PR and GT to background at ignored voxels\n # (Done on float/int volumes before any binarization or topology)\n\n    pr_eval = np.where(ign, 0, pr_raw)\n    gt_eval = np.where(ign, 0, gt_raw)\n```\n\nInstead of zeroing the models predictions and label we should switch to a more elegant solution. Instead of tampering with the prediction and causing unwanted and unregulated side effects (we don't know how much components/holes we will spawn if the prediction intersects the mask) we simply **ignore the predicted components if their birth coordinates lie within the ignore mask**. \n\nWe simply do two steps:\n 1. We calculate the topology for the **full prediction volume**.\n 2. Discard the **unmatched components from the list** if they lie within the **ignore region**. (there could be NO matched components in the ignore region as gt contains nothing there, so we just filter the unmatched list)\n 3. Our topology score now contains matched, predicted, ground truth counts **only from the valid regions**. And we didn't spawn any extra holes/components in doing so.\n\nThis is the original function. Right now we return the unmatched birth coordinates without any filtration.\n```\n    # @classmethod\n    # def _counts_for_dim(cls, result: Any, k: int) -> Tuple[int, int, int]:\n    #     m_k = cls._count_points(cls._safe_index(getattr(result, \"input1_matched_birth_coordinates\", []), k))\n    #     p_un_k = cls._count_points(cls._safe_index(getattr(result, \"input1_unmatched_birth_coordinates\", []), k))\n    #     g_un_k = cls._count_points(cls._safe_index(getattr(result, \"input2_unmatched_birth_coordinates\", []), k))\n    #     p_k = m_k + p_un_k\n    #     g_k = m_k + g_un_k\n    #     return m_k, p_k, g_k\n```\n\nAnd this is the solution with few lines of changed code - basic if else checks. We calculate the topology for the whole prediction and we don't tamper it or zero it out causing a mess, instead we simply check if the [z,y,x] birth coordinates lie in the ignore region mask! If they do, then discard them - they should not contribute to the total metric, as they should simply be ignored.\n\n```\n@classmethod\n    def _counts_for_dim(\n        cls, \n        result: Any, \n        k: int, \n        ignore_mask: Optional[np.ndarray]\n    ) -> Tuple[int, int, int]:\n        \n        # 1. Fetch raw lists\n        # Note: These are lists of \"birth coordinates\"\n        matched = cls._safe_index(getattr(result, \"input1_matched_birth_coordinates\", []), k)\n        p_unmatched = cls._safe_index(getattr(result, \"input1_unmatched_birth_coordinates\", []), k)\n        g_unmatched = cls._safe_index(getattr(result, \"input2_unmatched_birth_coordinates\", []), k)\n\n        # 2.Filter Helper\n        def count_valid(coords, mask):\n            \"\"\"Returns count of coords NOT in the ignore mask.\"\"\"\n            if coords is None: return 0\n            n_total = cls._count_points(coords)\n            if mask is None or n_total == 0:\n                return n_total\n\n            arr = np.asarray(coords)\n            if arr.ndim != 2 or arr.shape[1] != 3:\n                return n_total \n\n            valid_count = 0\n            for i in range(arr.shape[0]):\n                c0, c1, c2 = int(arr[i, 0]), int(arr[i, 1]), int(arr[i, 2])\n                \n                if (0 <= c0 < mask.shape[0] and \n                    0 <= c1 < mask.shape[1] and \n                    0 <= c2 < mask.shape[2]):\n                    \n                    # If mask is True, it is IGNORED -> Do not count\n                    if not mask[c0, c1, c2]:\n                        valid_count += 1\n                else:\n                    valid_count += 1\n            return valid_count\n\n        # 3. Calculate Counts\n        m_k_count = cls._count_points(matched)\n        g_un_k_count = count_valid(g_unmatched, ignore_mask) \n        p_un_k_count = count_valid(p_unmatched, ignore_mask)\n\n        p_k = m_k_count + p_un_k_count\n        g_k = m_k_count + g_un_k_count\n        \n        return m_k_count, p_k, g_k\n```\n\nWe just pass the ignore mask down to this function and use it as is.(Note: the mask should be sliced as well, because we are using (2,2,2) splits for the pred/gt)\n\nIn hindsight, the solution to this was simple, but maybe we dimissed the possibility that something can be done to fix it.\n\nIf there are no edge-cases or bugs in the fix (I encourage competitors to think on it as well, so we can find mistakes if there are any), I think this will also solve a long known issue that surfaced over two months ago, making the metric waterproof - and taking chance and lucky (unlucky) mask hitting out of the equation.",
    "3405061": "@giorgioangelotti @ryanholbrook Hi, Could you please clarify what exactly was fixed in the test set and how the evaluation metric was updated? For example, were holes in the labels filled algorithmically?\nThis information is important for fairness, especially since the training set remains unchanged while the test set was modified. Without knowing the specific fixes, it’s difficult to properly adapt and validate our models locally.\n\nThank you.",
    "3404272": "haha, totally not cool with the fact that the new timeline now extends into the lunar new yr holiday. ",
    "3404763": "Seeing everyone’s scores getting so high, I just realized I’m the one swimming naked. 😀",
    "3404468": "Has anyone gotten their updated scores after the re-evaluation? It looks like my recent submissions are still being assessed using the previous grading criteria.\n",
    "3404103": "You do not change the train set but if one would independently retrain on fixed labels, that would give a big advantage on the test set.",
    "3404261": "I have a small question - do both Betti number 1 and Betti number 2 get corrected, or only Betti number 2?",
    "3404222": "Hi @ryanholbrook @giorgioangelotti , What will happen to the submissions that was submitted before the update, but complete after the update. Is the score old or new?",
    "3403993": "Hi Giorgio — thanks again for handling the rescore, and for all the work running the competition.\n\nQuick question for clarity: for this rescore, will you re-evaluate all submissions, or only a subset (e.g., each team’s best / final-selected submissions)? If it’s a subset, could you share how submissions are chosen?\n\nThanks a lot!",
    "3405130": "Thank you for the update @giorgioangelotti \n",
    "3404354": "The worst orgs ever.",
    "3407003": "its just some ever-some ",
    "3406677": "\"Great update, thanks for keeping the competition fair for everyone. Let's keep pushing!\"",
    "3406035": "im a beginner how can i learn fast and efficient",
    "3404812": "Have the scores been updated?\n",
    "3408579": "Thanks for the transparency and quick action on this.\nReally appreciate the coordination with Kaggle Support and the deadline extension that definitely helps teams recalibrate fairly.\nLooking forward to seeing how the leaderboard evolves after the rescore!",
    "3404170": "My model performs better locally but worse on the leaderboard—it's frustrating. Thankfully, the Host ultimately opted to re-evaluate the scores.",
    "3403995": "so all submission will be rescored again which we submitted ",
    "3404412": "",
    "3404223": "",
    "3404155": "",
    "3404154": "",
    "3403998": ""
  }
}