{
  "id": 660123,
  "title": "Some thoughts on challenge design",
  "url": "/competitions/vesuvius-challenge-surface-detection/discussion/660123",
  "author_name": "",
  "post_date": "2025-12-12T11:06:15.463886700Z",
  "votes": 19,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Hi all,</p>\n<p>I wanted to share a few thoughts regarding the overall challenge design. I believe there are some aspects that could be improved. Since we are still relatively early in the competition, it seems like a good moment to discuss this and, if possible, adjust certain design choices so the challenge is better aligned with the downstream task.</p>\n<p>I think this should also be in the organizers’ interest. A better task formulation has the potential to significantly improve the final outcome when it comes to scroll unrolling, which is ultimately what this competition is about.</p>\n<p>Below are some thoughts that came up while working with the data and thinking about the evaluation setup. They are not fully polished ideas, but I hope they can serve as a starting point for discussion.</p>\n<p>(tagging <a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> :-) )</p>\n<hr>\n<h3>1. Formulation as semantic segmentation unnecessarily biases the solution space</h3>\n<p>I understand that the organizers have run preliminary experiments and that, in their internal pipelines, framing the problem as a 3D segmentation task has yielded promising results. However, this formulation comes with inherent issues.</p>\n<p>Examples include touching sheets, which can also occur in the ground truth (see hengck23’s comment: <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/651532#3370748)\" target=\"_blank\">https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/651532#3370748)</a>, inconsistent and seemingly arbitrary sheet thicknesses, and sheets that appear to float in empty space (see <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/651532#3369783)\" target=\"_blank\">https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/651532#3369783)</a>.</p>\n<p>The organizers have confirmed that annotation is performed using meshes that are manually adjusted to follow the sheets as well as possible. Meshes are a very natural representation for this task. They provide explicit sheet instances, and their exact location does not need to be voxel-perfect as long as it is good enough for virtual unwrapping.</p>\n<p>Given this, I find it somewhat confusing that only voxelized semantic segmentation masks are provided. By constraining participants to a 3D segmentation formulation, or requiring them to derive better representations such as instance maps or sheet surfaces from this unfitting formulation, the challenge may be unnecessarily limiting the space of possible solutions. This conversion step is inherently error prone and also wasteful, since instance maps or meshes could be provided directly. Participants might be able to explore more creative and potentially more effective approaches if they were not forced into a segmentation-only paradigm.</p>\n<hr>\n<h3>2. The use of Surface Dice as a metric</h3>\n<p>I want to emphasize that I generally like Surface Dice as a metric. However, in this particular setting, its use is difficult to justify.</p>\n<p>The ground-truth sheets exhibit varying thicknesses that do not seem to follow a deterministic or physically meaningful pattern. My impression is that the mesh-to-segmentation conversion script thins sheets in regions where they would otherwise touch. This is a sensible engineering decision, but it directly contradicts the framing of the task as a semantic segmentation problem.</p>\n<p>As a participant, it becomes unclear what sheet thickness one should aim for at a given location. A tolerance of two voxels does not seem sufficiently forgiving to absorb this ambiguity. Additionally, since some ground-truth sheets appear to float in the background, it may be fundamentally impossible to reliably hit the tolerance threshold in those regions, regardless of how reasonable the predicted surface is from an unwrapping perspective.</p>\n<hr>\n<h3>Possible directions for improvement</h3>\n<p>Below are several possible changes, ordered roughly from most radical to least intrusive. Some of these options are mutually exclusive. These are meant as ideas rather than concrete proposals.</p>\n<h4>1. Drop everything related to segmentation entirely (most radical)</h4>\n<p>This is fundamentally a mesh generation problem, not a segmentation problem. Provide meshes as ground truth and evaluate predicted meshes directly. Meshes should be provided as instances, one mesh per sheet.</p>\n<p>Metrics could be based on mesh pairing combined with average symmetric mesh distance, along with VOI-like and TopoScore-inspired measures adapted to mesh representations. This would require careful design, but it seems feasible.</p>\n<p>Participants could still choose to solve the problem via voxel segmentation if they wish. Organizers could provide convenience utilities for converting between meshes and voxel masks in both directions.</p>\n<h4>2. Reformulate as instance segmentation and drop thickness-sensitive metrics (balanced)</h4>\n<p>Since meshes are already available as annotations, it should be possible to provide instance maps rather than a single semantic foreground mask. Participants would be expected to submit instance segmentations.</p>\n<p>This would eliminate connected-components ambiguities in both ground truth and predictions, which I believe is crucial. Participants could still choose to predict semantic masks internally, but the responsibility of converting these into instances would lie with them, for example via connected components.</p>\n<p>Given that sheet thickness is arbitrary and sheets may appear detached from surrounding material, Surface Dice could be dropped. The evaluation could instead focus on VOI and TopoScore. It may also be worth investigating whether ASSD, computed either on binarized instance maps or per instance, could serve as a more appropriate replacement for Surface Dice.</p>\n<h4>3. Reformulate as instance segmentation while keeping the current metrics (least intrusive)</h4>\n<p>This would be the minimal change. Provide instance maps as ground truth and expect instance predictions, but keep the current metric suite unchanged.</p>\n<hr>\n<p>I would be very interested to hear what other participants think about these points, and especially how the organizers view the trade-offs involved. Any feedback, counterarguments, or alternative ideas would be very welcome.</p>",
  "messages": [
    {
      "id": "3373174",
      "postDate": "12/12/2025 11:06:15",
      "content": "<p>Hi all,</p>\n<p>I wanted to share a few thoughts regarding the overall challenge design. I believe there are some aspects that could be improved. Since we are still relatively early in the competition, it seems like a good moment to discuss this and, if possible, adjust certain design choices so the challenge is better aligned with the downstream task.</p>\n<p>I think this should also be in the organizers’ interest. A better task formulation has the potential to significantly improve the final outcome when it comes to scroll unrolling, which is ultimately what this competition is about.</p>\n<p>Below are some thoughts that came up while working with the data and thinking about the evaluation setup. They are not fully polished ideas, but I hope they can serve as a starting point for discussion.</p>\n<p>(tagging <a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> :-) )</p>\n<hr>\n<h3>1. Formulation as semantic segmentation unnecessarily biases the solution space</h3>\n<p>I understand that the organizers have run preliminary experiments and that, in their internal pipelines, framing the problem as a 3D segmentation task has yielded promising results. However, this formulation comes with inherent issues.</p>\n<p>Examples include touching sheets, which can also occur in the ground truth (see hengck23’s comment: <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/651532#3370748)\" target=\"_blank\">https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/651532#3370748)</a>, inconsistent and seemingly arbitrary sheet thicknesses, and sheets that appear to float in empty space (see <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/651532#3369783)\" target=\"_blank\">https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/651532#3369783)</a>.</p>\n<p>The organizers have confirmed that annotation is performed using meshes that are manually adjusted to follow the sheets as well as possible. Meshes are a very natural representation for this task. They provide explicit sheet instances, and their exact location does not need to be voxel-perfect as long as it is good enough for virtual unwrapping.</p>\n<p>Given this, I find it somewhat confusing that only voxelized semantic segmentation masks are provided. By constraining participants to a 3D segmentation formulation, or requiring them to derive better representations such as instance maps or sheet surfaces from this unfitting formulation, the challenge may be unnecessarily limiting the space of possible solutions. This conversion step is inherently error prone and also wasteful, since instance maps or meshes could be provided directly. Participants might be able to explore more creative and potentially more effective approaches if they were not forced into a segmentation-only paradigm.</p>\n<hr>\n<h3>2. The use of Surface Dice as a metric</h3>\n<p>I want to emphasize that I generally like Surface Dice as a metric. However, in this particular setting, its use is difficult to justify.</p>\n<p>The ground-truth sheets exhibit varying thicknesses that do not seem to follow a deterministic or physically meaningful pattern. My impression is that the mesh-to-segmentation conversion script thins sheets in regions where they would otherwise touch. This is a sensible engineering decision, but it directly contradicts the framing of the task as a semantic segmentation problem.</p>\n<p>As a participant, it becomes unclear what sheet thickness one should aim for at a given location. A tolerance of two voxels does not seem sufficiently forgiving to absorb this ambiguity. Additionally, since some ground-truth sheets appear to float in the background, it may be fundamentally impossible to reliably hit the tolerance threshold in those regions, regardless of how reasonable the predicted surface is from an unwrapping perspective.</p>\n<hr>\n<h3>Possible directions for improvement</h3>\n<p>Below are several possible changes, ordered roughly from most radical to least intrusive. Some of these options are mutually exclusive. These are meant as ideas rather than concrete proposals.</p>\n<h4>1. Drop everything related to segmentation entirely (most radical)</h4>\n<p>This is fundamentally a mesh generation problem, not a segmentation problem. Provide meshes as ground truth and evaluate predicted meshes directly. Meshes should be provided as instances, one mesh per sheet.</p>\n<p>Metrics could be based on mesh pairing combined with average symmetric mesh distance, along with VOI-like and TopoScore-inspired measures adapted to mesh representations. This would require careful design, but it seems feasible.</p>\n<p>Participants could still choose to solve the problem via voxel segmentation if they wish. Organizers could provide convenience utilities for converting between meshes and voxel masks in both directions.</p>\n<h4>2. Reformulate as instance segmentation and drop thickness-sensitive metrics (balanced)</h4>\n<p>Since meshes are already available as annotations, it should be possible to provide instance maps rather than a single semantic foreground mask. Participants would be expected to submit instance segmentations.</p>\n<p>This would eliminate connected-components ambiguities in both ground truth and predictions, which I believe is crucial. Participants could still choose to predict semantic masks internally, but the responsibility of converting these into instances would lie with them, for example via connected components.</p>\n<p>Given that sheet thickness is arbitrary and sheets may appear detached from surrounding material, Surface Dice could be dropped. The evaluation could instead focus on VOI and TopoScore. It may also be worth investigating whether ASSD, computed either on binarized instance maps or per instance, could serve as a more appropriate replacement for Surface Dice.</p>\n<h4>3. Reformulate as instance segmentation while keeping the current metrics (least intrusive)</h4>\n<p>This would be the minimal change. Provide instance maps as ground truth and expect instance predictions, but keep the current metric suite unchanged.</p>\n<hr>\n<p>I would be very interested to hear what other participants think about these points, and especially how the organizers view the trade-offs involved. Any feedback, counterarguments, or alternative ideas would be very welcome.</p>",
      "rawMarkdown": "Hi all,\n\nI wanted to share a few thoughts regarding the overall challenge design. I believe there are some aspects that could be improved. Since we are still relatively early in the competition, it seems like a good moment to discuss this and, if possible, adjust certain design choices so the challenge is better aligned with the downstream task.\n\nI think this should also be in the organizers’ interest. A better task formulation has the potential to significantly improve the final outcome when it comes to scroll unrolling, which is ultimately what this competition is about.\n\nBelow are some thoughts that came up while working with the data and thinking about the evaluation setup. They are not fully polished ideas, but I hope they can serve as a starting point for discussion.\n\n(tagging @giorgioangelotti :-) )\n\n---\n\n### 1. Formulation as semantic segmentation unnecessarily biases the solution space\n\nI understand that the organizers have run preliminary experiments and that, in their internal pipelines, framing the problem as a 3D segmentation task has yielded promising results. However, this formulation comes with inherent issues.\n\nExamples include touching sheets, which can also occur in the ground truth (see hengck23’s comment: https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/651532#3370748), inconsistent and seemingly arbitrary sheet thicknesses, and sheets that appear to float in empty space (see https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/651532#3369783).\n\nThe organizers have confirmed that annotation is performed using meshes that are manually adjusted to follow the sheets as well as possible. Meshes are a very natural representation for this task. They provide explicit sheet instances, and their exact location does not need to be voxel-perfect as long as it is good enough for virtual unwrapping.\n\nGiven this, I find it somewhat confusing that only voxelized semantic segmentation masks are provided. By constraining participants to a 3D segmentation formulation, or requiring them to derive better representations such as instance maps or sheet surfaces from this unfitting formulation, the challenge may be unnecessarily limiting the space of possible solutions. This conversion step is inherently error prone and also wasteful, since instance maps or meshes could be provided directly. Participants might be able to explore more creative and potentially more effective approaches if they were not forced into a segmentation-only paradigm.\n\n---\n\n### 2. The use of Surface Dice as a metric\n\nI want to emphasize that I generally like Surface Dice as a metric. However, in this particular setting, its use is difficult to justify.\n\nThe ground-truth sheets exhibit varying thicknesses that do not seem to follow a deterministic or physically meaningful pattern. My impression is that the mesh-to-segmentation conversion script thins sheets in regions where they would otherwise touch. This is a sensible engineering decision, but it directly contradicts the framing of the task as a semantic segmentation problem.\n\nAs a participant, it becomes unclear what sheet thickness one should aim for at a given location. A tolerance of two voxels does not seem sufficiently forgiving to absorb this ambiguity. Additionally, since some ground-truth sheets appear to float in the background, it may be fundamentally impossible to reliably hit the tolerance threshold in those regions, regardless of how reasonable the predicted surface is from an unwrapping perspective.\n\n---\n\n### Possible directions for improvement\n\nBelow are several possible changes, ordered roughly from most radical to least intrusive. Some of these options are mutually exclusive. These are meant as ideas rather than concrete proposals.\n\n#### 1. Drop everything related to segmentation entirely (most radical)\n\nThis is fundamentally a mesh generation problem, not a segmentation problem. Provide meshes as ground truth and evaluate predicted meshes directly. Meshes should be provided as instances, one mesh per sheet.\n\nMetrics could be based on mesh pairing combined with average symmetric mesh distance, along with VOI-like and TopoScore-inspired measures adapted to mesh representations. This would require careful design, but it seems feasible.\n\nParticipants could still choose to solve the problem via voxel segmentation if they wish. Organizers could provide convenience utilities for converting between meshes and voxel masks in both directions.\n\n#### 2. Reformulate as instance segmentation and drop thickness-sensitive metrics (balanced)\n\nSince meshes are already available as annotations, it should be possible to provide instance maps rather than a single semantic foreground mask. Participants would be expected to submit instance segmentations.\n\nThis would eliminate connected-components ambiguities in both ground truth and predictions, which I believe is crucial. Participants could still choose to predict semantic masks internally, but the responsibility of converting these into instances would lie with them, for example via connected components.\n\nGiven that sheet thickness is arbitrary and sheets may appear detached from surrounding material, Surface Dice could be dropped. The evaluation could instead focus on VOI and TopoScore. It may also be worth investigating whether ASSD, computed either on binarized instance maps or per instance, could serve as a more appropriate replacement for Surface Dice.\n\n#### 3. Reformulate as instance segmentation while keeping the current metrics (least intrusive)\n\nThis would be the minimal change. Provide instance maps as ground truth and expect instance predictions, but keep the current metric suite unchanged.\n\n---\n\nI would be very interested to hear what other participants think about these points, and especially how the organizers view the trade-offs involved. Any feedback, counterarguments, or alternative ideas would be very welcome.",
      "votes": null
    },
    {
      "id": "3373305",
      "postDate": "12/12/2025 12:04:09",
      "content": "<p>Initial a grid mesh=&gt;manipulate mesh using Inverse Composition Lucas-Kanade Method with GAN judgement =&gt; fitted object topology</p>\n<p>Need to do mask * vol first if host does not provide more information.</p>",
      "rawMarkdown": "Initial a grid mesh=>manipulate mesh using Inverse Composition Lucas-Kanade Method with GAN judgement => fitted object topology\n\nNeed to do mask * vol first if host does not provide more information.",
      "votes": null
    },
    {
      "id": "3373367",
      "postDate": "12/12/2025 12:36:02",
      "content": "<p>Sir, you have me confused. Could you unpack that a bit?</p>",
      "rawMarkdown": "Sir, you have me confused. Could you unpack that a bit?",
      "votes": null
    },
    {
      "id": "3374231",
      "postDate": "12/12/2025 17:07:10",
      "content": "<p>Thanks a lot Fabian, this is thoughtful feedback, and it’s very helpful to have you (and others) stress‑testing the task design against the real downstream goal.</p>\n<p>You’re absolutely right that the end product of our unwrapping pipeline is a surface mesh. Our “Virtual Unwrapping” pipeline is essentially: (1) obtain an intermediate surface representation (currently a binary voxel mask), then (2) turn it into a mesh and flatten/render it.</p>\n<p>For Kaggle, we chose to frame the challenge on voxel mask predictions mainly because it keeps the barrier to entry low and fits the platform constraints (submission format, evaluation robustness/runtime, etc.). A mesh‑submission task would require instance handling + mesh correspondence/pairing and a fairly careful evaluation design (check for manifoldness, self-intersections, intersections across meshes), which is doable but substantially more complex to deploy and to make stable. We are currently <a href=\"https://github.com/ScrollPrize/villa/tree/neural-tracing/vesuvius/src/vesuvius/neural_tracing\" target=\"_blank\">researching on mesh neural tracing</a> approaches, but the formulation is not straightforward! </p>\n<p>That said, we agree with your core point: \"binary\" voxelization introduces style/thickness ambiguity, especially where sheets come close or touch. More broadly, while the underlying annotated surface (plus some diversion in air that is deemed acceptable for the unwrapping) is the real target, the voxelized representation is not uniquely determined by the data (local thickness choices, tie‑breaking near contacts), so we want the score to primarily reflect the underlying surface geometry/topology rather than quirks of the mesh→voxelization pipeline. In practice, our voxelization/cleanup does include heuristics in near‑contact regions (including thinning/separation choices) to avoid gross mergers; this is useful downstream, but it does increase thickness/style ambiguity in the masks.</p>\n<p>We acknowledge there are edge cases where the current tolerance in Surface Dice may be too tight relative to label uncertainty. Rather than completely changing the current metric suite, we are open to consider loosening the tolerance (and we recognize metric changes are sensitive, so we’d want clear community agreement and consistent application). Is there a specific tolerance range that you, or anyone in the community, think would better suit the task? We’re also open to suggestions for more thickness‑invariant surface evaluation (e.g., ASSD), as long as they remain stable and feasible for Kaggle scoring.</p>\n<p>I think that the bias introduced by choosing a semantic segmentation formulation does not necessarily block the utilization of different approaches, like instance segmentation methods, provided that the ground truth (and eventually the predictions) are accurate enough to allow for a change. In particular, if sheets are cleanly separated, semantic→instances is often just connected components (with the usual caveats where a sheet enters/exits the cropped ROI). </p>\n<p>This is why on the data/label side, prompted by community reports and our own quality controls, we’ve identified a few recurring artifact classes in a non-trivial subset of the labels. Mainly (1) unintended mergers between nearby sheets, (2) small spurious “dust” components that should have been removed, and (3) small holes/gaps that can be introduced by the mesh→voxelization step. Because the evaluation blends VOI + Surface Dice + a Betti-matching topological score, label quality matters a lot here, so we’re treating this as high priority and we intend to release soon an updated version of the dataset. <a href=\"https://www.kaggle.com/seanjohnsonsp\" target=\"_blank\">@seanjohnsonsp</a> and the team already spent the last 10 hours on it!</p>\n<p>Importantly, this is not meant to discourage anyone: many of the strongest Kaggle approaches we’ve seen already use exactly the kinds of robustness + topology-aware processing that are aligned with the downstream unwrapping pipeline. Those ideas remain correct, and improving label artifacts should make training easier, not invalidate progress and will surely help us to read the scrolls!</p>",
      "rawMarkdown": "Thanks a lot Fabian, this is thoughtful feedback, and it’s very helpful to have you (and others) stress‑testing the task design against the real downstream goal.\n\nYou’re absolutely right that the end product of our unwrapping pipeline is a surface mesh. Our “Virtual Unwrapping” pipeline is essentially: (1) obtain an intermediate surface representation (currently a binary voxel mask), then (2) turn it into a mesh and flatten/render it.\n\nFor Kaggle, we chose to frame the challenge on voxel mask predictions mainly because it keeps the barrier to entry low and fits the platform constraints (submission format, evaluation robustness/runtime, etc.). A mesh‑submission task would require instance handling + mesh correspondence/pairing and a fairly careful evaluation design (check for manifoldness, self-intersections, intersections across meshes), which is doable but substantially more complex to deploy and to make stable. We are currently [researching on mesh neural tracing](https://github.com/ScrollPrize/villa/tree/neural-tracing/vesuvius/src/vesuvius/neural_tracing) approaches, but the formulation is not straightforward! \n\nThat said, we agree with your core point: \"binary\" voxelization introduces style/thickness ambiguity, especially where sheets come close or touch. More broadly, while the underlying annotated surface (plus some diversion in air that is deemed acceptable for the unwrapping) is the real target, the voxelized representation is not uniquely determined by the data (local thickness choices, tie‑breaking near contacts), so we want the score to primarily reflect the underlying surface geometry/topology rather than quirks of the mesh→voxelization pipeline. In practice, our voxelization/cleanup does include heuristics in near‑contact regions (including thinning/separation choices) to avoid gross mergers; this is useful downstream, but it does increase thickness/style ambiguity in the masks.\n\nWe acknowledge there are edge cases where the current tolerance in Surface Dice may be too tight relative to label uncertainty. Rather than completely changing the current metric suite, we are open to consider loosening the tolerance (and we recognize metric changes are sensitive, so we’d want clear community agreement and consistent application). Is there a specific tolerance range that you, or anyone in the community, think would better suit the task? We’re also open to suggestions for more thickness‑invariant surface evaluation (e.g., ASSD), as long as they remain stable and feasible for Kaggle scoring.\n\nI think that the bias introduced by choosing a semantic segmentation formulation does not necessarily block the utilization of different approaches, like instance segmentation methods, provided that the ground truth (and eventually the predictions) are accurate enough to allow for a change. In particular, if sheets are cleanly separated, semantic→instances is often just connected components (with the usual caveats where a sheet enters/exits the cropped ROI). \n\nThis is why on the data/label side, prompted by community reports and our own quality controls, we’ve identified a few recurring artifact classes in a non-trivial subset of the labels. Mainly (1) unintended mergers between nearby sheets, (2) small spurious “dust” components that should have been removed, and (3) small holes/gaps that can be introduced by the mesh→voxelization step. Because the evaluation blends VOI + Surface Dice + a Betti-matching topological score, label quality matters a lot here, so we’re treating this as high priority and we intend to release soon an updated version of the dataset. @seanjohnsonsp and the team already spent the last 10 hours on it!\n\nImportantly, this is not meant to discourage anyone: many of the strongest Kaggle approaches we’ve seen already use exactly the kinds of robustness + topology-aware processing that are aligned with the downstream unwrapping pipeline. Those ideas remain correct, and improving label artifacts should make training easier, not invalidate progress and will surely help us to read the scrolls!",
      "votes": null
    },
    {
      "id": "3375074",
      "postDate": "12/12/2025 20:39:30",
      "content": "<p>Thanks for the quick and detailed response. It is great to hear that you are treating GT quality as a high-priority item, and I really appreciate that you are actively revising the problematic cases and plan to release an updated dataset soon.</p>\n<p>I understand why you chose a voxel-mask submission format given Kaggle constraints and the need for a stable evaluation pipeline. I probably agree a bit less on the approachability argument, though. From my perspective, approachability and flexibility can coexist if you provide reliable reference conversions between representations (mesh ↔ instance mask ↔ semantic mask), while keeping the submission format simple.</p>\n<p>The point I feel most strongly about is the instance question. Even if GT is cleaned up such that connected components yields reasonable instances, the current pipeline still implicitly requires participants to produce a semantic mask that (a) stays close to the GT voxelization style and thickness for Surface Dice, and (b) can be separated into instances via connected components without accidental connections. For teams that naturally model sheets as instances or surfaces, this introduces a round trip that can be both error-prone and wasteful: instance or surface representation → semantic mask for submission → instances again via connected components on the evaluation side.</p>\n<p>In particular, avoiding inadvertent connections in the semantic mask becomes an extra constraint that is not really part of the downstream goal, but can heavily affect VOI and TopoScore, and which would largely disappear if instance- or surface-based representations were accepted directly.</p>\n<p>Because of this, I think it would be a meaningful improvement to define instances as the canonical representation: provide instance GT, expect instance submissions, and provide a small reference snippet showing how to convert a semantic mask to instances via connected components (cc3d is very fast) for participants who prefer to stay in semantic segmentation. This would reduce duplicated effort and make the evaluation semantics more explicit.</p>\n<p>On Surface Dice, I would not feel comfortable recommending a new tolerance value in the current setup. Increasing τ while computing Surface Dice on the global semantic foreground can introduce unintended behavior. For example, when one sheet is missed but lies close to another sheet, a high tolerance can still give partial credit simply due to proximity. More generally, thickness and style ambiguity in voxelization makes Surface Dice sensitive in ways that are not necessarily aligned with the underlying surface target.</p>\n<p>One possible way around this, without fully changing the metric suite, could be to compute Surface Dice at the instance level rather than on the global semantic mask. In rough terms:</p>\n<ul>\n<li><p>Generate instance maps for GT and predictions (or ideally make GT an instance map and expect teams to submit instance maps as well).</p></li>\n<li><p>Match GT and predicted instances, for example via pairwise IoU with greedy or Hungarian matching.</p></li>\n<li><p>For each matched pair, compute Surface Dice with a more permissive tolerance (for example τ ≈ 5), since neighboring sheets are no longer part of the comparison.</p></li>\n<li><p>Assign Surface Dice = 0 for unmatched GT instances (false negatives) and unmatched predicted instances (false positives).</p></li>\n<li><p>Aggregate per volume (for example average across instances), then average over volumes, so volumes with many sheets do not dominate.</p></li>\n</ul>\n<p>A nice side effect is that instance-based Surface Dice with a generous tolerance would largely remove the need to match an arbitrary voxel thickness. This would likely make it easier for creative approaches like <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> 's line completion ideas or the crazy methods <a href=\"https://www.kaggle.com/tom99763\" target=\"_blank\">@tom99763</a> is pursuing, without forcing everyone to tune postprocessing to a particular voxelization style.</p>\n<p>Overall, if you were willing to move to instance ground truth and instance submissions, provide conversion code for those who want to stay in semantic segmentation, and compute Surface Dice in an instance-aware manner as outlined above (while keeping VOI and TopoScore), I think this would lead to a setup that remains approachable while aligning more directly with the downstream unwrapping objective. From my perspective, this would be a relatively small change on the evaluation side, but of course one that the community would need to be comfortable with.</p>\n<p>I would be very interested to hear not only what you think about this direction, but also what other participants think.</p>",
      "rawMarkdown": "Thanks for the quick and detailed response. It is great to hear that you are treating GT quality as a high-priority item, and I really appreciate that you are actively revising the problematic cases and plan to release an updated dataset soon.\n\nI understand why you chose a voxel-mask submission format given Kaggle constraints and the need for a stable evaluation pipeline. I probably agree a bit less on the approachability argument, though. From my perspective, approachability and flexibility can coexist if you provide reliable reference conversions between representations (mesh ↔ instance mask ↔ semantic mask), while keeping the submission format simple.\n\nThe point I feel most strongly about is the instance question. Even if GT is cleaned up such that connected components yields reasonable instances, the current pipeline still implicitly requires participants to produce a semantic mask that (a) stays close to the GT voxelization style and thickness for Surface Dice, and (b) can be separated into instances via connected components without accidental connections. For teams that naturally model sheets as instances or surfaces, this introduces a round trip that can be both error-prone and wasteful: instance or surface representation → semantic mask for submission → instances again via connected components on the evaluation side.\n\nIn particular, avoiding inadvertent connections in the semantic mask becomes an extra constraint that is not really part of the downstream goal, but can heavily affect VOI and TopoScore, and which would largely disappear if instance- or surface-based representations were accepted directly.\n\nBecause of this, I think it would be a meaningful improvement to define instances as the canonical representation: provide instance GT, expect instance submissions, and provide a small reference snippet showing how to convert a semantic mask to instances via connected components (cc3d is very fast) for participants who prefer to stay in semantic segmentation. This would reduce duplicated effort and make the evaluation semantics more explicit.\n\nOn Surface Dice, I would not feel comfortable recommending a new tolerance value in the current setup. Increasing τ while computing Surface Dice on the global semantic foreground can introduce unintended behavior. For example, when one sheet is missed but lies close to another sheet, a high tolerance can still give partial credit simply due to proximity. More generally, thickness and style ambiguity in voxelization makes Surface Dice sensitive in ways that are not necessarily aligned with the underlying surface target.\n\nOne possible way around this, without fully changing the metric suite, could be to compute Surface Dice at the instance level rather than on the global semantic mask. In rough terms:\n\n- Generate instance maps for GT and predictions (or ideally make GT an instance map and expect teams to submit instance maps as well).\n\n- Match GT and predicted instances, for example via pairwise IoU with greedy or Hungarian matching.\n\n- For each matched pair, compute Surface Dice with a more permissive tolerance (for example τ ≈ 5), since neighboring sheets are no longer part of the comparison.\n\n- Assign Surface Dice = 0 for unmatched GT instances (false negatives) and unmatched predicted instances (false positives).\n\n- Aggregate per volume (for example average across instances), then average over volumes, so volumes with many sheets do not dominate.\n\nA nice side effect is that instance-based Surface Dice with a generous tolerance would largely remove the need to match an arbitrary voxel thickness. This would likely make it easier for creative approaches like @hengck23 's line completion ideas or the crazy methods @tom99763 is pursuing, without forcing everyone to tune postprocessing to a particular voxelization style.\n\nOverall, if you were willing to move to instance ground truth and instance submissions, provide conversion code for those who want to stay in semantic segmentation, and compute Surface Dice in an instance-aware manner as outlined above (while keeping VOI and TopoScore), I think this would lead to a setup that remains approachable while aligning more directly with the downstream unwrapping objective. From my perspective, this would be a relatively small change on the evaluation side, but of course one that the community would need to be comfortable with.\n\nI would be very interested to hear not only what you think about this direction, but also what other participants think.",
      "votes": null
    },
    {
      "id": "3375799",
      "postDate": "12/13/2025 06:12:00",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/fabianisensee\" target=\"_blank\">@fabianisensee</a> , I believe your perspective is very insightful, and the organizer’s <a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> reply is also well-reasoned. Therefore, I would like to discuss this issue from the standpoint of <strong>competition fairness</strong>:</p>\n<ol>\n<li><p>On Kaggle, unless there are serious flaws in the evaluation metric, <strong>it should not be changed</strong> during the competition. In my view, changing the metric is essentially equivalent to starting a new competition.</p></li>\n<li><p>Changing the metric affects various participants differently. Some have limited computational resources, while others may lack sufficient time in the later stages. In any case, this can easily <strong>undermine the efforts of participants who worked hard before the change,</strong> reducing their motivation and potentially causing the competition to miss out on many excellent solutions.</p></li>\n<li><p>Considering that the organizer intends to <strong>make changes to the dataset</strong>, <strong>this already requires additional effort from participants</strong>. Changing the metric on top of that would only add to the workload for many, making an already challenging situation even more difficult.</p></li>\n</ol>\n<p>Thank you for your time, and I hope the organizer achieves the solutions they are looking for.</p>",
      "rawMarkdown": "Hi @fabianisensee , I believe your perspective is very insightful, and the organizer’s @giorgioangelotti reply is also well-reasoned. Therefore, I would like to discuss this issue from the standpoint of **competition fairness**:\n\n1. On Kaggle, unless there are serious flaws in the evaluation metric, **it should not be changed** during the competition. In my view, changing the metric is essentially equivalent to starting a new competition.\n\n2. Changing the metric affects various participants differently. Some have limited computational resources, while others may lack sufficient time in the later stages. In any case, this can easily **undermine the efforts of participants who worked hard before the change,** reducing their motivation and potentially causing the competition to miss out on many excellent solutions.\n\n3. Considering that the organizer intends to **make changes to the dataset**, **this already requires additional effort from participants**. Changing the metric on top of that would only add to the workload for many, making an already challenging situation even more difficult.\n\nThank you for your time, and I hope the organizer achieves the solutions they are looking for.",
      "votes": null
    },
    {
      "id": "3375843",
      "postDate": "12/13/2025 08:14:08",
      "content": "<p>I am not 100% sure, but this could be the mesh to voxel conversion code <a href=\"https://github.com/ScrollPrize/villa/blob/3de44b8d70e49886295b310e54f1f9962ab3fe5c/vesuvius/src/vesuvius/models/datasets/mesh/voxelize.py\" target=\"_blank\">https://github.com/ScrollPrize/villa/blob/3de44b8d70e49886295b310e54f1f9962ab3fe5c/vesuvius/src/vesuvius/models/datasets/mesh/voxelize.py</a></p>\n<p>You can learn a network to do the reverse </p>",
      "rawMarkdown": "I am not 100% sure, but this could be the mesh to voxel conversion code https://github.com/ScrollPrize/villa/blob/3de44b8d70e49886295b310e54f1f9962ab3fe5c/vesuvius/src/vesuvius/models/datasets/mesh/voxelize.py\n\nYou can learn a network to do the reverse",
      "votes": null
    },
    {
      "id": "3375848",
      "postDate": "12/13/2025 08:25:06",
      "content": "<p><a href=\"https://www.kaggle.com/seanjohnsonsp\" target=\"_blank\">@seanjohnsonsp</a> annotated the meshes, voxelized them and then refined them manually. I don't think he used that code, or only that code.</p>",
      "rawMarkdown": "seanjohnsonsp annotated the meshes, voxelized them and then refined them manually. I don't think he used that code, or only that code.",
      "votes": null
    },
    {
      "id": "3376817",
      "postDate": "12/15/2025 07:08:16",
      "content": "<p>Thanks for raising this perspective. I fully agree that major metric changes during a competition should be avoided, as they can raise fairness concerns and undermine prior effort. If any changes were to be made, they should clearly have broad community support and, of course, would ultimately be at the discretion of the organizers.</p>\n<p>I would like to clarify that my most recent suggestion is not intended as a major metric redesign. Switching from a semantic foreground mask to an explicit instance representation for GT and submission, and computing Surface Dice on instances rather than on the global semantic mask, would in my view be a relatively minor adjustment that also elegantly addresses the arbitrary thickness issue, provided a permissive tolerance is used. VOI and TopoScore are already instance-based in practice, since connected components are run internally on submissions, so this would mainly make the instance notion explicit and consistent across all metrics, while at the same time providing a much richer and more task-aligned representation for participants to work with and explore creative ideas.</p>\n<p>Existing approaches would continue to work with minimal changes, and given that we are still about two months out, I personally would not see a need for a competition extension if such a change were agreed upon. That said, I fully agree that fairness comes first, and any decision here should be guided by community consensus and the organizers’ judgment.</p>",
      "rawMarkdown": "Thanks for raising this perspective. I fully agree that major metric changes during a competition should be avoided, as they can raise fairness concerns and undermine prior effort. If any changes were to be made, they should clearly have broad community support and, of course, would ultimately be at the discretion of the organizers.\n\nI would like to clarify that my most recent suggestion is not intended as a major metric redesign. Switching from a semantic foreground mask to an explicit instance representation for GT and submission, and computing Surface Dice on instances rather than on the global semantic mask, would in my view be a relatively minor adjustment that also elegantly addresses the arbitrary thickness issue, provided a permissive tolerance is used. VOI and TopoScore are already instance-based in practice, since connected components are run internally on submissions, so this would mainly make the instance notion explicit and consistent across all metrics, while at the same time providing a much richer and more task-aligned representation for participants to work with and explore creative ideas.\n\nExisting approaches would continue to work with minimal changes, and given that we are still about two months out, I personally would not see a need for a competition extension if such a change were agreed upon. That said, I fully agree that fairness comes first, and any decision here should be guided by community consensus and the organizers’ judgment.",
      "votes": null
    },
    {
      "id": "3377112",
      "postDate": "12/15/2025 18:24:26",
      "content": "<p>With the current metric, what (private) score would you consider/approximate a pipeline to be successful?\nSuccessful as in your target for this competition or capability to move to the next stage of decoding the scroll contents.</p>",
      "rawMarkdown": "With the current metric, what (private) score would you consider/approximate a pipeline to be successful?\nSuccessful as in your target for this competition or capability to move to the next stage of decoding the scroll contents.",
      "votes": null
    },
    {
      "id": "3377141",
      "postDate": "12/15/2025 20:13:09",
      "content": "<p>I would say that with a LB &gt; 0.7 we would be in a very good position. Our current pipeline is basically a \"mesher\" on top of these predictions. The pipe is already capable of automatically solving some \"easy\" mistakes, but it really struggles in areas where the sheets are densely packed. Those mistakes are usually fixed with manual intervention, but this requires a lot of effort. This is why creating these labels is no easy.</p>",
      "rawMarkdown": "I would say that with a LB > 0.7 we would be in a very good position. Our current pipeline is basically a \"mesher\" on top of these predictions. The pipe is already capable of automatically solving some \"easy\" mistakes, but it really struggles in areas where the sheets are densely packed. Those mistakes are usually fixed with manual intervention, but this requires a lot of effort. This is why creating these labels is no easy.",
      "votes": null
    },
    {
      "id": "3377988",
      "postDate": "12/17/2025 09:52:32",
      "content": "<p>In fact I would agree with all three points above but as also mentioned, in such case, it shall be bright new competition then pivoting data and metrics (all key aspects defining the competition) in middle of the race… </p>",
      "rawMarkdown": "In fact I would agree with all three points above but as also mentioned, in such case, it shall be bright new competition then pivoting data and metrics (all key aspects defining the competition) in middle of the race...",
      "votes": null
    },
    {
      "id": "3377989",
      "postDate": "12/17/2025 09:58:36",
      "content": "<blockquote>\n  <p>instances rather than on the global semantic mask</p>\n</blockquote>\n<p>That is major task shift and completely redefine the competition, although I agree that it shall be instance segmentation, this would disqualify all previous work and invested compute, as for example me with limited communication, noone reset 200+hours of compute I already used…</p>\n<p>Cc <a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a></p>",
      "rawMarkdown": "> instances rather than on the global semantic mask\n\nThat is major task shift and completely redefine the competition, although I agree that it shall be instance segmentation, this would disqualify all previous work and invested compute, as for example me with limited communication, noone reset 200+hours of compute I already used...\n\nCc @giorgioangelotti",
      "votes": null
    },
    {
      "id": "3377990",
      "postDate": "12/17/2025 10:03:53",
      "content": "<p>There is not such submission yet… The best one is 0.58 likely segmentation model 🤔</p>",
      "rawMarkdown": "There is not such submission yet... The best one is 0.58 likely segmentation model 🤔",
      "votes": null
    },
    {
      "id": "3378447",
      "postDate": "12/18/2025 08:14:38",
      "content": "<p>please start your performance(mislead) !</p>",
      "rawMarkdown": "please start your performance(mislead) !",
      "votes": null
    },
    {
      "id": "3379490",
      "postDate": "12/20/2025 00:27:36",
      "content": "<p><a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> \nThank you for information. I learn a lot of things in this competition.<br>\nI have a question. When will you change dataset?<br>\nyou said train/test label annotation will be changed. We need to retrain dataset. this competition is 3d volume, huge calculation cost.<br>\ncompetitor need enough time.  </p>",
      "rawMarkdown": "giorgioangelotti \nThank you for information. I learn a lot of things in this competition.  \nI have a question. When will you change dataset?  \nyou said train/test label annotation will be changed. We need to retrain dataset. this competition is 3d volume, huge calculation cost.  \ncompetitor need enough time.",
      "votes": null
    },
    {
      "id": "3380848",
      "postDate": "12/23/2025 07:24:02",
      "content": "<p>It's now live</p>",
      "rawMarkdown": "It's now live",
      "votes": null
    },
    {
      "id": "3381193",
      "postDate": "12/24/2025 02:01:57",
      "content": "<p>Thank you for announcement. I'm checking now.</p>",
      "rawMarkdown": "Thank you for announcement. I'm checking now.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3373305,
      "author_name": "tom99763",
      "author_url": "",
      "post_date": "12/12/2025 12:04:09",
      "content": "<p>Initial a grid mesh=&gt;manipulate mesh using Inverse Composition Lucas-Kanade Method with GAN judgement =&gt; fitted object topology</p>\n<p>Need to do mask * vol first if host does not provide more information.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3373367,
          "author_name": "fabianisensee",
          "author_url": "",
          "post_date": "12/12/2025 12:36:02",
          "content": "<p>Sir, you have me confused. Could you unpack that a bit?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3378447,
          "author_name": "clwwlc",
          "author_url": "",
          "post_date": "12/18/2025 08:14:38",
          "content": "<p>please start your performance(mislead) !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3374231,
      "author_name": "giorgioangelotti",
      "author_url": "",
      "post_date": "12/12/2025 17:07:10",
      "content": "<p>Thanks a lot Fabian, this is thoughtful feedback, and it’s very helpful to have you (and others) stress‑testing the task design against the real downstream goal.</p>\n<p>You’re absolutely right that the end product of our unwrapping pipeline is a surface mesh. Our “Virtual Unwrapping” pipeline is essentially: (1) obtain an intermediate surface representation (currently a binary voxel mask), then (2) turn it into a mesh and flatten/render it.</p>\n<p>For Kaggle, we chose to frame the challenge on voxel mask predictions mainly because it keeps the barrier to entry low and fits the platform constraints (submission format, evaluation robustness/runtime, etc.). A mesh‑submission task would require instance handling + mesh correspondence/pairing and a fairly careful evaluation design (check for manifoldness, self-intersections, intersections across meshes), which is doable but substantially more complex to deploy and to make stable. We are currently <a href=\"https://github.com/ScrollPrize/villa/tree/neural-tracing/vesuvius/src/vesuvius/neural_tracing\" target=\"_blank\">researching on mesh neural tracing</a> approaches, but the formulation is not straightforward! </p>\n<p>That said, we agree with your core point: \"binary\" voxelization introduces style/thickness ambiguity, especially where sheets come close or touch. More broadly, while the underlying annotated surface (plus some diversion in air that is deemed acceptable for the unwrapping) is the real target, the voxelized representation is not uniquely determined by the data (local thickness choices, tie‑breaking near contacts), so we want the score to primarily reflect the underlying surface geometry/topology rather than quirks of the mesh→voxelization pipeline. In practice, our voxelization/cleanup does include heuristics in near‑contact regions (including thinning/separation choices) to avoid gross mergers; this is useful downstream, but it does increase thickness/style ambiguity in the masks.</p>\n<p>We acknowledge there are edge cases where the current tolerance in Surface Dice may be too tight relative to label uncertainty. Rather than completely changing the current metric suite, we are open to consider loosening the tolerance (and we recognize metric changes are sensitive, so we’d want clear community agreement and consistent application). Is there a specific tolerance range that you, or anyone in the community, think would better suit the task? We’re also open to suggestions for more thickness‑invariant surface evaluation (e.g., ASSD), as long as they remain stable and feasible for Kaggle scoring.</p>\n<p>I think that the bias introduced by choosing a semantic segmentation formulation does not necessarily block the utilization of different approaches, like instance segmentation methods, provided that the ground truth (and eventually the predictions) are accurate enough to allow for a change. In particular, if sheets are cleanly separated, semantic→instances is often just connected components (with the usual caveats where a sheet enters/exits the cropped ROI). </p>\n<p>This is why on the data/label side, prompted by community reports and our own quality controls, we’ve identified a few recurring artifact classes in a non-trivial subset of the labels. Mainly (1) unintended mergers between nearby sheets, (2) small spurious “dust” components that should have been removed, and (3) small holes/gaps that can be introduced by the mesh→voxelization step. Because the evaluation blends VOI + Surface Dice + a Betti-matching topological score, label quality matters a lot here, so we’re treating this as high priority and we intend to release soon an updated version of the dataset. <a href=\"https://www.kaggle.com/seanjohnsonsp\" target=\"_blank\">@seanjohnsonsp</a> and the team already spent the last 10 hours on it!</p>\n<p>Importantly, this is not meant to discourage anyone: many of the strongest Kaggle approaches we’ve seen already use exactly the kinds of robustness + topology-aware processing that are aligned with the downstream unwrapping pipeline. Those ideas remain correct, and improving label artifacts should make training easier, not invalidate progress and will surely help us to read the scrolls!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3375074,
          "author_name": "fabianisensee",
          "author_url": "",
          "post_date": "12/12/2025 20:39:30",
          "content": "<p>Thanks for the quick and detailed response. It is great to hear that you are treating GT quality as a high-priority item, and I really appreciate that you are actively revising the problematic cases and plan to release an updated dataset soon.</p>\n<p>I understand why you chose a voxel-mask submission format given Kaggle constraints and the need for a stable evaluation pipeline. I probably agree a bit less on the approachability argument, though. From my perspective, approachability and flexibility can coexist if you provide reliable reference conversions between representations (mesh ↔ instance mask ↔ semantic mask), while keeping the submission format simple.</p>\n<p>The point I feel most strongly about is the instance question. Even if GT is cleaned up such that connected components yields reasonable instances, the current pipeline still implicitly requires participants to produce a semantic mask that (a) stays close to the GT voxelization style and thickness for Surface Dice, and (b) can be separated into instances via connected components without accidental connections. For teams that naturally model sheets as instances or surfaces, this introduces a round trip that can be both error-prone and wasteful: instance or surface representation → semantic mask for submission → instances again via connected components on the evaluation side.</p>\n<p>In particular, avoiding inadvertent connections in the semantic mask becomes an extra constraint that is not really part of the downstream goal, but can heavily affect VOI and TopoScore, and which would largely disappear if instance- or surface-based representations were accepted directly.</p>\n<p>Because of this, I think it would be a meaningful improvement to define instances as the canonical representation: provide instance GT, expect instance submissions, and provide a small reference snippet showing how to convert a semantic mask to instances via connected components (cc3d is very fast) for participants who prefer to stay in semantic segmentation. This would reduce duplicated effort and make the evaluation semantics more explicit.</p>\n<p>On Surface Dice, I would not feel comfortable recommending a new tolerance value in the current setup. Increasing τ while computing Surface Dice on the global semantic foreground can introduce unintended behavior. For example, when one sheet is missed but lies close to another sheet, a high tolerance can still give partial credit simply due to proximity. More generally, thickness and style ambiguity in voxelization makes Surface Dice sensitive in ways that are not necessarily aligned with the underlying surface target.</p>\n<p>One possible way around this, without fully changing the metric suite, could be to compute Surface Dice at the instance level rather than on the global semantic mask. In rough terms:</p>\n<ul>\n<li><p>Generate instance maps for GT and predictions (or ideally make GT an instance map and expect teams to submit instance maps as well).</p></li>\n<li><p>Match GT and predicted instances, for example via pairwise IoU with greedy or Hungarian matching.</p></li>\n<li><p>For each matched pair, compute Surface Dice with a more permissive tolerance (for example τ ≈ 5), since neighboring sheets are no longer part of the comparison.</p></li>\n<li><p>Assign Surface Dice = 0 for unmatched GT instances (false negatives) and unmatched predicted instances (false positives).</p></li>\n<li><p>Aggregate per volume (for example average across instances), then average over volumes, so volumes with many sheets do not dominate.</p></li>\n</ul>\n<p>A nice side effect is that instance-based Surface Dice with a generous tolerance would largely remove the need to match an arbitrary voxel thickness. This would likely make it easier for creative approaches like <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> 's line completion ideas or the crazy methods <a href=\"https://www.kaggle.com/tom99763\" target=\"_blank\">@tom99763</a> is pursuing, without forcing everyone to tune postprocessing to a particular voxelization style.</p>\n<p>Overall, if you were willing to move to instance ground truth and instance submissions, provide conversion code for those who want to stay in semantic segmentation, and compute Surface Dice in an instance-aware manner as outlined above (while keeping VOI and TopoScore), I think this would lead to a setup that remains approachable while aligning more directly with the downstream unwrapping objective. From my perspective, this would be a relatively small change on the evaluation side, but of course one that the community would need to be comfortable with.</p>\n<p>I would be very interested to hear not only what you think about this direction, but also what other participants think.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3377112,
          "author_name": "sroger",
          "author_url": "",
          "post_date": "12/15/2025 18:24:26",
          "content": "<p>With the current metric, what (private) score would you consider/approximate a pipeline to be successful?\nSuccessful as in your target for this competition or capability to move to the next stage of decoding the scroll contents.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3377141,
              "author_name": "giorgioangelotti",
              "author_url": "",
              "post_date": "12/15/2025 20:13:09",
              "content": "<p>I would say that with a LB &gt; 0.7 we would be in a very good position. Our current pipeline is basically a \"mesher\" on top of these predictions. The pipe is already capable of automatically solving some \"easy\" mistakes, but it really struggles in areas where the sheets are densely packed. Those mistakes are usually fixed with manual intervention, but this requires a lot of effort. This is why creating these labels is no easy.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3377990,
                  "author_name": "jirkaborovec",
                  "author_url": "",
                  "post_date": "12/17/2025 10:03:53",
                  "content": "<p>There is not such submission yet… The best one is 0.58 likely segmentation model 🤔</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        },
        {
          "id": 3379490,
          "author_name": "tereka",
          "author_url": "",
          "post_date": "12/20/2025 00:27:36",
          "content": "<p><a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> \nThank you for information. I learn a lot of things in this competition.<br>\nI have a question. When will you change dataset?<br>\nyou said train/test label annotation will be changed. We need to retrain dataset. this competition is 3d volume, huge calculation cost.<br>\ncompetitor need enough time.  </p>",
          "votes": null,
          "replies": [
            {
              "id": 3380848,
              "author_name": "giorgioangelotti",
              "author_url": "",
              "post_date": "12/23/2025 07:24:02",
              "content": "<p>It's now live</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3381193,
                  "author_name": "tereka",
                  "author_url": "",
                  "post_date": "12/24/2025 02:01:57",
                  "content": "<p>Thank you for announcement. I'm checking now.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3375799,
      "author_name": "forcewithme",
      "author_url": "",
      "post_date": "12/13/2025 06:12:00",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/fabianisensee\" target=\"_blank\">@fabianisensee</a> , I believe your perspective is very insightful, and the organizer’s <a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> reply is also well-reasoned. Therefore, I would like to discuss this issue from the standpoint of <strong>competition fairness</strong>:</p>\n<ol>\n<li><p>On Kaggle, unless there are serious flaws in the evaluation metric, <strong>it should not be changed</strong> during the competition. In my view, changing the metric is essentially equivalent to starting a new competition.</p></li>\n<li><p>Changing the metric affects various participants differently. Some have limited computational resources, while others may lack sufficient time in the later stages. In any case, this can easily <strong>undermine the efforts of participants who worked hard before the change,</strong> reducing their motivation and potentially causing the competition to miss out on many excellent solutions.</p></li>\n<li><p>Considering that the organizer intends to <strong>make changes to the dataset</strong>, <strong>this already requires additional effort from participants</strong>. Changing the metric on top of that would only add to the workload for many, making an already challenging situation even more difficult.</p></li>\n</ol>\n<p>Thank you for your time, and I hope the organizer achieves the solutions they are looking for.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3376817,
          "author_name": "fabianisensee",
          "author_url": "",
          "post_date": "12/15/2025 07:08:16",
          "content": "<p>Thanks for raising this perspective. I fully agree that major metric changes during a competition should be avoided, as they can raise fairness concerns and undermine prior effort. If any changes were to be made, they should clearly have broad community support and, of course, would ultimately be at the discretion of the organizers.</p>\n<p>I would like to clarify that my most recent suggestion is not intended as a major metric redesign. Switching from a semantic foreground mask to an explicit instance representation for GT and submission, and computing Surface Dice on instances rather than on the global semantic mask, would in my view be a relatively minor adjustment that also elegantly addresses the arbitrary thickness issue, provided a permissive tolerance is used. VOI and TopoScore are already instance-based in practice, since connected components are run internally on submissions, so this would mainly make the instance notion explicit and consistent across all metrics, while at the same time providing a much richer and more task-aligned representation for participants to work with and explore creative ideas.</p>\n<p>Existing approaches would continue to work with minimal changes, and given that we are still about two months out, I personally would not see a need for a competition extension if such a change were agreed upon. That said, I fully agree that fairness comes first, and any decision here should be guided by community consensus and the organizers’ judgment.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3377989,
              "author_name": "jirkaborovec",
              "author_url": "",
              "post_date": "12/17/2025 09:58:36",
              "content": "<blockquote>\n  <p>instances rather than on the global semantic mask</p>\n</blockquote>\n<p>That is major task shift and completely redefine the competition, although I agree that it shall be instance segmentation, this would disqualify all previous work and invested compute, as for example me with limited communication, noone reset 200+hours of compute I already used…</p>\n<p>Cc <a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a></p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3375843,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "12/13/2025 08:14:08",
      "content": "<p>I am not 100% sure, but this could be the mesh to voxel conversion code <a href=\"https://github.com/ScrollPrize/villa/blob/3de44b8d70e49886295b310e54f1f9962ab3fe5c/vesuvius/src/vesuvius/models/datasets/mesh/voxelize.py\" target=\"_blank\">https://github.com/ScrollPrize/villa/blob/3de44b8d70e49886295b310e54f1f9962ab3fe5c/vesuvius/src/vesuvius/models/datasets/mesh/voxelize.py</a></p>\n<p>You can learn a network to do the reverse </p>",
      "votes": null,
      "replies": [
        {
          "id": 3375848,
          "author_name": "giorgioangelotti",
          "author_url": "",
          "post_date": "12/13/2025 08:25:06",
          "content": "<p><a href=\"https://www.kaggle.com/seanjohnsonsp\" target=\"_blank\">@seanjohnsonsp</a> annotated the meshes, voxelized them and then refined them manually. I don't think he used that code, or only that code.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3377988,
      "author_name": "jirkaborovec",
      "author_url": "",
      "post_date": "12/17/2025 09:52:32",
      "content": "<p>In fact I would agree with all three points above but as also mentioned, in such case, it shall be bright new competition then pivoting data and metrics (all key aspects defining the competition) in middle of the race… </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3373174": "Hi all,\n\nI wanted to share a few thoughts regarding the overall challenge design. I believe there are some aspects that could be improved. Since we are still relatively early in the competition, it seems like a good moment to discuss this and, if possible, adjust certain design choices so the challenge is better aligned with the downstream task.\n\nI think this should also be in the organizers’ interest. A better task formulation has the potential to significantly improve the final outcome when it comes to scroll unrolling, which is ultimately what this competition is about.\n\nBelow are some thoughts that came up while working with the data and thinking about the evaluation setup. They are not fully polished ideas, but I hope they can serve as a starting point for discussion.\n\n(tagging @giorgioangelotti :-) )\n\n---\n\n### 1. Formulation as semantic segmentation unnecessarily biases the solution space\n\nI understand that the organizers have run preliminary experiments and that, in their internal pipelines, framing the problem as a 3D segmentation task has yielded promising results. However, this formulation comes with inherent issues.\n\nExamples include touching sheets, which can also occur in the ground truth (see hengck23’s comment: https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/651532#3370748), inconsistent and seemingly arbitrary sheet thicknesses, and sheets that appear to float in empty space (see https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/651532#3369783).\n\nThe organizers have confirmed that annotation is performed using meshes that are manually adjusted to follow the sheets as well as possible. Meshes are a very natural representation for this task. They provide explicit sheet instances, and their exact location does not need to be voxel-perfect as long as it is good enough for virtual unwrapping.\n\nGiven this, I find it somewhat confusing that only voxelized semantic segmentation masks are provided. By constraining participants to a 3D segmentation formulation, or requiring them to derive better representations such as instance maps or sheet surfaces from this unfitting formulation, the challenge may be unnecessarily limiting the space of possible solutions. This conversion step is inherently error prone and also wasteful, since instance maps or meshes could be provided directly. Participants might be able to explore more creative and potentially more effective approaches if they were not forced into a segmentation-only paradigm.\n\n---\n\n### 2. The use of Surface Dice as a metric\n\nI want to emphasize that I generally like Surface Dice as a metric. However, in this particular setting, its use is difficult to justify.\n\nThe ground-truth sheets exhibit varying thicknesses that do not seem to follow a deterministic or physically meaningful pattern. My impression is that the mesh-to-segmentation conversion script thins sheets in regions where they would otherwise touch. This is a sensible engineering decision, but it directly contradicts the framing of the task as a semantic segmentation problem.\n\nAs a participant, it becomes unclear what sheet thickness one should aim for at a given location. A tolerance of two voxels does not seem sufficiently forgiving to absorb this ambiguity. Additionally, since some ground-truth sheets appear to float in the background, it may be fundamentally impossible to reliably hit the tolerance threshold in those regions, regardless of how reasonable the predicted surface is from an unwrapping perspective.\n\n---\n\n### Possible directions for improvement\n\nBelow are several possible changes, ordered roughly from most radical to least intrusive. Some of these options are mutually exclusive. These are meant as ideas rather than concrete proposals.\n\n#### 1. Drop everything related to segmentation entirely (most radical)\n\nThis is fundamentally a mesh generation problem, not a segmentation problem. Provide meshes as ground truth and evaluate predicted meshes directly. Meshes should be provided as instances, one mesh per sheet.\n\nMetrics could be based on mesh pairing combined with average symmetric mesh distance, along with VOI-like and TopoScore-inspired measures adapted to mesh representations. This would require careful design, but it seems feasible.\n\nParticipants could still choose to solve the problem via voxel segmentation if they wish. Organizers could provide convenience utilities for converting between meshes and voxel masks in both directions.\n\n#### 2. Reformulate as instance segmentation and drop thickness-sensitive metrics (balanced)\n\nSince meshes are already available as annotations, it should be possible to provide instance maps rather than a single semantic foreground mask. Participants would be expected to submit instance segmentations.\n\nThis would eliminate connected-components ambiguities in both ground truth and predictions, which I believe is crucial. Participants could still choose to predict semantic masks internally, but the responsibility of converting these into instances would lie with them, for example via connected components.\n\nGiven that sheet thickness is arbitrary and sheets may appear detached from surrounding material, Surface Dice could be dropped. The evaluation could instead focus on VOI and TopoScore. It may also be worth investigating whether ASSD, computed either on binarized instance maps or per instance, could serve as a more appropriate replacement for Surface Dice.\n\n#### 3. Reformulate as instance segmentation while keeping the current metrics (least intrusive)\n\nThis would be the minimal change. Provide instance maps as ground truth and expect instance predictions, but keep the current metric suite unchanged.\n\n---\n\nI would be very interested to hear what other participants think about these points, and especially how the organizers view the trade-offs involved. Any feedback, counterarguments, or alternative ideas would be very welcome.",
    "3373305": "Initial a grid mesh=>manipulate mesh using Inverse Composition Lucas-Kanade Method with GAN judgement => fitted object topology\n\nNeed to do mask * vol first if host does not provide more information.",
    "3373367": "Sir, you have me confused. Could you unpack that a bit?",
    "3374231": "Thanks a lot Fabian, this is thoughtful feedback, and it’s very helpful to have you (and others) stress‑testing the task design against the real downstream goal.\n\nYou’re absolutely right that the end product of our unwrapping pipeline is a surface mesh. Our “Virtual Unwrapping” pipeline is essentially: (1) obtain an intermediate surface representation (currently a binary voxel mask), then (2) turn it into a mesh and flatten/render it.\n\nFor Kaggle, we chose to frame the challenge on voxel mask predictions mainly because it keeps the barrier to entry low and fits the platform constraints (submission format, evaluation robustness/runtime, etc.). A mesh‑submission task would require instance handling + mesh correspondence/pairing and a fairly careful evaluation design (check for manifoldness, self-intersections, intersections across meshes), which is doable but substantially more complex to deploy and to make stable. We are currently [researching on mesh neural tracing](https://github.com/ScrollPrize/villa/tree/neural-tracing/vesuvius/src/vesuvius/neural_tracing) approaches, but the formulation is not straightforward! \n\nThat said, we agree with your core point: \"binary\" voxelization introduces style/thickness ambiguity, especially where sheets come close or touch. More broadly, while the underlying annotated surface (plus some diversion in air that is deemed acceptable for the unwrapping) is the real target, the voxelized representation is not uniquely determined by the data (local thickness choices, tie‑breaking near contacts), so we want the score to primarily reflect the underlying surface geometry/topology rather than quirks of the mesh→voxelization pipeline. In practice, our voxelization/cleanup does include heuristics in near‑contact regions (including thinning/separation choices) to avoid gross mergers; this is useful downstream, but it does increase thickness/style ambiguity in the masks.\n\nWe acknowledge there are edge cases where the current tolerance in Surface Dice may be too tight relative to label uncertainty. Rather than completely changing the current metric suite, we are open to consider loosening the tolerance (and we recognize metric changes are sensitive, so we’d want clear community agreement and consistent application). Is there a specific tolerance range that you, or anyone in the community, think would better suit the task? We’re also open to suggestions for more thickness‑invariant surface evaluation (e.g., ASSD), as long as they remain stable and feasible for Kaggle scoring.\n\nI think that the bias introduced by choosing a semantic segmentation formulation does not necessarily block the utilization of different approaches, like instance segmentation methods, provided that the ground truth (and eventually the predictions) are accurate enough to allow for a change. In particular, if sheets are cleanly separated, semantic→instances is often just connected components (with the usual caveats where a sheet enters/exits the cropped ROI). \n\nThis is why on the data/label side, prompted by community reports and our own quality controls, we’ve identified a few recurring artifact classes in a non-trivial subset of the labels. Mainly (1) unintended mergers between nearby sheets, (2) small spurious “dust” components that should have been removed, and (3) small holes/gaps that can be introduced by the mesh→voxelization step. Because the evaluation blends VOI + Surface Dice + a Betti-matching topological score, label quality matters a lot here, so we’re treating this as high priority and we intend to release soon an updated version of the dataset. @seanjohnsonsp and the team already spent the last 10 hours on it!\n\nImportantly, this is not meant to discourage anyone: many of the strongest Kaggle approaches we’ve seen already use exactly the kinds of robustness + topology-aware processing that are aligned with the downstream unwrapping pipeline. Those ideas remain correct, and improving label artifacts should make training easier, not invalidate progress and will surely help us to read the scrolls!",
    "3375074": "Thanks for the quick and detailed response. It is great to hear that you are treating GT quality as a high-priority item, and I really appreciate that you are actively revising the problematic cases and plan to release an updated dataset soon.\n\nI understand why you chose a voxel-mask submission format given Kaggle constraints and the need for a stable evaluation pipeline. I probably agree a bit less on the approachability argument, though. From my perspective, approachability and flexibility can coexist if you provide reliable reference conversions between representations (mesh ↔ instance mask ↔ semantic mask), while keeping the submission format simple.\n\nThe point I feel most strongly about is the instance question. Even if GT is cleaned up such that connected components yields reasonable instances, the current pipeline still implicitly requires participants to produce a semantic mask that (a) stays close to the GT voxelization style and thickness for Surface Dice, and (b) can be separated into instances via connected components without accidental connections. For teams that naturally model sheets as instances or surfaces, this introduces a round trip that can be both error-prone and wasteful: instance or surface representation → semantic mask for submission → instances again via connected components on the evaluation side.\n\nIn particular, avoiding inadvertent connections in the semantic mask becomes an extra constraint that is not really part of the downstream goal, but can heavily affect VOI and TopoScore, and which would largely disappear if instance- or surface-based representations were accepted directly.\n\nBecause of this, I think it would be a meaningful improvement to define instances as the canonical representation: provide instance GT, expect instance submissions, and provide a small reference snippet showing how to convert a semantic mask to instances via connected components (cc3d is very fast) for participants who prefer to stay in semantic segmentation. This would reduce duplicated effort and make the evaluation semantics more explicit.\n\nOn Surface Dice, I would not feel comfortable recommending a new tolerance value in the current setup. Increasing τ while computing Surface Dice on the global semantic foreground can introduce unintended behavior. For example, when one sheet is missed but lies close to another sheet, a high tolerance can still give partial credit simply due to proximity. More generally, thickness and style ambiguity in voxelization makes Surface Dice sensitive in ways that are not necessarily aligned with the underlying surface target.\n\nOne possible way around this, without fully changing the metric suite, could be to compute Surface Dice at the instance level rather than on the global semantic mask. In rough terms:\n\n- Generate instance maps for GT and predictions (or ideally make GT an instance map and expect teams to submit instance maps as well).\n\n- Match GT and predicted instances, for example via pairwise IoU with greedy or Hungarian matching.\n\n- For each matched pair, compute Surface Dice with a more permissive tolerance (for example τ ≈ 5), since neighboring sheets are no longer part of the comparison.\n\n- Assign Surface Dice = 0 for unmatched GT instances (false negatives) and unmatched predicted instances (false positives).\n\n- Aggregate per volume (for example average across instances), then average over volumes, so volumes with many sheets do not dominate.\n\nA nice side effect is that instance-based Surface Dice with a generous tolerance would largely remove the need to match an arbitrary voxel thickness. This would likely make it easier for creative approaches like @hengck23 's line completion ideas or the crazy methods @tom99763 is pursuing, without forcing everyone to tune postprocessing to a particular voxelization style.\n\nOverall, if you were willing to move to instance ground truth and instance submissions, provide conversion code for those who want to stay in semantic segmentation, and compute Surface Dice in an instance-aware manner as outlined above (while keeping VOI and TopoScore), I think this would lead to a setup that remains approachable while aligning more directly with the downstream unwrapping objective. From my perspective, this would be a relatively small change on the evaluation side, but of course one that the community would need to be comfortable with.\n\nI would be very interested to hear not only what you think about this direction, but also what other participants think.",
    "3375799": "Hi @fabianisensee , I believe your perspective is very insightful, and the organizer’s @giorgioangelotti reply is also well-reasoned. Therefore, I would like to discuss this issue from the standpoint of **competition fairness**:\n\n1. On Kaggle, unless there are serious flaws in the evaluation metric, **it should not be changed** during the competition. In my view, changing the metric is essentially equivalent to starting a new competition.\n\n2. Changing the metric affects various participants differently. Some have limited computational resources, while others may lack sufficient time in the later stages. In any case, this can easily **undermine the efforts of participants who worked hard before the change,** reducing their motivation and potentially causing the competition to miss out on many excellent solutions.\n\n3. Considering that the organizer intends to **make changes to the dataset**, **this already requires additional effort from participants**. Changing the metric on top of that would only add to the workload for many, making an already challenging situation even more difficult.\n\nThank you for your time, and I hope the organizer achieves the solutions they are looking for.",
    "3375843": "I am not 100% sure, but this could be the mesh to voxel conversion code https://github.com/ScrollPrize/villa/blob/3de44b8d70e49886295b310e54f1f9962ab3fe5c/vesuvius/src/vesuvius/models/datasets/mesh/voxelize.py\n\nYou can learn a network to do the reverse",
    "3375848": "seanjohnsonsp annotated the meshes, voxelized them and then refined them manually. I don't think he used that code, or only that code.",
    "3376817": "Thanks for raising this perspective. I fully agree that major metric changes during a competition should be avoided, as they can raise fairness concerns and undermine prior effort. If any changes were to be made, they should clearly have broad community support and, of course, would ultimately be at the discretion of the organizers.\n\nI would like to clarify that my most recent suggestion is not intended as a major metric redesign. Switching from a semantic foreground mask to an explicit instance representation for GT and submission, and computing Surface Dice on instances rather than on the global semantic mask, would in my view be a relatively minor adjustment that also elegantly addresses the arbitrary thickness issue, provided a permissive tolerance is used. VOI and TopoScore are already instance-based in practice, since connected components are run internally on submissions, so this would mainly make the instance notion explicit and consistent across all metrics, while at the same time providing a much richer and more task-aligned representation for participants to work with and explore creative ideas.\n\nExisting approaches would continue to work with minimal changes, and given that we are still about two months out, I personally would not see a need for a competition extension if such a change were agreed upon. That said, I fully agree that fairness comes first, and any decision here should be guided by community consensus and the organizers’ judgment.",
    "3377112": "With the current metric, what (private) score would you consider/approximate a pipeline to be successful?\nSuccessful as in your target for this competition or capability to move to the next stage of decoding the scroll contents.",
    "3377141": "I would say that with a LB > 0.7 we would be in a very good position. Our current pipeline is basically a \"mesher\" on top of these predictions. The pipe is already capable of automatically solving some \"easy\" mistakes, but it really struggles in areas where the sheets are densely packed. Those mistakes are usually fixed with manual intervention, but this requires a lot of effort. This is why creating these labels is no easy.",
    "3377988": "In fact I would agree with all three points above but as also mentioned, in such case, it shall be bright new competition then pivoting data and metrics (all key aspects defining the competition) in middle of the race...",
    "3377989": "> instances rather than on the global semantic mask\n\nThat is major task shift and completely redefine the competition, although I agree that it shall be instance segmentation, this would disqualify all previous work and invested compute, as for example me with limited communication, noone reset 200+hours of compute I already used...\n\nCc @giorgioangelotti",
    "3377990": "There is not such submission yet... The best one is 0.58 likely segmentation model 🤔",
    "3378447": "please start your performance(mislead) !",
    "3379490": "giorgioangelotti \nThank you for information. I learn a lot of things in this competition.  \nI have a question. When will you change dataset?  \nyou said train/test label annotation will be changed. We need to retrain dataset. this competition is 3d volume, huge calculation cost.  \ncompetitor need enough time.",
    "3380848": "It's now live",
    "3381193": "Thank you for announcement. I'm checking now."
  },
  "source": "meta"
}