{
  "id": 635357,
  "title": "Quick Clarification on Annotation Detail and Score Ceiling",
  "url": "/competitions/vesuvius-challenge-surface-detection/discussion/635357",
  "author_name": "",
  "post_date": "2025-11-20T10:12:41.067474Z",
  "votes": 11,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I've been looking closely at the annotations, especially when viewing them in the 3D tool (like the one linked <a href=\"https://www.kaggle.com/code/jirkaborovec/surface-detect-interactive-img-mask-3d-view\" target=\"_blank\">here</a>).</p>\n<p>My perception is that the provided ground truth effectively captures the main direction and general shape of the scroll surface, but it seems to \"smooth out\" some of the very minor, local wraps, folds, and irregularities present in the raw CT data.</p>\n<p>This isn't a critique of the exhaustive annotation effort at all! Given how complex the geometry is, the ground truth is an incredible resource.</p>\n<p>I just wanted to spark a discussion:</p>\n<ol>\n<li>Does the community agree that the true, physical surface of the papyrus likely contains more \"minor local wraps\" than what's reflected in the simplified/smoothed annotations?</li>\n<li>If so, does this inherently set a realistic ceiling for the maximum achievable score (say, below 1.0), even for a \"perfect\" model that precisely segments the true, highly detailed physical surface? The perfect segmentation would be slightly penalized for mismatching the smoothed ground truth.</li>\n</ol>\n<p>Given that the SurfaceDice is tolerant (with $\\tau=2.0$), this might mitigate the issue, but I'd appreciate any insights from the hosts or experts on how to think about this trade-off between segmenting the <strong>physical reality</strong> vs. segmenting the <strong>annotated truth</strong>.</p>",
  "messages": [
    {
      "id": "3341620",
      "postDate": "11/20/2025 10:12:41",
      "content": "<p>I've been looking closely at the annotations, especially when viewing them in the 3D tool (like the one linked <a href=\"https://www.kaggle.com/code/jirkaborovec/surface-detect-interactive-img-mask-3d-view\" target=\"_blank\">here</a>).</p>\n<p>My perception is that the provided ground truth effectively captures the main direction and general shape of the scroll surface, but it seems to \"smooth out\" some of the very minor, local wraps, folds, and irregularities present in the raw CT data.</p>\n<p>This isn't a critique of the exhaustive annotation effort at all! Given how complex the geometry is, the ground truth is an incredible resource.</p>\n<p>I just wanted to spark a discussion:</p>\n<ol>\n<li>Does the community agree that the true, physical surface of the papyrus likely contains more \"minor local wraps\" than what's reflected in the simplified/smoothed annotations?</li>\n<li>If so, does this inherently set a realistic ceiling for the maximum achievable score (say, below 1.0), even for a \"perfect\" model that precisely segments the true, highly detailed physical surface? The perfect segmentation would be slightly penalized for mismatching the smoothed ground truth.</li>\n</ol>\n<p>Given that the SurfaceDice is tolerant (with $\\tau=2.0$), this might mitigate the issue, but I'd appreciate any insights from the hosts or experts on how to think about this trade-off between segmenting the <strong>physical reality</strong> vs. segmenting the <strong>annotated truth</strong>.</p>",
      "rawMarkdown": "I've been looking closely at the annotations, especially when viewing them in the 3D tool (like the one linked [here](https://www.kaggle.com/code/jirkaborovec/surface-detect-interactive-img-mask-3d-view)).\n\nMy perception is that the provided ground truth effectively captures the main direction and general shape of the scroll surface, but it seems to \"smooth out\" some of the very minor, local wraps, folds, and irregularities present in the raw CT data.\n\nThis isn't a critique of the exhaustive annotation effort at all! Given how complex the geometry is, the ground truth is an incredible resource.\n\nI just wanted to spark a discussion:\n\n1.  Does the community agree that the true, physical surface of the papyrus likely contains more \"minor local wraps\" than what's reflected in the simplified/smoothed annotations?\n2.  If so, does this inherently set a realistic ceiling for the maximum achievable score (say, below 1.0), even for a \"perfect\" model that precisely segments the true, highly detailed physical surface? The perfect segmentation would be slightly penalized for mismatching the smoothed ground truth.\n\nGiven that the SurfaceDice is tolerant (with $\\tau=2.0$), this might mitigate the issue, but I'd appreciate any insights from the hosts or experts on how to think about this trade-off between segmenting the **physical reality** vs. segmenting the **annotated truth**.",
      "votes": null
    },
    {
      "id": "3341853",
      "postDate": "11/20/2025 13:28:10",
      "content": "<p>Hey! Thanks for participating!</p>\n<p>You are correct in this assessment. We don't aim to smooth out the minor warps this unfortunately just is a result of the dataset creation pipeline as it stands now (it used to be much worse!). For the most part, the SurfaceDice should have enough tolerance to account for this discrepancy, but if you see any egregious examples please let me know. I've done my best to select samples which have minimal divergence between \"true\" and \"labeled\" topology. </p>\n<p>This does also bring up another interesting point however in that heavily penalizing per-voxel losses might benefit from some smoothing or other terms which do not penalize the placement as harshly as the general topology. </p>",
      "rawMarkdown": "Hey! Thanks for participating!\n\nYou are correct in this assessment. We don't aim to smooth out the minor warps this unfortunately just is a result of the dataset creation pipeline as it stands now (it used to be much worse!). For the most part, the SurfaceDice should have enough tolerance to account for this discrepancy, but if you see any egregious examples please let me know. I've done my best to select samples which have minimal divergence between \"true\" and \"labeled\" topology. \n\nThis does also bring up another interesting point however in that heavily penalizing per-voxel losses might benefit from some smoothing or other terms which do not penalize the placement as harshly as the general topology.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3341853,
      "author_name": "seanjohnsonsp",
      "author_url": "",
      "post_date": "11/20/2025 13:28:10",
      "content": "<p>Hey! Thanks for participating!</p>\n<p>You are correct in this assessment. We don't aim to smooth out the minor warps this unfortunately just is a result of the dataset creation pipeline as it stands now (it used to be much worse!). For the most part, the SurfaceDice should have enough tolerance to account for this discrepancy, but if you see any egregious examples please let me know. I've done my best to select samples which have minimal divergence between \"true\" and \"labeled\" topology. </p>\n<p>This does also bring up another interesting point however in that heavily penalizing per-voxel losses might benefit from some smoothing or other terms which do not penalize the placement as harshly as the general topology. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3341620": "I've been looking closely at the annotations, especially when viewing them in the 3D tool (like the one linked [here](https://www.kaggle.com/code/jirkaborovec/surface-detect-interactive-img-mask-3d-view)).\n\nMy perception is that the provided ground truth effectively captures the main direction and general shape of the scroll surface, but it seems to \"smooth out\" some of the very minor, local wraps, folds, and irregularities present in the raw CT data.\n\nThis isn't a critique of the exhaustive annotation effort at all! Given how complex the geometry is, the ground truth is an incredible resource.\n\nI just wanted to spark a discussion:\n\n1.  Does the community agree that the true, physical surface of the papyrus likely contains more \"minor local wraps\" than what's reflected in the simplified/smoothed annotations?\n2.  If so, does this inherently set a realistic ceiling for the maximum achievable score (say, below 1.0), even for a \"perfect\" model that precisely segments the true, highly detailed physical surface? The perfect segmentation would be slightly penalized for mismatching the smoothed ground truth.\n\nGiven that the SurfaceDice is tolerant (with $\\tau=2.0$), this might mitigate the issue, but I'd appreciate any insights from the hosts or experts on how to think about this trade-off between segmenting the **physical reality** vs. segmenting the **annotated truth**.",
    "3341853": "Hey! Thanks for participating!\n\nYou are correct in this assessment. We don't aim to smooth out the minor warps this unfortunately just is a result of the dataset creation pipeline as it stands now (it used to be much worse!). For the most part, the SurfaceDice should have enough tolerance to account for this discrepancy, but if you see any egregious examples please let me know. I've done my best to select samples which have minimal divergence between \"true\" and \"labeled\" topology. \n\nThis does also bring up another interesting point however in that heavily penalizing per-voxel losses might benefit from some smoothing or other terms which do not penalize the placement as harshly as the general topology."
  },
  "source": "meta"
}