{
  "id": 673114,
  "title": "Do binary hole filling algorithms fail?",
  "url": "/competitions/vesuvius-challenge-surface-detection/discussion/673114",
  "author_name": "",
  "post_date": "2026-02-12T12:40:16.651052600Z",
  "votes": 4,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I'm working with 3D binary segmentation masks and noticed that different hole-filling implementations produce unexpected results, basically it does not work on 3D. what is wrong here? any idea for surface aware smoothing or hole filling?</p>\n<h1>Method 1: scipy.ndimage</h1>\n<p>from scipy.ndimage import binary_fill_holes\nfilled_scipy = binary_fill_holes(mask)</p>\n<h1>Method 2: scikit-image</h1>\n<p>from skimage.morphology import remove_small_holes, binary_closing, ball\nfilled_skimage = remove_small_holes(mask, area_threshold=1000)</p>\n<h1>or</h1>\n<p>filled_skimage = binary_closing(mask, footprint=ball(3))</p>\n<h1>Method 3: SimpleITK</h1>\n<p>import SimpleITK as sitk\nsitk_img = sitk.GetImageFromArray(mask.astype(np.uint8))\nclosed = sitk.BinaryMorphologicalClosing(sitk_img, kernelRadius=[2,2,2])\nfilled_sitk = sitk.BinaryFillhole(closed, fullyConnected=True)\nfilled_sitk = sitk.GetArrayFromImage(filled_sitk)</p>",
  "messages": [
    {
      "id": "3405224",
      "postDate": "02/12/2026 12:40:16",
      "content": "<p>I'm working with 3D binary segmentation masks and noticed that different hole-filling implementations produce unexpected results, basically it does not work on 3D. what is wrong here? any idea for surface aware smoothing or hole filling?</p>\n<h1>Method 1: scipy.ndimage</h1>\n<p>from scipy.ndimage import binary_fill_holes\nfilled_scipy = binary_fill_holes(mask)</p>\n<h1>Method 2: scikit-image</h1>\n<p>from skimage.morphology import remove_small_holes, binary_closing, ball\nfilled_skimage = remove_small_holes(mask, area_threshold=1000)</p>\n<h1>or</h1>\n<p>filled_skimage = binary_closing(mask, footprint=ball(3))</p>\n<h1>Method 3: SimpleITK</h1>\n<p>import SimpleITK as sitk\nsitk_img = sitk.GetImageFromArray(mask.astype(np.uint8))\nclosed = sitk.BinaryMorphologicalClosing(sitk_img, kernelRadius=[2,2,2])\nfilled_sitk = sitk.BinaryFillhole(closed, fullyConnected=True)\nfilled_sitk = sitk.GetArrayFromImage(filled_sitk)</p>",
      "rawMarkdown": "I'm working with 3D binary segmentation masks and noticed that different hole-filling implementations produce unexpected results, basically it does not work on 3D. what is wrong here? any idea for surface aware smoothing or hole filling?\n\n# Method 1: scipy.ndimage\nfrom scipy.ndimage import binary_fill_holes\nfilled_scipy = binary_fill_holes(mask)\n\n# Method 2: scikit-image\nfrom skimage.morphology import remove_small_holes, binary_closing, ball\nfilled_skimage = remove_small_holes(mask, area_threshold=1000)\n# or\nfilled_skimage = binary_closing(mask, footprint=ball(3))\n\n# Method 3: SimpleITK\nimport SimpleITK as sitk\nsitk_img = sitk.GetImageFromArray(mask.astype(np.uint8))\nclosed = sitk.BinaryMorphologicalClosing(sitk_img, kernelRadius=[2,2,2])\nfilled_sitk = sitk.BinaryFillhole(closed, fullyConnected=True)\nfilled_sitk = sitk.GetArrayFromImage(filled_sitk)",
      "votes": null
    },
    {
      "id": "3405230",
      "postDate": "02/12/2026 13:04:10",
      "content": "<p>Have you tested these on the updated LB? Because i tried scipy binary_closing, and it did improved the local val. Although i haven't submitted it yet on the LB. </p>\n<p>What i am more certain is that these post processing after the inference won't be better than fixing the dataset on our own and then retraining the model again. As the public/priv LB is fixed with the toposcore issue. So making a robust model that would not-penaltize the holes while training would result in better on LB. These are just my understanding yet, i haven't done any testing on this as i was a little late in this competition and still figuring out most of the things. </p>",
      "rawMarkdown": "Have you tested these on the updated LB? Because i tried scipy binary_closing, and it did improved the local val. Although i haven't submitted it yet on the LB. \n\nWhat i am more certain is that these post processing after the inference won't be better than fixing the dataset on our own and then retraining the model again. As the public/priv LB is fixed with the toposcore issue. So making a robust model that would not-penaltize the holes while training would result in better on LB. These are just my understanding yet, i haven't done any testing on this as i was a little late in this competition and still figuring out most of the things.",
      "votes": null
    },
    {
      "id": "3405257",
      "postDate": "02/12/2026 14:48:45",
      "content": "<p>a 3d \"hole\" (at least in the majority of examples we have seen in this data) in a thin surface might be better described as a \"tunnel\", and most \"hole\" filling algorithms when used in 3d will fail to properly fix or detect it. a 3d hole fill on this data address \"voids\" more than it does \"tunnels\". </p>\n<p>you have a number of choices, all which have some downside. </p>\n<p>if you compute the betti dim 1 representative cycles, you can somewhat localize these by using the death voxel for the representative cycle of the betti dim1/h1 \"hole\" feature. you could create a mask of this death voxel or the entire representative cycle. if you have identified a coordinate or coordinates, you can address it a few ways. note that almost all of these have the potential to merge components, so doing a 26-connected components count before/after would be a good idea.</p>\n<ul>\n<li><p>a 2d binary hole fill, computed over slices (on each dim)</p>\n<ul>\n<li>this will miss some and may fill some things which in 3d are not holes at all but \"pockets\" or \"shelves\"</li></ul></li>\n<li><p>morphological operations (closing, opening, dilation). closing may be your best bet of these choices</p>\n<ul>\n<li>a small kernel will get some, but again not all, and a large kernel will likely alter topology</li></ul></li>\n<li><p>other, more complex methods, of which there are many ideas</p>\n<ul>\n<li>re-voxelization (heavy, kind of complex) via line interpolation<ul>\n<li>changes topology, heuristic heavy</li></ul></li>\n<li>inspirations from mesh-land<ul>\n<li>alpha shapes of the representative cycle, filled <ul>\n<li>need to tune alpha</li></ul></li>\n<li>poisson surface reconstruction</li></ul></li></ul></li>\n</ul>\n<p>an ideal result would: </p>\n<ul>\n<li>add as few voxels as possible to the gt</li>\n<li>not alter the 26-connected component count</li>\n<li>result in zero betti dim1/h1 features  (down from whatever you started with)</li>\n</ul>",
      "rawMarkdown": "a 3d \"hole\" (at least in the majority of examples we have seen in this data) in a thin surface might be better described as a \"tunnel\", and most \"hole\" filling algorithms when used in 3d will fail to properly fix or detect it. a 3d hole fill on this data address \"voids\" more than it does \"tunnels\". \n\nyou have a number of choices, all which have some downside. \n\nif you compute the betti dim 1 representative cycles, you can somewhat localize these by using the death voxel for the representative cycle of the betti dim1/h1 \"hole\" feature. you could create a mask of this death voxel or the entire representative cycle. if you have identified a coordinate or coordinates, you can address it a few ways. note that almost all of these have the potential to merge components, so doing a 26-connected components count before/after would be a good idea.\n\n- a 2d binary hole fill, computed over slices (on each dim)\n\t- this will miss some and may fill some things which in 3d are not holes at all but \"pockets\" or \"shelves\"\n\n- morphological operations (closing, opening, dilation). closing may be your best bet of these choices\n\t- a small kernel will get some, but again not all, and a large kernel will likely alter topology\n\t\n- other, more complex methods, of which there are many ideas\n\t- re-voxelization (heavy, kind of complex) via line interpolation\n\t\t- changes topology, heuristic heavy\n\t- inspirations from mesh-land\n\t\t- alpha shapes of the representative cycle, filled \n\t\t\t- need to tune alpha\n\t\t- poisson surface reconstruction\n\t\nan ideal result would: \n- add as few voxels as possible to the gt\n- not alter the 26-connected component count\n- result in zero betti dim1/h1 features  (down from whatever you started with)",
      "votes": null
    },
    {
      "id": "3405324",
      "postDate": "02/12/2026 18:31:10",
      "content": "<p>Yes, void filling and surface smoothing are likely the key factors for winning this competition. However, the labeling appears overly optimistic and does not consistently follow the actual surface in some areas.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3842406%2F5356766476bd23e8799e17cf97016369%2Fres.png?generation=1770921306629158&amp;alt=media\" alt=\"\"> <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3842406%2Fc34f2a0545a11b868b3f85c5ee9432f1%2FScreenshot%202026-02-12%20193034.png?generation=1770921289671879&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Yes, void filling and surface smoothing are likely the key factors for winning this competition. However, the labeling appears overly optimistic and does not consistently follow the actual surface in some areas.![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3842406%2F5356766476bd23e8799e17cf97016369%2Fres.png?generation=1770921306629158&alt=media) ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3842406%2Fc34f2a0545a11b868b3f85c5ee9432f1%2FScreenshot%202026-02-12%20193034.png?generation=1770921289671879&alt=media)",
      "votes": null
    },
    {
      "id": "3405327",
      "postDate": "02/12/2026 18:41:19",
      "content": "<p>I'm not 100% sure what this image is meant to show, but just as some background on the data: a sheet of papyrus is composed of two separate sheets of papyrus stuck together, with the strands of papyrus oriented horizontally on the front and vertically in the back. These sheets (and the fibers which they are composed of) split and fray often. The labels for this competition are meant to label the front surface. </p>\n<p>There are unfortunately areas they are not perfectly voxel accurate, due to the difficulty of labeling this data (which will hopefully be bootstrapped by the excellent work you guys have all put towards this problem). </p>",
      "rawMarkdown": "I'm not 100% sure what this image is meant to show, but just as some background on the data: a sheet of papyrus is composed of two separate sheets of papyrus stuck together, with the strands of papyrus oriented horizontally on the front and vertically in the back. These sheets (and the fibers which they are composed of) split and fray often. The labels for this competition are meant to label the front surface. \n\nThere are unfortunately areas they are not perfectly voxel accurate, due to the difficulty of labeling this data (which will hopefully be bootstrapped by the excellent work you guys have all put towards this problem).",
      "votes": null
    },
    {
      "id": "3405330",
      "postDate": "02/12/2026 18:46:07",
      "content": "<p>One more piece of information which may clear up some confusing areas in the data -- to compose a sheet of papyrus which is the length of a scroll, many sheets of individual papyrus (called kollema) are glued together (the seam itself being known as a koelleisis). This overlap part is very disorienting for manual annotation, but models seem relatively robust to it given enough samples. </p>",
      "rawMarkdown": "One more piece of information which may clear up some confusing areas in the data -- to compose a sheet of papyrus which is the length of a scroll, many sheets of individual papyrus (called kollema) are glued together (the seam itself being known as a koelleisis). This overlap part is very disorienting for manual annotation, but models seem relatively robust to it given enough samples.",
      "votes": null
    },
    {
      "id": "3405336",
      "postDate": "02/12/2026 19:07:33",
      "content": "<p>thank you. the figures are from left (prediction) and right gt. </p>",
      "rawMarkdown": "thank you. the figures are from left (prediction) and right gt.",
      "votes": null
    },
    {
      "id": "3406621",
      "postDate": "02/16/2026 07:59:16",
      "content": "<p>It's not just optimism—it's fundamentally different. Ignore mask is just like playing the lottery.</p>",
      "rawMarkdown": "It's not just optimism—it's fundamentally different. Ignore mask is just like playing the lottery.",
      "votes": null
    },
    {
      "id": "3406693",
      "postDate": "02/16/2026 11:51:04",
      "content": "<p>Give me a hint! per favor.</p>",
      "rawMarkdown": "Give me a hint! per favor.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3405230,
      "author_name": "muhammadibrahim3093",
      "author_url": "",
      "post_date": "02/12/2026 13:04:10",
      "content": "<p>Have you tested these on the updated LB? Because i tried scipy binary_closing, and it did improved the local val. Although i haven't submitted it yet on the LB. </p>\n<p>What i am more certain is that these post processing after the inference won't be better than fixing the dataset on our own and then retraining the model again. As the public/priv LB is fixed with the toposcore issue. So making a robust model that would not-penaltize the holes while training would result in better on LB. These are just my understanding yet, i haven't done any testing on this as i was a little late in this competition and still figuring out most of the things. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3405257,
      "author_name": "seanjohnsonsp",
      "author_url": "",
      "post_date": "02/12/2026 14:48:45",
      "content": "<p>a 3d \"hole\" (at least in the majority of examples we have seen in this data) in a thin surface might be better described as a \"tunnel\", and most \"hole\" filling algorithms when used in 3d will fail to properly fix or detect it. a 3d hole fill on this data address \"voids\" more than it does \"tunnels\". </p>\n<p>you have a number of choices, all which have some downside. </p>\n<p>if you compute the betti dim 1 representative cycles, you can somewhat localize these by using the death voxel for the representative cycle of the betti dim1/h1 \"hole\" feature. you could create a mask of this death voxel or the entire representative cycle. if you have identified a coordinate or coordinates, you can address it a few ways. note that almost all of these have the potential to merge components, so doing a 26-connected components count before/after would be a good idea.</p>\n<ul>\n<li><p>a 2d binary hole fill, computed over slices (on each dim)</p>\n<ul>\n<li>this will miss some and may fill some things which in 3d are not holes at all but \"pockets\" or \"shelves\"</li></ul></li>\n<li><p>morphological operations (closing, opening, dilation). closing may be your best bet of these choices</p>\n<ul>\n<li>a small kernel will get some, but again not all, and a large kernel will likely alter topology</li></ul></li>\n<li><p>other, more complex methods, of which there are many ideas</p>\n<ul>\n<li>re-voxelization (heavy, kind of complex) via line interpolation<ul>\n<li>changes topology, heuristic heavy</li></ul></li>\n<li>inspirations from mesh-land<ul>\n<li>alpha shapes of the representative cycle, filled <ul>\n<li>need to tune alpha</li></ul></li>\n<li>poisson surface reconstruction</li></ul></li></ul></li>\n</ul>\n<p>an ideal result would: </p>\n<ul>\n<li>add as few voxels as possible to the gt</li>\n<li>not alter the 26-connected component count</li>\n<li>result in zero betti dim1/h1 features  (down from whatever you started with)</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 3405324,
          "author_name": "jamalsaeedi",
          "author_url": "",
          "post_date": "02/12/2026 18:31:10",
          "content": "<p>Yes, void filling and surface smoothing are likely the key factors for winning this competition. However, the labeling appears overly optimistic and does not consistently follow the actual surface in some areas.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3842406%2F5356766476bd23e8799e17cf97016369%2Fres.png?generation=1770921306629158&amp;alt=media\" alt=\"\"> <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3842406%2Fc34f2a0545a11b868b3f85c5ee9432f1%2FScreenshot%202026-02-12%20193034.png?generation=1770921289671879&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": [
            {
              "id": 3405327,
              "author_name": "seanjohnsonsp",
              "author_url": "",
              "post_date": "02/12/2026 18:41:19",
              "content": "<p>I'm not 100% sure what this image is meant to show, but just as some background on the data: a sheet of papyrus is composed of two separate sheets of papyrus stuck together, with the strands of papyrus oriented horizontally on the front and vertically in the back. These sheets (and the fibers which they are composed of) split and fray often. The labels for this competition are meant to label the front surface. </p>\n<p>There are unfortunately areas they are not perfectly voxel accurate, due to the difficulty of labeling this data (which will hopefully be bootstrapped by the excellent work you guys have all put towards this problem). </p>",
              "votes": null,
              "replies": [
                {
                  "id": 3405330,
                  "author_name": "seanjohnsonsp",
                  "author_url": "",
                  "post_date": "02/12/2026 18:46:07",
                  "content": "<p>One more piece of information which may clear up some confusing areas in the data -- to compose a sheet of papyrus which is the length of a scroll, many sheets of individual papyrus (called kollema) are glued together (the seam itself being known as a koelleisis). This overlap part is very disorienting for manual annotation, but models seem relatively robust to it given enough samples. </p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 3405336,
                  "author_name": "jamalsaeedi",
                  "author_url": "",
                  "post_date": "02/12/2026 19:07:33",
                  "content": "<p>thank you. the figures are from left (prediction) and right gt. </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            },
            {
              "id": 3406621,
              "author_name": "ggayoayogg",
              "author_url": "",
              "post_date": "02/16/2026 07:59:16",
              "content": "<p>It's not just optimism—it's fundamentally different. Ignore mask is just like playing the lottery.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3406693,
                  "author_name": "jamalsaeedi",
                  "author_url": "",
                  "post_date": "02/16/2026 11:51:04",
                  "content": "<p>Give me a hint! per favor.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3405224": "I'm working with 3D binary segmentation masks and noticed that different hole-filling implementations produce unexpected results, basically it does not work on 3D. what is wrong here? any idea for surface aware smoothing or hole filling?\n\n# Method 1: scipy.ndimage\nfrom scipy.ndimage import binary_fill_holes\nfilled_scipy = binary_fill_holes(mask)\n\n# Method 2: scikit-image\nfrom skimage.morphology import remove_small_holes, binary_closing, ball\nfilled_skimage = remove_small_holes(mask, area_threshold=1000)\n# or\nfilled_skimage = binary_closing(mask, footprint=ball(3))\n\n# Method 3: SimpleITK\nimport SimpleITK as sitk\nsitk_img = sitk.GetImageFromArray(mask.astype(np.uint8))\nclosed = sitk.BinaryMorphologicalClosing(sitk_img, kernelRadius=[2,2,2])\nfilled_sitk = sitk.BinaryFillhole(closed, fullyConnected=True)\nfilled_sitk = sitk.GetArrayFromImage(filled_sitk)",
    "3405230": "Have you tested these on the updated LB? Because i tried scipy binary_closing, and it did improved the local val. Although i haven't submitted it yet on the LB. \n\nWhat i am more certain is that these post processing after the inference won't be better than fixing the dataset on our own and then retraining the model again. As the public/priv LB is fixed with the toposcore issue. So making a robust model that would not-penaltize the holes while training would result in better on LB. These are just my understanding yet, i haven't done any testing on this as i was a little late in this competition and still figuring out most of the things.",
    "3405257": "a 3d \"hole\" (at least in the majority of examples we have seen in this data) in a thin surface might be better described as a \"tunnel\", and most \"hole\" filling algorithms when used in 3d will fail to properly fix or detect it. a 3d hole fill on this data address \"voids\" more than it does \"tunnels\". \n\nyou have a number of choices, all which have some downside. \n\nif you compute the betti dim 1 representative cycles, you can somewhat localize these by using the death voxel for the representative cycle of the betti dim1/h1 \"hole\" feature. you could create a mask of this death voxel or the entire representative cycle. if you have identified a coordinate or coordinates, you can address it a few ways. note that almost all of these have the potential to merge components, so doing a 26-connected components count before/after would be a good idea.\n\n- a 2d binary hole fill, computed over slices (on each dim)\n\t- this will miss some and may fill some things which in 3d are not holes at all but \"pockets\" or \"shelves\"\n\n- morphological operations (closing, opening, dilation). closing may be your best bet of these choices\n\t- a small kernel will get some, but again not all, and a large kernel will likely alter topology\n\t\n- other, more complex methods, of which there are many ideas\n\t- re-voxelization (heavy, kind of complex) via line interpolation\n\t\t- changes topology, heuristic heavy\n\t- inspirations from mesh-land\n\t\t- alpha shapes of the representative cycle, filled \n\t\t\t- need to tune alpha\n\t\t- poisson surface reconstruction\n\t\nan ideal result would: \n- add as few voxels as possible to the gt\n- not alter the 26-connected component count\n- result in zero betti dim1/h1 features  (down from whatever you started with)",
    "3405324": "Yes, void filling and surface smoothing are likely the key factors for winning this competition. However, the labeling appears overly optimistic and does not consistently follow the actual surface in some areas.![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3842406%2F5356766476bd23e8799e17cf97016369%2Fres.png?generation=1770921306629158&alt=media) ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3842406%2Fc34f2a0545a11b868b3f85c5ee9432f1%2FScreenshot%202026-02-12%20193034.png?generation=1770921289671879&alt=media)",
    "3405327": "I'm not 100% sure what this image is meant to show, but just as some background on the data: a sheet of papyrus is composed of two separate sheets of papyrus stuck together, with the strands of papyrus oriented horizontally on the front and vertically in the back. These sheets (and the fibers which they are composed of) split and fray often. The labels for this competition are meant to label the front surface. \n\nThere are unfortunately areas they are not perfectly voxel accurate, due to the difficulty of labeling this data (which will hopefully be bootstrapped by the excellent work you guys have all put towards this problem).",
    "3405330": "One more piece of information which may clear up some confusing areas in the data -- to compose a sheet of papyrus which is the length of a scroll, many sheets of individual papyrus (called kollema) are glued together (the seam itself being known as a koelleisis). This overlap part is very disorienting for manual annotation, but models seem relatively robust to it given enough samples.",
    "3405336": "thank you. the figures are from left (prediction) and right gt.",
    "3406621": "It's not just optimism—it's fundamentally different. Ignore mask is just like playing the lottery.",
    "3406693": "Give me a hint! per favor."
  },
  "source": "meta"
}