{
  "id": 672447,
  "title": "Questions for the host about the metric.",
  "url": "/competitions/vesuvius-challenge-surface-detection/discussion/672447",
  "author_name": "",
  "post_date": "2026-02-08T10:57:40.985015Z",
  "votes": 6,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Apart from the two major issues discussed <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/653482\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/671160\" target=\"_blank\">here </a> there are more errors that needs the hosts attention. I don't know if it has been discussed before or if the host is aware, but the evaluation algorithm spawns multiple components/holes <strong>even though our 3D sheet is 26 connected</strong>. Is the host aware of this?</p>\n<p>This is concrete example:</p>\n<p>these are Z,Y,X slices of the center coordinate of a \"hole\" that the algorithm found.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2Fd79b5b80a8ced47f4f9b1447e5aa6713%2FScreenshot_3.png?generation=1770544036505032&amp;alt=media\" alt=\"\"></p>\n<p>Like you, I don't see any holes from any of the 3 dimensions point of view. At first I thought that it is some kind of issue with the spurious ending. When inverted to run the births/deaths algorithm somehow is calculated as a hole. However upon inspecting it in 3D and zooming into the crop, we can see what kind of hole the algorithm found:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2F867a52a1c65b8b0bc6045ab565d9c423%2Faliasing3.png?generation=1770544126701329&amp;alt=media\" alt=\"\"></p>\n<p>From gemini:</p>\n<blockquote>\n  <p>In a \"staircase\" of voxels, 26-connectivity (which connects diagonals) allows a path to exist through the corners of the cubes. However, if there is a tiny gap (a 0) underneath the overhang, the algorithm sees a \"hole\" passing through that diagonal gap. This creates a false handle (Betti-1 loop)</p>\n</blockquote>\n<p>So even though our sheet is completely connected, there are NUMEROUS instances of these \"holes\" detected.</p>\n<p>Following are some more examples of the <strong>actual</strong> coordinates of the holes/tunnels that the algorithm detected:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2F4ef5940ff8d3745c26a8e0108dc9dab8%2Faliasing2.png?generation=1770544300377788&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2Fd3b07c1d12afdf790ffa7fe67a01e006%2FScreenshot_1.png?generation=1770544311462087&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2F19e8c5c1a84bdbe34027dbc5315e2fa5%2Faliasing.png?generation=1770544678119918&amp;alt=media\" alt=\"\"></p>\n<p>And some might say that this is not that bad and predicting few extra holes doesnt matter. But if you look closely, the algorithm severely punishes holes - my assumption is that they picked the scaling because they deeply care about holes not being present in the final prediction. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2Fe577a3c552fa59fef1d62ced8382b742%2FScreenshot_2.png?generation=1770545358429775&amp;alt=media\" alt=\"\"></p>\n<p>These are the hypothetical differences in leaderboard if you predict 1 hole vs 2 holes. Lets say in both scenarios we have some defautl values of 0.8 surface dice and 0.5 VOI. Lets also say we have 0.5 betti-0 score for stable comparison of how holes will affect the score in both cases.</p>\n<h2>1 hole scenario:</h2>\n<p>p_k = 1 &lt;--- we assume 1 hole predicted by our model</p>\n<p>g_k = 0 &lt;--- we assume 0 holes in the ground truth</p>\n<pre><code>denom = p_k + g_k\n            if denom &gt; 0:\n                if g_k != 0:\n                    topoF1_k = (2.0 * m_k) / float(denom)\n                else:\n                    topoF1_k = 0.5 / (float(denom) + 0.5)\n</code></pre>\n<p>topoF1_k = 0.5 / (1.0 + 0.5) = 0.3333</p>\n<p>TopoScore = (0.5 (betti-0) + 0.3333 (betti-1)) / 2 = 0.416665</p>\n<p>Score = 0.30 × TopoScore + 0.35 × SurfaceDice@τ + 0.35 × VOI_score</p>\n<blockquote>\n  <p>(0.3 * 0.416665) + (0.35 * 0.8) + (0.35 * 0.5) =  0.579</p>\n</blockquote>\n<p>We get 0.579 leaderboard score</p>\n<h2>2 holes scenario:</h2>\n<p>p_k = 2 &lt;--- we assume 2 holes predicted by our mode</p>\n<p>g_k = 0 &lt;--- we assume 0 holes in the ground truth</p>\n<p>topoF1_k = 0.5 / (2.0 + 0.5) = 0.2</p>\n<p>TopoScore = (0.5 (betti-0) + 0.2 (betti-1))  / 2 = 0.35</p>\n<p>Score = 0.30 × TopoScore + 0.35 × SurfaceDice@τ + 0.35 × VOI_score</p>\n<blockquote>\n  <p>(0.3 * 0.35) + (0.35 * 0.8) + (0.35 * 0.5) =  0.559</p>\n</blockquote>\n<p>We get 0.559 leaderboard score.</p>\n<p>Thats 0.02! score difference. We can clearly see that now the leaderboard is not gameable and deeply cares if you do/don't output holes even when the difference is 1 hole. The final solutions will converge to the ones that are most topologically sound, and effort will be shifted from dice score optimization -&gt; topological optimization, since now topology is worth much more, and instead of optimizing dice for marginal gains 0.00X you can get 0.0X improvements as shown in the example above, since now topology is worth more, more time will be spent on that. Few holes difference might've not sounded like a lot, but after doing the calculation you can see how this changes the convergence of the final solutions, as you will be rewarded for <em>balancing ALL of the terms in the score equation and topology will mater just as much as dice</em>. Right now topology has little to no say in the final score.</p>\n<p>Ultimately this is down to the hosts and what they want to get out of the competition. So its up to them to decide if they will fix the issues and ask for competition extension/help from the Kaggle staff, or the competition will conclude as is.</p>",
  "messages": [
    {
      "id": "3403374",
      "postDate": "02/08/2026 10:57:40",
      "content": "<p>Apart from the two major issues discussed <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/653482\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/671160\" target=\"_blank\">here </a> there are more errors that needs the hosts attention. I don't know if it has been discussed before or if the host is aware, but the evaluation algorithm spawns multiple components/holes <strong>even though our 3D sheet is 26 connected</strong>. Is the host aware of this?</p>\n<p>This is concrete example:</p>\n<p>these are Z,Y,X slices of the center coordinate of a \"hole\" that the algorithm found.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2Fd79b5b80a8ced47f4f9b1447e5aa6713%2FScreenshot_3.png?generation=1770544036505032&amp;alt=media\" alt=\"\"></p>\n<p>Like you, I don't see any holes from any of the 3 dimensions point of view. At first I thought that it is some kind of issue with the spurious ending. When inverted to run the births/deaths algorithm somehow is calculated as a hole. However upon inspecting it in 3D and zooming into the crop, we can see what kind of hole the algorithm found:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2F867a52a1c65b8b0bc6045ab565d9c423%2Faliasing3.png?generation=1770544126701329&amp;alt=media\" alt=\"\"></p>\n<p>From gemini:</p>\n<blockquote>\n  <p>In a \"staircase\" of voxels, 26-connectivity (which connects diagonals) allows a path to exist through the corners of the cubes. However, if there is a tiny gap (a 0) underneath the overhang, the algorithm sees a \"hole\" passing through that diagonal gap. This creates a false handle (Betti-1 loop)</p>\n</blockquote>\n<p>So even though our sheet is completely connected, there are NUMEROUS instances of these \"holes\" detected.</p>\n<p>Following are some more examples of the <strong>actual</strong> coordinates of the holes/tunnels that the algorithm detected:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2F4ef5940ff8d3745c26a8e0108dc9dab8%2Faliasing2.png?generation=1770544300377788&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2Fd3b07c1d12afdf790ffa7fe67a01e006%2FScreenshot_1.png?generation=1770544311462087&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2F19e8c5c1a84bdbe34027dbc5315e2fa5%2Faliasing.png?generation=1770544678119918&amp;alt=media\" alt=\"\"></p>\n<p>And some might say that this is not that bad and predicting few extra holes doesnt matter. But if you look closely, the algorithm severely punishes holes - my assumption is that they picked the scaling because they deeply care about holes not being present in the final prediction. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2Fe577a3c552fa59fef1d62ced8382b742%2FScreenshot_2.png?generation=1770545358429775&amp;alt=media\" alt=\"\"></p>\n<p>These are the hypothetical differences in leaderboard if you predict 1 hole vs 2 holes. Lets say in both scenarios we have some defautl values of 0.8 surface dice and 0.5 VOI. Lets also say we have 0.5 betti-0 score for stable comparison of how holes will affect the score in both cases.</p>\n<h2>1 hole scenario:</h2>\n<p>p_k = 1 &lt;--- we assume 1 hole predicted by our model</p>\n<p>g_k = 0 &lt;--- we assume 0 holes in the ground truth</p>\n<pre><code>denom = p_k + g_k\n            if denom &gt; 0:\n                if g_k != 0:\n                    topoF1_k = (2.0 * m_k) / float(denom)\n                else:\n                    topoF1_k = 0.5 / (float(denom) + 0.5)\n</code></pre>\n<p>topoF1_k = 0.5 / (1.0 + 0.5) = 0.3333</p>\n<p>TopoScore = (0.5 (betti-0) + 0.3333 (betti-1)) / 2 = 0.416665</p>\n<p>Score = 0.30 × TopoScore + 0.35 × SurfaceDice@τ + 0.35 × VOI_score</p>\n<blockquote>\n  <p>(0.3 * 0.416665) + (0.35 * 0.8) + (0.35 * 0.5) =  0.579</p>\n</blockquote>\n<p>We get 0.579 leaderboard score</p>\n<h2>2 holes scenario:</h2>\n<p>p_k = 2 &lt;--- we assume 2 holes predicted by our mode</p>\n<p>g_k = 0 &lt;--- we assume 0 holes in the ground truth</p>\n<p>topoF1_k = 0.5 / (2.0 + 0.5) = 0.2</p>\n<p>TopoScore = (0.5 (betti-0) + 0.2 (betti-1))  / 2 = 0.35</p>\n<p>Score = 0.30 × TopoScore + 0.35 × SurfaceDice@τ + 0.35 × VOI_score</p>\n<blockquote>\n  <p>(0.3 * 0.35) + (0.35 * 0.8) + (0.35 * 0.5) =  0.559</p>\n</blockquote>\n<p>We get 0.559 leaderboard score.</p>\n<p>Thats 0.02! score difference. We can clearly see that now the leaderboard is not gameable and deeply cares if you do/don't output holes even when the difference is 1 hole. The final solutions will converge to the ones that are most topologically sound, and effort will be shifted from dice score optimization -&gt; topological optimization, since now topology is worth much more, and instead of optimizing dice for marginal gains 0.00X you can get 0.0X improvements as shown in the example above, since now topology is worth more, more time will be spent on that. Few holes difference might've not sounded like a lot, but after doing the calculation you can see how this changes the convergence of the final solutions, as you will be rewarded for <em>balancing ALL of the terms in the score equation and topology will mater just as much as dice</em>. Right now topology has little to no say in the final score.</p>\n<p>Ultimately this is down to the hosts and what they want to get out of the competition. So its up to them to decide if they will fix the issues and ask for competition extension/help from the Kaggle staff, or the competition will conclude as is.</p>",
      "rawMarkdown": "Apart from the two major issues discussed [here](https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/653482) and [here ](https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/671160) there are more errors that needs the hosts attention. I don't know if it has been discussed before or if the host is aware, but the evaluation algorithm spawns multiple components/holes **even though our 3D sheet is 26 connected**. Is the host aware of this?\n\nThis is concrete example:\n\nthese are Z,Y,X slices of the center coordinate of a \"hole\" that the algorithm found.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2Fd79b5b80a8ced47f4f9b1447e5aa6713%2FScreenshot_3.png?generation=1770544036505032&alt=media)\n\nLike you, I don't see any holes from any of the 3 dimensions point of view. At first I thought that it is some kind of issue with the spurious ending. When inverted to run the births/deaths algorithm somehow is calculated as a hole. However upon inspecting it in 3D and zooming into the crop, we can see what kind of hole the algorithm found:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2F867a52a1c65b8b0bc6045ab565d9c423%2Faliasing3.png?generation=1770544126701329&alt=media)\n\nFrom gemini:\n>In a \"staircase\" of voxels, 26-connectivity (which connects diagonals) allows a path to exist through the corners of the cubes. However, if there is a tiny gap (a 0) underneath the overhang, the algorithm sees a \"hole\" passing through that diagonal gap. This creates a false handle (Betti-1 loop)\n\nSo even though our sheet is completely connected, there are NUMEROUS instances of these \"holes\" detected.\n\nFollowing are some more examples of the **actual** coordinates of the holes/tunnels that the algorithm detected:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2F4ef5940ff8d3745c26a8e0108dc9dab8%2Faliasing2.png?generation=1770544300377788&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2Fd3b07c1d12afdf790ffa7fe67a01e006%2FScreenshot_1.png?generation=1770544311462087&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2F19e8c5c1a84bdbe34027dbc5315e2fa5%2Faliasing.png?generation=1770544678119918&alt=media)\n\n\nAnd some might say that this is not that bad and predicting few extra holes doesnt matter. But if you look closely, the algorithm severely punishes holes - my assumption is that they picked the scaling because they deeply care about holes not being present in the final prediction. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2Fe577a3c552fa59fef1d62ced8382b742%2FScreenshot_2.png?generation=1770545358429775&alt=media)\n\nThese are the hypothetical differences in leaderboard if you predict 1 hole vs 2 holes. Lets say in both scenarios we have some defautl values of 0.8 surface dice and 0.5 VOI. Lets also say we have 0.5 betti-0 score for stable comparison of how holes will affect the score in both cases.\n\n## 1 hole scenario:\n\np_k = 1 <--- we assume 1 hole predicted by our model\n\ng_k = 0 <--- we assume 0 holes in the ground truth\n\n``` \ndenom = p_k + g_k\n            if denom > 0:\n                if g_k != 0:\n                    topoF1_k = (2.0 * m_k) / float(denom)\n                else:\n                    topoF1_k = 0.5 / (float(denom) + 0.5)\n```\ntopoF1_k = 0.5 / (1.0 + 0.5) = 0.3333\n\nTopoScore = (0.5 (betti-0) + 0.3333 (betti-1)) / 2 = 0.416665\n\nScore = 0.30 × TopoScore + 0.35 × SurfaceDice@τ + 0.35 × VOI_score\n\n>(0.3 * 0.416665) + (0.35 * 0.8) + (0.35 * 0.5) =  0.579\n\nWe get 0.579 leaderboard score\n\n## 2 holes scenario:\np_k = 2 <--- we assume 2 holes predicted by our mode\n\ng_k = 0 <--- we assume 0 holes in the ground truth\n\ntopoF1_k = 0.5 / (2.0 + 0.5) = 0.2\n\nTopoScore = (0.5 (betti-0) + 0.2 (betti-1))  / 2 = 0.35\n\nScore = 0.30 × TopoScore + 0.35 × SurfaceDice@τ + 0.35 × VOI_score\n\n>(0.3 * 0.35) + (0.35 * 0.8) + (0.35 * 0.5) =  0.559\n\nWe get 0.559 leaderboard score.\n\nThats 0.02! score difference. We can clearly see that now the leaderboard is not gameable and deeply cares if you do/don't output holes even when the difference is 1 hole. The final solutions will converge to the ones that are most topologically sound, and effort will be shifted from dice score optimization -> topological optimization, since now topology is worth much more, and instead of optimizing dice for marginal gains 0.00X you can get 0.0X improvements as shown in the example above, since now topology is worth more, more time will be spent on that. Few holes difference might've not sounded like a lot, but after doing the calculation you can see how this changes the convergence of the final solutions, as you will be rewarded for *balancing ALL of the terms in the score equation and topology will mater just as much as dice*. Right now topology has little to no say in the final score.\n\nUltimately this is down to the hosts and what they want to get out of the competition. So its up to them to decide if they will fix the issues and ask for competition extension/help from the Kaggle staff, or the competition will conclude as is.",
      "votes": null
    },
    {
      "id": "3403383",
      "postDate": "02/08/2026 11:13:41",
      "content": "<p>Thanks for mentioning this. These are likely artifacts from the voxelization procedure.\nThe fix of the test set that we already produced aimed to tackle the sources of the spurious Betti-1 and Betti-2 discrepancies detectable by the evaluation metric.</p>",
      "rawMarkdown": "Thanks for mentioning this. These are likely artifacts from the voxelization procedure.\nThe fix of the test set that we already produced aimed to tackle the sources of the spurious Betti-1 and Betti-2 discrepancies detectable by the evaluation metric.",
      "votes": null
    },
    {
      "id": "3403387",
      "postDate": "02/08/2026 11:19:30",
      "content": "<p>If I understand correctly, the fixes are done in the metric and the betti matching algorithm? (So we are not penalized for these edge connection aliasing tunnels)</p>",
      "rawMarkdown": "If I understand correctly, the fixes are done in the metric and the betti matching algorithm? (So we are not penalized for these edge connection aliasing tunnels)",
      "votes": null
    },
    {
      "id": "3403388",
      "postDate": "02/08/2026 11:22:29",
      "content": "<p>No, the fixes are produced in the test set, so that individual connected components no longer have these artifacts.</p>",
      "rawMarkdown": "No, the fixes are produced in the test set, so that individual connected components no longer have these artifacts.",
      "votes": null
    },
    {
      "id": "3403389",
      "postDate": "02/08/2026 11:27:11",
      "content": "<p>But this issue is with the <strong>predictions submitted</strong> that the metric evaluates. The images above are predictions from a model. As we can see in the 3-axis pyplot that there is <strong>no</strong> actual hole anywhere, yet the metric detected it as a hole. The point is that aliasing is happening on the evaluation level in the C++ code.</p>\n<p>Should this be counted as a hole?\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2F473934adb302db92fb7d7512514982c5%2FScreenshot_4.png?generation=1770550013781740&amp;alt=media\" alt=\"\"></p>\n<p>there are many such examples in a single volume. And its not a problem with the model, we can clearly see that everything is 26 connected and sound.</p>",
      "rawMarkdown": "But this issue is with the **predictions submitted** that the metric evaluates. The images above are predictions from a model. As we can see in the 3-axis pyplot that there is **no** actual hole anywhere, yet the metric detected it as a hole. The point is that aliasing is happening on the evaluation level in the C++ code.\n\nShould this be counted as a hole?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2F473934adb302db92fb7d7512514982c5%2FScreenshot_4.png?generation=1770550013781740&alt=media)\n\nthere are many such examples in a single volume. And its not a problem with the model, we can clearly see that everything is 26 connected and sound.",
      "votes": null
    },
    {
      "id": "3403393",
      "postDate": "02/08/2026 11:38:14",
      "content": "<p>Ah, yes, we intentionally wanted to penalize these artifacts. While it's true that we don't care for voxel accurate predictions, ideally we don't want cracks in them.</p>",
      "rawMarkdown": "Ah, yes, we intentionally wanted to penalize these artifacts. While it's true that we don't care for voxel accurate predictions, ideally we don't want cracks in them.",
      "votes": null
    },
    {
      "id": "3403395",
      "postDate": "02/08/2026 11:44:36",
      "content": "<blockquote>\n  <p>… we intentionally wanted to penalize these artifacts …</p>\n  <p>we don't want cracks in them.</p>\n</blockquote>\n<p>But isnt this fully 26 connected? And each voxel is connected to its neighbor by an edge/face so there should be 0 holes or cracks. We can see that the 4 voxels above are fully connected, and the vertex where they meet is causing the issue in the C++.  We are <em>physically</em> incapable of predicting a <strong>better</strong> connection between the voxels. I seem to be missing something by your comment that this is intentionally penalized? </p>",
      "rawMarkdown": "> ... we intentionally wanted to penalize these artifacts ...\n\n>we don't want cracks in them.\n\n\n\nBut isnt this fully 26 connected? And each voxel is connected to its neighbor by an edge/face so there should be 0 holes or cracks. We can see that the 4 voxels above are fully connected, and the vertex where they meet is causing the issue in the C++.  We are *physically* incapable of predicting a **better** connection between the voxels. I seem to be missing something by your comment that this is intentionally penalized?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3403383,
      "author_name": "giorgioangelotti",
      "author_url": "",
      "post_date": "02/08/2026 11:13:41",
      "content": "<p>Thanks for mentioning this. These are likely artifacts from the voxelization procedure.\nThe fix of the test set that we already produced aimed to tackle the sources of the spurious Betti-1 and Betti-2 discrepancies detectable by the evaluation metric.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3403387,
          "author_name": "dankrstev",
          "author_url": "",
          "post_date": "02/08/2026 11:19:30",
          "content": "<p>If I understand correctly, the fixes are done in the metric and the betti matching algorithm? (So we are not penalized for these edge connection aliasing tunnels)</p>",
          "votes": null,
          "replies": [
            {
              "id": 3403388,
              "author_name": "giorgioangelotti",
              "author_url": "",
              "post_date": "02/08/2026 11:22:29",
              "content": "<p>No, the fixes are produced in the test set, so that individual connected components no longer have these artifacts.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3403389,
                  "author_name": "dankrstev",
                  "author_url": "",
                  "post_date": "02/08/2026 11:27:11",
                  "content": "<p>But this issue is with the <strong>predictions submitted</strong> that the metric evaluates. The images above are predictions from a model. As we can see in the 3-axis pyplot that there is <strong>no</strong> actual hole anywhere, yet the metric detected it as a hole. The point is that aliasing is happening on the evaluation level in the C++ code.</p>\n<p>Should this be counted as a hole?\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2F473934adb302db92fb7d7512514982c5%2FScreenshot_4.png?generation=1770550013781740&amp;alt=media\" alt=\"\"></p>\n<p>there are many such examples in a single volume. And its not a problem with the model, we can clearly see that everything is 26 connected and sound.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3403393,
                      "author_name": "giorgioangelotti",
                      "author_url": "",
                      "post_date": "02/08/2026 11:38:14",
                      "content": "<p>Ah, yes, we intentionally wanted to penalize these artifacts. While it's true that we don't care for voxel accurate predictions, ideally we don't want cracks in them.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3403395,
                          "author_name": "dankrstev",
                          "author_url": "",
                          "post_date": "02/08/2026 11:44:36",
                          "content": "<blockquote>\n  <p>… we intentionally wanted to penalize these artifacts …</p>\n  <p>we don't want cracks in them.</p>\n</blockquote>\n<p>But isnt this fully 26 connected? And each voxel is connected to its neighbor by an edge/face so there should be 0 holes or cracks. We can see that the 4 voxels above are fully connected, and the vertex where they meet is causing the issue in the C++.  We are <em>physically</em> incapable of predicting a <strong>better</strong> connection between the voxels. I seem to be missing something by your comment that this is intentionally penalized? </p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3403374": "Apart from the two major issues discussed [here](https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/653482) and [here ](https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/671160) there are more errors that needs the hosts attention. I don't know if it has been discussed before or if the host is aware, but the evaluation algorithm spawns multiple components/holes **even though our 3D sheet is 26 connected**. Is the host aware of this?\n\nThis is concrete example:\n\nthese are Z,Y,X slices of the center coordinate of a \"hole\" that the algorithm found.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2Fd79b5b80a8ced47f4f9b1447e5aa6713%2FScreenshot_3.png?generation=1770544036505032&alt=media)\n\nLike you, I don't see any holes from any of the 3 dimensions point of view. At first I thought that it is some kind of issue with the spurious ending. When inverted to run the births/deaths algorithm somehow is calculated as a hole. However upon inspecting it in 3D and zooming into the crop, we can see what kind of hole the algorithm found:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2F867a52a1c65b8b0bc6045ab565d9c423%2Faliasing3.png?generation=1770544126701329&alt=media)\n\nFrom gemini:\n>In a \"staircase\" of voxels, 26-connectivity (which connects diagonals) allows a path to exist through the corners of the cubes. However, if there is a tiny gap (a 0) underneath the overhang, the algorithm sees a \"hole\" passing through that diagonal gap. This creates a false handle (Betti-1 loop)\n\nSo even though our sheet is completely connected, there are NUMEROUS instances of these \"holes\" detected.\n\nFollowing are some more examples of the **actual** coordinates of the holes/tunnels that the algorithm detected:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2F4ef5940ff8d3745c26a8e0108dc9dab8%2Faliasing2.png?generation=1770544300377788&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2Fd3b07c1d12afdf790ffa7fe67a01e006%2FScreenshot_1.png?generation=1770544311462087&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2F19e8c5c1a84bdbe34027dbc5315e2fa5%2Faliasing.png?generation=1770544678119918&alt=media)\n\n\nAnd some might say that this is not that bad and predicting few extra holes doesnt matter. But if you look closely, the algorithm severely punishes holes - my assumption is that they picked the scaling because they deeply care about holes not being present in the final prediction. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2Fe577a3c552fa59fef1d62ced8382b742%2FScreenshot_2.png?generation=1770545358429775&alt=media)\n\nThese are the hypothetical differences in leaderboard if you predict 1 hole vs 2 holes. Lets say in both scenarios we have some defautl values of 0.8 surface dice and 0.5 VOI. Lets also say we have 0.5 betti-0 score for stable comparison of how holes will affect the score in both cases.\n\n## 1 hole scenario:\n\np_k = 1 <--- we assume 1 hole predicted by our model\n\ng_k = 0 <--- we assume 0 holes in the ground truth\n\n``` \ndenom = p_k + g_k\n            if denom > 0:\n                if g_k != 0:\n                    topoF1_k = (2.0 * m_k) / float(denom)\n                else:\n                    topoF1_k = 0.5 / (float(denom) + 0.5)\n```\ntopoF1_k = 0.5 / (1.0 + 0.5) = 0.3333\n\nTopoScore = (0.5 (betti-0) + 0.3333 (betti-1)) / 2 = 0.416665\n\nScore = 0.30 × TopoScore + 0.35 × SurfaceDice@τ + 0.35 × VOI_score\n\n>(0.3 * 0.416665) + (0.35 * 0.8) + (0.35 * 0.5) =  0.579\n\nWe get 0.579 leaderboard score\n\n## 2 holes scenario:\np_k = 2 <--- we assume 2 holes predicted by our mode\n\ng_k = 0 <--- we assume 0 holes in the ground truth\n\ntopoF1_k = 0.5 / (2.0 + 0.5) = 0.2\n\nTopoScore = (0.5 (betti-0) + 0.2 (betti-1))  / 2 = 0.35\n\nScore = 0.30 × TopoScore + 0.35 × SurfaceDice@τ + 0.35 × VOI_score\n\n>(0.3 * 0.35) + (0.35 * 0.8) + (0.35 * 0.5) =  0.559\n\nWe get 0.559 leaderboard score.\n\nThats 0.02! score difference. We can clearly see that now the leaderboard is not gameable and deeply cares if you do/don't output holes even when the difference is 1 hole. The final solutions will converge to the ones that are most topologically sound, and effort will be shifted from dice score optimization -> topological optimization, since now topology is worth much more, and instead of optimizing dice for marginal gains 0.00X you can get 0.0X improvements as shown in the example above, since now topology is worth more, more time will be spent on that. Few holes difference might've not sounded like a lot, but after doing the calculation you can see how this changes the convergence of the final solutions, as you will be rewarded for *balancing ALL of the terms in the score equation and topology will mater just as much as dice*. Right now topology has little to no say in the final score.\n\nUltimately this is down to the hosts and what they want to get out of the competition. So its up to them to decide if they will fix the issues and ask for competition extension/help from the Kaggle staff, or the competition will conclude as is.",
    "3403383": "Thanks for mentioning this. These are likely artifacts from the voxelization procedure.\nThe fix of the test set that we already produced aimed to tackle the sources of the spurious Betti-1 and Betti-2 discrepancies detectable by the evaluation metric.",
    "3403387": "If I understand correctly, the fixes are done in the metric and the betti matching algorithm? (So we are not penalized for these edge connection aliasing tunnels)",
    "3403388": "No, the fixes are produced in the test set, so that individual connected components no longer have these artifacts.",
    "3403389": "But this issue is with the **predictions submitted** that the metric evaluates. The images above are predictions from a model. As we can see in the 3-axis pyplot that there is **no** actual hole anywhere, yet the metric detected it as a hole. The point is that aliasing is happening on the evaluation level in the C++ code.\n\nShould this be counted as a hole?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2F473934adb302db92fb7d7512514982c5%2FScreenshot_4.png?generation=1770550013781740&alt=media)\n\nthere are many such examples in a single volume. And its not a problem with the model, we can clearly see that everything is 26 connected and sound.",
    "3403393": "Ah, yes, we intentionally wanted to penalize these artifacts. While it's true that we don't care for voxel accurate predictions, ideally we don't want cracks in them.",
    "3403395": "> ... we intentionally wanted to penalize these artifacts ...\n\n>we don't want cracks in them.\n\n\n\nBut isnt this fully 26 connected? And each voxel is connected to its neighbor by an edge/face so there should be 0 holes or cracks. We can see that the 4 voxels above are fully connected, and the vertex where they meet is causing the issue in the C++.  We are *physically* incapable of predicting a **better** connection between the voxels. I seem to be missing something by your comment that this is intentionally penalized?"
  },
  "source": "meta"
}