{
  "id": 671160,
  "title": "[CRITICAL] Why the local validation is invalid, and potentially the whole leaderboard.",
  "url": "/competitions/vesuvius-challenge-surface-detection/discussion/671160",
  "author_name": "dan4o",
  "post_date": "2026-01-31T11:40:47.372000",
  "votes": 55,
  "comment_count": 45,
  "views": 0,
  "content": "<p>There is a <strong>critical error</strong> when calculating the TopoScore on the training labels provided. Some ground truth labels are porous and the topo score is dragged down dramatically by imaginary (porous) holes. In some instances in my local validation and also in the Kaggle metric the scoring function finds over 400! holes in the seemingly perfect ground truths. I tried using both _SPLITS= (1,1,1) and the default (2,2,2) suspecting there is some artifacts when splitting, however with both parameters the scoring function yields enormous number of holes.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2F6bc2f7406041151e965665114bf6dbec%2FScreenshot_4.jpg?generation=1769853751110753&amp;alt=media\" alt=\"\"></p>\n<p>This is a <a href=\"https://www.kaggle.com/code/dankrstev/metric-failure-demo\" target=\"_blank\">notebook </a>which demonstrates <em>that the current leaderboard is also potentially invalid</em> <strong>IF</strong> the test data resembles the training data and they underwent similar processing. Even if your model is predicting perfect sheets that have 0 holes in them the metric will severely punish this as it will compare your 0 holes with the double, or even triple digit holes in the labels and bring your score to 0. Hence the metric score doesn't care if your model predicts good topology and no holes. </p>\n<p>There is a possibility that this is a behavior present <strong>ONLY</strong> in the training set and <strong>NOT</strong> in the test set, in which case only the local validation is invalid. However if they both underwent the same processing (thinning) I highly suspect that this also contaminated the test set. I would ask the hosts and kaggle staff ( <a href=\"https://www.kaggle.com/seanjohnsonsp\" target=\"_blank\">@seanjohnsonsp</a>, <a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a>, <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> ) to confirm by running the simple test shown in the notebook and inform is the leaderboard is indeed calculated with non-porous labels or if this is the reason why so many competitors were demotivated not seeing any improvement on the leaderboard despite having visually superior predictions.</p>\n<p>If the issue persists in the test set data as well, then the hosts should inform us and take the necessary actions (leaderboard rescoring, continue the competition for X more days/weeks to allow competitors to revisit approaches that were invalidated because they worked on faulty metric). Many competitors turned away from potentially good solutions from the fact that they received negative feedback on their topology algorithms and fixes despite having better visual and logical consistency. If the issue is indeed present in the test set, then half of the topology metric is invalid and the competition remains just dice score, which I think the hosts tried to avoid by incorporating the Topology Score in the first place!</p>",
  "messages": [
    {
      "id": 3399828,
      "postDate": "2026-01-31T11:40:47.373Z",
      "content": "<p>There is a <strong>critical error</strong> when calculating the TopoScore on the training labels provided. Some ground truth labels are porous and the topo score is dragged down dramatically by imaginary (porous) holes. In some instances in my local validation and also in the Kaggle metric the scoring function finds over 400! holes in the seemingly perfect ground truths. I tried using both _SPLITS= (1,1,1) and the default (2,2,2) suspecting there is some artifacts when splitting, however with both parameters the scoring function yields enormous number of holes.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2F6bc2f7406041151e965665114bf6dbec%2FScreenshot_4.jpg?generation=1769853751110753&amp;alt=media\" alt=\"\"></p>\n<p>This is a <a href=\"https://www.kaggle.com/code/dankrstev/metric-failure-demo\" target=\"_blank\">notebook </a>which demonstrates <em>that the current leaderboard is also potentially invalid</em> <strong>IF</strong> the test data resembles the training data and they underwent similar processing. Even if your model is predicting perfect sheets that have 0 holes in them the metric will severely punish this as it will compare your 0 holes with the double, or even triple digit holes in the labels and bring your score to 0. Hence the metric score doesn't care if your model predicts good topology and no holes. </p>\n<p>There is a possibility that this is a behavior present <strong>ONLY</strong> in the training set and <strong>NOT</strong> in the test set, in which case only the local validation is invalid. However if they both underwent the same processing (thinning) I highly suspect that this also contaminated the test set. I would ask the hosts and kaggle staff ( <a href=\"https://www.kaggle.com/seanjohnsonsp\" target=\"_blank\">@seanjohnsonsp</a>, <a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a>, <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> ) to confirm by running the simple test shown in the notebook and inform is the leaderboard is indeed calculated with non-porous labels or if this is the reason why so many competitors were demotivated not seeing any improvement on the leaderboard despite having visually superior predictions.</p>\n<p>If the issue persists in the test set data as well, then the hosts should inform us and take the necessary actions (leaderboard rescoring, continue the competition for X more days/weeks to allow competitors to revisit approaches that were invalidated because they worked on faulty metric). Many competitors turned away from potentially good solutions from the fact that they received negative feedback on their topology algorithms and fixes despite having better visual and logical consistency. If the issue is indeed present in the test set, then half of the topology metric is invalid and the competition remains just dice score, which I think the hosts tried to avoid by incorporating the Topology Score in the first place!</p>",
      "rawMarkdown": "There is a **critical error** when calculating the TopoScore on the training labels provided. Some ground truth labels are porous and the topo score is dragged down dramatically by imaginary (porous) holes. In some instances in my local validation and also in the Kaggle metric the scoring function finds over 400! holes in the seemingly perfect ground truths. I tried using both _SPLITS= (1,1,1) and the default (2,2,2) suspecting there is some artifacts when splitting, however with both parameters the scoring function yields enormous number of holes.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2F6bc2f7406041151e965665114bf6dbec%2FScreenshot_4.jpg?generation=1769853751110753&alt=media)\n\nThis is a [notebook ](https://www.kaggle.com/code/dankrstev/metric-failure-demo)which demonstrates *that the current leaderboard is also potentially invalid* **IF** the test data resembles the training data and they underwent similar processing. Even if your model is predicting perfect sheets that have 0 holes in them the metric will severely punish this as it will compare your 0 holes with the double, or even triple digit holes in the labels and bring your score to 0. Hence the metric score doesn't care if your model predicts good topology and no holes. \n\nThere is a possibility that this is a behavior present **ONLY** in the training set and **NOT** in the test set, in which case only the local validation is invalid. However if they both underwent the same processing (thinning) I highly suspect that this also contaminated the test set. I would ask the hosts and kaggle staff ( @seanjohnsonsp, @giorgioangelotti, @sohier ) to confirm by running the simple test shown in the notebook and inform is the leaderboard is indeed calculated with non-porous labels or if this is the reason why so many competitors were demotivated not seeing any improvement on the leaderboard despite having visually superior predictions.\n\nIf the issue persists in the test set data as well, then the hosts should inform us and take the necessary actions (leaderboard rescoring, continue the competition for X more days/weeks to allow competitors to revisit approaches that were invalidated because they worked on faulty metric). Many competitors turned away from potentially good solutions from the fact that they received negative feedback on their topology algorithms and fixes despite having better visual and logical consistency. If the issue is indeed present in the test set, then half of the topology metric is invalid and the competition remains just dice score, which I think the hosts tried to avoid by incorporating the Topology Score in the first place!\n\n\n",
      "votes": 55
    },
    {
      "id": 3402940,
      "postDate": "2026-02-07T08:18:51.457Z",
      "content": "<p>You’re right: we’ve confirmed the dataset contains many microscopic holes/cavities.</p>\n<p>For full transparency: during our pre-release validation of the updated data (before Christmas), an automated check (which I ran) raised this as a potential issue. We discussed it internally, but evidence and opinions in my team were mixed and we ultimately concluded it was likely a false positive (i.e., a limitation/bug in the check rather than a real problem in the data). In hindsight, that conclusion was wrong. As competition host and project lead, I take responsibility for not insisting on an additional independent verification steps and stronger automated topology checks before release.</p>\n<p>We have since regenerated the affected test set and submitted an updated version to Kaggle two days ago. Because changing evaluation data late in a competition has fairness implications, it’s not automatic that it will be applied: Kaggle will decide whether/how/when to use it.</p>",
      "rawMarkdown": "You’re right: we’ve confirmed the dataset contains many microscopic holes/cavities.\n\nFor full transparency: during our pre-release validation of the updated data (before Christmas), an automated check (which I ran) raised this as a potential issue. We discussed it internally, but evidence and opinions in my team were mixed and we ultimately concluded it was likely a false positive (i.e., a limitation/bug in the check rather than a real problem in the data). In hindsight, that conclusion was wrong. As competition host and project lead, I take responsibility for not insisting on an additional independent verification steps and stronger automated topology checks before release.\n\nWe have since regenerated the affected test set and submitted an updated version to Kaggle two days ago. Because changing evaluation data late in a competition has fairness implications, it’s not automatic that it will be applied: Kaggle will decide whether/how/when to use it.",
      "votes": 17,
      "replies": [
        {
          "id": 3402957,
          "postDate": "2026-02-07T08:52:10.297Z",
          "content": "<p>Thank you for the response, and correcting the data. Is it possible for us to know, how much of the test set was affected? And is this just the public test set, or does it include private test set as well? And regarding the corrected data being used for the test set, can we get a confirmation if this data will be used or not (since you mentioned it is up to Kaggle staff whether to use it or not) ?</p>",
          "rawMarkdown": "Thank you for the response, and correcting the data. Is it possible for us to know, how much of the test set was affected? And is this just the public test set, or does it include private test set as well? And regarding the corrected data being used for the test set, can we get a confirmation if this data will be used or not (since you mentioned it is up to Kaggle staff whether to use it or not) ?",
          "votes": 9
        },
        {
          "id": 3402994,
          "postDate": "2026-02-07T10:38:38.453Z",
          "content": "<p><a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a>  Many thanks to the host team for putting so much effort into this competition. I understand that, in order to achieve the project’s goals, the organizers selected a metric that is quite challenging to use as a competition evaluation standard.\nFrom a personal perspective, I hope it would be possible to rescore the results on both the public and private datasets. Our team worked extremely hard to develop what we believe is a strong and elegant approach. We are truly eager for our work to contribute to the Herculaneum Papyri Project, helping extract ancient texts and uncover the knowledge hidden within these historical scrolls.</p>",
          "rawMarkdown": "@giorgioangelotti  Many thanks to the host team for putting so much effort into this competition. I understand that, in order to achieve the project’s goals, the organizers selected a metric that is quite challenging to use as a competition evaluation standard.\nFrom a personal perspective, I hope it would be possible to rescore the results on both the public and private datasets. Our team worked extremely hard to develop what we believe is a strong and elegant approach. We are truly eager for our work to contribute to the Herculaneum Papyri Project, helping extract ancient texts and uncover the knowledge hidden within these historical scrolls.",
          "votes": 3,
          "replies": [
            {
              "id": 3403171,
              "postDate": "2026-02-07T20:07:11.243Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        },
        {
          "id": 3403175,
          "postDate": "2026-02-07T20:12:54.477Z",
          "content": "<blockquote>\n  <p>Because changing evaluation data late in a competition has fairness implications</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a>  There is no need to change the data. There is a fix where the metric is slightly changed on your end, and the leaderboard is rescored. <strong>We set the g_k = 0, and the current score is calculated as if the current data had no holes</strong>. Then the results will be fair for everyone. The more holes your model predicts, the lower your score will be.</p>\n<pre><code>           g_k = 0  ### &lt;--- FIX\n            counts_by_dim[k] = (m_k, p_k, g_k)\n            denom = p_k + g_k\n            if denom &gt; 0:\n                if g_k != 0:\n                    topoF1_k = (2.0 * m_k) / float(denom)\n                else:\n                    topoF1_k = 0.5 / (float(denom) + 0.5) \n                topoF1_by_dim[k] = float(topoF1_k)\n                active_dims.append(k)  \n            else:\n                topoF1_by_dim[k] = float(\"nan\")\n</code></pre>\n<p>This satisfies the fairness aspect of rescoring, as <strong>no</strong> data will be changed. The dynamics will be exactly the same and the models will see exactly the same data. Each teams model will be penalized only for the amount of holes they predicted <em>on the same data as it is now</em>, where teams that predicted better topology will be rewarded. Otherwise the current leaderboard is not truthful.</p>",
          "rawMarkdown": "> Because changing evaluation data late in a competition has fairness implications\n\n@giorgioangelotti  There is no need to change the data. There is a fix where the metric is slightly changed on your end, and the leaderboard is rescored. **We set the g_k = 0, and the current score is calculated as if the current data had no holes**. Then the results will be fair for everyone. The more holes your model predicts, the lower your score will be.\n\n```\n           g_k = 0  ### <--- FIX\n            counts_by_dim[k] = (m_k, p_k, g_k)\n            denom = p_k + g_k\n            if denom > 0:\n                if g_k != 0:\n                    topoF1_k = (2.0 * m_k) / float(denom)\n                else:\n                    topoF1_k = 0.5 / (float(denom) + 0.5) \n                topoF1_by_dim[k] = float(topoF1_k)\n                active_dims.append(k)  \n            else:\n                topoF1_by_dim[k] = float(\"nan\")\n```\n\nThis satisfies the fairness aspect of rescoring, as **no** data will be changed. The dynamics will be exactly the same and the models will see exactly the same data. Each teams model will be penalized only for the amount of holes they predicted *on the same data as it is now*, where teams that predicted better topology will be rewarded. Otherwise the current leaderboard is not truthful.",
          "votes": 3,
          "replies": [
            {
              "id": 3403179,
              "postDate": "2026-02-07T20:24:17.580Z",
              "content": "<p>This is exactly the same as changing the evaluation data. It's not about changing the training data, just the evaluation data as far as I understand. So if you set the g_k=0 or fill all holes, it is pretty much the same. So I doubt that this will satisfy the fairness aspect more than changing the evaluation data. Since this has the same effect on the competitors, a change in LB score.</p>\n<p>But I'd be happy about a fix and extending the competition. Otherwise it is unlikely that the best model (for the task) will win / place high.</p>",
              "rawMarkdown": "This is exactly the same as changing the evaluation data. It's not about changing the training data, just the evaluation data as far as I understand. So if you set the g_k=0 or fill all holes, it is pretty much the same. So I doubt that this will satisfy the fairness aspect more than changing the evaluation data. Since this has the same effect on the competitors, a change in LB score.\n\nBut I'd be happy about a fix and extending the competition. Otherwise it is unlikely that the best model (for the task) will win / place high."
            },
            {
              "id": 3403213,
              "postDate": "2026-02-07T22:44:02.783Z",
              "content": "<blockquote>\n  <p>But I'd be happy about a fix and extending the competition. Otherwise it is unlikely that the best model (for the task) will win / place high.</p>\n</blockquote>\n<p>I totally agree. Its kinda strange if there is no action taken <em>even when we know</em> that there is an error in the evaluation of the solutions. The actual reading of the scrolls doesn't care about if we are fair or not fair, it only cares about having the best solution (topologically sound) so it can work downstream. </p>",
              "rawMarkdown": "> But I'd be happy about a fix and extending the competition. Otherwise it is unlikely that the best model (for the task) will win / place high.\n\nI totally agree. Its kinda strange if there is no action taken *even when we know* that there is an error in the evaluation of the solutions. The actual reading of the scrolls doesn't care about if we are fair or not fair, it only cares about having the best solution (topologically sound) so it can work downstream. ",
              "votes": 1
            }
          ]
        },
        {
          "id": 3403303,
          "postDate": "2026-02-08T07:19:21.973Z",
          "content": "<p>Will old submissions be re-evaluated?</p>",
          "rawMarkdown": "Will old submissions be re-evaluated?"
        },
        {
          "id": 3403715,
          "postDate": "2026-02-09T05:33:41.393Z",
          "content": "<p>That’s why my notebook was penalized it was being treated as a false positive.\nI did join very late, though, so let’s see how it goes.\nNow i dont have time to retrain the model again</p>",
          "rawMarkdown": "That’s why my notebook was penalized it was being treated as a false positive.\nI did join very late, though, so let’s see how it goes.\nNow i dont have time to retrain the model again"
        },
        {
          "id": 3403825,
          "postDate": "2026-02-09T11:01:04.273Z",
          "content": "<blockquote>\n  <p>Kaggle will decide whether/how/when to use it.</p>\n</blockquote>\n<p>You are the one to make that decision, you are the host. The rules are a contract between you and us, not between Kaggle and us.</p>\n<p>If there is a change in test data then the competition should be extended. People have spent weeks, months to improve their score using a given test harness. Changing that is totally unfair to them, and they should have the opportunity to adjust what they did. But extending the competition also has its drawbacks, and will probably be very badly received.</p>\n<p>TL;DR Changing the metric or test data in the last week should not happen.</p>\n<p>I also found another type of annotation issue, I'll share privately as I don't want to share new ideas publicly during the last week. No one should share new ideas during the last week actually.</p>",
          "rawMarkdown": "> Kaggle will decide whether/how/when to use it.\n\nYou are the one to make that decision, you are the host. The rules are a contract between you and us, not between Kaggle and us.\n\nIf there is a change in test data then the competition should be extended. People have spent weeks, months to improve their score using a given test harness. Changing that is totally unfair to them, and they should have the opportunity to adjust what they did. But extending the competition also has its drawbacks, and will probably be very badly received.\n\nTL;DR Changing the metric or test data in the last week should not happen.\n\nI also found another type of annotation issue, I'll share privately as I don't want to share new ideas publicly during the last week. No one should share new ideas during the last week actually.",
          "votes": 10,
          "replies": [
            {
              "id": 3403847,
              "postDate": "2026-02-09T12:16:38.550Z",
              "content": "<blockquote>\n  <p>If there is a change in test data then the competition should be extended. People have spent weeks, months to improve their score using a given test harness.</p>\n</blockquote>\n<p>I agree. </p>\n<p>My personal opinion is this. I understand that its late, but a fix is needed.</p>\n<p>The current leaderboard doesn't reflect the true goals of this competition. Because participants optimize their models specifically for the provided metric—which unfortunatelly is flawed—the best solutions for topology, Dice, and VOI won't actually rise to the top. </p>\n<p>If the error remains unfixed, those solutions that would've scored higher on what the host wanted, might not surface in the top prize slots and teams won't bother to share their methodology or code, even when in a parallel universe where the metric worked as intended they would've solved the problem better (which is the universe where we should be evaluating in, and where the problem of reading the papyrus actually exists). As is, the host will receive the best solutions that maximize the current score on the metric — whether that metric is flawed or not.</p>",
              "rawMarkdown": ">If there is a change in test data then the competition should be extended. People have spent weeks, months to improve their score using a given test harness.\n\nI agree. \n\nMy personal opinion is this. I understand that its late, but a fix is needed.\n\nThe current leaderboard doesn't reflect the true goals of this competition. Because participants optimize their models specifically for the provided metric—which unfortunatelly is flawed—the best solutions for topology, Dice, and VOI won't actually rise to the top. \n\nIf the error remains unfixed, those solutions that would've scored higher on what the host wanted, might not surface in the top prize slots and teams won't bother to share their methodology or code, even when in a parallel universe where the metric worked as intended they would've solved the problem better (which is the universe where we should be evaluating in, and where the problem of reading the papyrus actually exists). As is, the host will receive the best solutions that maximize the current score on the metric — whether that metric is flawed or not.",
              "votes": 8
            },
            {
              "id": 3404151,
              "postDate": "2026-02-09T23:57:46.337Z",
              "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8722753%2Fc729c2f5ef77b01f0900085db719661c%2FSin%20ttulo.jpg?generation=1770681463824241&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8722753%2Fc729c2f5ef77b01f0900085db719661c%2FSin%20ttulo.jpg?generation=1770681463824241&alt=media)"
            }
          ]
        }
      ]
    },
    {
      "id": 3402376,
      "postDate": "2026-02-05T22:26:17.200Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/dankrstev\" target=\"_blank\">@dankrstev</a> -- thank you for surfacing this. The host is looking into it and we will provide an update soon. Stay tuned!</p>",
      "rawMarkdown": "Hi @dankrstev -- thank you for surfacing this. The host is looking into it and we will provide an update soon. Stay tuned!",
      "votes": 7,
      "replies": [
        {
          "id": 3402384,
          "postDate": "2026-02-05T23:23:06.747Z",
          "content": "<p>Its really late already, merging and entry ends tommorow</p>\n<p>And retraining might take couple of days for some of us, not sure if recalculating the metric again this time would really be a good idea</p>",
          "rawMarkdown": "Its really late already, merging and entry ends tommorow\n\nAnd retraining might take couple of days for some of us, not sure if recalculating the metric again this time would really be a good idea"
        },
        {
          "id": 3402402,
          "postDate": "2026-02-06T01:43:20.307Z",
          "content": "<p>Can we reach a conclusion today? Should we wait until a conclusion is reached before submitting the results?</p>",
          "rawMarkdown": "Can we reach a conclusion today? Should we wait until a conclusion is reached before submitting the results?",
          "votes": 5
        },
        {
          "id": 3402434,
          "postDate": "2026-02-06T04:12:46.580Z",
          "content": "<p>Will the competition’s end time being postponed? It’s too late—we’ve already wasted a lot of computing resources</p>",
          "rawMarkdown": "Will the competition’s end time being postponed? It’s too late—we’ve already wasted a lot of computing resources",
          "votes": 1
        },
        {
          "id": 3402717,
          "postDate": "2026-02-06T15:56:11.437Z",
          "content": "<p>Score can be re-scored. It doesn't matter if the competition will end soon. Its just the score and not the dataset. So,  it will not valid concern if anyone thinks score should not be rescored. </p>",
          "rawMarkdown": "Score can be re-scored. It doesn't matter if the competition will end soon. Its just the score and not the dataset. So,  it will not valid concern if anyone thinks score should not be rescored. ",
          "votes": -2
        }
      ]
    },
    {
      "id": 3404168,
      "postDate": "2026-02-10T01:01:57.423Z",
      "content": "<p>My first reaction was, “Why didn’t I catch this earlier, and what can I learn from it?”</p>\n<p>1) I realize I didn’t analyze the failure cases deeply enough. The code and concepts felt difficult, so I have skipped the details in the Betti matching part.</p>\n<p>2) As a result, some of my experiments were actually counterproductive. For example, when I replaced my prediction with the merged prediction and ground truth, the TopoScore became worse. I should have investigated this more carefully.</p>\n<p>As a machine learning practitioner, this is a very valuable lesson for me.</p>",
      "rawMarkdown": "My first reaction was, “Why didn’t I catch this earlier, and what can I learn from it?”\n\n1) I realize I didn’t analyze the failure cases deeply enough. The code and concepts felt difficult, so I have skipped the details in the Betti matching part.\n\n2) As a result, some of my experiments were actually counterproductive. For example, when I replaced my prediction with the merged prediction and ground truth, the TopoScore became worse. I should have investigated this more carefully.\n\nAs a machine learning practitioner, this is a very valuable lesson for me.",
      "votes": 6,
      "replies": [
        {
          "id": 3404182,
          "postDate": "2026-02-10T02:09:33.077Z",
          "content": "<p>There are even more problems with the current metric. I am writing a fix for the host to the issue you raised <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/653482\" target=\"_blank\">here</a>. </p>\n<p>It is <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/672634#3404181\" target=\"_blank\">here</a>, please review it and look at it from another perspective - I hope I didnt miss some obvious scenario where the fix would fail.</p>",
          "rawMarkdown": "There are even more problems with the current metric. I am writing a fix for the host to the issue you raised [here](https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/653482). \n\nIt is [here](https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/672634#3404181), please review it and look at it from another perspective - I hope I didnt miss some obvious scenario where the fix would fail."
        }
      ]
    },
    {
      "id": 3399896,
      "postDate": "2026-01-31T13:38:58.437Z",
      "content": "<p>Thank you for bringing this forward. I was also analyzing topo score since it didn't seem to add up. I have another weird situation, while both predictions are very similar (just tiny difference in post processing). I get for the same component a topo score of either 0.0 or 1.0. I suspected that there is a tiny voxel hole i just didn't see, but now I ran it with your notebook and got the following:</p>\n<p>The perfect one:\ntopo=TopoReport(toposcore=1.0, topoF1_by_dim={0: 1.0, 1: nan, 2: nan}, counts_by_dim={0: (1, 1, 1), 1: (0, 0, 0), 2: (0, 0, 0)}, dims=[0, 1, 2])</p>\n<p>The failure:\ntopo=TopoReport(toposcore=0.0, topoF1_by_dim={0: 0.0, 1: nan, 2: nan}, counts_by_dim={0: (0, 0, 1), 1: (0, 0, 0), 2: (0, 0, 0)}, dims=[0, 1, 2])</p>\n<p>How is it possible to get 0 on prediction in dim=0? \nMy prediction isn't empty, it is even very similar to the one which works.</p>\n<p>Here the visuals:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20325352%2F13929200bf07ccdd0e2fe6c2820cbdd9%2Ftopo.png?generation=1769866721033758&amp;alt=media\" alt=\"\"></p>\n<p>Here is the dataset:\n<a href=\"https://www.kaggle.com/datasets/mariusheuser/single-component-result/\" target=\"_blank\">https://www.kaggle.com/datasets/mariusheuser/single-component-result/</a></p>\n<p>The only way I see this happening is, if my prediction was invalid for some reason. But the dice score is valid (very high) for both predictions.</p>",
      "rawMarkdown": "Thank you for bringing this forward. I was also analyzing topo score since it didn't seem to add up. I have another weird situation, while both predictions are very similar (just tiny difference in post processing). I get for the same component a topo score of either 0.0 or 1.0. I suspected that there is a tiny voxel hole i just didn't see, but now I ran it with your notebook and got the following:\n\nThe perfect one:\ntopo=TopoReport(toposcore=1.0, topoF1_by_dim={0: 1.0, 1: nan, 2: nan}, counts_by_dim={0: (1, 1, 1), 1: (0, 0, 0), 2: (0, 0, 0)}, dims=[0, 1, 2])\n\nThe failure:\ntopo=TopoReport(toposcore=0.0, topoF1_by_dim={0: 0.0, 1: nan, 2: nan}, counts_by_dim={0: (0, 0, 1), 1: (0, 0, 0), 2: (0, 0, 0)}, dims=[0, 1, 2])\n\nHow is it possible to get 0 on prediction in dim=0? \nMy prediction isn't empty, it is even very similar to the one which works.\n\nHere the visuals:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20325352%2F13929200bf07ccdd0e2fe6c2820cbdd9%2Ftopo.png?generation=1769866721033758&alt=media)\n\nHere is the dataset:\n[https://www.kaggle.com/datasets/mariusheuser/single-component-result/](https://www.kaggle.com/datasets/mariusheuser/single-component-result/)\n\nThe only way I see this happening is, if my prediction was invalid for some reason. But the dice score is valid (very high) for both predictions.",
      "votes": 5
    },
    {
      "id": 3399961,
      "postDate": "2026-01-31T15:59:38.663Z",
      "content": "<p><a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a>  is there  any update regarding this</p>",
      "rawMarkdown": "@giorgioangelotti  is there  any update regarding this",
      "votes": 4
    },
    {
      "id": 3402597,
      "postDate": "2026-02-06T11:54:07.360Z",
      "content": "<p>Haha, when I previously used the cases in deprecated_train_images as external test data, I couldn't understand why models with stronger mask connectivity (fewer holes) resulted in lower LB scores. Thanks for helping me realize this—let’s just let the model play a game of 'lucky hole-hitting' then~</p>",
      "rawMarkdown": "Haha, when I previously used the cases in deprecated_train_images as external test data, I couldn't understand why models with stronger mask connectivity (fewer holes) resulted in lower LB scores. Thanks for helping me realize this—let’s just let the model play a game of 'lucky hole-hitting' then~",
      "votes": 2
    },
    {
      "id": 3404338,
      "postDate": "2026-02-10T10:40:35.160Z",
      "content": "<p>I will continue reading related topics about that. But if someone can confirm or fix my underastanding till now I will really appreciate it. </p>\n<ul>\n<li><p>Metric penalizes holes in masks, but training actual masks have already holes.</p></li>\n<li><p>Test has been fixed, so holes in masks have been removed, but training remains with them.</p></li>\n<li><p>We should fix holes in train masks by our side before properly train.</p></li>\n</ul>",
      "rawMarkdown": "I will continue reading related topics about that. But if someone can confirm or fix my underastanding till now I will really appreciate it. \n\n* Metric penalizes holes in masks, but training actual masks have already holes.\n\n* Test has been fixed, so holes in masks have been removed, but training remains with them.\n\n* We should fix holes in train masks by our side before properly train.",
      "votes": 1
    },
    {
      "id": 3404221,
      "postDate": "2026-02-10T04:16:16.563Z",
      "content": "<p>you filled by public code given by the author:</p>\n<pre><code>def fill_small_holes_by_closing(mask: np.ndarray, iters: int = 1, conn: int = 3):\n    \"\"\"\n    Fast but can create bridges/handles. Use carefully.\n    conn: 1-&gt;6 neigh, 2-&gt;18, 3-&gt;26\n    \"\"\"\n    m = mask.astype(bool)\n    struct = ndi.generate_binary_structure(3, conn)\n    closed = ndi.binary_closing(m, structure=struct, iterations=iters)\n    return closed.astype(np.uint8)\n</code></pre>\n<p>only connectivity=26 works (6 and 16 don't)<br>\nbut voxels are added to other normal \"non-hole surface\" as well. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F38d7fe7f157b8f94dca09db302b61ad7%2FSelection_2430.png?generation=1770696868259779&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F30ee7e9ffcd9da77d72f4d8ce023f464%2FSelection_2432.png?generation=1770696890377629&amp;alt=media\" alt=\"\"></p>\n<pre><code>ID 1294570892 : 8 components\n\n#holes_after, _ = dim1_holes_self(gt_one)\n  1294570892 lb 1 fg 337197 has dim1 holes = 399\n\n#holes_after, _ = dim1_holes_self(filled)\n    changed voxels: 5147\n    after filling: 1294570892 lb 1 has dim1 holes = 0\n\n#r_cmp = do_one_lb(filled, gt_one)\n    compare filled vs gt_one:\n      lb=0.688566  surfDice=1.000000  VOI=0.967332  topo=0.000000\n      topo details: TopoReport(toposcore=0.0, topoF1_by_dim={0: nan, 1: 0.0, 2: nan}, counts_by_dim={0: (0, 0, 0), 1: (0, 0, 399), 2: (0, 0, 0)}, dims=[0, 1, 2])\n</code></pre>",
      "rawMarkdown": "you filled by public code given by the author:\n\n```\ndef fill_small_holes_by_closing(mask: np.ndarray, iters: int = 1, conn: int = 3):\n    \"\"\"\n    Fast but can create bridges/handles. Use carefully.\n    conn: 1->6 neigh, 2->18, 3->26\n    \"\"\"\n    m = mask.astype(bool)\n    struct = ndi.generate_binary_structure(3, conn)\n    closed = ndi.binary_closing(m, structure=struct, iterations=iters)\n    return closed.astype(np.uint8)\n\n```\n\nonly connectivity=26 works (6 and 16 don't)   \nbut voxels are added to other normal \"non-hole surface\" as well. \n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F38d7fe7f157b8f94dca09db302b61ad7%2FSelection_2430.png?generation=1770696868259779&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F30ee7e9ffcd9da77d72f4d8ce023f464%2FSelection_2432.png?generation=1770696890377629&alt=media)\n\n\n```\nID 1294570892 : 8 components\n\n#holes_after, _ = dim1_holes_self(gt_one)\n  1294570892 lb 1 fg 337197 has dim1 holes = 399\n\n#holes_after, _ = dim1_holes_self(filled)\n    changed voxels: 5147\n    after filling: 1294570892 lb 1 has dim1 holes = 0\n\n#r_cmp = do_one_lb(filled, gt_one)\n    compare filled vs gt_one:\n      lb=0.688566  surfDice=1.000000  VOI=0.967332  topo=0.000000\n      topo details: TopoReport(toposcore=0.0, topoF1_by_dim={0: nan, 1: 0.0, 2: nan}, counts_by_dim={0: (0, 0, 0), 1: (0, 0, 399), 2: (0, 0, 0)}, dims=[0, 1, 2])\n\n\n```",
      "votes": 1
    },
    {
      "id": 3399988,
      "postDate": "2026-01-31T16:40:53.777Z",
      "content": "<p>gt id 3742893488 visualization result from different angles :<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15801225%2Fa112911e8b0800beabcae7aa8700e90b%2Fnewplot.png?generation=1769877606266448&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15801225%2Ff942d8fbf133584856e7d1ff0771524a%2Fnewplot%20(1).png?generation=1769877611737440&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15801225%2F0d5d9d855577aac2af6fe4b292898b04%2Fnewplot%20(2).png?generation=1769877636618099&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "gt id 3742893488 visualization result from different angles :![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15801225%2Fa112911e8b0800beabcae7aa8700e90b%2Fnewplot.png?generation=1769877606266448&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15801225%2Ff942d8fbf133584856e7d1ff0771524a%2Fnewplot%20(1).png?generation=1769877611737440&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15801225%2F0d5d9d855577aac2af6fe4b292898b04%2Fnewplot%20(2).png?generation=1769877636618099&alt=media)",
      "votes": 1
    },
    {
      "id": 3404217,
      "postDate": "2026-02-10T04:04:19.300Z",
      "content": "<p>the holes are actually surface discretization error:</p>\n<p>ID 1294570892 : 8 components<br>\n  1294570892 lb 1 fg 337197 has dim1 holes = 399   </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F775c94b7156c68d215d732b9bc521f41%2FSelection_2431.png?generation=1770696170658795&amp;alt=media\" alt=\"\"></p>\n<p>some are \"visible\", but some are not (e.g. truth instance with only 1,2,4 10 holes)</p>\n<pre><code>def dim1_holes_self(mask_bin: np.ndarray):\n    \"\"\"\n    Compute dim1 count using the competition topo scorer, by scoring mask vs itself.\n    For self-vs-self, counts_by_dim[1] entries should be equal, so max() is safe.\n    \"\"\"\n    r = do_one_lb(mask_bin, mask_bin)\n    c = r[\"more_topo\"].counts_by_dim[1]   # tuple-like\n    return int(max(c)), r\n\n\n#---------------------------------------------------------------\n        gt = tifffile.imread(f\"{kaggle_dir}/train_labels/{id}.tif\") \n        fg = (gt == 1)\n\n        cc = cc3d.connected_components(fg, connectivity=26)\n        ncomp = int(cc.max())\n        print(f\"\\nID {id} : {ncomp} components\")\n\n        for lb in range(1, ncomp + 1):\n            gt_one = (cc == lb).astype(np.uint8)\n            fg_count = int(gt_one.sum())\n            if fg_count &lt; min_fg:\n                continue\n\n            holes_before, rep_self = dim1_holes_self(gt_one)\n            print(f\"  {id} lb {lb} fg {fg_count} has dim1 holes = {holes_before}\")\n</code></pre>",
      "rawMarkdown": "the holes are actually surface discretization error:\n\nID 1294570892 : 8 components   \n  1294570892 lb 1 fg 337197 has dim1 holes = 399   \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F775c94b7156c68d215d732b9bc521f41%2FSelection_2431.png?generation=1770696170658795&alt=media)\n\nsome are \"visible\", but some are not (e.g. truth instance with only 1,2,4 10 holes)\n\n\n```\n\ndef dim1_holes_self(mask_bin: np.ndarray):\n    \"\"\"\n    Compute dim1 count using the competition topo scorer, by scoring mask vs itself.\n    For self-vs-self, counts_by_dim[1] entries should be equal, so max() is safe.\n    \"\"\"\n    r = do_one_lb(mask_bin, mask_bin)\n    c = r[\"more_topo\"].counts_by_dim[1]   # tuple-like\n    return int(max(c)), r\n\n\n#---------------------------------------------------------------\n        gt = tifffile.imread(f\"{kaggle_dir}/train_labels/{id}.tif\") \n        fg = (gt == 1)\n\n        cc = cc3d.connected_components(fg, connectivity=26)\n        ncomp = int(cc.max())\n        print(f\"\\nID {id} : {ncomp} components\")\n\n        for lb in range(1, ncomp + 1):\n            gt_one = (cc == lb).astype(np.uint8)\n            fg_count = int(gt_one.sum())\n            if fg_count < min_fg:\n                continue\n\n            holes_before, rep_self = dim1_holes_self(gt_one)\n            print(f\"  {id} lb {lb} fg {fg_count} has dim1 holes = {holes_before}\")\n\n```",
      "votes": 2,
      "replies": [
        {
          "id": 3404327,
          "postDate": "2026-02-10T10:15:27.553Z",
          "content": "<p>did you take time to review the other problem found by <a href=\"https://www.kaggle.com/dankrstev\" target=\"_blank\">@dankrstev</a> <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/672447\" target=\"_blank\">discussion</a>?</p>\n<blockquote>\n  <p>And each voxel is connected to its neighbor by an edge/face so there should be 0 holes or cracks. We can see that the 4 voxels above are fully connected, and the vertex where they meet is causing the issue in the C++. We are physically incapable of predicting a better connection between the voxels.</p>\n</blockquote>",
          "rawMarkdown": "did you take time to review the other problem found by @dankrstev [discussion](https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/672447)?\n\n> And each voxel is connected to its neighbor by an edge/face so there should be 0 holes or cracks. We can see that the 4 voxels above are fully connected, and the vertex where they meet is causing the issue in the C++. We are physically incapable of predicting a better connection between the voxels."
        }
      ]
    },
    {
      "id": 3399932,
      "postDate": "2026-01-31T14:44:57.927Z",
      "content": "<p>Indeed, this critical error had been <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/668313#3393314\" target=\"_blank\">discussed</a> before. However, it's disappointed that there are not any change yet (change the metric is also unpracticable as the deadline is coming). I think this issue will lead to the huge randomness and the huge shaking is coming in the way. </p>",
      "rawMarkdown": "Indeed, this critical error had been [discussed](https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/668313#3393314) before. However, it's disappointed that there are not any change yet (change the metric is also unpracticable as the deadline is coming). I think this issue will lead to the huge randomness and the huge shaking is coming in the way. ",
      "replies": [
        {
          "id": 3399934,
          "postDate": "2026-01-31T14:50:28.423Z",
          "content": "<p>This has nothing to do with the ignore label discussed in the post you linked. The issue in that discussion is about ignore mask producing more connected components by trimming the prediction. </p>\n<p>This post is about the ground truth having <strong>wrong number of porous holes</strong> that bring the TopoScore down regardless of what you do or how you fix your holes. Please read the post and the notebook.</p>",
          "rawMarkdown": "This has nothing to do with the ignore label discussed in the post you linked. The issue in that discussion is about ignore mask producing more connected components by trimming the prediction. \n\nThis post is about the ground truth having **wrong number of porous holes** that bring the TopoScore down regardless of what you do or how you fix your holes. Please read the post and the notebook.",
          "votes": 2
        },
        {
          "id": 3399935,
          "postDate": "2026-01-31T14:52:10.647Z",
          "content": "<p>That's a different issue. Here the GT is compared with the GT, so the ignore label shouldn't influence/change the prediction/result.</p>",
          "rawMarkdown": "That's a different issue. Here the GT is compared with the GT, so the ignore label shouldn't influence/change the prediction/result.",
          "votes": 3
        },
        {
          "id": 3399958,
          "postDate": "2026-01-31T15:54:14.537Z",
          "content": "<p>Oh, I am sorry for misunderstanding. In fact, I also find a few cases contrain betti-1 error few days ago which should not be expected in training data labels. I run your code on 30 random data just now, even 20 of 30 contains unexpected betti-1 error. And 3 of 30 even contains more than 50 betti-1 error. The betti-1 score is meaningless if the test set is also like this.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7796170%2F9a76d98c51d323fa70fea8335769bcea%2Fa6a7fffe-a398-47aa-8f22-8c13f4b7473a.png?generation=1769874741171850&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Oh, I am sorry for misunderstanding. In fact, I also find a few cases contrain betti-1 error few days ago which should not be expected in training data labels. I run your code on 30 random data just now, even 20 of 30 contains unexpected betti-1 error. And 3 of 30 even contains more than 50 betti-1 error. The betti-1 score is meaningless if the test set is also like this.![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7796170%2F9a76d98c51d323fa70fea8335769bcea%2Fa6a7fffe-a398-47aa-8f22-8c13f4b7473a.png?generation=1769874741171850&alt=media)",
          "replies": [
            {
              "id": 3399962,
              "postDate": "2026-01-31T16:00:42.700Z",
              "content": "<p>I'm not sure if I understood your post correctly, but the \"NaN\" values are supposed to be there. That simply means that betti-1 is ignored for the calculation, just like betti-2.\nThis is expected, since most GT labels have zero holes in them.\nThe 450, 413 are clearly errors and the others depend on the specific case.</p>",
              "rawMarkdown": "I'm not sure if I understood your post correctly, but the \"NaN\" values are supposed to be there. That simply means that betti-1 is ignored for the calculation, just like betti-2.\nThis is expected, since most GT labels have zero holes in them.\nThe 450, 413 are clearly errors and the others depend on the specific case.",
              "votes": 2
            },
            {
              "id": 3399969,
              "postDate": "2026-01-31T16:14:00.323Z",
              "content": "<p>I mean topoF1_by_dim in the figure should be {0:1, 1:nan, 2:nan} since the gt_label should not contain any betti-1 error. </p>",
              "rawMarkdown": "I mean topoF1_by_dim in the figure should be {0:1, 1:nan, 2:nan} since the gt_label should not contain any betti-1 error. "
            },
            {
              "id": 3399974,
              "postDate": "2026-01-31T16:20:00.747Z",
              "content": "<p>And The 450 and 413 I think it is true label error. But as the figure show, there are other cases which betti-1 number are less than 10. I am not sure are these betti-1 number is caused by the label itself or the way of betti number calculation (by chunks). But anyway, it means at least this kind of calculation of the labels  would lead to the most gt_label contains unexpected betti-1 numbers which need to be clarifed by hosts. </p>",
              "rawMarkdown": "And The 450 and 413 I think it is true label error. But as the figure show, there are other cases which betti-1 number are less than 10. I am not sure are these betti-1 number is caused by the label itself or the way of betti number calculation (by chunks). But anyway, it means at least this kind of calculation of the labels  would lead to the most gt_label contains unexpected betti-1 numbers which need to be clarifed by hosts. "
            },
            {
              "id": 3399976,
              "postDate": "2026-01-31T16:29:10.493Z",
              "content": "<p>Yes that is correct. If there are 0 predicted, and 0 ground truth holes then betti-1 should be nan for that inactive dimension. However it is not nan because we are \"predicting\" a lot of holes when using the ground truth as our prediction. </p>\n<p>The algorithm finds the holes and it \"matches\" them correctly, thinking that we sucessfully found the correct holes in the ground truth. </p>\n<p>However when you use your own real model prediction, NONE of these holes will be correctly matched so you will end up with (0, YOUR_MODEL_HOLES, FAKE_GT_HOLES) where FAKE_GT_HOLES can be a large number. Even if your model predicts 0 holes, the corrupted ground truth holes will contribute in the denominator as seen below.</p>\n<p>The three numbers are (matched_k, predicted_k, ground_k).\nThis is the code that provides the counts_by_dim:</p>\n<pre><code>for k in dims_list:\n            m_k, p_k, g_k = cls._counts_for_dim(result, k)\n            counts_by_dim[k] = (m_k, p_k, g_k)\n            denom = p_k + g_k\n            if denom &gt; 0:\n                if g_k != 0:\n                    topoF1_k = (2.0 * m_k) / float(denom)\n                else:\n                    topoF1_k = 0.5 / (float(denom) + 0.5) # adding pseudocount to avoid 0\n                topoF1_by_dim[k] = float(topoF1_k)\n                active_dims.append(k)  \n            else:\n                topoF1_by_dim[k] = float(\"nan\")  # keep NaN for inactive\n</code></pre>\n<p>Most of the time during score calculation we are in this regime:</p>\n<pre><code>if g_k != 0:\n     topoF1_k = (2.0 * m_k) / float(denom)\n</code></pre>\n<p>which results in (2.0 * 0) / (your_model_holes + big_number_fake_holes) = 0.0….~</p>\n<p>Whereas if the sheets should be trully without holes we should be in this regime:</p>\n<pre><code>else:\n     topoF1_k = 0.5 / (float(denom) + 0.5) # adding pseudocount to avoid 0\n</code></pre>\n<p>where the more holes your model predicts the score decreases since its the only thing affecting the denominator denom = p_k + g_k(0)</p>",
              "rawMarkdown": "Yes that is correct. If there are 0 predicted, and 0 ground truth holes then betti-1 should be nan for that inactive dimension. However it is not nan because we are \"predicting\" a lot of holes when using the ground truth as our prediction. \n\nThe algorithm finds the holes and it \"matches\" them correctly, thinking that we sucessfully found the correct holes in the ground truth. \n\nHowever when you use your own real model prediction, NONE of these holes will be correctly matched so you will end up with (0, YOUR_MODEL_HOLES, FAKE_GT_HOLES) where FAKE_GT_HOLES can be a large number. Even if your model predicts 0 holes, the corrupted ground truth holes will contribute in the denominator as seen below.\n\nThe three numbers are (matched_k, predicted_k, ground_k).\nThis is the code that provides the counts_by_dim:\n\n```\nfor k in dims_list:\n            m_k, p_k, g_k = cls._counts_for_dim(result, k)\n            counts_by_dim[k] = (m_k, p_k, g_k)\n            denom = p_k + g_k\n            if denom > 0:\n                if g_k != 0:\n                    topoF1_k = (2.0 * m_k) / float(denom)\n                else:\n                    topoF1_k = 0.5 / (float(denom) + 0.5) # adding pseudocount to avoid 0\n                topoF1_by_dim[k] = float(topoF1_k)\n                active_dims.append(k)  \n            else:\n                topoF1_by_dim[k] = float(\"nan\")  # keep NaN for inactive\n```\n\nMost of the time during score calculation we are in this regime:\n```\nif g_k != 0:\n     topoF1_k = (2.0 * m_k) / float(denom)\n```\n\nwhich results in (2.0 * 0) / (your_model_holes + big_number_fake_holes) = 0.0....~\n\nWhereas if the sheets should be trully without holes we should be in this regime:\n```\nelse:\n     topoF1_k = 0.5 / (float(denom) + 0.5) # adding pseudocount to avoid 0\n```\nwhere the more holes your model predicts the score decreases since its the only thing affecting the denominator denom = p_k + g_k(0)",
              "votes": 2
            },
            {
              "id": 3399994,
              "postDate": "2026-01-31T16:51:10.947Z",
              "content": "<p>I suppose there shouldn't be any betti-1 number in gt_labels. If we treat the gt_label as the prediction the algorithms should also find 0 holes. And the topo score should be \"nan\" since the prediction and label are both 0. But as the figure show, the evaluation code find some holes and lead to the betti-1 number ≠ 0, which is also a  issue I think besides the true label error (hundreds of holes). </p>\n<p>Given such a gt_label of which the betti-1 number is not 0 (or at least the calculation of betti-1 number ≠0 using the evaluation code), suppose we have a perfect model which can prediction 0 holes of the whole volume. Even if we get such a good prediction, the topo-1 score is calculated by (0, 0, FAKE_GT_HOLES) or (0, FAKE_PREDICTION_HOLES, FAKE_GT_HOLES) which equal to 0. </p>",
              "rawMarkdown": "I suppose there shouldn't be any betti-1 number in gt_labels. If we treat the gt_label as the prediction the algorithms should also find 0 holes. And the topo score should be \"nan\" since the prediction and label are both 0. But as the figure show, the evaluation code find some holes and lead to the betti-1 number ≠ 0, which is also a  issue I think besides the true label error (hundreds of holes). \n\nGiven such a gt_label of which the betti-1 number is not 0 (or at least the calculation of betti-1 number ≠0 using the evaluation code), suppose we have a perfect model which can prediction 0 holes of the whole volume. Even if we get such a good prediction, the topo-1 score is calculated by (0, 0, FAKE_GT_HOLES) or (0, FAKE_PREDICTION_HOLES, FAKE_GT_HOLES) which equal to 0. ",
              "votes": 2
            },
            {
              "id": 3399998,
              "postDate": "2026-01-31T17:02:26.637Z",
              "content": "<p>Here's a specific case:  [0:160, 160:320, 0:160] of 1075217434.tif. The betti-1 number are found in [30, 0, 5] and [100, 0, 4] using the betti_matching code. Seems that in the edge of the slice. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7796170%2Fed322316960b8602c3442b3c2e8855ac%2F4429c755-cbd7-498d-952f-e2a476d072c2.png?generation=1769878895880181&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "Here's a specific case:  [0:160, 160:320, 0:160] of 1075217434.tif. The betti-1 number are found in [30, 0, 5] and [100, 0, 4] using the betti_matching code. Seems that in the edge of the slice. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7796170%2Fed322316960b8602c3442b3c2e8855ac%2F4429c755-cbd7-498d-952f-e2a476d072c2.png?generation=1769878895880181&alt=media)",
              "votes": 1
            },
            {
              "id": 3400001,
              "postDate": "2026-01-31T17:08:02.083Z",
              "content": "<p>Yes, I also included in the description of the problem what kind of problem this is (porous holes), which can be fixed by BinaryClose as shown in the Notebook. We are waiting now on the Hosts to confirm if this problem is in the training labels alone (which I don't believe) or in the public/private tests sets as well (which I suspect). Since the thinning of the labels should happen with the same code, I suspect that there is when the error happened and it most likely propagated to the private/public data as well - hence why nobody sees clear improvements on the topological side.</p>",
              "rawMarkdown": "Yes, I also included in the description of the problem what kind of problem this is (porous holes), which can be fixed by BinaryClose as shown in the Notebook. We are waiting now on the Hosts to confirm if this problem is in the training labels alone (which I don't believe) or in the public/private tests sets as well (which I suspect). Since the thinning of the labels should happen with the same code, I suspect that there is when the error happened and it most likely propagated to the private/public data as well - hence why nobody sees clear improvements on the topological side.",
              "votes": 4
            }
          ]
        }
      ]
    },
    {
      "id": 3404127,
      "postDate": "2026-02-09T23:08:00.420Z",
      "content": "<p>Is it me or the scoring logic faulty also in other ways? I might be wrong but as far as I understand</p>\n<ul>\n<li>B1 = 0 on most (good) ground truth samples.</li>\n<li>F1 score, and also the code I saw here shared for metrics, gives a 0 score if the gt has b1 = 0 (usually) and the prediction has b1 &gt; 0</li>\n<li>This means for most examples predicting an example with a million holes gives the same score in terms of b1 as predicting an example with one hole). So basically it's an all or nothing metric.</li>\n</ul>\n<p>I might have got it wrong though.\nEdit: I guess I am wrong. I thought this is the score code: <a href=\"https://www.kaggle.com/code/jirkaborovec/replicate-lb-score-topology-aware-3d-surface-seg/\" target=\"_blank\">https://www.kaggle.com/code/jirkaborovec/replicate-lb-score-topology-aware-3d-surface-seg/</a></p>",
      "rawMarkdown": "Is it me or the scoring logic faulty also in other ways? I might be wrong but as far as I understand\n* B1 = 0 on most (good) ground truth samples.\n* F1 score, and also the code I saw here shared for metrics, gives a 0 score if the gt has b1 = 0 (usually) and the prediction has b1 > 0\n* This means for most examples predicting an example with a million holes gives the same score in terms of b1 as predicting an example with one hole). So basically it's an all or nothing metric.\n\nI might have got it wrong though.\nEdit: I guess I am wrong. I thought this is the score code: https://www.kaggle.com/code/jirkaborovec/replicate-lb-score-topology-aware-3d-surface-seg/"
    },
    {
      "id": 3400724,
      "postDate": "2026-02-02T06:15:14.257Z",
      "content": "<p>leaderboard rescoring?</p>",
      "rawMarkdown": "leaderboard rescoring?"
    },
    {
      "id": 3400592,
      "postDate": "2026-02-01T21:09:37.987Z",
      "content": "<p>Yeah, could be anwser on why my fighting with TopoScore was tottaly anti-productive.</p>",
      "rawMarkdown": "Yeah, could be anwser on why my fighting with TopoScore was tottaly anti-productive."
    },
    {
      "id": 3404171,
      "postDate": "2026-02-10T01:28:36.567Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 3399971,
      "postDate": "2026-01-31T16:16:03.973Z",
      "content": "<p>Yes, we noticed that too. How about treating true betti1 as 0? <a href=\"https://www.kaggle.com/seanjohnsonsp\" target=\"_blank\">@seanjohnsonsp</a> , <a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> , <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>",
      "rawMarkdown": "\nYes, we noticed that too. How about treating true betti1 as 0? @seanjohnsonsp , @giorgioangelotti , @sohier ",
      "isDeleted": true,
      "replies": [
        {
          "id": 3399990,
          "postDate": "2026-01-31T16:44:30.123Z",
          "content": "<p>The train labels can be fixed with binary close as shown in the <a href=\"https://www.kaggle.com/code/dankrstev/metric-failure-demo\" target=\"_blank\">Notebook</a> for your local validation. The question remains if these errors are present in the public/test set as well.</p>",
          "rawMarkdown": "The train labels can be fixed with binary close as shown in the [Notebook](https://www.kaggle.com/code/dankrstev/metric-failure-demo) for your local validation. The question remains if these errors are present in the public/test set as well.",
          "replies": [
            {
              "id": 3400084,
              "postDate": "2026-01-31T22:32:44.213Z",
              "content": "<p>I mean, the organizers don't have to change the test labels to fix this problem, they can just consider that the correct betti1 is 0 <a href=\"https://www.kaggle.com/seanjohnsonsp\" target=\"_blank\">@seanjohnsonsp</a> , <a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a></p>",
              "rawMarkdown": "\nI mean, the organizers don't have to change the test labels to fix this problem, they can just consider that the correct betti1 is 0 @seanjohnsonsp , @giorgioangelotti",
              "votes": 1,
              "isDeleted": true
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3402940,
      "author_name": "Giorgio Angelotti",
      "author_url": "",
      "post_date": "2026-02-07T08:18:51.457000",
      "content": "<p>You’re right: we’ve confirmed the dataset contains many microscopic holes/cavities.</p>\n<p>For full transparency: during our pre-release validation of the updated data (before Christmas), an automated check (which I ran) raised this as a potential issue. We discussed it internally, but evidence and opinions in my team were mixed and we ultimately concluded it was likely a false positive (i.e., a limitation/bug in the check rather than a real problem in the data). In hindsight, that conclusion was wrong. As competition host and project lead, I take responsibility for not insisting on an additional independent verification steps and stronger automated topology checks before release.</p>\n<p>We have since regenerated the affected test set and submitted an updated version to Kaggle two days ago. Because changing evaluation data late in a competition has fairness implications, it’s not automatic that it will be applied: Kaggle will decide whether/how/when to use it.</p>",
      "votes": 17,
      "replies": [
        {
          "id": 3402957,
          "author_name": "Satwik",
          "author_url": "",
          "post_date": "2026-02-07T08:52:10.297000",
          "content": "<p>Thank you for the response, and correcting the data. Is it possible for us to know, how much of the test set was affected? And is this just the public test set, or does it include private test set as well? And regarding the corrected data being used for the test set, can we get a confirmation if this data will be used or not (since you mentioned it is up to Kaggle staff whether to use it or not) ?</p>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 3402994,
          "author_name": "Tom",
          "author_url": "",
          "post_date": "2026-02-07T10:38:38.453000",
          "content": "<p><a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a>  Many thanks to the host team for putting so much effort into this competition. I understand that, in order to achieve the project’s goals, the organizers selected a metric that is quite challenging to use as a competition evaluation standard.\nFrom a personal perspective, I hope it would be possible to rescore the results on both the public and private datasets. Our team worked extremely hard to develop what we believe is a strong and elegant approach. We are truly eager for our work to contribute to the Herculaneum Papyri Project, helping extract ancient texts and uncover the knowledge hidden within these historical scrolls.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 3403171,
              "author_name": "",
              "author_url": "",
              "post_date": "2026-02-07T20:07:11.243000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3403175,
          "author_name": "dan4o",
          "author_url": "",
          "post_date": "2026-02-07T20:12:54.477000",
          "content": "<blockquote>\n  <p>Because changing evaluation data late in a competition has fairness implications</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a>  There is no need to change the data. There is a fix where the metric is slightly changed on your end, and the leaderboard is rescored. <strong>We set the g_k = 0, and the current score is calculated as if the current data had no holes</strong>. Then the results will be fair for everyone. The more holes your model predicts, the lower your score will be.</p>\n<pre><code>           g_k = 0  ### &lt;--- FIX\n            counts_by_dim[k] = (m_k, p_k, g_k)\n            denom = p_k + g_k\n            if denom &gt; 0:\n                if g_k != 0:\n                    topoF1_k = (2.0 * m_k) / float(denom)\n                else:\n                    topoF1_k = 0.5 / (float(denom) + 0.5) \n                topoF1_by_dim[k] = float(topoF1_k)\n                active_dims.append(k)  \n            else:\n                topoF1_by_dim[k] = float(\"nan\")\n</code></pre>\n<p>This satisfies the fairness aspect of rescoring, as <strong>no</strong> data will be changed. The dynamics will be exactly the same and the models will see exactly the same data. Each teams model will be penalized only for the amount of holes they predicted <em>on the same data as it is now</em>, where teams that predicted better topology will be rewarded. Otherwise the current leaderboard is not truthful.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 3403179,
              "author_name": "Marius Heuser",
              "author_url": "",
              "post_date": "2026-02-07T20:24:17.580000",
              "content": "<p>This is exactly the same as changing the evaluation data. It's not about changing the training data, just the evaluation data as far as I understand. So if you set the g_k=0 or fill all holes, it is pretty much the same. So I doubt that this will satisfy the fairness aspect more than changing the evaluation data. Since this has the same effect on the competitors, a change in LB score.</p>\n<p>But I'd be happy about a fix and extending the competition. Otherwise it is unlikely that the best model (for the task) will win / place high.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3403213,
              "author_name": "dan4o",
              "author_url": "",
              "post_date": "2026-02-07T22:44:02.783000",
              "content": "<blockquote>\n  <p>But I'd be happy about a fix and extending the competition. Otherwise it is unlikely that the best model (for the task) will win / place high.</p>\n</blockquote>\n<p>I totally agree. Its kinda strange if there is no action taken <em>even when we know</em> that there is an error in the evaluation of the solutions. The actual reading of the scrolls doesn't care about if we are fair or not fair, it only cares about having the best solution (topologically sound) so it can work downstream. </p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 3403303,
          "author_name": "GG Ayo (AyoGG)",
          "author_url": "",
          "post_date": "2026-02-08T07:19:21.973000",
          "content": "<p>Will old submissions be re-evaluated?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3403715,
          "author_name": "PRIYANSHU PRAJAPATI",
          "author_url": "",
          "post_date": "2026-02-09T05:33:41.393000",
          "content": "<p>That’s why my notebook was penalized it was being treated as a false positive.\nI did join very late, though, so let’s see how it goes.\nNow i dont have time to retrain the model again</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3403825,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2026-02-09T11:01:04.273000",
          "content": "<blockquote>\n  <p>Kaggle will decide whether/how/when to use it.</p>\n</blockquote>\n<p>You are the one to make that decision, you are the host. The rules are a contract between you and us, not between Kaggle and us.</p>\n<p>If there is a change in test data then the competition should be extended. People have spent weeks, months to improve their score using a given test harness. Changing that is totally unfair to them, and they should have the opportunity to adjust what they did. But extending the competition also has its drawbacks, and will probably be very badly received.</p>\n<p>TL;DR Changing the metric or test data in the last week should not happen.</p>\n<p>I also found another type of annotation issue, I'll share privately as I don't want to share new ideas publicly during the last week. No one should share new ideas during the last week actually.</p>",
          "votes": 10,
          "replies": [
            {
              "id": 3403847,
              "author_name": "dan4o",
              "author_url": "",
              "post_date": "2026-02-09T12:16:38.550000",
              "content": "<blockquote>\n  <p>If there is a change in test data then the competition should be extended. People have spent weeks, months to improve their score using a given test harness.</p>\n</blockquote>\n<p>I agree. </p>\n<p>My personal opinion is this. I understand that its late, but a fix is needed.</p>\n<p>The current leaderboard doesn't reflect the true goals of this competition. Because participants optimize their models specifically for the provided metric—which unfortunatelly is flawed—the best solutions for topology, Dice, and VOI won't actually rise to the top. </p>\n<p>If the error remains unfixed, those solutions that would've scored higher on what the host wanted, might not surface in the top prize slots and teams won't bother to share their methodology or code, even when in a parallel universe where the metric worked as intended they would've solved the problem better (which is the universe where we should be evaluating in, and where the problem of reading the papyrus actually exists). As is, the host will receive the best solutions that maximize the current score on the metric — whether that metric is flawed or not.</p>",
              "votes": 8,
              "replies": []
            },
            {
              "id": 3404151,
              "author_name": "Ángel Jacinto Sánchez Ruiz",
              "author_url": "",
              "post_date": "2026-02-09T23:57:46.337000",
              "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8722753%2Fc729c2f5ef77b01f0900085db719661c%2FSin%20ttulo.jpg?generation=1770681463824241&amp;alt=media\" alt=\"\"></p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3402376,
      "author_name": "María Cruz",
      "author_url": "",
      "post_date": "2026-02-05T22:26:17.200000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/dankrstev\" target=\"_blank\">@dankrstev</a> -- thank you for surfacing this. The host is looking into it and we will provide an update soon. Stay tuned!</p>",
      "votes": 7,
      "replies": [
        {
          "id": 3402384,
          "author_name": "Taha_Alshatiri",
          "author_url": "",
          "post_date": "2026-02-05T23:23:06.747000",
          "content": "<p>Its really late already, merging and entry ends tommorow</p>\n<p>And retraining might take couple of days for some of us, not sure if recalculating the metric again this time would really be a good idea</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3402402,
          "author_name": "huoxu",
          "author_url": "",
          "post_date": "2026-02-06T01:43:20.307000",
          "content": "<p>Can we reach a conclusion today? Should we wait until a conclusion is reached before submitting the results?</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 3402434,
          "author_name": "Wang Zhiyao (王致尧)",
          "author_url": "",
          "post_date": "2026-02-06T04:12:46.580000",
          "content": "<p>Will the competition’s end time being postponed? It’s too late—we’ve already wasted a lot of computing resources</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 3402717,
          "author_name": "J. Ferdous",
          "author_url": "",
          "post_date": "2026-02-06T15:56:11.437000",
          "content": "<p>Score can be re-scored. It doesn't matter if the competition will end soon. Its just the score and not the dataset. So,  it will not valid concern if anyone thinks score should not be rescored. </p>",
          "votes": -2,
          "replies": []
        }
      ]
    },
    {
      "id": 3404168,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2026-02-10T01:01:57.423000",
      "content": "<p>My first reaction was, “Why didn’t I catch this earlier, and what can I learn from it?”</p>\n<p>1) I realize I didn’t analyze the failure cases deeply enough. The code and concepts felt difficult, so I have skipped the details in the Betti matching part.</p>\n<p>2) As a result, some of my experiments were actually counterproductive. For example, when I replaced my prediction with the merged prediction and ground truth, the TopoScore became worse. I should have investigated this more carefully.</p>\n<p>As a machine learning practitioner, this is a very valuable lesson for me.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 3404182,
          "author_name": "dan4o",
          "author_url": "",
          "post_date": "2026-02-10T02:09:33.077000",
          "content": "<p>There are even more problems with the current metric. I am writing a fix for the host to the issue you raised <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/653482\" target=\"_blank\">here</a>. </p>\n<p>It is <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/672634#3404181\" target=\"_blank\">here</a>, please review it and look at it from another perspective - I hope I didnt miss some obvious scenario where the fix would fail.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3399896,
      "author_name": "Marius Heuser",
      "author_url": "",
      "post_date": "2026-01-31T13:38:58.437000",
      "content": "<p>Thank you for bringing this forward. I was also analyzing topo score since it didn't seem to add up. I have another weird situation, while both predictions are very similar (just tiny difference in post processing). I get for the same component a topo score of either 0.0 or 1.0. I suspected that there is a tiny voxel hole i just didn't see, but now I ran it with your notebook and got the following:</p>\n<p>The perfect one:\ntopo=TopoReport(toposcore=1.0, topoF1_by_dim={0: 1.0, 1: nan, 2: nan}, counts_by_dim={0: (1, 1, 1), 1: (0, 0, 0), 2: (0, 0, 0)}, dims=[0, 1, 2])</p>\n<p>The failure:\ntopo=TopoReport(toposcore=0.0, topoF1_by_dim={0: 0.0, 1: nan, 2: nan}, counts_by_dim={0: (0, 0, 1), 1: (0, 0, 0), 2: (0, 0, 0)}, dims=[0, 1, 2])</p>\n<p>How is it possible to get 0 on prediction in dim=0? \nMy prediction isn't empty, it is even very similar to the one which works.</p>\n<p>Here the visuals:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20325352%2F13929200bf07ccdd0e2fe6c2820cbdd9%2Ftopo.png?generation=1769866721033758&amp;alt=media\" alt=\"\"></p>\n<p>Here is the dataset:\n<a href=\"https://www.kaggle.com/datasets/mariusheuser/single-component-result/\" target=\"_blank\">https://www.kaggle.com/datasets/mariusheuser/single-component-result/</a></p>\n<p>The only way I see this happening is, if my prediction was invalid for some reason. But the dice score is valid (very high) for both predictions.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 3399961,
      "author_name": "Taha_Alshatiri",
      "author_url": "",
      "post_date": "2026-01-31T15:59:38.663000",
      "content": "<p><a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a>  is there  any update regarding this</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 3402597,
      "author_name": "Boredom",
      "author_url": "",
      "post_date": "2026-02-06T11:54:07.360000",
      "content": "<p>Haha, when I previously used the cases in deprecated_train_images as external test data, I couldn't understand why models with stronger mask connectivity (fewer holes) resulted in lower LB scores. Thanks for helping me realize this—let’s just let the model play a game of 'lucky hole-hitting' then~</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3404338,
      "author_name": "Ángel Jacinto Sánchez Ruiz",
      "author_url": "",
      "post_date": "2026-02-10T10:40:35.160000",
      "content": "<p>I will continue reading related topics about that. But if someone can confirm or fix my underastanding till now I will really appreciate it. </p>\n<ul>\n<li><p>Metric penalizes holes in masks, but training actual masks have already holes.</p></li>\n<li><p>Test has been fixed, so holes in masks have been removed, but training remains with them.</p></li>\n<li><p>We should fix holes in train masks by our side before properly train.</p></li>\n</ul>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3404221,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2026-02-10T04:16:16.563000",
      "content": "<p>you filled by public code given by the author:</p>\n<pre><code>def fill_small_holes_by_closing(mask: np.ndarray, iters: int = 1, conn: int = 3):\n    \"\"\"\n    Fast but can create bridges/handles. Use carefully.\n    conn: 1-&gt;6 neigh, 2-&gt;18, 3-&gt;26\n    \"\"\"\n    m = mask.astype(bool)\n    struct = ndi.generate_binary_structure(3, conn)\n    closed = ndi.binary_closing(m, structure=struct, iterations=iters)\n    return closed.astype(np.uint8)\n</code></pre>\n<p>only connectivity=26 works (6 and 16 don't)<br>\nbut voxels are added to other normal \"non-hole surface\" as well. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F38d7fe7f157b8f94dca09db302b61ad7%2FSelection_2430.png?generation=1770696868259779&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F30ee7e9ffcd9da77d72f4d8ce023f464%2FSelection_2432.png?generation=1770696890377629&amp;alt=media\" alt=\"\"></p>\n<pre><code>ID 1294570892 : 8 components\n\n#holes_after, _ = dim1_holes_self(gt_one)\n  1294570892 lb 1 fg 337197 has dim1 holes = 399\n\n#holes_after, _ = dim1_holes_self(filled)\n    changed voxels: 5147\n    after filling: 1294570892 lb 1 has dim1 holes = 0\n\n#r_cmp = do_one_lb(filled, gt_one)\n    compare filled vs gt_one:\n      lb=0.688566  surfDice=1.000000  VOI=0.967332  topo=0.000000\n      topo details: TopoReport(toposcore=0.0, topoF1_by_dim={0: nan, 1: 0.0, 2: nan}, counts_by_dim={0: (0, 0, 0), 1: (0, 0, 399), 2: (0, 0, 0)}, dims=[0, 1, 2])\n</code></pre>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3399988,
      "author_name": "Wayne_127",
      "author_url": "",
      "post_date": "2026-01-31T16:40:53.777000",
      "content": "<p>gt id 3742893488 visualization result from different angles :<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15801225%2Fa112911e8b0800beabcae7aa8700e90b%2Fnewplot.png?generation=1769877606266448&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15801225%2Ff942d8fbf133584856e7d1ff0771524a%2Fnewplot%20(1).png?generation=1769877611737440&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15801225%2F0d5d9d855577aac2af6fe4b292898b04%2Fnewplot%20(2).png?generation=1769877636618099&amp;alt=media\" alt=\"\"></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3404217,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2026-02-10T04:04:19.300000",
      "content": "<p>the holes are actually surface discretization error:</p>\n<p>ID 1294570892 : 8 components<br>\n  1294570892 lb 1 fg 337197 has dim1 holes = 399   </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F775c94b7156c68d215d732b9bc521f41%2FSelection_2431.png?generation=1770696170658795&amp;alt=media\" alt=\"\"></p>\n<p>some are \"visible\", but some are not (e.g. truth instance with only 1,2,4 10 holes)</p>\n<pre><code>def dim1_holes_self(mask_bin: np.ndarray):\n    \"\"\"\n    Compute dim1 count using the competition topo scorer, by scoring mask vs itself.\n    For self-vs-self, counts_by_dim[1] entries should be equal, so max() is safe.\n    \"\"\"\n    r = do_one_lb(mask_bin, mask_bin)\n    c = r[\"more_topo\"].counts_by_dim[1]   # tuple-like\n    return int(max(c)), r\n\n\n#---------------------------------------------------------------\n        gt = tifffile.imread(f\"{kaggle_dir}/train_labels/{id}.tif\") \n        fg = (gt == 1)\n\n        cc = cc3d.connected_components(fg, connectivity=26)\n        ncomp = int(cc.max())\n        print(f\"\\nID {id} : {ncomp} components\")\n\n        for lb in range(1, ncomp + 1):\n            gt_one = (cc == lb).astype(np.uint8)\n            fg_count = int(gt_one.sum())\n            if fg_count &lt; min_fg:\n                continue\n\n            holes_before, rep_self = dim1_holes_self(gt_one)\n            print(f\"  {id} lb {lb} fg {fg_count} has dim1 holes = {holes_before}\")\n</code></pre>",
      "votes": 2,
      "replies": [
        {
          "id": 3404327,
          "author_name": "cm391",
          "author_url": "",
          "post_date": "2026-02-10T10:15:27.553000",
          "content": "<p>did you take time to review the other problem found by <a href=\"https://www.kaggle.com/dankrstev\" target=\"_blank\">@dankrstev</a> <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/672447\" target=\"_blank\">discussion</a>?</p>\n<blockquote>\n  <p>And each voxel is connected to its neighbor by an edge/face so there should be 0 holes or cracks. We can see that the 4 voxels above are fully connected, and the vertex where they meet is causing the issue in the C++. We are physically incapable of predicting a better connection between the voxels.</p>\n</blockquote>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3399932,
      "author_name": "Starry",
      "author_url": "",
      "post_date": "2026-01-31T14:44:57.927000",
      "content": "<p>Indeed, this critical error had been <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/668313#3393314\" target=\"_blank\">discussed</a> before. However, it's disappointed that there are not any change yet (change the metric is also unpracticable as the deadline is coming). I think this issue will lead to the huge randomness and the huge shaking is coming in the way. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 3399934,
          "author_name": "dan4o",
          "author_url": "",
          "post_date": "2026-01-31T14:50:28.423000",
          "content": "<p>This has nothing to do with the ignore label discussed in the post you linked. The issue in that discussion is about ignore mask producing more connected components by trimming the prediction. </p>\n<p>This post is about the ground truth having <strong>wrong number of porous holes</strong> that bring the TopoScore down regardless of what you do or how you fix your holes. Please read the post and the notebook.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 3399935,
          "author_name": "Marius Heuser",
          "author_url": "",
          "post_date": "2026-01-31T14:52:10.647000",
          "content": "<p>That's a different issue. Here the GT is compared with the GT, so the ignore label shouldn't influence/change the prediction/result.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 3399958,
          "author_name": "Starry",
          "author_url": "",
          "post_date": "2026-01-31T15:54:14.537000",
          "content": "<p>Oh, I am sorry for misunderstanding. In fact, I also find a few cases contrain betti-1 error few days ago which should not be expected in training data labels. I run your code on 30 random data just now, even 20 of 30 contains unexpected betti-1 error. And 3 of 30 even contains more than 50 betti-1 error. The betti-1 score is meaningless if the test set is also like this.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7796170%2F9a76d98c51d323fa70fea8335769bcea%2Fa6a7fffe-a398-47aa-8f22-8c13f4b7473a.png?generation=1769874741171850&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": [
            {
              "id": 3399962,
              "author_name": "Marius Heuser",
              "author_url": "",
              "post_date": "2026-01-31T16:00:42.700000",
              "content": "<p>I'm not sure if I understood your post correctly, but the \"NaN\" values are supposed to be there. That simply means that betti-1 is ignored for the calculation, just like betti-2.\nThis is expected, since most GT labels have zero holes in them.\nThe 450, 413 are clearly errors and the others depend on the specific case.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3399969,
              "author_name": "Starry",
              "author_url": "",
              "post_date": "2026-01-31T16:14:00.323000",
              "content": "<p>I mean topoF1_by_dim in the figure should be {0:1, 1:nan, 2:nan} since the gt_label should not contain any betti-1 error. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3399974,
              "author_name": "Starry",
              "author_url": "",
              "post_date": "2026-01-31T16:20:00.747000",
              "content": "<p>And The 450 and 413 I think it is true label error. But as the figure show, there are other cases which betti-1 number are less than 10. I am not sure are these betti-1 number is caused by the label itself or the way of betti number calculation (by chunks). But anyway, it means at least this kind of calculation of the labels  would lead to the most gt_label contains unexpected betti-1 numbers which need to be clarifed by hosts. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3399976,
              "author_name": "dan4o",
              "author_url": "",
              "post_date": "2026-01-31T16:29:10.493000",
              "content": "<p>Yes that is correct. If there are 0 predicted, and 0 ground truth holes then betti-1 should be nan for that inactive dimension. However it is not nan because we are \"predicting\" a lot of holes when using the ground truth as our prediction. </p>\n<p>The algorithm finds the holes and it \"matches\" them correctly, thinking that we sucessfully found the correct holes in the ground truth. </p>\n<p>However when you use your own real model prediction, NONE of these holes will be correctly matched so you will end up with (0, YOUR_MODEL_HOLES, FAKE_GT_HOLES) where FAKE_GT_HOLES can be a large number. Even if your model predicts 0 holes, the corrupted ground truth holes will contribute in the denominator as seen below.</p>\n<p>The three numbers are (matched_k, predicted_k, ground_k).\nThis is the code that provides the counts_by_dim:</p>\n<pre><code>for k in dims_list:\n            m_k, p_k, g_k = cls._counts_for_dim(result, k)\n            counts_by_dim[k] = (m_k, p_k, g_k)\n            denom = p_k + g_k\n            if denom &gt; 0:\n                if g_k != 0:\n                    topoF1_k = (2.0 * m_k) / float(denom)\n                else:\n                    topoF1_k = 0.5 / (float(denom) + 0.5) # adding pseudocount to avoid 0\n                topoF1_by_dim[k] = float(topoF1_k)\n                active_dims.append(k)  \n            else:\n                topoF1_by_dim[k] = float(\"nan\")  # keep NaN for inactive\n</code></pre>\n<p>Most of the time during score calculation we are in this regime:</p>\n<pre><code>if g_k != 0:\n     topoF1_k = (2.0 * m_k) / float(denom)\n</code></pre>\n<p>which results in (2.0 * 0) / (your_model_holes + big_number_fake_holes) = 0.0….~</p>\n<p>Whereas if the sheets should be trully without holes we should be in this regime:</p>\n<pre><code>else:\n     topoF1_k = 0.5 / (float(denom) + 0.5) # adding pseudocount to avoid 0\n</code></pre>\n<p>where the more holes your model predicts the score decreases since its the only thing affecting the denominator denom = p_k + g_k(0)</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3399994,
              "author_name": "Starry",
              "author_url": "",
              "post_date": "2026-01-31T16:51:10.947000",
              "content": "<p>I suppose there shouldn't be any betti-1 number in gt_labels. If we treat the gt_label as the prediction the algorithms should also find 0 holes. And the topo score should be \"nan\" since the prediction and label are both 0. But as the figure show, the evaluation code find some holes and lead to the betti-1 number ≠ 0, which is also a  issue I think besides the true label error (hundreds of holes). </p>\n<p>Given such a gt_label of which the betti-1 number is not 0 (or at least the calculation of betti-1 number ≠0 using the evaluation code), suppose we have a perfect model which can prediction 0 holes of the whole volume. Even if we get such a good prediction, the topo-1 score is calculated by (0, 0, FAKE_GT_HOLES) or (0, FAKE_PREDICTION_HOLES, FAKE_GT_HOLES) which equal to 0. </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3399998,
              "author_name": "Starry",
              "author_url": "",
              "post_date": "2026-01-31T17:02:26.637000",
              "content": "<p>Here's a specific case:  [0:160, 160:320, 0:160] of 1075217434.tif. The betti-1 number are found in [30, 0, 5] and [100, 0, 4] using the betti_matching code. Seems that in the edge of the slice. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7796170%2Fed322316960b8602c3442b3c2e8855ac%2F4429c755-cbd7-498d-952f-e2a476d072c2.png?generation=1769878895880181&amp;alt=media\" alt=\"\"></p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3400001,
              "author_name": "dan4o",
              "author_url": "",
              "post_date": "2026-01-31T17:08:02.083000",
              "content": "<p>Yes, I also included in the description of the problem what kind of problem this is (porous holes), which can be fixed by BinaryClose as shown in the Notebook. We are waiting now on the Hosts to confirm if this problem is in the training labels alone (which I don't believe) or in the public/private tests sets as well (which I suspect). Since the thinning of the labels should happen with the same code, I suspect that there is when the error happened and it most likely propagated to the private/public data as well - hence why nobody sees clear improvements on the topological side.</p>",
              "votes": 4,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3404127,
      "author_name": "MlMl",
      "author_url": "",
      "post_date": "2026-02-09T23:08:00.420000",
      "content": "<p>Is it me or the scoring logic faulty also in other ways? I might be wrong but as far as I understand</p>\n<ul>\n<li>B1 = 0 on most (good) ground truth samples.</li>\n<li>F1 score, and also the code I saw here shared for metrics, gives a 0 score if the gt has b1 = 0 (usually) and the prediction has b1 &gt; 0</li>\n<li>This means for most examples predicting an example with a million holes gives the same score in terms of b1 as predicting an example with one hole). So basically it's an all or nothing metric.</li>\n</ul>\n<p>I might have got it wrong though.\nEdit: I guess I am wrong. I thought this is the score code: <a href=\"https://www.kaggle.com/code/jirkaborovec/replicate-lb-score-topology-aware-3d-surface-seg/\" target=\"_blank\">https://www.kaggle.com/code/jirkaborovec/replicate-lb-score-topology-aware-3d-surface-seg/</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3400724,
      "author_name": "Navneet",
      "author_url": "",
      "post_date": "2026-02-02T06:15:14.257000",
      "content": "<p>leaderboard rescoring?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3400592,
      "author_name": "The Loose Goose",
      "author_url": "",
      "post_date": "2026-02-01T21:09:37.987000",
      "content": "<p>Yeah, could be anwser on why my fighting with TopoScore was tottaly anti-productive.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3404171,
      "author_name": "",
      "author_url": "",
      "post_date": "2026-02-10T01:28:36.567000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3399971,
      "author_name": "",
      "author_url": "",
      "post_date": "2026-01-31T16:16:03.973000",
      "content": "<p>Yes, we noticed that too. How about treating true betti1 as 0? <a href=\"https://www.kaggle.com/seanjohnsonsp\" target=\"_blank\">@seanjohnsonsp</a> , <a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> , <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>",
      "votes": 0,
      "replies": [
        {
          "id": 3399990,
          "author_name": "dan4o",
          "author_url": "",
          "post_date": "2026-01-31T16:44:30.123000",
          "content": "<p>The train labels can be fixed with binary close as shown in the <a href=\"https://www.kaggle.com/code/dankrstev/metric-failure-demo\" target=\"_blank\">Notebook</a> for your local validation. The question remains if these errors are present in the public/test set as well.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3400084,
              "author_name": "",
              "author_url": "",
              "post_date": "2026-01-31T22:32:44.213000",
              "content": "<p>I mean, the organizers don't have to change the test labels to fix this problem, they can just consider that the correct betti1 is 0 <a href=\"https://www.kaggle.com/seanjohnsonsp\" target=\"_blank\">@seanjohnsonsp</a> , <a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a></p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3399828": "There is a **critical error** when calculating the TopoScore on the training labels provided. Some ground truth labels are porous and the topo score is dragged down dramatically by imaginary (porous) holes. In some instances in my local validation and also in the Kaggle metric the scoring function finds over 400! holes in the seemingly perfect ground truths. I tried using both _SPLITS= (1,1,1) and the default (2,2,2) suspecting there is some artifacts when splitting, however with both parameters the scoring function yields enormous number of holes.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2F6bc2f7406041151e965665114bf6dbec%2FScreenshot_4.jpg?generation=1769853751110753&alt=media)\n\nThis is a [notebook ](https://www.kaggle.com/code/dankrstev/metric-failure-demo)which demonstrates *that the current leaderboard is also potentially invalid* **IF** the test data resembles the training data and they underwent similar processing. Even if your model is predicting perfect sheets that have 0 holes in them the metric will severely punish this as it will compare your 0 holes with the double, or even triple digit holes in the labels and bring your score to 0. Hence the metric score doesn't care if your model predicts good topology and no holes. \n\nThere is a possibility that this is a behavior present **ONLY** in the training set and **NOT** in the test set, in which case only the local validation is invalid. However if they both underwent the same processing (thinning) I highly suspect that this also contaminated the test set. I would ask the hosts and kaggle staff ( @seanjohnsonsp, @giorgioangelotti, @sohier ) to confirm by running the simple test shown in the notebook and inform is the leaderboard is indeed calculated with non-porous labels or if this is the reason why so many competitors were demotivated not seeing any improvement on the leaderboard despite having visually superior predictions.\n\nIf the issue persists in the test set data as well, then the hosts should inform us and take the necessary actions (leaderboard rescoring, continue the competition for X more days/weeks to allow competitors to revisit approaches that were invalidated because they worked on faulty metric). Many competitors turned away from potentially good solutions from the fact that they received negative feedback on their topology algorithms and fixes despite having better visual and logical consistency. If the issue is indeed present in the test set, then half of the topology metric is invalid and the competition remains just dice score, which I think the hosts tried to avoid by incorporating the Topology Score in the first place!\n\n\n",
    "3402940": "You’re right: we’ve confirmed the dataset contains many microscopic holes/cavities.\n\nFor full transparency: during our pre-release validation of the updated data (before Christmas), an automated check (which I ran) raised this as a potential issue. We discussed it internally, but evidence and opinions in my team were mixed and we ultimately concluded it was likely a false positive (i.e., a limitation/bug in the check rather than a real problem in the data). In hindsight, that conclusion was wrong. As competition host and project lead, I take responsibility for not insisting on an additional independent verification steps and stronger automated topology checks before release.\n\nWe have since regenerated the affected test set and submitted an updated version to Kaggle two days ago. Because changing evaluation data late in a competition has fairness implications, it’s not automatic that it will be applied: Kaggle will decide whether/how/when to use it.",
    "3402376": "Hi @dankrstev -- thank you for surfacing this. The host is looking into it and we will provide an update soon. Stay tuned!",
    "3404168": "My first reaction was, “Why didn’t I catch this earlier, and what can I learn from it?”\n\n1) I realize I didn’t analyze the failure cases deeply enough. The code and concepts felt difficult, so I have skipped the details in the Betti matching part.\n\n2) As a result, some of my experiments were actually counterproductive. For example, when I replaced my prediction with the merged prediction and ground truth, the TopoScore became worse. I should have investigated this more carefully.\n\nAs a machine learning practitioner, this is a very valuable lesson for me.",
    "3399896": "Thank you for bringing this forward. I was also analyzing topo score since it didn't seem to add up. I have another weird situation, while both predictions are very similar (just tiny difference in post processing). I get for the same component a topo score of either 0.0 or 1.0. I suspected that there is a tiny voxel hole i just didn't see, but now I ran it with your notebook and got the following:\n\nThe perfect one:\ntopo=TopoReport(toposcore=1.0, topoF1_by_dim={0: 1.0, 1: nan, 2: nan}, counts_by_dim={0: (1, 1, 1), 1: (0, 0, 0), 2: (0, 0, 0)}, dims=[0, 1, 2])\n\nThe failure:\ntopo=TopoReport(toposcore=0.0, topoF1_by_dim={0: 0.0, 1: nan, 2: nan}, counts_by_dim={0: (0, 0, 1), 1: (0, 0, 0), 2: (0, 0, 0)}, dims=[0, 1, 2])\n\nHow is it possible to get 0 on prediction in dim=0? \nMy prediction isn't empty, it is even very similar to the one which works.\n\nHere the visuals:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20325352%2F13929200bf07ccdd0e2fe6c2820cbdd9%2Ftopo.png?generation=1769866721033758&alt=media)\n\nHere is the dataset:\n[https://www.kaggle.com/datasets/mariusheuser/single-component-result/](https://www.kaggle.com/datasets/mariusheuser/single-component-result/)\n\nThe only way I see this happening is, if my prediction was invalid for some reason. But the dice score is valid (very high) for both predictions.",
    "3399961": "@giorgioangelotti  is there  any update regarding this",
    "3402597": "Haha, when I previously used the cases in deprecated_train_images as external test data, I couldn't understand why models with stronger mask connectivity (fewer holes) resulted in lower LB scores. Thanks for helping me realize this—let’s just let the model play a game of 'lucky hole-hitting' then~",
    "3404338": "I will continue reading related topics about that. But if someone can confirm or fix my underastanding till now I will really appreciate it. \n\n* Metric penalizes holes in masks, but training actual masks have already holes.\n\n* Test has been fixed, so holes in masks have been removed, but training remains with them.\n\n* We should fix holes in train masks by our side before properly train.",
    "3404221": "you filled by public code given by the author:\n\n```\ndef fill_small_holes_by_closing(mask: np.ndarray, iters: int = 1, conn: int = 3):\n    \"\"\"\n    Fast but can create bridges/handles. Use carefully.\n    conn: 1->6 neigh, 2->18, 3->26\n    \"\"\"\n    m = mask.astype(bool)\n    struct = ndi.generate_binary_structure(3, conn)\n    closed = ndi.binary_closing(m, structure=struct, iterations=iters)\n    return closed.astype(np.uint8)\n\n```\n\nonly connectivity=26 works (6 and 16 don't)   \nbut voxels are added to other normal \"non-hole surface\" as well. \n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F38d7fe7f157b8f94dca09db302b61ad7%2FSelection_2430.png?generation=1770696868259779&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F30ee7e9ffcd9da77d72f4d8ce023f464%2FSelection_2432.png?generation=1770696890377629&alt=media)\n\n\n```\nID 1294570892 : 8 components\n\n#holes_after, _ = dim1_holes_self(gt_one)\n  1294570892 lb 1 fg 337197 has dim1 holes = 399\n\n#holes_after, _ = dim1_holes_self(filled)\n    changed voxels: 5147\n    after filling: 1294570892 lb 1 has dim1 holes = 0\n\n#r_cmp = do_one_lb(filled, gt_one)\n    compare filled vs gt_one:\n      lb=0.688566  surfDice=1.000000  VOI=0.967332  topo=0.000000\n      topo details: TopoReport(toposcore=0.0, topoF1_by_dim={0: nan, 1: 0.0, 2: nan}, counts_by_dim={0: (0, 0, 0), 1: (0, 0, 399), 2: (0, 0, 0)}, dims=[0, 1, 2])\n\n\n```",
    "3399988": "gt id 3742893488 visualization result from different angles :![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15801225%2Fa112911e8b0800beabcae7aa8700e90b%2Fnewplot.png?generation=1769877606266448&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15801225%2Ff942d8fbf133584856e7d1ff0771524a%2Fnewplot%20(1).png?generation=1769877611737440&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15801225%2F0d5d9d855577aac2af6fe4b292898b04%2Fnewplot%20(2).png?generation=1769877636618099&alt=media)",
    "3404217": "the holes are actually surface discretization error:\n\nID 1294570892 : 8 components   \n  1294570892 lb 1 fg 337197 has dim1 holes = 399   \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F775c94b7156c68d215d732b9bc521f41%2FSelection_2431.png?generation=1770696170658795&alt=media)\n\nsome are \"visible\", but some are not (e.g. truth instance with only 1,2,4 10 holes)\n\n\n```\n\ndef dim1_holes_self(mask_bin: np.ndarray):\n    \"\"\"\n    Compute dim1 count using the competition topo scorer, by scoring mask vs itself.\n    For self-vs-self, counts_by_dim[1] entries should be equal, so max() is safe.\n    \"\"\"\n    r = do_one_lb(mask_bin, mask_bin)\n    c = r[\"more_topo\"].counts_by_dim[1]   # tuple-like\n    return int(max(c)), r\n\n\n#---------------------------------------------------------------\n        gt = tifffile.imread(f\"{kaggle_dir}/train_labels/{id}.tif\") \n        fg = (gt == 1)\n\n        cc = cc3d.connected_components(fg, connectivity=26)\n        ncomp = int(cc.max())\n        print(f\"\\nID {id} : {ncomp} components\")\n\n        for lb in range(1, ncomp + 1):\n            gt_one = (cc == lb).astype(np.uint8)\n            fg_count = int(gt_one.sum())\n            if fg_count < min_fg:\n                continue\n\n            holes_before, rep_self = dim1_holes_self(gt_one)\n            print(f\"  {id} lb {lb} fg {fg_count} has dim1 holes = {holes_before}\")\n\n```",
    "3399932": "Indeed, this critical error had been [discussed](https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/668313#3393314) before. However, it's disappointed that there are not any change yet (change the metric is also unpracticable as the deadline is coming). I think this issue will lead to the huge randomness and the huge shaking is coming in the way. ",
    "3404127": "Is it me or the scoring logic faulty also in other ways? I might be wrong but as far as I understand\n* B1 = 0 on most (good) ground truth samples.\n* F1 score, and also the code I saw here shared for metrics, gives a 0 score if the gt has b1 = 0 (usually) and the prediction has b1 > 0\n* This means for most examples predicting an example with a million holes gives the same score in terms of b1 as predicting an example with one hole). So basically it's an all or nothing metric.\n\nI might have got it wrong though.\nEdit: I guess I am wrong. I thought this is the score code: https://www.kaggle.com/code/jirkaborovec/replicate-lb-score-topology-aware-3d-surface-seg/",
    "3400724": "leaderboard rescoring?",
    "3400592": "Yeah, could be anwser on why my fighting with TopoScore was tottaly anti-productive.",
    "3404171": "",
    "3399971": "\nYes, we noticed that too. How about treating true betti1 as 0? @seanjohnsonsp , @giorgioangelotti , @sohier "
  }
}