{
  "id": 668313,
  "title": "Ignore label is a trap",
  "url": "/competitions/vesuvius-challenge-surface-detection/discussion/668313",
  "author_name": "tingyi",
  "post_date": "2026-01-16T08:04:17.154000",
  "votes": 11,
  "comment_count": 21,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8788200%2F9fb19904fffd5fef6786098944dba2f7%2Fimage.webp?generation=1768550460239610&amp;alt=media\" alt=\")\"></p>\n<p>If you set the overlap between pred_label and ignore to background, and then calculate the 3D connected components, something interesting happens</p>",
  "messages": [
    {
      "id": 3392043,
      "postDate": "2026-01-16T08:04:17.153Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8788200%2F9fb19904fffd5fef6786098944dba2f7%2Fimage.webp?generation=1768550460239610&amp;alt=media\" alt=\")\"></p>\n<p>If you set the overlap between pred_label and ignore to background, and then calculate the 3D connected components, something interesting happens</p>",
      "rawMarkdown": "![)](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8788200%2F9fb19904fffd5fef6786098944dba2f7%2Fimage.webp?generation=1768550460239610&alt=media)\n\nIf you set the overlap between pred_label and ignore to background, and then calculate the 3D connected components, something interesting happens",
      "votes": 11
    },
    {
      "id": 3393107,
      "postDate": "2026-01-18T08:28:00.943Z",
      "content": "<p>I am actually very happy to see this post as this issue is something I observed myself on Friday when looking at our predictions on the cross-validation. This seems to be an underlying flaw with ignore label handling that has major effects on the metric computation!</p>\n<p>So I had a deeper look into the metrics today to understand what is happening and how the ignore label interacts with the prediction. The main issue is (I think) the voi metric which is computed on connected components. What is being done in the script is</p>\n<ol>\n<li>mask gt and pred with ignore label (set to 0)</li>\n<li>run cc on the rest</li>\n</ol>\n<p>This causes exactly the issue reported here. Previously connected components in the prediction are split, get assigned different ids and thus lower the score.</p>\n<p>H0 in topo score is probably getting messed up because of this as well.</p>\n<p>The recent change in labels amplifies this issue because the labels are now much narrower and the ignore label (at least on the train set) is much closer to the sheets. Given that the ignore label typically follows the nearest gt sheet with a fixed (small!) distance and that gt sheets like to float in the ether sometimes, it is easy to see that perfectly fine predictions will get killed as a result. </p>\n<p>So how could this be addressed?</p>\n<p>I think the eval script should act like this instead:</p>\n<ul>\n<li>core idea: do not apply ignore masking to the prediction anymore!</li>\n<li>compute CC in GT and PRED</li>\n<li>match GT and PRED instances (surface dice t&gt;=5?, hungarian matching?)</li>\n<li>for all unmatched pred instances, determine whether they predominantly lie in ignore mask (compute precision vs ignore label)<ul>\n<li>remove unmatched PRED sheets that are predominantly in ignore label</li></ul></li>\n<li>proceed as before.</li>\n</ul>\n<p>Or you know what would solve this issue even better? <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/660123\" target=\"_blank\">Representing sheets as instance maps instead of semantic segmentation, thus letting participants decide what their sheet instances are supposed to be</a>…</p>\n<p>Given how far along this challenge is I would be surprised if this was fixed. My proposal also introduces additional parameters (precision cutoff, matching strategy) that would need to be defined carefully. On a positive note, all participants are affected equally by this issue. </p>\n<p>Knowing where the ignore label in the test set is could be dangerous, that should remain a secret! Because we know that sheets are always annotated entirely so one could easily use this as additional signal for selecting/discarding sheet proposals. And don't forget that the goal of this challenge is to have an algorithm which can be applied to new unlabeled data (no ignore label available)…</p>",
      "rawMarkdown": "I am actually very happy to see this post as this issue is something I observed myself on Friday when looking at our predictions on the cross-validation. This seems to be an underlying flaw with ignore label handling that has major effects on the metric computation!\n\nSo I had a deeper look into the metrics today to understand what is happening and how the ignore label interacts with the prediction. The main issue is (I think) the voi metric which is computed on connected components. What is being done in the script is\n1. mask gt and pred with ignore label (set to 0)\n2. run cc on the rest\n\nThis causes exactly the issue reported here. Previously connected components in the prediction are split, get assigned different ids and thus lower the score.\n\nH0 in topo score is probably getting messed up because of this as well.\n\nThe recent change in labels amplifies this issue because the labels are now much narrower and the ignore label (at least on the train set) is much closer to the sheets. Given that the ignore label typically follows the nearest gt sheet with a fixed (small!) distance and that gt sheets like to float in the ether sometimes, it is easy to see that perfectly fine predictions will get killed as a result. \n\nSo how could this be addressed?\n\nI think the eval script should act like this instead:\n- core idea: do not apply ignore masking to the prediction anymore!\n- compute CC in GT and PRED\n- match GT and PRED instances (surface dice t>=5?, hungarian matching?)\n- for all unmatched pred instances, determine whether they predominantly lie in ignore mask (compute precision vs ignore label)\n  - remove unmatched PRED sheets that are predominantly in ignore label\n- proceed as before.\n\nOr you know what would solve this issue even better? [Representing sheets as instance maps instead of semantic segmentation, thus letting participants decide what their sheet instances are supposed to be](https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/660123)...\n\nGiven how far along this challenge is I would be surprised if this was fixed. My proposal also introduces additional parameters (precision cutoff, matching strategy) that would need to be defined carefully. On a positive note, all participants are affected equally by this issue. \n\nKnowing where the ignore label in the test set is could be dangerous, that should remain a secret! Because we know that sheets are always annotated entirely so one could easily use this as additional signal for selecting/discarding sheet proposals. And don't forget that the goal of this challenge is to have an algorithm which can be applied to new unlabeled data (no ignore label available)...\n",
      "votes": 8,
      "replies": [
        {
          "id": 3393132,
          "postDate": "2026-01-18T09:57:58.203Z",
          "content": "<blockquote>\n  <p>Knowing where the ignore label in the test set is could be dangerous, that should remain a secret! Because we know that sheets are always annotated entirely so one could easily use this as additional signal for selecting/discarding sheet proposals. \n  And don't forget that the goal of this challenge is to have an algorithm which can be applied to new unlabeled data (no ignore label available)…</p>\n</blockquote>\n<p>You're totally right, I didn't see it from the host's perspective.</p>\n<blockquote>\n  <p>On a positive note, all participants are affected equally by this issue. </p>\n</blockquote>\n<p>While this is true, I still feel like this issue could be \"unfair\". The way I see it: This problem could introduce a \"distribution shift\". Let's say the public LB has examples where this issue is not as common (less label==2 masks / examples with easier pathing) as for the hidden dataset or vise versa. Then there will be a shake up.</p>",
          "rawMarkdown": ">Knowing where the ignore label in the test set is could be dangerous, that should remain a secret! Because we know that sheets are always annotated entirely so one could easily use this as additional signal for selecting/discarding sheet proposals. \nAnd don't forget that the goal of this challenge is to have an algorithm which can be applied to new unlabeled data (no ignore label available)…\n\nYou're totally right, I didn't see it from the host's perspective.\n\n>On a positive note, all participants are affected equally by this issue. \n\nWhile this is true, I still feel like this issue could be \"unfair\". The way I see it: This problem could introduce a \"distribution shift\". Let's say the public LB has examples where this issue is not as common (less label==2 masks / examples with easier pathing) as for the hidden dataset or vise versa. Then there will be a shake up.\n",
          "votes": 2
        }
      ]
    },
    {
      "id": 3393314,
      "postDate": "2026-01-18T19:06:18.973Z",
      "content": "<p>I read all the comments posted so far. Unfortunately, we were aware of this issue from the beginning, and another participant ( <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> ) talked about this a month ago or so ( <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/653482\" target=\"_blank\">https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/653482</a> ).</p>\n<p>When we discussed this potential issue with the Kaggle support team even before launching the competition, the direction we received is that as long as an effect has an impact on any team equally, then the competition would be still fair, and I think this is the case.</p>\n<p>Nevertheless, on the old dataset we suggested just predicting thinner labels, but this would of course just mitigate and not fix the issue.</p>\n<p>Unfortunately, producing the labels for the missing sheets in the ignore region, and hence having a test set without the ignore region, is not practical. If we were able to produce these labels in a reasonable timeframe, we would have already had the model that we hope we can get from this challenge.</p>\n<p>One way to address the problem more systematically could be what <a href=\"https://www.kaggle.com/fabianisensee\" target=\"_blank\">@fabianisensee</a> suggests (modifying the metrics).</p>\n<p>This would add additional hyperparameters. As a host, I think these suggestions are smart and would guarantee the winning solution to be closer to what we would like to achieve from this challenge: a model that can predict separate sheets. However, there are some issues to consider:</p>\n<ol>\n<li>changing the metrics will likely cause a little shuffle in the leaderboard, with some teams that could be penalized by the choice of hyperparameters / thresholds, while others could benefit from them</li>\n<li>we are already very close to the limits of what can \"reasonably\" run on Kaggle as a metrics in terms of compute time and resources, since the Betti part of the loss is computational intensive. If you look at the script, you will notice that we had to split it in subchunks otherwise it wouldn't run. Adding more steps to the metrics computation could possibly cause some stability issues on Kaggle.</li>\n</ol>",
      "rawMarkdown": "I read all the comments posted so far. Unfortunately, we were aware of this issue from the beginning, and another participant ( @hengck23 ) talked about this a month ago or so ( https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/653482 ).\n\nWhen we discussed this potential issue with the Kaggle support team even before launching the competition, the direction we received is that as long as an effect has an impact on any team equally, then the competition would be still fair, and I think this is the case.\n\nNevertheless, on the old dataset we suggested just predicting thinner labels, but this would of course just mitigate and not fix the issue.\n\nUnfortunately, producing the labels for the missing sheets in the ignore region, and hence having a test set without the ignore region, is not practical. If we were able to produce these labels in a reasonable timeframe, we would have already had the model that we hope we can get from this challenge.\n\nOne way to address the problem more systematically could be what @fabianisensee suggests (modifying the metrics).\n\nThis would add additional hyperparameters. As a host, I think these suggestions are smart and would guarantee the winning solution to be closer to what we would like to achieve from this challenge: a model that can predict separate sheets. However, there are some issues to consider:\n\n1. changing the metrics will likely cause a little shuffle in the leaderboard, with some teams that could be penalized by the choice of hyperparameters / thresholds, while others could benefit from them\n2. we are already very close to the limits of what can \"reasonably\" run on Kaggle as a metrics in terms of compute time and resources, since the Betti part of the loss is computational intensive. If you look at the script, you will notice that we had to split it in subchunks otherwise it wouldn't run. Adding more steps to the metrics computation could possibly cause some stability issues on Kaggle.",
      "votes": 6,
      "replies": [
        {
          "id": 3393363,
          "postDate": "2026-01-18T21:33:45.113Z",
          "content": "<p>I think you significantly underestimate how much any change in metric is disrupting a competition. I dont really see the upside as I dont think core of current solutions would change, but I see a lot of risk and additional work for participants. </p>",
          "rawMarkdown": "I think you significantly underestimate how much any change in metric is disrupting a competition. I dont really see the upside as I dont think core of current solutions would change, but I see a lot of risk and additional work for participants. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 3393782,
      "postDate": "2026-01-19T18:55:09.417Z",
      "content": "<p>This topic is very interesting and a little frustrating as well.</p>\n<p>I want to chip in with my two cents, which I hope could be useful.</p>\n<p>First cent, can the leaderboard inference be done by feeding the image with ignored region erased to 0? In this way, a model is very less likely to predict a surface for the remaining fabric around the ignored region.</p>\n<p>Second, the host could spend time to curate the ignore mask for the leaderboard evaluation. It seems feasible within the timeline, I guess, since each ignore region is a huge chunk and easy to modify?</p>",
      "rawMarkdown": "This topic is very interesting and a little frustrating as well.\n\nI want to chip in with my two cents, which I hope could be useful.\n\nFirst cent, can the leaderboard inference be done by feeding the image with ignored region erased to 0? In this way, a model is very less likely to predict a surface for the remaining fabric around the ignored region.\n\nSecond, the host could spend time to curate the ignore mask for the leaderboard evaluation. It seems feasible within the timeline, I guess, since each ignore region is a huge chunk and easy to modify?\n",
      "votes": 1,
      "replies": [
        {
          "id": 3394253,
          "postDate": "2026-01-20T18:11:10.143Z",
          "content": "<p>The first is an interesting suggestion. We are inspecting the results of our baselines on this.\nUPDATE: public LB (without post processing) 0.478 . This seems to be detrimental to our baseline.</p>\n<p>The second is not feasible.</p>",
          "rawMarkdown": "The first is an interesting suggestion. We are inspecting the results of our baselines on this.\nUPDATE: public LB (without post processing) 0.478 . This seems to be detrimental to our baseline.\n\nThe second is not feasible.\n",
          "votes": 2,
          "replies": [
            {
              "id": 3394287,
              "postDate": "2026-01-20T19:33:26.753Z",
              "content": "<p>Yeah because this messes up the normalization during preprocessing (assuming the baseline is nnU-Net) and it will also mess up any batchnorm layers that were trained on non-zeroed out data. So it's not a good idea because it would require everyone to retrain their models.\nPlus, this challenge would yield models that expect large parts of the image to be blackened out which is not the case when you apply the best methods on unseen data post-challenge, so you'd likely be shooting yourself in the foot with that ;-)\nThe second suggestion is I think interesting. In many cases the ignore label is quite close to the last gt sheet even if sheets are far apart. It could indeed be moved in many cases to expose more background</p>",
              "rawMarkdown": "Yeah because this messes up the normalization during preprocessing (assuming the baseline is nnU-Net) and it will also mess up any batchnorm layers that were trained on non-zeroed out data. So it's not a good idea because it would require everyone to retrain their models.\nPlus, this challenge would yield models that expect large parts of the image to be blackened out which is not the case when you apply the best methods on unseen data post-challenge, so you'd likely be shooting yourself in the foot with that ;-)\nThe second suggestion is I think interesting. In many cases the ignore label is quite close to the last gt sheet even if sheets are far apart. It could indeed be moved in many cases to expose more background",
              "votes": 2
            },
            {
              "id": 3394402,
              "postDate": "2026-01-21T00:47:09.503Z",
              "content": "<blockquote>\n  <p>it would require everyone to retrain their models</p>\n</blockquote>\n<p>That's true that the first idea requires the model seeing erased samples during training.</p>",
              "rawMarkdown": ">it would require everyone to retrain their models\n\nThat's true that the first idea requires the model seeing erased samples during training."
            }
          ]
        }
      ]
    },
    {
      "id": 3392064,
      "postDate": "2026-01-16T08:52:21.460Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8788200%2F69697c1f66ed38c37be65d075fdbf384%2F.png?generation=1768553506758217&amp;alt=media\" alt=\"\"></p>\n<p>Observe this interesting phenomenon up close</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8788200%2F69697c1f66ed38c37be65d075fdbf384%2F.png?generation=1768553506758217&alt=media)\n\nObserve this interesting phenomenon up close",
      "votes": 1,
      "replies": [
        {
          "id": 3392119,
          "postDate": "2026-01-16T10:09:21.213Z",
          "content": "<p>Could you shed some more light on what is what</p>",
          "rawMarkdown": "Could you shed some more light on what is what",
          "replies": [
            {
              "id": 3392134,
              "postDate": "2026-01-16T10:32:32.623Z",
              "content": "<p>Oh, this is a little guy you know you can't avoid. These little guys are incredibly mischievous, and they only take a liking to the lucky ones.\"</p>",
              "rawMarkdown": "Oh, this is a little guy you know you can't avoid. These little guys are incredibly mischievous, and they only take a liking to the lucky ones.\""
            },
            {
              "id": 3392149,
              "postDate": "2026-01-16T11:01:18.710Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 3392045,
      "postDate": "2026-01-16T08:10:09.343Z",
      "content": "<p>1044587645\nLocal scores obtained using the Jirka code\n === 結果比較 ===</p>\n<p>--- Timing Breakdown ---\nPreprocess:   0.0859s\nSurface Dice: 10.4916s\nVOI Score:    1.0406s\nTopo Score:   1.7105s\n--- Particular Metrics ---\nSurfaceDice@2.0: 0.958988\nVOI_score:        0.559361\nTopoScore:        0.355224\n  Pred β: (13, 27, 0), GT β: (7, 0, 0)</p>\n<blockquote>\n  <blockquote>\n    <blockquote>\n      <p>SCORE:        0.637989</p>\n    </blockquote>\n  </blockquote>\n</blockquote>\n<p>原始: 13 塊</p>\n<p>--- Timing Breakdown ---\nPreprocess:   0.0873s\nSurface Dice: 9.5518s\nVOI Score:    1.0656s\nTopo Score:   1.7192s\n--- Particular Metrics ---\nSurfaceDice@2.0: 0.961063\nVOI_score:        0.559134\nTopoScore:        0.244982\n  Pred β: (22, 37, 0), GT β: (7, 0, 0)</p>\n<blockquote>\n  <blockquote>\n    <blockquote>\n      <p>SCORE:        0.605564</p>\n    </blockquote>\n  </blockquote>\n</blockquote>\n<p>處理後: 22 塊\nLabel: 7 塊</p>",
      "rawMarkdown": "1044587645\nLocal scores obtained using the Jirka code\n === 結果比較 ===\n\n--- Timing Breakdown ---\nPreprocess:   0.0859s\nSurface Dice: 10.4916s\nVOI Score:    1.0406s\nTopo Score:   1.7105s\n--- Particular Metrics ---\nSurfaceDice@2.0: 0.958988\nVOI_score:        0.559361\nTopoScore:        0.355224\n  Pred β: (13, 27, 0), GT β: (7, 0, 0)\n>>> SCORE:        0.637989\n\n原始: 13 塊\n\n--- Timing Breakdown ---\nPreprocess:   0.0873s\nSurface Dice: 9.5518s\nVOI Score:    1.0656s\nTopo Score:   1.7192s\n--- Particular Metrics ---\nSurfaceDice@2.0: 0.961063\nVOI_score:        0.559134\nTopoScore:        0.244982\n  Pred β: (22, 37, 0), GT β: (7, 0, 0)\n>>> SCORE:        0.605564\n\n處理後: 22 塊\nLabel: 7 塊",
      "votes": 1
    },
    {
      "id": 3392866,
      "postDate": "2026-01-17T17:06:05.160Z",
      "content": "<p>I think this is a valid concern and I would be curious to hear from the organizers how that is handled when evaluating the submissions. If everything in the ignore mask is removed, that may cause correctly predicted sheets to now be fragmented which will penalize the prediction very hard. This is particularly concerning with the new ground truth where the gap between sheets and ignore label is much narrower than before.</p>",
      "rawMarkdown": "I think this is a valid concern and I would be curious to hear from the organizers how that is handled when evaluating the submissions. If everything in the ignore mask is removed, that may cause correctly predicted sheets to now be fragmented which will penalize the prediction very hard. This is particularly concerning with the new ground truth where the gap between sheets and ignore label is much narrower than before.",
      "votes": 2
    },
    {
      "id": 3392842,
      "postDate": "2026-01-17T16:02:05.903Z",
      "content": "<p>I'm having the same problem. A prediction, which looks topologically promising, can still get a terrible betti-score.\nThe only way I see to prevent this is, if we had the label==2 masks available for the test set, which is probably too big of a change at this stage. And I'm not sure if that could be abused in any way, but otherwise it seems a bit random if you score high or low on betti:</p>\n<p>Prediction\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20325352%2Fd0a305ad44045745b2b3f7cc77763002%2Fprediction.png?generation=1768665691622237&amp;alt=media\" alt=\"\">\nPrediction with label==2 removed\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20325352%2Fa04aa550388fe9c98c0564a57ee42257%2Fprediction_masked.png?generation=1768665721932123&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I'm having the same problem. A prediction, which looks topologically promising, can still get a terrible betti-score.\nThe only way I see to prevent this is, if we had the label==2 masks available for the test set, which is probably too big of a change at this stage. And I'm not sure if that could be abused in any way, but otherwise it seems a bit random if you score high or low on betti:\n\nPrediction\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20325352%2Fd0a305ad44045745b2b3f7cc77763002%2Fprediction.png?generation=1768665691622237&alt=media)\nPrediction with label==2 removed\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20325352%2Fa04aa550388fe9c98c0564a57ee42257%2Fprediction_masked.png?generation=1768665721932123&alt=media)",
      "votes": 2
    },
    {
      "id": 3393057,
      "postDate": "2026-01-18T06:57:14.027Z",
      "content": "<p><a href=\"https://www.kaggle.com/fabianisensee\" target=\"_blank\">@fabianisensee</a> <a href=\"https://www.kaggle.com/mariusheuser\" target=\"_blank\">@mariusheuser</a> <a href=\"https://www.kaggle.com/shtljw\" target=\"_blank\">@shtljw</a> \nOh, my poor friend, let's ask the Competition Host to fix this ridiculous problem together.</p>",
      "rawMarkdown": "@fabianisensee @mariusheuser @shtljw \nOh, my poor friend, let's ask the Competition Host to fix this ridiculous problem together."
    },
    {
      "id": 3392146,
      "postDate": "2026-01-16T10:57:38.203Z",
      "content": "<p>How can you do this on the test set?</p>",
      "rawMarkdown": "How can you do this on the test set?",
      "replies": [
        {
          "id": 3392160,
          "postDate": "2026-01-16T11:19:15.150Z",
          "content": "<p>First, identify cases with extremely low Topo scores in local validation. Then, attempt to simulate the scoring by ignoring voxels in the Pred that overlap with the Ignore mask—our approach is to set these overlapping regions to background.</p>\n<p>Next, label the independent 3D connected components of the Pred using different colors. You will then notice these tiny 'isolated islands' truncated by the Ignore mask; it seems they just don't want to be 'friends' with me.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8788200%2F0b28ec69e33f5b460273a354cd5579da%2F2026-01-16%20190156.png?generation=1768561890399736&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8788200%2Fbd2ef6d9e12eb160ad8e22a7d33c092a%2Fimage.png?generation=1768562353348928&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "First, identify cases with extremely low Topo scores in local validation. Then, attempt to simulate the scoring by ignoring voxels in the Pred that overlap with the Ignore mask—our approach is to set these overlapping regions to background.\n\nNext, label the independent 3D connected components of the Pred using different colors. You will then notice these tiny 'isolated islands' truncated by the Ignore mask; it seems they just don't want to be 'friends' with me.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8788200%2F0b28ec69e33f5b460273a354cd5579da%2F2026-01-16%20190156.png?generation=1768561890399736&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8788200%2Fbd2ef6d9e12eb160ad8e22a7d33c092a%2Fimage.png?generation=1768562353348928&alt=media)",
          "replies": [
            {
              "id": 3392277,
              "postDate": "2026-01-16T15:43:49.803Z",
              "content": "<p>But we don't know whether the test socre is caculated using all label or not (using label+unlabel like training data). If it is calculated using the latter. The situation would be more complicated: there would be some cases that topo score would be very low since these isolated islands(truncted by unlabel mask) and it hard to resolve as we can't know the unlabel region.</p>",
              "rawMarkdown": "But we don't know whether the test socre is caculated using all label or not (using label+unlabel like training data). If it is calculated using the latter. The situation would be more complicated: there would be some cases that topo score would be very low since these isolated islands(truncted by unlabel mask) and it hard to resolve as we can't know the unlabel region.",
              "votes": 1
            }
          ]
        },
        {
          "id": 3393053,
          "postDate": "2026-01-18T06:49:56.203Z",
          "content": "<p><a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> I have a good idea: we could provide the ignore label regions along with the inference images. This would let participants decide for themselves how to handle them.</p>",
          "rawMarkdown": "@giorgioangelotti I have a good idea: we could provide the ignore label regions along with the inference images. This would let participants decide for themselves how to handle them."
        },
        {
          "id": 3393067,
          "postDate": "2026-01-18T07:24:14.100Z",
          "content": "<p>Due to the influence of the ignore label, our local evaluations frequently exhibit a scenario where Surface and VOI scores improve significantly while Topo scores decline sharply. We have determined that this situation is the cause, constituting an error in the scoring process.</p>",
          "rawMarkdown": "Due to the influence of the ignore label, our local evaluations frequently exhibit a scenario where Surface and VOI scores improve significantly while Topo scores decline sharply. We have determined that this situation is the cause, constituting an error in the scoring process."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3393107,
      "author_name": "FabianIsensee",
      "author_url": "",
      "post_date": "2026-01-18T08:28:00.943000",
      "content": "<p>I am actually very happy to see this post as this issue is something I observed myself on Friday when looking at our predictions on the cross-validation. This seems to be an underlying flaw with ignore label handling that has major effects on the metric computation!</p>\n<p>So I had a deeper look into the metrics today to understand what is happening and how the ignore label interacts with the prediction. The main issue is (I think) the voi metric which is computed on connected components. What is being done in the script is</p>\n<ol>\n<li>mask gt and pred with ignore label (set to 0)</li>\n<li>run cc on the rest</li>\n</ol>\n<p>This causes exactly the issue reported here. Previously connected components in the prediction are split, get assigned different ids and thus lower the score.</p>\n<p>H0 in topo score is probably getting messed up because of this as well.</p>\n<p>The recent change in labels amplifies this issue because the labels are now much narrower and the ignore label (at least on the train set) is much closer to the sheets. Given that the ignore label typically follows the nearest gt sheet with a fixed (small!) distance and that gt sheets like to float in the ether sometimes, it is easy to see that perfectly fine predictions will get killed as a result. </p>\n<p>So how could this be addressed?</p>\n<p>I think the eval script should act like this instead:</p>\n<ul>\n<li>core idea: do not apply ignore masking to the prediction anymore!</li>\n<li>compute CC in GT and PRED</li>\n<li>match GT and PRED instances (surface dice t&gt;=5?, hungarian matching?)</li>\n<li>for all unmatched pred instances, determine whether they predominantly lie in ignore mask (compute precision vs ignore label)<ul>\n<li>remove unmatched PRED sheets that are predominantly in ignore label</li></ul></li>\n<li>proceed as before.</li>\n</ul>\n<p>Or you know what would solve this issue even better? <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/660123\" target=\"_blank\">Representing sheets as instance maps instead of semantic segmentation, thus letting participants decide what their sheet instances are supposed to be</a>…</p>\n<p>Given how far along this challenge is I would be surprised if this was fixed. My proposal also introduces additional parameters (precision cutoff, matching strategy) that would need to be defined carefully. On a positive note, all participants are affected equally by this issue. </p>\n<p>Knowing where the ignore label in the test set is could be dangerous, that should remain a secret! Because we know that sheets are always annotated entirely so one could easily use this as additional signal for selecting/discarding sheet proposals. And don't forget that the goal of this challenge is to have an algorithm which can be applied to new unlabeled data (no ignore label available)…</p>",
      "votes": 8,
      "replies": [
        {
          "id": 3393132,
          "author_name": "Marius Heuser",
          "author_url": "",
          "post_date": "2026-01-18T09:57:58.203000",
          "content": "<blockquote>\n  <p>Knowing where the ignore label in the test set is could be dangerous, that should remain a secret! Because we know that sheets are always annotated entirely so one could easily use this as additional signal for selecting/discarding sheet proposals. \n  And don't forget that the goal of this challenge is to have an algorithm which can be applied to new unlabeled data (no ignore label available)…</p>\n</blockquote>\n<p>You're totally right, I didn't see it from the host's perspective.</p>\n<blockquote>\n  <p>On a positive note, all participants are affected equally by this issue. </p>\n</blockquote>\n<p>While this is true, I still feel like this issue could be \"unfair\". The way I see it: This problem could introduce a \"distribution shift\". Let's say the public LB has examples where this issue is not as common (less label==2 masks / examples with easier pathing) as for the hidden dataset or vise versa. Then there will be a shake up.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 3393314,
      "author_name": "Giorgio Angelotti",
      "author_url": "",
      "post_date": "2026-01-18T19:06:18.973000",
      "content": "<p>I read all the comments posted so far. Unfortunately, we were aware of this issue from the beginning, and another participant ( <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> ) talked about this a month ago or so ( <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/653482\" target=\"_blank\">https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/653482</a> ).</p>\n<p>When we discussed this potential issue with the Kaggle support team even before launching the competition, the direction we received is that as long as an effect has an impact on any team equally, then the competition would be still fair, and I think this is the case.</p>\n<p>Nevertheless, on the old dataset we suggested just predicting thinner labels, but this would of course just mitigate and not fix the issue.</p>\n<p>Unfortunately, producing the labels for the missing sheets in the ignore region, and hence having a test set without the ignore region, is not practical. If we were able to produce these labels in a reasonable timeframe, we would have already had the model that we hope we can get from this challenge.</p>\n<p>One way to address the problem more systematically could be what <a href=\"https://www.kaggle.com/fabianisensee\" target=\"_blank\">@fabianisensee</a> suggests (modifying the metrics).</p>\n<p>This would add additional hyperparameters. As a host, I think these suggestions are smart and would guarantee the winning solution to be closer to what we would like to achieve from this challenge: a model that can predict separate sheets. However, there are some issues to consider:</p>\n<ol>\n<li>changing the metrics will likely cause a little shuffle in the leaderboard, with some teams that could be penalized by the choice of hyperparameters / thresholds, while others could benefit from them</li>\n<li>we are already very close to the limits of what can \"reasonably\" run on Kaggle as a metrics in terms of compute time and resources, since the Betti part of the loss is computational intensive. If you look at the script, you will notice that we had to split it in subchunks otherwise it wouldn't run. Adding more steps to the metrics computation could possibly cause some stability issues on Kaggle.</li>\n</ol>",
      "votes": 6,
      "replies": [
        {
          "id": 3393363,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2026-01-18T21:33:45.113000",
          "content": "<p>I think you significantly underestimate how much any change in metric is disrupting a competition. I dont really see the upside as I dont think core of current solutions would change, but I see a lot of risk and additional work for participants. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 3393782,
      "author_name": "HOUJING HUANG",
      "author_url": "",
      "post_date": "2026-01-19T18:55:09.417000",
      "content": "<p>This topic is very interesting and a little frustrating as well.</p>\n<p>I want to chip in with my two cents, which I hope could be useful.</p>\n<p>First cent, can the leaderboard inference be done by feeding the image with ignored region erased to 0? In this way, a model is very less likely to predict a surface for the remaining fabric around the ignored region.</p>\n<p>Second, the host could spend time to curate the ignore mask for the leaderboard evaluation. It seems feasible within the timeline, I guess, since each ignore region is a huge chunk and easy to modify?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3394253,
          "author_name": "Giorgio Angelotti",
          "author_url": "",
          "post_date": "2026-01-20T18:11:10.143000",
          "content": "<p>The first is an interesting suggestion. We are inspecting the results of our baselines on this.\nUPDATE: public LB (without post processing) 0.478 . This seems to be detrimental to our baseline.</p>\n<p>The second is not feasible.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3394287,
              "author_name": "FabianIsensee",
              "author_url": "",
              "post_date": "2026-01-20T19:33:26.753000",
              "content": "<p>Yeah because this messes up the normalization during preprocessing (assuming the baseline is nnU-Net) and it will also mess up any batchnorm layers that were trained on non-zeroed out data. So it's not a good idea because it would require everyone to retrain their models.\nPlus, this challenge would yield models that expect large parts of the image to be blackened out which is not the case when you apply the best methods on unseen data post-challenge, so you'd likely be shooting yourself in the foot with that ;-)\nThe second suggestion is I think interesting. In many cases the ignore label is quite close to the last gt sheet even if sheets are far apart. It could indeed be moved in many cases to expose more background</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3394402,
              "author_name": "HOUJING HUANG",
              "author_url": "",
              "post_date": "2026-01-21T00:47:09.503000",
              "content": "<blockquote>\n  <p>it would require everyone to retrain their models</p>\n</blockquote>\n<p>That's true that the first idea requires the model seeing erased samples during training.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3392064,
      "author_name": "tingyi",
      "author_url": "",
      "post_date": "2026-01-16T08:52:21.460000",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8788200%2F69697c1f66ed38c37be65d075fdbf384%2F.png?generation=1768553506758217&amp;alt=media\" alt=\"\"></p>\n<p>Observe this interesting phenomenon up close</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3392119,
          "author_name": "ArjunB",
          "author_url": "",
          "post_date": "2026-01-16T10:09:21.213000",
          "content": "<p>Could you shed some more light on what is what</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3392134,
              "author_name": "tingyi",
              "author_url": "",
              "post_date": "2026-01-16T10:32:32.623000",
              "content": "<p>Oh, this is a little guy you know you can't avoid. These little guys are incredibly mischievous, and they only take a liking to the lucky ones.\"</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3392149,
              "author_name": "",
              "author_url": "",
              "post_date": "2026-01-16T11:01:18.710000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3392045,
      "author_name": "GG Ayo (AyoGG)",
      "author_url": "",
      "post_date": "2026-01-16T08:10:09.343000",
      "content": "<p>1044587645\nLocal scores obtained using the Jirka code\n === 結果比較 ===</p>\n<p>--- Timing Breakdown ---\nPreprocess:   0.0859s\nSurface Dice: 10.4916s\nVOI Score:    1.0406s\nTopo Score:   1.7105s\n--- Particular Metrics ---\nSurfaceDice@2.0: 0.958988\nVOI_score:        0.559361\nTopoScore:        0.355224\n  Pred β: (13, 27, 0), GT β: (7, 0, 0)</p>\n<blockquote>\n  <blockquote>\n    <blockquote>\n      <p>SCORE:        0.637989</p>\n    </blockquote>\n  </blockquote>\n</blockquote>\n<p>原始: 13 塊</p>\n<p>--- Timing Breakdown ---\nPreprocess:   0.0873s\nSurface Dice: 9.5518s\nVOI Score:    1.0656s\nTopo Score:   1.7192s\n--- Particular Metrics ---\nSurfaceDice@2.0: 0.961063\nVOI_score:        0.559134\nTopoScore:        0.244982\n  Pred β: (22, 37, 0), GT β: (7, 0, 0)</p>\n<blockquote>\n  <blockquote>\n    <blockquote>\n      <p>SCORE:        0.605564</p>\n    </blockquote>\n  </blockquote>\n</blockquote>\n<p>處理後: 22 塊\nLabel: 7 塊</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3392866,
      "author_name": "FabianIsensee",
      "author_url": "",
      "post_date": "2026-01-17T17:06:05.160000",
      "content": "<p>I think this is a valid concern and I would be curious to hear from the organizers how that is handled when evaluating the submissions. If everything in the ignore mask is removed, that may cause correctly predicted sheets to now be fragmented which will penalize the prediction very hard. This is particularly concerning with the new ground truth where the gap between sheets and ignore label is much narrower than before.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3392842,
      "author_name": "Marius Heuser",
      "author_url": "",
      "post_date": "2026-01-17T16:02:05.903000",
      "content": "<p>I'm having the same problem. A prediction, which looks topologically promising, can still get a terrible betti-score.\nThe only way I see to prevent this is, if we had the label==2 masks available for the test set, which is probably too big of a change at this stage. And I'm not sure if that could be abused in any way, but otherwise it seems a bit random if you score high or low on betti:</p>\n<p>Prediction\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20325352%2Fd0a305ad44045745b2b3f7cc77763002%2Fprediction.png?generation=1768665691622237&amp;alt=media\" alt=\"\">\nPrediction with label==2 removed\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20325352%2Fa04aa550388fe9c98c0564a57ee42257%2Fprediction_masked.png?generation=1768665721932123&amp;alt=media\" alt=\"\"></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3393057,
      "author_name": "tingyi",
      "author_url": "",
      "post_date": "2026-01-18T06:57:14.027000",
      "content": "<p><a href=\"https://www.kaggle.com/fabianisensee\" target=\"_blank\">@fabianisensee</a> <a href=\"https://www.kaggle.com/mariusheuser\" target=\"_blank\">@mariusheuser</a> <a href=\"https://www.kaggle.com/shtljw\" target=\"_blank\">@shtljw</a> \nOh, my poor friend, let's ask the Competition Host to fix this ridiculous problem together.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3392146,
      "author_name": "Giorgio Angelotti",
      "author_url": "",
      "post_date": "2026-01-16T10:57:38.203000",
      "content": "<p>How can you do this on the test set?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3392160,
          "author_name": "tingyi",
          "author_url": "",
          "post_date": "2026-01-16T11:19:15.150000",
          "content": "<p>First, identify cases with extremely low Topo scores in local validation. Then, attempt to simulate the scoring by ignoring voxels in the Pred that overlap with the Ignore mask—our approach is to set these overlapping regions to background.</p>\n<p>Next, label the independent 3D connected components of the Pred using different colors. You will then notice these tiny 'isolated islands' truncated by the Ignore mask; it seems they just don't want to be 'friends' with me.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8788200%2F0b28ec69e33f5b460273a354cd5579da%2F2026-01-16%20190156.png?generation=1768561890399736&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8788200%2Fbd2ef6d9e12eb160ad8e22a7d33c092a%2Fimage.png?generation=1768562353348928&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": [
            {
              "id": 3392277,
              "author_name": "Starry",
              "author_url": "",
              "post_date": "2026-01-16T15:43:49.803000",
              "content": "<p>But we don't know whether the test socre is caculated using all label or not (using label+unlabel like training data). If it is calculated using the latter. The situation would be more complicated: there would be some cases that topo score would be very low since these isolated islands(truncted by unlabel mask) and it hard to resolve as we can't know the unlabel region.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 3393053,
          "author_name": "tingyi",
          "author_url": "",
          "post_date": "2026-01-18T06:49:56.203000",
          "content": "<p><a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> I have a good idea: we could provide the ignore label regions along with the inference images. This would let participants decide for themselves how to handle them.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3393067,
          "author_name": "GG Ayo (AyoGG)",
          "author_url": "",
          "post_date": "2026-01-18T07:24:14.100000",
          "content": "<p>Due to the influence of the ignore label, our local evaluations frequently exhibit a scenario where Surface and VOI scores improve significantly while Topo scores decline sharply. We have determined that this situation is the cause, constituting an error in the scoring process.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3392043": "![)](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8788200%2F9fb19904fffd5fef6786098944dba2f7%2Fimage.webp?generation=1768550460239610&alt=media)\n\nIf you set the overlap between pred_label and ignore to background, and then calculate the 3D connected components, something interesting happens",
    "3393107": "I am actually very happy to see this post as this issue is something I observed myself on Friday when looking at our predictions on the cross-validation. This seems to be an underlying flaw with ignore label handling that has major effects on the metric computation!\n\nSo I had a deeper look into the metrics today to understand what is happening and how the ignore label interacts with the prediction. The main issue is (I think) the voi metric which is computed on connected components. What is being done in the script is\n1. mask gt and pred with ignore label (set to 0)\n2. run cc on the rest\n\nThis causes exactly the issue reported here. Previously connected components in the prediction are split, get assigned different ids and thus lower the score.\n\nH0 in topo score is probably getting messed up because of this as well.\n\nThe recent change in labels amplifies this issue because the labels are now much narrower and the ignore label (at least on the train set) is much closer to the sheets. Given that the ignore label typically follows the nearest gt sheet with a fixed (small!) distance and that gt sheets like to float in the ether sometimes, it is easy to see that perfectly fine predictions will get killed as a result. \n\nSo how could this be addressed?\n\nI think the eval script should act like this instead:\n- core idea: do not apply ignore masking to the prediction anymore!\n- compute CC in GT and PRED\n- match GT and PRED instances (surface dice t>=5?, hungarian matching?)\n- for all unmatched pred instances, determine whether they predominantly lie in ignore mask (compute precision vs ignore label)\n  - remove unmatched PRED sheets that are predominantly in ignore label\n- proceed as before.\n\nOr you know what would solve this issue even better? [Representing sheets as instance maps instead of semantic segmentation, thus letting participants decide what their sheet instances are supposed to be](https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/660123)...\n\nGiven how far along this challenge is I would be surprised if this was fixed. My proposal also introduces additional parameters (precision cutoff, matching strategy) that would need to be defined carefully. On a positive note, all participants are affected equally by this issue. \n\nKnowing where the ignore label in the test set is could be dangerous, that should remain a secret! Because we know that sheets are always annotated entirely so one could easily use this as additional signal for selecting/discarding sheet proposals. And don't forget that the goal of this challenge is to have an algorithm which can be applied to new unlabeled data (no ignore label available)...\n",
    "3393314": "I read all the comments posted so far. Unfortunately, we were aware of this issue from the beginning, and another participant ( @hengck23 ) talked about this a month ago or so ( https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/653482 ).\n\nWhen we discussed this potential issue with the Kaggle support team even before launching the competition, the direction we received is that as long as an effect has an impact on any team equally, then the competition would be still fair, and I think this is the case.\n\nNevertheless, on the old dataset we suggested just predicting thinner labels, but this would of course just mitigate and not fix the issue.\n\nUnfortunately, producing the labels for the missing sheets in the ignore region, and hence having a test set without the ignore region, is not practical. If we were able to produce these labels in a reasonable timeframe, we would have already had the model that we hope we can get from this challenge.\n\nOne way to address the problem more systematically could be what @fabianisensee suggests (modifying the metrics).\n\nThis would add additional hyperparameters. As a host, I think these suggestions are smart and would guarantee the winning solution to be closer to what we would like to achieve from this challenge: a model that can predict separate sheets. However, there are some issues to consider:\n\n1. changing the metrics will likely cause a little shuffle in the leaderboard, with some teams that could be penalized by the choice of hyperparameters / thresholds, while others could benefit from them\n2. we are already very close to the limits of what can \"reasonably\" run on Kaggle as a metrics in terms of compute time and resources, since the Betti part of the loss is computational intensive. If you look at the script, you will notice that we had to split it in subchunks otherwise it wouldn't run. Adding more steps to the metrics computation could possibly cause some stability issues on Kaggle.",
    "3393782": "This topic is very interesting and a little frustrating as well.\n\nI want to chip in with my two cents, which I hope could be useful.\n\nFirst cent, can the leaderboard inference be done by feeding the image with ignored region erased to 0? In this way, a model is very less likely to predict a surface for the remaining fabric around the ignored region.\n\nSecond, the host could spend time to curate the ignore mask for the leaderboard evaluation. It seems feasible within the timeline, I guess, since each ignore region is a huge chunk and easy to modify?\n",
    "3392064": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8788200%2F69697c1f66ed38c37be65d075fdbf384%2F.png?generation=1768553506758217&alt=media)\n\nObserve this interesting phenomenon up close",
    "3392045": "1044587645\nLocal scores obtained using the Jirka code\n === 結果比較 ===\n\n--- Timing Breakdown ---\nPreprocess:   0.0859s\nSurface Dice: 10.4916s\nVOI Score:    1.0406s\nTopo Score:   1.7105s\n--- Particular Metrics ---\nSurfaceDice@2.0: 0.958988\nVOI_score:        0.559361\nTopoScore:        0.355224\n  Pred β: (13, 27, 0), GT β: (7, 0, 0)\n>>> SCORE:        0.637989\n\n原始: 13 塊\n\n--- Timing Breakdown ---\nPreprocess:   0.0873s\nSurface Dice: 9.5518s\nVOI Score:    1.0656s\nTopo Score:   1.7192s\n--- Particular Metrics ---\nSurfaceDice@2.0: 0.961063\nVOI_score:        0.559134\nTopoScore:        0.244982\n  Pred β: (22, 37, 0), GT β: (7, 0, 0)\n>>> SCORE:        0.605564\n\n處理後: 22 塊\nLabel: 7 塊",
    "3392866": "I think this is a valid concern and I would be curious to hear from the organizers how that is handled when evaluating the submissions. If everything in the ignore mask is removed, that may cause correctly predicted sheets to now be fragmented which will penalize the prediction very hard. This is particularly concerning with the new ground truth where the gap between sheets and ignore label is much narrower than before.",
    "3392842": "I'm having the same problem. A prediction, which looks topologically promising, can still get a terrible betti-score.\nThe only way I see to prevent this is, if we had the label==2 masks available for the test set, which is probably too big of a change at this stage. And I'm not sure if that could be abused in any way, but otherwise it seems a bit random if you score high or low on betti:\n\nPrediction\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20325352%2Fd0a305ad44045745b2b3f7cc77763002%2Fprediction.png?generation=1768665691622237&alt=media)\nPrediction with label==2 removed\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20325352%2Fa04aa550388fe9c98c0564a57ee42257%2Fprediction_masked.png?generation=1768665721932123&alt=media)",
    "3393057": "@fabianisensee @mariusheuser @shtljw \nOh, my poor friend, let's ask the Competition Host to fix this ridiculous problem together.",
    "3392146": "How can you do this on the test set?"
  }
}