{
  "id": 678904,
  "title": "Is the hidden dataset different from the training data?",
  "url": "/competitions/vesuvius-challenge-surface-detection/discussion/678904",
  "author_name": "lllleeeo",
  "post_date": "2026-02-25T17:45:31.410000",
  "votes": 9,
  "comment_count": 52,
  "views": 0,
  "content": "<p>I recently developed a new model that achieves better CV locally with fewer holes and adhesions. However, after submission, its score fell significantly below other models with poorer CV. After each submission, I examine the model's prediction for a provided test sample. While a single sample lacks representativeness, it still offers some insight into the model's predictive capability. The output from this test sample also reveals that the model possesses a superior topological structure. Yet, its score is 0.557, even lower than the best public notebook. Below are the predictions from this model and the public notebook.</p>\n<p>Public Notebook LB 0.559\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19164179%2F0bfab6d3d687dfcdea127b1eefcd1a42%2FScreenShot_2026-02-25_171949_505.png?generation=1772040934339296&amp;alt=media\" alt=\"\"></p>\n<p>My Model LB 0.557\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19164179%2Fd8368e576c918cfcdcefc48befcdd558%2FScreenShot_2026-02-25_173612_063.png?generation=1772041002339142&amp;alt=media\" alt=\"\"></p>\n<p>Achieving a better topological structure has always been my goal in model optimization. Historically, improved topology consistently yielded higher scores. Yet this time, the model trained under the new framework with superior topology resulted in a dramatic score drop. Frankly, this has left me somewhat perplexed. Perhaps the performance on the training data doesn't fully align with private data. Could optimizing for a better DIC be the true objective?</p>\n<p>Has anyone encountered a situation where the topology improved but the LB worsened?</p>",
  "messages": [
    {
      "id": 3413997,
      "postDate": "2026-02-25T17:45:31.410Z",
      "content": "<p>I recently developed a new model that achieves better CV locally with fewer holes and adhesions. However, after submission, its score fell significantly below other models with poorer CV. After each submission, I examine the model's prediction for a provided test sample. While a single sample lacks representativeness, it still offers some insight into the model's predictive capability. The output from this test sample also reveals that the model possesses a superior topological structure. Yet, its score is 0.557, even lower than the best public notebook. Below are the predictions from this model and the public notebook.</p>\n<p>Public Notebook LB 0.559\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19164179%2F0bfab6d3d687dfcdea127b1eefcd1a42%2FScreenShot_2026-02-25_171949_505.png?generation=1772040934339296&amp;alt=media\" alt=\"\"></p>\n<p>My Model LB 0.557\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19164179%2Fd8368e576c918cfcdcefc48befcdd558%2FScreenShot_2026-02-25_173612_063.png?generation=1772041002339142&amp;alt=media\" alt=\"\"></p>\n<p>Achieving a better topological structure has always been my goal in model optimization. Historically, improved topology consistently yielded higher scores. Yet this time, the model trained under the new framework with superior topology resulted in a dramatic score drop. Frankly, this has left me somewhat perplexed. Perhaps the performance on the training data doesn't fully align with private data. Could optimizing for a better DIC be the true objective?</p>\n<p>Has anyone encountered a situation where the topology improved but the LB worsened?</p>",
      "rawMarkdown": "I recently developed a new model that achieves better CV locally with fewer holes and adhesions. However, after submission, its score fell significantly below other models with poorer CV. After each submission, I examine the model's prediction for a provided test sample. While a single sample lacks representativeness, it still offers some insight into the model's predictive capability. The output from this test sample also reveals that the model possesses a superior topological structure. Yet, its score is 0.557, even lower than the best public notebook. Below are the predictions from this model and the public notebook.\n\nPublic Notebook LB 0.559\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19164179%2F0bfab6d3d687dfcdea127b1eefcd1a42%2FScreenShot_2026-02-25_171949_505.png?generation=1772040934339296&alt=media)\n\nMy Model LB 0.557\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19164179%2Fd8368e576c918cfcdcefc48befcdd558%2FScreenShot_2026-02-25_173612_063.png?generation=1772041002339142&alt=media)\n\nAchieving a better topological structure has always been my goal in model optimization. Historically, improved topology consistently yielded higher scores. Yet this time, the model trained under the new framework with superior topology resulted in a dramatic score drop. Frankly, this has left me somewhat perplexed. Perhaps the performance on the training data doesn't fully align with private data. Could optimizing for a better DIC be the true objective?\n\nHas anyone encountered a situation where the topology improved but the LB worsened?",
      "votes": 9
    },
    {
      "id": 3414592,
      "postDate": "2026-02-27T08:17:32.270Z",
      "content": "<p>Why have a few assassins suddenly appeared these last few days?</p>",
      "rawMarkdown": "Why have a few assassins suddenly appeared these last few days?",
      "votes": 6,
      "replies": [
        {
          "id": 3414689,
          "postDate": "2026-02-27T12:12:17.780Z",
          "content": "<p>I also noticed this! I wish I could have gone on holidays and then reach the top of the leaderboard a day before the competition ends… I'm sure the Kaggle team will check that everything is \"fair\".</p>",
          "rawMarkdown": "I also noticed this! I wish I could have gone on holidays and then reach the top of the leaderboard a day before the competition ends... I'm sure the Kaggle team will check that everything is \"fair\".",
          "votes": 3
        },
        {
          "id": 3414728,
          "postDate": "2026-02-27T14:40:39.127Z",
          "content": "<p>It seems weird. Because the public notebooks LB 0.566 and 0.558 with the private notebooks, still so many teams can get great scores with less than 10 submissions </p>",
          "rawMarkdown": "It seems weird. Because the public notebooks LB 0.566 and 0.558 with the private notebooks, still so many teams can get great scores with less than 10 submissions ",
          "votes": 2,
          "replies": [
            {
              "id": 3414748,
              "postDate": "2026-02-27T15:26:19.300Z",
              "content": "<p>Can they with the .566? The models are hidden behind unshared datasets, although reproduction might be straightforward.</p>",
              "rawMarkdown": "Can they with the .566? The models are hidden behind unshared datasets, although reproduction might be straightforward."
            },
            {
              "id": 3414754,
              "postDate": "2026-02-27T15:34:11.843Z",
              "content": "<p>This is also what puzzles me, because other competitions are copied quickly because the resources are publicly available.</p>",
              "rawMarkdown": "This is also what puzzles me, because other competitions are copied quickly because the resources are publicly available."
            },
            {
              "id": 3414758,
              "postDate": "2026-02-27T15:39:23.133Z",
              "content": "<p>I assume this is a clever technique with the auto +1 on editing your copy to gain status</p>",
              "rawMarkdown": "I assume this is a clever technique with the auto +1 on editing your copy to gain status",
              "votes": 1
            },
            {
              "id": 3414764,
              "postDate": "2026-02-27T15:51:44.233Z",
              "content": "<p>I previously confirmed that teams that suddenly appeared with a 0.559 score (with fewer than 15 submissions) had upvotes on the LB 0.559 public notebook. This raises suspicions that some private transaction notebooks were involved.</p>",
              "rawMarkdown": "I previously confirmed that teams that suddenly appeared with a 0.559 score (with fewer than 15 submissions) had upvotes on the LB 0.559 public notebook. This raises suspicions that some private transaction notebooks were involved.",
              "votes": 3
            },
            {
              "id": 3414765,
              "postDate": "2026-02-27T15:57:16.170Z",
              "content": "<p>I've encountered this situation many times before.</p>",
              "rawMarkdown": "I've encountered this situation many times before.",
              "votes": 3
            },
            {
              "id": 3414767,
              "postDate": "2026-02-27T16:03:50.420Z",
              "content": "<p>If one goes snooping for Vesuvius content in the overall public Datasets who knows what they will find (yes some opsec issues for legit teams too but not really what I’m driving at)</p>",
              "rawMarkdown": "If one goes snooping for Vesuvius content in the overall public Datasets who knows what they will find (yes some opsec issues for legit teams too but not really what I’m driving at)",
              "votes": 1
            }
          ]
        },
        {
          "id": 3414772,
          "postDate": "2026-02-27T16:16:33.990Z",
          "content": "<p>Private share with/without payments and account selling. It's hard to ban them all in technical way. </p>",
          "rawMarkdown": "Private share with/without payments and account selling. It's hard to ban them all in technical way. ",
          "votes": 4,
          "replies": [
            {
              "id": 3414773,
              "postDate": "2026-02-27T16:21:05.333Z",
              "content": "<p>Those who made those deals were caught during the Santa Claus competition.</p>",
              "rawMarkdown": "Those who made those deals were caught during the Santa Claus competition.",
              "votes": 1
            },
            {
              "id": 3414918,
              "postDate": "2026-02-27T23:32:38.787Z",
              "content": "<p>Some of them are caught, But there are always survivers.</p>",
              "rawMarkdown": "Some of them are caught, But there are always survivers.",
              "votes": 4
            }
          ]
        },
        {
          "id": 3415684,
          "postDate": "2026-03-01T06:20:00.897Z",
          "content": "<p>I'm really looking forward to seeing #hui amazing performance.</p>",
          "rawMarkdown": "I'm really looking forward to seeing #hui amazing performance."
        }
      ]
    },
    {
      "id": 3414121,
      "postDate": "2026-02-26T01:28:27.750Z",
      "content": "<p>My model performs better in local cv and significantly outperforms other models on the host test samples, yet its score drops by 10 points on the LB. This issue has been troubling me for nearly a month and persists even after updating the metric data.</p>",
      "rawMarkdown": "My model performs better in local cv and significantly outperforms other models on the host test samples, yet its score drops by 10 points on the LB. This issue has been troubling me for nearly a month and persists even after updating the metric data.",
      "votes": 3,
      "replies": [
        {
          "id": 3414124,
          "postDate": "2026-02-26T01:46:48.340Z",
          "content": "<p>However, with the same model, if I only modify the post-processing, the scores on local CV and public LB seem to correspond basically.</p>",
          "rawMarkdown": "However, with the same model, if I only modify the post-processing, the scores on local CV and public LB seem to correspond basically.",
          "votes": 2,
          "replies": [
            {
              "id": 3414193,
              "postDate": "2026-02-26T08:10:27.833Z",
              "content": "<p>We've also encountered this situation.But it's not as high as 10%.</p>",
              "rawMarkdown": "We've also encountered this situation.But it's not as high as 10%.",
              "votes": 1
            },
            {
              "id": 3414236,
              "postDate": "2026-02-26T09:54:14.013Z",
              "content": "<blockquote>\n  <p>the scores on local CV and public LB seem to correspond basically.</p>\n</blockquote>\n<p>Thanks for sharing.</p>\n<p>I have 6% gap between CV and LB (CV above 0.61 with 1/6 of samples as valid). Looks like I should just tune post processing on the LB… Which I can't in two days, LOL.</p>",
              "rawMarkdown": "> the scores on local CV and public LB seem to correspond basically.\n\nThanks for sharing.\n\nI have 6% gap between CV and LB (CV above 0.61 with 1/6 of samples as valid). Looks like I should just tune post processing on the LB... Which I can't in two days, LOL."
            },
            {
              "id": 3414241,
              "postDate": "2026-02-26T10:06:45.967Z",
              "content": "<p>We did not draw a specific sample for verification, but I estimate the top-ranked CV exceeds 0.65</p>",
              "rawMarkdown": "We did not draw a specific sample for verification, but I estimate the top-ranked CV exceeds 0.65"
            },
            {
              "id": 3414272,
              "postDate": "2026-02-26T11:33:48.203Z",
              "content": "<p>What is yours?</p>",
              "rawMarkdown": "What is yours?"
            },
            {
              "id": 3414357,
              "postDate": "2026-02-26T15:49:19.100Z",
              "content": "<p>We have limited post processing that actually works. Our CV is quite high but nothing seems to impact LB these last 2 weeks. We can also see visually much better predictions that get much worse scores. The metric is just not bulletproof and 2 weeks was not nearly enough time to reevaluate exactly how everything works and retest all of our experiments. We went from first to fighting for the bottom of gold and our predictions look worse than they did 2 weeks ago visually. I fear this update is going to give the opposite of their intention unless the #1 team has something special and its not just an overlap tune. Excited to see the results. I am sure my family cannot wait for this competition to be over. </p>",
              "rawMarkdown": "We have limited post processing that actually works. Our CV is quite high but nothing seems to impact LB these last 2 weeks. We can also see visually much better predictions that get much worse scores. The metric is just not bulletproof and 2 weeks was not nearly enough time to reevaluate exactly how everything works and retest all of our experiments. We went from first to fighting for the bottom of gold and our predictions look worse than they did 2 weeks ago visually. I fear this update is going to give the opposite of their intention unless the #1 team has something special and its not just an overlap tune. Excited to see the results. I am sure my family cannot wait for this competition to be over. ",
              "votes": 3
            },
            {
              "id": 3414358,
              "postDate": "2026-02-26T15:54:24.220Z",
              "content": "<p><a href=\"https://www.kaggle.com/cody11null\" target=\"_blank\">@cody11null</a> I'm sure I'll learn from your solution. I hope I worked enough in this comp to understand what top teams did. I only scratched the surface on my side.</p>",
              "rawMarkdown": "@cody11null I'm sure I'll learn from your solution. I hope I worked enough in this comp to understand what top teams did. I only scratched the surface on my side.",
              "votes": 2
            },
            {
              "id": 3414389,
              "postDate": "2026-02-26T16:52:58.617Z",
              "content": "<p>Means a lot coming from you <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> I think we have something creative but we will see!</p>",
              "rawMarkdown": "Means a lot coming from you @cpmpml I think we have something creative but we will see!"
            },
            {
              "id": 3414552,
              "postDate": "2026-02-27T05:16:23.427Z",
              "content": "<p>the #1 team may just heavily overfit the public LB 😂\nWill share everything if not drop a lot</p>",
              "rawMarkdown": "the #1 team may just heavily overfit the public LB 😂\nWill share everything if not drop a lot",
              "votes": 2
            },
            {
              "id": 3414553,
              "postDate": "2026-02-27T05:20:42.993Z",
              "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> you might need to wait more days because I'm traveling for 1 weeks. I'm not sure my team can completely writeup the findings and developments.</p>",
              "rawMarkdown": "@cpmpml you might need to wait more days because I'm traveling for 1 weeks. I'm not sure my team can completely writeup the findings and developments.",
              "votes": 1
            },
            {
              "id": 3414554,
              "postDate": "2026-02-27T05:21:18.580Z",
              "content": "<p>We are too.</p>",
              "rawMarkdown": "We are too."
            },
            {
              "id": 3414757,
              "postDate": "2026-02-27T15:38:23.083Z",
              "content": "<p>Anyone else finding their post processing thresholds looking different for CV and LB? Makes me nervous the current values I’m using for LB are overfitting since I’m not using the same approach from before the test update</p>",
              "rawMarkdown": "Anyone else finding their post processing thresholds looking different for CV and LB? Makes me nervous the current values I’m using for LB are overfitting since I’m not using the same approach from before the test update",
              "votes": 2
            },
            {
              "id": 3414762,
              "postDate": "2026-02-27T15:44:09.040Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 3414763,
              "postDate": "2026-02-27T15:49:30.267Z",
              "content": "<p>u are right</p>",
              "rawMarkdown": "u are right",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3414010,
      "postDate": "2026-02-25T18:17:20.333Z",
      "content": "<p>Test data has been updated recently to remove holes in ground truth. I don't know if it explains what you are seeing.</p>\n<p>I can't comment on better topology effect as I only started submitting yesterday, I don't have history at all. Let's hope others will chime in to answer that.</p>",
      "rawMarkdown": "Test data has been updated recently to remove holes in ground truth. I don't know if it explains what you are seeing.\n\nI can't comment on better topology effect as I only started submitting yesterday, I don't have history at all. Let's hope others will chime in to answer that.",
      "votes": 3,
      "replies": [
        {
          "id": 3414029,
          "postDate": "2026-02-25T19:10:34.317Z",
          "content": "<p>Thanks! This does remind me that my local labels weren't patched for holes, which might explain why CV couldn't align. However, my main change was reducing adhesion, so issues about holes probably aren't the primary cause. My current hypothesis is that VOI scores penalize boundary misalignment in addition to stickiness. Perhaps my approach smoothed boundaries too much on the hidden dataset, incurring penalties. But I can't confirm this since VOI scores actually improved during local validation. I'd appreciate if experienced teams can share some insights.</p>",
          "rawMarkdown": "Thanks! This does remind me that my local labels weren't patched for holes, which might explain why CV couldn't align. However, my main change was reducing adhesion, so issues about holes probably aren't the primary cause. My current hypothesis is that VOI scores penalize boundary misalignment in addition to stickiness. Perhaps my approach smoothed boundaries too much on the hidden dataset, incurring penalties. But I can't confirm this since VOI scores actually improved during local validation. I'd appreciate if experienced teams can share some insights.",
          "replies": [
            {
              "id": 3414033,
              "postDate": "2026-02-25T19:16:19.560Z",
              "content": "<p>Others reported in this forum that removing merges decreased their LB. That's all I know :D</p>",
              "rawMarkdown": "Others reported in this forum that removing merges decreased their LB. That's all I know :D",
              "votes": 1
            },
            {
              "id": 3414526,
              "postDate": "2026-02-27T03:12:42.010Z",
              "content": "<p>There are a few issues here. A major one is that lots of multi-sheet components are actually a single papyrus sheet coming apart, so there's only a gt pred on the first one. </p>",
              "rawMarkdown": "There are a few issues here. A major one is that lots of multi-sheet components are actually a single papyrus sheet coming apart, so there's only a gt pred on the first one. \n",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3414245,
      "postDate": "2026-02-26T10:18:48.277Z",
      "content": "<p>There are two chances for final submission. You may now begin placing your bets.</p>\n<p>The ignore mask is way too random.</p>",
      "rawMarkdown": "There are two chances for final submission. You may now begin placing your bets.\n\nThe ignore mask is way too random.",
      "votes": 2,
      "replies": [
        {
          "id": 3414585,
          "postDate": "2026-02-27T08:09:46.210Z",
          "content": "<p>The last three times<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15255201%2Fd02acc80a5e17a901dcab3de7aa25a46%2FScreenshot%202026-02-27%20160820.png?generation=1772179782726603&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "The last three times![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15255201%2Fd02acc80a5e17a901dcab3de7aa25a46%2FScreenshot%202026-02-27%20160820.png?generation=1772179782726603&alt=media)",
          "votes": 2,
          "replies": [
            {
              "id": 3414930,
              "postDate": "2026-02-28T00:07:19.707Z",
              "content": "<p>Unfortunately didn't defeat dietr, but congrats on gold💪</p>",
              "rawMarkdown": "Unfortunately didn't defeat dietr, but congrats on gold💪",
              "votes": 1
            },
            {
              "id": 3415227,
              "postDate": "2026-02-28T13:05:34.197Z",
              "content": "<p>dream come true</p>",
              "rawMarkdown": "dream come true"
            }
          ]
        }
      ]
    },
    {
      "id": 3414046,
      "postDate": "2026-02-25T19:57:10.450Z",
      "content": "<p>Typically reducing merges and reducing holes should produce a higher LB score, because the LB metrics toposcore penalizes even minor errors quite heavily. It is however only a portion of the total LB score. </p>\n<p>Because surface dice in particular allows some wiggle room, you can often cheat a little towards maximizing topo vs surface dice and end up not taking a very large hit on the surface dice score, but this has some limits. </p>\n<p>This very difficult combination of metrics is extremely annoying, trust me when i say i have a ton of sympathy for the participants in this challenge as i too have been fighting these kinds of metrics on this data too, but this combination was chosen because each metric individually doesn't result in readability. </p>\n<ul>\n<li><p>A perfect toposcore gives the same number of components, and outputs with no holes or voids, but it does not <em>localize</em> the sheet , which means a model with a perfect toposcore but bad surface dice produces meshes which are too far from the written surface to run accurate ink detection on</p></li>\n<li><p>On the other hand, a model with good surface dice but bad topo metrics mean we may have good sheet localization, but we cannot mesh it and flatten it without complex algorithms for dealing with noisy surfaces. </p></li>\n</ul>\n<p>Our goal is to produce surface predictions which are well localised and topologically accurate (sane). A \"perfect\" model output of surface predictions ran on an entire scroll would be trivially \"meshable\"/\"flattenable\" with even the most simple algorithms. Even something like active contour models could resolve it alone.  Of course we will never have perfect predictions, but you could see how this interplay between how complex our unrolling software must be and how good the surface predictions are affects the metrics we have chosen, however imperfect they may be -- and they are imperfect, they are however the best we have been able to come up with for this complex task so far. </p>",
      "rawMarkdown": "Typically reducing merges and reducing holes should produce a higher LB score, because the LB metrics toposcore penalizes even minor errors quite heavily. It is however only a portion of the total LB score. \n\nBecause surface dice in particular allows some wiggle room, you can often cheat a little towards maximizing topo vs surface dice and end up not taking a very large hit on the surface dice score, but this has some limits. \n\nThis very difficult combination of metrics is extremely annoying, trust me when i say i have a ton of sympathy for the participants in this challenge as i too have been fighting these kinds of metrics on this data too, but this combination was chosen because each metric individually doesn't result in readability. \n\n- A perfect toposcore gives the same number of components, and outputs with no holes or voids, but it does not _localize_ the sheet , which means a model with a perfect toposcore but bad surface dice produces meshes which are too far from the written surface to run accurate ink detection on\n\n- On the other hand, a model with good surface dice but bad topo metrics mean we may have good sheet localization, but we cannot mesh it and flatten it without complex algorithms for dealing with noisy surfaces. \n\nOur goal is to produce surface predictions which are well localised and topologically accurate (sane). A \"perfect\" model output of surface predictions ran on an entire scroll would be trivially \"meshable\"/\"flattenable\" with even the most simple algorithms. Even something like active contour models could resolve it alone.  Of course we will never have perfect predictions, but you could see how this interplay between how complex our unrolling software must be and how good the surface predictions are affects the metrics we have chosen, however imperfect they may be -- and they are imperfect, they are however the best we have been able to come up with for this complex task so far. ",
      "votes": 2,
      "replies": [
        {
          "id": 3414107,
          "postDate": "2026-02-26T00:14:49.177Z",
          "content": "<p><a href=\"https://www.kaggle.com/seanjohnsonsp\" target=\"_blank\">@seanjohnsonsp</a> Its not that the \"This very difficult combination of metrics is extremely annoying\" its that the actual metric is still extremely broken. I've given up trying to help this competition since my other post got ignored by your colleague. There are layers of bugs in the metric, and its was clearly rushed and very lightly tested.</p>\n<p>It's not that the labels that you fixed were wrong, its that your metric actually has a flaw with discretization/aliasing - posted <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/672447\" target=\"_blank\">here</a>, so you might've fixed the errors in the labels after my post, but still the predictions that the competitors submit have a chance to have the same \"porous\" holes from the discretization, detected in them. I talk about this in the post and Georgio said that you <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/672447#3403393\" target=\"_blank\">intentionally wanted to punish these artifacts</a>? \"Ah, yes, we intentionally wanted to penalize these artifacts\". This made me realize that either he doesn't know what he is talking about, or he didnt bother to read/understand my post, since we are incapable of producing better predictions, and the problem is the metric code which suffers from aliasing, which is fixable. I've fixed it locally for my environment - or rather patched it.</p>\n<p>Second, I've spent enormous time (and I mean enormous) trying to come up with a model which solves the topology, but the current metric flaw with the ignore mask cutting the prediction and producing double digits components and holes, makes the competition a random dice throw on the private set, a game of \"lets play who will be unlucky to hit the mask harder\".  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2F711014c180190cc2489f7947865cd280%2Fsheet_holes.png?generation=1772061050836046&amp;alt=media\" alt=\"\"></p>\n<p>This sheet apparently have (0, 359, 0) 359 holes after being hit by the mask. From perfect sheet to 359 holes… honestly? Really?</p>\n<p>And what is tragic about this is that this probelm was mentioned 2 months ago, and there are practical things you can do as hosts to fix it. I explained one such solution to this <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/672634#3404181\" target=\"_blank\">here</a>. Sure it might not be perfect since I've thought about it for about 15 minutes by myself, but you are a team and in 2 months you would've done a much better job than me.</p>\n<p>Here is another clear example of how a model which 90% of the time produces almost perfect topology is affected by the ignore mask. Mind you, going from going from 0 holes to 1 hole, or 1 hole to 2 holes is a ENORMOUS difference in the leaderboard. I give one such calculation in <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/672447\" target=\"_blank\">this post</a>. But the bottom line is that going from 1 hole to 2 holes is going from 0.333 to 0.2 betti-1 score which when averaged with the betti-0 and multiplied by the metric topo weight of 0.3 is <strong>0.05</strong> vs <strong>0.03</strong> of <strong>leaderboard difference</strong>! I don't think many of the competitiors had time or interest to do the calculations and haven't realized what is the actual difference in the leaderboard is if everything worked as intended. So right now it <strong>doesnt matter</strong> if your model/algorithm produces perfect sheets. You thought you fixed the competition/leaderboard with the ground truth fix, but that was just the first layer of much worse condition. If previously the competition was deemed unfair, the same condition applies now.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2Fb2f5c208c27ee168da548dd56c175728%2FUntitled-2.jpg?generation=1772061706902834&amp;alt=media\" alt=\"\"></p>\n<p>This prediction is penalized heavily for being perfect and having no holes. And I am not saying that only my model is affected by this, everyone is also affected - <em>randomly</em> - but having random chance influence the leaderboard is not a truthful ML competition and the best solutions which produce best topology won't surface to the top. And if someone predicts perfect 0 hole sheets across all predictions, they are actually having a handicap of HALF the total topo score, which in terms of dice score equivalent is 0.5~ dice score difference. Imagine if teams randomly were penalized 0.5 dice score??  Right now all teams are clustered with roughly similar dice and are fighting for small percentages and imagine if some team predicted perfect sheet, they would BLOW past everyone else that predicted even few holes. I just want you to realize the MAGNITUDE of the error at hand. This is not something light, is actually mind boggling how wrong the table is right now, than if everything worked as intended. Please think about it for a second, grab a pen and paper and scribe down few scenarios of different teams predicting different holes, and do the small calculation.</p>\n<p>This is also one of the reasons why the author of the current post is having worse score even when the topology is better. Its because we are playing dice, on who will be less affected by the mask. My prediction is that there will be major shake up in the final standings, and the host can probably see it now in the private tab, since the variance in hitting the mask in the hidden 80% data in the test will affect the topo score more than the score difference between teams, and theres nothing the teams can do about it.</p>",
          "rawMarkdown": "@seanjohnsonsp Its not that the \"This very difficult combination of metrics is extremely annoying\" its that the actual metric is still extremely broken. I've given up trying to help this competition since my other post got ignored by your colleague. There are layers of bugs in the metric, and its was clearly rushed and very lightly tested.\n\nIt's not that the labels that you fixed were wrong, its that your metric actually has a flaw with discretization/aliasing - posted [here](https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/672447), so you might've fixed the errors in the labels after my post, but still the predictions that the competitors submit have a chance to have the same \"porous\" holes from the discretization, detected in them. I talk about this in the post and Georgio said that you [intentionally wanted to punish these artifacts](https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/672447#3403393)? \"Ah, yes, we intentionally wanted to penalize these artifacts\". This made me realize that either he doesn't know what he is talking about, or he didnt bother to read/understand my post, since we are incapable of producing better predictions, and the problem is the metric code which suffers from aliasing, which is fixable. I've fixed it locally for my environment - or rather patched it.\n\nSecond, I've spent enormous time (and I mean enormous) trying to come up with a model which solves the topology, but the current metric flaw with the ignore mask cutting the prediction and producing double digits components and holes, makes the competition a random dice throw on the private set, a game of \"lets play who will be unlucky to hit the mask harder\".  ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2F711014c180190cc2489f7947865cd280%2Fsheet_holes.png?generation=1772061050836046&alt=media)\n\nThis sheet apparently have (0, 359, 0) 359 holes after being hit by the mask. From perfect sheet to 359 holes... honestly? Really?\n\nAnd what is tragic about this is that this probelm was mentioned 2 months ago, and there are practical things you can do as hosts to fix it. I explained one such solution to this [here](https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/672634#3404181). Sure it might not be perfect since I've thought about it for about 15 minutes by myself, but you are a team and in 2 months you would've done a much better job than me.\n\nHere is another clear example of how a model which 90% of the time produces almost perfect topology is affected by the ignore mask. Mind you, going from going from 0 holes to 1 hole, or 1 hole to 2 holes is a ENORMOUS difference in the leaderboard. I give one such calculation in [this post](https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/672447). But the bottom line is that going from 1 hole to 2 holes is going from 0.333 to 0.2 betti-1 score which when averaged with the betti-0 and multiplied by the metric topo weight of 0.3 is **0.05** vs **0.03** of **leaderboard difference**! I don't think many of the competitiors had time or interest to do the calculations and haven't realized what is the actual difference in the leaderboard is if everything worked as intended. So right now it **doesnt matter** if your model/algorithm produces perfect sheets. You thought you fixed the competition/leaderboard with the ground truth fix, but that was just the first layer of much worse condition. If previously the competition was deemed unfair, the same condition applies now.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2Fb2f5c208c27ee168da548dd56c175728%2FUntitled-2.jpg?generation=1772061706902834&alt=media)\n\nThis prediction is penalized heavily for being perfect and having no holes. And I am not saying that only my model is affected by this, everyone is also affected - *randomly* - but having random chance influence the leaderboard is not a truthful ML competition and the best solutions which produce best topology won't surface to the top. And if someone predicts perfect 0 hole sheets across all predictions, they are actually having a handicap of HALF the total topo score, which in terms of dice score equivalent is 0.5~ dice score difference. Imagine if teams randomly were penalized 0.5 dice score??  Right now all teams are clustered with roughly similar dice and are fighting for small percentages and imagine if some team predicted perfect sheet, they would BLOW past everyone else that predicted even few holes. I just want you to realize the MAGNITUDE of the error at hand. This is not something light, is actually mind boggling how wrong the table is right now, than if everything worked as intended. Please think about it for a second, grab a pen and paper and scribe down few scenarios of different teams predicting different holes, and do the small calculation.\n\nThis is also one of the reasons why the author of the current post is having worse score even when the topology is better. Its because we are playing dice, on who will be less affected by the mask. My prediction is that there will be major shake up in the final standings, and the host can probably see it now in the private tab, since the variance in hitting the mask in the hidden 80% data in the test will affect the topo score more than the score difference between teams, and theres nothing the teams can do about it.\n\n\n\n\n",
          "votes": 9
        },
        {
          "id": 3414128,
          "postDate": "2026-02-26T02:02:13.827Z",
          "content": "<p>On top of all this, you have a <strong>basic</strong> bug which makes no intuitive sense, in the calculation of the Topo Score by taking into account only the active dimensions and treating inactive dimensions with no weight. This is a simple scenario.</p>\n<p>Scenario 1: You score 0.2 on betti-0, and nan (0 matched, 0 predicted, 0 gt holes) on betti-1, you get 0.2 topo score.</p>\n<p>Scenario 2:  You score 0.2 on betti-0, and 0.333 (0 matched, 1 predicted, 0 gt holes) on betti-1, you get (0.2+0.333) / 2 = 0.266 topo score</p>\n<p>So you actually get HIGHER score by making worse prediction? <strong>You should get HIGHER SCORE if you predict less holes!</strong> If 1.0 is the maximum topo score, then each betti dimension should have influence equally in the final score, and right now if you predict perfect betti-1 in the scenario above you actually get worse score than if you predicted just 1 hole.</p>",
          "rawMarkdown": "On top of all this, you have a **basic** bug which makes no intuitive sense, in the calculation of the Topo Score by taking into account only the active dimensions and treating inactive dimensions with no weight. This is a simple scenario.\n\nScenario 1: You score 0.2 on betti-0, and nan (0 matched, 0 predicted, 0 gt holes) on betti-1, you get 0.2 topo score.\n\nScenario 2:  You score 0.2 on betti-0, and 0.333 (0 matched, 1 predicted, 0 gt holes) on betti-1, you get (0.2+0.333) / 2 = 0.266 topo score\n\nSo you actually get HIGHER score by making worse prediction? **You should get HIGHER SCORE if you predict less holes!** If 1.0 is the maximum topo score, then each betti dimension should have influence equally in the final score, and right now if you predict perfect betti-1 in the scenario above you actually get worse score than if you predicted just 1 hole.",
          "votes": 7,
          "replies": [
            {
              "id": 3414194,
              "postDate": "2026-02-26T08:11:04.383Z",
              "content": "<p>Hello Dan,</p>\n<p>Thanks for your thoughtful messages and for taking the time to flag these points.</p>\n<p>As I mentioned earlier, there are a couple of different things going on: some behaviors were tolerated as part of the design trade-offs, while others were not intended and only became clear after the competition started. The latter involved the dataset and are the ones we decided to fix.</p>\n<p>On the “double scenario” you described: we were already aware of it. In practice it leads to a counterintuitive score improvement only when one other score is below ~0.33. Our expectation was that top leaderboard entries would generally be above that range, so the practical impact at the top end should be limited.</p>\n<p>On the ignore-mask concern: this was raised by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> a few weeks after launch. For our annotation team, using an ignore mask was necessary to build a dataset with enough coverage in the challenging regions we care about most. The goal of the competition isn’t only to run a challenge for its own sake (or to offer prizes), but to get help solving a real technical problem. If we could reliably produce large amounts of fully clean data in those difficult areas, we wouldn’t have needed to run a competition in the first place.</p>\n<p>Regarding this point, we also consulted Kaggle support before launch; the guidance we received was that if an issue cannot be fully fixed and it affects everyone comparably, the competition can still be considered fair.</p>\n<p>On your concerns about possible shake-ups: some movement is normal in most Kaggle competitions. Based on what we’ve seen so far, we don’t expect anything unusually extreme, and likely not more than what we’ve seen in prior 3D segmentation competitions (e.g., <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/leaderboard\" target=\"_blank\">SenNet + HOA</a> ).</p>\n<p>More broadly, I understand the worry about scenarios where small details dominate the score. That said, we’ve (only visually) inspected a sample of predictions from a range of submissions. They look much better than our baseline, but we’re not seeing “perfect sheets everywhere\".</p>\n<p>If many competitors are worried about a “dice roll” outcome, the biggest risk I’d caution against is actually heavy overfitting.</p>\n<p>To be transparent: we know the metric isn’t perfect, and we know the dataset isn’t perfect. We did our best to make the setup useful and as fair as possible within technical constraints. In general, we’re happy to keep the discussion constructive.</p>\n<p>We are immensely grateful to all the participants who helped making this competition better, including you, and we’ll take lessons learned into account for future work.</p>",
              "rawMarkdown": "Hello Dan,\n\nThanks for your thoughtful messages and for taking the time to flag these points.\n\nAs I mentioned earlier, there are a couple of different things going on: some behaviors were tolerated as part of the design trade-offs, while others were not intended and only became clear after the competition started. The latter involved the dataset and are the ones we decided to fix.\n\nOn the “double scenario” you described: we were already aware of it. In practice it leads to a counterintuitive score improvement only when one other score is below ~0.33. Our expectation was that top leaderboard entries would generally be above that range, so the practical impact at the top end should be limited.\n\nOn the ignore-mask concern: this was raised by @hengck23 a few weeks after launch. For our annotation team, using an ignore mask was necessary to build a dataset with enough coverage in the challenging regions we care about most. The goal of the competition isn’t only to run a challenge for its own sake (or to offer prizes), but to get help solving a real technical problem. If we could reliably produce large amounts of fully clean data in those difficult areas, we wouldn’t have needed to run a competition in the first place.\n\nRegarding this point, we also consulted Kaggle support before launch; the guidance we received was that if an issue cannot be fully fixed and it affects everyone comparably, the competition can still be considered fair.\n\nOn your concerns about possible shake-ups: some movement is normal in most Kaggle competitions. Based on what we’ve seen so far, we don’t expect anything unusually extreme, and likely not more than what we’ve seen in prior 3D segmentation competitions (e.g., [SenNet + HOA](https://www.kaggle.com/competitions/blood-vessel-segmentation/leaderboard) ).\n\nMore broadly, I understand the worry about scenarios where small details dominate the score. That said, we’ve (only visually) inspected a sample of predictions from a range of submissions. They look much better than our baseline, but we’re not seeing “perfect sheets everywhere\".\n\nIf many competitors are worried about a “dice roll” outcome, the biggest risk I’d caution against is actually heavy overfitting.\n\nTo be transparent: we know the metric isn’t perfect, and we know the dataset isn’t perfect. We did our best to make the setup useful and as fair as possible within technical constraints. In general, we’re happy to keep the discussion constructive.\n\nWe are immensely grateful to all the participants who helped making this competition better, including you, and we’ll take lessons learned into account for future work.",
              "votes": 7
            },
            {
              "id": 3414315,
              "postDate": "2026-02-26T14:13:38.950Z",
              "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18473244%2Fd784a16ddd488cbe950c969ccafbdd8f%2FScreenshot%202026-02-26%20170944.png?generation=1772115002559383&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18473244%2Ff228f7127258fa7bf59c094d04061238%2F1h37zl.jpg?generation=1772115189602671&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18473244%2Fd784a16ddd488cbe950c969ccafbdd8f%2FScreenshot%202026-02-26%20170944.png?generation=1772115002559383&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18473244%2Ff228f7127258fa7bf59c094d04061238%2F1h37zl.jpg?generation=1772115189602671&alt=media)\n",
              "votes": 5
            },
            {
              "id": 3414343,
              "postDate": "2026-02-26T15:06:01.813Z",
              "content": "<p>It's possible.</p>",
              "rawMarkdown": "It's possible.",
              "votes": 2
            },
            {
              "id": 3414428,
              "postDate": "2026-02-26T18:41:36.617Z",
              "content": "<p>Is February 28, 2026, 00:00 AM UTC when the Private Leaderboard will be made public, given that February 27, 2026, 11:59 PM UTC is the final submission deadline?</p>",
              "rawMarkdown": "Is February 28, 2026, 00:00 AM UTC when the Private Leaderboard will be made public, given that February 27, 2026, 11:59 PM UTC is the final submission deadline?",
              "votes": 1
            },
            {
              "id": 3414490,
              "postDate": "2026-02-26T22:49:05.357Z",
              "content": "<p>the moment the competition  finish  the private lb will show up, unless mentioned in the overview page</p>",
              "rawMarkdown": "the moment the competition  finish  the private lb will show up, unless mentioned in the overview page",
              "votes": 2
            },
            {
              "id": 3414548,
              "postDate": "2026-02-27T05:14:28.757Z",
              "content": "<p>based on my experiences, will see results at 11:59:00 UTC</p>",
              "rawMarkdown": "based on my experiences, will see results at 11:59:00 UTC",
              "votes": 3
            },
            {
              "id": 3414852,
              "postDate": "2026-02-27T19:34:14.930Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 3414862,
              "postDate": "2026-02-27T20:15:10.727Z",
              "content": "<p>Isn't the deadline for <em>starting</em> a submission 11:59pm UTC? Some submissions will run for hours. </p>",
              "rawMarkdown": "Isn't the deadline for *starting* a submission 11:59pm UTC? Some submissions will run for hours. "
            },
            {
              "id": 3414863,
              "postDate": "2026-02-27T20:16:13.960Z",
              "content": "<p>I'm new to kaggle. Isn't that time the deadline for <em>starting</em> a submission? A run could take hours. </p>",
              "rawMarkdown": "I'm new to kaggle. Isn't that time the deadline for *starting* a submission? A run could take hours. "
            }
          ]
        },
        {
          "id": 3414532,
          "postDate": "2026-02-27T04:07:36.240Z",
          "content": "<p>Just use Hausdorff properly, you'll implicitly get everything you're looking for. You should fire this up as a sanity-check against the current metric. <a href=\"https://github.com/cnr-isti-vclab/vcglib/tree/main/apps/metro\" target=\"_blank\">https://github.com/cnr-isti-vclab/vcglib/tree/main/apps/metro</a></p>\n<p>It's a standard metric for the application you have here. Have a read: <a href=\"https://www.researchgate.net/publication/227611629_METRO_Measuring_error_on_simplified_surfaces\" target=\"_blank\">https://www.researchgate.net/publication/227611629_METRO_Measuring_error_on_simplified_surfaces</a></p>\n<p>It would be illuminating to run metro on the LB. </p>\n<p>edit - you could round metro out with running Betti numbers on a mesh, that's decent. But once you leave mesh-land you lose your simplicial homology and everything is a disaster approximation. Convert everything to meshes for coherent metrics. </p>",
          "rawMarkdown": "Just use Hausdorff properly, you'll implicitly get everything you're looking for. You should fire this up as a sanity-check against the current metric. https://github.com/cnr-isti-vclab/vcglib/tree/main/apps/metro\n\nIt's a standard metric for the application you have here. Have a read: https://www.researchgate.net/publication/227611629_METRO_Measuring_error_on_simplified_surfaces\n\nIt would be illuminating to run metro on the LB. \n\nedit - you could round metro out with running Betti numbers on a mesh, that's decent. But once you leave mesh-land you lose your simplicial homology and everything is a disaster approximation. Convert everything to meshes for coherent metrics. "
        }
      ]
    },
    {
      "id": 3414195,
      "postDate": "2026-02-26T08:14:49.667Z",
      "content": "<p>Whenever your prediction hits an ignore mask, your score will plummet—and it's not your fault.</p>",
      "rawMarkdown": "Whenever your prediction hits an ignore mask, your score will plummet—and it's not your fault.",
      "replies": [
        {
          "id": 3414523,
          "postDate": "2026-02-27T02:50:48.377Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 3414850,
      "postDate": "2026-02-27T19:30:38.883Z",
      "content": "<p>That’s a very interesting point. If topology improves but LB drops, it suggests the metric may be rewarding something slightly different than what we expect. Would be great to hear whether others have observed this trade-off between DIC optimization and leaderboard performance.</p>",
      "rawMarkdown": "That’s a very interesting point. If topology improves but LB drops, it suggests the metric may be rewarding something slightly different than what we expect. Would be great to hear whether others have observed this trade-off between DIC optimization and leaderboard performance."
    }
  ],
  "comments": [
    {
      "id": 3414592,
      "author_name": "tingyi",
      "author_url": "",
      "post_date": "2026-02-27T08:17:32.270000",
      "content": "<p>Why have a few assassins suddenly appeared these last few days?</p>",
      "votes": 6,
      "replies": [
        {
          "id": 3414689,
          "author_name": "Francois Lemarchand",
          "author_url": "",
          "post_date": "2026-02-27T12:12:17.780000",
          "content": "<p>I also noticed this! I wish I could have gone on holidays and then reach the top of the leaderboard a day before the competition ends… I'm sure the Kaggle team will check that everything is \"fair\".</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 3414728,
          "author_name": "JasoOKA",
          "author_url": "",
          "post_date": "2026-02-27T14:40:39.127000",
          "content": "<p>It seems weird. Because the public notebooks LB 0.566 and 0.558 with the private notebooks, still so many teams can get great scores with less than 10 submissions </p>",
          "votes": 2,
          "replies": [
            {
              "id": 3414748,
              "author_name": "Rob Freeman",
              "author_url": "",
              "post_date": "2026-02-27T15:26:19.300000",
              "content": "<p>Can they with the .566? The models are hidden behind unshared datasets, although reproduction might be straightforward.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3414754,
              "author_name": "JasoOKA",
              "author_url": "",
              "post_date": "2026-02-27T15:34:11.843000",
              "content": "<p>This is also what puzzles me, because other competitions are copied quickly because the resources are publicly available.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3414758,
              "author_name": "Rob Freeman",
              "author_url": "",
              "post_date": "2026-02-27T15:39:23.133000",
              "content": "<p>I assume this is a clever technique with the auto +1 on editing your copy to gain status</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3414764,
              "author_name": "JasoOKA",
              "author_url": "",
              "post_date": "2026-02-27T15:51:44.233000",
              "content": "<p>I previously confirmed that teams that suddenly appeared with a 0.559 score (with fewer than 15 submissions) had upvotes on the LB 0.559 public notebook. This raises suspicions that some private transaction notebooks were involved.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 3414765,
              "author_name": "GG Ayo (AyoGG)",
              "author_url": "",
              "post_date": "2026-02-27T15:57:16.170000",
              "content": "<p>I've encountered this situation many times before.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 3414767,
              "author_name": "Rob Freeman",
              "author_url": "",
              "post_date": "2026-02-27T16:03:50.420000",
              "content": "<p>If one goes snooping for Vesuvius content in the overall public Datasets who knows what they will find (yes some opsec issues for legit teams too but not really what I’m driving at)</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 3414772,
          "author_name": "ForcewithMe",
          "author_url": "",
          "post_date": "2026-02-27T16:16:33.990000",
          "content": "<p>Private share with/without payments and account selling. It's hard to ban them all in technical way. </p>",
          "votes": 4,
          "replies": [
            {
              "id": 3414773,
              "author_name": "JasoOKA",
              "author_url": "",
              "post_date": "2026-02-27T16:21:05.333000",
              "content": "<p>Those who made those deals were caught during the Santa Claus competition.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3414918,
              "author_name": "ForcewithMe",
              "author_url": "",
              "post_date": "2026-02-27T23:32:38.787000",
              "content": "<p>Some of them are caught, But there are always survivers.</p>",
              "votes": 4,
              "replies": []
            }
          ]
        },
        {
          "id": 3415684,
          "author_name": "GG Ayo (AyoGG)",
          "author_url": "",
          "post_date": "2026-03-01T06:20:00.897000",
          "content": "<p>I'm really looking forward to seeing #hui amazing performance.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3414121,
      "author_name": "huoxu",
      "author_url": "",
      "post_date": "2026-02-26T01:28:27.750000",
      "content": "<p>My model performs better in local cv and significantly outperforms other models on the host test samples, yet its score drops by 10 points on the LB. This issue has been troubling me for nearly a month and persists even after updating the metric data.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3414124,
          "author_name": "huoxu",
          "author_url": "",
          "post_date": "2026-02-26T01:46:48.340000",
          "content": "<p>However, with the same model, if I only modify the post-processing, the scores on local CV and public LB seem to correspond basically.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3414193,
              "author_name": "GG Ayo (AyoGG)",
              "author_url": "",
              "post_date": "2026-02-26T08:10:27.833000",
              "content": "<p>We've also encountered this situation.But it's not as high as 10%.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3414236,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2026-02-26T09:54:14.013000",
              "content": "<blockquote>\n  <p>the scores on local CV and public LB seem to correspond basically.</p>\n</blockquote>\n<p>Thanks for sharing.</p>\n<p>I have 6% gap between CV and LB (CV above 0.61 with 1/6 of samples as valid). Looks like I should just tune post processing on the LB… Which I can't in two days, LOL.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3414241,
              "author_name": "GG Ayo (AyoGG)",
              "author_url": "",
              "post_date": "2026-02-26T10:06:45.967000",
              "content": "<p>We did not draw a specific sample for verification, but I estimate the top-ranked CV exceeds 0.65</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3414272,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2026-02-26T11:33:48.203000",
              "content": "<p>What is yours?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3414357,
              "author_name": "Cody_Null",
              "author_url": "",
              "post_date": "2026-02-26T15:49:19.100000",
              "content": "<p>We have limited post processing that actually works. Our CV is quite high but nothing seems to impact LB these last 2 weeks. We can also see visually much better predictions that get much worse scores. The metric is just not bulletproof and 2 weeks was not nearly enough time to reevaluate exactly how everything works and retest all of our experiments. We went from first to fighting for the bottom of gold and our predictions look worse than they did 2 weeks ago visually. I fear this update is going to give the opposite of their intention unless the #1 team has something special and its not just an overlap tune. Excited to see the results. I am sure my family cannot wait for this competition to be over. </p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 3414358,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2026-02-26T15:54:24.220000",
              "content": "<p><a href=\"https://www.kaggle.com/cody11null\" target=\"_blank\">@cody11null</a> I'm sure I'll learn from your solution. I hope I worked enough in this comp to understand what top teams did. I only scratched the surface on my side.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3414389,
              "author_name": "Cody_Null",
              "author_url": "",
              "post_date": "2026-02-26T16:52:58.617000",
              "content": "<p>Means a lot coming from you <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> I think we have something creative but we will see!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3414552,
              "author_name": "Yiheng Wang",
              "author_url": "",
              "post_date": "2026-02-27T05:16:23.427000",
              "content": "<p>the #1 team may just heavily overfit the public LB 😂\nWill share everything if not drop a lot</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3414553,
              "author_name": "Tom",
              "author_url": "",
              "post_date": "2026-02-27T05:20:42.993000",
              "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> you might need to wait more days because I'm traveling for 1 weeks. I'm not sure my team can completely writeup the findings and developments.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3414554,
              "author_name": "GG Ayo (AyoGG)",
              "author_url": "",
              "post_date": "2026-02-27T05:21:18.580000",
              "content": "<p>We are too.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3414757,
              "author_name": "Rob Freeman",
              "author_url": "",
              "post_date": "2026-02-27T15:38:23.083000",
              "content": "<p>Anyone else finding their post processing thresholds looking different for CV and LB? Makes me nervous the current values I’m using for LB are overfitting since I’m not using the same approach from before the test update</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3414762,
              "author_name": "",
              "author_url": "",
              "post_date": "2026-02-27T15:44:09.040000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3414763,
              "author_name": "GG Ayo (AyoGG)",
              "author_url": "",
              "post_date": "2026-02-27T15:49:30.267000",
              "content": "<p>u are right</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3414010,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2026-02-25T18:17:20.333000",
      "content": "<p>Test data has been updated recently to remove holes in ground truth. I don't know if it explains what you are seeing.</p>\n<p>I can't comment on better topology effect as I only started submitting yesterday, I don't have history at all. Let's hope others will chime in to answer that.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3414029,
          "author_name": "lllleeeo",
          "author_url": "",
          "post_date": "2026-02-25T19:10:34.317000",
          "content": "<p>Thanks! This does remind me that my local labels weren't patched for holes, which might explain why CV couldn't align. However, my main change was reducing adhesion, so issues about holes probably aren't the primary cause. My current hypothesis is that VOI scores penalize boundary misalignment in addition to stickiness. Perhaps my approach smoothed boundaries too much on the hidden dataset, incurring penalties. But I can't confirm this since VOI scores actually improved during local validation. I'd appreciate if experienced teams can share some insights.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3414033,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2026-02-25T19:16:19.560000",
              "content": "<p>Others reported in this forum that removing merges decreased their LB. That's all I know :D</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3414526,
              "author_name": "/|\\^||^\\|/",
              "author_url": "",
              "post_date": "2026-02-27T03:12:42.010000",
              "content": "<p>There are a few issues here. A major one is that lots of multi-sheet components are actually a single papyrus sheet coming apart, so there's only a gt pred on the first one. </p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3414245,
      "author_name": "GG Ayo (AyoGG)",
      "author_url": "",
      "post_date": "2026-02-26T10:18:48.277000",
      "content": "<p>There are two chances for final submission. You may now begin placing your bets.</p>\n<p>The ignore mask is way too random.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3414585,
          "author_name": "GG Ayo (AyoGG)",
          "author_url": "",
          "post_date": "2026-02-27T08:09:46.210000",
          "content": "<p>The last three times<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15255201%2Fd02acc80a5e17a901dcab3de7aa25a46%2FScreenshot%202026-02-27%20160820.png?generation=1772179782726603&amp;alt=media\" alt=\"\"></p>",
          "votes": 2,
          "replies": [
            {
              "id": 3414930,
              "author_name": "Taha_Alshatiri",
              "author_url": "",
              "post_date": "2026-02-28T00:07:19.707000",
              "content": "<p>Unfortunately didn't defeat dietr, but congrats on gold💪</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3415227,
              "author_name": "GG Ayo (AyoGG)",
              "author_url": "",
              "post_date": "2026-02-28T13:05:34.197000",
              "content": "<p>dream come true</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3414046,
      "author_name": "Sean Johnson_SP",
      "author_url": "",
      "post_date": "2026-02-25T19:57:10.450000",
      "content": "<p>Typically reducing merges and reducing holes should produce a higher LB score, because the LB metrics toposcore penalizes even minor errors quite heavily. It is however only a portion of the total LB score. </p>\n<p>Because surface dice in particular allows some wiggle room, you can often cheat a little towards maximizing topo vs surface dice and end up not taking a very large hit on the surface dice score, but this has some limits. </p>\n<p>This very difficult combination of metrics is extremely annoying, trust me when i say i have a ton of sympathy for the participants in this challenge as i too have been fighting these kinds of metrics on this data too, but this combination was chosen because each metric individually doesn't result in readability. </p>\n<ul>\n<li><p>A perfect toposcore gives the same number of components, and outputs with no holes or voids, but it does not <em>localize</em> the sheet , which means a model with a perfect toposcore but bad surface dice produces meshes which are too far from the written surface to run accurate ink detection on</p></li>\n<li><p>On the other hand, a model with good surface dice but bad topo metrics mean we may have good sheet localization, but we cannot mesh it and flatten it without complex algorithms for dealing with noisy surfaces. </p></li>\n</ul>\n<p>Our goal is to produce surface predictions which are well localised and topologically accurate (sane). A \"perfect\" model output of surface predictions ran on an entire scroll would be trivially \"meshable\"/\"flattenable\" with even the most simple algorithms. Even something like active contour models could resolve it alone.  Of course we will never have perfect predictions, but you could see how this interplay between how complex our unrolling software must be and how good the surface predictions are affects the metrics we have chosen, however imperfect they may be -- and they are imperfect, they are however the best we have been able to come up with for this complex task so far. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 3414107,
          "author_name": "dan4o",
          "author_url": "",
          "post_date": "2026-02-26T00:14:49.177000",
          "content": "<p><a href=\"https://www.kaggle.com/seanjohnsonsp\" target=\"_blank\">@seanjohnsonsp</a> Its not that the \"This very difficult combination of metrics is extremely annoying\" its that the actual metric is still extremely broken. I've given up trying to help this competition since my other post got ignored by your colleague. There are layers of bugs in the metric, and its was clearly rushed and very lightly tested.</p>\n<p>It's not that the labels that you fixed were wrong, its that your metric actually has a flaw with discretization/aliasing - posted <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/672447\" target=\"_blank\">here</a>, so you might've fixed the errors in the labels after my post, but still the predictions that the competitors submit have a chance to have the same \"porous\" holes from the discretization, detected in them. I talk about this in the post and Georgio said that you <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/672447#3403393\" target=\"_blank\">intentionally wanted to punish these artifacts</a>? \"Ah, yes, we intentionally wanted to penalize these artifacts\". This made me realize that either he doesn't know what he is talking about, or he didnt bother to read/understand my post, since we are incapable of producing better predictions, and the problem is the metric code which suffers from aliasing, which is fixable. I've fixed it locally for my environment - or rather patched it.</p>\n<p>Second, I've spent enormous time (and I mean enormous) trying to come up with a model which solves the topology, but the current metric flaw with the ignore mask cutting the prediction and producing double digits components and holes, makes the competition a random dice throw on the private set, a game of \"lets play who will be unlucky to hit the mask harder\".  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2F711014c180190cc2489f7947865cd280%2Fsheet_holes.png?generation=1772061050836046&amp;alt=media\" alt=\"\"></p>\n<p>This sheet apparently have (0, 359, 0) 359 holes after being hit by the mask. From perfect sheet to 359 holes… honestly? Really?</p>\n<p>And what is tragic about this is that this probelm was mentioned 2 months ago, and there are practical things you can do as hosts to fix it. I explained one such solution to this <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/672634#3404181\" target=\"_blank\">here</a>. Sure it might not be perfect since I've thought about it for about 15 minutes by myself, but you are a team and in 2 months you would've done a much better job than me.</p>\n<p>Here is another clear example of how a model which 90% of the time produces almost perfect topology is affected by the ignore mask. Mind you, going from going from 0 holes to 1 hole, or 1 hole to 2 holes is a ENORMOUS difference in the leaderboard. I give one such calculation in <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/672447\" target=\"_blank\">this post</a>. But the bottom line is that going from 1 hole to 2 holes is going from 0.333 to 0.2 betti-1 score which when averaged with the betti-0 and multiplied by the metric topo weight of 0.3 is <strong>0.05</strong> vs <strong>0.03</strong> of <strong>leaderboard difference</strong>! I don't think many of the competitiors had time or interest to do the calculations and haven't realized what is the actual difference in the leaderboard is if everything worked as intended. So right now it <strong>doesnt matter</strong> if your model/algorithm produces perfect sheets. You thought you fixed the competition/leaderboard with the ground truth fix, but that was just the first layer of much worse condition. If previously the competition was deemed unfair, the same condition applies now.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14754958%2Fb2f5c208c27ee168da548dd56c175728%2FUntitled-2.jpg?generation=1772061706902834&amp;alt=media\" alt=\"\"></p>\n<p>This prediction is penalized heavily for being perfect and having no holes. And I am not saying that only my model is affected by this, everyone is also affected - <em>randomly</em> - but having random chance influence the leaderboard is not a truthful ML competition and the best solutions which produce best topology won't surface to the top. And if someone predicts perfect 0 hole sheets across all predictions, they are actually having a handicap of HALF the total topo score, which in terms of dice score equivalent is 0.5~ dice score difference. Imagine if teams randomly were penalized 0.5 dice score??  Right now all teams are clustered with roughly similar dice and are fighting for small percentages and imagine if some team predicted perfect sheet, they would BLOW past everyone else that predicted even few holes. I just want you to realize the MAGNITUDE of the error at hand. This is not something light, is actually mind boggling how wrong the table is right now, than if everything worked as intended. Please think about it for a second, grab a pen and paper and scribe down few scenarios of different teams predicting different holes, and do the small calculation.</p>\n<p>This is also one of the reasons why the author of the current post is having worse score even when the topology is better. Its because we are playing dice, on who will be less affected by the mask. My prediction is that there will be major shake up in the final standings, and the host can probably see it now in the private tab, since the variance in hitting the mask in the hidden 80% data in the test will affect the topo score more than the score difference between teams, and theres nothing the teams can do about it.</p>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 3414128,
          "author_name": "dan4o",
          "author_url": "",
          "post_date": "2026-02-26T02:02:13.827000",
          "content": "<p>On top of all this, you have a <strong>basic</strong> bug which makes no intuitive sense, in the calculation of the Topo Score by taking into account only the active dimensions and treating inactive dimensions with no weight. This is a simple scenario.</p>\n<p>Scenario 1: You score 0.2 on betti-0, and nan (0 matched, 0 predicted, 0 gt holes) on betti-1, you get 0.2 topo score.</p>\n<p>Scenario 2:  You score 0.2 on betti-0, and 0.333 (0 matched, 1 predicted, 0 gt holes) on betti-1, you get (0.2+0.333) / 2 = 0.266 topo score</p>\n<p>So you actually get HIGHER score by making worse prediction? <strong>You should get HIGHER SCORE if you predict less holes!</strong> If 1.0 is the maximum topo score, then each betti dimension should have influence equally in the final score, and right now if you predict perfect betti-1 in the scenario above you actually get worse score than if you predicted just 1 hole.</p>",
          "votes": 7,
          "replies": [
            {
              "id": 3414194,
              "author_name": "Giorgio Angelotti",
              "author_url": "",
              "post_date": "2026-02-26T08:11:04.383000",
              "content": "<p>Hello Dan,</p>\n<p>Thanks for your thoughtful messages and for taking the time to flag these points.</p>\n<p>As I mentioned earlier, there are a couple of different things going on: some behaviors were tolerated as part of the design trade-offs, while others were not intended and only became clear after the competition started. The latter involved the dataset and are the ones we decided to fix.</p>\n<p>On the “double scenario” you described: we were already aware of it. In practice it leads to a counterintuitive score improvement only when one other score is below ~0.33. Our expectation was that top leaderboard entries would generally be above that range, so the practical impact at the top end should be limited.</p>\n<p>On the ignore-mask concern: this was raised by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> a few weeks after launch. For our annotation team, using an ignore mask was necessary to build a dataset with enough coverage in the challenging regions we care about most. The goal of the competition isn’t only to run a challenge for its own sake (or to offer prizes), but to get help solving a real technical problem. If we could reliably produce large amounts of fully clean data in those difficult areas, we wouldn’t have needed to run a competition in the first place.</p>\n<p>Regarding this point, we also consulted Kaggle support before launch; the guidance we received was that if an issue cannot be fully fixed and it affects everyone comparably, the competition can still be considered fair.</p>\n<p>On your concerns about possible shake-ups: some movement is normal in most Kaggle competitions. Based on what we’ve seen so far, we don’t expect anything unusually extreme, and likely not more than what we’ve seen in prior 3D segmentation competitions (e.g., <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/leaderboard\" target=\"_blank\">SenNet + HOA</a> ).</p>\n<p>More broadly, I understand the worry about scenarios where small details dominate the score. That said, we’ve (only visually) inspected a sample of predictions from a range of submissions. They look much better than our baseline, but we’re not seeing “perfect sheets everywhere\".</p>\n<p>If many competitors are worried about a “dice roll” outcome, the biggest risk I’d caution against is actually heavy overfitting.</p>\n<p>To be transparent: we know the metric isn’t perfect, and we know the dataset isn’t perfect. We did our best to make the setup useful and as fair as possible within technical constraints. In general, we’re happy to keep the discussion constructive.</p>\n<p>We are immensely grateful to all the participants who helped making this competition better, including you, and we’ll take lessons learned into account for future work.</p>",
              "votes": 7,
              "replies": []
            },
            {
              "id": 3414315,
              "author_name": "Taha_Alshatiri",
              "author_url": "",
              "post_date": "2026-02-26T14:13:38.950000",
              "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18473244%2Fd784a16ddd488cbe950c969ccafbdd8f%2FScreenshot%202026-02-26%20170944.png?generation=1772115002559383&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18473244%2Ff228f7127258fa7bf59c094d04061238%2F1h37zl.jpg?generation=1772115189602671&amp;alt=media\" alt=\"\"></p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 3414343,
              "author_name": "ForcewithMe",
              "author_url": "",
              "post_date": "2026-02-26T15:06:01.813000",
              "content": "<p>It's possible.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3414428,
              "author_name": "GG Ayo (AyoGG)",
              "author_url": "",
              "post_date": "2026-02-26T18:41:36.617000",
              "content": "<p>Is February 28, 2026, 00:00 AM UTC when the Private Leaderboard will be made public, given that February 27, 2026, 11:59 PM UTC is the final submission deadline?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3414490,
              "author_name": "Taha_Alshatiri",
              "author_url": "",
              "post_date": "2026-02-26T22:49:05.357000",
              "content": "<p>the moment the competition  finish  the private lb will show up, unless mentioned in the overview page</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3414548,
              "author_name": "Yiheng Wang",
              "author_url": "",
              "post_date": "2026-02-27T05:14:28.757000",
              "content": "<p>based on my experiences, will see results at 11:59:00 UTC</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 3414852,
              "author_name": "",
              "author_url": "",
              "post_date": "2026-02-27T19:34:14.930000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3414862,
              "author_name": "/|\\^||^\\|/",
              "author_url": "",
              "post_date": "2026-02-27T20:15:10.727000",
              "content": "<p>Isn't the deadline for <em>starting</em> a submission 11:59pm UTC? Some submissions will run for hours. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3414863,
              "author_name": "/|\\^||^\\|/",
              "author_url": "",
              "post_date": "2026-02-27T20:16:13.960000",
              "content": "<p>I'm new to kaggle. Isn't that time the deadline for <em>starting</em> a submission? A run could take hours. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3414532,
          "author_name": "/|\\^||^\\|/",
          "author_url": "",
          "post_date": "2026-02-27T04:07:36.240000",
          "content": "<p>Just use Hausdorff properly, you'll implicitly get everything you're looking for. You should fire this up as a sanity-check against the current metric. <a href=\"https://github.com/cnr-isti-vclab/vcglib/tree/main/apps/metro\" target=\"_blank\">https://github.com/cnr-isti-vclab/vcglib/tree/main/apps/metro</a></p>\n<p>It's a standard metric for the application you have here. Have a read: <a href=\"https://www.researchgate.net/publication/227611629_METRO_Measuring_error_on_simplified_surfaces\" target=\"_blank\">https://www.researchgate.net/publication/227611629_METRO_Measuring_error_on_simplified_surfaces</a></p>\n<p>It would be illuminating to run metro on the LB. </p>\n<p>edit - you could round metro out with running Betti numbers on a mesh, that's decent. But once you leave mesh-land you lose your simplicial homology and everything is a disaster approximation. Convert everything to meshes for coherent metrics. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3414195,
      "author_name": "GG Ayo (AyoGG)",
      "author_url": "",
      "post_date": "2026-02-26T08:14:49.667000",
      "content": "<p>Whenever your prediction hits an ignore mask, your score will plummet—and it's not your fault.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3414523,
          "author_name": "",
          "author_url": "",
          "post_date": "2026-02-27T02:50:48.377000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3414850,
      "author_name": "Muhammad Ayaz",
      "author_url": "",
      "post_date": "2026-02-27T19:30:38.883000",
      "content": "<p>That’s a very interesting point. If topology improves but LB drops, it suggests the metric may be rewarding something slightly different than what we expect. Would be great to hear whether others have observed this trade-off between DIC optimization and leaderboard performance.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3413997": "I recently developed a new model that achieves better CV locally with fewer holes and adhesions. However, after submission, its score fell significantly below other models with poorer CV. After each submission, I examine the model's prediction for a provided test sample. While a single sample lacks representativeness, it still offers some insight into the model's predictive capability. The output from this test sample also reveals that the model possesses a superior topological structure. Yet, its score is 0.557, even lower than the best public notebook. Below are the predictions from this model and the public notebook.\n\nPublic Notebook LB 0.559\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19164179%2F0bfab6d3d687dfcdea127b1eefcd1a42%2FScreenShot_2026-02-25_171949_505.png?generation=1772040934339296&alt=media)\n\nMy Model LB 0.557\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19164179%2Fd8368e576c918cfcdcefc48befcdd558%2FScreenShot_2026-02-25_173612_063.png?generation=1772041002339142&alt=media)\n\nAchieving a better topological structure has always been my goal in model optimization. Historically, improved topology consistently yielded higher scores. Yet this time, the model trained under the new framework with superior topology resulted in a dramatic score drop. Frankly, this has left me somewhat perplexed. Perhaps the performance on the training data doesn't fully align with private data. Could optimizing for a better DIC be the true objective?\n\nHas anyone encountered a situation where the topology improved but the LB worsened?",
    "3414592": "Why have a few assassins suddenly appeared these last few days?",
    "3414121": "My model performs better in local cv and significantly outperforms other models on the host test samples, yet its score drops by 10 points on the LB. This issue has been troubling me for nearly a month and persists even after updating the metric data.",
    "3414010": "Test data has been updated recently to remove holes in ground truth. I don't know if it explains what you are seeing.\n\nI can't comment on better topology effect as I only started submitting yesterday, I don't have history at all. Let's hope others will chime in to answer that.",
    "3414245": "There are two chances for final submission. You may now begin placing your bets.\n\nThe ignore mask is way too random.",
    "3414046": "Typically reducing merges and reducing holes should produce a higher LB score, because the LB metrics toposcore penalizes even minor errors quite heavily. It is however only a portion of the total LB score. \n\nBecause surface dice in particular allows some wiggle room, you can often cheat a little towards maximizing topo vs surface dice and end up not taking a very large hit on the surface dice score, but this has some limits. \n\nThis very difficult combination of metrics is extremely annoying, trust me when i say i have a ton of sympathy for the participants in this challenge as i too have been fighting these kinds of metrics on this data too, but this combination was chosen because each metric individually doesn't result in readability. \n\n- A perfect toposcore gives the same number of components, and outputs with no holes or voids, but it does not _localize_ the sheet , which means a model with a perfect toposcore but bad surface dice produces meshes which are too far from the written surface to run accurate ink detection on\n\n- On the other hand, a model with good surface dice but bad topo metrics mean we may have good sheet localization, but we cannot mesh it and flatten it without complex algorithms for dealing with noisy surfaces. \n\nOur goal is to produce surface predictions which are well localised and topologically accurate (sane). A \"perfect\" model output of surface predictions ran on an entire scroll would be trivially \"meshable\"/\"flattenable\" with even the most simple algorithms. Even something like active contour models could resolve it alone.  Of course we will never have perfect predictions, but you could see how this interplay between how complex our unrolling software must be and how good the surface predictions are affects the metrics we have chosen, however imperfect they may be -- and they are imperfect, they are however the best we have been able to come up with for this complex task so far. ",
    "3414195": "Whenever your prediction hits an ignore mask, your score will plummet—and it's not your fault.",
    "3414850": "That’s a very interesting point. If topology improves but LB drops, it suggests the metric may be rewarding something slightly different than what we expect. Would be great to hear whether others have observed this trade-off between DIC optimization and leaderboard performance."
  }
}