{
  "id": 672996,
  "title": "Request to host - Transparency on new leaderboard scoring",
  "url": "/competitions/vesuvius-challenge-surface-detection/discussion/672996",
  "author_name": "Nischay Dhankhar",
  "post_date": "2026-02-11T19:15:20.437000",
  "votes": 37,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hi <a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> ,</p>\n<p>First of all, thank you for addressing the issue with Topology score and committing to rescoring all submissions. We appreciate the transparency around the process and the deadline extension to work on the problem.</p>\n<p>That said, many of us are currently facing a critical challenge:\n<strong>Without understanding what exactly was modified in the test set / evaluation, it becomes very difficult to adapt our local validation pipelines in a principled way.</strong></p>\n<p>To help the community react fairly, could you please clarify if the fix applied to:\n(A) Metric computation only?\n(B) Test labels only?\n(C) Both?</p>\n<p>If micro holes were filled in the test labels but not in train, we now may have a structural distribution mismatch. This shifts the task from pure 3D segmentation towards reverse engineering preproceessing / post processing pipelines. </p>\n<p>Since the rescoring has now completed, we observe that some teams gained ~1–1.5% while others saw negligible change. This further suggests that the update may have affected certain prediction characteristics more than others or maybe some teams already figured what the changes were. </p>\n<p>Maybe more clarity on the nature of the fix would allow us to focus on improving true generalization rather than guessing hidden changes. </p>",
  "messages": [
    {
      "id": 3404994,
      "postDate": "2026-02-11T19:15:20.437Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> ,</p>\n<p>First of all, thank you for addressing the issue with Topology score and committing to rescoring all submissions. We appreciate the transparency around the process and the deadline extension to work on the problem.</p>\n<p>That said, many of us are currently facing a critical challenge:\n<strong>Without understanding what exactly was modified in the test set / evaluation, it becomes very difficult to adapt our local validation pipelines in a principled way.</strong></p>\n<p>To help the community react fairly, could you please clarify if the fix applied to:\n(A) Metric computation only?\n(B) Test labels only?\n(C) Both?</p>\n<p>If micro holes were filled in the test labels but not in train, we now may have a structural distribution mismatch. This shifts the task from pure 3D segmentation towards reverse engineering preproceessing / post processing pipelines. </p>\n<p>Since the rescoring has now completed, we observe that some teams gained ~1–1.5% while others saw negligible change. This further suggests that the update may have affected certain prediction characteristics more than others or maybe some teams already figured what the changes were. </p>\n<p>Maybe more clarity on the nature of the fix would allow us to focus on improving true generalization rather than guessing hidden changes. </p>",
      "rawMarkdown": "Hi @giorgioangelotti ,\n\nFirst of all, thank you for addressing the issue with Topology score and committing to rescoring all submissions. We appreciate the transparency around the process and the deadline extension to work on the problem.\n\nThat said, many of us are currently facing a critical challenge:\n**Without understanding what exactly was modified in the test set / evaluation, it becomes very difficult to adapt our local validation pipelines in a principled way.**\n\nTo help the community react fairly, could you please clarify if the fix applied to:\n(A) Metric computation only?\n(B) Test labels only?\n(C) Both?\n\nIf micro holes were filled in the test labels but not in train, we now may have a structural distribution mismatch. This shifts the task from pure 3D segmentation towards reverse engineering preproceessing / post processing pipelines. \n\nSince the rescoring has now completed, we observe that some teams gained ~1–1.5% while others saw negligible change. This further suggests that the update may have affected certain prediction characteristics more than others or maybe some teams already figured what the changes were. \n\nMaybe more clarity on the nature of the fix would allow us to focus on improving true generalization rather than guessing hidden changes. ",
      "votes": 37
    },
    {
      "id": 3405125,
      "postDate": "2026-02-12T07:01:42.557Z",
      "content": "<p>No changes were done to the metrics suite, only the known critical issues in the test set were fixed.\nThe fix was applied with the prescription that the number of changed voxels should be minimal to prevent distribution shifts. Nevertheless, some of the fixed issues involved a human in the loop.</p>\n<p>This was the procedure applied by the team:</p>\n<ol>\n<li>Automated detection of regions of interest that are the source of an anomalous Betti-numbers\n2a. For small cavities, automated hole filling\n2b. For tunnels, automated filling if not disrupting, otherwise, semi-manual filling. The second involved first closing the tunnel, and then automatically filling it.\n2c. Other sources of anomalies were manually fixed.</li>\n</ol>\n<p>The great majority of the issues could be automatically fixed but there were still some that required some manual effort.</p>\n<p>We haven't shared a training set fix because, reading the conversations, many teams had already spotted some issues in the data and produced their own independent preprocessing steps to address some (or all) of them. While we understand that a big fraction of the community would like to have a completely fixed dataset (both train and test), sharing a complete dataset fix so late in the competition would invalidate the work of those teams who spent time on this part of the pipe.</p>\n<p>However, for purely research purposes, we aim sharing a complete dataset after the competition, to promote further developments in this field. As many of you could see, having a 3D vision network (working on voxel space) outputting simple \"lines\" is no easy task.</p>",
      "rawMarkdown": "No changes were done to the metrics suite, only the known critical issues in the test set were fixed.\nThe fix was applied with the prescription that the number of changed voxels should be minimal to prevent distribution shifts. Nevertheless, some of the fixed issues involved a human in the loop.\n\nThis was the procedure applied by the team:\n1. Automated detection of regions of interest that are the source of an anomalous Betti-numbers\n2a. For small cavities, automated hole filling\n2b. For tunnels, automated filling if not disrupting, otherwise, semi-manual filling. The second involved first closing the tunnel, and then automatically filling it.\n2c. Other sources of anomalies were manually fixed.\n\nThe great majority of the issues could be automatically fixed but there were still some that required some manual effort.\n\nWe haven't shared a training set fix because, reading the conversations, many teams had already spotted some issues in the data and produced their own independent preprocessing steps to address some (or all) of them. While we understand that a big fraction of the community would like to have a completely fixed dataset (both train and test), sharing a complete dataset fix so late in the competition would invalidate the work of those teams who spent time on this part of the pipe.\n\nHowever, for purely research purposes, we aim sharing a complete dataset after the competition, to promote further developments in this field. As many of you could see, having a 3D vision network (working on voxel space) outputting simple \"lines\" is no easy task.",
      "votes": 18,
      "isPinned": true,
      "replies": [
        {
          "id": 3405127,
          "postDate": "2026-02-12T07:09:13.917Z",
          "content": "<p>It's clear and fair, thanks. Greatly Respect 💪 <a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> </p>",
          "rawMarkdown": "It's clear and fair, thanks. Greatly Respect 💪 @giorgioangelotti ",
          "votes": 1
        },
        {
          "id": 3405132,
          "postDate": "2026-02-12T07:23:56.143Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> </p>",
          "rawMarkdown": "Thanks @giorgioangelotti ",
          "votes": 1
        }
      ]
    },
    {
      "id": 3405011,
      "postDate": "2026-02-11T20:12:18.863Z",
      "content": "<p>Agreed, we have guesses for what we think changed in the metric but it not being explicitly stated or just updated within the existing metric gives me a lot more questions. We are now going to need to work for 2 more weeks, so I think clarity is only fair</p>",
      "rawMarkdown": "Agreed, we have guesses for what we think changed in the metric but it not being explicitly stated or just updated within the existing metric gives me a lot more questions. We are now going to need to work for 2 more weeks, so I think clarity is only fair",
      "votes": 7,
      "replies": [
        {
          "id": 3405110,
          "postDate": "2026-02-12T05:54:15.597Z",
          "content": "<p>Was the metric changes? Wasn't just the test set updated</p>",
          "rawMarkdown": "Was the metric changes? Wasn't just the test set updated",
          "votes": 2,
          "replies": [
            {
              "id": 3405112,
              "postDate": "2026-02-12T05:58:02.113Z",
              "content": "<p>I thought it was a metric change but at this point I don’t even know! I don’t really understand why it hasn’t been made clear yet because at this point it’s been days since the change was announced. Clearly one team has figured it out haha but I would imagine since this is all about getting the best answers it would be in the best interest of the competition hosts/kaggle admins to clarify it</p>",
              "rawMarkdown": "I thought it was a metric change but at this point I don’t even know! I don’t really understand why it hasn’t been made clear yet because at this point it’s been days since the change was announced. Clearly one team has figured it out haha but I would imagine since this is all about getting the best answers it would be in the best interest of the competition hosts/kaggle admins to clarify it"
            },
            {
              "id": 3405262,
              "postDate": "2026-02-12T15:04:13.887Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 3406884,
      "postDate": "2026-02-16T22:59:43.957Z",
      "content": "<p>Thanks for addressing the Topology score issue and committing to a full rescoring. The transparency and deadline extension are appreciated.</p>\n<p>One open point is how participants should adapt local validation. Without clarity on what exactly changed, it’s hard to know whether improvements come from better generalization or from alignment with the updated evaluation.</p>\n<p>Could you clarify whether the fix affected (A) metric computation, (B) test labels, or (C) both? Some teams saw noticeable gains after rescoring while others saw little change, which suggests certain prediction characteristics may now be favored. More detail on the nature of the fix would help focus efforts on real modeling improvements rather than guessing evaluation changes.</p>",
      "rawMarkdown": "Thanks for addressing the Topology score issue and committing to a full rescoring. The transparency and deadline extension are appreciated.\n\nOne open point is how participants should adapt local validation. Without clarity on what exactly changed, it’s hard to know whether improvements come from better generalization or from alignment with the updated evaluation.\n\nCould you clarify whether the fix affected (A) metric computation, (B) test labels, or (C) both? Some teams saw noticeable gains after rescoring while others saw little change, which suggests certain prediction characteristics may now be favored. More detail on the nature of the fix would help focus efforts on real modeling improvements rather than guessing evaluation changes.",
      "votes": -5,
      "replies": [
        {
          "id": 3407187,
          "postDate": "2026-02-17T19:45:49.350Z",
          "content": "<p>It was clearly stated by host that only test labels were modified.</p>",
          "rawMarkdown": "It was clearly stated by host that only test labels were modified.",
          "votes": 2
        }
      ]
    },
    {
      "id": 3405490,
      "postDate": "2026-02-13T07:41:12.910Z",
      "content": "<p>Teams gained ~1–1.5% while others saw negligible change? <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> </p>",
      "rawMarkdown": "Teams gained ~1–1.5% while others saw negligible change? @nischaydnk "
    }
  ],
  "comments": [
    {
      "id": 3405125,
      "author_name": "Giorgio Angelotti",
      "author_url": "",
      "post_date": "2026-02-12T07:01:42.557000",
      "content": "<p>No changes were done to the metrics suite, only the known critical issues in the test set were fixed.\nThe fix was applied with the prescription that the number of changed voxels should be minimal to prevent distribution shifts. Nevertheless, some of the fixed issues involved a human in the loop.</p>\n<p>This was the procedure applied by the team:</p>\n<ol>\n<li>Automated detection of regions of interest that are the source of an anomalous Betti-numbers\n2a. For small cavities, automated hole filling\n2b. For tunnels, automated filling if not disrupting, otherwise, semi-manual filling. The second involved first closing the tunnel, and then automatically filling it.\n2c. Other sources of anomalies were manually fixed.</li>\n</ol>\n<p>The great majority of the issues could be automatically fixed but there were still some that required some manual effort.</p>\n<p>We haven't shared a training set fix because, reading the conversations, many teams had already spotted some issues in the data and produced their own independent preprocessing steps to address some (or all) of them. While we understand that a big fraction of the community would like to have a completely fixed dataset (both train and test), sharing a complete dataset fix so late in the competition would invalidate the work of those teams who spent time on this part of the pipe.</p>\n<p>However, for purely research purposes, we aim sharing a complete dataset after the competition, to promote further developments in this field. As many of you could see, having a 3D vision network (working on voxel space) outputting simple \"lines\" is no easy task.</p>",
      "votes": 18,
      "replies": [
        {
          "id": 3405127,
          "author_name": "ForcewithMe",
          "author_url": "",
          "post_date": "2026-02-12T07:09:13.917000",
          "content": "<p>It's clear and fair, thanks. Greatly Respect 💪 <a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 3405132,
          "author_name": "Tom",
          "author_url": "",
          "post_date": "2026-02-12T07:23:56.143000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3405011,
      "author_name": "Cody_Null",
      "author_url": "",
      "post_date": "2026-02-11T20:12:18.863000",
      "content": "<p>Agreed, we have guesses for what we think changed in the metric but it not being explicitly stated or just updated within the existing metric gives me a lot more questions. We are now going to need to work for 2 more weeks, so I think clarity is only fair</p>",
      "votes": 7,
      "replies": [
        {
          "id": 3405110,
          "author_name": "Manas Choudhary",
          "author_url": "",
          "post_date": "2026-02-12T05:54:15.597000",
          "content": "<p>Was the metric changes? Wasn't just the test set updated</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3405112,
              "author_name": "Cody_Null",
              "author_url": "",
              "post_date": "2026-02-12T05:58:02.113000",
              "content": "<p>I thought it was a metric change but at this point I don’t even know! I don’t really understand why it hasn’t been made clear yet because at this point it’s been days since the change was announced. Clearly one team has figured it out haha but I would imagine since this is all about getting the best answers it would be in the best interest of the competition hosts/kaggle admins to clarify it</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3405262,
              "author_name": "",
              "author_url": "",
              "post_date": "2026-02-12T15:04:13.887000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3406884,
      "author_name": "Muhammad Ayaz",
      "author_url": "",
      "post_date": "2026-02-16T22:59:43.957000",
      "content": "<p>Thanks for addressing the Topology score issue and committing to a full rescoring. The transparency and deadline extension are appreciated.</p>\n<p>One open point is how participants should adapt local validation. Without clarity on what exactly changed, it’s hard to know whether improvements come from better generalization or from alignment with the updated evaluation.</p>\n<p>Could you clarify whether the fix affected (A) metric computation, (B) test labels, or (C) both? Some teams saw noticeable gains after rescoring while others saw little change, which suggests certain prediction characteristics may now be favored. More detail on the nature of the fix would help focus efforts on real modeling improvements rather than guessing evaluation changes.</p>",
      "votes": -5,
      "replies": [
        {
          "id": 3407187,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2026-02-17T19:45:49.350000",
          "content": "<p>It was clearly stated by host that only test labels were modified.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 3405490,
      "author_name": "Navneet",
      "author_url": "",
      "post_date": "2026-02-13T07:41:12.910000",
      "content": "<p>Teams gained ~1–1.5% while others saw negligible change? <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3404994": "Hi @giorgioangelotti ,\n\nFirst of all, thank you for addressing the issue with Topology score and committing to rescoring all submissions. We appreciate the transparency around the process and the deadline extension to work on the problem.\n\nThat said, many of us are currently facing a critical challenge:\n**Without understanding what exactly was modified in the test set / evaluation, it becomes very difficult to adapt our local validation pipelines in a principled way.**\n\nTo help the community react fairly, could you please clarify if the fix applied to:\n(A) Metric computation only?\n(B) Test labels only?\n(C) Both?\n\nIf micro holes were filled in the test labels but not in train, we now may have a structural distribution mismatch. This shifts the task from pure 3D segmentation towards reverse engineering preproceessing / post processing pipelines. \n\nSince the rescoring has now completed, we observe that some teams gained ~1–1.5% while others saw negligible change. This further suggests that the update may have affected certain prediction characteristics more than others or maybe some teams already figured what the changes were. \n\nMaybe more clarity on the nature of the fix would allow us to focus on improving true generalization rather than guessing hidden changes. ",
    "3405125": "No changes were done to the metrics suite, only the known critical issues in the test set were fixed.\nThe fix was applied with the prescription that the number of changed voxels should be minimal to prevent distribution shifts. Nevertheless, some of the fixed issues involved a human in the loop.\n\nThis was the procedure applied by the team:\n1. Automated detection of regions of interest that are the source of an anomalous Betti-numbers\n2a. For small cavities, automated hole filling\n2b. For tunnels, automated filling if not disrupting, otherwise, semi-manual filling. The second involved first closing the tunnel, and then automatically filling it.\n2c. Other sources of anomalies were manually fixed.\n\nThe great majority of the issues could be automatically fixed but there were still some that required some manual effort.\n\nWe haven't shared a training set fix because, reading the conversations, many teams had already spotted some issues in the data and produced their own independent preprocessing steps to address some (or all) of them. While we understand that a big fraction of the community would like to have a completely fixed dataset (both train and test), sharing a complete dataset fix so late in the competition would invalidate the work of those teams who spent time on this part of the pipe.\n\nHowever, for purely research purposes, we aim sharing a complete dataset after the competition, to promote further developments in this field. As many of you could see, having a 3D vision network (working on voxel space) outputting simple \"lines\" is no easy task.",
    "3405011": "Agreed, we have guesses for what we think changed in the metric but it not being explicitly stated or just updated within the existing metric gives me a lot more questions. We are now going to need to work for 2 more weeks, so I think clarity is only fair",
    "3406884": "Thanks for addressing the Topology score issue and committing to a full rescoring. The transparency and deadline extension are appreciated.\n\nOne open point is how participants should adapt local validation. Without clarity on what exactly changed, it’s hard to know whether improvements come from better generalization or from alignment with the updated evaluation.\n\nCould you clarify whether the fix affected (A) metric computation, (B) test labels, or (C) both? Some teams saw noticeable gains after rescoring while others saw little change, which suggests certain prediction characteristics may now be favored. More detail on the nature of the fix would help focus efforts on real modeling improvements rather than guessing evaluation changes.",
    "3405490": "Teams gained ~1–1.5% while others saw negligible change? @nischaydnk "
  }
}