{
  "id": 674052,
  "title": "VOI Metric - Critical Issue?",
  "url": "/competitions/vesuvius-challenge-surface-detection/discussion/674052",
  "author_name": "",
  "post_date": "2026-02-18T12:26:33.448299500Z",
  "votes": 8,
  "comment_count": 11,
  "views": 0,
  "content": "<h1>Brief</h1>\n<p>I check the metrics code, and I believe I may have stumbled upon another issue, this time with the VOI metric.  </p>\n<h1>Demonstration</h1>\n<p>I have prepared a notebook that demonstrates the issue.  Kaggle doesn't let me share it this late in the competition, but I can show it to the organizers.</p>\n<p>Below are some screenshots from the notebook.  I fed the score_single_tif function three manufactured predictions, for volume 327851248.</p>\n<ul>\n<li><p>\"Exact\" means feeding the ground truth as the prediction (with \"ignore\" replaced by \"background\")</p></li>\n<li><p>\"Shifted 2 Voxels\" means taking the original image, and simply shifting it 2 voxels on each axis.  It gets a perfect surface dice, but a low VOI score (0.527).</p></li>\n<li><p>\"Random Blocks\" means generating a random image from 64x64x64 blocks, with the same size as the ground truth. This does much better, in terms of the VOI score, than the shifted version (it gets 0.774).  </p></li>\n</ul>\n<p>The images of the three predictions are also provided below.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F907884%2Fc7f30b6413c2fcee2c556db9fa9a27ef%2FScreenshot%202026-02-18%20at%2013.57.50.png?generation=1771417399802989&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F907884%2F519fda716589637952146468dac5e4fd%2FScreenshot%202026-02-18%20at%2014.20.33.png?generation=1771417414040343&amp;alt=media\" alt=\"\"></p>\n<h1>Analysis:</h1>\n<p>Having investigated the issue, I have a good idea what probably caused it, and a quick fix.  </p>\n<p>The problem appears to arise from the background class.  Looking at the function compute_voi_metrics, we see the line:</p>\n<pre><code>m = (gt_lab &gt; 0) | (pr_lab &gt; 0)\n</code></pre>\n<p>This means that the variation_of_information receives a mix that includes the background class.  Continuing downstream, we have, in the function _vi_tables of skimage.metrics._variation_of_information, the instructions:</p>\n<pre><code>hygx = -px @ _xlogx(px_inv @ pxy).sum(axis=1)\nhxgy = -_xlogx(pxy @ py_inv).sum(axis=0) @ py\n</code></pre>\n<p>Examining the terms of the sum, we can see that the zero (background) class dominates the result.</p>\n<p>A simple fix would be to simple replace the mask with:</p>\n<pre><code>m = (gt_lab &gt; 0) &amp; (pr_lab &gt; 0)\n</code></pre>\n<h1>Finally</h1>\n<p>I have a nice approach that I am excited about, but currently does not deliver a good VOI metric.  If the problem I described is real, I would appreciate if the organizers (@seanjohnsonsp, <a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a>, <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>) could address it.</p>",
  "messages": [
    {
      "id": "3407477",
      "postDate": "02/18/2026 12:26:33",
      "content": "<h1>Brief</h1>\n<p>I check the metrics code, and I believe I may have stumbled upon another issue, this time with the VOI metric.  </p>\n<h1>Demonstration</h1>\n<p>I have prepared a notebook that demonstrates the issue.  Kaggle doesn't let me share it this late in the competition, but I can show it to the organizers.</p>\n<p>Below are some screenshots from the notebook.  I fed the score_single_tif function three manufactured predictions, for volume 327851248.</p>\n<ul>\n<li><p>\"Exact\" means feeding the ground truth as the prediction (with \"ignore\" replaced by \"background\")</p></li>\n<li><p>\"Shifted 2 Voxels\" means taking the original image, and simply shifting it 2 voxels on each axis.  It gets a perfect surface dice, but a low VOI score (0.527).</p></li>\n<li><p>\"Random Blocks\" means generating a random image from 64x64x64 blocks, with the same size as the ground truth. This does much better, in terms of the VOI score, than the shifted version (it gets 0.774).  </p></li>\n</ul>\n<p>The images of the three predictions are also provided below.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F907884%2Fc7f30b6413c2fcee2c556db9fa9a27ef%2FScreenshot%202026-02-18%20at%2013.57.50.png?generation=1771417399802989&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F907884%2F519fda716589637952146468dac5e4fd%2FScreenshot%202026-02-18%20at%2014.20.33.png?generation=1771417414040343&amp;alt=media\" alt=\"\"></p>\n<h1>Analysis:</h1>\n<p>Having investigated the issue, I have a good idea what probably caused it, and a quick fix.  </p>\n<p>The problem appears to arise from the background class.  Looking at the function compute_voi_metrics, we see the line:</p>\n<pre><code>m = (gt_lab &gt; 0) | (pr_lab &gt; 0)\n</code></pre>\n<p>This means that the variation_of_information receives a mix that includes the background class.  Continuing downstream, we have, in the function _vi_tables of skimage.metrics._variation_of_information, the instructions:</p>\n<pre><code>hygx = -px @ _xlogx(px_inv @ pxy).sum(axis=1)\nhxgy = -_xlogx(pxy @ py_inv).sum(axis=0) @ py\n</code></pre>\n<p>Examining the terms of the sum, we can see that the zero (background) class dominates the result.</p>\n<p>A simple fix would be to simple replace the mask with:</p>\n<pre><code>m = (gt_lab &gt; 0) &amp; (pr_lab &gt; 0)\n</code></pre>\n<h1>Finally</h1>\n<p>I have a nice approach that I am excited about, but currently does not deliver a good VOI metric.  If the problem I described is real, I would appreciate if the organizers (@seanjohnsonsp, <a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a>, <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>) could address it.</p>",
      "rawMarkdown": "# Brief\nI check the metrics code, and I believe I may have stumbled upon another issue, this time with the VOI metric.  \n\n# Demonstration\nI have prepared a notebook that demonstrates the issue.  Kaggle doesn't let me share it this late in the competition, but I can show it to the organizers.\n\nBelow are some screenshots from the notebook.  I fed the score_single_tif function three manufactured predictions, for volume 327851248.\n\n- \"Exact\" means feeding the ground truth as the prediction (with \"ignore\" replaced by \"background\")\n\n- \"Shifted 2 Voxels\" means taking the original image, and simply shifting it 2 voxels on each axis.  It gets a perfect surface dice, but a low VOI score (0.527).\n\n- \"Random Blocks\" means generating a random image from 64x64x64 blocks, with the same size as the ground truth. This does much better, in terms of the VOI score, than the shifted version (it gets 0.774).  \n\nThe images of the three predictions are also provided below.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F907884%2Fc7f30b6413c2fcee2c556db9fa9a27ef%2FScreenshot%202026-02-18%20at%2013.57.50.png?generation=1771417399802989&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F907884%2F519fda716589637952146468dac5e4fd%2FScreenshot%202026-02-18%20at%2014.20.33.png?generation=1771417414040343&alt=media)\n\n# Analysis:\nHaving investigated the issue, I have a good idea what probably caused it, and a quick fix.  \n\nThe problem appears to arise from the background class.  Looking at the function compute_voi_metrics, we see the line:\n\n    m = (gt_lab > 0) | (pr_lab > 0)\n\nThis means that the variation_of_information receives a mix that includes the background class.  Continuing downstream, we have, in the function _vi_tables of skimage.metrics._variation_of_information, the instructions:\n\n    hygx = -px @ _xlogx(px_inv @ pxy).sum(axis=1)\n    hxgy = -_xlogx(pxy @ py_inv).sum(axis=0) @ py\n\nExamining the terms of the sum, we can see that the zero (background) class dominates the result.\n\nA simple fix would be to simple replace the mask with:\n\n    m = (gt_lab > 0) & (pr_lab > 0)\n\n# Finally\nI have a nice approach that I am excited about, but currently does not deliver a good VOI metric.  If the problem I described is real, I would appreciate if the organizers (@seanjohnsonsp, @giorgioangelotti, @sohier) could address it.",
      "votes": null
    },
    {
      "id": "3407530",
      "postDate": "02/18/2026 14:20:00",
      "content": "<p>I don't think this is an \"critical\" issue because you drastically lowered the score of other two metrics. Here's another example: consider you are using a metric like dice + no of components, the if sample contains a lot of component then you can predict a giant blob which will give you a good dice score but bad component score. </p>\n<p>Similarly, if your solution is good then it should increase all the metric scores simultaneously. If I am wrong would you mind sharing an example instance of you prediction where your surface dice and topo score are very high but your voi score is low?</p>",
      "rawMarkdown": "I don't think this is an \"critical\" issue because you drastically lowered the score of other two metrics. Here's another example: consider you are using a metric like dice + no of components, the if sample contains a lot of component then you can predict a giant blob which will give you a good dice score but bad component score. \n\n\nSimilarly, if your solution is good then it should increase all the metric scores simultaneously. If I am wrong would you mind sharing an example instance of you prediction where your surface dice and topo score are very high but your voi score is low?",
      "votes": null
    },
    {
      "id": "3407535",
      "postDate": "02/18/2026 14:35:41",
      "content": "<p>That and if this competition changes or extends again and we slide on LB again I might lose my mind 😂</p>",
      "rawMarkdown": "That and if this competition changes or extends again and we slide on LB again I might lose my mind 😂",
      "votes": null
    },
    {
      "id": "3407544",
      "postDate": "02/18/2026 14:59:52",
      "content": "<p>Hello,</p>\n<p>Isn't this exactly how the VOI should work?</p>\n<p>VOI (not the VOI score) should be H(ground truth) + H(predictions) - 2 I, with I being the mutual information</p>\n<p>If your shift kills the mutual information term, the VOI for the shifted case is approximately 2*H(ground truth), since H(predictions) should be similar to H(ground truth)</p>\n<p>For the random blocks case, I expect the H(predictions) to be very low, therefore the VOI should be approximately H(ground truth), which is lower than the one for the shifted term.</p>\n<p>And hence VOI_score will be counterintuively higher for the random blocks case.</p>\n<p>However, the metric include also the Surface Dice and the TopoScore, which make the final LB reasonable.</p>",
      "rawMarkdown": "Hello,\n\nIsn't this exactly how the VOI should work?\n\nVOI (not the VOI score) should be H(ground truth) + H(predictions) - 2 I, with I being the mutual information\n\nIf your shift kills the mutual information term, the VOI for the shifted case is approximately 2*H(ground truth), since H(predictions) should be similar to H(ground truth)\n\nFor the random blocks case, I expect the H(predictions) to be very low, therefore the VOI should be approximately H(ground truth), which is lower than the one for the shifted term.\n\nAnd hence VOI_score will be counterintuively higher for the random blocks case.\n\nHowever, the metric include also the Surface Dice and the TopoScore, which make the final LB reasonable.",
      "votes": null
    },
    {
      "id": "3407548",
      "postDate": "02/18/2026 15:21:42",
      "content": "<p>Attached is a better example.  It is from volume 40625686.  Notice that the 2-voxel-shift gets perfect dice and nearly perfect toposcore.   It should also get a perfect VOI score, but it gets 0.577.</p>\n<p>What's going on?  The surface dice is designed to be agnostic to small shifts, and indeed it is.   We want this, because we are focusing on the grand topology, not individual voxels.  The VOI suffers under such small shifts.</p>\n<p>The conclusion is that the VOI is not truly meaningful.  </p>\n<p>My algorithm scores similarly on all three metrics for this volume.  Removing the background component in the scoring function makes my VOI score shoot up.  This is also evident when visualizing the prediction in 3D - it is nearly perfect.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F907884%2F896c133d467fc7b7e3a66f859d70f9df%2FScreenshot%202026-02-18%20at%2017.07.53.png?generation=1771427462081185&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Attached is a better example.  It is from volume 40625686.  Notice that the 2-voxel-shift gets perfect dice and nearly perfect toposcore.   It should also get a perfect VOI score, but it gets 0.577.\n\nWhat's going on?  The surface dice is designed to be agnostic to small shifts, and indeed it is.   We want this, because we are focusing on the grand topology, not individual voxels.  The VOI suffers under such small shifts.\n\nThe conclusion is that the VOI is not truly meaningful.  \n\nMy algorithm scores similarly on all three metrics for this volume.  Removing the background component in the scoring function makes my VOI score shoot up.  This is also evident when visualizing the prediction in 3D - it is nearly perfect.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F907884%2F896c133d467fc7b7e3a66f859d70f9df%2FScreenshot%202026-02-18%20at%2017.07.53.png?generation=1771427462081185&alt=media)",
      "votes": null
    },
    {
      "id": "3407554",
      "postDate": "02/18/2026 15:34:20",
      "content": "<p>Let's review the motivation behind the VOI expression.</p>\n<p>Did we want a 2 voxel shift to devastate the VOI score?  Does it indicate a greater number of splits or merges?      The surface dice was designed to be agnostic to such shifts, because they were deemed insignificant.</p>\n<p>Including the background in the expression means that we treat it as a <em>fragment</em>.   It dominates, because it is by far the largest.  A small conditional entropy (an intersection with other fragments) has an outsized impact on the VOI score.</p>\n<p>My algorithm's prediction, for this particular volume, is nearly perfect, but misses at occasional voxels around the labels.  </p>",
      "rawMarkdown": "Let's review the motivation behind the VOI expression.\n\nDid we want a 2 voxel shift to devastate the VOI score?  Does it indicate a greater number of splits or merges?      The surface dice was designed to be agnostic to such shifts, because they were deemed insignificant.\n\nIncluding the background in the expression means that we treat it as a *fragment*.   It dominates, because it is by far the largest.  A small conditional entropy (an intersection with other fragments) has an outsized impact on the VOI score.\n\nMy algorithm's prediction, for this particular volume, is nearly perfect, but misses at occasional voxels around the labels.",
      "votes": null
    },
    {
      "id": "3407892",
      "postDate": "02/19/2026 11:36:01",
      "content": "<p>There is no 100% perfect metrics. I think current mixture of metrics is fine. If you want to go into details, you can argue that the labels are not perfect and that would be endless.</p>\n<p>So long that we can differentiate the abilities of different models, and the host can identify the weaknesses of the model, the metric is good enough even though it is not 100% perfect.</p>",
      "rawMarkdown": "There is no 100% perfect metrics. I think current mixture of metrics is fine. If you want to go into details, you can argue that the labels are not perfect and that would be endless.\n\nSo long that we can differentiate the abilities of different models, and the host can identify the weaknesses of the model, the metric is good enough even though it is not 100% perfect.",
      "votes": null
    },
    {
      "id": "3408261",
      "postDate": "02/20/2026 07:28:02",
      "content": "<p>There will never be a perfect metric, the current metrics measure what they want to too a good extend, if there would have had been a perfect metric/loss, there wouldn't be research happening in ai.</p>",
      "rawMarkdown": "There will never be a perfect metric, the current metrics measure what they want to too a good extend, if there would have had been a perfect metric/loss, there wouldn't be research happening in ai.",
      "votes": null
    },
    {
      "id": "3409039",
      "postDate": "02/22/2026 00:15:42",
      "content": "<p>This is a very important observation.\nIf a 2-voxel spatial shift (which preserves structure) performs significantly worse in VOI than random block noise, that suggests the metric may be behaving counterintuitively.\nI would strongly support a deeper review of the VOI implementation, especially given how leaderboard positions can depend heavily on it.</p>",
      "rawMarkdown": "This is a very important observation.\nIf a 2-voxel spatial shift (which preserves structure) performs significantly worse in VOI than random block noise, that suggests the metric may be behaving counterintuitively.\nI would strongly support a deeper review of the VOI implementation, especially given how leaderboard positions can depend heavily on it.",
      "votes": null
    },
    {
      "id": "3410556",
      "postDate": "02/23/2026 09:07:23",
      "content": "<p>Thank you for the analysis <a href=\"https://www.kaggle.com/bennatanamir\" target=\"_blank\">@bennatanamir</a> </p>",
      "rawMarkdown": "Thank you for the analysis @bennatanamir",
      "votes": null
    },
    {
      "id": "3413091",
      "postDate": "02/24/2026 08:34:31",
      "content": "<p>Your ad hoc metric is a can of worms for several reasons. There's a simple standard algorithm for computing the distance between arbitrary meshes that intrinsically has all the information you are trying to measure.  It is trivial to code out the metric, find it here: <a href=\"http://vcglib.net/metro.html\" target=\"_blank\">http://vcglib.net/metro.html</a></p>",
      "rawMarkdown": "Your ad hoc metric is a can of worms for several reasons. There's a simple standard algorithm for computing the distance between arbitrary meshes that intrinsically has all the information you are trying to measure.  It is trivial to code out the metric, find it here: http://vcglib.net/metro.html",
      "votes": null
    },
    {
      "id": "3413101",
      "postDate": "02/24/2026 09:19:59",
      "content": "<p>The Metro algorithm you’re suggesting seems to achieve, in mesh space, something similar to what SurfaceDice already does in voxel space, without an explicit distance threshold.</p>\n<p>I agree that a mesh-to-mesh distance can intrinsically reflect many of the geometric characteristics we care about. However, for our purposes it is preferable to encode those priorities explicitly by enforcing a strict ordering and applying targeted penalties (in particular for topology), rather than relying on those properties to emerge implicitly from a generic distance measure.</p>",
      "rawMarkdown": "The Metro algorithm you’re suggesting seems to achieve, in mesh space, something similar to what SurfaceDice already does in voxel space, without an explicit distance threshold.\n\nI agree that a mesh-to-mesh distance can intrinsically reflect many of the geometric characteristics we care about. However, for our purposes it is preferable to encode those priorities explicitly by enforcing a strict ordering and applying targeted penalties (in particular for topology), rather than relying on those properties to emerge implicitly from a generic distance measure.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3407530,
      "author_name": "iamparadox",
      "author_url": "",
      "post_date": "02/18/2026 14:20:00",
      "content": "<p>I don't think this is an \"critical\" issue because you drastically lowered the score of other two metrics. Here's another example: consider you are using a metric like dice + no of components, the if sample contains a lot of component then you can predict a giant blob which will give you a good dice score but bad component score. </p>\n<p>Similarly, if your solution is good then it should increase all the metric scores simultaneously. If I am wrong would you mind sharing an example instance of you prediction where your surface dice and topo score are very high but your voi score is low?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3407535,
          "author_name": "cody11null",
          "author_url": "",
          "post_date": "02/18/2026 14:35:41",
          "content": "<p>That and if this competition changes or extends again and we slide on LB again I might lose my mind 😂</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3407548,
          "author_name": "bennatanamir",
          "author_url": "",
          "post_date": "02/18/2026 15:21:42",
          "content": "<p>Attached is a better example.  It is from volume 40625686.  Notice that the 2-voxel-shift gets perfect dice and nearly perfect toposcore.   It should also get a perfect VOI score, but it gets 0.577.</p>\n<p>What's going on?  The surface dice is designed to be agnostic to small shifts, and indeed it is.   We want this, because we are focusing on the grand topology, not individual voxels.  The VOI suffers under such small shifts.</p>\n<p>The conclusion is that the VOI is not truly meaningful.  </p>\n<p>My algorithm scores similarly on all three metrics for this volume.  Removing the background component in the scoring function makes my VOI score shoot up.  This is also evident when visualizing the prediction in 3D - it is nearly perfect.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F907884%2F896c133d467fc7b7e3a66f859d70f9df%2FScreenshot%202026-02-18%20at%2017.07.53.png?generation=1771427462081185&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3407544,
      "author_name": "giorgioangelotti",
      "author_url": "",
      "post_date": "02/18/2026 14:59:52",
      "content": "<p>Hello,</p>\n<p>Isn't this exactly how the VOI should work?</p>\n<p>VOI (not the VOI score) should be H(ground truth) + H(predictions) - 2 I, with I being the mutual information</p>\n<p>If your shift kills the mutual information term, the VOI for the shifted case is approximately 2*H(ground truth), since H(predictions) should be similar to H(ground truth)</p>\n<p>For the random blocks case, I expect the H(predictions) to be very low, therefore the VOI should be approximately H(ground truth), which is lower than the one for the shifted term.</p>\n<p>And hence VOI_score will be counterintuively higher for the random blocks case.</p>\n<p>However, the metric include also the Surface Dice and the TopoScore, which make the final LB reasonable.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3407554,
          "author_name": "bennatanamir",
          "author_url": "",
          "post_date": "02/18/2026 15:34:20",
          "content": "<p>Let's review the motivation behind the VOI expression.</p>\n<p>Did we want a 2 voxel shift to devastate the VOI score?  Does it indicate a greater number of splits or merges?      The surface dice was designed to be agnostic to such shifts, because they were deemed insignificant.</p>\n<p>Including the background in the expression means that we treat it as a <em>fragment</em>.   It dominates, because it is by far the largest.  A small conditional entropy (an intersection with other fragments) has an outsized impact on the VOI score.</p>\n<p>My algorithm's prediction, for this particular volume, is nearly perfect, but misses at occasional voxels around the labels.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3413091,
          "author_name": "yugensan",
          "author_url": "",
          "post_date": "02/24/2026 08:34:31",
          "content": "<p>Your ad hoc metric is a can of worms for several reasons. There's a simple standard algorithm for computing the distance between arbitrary meshes that intrinsically has all the information you are trying to measure.  It is trivial to code out the metric, find it here: <a href=\"http://vcglib.net/metro.html\" target=\"_blank\">http://vcglib.net/metro.html</a></p>",
          "votes": null,
          "replies": [
            {
              "id": 3413101,
              "author_name": "giorgioangelotti",
              "author_url": "",
              "post_date": "02/24/2026 09:19:59",
              "content": "<p>The Metro algorithm you’re suggesting seems to achieve, in mesh space, something similar to what SurfaceDice already does in voxel space, without an explicit distance threshold.</p>\n<p>I agree that a mesh-to-mesh distance can intrinsically reflect many of the geometric characteristics we care about. However, for our purposes it is preferable to encode those priorities explicitly by enforcing a strict ordering and applying targeted penalties (in particular for topology), rather than relying on those properties to emerge implicitly from a generic distance measure.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3407892,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "02/19/2026 11:36:01",
      "content": "<p>There is no 100% perfect metrics. I think current mixture of metrics is fine. If you want to go into details, you can argue that the labels are not perfect and that would be endless.</p>\n<p>So long that we can differentiate the abilities of different models, and the host can identify the weaknesses of the model, the metric is good enough even though it is not 100% perfect.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3408261,
      "author_name": "choudharymanas",
      "author_url": "",
      "post_date": "02/20/2026 07:28:02",
      "content": "<p>There will never be a perfect metric, the current metrics measure what they want to too a good extend, if there would have had been a perfect metric/loss, there wouldn't be research happening in ai.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3409039,
      "author_name": "ayazi137",
      "author_url": "",
      "post_date": "02/22/2026 00:15:42",
      "content": "<p>This is a very important observation.\nIf a 2-voxel spatial shift (which preserves structure) performs significantly worse in VOI than random block noise, that suggests the metric may be behaving counterintuitively.\nI would strongly support a deeper review of the VOI implementation, especially given how leaderboard positions can depend heavily on it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3410556,
      "author_name": "navneetbende",
      "author_url": "",
      "post_date": "02/23/2026 09:07:23",
      "content": "<p>Thank you for the analysis <a href=\"https://www.kaggle.com/bennatanamir\" target=\"_blank\">@bennatanamir</a> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3407477": "# Brief\nI check the metrics code, and I believe I may have stumbled upon another issue, this time with the VOI metric.  \n\n# Demonstration\nI have prepared a notebook that demonstrates the issue.  Kaggle doesn't let me share it this late in the competition, but I can show it to the organizers.\n\nBelow are some screenshots from the notebook.  I fed the score_single_tif function three manufactured predictions, for volume 327851248.\n\n- \"Exact\" means feeding the ground truth as the prediction (with \"ignore\" replaced by \"background\")\n\n- \"Shifted 2 Voxels\" means taking the original image, and simply shifting it 2 voxels on each axis.  It gets a perfect surface dice, but a low VOI score (0.527).\n\n- \"Random Blocks\" means generating a random image from 64x64x64 blocks, with the same size as the ground truth. This does much better, in terms of the VOI score, than the shifted version (it gets 0.774).  \n\nThe images of the three predictions are also provided below.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F907884%2Fc7f30b6413c2fcee2c556db9fa9a27ef%2FScreenshot%202026-02-18%20at%2013.57.50.png?generation=1771417399802989&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F907884%2F519fda716589637952146468dac5e4fd%2FScreenshot%202026-02-18%20at%2014.20.33.png?generation=1771417414040343&alt=media)\n\n# Analysis:\nHaving investigated the issue, I have a good idea what probably caused it, and a quick fix.  \n\nThe problem appears to arise from the background class.  Looking at the function compute_voi_metrics, we see the line:\n\n    m = (gt_lab > 0) | (pr_lab > 0)\n\nThis means that the variation_of_information receives a mix that includes the background class.  Continuing downstream, we have, in the function _vi_tables of skimage.metrics._variation_of_information, the instructions:\n\n    hygx = -px @ _xlogx(px_inv @ pxy).sum(axis=1)\n    hxgy = -_xlogx(pxy @ py_inv).sum(axis=0) @ py\n\nExamining the terms of the sum, we can see that the zero (background) class dominates the result.\n\nA simple fix would be to simple replace the mask with:\n\n    m = (gt_lab > 0) & (pr_lab > 0)\n\n# Finally\nI have a nice approach that I am excited about, but currently does not deliver a good VOI metric.  If the problem I described is real, I would appreciate if the organizers (@seanjohnsonsp, @giorgioangelotti, @sohier) could address it.",
    "3407530": "I don't think this is an \"critical\" issue because you drastically lowered the score of other two metrics. Here's another example: consider you are using a metric like dice + no of components, the if sample contains a lot of component then you can predict a giant blob which will give you a good dice score but bad component score. \n\n\nSimilarly, if your solution is good then it should increase all the metric scores simultaneously. If I am wrong would you mind sharing an example instance of you prediction where your surface dice and topo score are very high but your voi score is low?",
    "3407535": "That and if this competition changes or extends again and we slide on LB again I might lose my mind 😂",
    "3407544": "Hello,\n\nIsn't this exactly how the VOI should work?\n\nVOI (not the VOI score) should be H(ground truth) + H(predictions) - 2 I, with I being the mutual information\n\nIf your shift kills the mutual information term, the VOI for the shifted case is approximately 2*H(ground truth), since H(predictions) should be similar to H(ground truth)\n\nFor the random blocks case, I expect the H(predictions) to be very low, therefore the VOI should be approximately H(ground truth), which is lower than the one for the shifted term.\n\nAnd hence VOI_score will be counterintuively higher for the random blocks case.\n\nHowever, the metric include also the Surface Dice and the TopoScore, which make the final LB reasonable.",
    "3407548": "Attached is a better example.  It is from volume 40625686.  Notice that the 2-voxel-shift gets perfect dice and nearly perfect toposcore.   It should also get a perfect VOI score, but it gets 0.577.\n\nWhat's going on?  The surface dice is designed to be agnostic to small shifts, and indeed it is.   We want this, because we are focusing on the grand topology, not individual voxels.  The VOI suffers under such small shifts.\n\nThe conclusion is that the VOI is not truly meaningful.  \n\nMy algorithm scores similarly on all three metrics for this volume.  Removing the background component in the scoring function makes my VOI score shoot up.  This is also evident when visualizing the prediction in 3D - it is nearly perfect.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F907884%2F896c133d467fc7b7e3a66f859d70f9df%2FScreenshot%202026-02-18%20at%2017.07.53.png?generation=1771427462081185&alt=media)",
    "3407554": "Let's review the motivation behind the VOI expression.\n\nDid we want a 2 voxel shift to devastate the VOI score?  Does it indicate a greater number of splits or merges?      The surface dice was designed to be agnostic to such shifts, because they were deemed insignificant.\n\nIncluding the background in the expression means that we treat it as a *fragment*.   It dominates, because it is by far the largest.  A small conditional entropy (an intersection with other fragments) has an outsized impact on the VOI score.\n\nMy algorithm's prediction, for this particular volume, is nearly perfect, but misses at occasional voxels around the labels.",
    "3407892": "There is no 100% perfect metrics. I think current mixture of metrics is fine. If you want to go into details, you can argue that the labels are not perfect and that would be endless.\n\nSo long that we can differentiate the abilities of different models, and the host can identify the weaknesses of the model, the metric is good enough even though it is not 100% perfect.",
    "3408261": "There will never be a perfect metric, the current metrics measure what they want to too a good extend, if there would have had been a perfect metric/loss, there wouldn't be research happening in ai.",
    "3409039": "This is a very important observation.\nIf a 2-voxel spatial shift (which preserves structure) performs significantly worse in VOI than random block noise, that suggests the metric may be behaving counterintuitively.\nI would strongly support a deeper review of the VOI implementation, especially given how leaderboard positions can depend heavily on it.",
    "3410556": "Thank you for the analysis @bennatanamir",
    "3413091": "Your ad hoc metric is a can of worms for several reasons. There's a simple standard algorithm for computing the distance between arbitrary meshes that intrinsically has all the information you are trying to measure.  It is trivial to code out the metric, find it here: http://vcglib.net/metro.html",
    "3413101": "The Metro algorithm you’re suggesting seems to achieve, in mesh space, something similar to what SurfaceDice already does in voxel space, without an explicit distance threshold.\n\nI agree that a mesh-to-mesh distance can intrinsically reflect many of the geometric characteristics we care about. However, for our purposes it is preferable to encode those priorities explicitly by enforcing a strict ordering and applying targeted penalties (in particular for topology), rather than relying on those properties to emerge implicitly from a generic distance measure."
  },
  "source": "meta"
}