{
  "id": 472209,
  "title": "is it reasonable that surface_dice(random predict, random truth)=0.98?",
  "url": "/competitions/blood-vessel-segmentation/discussion/472209",
  "author_name": "",
  "post_date": "2024-01-31T07:06:35.111466400Z",
  "votes": 12,
  "comment_count": 15,
  "views": 0,
  "content": "<p>i use the code from <br>\n<a href=\"https://www.kaggle.com/code/junkoda/fast-surface-dice-computation\" target=\"_blank\">https://www.kaggle.com/code/junkoda/fast-surface-dice-computation</a></p>\n<p>please verify the following at your side:</p>\n<pre><code>D,H,W = ,,\npredict = np(,(D,H,W),p=)(np.uint8)\ntruth   = np(,(D,H,W),p=)(np.uint8)\n\nsurface_dice = (\n        predict, \n        truth,\n    )\n\n\n\n got surface_dice= !!!!\n</code></pre>\n<p>in the extreme case</p>\n<pre><code>=1-truth # truth is random as above\nsurface_dice = fast_compute_surface_dice_score_from_tensor(\n        predict, \n        truth,\n    )\n\n(surface_dice)\n\nstill got =0.98 !!!!\n</code></pre>\n<p>```</p>",
  "messages": [
    {
      "id": "2628219",
      "postDate": "01/31/2024 07:06:35",
      "content": "<p>i use the code from <br>\n<a href=\"https://www.kaggle.com/code/junkoda/fast-surface-dice-computation\" target=\"_blank\">https://www.kaggle.com/code/junkoda/fast-surface-dice-computation</a></p>\n<p>please verify the following at your side:</p>\n<pre><code>D,H,W = ,,\npredict = np(,(D,H,W),p=)(np.uint8)\ntruth   = np(,(D,H,W),p=)(np.uint8)\n\nsurface_dice = (\n        predict, \n        truth,\n    )\n\n\n\n got surface_dice= !!!!\n</code></pre>\n<p>in the extreme case</p>\n<pre><code>=1-truth # truth is random as above\nsurface_dice = fast_compute_surface_dice_score_from_tensor(\n        predict, \n        truth,\n    )\n\n(surface_dice)\n\nstill got =0.98 !!!!\n</code></pre>\n<p>```</p>",
      "rawMarkdown": "i use the code from \nhttps://www.kaggle.com/code/junkoda/fast-surface-dice-computation\n\nplease verify the following at your side:\n\n```\nD,H,W = 15,128,128\npredict = np.random.choice(2,(D,H,W),p=[0.5,0.5]).astype(np.uint8)\ntruth   = np.random.choice(2,(D,H,W),p=[0.5,0.5]).astype(np.uint8)\n\nsurface_dice = fast_compute_surface_dice_score_from_tensor(\n\t\tpredict, \n\t\ttruth,\n\t)\n\nprint(surface_dice)\n\ni got surface_dice=0.98 !!!!\n```\n\nin the extreme case\n\n```\n\npredict=1-truth # truth is random as above\nsurface_dice = fast_compute_surface_dice_score_from_tensor(\n\t\tpredict, \n\t\ttruth,\n\t)\n\nprint(surface_dice)\n\nstill got surface_dice=0.98 !!!!\n```\n```",
      "votes": null
    },
    {
      "id": "2628269",
      "postDate": "01/31/2024 07:44:18",
      "content": "<p>there is a serious bug in the computation metric.<br>\nit uses marching cube, the smalleset unit is 2x2x2.</p>\n<p>hence we cannot consider object that is smaller than this (i.e. it accuracy measurement vessel with 1 voxel thick is not accurate)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F79a52da1f639c9b53c01565b60fc280f%2FSelection_999(4717).png?generation=1706687008833090&amp;alt=media\"></p>\n<p>as an example,</p>\n<pre><code>if \nmc3\nmc4\nscore from the code  (which is wrong)\n</code></pre>",
      "rawMarkdown": "there is a serious bug in the computation metric.\nit uses marching cube, the smalleset unit is 2x2x2.\n\nhence we cannot consider object that is smaller than this (i.e. it accuracy measurement vessel with 1 voxel thick is not accurate)\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F79a52da1f639c9b53c01565b60fc280f%2FSelection_999(4717).png?generation=1706687008833090&alt=media)\n\nas an example,\n\n```\nif \ntruth=mc3\npredict=mc4\nscore from the code =1 (which is wrong)\n\n```",
      "votes": null
    },
    {
      "id": "2628271",
      "postDate": "01/31/2024 07:50:15",
      "content": "<p>why you have score=0.98 in score?</p>\n<p>since each voxel is randomly turn on of off which p=0.5,<br>\neach 2x2x2 cube will have 4 voxel turned on.</p>\n<p>hence you end up with</p>\n<pre><code>.g.\n = any of mc8, ,,,,, ...\n  = any of mc8, ,,,,,... \n = high value (even of the mc doesn't matched)\n</code></pre>\n<p>the code only check surfce area of 2x2x2 cube.<br>\nit don't care which of the  voxels in 2x2x2 cube are turned on or off.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2c2038210ea20f1e0d9b054b47f82cc6%2FSelection_999(4721).png?generation=1706688504698837&amp;alt=media\"></p>",
      "rawMarkdown": "why you have score=0.98 in score?\n\nsince each voxel is randomly turn on of off which p=0.5,\neach 2x2x2 cube will have 4 voxel turned on.\n\nhence you end up with\n\n```\ne.g.\ntruth = any of mc8, 9,10,11,12,13, ...\npredict  = any of mc8, 9,10,11,12,13,... \nscore = high value (even of the mc doesn't matched)\n\n```\n\nthe code only check surfce area of 2x2x2 cube.\nit don't care which of the  voxels in 2x2x2 cube are turned on or off.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2c2038210ea20f1e0d9b054b47f82cc6%2FSelection_999(4721).png?generation=1706688504698837&alt=media)",
      "votes": null
    },
    {
      "id": "2628363",
      "postDate": "01/31/2024 08:51:43",
      "content": "<p>Yes, the random ground truth and predictions that you create will actually have huge surface!</p>\n<p>If we consider two consecutive slices of shape (2, 128, 128), the corresponding surface lies in a single matrix of size 129x129 corresponding to each summit of your voxels.</p>\n<p>For an element of your surface to be 0 you need either all surrounding elements to be 0s or 1s. Since you are randomly sampling between 0 and 1 and that there are 8 pixels linked to 1 summit this has a chance to happen 1/256 (only 0s) + 1/256 (only 1s) = 1/128. So your surface is 1/128 zeros and all the rest are ones: &gt;99.2% of your pixels belong to the surface.</p>\n<p>The same holds true for both the labels and predictions you created.<br>\nComputing dice score on two images that mostly contains 1s will give you a very high dice score.</p>",
      "rawMarkdown": "Yes, the random ground truth and predictions that you create will actually have huge surface!\n\nIf we consider two consecutive slices of shape (2, 128, 128), the corresponding surface lies in a single matrix of size 129x129 corresponding to each summit of your voxels.\n\nFor an element of your surface to be 0 you need either all surrounding elements to be 0s or 1s. Since you are randomly sampling between 0 and 1 and that there are 8 pixels linked to 1 summit this has a chance to happen 1/256 (only 0s) + 1/256 (only 1s) = 1/128. So your surface is 1/128 zeros and all the rest are ones: >99.2% of your pixels belong to the surface.\n\nThe same holds true for both the labels and predictions you created.\nComputing dice score on two images that mostly contains 1s will give you a very high dice score.",
      "votes": null
    },
    {
      "id": "2628369",
      "postDate": "01/31/2024 08:55:36",
      "content": "<p>I am still surprised that with tolerance 0 we can have such high scores both on CV and LB, I think this is due to the fact that labels have been drawn using an algorithm, so the model can learn exactly where the algorithm would have stopped. If labels had been drawn by hand, everyone would have a score close to 0.<br>\nSee the discussion here: <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/461041\" target=\"_blank\">https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/461041</a></p>",
      "rawMarkdown": "I am still surprised that with tolerance 0 we can have such high scores both on CV and LB, I think this is due to the fact that labels have been drawn using an algorithm, so the model can learn exactly where the algorithm would have stopped. If labels had been drawn by hand, everyone would have a score close to 0.\nSee the discussion here: https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/461041",
      "votes": null
    },
    {
      "id": "2628427",
      "postDate": "01/31/2024 09:45:25",
      "content": "<p>Whilst I agree to some extent with what you say here about the labels, I think it is important to understand that the labelling process is incredibly manual here. <br>\nThere are many thousands of seed points selected by the annotators, and the choice between whether to use an intensity, contrast or hard limits on the 3D region growing, needs to be chosen by the annotator and then the thresholds for each type of limit needs to be altered just about every time a new point is selected. On top of this there is some voxel painting that has to be done afterwards to correct areas that are just never able to be filled by the semi-automated approach. So while you may be correct that this semi-automated approach enables higher scores on CV and LB, it is the only way this amount of data can be generated to even begin to explore these types of models. Note that even with this approach which is massively faster than voxel painting on consecutive slices, each of these datasets takes on the order of ~200 hrs of segmentation to make, and requires three expert annotators for our validation process. </p>",
      "rawMarkdown": "Whilst I agree to some extent with what you say here about the labels, I think it is important to understand that the labelling process is incredibly manual here. \nThere are many thousands of seed points selected by the annotators, and the choice between whether to use an intensity, contrast or hard limits on the 3D region growing, needs to be chosen by the annotator and then the thresholds for each type of limit needs to be altered just about every time a new point is selected. On top of this there is some voxel painting that has to be done afterwards to correct areas that are just never able to be filled by the semi-automated approach. So while you may be correct that this semi-automated approach enables higher scores on CV and LB, it is the only way this amount of data can be generated to even begin to explore these types of models. Note that even with this approach which is massively faster than voxel painting on consecutive slices, each of these datasets takes on the order of ~200 hrs of segmentation to make, and requires three expert annotators for our validation process.",
      "votes": null
    },
    {
      "id": "2628442",
      "postDate": "01/31/2024 10:05:53",
      "content": "<p>I have annotated 3D medical images myself for work, I know the struggle.</p>\n<p>DeepLearning algorithm can probably learn to ignore local noise to perform a similar segmentation as floodfill. On the parts drawn by hand, I think that a tolerance of 0 make things much more difficult, because humans can not annotate at the pixel level, so it's unlikely that the model can learn exactly where to stop and match the human annotators. Moreover, the surface dice score between two different annotators on a part drawn by hand will also be extremely low.</p>\n<p>This would have been different with a tolerance of a few pixels, that's all I am saying.</p>\n<p>I don't know if you'll have the chance to do this (would probably cost a lot of computation time for Kaggle): compute a few different metrics on all submissions (surface dice with tolerance 1, 2, 3, 5 and 3D volumic dice score) and see how much it shakes the leaderboard. If there is a large shake, the metric probably does not reflect the quality of segmentations, if there is only a small shake this means that the surface dice with tolerance 0 is just fine. Maybe you'll be able to do this experiment only on selected submission of the top 100 competitors to save computational time!</p>",
      "rawMarkdown": "I have annotated 3D medical images myself for work, I know the struggle.\n\nDeepLearning algorithm can probably learn to ignore local noise to perform a similar segmentation as floodfill. On the parts drawn by hand, I think that a tolerance of 0 make things much more difficult, because humans can not annotate at the pixel level, so it's unlikely that the model can learn exactly where to stop and match the human annotators. Moreover, the surface dice score between two different annotators on a part drawn by hand will also be extremely low.\n\nThis would have been different with a tolerance of a few pixels, that's all I am saying.\n\nI don't know if you'll have the chance to do this (would probably cost a lot of computation time for Kaggle): compute a few different metrics on all submissions (surface dice with tolerance 1, 2, 3, 5 and 3D volumic dice score) and see how much it shakes the leaderboard. If there is a large shake, the metric probably does not reflect the quality of segmentations, if there is only a small shake this means that the surface dice with tolerance 0 is just fine. Maybe you'll be able to do this experiment only on selected submission of the top 100 competitors to save computational time!",
      "votes": null
    },
    {
      "id": "2628462",
      "postDate": "01/31/2024 10:15:20",
      "content": "<p>\"I am still surprised that with tolerance 0 we can have such high scores both on CV and LB,\"</p>\n<p>the reason is:</p>\n<ul>\n<li>consider kidney3/dense and kidney3/sparse annotation</li>\n<li>sparsity is 0.85 but suface dice is 0.98 (contribution of small vessels are low?)</li>\n</ul>\n<p>Now, vessel(dark) is surrounded by vessel wall (bright). despite manual flood fill, i think results are pretty consistent (at least for the big vessels).</p>\n<p>Then there is also the metric issue: consider 2x2x2 as smallest unit. although tolerance is zero, it is not really zero (becuase of 2x2x2)</p>\n<hr>\n<p>you can measure the following 2d dice (i.e. computation at each image) for your model:</p>\n<ul>\n<li>normal 2d area dice</li>\n<li>do edge detection and measure boundary(1 pixel) dice</li>\n<li>same as above, but boundary(2 pixel) dice</li>\n</ul>",
      "rawMarkdown": "\"I am still surprised that with tolerance 0 we can have such high scores both on CV and LB,\"\n\nthe reason is:\n- consider kidney3/dense and kidney3/sparse annotation\n- sparsity is 0.85 but suface dice is 0.98 (contribution of small vessels are low?)\n\nNow, vessel(dark) is surrounded by vessel wall (bright). despite manual flood fill, i think results are pretty consistent (at least for the big vessels).\n\nThen there is also the metric issue: consider 2x2x2 as smallest unit. although tolerance is zero, it is not really zero (becuase of 2x2x2)\n\n---\n\nyou can measure the following 2d dice (i.e. computation at each image) for your model:\n- normal 2d area dice\n- do edge detection and measure boundary(1 pixel) dice\n- same as above, but boundary(2 pixel) dice",
      "votes": null
    },
    {
      "id": "2628507",
      "postDate": "01/31/2024 10:48:39",
      "content": "<p>Yep I definitely see your points, I think there need to be more metrics out there specifically for blood vessel networks that take into consideration how the \"functional\" behaviour of the network is affected by the segmentation. Thanks for the idea for the various tolerances it is a good one. I will discuss with the rest of the team. </p>",
      "rawMarkdown": "Yep I definitely see your points, I think there need to be more metrics out there specifically for blood vessel networks that take into consideration how the \"functional\" behaviour of the network is affected by the segmentation. Thanks for the idea for the various tolerances it is a good one. I will discuss with the rest of the team.",
      "votes": null
    },
    {
      "id": "2628711",
      "postDate": "01/31/2024 13:00:14",
      "content": "<p>That's a good insight. Do you think it will get patched?</p>",
      "rawMarkdown": "That's a good insight. Do you think it will get patched?",
      "votes": null
    },
    {
      "id": "2628786",
      "postDate": "01/31/2024 13:46:42",
      "content": "<p>Unlikely. </p>\n<p>Making dramatic changes in the last week of the competition is almost always a bad idea. </p>",
      "rawMarkdown": "Unlikely. \n\nMaking dramatic changes in the last week of the competition is almost always a bad idea.",
      "votes": null
    },
    {
      "id": "2628794",
      "postDate": "01/31/2024 13:49:18",
      "content": "<p>That's not a bug, that's the expected behavior.</p>",
      "rawMarkdown": "That's not a bug, that's the expected behavior.",
      "votes": null
    },
    {
      "id": "2629685",
      "postDate": "02/01/2024 00:45:20",
      "content": "<p>I think it's expected that you see high scores.</p>\n<p>The current surface dice score code doesn't care if the marched cube aligns in the same \"byte code\" that they compute. It simply computes the surface area of the marching cube, then compute the intersection of those area with a very naive intersection mask of <code>idx = torch.logical_and(area_pred &gt; 0, area_true &gt; 0)</code></p>\n<p>this means in a simple case where the ground truth is a 2x2x2 cube and a prediction is a 2x2x2 cube:<br>\npred = 00000001, label = 00010000 gets you 1.0 surface dice, despite not really making a lot of sense<br>\n(once again, the reason is because these 2 different marched cube yields the same surface area, despite not intersecting)</p>\n<p>note that if you plug some 2x2x2 cube into the surface dice code you'll not get 1.0 out because the metric pads the start and the end slice with zeros, but my point still stands - that it does naively compute the sum of the areas where the surface area is not 0, instead of actually computing the intersection of the surface area.</p>",
      "rawMarkdown": "I think it's expected that you see high scores.\n\nThe current surface dice score code doesn't care if the marched cube aligns in the same \"byte code\" that they compute. It simply computes the surface area of the marching cube, then compute the intersection of those area with a very naive intersection mask of `idx = torch.logical_and(area_pred > 0, area_true > 0)`\n\nthis means in a simple case where the ground truth is a 2x2x2 cube and a prediction is a 2x2x2 cube:\npred = 00000001, label = 00010000 gets you 1.0 surface dice, despite not really making a lot of sense\n(once again, the reason is because these 2 different marched cube yields the same surface area, despite not intersecting)\n\nnote that if you plug some 2x2x2 cube into the surface dice code you'll not get 1.0 out because the metric pads the start and the end slice with zeros, but my point still stands - that it does naively compute the sum of the areas where the surface area is not 0, instead of actually computing the intersection of the surface area.",
      "votes": null
    },
    {
      "id": "2629694",
      "postDate": "02/01/2024 00:53:09",
      "content": "<p>whoops, just saw that it has already been figured out! 😅</p>",
      "rawMarkdown": "whoops, just saw that it has already been figured out! 😅",
      "votes": null
    },
    {
      "id": "2630039",
      "postDate": "02/01/2024 06:04:19",
      "content": "<p>\"Making dramatic changes in the last week of the competition is almost always a bad idea.\"</p>\n<p>maybe this can be done after the competition.<br>\na simple fix is just to upscale the truth and prediction by 4x (using nearest neighbour). then apply the surface dice metric code as usual</p>",
      "rawMarkdown": "\"Making dramatic changes in the last week of the competition is almost always a bad idea.\"\n\nmaybe this can be done after the competition.\na simple fix is just to upscale the truth and prediction by 4x (using nearest neighbour). then apply the surface dice metric code as usual",
      "votes": null
    },
    {
      "id": "2630826",
      "postDate": "02/01/2024 13:44:08",
      "content": "<p>From very beginning we know that It is either metric sensitive or insensitive competition.</p>",
      "rawMarkdown": "From very beginning we know that It is either metric sensitive or insensitive competition.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2628269,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "01/31/2024 07:44:18",
      "content": "<p>there is a serious bug in the computation metric.<br>\nit uses marching cube, the smalleset unit is 2x2x2.</p>\n<p>hence we cannot consider object that is smaller than this (i.e. it accuracy measurement vessel with 1 voxel thick is not accurate)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F79a52da1f639c9b53c01565b60fc280f%2FSelection_999(4717).png?generation=1706687008833090&amp;alt=media\"></p>\n<p>as an example,</p>\n<pre><code>if \nmc3\nmc4\nscore from the code  (which is wrong)\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 2628271,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "01/31/2024 07:50:15",
          "content": "<p>why you have score=0.98 in score?</p>\n<p>since each voxel is randomly turn on of off which p=0.5,<br>\neach 2x2x2 cube will have 4 voxel turned on.</p>\n<p>hence you end up with</p>\n<pre><code>.g.\n = any of mc8, ,,,,, ...\n  = any of mc8, ,,,,,... \n = high value (even of the mc doesn't matched)\n</code></pre>\n<p>the code only check surfce area of 2x2x2 cube.<br>\nit don't care which of the  voxels in 2x2x2 cube are turned on or off.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2c2038210ea20f1e0d9b054b47f82cc6%2FSelection_999(4721).png?generation=1706688504698837&amp;alt=media\"></p>",
          "votes": null,
          "replies": [
            {
              "id": 2628711,
              "author_name": "yassinealouini",
              "author_url": "",
              "post_date": "01/31/2024 13:00:14",
              "content": "<p>That's a good insight. Do you think it will get patched?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2628786,
                  "author_name": "ivanpan",
                  "author_url": "",
                  "post_date": "01/31/2024 13:46:42",
                  "content": "<p>Unlikely. </p>\n<p>Making dramatic changes in the last week of the competition is almost always a bad idea. </p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2628794,
                      "author_name": "optimo",
                      "author_url": "",
                      "post_date": "01/31/2024 13:49:18",
                      "content": "<p>That's not a bug, that's the expected behavior.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2630039,
                          "author_name": "hengck23",
                          "author_url": "",
                          "post_date": "02/01/2024 06:04:19",
                          "content": "<p>\"Making dramatic changes in the last week of the competition is almost always a bad idea.\"</p>\n<p>maybe this can be done after the competition.<br>\na simple fix is just to upscale the truth and prediction by 4x (using nearest neighbour). then apply the surface dice metric code as usual</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2628363,
      "author_name": "optimo",
      "author_url": "",
      "post_date": "01/31/2024 08:51:43",
      "content": "<p>Yes, the random ground truth and predictions that you create will actually have huge surface!</p>\n<p>If we consider two consecutive slices of shape (2, 128, 128), the corresponding surface lies in a single matrix of size 129x129 corresponding to each summit of your voxels.</p>\n<p>For an element of your surface to be 0 you need either all surrounding elements to be 0s or 1s. Since you are randomly sampling between 0 and 1 and that there are 8 pixels linked to 1 summit this has a chance to happen 1/256 (only 0s) + 1/256 (only 1s) = 1/128. So your surface is 1/128 zeros and all the rest are ones: &gt;99.2% of your pixels belong to the surface.</p>\n<p>The same holds true for both the labels and predictions you created.<br>\nComputing dice score on two images that mostly contains 1s will give you a very high dice score.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2628369,
          "author_name": "optimo",
          "author_url": "",
          "post_date": "01/31/2024 08:55:36",
          "content": "<p>I am still surprised that with tolerance 0 we can have such high scores both on CV and LB, I think this is due to the fact that labels have been drawn using an algorithm, so the model can learn exactly where the algorithm would have stopped. If labels had been drawn by hand, everyone would have a score close to 0.<br>\nSee the discussion here: <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/461041\" target=\"_blank\">https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/461041</a></p>",
          "votes": null,
          "replies": [
            {
              "id": 2628427,
              "author_name": "clairewalsh",
              "author_url": "",
              "post_date": "01/31/2024 09:45:25",
              "content": "<p>Whilst I agree to some extent with what you say here about the labels, I think it is important to understand that the labelling process is incredibly manual here. <br>\nThere are many thousands of seed points selected by the annotators, and the choice between whether to use an intensity, contrast or hard limits on the 3D region growing, needs to be chosen by the annotator and then the thresholds for each type of limit needs to be altered just about every time a new point is selected. On top of this there is some voxel painting that has to be done afterwards to correct areas that are just never able to be filled by the semi-automated approach. So while you may be correct that this semi-automated approach enables higher scores on CV and LB, it is the only way this amount of data can be generated to even begin to explore these types of models. Note that even with this approach which is massively faster than voxel painting on consecutive slices, each of these datasets takes on the order of ~200 hrs of segmentation to make, and requires three expert annotators for our validation process. </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2628442,
                  "author_name": "optimo",
                  "author_url": "",
                  "post_date": "01/31/2024 10:05:53",
                  "content": "<p>I have annotated 3D medical images myself for work, I know the struggle.</p>\n<p>DeepLearning algorithm can probably learn to ignore local noise to perform a similar segmentation as floodfill. On the parts drawn by hand, I think that a tolerance of 0 make things much more difficult, because humans can not annotate at the pixel level, so it's unlikely that the model can learn exactly where to stop and match the human annotators. Moreover, the surface dice score between two different annotators on a part drawn by hand will also be extremely low.</p>\n<p>This would have been different with a tolerance of a few pixels, that's all I am saying.</p>\n<p>I don't know if you'll have the chance to do this (would probably cost a lot of computation time for Kaggle): compute a few different metrics on all submissions (surface dice with tolerance 1, 2, 3, 5 and 3D volumic dice score) and see how much it shakes the leaderboard. If there is a large shake, the metric probably does not reflect the quality of segmentations, if there is only a small shake this means that the surface dice with tolerance 0 is just fine. Maybe you'll be able to do this experiment only on selected submission of the top 100 competitors to save computational time!</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2628462,
                      "author_name": "hengck23",
                      "author_url": "",
                      "post_date": "01/31/2024 10:15:20",
                      "content": "<p>\"I am still surprised that with tolerance 0 we can have such high scores both on CV and LB,\"</p>\n<p>the reason is:</p>\n<ul>\n<li>consider kidney3/dense and kidney3/sparse annotation</li>\n<li>sparsity is 0.85 but suface dice is 0.98 (contribution of small vessels are low?)</li>\n</ul>\n<p>Now, vessel(dark) is surrounded by vessel wall (bright). despite manual flood fill, i think results are pretty consistent (at least for the big vessels).</p>\n<p>Then there is also the metric issue: consider 2x2x2 as smallest unit. although tolerance is zero, it is not really zero (becuase of 2x2x2)</p>\n<hr>\n<p>you can measure the following 2d dice (i.e. computation at each image) for your model:</p>\n<ul>\n<li>normal 2d area dice</li>\n<li>do edge detection and measure boundary(1 pixel) dice</li>\n<li>same as above, but boundary(2 pixel) dice</li>\n</ul>",
                      "votes": null,
                      "replies": []
                    },
                    {
                      "id": 2628507,
                      "author_name": "clairewalsh",
                      "author_url": "",
                      "post_date": "01/31/2024 10:48:39",
                      "content": "<p>Yep I definitely see your points, I think there need to be more metrics out there specifically for blood vessel networks that take into consideration how the \"functional\" behaviour of the network is affected by the segmentation. Thanks for the idea for the various tolerances it is a good one. I will discuss with the rest of the team. </p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2629685,
      "author_name": "sirapoabchaikunsaeng",
      "author_url": "",
      "post_date": "02/01/2024 00:45:20",
      "content": "<p>I think it's expected that you see high scores.</p>\n<p>The current surface dice score code doesn't care if the marched cube aligns in the same \"byte code\" that they compute. It simply computes the surface area of the marching cube, then compute the intersection of those area with a very naive intersection mask of <code>idx = torch.logical_and(area_pred &gt; 0, area_true &gt; 0)</code></p>\n<p>this means in a simple case where the ground truth is a 2x2x2 cube and a prediction is a 2x2x2 cube:<br>\npred = 00000001, label = 00010000 gets you 1.0 surface dice, despite not really making a lot of sense<br>\n(once again, the reason is because these 2 different marched cube yields the same surface area, despite not intersecting)</p>\n<p>note that if you plug some 2x2x2 cube into the surface dice code you'll not get 1.0 out because the metric pads the start and the end slice with zeros, but my point still stands - that it does naively compute the sum of the areas where the surface area is not 0, instead of actually computing the intersection of the surface area.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2629694,
          "author_name": "sirapoabchaikunsaeng",
          "author_url": "",
          "post_date": "02/01/2024 00:53:09",
          "content": "<p>whoops, just saw that it has already been figured out! 😅</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2630826,
      "author_name": "dragonzhang",
      "author_url": "",
      "post_date": "02/01/2024 13:44:08",
      "content": "<p>From very beginning we know that It is either metric sensitive or insensitive competition.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2628219": "i use the code from \nhttps://www.kaggle.com/code/junkoda/fast-surface-dice-computation\n\nplease verify the following at your side:\n\n```\nD,H,W = 15,128,128\npredict = np.random.choice(2,(D,H,W),p=[0.5,0.5]).astype(np.uint8)\ntruth   = np.random.choice(2,(D,H,W),p=[0.5,0.5]).astype(np.uint8)\n\nsurface_dice = fast_compute_surface_dice_score_from_tensor(\n\t\tpredict, \n\t\ttruth,\n\t)\n\nprint(surface_dice)\n\ni got surface_dice=0.98 !!!!\n```\n\nin the extreme case\n\n```\n\npredict=1-truth # truth is random as above\nsurface_dice = fast_compute_surface_dice_score_from_tensor(\n\t\tpredict, \n\t\ttruth,\n\t)\n\nprint(surface_dice)\n\nstill got surface_dice=0.98 !!!!\n```\n```",
    "2628269": "there is a serious bug in the computation metric.\nit uses marching cube, the smalleset unit is 2x2x2.\n\nhence we cannot consider object that is smaller than this (i.e. it accuracy measurement vessel with 1 voxel thick is not accurate)\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F79a52da1f639c9b53c01565b60fc280f%2FSelection_999(4717).png?generation=1706687008833090&alt=media)\n\nas an example,\n\n```\nif \ntruth=mc3\npredict=mc4\nscore from the code =1 (which is wrong)\n\n```",
    "2628271": "why you have score=0.98 in score?\n\nsince each voxel is randomly turn on of off which p=0.5,\neach 2x2x2 cube will have 4 voxel turned on.\n\nhence you end up with\n\n```\ne.g.\ntruth = any of mc8, 9,10,11,12,13, ...\npredict  = any of mc8, 9,10,11,12,13,... \nscore = high value (even of the mc doesn't matched)\n\n```\n\nthe code only check surfce area of 2x2x2 cube.\nit don't care which of the  voxels in 2x2x2 cube are turned on or off.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2c2038210ea20f1e0d9b054b47f82cc6%2FSelection_999(4721).png?generation=1706688504698837&alt=media)",
    "2628363": "Yes, the random ground truth and predictions that you create will actually have huge surface!\n\nIf we consider two consecutive slices of shape (2, 128, 128), the corresponding surface lies in a single matrix of size 129x129 corresponding to each summit of your voxels.\n\nFor an element of your surface to be 0 you need either all surrounding elements to be 0s or 1s. Since you are randomly sampling between 0 and 1 and that there are 8 pixels linked to 1 summit this has a chance to happen 1/256 (only 0s) + 1/256 (only 1s) = 1/128. So your surface is 1/128 zeros and all the rest are ones: >99.2% of your pixels belong to the surface.\n\nThe same holds true for both the labels and predictions you created.\nComputing dice score on two images that mostly contains 1s will give you a very high dice score.",
    "2628369": "I am still surprised that with tolerance 0 we can have such high scores both on CV and LB, I think this is due to the fact that labels have been drawn using an algorithm, so the model can learn exactly where the algorithm would have stopped. If labels had been drawn by hand, everyone would have a score close to 0.\nSee the discussion here: https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/461041",
    "2628427": "Whilst I agree to some extent with what you say here about the labels, I think it is important to understand that the labelling process is incredibly manual here. \nThere are many thousands of seed points selected by the annotators, and the choice between whether to use an intensity, contrast or hard limits on the 3D region growing, needs to be chosen by the annotator and then the thresholds for each type of limit needs to be altered just about every time a new point is selected. On top of this there is some voxel painting that has to be done afterwards to correct areas that are just never able to be filled by the semi-automated approach. So while you may be correct that this semi-automated approach enables higher scores on CV and LB, it is the only way this amount of data can be generated to even begin to explore these types of models. Note that even with this approach which is massively faster than voxel painting on consecutive slices, each of these datasets takes on the order of ~200 hrs of segmentation to make, and requires three expert annotators for our validation process.",
    "2628442": "I have annotated 3D medical images myself for work, I know the struggle.\n\nDeepLearning algorithm can probably learn to ignore local noise to perform a similar segmentation as floodfill. On the parts drawn by hand, I think that a tolerance of 0 make things much more difficult, because humans can not annotate at the pixel level, so it's unlikely that the model can learn exactly where to stop and match the human annotators. Moreover, the surface dice score between two different annotators on a part drawn by hand will also be extremely low.\n\nThis would have been different with a tolerance of a few pixels, that's all I am saying.\n\nI don't know if you'll have the chance to do this (would probably cost a lot of computation time for Kaggle): compute a few different metrics on all submissions (surface dice with tolerance 1, 2, 3, 5 and 3D volumic dice score) and see how much it shakes the leaderboard. If there is a large shake, the metric probably does not reflect the quality of segmentations, if there is only a small shake this means that the surface dice with tolerance 0 is just fine. Maybe you'll be able to do this experiment only on selected submission of the top 100 competitors to save computational time!",
    "2628462": "\"I am still surprised that with tolerance 0 we can have such high scores both on CV and LB,\"\n\nthe reason is:\n- consider kidney3/dense and kidney3/sparse annotation\n- sparsity is 0.85 but suface dice is 0.98 (contribution of small vessels are low?)\n\nNow, vessel(dark) is surrounded by vessel wall (bright). despite manual flood fill, i think results are pretty consistent (at least for the big vessels).\n\nThen there is also the metric issue: consider 2x2x2 as smallest unit. although tolerance is zero, it is not really zero (becuase of 2x2x2)\n\n---\n\nyou can measure the following 2d dice (i.e. computation at each image) for your model:\n- normal 2d area dice\n- do edge detection and measure boundary(1 pixel) dice\n- same as above, but boundary(2 pixel) dice",
    "2628507": "Yep I definitely see your points, I think there need to be more metrics out there specifically for blood vessel networks that take into consideration how the \"functional\" behaviour of the network is affected by the segmentation. Thanks for the idea for the various tolerances it is a good one. I will discuss with the rest of the team.",
    "2628711": "That's a good insight. Do you think it will get patched?",
    "2628786": "Unlikely. \n\nMaking dramatic changes in the last week of the competition is almost always a bad idea.",
    "2628794": "That's not a bug, that's the expected behavior.",
    "2629685": "I think it's expected that you see high scores.\n\nThe current surface dice score code doesn't care if the marched cube aligns in the same \"byte code\" that they compute. It simply computes the surface area of the marching cube, then compute the intersection of those area with a very naive intersection mask of `idx = torch.logical_and(area_pred > 0, area_true > 0)`\n\nthis means in a simple case where the ground truth is a 2x2x2 cube and a prediction is a 2x2x2 cube:\npred = 00000001, label = 00010000 gets you 1.0 surface dice, despite not really making a lot of sense\n(once again, the reason is because these 2 different marched cube yields the same surface area, despite not intersecting)\n\nnote that if you plug some 2x2x2 cube into the surface dice code you'll not get 1.0 out because the metric pads the start and the end slice with zeros, but my point still stands - that it does naively compute the sum of the areas where the surface area is not 0, instead of actually computing the intersection of the surface area.",
    "2629694": "whoops, just saw that it has already been figured out! 😅",
    "2630039": "\"Making dramatic changes in the last week of the competition is almost always a bad idea.\"\n\nmaybe this can be done after the competition.\na simple fix is just to upscale the truth and prediction by 4x (using nearest neighbour). then apply the surface dice metric code as usual",
    "2630826": "From very beginning we know that It is either metric sensitive or insensitive competition."
  },
  "source": "meta"
}