{
  "id": 461041,
  "title": "Right choice of metric?",
  "url": "/competitions/blood-vessel-segmentation/discussion/461041",
  "author_name": "",
  "post_date": "2023-12-12T11:16:24.822855300Z",
  "votes": 9,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I understand that Surface Dice seemed better to focus only on the surface of predictions, which gives a similar weight to small vessels and large vessels.</p>\n<p>However, if computing the competition metrics on one kidney and emulating predictions as the mask with a simple dilation or erosion in 2D (for each slice independently) here are the results:</p>\n<blockquote>\n  <p>erosion score 0.6113542914390564<br>\n  dilation score 0.5309284329414368<br>\n  Num iterations: 2<br>\n  erosion score 0.18636780977249146<br>\n  dilation score 0.10018275678157806<br>\n  Num iterations: 3<br>\n  erosion score 0.08591027557849884<br>\n  dilation score 0.026479000225663185</p>\n</blockquote>\n<p>With a small 2D dilation or erosion of a single pixel the competition score drops to 0.53 and 0.611 respectively. Things get worse when the number of dilations or erosions increases.</p>\n<p>These sudden drops happen because the tolerance here has been set to 0, but are we confident that the annotations are correct at the pixel level ? As <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> showed <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/456118#2543402\" target=\"_blank\">here</a> that labels can have shifts. It is very likely that two annotators would not set the exact same annotations at pixel level. I am actually quite surprised that models can beat the 1 pixel dilation or erosion benchmark, does it mean that we already have excellent models? Or that the metric used is not so good?</p>\n<p>I also don't understand why after 3 iterations, the score is not exactly 0 ? <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> do you understand why ?</p>\n<p><a href=\"https://www.kaggle.com/clairewalsh\" target=\"_blank\">@clairewalsh</a> don't you think that some level of tolerance would yield more accurate results to spot the best solution ?</p>",
  "messages": [
    {
      "id": "2558807",
      "postDate": "12/12/2023 11:16:24",
      "content": "<p>I understand that Surface Dice seemed better to focus only on the surface of predictions, which gives a similar weight to small vessels and large vessels.</p>\n<p>However, if computing the competition metrics on one kidney and emulating predictions as the mask with a simple dilation or erosion in 2D (for each slice independently) here are the results:</p>\n<blockquote>\n  <p>erosion score 0.6113542914390564<br>\n  dilation score 0.5309284329414368<br>\n  Num iterations: 2<br>\n  erosion score 0.18636780977249146<br>\n  dilation score 0.10018275678157806<br>\n  Num iterations: 3<br>\n  erosion score 0.08591027557849884<br>\n  dilation score 0.026479000225663185</p>\n</blockquote>\n<p>With a small 2D dilation or erosion of a single pixel the competition score drops to 0.53 and 0.611 respectively. Things get worse when the number of dilations or erosions increases.</p>\n<p>These sudden drops happen because the tolerance here has been set to 0, but are we confident that the annotations are correct at the pixel level ? As <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> showed <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/456118#2543402\" target=\"_blank\">here</a> that labels can have shifts. It is very likely that two annotators would not set the exact same annotations at pixel level. I am actually quite surprised that models can beat the 1 pixel dilation or erosion benchmark, does it mean that we already have excellent models? Or that the metric used is not so good?</p>\n<p>I also don't understand why after 3 iterations, the score is not exactly 0 ? <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> do you understand why ?</p>\n<p><a href=\"https://www.kaggle.com/clairewalsh\" target=\"_blank\">@clairewalsh</a> don't you think that some level of tolerance would yield more accurate results to spot the best solution ?</p>",
      "rawMarkdown": "I understand that Surface Dice seemed better to focus only on the surface of predictions, which gives a similar weight to small vessels and large vessels.\n\nHowever, if computing the competition metrics on one kidney and emulating predictions as the mask with a simple dilation or erosion in 2D (for each slice independently) here are the results:\n>erosion score 0.6113542914390564\ndilation score 0.5309284329414368\nNum iterations: 2\nerosion score 0.18636780977249146\ndilation score 0.10018275678157806\nNum iterations: 3\nerosion score 0.08591027557849884\ndilation score 0.026479000225663185\n\nWith a small 2D dilation or erosion of a single pixel the competition score drops to 0.53 and 0.611 respectively. Things get worse when the number of dilations or erosions increases.\n\nThese sudden drops happen because the tolerance here has been set to 0, but are we confident that the annotations are correct at the pixel level ? As @hengck23 showed [here](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/456118#2543402) that labels can have shifts. It is very likely that two annotators would not set the exact same annotations at pixel level. I am actually quite surprised that models can beat the 1 pixel dilation or erosion benchmark, does it mean that we already have excellent models? Or that the metric used is not so good?\n\nI also don't understand why after 3 iterations, the score is not exactly 0 ? @junkoda do you understand why ?\n\n@clairewalsh don't you think that some level of tolerance would yield more accurate results to spot the best solution ?",
      "votes": null
    },
    {
      "id": "2559529",
      "postDate": "12/12/2023 22:45:51",
      "content": "<p>I agree that surface Dice with tolerance zero is difficult to generalize to different resolutions. I just started and don't know if that's reasonable or unreasonable, though. Center-line Dice, in organizers review, seems to be more resolution independent, but I don't know any medical meaning of line vs surface. </p>\n<p>I don't know if the score should be mathematically zero. I guess random match of true positives can be possible.</p>",
      "rawMarkdown": "I agree that surface Dice with tolerance zero is difficult to generalize to different resolutions. I just started and don't know if that's reasonable or unreasonable, though. Center-line Dice, in organizers review, seems to be more resolution independent, but I don't know any medical meaning of line vs surface. \n\nI don't know if the score should be mathematically zero. I guess random match of true positives can be possible.",
      "votes": null
    },
    {
      "id": "2560365",
      "postDate": "12/13/2023 14:19:31",
      "content": "<p>agree. one of the worst metric I ever met in competitions.</p>",
      "rawMarkdown": "agree. one of the worst metric I ever met in competitions.",
      "votes": null
    },
    {
      "id": "2561703",
      "postDate": "12/14/2023 18:18:24",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a> thanks for your insights into this metric. <br>\nI know the metric has caused challenges throughout the competition so here is a bit about what and why we used it: <br>\nWe spent a considerable amount of time in the lead up to this competition trying to find a metric that would provide a good measure the output that we are aiming for in terms of the end use of these types of segmentations.<br>\nWith regard to tolerance, we did investigate how changing the tolerance changes the output. The challenge we found is that increasing the tolerance even to 1 for the surface dice, would give a high metric output (high 0.9), despite the very small vasculature being highly disconnected, something which causes fundamentally different results in downstream analyses such a flow modelling. <br>\nRegarding your comment about different annotators having different annotations not within 1 pixel, this again is a challenge and is why we had the three annotator validation process (two annotators working on the same dataset sequentially) then a third doing a quantitative scoring system. This was our best approach to getting a consensus multi-annotator ground truth. Such a process takes enormous effort on the part of the annotators particularly with such multi-scale structures.  Other metrics such as the  cl-dice which we also explored whilst interesting for preserving connectivity has a skeletonisation component to it which is also somewhat computationally intensive and can cause artefacts (ball like structures) at complicated junctions i. more than simple bifurcation or in the collapsed vessel structures. <br>\nThe metrics is indeed not a simple topic ,and what is the 'best' metric for multiscale vascular segmentations is, I believe an open and interesting research question in and of itself. I think this competition, shows this to be the case.   </p>",
      "rawMarkdown": "Hi @optimo thanks for your insights into this metric. \nI know the metric has caused challenges throughout the competition so here is a bit about what and why we used it: \nWe spent a considerable amount of time in the lead up to this competition trying to find a metric that would provide a good measure the output that we are aiming for in terms of the end use of these types of segmentations.\nWith regard to tolerance, we did investigate how changing the tolerance changes the output. The challenge we found is that increasing the tolerance even to 1 for the surface dice, would give a high metric output (high 0.9), despite the very small vasculature being highly disconnected, something which causes fundamentally different results in downstream analyses such a flow modelling. \nRegarding your comment about different annotators having different annotations not within 1 pixel, this again is a challenge and is why we had the three annotator validation process (two annotators working on the same dataset sequentially) then a third doing a quantitative scoring system. This was our best approach to getting a consensus multi-annotator ground truth. Such a process takes enormous effort on the part of the annotators particularly with such multi-scale structures.  Other metrics such as the  cl-dice which we also explored whilst interesting for preserving connectivity has a skeletonisation component to it which is also somewhat computationally intensive and can cause artefacts (ball like structures) at complicated junctions i. more than simple bifurcation or in the collapsed vessel structures. \nThe metrics is indeed not a simple topic ,and what is the 'best' metric for multiscale vascular segmentations is, I believe an open and interesting research question in and of itself. I think this competition, shows this to be the case.",
      "votes": null
    },
    {
      "id": "2561854",
      "postDate": "12/14/2023 21:53:34",
      "content": "<p>Thanks for the clarifications <a href=\"https://www.kaggle.com/clairewalsh\" target=\"_blank\">@clairewalsh</a>, it is indeed an open problem and I'm glad to know you have thought about the different possibilities and chosen the best metric for your final applications.<br>\n<a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> since we have the right metric, could you please update the code with <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a>'s implementation as it is much faster and uses far less memory? This will be a big win for everyone in the competition!</p>",
      "rawMarkdown": "Thanks for the clarifications @clairewalsh, it is indeed an open problem and I'm glad to know you have thought about the different possibilities and chosen the best metric for your final applications.\n@ryanholbrook since we have the right metric, could you please update the code with @junkoda's implementation as it is much faster and uses far less memory? This will be a big win for everyone in the competition!",
      "votes": null
    },
    {
      "id": "2561882",
      "postDate": "12/14/2023 23:14:50",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a>,</p>\n<p>I am currently heads down on the (very much overdue) <a href=\"https://www.kaggle.com/discussions/competition-hosting/459827\" target=\"_blank\">Santa 2023</a> competition and won't have time to spare for a bit. I'll try to take a look at the new implementation, but it's unlikely I'll be able to make any changes until after the new year. Evaluation metrics in running competitions we need to treat as production code and changes thus require considerable care. (Breaking a competition over the holidays would be very bad for everyone.)</p>",
      "rawMarkdown": "Hi @optimo,\n\nI am currently heads down on the (very much overdue) [Santa 2023](https://www.kaggle.com/discussions/competition-hosting/459827) competition and won't have time to spare for a bit. I'll try to take a look at the new implementation, but it's unlikely I'll be able to make any changes until after the new year. Evaluation metrics in running competitions we need to treat as production code and changes thus require considerable care. (Breaking a competition over the holidays would be very bad for everyone.)",
      "votes": null
    },
    {
      "id": "2561917",
      "postDate": "12/15/2023 00:36:05",
      "content": "<p>\" (very much overdue) Santa 2023 \"</p>\n<p>is there still a santa 2023? i am still lokking forward it :)</p>",
      "rawMarkdown": "\" (very much overdue) Santa 2023 \"\n\nis there still a santa 2023? i am still lokking forward it :)",
      "votes": null
    },
    {
      "id": "2562107",
      "postDate": "12/15/2023 05:43:00",
      "content": "<p>Indeed - see this post <a href=\"https://www.kaggle.com/discussions/competition-hosting/459827\" target=\"_blank\">Santa, we miss you (2023)</a>  <br>\ngood luck <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> with your hard work, Santa will surely have you on the nice list!</p>",
      "rawMarkdown": "Indeed - see this post [Santa, we miss you (2023)](https://www.kaggle.com/discussions/competition-hosting/459827)  \ngood luck @ryanholbrook with your hard work, Santa will surely have you on the nice list!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2559529,
      "author_name": "junkoda",
      "author_url": "",
      "post_date": "12/12/2023 22:45:51",
      "content": "<p>I agree that surface Dice with tolerance zero is difficult to generalize to different resolutions. I just started and don't know if that's reasonable or unreasonable, though. Center-line Dice, in organizers review, seems to be more resolution independent, but I don't know any medical meaning of line vs surface. </p>\n<p>I don't know if the score should be mathematically zero. I guess random match of true positives can be possible.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2560365,
      "author_name": "dragonzhang",
      "author_url": "",
      "post_date": "12/13/2023 14:19:31",
      "content": "<p>agree. one of the worst metric I ever met in competitions.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2561703,
      "author_name": "clairewalsh",
      "author_url": "",
      "post_date": "12/14/2023 18:18:24",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a> thanks for your insights into this metric. <br>\nI know the metric has caused challenges throughout the competition so here is a bit about what and why we used it: <br>\nWe spent a considerable amount of time in the lead up to this competition trying to find a metric that would provide a good measure the output that we are aiming for in terms of the end use of these types of segmentations.<br>\nWith regard to tolerance, we did investigate how changing the tolerance changes the output. The challenge we found is that increasing the tolerance even to 1 for the surface dice, would give a high metric output (high 0.9), despite the very small vasculature being highly disconnected, something which causes fundamentally different results in downstream analyses such a flow modelling. <br>\nRegarding your comment about different annotators having different annotations not within 1 pixel, this again is a challenge and is why we had the three annotator validation process (two annotators working on the same dataset sequentially) then a third doing a quantitative scoring system. This was our best approach to getting a consensus multi-annotator ground truth. Such a process takes enormous effort on the part of the annotators particularly with such multi-scale structures.  Other metrics such as the  cl-dice which we also explored whilst interesting for preserving connectivity has a skeletonisation component to it which is also somewhat computationally intensive and can cause artefacts (ball like structures) at complicated junctions i. more than simple bifurcation or in the collapsed vessel structures. <br>\nThe metrics is indeed not a simple topic ,and what is the 'best' metric for multiscale vascular segmentations is, I believe an open and interesting research question in and of itself. I think this competition, shows this to be the case.   </p>",
      "votes": null,
      "replies": [
        {
          "id": 2561854,
          "author_name": "optimo",
          "author_url": "",
          "post_date": "12/14/2023 21:53:34",
          "content": "<p>Thanks for the clarifications <a href=\"https://www.kaggle.com/clairewalsh\" target=\"_blank\">@clairewalsh</a>, it is indeed an open problem and I'm glad to know you have thought about the different possibilities and chosen the best metric for your final applications.<br>\n<a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> since we have the right metric, could you please update the code with <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a>'s implementation as it is much faster and uses far less memory? This will be a big win for everyone in the competition!</p>",
          "votes": null,
          "replies": [
            {
              "id": 2561882,
              "author_name": "ryanholbrook",
              "author_url": "",
              "post_date": "12/14/2023 23:14:50",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a>,</p>\n<p>I am currently heads down on the (very much overdue) <a href=\"https://www.kaggle.com/discussions/competition-hosting/459827\" target=\"_blank\">Santa 2023</a> competition and won't have time to spare for a bit. I'll try to take a look at the new implementation, but it's unlikely I'll be able to make any changes until after the new year. Evaluation metrics in running competitions we need to treat as production code and changes thus require considerable care. (Breaking a competition over the holidays would be very bad for everyone.)</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2561917,
                  "author_name": "hengck23",
                  "author_url": "",
                  "post_date": "12/15/2023 00:36:05",
                  "content": "<p>\" (very much overdue) Santa 2023 \"</p>\n<p>is there still a santa 2023? i am still lokking forward it :)</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2562107,
                      "author_name": "something4kag",
                      "author_url": "",
                      "post_date": "12/15/2023 05:43:00",
                      "content": "<p>Indeed - see this post <a href=\"https://www.kaggle.com/discussions/competition-hosting/459827\" target=\"_blank\">Santa, we miss you (2023)</a>  <br>\ngood luck <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> with your hard work, Santa will surely have you on the nice list!</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2558807": "I understand that Surface Dice seemed better to focus only on the surface of predictions, which gives a similar weight to small vessels and large vessels.\n\nHowever, if computing the competition metrics on one kidney and emulating predictions as the mask with a simple dilation or erosion in 2D (for each slice independently) here are the results:\n>erosion score 0.6113542914390564\ndilation score 0.5309284329414368\nNum iterations: 2\nerosion score 0.18636780977249146\ndilation score 0.10018275678157806\nNum iterations: 3\nerosion score 0.08591027557849884\ndilation score 0.026479000225663185\n\nWith a small 2D dilation or erosion of a single pixel the competition score drops to 0.53 and 0.611 respectively. Things get worse when the number of dilations or erosions increases.\n\nThese sudden drops happen because the tolerance here has been set to 0, but are we confident that the annotations are correct at the pixel level ? As @hengck23 showed [here](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/456118#2543402) that labels can have shifts. It is very likely that two annotators would not set the exact same annotations at pixel level. I am actually quite surprised that models can beat the 1 pixel dilation or erosion benchmark, does it mean that we already have excellent models? Or that the metric used is not so good?\n\nI also don't understand why after 3 iterations, the score is not exactly 0 ? @junkoda do you understand why ?\n\n@clairewalsh don't you think that some level of tolerance would yield more accurate results to spot the best solution ?",
    "2559529": "I agree that surface Dice with tolerance zero is difficult to generalize to different resolutions. I just started and don't know if that's reasonable or unreasonable, though. Center-line Dice, in organizers review, seems to be more resolution independent, but I don't know any medical meaning of line vs surface. \n\nI don't know if the score should be mathematically zero. I guess random match of true positives can be possible.",
    "2560365": "agree. one of the worst metric I ever met in competitions.",
    "2561703": "Hi @optimo thanks for your insights into this metric. \nI know the metric has caused challenges throughout the competition so here is a bit about what and why we used it: \nWe spent a considerable amount of time in the lead up to this competition trying to find a metric that would provide a good measure the output that we are aiming for in terms of the end use of these types of segmentations.\nWith regard to tolerance, we did investigate how changing the tolerance changes the output. The challenge we found is that increasing the tolerance even to 1 for the surface dice, would give a high metric output (high 0.9), despite the very small vasculature being highly disconnected, something which causes fundamentally different results in downstream analyses such a flow modelling. \nRegarding your comment about different annotators having different annotations not within 1 pixel, this again is a challenge and is why we had the three annotator validation process (two annotators working on the same dataset sequentially) then a third doing a quantitative scoring system. This was our best approach to getting a consensus multi-annotator ground truth. Such a process takes enormous effort on the part of the annotators particularly with such multi-scale structures.  Other metrics such as the  cl-dice which we also explored whilst interesting for preserving connectivity has a skeletonisation component to it which is also somewhat computationally intensive and can cause artefacts (ball like structures) at complicated junctions i. more than simple bifurcation or in the collapsed vessel structures. \nThe metrics is indeed not a simple topic ,and what is the 'best' metric for multiscale vascular segmentations is, I believe an open and interesting research question in and of itself. I think this competition, shows this to be the case.",
    "2561854": "Thanks for the clarifications @clairewalsh, it is indeed an open problem and I'm glad to know you have thought about the different possibilities and chosen the best metric for your final applications.\n@ryanholbrook since we have the right metric, could you please update the code with @junkoda's implementation as it is much faster and uses far less memory? This will be a big win for everyone in the competition!",
    "2561882": "Hi @optimo,\n\nI am currently heads down on the (very much overdue) [Santa 2023](https://www.kaggle.com/discussions/competition-hosting/459827) competition and won't have time to spare for a bit. I'll try to take a look at the new implementation, but it's unlikely I'll be able to make any changes until after the new year. Evaluation metrics in running competitions we need to treat as production code and changes thus require considerable care. (Breaking a competition over the holidays would be very bad for everyone.)",
    "2561917": "\" (very much overdue) Santa 2023 \"\n\nis there still a santa 2023? i am still lokking forward it :)",
    "2562107": "Indeed - see this post [Santa, we miss you (2023)](https://www.kaggle.com/discussions/competition-hosting/459827)  \ngood luck @ryanholbrook with your hard work, Santa will surely have you on the nice list!"
  },
  "source": "meta"
}