{
  "id": 455062,
  "title": "Could the Implementation of the Surface-Dice Metric lead to the Submission Problem?",
  "url": "/competitions/blood-vessel-segmentation/discussion/455062",
  "author_name": "",
  "post_date": "2023-11-13T09:04:08.276849500Z",
  "votes": 3,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Many people have noted the issue of a Submission- / Scoring-error when submitting a dummy prediction, such as '1 0', which corresponds to an empty mask, see for example <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/455001\" target=\"_blank\">this post by Harshit</a> or <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/455008\" target=\"_blank\">this post by Djata</a>.  </p>\n<p>I have taken a look at the <a href=\"https://www.kaggle.com/code/metric/surface-dice-metric/notebook\" target=\"_blank\">implementation of the surface dice metric</a> linked in the competition overview and it seems that the formula they use is $$\\text{DICE}(y,y_{pred})=\\frac{y\\cdot y_{pred}}{y+y_{pred}}$$ (applied componentwise, then summed over all entries). </p>\n<p>However, if both $y$ and $y_{pred}$ are $0$, then we may face an issue with this metric, since the expression evaluates to $0/0$. Typically, you'd add a \"smoothing parameter\" $\\epsilon$ and compute the metric via $$\\text{DICE}(y,y_{pred})=\\frac{y\\cdot y_{pred}+\\epsilon}{y+y_{pred}+\\epsilon}$$ to avoid division by $0$. </p>\n<p>I'm not 100% certain that this is the issue, I just wanted to throw it into the discussion to hear what you guys think. </p>\n<p>Here's the relevant snippet of the code linked above: </p>\n<blockquote>\n  <p>distances_gt_to_pred = surface_distances[\"distances_gt_to_pred\"]<br>\n      distances_pred_to_gt = surface_distances[\"distances_pred_to_gt\"]<br>\n      surfel_areas_gt = surface_distances[\"surfel_areas_gt\"]<br>\n      surfel_areas_pred = surface_distances[\"surfel_areas_pred\"]<br>\n      overlap_gt = np.sum(surfel_areas_gt[distances_gt_to_pred &lt;= tolerance_mm])<br>\n      overlap_pred = np.sum(<br>\n          surfel_areas_pred[distances_pred_to_gt &lt;= tolerance_mm])<br>\n      surface_dice = (overlap_gt + overlap_pred) / (np.sum(surfel_areas_gt) +<br>\n                                                    np.sum(surfel_areas_pred))<br>\n      return surface_dice</p>\n</blockquote>",
  "messages": [
    {
      "id": "2523108",
      "postDate": "11/13/2023 09:04:08",
      "content": "<p>Many people have noted the issue of a Submission- / Scoring-error when submitting a dummy prediction, such as '1 0', which corresponds to an empty mask, see for example <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/455001\" target=\"_blank\">this post by Harshit</a> or <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/455008\" target=\"_blank\">this post by Djata</a>.  </p>\n<p>I have taken a look at the <a href=\"https://www.kaggle.com/code/metric/surface-dice-metric/notebook\" target=\"_blank\">implementation of the surface dice metric</a> linked in the competition overview and it seems that the formula they use is $$\\text{DICE}(y,y_{pred})=\\frac{y\\cdot y_{pred}}{y+y_{pred}}$$ (applied componentwise, then summed over all entries). </p>\n<p>However, if both $y$ and $y_{pred}$ are $0$, then we may face an issue with this metric, since the expression evaluates to $0/0$. Typically, you'd add a \"smoothing parameter\" $\\epsilon$ and compute the metric via $$\\text{DICE}(y,y_{pred})=\\frac{y\\cdot y_{pred}+\\epsilon}{y+y_{pred}+\\epsilon}$$ to avoid division by $0$. </p>\n<p>I'm not 100% certain that this is the issue, I just wanted to throw it into the discussion to hear what you guys think. </p>\n<p>Here's the relevant snippet of the code linked above: </p>\n<blockquote>\n  <p>distances_gt_to_pred = surface_distances[\"distances_gt_to_pred\"]<br>\n      distances_pred_to_gt = surface_distances[\"distances_pred_to_gt\"]<br>\n      surfel_areas_gt = surface_distances[\"surfel_areas_gt\"]<br>\n      surfel_areas_pred = surface_distances[\"surfel_areas_pred\"]<br>\n      overlap_gt = np.sum(surfel_areas_gt[distances_gt_to_pred &lt;= tolerance_mm])<br>\n      overlap_pred = np.sum(<br>\n          surfel_areas_pred[distances_pred_to_gt &lt;= tolerance_mm])<br>\n      surface_dice = (overlap_gt + overlap_pred) / (np.sum(surfel_areas_gt) +<br>\n                                                    np.sum(surfel_areas_pred))<br>\n      return surface_dice</p>\n</blockquote>",
      "rawMarkdown": "Many people have noted the issue of a Submission- / Scoring-error when submitting a dummy prediction, such as '1 0', which corresponds to an empty mask, see for example [this post by Harshit](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/455001) or [this post by Djata](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/455008).  \n\nI have taken a look at the [implementation of the surface dice metric](https://www.kaggle.com/code/metric/surface-dice-metric/notebook) linked in the competition overview and it seems that the formula they use is $$\\text{DICE}(y,y_{pred})=\\frac{y\\cdot y_{pred}}{y+y_{pred}}$$ (applied componentwise, then summed over all entries). \n\nHowever, if both $y$ and $y_{pred}$ are $0$, then we may face an issue with this metric, since the expression evaluates to $0/0$. Typically, you'd add a \"smoothing parameter\" $\\epsilon$ and compute the metric via $$\\text{DICE}(y,y_{pred})=\\frac{y\\cdot y_{pred}+\\epsilon}{y+y_{pred}+\\epsilon}$$ to avoid division by $0$. \n\nI'm not 100% certain that this is the issue, I just wanted to throw it into the discussion to hear what you guys think. \n\nHere's the relevant snippet of the code linked above: \n\n>    distances_gt_to_pred = surface_distances[\"distances_gt_to_pred\"]\n    distances_pred_to_gt = surface_distances[\"distances_pred_to_gt\"]\n    surfel_areas_gt = surface_distances[\"surfel_areas_gt\"]\n    surfel_areas_pred = surface_distances[\"surfel_areas_pred\"]\n    overlap_gt = np.sum(surfel_areas_gt[distances_gt_to_pred <= tolerance_mm])\n    overlap_pred = np.sum(\n        surfel_areas_pred[distances_pred_to_gt <= tolerance_mm])\n    surface_dice = (overlap_gt + overlap_pred) / (np.sum(surfel_areas_gt) +\n                                                  np.sum(surfel_areas_pred))\n    return surface_dice",
      "votes": null
    },
    {
      "id": "2529363",
      "postDate": "11/18/2023 06:55:50",
      "content": "<p>The implementation of the Surface-Dice metric in this competition could potentially lead to submission challenges, particularly due to its high sensitivity to precise segmentation of intricate vascular structures. Given the complexity of 3D Hierarchical Phase-Contrast Tomography (HiP-CT) scans of human kidneys, any minor inaccuracies in segmenting tiny blood vessels can significantly impact the metric, thereby affecting the evaluation of submissions. Additionally, the need for high precision in segmentation underlines the importance of developing robust models that can handle the variability inherent in the provided datasets.</p>",
      "rawMarkdown": "The implementation of the Surface-Dice metric in this competition could potentially lead to submission challenges, particularly due to its high sensitivity to precise segmentation of intricate vascular structures. Given the complexity of 3D Hierarchical Phase-Contrast Tomography (HiP-CT) scans of human kidneys, any minor inaccuracies in segmenting tiny blood vessels can significantly impact the metric, thereby affecting the evaluation of submissions. Additionally, the need for high precision in segmentation underlines the importance of developing robust models that can handle the variability inherent in the provided datasets.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2529363,
      "author_name": "dhinkris",
      "author_url": "",
      "post_date": "11/18/2023 06:55:50",
      "content": "<p>The implementation of the Surface-Dice metric in this competition could potentially lead to submission challenges, particularly due to its high sensitivity to precise segmentation of intricate vascular structures. Given the complexity of 3D Hierarchical Phase-Contrast Tomography (HiP-CT) scans of human kidneys, any minor inaccuracies in segmenting tiny blood vessels can significantly impact the metric, thereby affecting the evaluation of submissions. Additionally, the need for high precision in segmentation underlines the importance of developing robust models that can handle the variability inherent in the provided datasets.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2523108": "Many people have noted the issue of a Submission- / Scoring-error when submitting a dummy prediction, such as '1 0', which corresponds to an empty mask, see for example [this post by Harshit](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/455001) or [this post by Djata](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/455008).  \n\nI have taken a look at the [implementation of the surface dice metric](https://www.kaggle.com/code/metric/surface-dice-metric/notebook) linked in the competition overview and it seems that the formula they use is $$\\text{DICE}(y,y_{pred})=\\frac{y\\cdot y_{pred}}{y+y_{pred}}$$ (applied componentwise, then summed over all entries). \n\nHowever, if both $y$ and $y_{pred}$ are $0$, then we may face an issue with this metric, since the expression evaluates to $0/0$. Typically, you'd add a \"smoothing parameter\" $\\epsilon$ and compute the metric via $$\\text{DICE}(y,y_{pred})=\\frac{y\\cdot y_{pred}+\\epsilon}{y+y_{pred}+\\epsilon}$$ to avoid division by $0$. \n\nI'm not 100% certain that this is the issue, I just wanted to throw it into the discussion to hear what you guys think. \n\nHere's the relevant snippet of the code linked above: \n\n>    distances_gt_to_pred = surface_distances[\"distances_gt_to_pred\"]\n    distances_pred_to_gt = surface_distances[\"distances_pred_to_gt\"]\n    surfel_areas_gt = surface_distances[\"surfel_areas_gt\"]\n    surfel_areas_pred = surface_distances[\"surfel_areas_pred\"]\n    overlap_gt = np.sum(surfel_areas_gt[distances_gt_to_pred <= tolerance_mm])\n    overlap_pred = np.sum(\n        surfel_areas_pred[distances_pred_to_gt <= tolerance_mm])\n    surface_dice = (overlap_gt + overlap_pred) / (np.sum(surfel_areas_gt) +\n                                                  np.sum(surfel_areas_pred))\n    return surface_dice",
    "2529363": "The implementation of the Surface-Dice metric in this competition could potentially lead to submission challenges, particularly due to its high sensitivity to precise segmentation of intricate vascular structures. Given the complexity of 3D Hierarchical Phase-Contrast Tomography (HiP-CT) scans of human kidneys, any minor inaccuracies in segmenting tiny blood vessels can significantly impact the metric, thereby affecting the evaluation of submissions. Additionally, the need for high precision in segmentation underlines the importance of developing robust models that can handle the variability inherent in the provided datasets."
  },
  "source": "meta"
}