{
  "id": 456254,
  "title": "Surface Dice Metric: NaN for empty ground truth???",
  "url": "/competitions/blood-vessel-segmentation/discussion/456254",
  "author_name": "",
  "post_date": "2023-11-19T00:46:34.091061500Z",
  "votes": 10,
  "comment_count": 6,
  "views": 0,
  "content": "<h1>EDIT:</h1>\n<h1>The evaluation occurs in 3D, over the entire kidney, not over the individual 2D slices. There shouldn't be a case when the ground truth is empty.</h1>\n<p>(see comment from kaggle organizer below)</p>\n<hr>\n<p>It is not very clear on how the competition metric handles cases for images when ground truth is empty (i.e. no mask).</p>\n<p>the default surface_distance code gives NaN</p>\n<pre><code> surface_distance  *\n\n\n\nmask_gt = np.asarray(\n,\n,\n,\n,\n,\n,\n    ],\n    dtype=)\n\nmask_pred = np.asarray(\n,\n,\n,\n,\n,\n,\n    ],\n    dtype=)\n\nvertical = \nhorizontal = \ndiag =  * math.sqrt(horizontal **  + vertical ** )\nsurface_distances = compute_surface_distances(\n    mask_gt, mask_pred, spacing_mm=(vertical, horizontal))\n#print(surface_distances)\nsurface_dice = compute_surface_dice_at_tolerance(surface_distances, tolerance_mm=)\nprint(surface_dice)\n\n# NaN  printed\n</code></pre>",
  "messages": [
    {
      "id": "2530280",
      "postDate": "11/19/2023 00:46:34",
      "content": "<h1>EDIT:</h1>\n<h1>The evaluation occurs in 3D, over the entire kidney, not over the individual 2D slices. There shouldn't be a case when the ground truth is empty.</h1>\n<p>(see comment from kaggle organizer below)</p>\n<hr>\n<p>It is not very clear on how the competition metric handles cases for images when ground truth is empty (i.e. no mask).</p>\n<p>the default surface_distance code gives NaN</p>\n<pre><code> surface_distance  *\n\n\n\nmask_gt = np.asarray(\n,\n,\n,\n,\n,\n,\n    ],\n    dtype=)\n\nmask_pred = np.asarray(\n,\n,\n,\n,\n,\n,\n    ],\n    dtype=)\n\nvertical = \nhorizontal = \ndiag =  * math.sqrt(horizontal **  + vertical ** )\nsurface_distances = compute_surface_distances(\n    mask_gt, mask_pred, spacing_mm=(vertical, horizontal))\n#print(surface_distances)\nsurface_dice = compute_surface_dice_at_tolerance(surface_distances, tolerance_mm=)\nprint(surface_dice)\n\n# NaN  printed\n</code></pre>",
      "rawMarkdown": "# EDIT:\n# The evaluation occurs in 3D, over the entire kidney, not over the individual 2D slices. There shouldn't be a case when the ground truth is empty.\n(see comment from kaggle organizer below)\n\n\n---\n\n\nIt is not very clear on how the competition metric handles cases for images when ground truth is empty (i.e. no mask).\n\nthe default surface_distance code gives NaN\n\n```\n\n\nfrom surface_distance import *\n\n\n\nmask_gt = np.asarray(\n\t[\n\t\t[0, 0, 0, 0, 0, 0],\n\t\t[0, 0, 0, 0, 0, 0],\n\t\t[0, 0, 0, 0, 0, 0],\n\t\t[0, 0, 0, 0, 0, 0],\n\t\t[0, 0, 0, 0, 0, 0],\n\t\t[0, 0, 0, 0, 0, 0],\n\t],\n\tdtype=bool)\n\nmask_pred = np.asarray(\n\t[\n\t\t[0, 0, 0, 0, 0, 0],\n\t\t[0, 0, 0, 0, 0, 0],\n\t\t[0, 0, 0, 0, 0, 0],\n\t\t[0, 0, 0, 0, 0, 0],\n\t\t[0, 0, 0, 0, 0, 0],\n\t\t[0, 0, 0, 0, 0, 0],\n\t],\n\tdtype=bool)\n\nvertical = 1\nhorizontal = 1\ndiag = 0.5 * math.sqrt(horizontal ** 2 + vertical ** 2)\nsurface_distances = compute_surface_distances(\n\tmask_gt, mask_pred, spacing_mm=(vertical, horizontal))\n#print(surface_distances)\nsurface_dice = compute_surface_dice_at_tolerance(surface_distances, tolerance_mm=0)\nprint(surface_dice)\n\n# NaN is printed\n```",
      "votes": null
    },
    {
      "id": "2530315",
      "postDate": "11/19/2023 02:16:33",
      "content": "<p>This is probably the reason why there is now a nan_to_num function happening in the metric, which is the cause of countless people suffering from \"Submission Scoring Error\" when their submissions were running fine before the metric was updated to handle this NaN output<br>\nIs there an alternative way you have in mind to calculate so that we do not need nan_to_num</p>",
      "rawMarkdown": "This is probably the reason why there is now a nan_to_num function happening in the metric, which is the cause of countless people suffering from \"Submission Scoring Error\" when their submissions were running fine before the metric was updated to handle this NaN output\nIs there an alternative way you have in mind to calculate so that we do not need nan_to_num",
      "votes": null
    },
    {
      "id": "2530317",
      "postDate": "11/19/2023 02:35:51",
      "content": "<p>If the surface metric is calculated based on the data frame, then the <code>rle_encode</code> and <code>rle_decode</code> functions will handle this issue </p>\n<p>It assigns a negligible value to the prediction if the whole array is zero</p>\n<p>based on the mask_pred matrix, the result of:</p>\n<pre><code>mask_pred = np.asarray(\n    [\n        [, , , , , ],\n        [, , , , , ],\n        [, , , , , ],\n        [, , , , , ],\n        [, , , , , ],\n        [, , , , , ],\n    ],\n    dtype=)\n\nrle_decode(rle_encode(mask_pred), mask_pred.shape)\n\n\n\n\n\n\n\n</code></pre>",
      "rawMarkdown": "If the surface metric is calculated based on the data frame, then the `rle_encode` and `rle_decode` functions will handle this issue \n\nIt assigns a negligible value to the prediction if the whole array is zero\n\nbased on the mask_pred matrix, the result of:\n\n```python\nmask_pred = np.asarray(\n    [\n        [0, 0, 0, 0, 0, 0],\n        [0, 0, 0, 0, 0, 0],\n        [0, 0, 0, 0, 0, 0],\n        [0, 0, 0, 0, 0, 0],\n        [0, 0, 0, 0, 0, 0],\n        [0, 0, 0, 0, 0, 0],\n    ],\n    dtype=bool)\n\nrle_decode(rle_encode(mask_pred), mask_pred.shape)\n\n# [[1, 0, 0, 0, 0, 0],\n# [0, 0, 0, 0, 0, 0],\n# [0, 0, 0, 0, 0, 0],\n# [0, 0, 0, 0, 0, 0],\n# [0, 0, 0, 0, 0, 0],\n# [0, 0, 0, 0, 0, 0]]\n```",
      "votes": null
    },
    {
      "id": "2530359",
      "postDate": "11/19/2023 05:10:57",
      "content": "<p>One problem is most people use : <code>arr/np.max(arr)</code> as normalization. If the max is 0, you will get nans you should use <code>arr/(np.max(arr) + epsilon)</code> where epsilon is a small number ie: 1e-6 … You shouldn't have much problem during training (because frameworks mitigate that problem using tricks) but it will destroy your submission…</p>",
      "rawMarkdown": "One problem is most people use : `arr/np.max(arr)` as normalization. If the max is 0, you will get nans you should use `arr/(np.max(arr) + epsilon)` where epsilon is a small number ie: 1e-6 ... You shouldn't have much problem during training (because frameworks mitigate that problem using tricks) but it will destroy your submission...",
      "votes": null
    },
    {
      "id": "2530379",
      "postDate": "11/19/2023 05:31:49",
      "content": "<p>Most of your nans will come from your post-processings steps.</p>",
      "rawMarkdown": "Most of your nans will come from your post-processings steps.",
      "votes": null
    },
    {
      "id": "2531723",
      "postDate": "11/20/2023 13:13:55",
      "content": "<p>The evaluation occurs in 3D, over the entire kidney, not over the individual 2D slices. There shouldn't be a case when the ground truth is empty.</p>",
      "rawMarkdown": "The evaluation occurs in 3D, over the entire kidney, not over the individual 2D slices. There shouldn't be a case when the ground truth is empty.",
      "votes": null
    },
    {
      "id": "2599635",
      "postDate": "01/13/2024 03:10:06",
      "content": "<p>Does it mean that the evaluation score notebook isn't the one that is being used for evaluation for this competition? I think the current \"score\" function is corresponding to the 2D calculation (grouping happens to \"id\" column, which is 2D slice, and surface dice is calculated for each slice then incremented)</p>",
      "rawMarkdown": "Does it mean that the evaluation score notebook isn't the one that is being used for evaluation for this competition? I think the current \"score\" function is corresponding to the 2D calculation (grouping happens to \"id\" column, which is 2D slice, and surface dice is calculated for each slice then incremented)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2530315,
      "author_name": "harshitsheoran",
      "author_url": "",
      "post_date": "11/19/2023 02:16:33",
      "content": "<p>This is probably the reason why there is now a nan_to_num function happening in the metric, which is the cause of countless people suffering from \"Submission Scoring Error\" when their submissions were running fine before the metric was updated to handle this NaN output<br>\nIs there an alternative way you have in mind to calculate so that we do not need nan_to_num</p>",
      "votes": null,
      "replies": [
        {
          "id": 2530359,
          "author_name": "abdrah",
          "author_url": "",
          "post_date": "11/19/2023 05:10:57",
          "content": "<p>One problem is most people use : <code>arr/np.max(arr)</code> as normalization. If the max is 0, you will get nans you should use <code>arr/(np.max(arr) + epsilon)</code> where epsilon is a small number ie: 1e-6 … You shouldn't have much problem during training (because frameworks mitigate that problem using tricks) but it will destroy your submission…</p>",
          "votes": null,
          "replies": [
            {
              "id": 2530379,
              "author_name": "abdrah",
              "author_url": "",
              "post_date": "11/19/2023 05:31:49",
              "content": "<p>Most of your nans will come from your post-processings steps.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2530317,
      "author_name": "dhinkris",
      "author_url": "",
      "post_date": "11/19/2023 02:35:51",
      "content": "<p>If the surface metric is calculated based on the data frame, then the <code>rle_encode</code> and <code>rle_decode</code> functions will handle this issue </p>\n<p>It assigns a negligible value to the prediction if the whole array is zero</p>\n<p>based on the mask_pred matrix, the result of:</p>\n<pre><code>mask_pred = np.asarray(\n    [\n        [, , , , , ],\n        [, , , , , ],\n        [, , , , , ],\n        [, , , , , ],\n        [, , , , , ],\n        [, , , , , ],\n    ],\n    dtype=)\n\nrle_decode(rle_encode(mask_pred), mask_pred.shape)\n\n\n\n\n\n\n\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2531723,
      "author_name": "ryanholbrook",
      "author_url": "",
      "post_date": "11/20/2023 13:13:55",
      "content": "<p>The evaluation occurs in 3D, over the entire kidney, not over the individual 2D slices. There shouldn't be a case when the ground truth is empty.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2599635,
          "author_name": "juhha1",
          "author_url": "",
          "post_date": "01/13/2024 03:10:06",
          "content": "<p>Does it mean that the evaluation score notebook isn't the one that is being used for evaluation for this competition? I think the current \"score\" function is corresponding to the 2D calculation (grouping happens to \"id\" column, which is 2D slice, and surface dice is calculated for each slice then incremented)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2530280": "# EDIT:\n# The evaluation occurs in 3D, over the entire kidney, not over the individual 2D slices. There shouldn't be a case when the ground truth is empty.\n(see comment from kaggle organizer below)\n\n\n---\n\n\nIt is not very clear on how the competition metric handles cases for images when ground truth is empty (i.e. no mask).\n\nthe default surface_distance code gives NaN\n\n```\n\n\nfrom surface_distance import *\n\n\n\nmask_gt = np.asarray(\n\t[\n\t\t[0, 0, 0, 0, 0, 0],\n\t\t[0, 0, 0, 0, 0, 0],\n\t\t[0, 0, 0, 0, 0, 0],\n\t\t[0, 0, 0, 0, 0, 0],\n\t\t[0, 0, 0, 0, 0, 0],\n\t\t[0, 0, 0, 0, 0, 0],\n\t],\n\tdtype=bool)\n\nmask_pred = np.asarray(\n\t[\n\t\t[0, 0, 0, 0, 0, 0],\n\t\t[0, 0, 0, 0, 0, 0],\n\t\t[0, 0, 0, 0, 0, 0],\n\t\t[0, 0, 0, 0, 0, 0],\n\t\t[0, 0, 0, 0, 0, 0],\n\t\t[0, 0, 0, 0, 0, 0],\n\t],\n\tdtype=bool)\n\nvertical = 1\nhorizontal = 1\ndiag = 0.5 * math.sqrt(horizontal ** 2 + vertical ** 2)\nsurface_distances = compute_surface_distances(\n\tmask_gt, mask_pred, spacing_mm=(vertical, horizontal))\n#print(surface_distances)\nsurface_dice = compute_surface_dice_at_tolerance(surface_distances, tolerance_mm=0)\nprint(surface_dice)\n\n# NaN is printed\n```",
    "2530315": "This is probably the reason why there is now a nan_to_num function happening in the metric, which is the cause of countless people suffering from \"Submission Scoring Error\" when their submissions were running fine before the metric was updated to handle this NaN output\nIs there an alternative way you have in mind to calculate so that we do not need nan_to_num",
    "2530317": "If the surface metric is calculated based on the data frame, then the `rle_encode` and `rle_decode` functions will handle this issue \n\nIt assigns a negligible value to the prediction if the whole array is zero\n\nbased on the mask_pred matrix, the result of:\n\n```python\nmask_pred = np.asarray(\n    [\n        [0, 0, 0, 0, 0, 0],\n        [0, 0, 0, 0, 0, 0],\n        [0, 0, 0, 0, 0, 0],\n        [0, 0, 0, 0, 0, 0],\n        [0, 0, 0, 0, 0, 0],\n        [0, 0, 0, 0, 0, 0],\n    ],\n    dtype=bool)\n\nrle_decode(rle_encode(mask_pred), mask_pred.shape)\n\n# [[1, 0, 0, 0, 0, 0],\n# [0, 0, 0, 0, 0, 0],\n# [0, 0, 0, 0, 0, 0],\n# [0, 0, 0, 0, 0, 0],\n# [0, 0, 0, 0, 0, 0],\n# [0, 0, 0, 0, 0, 0]]\n```",
    "2530359": "One problem is most people use : `arr/np.max(arr)` as normalization. If the max is 0, you will get nans you should use `arr/(np.max(arr) + epsilon)` where epsilon is a small number ie: 1e-6 ... You shouldn't have much problem during training (because frameworks mitigate that problem using tricks) but it will destroy your submission...",
    "2530379": "Most of your nans will come from your post-processings steps.",
    "2531723": "The evaluation occurs in 3D, over the entire kidney, not over the individual 2D slices. There shouldn't be a case when the ground truth is empty.",
    "2599635": "Does it mean that the evaluation score notebook isn't the one that is being used for evaluation for this competition? I think the current \"score\" function is corresponding to the 2D calculation (grouping happens to \"id\" column, which is 2D slice, and surface dice is calculated for each slice then incremented)"
  },
  "source": "meta"
}