{
  "id": 422192,
  "title": "How to calculate the metric?",
  "url": "/competitions/hubmap-hacking-the-human-vasculature/discussion/422192",
  "author_name": "",
  "post_date": "2023-07-08T17:14:08.620520600Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I'm stuck on understanding the metric of this comp, can we discuss it here?</p>\n<p>This sentence in the evaluation page confuses me:</p>\n<pre><code>Segmentation is calculated using IoU with a threshold of 0.6.\n</code></pre>\n<p>Does this mean that you don't actually calculate the precision-recall curve, but rather you use this threshold when calculating iou between prediction and target masks, after which you calculate the precision over a set of images?</p>\n<p>Below is a code snippet where I calculate the precision and recall for a given image. <br>\n<code>pred_masks</code> and <code>target_masks</code> are lists of prediction and target masks.</p>\n<pre><code>n_gts = \nn_true_positives = \nn_false_positives = \n\nmatched_gt_ids = []\nmatched_pred_ids = []\nfalse_pred_ids = []\n\n (target_masks) &gt; :\n     i  ((pred_masks)):\n        j = np.argmax(ious[i])  \n         ious[i, j] &gt;   j   matched_gt_ids  i   matched_pred_ids:\n            matched_gt_ids.append(j)\n            matched_pred_ids.append(i)\n        :\n            false_pred_ids.append(i)\n\n (target_masks) &gt; :\n    n_gts = (target_masks)\n    n_true_positives = (matched_gt_ids)\n    n_false_positives = (false_pred_ids)\n:\n    n_false_positives = (pred_masks)        \n\nprecision = n_true_positives / (n_true_positives + n_false_positives + )\nrecall = n_true_positives / (n_gts + )\n</code></pre>\n<p>Is this a correct implementation of the metric? Or does one need to calculate number of true positives, false positives and number of ground truth masks for all images and then calculate a global precision and recall?</p>\n<p>Thanks everyone! 🙏</p>",
  "messages": [
    {
      "id": "2335640",
      "postDate": "07/08/2023 17:14:08",
      "content": "<p>I'm stuck on understanding the metric of this comp, can we discuss it here?</p>\n<p>This sentence in the evaluation page confuses me:</p>\n<pre><code>Segmentation is calculated using IoU with a threshold of 0.6.\n</code></pre>\n<p>Does this mean that you don't actually calculate the precision-recall curve, but rather you use this threshold when calculating iou between prediction and target masks, after which you calculate the precision over a set of images?</p>\n<p>Below is a code snippet where I calculate the precision and recall for a given image. <br>\n<code>pred_masks</code> and <code>target_masks</code> are lists of prediction and target masks.</p>\n<pre><code>n_gts = \nn_true_positives = \nn_false_positives = \n\nmatched_gt_ids = []\nmatched_pred_ids = []\nfalse_pred_ids = []\n\n (target_masks) &gt; :\n     i  ((pred_masks)):\n        j = np.argmax(ious[i])  \n         ious[i, j] &gt;   j   matched_gt_ids  i   matched_pred_ids:\n            matched_gt_ids.append(j)\n            matched_pred_ids.append(i)\n        :\n            false_pred_ids.append(i)\n\n (target_masks) &gt; :\n    n_gts = (target_masks)\n    n_true_positives = (matched_gt_ids)\n    n_false_positives = (false_pred_ids)\n:\n    n_false_positives = (pred_masks)        \n\nprecision = n_true_positives / (n_true_positives + n_false_positives + )\nrecall = n_true_positives / (n_gts + )\n</code></pre>\n<p>Is this a correct implementation of the metric? Or does one need to calculate number of true positives, false positives and number of ground truth masks for all images and then calculate a global precision and recall?</p>\n<p>Thanks everyone! 🙏</p>",
      "rawMarkdown": "I'm stuck on understanding the metric of this comp, can we discuss it here?\n\nThis sentence in the evaluation page confuses me:\n```text\nSegmentation is calculated using IoU with a threshold of 0.6.\n```\n\nDoes this mean that you don't actually calculate the precision-recall curve, but rather you use this threshold when calculating iou between prediction and target masks, after which you calculate the precision over a set of images?\n\nBelow is a code snippet where I calculate the precision and recall for a given image. \n`pred_masks` and `target_masks` are lists of prediction and target masks.\n\n\n\n```python\n\n\nn_gts = 0\nn_true_positives = 0\nn_false_positives = 0\n\nmatched_gt_ids = []\nmatched_pred_ids = []\nfalse_pred_ids = []\n        \nif len(target_masks) > 0:\n    for i in range(len(pred_masks)):\n        j = np.argmax(ious[i])  \n        if ious[i, j] > 0.6 and j not in matched_gt_ids and i not in matched_pred_ids:\n            matched_gt_ids.append(j)\n            matched_pred_ids.append(i)\n        else:\n            false_pred_ids.append(i)\n\nif len(target_masks) > 0:\n    n_gts = len(target_masks)\n    n_true_positives = len(matched_gt_ids)\n    n_false_positives = len(false_pred_ids)\nelse:\n    n_false_positives = len(pred_masks)        \n    \nprecision = n_true_positives / (n_true_positives + n_false_positives + 1e-6)\nrecall = n_true_positives / (n_gts + 1e-6)\n\n```\n\n\nIs this a correct implementation of the metric? Or does one need to calculate number of true positives, false positives and number of ground truth masks for all images and then calculate a global precision and recall?\n\nThanks everyone! 🙏",
      "votes": null
    },
    {
      "id": "2335858",
      "postDate": "07/08/2023 21:45:26",
      "content": "<p>The following discussion may be helpful.</p>\n<p><a href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/415508\" target=\"_blank\">https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/415508</a><br>\n<a href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/414877\" target=\"_blank\">https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/414877</a><br>\n<a href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/418383\" target=\"_blank\">https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/418383</a></p>",
      "rawMarkdown": "The following discussion may be helpful.\n\nhttps://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/415508\nhttps://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/414877\nhttps://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/418383",
      "votes": null
    },
    {
      "id": "2335887",
      "postDate": "07/08/2023 23:03:58",
      "content": "<p>If I understand the method correctly, we should, (simplified):</p>\n<ul>\n<li>Sort all predicted instances from maximum to minimum by score</li>\n<li>Take one predicted instance (larger score first) and compare with GT masks</li>\n<li>If IOU &gt; 0.6 and GT is mask not used, considering as True and mark GT mask as 'used'. Otherwise as False.</li>\n</ul>\n<p>We then have to calculate the accumulated tp/fp and precision/recall for each point:</p>\n<pre><code>precision_list,recall_list = [],[]\n result  all_instances_matched_with_gt_masks:\n         result == \n            tp_acc += \n        :\n            fp_acc += \n\n        precision = tp_acc / (tp_acc + fp_acc)\n        recall = tp_acc / gt_total\n        precision_list.append(precision)\n        recall_list.append(recall)\n</code></pre>\n<p>And then we can calculate the AUC based on precision/recall.</p>\n<p>Please, correct me if I'm wrong.</p>\n<p>PS:<br>\nShould we apply interpolation, as mentioned in this <a href=\"https://kharshit.github.io/blog/2019/09/20/evaluation-metrics-for-object-detection-and-segmentation\" target=\"_blank\">https://kharshit.github.io/blog/2019/09/20/evaluation-metrics-for-object-detection-and-segmentation</a> article? </p>",
      "rawMarkdown": "If I understand the method correctly, we should, (simplified):\n* Sort all predicted instances from maximum to minimum by score\n* Take one predicted instance (larger score first) and compare with GT masks\n* If IOU > 0.6 and GT is mask not used, considering as True and mark GT mask as 'used'. Otherwise as False.\n\nWe then have to calculate the accumulated tp/fp and precision/recall for each point:\n```python\nprecision_list,recall_list = [],[]\nfor result in all_instances_matched_with_gt_masks:\n        if result == True\n            tp_acc += 1\n        else:\n            fp_acc += 1\n\n        precision = tp_acc / (tp_acc + fp_acc)\n        recall = tp_acc / gt_total\n        precision_list.append(precision)\n        recall_list.append(recall)\n```\nAnd then we can calculate the AUC based on precision/recall.\n\nPlease, correct me if I'm wrong.\n\nPS:\nShould we apply interpolation, as mentioned in this https://kharshit.github.io/blog/2019/09/20/evaluation-metrics-for-object-detection-and-segmentation article?",
      "votes": null
    },
    {
      "id": "2337056",
      "postDate": "07/09/2023 20:32:03",
      "content": "<p>That's an interesting take! <br>\nDid I understand correctly, <code>all_instances_matched_with_gt_masks</code> is an array that contains true/false values depending on whether your model predicts <em>all</em> the blood vessel masks in a single image? <br>\nWhen you calculate the metric the local CV, how well it compares with the public LB?</p>",
      "rawMarkdown": "That's an interesting take! \nDid I understand correctly, `all_instances_matched_with_gt_masks` is an array that contains true/false values depending on whether your model predicts *all* the blood vessel masks in a single image? \nWhen you calculate the metric the local CV, how well it compares with the public LB?",
      "votes": null
    },
    {
      "id": "2337065",
      "postDate": "07/09/2023 20:46:41",
      "content": "<p>No, all_instances_matched_with_gt_masks is an array that contains True/False values of all predicted instances across all images in test set sorted by confidence score. Highest score first. One value - one instance. Sorry if the name of the variable misleads you.</p>",
      "rawMarkdown": "No, all_instances_matched_with_gt_masks is an array that contains True/False values of all predicted instances across all images in test set sorted by confidence score. Highest score first. One value - one instance. Sorry if the name of the variable misleads you.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2335858,
      "author_name": "mitsuyasuhoshino",
      "author_url": "",
      "post_date": "07/08/2023 21:45:26",
      "content": "<p>The following discussion may be helpful.</p>\n<p><a href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/415508\" target=\"_blank\">https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/415508</a><br>\n<a href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/414877\" target=\"_blank\">https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/414877</a><br>\n<a href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/418383\" target=\"_blank\">https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/418383</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2335887,
      "author_name": "tsobolev",
      "author_url": "",
      "post_date": "07/08/2023 23:03:58",
      "content": "<p>If I understand the method correctly, we should, (simplified):</p>\n<ul>\n<li>Sort all predicted instances from maximum to minimum by score</li>\n<li>Take one predicted instance (larger score first) and compare with GT masks</li>\n<li>If IOU &gt; 0.6 and GT is mask not used, considering as True and mark GT mask as 'used'. Otherwise as False.</li>\n</ul>\n<p>We then have to calculate the accumulated tp/fp and precision/recall for each point:</p>\n<pre><code>precision_list,recall_list = [],[]\n result  all_instances_matched_with_gt_masks:\n         result == \n            tp_acc += \n        :\n            fp_acc += \n\n        precision = tp_acc / (tp_acc + fp_acc)\n        recall = tp_acc / gt_total\n        precision_list.append(precision)\n        recall_list.append(recall)\n</code></pre>\n<p>And then we can calculate the AUC based on precision/recall.</p>\n<p>Please, correct me if I'm wrong.</p>\n<p>PS:<br>\nShould we apply interpolation, as mentioned in this <a href=\"https://kharshit.github.io/blog/2019/09/20/evaluation-metrics-for-object-detection-and-segmentation\" target=\"_blank\">https://kharshit.github.io/blog/2019/09/20/evaluation-metrics-for-object-detection-and-segmentation</a> article? </p>",
      "votes": null,
      "replies": [
        {
          "id": 2337056,
          "author_name": "viktorcikojevic",
          "author_url": "",
          "post_date": "07/09/2023 20:32:03",
          "content": "<p>That's an interesting take! <br>\nDid I understand correctly, <code>all_instances_matched_with_gt_masks</code> is an array that contains true/false values depending on whether your model predicts <em>all</em> the blood vessel masks in a single image? <br>\nWhen you calculate the metric the local CV, how well it compares with the public LB?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2337065,
              "author_name": "tsobolev",
              "author_url": "",
              "post_date": "07/09/2023 20:46:41",
              "content": "<p>No, all_instances_matched_with_gt_masks is an array that contains True/False values of all predicted instances across all images in test set sorted by confidence score. Highest score first. One value - one instance. Sorry if the name of the variable misleads you.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2335640": "I'm stuck on understanding the metric of this comp, can we discuss it here?\n\nThis sentence in the evaluation page confuses me:\n```text\nSegmentation is calculated using IoU with a threshold of 0.6.\n```\n\nDoes this mean that you don't actually calculate the precision-recall curve, but rather you use this threshold when calculating iou between prediction and target masks, after which you calculate the precision over a set of images?\n\nBelow is a code snippet where I calculate the precision and recall for a given image. \n`pred_masks` and `target_masks` are lists of prediction and target masks.\n\n\n\n```python\n\n\nn_gts = 0\nn_true_positives = 0\nn_false_positives = 0\n\nmatched_gt_ids = []\nmatched_pred_ids = []\nfalse_pred_ids = []\n        \nif len(target_masks) > 0:\n    for i in range(len(pred_masks)):\n        j = np.argmax(ious[i])  \n        if ious[i, j] > 0.6 and j not in matched_gt_ids and i not in matched_pred_ids:\n            matched_gt_ids.append(j)\n            matched_pred_ids.append(i)\n        else:\n            false_pred_ids.append(i)\n\nif len(target_masks) > 0:\n    n_gts = len(target_masks)\n    n_true_positives = len(matched_gt_ids)\n    n_false_positives = len(false_pred_ids)\nelse:\n    n_false_positives = len(pred_masks)        \n    \nprecision = n_true_positives / (n_true_positives + n_false_positives + 1e-6)\nrecall = n_true_positives / (n_gts + 1e-6)\n\n```\n\n\nIs this a correct implementation of the metric? Or does one need to calculate number of true positives, false positives and number of ground truth masks for all images and then calculate a global precision and recall?\n\nThanks everyone! 🙏",
    "2335858": "The following discussion may be helpful.\n\nhttps://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/415508\nhttps://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/414877\nhttps://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/418383",
    "2335887": "If I understand the method correctly, we should, (simplified):\n* Sort all predicted instances from maximum to minimum by score\n* Take one predicted instance (larger score first) and compare with GT masks\n* If IOU > 0.6 and GT is mask not used, considering as True and mark GT mask as 'used'. Otherwise as False.\n\nWe then have to calculate the accumulated tp/fp and precision/recall for each point:\n```python\nprecision_list,recall_list = [],[]\nfor result in all_instances_matched_with_gt_masks:\n        if result == True\n            tp_acc += 1\n        else:\n            fp_acc += 1\n\n        precision = tp_acc / (tp_acc + fp_acc)\n        recall = tp_acc / gt_total\n        precision_list.append(precision)\n        recall_list.append(recall)\n```\nAnd then we can calculate the AUC based on precision/recall.\n\nPlease, correct me if I'm wrong.\n\nPS:\nShould we apply interpolation, as mentioned in this https://kharshit.github.io/blog/2019/09/20/evaluation-metrics-for-object-detection-and-segmentation article?",
    "2337056": "That's an interesting take! \nDid I understand correctly, `all_instances_matched_with_gt_masks` is an array that contains true/false values depending on whether your model predicts *all* the blood vessel masks in a single image? \nWhen you calculate the metric the local CV, how well it compares with the public LB?",
    "2337065": "No, all_instances_matched_with_gt_masks is an array that contains True/False values of all predicted instances across all images in test set sorted by confidence score. Highest score first. One value - one instance. Sorry if the name of the variable misleads you."
  },
  "source": "meta"
}