{
  "id": 414877,
  "title": "Confidence Prediction in evaluation score",
  "url": "/competitions/hubmap-hacking-the-human-vasculature/discussion/414877",
  "author_name": "",
  "post_date": "2023-06-03T19:40:13.259716900Z",
  "votes": 5,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I have read the open images evaluation metric but still can't figure how would box prediction confidence in the prediction string impact the overall score. If you have any information in this regard please advise.</p>",
  "messages": [
    {
      "id": "2286807",
      "postDate": "06/03/2023 19:40:13",
      "content": "<p>I have read the open images evaluation metric but still can't figure how would box prediction confidence in the prediction string impact the overall score. If you have any information in this regard please advise.</p>",
      "rawMarkdown": "I have read the open images evaluation metric but still can't figure how would box prediction confidence in the prediction string impact the overall score. If you have any information in this regard please advise.",
      "votes": null
    },
    {
      "id": "2287368",
      "postDate": "06/04/2023 11:35:03",
      "content": "<p>pls take a look at this explanation of AP calculation (open images metric in general is the same but with some small changes) - <a href=\"https://www.v7labs.com/blog/mean-average-precision#:~:text=Mean%20Average%20Precision(mAP)%20is%20a%20metric%20used%20to%20evaluate,values%20from%200%20to%201\" target=\"_blank\">https://www.v7labs.com/blog/mean-average-precision#:~:text=Mean%20Average%20Precision(mAP)%20is%20a%20metric%20used%20to%20evaluate,values%20from%200%20to%201</a>.</p>",
      "rawMarkdown": "pls take a look at this explanation of AP calculation (open images metric in general is the same but with some small changes) - https://www.v7labs.com/blog/mean-average-precision#:~:text=Mean%20Average%20Precision(mAP)%20is%20a%20metric%20used%20to%20evaluate,values%20from%200%20to%201.",
      "votes": null
    },
    {
      "id": "2287409",
      "postDate": "06/04/2023 12:28:59",
      "content": "<p>I like to think of the basic case below for some intuition when you are doing AP calculation. There are a few stages. </p>\n<hr>\n<ol>\n<li><p>Choose an IoU (you will perhaps choose a number of them, then you repeat the calculations below and average). Let us suppose it is some value 0&lt;I&lt;1.   This is often denoted as mAP@I.  </p></li>\n<li><p>Plot precision-recall graph. </p></li>\n</ol>\n<p>Let us begin in the case <strong><em>without any confidence or IoU.</em></strong> </p>\n<p>--- The case <em>without confidence or IoU</em> --- <br>\nSuppose you had no confidence attached. Your model did the following: <br>\nInputs: an image, whose ground truth has 10 blood vessels. <br>\nOutputs:  </p>\n<ul>\n<li>9 blood vessels polygons. </li>\n</ul>\n<p>The key point: <br>\nWe have to say what it means for a polygon to be \"(un)successful prediction\". Let us suppose (un)successful means your prediction polygon intersection has (no) an intersection with the of some blood vessel in ground truth. We will see later, we use the confidence values to define what is successful. I'm sure you know the term for this is just True Positive. </p>\n<p>For example: we may deduce</p>\n<ul>\n<li><p><strong><em>successful/True Positive (TP)</em></strong> 5 blood vessels polygons</p></li>\n<li><p><strong><em>unsuccessful/False Positive (FP)</em></strong> 4 blood vessels (FP=False Positive)  polygons</p></li>\n</ul>\n<p>Then: <br>\nPrecision=TP/(TP+FP)=5/(5+4) ~56%<br>\nRecall =TP/(TP+FN)=  5/10 =50%</p>\n<p>Note one can think of Recall as <br>\n number of predictions/number of vessels in the ground truth. </p>\n<hr>\n<p>--- The case <strong><em>with confidence and IoU threshold:=I</em></strong> ---- <br>\nIn this case  you consider the confidence attached. Then we plot Precision-recall(PR) as a *function of choices of confidence threshold. *</p>\n<p>Note how I italicized \"successfully, unsuccessfully\" in the case without confidence. <br>\nHere are your<br>\nOutputs:  </p>\n<ul>\n<li>9 polygons for blood vessels with confidence (0.9,0.87,0.8,0.7,0.7,0.3,0.2,0.1)</li>\n</ul>\n<p>Now suppose we have the following IoUs (one way we can define this is say, for each predicted polygon, pick the largest IoU with a polygon of blood vessel in the ground truth)  (0.8,0.7,0.6,0.5,0.5,0.9,0.1,0.5,0.1)</p>\n<p>For a choice of confidence threshold, C, a prediction is considered <br>\nTP :: confidence &gt;C and IoU &gt;= I<br>\nFP :: confidence &gt;C but IoU&lt; I </p>\n<p>Now we can compute PR for each choice of confidence. <br>\nPR for 0.9: (TP,FP)=(1,0): Precision=100%, Recall = 10%<br>\nPR for 0.1:(TP,FP,FN)= (7,2) : Precision ~ 78%, Recall ~ 70%</p>\n<p>The last step is to get an \"averaged\" value by considering a weighted sum. (Other ways to encode all of this data  could be F1 and AUC(area under curve), but AP seems to be the convention these days.)</p>\n<hr>\n<p>Observe if the confidence threshold gets smaller; generally, your prediction  precision decreases (since you make more predictions), and your recall increases (you potentially detect more predictions)</p>\n<p>Hope this helps! </p>",
      "rawMarkdown": "I like to think of the basic case below for some intuition when you are doing AP calculation. There are a few stages. \n\n--- \n1. Choose an IoU (you will perhaps choose a number of them, then you repeat the calculations below and average). Let us suppose it is some value 0<I<1.   This is often denoted as mAP@I.  \n\n2. Plot precision-recall graph. \n\nLet us begin in the case ***without any confidence or IoU.*** \n\n--- The case *without confidence or IoU* --- \nSuppose you had no confidence attached. Your model did the following: \nInputs: an image, whose ground truth has 10 blood vessels. \nOutputs:  \n- 9 blood vessels polygons. \n\n\nThe key point: \nWe have to say what it means for a polygon to be \"(un)successful prediction\". Let us suppose (un)successful means your prediction polygon intersection has (no) an intersection with the of some blood vessel in ground truth. We will see later, we use the confidence values to define what is successful. I'm sure you know the term for this is just True Positive. \n\nFor example: we may deduce\n\n-  ***successful/True Positive (TP)*** 5 blood vessels polygons\n\n-  ***unsuccessful/False Positive (FP)*** 4 blood vessels (FP=False Positive)  polygons\n\n\nThen: \nPrecision=TP/(TP+FP)=5/(5+4) ~56%\nRecall =TP/(TP+FN)=  5/10 =50%\n\nNote one can think of Recall as \n number of predictions/number of vessels in the ground truth. \n\n--- \n--- The case ***with confidence and IoU threshold:=I*** ---- \nIn this case  you consider the confidence attached. Then we plot Precision-recall(PR) as a *function of choices of confidence threshold. *\n\nNote how I italicized \"successfully, unsuccessfully\" in the case without confidence. \nHere are your\nOutputs:  \n- 9 polygons for blood vessels with confidence (0.9,0.87,0.8,0.7,0.7,0.3,0.2,0.1)\n\n\nNow suppose we have the following IoUs (one way we can define this is say, for each predicted polygon, pick the largest IoU with a polygon of blood vessel in the ground truth)  (0.8,0.7,0.6,0.5,0.5,0.9,0.1,0.5,0.1)\n\n\nFor a choice of confidence threshold, C, a prediction is considered \nTP :: confidence >C and IoU >= I\nFP :: confidence >C but IoU< I \n\nNow we can compute PR for each choice of confidence. \nPR for 0.9: (TP,FP)=(1,0): Precision=100%, Recall = 10%\nPR for 0.1:(TP,FP,FN)= (7,2) : Precision ~ 78%, Recall ~ 70%\n\nThe last step is to get an \"averaged\" value by considering a weighted sum. (Other ways to encode all of this data  could be F1 and AUC(area under curve), but AP seems to be the convention these days.)\n \n--- \nObserve if the confidence threshold gets smaller; generally, your prediction  precision decreases (since you make more predictions), and your recall increases (you potentially detect more predictions)\n\nHope this helps!",
      "votes": null
    },
    {
      "id": "2293714",
      "postDate": "06/09/2023 12:27:50",
      "content": "<p><a href=\"https://www.kaggle.com/eatalittlel\" target=\"_blank\">@eatalittlel</a> Thank you so much, so what you mean is confidence is incorporated the same way iou threshold is? or this is just some sort of analogy?</p>",
      "rawMarkdown": "eatalittlel Thank you so much, so what you mean is confidence is incorporated the same way iou threshold is? or this is just some sort of analogy?",
      "votes": null
    },
    {
      "id": "2293881",
      "postDate": "06/09/2023 14:56:40",
      "content": "<p>Yes! It is not an analogy, but is exactly what IoU and confidence are doing: they are two types of thresholds for saying what makes a classification successful or unsuccessful.  (Hence allowing one to compute recall and precision)</p>",
      "rawMarkdown": "Yes! It is not an analogy, but is exactly what IoU and confidence are doing: they are two types of thresholds for saying what makes a classification successful or unsuccessful.  (Hence allowing one to compute recall and precision)",
      "votes": null
    },
    {
      "id": "2294026",
      "postDate": "06/09/2023 17:03:15",
      "content": "<p>I found this article helpful: <a href=\"url\" target=\"_blank\">https://kharshit.github.io/blog/2019/09/20/evaluation-metrics-for-object-detection-and-segmentation</a></p>",
      "rawMarkdown": "I found this article helpful: [https://kharshit.github.io/blog/2019/09/20/evaluation-metrics-for-object-detection-and-segmentation](url)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2287368,
      "author_name": "maksimovka",
      "author_url": "",
      "post_date": "06/04/2023 11:35:03",
      "content": "<p>pls take a look at this explanation of AP calculation (open images metric in general is the same but with some small changes) - <a href=\"https://www.v7labs.com/blog/mean-average-precision#:~:text=Mean%20Average%20Precision(mAP)%20is%20a%20metric%20used%20to%20evaluate,values%20from%200%20to%201\" target=\"_blank\">https://www.v7labs.com/blog/mean-average-precision#:~:text=Mean%20Average%20Precision(mAP)%20is%20a%20metric%20used%20to%20evaluate,values%20from%200%20to%201</a>.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2287409,
      "author_name": "eatalittlel",
      "author_url": "",
      "post_date": "06/04/2023 12:28:59",
      "content": "<p>I like to think of the basic case below for some intuition when you are doing AP calculation. There are a few stages. </p>\n<hr>\n<ol>\n<li><p>Choose an IoU (you will perhaps choose a number of them, then you repeat the calculations below and average). Let us suppose it is some value 0&lt;I&lt;1.   This is often denoted as mAP@I.  </p></li>\n<li><p>Plot precision-recall graph. </p></li>\n</ol>\n<p>Let us begin in the case <strong><em>without any confidence or IoU.</em></strong> </p>\n<p>--- The case <em>without confidence or IoU</em> --- <br>\nSuppose you had no confidence attached. Your model did the following: <br>\nInputs: an image, whose ground truth has 10 blood vessels. <br>\nOutputs:  </p>\n<ul>\n<li>9 blood vessels polygons. </li>\n</ul>\n<p>The key point: <br>\nWe have to say what it means for a polygon to be \"(un)successful prediction\". Let us suppose (un)successful means your prediction polygon intersection has (no) an intersection with the of some blood vessel in ground truth. We will see later, we use the confidence values to define what is successful. I'm sure you know the term for this is just True Positive. </p>\n<p>For example: we may deduce</p>\n<ul>\n<li><p><strong><em>successful/True Positive (TP)</em></strong> 5 blood vessels polygons</p></li>\n<li><p><strong><em>unsuccessful/False Positive (FP)</em></strong> 4 blood vessels (FP=False Positive)  polygons</p></li>\n</ul>\n<p>Then: <br>\nPrecision=TP/(TP+FP)=5/(5+4) ~56%<br>\nRecall =TP/(TP+FN)=  5/10 =50%</p>\n<p>Note one can think of Recall as <br>\n number of predictions/number of vessels in the ground truth. </p>\n<hr>\n<p>--- The case <strong><em>with confidence and IoU threshold:=I</em></strong> ---- <br>\nIn this case  you consider the confidence attached. Then we plot Precision-recall(PR) as a *function of choices of confidence threshold. *</p>\n<p>Note how I italicized \"successfully, unsuccessfully\" in the case without confidence. <br>\nHere are your<br>\nOutputs:  </p>\n<ul>\n<li>9 polygons for blood vessels with confidence (0.9,0.87,0.8,0.7,0.7,0.3,0.2,0.1)</li>\n</ul>\n<p>Now suppose we have the following IoUs (one way we can define this is say, for each predicted polygon, pick the largest IoU with a polygon of blood vessel in the ground truth)  (0.8,0.7,0.6,0.5,0.5,0.9,0.1,0.5,0.1)</p>\n<p>For a choice of confidence threshold, C, a prediction is considered <br>\nTP :: confidence &gt;C and IoU &gt;= I<br>\nFP :: confidence &gt;C but IoU&lt; I </p>\n<p>Now we can compute PR for each choice of confidence. <br>\nPR for 0.9: (TP,FP)=(1,0): Precision=100%, Recall = 10%<br>\nPR for 0.1:(TP,FP,FN)= (7,2) : Precision ~ 78%, Recall ~ 70%</p>\n<p>The last step is to get an \"averaged\" value by considering a weighted sum. (Other ways to encode all of this data  could be F1 and AUC(area under curve), but AP seems to be the convention these days.)</p>\n<hr>\n<p>Observe if the confidence threshold gets smaller; generally, your prediction  precision decreases (since you make more predictions), and your recall increases (you potentially detect more predictions)</p>\n<p>Hope this helps! </p>",
      "votes": null,
      "replies": [
        {
          "id": 2293714,
          "author_name": "mhmdsab",
          "author_url": "",
          "post_date": "06/09/2023 12:27:50",
          "content": "<p><a href=\"https://www.kaggle.com/eatalittlel\" target=\"_blank\">@eatalittlel</a> Thank you so much, so what you mean is confidence is incorporated the same way iou threshold is? or this is just some sort of analogy?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2293881,
              "author_name": "eatalittlel",
              "author_url": "",
              "post_date": "06/09/2023 14:56:40",
              "content": "<p>Yes! It is not an analogy, but is exactly what IoU and confidence are doing: they are two types of thresholds for saying what makes a classification successful or unsuccessful.  (Hence allowing one to compute recall and precision)</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2294026,
      "author_name": "tsobolev",
      "author_url": "",
      "post_date": "06/09/2023 17:03:15",
      "content": "<p>I found this article helpful: <a href=\"url\" target=\"_blank\">https://kharshit.github.io/blog/2019/09/20/evaluation-metrics-for-object-detection-and-segmentation</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2286807": "I have read the open images evaluation metric but still can't figure how would box prediction confidence in the prediction string impact the overall score. If you have any information in this regard please advise.",
    "2287368": "pls take a look at this explanation of AP calculation (open images metric in general is the same but with some small changes) - https://www.v7labs.com/blog/mean-average-precision#:~:text=Mean%20Average%20Precision(mAP)%20is%20a%20metric%20used%20to%20evaluate,values%20from%200%20to%201.",
    "2287409": "I like to think of the basic case below for some intuition when you are doing AP calculation. There are a few stages. \n\n--- \n1. Choose an IoU (you will perhaps choose a number of them, then you repeat the calculations below and average). Let us suppose it is some value 0<I<1.   This is often denoted as mAP@I.  \n\n2. Plot precision-recall graph. \n\nLet us begin in the case ***without any confidence or IoU.*** \n\n--- The case *without confidence or IoU* --- \nSuppose you had no confidence attached. Your model did the following: \nInputs: an image, whose ground truth has 10 blood vessels. \nOutputs:  \n- 9 blood vessels polygons. \n\n\nThe key point: \nWe have to say what it means for a polygon to be \"(un)successful prediction\". Let us suppose (un)successful means your prediction polygon intersection has (no) an intersection with the of some blood vessel in ground truth. We will see later, we use the confidence values to define what is successful. I'm sure you know the term for this is just True Positive. \n\nFor example: we may deduce\n\n-  ***successful/True Positive (TP)*** 5 blood vessels polygons\n\n-  ***unsuccessful/False Positive (FP)*** 4 blood vessels (FP=False Positive)  polygons\n\n\nThen: \nPrecision=TP/(TP+FP)=5/(5+4) ~56%\nRecall =TP/(TP+FN)=  5/10 =50%\n\nNote one can think of Recall as \n number of predictions/number of vessels in the ground truth. \n\n--- \n--- The case ***with confidence and IoU threshold:=I*** ---- \nIn this case  you consider the confidence attached. Then we plot Precision-recall(PR) as a *function of choices of confidence threshold. *\n\nNote how I italicized \"successfully, unsuccessfully\" in the case without confidence. \nHere are your\nOutputs:  \n- 9 polygons for blood vessels with confidence (0.9,0.87,0.8,0.7,0.7,0.3,0.2,0.1)\n\n\nNow suppose we have the following IoUs (one way we can define this is say, for each predicted polygon, pick the largest IoU with a polygon of blood vessel in the ground truth)  (0.8,0.7,0.6,0.5,0.5,0.9,0.1,0.5,0.1)\n\n\nFor a choice of confidence threshold, C, a prediction is considered \nTP :: confidence >C and IoU >= I\nFP :: confidence >C but IoU< I \n\nNow we can compute PR for each choice of confidence. \nPR for 0.9: (TP,FP)=(1,0): Precision=100%, Recall = 10%\nPR for 0.1:(TP,FP,FN)= (7,2) : Precision ~ 78%, Recall ~ 70%\n\nThe last step is to get an \"averaged\" value by considering a weighted sum. (Other ways to encode all of this data  could be F1 and AUC(area under curve), but AP seems to be the convention these days.)\n \n--- \nObserve if the confidence threshold gets smaller; generally, your prediction  precision decreases (since you make more predictions), and your recall increases (you potentially detect more predictions)\n\nHope this helps!",
    "2293714": "eatalittlel Thank you so much, so what you mean is confidence is incorporated the same way iou threshold is? or this is just some sort of analogy?",
    "2293881": "Yes! It is not an analogy, but is exactly what IoU and confidence are doing: they are two types of thresholds for saying what makes a classification successful or unsuccessful.  (Hence allowing one to compute recall and precision)",
    "2294026": "I found this article helpful: [https://kharshit.github.io/blog/2019/09/20/evaluation-metrics-for-object-detection-and-segmentation](url)"
  },
  "source": "meta"
}