{
  "id": 121499,
  "title": "Understanding the mAP Evaluation Metric for Object Detection",
  "url": "/competitions/pku-autonomous-driving/discussion/121499",
  "author_name": "",
  "post_date": "2019-12-13T17:46:37.322265500Z",
  "votes": 26,
  "comment_count": 11,
  "views": 0,
  "content": "<h1>Evaluating Object Detectors</h1>\n\n<p>In object detection, evaluation is non trivial, because there are two distinct tasks to measure:\n1. Determining whether an object exists in the image (classification)\n2. Determining the location of the object (localization, a regression task).</p>\n\n<p>Furthermore, in a typical data set there will be many classes and their distribution is non-uniform (for example there might be many more dogs than ice cream cones). So a simple accuracy-based metric will introduce biases. It is also important to assess the risk of misclassifications. Thus, there is the need to associate a “confidence score” or model score with each bounding box detected and to assess the model at various level of confidence.</p>\n\n<p>In order to address these needs, the Average Precision (AP) was introduced. To understand the AP, it is necessary to understand the precision and recall of a classifier. For a more comprehensive explanation of these terms, the <strong><a href=\"https://en.wikipedia.org/wiki/Precision_and_recall\">wikipedia article</a></strong> is a nice place to start. Briefly, in this context, precision measures the “false positive rate” or the ratio of true object detections to the total number of objects that the classifier predicted. If you have a precision score of close to 1.0 then there is a high likelihood that whatever the classifier predicts as a positive detection is in fact a correct prediction. Recall measures the “false negative rate” or the ratio of true object detections to the total number of objects in the data set. If you have a recall score close to 1.0 then almost all objects that are in your dataset will be positively detected by the model. Finally, it is very important to note that the there is an inverse relationship between precision and recall and that these metrics are dependent on the model score threshold that you set (as well as of course, the quality of the model).</p>\n\n<p>To calculate the AP, for a specific class (say a “person”) the precision-recall curve is computed from the model’s detection output, by varying the model score threshold that determines what is counted as a model-predicted positive detection of the class. An example precision-recall curve may look something like this for a given classifier:</p>\n\n<p></p>\n\n<p>The final step to calculating the AP score is to take the average value of the precision across all recall values</p>\n\n<h1>Localization and Intersection over Union</h1>\n\n<p>In order to evaluate the model on the task of object localization, we must first determine how well the model predicted the location of the object. Usually, this is done by drawing a bounding box around the object of interest, but in some cases it is an N-sided polygon or even pixel by pixel segmentation. For all of these cases, the localization task is typically evaluated on the Intersection over Union threshold (IoU). For definiteness, throughout the rest of the article, I’ll assume that the model predicts bounding boxes, but almost everything said will also apply to pixel-wise segmentation or N-sided polygons. Many good explanations of IoU exist, (see <a href=\"https://www.pyimagesearch.com/2016/11/07/intersection-over-union-iou-for-object-detection/\">this one </a>for example), but the basic idea is that it summarizes how well the ground truth object overlaps the object boundary predicted by the model.</p>\n\n<p>Model object detections are determined to be true or false depending upon the IoU threshold. This IoU threshold(s) for each competition vary, but in the <a href=\"http://cocodataset.org/#detections-eval\">COCO challenge</a>, for example, 10 different IoU thresholds are considered, from 0.5 to 0.95 in steps of 0.05. For a specific object (say, ‘person’) this is what the precision-recall curves may look like when calculated at the different IoU thresholds of the COCO challenge:</p>\n\n<p></p>\n\n<p><strong>Now that we’ve defined Average Precision (AP) and seen how the IoU threshold affects it, the mean Average Precision or mAP score is calculated by taking the mean AP over all classes and/or over all IoU thresholds, depending on the competition</strong></p>\n\n<p><a href=\"https://gist.github.com/tarlen5/008809c3decf19313de216b9208f3734\"># Code for Calculating the mean Average Precision</a></p>\n\n<p>Ref : <a href=\"https://medium.com/@timothycarlen/understanding-the-map-evaluation-metric-for-object-detection-a07fe6962cf3\">Link </a></p>",
  "messages": [
    {
      "id": "694498",
      "postDate": "12/13/2019 17:46:37",
      "content": "<h1>Evaluating Object Detectors</h1>\n\n<p>In object detection, evaluation is non trivial, because there are two distinct tasks to measure:\n1. Determining whether an object exists in the image (classification)\n2. Determining the location of the object (localization, a regression task).</p>\n\n<p>Furthermore, in a typical data set there will be many classes and their distribution is non-uniform (for example there might be many more dogs than ice cream cones). So a simple accuracy-based metric will introduce biases. It is also important to assess the risk of misclassifications. Thus, there is the need to associate a “confidence score” or model score with each bounding box detected and to assess the model at various level of confidence.</p>\n\n<p>In order to address these needs, the Average Precision (AP) was introduced. To understand the AP, it is necessary to understand the precision and recall of a classifier. For a more comprehensive explanation of these terms, the <strong><a href=\"https://en.wikipedia.org/wiki/Precision_and_recall\">wikipedia article</a></strong> is a nice place to start. Briefly, in this context, precision measures the “false positive rate” or the ratio of true object detections to the total number of objects that the classifier predicted. If you have a precision score of close to 1.0 then there is a high likelihood that whatever the classifier predicts as a positive detection is in fact a correct prediction. Recall measures the “false negative rate” or the ratio of true object detections to the total number of objects in the data set. If you have a recall score close to 1.0 then almost all objects that are in your dataset will be positively detected by the model. Finally, it is very important to note that the there is an inverse relationship between precision and recall and that these metrics are dependent on the model score threshold that you set (as well as of course, the quality of the model).</p>\n\n<p>To calculate the AP, for a specific class (say a “person”) the precision-recall curve is computed from the model’s detection output, by varying the model score threshold that determines what is counted as a model-predicted positive detection of the class. An example precision-recall curve may look something like this for a given classifier:</p>\n\n<p></p>\n\n<p>The final step to calculating the AP score is to take the average value of the precision across all recall values</p>\n\n<h1>Localization and Intersection over Union</h1>\n\n<p>In order to evaluate the model on the task of object localization, we must first determine how well the model predicted the location of the object. Usually, this is done by drawing a bounding box around the object of interest, but in some cases it is an N-sided polygon or even pixel by pixel segmentation. For all of these cases, the localization task is typically evaluated on the Intersection over Union threshold (IoU). For definiteness, throughout the rest of the article, I’ll assume that the model predicts bounding boxes, but almost everything said will also apply to pixel-wise segmentation or N-sided polygons. Many good explanations of IoU exist, (see <a href=\"https://www.pyimagesearch.com/2016/11/07/intersection-over-union-iou-for-object-detection/\">this one </a>for example), but the basic idea is that it summarizes how well the ground truth object overlaps the object boundary predicted by the model.</p>\n\n<p>Model object detections are determined to be true or false depending upon the IoU threshold. This IoU threshold(s) for each competition vary, but in the <a href=\"http://cocodataset.org/#detections-eval\">COCO challenge</a>, for example, 10 different IoU thresholds are considered, from 0.5 to 0.95 in steps of 0.05. For a specific object (say, ‘person’) this is what the precision-recall curves may look like when calculated at the different IoU thresholds of the COCO challenge:</p>\n\n<p></p>\n\n<p><strong>Now that we’ve defined Average Precision (AP) and seen how the IoU threshold affects it, the mean Average Precision or mAP score is calculated by taking the mean AP over all classes and/or over all IoU thresholds, depending on the competition</strong></p>\n\n<p><a href=\"https://gist.github.com/tarlen5/008809c3decf19313de216b9208f3734\"># Code for Calculating the mean Average Precision</a></p>\n\n<p>Ref : <a href=\"https://medium.com/@timothycarlen/understanding-the-map-evaluation-metric-for-object-detection-a07fe6962cf3\">Link </a></p>",
      "rawMarkdown": "# Evaluating Object Detectors\nIn object detection, evaluation is non trivial, because there are two distinct tasks to measure:\n1. Determining whether an object exists in the image (classification)\n2. Determining the location of the object (localization, a regression task).\n\nFurthermore, in a typical data set there will be many classes and their distribution is non-uniform (for example there might be many more dogs than ice cream cones). So a simple accuracy-based metric will introduce biases. It is also important to assess the risk of misclassifications. Thus, there is the need to associate a “confidence score” or model score with each bounding box detected and to assess the model at various level of confidence.\n\nIn order to address these needs, the Average Precision (AP) was introduced. To understand the AP, it is necessary to understand the precision and recall of a classifier. For a more comprehensive explanation of these terms, the **[wikipedia article](https://en.wikipedia.org/wiki/Precision_and_recall)** is a nice place to start. Briefly, in this context, precision measures the “false positive rate” or the ratio of true object detections to the total number of objects that the classifier predicted. If you have a precision score of close to 1.0 then there is a high likelihood that whatever the classifier predicts as a positive detection is in fact a correct prediction. Recall measures the “false negative rate” or the ratio of true object detections to the total number of objects in the data set. If you have a recall score close to 1.0 then almost all objects that are in your dataset will be positively detected by the model. Finally, it is very important to note that the there is an inverse relationship between precision and recall and that these metrics are dependent on the model score threshold that you set (as well as of course, the quality of the model).\n\nTo calculate the AP, for a specific class (say a “person”) the precision-recall curve is computed from the model’s detection output, by varying the model score threshold that determines what is counted as a model-predicted positive detection of the class. An example precision-recall curve may look something like this for a given classifier:\n\n![](https://miro.medium.com/max/1094/1*ceYw4cXV4BtB2s7S9elExg.png)\n\nThe final step to calculating the AP score is to take the average value of the precision across all recall values\n\n# Localization and Intersection over Union\n\nIn order to evaluate the model on the task of object localization, we must first determine how well the model predicted the location of the object. Usually, this is done by drawing a bounding box around the object of interest, but in some cases it is an N-sided polygon or even pixel by pixel segmentation. For all of these cases, the localization task is typically evaluated on the Intersection over Union threshold (IoU). For definiteness, throughout the rest of the article, I’ll assume that the model predicts bounding boxes, but almost everything said will also apply to pixel-wise segmentation or N-sided polygons. Many good explanations of IoU exist, (see [this one ](https://www.pyimagesearch.com/2016/11/07/intersection-over-union-iou-for-object-detection/)for example), but the basic idea is that it summarizes how well the ground truth object overlaps the object boundary predicted by the model.\n\nModel object detections are determined to be true or false depending upon the IoU threshold. This IoU threshold(s) for each competition vary, but in the [COCO challenge](http://cocodataset.org/#detections-eval), for example, 10 different IoU thresholds are considered, from 0.5 to 0.95 in steps of 0.05. For a specific object (say, ‘person’) this is what the precision-recall curves may look like when calculated at the different IoU thresholds of the COCO challenge:\n\n![](https://miro.medium.com/max/1140/1*kmAon_ut7ZU3h0XkGbqR6g.png)\n\n**Now that we’ve defined Average Precision (AP) and seen how the IoU threshold affects it, the mean Average Precision or mAP score is calculated by taking the mean AP over all classes and/or over all IoU thresholds, depending on the competition**\n\n[# Code for Calculating the mean Average Precision](https://gist.github.com/tarlen5/008809c3decf19313de216b9208f3734)\n\nRef : [Link ](https://medium.com/@timothycarlen/understanding-the-map-evaluation-metric-for-object-detection-a07fe6962cf3)",
      "votes": null
    },
    {
      "id": "694740",
      "postDate": "12/14/2019 02:51:47",
      "content": "<p>Excellent topic. Unfortunately I have only 1 to give. </p>",
      "rawMarkdown": "Excellent topic. Unfortunately I have only 1 to give.",
      "votes": null
    },
    {
      "id": "694742",
      "postDate": "12/14/2019 02:53:14",
      "content": "<p>Great Explanation <a href=\"/mobassir\">@mobassir</a> </p>",
      "rawMarkdown": "Great Explanation @mobassir",
      "votes": null
    },
    {
      "id": "694756",
      "postDate": "12/14/2019 03:18:56",
      "content": "<p>thanks dear <a href=\"/veeralakrishna\">@veeralakrishna</a> </p>",
      "rawMarkdown": "thanks dear @veeralakrishna",
      "votes": null
    },
    {
      "id": "694757",
      "postDate": "12/14/2019 03:19:10",
      "content": "<p>hahahha thank you <a href=\"/mpwolke\">@mpwolke</a> </p>",
      "rawMarkdown": "hahahha thank you @mpwolke",
      "votes": null
    },
    {
      "id": "694913",
      "postDate": "12/14/2019 10:46:04",
      "content": "<p>Great. Thanks for sharing <a href=\"/mobassir\">@mobassir</a>... Could see that you have done a lot of research on this topic. Excellent work..... </p>",
      "rawMarkdown": "Great. Thanks for sharing @mobassir... Could see that you have done a lot of research on this topic. Excellent work.....",
      "votes": null
    },
    {
      "id": "694914",
      "postDate": "12/14/2019 10:46:59",
      "content": "<p>Thank you <a href=\"/manojprabhaakr\">@manojprabhaakr</a> :)</p>",
      "rawMarkdown": "Thank you @manojprabhaakr :)",
      "votes": null
    },
    {
      "id": "695386",
      "postDate": "12/15/2019 03:36:49",
      "content": "<p>Good deep understanding👍 </p>",
      "rawMarkdown": "Good deep understanding👍",
      "votes": null
    },
    {
      "id": "695418",
      "postDate": "12/15/2019 05:46:00",
      "content": "<p>Nice one ! very comprehensive explanation. Those One liners for Precision-Recall are perfect.</p>",
      "rawMarkdown": "Nice one ! very comprehensive explanation. Those One liners for Precision-Recall are perfect.",
      "votes": null
    },
    {
      "id": "696128",
      "postDate": "12/16/2019 06:43:34",
      "content": "<p>Hi! Thanks for your explaination! I am confused about how does false negative play a role in mAP. Also, if two bounding boxes are accurate enough to be consider as true positives for one car, how is mAP like now.</p>",
      "rawMarkdown": "Hi! Thanks for your explaination! I am confused about how does false negative play a role in mAP. Also, if two bounding boxes are accurate enough to be consider as true positives for one car, how is mAP like now.",
      "votes": null
    },
    {
      "id": "696136",
      "postDate": "12/16/2019 07:03:15",
      "content": "<p>hi <a href=\"/tonychenxyz\">@tonychenxyz</a> \n if two bounding boxes are accurate enough to be consider as true positives for one car then recall score will be high.\nwe calculate  Average Precision (AP) before calculating mAP\nand inside AP calculation we get scores like fp,fn \nRecall measures the “false negative rate” or the ratio of true object detections to the total number of objects in the data set. If you have a recall score close to 1.0 then almost all objects that are in your dataset will be positively detected by the model.</p>",
      "rawMarkdown": "hi @tonychenxyz \n if two bounding boxes are accurate enough to be consider as true positives for one car then recall score will be high.\nwe calculate  Average Precision (AP) before calculating mAP\nand inside AP calculation we get scores like fp,fn \nRecall measures the “false negative rate” or the ratio of true object detections to the total number of objects in the data set. If you have a recall score close to 1.0 then almost all objects that are in your dataset will be positively detected by the model.",
      "votes": null
    },
    {
      "id": "696246",
      "postDate": "12/16/2019 10:36:54",
      "content": "<p>Agree ! had a similar thought in mind.\n<strong>Precision</strong> : Measure of Exactness ( Quality)\n<strong>Recall</strong> : Measure of Completeness ( Quantity)</p>",
      "rawMarkdown": "Agree ! had a similar thought in mind.\n**Precision** : Measure of Exactness ( Quality)\n**Recall** : Measure of Completeness ( Quantity)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 694740,
      "author_name": "mpwolke",
      "author_url": "",
      "post_date": "12/14/2019 02:51:47",
      "content": "<p>Excellent topic. Unfortunately I have only 1 to give. </p>",
      "votes": null,
      "replies": [
        {
          "id": 694757,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "12/14/2019 03:19:10",
          "content": "<p>hahahha thank you <a href=\"/mpwolke\">@mpwolke</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 694742,
      "author_name": "veeralakrishna",
      "author_url": "",
      "post_date": "12/14/2019 02:53:14",
      "content": "<p>Great Explanation <a href=\"/mobassir\">@mobassir</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 694756,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "12/14/2019 03:18:56",
          "content": "<p>thanks dear <a href=\"/veeralakrishna\">@veeralakrishna</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 694913,
      "author_name": "manojprabhaakr",
      "author_url": "",
      "post_date": "12/14/2019 10:46:04",
      "content": "<p>Great. Thanks for sharing <a href=\"/mobassir\">@mobassir</a>... Could see that you have done a lot of research on this topic. Excellent work..... </p>",
      "votes": null,
      "replies": [
        {
          "id": 694914,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "12/14/2019 10:46:59",
          "content": "<p>Thank you <a href=\"/manojprabhaakr\">@manojprabhaakr</a> :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 695386,
      "author_name": "diegojohnson",
      "author_url": "",
      "post_date": "12/15/2019 03:36:49",
      "content": "<p>Good deep understanding👍 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 695418,
      "author_name": "atulanandjha",
      "author_url": "",
      "post_date": "12/15/2019 05:46:00",
      "content": "<p>Nice one ! very comprehensive explanation. Those One liners for Precision-Recall are perfect.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 696128,
      "author_name": "tonychenxyz",
      "author_url": "",
      "post_date": "12/16/2019 06:43:34",
      "content": "<p>Hi! Thanks for your explaination! I am confused about how does false negative play a role in mAP. Also, if two bounding boxes are accurate enough to be consider as true positives for one car, how is mAP like now.</p>",
      "votes": null,
      "replies": [
        {
          "id": 696136,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "12/16/2019 07:03:15",
          "content": "<p>hi <a href=\"/tonychenxyz\">@tonychenxyz</a> \n if two bounding boxes are accurate enough to be consider as true positives for one car then recall score will be high.\nwe calculate  Average Precision (AP) before calculating mAP\nand inside AP calculation we get scores like fp,fn \nRecall measures the “false negative rate” or the ratio of true object detections to the total number of objects in the data set. If you have a recall score close to 1.0 then almost all objects that are in your dataset will be positively detected by the model.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 696246,
          "author_name": "atulanandjha",
          "author_url": "",
          "post_date": "12/16/2019 10:36:54",
          "content": "<p>Agree ! had a similar thought in mind.\n<strong>Precision</strong> : Measure of Exactness ( Quality)\n<strong>Recall</strong> : Measure of Completeness ( Quantity)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "694498": "# Evaluating Object Detectors\nIn object detection, evaluation is non trivial, because there are two distinct tasks to measure:\n1. Determining whether an object exists in the image (classification)\n2. Determining the location of the object (localization, a regression task).\n\nFurthermore, in a typical data set there will be many classes and their distribution is non-uniform (for example there might be many more dogs than ice cream cones). So a simple accuracy-based metric will introduce biases. It is also important to assess the risk of misclassifications. Thus, there is the need to associate a “confidence score” or model score with each bounding box detected and to assess the model at various level of confidence.\n\nIn order to address these needs, the Average Precision (AP) was introduced. To understand the AP, it is necessary to understand the precision and recall of a classifier. For a more comprehensive explanation of these terms, the **[wikipedia article](https://en.wikipedia.org/wiki/Precision_and_recall)** is a nice place to start. Briefly, in this context, precision measures the “false positive rate” or the ratio of true object detections to the total number of objects that the classifier predicted. If you have a precision score of close to 1.0 then there is a high likelihood that whatever the classifier predicts as a positive detection is in fact a correct prediction. Recall measures the “false negative rate” or the ratio of true object detections to the total number of objects in the data set. If you have a recall score close to 1.0 then almost all objects that are in your dataset will be positively detected by the model. Finally, it is very important to note that the there is an inverse relationship between precision and recall and that these metrics are dependent on the model score threshold that you set (as well as of course, the quality of the model).\n\nTo calculate the AP, for a specific class (say a “person”) the precision-recall curve is computed from the model’s detection output, by varying the model score threshold that determines what is counted as a model-predicted positive detection of the class. An example precision-recall curve may look something like this for a given classifier:\n\n![](https://miro.medium.com/max/1094/1*ceYw4cXV4BtB2s7S9elExg.png)\n\nThe final step to calculating the AP score is to take the average value of the precision across all recall values\n\n# Localization and Intersection over Union\n\nIn order to evaluate the model on the task of object localization, we must first determine how well the model predicted the location of the object. Usually, this is done by drawing a bounding box around the object of interest, but in some cases it is an N-sided polygon or even pixel by pixel segmentation. For all of these cases, the localization task is typically evaluated on the Intersection over Union threshold (IoU). For definiteness, throughout the rest of the article, I’ll assume that the model predicts bounding boxes, but almost everything said will also apply to pixel-wise segmentation or N-sided polygons. Many good explanations of IoU exist, (see [this one ](https://www.pyimagesearch.com/2016/11/07/intersection-over-union-iou-for-object-detection/)for example), but the basic idea is that it summarizes how well the ground truth object overlaps the object boundary predicted by the model.\n\nModel object detections are determined to be true or false depending upon the IoU threshold. This IoU threshold(s) for each competition vary, but in the [COCO challenge](http://cocodataset.org/#detections-eval), for example, 10 different IoU thresholds are considered, from 0.5 to 0.95 in steps of 0.05. For a specific object (say, ‘person’) this is what the precision-recall curves may look like when calculated at the different IoU thresholds of the COCO challenge:\n\n![](https://miro.medium.com/max/1140/1*kmAon_ut7ZU3h0XkGbqR6g.png)\n\n**Now that we’ve defined Average Precision (AP) and seen how the IoU threshold affects it, the mean Average Precision or mAP score is calculated by taking the mean AP over all classes and/or over all IoU thresholds, depending on the competition**\n\n[# Code for Calculating the mean Average Precision](https://gist.github.com/tarlen5/008809c3decf19313de216b9208f3734)\n\nRef : [Link ](https://medium.com/@timothycarlen/understanding-the-map-evaluation-metric-for-object-detection-a07fe6962cf3)",
    "694740": "Excellent topic. Unfortunately I have only 1 to give.",
    "694742": "Great Explanation @mobassir",
    "694756": "thanks dear @veeralakrishna",
    "694757": "hahahha thank you @mpwolke",
    "694913": "Great. Thanks for sharing @mobassir... Could see that you have done a lot of research on this topic. Excellent work.....",
    "694914": "Thank you @manojprabhaakr :)",
    "695386": "Good deep understanding👍",
    "695418": "Nice one ! very comprehensive explanation. Those One liners for Precision-Recall are perfect.",
    "696128": "Hi! Thanks for your explaination! I am confused about how does false negative play a role in mAP. Also, if two bounding boxes are accurate enough to be consider as true positives for one car, how is mAP like now.",
    "696136": "hi @tonychenxyz \n if two bounding boxes are accurate enough to be consider as true positives for one car then recall score will be high.\nwe calculate  Average Precision (AP) before calculating mAP\nand inside AP calculation we get scores like fp,fn \nRecall measures the “false negative rate” or the ratio of true object detections to the total number of objects in the data set. If you have a recall score close to 1.0 then almost all objects that are in your dataset will be positively detected by the model.",
    "696246": "Agree ! had a similar thought in mind.\n**Precision** : Measure of Exactness ( Quality)\n**Recall** : Measure of Completeness ( Quantity)"
  },
  "source": "meta"
}