{
  "id": 253345,
  "title": "Understand the competition metric!",
  "url": "/competitions/siim-covid19-detection/discussion/253345",
  "author_name": "",
  "post_date": "2021-07-16T07:42:04.273368100Z",
  "votes": 70,
  "comment_count": 23,
  "views": 0,
  "content": "<p><strong>TLDR:</strong></p>\n<ul>\n<li>The metric calculates the Area under the Precision/Recall curve for each class and averages the results</li>\n<li>All boxes with an IoU &gt; 0.5 with a ground truth box of the same class are considered a true positive</li>\n<li>You can only predict each ground truth once, only the first prediction of a box will be counted as a true positive, the rest are false positives</li>\n<li>Lowering your bbox threshold can only increase the score.<br>\nThe confidence score dictates the order in which the predictions are judged by the metric</li>\n</ul>\n<p><strong>I would like to thank these resources for providing great information about the topic:</strong></p>\n<p><a href=\"url\" target=\"_blank\">https://jonathan-hui.medium.com/map-mean-average-precision-for-object-detection-45c121a31173</a><br>\n<a href=\"url\" target=\"_blank\">https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229637</a></p>\n<p>Hi everyone! I’ve been struggling with the metric of this competition and was unable to understand it’s behavior on my submissions. Now that I’ve spent some more time learning about it I would like to share my understanding of it. Any feedback and corrections are very welcome!</p>\n<h3>6 Classes</h3>\n<p>In this competition we predict 6 Classes: negative, typical, indeterminate, atypical, none, opacity<br>\nThe metric calculates the area under the precision/recall curve for each of these classes independently and averages them to get the result. It is unknown if there is any weighting going on in this competition, I believe.<br>\nSo lets look at what the metric is doing for each class:</p>\n<h3>Intersection over Union</h3>\n<p><img src=\"https://pyimagesearch.com/wp-content/uploads/2016/09/iou_equation.png\" alt=\"\"><br>\nA box will be considered a true positive if it has an IoU with a ground truth box greater than 0.5<br>\nThe intersection over Union divides the shared area of two boxes by their combined Area. So if two boxes shared Area is greater than half of their combined area we will consider it a true positive.<br>\nBut: We can only correctly predict each bounding box once, meaning that all predictions of one bounding box after the first one will be counted as a false positive!</p>\n<h3>Area under Curve</h3>\n<p>Precision describes how many of our predictions were correct and Recall describes how many of the ground truth positives we were able to find.<br>\nThe Precision/Recall curve looks at each box one by one and calculates the new overall precision and recall and plots the values. It starts at the box with the highest confidence value and stops at the box with the lowest confidence value.<br>\n<img src=\"https://www.ccom.ucsd.edu/~cdeotte/Kaggle/map1.png\" alt=\"\"><br>\nYou can see that for each true positive we find the recall will only increase because we are increasing the fraction of ground truth boxes we correctly identified. Since the metric uses the area under this curve, we want to keep the precision as high as possible for as long as possible. That is why we need to make sure that the boxes with a high confidence score are highly likely to be true positives. As Chris Deotte points out in the link above you can already increase your score with a better order of the boxes by adjusting the confidence values:<br>\n<img src=\"https://www.ccom.ucsd.edu/~cdeotte/Kaggle/map2.png\" alt=\"\"><br>\nNow you can see why lowering the threshold can only increase the score: Even if all the new boxes were false positives our precision would simply drop towards zero and the recall wouldn’t change, meaning the area under curve would be exactly the same. But since there are likely some true positives in the low confidence boxes our recall will increase and even if our precision is very low it will still cause an increase in Area!<br>\n<img src=\"https://www.ccom.ucsd.edu/~cdeotte/Kaggle/map3.png\" alt=\"\"><br>\nFor the classes without bounding boxes the metric works basically the same. If we predict the correct class, it is considered a true positive and if we predict an incorrect class, it is a false positive the only difference being you do not have to worry about the position of the box. You can probably see how important it is for your model to make good confidence scores instead of simply predicting 1 or 0.</p>\n<p><strong>Only the order of the confidence score within its class matters not the values</strong></p>\n<p>One more thing about this metric is that the P/R curve is monotonically decreasing. If we were to plot the actual P/R curve it would look like the B graph but in this metric it only drops to the highest value to the right of the point as seen in C.<br>\n<img src=\"https://miro.medium.com/max/700/1*zqTL1KW1gwzion9jY8SjHA.png\" alt=\"\"></p>\n<p>One problem from my own experience was that my lb score would drop if I increased the threshold of my detection model… By now you can probably understand the problem yourself! When you’re only doing detection without a 2-class model to predict the none class then by lowering your threshold you will be less likely to predict the none class. If you use a threshold of for example 0.001 you will never be predicting none which means that the score for the none class will be 0! And since we average the scores for all classes the reason for the lb score dropping is that the score of the none class drops!</p>\n<p>If you've made it this far I want to honestly thank you for reading my kernel, I hope you got something valuable out of it😄<br>\nIf you have any question be sure to post them below and I will make sure to answer them or include the answers in the post.</p>",
  "messages": [
    {
      "id": "1389908",
      "postDate": "07/16/2021 07:42:04",
      "content": "<p><strong>TLDR:</strong></p>\n<ul>\n<li>The metric calculates the Area under the Precision/Recall curve for each class and averages the results</li>\n<li>All boxes with an IoU &gt; 0.5 with a ground truth box of the same class are considered a true positive</li>\n<li>You can only predict each ground truth once, only the first prediction of a box will be counted as a true positive, the rest are false positives</li>\n<li>Lowering your bbox threshold can only increase the score.<br>\nThe confidence score dictates the order in which the predictions are judged by the metric</li>\n</ul>\n<p><strong>I would like to thank these resources for providing great information about the topic:</strong></p>\n<p><a href=\"url\" target=\"_blank\">https://jonathan-hui.medium.com/map-mean-average-precision-for-object-detection-45c121a31173</a><br>\n<a href=\"url\" target=\"_blank\">https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229637</a></p>\n<p>Hi everyone! I’ve been struggling with the metric of this competition and was unable to understand it’s behavior on my submissions. Now that I’ve spent some more time learning about it I would like to share my understanding of it. Any feedback and corrections are very welcome!</p>\n<h3>6 Classes</h3>\n<p>In this competition we predict 6 Classes: negative, typical, indeterminate, atypical, none, opacity<br>\nThe metric calculates the area under the precision/recall curve for each of these classes independently and averages them to get the result. It is unknown if there is any weighting going on in this competition, I believe.<br>\nSo lets look at what the metric is doing for each class:</p>\n<h3>Intersection over Union</h3>\n<p><img src=\"https://pyimagesearch.com/wp-content/uploads/2016/09/iou_equation.png\" alt=\"\"><br>\nA box will be considered a true positive if it has an IoU with a ground truth box greater than 0.5<br>\nThe intersection over Union divides the shared area of two boxes by their combined Area. So if two boxes shared Area is greater than half of their combined area we will consider it a true positive.<br>\nBut: We can only correctly predict each bounding box once, meaning that all predictions of one bounding box after the first one will be counted as a false positive!</p>\n<h3>Area under Curve</h3>\n<p>Precision describes how many of our predictions were correct and Recall describes how many of the ground truth positives we were able to find.<br>\nThe Precision/Recall curve looks at each box one by one and calculates the new overall precision and recall and plots the values. It starts at the box with the highest confidence value and stops at the box with the lowest confidence value.<br>\n<img src=\"https://www.ccom.ucsd.edu/~cdeotte/Kaggle/map1.png\" alt=\"\"><br>\nYou can see that for each true positive we find the recall will only increase because we are increasing the fraction of ground truth boxes we correctly identified. Since the metric uses the area under this curve, we want to keep the precision as high as possible for as long as possible. That is why we need to make sure that the boxes with a high confidence score are highly likely to be true positives. As Chris Deotte points out in the link above you can already increase your score with a better order of the boxes by adjusting the confidence values:<br>\n<img src=\"https://www.ccom.ucsd.edu/~cdeotte/Kaggle/map2.png\" alt=\"\"><br>\nNow you can see why lowering the threshold can only increase the score: Even if all the new boxes were false positives our precision would simply drop towards zero and the recall wouldn’t change, meaning the area under curve would be exactly the same. But since there are likely some true positives in the low confidence boxes our recall will increase and even if our precision is very low it will still cause an increase in Area!<br>\n<img src=\"https://www.ccom.ucsd.edu/~cdeotte/Kaggle/map3.png\" alt=\"\"><br>\nFor the classes without bounding boxes the metric works basically the same. If we predict the correct class, it is considered a true positive and if we predict an incorrect class, it is a false positive the only difference being you do not have to worry about the position of the box. You can probably see how important it is for your model to make good confidence scores instead of simply predicting 1 or 0.</p>\n<p><strong>Only the order of the confidence score within its class matters not the values</strong></p>\n<p>One more thing about this metric is that the P/R curve is monotonically decreasing. If we were to plot the actual P/R curve it would look like the B graph but in this metric it only drops to the highest value to the right of the point as seen in C.<br>\n<img src=\"https://miro.medium.com/max/700/1*zqTL1KW1gwzion9jY8SjHA.png\" alt=\"\"></p>\n<p>One problem from my own experience was that my lb score would drop if I increased the threshold of my detection model… By now you can probably understand the problem yourself! When you’re only doing detection without a 2-class model to predict the none class then by lowering your threshold you will be less likely to predict the none class. If you use a threshold of for example 0.001 you will never be predicting none which means that the score for the none class will be 0! And since we average the scores for all classes the reason for the lb score dropping is that the score of the none class drops!</p>\n<p>If you've made it this far I want to honestly thank you for reading my kernel, I hope you got something valuable out of it😄<br>\nIf you have any question be sure to post them below and I will make sure to answer them or include the answers in the post.</p>",
      "rawMarkdown": "**TLDR:**\n- The metric calculates the Area under the Precision/Recall curve for each class and averages the results\n- All boxes with an IoU > 0.5 with a ground truth box of the same class are considered a true positive\n- You can only predict each ground truth once, only the first prediction of a box will be counted as a true positive, the rest are false positives\n- Lowering your bbox threshold can only increase the score.\nThe confidence score dictates the order in which the predictions are judged by the metric\n\n\n**I would like to thank these resources for providing great information about the topic:**\n\n[https://jonathan-hui.medium.com/map-mean-average-precision-for-object-detection-45c121a31173](url)\n[https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229637](url)\n\n\nHi everyone! I’ve been struggling with the metric of this competition and was unable to understand it’s behavior on my submissions. Now that I’ve spent some more time learning about it I would like to share my understanding of it. Any feedback and corrections are very welcome!\n###6 Classes\nIn this competition we predict 6 Classes: negative, typical, indeterminate, atypical, none, opacity\nThe metric calculates the area under the precision/recall curve for each of these classes independently and averages them to get the result. It is unknown if there is any weighting going on in this competition, I believe.\nSo lets look at what the metric is doing for each class:\n###Intersection over Union\n![](https://pyimagesearch.com/wp-content/uploads/2016/09/iou_equation.png)\nA box will be considered a true positive if it has an IoU with a ground truth box greater than 0.5\nThe intersection over Union divides the shared area of two boxes by their combined Area. So if two boxes shared Area is greater than half of their combined area we will consider it a true positive.\nBut: We can only correctly predict each bounding box once, meaning that all predictions of one bounding box after the first one will be counted as a false positive!\n###Area under Curve\nPrecision describes how many of our predictions were correct and Recall describes how many of the ground truth positives we were able to find.\nThe Precision/Recall curve looks at each box one by one and calculates the new overall precision and recall and plots the values. It starts at the box with the highest confidence value and stops at the box with the lowest confidence value.\n![](https://www.ccom.ucsd.edu/~cdeotte/Kaggle/map1.png)\nYou can see that for each true positive we find the recall will only increase because we are increasing the fraction of ground truth boxes we correctly identified. Since the metric uses the area under this curve, we want to keep the precision as high as possible for as long as possible. That is why we need to make sure that the boxes with a high confidence score are highly likely to be true positives. As Chris Deotte points out in the link above you can already increase your score with a better order of the boxes by adjusting the confidence values:\n![](https://www.ccom.ucsd.edu/~cdeotte/Kaggle/map2.png)\nNow you can see why lowering the threshold can only increase the score: Even if all the new boxes were false positives our precision would simply drop towards zero and the recall wouldn’t change, meaning the area under curve would be exactly the same. But since there are likely some true positives in the low confidence boxes our recall will increase and even if our precision is very low it will still cause an increase in Area!\n![](https://www.ccom.ucsd.edu/~cdeotte/Kaggle/map3.png)\nFor the classes without bounding boxes the metric works basically the same. If we predict the correct class, it is considered a true positive and if we predict an incorrect class, it is a false positive the only difference being you do not have to worry about the position of the box. You can probably see how important it is for your model to make good confidence scores instead of simply predicting 1 or 0.\n\n**Only the order of the confidence score within its class matters not the values**\n\nOne more thing about this metric is that the P/R curve is monotonically decreasing. If we were to plot the actual P/R curve it would look like the B graph but in this metric it only drops to the highest value to the right of the point as seen in C.\n![](https://miro.medium.com/max/700/1*zqTL1KW1gwzion9jY8SjHA.png)\n\nOne problem from my own experience was that my lb score would drop if I increased the threshold of my detection model… By now you can probably understand the problem yourself! When you’re only doing detection without a 2-class model to predict the none class then by lowering your threshold you will be less likely to predict the none class. If you use a threshold of for example 0.001 you will never be predicting none which means that the score for the none class will be 0! And since we average the scores for all classes the reason for the lb score dropping is that the score of the none class drops!\n\nIf you've made it this far I want to honestly thank you for reading my kernel, I hope you got something valuable out of it😄\nIf you have any question be sure to post them below and I will make sure to answer them or include the answers in the post.",
      "votes": null
    },
    {
      "id": "1396147",
      "postDate": "07/21/2021 20:36:48",
      "content": "<p>Nice article!</p>",
      "rawMarkdown": "Nice article!",
      "votes": null
    },
    {
      "id": "1396171",
      "postDate": "07/21/2021 21:23:10",
      "content": "<p>Thank you very much</p>",
      "rawMarkdown": "Thank you very much",
      "votes": null
    },
    {
      "id": "1396634",
      "postDate": "07/22/2021 10:11:18",
      "content": "<p>If randomly add ROI with 0.001 this should increase the scores?</p>",
      "rawMarkdown": "If randomly add ROI with 0.001 this should increase the scores?",
      "votes": null
    },
    {
      "id": "1396640",
      "postDate": "07/22/2021 10:17:19",
      "content": "<p>You mean randomly adding opacity boxes with 0.001 confidence? I think that should work yes,but make sure the boxes dont overlap with previous boxes. It is similar to lowering the threshold on the detection model.</p>",
      "rawMarkdown": "You mean randomly adding opacity boxes with 0.001 confidence? I think that should work yes,but make sure the boxes dont overlap with previous boxes. It is similar to lowering the threshold on the detection model.",
      "votes": null
    },
    {
      "id": "1396644",
      "postDate": "07/22/2021 10:24:26",
      "content": "<p>Yes. I'll try if I take the time to do it</p>",
      "rawMarkdown": "Yes. I'll try if I take the time to do it",
      "votes": null
    },
    {
      "id": "1397165",
      "postDate": "07/22/2021 21:13:28",
      "content": "<p>Wow thanks for this! Great read!</p>",
      "rawMarkdown": "Wow thanks for this! Great read!",
      "votes": null
    },
    {
      "id": "1397168",
      "postDate": "07/22/2021 21:23:34",
      "content": "<p>Thank you very much</p>",
      "rawMarkdown": "Thank you very much",
      "votes": null
    },
    {
      "id": "1397733",
      "postDate": "07/23/2021 13:01:58",
      "content": "<p>Nicely explained. Thanks for sharing <a href=\"https://www.kaggle.com/simon111\" target=\"_blank\">@simon111</a> </p>",
      "rawMarkdown": "Nicely explained. Thanks for sharing @simon111",
      "votes": null
    },
    {
      "id": "1397760",
      "postDate": "07/23/2021 13:28:05",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    },
    {
      "id": "1398400",
      "postDate": "07/24/2021 06:35:58",
      "content": "<p>Thank you for explaining this metric. it's so helpful for me</p>",
      "rawMarkdown": "Thank you for explaining this metric. it's so helpful for me",
      "votes": null
    },
    {
      "id": "1398529",
      "postDate": "07/24/2021 08:54:45",
      "content": "<p>My pleasure!</p>",
      "rawMarkdown": "My pleasure!",
      "votes": null
    },
    {
      "id": "1398846",
      "postDate": "07/24/2021 14:53:29",
      "content": "<p>So just want to clarify, you can increase your score by:</p>\n<ol>\n<li>Decreasing threshold of object detection model to increase number of predictions</li>\n<li>Removing excess of predictions that have high overlap with each other (IoU &gt;0.5)</li>\n<li>Reordering your predictions in descending order of confidence score</li>\n</ol>\n<p>Hope I understood this right. Thank you for the great post! Going to try this out when I get time. </p>",
      "rawMarkdown": "So just want to clarify, you can increase your score by:\n\n1. Decreasing threshold of object detection model to increase number of predictions\n2. Removing excess of predictions that have high overlap with each other (IoU >0.5)\n3. Reordering your predictions in descending order of confidence score\n\nHope I understood this right. Thank you for the great post! Going to try this out when I get time.",
      "votes": null
    },
    {
      "id": "1398897",
      "postDate": "07/24/2021 16:07:21",
      "content": "<p>Hi the first two points are correct. I have a small correction for the 3. : You dont have to sort your predictions by confidence. The metric sorts all your predictions for each class by confidence, so what you want to do is to establish a good order by having good meaning correct confidence scores. I hope that makes sense :)</p>",
      "rawMarkdown": "Hi the first two points are correct. I have a small correction for the 3. : You dont have to sort your predictions by confidence. The metric sorts all your predictions for each class by confidence, so what you want to do is to establish a good order by having good meaning correct confidence scores. I hope that makes sense :)",
      "votes": null
    },
    {
      "id": "1399026",
      "postDate": "07/24/2021 19:25:19",
      "content": "<p>Ah ok, makes sense, I looked closer at your examples above. Thanks for clarifying!</p>",
      "rawMarkdown": "Ah ok, makes sense, I looked closer at your examples above. Thanks for clarifying!",
      "votes": null
    },
    {
      "id": "1399633",
      "postDate": "07/25/2021 14:25:16",
      "content": "<p>Really nice post, Ty so much</p>",
      "rawMarkdown": "Really nice post, Ty so much",
      "votes": null
    },
    {
      "id": "1399640",
      "postDate": "07/25/2021 14:31:11",
      "content": "<p>Thank you🙏</p>",
      "rawMarkdown": "Thank you🙏",
      "votes": null
    },
    {
      "id": "1400384",
      "postDate": "07/26/2021 08:46:18",
      "content": "<p>wow this is awesome! Thank you for sharing the information!!</p>",
      "rawMarkdown": "wow this is awesome! Thank you for sharing the information!!",
      "votes": null
    },
    {
      "id": "1400409",
      "postDate": "07/26/2021 09:04:46",
      "content": "<p>Thank you very much!</p>",
      "rawMarkdown": "Thank you very much!",
      "votes": null
    },
    {
      "id": "1400774",
      "postDate": "07/26/2021 15:10:06",
      "content": "<p>Great post! I'm wondering about one thing though: do they calculate the mAP scores for studies and images separately or altogether? So for instance if you had study predictions with confidence scores 0.8, 0.6, 0.4 and image predictions with confidence scores 0.7, 0.5. Would they sort them like: 0.8, 0.7, 0.6, 0.5, 0.4, and then calculate mAP in that order? Or would they sort them separately like 0.8, 0.6, 0.4 &amp; 0.7, 0.5. Calculate the two mAP scores separately and then add them up? </p>\n<p>I've seen some posts suggesting that they compute it separately, while your post seems to suggest that they compute it altogether. Does anyone have any insights on this?</p>",
      "rawMarkdown": "Great post! I'm wondering about one thing though: do they calculate the mAP scores for studies and images separately or altogether? So for instance if you had study predictions with confidence scores 0.8, 0.6, 0.4 and image predictions with confidence scores 0.7, 0.5. Would they sort them like: 0.8, 0.7, 0.6, 0.5, 0.4, and then calculate mAP in that order? Or would they sort them separately like 0.8, 0.6, 0.4 & 0.7, 0.5. Calculate the two mAP scores separately and then add them up? \n\nI've seen some posts suggesting that they compute it separately, while your post seems to suggest that they compute it altogether. Does anyone have any insights on this?",
      "votes": null
    },
    {
      "id": "1400833",
      "postDate": "07/26/2021 16:13:31",
      "content": "<p>Since the classes none and opacity only appear in images and the 4 other classes only appear in studies it doesnt make a difference. I'm not sure if they for example filter none predictions out of study rows but if you only predict the 2 classes on image and the 4 classes on study it shouldnt make a difference</p>",
      "rawMarkdown": "Since the classes none and opacity only appear in images and the 4 other classes only appear in studies it doesnt make a difference. I'm not sure if they for example filter none predictions out of study rows but if you only predict the 2 classes on image and the 4 classes on study it shouldnt make a difference",
      "votes": null
    },
    {
      "id": "1400840",
      "postDate": "07/26/2021 16:23:26",
      "content": "<p>Yes you are absolutely right, thanks for the explanation!</p>",
      "rawMarkdown": "Yes you are absolutely right, thanks for the explanation!",
      "votes": null
    },
    {
      "id": "1402233",
      "postDate": "07/28/2021 03:10:51",
      "content": "<p>If 2 boxes have an IOU &gt; 0.5 how do you decide which to keep and which not to?</p>",
      "rawMarkdown": "If 2 boxes have an IOU > 0.5 how do you decide which to keep and which not to?",
      "votes": null
    },
    {
      "id": "1402483",
      "postDate": "07/28/2021 08:28:23",
      "content": "<p>Always the first one meaning the one with the higher confidence gets the true positive and all after that get false positive</p>",
      "rawMarkdown": "Always the first one meaning the one with the higher confidence gets the true positive and all after that get false positive",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1396147,
      "author_name": "artanddav",
      "author_url": "",
      "post_date": "07/21/2021 20:36:48",
      "content": "<p>Nice article!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1396171,
          "author_name": "simon111",
          "author_url": "",
          "post_date": "07/21/2021 21:23:10",
          "content": "<p>Thank you very much</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1396634,
      "author_name": "aleksandrkruchinin",
      "author_url": "",
      "post_date": "07/22/2021 10:11:18",
      "content": "<p>If randomly add ROI with 0.001 this should increase the scores?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1396640,
          "author_name": "simon111",
          "author_url": "",
          "post_date": "07/22/2021 10:17:19",
          "content": "<p>You mean randomly adding opacity boxes with 0.001 confidence? I think that should work yes,but make sure the boxes dont overlap with previous boxes. It is similar to lowering the threshold on the detection model.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1396644,
          "author_name": "aleksandrkruchinin",
          "author_url": "",
          "post_date": "07/22/2021 10:24:26",
          "content": "<p>Yes. I'll try if I take the time to do it</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1397165,
      "author_name": "neelgajare",
      "author_url": "",
      "post_date": "07/22/2021 21:13:28",
      "content": "<p>Wow thanks for this! Great read!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1397168,
          "author_name": "simon111",
          "author_url": "",
          "post_date": "07/22/2021 21:23:34",
          "content": "<p>Thank you very much</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1397733,
      "author_name": "amritpal333",
      "author_url": "",
      "post_date": "07/23/2021 13:01:58",
      "content": "<p>Nicely explained. Thanks for sharing <a href=\"https://www.kaggle.com/simon111\" target=\"_blank\">@simon111</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 1397760,
          "author_name": "simon111",
          "author_url": "",
          "post_date": "07/23/2021 13:28:05",
          "content": "<p>Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1398400,
      "author_name": "adldotori",
      "author_url": "",
      "post_date": "07/24/2021 06:35:58",
      "content": "<p>Thank you for explaining this metric. it's so helpful for me</p>",
      "votes": null,
      "replies": [
        {
          "id": 1398529,
          "author_name": "simon111",
          "author_url": "",
          "post_date": "07/24/2021 08:54:45",
          "content": "<p>My pleasure!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1398846,
      "author_name": "illidan7",
      "author_url": "",
      "post_date": "07/24/2021 14:53:29",
      "content": "<p>So just want to clarify, you can increase your score by:</p>\n<ol>\n<li>Decreasing threshold of object detection model to increase number of predictions</li>\n<li>Removing excess of predictions that have high overlap with each other (IoU &gt;0.5)</li>\n<li>Reordering your predictions in descending order of confidence score</li>\n</ol>\n<p>Hope I understood this right. Thank you for the great post! Going to try this out when I get time. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1398897,
          "author_name": "simon111",
          "author_url": "",
          "post_date": "07/24/2021 16:07:21",
          "content": "<p>Hi the first two points are correct. I have a small correction for the 3. : You dont have to sort your predictions by confidence. The metric sorts all your predictions for each class by confidence, so what you want to do is to establish a good order by having good meaning correct confidence scores. I hope that makes sense :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1399026,
          "author_name": "illidan7",
          "author_url": "",
          "post_date": "07/24/2021 19:25:19",
          "content": "<p>Ah ok, makes sense, I looked closer at your examples above. Thanks for clarifying!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1399633,
      "author_name": "ishaananand77",
      "author_url": "",
      "post_date": "07/25/2021 14:25:16",
      "content": "<p>Really nice post, Ty so much</p>",
      "votes": null,
      "replies": [
        {
          "id": 1399640,
          "author_name": "simon111",
          "author_url": "",
          "post_date": "07/25/2021 14:31:11",
          "content": "<p>Thank you🙏</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1400384,
      "author_name": "heejaemon",
      "author_url": "",
      "post_date": "07/26/2021 08:46:18",
      "content": "<p>wow this is awesome! Thank you for sharing the information!!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1400409,
          "author_name": "simon111",
          "author_url": "",
          "post_date": "07/26/2021 09:04:46",
          "content": "<p>Thank you very much!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1400774,
      "author_name": "diningeachox",
      "author_url": "",
      "post_date": "07/26/2021 15:10:06",
      "content": "<p>Great post! I'm wondering about one thing though: do they calculate the mAP scores for studies and images separately or altogether? So for instance if you had study predictions with confidence scores 0.8, 0.6, 0.4 and image predictions with confidence scores 0.7, 0.5. Would they sort them like: 0.8, 0.7, 0.6, 0.5, 0.4, and then calculate mAP in that order? Or would they sort them separately like 0.8, 0.6, 0.4 &amp; 0.7, 0.5. Calculate the two mAP scores separately and then add them up? </p>\n<p>I've seen some posts suggesting that they compute it separately, while your post seems to suggest that they compute it altogether. Does anyone have any insights on this?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1400833,
          "author_name": "simon111",
          "author_url": "",
          "post_date": "07/26/2021 16:13:31",
          "content": "<p>Since the classes none and opacity only appear in images and the 4 other classes only appear in studies it doesnt make a difference. I'm not sure if they for example filter none predictions out of study rows but if you only predict the 2 classes on image and the 4 classes on study it shouldnt make a difference</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1400840,
          "author_name": "diningeachox",
          "author_url": "",
          "post_date": "07/26/2021 16:23:26",
          "content": "<p>Yes you are absolutely right, thanks for the explanation!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1402233,
      "author_name": "khubchandani",
      "author_url": "",
      "post_date": "07/28/2021 03:10:51",
      "content": "<p>If 2 boxes have an IOU &gt; 0.5 how do you decide which to keep and which not to?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1402483,
          "author_name": "simon111",
          "author_url": "",
          "post_date": "07/28/2021 08:28:23",
          "content": "<p>Always the first one meaning the one with the higher confidence gets the true positive and all after that get false positive</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1389908": "**TLDR:**\n- The metric calculates the Area under the Precision/Recall curve for each class and averages the results\n- All boxes with an IoU > 0.5 with a ground truth box of the same class are considered a true positive\n- You can only predict each ground truth once, only the first prediction of a box will be counted as a true positive, the rest are false positives\n- Lowering your bbox threshold can only increase the score.\nThe confidence score dictates the order in which the predictions are judged by the metric\n\n\n**I would like to thank these resources for providing great information about the topic:**\n\n[https://jonathan-hui.medium.com/map-mean-average-precision-for-object-detection-45c121a31173](url)\n[https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229637](url)\n\n\nHi everyone! I’ve been struggling with the metric of this competition and was unable to understand it’s behavior on my submissions. Now that I’ve spent some more time learning about it I would like to share my understanding of it. Any feedback and corrections are very welcome!\n###6 Classes\nIn this competition we predict 6 Classes: negative, typical, indeterminate, atypical, none, opacity\nThe metric calculates the area under the precision/recall curve for each of these classes independently and averages them to get the result. It is unknown if there is any weighting going on in this competition, I believe.\nSo lets look at what the metric is doing for each class:\n###Intersection over Union\n![](https://pyimagesearch.com/wp-content/uploads/2016/09/iou_equation.png)\nA box will be considered a true positive if it has an IoU with a ground truth box greater than 0.5\nThe intersection over Union divides the shared area of two boxes by their combined Area. So if two boxes shared Area is greater than half of their combined area we will consider it a true positive.\nBut: We can only correctly predict each bounding box once, meaning that all predictions of one bounding box after the first one will be counted as a false positive!\n###Area under Curve\nPrecision describes how many of our predictions were correct and Recall describes how many of the ground truth positives we were able to find.\nThe Precision/Recall curve looks at each box one by one and calculates the new overall precision and recall and plots the values. It starts at the box with the highest confidence value and stops at the box with the lowest confidence value.\n![](https://www.ccom.ucsd.edu/~cdeotte/Kaggle/map1.png)\nYou can see that for each true positive we find the recall will only increase because we are increasing the fraction of ground truth boxes we correctly identified. Since the metric uses the area under this curve, we want to keep the precision as high as possible for as long as possible. That is why we need to make sure that the boxes with a high confidence score are highly likely to be true positives. As Chris Deotte points out in the link above you can already increase your score with a better order of the boxes by adjusting the confidence values:\n![](https://www.ccom.ucsd.edu/~cdeotte/Kaggle/map2.png)\nNow you can see why lowering the threshold can only increase the score: Even if all the new boxes were false positives our precision would simply drop towards zero and the recall wouldn’t change, meaning the area under curve would be exactly the same. But since there are likely some true positives in the low confidence boxes our recall will increase and even if our precision is very low it will still cause an increase in Area!\n![](https://www.ccom.ucsd.edu/~cdeotte/Kaggle/map3.png)\nFor the classes without bounding boxes the metric works basically the same. If we predict the correct class, it is considered a true positive and if we predict an incorrect class, it is a false positive the only difference being you do not have to worry about the position of the box. You can probably see how important it is for your model to make good confidence scores instead of simply predicting 1 or 0.\n\n**Only the order of the confidence score within its class matters not the values**\n\nOne more thing about this metric is that the P/R curve is monotonically decreasing. If we were to plot the actual P/R curve it would look like the B graph but in this metric it only drops to the highest value to the right of the point as seen in C.\n![](https://miro.medium.com/max/700/1*zqTL1KW1gwzion9jY8SjHA.png)\n\nOne problem from my own experience was that my lb score would drop if I increased the threshold of my detection model… By now you can probably understand the problem yourself! When you’re only doing detection without a 2-class model to predict the none class then by lowering your threshold you will be less likely to predict the none class. If you use a threshold of for example 0.001 you will never be predicting none which means that the score for the none class will be 0! And since we average the scores for all classes the reason for the lb score dropping is that the score of the none class drops!\n\nIf you've made it this far I want to honestly thank you for reading my kernel, I hope you got something valuable out of it😄\nIf you have any question be sure to post them below and I will make sure to answer them or include the answers in the post.",
    "1396147": "Nice article!",
    "1396171": "Thank you very much",
    "1396634": "If randomly add ROI with 0.001 this should increase the scores?",
    "1396640": "You mean randomly adding opacity boxes with 0.001 confidence? I think that should work yes,but make sure the boxes dont overlap with previous boxes. It is similar to lowering the threshold on the detection model.",
    "1396644": "Yes. I'll try if I take the time to do it",
    "1397165": "Wow thanks for this! Great read!",
    "1397168": "Thank you very much",
    "1397733": "Nicely explained. Thanks for sharing @simon111",
    "1397760": "Thank you!",
    "1398400": "Thank you for explaining this metric. it's so helpful for me",
    "1398529": "My pleasure!",
    "1398846": "So just want to clarify, you can increase your score by:\n\n1. Decreasing threshold of object detection model to increase number of predictions\n2. Removing excess of predictions that have high overlap with each other (IoU >0.5)\n3. Reordering your predictions in descending order of confidence score\n\nHope I understood this right. Thank you for the great post! Going to try this out when I get time.",
    "1398897": "Hi the first two points are correct. I have a small correction for the 3. : You dont have to sort your predictions by confidence. The metric sorts all your predictions for each class by confidence, so what you want to do is to establish a good order by having good meaning correct confidence scores. I hope that makes sense :)",
    "1399026": "Ah ok, makes sense, I looked closer at your examples above. Thanks for clarifying!",
    "1399633": "Really nice post, Ty so much",
    "1399640": "Thank you🙏",
    "1400384": "wow this is awesome! Thank you for sharing the information!!",
    "1400409": "Thank you very much!",
    "1400774": "Great post! I'm wondering about one thing though: do they calculate the mAP scores for studies and images separately or altogether? So for instance if you had study predictions with confidence scores 0.8, 0.6, 0.4 and image predictions with confidence scores 0.7, 0.5. Would they sort them like: 0.8, 0.7, 0.6, 0.5, 0.4, and then calculate mAP in that order? Or would they sort them separately like 0.8, 0.6, 0.4 & 0.7, 0.5. Calculate the two mAP scores separately and then add them up? \n\nI've seen some posts suggesting that they compute it separately, while your post seems to suggest that they compute it altogether. Does anyone have any insights on this?",
    "1400833": "Since the classes none and opacity only appear in images and the 4 other classes only appear in studies it doesnt make a difference. I'm not sure if they for example filter none predictions out of study rows but if you only predict the 2 classes on image and the 4 classes on study it shouldnt make a difference",
    "1400840": "Yes you are absolutely right, thanks for the explanation!",
    "1402233": "If 2 boxes have an IOU > 0.5 how do you decide which to keep and which not to?",
    "1402483": "Always the first one meaning the one with the higher confidence gets the true positive and all after that get false positive"
  },
  "source": "meta"
}