{
  "id": 248467,
  "title": "Towards clarification of competition metric",
  "url": "/competitions/siim-covid19-detection/discussion/248467",
  "author_name": "",
  "post_date": "2021-06-23T16:09:09.444419900Z",
  "votes": 6,
  "comment_count": 4,
  "views": 0,
  "content": "<p>As this is my first time to use mAP metric, I am striving to implement the exact competition metric in my local CV settings as close as possible to LB. I would like to share what I understand til now and to discuss confusing points, especially about study-level mAP. Would be appreciated to correct me if anything wrongs below and clarify few things.</p>\n<h2><strong>Image-level mean Average Precision</strong></h2>\n<p>According to the hosts, the calculation is based on PASCAL VOC 2010 competition metric. The PASCAL VOC 2010 is one version of several mAP versions. The mAP metric is modified over the past competitions and literatures. The basic common concept of mAP is nothing but <strong>the area under the recall-precision curve</strong>. However, we cannot draw this curve continuously. Depending on how to deal with this approximation of recall-precision curve and its auc values, the mAP versions have their own methods. <br>\nOur competition metric mAP is drawn by interpolating (recall, precision) points at some IoU values. These IoU values usually between 0.5 and 0.75 with step size 0.05. (i.e. [0.5, 0.55, 0.60, 0.65, 0.7, 0.75]) However, in this competition, we use IoU values &gt; 0.5 as described in evaluation section. Step size is not mentioned, but mostly, the step size is 0.05. I am not sure how much the gaps between mAP values with different step sizes (0.01, 0.05, 0.1, ,,,), but It would be grateful if the hosts clarify the step size. I am taking step size 0.05 and the IoU values to calculate recall-precision curve are [0.5, 0.55, ,,, 0.95, 1.0]. Note that this mAP in PASCAL VOC 2010 make the recall-precision curve monotone decreasing form when calculating AUC of recall-precision curve. (By just taking the maximum recall among recall values at next IoU steps) The above is done for each class and the mean value of auc values for each class is the final mAP of our model.</p>\n<p>One can find a nice detailed explanation here (<a href=\"https://jonathan-hui.medium.com/map-mean-average-precision-for-object-detection-45c121a31173\" target=\"_blank\">https://jonathan-hui.medium.com/map-mean-average-precision-for-object-detection-45c121a31173</a>) and a great kernel (<a href=\"https://www.kaggle.com/pestipeti/competition-metric-details-script\" target=\"_blank\">https://www.kaggle.com/pestipeti/competition-metric-details-script</a>).</p>\n<p>For evaluating image-level labels, the above kernel's mAP is valid, but, for this competition, it is needed  additional implementation of monotone decrease function for recall values would be enough to implement the right mAP implementation for this competition. Note that the monotone decreased recall is not included in most mAP calculation libraries sklearn and ZFTurbo's <a href=\"https://github.com/ZFTurbo/Mean-Average-Precision-for-Boxes\" target=\"_blank\">Mean-Average-Precision-for-Boxes</a>.</p>\n<h2><strong>Study-level mean Average Precision?</strong></h2>\n<p>I am still not getting how to calculate this and more detailed explanations or formulas need to be revealed by the host IHMO. <br>\nCurrently, I am assuming that the IoU values is replaced to confidence values and calculated in one-versus-rest manner since no IoU can be calculated. That is, for each class, we can draw recall-precision curve using points at each confidence(correspond to IoU in PASCAL VOC 2010 mAP) step. [0.5, 0.55, ,,, 0.95, 1.0] If this is true, we should make prediction results not with a single-label prediction with confidence 1, (i.e. [0.9, 0.6, 0.1, 0.3] -&gt; \"negative 1 0 0 1 1\") but multi-label predictions (i.e. [0.9, 0.6, 0.1, 0.3] -&gt; \"negative 0.9 0 0 1 1 typical 0.6 0 0 1 1\")</p>\n<p>Also, I am wondering there could be multiple labels for a single study in test set. If so, are the test set labeled with confidence or just 1? From the evaluation section in overview tab, the example says it is possible to make PredictionString contains multi labels. Before the LB updates, some people say a study can have multiple images and thus, multiple study-level labels can be predicted for a study (from discussion <a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/240329\" target=\"_blank\">here</a>). However, now in the test set there are only one image according to the host <a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/246597\" target=\"_blank\">here</a>. Then, do we have single study-level label in test set as the training set does?</p>\n<p>Thank you for reading and any comments will be welcomed :)</p>",
  "messages": [
    {
      "id": "1362763",
      "postDate": "06/23/2021 16:09:09",
      "content": "<p>As this is my first time to use mAP metric, I am striving to implement the exact competition metric in my local CV settings as close as possible to LB. I would like to share what I understand til now and to discuss confusing points, especially about study-level mAP. Would be appreciated to correct me if anything wrongs below and clarify few things.</p>\n<h2><strong>Image-level mean Average Precision</strong></h2>\n<p>According to the hosts, the calculation is based on PASCAL VOC 2010 competition metric. The PASCAL VOC 2010 is one version of several mAP versions. The mAP metric is modified over the past competitions and literatures. The basic common concept of mAP is nothing but <strong>the area under the recall-precision curve</strong>. However, we cannot draw this curve continuously. Depending on how to deal with this approximation of recall-precision curve and its auc values, the mAP versions have their own methods. <br>\nOur competition metric mAP is drawn by interpolating (recall, precision) points at some IoU values. These IoU values usually between 0.5 and 0.75 with step size 0.05. (i.e. [0.5, 0.55, 0.60, 0.65, 0.7, 0.75]) However, in this competition, we use IoU values &gt; 0.5 as described in evaluation section. Step size is not mentioned, but mostly, the step size is 0.05. I am not sure how much the gaps between mAP values with different step sizes (0.01, 0.05, 0.1, ,,,), but It would be grateful if the hosts clarify the step size. I am taking step size 0.05 and the IoU values to calculate recall-precision curve are [0.5, 0.55, ,,, 0.95, 1.0]. Note that this mAP in PASCAL VOC 2010 make the recall-precision curve monotone decreasing form when calculating AUC of recall-precision curve. (By just taking the maximum recall among recall values at next IoU steps) The above is done for each class and the mean value of auc values for each class is the final mAP of our model.</p>\n<p>One can find a nice detailed explanation here (<a href=\"https://jonathan-hui.medium.com/map-mean-average-precision-for-object-detection-45c121a31173\" target=\"_blank\">https://jonathan-hui.medium.com/map-mean-average-precision-for-object-detection-45c121a31173</a>) and a great kernel (<a href=\"https://www.kaggle.com/pestipeti/competition-metric-details-script\" target=\"_blank\">https://www.kaggle.com/pestipeti/competition-metric-details-script</a>).</p>\n<p>For evaluating image-level labels, the above kernel's mAP is valid, but, for this competition, it is needed  additional implementation of monotone decrease function for recall values would be enough to implement the right mAP implementation for this competition. Note that the monotone decreased recall is not included in most mAP calculation libraries sklearn and ZFTurbo's <a href=\"https://github.com/ZFTurbo/Mean-Average-Precision-for-Boxes\" target=\"_blank\">Mean-Average-Precision-for-Boxes</a>.</p>\n<h2><strong>Study-level mean Average Precision?</strong></h2>\n<p>I am still not getting how to calculate this and more detailed explanations or formulas need to be revealed by the host IHMO. <br>\nCurrently, I am assuming that the IoU values is replaced to confidence values and calculated in one-versus-rest manner since no IoU can be calculated. That is, for each class, we can draw recall-precision curve using points at each confidence(correspond to IoU in PASCAL VOC 2010 mAP) step. [0.5, 0.55, ,,, 0.95, 1.0] If this is true, we should make prediction results not with a single-label prediction with confidence 1, (i.e. [0.9, 0.6, 0.1, 0.3] -&gt; \"negative 1 0 0 1 1\") but multi-label predictions (i.e. [0.9, 0.6, 0.1, 0.3] -&gt; \"negative 0.9 0 0 1 1 typical 0.6 0 0 1 1\")</p>\n<p>Also, I am wondering there could be multiple labels for a single study in test set. If so, are the test set labeled with confidence or just 1? From the evaluation section in overview tab, the example says it is possible to make PredictionString contains multi labels. Before the LB updates, some people say a study can have multiple images and thus, multiple study-level labels can be predicted for a study (from discussion <a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/240329\" target=\"_blank\">here</a>). However, now in the test set there are only one image according to the host <a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/246597\" target=\"_blank\">here</a>. Then, do we have single study-level label in test set as the training set does?</p>\n<p>Thank you for reading and any comments will be welcomed :)</p>",
      "rawMarkdown": "As this is my first time to use mAP metric, I am striving to implement the exact competition metric in my local CV settings as close as possible to LB. I would like to share what I understand til now and to discuss confusing points, especially about study-level mAP. Would be appreciated to correct me if anything wrongs below and clarify few things.\n\n## **Image-level mean Average Precision**\n\nAccording to the hosts, the calculation is based on PASCAL VOC 2010 competition metric. The PASCAL VOC 2010 is one version of several mAP versions. The mAP metric is modified over the past competitions and literatures. The basic common concept of mAP is nothing but **the area under the recall-precision curve**. However, we cannot draw this curve continuously. Depending on how to deal with this approximation of recall-precision curve and its auc values, the mAP versions have their own methods. \nOur competition metric mAP is drawn by interpolating (recall, precision) points at some IoU values. These IoU values usually between 0.5 and 0.75 with step size 0.05. (i.e. [0.5, 0.55, 0.60, 0.65, 0.7, 0.75]) However, in this competition, we use IoU values > 0.5 as described in evaluation section. Step size is not mentioned, but mostly, the step size is 0.05. I am not sure how much the gaps between mAP values with different step sizes (0.01, 0.05, 0.1, ,,,), but It would be grateful if the hosts clarify the step size. I am taking step size 0.05 and the IoU values to calculate recall-precision curve are [0.5, 0.55, ,,, 0.95, 1.0]. Note that this mAP in PASCAL VOC 2010 make the recall-precision curve monotone decreasing form when calculating AUC of recall-precision curve. (By just taking the maximum recall among recall values at next IoU steps) The above is done for each class and the mean value of auc values for each class is the final mAP of our model.\n\nOne can find a nice detailed explanation here ([https://jonathan-hui.medium.com/map-mean-average-precision-for-object-detection-45c121a31173](https://jonathan-hui.medium.com/map-mean-average-precision-for-object-detection-45c121a31173)) and a great kernel ([https://www.kaggle.com/pestipeti/competition-metric-details-script]( https://www.kaggle.com/pestipeti/competition-metric-details-script)).\n\nFor evaluating image-level labels, the above kernel's mAP is valid, but, for this competition, it is needed  additional implementation of monotone decrease function for recall values would be enough to implement the right mAP implementation for this competition. Note that the monotone decreased recall is not included in most mAP calculation libraries sklearn and ZFTurbo's [Mean-Average-Precision-for-Boxes](https://github.com/ZFTurbo/Mean-Average-Precision-for-Boxes).\n\n## **Study-level mean Average Precision?**\n\nI am still not getting how to calculate this and more detailed explanations or formulas need to be revealed by the host IHMO. \nCurrently, I am assuming that the IoU values is replaced to confidence values and calculated in one-versus-rest manner since no IoU can be calculated. That is, for each class, we can draw recall-precision curve using points at each confidence(correspond to IoU in PASCAL VOC 2010 mAP) step. [0.5, 0.55, ,,, 0.95, 1.0] If this is true, we should make prediction results not with a single-label prediction with confidence 1, (i.e. [0.9, 0.6, 0.1, 0.3] -> \"negative 1 0 0 1 1\") but multi-label predictions (i.e. [0.9, 0.6, 0.1, 0.3] -> \"negative 0.9 0 0 1 1 typical 0.6 0 0 1 1\")\n\nAlso, I am wondering there could be multiple labels for a single study in test set. If so, are the test set labeled with confidence or just 1? From the evaluation section in overview tab, the example says it is possible to make PredictionString contains multi labels. Before the LB updates, some people say a study can have multiple images and thus, multiple study-level labels can be predicted for a study (from discussion [here](https://www.kaggle.com/c/siim-covid19-detection/discussion/240329)). However, now in the test set there are only one image according to the host [here](https://www.kaggle.com/c/siim-covid19-detection/discussion/246597). Then, do we have single study-level label in test set as the training set does?\n\nThank you for reading and any comments will be welcomed :)",
      "votes": null
    },
    {
      "id": "1362916",
      "postDate": "06/23/2021 18:13:12",
      "content": "<p>Hi! Good questions.</p>\n<p>1) There is only one IoU threshold - <code>0.5</code>. The AP calculation is done using a P/R curve - specifically, counting precision at each unique recall value. There is clarification around this in <a href=\"https://storage.googleapis.com/kaggle-media/competitions/SIIM2021/VOC2012_doc.pdf\" target=\"_blank\">this document</a>, which is also linked in the Evaluation page.</p>\n<p>2) For the study-level labels, to maximize your score, you should be making ONE correct prediction. Making multiple predictions will result in false positives. Confidence scores are used only for sorting predictions (i.e. so that your most confident prediction will be tried against the ground truth first) and have no bearing on the correctness of the prediction outside of that, or on additional predictions being counted as FPs. Because we use placeholder bounding boxes the IoU should always be 1.0, so getting the class correct is key.</p>\n<p>Hopefully the linked document helps. Note that there is also MATLAB evaluation code associated with that document, which may help clarify some of the issues in a more concrete way.</p>",
      "rawMarkdown": "Hi! Good questions.\n\n1) There is only one IoU threshold - `0.5`. The AP calculation is done using a P/R curve - specifically, counting precision at each unique recall value. There is clarification around this in [this document](https://storage.googleapis.com/kaggle-media/competitions/SIIM2021/VOC2012_doc.pdf), which is also linked in the Evaluation page.\n\n2) For the study-level labels, to maximize your score, you should be making ONE correct prediction. Making multiple predictions will result in false positives. Confidence scores are used only for sorting predictions (i.e. so that your most confident prediction will be tried against the ground truth first) and have no bearing on the correctness of the prediction outside of that, or on additional predictions being counted as FPs. Because we use placeholder bounding boxes the IoU should always be 1.0, so getting the class correct is key.\n\nHopefully the linked document helps. Note that there is also MATLAB evaluation code associated with that document, which may help clarify some of the issues in a more concrete way.",
      "votes": null
    },
    {
      "id": "1363379",
      "postDate": "06/24/2021 06:09:15",
      "content": "<p>Thank you for the kind and quick response. <br>\nI thoroughly re-read the document in the link and your comment, I could realize which part is missing in my understanding. I was confused the role of IoU during mAP calculation and PR-curve calculation. <br>\nNow I think I finally get the right PR-curve and AP implementations. thank you so much :))</p>",
      "rawMarkdown": "Thank you for the kind and quick response. \nI thoroughly re-read the document in the link and your comment, I could realize which part is missing in my understanding. I was confused the role of IoU during mAP calculation and PR-curve calculation. \nNow I think I finally get the right PR-curve and AP implementations. thank you so much :))",
      "votes": null
    },
    {
      "id": "1364453",
      "postDate": "06/25/2021 03:10:44",
      "content": "<p>Hello! after some submissions, I have another question.<br>\nPublic LB gives better results when I gave prediction for all four classes with sigmoid confidence values (LB 0.442) than argmax single prediction with confidence value 1 (LB 0.318).</p>\n<p>for instance, if we get [0.7, 0.5, 0.4, 0.1],</p>\n<pre><code>negative 0.7 0 0 1 1 typical 0.5 0 0 1 1 indeterminate 0.4 0 0 1 1 atypical 0.1 0 0 1 1 # Higher LB with this row's format\nnegative 1 0 0 1 1\n</code></pre>\n<p>are there any internal thresholding process with confidence value for study-level prediction? According to your comment, the false positive will be always 3. and IMHO it should be harmful to the prediction results.</p>",
      "rawMarkdown": "Hello! after some submissions, I have another question.\nPublic LB gives better results when I gave prediction for all four classes with sigmoid confidence values (LB 0.442) than argmax single prediction with confidence value 1 (LB 0.318).\n\nfor instance, if we get [0.7, 0.5, 0.4, 0.1],\n```\nnegative 0.7 0 0 1 1 typical 0.5 0 0 1 1 indeterminate 0.4 0 0 1 1 atypical 0.1 0 0 1 1 # Higher LB with this row's format\nnegative 1 0 0 1 1\n```\n\nare there any internal thresholding process with confidence value for study-level prediction? According to your comment, the false positive will be always 3. and IMHO it should be harmful to the prediction results.",
      "votes": null
    },
    {
      "id": "1398775",
      "postDate": "07/24/2021 13:28:47",
      "content": "<p>Hi, </p>\n<p>Originally I thought the so-called mAP is calculated as the mean of the precision values at each IoU threshold, but now for this competition, mAP calculation is actually using one single loU threshold 0.5 right? This is called mean average precision just becuase we are averaging the precision when we are selecting different recall values, instead of different loU, do I get it right? Thanks!</p>",
      "rawMarkdown": "Hi, \n\nOriginally I thought the so-called mAP is calculated as the mean of the precision values at each IoU threshold, but now for this competition, mAP calculation is actually using one single loU threshold 0.5 right? This is called mean average precision just becuase we are averaging the precision when we are selecting different recall values, instead of different loU, do I get it right? Thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1362916,
      "author_name": "philculliton",
      "author_url": "",
      "post_date": "06/23/2021 18:13:12",
      "content": "<p>Hi! Good questions.</p>\n<p>1) There is only one IoU threshold - <code>0.5</code>. The AP calculation is done using a P/R curve - specifically, counting precision at each unique recall value. There is clarification around this in <a href=\"https://storage.googleapis.com/kaggle-media/competitions/SIIM2021/VOC2012_doc.pdf\" target=\"_blank\">this document</a>, which is also linked in the Evaluation page.</p>\n<p>2) For the study-level labels, to maximize your score, you should be making ONE correct prediction. Making multiple predictions will result in false positives. Confidence scores are used only for sorting predictions (i.e. so that your most confident prediction will be tried against the ground truth first) and have no bearing on the correctness of the prediction outside of that, or on additional predictions being counted as FPs. Because we use placeholder bounding boxes the IoU should always be 1.0, so getting the class correct is key.</p>\n<p>Hopefully the linked document helps. Note that there is also MATLAB evaluation code associated with that document, which may help clarify some of the issues in a more concrete way.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1363379,
          "author_name": "jihunlorenzopark",
          "author_url": "",
          "post_date": "06/24/2021 06:09:15",
          "content": "<p>Thank you for the kind and quick response. <br>\nI thoroughly re-read the document in the link and your comment, I could realize which part is missing in my understanding. I was confused the role of IoU during mAP calculation and PR-curve calculation. <br>\nNow I think I finally get the right PR-curve and AP implementations. thank you so much :))</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1364453,
          "author_name": "jihunlorenzopark",
          "author_url": "",
          "post_date": "06/25/2021 03:10:44",
          "content": "<p>Hello! after some submissions, I have another question.<br>\nPublic LB gives better results when I gave prediction for all four classes with sigmoid confidence values (LB 0.442) than argmax single prediction with confidence value 1 (LB 0.318).</p>\n<p>for instance, if we get [0.7, 0.5, 0.4, 0.1],</p>\n<pre><code>negative 0.7 0 0 1 1 typical 0.5 0 0 1 1 indeterminate 0.4 0 0 1 1 atypical 0.1 0 0 1 1 # Higher LB with this row's format\nnegative 1 0 0 1 1\n</code></pre>\n<p>are there any internal thresholding process with confidence value for study-level prediction? According to your comment, the false positive will be always 3. and IMHO it should be harmful to the prediction results.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1398775,
          "author_name": "yus002",
          "author_url": "",
          "post_date": "07/24/2021 13:28:47",
          "content": "<p>Hi, </p>\n<p>Originally I thought the so-called mAP is calculated as the mean of the precision values at each IoU threshold, but now for this competition, mAP calculation is actually using one single loU threshold 0.5 right? This is called mean average precision just becuase we are averaging the precision when we are selecting different recall values, instead of different loU, do I get it right? Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1362763": "As this is my first time to use mAP metric, I am striving to implement the exact competition metric in my local CV settings as close as possible to LB. I would like to share what I understand til now and to discuss confusing points, especially about study-level mAP. Would be appreciated to correct me if anything wrongs below and clarify few things.\n\n## **Image-level mean Average Precision**\n\nAccording to the hosts, the calculation is based on PASCAL VOC 2010 competition metric. The PASCAL VOC 2010 is one version of several mAP versions. The mAP metric is modified over the past competitions and literatures. The basic common concept of mAP is nothing but **the area under the recall-precision curve**. However, we cannot draw this curve continuously. Depending on how to deal with this approximation of recall-precision curve and its auc values, the mAP versions have their own methods. \nOur competition metric mAP is drawn by interpolating (recall, precision) points at some IoU values. These IoU values usually between 0.5 and 0.75 with step size 0.05. (i.e. [0.5, 0.55, 0.60, 0.65, 0.7, 0.75]) However, in this competition, we use IoU values > 0.5 as described in evaluation section. Step size is not mentioned, but mostly, the step size is 0.05. I am not sure how much the gaps between mAP values with different step sizes (0.01, 0.05, 0.1, ,,,), but It would be grateful if the hosts clarify the step size. I am taking step size 0.05 and the IoU values to calculate recall-precision curve are [0.5, 0.55, ,,, 0.95, 1.0]. Note that this mAP in PASCAL VOC 2010 make the recall-precision curve monotone decreasing form when calculating AUC of recall-precision curve. (By just taking the maximum recall among recall values at next IoU steps) The above is done for each class and the mean value of auc values for each class is the final mAP of our model.\n\nOne can find a nice detailed explanation here ([https://jonathan-hui.medium.com/map-mean-average-precision-for-object-detection-45c121a31173](https://jonathan-hui.medium.com/map-mean-average-precision-for-object-detection-45c121a31173)) and a great kernel ([https://www.kaggle.com/pestipeti/competition-metric-details-script]( https://www.kaggle.com/pestipeti/competition-metric-details-script)).\n\nFor evaluating image-level labels, the above kernel's mAP is valid, but, for this competition, it is needed  additional implementation of monotone decrease function for recall values would be enough to implement the right mAP implementation for this competition. Note that the monotone decreased recall is not included in most mAP calculation libraries sklearn and ZFTurbo's [Mean-Average-Precision-for-Boxes](https://github.com/ZFTurbo/Mean-Average-Precision-for-Boxes).\n\n## **Study-level mean Average Precision?**\n\nI am still not getting how to calculate this and more detailed explanations or formulas need to be revealed by the host IHMO. \nCurrently, I am assuming that the IoU values is replaced to confidence values and calculated in one-versus-rest manner since no IoU can be calculated. That is, for each class, we can draw recall-precision curve using points at each confidence(correspond to IoU in PASCAL VOC 2010 mAP) step. [0.5, 0.55, ,,, 0.95, 1.0] If this is true, we should make prediction results not with a single-label prediction with confidence 1, (i.e. [0.9, 0.6, 0.1, 0.3] -> \"negative 1 0 0 1 1\") but multi-label predictions (i.e. [0.9, 0.6, 0.1, 0.3] -> \"negative 0.9 0 0 1 1 typical 0.6 0 0 1 1\")\n\nAlso, I am wondering there could be multiple labels for a single study in test set. If so, are the test set labeled with confidence or just 1? From the evaluation section in overview tab, the example says it is possible to make PredictionString contains multi labels. Before the LB updates, some people say a study can have multiple images and thus, multiple study-level labels can be predicted for a study (from discussion [here](https://www.kaggle.com/c/siim-covid19-detection/discussion/240329)). However, now in the test set there are only one image according to the host [here](https://www.kaggle.com/c/siim-covid19-detection/discussion/246597). Then, do we have single study-level label in test set as the training set does?\n\nThank you for reading and any comments will be welcomed :)",
    "1362916": "Hi! Good questions.\n\n1) There is only one IoU threshold - `0.5`. The AP calculation is done using a P/R curve - specifically, counting precision at each unique recall value. There is clarification around this in [this document](https://storage.googleapis.com/kaggle-media/competitions/SIIM2021/VOC2012_doc.pdf), which is also linked in the Evaluation page.\n\n2) For the study-level labels, to maximize your score, you should be making ONE correct prediction. Making multiple predictions will result in false positives. Confidence scores are used only for sorting predictions (i.e. so that your most confident prediction will be tried against the ground truth first) and have no bearing on the correctness of the prediction outside of that, or on additional predictions being counted as FPs. Because we use placeholder bounding boxes the IoU should always be 1.0, so getting the class correct is key.\n\nHopefully the linked document helps. Note that there is also MATLAB evaluation code associated with that document, which may help clarify some of the issues in a more concrete way.",
    "1363379": "Thank you for the kind and quick response. \nI thoroughly re-read the document in the link and your comment, I could realize which part is missing in my understanding. I was confused the role of IoU during mAP calculation and PR-curve calculation. \nNow I think I finally get the right PR-curve and AP implementations. thank you so much :))",
    "1364453": "Hello! after some submissions, I have another question.\nPublic LB gives better results when I gave prediction for all four classes with sigmoid confidence values (LB 0.442) than argmax single prediction with confidence value 1 (LB 0.318).\n\nfor instance, if we get [0.7, 0.5, 0.4, 0.1],\n```\nnegative 0.7 0 0 1 1 typical 0.5 0 0 1 1 indeterminate 0.4 0 0 1 1 atypical 0.1 0 0 1 1 # Higher LB with this row's format\nnegative 1 0 0 1 1\n```\n\nare there any internal thresholding process with confidence value for study-level prediction? According to your comment, the false positive will be always 3. and IMHO it should be harmful to the prediction results.",
    "1398775": "Hi, \n\nOriginally I thought the so-called mAP is calculated as the mean of the precision values at each IoU threshold, but now for this competition, mAP calculation is actually using one single loU threshold 0.5 right? This is called mean average precision just becuase we are averaging the precision when we are selecting different recall values, instead of different loU, do I get it right? Thanks!"
  },
  "source": "meta"
}