{
  "id": 304143,
  "title": "Difference Between AP and this Competition's Evaluation Metrics",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/304143",
  "author_name": "",
  "post_date": "2022-01-31T00:33:44.179607200Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<h1>In short</h1>\n<p>AP and this competition's evaluation metrics differs as follows:</p>\n<ul>\n<li>AP considers prediction order of your model, but this competition's metrics doesn't (at least in a limited way)</li>\n<li>AP don't care FPs if the confidence of them are sufficiently less. AP only considers your model's prediction's order of confidence. This means AP is indifferent to confidence threshold.</li>\n</ul>\n<h1>Example</h1>\n<p>Assume there are two models named A and B.<br>\nThe first model predicts TPs with higher confidence than that of FPs, whereas the second model predicts TPs with lower confidence than that of FPs (Table 1). </p>\n<p>The APs calculated for model A and B are 1.0 and 0.5 respectively (Table 2 right and Fig 1-2).<br>\nSince AP considers your model's prediction order, it evaluates higher score for model A than model B.</p>\n<p>On the contrary, the calculated F2 score for A and B are the same, i.e. F2=0.83 (Table 2 left).</p>\n<p>Thus, this competition metric does not take into account the order of prediction of your model. This means the high-mAP model doesn't necessarily performs better on the competition LB.</p>\n<p>(Strictly speaking, the higher mAP model tends to perform better on LB, but the model of enhanced mAP doesn't necessarily performs well on LB if the enhancement is about the order of your model's prediction.)</p>\n<h1>Reference</h1>\n<ul>\n<li>[1] calculation of AP/mAP: <a href=\"https://jonathan-hui.medium.com/map-mean-average-precision-for-object-detection-45c121a31173\" target=\"_blank\">https://jonathan-hui.medium.com/map-mean-average-precision-for-object-detection-45c121a31173</a></li>\n</ul>\n<hr>\n<p><strong>Table 1.</strong>: The model's prediction. It only shows order of the prediction instead of confidence for convenience.<br>\n<a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/6Rg2vrj/Screen-Shot-2022-01-31-at-8-59-33.png\" alt=\"Screen-Shot-2022-01-31-at-8-59-33\"></a></p>\n<p><strong>Table 2.</strong>: left: calculated AP for model A and B. right: calculated F2 for both model A and B.<br>\n<a href=\"https://ibb.co/smnQ295\"><img src=\"https://i.ibb.co/Gxyk0Qs/Screen-Shot-2022-01-31-at-8-59-48.png\" alt=\"Screen-Shot-2022-01-31-at-8-59-48\"></a></p>\n<p><strong>Fig 1.</strong>: AP for model A visualized. The AP is area under the red line.</p>\n<p><a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/d2twD4b/AP-for-Model-A.png\" alt=\"AP-for-Model-A\"></a></p>\n<p><strong>Fig 2.</strong>: AP for model B visualized. The AP is area under the red line.</p>\n<p><a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/5ThKCQM/AP-for-Model-B.png\" alt=\"AP-for-Model-B\"></a></p>",
  "messages": [
    {
      "id": "1669802",
      "postDate": "01/31/2022 00:33:44",
      "content": "<h1>In short</h1>\n<p>AP and this competition's evaluation metrics differs as follows:</p>\n<ul>\n<li>AP considers prediction order of your model, but this competition's metrics doesn't (at least in a limited way)</li>\n<li>AP don't care FPs if the confidence of them are sufficiently less. AP only considers your model's prediction's order of confidence. This means AP is indifferent to confidence threshold.</li>\n</ul>\n<h1>Example</h1>\n<p>Assume there are two models named A and B.<br>\nThe first model predicts TPs with higher confidence than that of FPs, whereas the second model predicts TPs with lower confidence than that of FPs (Table 1). </p>\n<p>The APs calculated for model A and B are 1.0 and 0.5 respectively (Table 2 right and Fig 1-2).<br>\nSince AP considers your model's prediction order, it evaluates higher score for model A than model B.</p>\n<p>On the contrary, the calculated F2 score for A and B are the same, i.e. F2=0.83 (Table 2 left).</p>\n<p>Thus, this competition metric does not take into account the order of prediction of your model. This means the high-mAP model doesn't necessarily performs better on the competition LB.</p>\n<p>(Strictly speaking, the higher mAP model tends to perform better on LB, but the model of enhanced mAP doesn't necessarily performs well on LB if the enhancement is about the order of your model's prediction.)</p>\n<h1>Reference</h1>\n<ul>\n<li>[1] calculation of AP/mAP: <a href=\"https://jonathan-hui.medium.com/map-mean-average-precision-for-object-detection-45c121a31173\" target=\"_blank\">https://jonathan-hui.medium.com/map-mean-average-precision-for-object-detection-45c121a31173</a></li>\n</ul>\n<hr>\n<p><strong>Table 1.</strong>: The model's prediction. It only shows order of the prediction instead of confidence for convenience.<br>\n<a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/6Rg2vrj/Screen-Shot-2022-01-31-at-8-59-33.png\" alt=\"Screen-Shot-2022-01-31-at-8-59-33\"></a></p>\n<p><strong>Table 2.</strong>: left: calculated AP for model A and B. right: calculated F2 for both model A and B.<br>\n<a href=\"https://ibb.co/smnQ295\"><img src=\"https://i.ibb.co/Gxyk0Qs/Screen-Shot-2022-01-31-at-8-59-48.png\" alt=\"Screen-Shot-2022-01-31-at-8-59-48\"></a></p>\n<p><strong>Fig 1.</strong>: AP for model A visualized. The AP is area under the red line.</p>\n<p><a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/d2twD4b/AP-for-Model-A.png\" alt=\"AP-for-Model-A\"></a></p>\n<p><strong>Fig 2.</strong>: AP for model B visualized. The AP is area under the red line.</p>\n<p><a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/5ThKCQM/AP-for-Model-B.png\" alt=\"AP-for-Model-B\"></a></p>",
      "rawMarkdown": "# In short \n\nAP and this competition's evaluation metrics differs as follows:\n\n* AP considers prediction order of your model, but this competition's metrics doesn't (at least in a limited way)\n* AP don't care FPs if the confidence of them are sufficiently less. AP only considers your model's prediction's order of confidence. This means AP is indifferent to confidence threshold.\n\n# Example\n\nAssume there are two models named A and B.\nThe first model predicts TPs with higher confidence than that of FPs, whereas the second model predicts TPs with lower confidence than that of FPs (Table 1). \n\nThe APs calculated for model A and B are 1.0 and 0.5 respectively (Table 2 right and Fig 1-2).\nSince AP considers your model's prediction order, it evaluates higher score for model A than model B.\n\nOn the contrary, the calculated F2 score for A and B are the same, i.e. F2=0.83 (Table 2 left).\n\nThus, this competition metric does not take into account the order of prediction of your model. This means the high-mAP model doesn't necessarily performs better on the competition LB.\n\n(Strictly speaking, the higher mAP model tends to perform better on LB, but the model of enhanced mAP doesn't necessarily performs well on LB if the enhancement is about the order of your model's prediction.)\n\n# Reference\n\n* [1] calculation of AP/mAP: https://jonathan-hui.medium.com/map-mean-average-precision-for-object-detection-45c121a31173\n\n----\n\n**Table 1.**: The model's prediction. It only shows order of the prediction instead of confidence for convenience.\n<a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/6Rg2vrj/Screen-Shot-2022-01-31-at-8-59-33.png\" alt=\"Screen-Shot-2022-01-31-at-8-59-33\" border=\"0\"></a>\n\n**Table 2.**: left: calculated AP for model A and B. right: calculated F2 for both model A and B.\n<a href=\"https://ibb.co/smnQ295\"><img src=\"https://i.ibb.co/Gxyk0Qs/Screen-Shot-2022-01-31-at-8-59-48.png\" alt=\"Screen-Shot-2022-01-31-at-8-59-48\" border=\"0\"></a>\n\n**Fig 1.**: AP for model A visualized. The AP is area under the red line.\n\n<a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/d2twD4b/AP-for-Model-A.png\" alt=\"AP-for-Model-A\" border=\"0\"></a>\n\n**Fig 2.**: AP for model B visualized. The AP is area under the red line.\n\n<a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/5ThKCQM/AP-for-Model-B.png\" alt=\"AP-for-Model-B\" border=\"0\"></a>",
      "votes": null
    },
    {
      "id": "1669807",
      "postDate": "01/31/2022 00:47:37",
      "content": "<p>What I really want to prove is \"Does mAP well-reflects the performance of our model on LB?\".</p>\n<p>You know, mAP is convenient metric because we doesn't have to consider confidence threshold while training our model.</p>\n<p>But one question occurs: \"Does the model selected by mAP really performs well on the LB?\".<br>\nIf it doesn't, we have to consider better model-selection criterion instead of mAP.</p>\n<p>Adopting this competition metrics (based on F2) is one option. But then, we have to choose proper confidence threshold before we train the model. That is a nuisance.</p>",
      "rawMarkdown": "What I really want to prove is \"Does mAP well-reflects the performance of our model on LB?\".\n\nYou know, mAP is convenient metric because we doesn't have to consider confidence threshold while training our model.\n\nBut one question occurs: \"Does the model selected by mAP really performs well on the LB?\".\nIf it doesn't, we have to consider better model-selection criterion instead of mAP.\n\nAdopting this competition metrics (based on F2) is one option. But then, we have to choose proper confidence threshold before we train the model. That is a nuisance.",
      "votes": null
    },
    {
      "id": "1669817",
      "postDate": "01/31/2022 01:10:16",
      "content": "<h1>Candidates of better model-selection criterion</h1>\n<h2>Max value of F2 on the F2-confidence curve</h2>\n<p>This is a high-return strategy because it tries to maximize the maximum return: maximizing F2 if your confidence-threshold is optimal. It's also a high-risk strategy because this metrics won't assure your model performs also well if the confidence threshold is not optimal.</p>\n<h2>Area under the F2-confidence curve</h2>\n<p>That strategy tries to maximize the expected value of F2 with respect to confidence threshold. <br>\nIf F2 is higher for all the confidence threshold, the model is expected to perform well on the LB if your confidence threshold is not optimal.<br>\nOn the contrary, it doesn't assure your model performs the best on the LB even if the confidence threshold is the optimal value.<br>\nThus is a low-risk, low-return strategy.</p>",
      "rawMarkdown": "# Candidates of better model-selection criterion\n\n## Max value of F2 on the F2-confidence curve\n\nThis is a high-return strategy because it tries to maximize the maximum return: maximizing F2 if your confidence-threshold is optimal. It's also a high-risk strategy because this metrics won't assure your model performs also well if the confidence threshold is not optimal.\n\n## Area under the F2-confidence curve\n\nThat strategy tries to maximize the expected value of F2 with respect to confidence threshold. \nIf F2 is higher for all the confidence threshold, the model is expected to perform well on the LB if your confidence threshold is not optimal.\nOn the contrary, it doesn't assure your model performs the best on the LB even if the confidence threshold is the optimal value.\nThus is a low-risk, low-return strategy.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1669807,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "01/31/2022 00:47:37",
      "content": "<p>What I really want to prove is \"Does mAP well-reflects the performance of our model on LB?\".</p>\n<p>You know, mAP is convenient metric because we doesn't have to consider confidence threshold while training our model.</p>\n<p>But one question occurs: \"Does the model selected by mAP really performs well on the LB?\".<br>\nIf it doesn't, we have to consider better model-selection criterion instead of mAP.</p>\n<p>Adopting this competition metrics (based on F2) is one option. But then, we have to choose proper confidence threshold before we train the model. That is a nuisance.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1669817,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "01/31/2022 01:10:16",
          "content": "<h1>Candidates of better model-selection criterion</h1>\n<h2>Max value of F2 on the F2-confidence curve</h2>\n<p>This is a high-return strategy because it tries to maximize the maximum return: maximizing F2 if your confidence-threshold is optimal. It's also a high-risk strategy because this metrics won't assure your model performs also well if the confidence threshold is not optimal.</p>\n<h2>Area under the F2-confidence curve</h2>\n<p>That strategy tries to maximize the expected value of F2 with respect to confidence threshold. <br>\nIf F2 is higher for all the confidence threshold, the model is expected to perform well on the LB if your confidence threshold is not optimal.<br>\nOn the contrary, it doesn't assure your model performs the best on the LB even if the confidence threshold is the optimal value.<br>\nThus is a low-risk, low-return strategy.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1669802": "# In short \n\nAP and this competition's evaluation metrics differs as follows:\n\n* AP considers prediction order of your model, but this competition's metrics doesn't (at least in a limited way)\n* AP don't care FPs if the confidence of them are sufficiently less. AP only considers your model's prediction's order of confidence. This means AP is indifferent to confidence threshold.\n\n# Example\n\nAssume there are two models named A and B.\nThe first model predicts TPs with higher confidence than that of FPs, whereas the second model predicts TPs with lower confidence than that of FPs (Table 1). \n\nThe APs calculated for model A and B are 1.0 and 0.5 respectively (Table 2 right and Fig 1-2).\nSince AP considers your model's prediction order, it evaluates higher score for model A than model B.\n\nOn the contrary, the calculated F2 score for A and B are the same, i.e. F2=0.83 (Table 2 left).\n\nThus, this competition metric does not take into account the order of prediction of your model. This means the high-mAP model doesn't necessarily performs better on the competition LB.\n\n(Strictly speaking, the higher mAP model tends to perform better on LB, but the model of enhanced mAP doesn't necessarily performs well on LB if the enhancement is about the order of your model's prediction.)\n\n# Reference\n\n* [1] calculation of AP/mAP: https://jonathan-hui.medium.com/map-mean-average-precision-for-object-detection-45c121a31173\n\n----\n\n**Table 1.**: The model's prediction. It only shows order of the prediction instead of confidence for convenience.\n<a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/6Rg2vrj/Screen-Shot-2022-01-31-at-8-59-33.png\" alt=\"Screen-Shot-2022-01-31-at-8-59-33\" border=\"0\"></a>\n\n**Table 2.**: left: calculated AP for model A and B. right: calculated F2 for both model A and B.\n<a href=\"https://ibb.co/smnQ295\"><img src=\"https://i.ibb.co/Gxyk0Qs/Screen-Shot-2022-01-31-at-8-59-48.png\" alt=\"Screen-Shot-2022-01-31-at-8-59-48\" border=\"0\"></a>\n\n**Fig 1.**: AP for model A visualized. The AP is area under the red line.\n\n<a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/d2twD4b/AP-for-Model-A.png\" alt=\"AP-for-Model-A\" border=\"0\"></a>\n\n**Fig 2.**: AP for model B visualized. The AP is area under the red line.\n\n<a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/5ThKCQM/AP-for-Model-B.png\" alt=\"AP-for-Model-B\" border=\"0\"></a>",
    "1669807": "What I really want to prove is \"Does mAP well-reflects the performance of our model on LB?\".\n\nYou know, mAP is convenient metric because we doesn't have to consider confidence threshold while training our model.\n\nBut one question occurs: \"Does the model selected by mAP really performs well on the LB?\".\nIf it doesn't, we have to consider better model-selection criterion instead of mAP.\n\nAdopting this competition metrics (based on F2) is one option. But then, we have to choose proper confidence threshold before we train the model. That is a nuisance.",
    "1669817": "# Candidates of better model-selection criterion\n\n## Max value of F2 on the F2-confidence curve\n\nThis is a high-return strategy because it tries to maximize the maximum return: maximizing F2 if your confidence-threshold is optimal. It's also a high-risk strategy because this metrics won't assure your model performs also well if the confidence threshold is not optimal.\n\n## Area under the F2-confidence curve\n\nThat strategy tries to maximize the expected value of F2 with respect to confidence threshold. \nIf F2 is higher for all the confidence threshold, the model is expected to perform well on the LB if your confidence threshold is not optimal.\nOn the contrary, it doesn't assure your model performs the best on the LB even if the confidence threshold is the optimal value.\nThus is a low-risk, low-return strategy."
  },
  "source": "meta"
}