{
  "id": 315274,
  "title": "Why do we sort confidence scores when calculating MAP in Object Detection?",
  "url": "/competitions/happy-whale-and-dolphin/discussion/315274",
  "author_name": "",
  "post_date": "2022-03-27T08:37:39.413642200Z",
  "votes": 3,
  "comment_count": 7,
  "views": 0,
  "content": "<p>This is slightly different from the metric in this competition, but since I wanted to read up more on MAP as a metric in general, I sieved through multiple articles like <a href=\"https://jonathan-hui.medium.com/map-mean-average-precision-for-object-detection-45c121a31173\" target=\"_blank\">the one by Jonathon</a> and <a href=\"https://www.kaggle.com/code/its7171/map-understanding-with-code-and-its-tips?scriptVersionId=54650292\" target=\"_blank\">here by Kaggler Tito</a>.</p>\n<p>One part I cannot grasp immediately is the need to sort the confidence scores of the bounding box in descending order. What is the reason behind sorting here? Amongst the &gt; 10 articles I read, only <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> notebook mentioned about sorting, but I cannot grasp it also.</p>\n<p>Is it because we want to have a cumulative precision and recall based off the cut-off threshold by the confidence scores since we are discarding all predictions under a certain confidence score? </p>",
  "messages": [
    {
      "id": "1736385",
      "postDate": "03/27/2022 08:37:39",
      "content": "<p>This is slightly different from the metric in this competition, but since I wanted to read up more on MAP as a metric in general, I sieved through multiple articles like <a href=\"https://jonathan-hui.medium.com/map-mean-average-precision-for-object-detection-45c121a31173\" target=\"_blank\">the one by Jonathon</a> and <a href=\"https://www.kaggle.com/code/its7171/map-understanding-with-code-and-its-tips?scriptVersionId=54650292\" target=\"_blank\">here by Kaggler Tito</a>.</p>\n<p>One part I cannot grasp immediately is the need to sort the confidence scores of the bounding box in descending order. What is the reason behind sorting here? Amongst the &gt; 10 articles I read, only <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> notebook mentioned about sorting, but I cannot grasp it also.</p>\n<p>Is it because we want to have a cumulative precision and recall based off the cut-off threshold by the confidence scores since we are discarding all predictions under a certain confidence score? </p>",
      "rawMarkdown": "This is slightly different from the metric in this competition, but since I wanted to read up more on MAP as a metric in general, I sieved through multiple articles like [the one by Jonathon](https://jonathan-hui.medium.com/map-mean-average-precision-for-object-detection-45c121a31173) and [here by Kaggler Tito](https://www.kaggle.com/code/its7171/map-understanding-with-code-and-its-tips?scriptVersionId=54650292).\n\nOne part I cannot grasp immediately is the need to sort the confidence scores of the bounding box in descending order. What is the reason behind sorting here? Amongst the > 10 articles I read, only @its7171 notebook mentioned about sorting, but I cannot grasp it also.\n\nIs it because we want to have a cumulative precision and recall based off the cut-off threshold by the confidence scores since we are discarding all predictions under a certain confidence score?",
      "votes": null
    },
    {
      "id": "1736396",
      "postDate": "03/27/2022 09:00:51",
      "content": "<p>I feel the answer lies in the statement: 'The lower the confidence threshold, the higher the mAP we will always get', and your doubt is infact the answer to your question.</p>",
      "rawMarkdown": "I feel the answer lies in the statement: 'The lower the confidence threshold, the higher the mAP we will always get', and your doubt is infact the answer to your question.",
      "votes": null
    },
    {
      "id": "1736419",
      "postDate": "03/27/2022 09:24:19",
      "content": "<p>Thanks, this is kind of confusing for me since, amongst the 10+ articles I read, none actually mentioned the rationale of sorting. Tito's notebook mentioned though, but I cannot grasp it also, I read through a youtuber's source code on this, but they indeed sorted them (without explaining).</p>",
      "rawMarkdown": "Thanks, this is kind of confusing for me since, amongst the 10+ articles I read, none actually mentioned the rationale of sorting. Tito's notebook mentioned though, but I cannot grasp it also, I read through a youtuber's source code on this, but they indeed sorted them (without explaining).",
      "votes": null
    },
    {
      "id": "1736602",
      "postDate": "03/27/2022 14:15:06",
      "content": "<p>The metric MAP answers the questions, <br>\n\"how good is your prediction if you only guess 1 bounding box\"?, <br>\n\"how good is your prediction if you only guess 2 bounding boxes\"?<br>\n\"how good is your prediction if you only guess 3 bounding boxes\"?<br>\n\"how good is your prediction if you only guess 4 bounding boxes\"?<br>\netc etc etc<br>\nthen MAP takes the average of these questions</p>\n<p>When evaluating \"how good is your prediction if you only guess X bound boxes\", then of course we would like to use our best X bounding boxes. Our best X bounding boxes are the X bounding boxes with best confidence. Therefore we sort by confidence before answering the above questions. </p>",
      "rawMarkdown": "The metric MAP answers the questions, \n\"how good is your prediction if you only guess 1 bounding box\"?, \n\"how good is your prediction if you only guess 2 bounding boxes\"?\n\"how good is your prediction if you only guess 3 bounding boxes\"?\n\"how good is your prediction if you only guess 4 bounding boxes\"?\netc etc etc\nthen MAP takes the average of these questions\n\nWhen evaluating \"how good is your prediction if you only guess X bound boxes\", then of course we would like to use our best X bounding boxes. Our best X bounding boxes are the X bounding boxes with best confidence. Therefore we sort by confidence before answering the above questions.",
      "votes": null
    },
    {
      "id": "1736725",
      "postDate": "03/27/2022 16:45:47",
      "content": "<p>It gives the precision value from the rank of outcomes we derive.</p>",
      "rawMarkdown": "It gives the precision value from the rank of outcomes we derive.",
      "votes": null
    },
    {
      "id": "1737045",
      "postDate": "03/28/2022 04:27:39",
      "content": "<p>Thanks Chris for the clear answer. In that case, are we able to ask</p>\n<p>\"at each confidence level (threshold), what is the precision-recall score of the predictions of all bounding boxes while discarding off those below the threshold, then average them over all thresholds?\"</p>",
      "rawMarkdown": "Thanks Chris for the clear answer. In that case, are we able to ask\n\n\"at each confidence level (threshold), what is the precision-recall score of the predictions of all bounding boxes while discarding off those below the threshold, then average them over all thresholds?\"",
      "votes": null
    },
    {
      "id": "1737049",
      "postDate": "03/28/2022 04:33:20",
      "content": "<p>Yes basically yes. The exact statement is \"\"at each confidence level (threshold), if recall increases from previous confidence level, then include this precision and average all previous included precisions\". But what you said, is basically correct.</p>",
      "rawMarkdown": "Yes basically yes. The exact statement is \"\"at each confidence level (threshold), if recall increases from previous confidence level, then include this precision and average all previous included precisions\". But what you said, is basically correct.",
      "votes": null
    },
    {
      "id": "1737147",
      "postDate": "03/28/2022 06:30:57",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks a lot for this, it clears some air. I will go and read the source code again and update my understanding… </p>",
      "rawMarkdown": "cdeotte Thanks a lot for this, it clears some air. I will go and read the source code again and update my understanding...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1736396,
      "author_name": "gianetan",
      "author_url": "",
      "post_date": "03/27/2022 09:00:51",
      "content": "<p>I feel the answer lies in the statement: 'The lower the confidence threshold, the higher the mAP we will always get', and your doubt is infact the answer to your question.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1736419,
          "author_name": "reighns",
          "author_url": "",
          "post_date": "03/27/2022 09:24:19",
          "content": "<p>Thanks, this is kind of confusing for me since, amongst the 10+ articles I read, none actually mentioned the rationale of sorting. Tito's notebook mentioned though, but I cannot grasp it also, I read through a youtuber's source code on this, but they indeed sorted them (without explaining).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1736602,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "03/27/2022 14:15:06",
      "content": "<p>The metric MAP answers the questions, <br>\n\"how good is your prediction if you only guess 1 bounding box\"?, <br>\n\"how good is your prediction if you only guess 2 bounding boxes\"?<br>\n\"how good is your prediction if you only guess 3 bounding boxes\"?<br>\n\"how good is your prediction if you only guess 4 bounding boxes\"?<br>\netc etc etc<br>\nthen MAP takes the average of these questions</p>\n<p>When evaluating \"how good is your prediction if you only guess X bound boxes\", then of course we would like to use our best X bounding boxes. Our best X bounding boxes are the X bounding boxes with best confidence. Therefore we sort by confidence before answering the above questions. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1737045,
          "author_name": "reighns",
          "author_url": "",
          "post_date": "03/28/2022 04:27:39",
          "content": "<p>Thanks Chris for the clear answer. In that case, are we able to ask</p>\n<p>\"at each confidence level (threshold), what is the precision-recall score of the predictions of all bounding boxes while discarding off those below the threshold, then average them over all thresholds?\"</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1737049,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "03/28/2022 04:33:20",
          "content": "<p>Yes basically yes. The exact statement is \"\"at each confidence level (threshold), if recall increases from previous confidence level, then include this precision and average all previous included precisions\". But what you said, is basically correct.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1737147,
          "author_name": "reighns",
          "author_url": "",
          "post_date": "03/28/2022 06:30:57",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks a lot for this, it clears some air. I will go and read the source code again and update my understanding… </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1736725,
      "author_name": "jainishsavalia",
      "author_url": "",
      "post_date": "03/27/2022 16:45:47",
      "content": "<p>It gives the precision value from the rank of outcomes we derive.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1736385": "This is slightly different from the metric in this competition, but since I wanted to read up more on MAP as a metric in general, I sieved through multiple articles like [the one by Jonathon](https://jonathan-hui.medium.com/map-mean-average-precision-for-object-detection-45c121a31173) and [here by Kaggler Tito](https://www.kaggle.com/code/its7171/map-understanding-with-code-and-its-tips?scriptVersionId=54650292).\n\nOne part I cannot grasp immediately is the need to sort the confidence scores of the bounding box in descending order. What is the reason behind sorting here? Amongst the > 10 articles I read, only @its7171 notebook mentioned about sorting, but I cannot grasp it also.\n\nIs it because we want to have a cumulative precision and recall based off the cut-off threshold by the confidence scores since we are discarding all predictions under a certain confidence score?",
    "1736396": "I feel the answer lies in the statement: 'The lower the confidence threshold, the higher the mAP we will always get', and your doubt is infact the answer to your question.",
    "1736419": "Thanks, this is kind of confusing for me since, amongst the 10+ articles I read, none actually mentioned the rationale of sorting. Tito's notebook mentioned though, but I cannot grasp it also, I read through a youtuber's source code on this, but they indeed sorted them (without explaining).",
    "1736602": "The metric MAP answers the questions, \n\"how good is your prediction if you only guess 1 bounding box\"?, \n\"how good is your prediction if you only guess 2 bounding boxes\"?\n\"how good is your prediction if you only guess 3 bounding boxes\"?\n\"how good is your prediction if you only guess 4 bounding boxes\"?\netc etc etc\nthen MAP takes the average of these questions\n\nWhen evaluating \"how good is your prediction if you only guess X bound boxes\", then of course we would like to use our best X bounding boxes. Our best X bounding boxes are the X bounding boxes with best confidence. Therefore we sort by confidence before answering the above questions.",
    "1736725": "It gives the precision value from the rank of outcomes we derive.",
    "1737045": "Thanks Chris for the clear answer. In that case, are we able to ask\n\n\"at each confidence level (threshold), what is the precision-recall score of the predictions of all bounding boxes while discarding off those below the threshold, then average them over all thresholds?\"",
    "1737049": "Yes basically yes. The exact statement is \"\"at each confidence level (threshold), if recall increases from previous confidence level, then include this precision and average all previous included precisions\". But what you said, is basically correct.",
    "1737147": "cdeotte Thanks a lot for this, it clears some air. I will go and read the source code again and update my understanding..."
  },
  "source": "meta"
}