{
  "id": 64542,
  "title": "What is the role of the \"confidence\" score per bounding box?",
  "url": "/competitions/rsna-pneumonia-detection-challenge/discussion/64542",
  "author_name": "",
  "post_date": "2018-08-30T07:50:21.395800900Z",
  "votes": 16,
  "comment_count": 4,
  "views": 0,
  "content": "<p>It seems that the order of evaluating the predicted bounding boxes does not influence the precision. If we have two predictions overlapping the same ground truth bounding box, one of them will always be counted as false positive regardless of the order. \nThe only situation in which the order might play a part is when two ground truth objects are covered by one large predicted box, but in that case it is unlikely that both will have IoU greater than 0.4 (and never both greater than 0.5)\nAm I missing anything? Is it possible to publish the evaluation script?\nThanks</p>",
  "messages": [
    {
      "id": "378417",
      "postDate": "08/30/2018 07:50:21",
      "content": "<p>It seems that the order of evaluating the predicted bounding boxes does not influence the precision. If we have two predictions overlapping the same ground truth bounding box, one of them will always be counted as false positive regardless of the order. \nThe only situation in which the order might play a part is when two ground truth objects are covered by one large predicted box, but in that case it is unlikely that both will have IoU greater than 0.4 (and never both greater than 0.5)\nAm I missing anything? Is it possible to publish the evaluation script?\nThanks</p>",
      "rawMarkdown": "It seems that the order of evaluating the predicted bounding boxes does not influence the precision. If we have two predictions overlapping the same ground truth bounding box, one of them will always be counted as false positive regardless of the order. \nThe only situation in which the order might play a part is when two ground truth objects are covered by one large predicted box, but in that case it is unlikely that both will have IoU greater than 0.4 (and never both greater than 0.5)\nAm I missing anything? Is it possible to publish the evaluation script?\nThanks",
      "votes": null
    },
    {
      "id": "380168",
      "postDate": "09/01/2018 21:44:35",
      "content": "<p>Actually there is such quite general symmetrical situation. \nConsider two algorithms</p>\n\n<p><strong>The algorithm based on predicted boxes primarily</strong>\n0. TP = 0, FP = 0</p>\n\n<ol>\n<li><p>for each predicted box in the order of <code>confidence</code> get all IoU &gt; threshold target boxes - N</p></li>\n<li><p>if N == 0 : FP += 1 and exclude this particular predicted box from further consideration</p></li>\n<li><p>if N == 1  : TP += 1 and exclude these particular predicted box and target box from further consideration</p></li>\n<li><p>if N &gt; 1 : TP += 1 and exclude this particular predicted box and target box with the <code>maximum IoU</code> (?) from further consideration</p></li>\n<li><p>go to step (1)</p></li>\n</ol>\n\n<p><strong>The algorithm based on target boxes primarily</strong>\n0. TP = 0, FN = 0</p>\n\n<ol>\n<li><p>for each target box in the order of <code>area size</code> (?) get all IoU &gt; threshold predicted boxes - N</p></li>\n<li><p>if N == 0 : FN += 1 and exclude this particular target box from further consideration</p></li>\n<li><p>if N == 1  : TP += 1 and exclude these particular target box and predicted box from further consideration</p></li>\n<li><p>if N &gt; 1 : TP += 1 and exclude this particular target box and predicted box with the maximum <code>confidence</code> from further consideration</p></li>\n<li><p>go to step (1)</p></li>\n</ol>\n\n<p>Basically you have two slightly different ways to calculate TP. All the evaluation procedure and requirement to provide <code>confidence</code> for each predicted box make sense for the second algorithm of counting TP.</p>\n\n<p>But generally you are absolutely right - the script of evaluation is very needed. </p>",
      "rawMarkdown": "Actually there is such quite general symmetrical situation. \nConsider two algorithms\n\n**The algorithm based on predicted boxes primarily**\n0. TP = 0, FP = 0\n\n1. for each predicted box in the order of `confidence` get all IoU &gt; threshold target boxes - N\n\n2. if N == 0 : FP += 1 and exclude this particular predicted box from further consideration\n\n3. if N == 1  : TP += 1 and exclude these particular predicted box and target box from further consideration\n\n4. if N &gt; 1 : TP += 1 and exclude this particular predicted box and target box with the `maximum IoU` (?) from further consideration\n\n5. go to step (1)\n\n**The algorithm based on target boxes primarily**\n0. TP = 0, FN = 0\n\n1. for each target box in the order of `area size` (?) get all IoU &gt; threshold predicted boxes - N\n\n2. if N == 0 : FN += 1 and exclude this particular target box from further consideration\n\n3. if N == 1  : TP += 1 and exclude these particular target box and predicted box from further consideration\n\n4. if N &gt; 1 : TP += 1 and exclude this particular target box and predicted box with the maximum `confidence` from further consideration\n\n5. go to step (1)\n\nBasically you have two slightly different ways to calculate TP. All the evaluation procedure and requirement to provide `confidence` for each predicted box make sense for the second algorithm of counting TP.\n\nBut generally you are absolutely right - the script of evaluation is very needed.",
      "votes": null
    },
    {
      "id": "380252",
      "postDate": "09/02/2018 06:25:10",
      "content": "<p>Thanks for the detailed response. Still I cannot see what is the relevance of the confidence even in your second algorithm, because in this algorithm, after you exclude the highest confidence predicted box in stage 4, all the lower confidence predicted boxes are counted as FPs, isn't it so?\nAdministrators - please respond!</p>",
      "rawMarkdown": "Thanks for the detailed response. Still I cannot see what is the relevance of the confidence even in your second algorithm, because in this algorithm, after you exclude the highest confidence predicted box in stage 4, all the lower confidence predicted boxes are counted as FPs, isn't it so?\nAdministrators - please respond!",
      "votes": null
    },
    {
      "id": "380258",
      "postDate": "09/02/2018 06:54:36",
      "content": "<p>Actually there might be an evaluation algorithm in which the confidence does count: If the predicted bounding boxes are first ordered by their confidence and only after that are tested with the ground truth boxes, then there might be a difference in the scoring... For example assume that there are a predicted box A with high IoU and a predicted box B with low IoU (for the same ground truth of course). If A has a larger confidence than B, then we get one TP and one FP, but if A has a smaller confidence than B, we might get two FPs and one FN !. Anyways, again, the actual evaluation script is required..</p>",
      "rawMarkdown": "Actually there might be an evaluation algorithm in which the confidence does count: If the predicted bounding boxes are first ordered by their confidence and only after that are tested with the ground truth boxes, then there might be a difference in the scoring... For example assume that there are a predicted box A with high IoU and a predicted box B with low IoU (for the same ground truth of course). If A has a larger confidence than B, then we get one TP and one FP, but if A has a smaller confidence than B, we might get two FPs and one FN !. Anyways, again, the actual evaluation script is required..",
      "votes": null
    },
    {
      "id": "383342",
      "postDate": "09/08/2018 12:32:06",
      "content": "<p>I agree, this metric is tricky. <strong>We need official clarification</strong>.  </p>\n\n<p>Here is my <strong>assumption</strong>.</p>\n\n<h1>1. Threshold &gt; 0.5:</h1>\n\n<p>Then for each predicted BBox (<strong>pbox</strong>) we can have only one matching true\nBBox (<strong>tbox</strong>)(strictly speaking there is necessary condition: <em>each pair of\ntboxes must have zero intersection</em>. It is almost true for given data).</p>\n\n<p>In this case there is no ambiguity how to match pboxes with tboxes:  we can use second algorithm, proposed by Ivan (actually it doesn't matter, we can use first as well in this case).</p>\n\n<p>So we have correspondence 1 tbox to *(many) pboxes.  </p>\n\n<ul>\n<li>tbox has <code>0</code> matches --&gt; <code>TN += 1</code></li>\n<li>tbox has <code>more then 1</code> matches --&gt; <code>TP += 1</code> and other goes to <code>FP</code>, no matter which of them, because all of them correspond to 1 tbox. </li>\n</ul>\n\n<p>Thus in this case order of matching doesn't matter, as well as confidence.</p>\n\n<h1>2. Thresholds &lt;= 0.5:</h1>\n\n<p>In this case we deal with correspondence * tboxes to * pboxes. </p>\n\n<p>Theoretically can happen next situation (drawing is not accurate, rely more on inequalities):\n<img src=\"https://pp.userapi.com/c846321/v846321499/ed955/VdAJWSQ4rlI.jpg\" alt=\"situation\"></p>\n\n<p>You would like to match <code>A-B</code>, <code>C-D</code>. \nBut let's consider what thinks about it two proposed algorithms by Ivan.</p>\n\n<p><strong>First Algorithm.</strong>\nTakes first <code>A</code> (because of confidence) and matches it to <code>C</code> (because of IoU) is <code>TP</code>. <br>\nThen <code>D</code> is <code>FP</code> and <code>B</code> is <code>FN</code>. Result: <code>A-C</code></p>\n\n<p>If you swap confidence of <code>A</code> and <code>B</code>, then you will get desired result.</p>\n\n<p><strong>Second Algorithm.</strong>\nThe result of this algorithm depends on order of tboxes. <br>\nIf first <code>B</code>, then we will get desired result. <br>\nIf first <code>C</code>, then it will be matched with A (because of IoU), and we will get bad result.</p>\n\n<p>Actually it is difficult imagine, that this case will happen someday or that it can drastically affect overall result. </p>\n\n<p>I think it is more matter of design. Your result shouldn't depend on predefined order of tboxes. That is why I think they took first approach, <strong>in that case result depend only on your proposal, your algorithm</strong>. Result fully under your control.</p>\n\n<h1>Why do we need to match BBoxes at all?</h1>\n\n<p>Why can't we calculate some kind of <code>average intersection</code>?\nBecause in this case it is easy to fool metric by predicting 1 billion of identical bboxes in which you are pretty confident. So there should be instrument for punishing for different number of pboxes and tboxes.</p>\n\n<p>But there are still exists method to do it properly:\n\\( Score(tboxes, pboxes) = \\Big(\\bigcup\\limits_{tb \\in tboxes} tb \\Big) \\bigcap \\Big(\\bigcup\\limits_{pb \\in pboxes} pb \\Big) \\)</p>\n\n<p>It has one advantage, because one scrupulous radiologist will annotate 4 small bboxes, while another one -- 2 big bboxes. Score above doesn't depend on it.</p>\n\n<p>But this <code>score</code> has disadvantage too. Because bboxes contain background as well as object. It is already looks more like a segmentation, but we don't have precise annotation of object of interests and probably it is even impossible to get segmentation of opacities, where do they start?</p>\n\n<p>So from the design point of view it is better to use <code>First algorithm</code> proposed by Ivan, but intuitively I would say that last <code>Score</code> would work better for this task.</p>",
      "rawMarkdown": "I agree, this metric is tricky. **We need official clarification**.  \n\nHere is my **assumption**.\n\n1. Threshold &gt; 0.5:\n===================\n\nThen for each predicted BBox (**pbox**) we can have only one matching true\nBBox (**tbox**)(strictly speaking there is necessary condition: *each pair of\ntboxes must have zero intersection*. It is almost true for given data).\n\nIn this case there is no ambiguity how to match pboxes with tboxes:  we can use second algorithm, proposed by Ivan (actually it doesn't matter, we can use first as well in this case).\n\nSo we have correspondence 1 tbox to *(many) pboxes.  \n\n- tbox has `0` matches --&gt; `TN += 1`\n- tbox has `more then 1` matches --&gt; `TP += 1` and other goes to `FP`, no matter which of them, because all of them correspond to 1 tbox. \n\nThus in this case order of matching doesn't matter, as well as confidence.\n\n2. Thresholds &lt;= 0.5:\n====================\n\nIn this case we deal with correspondence * tboxes to * pboxes. \n\nTheoretically can happen next situation (drawing is not accurate, rely more on inequalities):\n![situation](https://pp.userapi.com/c846321/v846321499/ed955/VdAJWSQ4rlI.jpg)\n\nYou would like to match `A-B`, `C-D`. \nBut let's consider what thinks about it two proposed algorithms by Ivan.\n\n**First Algorithm.**\nTakes first `A` (because of confidence) and matches it to `C` (because of IoU) is `TP`.  \nThen `D` is `FP` and `B` is `FN`. Result: `A-C`\n\nIf you swap confidence of `A` and `B`, then you will get desired result.\n\n**Second Algorithm.**\nThe result of this algorithm depends on order of tboxes.  \nIf first `B`, then we will get desired result.  \nIf first `C`, then it will be matched with A (because of IoU), and we will get bad result.\n\nActually it is difficult imagine, that this case will happen someday or that it can drastically affect overall result. \n\nI think it is more matter of design. Your result shouldn't depend on predefined order of tboxes. That is why I think they took first approach, **in that case result depend only on your proposal, your algorithm**. Result fully under your control.\n\nWhy do we need to match BBoxes at all?\n===============================\nWhy can't we calculate some kind of `average intersection`?\nBecause in this case it is easy to fool metric by predicting 1 billion of identical bboxes in which you are pretty confident. So there should be instrument for punishing for different number of pboxes and tboxes.\n\nBut there are still exists method to do it properly:\n\\\\( Score(tboxes, pboxes) = \\Big(\\bigcup\\limits_{tb \\in tboxes} tb \\Big) \\bigcap \\Big(\\bigcup\\limits_{pb \\in pboxes} pb \\Big) \\\\)\n\nIt has one advantage, because one scrupulous radiologist will annotate 4 small bboxes, while another one -- 2 big bboxes. Score above doesn't depend on it.\n\nBut this `score` has disadvantage too. Because bboxes contain background as well as object. It is already looks more like a segmentation, but we don't have precise annotation of object of interests and probably it is even impossible to get segmentation of opacities, where do they start?\n\nSo from the design point of view it is better to use `First algorithm` proposed by Ivan, but intuitively I would say that last `Score` would work better for this task.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 380168,
      "author_name": "ivanpozd",
      "author_url": "",
      "post_date": "09/01/2018 21:44:35",
      "content": "<p>Actually there is such quite general symmetrical situation. \nConsider two algorithms</p>\n\n<p><strong>The algorithm based on predicted boxes primarily</strong>\n0. TP = 0, FP = 0</p>\n\n<ol>\n<li><p>for each predicted box in the order of <code>confidence</code> get all IoU &gt; threshold target boxes - N</p></li>\n<li><p>if N == 0 : FP += 1 and exclude this particular predicted box from further consideration</p></li>\n<li><p>if N == 1  : TP += 1 and exclude these particular predicted box and target box from further consideration</p></li>\n<li><p>if N &gt; 1 : TP += 1 and exclude this particular predicted box and target box with the <code>maximum IoU</code> (?) from further consideration</p></li>\n<li><p>go to step (1)</p></li>\n</ol>\n\n<p><strong>The algorithm based on target boxes primarily</strong>\n0. TP = 0, FN = 0</p>\n\n<ol>\n<li><p>for each target box in the order of <code>area size</code> (?) get all IoU &gt; threshold predicted boxes - N</p></li>\n<li><p>if N == 0 : FN += 1 and exclude this particular target box from further consideration</p></li>\n<li><p>if N == 1  : TP += 1 and exclude these particular target box and predicted box from further consideration</p></li>\n<li><p>if N &gt; 1 : TP += 1 and exclude this particular target box and predicted box with the maximum <code>confidence</code> from further consideration</p></li>\n<li><p>go to step (1)</p></li>\n</ol>\n\n<p>Basically you have two slightly different ways to calculate TP. All the evaluation procedure and requirement to provide <code>confidence</code> for each predicted box make sense for the second algorithm of counting TP.</p>\n\n<p>But generally you are absolutely right - the script of evaluation is very needed. </p>",
      "votes": null,
      "replies": [
        {
          "id": 380252,
          "author_name": "hadarpo",
          "author_url": "",
          "post_date": "09/02/2018 06:25:10",
          "content": "<p>Thanks for the detailed response. Still I cannot see what is the relevance of the confidence even in your second algorithm, because in this algorithm, after you exclude the highest confidence predicted box in stage 4, all the lower confidence predicted boxes are counted as FPs, isn't it so?\nAdministrators - please respond!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 380258,
      "author_name": "hadarpo",
      "author_url": "",
      "post_date": "09/02/2018 06:54:36",
      "content": "<p>Actually there might be an evaluation algorithm in which the confidence does count: If the predicted bounding boxes are first ordered by their confidence and only after that are tested with the ground truth boxes, then there might be a difference in the scoring... For example assume that there are a predicted box A with high IoU and a predicted box B with low IoU (for the same ground truth of course). If A has a larger confidence than B, then we get one TP and one FP, but if A has a smaller confidence than B, we might get two FPs and one FN !. Anyways, again, the actual evaluation script is required..</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 383342,
      "author_name": "pisarik",
      "author_url": "",
      "post_date": "09/08/2018 12:32:06",
      "content": "<p>I agree, this metric is tricky. <strong>We need official clarification</strong>.  </p>\n\n<p>Here is my <strong>assumption</strong>.</p>\n\n<h1>1. Threshold &gt; 0.5:</h1>\n\n<p>Then for each predicted BBox (<strong>pbox</strong>) we can have only one matching true\nBBox (<strong>tbox</strong>)(strictly speaking there is necessary condition: <em>each pair of\ntboxes must have zero intersection</em>. It is almost true for given data).</p>\n\n<p>In this case there is no ambiguity how to match pboxes with tboxes:  we can use second algorithm, proposed by Ivan (actually it doesn't matter, we can use first as well in this case).</p>\n\n<p>So we have correspondence 1 tbox to *(many) pboxes.  </p>\n\n<ul>\n<li>tbox has <code>0</code> matches --&gt; <code>TN += 1</code></li>\n<li>tbox has <code>more then 1</code> matches --&gt; <code>TP += 1</code> and other goes to <code>FP</code>, no matter which of them, because all of them correspond to 1 tbox. </li>\n</ul>\n\n<p>Thus in this case order of matching doesn't matter, as well as confidence.</p>\n\n<h1>2. Thresholds &lt;= 0.5:</h1>\n\n<p>In this case we deal with correspondence * tboxes to * pboxes. </p>\n\n<p>Theoretically can happen next situation (drawing is not accurate, rely more on inequalities):\n<img src=\"https://pp.userapi.com/c846321/v846321499/ed955/VdAJWSQ4rlI.jpg\" alt=\"situation\"></p>\n\n<p>You would like to match <code>A-B</code>, <code>C-D</code>. \nBut let's consider what thinks about it two proposed algorithms by Ivan.</p>\n\n<p><strong>First Algorithm.</strong>\nTakes first <code>A</code> (because of confidence) and matches it to <code>C</code> (because of IoU) is <code>TP</code>. <br>\nThen <code>D</code> is <code>FP</code> and <code>B</code> is <code>FN</code>. Result: <code>A-C</code></p>\n\n<p>If you swap confidence of <code>A</code> and <code>B</code>, then you will get desired result.</p>\n\n<p><strong>Second Algorithm.</strong>\nThe result of this algorithm depends on order of tboxes. <br>\nIf first <code>B</code>, then we will get desired result. <br>\nIf first <code>C</code>, then it will be matched with A (because of IoU), and we will get bad result.</p>\n\n<p>Actually it is difficult imagine, that this case will happen someday or that it can drastically affect overall result. </p>\n\n<p>I think it is more matter of design. Your result shouldn't depend on predefined order of tboxes. That is why I think they took first approach, <strong>in that case result depend only on your proposal, your algorithm</strong>. Result fully under your control.</p>\n\n<h1>Why do we need to match BBoxes at all?</h1>\n\n<p>Why can't we calculate some kind of <code>average intersection</code>?\nBecause in this case it is easy to fool metric by predicting 1 billion of identical bboxes in which you are pretty confident. So there should be instrument for punishing for different number of pboxes and tboxes.</p>\n\n<p>But there are still exists method to do it properly:\n\\( Score(tboxes, pboxes) = \\Big(\\bigcup\\limits_{tb \\in tboxes} tb \\Big) \\bigcap \\Big(\\bigcup\\limits_{pb \\in pboxes} pb \\Big) \\)</p>\n\n<p>It has one advantage, because one scrupulous radiologist will annotate 4 small bboxes, while another one -- 2 big bboxes. Score above doesn't depend on it.</p>\n\n<p>But this <code>score</code> has disadvantage too. Because bboxes contain background as well as object. It is already looks more like a segmentation, but we don't have precise annotation of object of interests and probably it is even impossible to get segmentation of opacities, where do they start?</p>\n\n<p>So from the design point of view it is better to use <code>First algorithm</code> proposed by Ivan, but intuitively I would say that last <code>Score</code> would work better for this task.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "378417": "It seems that the order of evaluating the predicted bounding boxes does not influence the precision. If we have two predictions overlapping the same ground truth bounding box, one of them will always be counted as false positive regardless of the order. \nThe only situation in which the order might play a part is when two ground truth objects are covered by one large predicted box, but in that case it is unlikely that both will have IoU greater than 0.4 (and never both greater than 0.5)\nAm I missing anything? Is it possible to publish the evaluation script?\nThanks",
    "380168": "Actually there is such quite general symmetrical situation. \nConsider two algorithms\n\n**The algorithm based on predicted boxes primarily**\n0. TP = 0, FP = 0\n\n1. for each predicted box in the order of `confidence` get all IoU &gt; threshold target boxes - N\n\n2. if N == 0 : FP += 1 and exclude this particular predicted box from further consideration\n\n3. if N == 1  : TP += 1 and exclude these particular predicted box and target box from further consideration\n\n4. if N &gt; 1 : TP += 1 and exclude this particular predicted box and target box with the `maximum IoU` (?) from further consideration\n\n5. go to step (1)\n\n**The algorithm based on target boxes primarily**\n0. TP = 0, FN = 0\n\n1. for each target box in the order of `area size` (?) get all IoU &gt; threshold predicted boxes - N\n\n2. if N == 0 : FN += 1 and exclude this particular target box from further consideration\n\n3. if N == 1  : TP += 1 and exclude these particular target box and predicted box from further consideration\n\n4. if N &gt; 1 : TP += 1 and exclude this particular target box and predicted box with the maximum `confidence` from further consideration\n\n5. go to step (1)\n\nBasically you have two slightly different ways to calculate TP. All the evaluation procedure and requirement to provide `confidence` for each predicted box make sense for the second algorithm of counting TP.\n\nBut generally you are absolutely right - the script of evaluation is very needed.",
    "380252": "Thanks for the detailed response. Still I cannot see what is the relevance of the confidence even in your second algorithm, because in this algorithm, after you exclude the highest confidence predicted box in stage 4, all the lower confidence predicted boxes are counted as FPs, isn't it so?\nAdministrators - please respond!",
    "380258": "Actually there might be an evaluation algorithm in which the confidence does count: If the predicted bounding boxes are first ordered by their confidence and only after that are tested with the ground truth boxes, then there might be a difference in the scoring... For example assume that there are a predicted box A with high IoU and a predicted box B with low IoU (for the same ground truth of course). If A has a larger confidence than B, then we get one TP and one FP, but if A has a smaller confidence than B, we might get two FPs and one FN !. Anyways, again, the actual evaluation script is required..",
    "383342": "I agree, this metric is tricky. **We need official clarification**.  \n\nHere is my **assumption**.\n\n1. Threshold &gt; 0.5:\n===================\n\nThen for each predicted BBox (**pbox**) we can have only one matching true\nBBox (**tbox**)(strictly speaking there is necessary condition: *each pair of\ntboxes must have zero intersection*. It is almost true for given data).\n\nIn this case there is no ambiguity how to match pboxes with tboxes:  we can use second algorithm, proposed by Ivan (actually it doesn't matter, we can use first as well in this case).\n\nSo we have correspondence 1 tbox to *(many) pboxes.  \n\n- tbox has `0` matches --&gt; `TN += 1`\n- tbox has `more then 1` matches --&gt; `TP += 1` and other goes to `FP`, no matter which of them, because all of them correspond to 1 tbox. \n\nThus in this case order of matching doesn't matter, as well as confidence.\n\n2. Thresholds &lt;= 0.5:\n====================\n\nIn this case we deal with correspondence * tboxes to * pboxes. \n\nTheoretically can happen next situation (drawing is not accurate, rely more on inequalities):\n![situation](https://pp.userapi.com/c846321/v846321499/ed955/VdAJWSQ4rlI.jpg)\n\nYou would like to match `A-B`, `C-D`. \nBut let's consider what thinks about it two proposed algorithms by Ivan.\n\n**First Algorithm.**\nTakes first `A` (because of confidence) and matches it to `C` (because of IoU) is `TP`.  \nThen `D` is `FP` and `B` is `FN`. Result: `A-C`\n\nIf you swap confidence of `A` and `B`, then you will get desired result.\n\n**Second Algorithm.**\nThe result of this algorithm depends on order of tboxes.  \nIf first `B`, then we will get desired result.  \nIf first `C`, then it will be matched with A (because of IoU), and we will get bad result.\n\nActually it is difficult imagine, that this case will happen someday or that it can drastically affect overall result. \n\nI think it is more matter of design. Your result shouldn't depend on predefined order of tboxes. That is why I think they took first approach, **in that case result depend only on your proposal, your algorithm**. Result fully under your control.\n\nWhy do we need to match BBoxes at all?\n===============================\nWhy can't we calculate some kind of `average intersection`?\nBecause in this case it is easy to fool metric by predicting 1 billion of identical bboxes in which you are pretty confident. So there should be instrument for punishing for different number of pboxes and tboxes.\n\nBut there are still exists method to do it properly:\n\\\\( Score(tboxes, pboxes) = \\Big(\\bigcup\\limits_{tb \\in tboxes} tb \\Big) \\bigcap \\Big(\\bigcup\\limits_{pb \\in pboxes} pb \\Big) \\\\)\n\nIt has one advantage, because one scrupulous radiologist will annotate 4 small bboxes, while another one -- 2 big bboxes. Score above doesn't depend on it.\n\nBut this `score` has disadvantage too. Because bboxes contain background as well as object. It is already looks more like a segmentation, but we don't have precise annotation of object of interests and probably it is even impossible to get segmentation of opacities, where do they start?\n\nSo from the design point of view it is better to use `First algorithm` proposed by Ivan, but intuitively I would say that last `Score` would work better for this task."
  },
  "source": "meta"
}