{
  "id": 291348,
  "title": "Problem with f2_score if len(pred_bboxes) > len(gt_boxes)?",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/291348",
  "author_name": "",
  "post_date": "2021-11-28T22:12:35.375530300Z",
  "votes": 4,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I copied the \"Metric implementation\" code from the \"competition metric implementation\" notebook and it works fine if I limit it to the most-likely predicted bbox, but it fails in practice for me, I believe when <code>len(pred_bboxes) &gt; len(gt_boxes)</code>. For instance:</p>\n<pre><code>pred_bboxes = np.array([\n    [.981, 1017, 633, 54, 54],\n    [.873, 547, 680, 66, 37],\n    [.832, 562, 689, 50, 29],\n    [.736, 549, 674, 53, 45],\n    [.713, 612, 643, 49, 29],\n    [.567, 609, 633, 54, 47]\n])\ngt_bboxes = np.array([[548, 664, 70, 55]])\nf2_score(gt_bboxes, pred_bboxes, True)\n</code></pre>\n<p>fails. Digging in, it seems that <code>gt_bboxes</code> becomes 0-length in calls to <code>ious = calc_iou(gt_bboxes, pred_bbox[None, 1:])</code> but I don't see anything obvious in the code that shortens <code>gt_bboxes</code>. </p>\n<p>For instance, if you add <code>print(bboxes_1)</code> as the first line in <code>calc_iou</code>, you get:</p>\n<pre><code>[[548 664  70  55]]\n[[548 664  70  55]]\n[]\n</code></pre>",
  "messages": [
    {
      "id": "1598825",
      "postDate": "11/28/2021 22:12:35",
      "content": "<p>I copied the \"Metric implementation\" code from the \"competition metric implementation\" notebook and it works fine if I limit it to the most-likely predicted bbox, but it fails in practice for me, I believe when <code>len(pred_bboxes) &gt; len(gt_boxes)</code>. For instance:</p>\n<pre><code>pred_bboxes = np.array([\n    [.981, 1017, 633, 54, 54],\n    [.873, 547, 680, 66, 37],\n    [.832, 562, 689, 50, 29],\n    [.736, 549, 674, 53, 45],\n    [.713, 612, 643, 49, 29],\n    [.567, 609, 633, 54, 47]\n])\ngt_bboxes = np.array([[548, 664, 70, 55]])\nf2_score(gt_bboxes, pred_bboxes, True)\n</code></pre>\n<p>fails. Digging in, it seems that <code>gt_bboxes</code> becomes 0-length in calls to <code>ious = calc_iou(gt_bboxes, pred_bbox[None, 1:])</code> but I don't see anything obvious in the code that shortens <code>gt_bboxes</code>. </p>\n<p>For instance, if you add <code>print(bboxes_1)</code> as the first line in <code>calc_iou</code>, you get:</p>\n<pre><code>[[548 664  70  55]]\n[[548 664  70  55]]\n[]\n</code></pre>",
      "rawMarkdown": "I copied the \"Metric implementation\" code from the \"competition metric implementation\" notebook and it works fine if I limit it to the most-likely predicted bbox, but it fails in practice for me, I believe when `len(pred_bboxes) > len(gt_boxes)`. For instance:\n\n```python\npred_bboxes = np.array([\n    [.981, 1017, 633, 54, 54],\n    [.873, 547, 680, 66, 37],\n    [.832, 562, 689, 50, 29],\n    [.736, 549, 674, 53, 45],\n    [.713, 612, 643, 49, 29],\n    [.567, 609, 633, 54, 47]\n])\ngt_bboxes = np.array([[548, 664, 70, 55]])\nf2_score(gt_bboxes, pred_bboxes, True)\n```\nfails. Digging in, it seems that `gt_bboxes` becomes 0-length in calls to `ious = calc_iou(gt_bboxes, pred_bbox[None, 1:])` but I don't see anything obvious in the code that shortens `gt_bboxes`. \n\nFor instance, if you add `print(bboxes_1)` as the first line in `calc_iou`, you get:\n\n```\n[[548 664  70  55]]\n[[548 664  70  55]]\n[]\n```",
      "votes": null
    },
    {
      "id": "1598873",
      "postDate": "11/28/2021 23:43:23",
      "content": "<p>I think you are referring <a href=\"https://www.kaggle.com/bamps53\" target=\"_blank\">@bamps53</a> notebook?  The first column of <code>pred_bboxes</code> should be probability. [<code>p</code>, <code>x1</code>, <code>y1</code>, <code>x2,</code>y2<code>]. Bassicly</code>pred_bboxes`  get sorted by probability and then the score is calculated. Check out this:</p>\n<p><code>In your submission, you are also asked to provide a confidence level for each bounding box. Bounding boxes are evaluated in order of their confidence levels. This means that bounding boxes with higher confidence will be checked first for matches against solutions, which determines what boxes are considered true and false positives.</code></p>",
      "rawMarkdown": "I think you are referring @bamps53 notebook?  The first column of `pred_bboxes` should be probability. [`p`, `x1`, `y1`, `x2, `y2`]. Bassicly  `pred_bboxes`  get sorted by probability and then the score is calculated. Check out this:\n\n`In your submission, you are also asked to provide a confidence level for each bounding box. Bounding boxes are evaluated in order of their confidence levels. This means that bounding boxes with higher confidence will be checked first for matches against solutions, which determines what boxes are considered true and false positives.`",
      "votes": null
    },
    {
      "id": "1598950",
      "postDate": "11/29/2021 02:58:33",
      "content": "<p>Thanks for pointing out. I'll check it out later;)<br>\nBut it's better to mention me directly or raise in the notebook so that I can respond quickly!<br>\n(I found this because DrHB pointed me:))</p>",
      "rawMarkdown": "Thanks for pointing out. I'll check it out later;)\nBut it's better to mention me directly or raise in the notebook so that I can respond quickly!\n(I found this because DrHB pointed me:))",
      "votes": null
    },
    {
      "id": "1599011",
      "postDate": "11/29/2021 04:21:01",
      "content": "<blockquote>\n  <p>The first column of pred_bboxes should be probability. [p, x1, y1, x2,y2]</p>\n</blockquote>\n<p>I think the shape of <code>pred_bboxes</code> is correct. I have, for instance, <code>[.981, 1017, 633, 54, 54]</code> -- 98% confidence of a box at [1017,633] of size [54,54]. Isn't that right? </p>",
      "rawMarkdown": "> The first column of pred_bboxes should be probability. [p, x1, y1, x2,y2]\n\nI think the shape of `pred_bboxes` is correct. I have, for instance, `[.981, 1017, 633, 54, 54]` -- 98% confidence of a box at [1017,633] of size [54,54]. Isn't that right?",
      "votes": null
    },
    {
      "id": "1599039",
      "postDate": "11/29/2021 04:43:07",
      "content": "<p>Sorry about that. I think the issue is if a true positive zeroes out the <code>gt_bboxes</code> but the loop doesn't end (because more <code>pred_bboxes</code>), the call to <code>calc_iou</code> is invalid :</p>\n<pre><code>for pred_bbox in pred_bboxes:\n        ious = calc_iou(gt_bboxes, pred_bbox[None, 1:]) # If `len(gt_bboxes) == 0` this throws\n        max_iou = ious.max()\n        if max_iou &gt; iou_th:\n            tp += 1\n            gt_bboxes = np.delete(gt_bboxes, ious.argmax(), axis=0) # This can zero out `gt_bboxes`\n        else:\n            fp += 1\n</code></pre>\n<p>If I understand the logic correctly, you could either alter <code>calc_iou</code> to return 0.0 if <code>len(gt_boxes) == 0</code> or test at the top of this loop, add <code>len(pred_bboxes) - len(gt_bboxes_original)</code> to <code>fp</code>, and break out of the loop. </p>",
      "rawMarkdown": "Sorry about that. I think the issue is if a true positive zeroes out the `gt_bboxes` but the loop doesn't end (because more `pred_bboxes`), the call to `calc_iou` is invalid :\n\n```\nfor pred_bbox in pred_bboxes:\n        ious = calc_iou(gt_bboxes, pred_bbox[None, 1:]) # If `len(gt_bboxes) == 0` this throws\n        max_iou = ious.max()\n        if max_iou > iou_th:\n            tp += 1\n            gt_bboxes = np.delete(gt_bboxes, ious.argmax(), axis=0) # This can zero out `gt_bboxes`\n        else:\n            fp += 1\n```\n\nIf I understand the logic correctly, you could either alter `calc_iou` to return 0.0 if `len(gt_boxes) == 0` or test at the top of this loop, add `len(pred_bboxes) - len(gt_bboxes_original)` to `fp`, and break out of the loop.",
      "votes": null
    },
    {
      "id": "1599078",
      "postDate": "11/29/2021 05:28:30",
      "content": "<p>that’s correct, my bad i did not notice :) </p>",
      "rawMarkdown": "that’s correct, my bad i did not notice :)",
      "votes": null
    },
    {
      "id": "1599477",
      "postDate": "11/29/2021 13:17:52",
      "content": "<p><a href=\"https://www.kaggle.com/lobrien\" target=\"_blank\">@lobrien</a> I've fixed it and also added correct metric. In my first notebook, it calculate image-wise metric but it might be not correct as competition metric is average over all frames. Please check it out!<br>\n<a href=\"https://www.kaggle.com/bamps53/competition-metric-implementation?scriptVersionId=81087805\" target=\"_blank\">https://www.kaggle.com/bamps53/competition-metric-implementation?scriptVersionId=81087805</a></p>",
      "rawMarkdown": "lobrien I've fixed it and also added correct metric. In my first notebook, it calculate image-wise metric but it might be not correct as competition metric is average over all frames. Please check it out!\nhttps://www.kaggle.com/bamps53/competition-metric-implementation?scriptVersionId=81087805",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1598873,
      "author_name": "drhabib",
      "author_url": "",
      "post_date": "11/28/2021 23:43:23",
      "content": "<p>I think you are referring <a href=\"https://www.kaggle.com/bamps53\" target=\"_blank\">@bamps53</a> notebook?  The first column of <code>pred_bboxes</code> should be probability. [<code>p</code>, <code>x1</code>, <code>y1</code>, <code>x2,</code>y2<code>]. Bassicly</code>pred_bboxes`  get sorted by probability and then the score is calculated. Check out this:</p>\n<p><code>In your submission, you are also asked to provide a confidence level for each bounding box. Bounding boxes are evaluated in order of their confidence levels. This means that bounding boxes with higher confidence will be checked first for matches against solutions, which determines what boxes are considered true and false positives.</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 1599011,
          "author_name": "lobrien",
          "author_url": "",
          "post_date": "11/29/2021 04:21:01",
          "content": "<blockquote>\n  <p>The first column of pred_bboxes should be probability. [p, x1, y1, x2,y2]</p>\n</blockquote>\n<p>I think the shape of <code>pred_bboxes</code> is correct. I have, for instance, <code>[.981, 1017, 633, 54, 54]</code> -- 98% confidence of a box at [1017,633] of size [54,54]. Isn't that right? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1599078,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "11/29/2021 05:28:30",
          "content": "<p>that’s correct, my bad i did not notice :) </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1598950,
      "author_name": "bamps53",
      "author_url": "",
      "post_date": "11/29/2021 02:58:33",
      "content": "<p>Thanks for pointing out. I'll check it out later;)<br>\nBut it's better to mention me directly or raise in the notebook so that I can respond quickly!<br>\n(I found this because DrHB pointed me:))</p>",
      "votes": null,
      "replies": [
        {
          "id": 1599039,
          "author_name": "lobrien",
          "author_url": "",
          "post_date": "11/29/2021 04:43:07",
          "content": "<p>Sorry about that. I think the issue is if a true positive zeroes out the <code>gt_bboxes</code> but the loop doesn't end (because more <code>pred_bboxes</code>), the call to <code>calc_iou</code> is invalid :</p>\n<pre><code>for pred_bbox in pred_bboxes:\n        ious = calc_iou(gt_bboxes, pred_bbox[None, 1:]) # If `len(gt_bboxes) == 0` this throws\n        max_iou = ious.max()\n        if max_iou &gt; iou_th:\n            tp += 1\n            gt_bboxes = np.delete(gt_bboxes, ious.argmax(), axis=0) # This can zero out `gt_bboxes`\n        else:\n            fp += 1\n</code></pre>\n<p>If I understand the logic correctly, you could either alter <code>calc_iou</code> to return 0.0 if <code>len(gt_boxes) == 0</code> or test at the top of this loop, add <code>len(pred_bboxes) - len(gt_bboxes_original)</code> to <code>fp</code>, and break out of the loop. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1599477,
      "author_name": "bamps53",
      "author_url": "",
      "post_date": "11/29/2021 13:17:52",
      "content": "<p><a href=\"https://www.kaggle.com/lobrien\" target=\"_blank\">@lobrien</a> I've fixed it and also added correct metric. In my first notebook, it calculate image-wise metric but it might be not correct as competition metric is average over all frames. Please check it out!<br>\n<a href=\"https://www.kaggle.com/bamps53/competition-metric-implementation?scriptVersionId=81087805\" target=\"_blank\">https://www.kaggle.com/bamps53/competition-metric-implementation?scriptVersionId=81087805</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1598825": "I copied the \"Metric implementation\" code from the \"competition metric implementation\" notebook and it works fine if I limit it to the most-likely predicted bbox, but it fails in practice for me, I believe when `len(pred_bboxes) > len(gt_boxes)`. For instance:\n\n```python\npred_bboxes = np.array([\n    [.981, 1017, 633, 54, 54],\n    [.873, 547, 680, 66, 37],\n    [.832, 562, 689, 50, 29],\n    [.736, 549, 674, 53, 45],\n    [.713, 612, 643, 49, 29],\n    [.567, 609, 633, 54, 47]\n])\ngt_bboxes = np.array([[548, 664, 70, 55]])\nf2_score(gt_bboxes, pred_bboxes, True)\n```\nfails. Digging in, it seems that `gt_bboxes` becomes 0-length in calls to `ious = calc_iou(gt_bboxes, pred_bbox[None, 1:])` but I don't see anything obvious in the code that shortens `gt_bboxes`. \n\nFor instance, if you add `print(bboxes_1)` as the first line in `calc_iou`, you get:\n\n```\n[[548 664  70  55]]\n[[548 664  70  55]]\n[]\n```",
    "1598873": "I think you are referring @bamps53 notebook?  The first column of `pred_bboxes` should be probability. [`p`, `x1`, `y1`, `x2, `y2`]. Bassicly  `pred_bboxes`  get sorted by probability and then the score is calculated. Check out this:\n\n`In your submission, you are also asked to provide a confidence level for each bounding box. Bounding boxes are evaluated in order of their confidence levels. This means that bounding boxes with higher confidence will be checked first for matches against solutions, which determines what boxes are considered true and false positives.`",
    "1598950": "Thanks for pointing out. I'll check it out later;)\nBut it's better to mention me directly or raise in the notebook so that I can respond quickly!\n(I found this because DrHB pointed me:))",
    "1599011": "> The first column of pred_bboxes should be probability. [p, x1, y1, x2,y2]\n\nI think the shape of `pred_bboxes` is correct. I have, for instance, `[.981, 1017, 633, 54, 54]` -- 98% confidence of a box at [1017,633] of size [54,54]. Isn't that right?",
    "1599039": "Sorry about that. I think the issue is if a true positive zeroes out the `gt_bboxes` but the loop doesn't end (because more `pred_bboxes`), the call to `calc_iou` is invalid :\n\n```\nfor pred_bbox in pred_bboxes:\n        ious = calc_iou(gt_bboxes, pred_bbox[None, 1:]) # If `len(gt_bboxes) == 0` this throws\n        max_iou = ious.max()\n        if max_iou > iou_th:\n            tp += 1\n            gt_bboxes = np.delete(gt_bboxes, ious.argmax(), axis=0) # This can zero out `gt_bboxes`\n        else:\n            fp += 1\n```\n\nIf I understand the logic correctly, you could either alter `calc_iou` to return 0.0 if `len(gt_boxes) == 0` or test at the top of this loop, add `len(pred_bboxes) - len(gt_bboxes_original)` to `fp`, and break out of the loop.",
    "1599078": "that’s correct, my bad i did not notice :)",
    "1599477": "lobrien I've fixed it and also added correct metric. In my first notebook, it calculate image-wise metric but it might be not correct as competition metric is average over all frames. Please check it out!\nhttps://www.kaggle.com/bamps53/competition-metric-implementation?scriptVersionId=81087805"
  },
  "source": "meta"
}