{
  "id": 90989,
  "title": "Couple potential mismatches in the detections",
  "url": "/competitions/iwildcam-2019-fgvc6/discussion/90989",
  "author_name": "",
  "post_date": "2019-04-29T21:56:24.012283600Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Can anyone confirm, shed light, on whether they have plotted the detections correctly. I was using the </p>\n\n<p>CCT_Detection_Results_2.p</p>\n\n<p>grabbing a reasonable image and using the render_boxes from the linked </p>\n\n<p><a href=\"https://github.com/Microsoft/CameraTraps/blob/master/detection/run_tf_detector.py#L276\">https://github.com/Microsoft/CameraTraps/blob/master/detection/run_tf_detector.py#L276</a></p>\n\n<p>yields a box that I'm pretty sure is upside-down. Hard to say without knowing how the data was created, but the code specifies a couple things that are not true. </p>\n\n<ol>\n<li>The link script expects coordinates relative to the image size, and then multiplies them by the shape\n<a href=\"https://github.com/Microsoft/CameraTraps/blob/b05e3a7a5499a32a226bbcf736ecf8d965ca8cac/detection/run_tf_detector.py#L269\">https://github.com/Microsoft/CameraTraps/blob/b05e3a7a5499a32a226bbcf736ecf8d965ca8cac/detection/run_tf_detector.py#L269</a></li>\n</ol>\n\n<p><code>\n            x = leftRel * imageWidth\n            y = topRel * imageHeight\n            w = (rightRel-leftRel) * imageWidth\n            h = (bottomRel-topRel) * imageHeight\n</code></p>\n\n<p>but the coordinates of the pickle object are clearly not relative, but in absolute terms.</p>\n\n<p><code>\nimage_detection\n[1365.7012939453125, 556.8687520623207, 1647.4263916015625, 1018.4135702848434]\n</code>\nwith an image shape of </p>\n\n<p><code>\nimage.shape\n(1494, 2048, 3)\n</code></p>\n\n<p>no problem, just remove those lines. But then plotting the image you see a weird box (Figure 2). You can tell pretty fast that its actually the right box, just x,y flipped. </p>\n\n<p>Changing the order of x,y to\n<code>\n            rect = patches.Rectangle((y,x),w,h,linewidth=linewidth,edgecolor='r',\n                                     facecolor='none')\n</code></p>\n\n<p>looks like it yields the correct box (Figure 1). Presumably. I haven't touched the code much, has anyone else noticed this? I'm assuming the order of boxes was somehow switched.</p>",
  "messages": [
    {
      "id": "524973",
      "postDate": "04/29/2019 21:56:24",
      "content": "<p>Can anyone confirm, shed light, on whether they have plotted the detections correctly. I was using the </p>\n\n<p>CCT_Detection_Results_2.p</p>\n\n<p>grabbing a reasonable image and using the render_boxes from the linked </p>\n\n<p><a href=\"https://github.com/Microsoft/CameraTraps/blob/master/detection/run_tf_detector.py#L276\">https://github.com/Microsoft/CameraTraps/blob/master/detection/run_tf_detector.py#L276</a></p>\n\n<p>yields a box that I'm pretty sure is upside-down. Hard to say without knowing how the data was created, but the code specifies a couple things that are not true. </p>\n\n<ol>\n<li>The link script expects coordinates relative to the image size, and then multiplies them by the shape\n<a href=\"https://github.com/Microsoft/CameraTraps/blob/b05e3a7a5499a32a226bbcf736ecf8d965ca8cac/detection/run_tf_detector.py#L269\">https://github.com/Microsoft/CameraTraps/blob/b05e3a7a5499a32a226bbcf736ecf8d965ca8cac/detection/run_tf_detector.py#L269</a></li>\n</ol>\n\n<p><code>\n            x = leftRel * imageWidth\n            y = topRel * imageHeight\n            w = (rightRel-leftRel) * imageWidth\n            h = (bottomRel-topRel) * imageHeight\n</code></p>\n\n<p>but the coordinates of the pickle object are clearly not relative, but in absolute terms.</p>\n\n<p><code>\nimage_detection\n[1365.7012939453125, 556.8687520623207, 1647.4263916015625, 1018.4135702848434]\n</code>\nwith an image shape of </p>\n\n<p><code>\nimage.shape\n(1494, 2048, 3)\n</code></p>\n\n<p>no problem, just remove those lines. But then plotting the image you see a weird box (Figure 2). You can tell pretty fast that its actually the right box, just x,y flipped. </p>\n\n<p>Changing the order of x,y to\n<code>\n            rect = patches.Rectangle((y,x),w,h,linewidth=linewidth,edgecolor='r',\n                                     facecolor='none')\n</code></p>\n\n<p>looks like it yields the correct box (Figure 1). Presumably. I haven't touched the code much, has anyone else noticed this? I'm assuming the order of boxes was somehow switched.</p>",
      "rawMarkdown": "Can anyone confirm, shed light, on whether they have plotted the detections correctly. I was using the \n\nCCT_Detection_Results_2.p\n\ngrabbing a reasonable image and using the render_boxes from the linked \n\nhttps://github.com/Microsoft/CameraTraps/blob/master/detection/run_tf_detector.py#L276\n\nyields a box that I'm pretty sure is upside-down. Hard to say without knowing how the data was created, but the code specifies a couple things that are not true. \n\n1. The link script expects coordinates relative to the image size, and then multiplies them by the shape\nhttps://github.com/Microsoft/CameraTraps/blob/b05e3a7a5499a32a226bbcf736ecf8d965ca8cac/detection/run_tf_detector.py#L269\n\n```\n            x = leftRel * imageWidth\n            y = topRel * imageHeight\n            w = (rightRel-leftRel) * imageWidth\n            h = (bottomRel-topRel) * imageHeight\n```\n\nbut the coordinates of the pickle object are clearly not relative, but in absolute terms.\n\n```\nimage_detection\n[1365.7012939453125, 556.8687520623207, 1647.4263916015625, 1018.4135702848434]\n```\nwith an image shape of \n\n```\nimage.shape\n(1494, 2048, 3)\n```\n\nno problem, just remove those lines. But then plotting the image you see a weird box (Figure 2). You can tell pretty fast that its actually the right box, just x,y flipped. \n\nChanging the order of x,y to\n```\n            rect = patches.Rectangle((y,x),w,h,linewidth=linewidth,edgecolor='r',\n                                     facecolor='none')\n```\n\nlooks like it yields the correct box (Figure 1). Presumably. I haven't touched the code much, has anyone else noticed this? I'm assuming the order of boxes was somehow switched.",
      "votes": null
    },
    {
      "id": "525330",
      "postDate": "04/30/2019 18:13:24",
      "content": "<p>Ben, good catch.  The detection result boxes I released are in the <a href=\"http://cocodataset.org/#format-data\">COCO format</a> (eg non-relative coordinates, (x,y,w,h) where x,y is the upper left point), I was unaware that the Microsoft code snippet was not expecting that format.  As to the x,y flip, that is based on how <a href=\"https://matplotlib.org/api/_as_gen/matplotlib.patches.Rectangle.html\">matplotlib Rectangle</a> works (this is a classic coordinate system issue in CV), it expects the box to be x,y,w,h with x,y in the lower left, and the data reference frame for matplotlib has 0,0 in the lower left hand corner. So this effects the visualization in the MS code, but the boxes are in the correct format to be input directly into the COCO-CameraTraps JSON files.</p>",
      "rawMarkdown": "Ben, good catch.  The detection result boxes I released are in the [COCO format](http://cocodataset.org/#format-data) (eg non-relative coordinates, (x,y,w,h) where x,y is the upper left point), I was unaware that the Microsoft code snippet was not expecting that format.  As to the x,y flip, that is based on how [matplotlib Rectangle](https://matplotlib.org/api/_as_gen/matplotlib.patches.Rectangle.html) works (this is a classic coordinate system issue in CV), it expects the box to be x,y,w,h with x,y in the lower left, and the data reference frame for matplotlib has 0,0 in the lower left hand corner. So this effects the visualization in the MS code, but the boxes are in the correct format to be input directly into the COCO-CameraTraps JSON files.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 525330,
      "author_name": "sbeery",
      "author_url": "",
      "post_date": "04/30/2019 18:13:24",
      "content": "<p>Ben, good catch.  The detection result boxes I released are in the <a href=\"http://cocodataset.org/#format-data\">COCO format</a> (eg non-relative coordinates, (x,y,w,h) where x,y is the upper left point), I was unaware that the Microsoft code snippet was not expecting that format.  As to the x,y flip, that is based on how <a href=\"https://matplotlib.org/api/_as_gen/matplotlib.patches.Rectangle.html\">matplotlib Rectangle</a> works (this is a classic coordinate system issue in CV), it expects the box to be x,y,w,h with x,y in the lower left, and the data reference frame for matplotlib has 0,0 in the lower left hand corner. So this effects the visualization in the MS code, but the boxes are in the correct format to be input directly into the COCO-CameraTraps JSON files.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "524973": "Can anyone confirm, shed light, on whether they have plotted the detections correctly. I was using the \n\nCCT_Detection_Results_2.p\n\ngrabbing a reasonable image and using the render_boxes from the linked \n\nhttps://github.com/Microsoft/CameraTraps/blob/master/detection/run_tf_detector.py#L276\n\nyields a box that I'm pretty sure is upside-down. Hard to say without knowing how the data was created, but the code specifies a couple things that are not true. \n\n1. The link script expects coordinates relative to the image size, and then multiplies them by the shape\nhttps://github.com/Microsoft/CameraTraps/blob/b05e3a7a5499a32a226bbcf736ecf8d965ca8cac/detection/run_tf_detector.py#L269\n\n```\n            x = leftRel * imageWidth\n            y = topRel * imageHeight\n            w = (rightRel-leftRel) * imageWidth\n            h = (bottomRel-topRel) * imageHeight\n```\n\nbut the coordinates of the pickle object are clearly not relative, but in absolute terms.\n\n```\nimage_detection\n[1365.7012939453125, 556.8687520623207, 1647.4263916015625, 1018.4135702848434]\n```\nwith an image shape of \n\n```\nimage.shape\n(1494, 2048, 3)\n```\n\nno problem, just remove those lines. But then plotting the image you see a weird box (Figure 2). You can tell pretty fast that its actually the right box, just x,y flipped. \n\nChanging the order of x,y to\n```\n            rect = patches.Rectangle((y,x),w,h,linewidth=linewidth,edgecolor='r',\n                                     facecolor='none')\n```\n\nlooks like it yields the correct box (Figure 1). Presumably. I haven't touched the code much, has anyone else noticed this? I'm assuming the order of boxes was somehow switched.",
    "525330": "Ben, good catch.  The detection result boxes I released are in the [COCO format](http://cocodataset.org/#format-data) (eg non-relative coordinates, (x,y,w,h) where x,y is the upper left point), I was unaware that the Microsoft code snippet was not expecting that format.  As to the x,y flip, that is based on how [matplotlib Rectangle](https://matplotlib.org/api/_as_gen/matplotlib.patches.Rectangle.html) works (this is a classic coordinate system issue in CV), it expects the box to be x,y,w,h with x,y in the lower left, and the data reference frame for matplotlib has 0,0 in the lower left hand corner. So this effects the visualization in the MS code, but the boxes are in the correct format to be input directly into the COCO-CameraTraps JSON files."
  },
  "source": "meta"
}