{
  "id": 108943,
  "title": "Having  problem of understanding the evaluation part~",
  "url": "/competitions/3d-object-detection-for-autonomous-vehicles/discussion/108943",
  "author_name": "",
  "post_date": "2019-09-15T04:38:24.913675300Z",
  "votes": 3,
  "comment_count": 8,
  "views": 0,
  "content": "<p>In the \"3D context\":\n    \"The difference between the 2D and 3D bounding volume contexts is small. In the 3D context we reduce the bounding volume to a ground bounding box and a height. <strong>The IoU is then the intersection of the ground bounding boxes * the intersection of the height differences</strong>, divided by the union of the bounding boxes.\"</p>\n\n<p>What's does it mean by \"ground bounding boxes\", \"<em>\" and \"the intersection of the height differences\"?(Does \"ground bounding boxes\" mean the lowest part of a 3D box? Does \"</em>\" mean \"multiply\"? Why will there be an intersection between the <strong>lowest square</strong> of a 3D box and <strong>height</strong> of that 3D box? )</p>\n\n<p>By the way I'm pretty curious about why they use the word \"reduce\" in \"reducing the bounding volume to a ground bounding box and a height\".(But this question is trivial)</p>",
  "messages": [
    {
      "id": "626901",
      "postDate": "09/15/2019 04:38:24",
      "content": "<p>In the \"3D context\":\n    \"The difference between the 2D and 3D bounding volume contexts is small. In the 3D context we reduce the bounding volume to a ground bounding box and a height. <strong>The IoU is then the intersection of the ground bounding boxes * the intersection of the height differences</strong>, divided by the union of the bounding boxes.\"</p>\n\n<p>What's does it mean by \"ground bounding boxes\", \"<em>\" and \"the intersection of the height differences\"?(Does \"ground bounding boxes\" mean the lowest part of a 3D box? Does \"</em>\" mean \"multiply\"? Why will there be an intersection between the <strong>lowest square</strong> of a 3D box and <strong>height</strong> of that 3D box? )</p>\n\n<p>By the way I'm pretty curious about why they use the word \"reduce\" in \"reducing the bounding volume to a ground bounding box and a height\".(But this question is trivial)</p>",
      "rawMarkdown": "In the \"3D context\":\n    \"The difference between the 2D and 3D bounding volume contexts is small. In the 3D context we reduce the bounding volume to a ground bounding box and a height. **The IoU is then the intersection of the ground bounding boxes * the intersection of the height differences**, divided by the union of the bounding boxes.\"\n\nWhat's does it mean by \"ground bounding boxes\", \"*\" and \"the intersection of the height differences\"?(Does \"ground bounding boxes\" mean the lowest part of a 3D box? Does \"*\" mean \"multiply\"? Why will there be an intersection between the **lowest square** of a 3D box and **height** of that 3D box? )\n\nBy the way I'm pretty curious about why they use the word \"reduce\" in \"reducing the bounding volume to a ground bounding box and a height\".(But this question is trivial)",
      "votes": null
    },
    {
      "id": "626952",
      "postDate": "09/15/2019 06:45:53",
      "content": "<p>Since bounding boxes are all in x,y plane and rotated around the z axis, you can reduce (hehe) calculation of 3d box IoU to calculating IoU along the z axis (vertical) and in x,y plane (on the \"ground\").\nEdit: in more detail, for such boxes, you can calculate IoU3d as multiplication of IoUGround and IoUHeight, where IoUGround would be usual 2d IoU of boxes projected on the ground, and IoUHeight would be the 1d IoU of boxes along the z axis. Now, why can we multiply IoUs here? Because we can multiply both intersection and union.</p>",
      "rawMarkdown": "Since bounding boxes are all in x,y plane and rotated around the z axis, you can reduce (hehe) calculation of 3d box IoU to calculating IoU along the z axis (vertical) and in x,y plane (on the \"ground\").\nEdit: in more detail, for such boxes, you can calculate IoU3d as multiplication of IoUGround and IoUHeight, where IoUGround would be usual 2d IoU of boxes projected on the ground, and IoUHeight would be the 1d IoU of boxes along the z axis. Now, why can we multiply IoUs here? Because we can multiply both intersection and union.",
      "votes": null
    },
    {
      "id": "627052",
      "postDate": "09/15/2019 10:15:01",
      "content": "<p>Modern cars are usually narrower on the roof. Hence, a proper 3D bounding box wouldn't be a standard rectangular box. However, when it comes to robotics (incl. self-driving cars), we are more interested in finding occupied spaces to avoid collisions. Since we use some safety margin anyhow, we are only interested in the maximum extend of a vehicle. The volumetric difference between a perfect 3D bounding box and a simplified cuboid is small and not relevant .</p>\n\n<p>You can think about the ground bounding box as looking on a vehicle from above and draw the boundaries. It doesn't really matter if the vehicle is widest at the top or the bottom, because we are projecting the maximum extend to 2D.  This reduces computational costs. Path planning in 2D is faster and more efficient. The 3D component is reduced to a simple height which allows to reconstruct a simple cuboid. The metric is reduced to compute IoUs of a line (height feature) and an area feature (ground box).</p>",
      "rawMarkdown": "Modern cars are usually narrower on the roof. Hence, a proper 3D bounding box wouldn't be a standard rectangular box. However, when it comes to robotics (incl. self-driving cars), we are more interested in finding occupied spaces to avoid collisions. Since we use some safety margin anyhow, we are only interested in the maximum extend of a vehicle. The volumetric difference between a perfect 3D bounding box and a simplified cuboid is small and not relevant .\n\nYou can think about the ground bounding box as looking on a vehicle from above and draw the boundaries. It doesn't really matter if the vehicle is widest at the top or the bottom, because we are projecting the maximum extend to 2D.  This reduces computational costs. Path planning in 2D is faster and more efficient. The 3D component is reduced to a simple height which allows to reconstruct a simple cuboid. The metric is reduced to compute IoUs of a line (height feature) and an area feature (ground box).",
      "votes": null
    },
    {
      "id": "628007",
      "postDate": "09/16/2019 17:31:06",
      "content": "<p>Sample metric implementation: <a href=\"https://github.com/lyft/nuscenes-devkit/blob/master/lyft_dataset_sdk/eval/detection/mAP_evaluation.py\">https://github.com/lyft/nuscenes-devkit/blob/master/lyft_dataset_sdk/eval/detection/mAP_evaluation.py</a></p>",
      "rawMarkdown": "Sample metric implementation: https://github.com/lyft/nuscenes-devkit/blob/master/lyft_dataset_sdk/eval/detection/mAP_evaluation.py",
      "votes": null
    },
    {
      "id": "628315",
      "postDate": "09/17/2019 06:45:25",
      "content": "<p>Hi, the github code doesn't match the evaluation described in the competitition. Github implements a classic AP (with area under precision-recall curve), while in Kaggle just an accuracy/Critical Success Index is described. Which one is really used here ? Since it is said that confidence scores are not important, I guess it is the metric described here, but it has little to do with the github implementation then.</p>",
      "rawMarkdown": "Hi, the github code doesn't match the evaluation described in the competitition. Github implements a classic AP (with area under precision-recall curve), while in Kaggle just an accuracy/Critical Success Index is described. Which one is really used here ? Since it is said that confidence scores are not important, I guess it is the metric described here, but it has little to do with the github implementation then.",
      "votes": null
    },
    {
      "id": "628446",
      "postDate": "09/17/2019 11:08:30",
      "content": "<p>Hi <a href=\"/thomasgilles\">@thomasgilles</a> - the Kaggle implementation is based on the implementation that Vladimir linked. Any discrepancies are in the explanation of the metric. I'll add a description of AP using area under the precision-recall curve, as that is definitely what we're using. Thanks!</p>",
      "rawMarkdown": "Hi @thomasgilles - the Kaggle implementation is based on the implementation that Vladimir linked. Any discrepancies are in the explanation of the metric. I'll add a description of AP using area under the precision-recall curve, as that is definitely what we're using. Thanks!",
      "votes": null
    },
    {
      "id": "633078",
      "postDate": "09/24/2019 12:21:15",
      "content": "<p>Thank you. ^^</p>",
      "rawMarkdown": "Thank you. ^^",
      "votes": null
    },
    {
      "id": "633080",
      "postDate": "09/24/2019 12:21:37",
      "content": "<p>Thank you~</p>",
      "rawMarkdown": "Thank you~",
      "votes": null
    },
    {
      "id": "633081",
      "postDate": "09/24/2019 12:21:54",
      "content": "<p>Thank you ~~ ^ ^</p>",
      "rawMarkdown": "Thank you ~~ ^ ^",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 626952,
      "author_name": "lopuhin",
      "author_url": "",
      "post_date": "09/15/2019 06:45:53",
      "content": "<p>Since bounding boxes are all in x,y plane and rotated around the z axis, you can reduce (hehe) calculation of 3d box IoU to calculating IoU along the z axis (vertical) and in x,y plane (on the \"ground\").\nEdit: in more detail, for such boxes, you can calculate IoU3d as multiplication of IoUGround and IoUHeight, where IoUGround would be usual 2d IoU of boxes projected on the ground, and IoUHeight would be the 1d IoU of boxes along the z axis. Now, why can we multiply IoUs here? Because we can multiply both intersection and union.</p>",
      "votes": null,
      "replies": [
        {
          "id": 633080,
          "author_name": "a6893676",
          "author_url": "",
          "post_date": "09/24/2019 12:21:37",
          "content": "<p>Thank you~</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 627052,
      "author_name": "simonwenkel",
      "author_url": "",
      "post_date": "09/15/2019 10:15:01",
      "content": "<p>Modern cars are usually narrower on the roof. Hence, a proper 3D bounding box wouldn't be a standard rectangular box. However, when it comes to robotics (incl. self-driving cars), we are more interested in finding occupied spaces to avoid collisions. Since we use some safety margin anyhow, we are only interested in the maximum extend of a vehicle. The volumetric difference between a perfect 3D bounding box and a simplified cuboid is small and not relevant .</p>\n\n<p>You can think about the ground bounding box as looking on a vehicle from above and draw the boundaries. It doesn't really matter if the vehicle is widest at the top or the bottom, because we are projecting the maximum extend to 2D.  This reduces computational costs. Path planning in 2D is faster and more efficient. The 3D component is reduced to a simple height which allows to reconstruct a simple cuboid. The metric is reduced to compute IoUs of a line (height feature) and an area feature (ground box).</p>",
      "votes": null,
      "replies": [
        {
          "id": 633081,
          "author_name": "a6893676",
          "author_url": "",
          "post_date": "09/24/2019 12:21:54",
          "content": "<p>Thank you ~~ ^ ^</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 628007,
      "author_name": "iglovikov",
      "author_url": "",
      "post_date": "09/16/2019 17:31:06",
      "content": "<p>Sample metric implementation: <a href=\"https://github.com/lyft/nuscenes-devkit/blob/master/lyft_dataset_sdk/eval/detection/mAP_evaluation.py\">https://github.com/lyft/nuscenes-devkit/blob/master/lyft_dataset_sdk/eval/detection/mAP_evaluation.py</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 628315,
          "author_name": "thomasgilles",
          "author_url": "",
          "post_date": "09/17/2019 06:45:25",
          "content": "<p>Hi, the github code doesn't match the evaluation described in the competitition. Github implements a classic AP (with area under precision-recall curve), while in Kaggle just an accuracy/Critical Success Index is described. Which one is really used here ? Since it is said that confidence scores are not important, I guess it is the metric described here, but it has little to do with the github implementation then.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 628446,
          "author_name": "philculliton",
          "author_url": "",
          "post_date": "09/17/2019 11:08:30",
          "content": "<p>Hi <a href=\"/thomasgilles\">@thomasgilles</a> - the Kaggle implementation is based on the implementation that Vladimir linked. Any discrepancies are in the explanation of the metric. I'll add a description of AP using area under the precision-recall curve, as that is definitely what we're using. Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 633078,
          "author_name": "a6893676",
          "author_url": "",
          "post_date": "09/24/2019 12:21:15",
          "content": "<p>Thank you. ^^</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "626901": "In the \"3D context\":\n    \"The difference between the 2D and 3D bounding volume contexts is small. In the 3D context we reduce the bounding volume to a ground bounding box and a height. **The IoU is then the intersection of the ground bounding boxes * the intersection of the height differences**, divided by the union of the bounding boxes.\"\n\nWhat's does it mean by \"ground bounding boxes\", \"*\" and \"the intersection of the height differences\"?(Does \"ground bounding boxes\" mean the lowest part of a 3D box? Does \"*\" mean \"multiply\"? Why will there be an intersection between the **lowest square** of a 3D box and **height** of that 3D box? )\n\nBy the way I'm pretty curious about why they use the word \"reduce\" in \"reducing the bounding volume to a ground bounding box and a height\".(But this question is trivial)",
    "626952": "Since bounding boxes are all in x,y plane and rotated around the z axis, you can reduce (hehe) calculation of 3d box IoU to calculating IoU along the z axis (vertical) and in x,y plane (on the \"ground\").\nEdit: in more detail, for such boxes, you can calculate IoU3d as multiplication of IoUGround and IoUHeight, where IoUGround would be usual 2d IoU of boxes projected on the ground, and IoUHeight would be the 1d IoU of boxes along the z axis. Now, why can we multiply IoUs here? Because we can multiply both intersection and union.",
    "627052": "Modern cars are usually narrower on the roof. Hence, a proper 3D bounding box wouldn't be a standard rectangular box. However, when it comes to robotics (incl. self-driving cars), we are more interested in finding occupied spaces to avoid collisions. Since we use some safety margin anyhow, we are only interested in the maximum extend of a vehicle. The volumetric difference between a perfect 3D bounding box and a simplified cuboid is small and not relevant .\n\nYou can think about the ground bounding box as looking on a vehicle from above and draw the boundaries. It doesn't really matter if the vehicle is widest at the top or the bottom, because we are projecting the maximum extend to 2D.  This reduces computational costs. Path planning in 2D is faster and more efficient. The 3D component is reduced to a simple height which allows to reconstruct a simple cuboid. The metric is reduced to compute IoUs of a line (height feature) and an area feature (ground box).",
    "628007": "Sample metric implementation: https://github.com/lyft/nuscenes-devkit/blob/master/lyft_dataset_sdk/eval/detection/mAP_evaluation.py",
    "628315": "Hi, the github code doesn't match the evaluation described in the competitition. Github implements a classic AP (with area under precision-recall curve), while in Kaggle just an accuracy/Critical Success Index is described. Which one is really used here ? Since it is said that confidence scores are not important, I guess it is the metric described here, but it has little to do with the github implementation then.",
    "628446": "Hi @thomasgilles - the Kaggle implementation is based on the implementation that Vladimir linked. Any discrepancies are in the explanation of the metric. I'll add a description of AP using area under the precision-recall curve, as that is definitely what we're using. Thanks!",
    "633078": "Thank you. ^^",
    "633080": "Thank you~",
    "633081": "Thank you ~~ ^ ^"
  },
  "source": "meta"
}