{
  "id": 108793,
  "title": "Current SOTA Nets (KITTI, etc.)",
  "url": "/competitions/3d-object-detection-for-autonomous-vehicles/discussion/108793",
  "author_name": "",
  "post_date": "2019-09-14T01:03:10.055308400Z",
  "votes": 15,
  "comment_count": 3,
  "views": 0,
  "content": "<p>It might be a good idea to put together all relevant SOTA for 3D detection here. \nKITTI leaderboard is here: <a href=\"http://www.cvlibs.net/datasets/kitti/eval_object.php?obj_benchmark=bev\">http://www.cvlibs.net/datasets/kitti/eval_object.php?obj_benchmark=bev</a></p>\n\n<p>Not all the nets can be found, some are here:</p>\n\n<p>STD: Sparse-toDenase 3D Object detector <a href=\"https://arxiv.org/abs/1907.10471\">https://arxiv.org/abs/1907.10471</a></p>\n\n<p>Voxel-FPN:  <a href=\"https://arxiv.org/abs/1907.05286\">https://arxiv.org/abs/1907.05286</a></p>\n\n<p>Papers with code are here: <a href=\"https://paperswithcode.com/task/3d-object-detection\">https://paperswithcode.com/task/3d-object-detection</a></p>\n\n<p>Feel free to add in the thread. </p>\n\n<p>Not that I have resources for that, but some have, just wondering:  is pre-train on KITII and nuScenes, etc, allowed?</p>",
  "messages": [
    {
      "id": "626173",
      "postDate": "09/14/2019 01:03:10",
      "content": "<p>It might be a good idea to put together all relevant SOTA for 3D detection here. \nKITTI leaderboard is here: <a href=\"http://www.cvlibs.net/datasets/kitti/eval_object.php?obj_benchmark=bev\">http://www.cvlibs.net/datasets/kitti/eval_object.php?obj_benchmark=bev</a></p>\n\n<p>Not all the nets can be found, some are here:</p>\n\n<p>STD: Sparse-toDenase 3D Object detector <a href=\"https://arxiv.org/abs/1907.10471\">https://arxiv.org/abs/1907.10471</a></p>\n\n<p>Voxel-FPN:  <a href=\"https://arxiv.org/abs/1907.05286\">https://arxiv.org/abs/1907.05286</a></p>\n\n<p>Papers with code are here: <a href=\"https://paperswithcode.com/task/3d-object-detection\">https://paperswithcode.com/task/3d-object-detection</a></p>\n\n<p>Feel free to add in the thread. </p>\n\n<p>Not that I have resources for that, but some have, just wondering:  is pre-train on KITII and nuScenes, etc, allowed?</p>",
      "rawMarkdown": "It might be a good idea to put together all relevant SOTA for 3D detection here. \nKITTI leaderboard is here: http://www.cvlibs.net/datasets/kitti/eval_object.php?obj_benchmark=bev\n\nNot all the nets can be found, some are here:\n\nSTD: Sparse-toDenase 3D Object detector https://arxiv.org/abs/1907.10471\n\nVoxel-FPN:  https://arxiv.org/abs/1907.05286\n\nPapers with code are here: https://paperswithcode.com/task/3d-object-detection\n\nFeel free to add in the thread. \n\nNot that I have resources for that, but some have, just wondering:  is pre-train on KITII and nuScenes, etc, allowed?",
      "votes": null
    },
    {
      "id": "626187",
      "postDate": "09/14/2019 01:39:35",
      "content": "<p><a href=\"https://arxiv.org/abs/1908.02990\">Fast Point R-CNN</a></p>\n\n<p>abstract: \nWe present a unified, efficient and effective framework for point-cloud based 3D object detection. Our two-stage approach utilizes both voxel representation and raw point cloud data to exploit respective advantages. The first stage network, with voxel representation as input, only consists of light convolutional operations, producing a small number of high-quality initial predictions. Coordinate and indexed convolutional feature of each point in initial prediction are effectively fused with the attention mechanism, preserving both accurate localization and context information. The second stage works on interior points with their fused feature for further refining the prediction. Our method is evaluated on KITTI dataset, in terms of both 3D and Bird's Eye View (BEV) detection, and achieves state-of-the-arts with a 15FPS detection rate.</p>",
      "rawMarkdown": "[Fast Point R-CNN](https://arxiv.org/abs/1908.02990)\n\nabstract: \nWe present a unified, efficient and effective framework for point-cloud based 3D object detection. Our two-stage approach utilizes both voxel representation and raw point cloud data to exploit respective advantages. The first stage network, with voxel representation as input, only consists of light convolutional operations, producing a small number of high-quality initial predictions. Coordinate and indexed convolutional feature of each point in initial prediction are effectively fused with the attention mechanism, preserving both accurate localization and context information. The second stage works on interior points with their fused feature for further refining the prediction. Our method is evaluated on KITTI dataset, in terms of both 3D and Bird's Eye View (BEV) detection, and achieves state-of-the-arts with a 15FPS detection rate.",
      "votes": null
    },
    {
      "id": "627736",
      "postDate": "09/16/2019 10:04:50",
      "content": "<p>PointRCNN: <a href=\"https://github.com/sshaoshuai/PointRCNN\">https://github.com/sshaoshuai/PointRCNN</a>\nSecond: <a href=\"https://github.com/traveller59/second.pytorch\">https://github.com/traveller59/second.pytorch</a>\nFrustum ConvNet: <a href=\"https://github.com/zhixinwang/frustum-convnet\">https://github.com/zhixinwang/frustum-convnet</a></p>",
      "rawMarkdown": "PointRCNN: https://github.com/sshaoshuai/PointRCNN\nSecond: https://github.com/traveller59/second.pytorch\nFrustum ConvNet: https://github.com/zhixinwang/frustum-convnet",
      "votes": null
    },
    {
      "id": "627739",
      "postDate": "09/16/2019 10:07:30",
      "content": "<p>Damn, I read Point-RCNN paper yesterday, now got to know about Fast Point-RCNN, maybe next is Faster Point-RCNN :p then maybe Fast and furious Point-RCNN :p</p>",
      "rawMarkdown": "Damn, I read Point-RCNN paper yesterday, now got to know about Fast Point-RCNN, maybe next is Faster Point-RCNN :p then maybe Fast and furious Point-RCNN :p",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 626187,
      "author_name": "seefun",
      "author_url": "",
      "post_date": "09/14/2019 01:39:35",
      "content": "<p><a href=\"https://arxiv.org/abs/1908.02990\">Fast Point R-CNN</a></p>\n\n<p>abstract: \nWe present a unified, efficient and effective framework for point-cloud based 3D object detection. Our two-stage approach utilizes both voxel representation and raw point cloud data to exploit respective advantages. The first stage network, with voxel representation as input, only consists of light convolutional operations, producing a small number of high-quality initial predictions. Coordinate and indexed convolutional feature of each point in initial prediction are effectively fused with the attention mechanism, preserving both accurate localization and context information. The second stage works on interior points with their fused feature for further refining the prediction. Our method is evaluated on KITTI dataset, in terms of both 3D and Bird's Eye View (BEV) detection, and achieves state-of-the-arts with a 15FPS detection rate.</p>",
      "votes": null,
      "replies": [
        {
          "id": 627739,
          "author_name": "rishabhiitbhu",
          "author_url": "",
          "post_date": "09/16/2019 10:07:30",
          "content": "<p>Damn, I read Point-RCNN paper yesterday, now got to know about Fast Point-RCNN, maybe next is Faster Point-RCNN :p then maybe Fast and furious Point-RCNN :p</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 627736,
      "author_name": "rishabhiitbhu",
      "author_url": "",
      "post_date": "09/16/2019 10:04:50",
      "content": "<p>PointRCNN: <a href=\"https://github.com/sshaoshuai/PointRCNN\">https://github.com/sshaoshuai/PointRCNN</a>\nSecond: <a href=\"https://github.com/traveller59/second.pytorch\">https://github.com/traveller59/second.pytorch</a>\nFrustum ConvNet: <a href=\"https://github.com/zhixinwang/frustum-convnet\">https://github.com/zhixinwang/frustum-convnet</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "626173": "It might be a good idea to put together all relevant SOTA for 3D detection here. \nKITTI leaderboard is here: http://www.cvlibs.net/datasets/kitti/eval_object.php?obj_benchmark=bev\n\nNot all the nets can be found, some are here:\n\nSTD: Sparse-toDenase 3D Object detector https://arxiv.org/abs/1907.10471\n\nVoxel-FPN:  https://arxiv.org/abs/1907.05286\n\nPapers with code are here: https://paperswithcode.com/task/3d-object-detection\n\nFeel free to add in the thread. \n\nNot that I have resources for that, but some have, just wondering:  is pre-train on KITII and nuScenes, etc, allowed?",
    "626187": "[Fast Point R-CNN](https://arxiv.org/abs/1908.02990)\n\nabstract: \nWe present a unified, efficient and effective framework for point-cloud based 3D object detection. Our two-stage approach utilizes both voxel representation and raw point cloud data to exploit respective advantages. The first stage network, with voxel representation as input, only consists of light convolutional operations, producing a small number of high-quality initial predictions. Coordinate and indexed convolutional feature of each point in initial prediction are effectively fused with the attention mechanism, preserving both accurate localization and context information. The second stage works on interior points with their fused feature for further refining the prediction. Our method is evaluated on KITTI dataset, in terms of both 3D and Bird's Eye View (BEV) detection, and achieves state-of-the-arts with a 15FPS detection rate.",
    "627736": "PointRCNN: https://github.com/sshaoshuai/PointRCNN\nSecond: https://github.com/traveller59/second.pytorch\nFrustum ConvNet: https://github.com/zhixinwang/frustum-convnet",
    "627739": "Damn, I read Point-RCNN paper yesterday, now got to know about Fast Point-RCNN, maybe next is Faster Point-RCNN :p then maybe Fast and furious Point-RCNN :p"
  },
  "source": "meta"
}