{
  "id": 123004,
  "title": "2nd Place Solution Summary (Private: 0.202 / Public: 0.205)",
  "url": "/competitions/3d-object-detection-for-autonomous-vehicles/discussion/123004",
  "author_name": "Kyle Lee",
  "post_date": "2019-12-24T04:26:04.361000",
  "votes": 22,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Before I describe&nbsp;my solution (which I will only summarize verbally), I'd like to thank the organizers and hosts for this interesting, well-organized competition which came with a rich, multi-data dataset that provided multiple solution avenues.&nbsp; And of course, I'd also like to thank the competitors for contributing to this competition and congratulations to all the winners and medal holders too.</p>\n\n<p>Also, apologies for not having written this sooner since I got caught up with multiple commitments after the competition - but better late than never.&nbsp; &nbsp;Particularly, my solution takes a slightly different approach wherein I actually did end up using 2D images effectively (compared to the other approaches which use only LIDAR), so it may be of interest.  </p>\n\n<h1>Methods</h1>\n\n<p>These are my primary methods:</p>\n\n<p><strong>Method 1.</strong>&nbsp; As with most of high scoring solutions which have been shared one method I used is based on Voxelnet with PointPillars (<a href=\"https://github.com/traveller59/second.pytorch\">https://github.com/traveller59/second.pytorch</a>).&nbsp; I started with this since it seemed to have the most fastest path to getting a first cut solution - NuScenes format, multi-class, similar classes, etc..&nbsp; In particular I re-used the multi-head version of the config (all.mhead.pp.config), and simply remapped existing NuScenes classes to Lyft classes (with the exception of some new classes which I just did a closest match, e.g. other_vehicle to trailer).&nbsp; Here I found that the following parameters made the most difference</p>\n\n<ul>\n<li><code>point_cloud_range</code> - [-100, -100, -5, 100, 100, 3]</li>\n<li><code>voxel_size</code> - in the final implementation I used an ensemble of models ranging from 0.1x0.1 to 0.25x0.25, with the best performing single model at 0.2x0.2.&nbsp;&nbsp;</li>\n<li><code>feature_map</code> - this has to be changed to match <code>voxel_size</code> above</li>\n<li><code>max_number_of_voxels</code> - I found that this was best set to as large as possible, e.g. &gt; 200k (depending on the amount of GPU memory available).</li>\n</ul>\n\n<p>Other than flip augmentations I did not find any additional advantage in additional augmentations, nor did I spend any time varying architecture parameters.&nbsp; I left the single model NMS parameters at the default and did not tune the thresholds, even though the downstream ensemble used rotated soft-NMS.</p>\n\n<p>In terms of external data, I did start my first model by training on the NuScenes dataset from scratch, which helped to (1) validate that the repository results were acceptable on NuScenes scoring (on average yes, but per class scores varied somewhat from what was reported) (2) accelerate subsequent training cycles as a pretrained model for Lyft dataset.&nbsp; &nbsp;Post-competition analysis showed that using the NuScenes dataset as an initial pretrained model might not have been necessary, but it was still helpful for me to validate the integrity of the repository.</p>\n\n<p>My best model had a private/public of 0.170/0.173 using a voxel size of 0.2x0.2.&nbsp; Including test time augmentation for this model.&nbsp; &nbsp;With test time augmentation (described below) this contributed to a private/public of 0.185/0.188.&nbsp;&nbsp;</p>\n\n<p>For the final ensemble I used a combination of 8 models of varying voxel sizes (and a few of them either used full data rather than train/val and one of them at some additional augs e.g. scaling).&nbsp; This LIDAR only ensemble gave a private/public of 0.188/0.191, which is not that much better from the best model TTA.</p>\n\n<p><strong>Method 2.</strong>&nbsp; Additionally, I used a 2D-&gt;3D approach that leveraged Frustum ConvNet (<a href=\"https://github.com/zhixinwang/frustum-convnet\">https://github.com/zhixinwang/frustum-convnet</a>) by using 2D boxes/proposals from conventional object detectors.&nbsp; Since this repository worked on KITTI format, this required a few corrections on the KITTI converter:</p>\n\n<ol>\n<li>The rotation had to be pi instead of pi/2 (this was corrected by the Kaggle community)</li>\n<li>Ego pose differences between LIDAR and cameras had to be taken into account (see PR:&nbsp;<a href=\"https://github.com/lyft/nuscenes-devkit/pull/75\">https://github.com/lyft/nuscenes-devkit/pull/75</a>)</li>\n</ol>\n\n<p>For the object detectors I used Detectron 2 and Tensorflow Object Detection API and trained the labels after converting from KITTI to COCO format.&nbsp; Pretrained models used from these object detectors include&nbsp;Faster-RCNN with X101-FPN for Detectron 2 and Faster-RCNN with Inception ResNet V2 (Atrous) for TF object detection API. </p>\n\n<p>I also wrote a de-converter from KITTI back to the Lyft submission format, and ignore overlapping boxes across camera views since they would be rectified by the ensemble process (which is rotated soft-NMS) described later.</p>\n\n<p>For the Frustum-ConvNet training process, I maintained 3 classes as per the original repository (instead of expanding this to all classes in a single training cycle, to minimize risk of modifications).&nbsp; I then mapped bicycle and motorcycle classes to 'Cyclist', pedestrian and animal classes to 'Pedestrian', and all other vehicle classes to 'Car', but modified their average sizes whenever the class was used.&nbsp; This required training each class mostly independently.</p>\n\n<p>Surprisingly, this approach by itself gave around private/public of 0.169/0.171 (ensemble of the 2 models from the frameworks x 2 checkpoints), which was pretty respectable - more importantly, combining this with approach (using SECOND) gave a final private/public of 0.202/0.205, which was would not have been achievable just using LIDAR alone since my LIDAR ensemble had already saturated.&nbsp;&nbsp;</p>\n\n<p>Further investigation showed that the best model for Frustum-ConvNet when compared to the SECOND approach has an improvement on pedestrian and bicycle classes, as shown below.&nbsp; Additionally, my guess is that this approach was able to detect some missed false negatives from the LIDAR approach for most of the classes which even though the overall score was lower.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F495818%2F6bb7a3583aee98d1efe199da3c8fb2e6%2F3d_vs_2d.png?generation=1577161276113708&amp;alt=media\" alt=\"\"></p>\n\n<h1>Ensembles</h1>\n\n<ol>\n<li><p>I used test time augmentation for the SECOND/LIDAR approach, which effectively is just orig, x flip, y flip, and x-y flip - preprocessing the point cloud by these flips, and then postprocessing the detected object boxes by the inverse.&nbsp; I did not apply TTA for the Frustum-ConvNet approach.</p></li>\n<li><p>I employed the same rotated soft-NMS for all ensembles, which include combining TTA predictions&nbsp;+ merging all LIDAR models and all Frustum-ConvNet models.&nbsp; In other words this acted as somewhat of a \"late fusion\" approach to both 3D and 2D-&gt;3D (via Frustum-ConvNet) models.&nbsp; A little bit of validation checking showed that using Gaussian function at a sigma of 0.85 gave the best validation score, so I stuck to this throughout the competition.&nbsp; I did not have success with 3D (i.e. including height intersection) but only IOU considering the rotated bounding boxes of the objects (it scored worse), and only considered soft-NMS within the same class.</p></li>\n</ol>\n\n<h1>What Didn't Work</h1>\n\n<p>I would be very interested if anyone enabled the following to work properly.&nbsp; It is also food for thought:</p>\n\n<ol>\n<li><p>Using high resolution maps to improve LIDAR/SECOND training - I had the idea of augmenting the point clouds with non road vs road points, but wasn't confident of this approach.&nbsp; I also tried weighing predicted objects directly using the map information to down-weigh non-road objects for certain classes, but I only got extremely marginal improvement.</p></li>\n<li><p>Using samples which were adjacent in time - I tried weighing predictions across pre and post adjacent frame samples for both non-road and road objects, and boost the predictions of objects which were seen across multiple adjacent frames.&nbsp; This did not help.&nbsp; Perhaps a tracking approach would have improved scores here.</p></li>\n<li><p>Animals - the 2D (Frustum-ConvNet) approach actually detected some animals (dogs) correctly, but the IOU evaluation was too sensitive for this to work well, so my score was 0 for both local and LB.</p></li>\n<li><p>PointRCNN - a last minute attempt to use this did not work well and I only got about slightly better than half the score for the car class (if I remember ~0.15 instead of 0.3+ for cars).</p></li>\n</ol>",
  "messages": [
    {
      "id": 701931,
      "postDate": "2019-12-24T04:26:04.363Z",
      "content": "<p>Before I describe&nbsp;my solution (which I will only summarize verbally), I'd like to thank the organizers and hosts for this interesting, well-organized competition which came with a rich, multi-data dataset that provided multiple solution avenues.&nbsp; And of course, I'd also like to thank the competitors for contributing to this competition and congratulations to all the winners and medal holders too.</p>\n\n<p>Also, apologies for not having written this sooner since I got caught up with multiple commitments after the competition - but better late than never.&nbsp; &nbsp;Particularly, my solution takes a slightly different approach wherein I actually did end up using 2D images effectively (compared to the other approaches which use only LIDAR), so it may be of interest.  </p>\n\n<h1>Methods</h1>\n\n<p>These are my primary methods:</p>\n\n<p><strong>Method 1.</strong>&nbsp; As with most of high scoring solutions which have been shared one method I used is based on Voxelnet with PointPillars (<a href=\"https://github.com/traveller59/second.pytorch\">https://github.com/traveller59/second.pytorch</a>).&nbsp; I started with this since it seemed to have the most fastest path to getting a first cut solution - NuScenes format, multi-class, similar classes, etc..&nbsp; In particular I re-used the multi-head version of the config (all.mhead.pp.config), and simply remapped existing NuScenes classes to Lyft classes (with the exception of some new classes which I just did a closest match, e.g. other_vehicle to trailer).&nbsp; Here I found that the following parameters made the most difference</p>\n\n<ul>\n<li><code>point_cloud_range</code> - [-100, -100, -5, 100, 100, 3]</li>\n<li><code>voxel_size</code> - in the final implementation I used an ensemble of models ranging from 0.1x0.1 to 0.25x0.25, with the best performing single model at 0.2x0.2.&nbsp;&nbsp;</li>\n<li><code>feature_map</code> - this has to be changed to match <code>voxel_size</code> above</li>\n<li><code>max_number_of_voxels</code> - I found that this was best set to as large as possible, e.g. &gt; 200k (depending on the amount of GPU memory available).</li>\n</ul>\n\n<p>Other than flip augmentations I did not find any additional advantage in additional augmentations, nor did I spend any time varying architecture parameters.&nbsp; I left the single model NMS parameters at the default and did not tune the thresholds, even though the downstream ensemble used rotated soft-NMS.</p>\n\n<p>In terms of external data, I did start my first model by training on the NuScenes dataset from scratch, which helped to (1) validate that the repository results were acceptable on NuScenes scoring (on average yes, but per class scores varied somewhat from what was reported) (2) accelerate subsequent training cycles as a pretrained model for Lyft dataset.&nbsp; &nbsp;Post-competition analysis showed that using the NuScenes dataset as an initial pretrained model might not have been necessary, but it was still helpful for me to validate the integrity of the repository.</p>\n\n<p>My best model had a private/public of 0.170/0.173 using a voxel size of 0.2x0.2.&nbsp; Including test time augmentation for this model.&nbsp; &nbsp;With test time augmentation (described below) this contributed to a private/public of 0.185/0.188.&nbsp;&nbsp;</p>\n\n<p>For the final ensemble I used a combination of 8 models of varying voxel sizes (and a few of them either used full data rather than train/val and one of them at some additional augs e.g. scaling).&nbsp; This LIDAR only ensemble gave a private/public of 0.188/0.191, which is not that much better from the best model TTA.</p>\n\n<p><strong>Method 2.</strong>&nbsp; Additionally, I used a 2D-&gt;3D approach that leveraged Frustum ConvNet (<a href=\"https://github.com/zhixinwang/frustum-convnet\">https://github.com/zhixinwang/frustum-convnet</a>) by using 2D boxes/proposals from conventional object detectors.&nbsp; Since this repository worked on KITTI format, this required a few corrections on the KITTI converter:</p>\n\n<ol>\n<li>The rotation had to be pi instead of pi/2 (this was corrected by the Kaggle community)</li>\n<li>Ego pose differences between LIDAR and cameras had to be taken into account (see PR:&nbsp;<a href=\"https://github.com/lyft/nuscenes-devkit/pull/75\">https://github.com/lyft/nuscenes-devkit/pull/75</a>)</li>\n</ol>\n\n<p>For the object detectors I used Detectron 2 and Tensorflow Object Detection API and trained the labels after converting from KITTI to COCO format.&nbsp; Pretrained models used from these object detectors include&nbsp;Faster-RCNN with X101-FPN for Detectron 2 and Faster-RCNN with Inception ResNet V2 (Atrous) for TF object detection API. </p>\n\n<p>I also wrote a de-converter from KITTI back to the Lyft submission format, and ignore overlapping boxes across camera views since they would be rectified by the ensemble process (which is rotated soft-NMS) described later.</p>\n\n<p>For the Frustum-ConvNet training process, I maintained 3 classes as per the original repository (instead of expanding this to all classes in a single training cycle, to minimize risk of modifications).&nbsp; I then mapped bicycle and motorcycle classes to 'Cyclist', pedestrian and animal classes to 'Pedestrian', and all other vehicle classes to 'Car', but modified their average sizes whenever the class was used.&nbsp; This required training each class mostly independently.</p>\n\n<p>Surprisingly, this approach by itself gave around private/public of 0.169/0.171 (ensemble of the 2 models from the frameworks x 2 checkpoints), which was pretty respectable - more importantly, combining this with approach (using SECOND) gave a final private/public of 0.202/0.205, which was would not have been achievable just using LIDAR alone since my LIDAR ensemble had already saturated.&nbsp;&nbsp;</p>\n\n<p>Further investigation showed that the best model for Frustum-ConvNet when compared to the SECOND approach has an improvement on pedestrian and bicycle classes, as shown below.&nbsp; Additionally, my guess is that this approach was able to detect some missed false negatives from the LIDAR approach for most of the classes which even though the overall score was lower.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F495818%2F6bb7a3583aee98d1efe199da3c8fb2e6%2F3d_vs_2d.png?generation=1577161276113708&amp;alt=media\" alt=\"\"></p>\n\n<h1>Ensembles</h1>\n\n<ol>\n<li><p>I used test time augmentation for the SECOND/LIDAR approach, which effectively is just orig, x flip, y flip, and x-y flip - preprocessing the point cloud by these flips, and then postprocessing the detected object boxes by the inverse.&nbsp; I did not apply TTA for the Frustum-ConvNet approach.</p></li>\n<li><p>I employed the same rotated soft-NMS for all ensembles, which include combining TTA predictions&nbsp;+ merging all LIDAR models and all Frustum-ConvNet models.&nbsp; In other words this acted as somewhat of a \"late fusion\" approach to both 3D and 2D-&gt;3D (via Frustum-ConvNet) models.&nbsp; A little bit of validation checking showed that using Gaussian function at a sigma of 0.85 gave the best validation score, so I stuck to this throughout the competition.&nbsp; I did not have success with 3D (i.e. including height intersection) but only IOU considering the rotated bounding boxes of the objects (it scored worse), and only considered soft-NMS within the same class.</p></li>\n</ol>\n\n<h1>What Didn't Work</h1>\n\n<p>I would be very interested if anyone enabled the following to work properly.&nbsp; It is also food for thought:</p>\n\n<ol>\n<li><p>Using high resolution maps to improve LIDAR/SECOND training - I had the idea of augmenting the point clouds with non road vs road points, but wasn't confident of this approach.&nbsp; I also tried weighing predicted objects directly using the map information to down-weigh non-road objects for certain classes, but I only got extremely marginal improvement.</p></li>\n<li><p>Using samples which were adjacent in time - I tried weighing predictions across pre and post adjacent frame samples for both non-road and road objects, and boost the predictions of objects which were seen across multiple adjacent frames.&nbsp; This did not help.&nbsp; Perhaps a tracking approach would have improved scores here.</p></li>\n<li><p>Animals - the 2D (Frustum-ConvNet) approach actually detected some animals (dogs) correctly, but the IOU evaluation was too sensitive for this to work well, so my score was 0 for both local and LB.</p></li>\n<li><p>PointRCNN - a last minute attempt to use this did not work well and I only got about slightly better than half the score for the car class (if I remember ~0.15 instead of 0.3+ for cars).</p></li>\n</ol>",
      "rawMarkdown": "Before I describe&nbsp;my solution (which I will only summarize verbally), I'd like to thank the organizers and hosts for this interesting, well-organized competition which came with a rich, multi-data dataset that provided multiple solution avenues.&nbsp; And of course, I'd also like to thank the competitors for contributing to this competition and congratulations to all the winners and medal holders too.\n\nAlso, apologies for not having written this sooner since I got caught up with multiple commitments after the competition - but better late than never.&nbsp; &nbsp;Particularly, my solution takes a slightly different approach wherein I actually did end up using 2D images effectively (compared to the other approaches which use only LIDAR), so it may be of interest.  \n\n# Methods\nThese are my primary methods:\n\n**Method 1.**&nbsp; As with most of high scoring solutions which have been shared one method I used is based on Voxelnet with PointPillars (https://github.com/traveller59/second.pytorch).&nbsp; I started with this since it seemed to have the most fastest path to getting a first cut solution - NuScenes format, multi-class, similar classes, etc..&nbsp; In particular I re-used the multi-head version of the config (all.mhead.pp.config), and simply remapped existing NuScenes classes to Lyft classes (with the exception of some new classes which I just did a closest match, e.g. other_vehicle to trailer).&nbsp; Here I found that the following parameters made the most difference\n\n- `point_cloud_range` - [-100, -100, -5, 100, 100, 3]\n- `voxel_size` - in the final implementation I used an ensemble of models ranging from 0.1x0.1 to 0.25x0.25, with the best performing single model at 0.2x0.2.&nbsp;&nbsp;\n- `feature_map` - this has to be changed to match `voxel_size` above\n- `max_number_of_voxels` - I found that this was best set to as large as possible, e.g. &gt; 200k (depending on the amount of GPU memory available).\n\nOther than flip augmentations I did not find any additional advantage in additional augmentations, nor did I spend any time varying architecture parameters.&nbsp; I left the single model NMS parameters at the default and did not tune the thresholds, even though the downstream ensemble used rotated soft-NMS.\n\nIn terms of external data, I did start my first model by training on the NuScenes dataset from scratch, which helped to (1) validate that the repository results were acceptable on NuScenes scoring (on average yes, but per class scores varied somewhat from what was reported) (2) accelerate subsequent training cycles as a pretrained model for Lyft dataset.&nbsp; &nbsp;Post-competition analysis showed that using the NuScenes dataset as an initial pretrained model might not have been necessary, but it was still helpful for me to validate the integrity of the repository.\n\nMy best model had a private/public of 0.170/0.173 using a voxel size of 0.2x0.2.&nbsp; Including test time augmentation for this model.&nbsp; &nbsp;With test time augmentation (described below) this contributed to a private/public of 0.185/0.188.&nbsp;&nbsp;\n\nFor the final ensemble I used a combination of 8 models of varying voxel sizes (and a few of them either used full data rather than train/val and one of them at some additional augs e.g. scaling).&nbsp; This LIDAR only ensemble gave a private/public of 0.188/0.191, which is not that much better from the best model TTA.\n\n**Method 2.**&nbsp; Additionally, I used a 2D-&gt;3D approach that leveraged Frustum ConvNet (https://github.com/zhixinwang/frustum-convnet) by using 2D boxes/proposals from conventional object detectors.&nbsp; Since this repository worked on KITTI format, this required a few corrections on the KITTI converter:\n\n1. The rotation had to be pi instead of pi/2 (this was corrected by the Kaggle community)\n2. Ego pose differences between LIDAR and cameras had to be taken into account (see PR:&nbsp;https://github.com/lyft/nuscenes-devkit/pull/75)\n\nFor the object detectors I used Detectron 2 and Tensorflow Object Detection API and trained the labels after converting from KITTI to COCO format.&nbsp; Pretrained models used from these object detectors include&nbsp;Faster-RCNN with X101-FPN for Detectron 2 and Faster-RCNN with Inception ResNet V2 (Atrous) for TF object detection API. \n\nI also wrote a de-converter from KITTI back to the Lyft submission format, and ignore overlapping boxes across camera views since they would be rectified by the ensemble process (which is rotated soft-NMS) described later.\n\nFor the Frustum-ConvNet training process, I maintained 3 classes as per the original repository (instead of expanding this to all classes in a single training cycle, to minimize risk of modifications).&nbsp; I then mapped bicycle and motorcycle classes to 'Cyclist', pedestrian and animal classes to 'Pedestrian', and all other vehicle classes to 'Car', but modified their average sizes whenever the class was used.&nbsp; This required training each class mostly independently.\n\nSurprisingly, this approach by itself gave around private/public of 0.169/0.171 (ensemble of the 2 models from the frameworks x 2 checkpoints), which was pretty respectable - more importantly, combining this with approach (using SECOND) gave a final private/public of 0.202/0.205, which was would not have been achievable just using LIDAR alone since my LIDAR ensemble had already saturated.&nbsp;&nbsp;\n\nFurther investigation showed that the best model for Frustum-ConvNet when compared to the SECOND approach has an improvement on pedestrian and bicycle classes, as shown below.&nbsp; Additionally, my guess is that this approach was able to detect some missed false negatives from the LIDAR approach for most of the classes which even though the overall score was lower.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F495818%2F6bb7a3583aee98d1efe199da3c8fb2e6%2F3d_vs_2d.png?generation=1577161276113708&amp;alt=media)\n\n# Ensembles\n\n1. I used test time augmentation for the SECOND/LIDAR approach, which effectively is just orig, x flip, y flip, and x-y flip - preprocessing the point cloud by these flips, and then postprocessing the detected object boxes by the inverse.&nbsp; I did not apply TTA for the Frustum-ConvNet approach.\n\n2. I employed the same rotated soft-NMS for all ensembles, which include combining TTA predictions&nbsp;+ merging all LIDAR models and all Frustum-ConvNet models.&nbsp; In other words this acted as somewhat of a \"late fusion\" approach to both 3D and 2D-&gt;3D (via Frustum-ConvNet) models.&nbsp; A little bit of validation checking showed that using Gaussian function at a sigma of 0.85 gave the best validation score, so I stuck to this throughout the competition.&nbsp; I did not have success with 3D (i.e. including height intersection) but only IOU considering the rotated bounding boxes of the objects (it scored worse), and only considered soft-NMS within the same class.\n\n# What Didn't Work\n\nI would be very interested if anyone enabled the following to work properly.&nbsp; It is also food for thought:\n\n1. Using high resolution maps to improve LIDAR/SECOND training - I had the idea of augmenting the point clouds with non road vs road points, but wasn't confident of this approach.&nbsp; I also tried weighing predicted objects directly using the map information to down-weigh non-road objects for certain classes, but I only got extremely marginal improvement.\n\n2. Using samples which were adjacent in time - I tried weighing predictions across pre and post adjacent frame samples for both non-road and road objects, and boost the predictions of objects which were seen across multiple adjacent frames.&nbsp; This did not help.&nbsp; Perhaps a tracking approach would have improved scores here.\n\n3. Animals - the 2D (Frustum-ConvNet) approach actually detected some animals (dogs) correctly, but the IOU evaluation was too sensitive for this to work well, so my score was 0 for both local and LB.\n\n4. PointRCNN - a last minute attempt to use this did not work well and I only got about slightly better than half the score for the car class (if I remember ~0.15 instead of 0.3+ for cars).",
      "votes": 22
    },
    {
      "id": 2213388,
      "postDate": "2023-04-07T14:43:23.510Z",
      "content": "<p>Nice work! </p>",
      "rawMarkdown": "Nice work! "
    },
    {
      "id": 716674,
      "postDate": "2020-01-12T04:58:18.107Z",
      "content": "<p>Thank you for sharing your approach! It's interesting to see how your solution incorporated 2-D image detection as opposed to the others I have seen that only used point-cloud models. </p>",
      "rawMarkdown": "Thank you for sharing your approach! It's interesting to see how your solution incorporated 2-D image detection as opposed to the others I have seen that only used point-cloud models. "
    },
    {
      "id": 780029,
      "postDate": "2020-03-19T23:09:26.633Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 701948,
      "postDate": "2019-12-24T05:05:44.697Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 703673,
      "postDate": "2019-12-26T13:32:42.747Z",
      "content": "<p>Nice work! Thanks</p>",
      "rawMarkdown": "Nice work! Thanks"
    }
  ],
  "comments": [
    {
      "id": 2213388,
      "author_name": "Tazria Helal",
      "author_url": "",
      "post_date": "2023-04-07T14:43:23.510000",
      "content": "<p>Nice work! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 716674,
      "author_name": "Dustin Z",
      "author_url": "",
      "post_date": "2020-01-12T04:58:18.107000",
      "content": "<p>Thank you for sharing your approach! It's interesting to see how your solution incorporated 2-D image detection as opposed to the others I have seen that only used point-cloud models. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 780029,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-19T23:09:26.633000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 701948,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-24T05:05:44.697000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 703673,
      "author_name": "Agastya Kommanamanchi",
      "author_url": "",
      "post_date": "2019-12-26T13:32:42.747000",
      "content": "<p>Nice work! Thanks</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "701931": "Before I describe&nbsp;my solution (which I will only summarize verbally), I'd like to thank the organizers and hosts for this interesting, well-organized competition which came with a rich, multi-data dataset that provided multiple solution avenues.&nbsp; And of course, I'd also like to thank the competitors for contributing to this competition and congratulations to all the winners and medal holders too.\n\nAlso, apologies for not having written this sooner since I got caught up with multiple commitments after the competition - but better late than never.&nbsp; &nbsp;Particularly, my solution takes a slightly different approach wherein I actually did end up using 2D images effectively (compared to the other approaches which use only LIDAR), so it may be of interest.  \n\n# Methods\nThese are my primary methods:\n\n**Method 1.**&nbsp; As with most of high scoring solutions which have been shared one method I used is based on Voxelnet with PointPillars (https://github.com/traveller59/second.pytorch).&nbsp; I started with this since it seemed to have the most fastest path to getting a first cut solution - NuScenes format, multi-class, similar classes, etc..&nbsp; In particular I re-used the multi-head version of the config (all.mhead.pp.config), and simply remapped existing NuScenes classes to Lyft classes (with the exception of some new classes which I just did a closest match, e.g. other_vehicle to trailer).&nbsp; Here I found that the following parameters made the most difference\n\n- `point_cloud_range` - [-100, -100, -5, 100, 100, 3]\n- `voxel_size` - in the final implementation I used an ensemble of models ranging from 0.1x0.1 to 0.25x0.25, with the best performing single model at 0.2x0.2.&nbsp;&nbsp;\n- `feature_map` - this has to be changed to match `voxel_size` above\n- `max_number_of_voxels` - I found that this was best set to as large as possible, e.g. &gt; 200k (depending on the amount of GPU memory available).\n\nOther than flip augmentations I did not find any additional advantage in additional augmentations, nor did I spend any time varying architecture parameters.&nbsp; I left the single model NMS parameters at the default and did not tune the thresholds, even though the downstream ensemble used rotated soft-NMS.\n\nIn terms of external data, I did start my first model by training on the NuScenes dataset from scratch, which helped to (1) validate that the repository results were acceptable on NuScenes scoring (on average yes, but per class scores varied somewhat from what was reported) (2) accelerate subsequent training cycles as a pretrained model for Lyft dataset.&nbsp; &nbsp;Post-competition analysis showed that using the NuScenes dataset as an initial pretrained model might not have been necessary, but it was still helpful for me to validate the integrity of the repository.\n\nMy best model had a private/public of 0.170/0.173 using a voxel size of 0.2x0.2.&nbsp; Including test time augmentation for this model.&nbsp; &nbsp;With test time augmentation (described below) this contributed to a private/public of 0.185/0.188.&nbsp;&nbsp;\n\nFor the final ensemble I used a combination of 8 models of varying voxel sizes (and a few of them either used full data rather than train/val and one of them at some additional augs e.g. scaling).&nbsp; This LIDAR only ensemble gave a private/public of 0.188/0.191, which is not that much better from the best model TTA.\n\n**Method 2.**&nbsp; Additionally, I used a 2D-&gt;3D approach that leveraged Frustum ConvNet (https://github.com/zhixinwang/frustum-convnet) by using 2D boxes/proposals from conventional object detectors.&nbsp; Since this repository worked on KITTI format, this required a few corrections on the KITTI converter:\n\n1. The rotation had to be pi instead of pi/2 (this was corrected by the Kaggle community)\n2. Ego pose differences between LIDAR and cameras had to be taken into account (see PR:&nbsp;https://github.com/lyft/nuscenes-devkit/pull/75)\n\nFor the object detectors I used Detectron 2 and Tensorflow Object Detection API and trained the labels after converting from KITTI to COCO format.&nbsp; Pretrained models used from these object detectors include&nbsp;Faster-RCNN with X101-FPN for Detectron 2 and Faster-RCNN with Inception ResNet V2 (Atrous) for TF object detection API. \n\nI also wrote a de-converter from KITTI back to the Lyft submission format, and ignore overlapping boxes across camera views since they would be rectified by the ensemble process (which is rotated soft-NMS) described later.\n\nFor the Frustum-ConvNet training process, I maintained 3 classes as per the original repository (instead of expanding this to all classes in a single training cycle, to minimize risk of modifications).&nbsp; I then mapped bicycle and motorcycle classes to 'Cyclist', pedestrian and animal classes to 'Pedestrian', and all other vehicle classes to 'Car', but modified their average sizes whenever the class was used.&nbsp; This required training each class mostly independently.\n\nSurprisingly, this approach by itself gave around private/public of 0.169/0.171 (ensemble of the 2 models from the frameworks x 2 checkpoints), which was pretty respectable - more importantly, combining this with approach (using SECOND) gave a final private/public of 0.202/0.205, which was would not have been achievable just using LIDAR alone since my LIDAR ensemble had already saturated.&nbsp;&nbsp;\n\nFurther investigation showed that the best model for Frustum-ConvNet when compared to the SECOND approach has an improvement on pedestrian and bicycle classes, as shown below.&nbsp; Additionally, my guess is that this approach was able to detect some missed false negatives from the LIDAR approach for most of the classes which even though the overall score was lower.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F495818%2F6bb7a3583aee98d1efe199da3c8fb2e6%2F3d_vs_2d.png?generation=1577161276113708&amp;alt=media)\n\n# Ensembles\n\n1. I used test time augmentation for the SECOND/LIDAR approach, which effectively is just orig, x flip, y flip, and x-y flip - preprocessing the point cloud by these flips, and then postprocessing the detected object boxes by the inverse.&nbsp; I did not apply TTA for the Frustum-ConvNet approach.\n\n2. I employed the same rotated soft-NMS for all ensembles, which include combining TTA predictions&nbsp;+ merging all LIDAR models and all Frustum-ConvNet models.&nbsp; In other words this acted as somewhat of a \"late fusion\" approach to both 3D and 2D-&gt;3D (via Frustum-ConvNet) models.&nbsp; A little bit of validation checking showed that using Gaussian function at a sigma of 0.85 gave the best validation score, so I stuck to this throughout the competition.&nbsp; I did not have success with 3D (i.e. including height intersection) but only IOU considering the rotated bounding boxes of the objects (it scored worse), and only considered soft-NMS within the same class.\n\n# What Didn't Work\n\nI would be very interested if anyone enabled the following to work properly.&nbsp; It is also food for thought:\n\n1. Using high resolution maps to improve LIDAR/SECOND training - I had the idea of augmenting the point clouds with non road vs road points, but wasn't confident of this approach.&nbsp; I also tried weighing predicted objects directly using the map information to down-weigh non-road objects for certain classes, but I only got extremely marginal improvement.\n\n2. Using samples which were adjacent in time - I tried weighing predictions across pre and post adjacent frame samples for both non-road and road objects, and boost the predictions of objects which were seen across multiple adjacent frames.&nbsp; This did not help.&nbsp; Perhaps a tracking approach would have improved scores here.\n\n3. Animals - the 2D (Frustum-ConvNet) approach actually detected some animals (dogs) correctly, but the IOU evaluation was too sensitive for this to work well, so my score was 0 for both local and LB.\n\n4. PointRCNN - a last minute attempt to use this did not work well and I only got about slightly better than half the score for the car class (if I remember ~0.15 instead of 0.3+ for cars).",
    "2213388": "Nice work! ",
    "716674": "Thank you for sharing your approach! It's interesting to see how your solution incorporated 2-D image detection as opposed to the others I have seen that only used point-cloud models. ",
    "780029": "",
    "701948": "",
    "703673": "Nice work! Thanks"
  }
}