{
  "id": 117269,
  "title": "3rd place solution [0.182 Private LB]",
  "url": "/competitions/3d-object-detection-for-autonomous-vehicles/discussion/117269",
  "author_name": "yukke42",
  "post_date": "2019-11-14T08:46:06.817000",
  "votes": 63,
  "comment_count": 20,
  "views": 0,
  "content": "<p>Thanks to organizers and Kaggle for this competition!</p>\n\n<p>Congratulations to all winners!</p>\n\n<p>Here is my solution:</p>\n\n<h3>Dataset &amp; Pre-processing</h3>\n\n<ul>\n<li>No external data and only point-cloud of this competition</li>\n<li>Animals and emergency vehicle classes were not used.</li>\n<li>The objects which have less than 5 points were ignored.</li>\n<li>The detection area was 100m x 75m for vehicle's classes (car, other_vehicle, truck and bus) and 100m x 50 for small object's classes (pedestrian, bicycle and motorcycle).</li>\n<li>Dataset was splitted bt StratifiedKFold using the number of objects for each class in a scenes.</li>\n</ul>\n\n<h3>Model</h3>\n\n<p>My network was the combination of VoxelNet [1] and PointPillars [2], and the implementation was based on <a href=\"https://github.com/traveller59/second.pytorch\">traveller59's second.pytorch</a>. The network utilized only FC and Conv2d, no Sparse Convolution or Deformable Convolution.</p>\n\n<p>Base model's pipeline:</p>\n\n<ol>\n<li>Point-cloud was splitted into voxels [0.25m x 0.25m x 0.75m]</li>\n<li>The same network of PointPillars' Pillar Feature Net was applied to each voxel and output channel size = 16 worked best for me.</li>\n<li>Voxel-representation features (C x D x H x W) were reshaped into pseudo-image features (C * D x H x W).</li>\n<li>Almost the same network of RPN in VoxelNet was used for vehicle's classes and this was based on <a href=\"https://github.com/traveller59/second.pytorch/blob/master/second/configs/nuscenes/all.pp.mhead.config\">this configuration file</a>. The first DeConv2d was replaced to Conv2d and other parameters were adjusted. <br>\nFor other small object's classes, three Conv2d were applied to the cropped feature map from the reshaped feature map. The cropped featuere map corresponding to the detection area.</li>\n<li>The prediction outputs were localization, classification and direction.</li>\n</ol>\n\n<p>Optional models:</p>\n\n<ol>\n<li>Pre-activation ResNet was used in RPN and this idea was from [3]</li>\n<li>RPN for small object's classes was changed. Conv2d x 1 (1st output) and Conv2d x 3 (2nd output) were applied to the cropped feature map then these outputs were concatenated. This was based on the original RPN in VoxelNet.</li>\n</ol>\n\n<h3>Post-processing</h3>\n\n<ul>\n<li>NMS was used to suppress the overlaps for each model.</li>\n<li>Three models were ensembled using Soft NMS. </li>\n<li>No score threshold.</li>\n</ul>\n\n<h3>Data Augmentation</h3>\n\n<ul>\n<li>global translation (x, y, z)</li>\n<li>global scaling</li>\n<li>rotation around z-axis</li>\n<li>mixup-augmentation <br>\nGround-truths were cropped and pasted to other samples (More explanation is in [4]). This augmentation raised the scores a lot for all classes. </li>\n</ul>\n\n<h3>No Improvement for me</h3>\n\n<ul>\n<li>I extended Weighted-boxes-fusion (<a href=\"https://www.kaggle.com/c/google-ai-open-images-object-detection-track/discussion/64633#latest-590283\">ZFTurbo's solution on the Open Image competition last year</a> [5]) to apply 3D Bonding  Boxes to ensemble different models, but the score was lower than Soft NMS. It was better than NMS.</li>\n<li>I tried to use a semantic map image (original, filtered) to filter the predictions, but both positive and false predictions were dropped.</li>\n<li>Point-cloud-coloring using raw images or feature maps extracted from 2D detection model did not work as additional feature for Point-cloud in my architectures.</li>\n</ul>\n\n<h3>Reference</h3>\n\n<ol>\n<li><p>VoxelNet: End-to-End Learning for Point Cloud Based 3D Object Detection [<a href=\"https://arxiv.org/abs/1711.06396\">arvix</a>]</p></li>\n<li><p>PointPillars: Fast Encoders for Object Detection from Point Clouds [<a href=\"https://arxiv.org/abs/1812.05784\">arvix</a>]</p></li>\n<li><p>End-to-End Multi-View Fusion for 3D Object Detection in LiDAR Point Clouds [<a href=\"https://arxiv.org/abs/1910.06528\">arvix</a>]</p></li>\n<li><p>Fast Point R-CNN [<a href=\"https://arxiv.org/abs/1908.02990\">arvix</a>]</p></li>\n<li><p>Weighted Boxes Fusion: ensembling boxes for object detection models [<a href=\"https://arxiv.org/abs/1910.13302\">arxiv</a>] [<a href=\"https://github.com/ZFTurbo/Weighted-Boxes-Fusion\">github</a>]</p></li>\n</ol>",
  "messages": [
    {
      "id": 672871,
      "postDate": "2019-11-14T08:46:06.817Z",
      "content": "<p>Thanks to organizers and Kaggle for this competition!</p>\n\n<p>Congratulations to all winners!</p>\n\n<p>Here is my solution:</p>\n\n<h3>Dataset &amp; Pre-processing</h3>\n\n<ul>\n<li>No external data and only point-cloud of this competition</li>\n<li>Animals and emergency vehicle classes were not used.</li>\n<li>The objects which have less than 5 points were ignored.</li>\n<li>The detection area was 100m x 75m for vehicle's classes (car, other_vehicle, truck and bus) and 100m x 50 for small object's classes (pedestrian, bicycle and motorcycle).</li>\n<li>Dataset was splitted bt StratifiedKFold using the number of objects for each class in a scenes.</li>\n</ul>\n\n<h3>Model</h3>\n\n<p>My network was the combination of VoxelNet [1] and PointPillars [2], and the implementation was based on <a href=\"https://github.com/traveller59/second.pytorch\">traveller59's second.pytorch</a>. The network utilized only FC and Conv2d, no Sparse Convolution or Deformable Convolution.</p>\n\n<p>Base model's pipeline:</p>\n\n<ol>\n<li>Point-cloud was splitted into voxels [0.25m x 0.25m x 0.75m]</li>\n<li>The same network of PointPillars' Pillar Feature Net was applied to each voxel and output channel size = 16 worked best for me.</li>\n<li>Voxel-representation features (C x D x H x W) were reshaped into pseudo-image features (C * D x H x W).</li>\n<li>Almost the same network of RPN in VoxelNet was used for vehicle's classes and this was based on <a href=\"https://github.com/traveller59/second.pytorch/blob/master/second/configs/nuscenes/all.pp.mhead.config\">this configuration file</a>. The first DeConv2d was replaced to Conv2d and other parameters were adjusted. <br>\nFor other small object's classes, three Conv2d were applied to the cropped feature map from the reshaped feature map. The cropped featuere map corresponding to the detection area.</li>\n<li>The prediction outputs were localization, classification and direction.</li>\n</ol>\n\n<p>Optional models:</p>\n\n<ol>\n<li>Pre-activation ResNet was used in RPN and this idea was from [3]</li>\n<li>RPN for small object's classes was changed. Conv2d x 1 (1st output) and Conv2d x 3 (2nd output) were applied to the cropped feature map then these outputs were concatenated. This was based on the original RPN in VoxelNet.</li>\n</ol>\n\n<h3>Post-processing</h3>\n\n<ul>\n<li>NMS was used to suppress the overlaps for each model.</li>\n<li>Three models were ensembled using Soft NMS. </li>\n<li>No score threshold.</li>\n</ul>\n\n<h3>Data Augmentation</h3>\n\n<ul>\n<li>global translation (x, y, z)</li>\n<li>global scaling</li>\n<li>rotation around z-axis</li>\n<li>mixup-augmentation <br>\nGround-truths were cropped and pasted to other samples (More explanation is in [4]). This augmentation raised the scores a lot for all classes. </li>\n</ul>\n\n<h3>No Improvement for me</h3>\n\n<ul>\n<li>I extended Weighted-boxes-fusion (<a href=\"https://www.kaggle.com/c/google-ai-open-images-object-detection-track/discussion/64633#latest-590283\">ZFTurbo's solution on the Open Image competition last year</a> [5]) to apply 3D Bonding  Boxes to ensemble different models, but the score was lower than Soft NMS. It was better than NMS.</li>\n<li>I tried to use a semantic map image (original, filtered) to filter the predictions, but both positive and false predictions were dropped.</li>\n<li>Point-cloud-coloring using raw images or feature maps extracted from 2D detection model did not work as additional feature for Point-cloud in my architectures.</li>\n</ul>\n\n<h3>Reference</h3>\n\n<ol>\n<li><p>VoxelNet: End-to-End Learning for Point Cloud Based 3D Object Detection [<a href=\"https://arxiv.org/abs/1711.06396\">arvix</a>]</p></li>\n<li><p>PointPillars: Fast Encoders for Object Detection from Point Clouds [<a href=\"https://arxiv.org/abs/1812.05784\">arvix</a>]</p></li>\n<li><p>End-to-End Multi-View Fusion for 3D Object Detection in LiDAR Point Clouds [<a href=\"https://arxiv.org/abs/1910.06528\">arvix</a>]</p></li>\n<li><p>Fast Point R-CNN [<a href=\"https://arxiv.org/abs/1908.02990\">arvix</a>]</p></li>\n<li><p>Weighted Boxes Fusion: ensembling boxes for object detection models [<a href=\"https://arxiv.org/abs/1910.13302\">arxiv</a>] [<a href=\"https://github.com/ZFTurbo/Weighted-Boxes-Fusion\">github</a>]</p></li>\n</ol>",
      "rawMarkdown": "Thanks to organizers and Kaggle for this competition!\n\nCongratulations to all winners!\n\n\n\nHere is my solution:\n\n### Dataset &amp; Pre-processing\n\n- No external data and only point-cloud of this competition\n- Animals and emergency vehicle classes were not used.\n- The objects which have less than 5 points were ignored.\n- The detection area was 100m x 75m for vehicle's classes (car, other_vehicle, truck and bus) and 100m x 50 for small object's classes (pedestrian, bicycle and motorcycle).\n- Dataset was splitted bt StratifiedKFold using the number of objects for each class in a scenes.\n\n\n\n### Model\n\nMy network was the combination of VoxelNet [1] and PointPillars [2], and the implementation was based on [traveller59's second.pytorch](https://github.com/traveller59/second.pytorch). The network utilized only FC and Conv2d, no Sparse Convolution or Deformable Convolution.\n\nBase model's pipeline:\n\n1. Point-cloud was splitted into voxels [0.25m x 0.25m x 0.75m]\n2. The same network of PointPillars' Pillar Feature Net was applied to each voxel and output channel size = 16 worked best for me.\n3. Voxel-representation features (C x D x H x W) were reshaped into pseudo-image features (C * D x H x W).\n4. Almost the same network of RPN in VoxelNet was used for vehicle's classes and this was based on [this configuration file](https://github.com/traveller59/second.pytorch/blob/master/second/configs/nuscenes/all.pp.mhead.config). The first DeConv2d was replaced to Conv2d and other parameters were adjusted.  \n   For other small object's classes, three Conv2d were applied to the cropped feature map from the reshaped feature map. The cropped featuere map corresponding to the detection area.\n5. The prediction outputs were localization, classification and direction.\n\n\n\nOptional models:\n\n1. Pre-activation ResNet was used in RPN and this idea was from [3]\n2.  RPN for small object's classes was changed. Conv2d x 1 (1st output) and Conv2d x 3 (2nd output) were applied to the cropped feature map then these outputs were concatenated. This was based on the original RPN in VoxelNet.\n\n\n\n### Post-processing\n\n- NMS was used to suppress the overlaps for each model.\n- Three models were ensembled using Soft NMS. \n- No score threshold.\n\n\n\n### Data Augmentation\n\n- global translation (x, y, z)\n- global scaling\n- rotation around z-axis\n- mixup-augmentation  \n  Ground-truths were cropped and pasted to other samples (More explanation is in [4]). This augmentation raised the scores a lot for all classes. \n\n\n\n### No Improvement for me\n\n- I extended Weighted-boxes-fusion ([ZFTurbo's solution on the Open Image competition last year](https://www.kaggle.com/c/google-ai-open-images-object-detection-track/discussion/64633#latest-590283) [5]) to apply 3D Bonding  Boxes to ensemble different models, but the score was lower than Soft NMS. It was better than NMS.\n- I tried to use a semantic map image (original, filtered) to filter the predictions, but both positive and false predictions were dropped.\n- Point-cloud-coloring using raw images or feature maps extracted from 2D detection model did not work as additional feature for Point-cloud in my architectures.\n\n\n\n### Reference\n\n1. VoxelNet: End-to-End Learning for Point Cloud Based 3D Object Detection [[arvix](https://arxiv.org/abs/1711.06396)]\n\n2. PointPillars: Fast Encoders for Object Detection from Point Clouds [[arvix](https://arxiv.org/abs/1812.05784)]\n\n3. End-to-End Multi-View Fusion for 3D Object Detection in LiDAR Point Clouds [[arvix](https://arxiv.org/abs/1910.06528)]\n\n4. Fast Point R-CNN [[arvix](https://arxiv.org/abs/1908.02990)]\n\n5. Weighted Boxes Fusion: ensembling boxes for object detection models [[arxiv](https://arxiv.org/abs/1910.13302)] [[github](https://github.com/ZFTurbo/Weighted-Boxes-Fusion)]\n",
      "votes": 63
    },
    {
      "id": 673685,
      "postDate": "2019-11-15T10:36:34.113Z",
      "content": "<p>Congrats!\nWould you mind to share your code or at least configs? </p>",
      "rawMarkdown": "Congrats!\nWould you mind to share your code or at least configs? ",
      "votes": 5
    },
    {
      "id": 673590,
      "postDate": "2019-11-15T07:39:14.920Z",
      "content": "<p>Congrats!</p>\n\n<p>You need to change the title to \"4th place solution\" now 😃 </p>",
      "rawMarkdown": "Congrats!\n\nYou need to change the title to \"4th place solution\" now 😃 ",
      "votes": 3
    },
    {
      "id": 679373,
      "postDate": "2019-11-22T17:41:58.610Z",
      "content": "<p>Maybe it will become the 1st solution soon :D</p>",
      "rawMarkdown": "Maybe it will become the 1st solution soon :D",
      "votes": 1
    },
    {
      "id": 678948,
      "postDate": "2019-11-22T05:04:38.017Z",
      "content": "<p><a href=\"/yukke42\">@yukke42</a> now this is 3rd place solution :). Congrats</p>",
      "rawMarkdown": "@yukke42 now this is 3rd place solution :). Congrats",
      "votes": 1
    },
    {
      "id": 673105,
      "postDate": "2019-11-14T14:19:45.810Z",
      "content": "<p>Congrats! Great result, my solution also used Point Pillars from second.pytorch. What did you have <code>point_cloud_range</code> range set to?</p>",
      "rawMarkdown": "Congrats! Great result, my solution also used Point Pillars from second.pytorch. What did you have `point_cloud_range` range set to?",
      "votes": 1,
      "replies": [
        {
          "id": 673533,
          "postDate": "2019-11-15T05:53:22.673Z",
          "content": "<p>Thanks!</p>\n\n<p>point_cloud_range = [-100, -75, -2, 100, 75, 5.5] and I used <code>flat_vehicle_coordinates=True</code> option in <code>get_sample_data</code>.</p>",
          "rawMarkdown": "Thanks!\n\npoint_cloud_range = [-100, -75, -2, 100, 75, 5.5] and I used `flat_vehicle_coordinates=True` option in `get_sample_data`.",
          "votes": 2
        }
      ]
    },
    {
      "id": 679422,
      "postDate": "2019-11-22T19:06:07.993Z",
      "content": "<p>Congrats to you <a href=\"/yukke42\">@yukke42</a> . I use almost the exactly same method. The only difference is that I forget to ignore the objects which have less than 5 points. I found Semantic Map help, which can filter out off-road vehicle with points as preprocessing and improve a little on motorcycle, bus, truck and other_vehicle. </p>",
      "rawMarkdown": "Congrats to you @yukke42 . I use almost the exactly same method. The only difference is that I forget to ignore the objects which have less than 5 points. I found Semantic Map help, which can filter out off-road vehicle with points as preprocessing and improve a little on motorcycle, bus, truck and other_vehicle. \n",
      "votes": 2
    },
    {
      "id": 676971,
      "postDate": "2019-11-19T17:04:12.533Z",
      "content": "<p>Congrats on your work and thank you for sharing! I am hoping to reproduce similar work, and I was wondering if you followed the nuscenes format or the KITTI format when implementing the pointpillars repo?</p>",
      "rawMarkdown": "Congrats on your work and thank you for sharing! I am hoping to reproduce similar work, and I was wondering if you followed the nuscenes format or the KITTI format when implementing the pointpillars repo?",
      "votes": 2
    },
    {
      "id": 672927,
      "postDate": "2019-11-14T09:48:45.447Z",
      "content": "<p>Thanks for the write-up!</p>\n\n<p>I thought Soft-NMS doesn't make any sense here since metric ignores confidence values! Basing on their formula from the Evaluation page, the score should be the same as without any NMS at all. But it seems that the actual metric implementation does not match that formula, am I correct?</p>",
      "rawMarkdown": "Thanks for the write-up!\n\nI thought Soft-NMS doesn't make any sense here since metric ignores confidence values! Basing on their formula from the Evaluation page, the score should be the same as without any NMS at all. But it seems that the actual metric implementation does not match that formula, am I correct?",
      "replies": [
        {
          "id": 672957,
          "postDate": "2019-11-14T10:37:17.307Z",
          "content": "<p>Thanks!</p>\n\n<p>In the evaluation page, </p>\n\n<p>&gt; In nearly all cases confidence will have no impact on scoring. </p>\n\n<p>\"In nearly all cases\" is very ambiguous and misleading. It doesn't say that the metric ignores the confidence values. At least, average precision has some impact on the evaluation metric.</p>\n\n<p>If this competition uses <a href=\"https://github.com/lyft/nuscenes-devkit/blob/master/lyft_dataset_sdk/eval/detection/mAP_evaluation.py\">the same metric implementation on official devkit</a>, the metric changes.\nIn my case, <code>0.837</code> drops to <code>0.783</code> by changing confidence from <code>c</code> to <code>1 - c</code></p>",
          "rawMarkdown": "Thanks!\n\nIn the evaluation page, \n\n&gt; In nearly all cases confidence will have no impact on scoring. \n\n\"In nearly all cases\" is very ambiguous and misleading. It doesn't say that the metric ignores the confidence values. At least, average precision has some impact on the evaluation metric.\n\nIf this competition uses [the same metric implementation on official devkit](https://github.com/lyft/nuscenes-devkit/blob/master/lyft_dataset_sdk/eval/detection/mAP_evaluation.py), the metric changes.\nIn my case, `0.837` drops to `0.783` by changing confidence from `c` to `1 - c`\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 679717,
      "postDate": "2019-11-23T08:47:23.243Z",
      "content": "<p>Congratulations</p>",
      "rawMarkdown": "Congratulations"
    },
    {
      "id": 673186,
      "postDate": "2019-11-14T16:20:25.627Z",
      "content": "<p>Congratulations!!🎉 Thank you very much for sharing!!😄 </p>",
      "rawMarkdown": "Congratulations!!🎉 Thank you very much for sharing!!😄 "
    },
    {
      "id": 672993,
      "postDate": "2019-11-14T11:19:20.737Z",
      "content": "<p>what were training / eval times like?</p>",
      "rawMarkdown": "what were training / eval times like?",
      "replies": [
        {
          "id": 673534,
          "postDate": "2019-11-15T05:54:02.837Z",
          "content": "<p>For 3 ~ 4 days with single GTX 1080Ti with batch_size = 2 or 3</p>",
          "rawMarkdown": "For 3 ~ 4 days with single GTX 1080Ti with batch_size = 2 or 3",
          "votes": 3
        }
      ]
    },
    {
      "id": 672873,
      "postDate": "2019-11-14T08:48:03.043Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 672959,
          "postDate": "2019-11-14T10:39:13.687Z",
          "content": "<p>Thanks!</p>",
          "rawMarkdown": "Thanks!",
          "votes": 1
        }
      ]
    },
    {
      "id": 676963,
      "postDate": "2019-11-19T16:53:07.437Z",
      "content": "<p>Wow ! Thanks for sharing! Congrats !</p>",
      "rawMarkdown": "Wow ! Thanks for sharing! Congrats !"
    },
    {
      "id": 673747,
      "postDate": "2019-11-15T12:33:59.080Z",
      "content": "<p>Thanks for sharing！</p>",
      "rawMarkdown": "Thanks for sharing！"
    },
    {
      "id": 673626,
      "postDate": "2019-11-15T08:37:25.610Z",
      "content": "<p>Thanks for sharing. It was very useful</p>",
      "rawMarkdown": "Thanks for sharing. It was very useful"
    },
    {
      "id": 673415,
      "postDate": "2019-11-15T00:09:51.483Z",
      "content": "<p>Congrats and thank you for sharing.</p>",
      "rawMarkdown": "Congrats and thank you for sharing."
    }
  ],
  "comments": [
    {
      "id": 673685,
      "author_name": "Azat Akhtyamov",
      "author_url": "",
      "post_date": "2019-11-15T10:36:34.113000",
      "content": "<p>Congrats!\nWould you mind to share your code or at least configs? </p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 673590,
      "author_name": "Alvor",
      "author_url": "",
      "post_date": "2019-11-15T07:39:14.920000",
      "content": "<p>Congrats!</p>\n\n<p>You need to change the title to \"4th place solution\" now 😃 </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 679373,
      "author_name": "Artyom Palvelev",
      "author_url": "",
      "post_date": "2019-11-22T17:41:58.610000",
      "content": "<p>Maybe it will become the 1st solution soon :D</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 678948,
      "author_name": "Ram Ramrakhya",
      "author_url": "",
      "post_date": "2019-11-22T05:04:38.017000",
      "content": "<p><a href=\"/yukke42\">@yukke42</a> now this is 3rd place solution :). Congrats</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 673105,
      "author_name": "Jack Vial",
      "author_url": "",
      "post_date": "2019-11-14T14:19:45.810000",
      "content": "<p>Congrats! Great result, my solution also used Point Pillars from second.pytorch. What did you have <code>point_cloud_range</code> range set to?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 673533,
          "author_name": "yukke42",
          "author_url": "",
          "post_date": "2019-11-15T05:53:22.673000",
          "content": "<p>Thanks!</p>\n\n<p>point_cloud_range = [-100, -75, -2, 100, 75, 5.5] and I used <code>flat_vehicle_coordinates=True</code> option in <code>get_sample_data</code>.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 679422,
      "author_name": "Hanxiao Deng",
      "author_url": "",
      "post_date": "2019-11-22T19:06:07.993000",
      "content": "<p>Congrats to you <a href=\"/yukke42\">@yukke42</a> . I use almost the exactly same method. The only difference is that I forget to ignore the objects which have less than 5 points. I found Semantic Map help, which can filter out off-road vehicle with points as preprocessing and improve a little on motorcycle, bus, truck and other_vehicle. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 676971,
      "author_name": "Brian Lee",
      "author_url": "",
      "post_date": "2019-11-19T17:04:12.533000",
      "content": "<p>Congrats on your work and thank you for sharing! I am hoping to reproduce similar work, and I was wondering if you followed the nuscenes format or the KITTI format when implementing the pointpillars repo?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 672927,
      "author_name": "Artyom Palvelev",
      "author_url": "",
      "post_date": "2019-11-14T09:48:45.447000",
      "content": "<p>Thanks for the write-up!</p>\n\n<p>I thought Soft-NMS doesn't make any sense here since metric ignores confidence values! Basing on their formula from the Evaluation page, the score should be the same as without any NMS at all. But it seems that the actual metric implementation does not match that formula, am I correct?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 672957,
          "author_name": "yukke42",
          "author_url": "",
          "post_date": "2019-11-14T10:37:17.307000",
          "content": "<p>Thanks!</p>\n\n<p>In the evaluation page, </p>\n\n<p>&gt; In nearly all cases confidence will have no impact on scoring. </p>\n\n<p>\"In nearly all cases\" is very ambiguous and misleading. It doesn't say that the metric ignores the confidence values. At least, average precision has some impact on the evaluation metric.</p>\n\n<p>If this competition uses <a href=\"https://github.com/lyft/nuscenes-devkit/blob/master/lyft_dataset_sdk/eval/detection/mAP_evaluation.py\">the same metric implementation on official devkit</a>, the metric changes.\nIn my case, <code>0.837</code> drops to <code>0.783</code> by changing confidence from <code>c</code> to <code>1 - c</code></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 679717,
      "author_name": "YangLei",
      "author_url": "",
      "post_date": "2019-11-23T08:47:23.243000",
      "content": "<p>Congratulations</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 673186,
      "author_name": "Miyabon",
      "author_url": "",
      "post_date": "2019-11-14T16:20:25.627000",
      "content": "<p>Congratulations!!🎉 Thank you very much for sharing!!😄 </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 672993,
      "author_name": "oarph",
      "author_url": "",
      "post_date": "2019-11-14T11:19:20.737000",
      "content": "<p>what were training / eval times like?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 673534,
          "author_name": "yukke42",
          "author_url": "",
          "post_date": "2019-11-15T05:54:02.837000",
          "content": "<p>For 3 ~ 4 days with single GTX 1080Ti with batch_size = 2 or 3</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 672873,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-14T08:48:03.043000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 672959,
          "author_name": "yukke42",
          "author_url": "",
          "post_date": "2019-11-14T10:39:13.687000",
          "content": "<p>Thanks!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 676963,
      "author_name": "Caesar Lupum",
      "author_url": "",
      "post_date": "2019-11-19T16:53:07.437000",
      "content": "<p>Wow ! Thanks for sharing! Congrats !</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 673747,
      "author_name": "Xianzhong",
      "author_url": "",
      "post_date": "2019-11-15T12:33:59.080000",
      "content": "<p>Thanks for sharing！</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 673626,
      "author_name": "Ochiroo",
      "author_url": "",
      "post_date": "2019-11-15T08:37:25.610000",
      "content": "<p>Thanks for sharing. It was very useful</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 673415,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2019-11-15T00:09:51.483000",
      "content": "<p>Congrats and thank you for sharing.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "672871": "Thanks to organizers and Kaggle for this competition!\n\nCongratulations to all winners!\n\n\n\nHere is my solution:\n\n### Dataset &amp; Pre-processing\n\n- No external data and only point-cloud of this competition\n- Animals and emergency vehicle classes were not used.\n- The objects which have less than 5 points were ignored.\n- The detection area was 100m x 75m for vehicle's classes (car, other_vehicle, truck and bus) and 100m x 50 for small object's classes (pedestrian, bicycle and motorcycle).\n- Dataset was splitted bt StratifiedKFold using the number of objects for each class in a scenes.\n\n\n\n### Model\n\nMy network was the combination of VoxelNet [1] and PointPillars [2], and the implementation was based on [traveller59's second.pytorch](https://github.com/traveller59/second.pytorch). The network utilized only FC and Conv2d, no Sparse Convolution or Deformable Convolution.\n\nBase model's pipeline:\n\n1. Point-cloud was splitted into voxels [0.25m x 0.25m x 0.75m]\n2. The same network of PointPillars' Pillar Feature Net was applied to each voxel and output channel size = 16 worked best for me.\n3. Voxel-representation features (C x D x H x W) were reshaped into pseudo-image features (C * D x H x W).\n4. Almost the same network of RPN in VoxelNet was used for vehicle's classes and this was based on [this configuration file](https://github.com/traveller59/second.pytorch/blob/master/second/configs/nuscenes/all.pp.mhead.config). The first DeConv2d was replaced to Conv2d and other parameters were adjusted.  \n   For other small object's classes, three Conv2d were applied to the cropped feature map from the reshaped feature map. The cropped featuere map corresponding to the detection area.\n5. The prediction outputs were localization, classification and direction.\n\n\n\nOptional models:\n\n1. Pre-activation ResNet was used in RPN and this idea was from [3]\n2.  RPN for small object's classes was changed. Conv2d x 1 (1st output) and Conv2d x 3 (2nd output) were applied to the cropped feature map then these outputs were concatenated. This was based on the original RPN in VoxelNet.\n\n\n\n### Post-processing\n\n- NMS was used to suppress the overlaps for each model.\n- Three models were ensembled using Soft NMS. \n- No score threshold.\n\n\n\n### Data Augmentation\n\n- global translation (x, y, z)\n- global scaling\n- rotation around z-axis\n- mixup-augmentation  \n  Ground-truths were cropped and pasted to other samples (More explanation is in [4]). This augmentation raised the scores a lot for all classes. \n\n\n\n### No Improvement for me\n\n- I extended Weighted-boxes-fusion ([ZFTurbo's solution on the Open Image competition last year](https://www.kaggle.com/c/google-ai-open-images-object-detection-track/discussion/64633#latest-590283) [5]) to apply 3D Bonding  Boxes to ensemble different models, but the score was lower than Soft NMS. It was better than NMS.\n- I tried to use a semantic map image (original, filtered) to filter the predictions, but both positive and false predictions were dropped.\n- Point-cloud-coloring using raw images or feature maps extracted from 2D detection model did not work as additional feature for Point-cloud in my architectures.\n\n\n\n### Reference\n\n1. VoxelNet: End-to-End Learning for Point Cloud Based 3D Object Detection [[arvix](https://arxiv.org/abs/1711.06396)]\n\n2. PointPillars: Fast Encoders for Object Detection from Point Clouds [[arvix](https://arxiv.org/abs/1812.05784)]\n\n3. End-to-End Multi-View Fusion for 3D Object Detection in LiDAR Point Clouds [[arvix](https://arxiv.org/abs/1910.06528)]\n\n4. Fast Point R-CNN [[arvix](https://arxiv.org/abs/1908.02990)]\n\n5. Weighted Boxes Fusion: ensembling boxes for object detection models [[arxiv](https://arxiv.org/abs/1910.13302)] [[github](https://github.com/ZFTurbo/Weighted-Boxes-Fusion)]\n",
    "673685": "Congrats!\nWould you mind to share your code or at least configs? ",
    "673590": "Congrats!\n\nYou need to change the title to \"4th place solution\" now 😃 ",
    "679373": "Maybe it will become the 1st solution soon :D",
    "678948": "@yukke42 now this is 3rd place solution :). Congrats",
    "673105": "Congrats! Great result, my solution also used Point Pillars from second.pytorch. What did you have `point_cloud_range` range set to?",
    "679422": "Congrats to you @yukke42 . I use almost the exactly same method. The only difference is that I forget to ignore the objects which have less than 5 points. I found Semantic Map help, which can filter out off-road vehicle with points as preprocessing and improve a little on motorcycle, bus, truck and other_vehicle. \n",
    "676971": "Congrats on your work and thank you for sharing! I am hoping to reproduce similar work, and I was wondering if you followed the nuscenes format or the KITTI format when implementing the pointpillars repo?",
    "672927": "Thanks for the write-up!\n\nI thought Soft-NMS doesn't make any sense here since metric ignores confidence values! Basing on their formula from the Evaluation page, the score should be the same as without any NMS at all. But it seems that the actual metric implementation does not match that formula, am I correct?",
    "679717": "Congratulations",
    "673186": "Congratulations!!🎉 Thank you very much for sharing!!😄 ",
    "672993": "what were training / eval times like?",
    "672873": "",
    "676963": "Wow ! Thanks for sharing! Congrats !",
    "673747": "Thanks for sharing！",
    "673626": "Thanks for sharing. It was very useful",
    "673415": "Congrats and thank you for sharing."
  }
}