{
  "id": 133895,
  "title": "Summary of the competition from the host",
  "url": "/competitions/3d-object-detection-for-autonomous-vehicles/discussion/133895",
  "author_name": "",
  "post_date": "2020-03-04T21:16:26.505561400Z",
  "votes": 88,
  "comment_count": 9,
  "views": 0,
  "content": "<h1>Introduction</h1>\n<p>Hello! My name is <a href=\"https://medium.com/lyftlevel5/employee-spotlight-vladimir-iglovikov-11b71ed02bc\" target=\"_blank\">Vladimir Iglovikov</a>, and I am a Computer Vision Engineer at Lyft Level 5. Here, I work on a diverse set of projects relating to building, testing, deploying, and refining deep learning models for mapping and perception related tasks. </p>\n<p>I am <a href=\"http://ternaus.blog/interview/2018/08/30/ama.html\" target=\"_blank\">Kaggle Grandmaster</a>, which means that I have some experience of <a href=\"http://ternaus.blog/career/2020/01/08/How-I-Found-my-current-job.html\" target=\"_blank\">participating</a> in machine learning competitions. As a result, I was very pleased to join a team of colleagues at Level 5 who wanted to create a new machine learning competition for Lyft.</p>\n<p>After carefully putting together a dataset and an objective, we finally put together the Lyft 3D Object Detection for Autonomous Vehicles Challenge. Over the course of this challenge, over 500 teams plugged away at achieving a top leaderboard spot. Now that the challenge has finished and the dust has settled, we would like to share with you the top solutions.</p>\n<p>I would also like to discuss what we have learned throughout the process as organizers of the competition. There are many <a href=\"https://docs.google.com/a/chalearn.org/viewer?a=v&amp;pid=sites&amp;srcid=Y2hhbGVhcm4ub3JnfHdvcmtzaG9wfGd4OjQwYzBhY2Q1N2YxNGE3YWI\" target=\"_blank\">difficulties</a> to overcome when developing a machine learning challenge. From data leaks to issues with leaderboard scoring, it is common to make critical mistakes. I have provided much feedback to organizers when I was a participant in different machine learning competitions. Now it was time to get such feedback myself.</p>\n<p>This blog post will have two parts. In the first introductory part, I will tell you how this project was started and the challenges we ran into. In the second technical part, I will describe the solutions from the top three teams in a level of depth and rigor that I would have wanted to learn from if I had been a participant.</p>\n<h1>Why does Lyft need a Machine Learning Competition?</h1>\n<p>When I go to academic conferences I see a lot of papers that have the keywords “autonomous” and “self-driving”. For sure, all these academic works are related to the autonomous industry in some way, but for many of them, the relationship is rather weak.</p>\n<p>One of the reasons for this gap may be that the research community does not have access to the same data as the industry. There are many academic datasets that rely solely on camera images. Meanwhile, companies, including Lyft, use a combination of camera, lidar, and radar sensors, all of which are crucial for the construction of high definition maps or the functioning of autonomy algorithms. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2Fd280f4af3e91991ad641606967cd0bbe%2FScreenshot%20from%202020-03-04%2011-28-24.png?generation=1583350172278513&amp;alt=media\" alt=\"\"></p>\n<p>Prior to 2019, the <a href=\"http://www.cvlibs.net/datasets/kitti/\" target=\"_blank\">KITTI</a> dataset was the most useful dataset within the research community, that provided both camera and lidar data. KITTI was far ahead of its time and has been a staple in the self-driving industry even today, but it has its own limitations. </p>\n<p>In 2019 the field shifted rapidly. Lyft, along with some other self-driving companies, released their own, much larger datasets. This was a great move, but it is not enough to just release the dataset. It is also necessary to promote it and engage with the community. By going these extra steps, a dataset is better able to foster collaboration within the machine learning community. This can lead to new GitHub repositories, academic papers, or blog posts. Organizing a competition is one of the most popular ways of realizing these benefits.</p>\n<p>When Aptiv released their <a href=\"https://www.nuscenes.org/overview\" target=\"_blank\">dataset</a> in March 2019, they launched a competition as part of the WAD workshop at CVPR 2019. In addition to this, Aptiv released an <a href=\"https://github.com/nutonomy/nuscenes-devkit\" target=\"_blank\">SDK</a> with useful helper functions and visualization code. For <a href=\"https://www.nuscenes.org/object-detection/?externalData=all&amp;mapData=all&amp;modalities=Any\" target=\"_blank\">evaluation</a>, organizers used a weighted sum of the metrics responsible for detection, orientation, velocity, and attribute estimation.</p>\n<p>Lyft <a href=\"https://medium.com/lyftlevel5/unlocking-access-to-self-driving-research-the-lyft-level-5-dataset-and-competition-d487c27b1b6c\" target=\"_blank\">released</a> the dataset slightly later. Since the data was structurally similar to those in the Aptiv dataset, we decided to use the same format.</p>\n<p>Preparing a good dataset is a big challenge. 90% of the success of the competition is defined by the quality of the data. I would like to thank <strong>Robert Kesten, Muhammad Usman, John Houston, Tirth Pandya, K. Nadhamuni, Ana Ferreira, Michael Yuan, Ben Low, Ashesh Jain, Peter Ondrushka, Sami Omari, Shaswat Shah, Amruta Kulkarni, Alex Kazakova, Charlotte Tao, Lucas Platinsky, William Jiang, and Vinay Shet</strong> for preparing and releasing the dataset.</p>\n<h1>How did we prepare for the competition?</h1>\n<p>From the beginning, there was a plan to organize a challenge as a part of the competition track of NeurIPS 2019. We needed to decide who would host the competition for us. There were two options that we considered:</p>\n<ol>\n<li>Host ourselves: release the dataset, release the code for the metric, write a blog post about the challenge, provide prize money, and ask participants to send their predictions via email.</li>\n<li>Host through Kaggle. (I am not affiliated with Kaggle, but at the moment is the most established platform for machine competitions in the world). Kaggle would host the platform, the dataset, forum, real-time leaderboard, and notebooks in which participants may perform exploratory data analysis and even train models. There are other choices besides Kaggle, but they are less popular.</li>\n</ol>\n<p>Both strategies have their own advantages, but the first one is more common. It is commonly used when organizers have a small budget or Kaggle cannot support some sort of functionality, like submitting a Docker container instead of a CSV file. Most of the challenges in competition tracks in conferences like NeurIPS, ICCV, ECCV, CVPR, ICML, KDD, MICCAI, etc follow the first path.</p>\n<p>We decided to go with Kaggle. Although we knew we could host a quality competition ourselves, we wanted to make community engagement our highest priority.</p>\n<p>I believe that the success of the competition from the standpoint of the organizers is measured not by the titles and affiliations of the participants, nor the strength of the solutions, but rather by the number of active participants. If you only get 3-20 teams, you did not do a very good job. If the number is in the range 100-1000 you are doing well. Over 1000 teams and you’re doing extremely well! The background of the participants should not matter; whether they are from academia, industry, or are self-taught, everyone should be welcome.</p>\n<p>Such a philosophy helps to further democratize the field, elevating the skill of the community as a whole while also encouraging the development of diverse new methods and applications.  </p>\n<p>In my experience, the number of participants in Kaggle competitions is quite large, on the order of hundreds or thousands. The knowledge accumulated in Kaggle kernels and discussion threads became a great reference point for people investigating similar problems in the future.</p>\n<p>Following this way of thinking, we reached out to Kaggle and our partnership started.</p>\n<p>There are many different tasks that one can try to solve on the Lyft dataset: 3D object detection, lidar point segmentation, 3D tracking, orientation, and velocity prediction, etc. All of them are interesting tasks.</p>\n<p>This competition was the first ML challenge with this data type hosted on Kaggle. Its purpose was exploratory by nature, providing a testing ground for the ML community to become accustomed to this format of data. It is nice to get winning solutions that are closely related to the tasks that you face in production, but for this challenge it was secondary. We decided to go with the simplest possible task that one can have on this type of data: 3D object detection with the <a href=\"https://github.com/lyft/nuscenes-devkit/blob/master/lyft_dataset_sdk/eval/detection/mAP_evaluation.py\" target=\"_blank\">3D version of the COCO mAP</a> as a metric. This turned out to be very convenient as it helped to bridge the gap between more familiar bounding box problems and this new challenge. It was also nice because it could be easily demonstrated in Python and reimplemented in the data scientist’s language of choice. We also had a limited timeline, so avoiding the potential for bugs was a great plus.</p>\n<h1>Sample model</h1>\n<p>We expected that most of the people joining the competition would not have hands-on experience with this type of data. Therefore we wanted to cater specifically to participants who were motivated by learning and skill growth. To me, machine Learning is an applied discipline, so it makes life much easier when there is <a href=\"https://towardsdatascience.com/ask-me-anything-session-with-a-kaggle-grandmaster-vladimir-i-iglovikov-942ad6a06acd\" target=\"_blank\">an example to learn from</a>. Taking someone else's end-to-end code and playing with it was always a good way to be less overwhelmed with the volume of new things you need to learn and implement correctly.</p>\n<p><a href=\"https://www.linkedin.com/in/guido-zuidhof-377b6947/\" target=\"_blank\">Guido Zuidhof</a>, a <a href=\"https://www.kaggle.com/gzuidhof\" target=\"_blank\">Kaggle Master</a> and colleague in the AV Research team at Lyft Level 5 prepared a <a href=\"https://www.kaggle.com/gzuidhof/reference-model\" target=\"_blank\">Kaggle kernel with a sample model</a> that many of the participants used in their solutions.</p>\n<h1>Model limitations</h1>\n<p>In this challenge, I decided to maximize the flexibility given to participants with respect to model development:</p>\n<ol>\n<li>No limitations on the hardware.</li>\n<li>No limitations on the model size and inference time.</li>\n<li>Any external data or pre-trained models as long as they are listed in the <a href=\"https://www.kaggle.com/c/3d-object-detection-for-autonomous-vehicles/discussion/109361\" target=\"_blank\">special thread at the discussion forum</a>.</li>\n</ol>\n<p>It is normally a good idea to enforce constraints on model development such that it is more compatible with a production environment, but I decided against this.</p>\n<p>I do not know any examples of when a winning ML competition solution made it all the way to production. I did not expect that this challenge would be any different. But this was not our goal. At the end of the day, the challenges that we face are different from the competition task. We were looking to foster innovation from the community. I was also worried that constraints may push some members of the community away.</p>\n<h1>Technical discussion</h1>\n<h2>The dataset</h2>\n<ul>\n<li>The dataset can be downloaded from our <a href=\"https://level5.lyft.com/dataset/\" target=\"_blank\">website</a> or on <a href=\"https://www.kaggle.com/c/3d-object-detection-for-autonomous-vehicles/data\" target=\"_blank\">Kaggle</a>.</li>\n<li>It consists of a raw camera, lidar data, and HD semantic map.</li>\n<li>180 scenes, 25s each</li>\n<li>638,000 2D and 3D annotations over 18,000 objects</li>\n<li>The dataset had nine classes with a large class imbalance.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2F856dd61314ae3420169e9db18a42a687%2FScreenshot%20from%202020-03-04%2011-39-23.png?generation=1583350814021061&amp;alt=media\" alt=\"\"></p>\n<h2>The problem</h2>\n<ul>\n<li>Metric: 3D version of the <a href=\"http://cocodataset.org/#detection-eval\" target=\"_blank\">2D detection COCO metric</a>: average mAP over 9 classes over thresholds [0.5, 0.55, 0.6, …, 0.9, 0.95]</li>\n<li>Helper code: <a href=\"https://github.com/lyft/nuscenes-devkit\" target=\"_blank\">Lyft SDK</a></li>\n<li>Duration two months: Sep 12 - Nov 12</li>\n<li>Train set <strong>40%</strong></li>\n<li>Test set: public <strong>30%</strong>, private <strong>30%</strong></li>\n</ul>\n<h2>Evaluation</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2F50e9bf683df416ec78397c72d8459d1f%2FScreenshot%20from%202020-03-04%2011-43-16.png?generation=1583351063905046&amp;alt=media\" alt=\"\"></p>\n<p>For evaluation, we used the standard method that the Kaggle platform uses to prevent overfitting on the test set: public and private test splits.</p>\n<p>The test set is split into two halves, the public test set, and the private test set. Participants did not know which samples in the test set belong to the public or private test sets.</p>\n<p>During the competition, participants were allowed to submit predictions on the whole test set twice per day. They would immediately receive their score with respect to the public test set, immortalized on the public leaderboard. This score is only a tentative measure of team performance. After the competition ends, the private leaderboard, made out of the predictions on the private part of the test set released and it is used to decide winners.</p>\n<h1>Solutions</h1>\n<h2>Summary</h2>\n<ul>\n<li>Hardware:<ul>\n<li>1st place: <strong>24</strong> GPUs</li>\n<li>2nd place: <strong>17</strong> GPUs</li>\n<li>3rd place: <strong>1</strong> GPU</li></ul></li>\n<li>All teams used the <a href=\"https://github.com/traveller59/second.pytorch\" target=\"_blank\">SECOND pytorch repo</a> in their solutions.</li>\n<li>The best submission for each team was based on the ensemble of different models or variations of the same model.</li>\n<li>Each team used different ensembling techniques.</li>\n<li>Each team had a single model that would lead them to the top 5.</li>\n<li>The strongest single model for all winning teams was based only on the lidar data.</li>\n</ul>\n<p>I asked winners to share their approach at the Kaggle discussion forum. If you have any questions, you may ask them there:</p>\n<ol>\n<li><a href=\"https://www.kaggle.com/c/3d-object-detection-for-autonomous-vehicles/discussion/122820\" target=\"_blank\">1st place solution</a></li>\n<li><a href=\"https://www.kaggle.com/c/3d-object-detection-for-autonomous-vehicles/discussion/123004\" target=\"_blank\">2nd place solution</a></li>\n<li><a href=\"https://www.kaggle.com/c/3d-object-detection-for-autonomous-vehicles/discussion/117269\" target=\"_blank\">3rd place solution</a></li>\n</ol>\n<h2>First place solution</h2>\n<h3><a href=\"https://www.kaggle.com/nywenjing\" target=\"_blank\">Wenjing Zhang</a></h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2Fb44b5a31462460a061dcc7674cf05971%2FScreenshot%20from%202020-03-04%2011-49-40.png?generation=1583351445181412&amp;alt=media\" alt=\"\"></p>\n<p>Wenjing is a Sr. AI Engineer at Ankobot, Singapore. He got his Ph.D. at Nanyang Technological University, where his research area was computer Graphics and 3D Geometry processing. He was a 2nd place winner at the <a href=\"https://sites.google.com/view/wad2019/challenge\" target=\"_blank\">CVPR 2019 WAD challenge</a>.</p>\n<h3><a href=\"https://www.kaggle.com/avsanjay\" target=\"_blank\">Sanjay Addicam</a></h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2F4cbef3f1320a442cebb4d50a054e836c%2FScreenshot%20from%202020-03-04%2011-49-43.png?generation=1583351577845147&amp;alt=media\" alt=\"\"></p>\n<p>Sanjay works as a CTO of the Visual Retail group at Intel. He is a Kaggle Master with 1 gold, 4 silvers, and 1 bronze medals.</p>\n<p>This team did not use camera images or HD map data. Only lidar was used. The team did not use pre-trained models or external data.</p>\n<p>The code was based on the <a href=\"http://openaccess.thecvf.com/content_CVPR_2019/papers/Lang_PointPillars_Fast_Encoders_for_Object_Detection_From_Point_Clouds_CVPR_2019_paper.pdf\" target=\"_blank\">PointPillars method by nuTonomy</a> implemented in <a href=\"https://github.com/traveller59/second.pytorch\" target=\"_blank\">SECOND</a>. </p>\n<p>The team trained seven different models with varying backbones and Voxel sizes.</p>\n<p>Backbones:</p>\n<ul>\n<li>PIllarFeatureNet</li>\n<li>PIllarFeatureNetRadius</li>\n<li>PIllarFeatureNetHeight</li>\n</ul>\n<p>Voxel Sizes: 0.1, 0.125, 0.2, 0.25</p>\n<p><strong>Parameters</strong></p>\n<ul>\n<li>Detection range: [-100,-100,-5,100,100,3]</li>\n<li>No direction classifier</li>\n<li>No specific post-processing, but score thresholding (0.1) and NMS</li>\n<li>Adam with one-cycle policy</li>\n<li>LR max:             1e-3</li>\n<li>Division factor:  10</li>\n<li>Weight decay:    0.01</li>\n<li>Batch size:         2</li>\n<li>Epochs:              30</li>\n<li>Train augmentations: original, flip X, flip Y, flip XY, random rotation, random scaling, and random translation</li>\n<li>Test augmentations: original, flip X, flip Y, flip XY</li>\n</ul>\n<p>All 4 TTA * 7 models = 28 predictions were ensembled.</p>\n<p>The best single model without TTA is <strong>0.177 (5th place)</strong><br>\nThe best single model with TTA is <strong>0.197 (3rd place)</strong><br>\nModel Ensemble <strong>0.220 (1st place)</strong></p>\n<p>They introduced a new method for 3D boxes fusion. Yaw angle had a 180-degree ambiguity for the IOU calculation. As a result, the direction of the predicted box is not reliable.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2Fa6324af55cfd06882aa92731cb58ed88%2FScreenshot%20from%202020-03-04%2011-56-10.png?generation=1583351839939577&amp;alt=media\" alt=\"\"></p>\n<p>In the example above, a simple angle averaging depicted on the left would not give the desired result, while the proposed method can take into account the invariance of the evaluation metric with respect to the 180-degree rotations.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2Fc9094868164720a528b70a1b6b734886%2FScreenshot%20from%202020-03-04%2011-56-13.png?generation=1583351871230849&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2Fe58cefac22bff9ec086c9cf17ee9ee32%2FScreenshot%20from%202020-03-04%2012-54-00.png?generation=1583355290969346&amp;alt=media\" alt=\"\"></p>\n<h3>Approaches that did not work or were not attempted</h3>\n<ul>\n<li>Increasing the maximum number of voxels worked until some point, but decreased after this.</li>\n<li>The team did not try other methods like PointRCNN or Frustrum PointNet</li>\n</ul>\n<h2>Second place solution</h2>\n<h3><a href=\"https://www.kaggle.com/kylelee\" target=\"_blank\">Kyle Lee</a></h3>\n<p>Kyle is a Senior AI Scientist at the Dishcraft Robotics. He got his B.SC. in Electrical and Computer Engineering at Cornell University. He has a working knowledge of ML/DL and perception concepts for object manipulation. Kyle is Kaggle Grandmaster.</p>\n<h3>General solution flow</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2F2aecec84e5e00c5399d3eeae21d53e6d%2FScreenshot%20from%202020-03-04%2012-57-10.png?generation=1583355476697920&amp;alt=media\" alt=\"\"></p>\n<p>The solution is an ensemble of two approaches:</p>\n<ol>\n<li>Lidar based using VoxelNet/PointPillars based on SECOND, with modifications on data loading, point cloud range, voxel sizes, and maximum voxel size parameters with test-time</li>\n<li>A frustum-based approach based on <a href=\"https://github.com/zhixinwang/frustum-convnet\" target=\"_blank\">Frustum ConvNet</a>, where 2D object detection boxes were inferred from various re-trained <a href=\"https://github.com/facebookresearch/detectron2\" target=\"_blank\">Detectron2</a> and <a href=\"https://github.com/tensorflow/models/tree/master/research/object_detection\" target=\"_blank\">Tensorflow Object Detection</a> API object detection frameworks and the aligned point cloud sampled as a sequence of frustums into a fully convolutional network (FCN).</li>\n</ol>\n<p>The first approach gave a public score of 0.191 or a private score of 0.188, which on its own would have been sufficient for ​2nd place​ on the leaderboard.</p>\n<p>The second approach gave a public score of 0.171 or a private score of 0.169, which on its own would have been sufficient for​ 5th place​ on the leaderboard.</p>\n<p>Soft-NMS using the Gaussian function was used to ensemble all the above<br>\npredictions together. The combination of the above boosted the LIDAR only approach<br>\nby approximately +0.014 on both public/private leaderboards to 0.205 (public) / 0.202</p>\n<p>Train time: 18 days for SECOND (simplified model: 3 days); 5 days for Frustum-ConvNet</p>\n<p>These two approaches worked well in an ensemble for the two main reasons:</p>\n<ol>\n<li>Both gave strong models. (You cannot ensemble weak models and expect to get something that is really good. Garbage in, garbage out)</li>\n<li>The approaches have different strengths and weaknesses. </li>\n</ol>\n<h3>Pros of 2D images for F-ConvNet:</h3>\n<ul>\n<li>Denser information for smaller classes</li>\n<li>Richer context/features can help distinguish confused classes</li>\n<li>Leverage mature 2D object detectors</li>\n</ul>\n<h3>Cons of 2D images:</h3>\n<ul>\n<li>Occluded/truncated objects are unrecognizable in 2D images</li>\n<li>Limited range</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2Ffe1187474323e715085869cc50128ee8%2FScreenshot%20from%202020-03-04%2013-00-57.png?generation=1583355710397955&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Point Cloud range  [-100, -100, -5, 100, 100, 3]</li>\n<li>Voxel size: different values in the range 0.1x0.1 to 0.25x0.25</li>\n<li>The number of voxels to be as large as possible.</li>\n<li>Optimizer: Adam</li>\n<li>Step-wise learning rate decay. At steps 100k, 200k learning rate was divided by 10.</li>\n<li>Batch size 1, since the purpose was to maximize the number of voxels.</li>\n<li>Train augmentations: original, flip X, flip Y, flip XY</li>\n<li>Test time augmentations (Lidar model only): flip X, flip Y, flip XY</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2Fc82c4cbe4a0a2ed52808eabeafced527%2FScreenshot%20from%202020-03-04%2013-03-12.png?generation=1583355836007959&amp;alt=media\" alt=\"\"></p>\n<p><strong>Best single model (SECOND/Lidar)</strong>:</p>\n<ul>\n<li>Voxel size: 0.2x0.2x8</li>\n<li>Private score: <ul>\n<li>Without TTA: 0.170 <strong>(top 5)</strong></li>\n<li>With TTA: 0.185 <strong>(top 3)</strong></li></ul></li>\n<li>Train time (~3 days)</li>\n<li>Test time (~3 hours) over 27,468 samples<ul>\n<li>2.54 samples / s</li>\n<li>0.39s per sample</li></ul></li>\n</ul>\n<h3>Approaches that did not work</h3>\n<p>Oversampling rare classes<br>\n• On-road / off-road filtering<br>\n• Adjacent timestamp object tracking/filtering<br>\n• Rotated soft-NMS with height (3D)<br>\n• Animals<br>\n• PointRCNN</p>\n<h2>3rd place solution</h2>\n<h3><a href=\"https://www.kaggle.com/yukke42\" target=\"_blank\">Yusuke Muramatsu</a></h3>\n<h3>Data:</h3>\n<ul>\n<li>Pointcloud only. Camera and HD map was not used.</li>\n<li>External data was not used.</li>\n<li>Animals and emergency vehicle classes were not used.</li>\n<li>Objects that had less than 5 points were ignored.</li>\n</ul>\n<h3>Model:</h3>\n<ul>\n<li>The combination of VoxelNet and PointPillars</li>\n<li>The implementation is based on <a href=\"https://github.com/traveller59/second.pytorch\" target=\"_blank\">SECOND</a></li>\n</ul>\n<h3>General flow</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2F9d4ea57b7b5a84d768ac5cc274a73f2b%2FScreenshot%20from%202020-03-04%2013-07-40.png?generation=1583356106582763&amp;alt=media\" alt=\"\"></p>\n<h3>Training details</h3>\n<ul>\n<li>50 epochs</li>\n<li>Batch size = 4</li>\n<li>Scheduler CosineAnnealingLR</li>\n<li>Train augmentation: translation, scaling, rotation around the z-axis, mixup augmentation (pasting - objects to the point cloud from a library of objects created in ).</li>\n<li>No Test Time Augmentation</li>\n<li>Training time: about 2 to 4 days<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2F0f3cc1a59663d8a4cebbaf7497641e91%2FScreenshot%20from%202020-03-04%2013-09-15.png?generation=1583356194029533&amp;alt=media\" alt=\"\"></li>\n</ul>\n<p>Detection range:</p>\n<ul>\n<li>Car, other vehicle, truck, bus =&gt; 100m x 75m</li>\n<li>Pedestrian, bicycle, motorcycle =&gt; 100m x 50m<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2F7cd509331e168a34e702c5faca898ca4%2FScreenshot%20from%202020-03-04%2013-10-35.png?generation=1583356294708470&amp;alt=media\" alt=\"\"></li>\n</ul>\n<p>The detection range along the z-axis is smaller by using the custom coordinate system.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2Fecf0a2f4931c4afd073ff4e1ae384753%2FScreenshot%20from%202020-03-04%2013-10-38.png?generation=1583356317234841&amp;alt=media\" alt=\"\"></p>\n<h3>Ideas that did not work:</h3>\n<ul>\n<li>Weighted-boxed-fusion. The score was lower than Soft NMS.</li>\n<li>HD Map for object filtering.</li>\n<li>Point cloud coloring using raw images or feature maps extracted from 2D detection.</li>\n</ul>\n<h1>Conclusion</h1>\n<p>My biggest fear entering the competition was that we would discover a critical blunder with the dataset after we had released it. Now that the competition is over, I am happy to say that everything went really well. I would say this is a success.<br>\nIn the end, we had:</p>\n<ul>\n<li>660 competitors from all over the world that formed 547 teams</li>\n<li>5707 submissions</li>\n<li>A dataset that was well prepared:<ul>\n<li>No data leaks we found.</li>\n<li>The difference between Public and Private leaderboards is minimal.</li></ul></li>\n<li>An active community that created a number of pull requests to the Lyft SDK with improvements and bug fixes.</li>\n<li>Plenty of positive feedback from emails and in-person thanking us for the dataset and competition.</li>\n</ul>\n<h1>Improvements and opportunities for next time</h1>\n<ol>\n<li>Longer duration: Many people gave me feedback that it takes some time to get used to the dataset and two months that we had for the challenge is not enough. I believe 3-4 months would work better.</li>\n<li>Build a new metric to encourage data fusion: Many teams were able to work exclusively off of lidar data. Next time we should create a metric that would be improved by using more radar, lidar, and camera images while also applying constraints to inference time. </li>\n</ol>\n<p>I would like to give special thanks to <a href=\"https://www.linkedin.com/in/christina-robertson/\" target=\"_blank\">Christy Robertson</a> who was driving the marketing part of the project and without whom it would not have any chances for success and <a href=\"https://www.linkedin.com/in/erikgaas/\" target=\"_blank\">Erik Gaasedelen</a> who helped me to prepare this blog post. And last but not least, I would like to thank all the enthusiastic Kagglers who competed in the challenge. We would not have been successful without your diligence and helpfulness within the community.</p>",
  "messages": [
    {
      "id": "763788",
      "postDate": "03/04/2020 21:16:26",
      "content": "<h1>Introduction</h1>\n<p>Hello! My name is <a href=\"https://medium.com/lyftlevel5/employee-spotlight-vladimir-iglovikov-11b71ed02bc\" target=\"_blank\">Vladimir Iglovikov</a>, and I am a Computer Vision Engineer at Lyft Level 5. Here, I work on a diverse set of projects relating to building, testing, deploying, and refining deep learning models for mapping and perception related tasks. </p>\n<p>I am <a href=\"http://ternaus.blog/interview/2018/08/30/ama.html\" target=\"_blank\">Kaggle Grandmaster</a>, which means that I have some experience of <a href=\"http://ternaus.blog/career/2020/01/08/How-I-Found-my-current-job.html\" target=\"_blank\">participating</a> in machine learning competitions. As a result, I was very pleased to join a team of colleagues at Level 5 who wanted to create a new machine learning competition for Lyft.</p>\n<p>After carefully putting together a dataset and an objective, we finally put together the Lyft 3D Object Detection for Autonomous Vehicles Challenge. Over the course of this challenge, over 500 teams plugged away at achieving a top leaderboard spot. Now that the challenge has finished and the dust has settled, we would like to share with you the top solutions.</p>\n<p>I would also like to discuss what we have learned throughout the process as organizers of the competition. There are many <a href=\"https://docs.google.com/a/chalearn.org/viewer?a=v&amp;pid=sites&amp;srcid=Y2hhbGVhcm4ub3JnfHdvcmtzaG9wfGd4OjQwYzBhY2Q1N2YxNGE3YWI\" target=\"_blank\">difficulties</a> to overcome when developing a machine learning challenge. From data leaks to issues with leaderboard scoring, it is common to make critical mistakes. I have provided much feedback to organizers when I was a participant in different machine learning competitions. Now it was time to get such feedback myself.</p>\n<p>This blog post will have two parts. In the first introductory part, I will tell you how this project was started and the challenges we ran into. In the second technical part, I will describe the solutions from the top three teams in a level of depth and rigor that I would have wanted to learn from if I had been a participant.</p>\n<h1>Why does Lyft need a Machine Learning Competition?</h1>\n<p>When I go to academic conferences I see a lot of papers that have the keywords “autonomous” and “self-driving”. For sure, all these academic works are related to the autonomous industry in some way, but for many of them, the relationship is rather weak.</p>\n<p>One of the reasons for this gap may be that the research community does not have access to the same data as the industry. There are many academic datasets that rely solely on camera images. Meanwhile, companies, including Lyft, use a combination of camera, lidar, and radar sensors, all of which are crucial for the construction of high definition maps or the functioning of autonomy algorithms. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2Fd280f4af3e91991ad641606967cd0bbe%2FScreenshot%20from%202020-03-04%2011-28-24.png?generation=1583350172278513&amp;alt=media\" alt=\"\"></p>\n<p>Prior to 2019, the <a href=\"http://www.cvlibs.net/datasets/kitti/\" target=\"_blank\">KITTI</a> dataset was the most useful dataset within the research community, that provided both camera and lidar data. KITTI was far ahead of its time and has been a staple in the self-driving industry even today, but it has its own limitations. </p>\n<p>In 2019 the field shifted rapidly. Lyft, along with some other self-driving companies, released their own, much larger datasets. This was a great move, but it is not enough to just release the dataset. It is also necessary to promote it and engage with the community. By going these extra steps, a dataset is better able to foster collaboration within the machine learning community. This can lead to new GitHub repositories, academic papers, or blog posts. Organizing a competition is one of the most popular ways of realizing these benefits.</p>\n<p>When Aptiv released their <a href=\"https://www.nuscenes.org/overview\" target=\"_blank\">dataset</a> in March 2019, they launched a competition as part of the WAD workshop at CVPR 2019. In addition to this, Aptiv released an <a href=\"https://github.com/nutonomy/nuscenes-devkit\" target=\"_blank\">SDK</a> with useful helper functions and visualization code. For <a href=\"https://www.nuscenes.org/object-detection/?externalData=all&amp;mapData=all&amp;modalities=Any\" target=\"_blank\">evaluation</a>, organizers used a weighted sum of the metrics responsible for detection, orientation, velocity, and attribute estimation.</p>\n<p>Lyft <a href=\"https://medium.com/lyftlevel5/unlocking-access-to-self-driving-research-the-lyft-level-5-dataset-and-competition-d487c27b1b6c\" target=\"_blank\">released</a> the dataset slightly later. Since the data was structurally similar to those in the Aptiv dataset, we decided to use the same format.</p>\n<p>Preparing a good dataset is a big challenge. 90% of the success of the competition is defined by the quality of the data. I would like to thank <strong>Robert Kesten, Muhammad Usman, John Houston, Tirth Pandya, K. Nadhamuni, Ana Ferreira, Michael Yuan, Ben Low, Ashesh Jain, Peter Ondrushka, Sami Omari, Shaswat Shah, Amruta Kulkarni, Alex Kazakova, Charlotte Tao, Lucas Platinsky, William Jiang, and Vinay Shet</strong> for preparing and releasing the dataset.</p>\n<h1>How did we prepare for the competition?</h1>\n<p>From the beginning, there was a plan to organize a challenge as a part of the competition track of NeurIPS 2019. We needed to decide who would host the competition for us. There were two options that we considered:</p>\n<ol>\n<li>Host ourselves: release the dataset, release the code for the metric, write a blog post about the challenge, provide prize money, and ask participants to send their predictions via email.</li>\n<li>Host through Kaggle. (I am not affiliated with Kaggle, but at the moment is the most established platform for machine competitions in the world). Kaggle would host the platform, the dataset, forum, real-time leaderboard, and notebooks in which participants may perform exploratory data analysis and even train models. There are other choices besides Kaggle, but they are less popular.</li>\n</ol>\n<p>Both strategies have their own advantages, but the first one is more common. It is commonly used when organizers have a small budget or Kaggle cannot support some sort of functionality, like submitting a Docker container instead of a CSV file. Most of the challenges in competition tracks in conferences like NeurIPS, ICCV, ECCV, CVPR, ICML, KDD, MICCAI, etc follow the first path.</p>\n<p>We decided to go with Kaggle. Although we knew we could host a quality competition ourselves, we wanted to make community engagement our highest priority.</p>\n<p>I believe that the success of the competition from the standpoint of the organizers is measured not by the titles and affiliations of the participants, nor the strength of the solutions, but rather by the number of active participants. If you only get 3-20 teams, you did not do a very good job. If the number is in the range 100-1000 you are doing well. Over 1000 teams and you’re doing extremely well! The background of the participants should not matter; whether they are from academia, industry, or are self-taught, everyone should be welcome.</p>\n<p>Such a philosophy helps to further democratize the field, elevating the skill of the community as a whole while also encouraging the development of diverse new methods and applications.  </p>\n<p>In my experience, the number of participants in Kaggle competitions is quite large, on the order of hundreds or thousands. The knowledge accumulated in Kaggle kernels and discussion threads became a great reference point for people investigating similar problems in the future.</p>\n<p>Following this way of thinking, we reached out to Kaggle and our partnership started.</p>\n<p>There are many different tasks that one can try to solve on the Lyft dataset: 3D object detection, lidar point segmentation, 3D tracking, orientation, and velocity prediction, etc. All of them are interesting tasks.</p>\n<p>This competition was the first ML challenge with this data type hosted on Kaggle. Its purpose was exploratory by nature, providing a testing ground for the ML community to become accustomed to this format of data. It is nice to get winning solutions that are closely related to the tasks that you face in production, but for this challenge it was secondary. We decided to go with the simplest possible task that one can have on this type of data: 3D object detection with the <a href=\"https://github.com/lyft/nuscenes-devkit/blob/master/lyft_dataset_sdk/eval/detection/mAP_evaluation.py\" target=\"_blank\">3D version of the COCO mAP</a> as a metric. This turned out to be very convenient as it helped to bridge the gap between more familiar bounding box problems and this new challenge. It was also nice because it could be easily demonstrated in Python and reimplemented in the data scientist’s language of choice. We also had a limited timeline, so avoiding the potential for bugs was a great plus.</p>\n<h1>Sample model</h1>\n<p>We expected that most of the people joining the competition would not have hands-on experience with this type of data. Therefore we wanted to cater specifically to participants who were motivated by learning and skill growth. To me, machine Learning is an applied discipline, so it makes life much easier when there is <a href=\"https://towardsdatascience.com/ask-me-anything-session-with-a-kaggle-grandmaster-vladimir-i-iglovikov-942ad6a06acd\" target=\"_blank\">an example to learn from</a>. Taking someone else's end-to-end code and playing with it was always a good way to be less overwhelmed with the volume of new things you need to learn and implement correctly.</p>\n<p><a href=\"https://www.linkedin.com/in/guido-zuidhof-377b6947/\" target=\"_blank\">Guido Zuidhof</a>, a <a href=\"https://www.kaggle.com/gzuidhof\" target=\"_blank\">Kaggle Master</a> and colleague in the AV Research team at Lyft Level 5 prepared a <a href=\"https://www.kaggle.com/gzuidhof/reference-model\" target=\"_blank\">Kaggle kernel with a sample model</a> that many of the participants used in their solutions.</p>\n<h1>Model limitations</h1>\n<p>In this challenge, I decided to maximize the flexibility given to participants with respect to model development:</p>\n<ol>\n<li>No limitations on the hardware.</li>\n<li>No limitations on the model size and inference time.</li>\n<li>Any external data or pre-trained models as long as they are listed in the <a href=\"https://www.kaggle.com/c/3d-object-detection-for-autonomous-vehicles/discussion/109361\" target=\"_blank\">special thread at the discussion forum</a>.</li>\n</ol>\n<p>It is normally a good idea to enforce constraints on model development such that it is more compatible with a production environment, but I decided against this.</p>\n<p>I do not know any examples of when a winning ML competition solution made it all the way to production. I did not expect that this challenge would be any different. But this was not our goal. At the end of the day, the challenges that we face are different from the competition task. We were looking to foster innovation from the community. I was also worried that constraints may push some members of the community away.</p>\n<h1>Technical discussion</h1>\n<h2>The dataset</h2>\n<ul>\n<li>The dataset can be downloaded from our <a href=\"https://level5.lyft.com/dataset/\" target=\"_blank\">website</a> or on <a href=\"https://www.kaggle.com/c/3d-object-detection-for-autonomous-vehicles/data\" target=\"_blank\">Kaggle</a>.</li>\n<li>It consists of a raw camera, lidar data, and HD semantic map.</li>\n<li>180 scenes, 25s each</li>\n<li>638,000 2D and 3D annotations over 18,000 objects</li>\n<li>The dataset had nine classes with a large class imbalance.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2F856dd61314ae3420169e9db18a42a687%2FScreenshot%20from%202020-03-04%2011-39-23.png?generation=1583350814021061&amp;alt=media\" alt=\"\"></p>\n<h2>The problem</h2>\n<ul>\n<li>Metric: 3D version of the <a href=\"http://cocodataset.org/#detection-eval\" target=\"_blank\">2D detection COCO metric</a>: average mAP over 9 classes over thresholds [0.5, 0.55, 0.6, …, 0.9, 0.95]</li>\n<li>Helper code: <a href=\"https://github.com/lyft/nuscenes-devkit\" target=\"_blank\">Lyft SDK</a></li>\n<li>Duration two months: Sep 12 - Nov 12</li>\n<li>Train set <strong>40%</strong></li>\n<li>Test set: public <strong>30%</strong>, private <strong>30%</strong></li>\n</ul>\n<h2>Evaluation</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2F50e9bf683df416ec78397c72d8459d1f%2FScreenshot%20from%202020-03-04%2011-43-16.png?generation=1583351063905046&amp;alt=media\" alt=\"\"></p>\n<p>For evaluation, we used the standard method that the Kaggle platform uses to prevent overfitting on the test set: public and private test splits.</p>\n<p>The test set is split into two halves, the public test set, and the private test set. Participants did not know which samples in the test set belong to the public or private test sets.</p>\n<p>During the competition, participants were allowed to submit predictions on the whole test set twice per day. They would immediately receive their score with respect to the public test set, immortalized on the public leaderboard. This score is only a tentative measure of team performance. After the competition ends, the private leaderboard, made out of the predictions on the private part of the test set released and it is used to decide winners.</p>\n<h1>Solutions</h1>\n<h2>Summary</h2>\n<ul>\n<li>Hardware:<ul>\n<li>1st place: <strong>24</strong> GPUs</li>\n<li>2nd place: <strong>17</strong> GPUs</li>\n<li>3rd place: <strong>1</strong> GPU</li></ul></li>\n<li>All teams used the <a href=\"https://github.com/traveller59/second.pytorch\" target=\"_blank\">SECOND pytorch repo</a> in their solutions.</li>\n<li>The best submission for each team was based on the ensemble of different models or variations of the same model.</li>\n<li>Each team used different ensembling techniques.</li>\n<li>Each team had a single model that would lead them to the top 5.</li>\n<li>The strongest single model for all winning teams was based only on the lidar data.</li>\n</ul>\n<p>I asked winners to share their approach at the Kaggle discussion forum. If you have any questions, you may ask them there:</p>\n<ol>\n<li><a href=\"https://www.kaggle.com/c/3d-object-detection-for-autonomous-vehicles/discussion/122820\" target=\"_blank\">1st place solution</a></li>\n<li><a href=\"https://www.kaggle.com/c/3d-object-detection-for-autonomous-vehicles/discussion/123004\" target=\"_blank\">2nd place solution</a></li>\n<li><a href=\"https://www.kaggle.com/c/3d-object-detection-for-autonomous-vehicles/discussion/117269\" target=\"_blank\">3rd place solution</a></li>\n</ol>\n<h2>First place solution</h2>\n<h3><a href=\"https://www.kaggle.com/nywenjing\" target=\"_blank\">Wenjing Zhang</a></h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2Fb44b5a31462460a061dcc7674cf05971%2FScreenshot%20from%202020-03-04%2011-49-40.png?generation=1583351445181412&amp;alt=media\" alt=\"\"></p>\n<p>Wenjing is a Sr. AI Engineer at Ankobot, Singapore. He got his Ph.D. at Nanyang Technological University, where his research area was computer Graphics and 3D Geometry processing. He was a 2nd place winner at the <a href=\"https://sites.google.com/view/wad2019/challenge\" target=\"_blank\">CVPR 2019 WAD challenge</a>.</p>\n<h3><a href=\"https://www.kaggle.com/avsanjay\" target=\"_blank\">Sanjay Addicam</a></h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2F4cbef3f1320a442cebb4d50a054e836c%2FScreenshot%20from%202020-03-04%2011-49-43.png?generation=1583351577845147&amp;alt=media\" alt=\"\"></p>\n<p>Sanjay works as a CTO of the Visual Retail group at Intel. He is a Kaggle Master with 1 gold, 4 silvers, and 1 bronze medals.</p>\n<p>This team did not use camera images or HD map data. Only lidar was used. The team did not use pre-trained models or external data.</p>\n<p>The code was based on the <a href=\"http://openaccess.thecvf.com/content_CVPR_2019/papers/Lang_PointPillars_Fast_Encoders_for_Object_Detection_From_Point_Clouds_CVPR_2019_paper.pdf\" target=\"_blank\">PointPillars method by nuTonomy</a> implemented in <a href=\"https://github.com/traveller59/second.pytorch\" target=\"_blank\">SECOND</a>. </p>\n<p>The team trained seven different models with varying backbones and Voxel sizes.</p>\n<p>Backbones:</p>\n<ul>\n<li>PIllarFeatureNet</li>\n<li>PIllarFeatureNetRadius</li>\n<li>PIllarFeatureNetHeight</li>\n</ul>\n<p>Voxel Sizes: 0.1, 0.125, 0.2, 0.25</p>\n<p><strong>Parameters</strong></p>\n<ul>\n<li>Detection range: [-100,-100,-5,100,100,3]</li>\n<li>No direction classifier</li>\n<li>No specific post-processing, but score thresholding (0.1) and NMS</li>\n<li>Adam with one-cycle policy</li>\n<li>LR max:             1e-3</li>\n<li>Division factor:  10</li>\n<li>Weight decay:    0.01</li>\n<li>Batch size:         2</li>\n<li>Epochs:              30</li>\n<li>Train augmentations: original, flip X, flip Y, flip XY, random rotation, random scaling, and random translation</li>\n<li>Test augmentations: original, flip X, flip Y, flip XY</li>\n</ul>\n<p>All 4 TTA * 7 models = 28 predictions were ensembled.</p>\n<p>The best single model without TTA is <strong>0.177 (5th place)</strong><br>\nThe best single model with TTA is <strong>0.197 (3rd place)</strong><br>\nModel Ensemble <strong>0.220 (1st place)</strong></p>\n<p>They introduced a new method for 3D boxes fusion. Yaw angle had a 180-degree ambiguity for the IOU calculation. As a result, the direction of the predicted box is not reliable.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2Fa6324af55cfd06882aa92731cb58ed88%2FScreenshot%20from%202020-03-04%2011-56-10.png?generation=1583351839939577&amp;alt=media\" alt=\"\"></p>\n<p>In the example above, a simple angle averaging depicted on the left would not give the desired result, while the proposed method can take into account the invariance of the evaluation metric with respect to the 180-degree rotations.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2Fc9094868164720a528b70a1b6b734886%2FScreenshot%20from%202020-03-04%2011-56-13.png?generation=1583351871230849&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2Fe58cefac22bff9ec086c9cf17ee9ee32%2FScreenshot%20from%202020-03-04%2012-54-00.png?generation=1583355290969346&amp;alt=media\" alt=\"\"></p>\n<h3>Approaches that did not work or were not attempted</h3>\n<ul>\n<li>Increasing the maximum number of voxels worked until some point, but decreased after this.</li>\n<li>The team did not try other methods like PointRCNN or Frustrum PointNet</li>\n</ul>\n<h2>Second place solution</h2>\n<h3><a href=\"https://www.kaggle.com/kylelee\" target=\"_blank\">Kyle Lee</a></h3>\n<p>Kyle is a Senior AI Scientist at the Dishcraft Robotics. He got his B.SC. in Electrical and Computer Engineering at Cornell University. He has a working knowledge of ML/DL and perception concepts for object manipulation. Kyle is Kaggle Grandmaster.</p>\n<h3>General solution flow</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2F2aecec84e5e00c5399d3eeae21d53e6d%2FScreenshot%20from%202020-03-04%2012-57-10.png?generation=1583355476697920&amp;alt=media\" alt=\"\"></p>\n<p>The solution is an ensemble of two approaches:</p>\n<ol>\n<li>Lidar based using VoxelNet/PointPillars based on SECOND, with modifications on data loading, point cloud range, voxel sizes, and maximum voxel size parameters with test-time</li>\n<li>A frustum-based approach based on <a href=\"https://github.com/zhixinwang/frustum-convnet\" target=\"_blank\">Frustum ConvNet</a>, where 2D object detection boxes were inferred from various re-trained <a href=\"https://github.com/facebookresearch/detectron2\" target=\"_blank\">Detectron2</a> and <a href=\"https://github.com/tensorflow/models/tree/master/research/object_detection\" target=\"_blank\">Tensorflow Object Detection</a> API object detection frameworks and the aligned point cloud sampled as a sequence of frustums into a fully convolutional network (FCN).</li>\n</ol>\n<p>The first approach gave a public score of 0.191 or a private score of 0.188, which on its own would have been sufficient for ​2nd place​ on the leaderboard.</p>\n<p>The second approach gave a public score of 0.171 or a private score of 0.169, which on its own would have been sufficient for​ 5th place​ on the leaderboard.</p>\n<p>Soft-NMS using the Gaussian function was used to ensemble all the above<br>\npredictions together. The combination of the above boosted the LIDAR only approach<br>\nby approximately +0.014 on both public/private leaderboards to 0.205 (public) / 0.202</p>\n<p>Train time: 18 days for SECOND (simplified model: 3 days); 5 days for Frustum-ConvNet</p>\n<p>These two approaches worked well in an ensemble for the two main reasons:</p>\n<ol>\n<li>Both gave strong models. (You cannot ensemble weak models and expect to get something that is really good. Garbage in, garbage out)</li>\n<li>The approaches have different strengths and weaknesses. </li>\n</ol>\n<h3>Pros of 2D images for F-ConvNet:</h3>\n<ul>\n<li>Denser information for smaller classes</li>\n<li>Richer context/features can help distinguish confused classes</li>\n<li>Leverage mature 2D object detectors</li>\n</ul>\n<h3>Cons of 2D images:</h3>\n<ul>\n<li>Occluded/truncated objects are unrecognizable in 2D images</li>\n<li>Limited range</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2Ffe1187474323e715085869cc50128ee8%2FScreenshot%20from%202020-03-04%2013-00-57.png?generation=1583355710397955&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Point Cloud range  [-100, -100, -5, 100, 100, 3]</li>\n<li>Voxel size: different values in the range 0.1x0.1 to 0.25x0.25</li>\n<li>The number of voxels to be as large as possible.</li>\n<li>Optimizer: Adam</li>\n<li>Step-wise learning rate decay. At steps 100k, 200k learning rate was divided by 10.</li>\n<li>Batch size 1, since the purpose was to maximize the number of voxels.</li>\n<li>Train augmentations: original, flip X, flip Y, flip XY</li>\n<li>Test time augmentations (Lidar model only): flip X, flip Y, flip XY</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2Fc82c4cbe4a0a2ed52808eabeafced527%2FScreenshot%20from%202020-03-04%2013-03-12.png?generation=1583355836007959&amp;alt=media\" alt=\"\"></p>\n<p><strong>Best single model (SECOND/Lidar)</strong>:</p>\n<ul>\n<li>Voxel size: 0.2x0.2x8</li>\n<li>Private score: <ul>\n<li>Without TTA: 0.170 <strong>(top 5)</strong></li>\n<li>With TTA: 0.185 <strong>(top 3)</strong></li></ul></li>\n<li>Train time (~3 days)</li>\n<li>Test time (~3 hours) over 27,468 samples<ul>\n<li>2.54 samples / s</li>\n<li>0.39s per sample</li></ul></li>\n</ul>\n<h3>Approaches that did not work</h3>\n<p>Oversampling rare classes<br>\n• On-road / off-road filtering<br>\n• Adjacent timestamp object tracking/filtering<br>\n• Rotated soft-NMS with height (3D)<br>\n• Animals<br>\n• PointRCNN</p>\n<h2>3rd place solution</h2>\n<h3><a href=\"https://www.kaggle.com/yukke42\" target=\"_blank\">Yusuke Muramatsu</a></h3>\n<h3>Data:</h3>\n<ul>\n<li>Pointcloud only. Camera and HD map was not used.</li>\n<li>External data was not used.</li>\n<li>Animals and emergency vehicle classes were not used.</li>\n<li>Objects that had less than 5 points were ignored.</li>\n</ul>\n<h3>Model:</h3>\n<ul>\n<li>The combination of VoxelNet and PointPillars</li>\n<li>The implementation is based on <a href=\"https://github.com/traveller59/second.pytorch\" target=\"_blank\">SECOND</a></li>\n</ul>\n<h3>General flow</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2F9d4ea57b7b5a84d768ac5cc274a73f2b%2FScreenshot%20from%202020-03-04%2013-07-40.png?generation=1583356106582763&amp;alt=media\" alt=\"\"></p>\n<h3>Training details</h3>\n<ul>\n<li>50 epochs</li>\n<li>Batch size = 4</li>\n<li>Scheduler CosineAnnealingLR</li>\n<li>Train augmentation: translation, scaling, rotation around the z-axis, mixup augmentation (pasting - objects to the point cloud from a library of objects created in ).</li>\n<li>No Test Time Augmentation</li>\n<li>Training time: about 2 to 4 days<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2F0f3cc1a59663d8a4cebbaf7497641e91%2FScreenshot%20from%202020-03-04%2013-09-15.png?generation=1583356194029533&amp;alt=media\" alt=\"\"></li>\n</ul>\n<p>Detection range:</p>\n<ul>\n<li>Car, other vehicle, truck, bus =&gt; 100m x 75m</li>\n<li>Pedestrian, bicycle, motorcycle =&gt; 100m x 50m<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2F7cd509331e168a34e702c5faca898ca4%2FScreenshot%20from%202020-03-04%2013-10-35.png?generation=1583356294708470&amp;alt=media\" alt=\"\"></li>\n</ul>\n<p>The detection range along the z-axis is smaller by using the custom coordinate system.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2Fecf0a2f4931c4afd073ff4e1ae384753%2FScreenshot%20from%202020-03-04%2013-10-38.png?generation=1583356317234841&amp;alt=media\" alt=\"\"></p>\n<h3>Ideas that did not work:</h3>\n<ul>\n<li>Weighted-boxed-fusion. The score was lower than Soft NMS.</li>\n<li>HD Map for object filtering.</li>\n<li>Point cloud coloring using raw images or feature maps extracted from 2D detection.</li>\n</ul>\n<h1>Conclusion</h1>\n<p>My biggest fear entering the competition was that we would discover a critical blunder with the dataset after we had released it. Now that the competition is over, I am happy to say that everything went really well. I would say this is a success.<br>\nIn the end, we had:</p>\n<ul>\n<li>660 competitors from all over the world that formed 547 teams</li>\n<li>5707 submissions</li>\n<li>A dataset that was well prepared:<ul>\n<li>No data leaks we found.</li>\n<li>The difference between Public and Private leaderboards is minimal.</li></ul></li>\n<li>An active community that created a number of pull requests to the Lyft SDK with improvements and bug fixes.</li>\n<li>Plenty of positive feedback from emails and in-person thanking us for the dataset and competition.</li>\n</ul>\n<h1>Improvements and opportunities for next time</h1>\n<ol>\n<li>Longer duration: Many people gave me feedback that it takes some time to get used to the dataset and two months that we had for the challenge is not enough. I believe 3-4 months would work better.</li>\n<li>Build a new metric to encourage data fusion: Many teams were able to work exclusively off of lidar data. Next time we should create a metric that would be improved by using more radar, lidar, and camera images while also applying constraints to inference time. </li>\n</ol>\n<p>I would like to give special thanks to <a href=\"https://www.linkedin.com/in/christina-robertson/\" target=\"_blank\">Christy Robertson</a> who was driving the marketing part of the project and without whom it would not have any chances for success and <a href=\"https://www.linkedin.com/in/erikgaas/\" target=\"_blank\">Erik Gaasedelen</a> who helped me to prepare this blog post. And last but not least, I would like to thank all the enthusiastic Kagglers who competed in the challenge. We would not have been successful without your diligence and helpfulness within the community.</p>",
      "rawMarkdown": "# Introduction\n\nHello! My name is [Vladimir Iglovikov](https://medium.com/lyftlevel5/employee-spotlight-vladimir-iglovikov-11b71ed02bc), and I am a Computer Vision Engineer at Lyft Level 5. Here, I work on a diverse set of projects relating to building, testing, deploying, and refining deep learning models for mapping and perception related tasks. \n\nI am [Kaggle Grandmaster](http://ternaus.blog/interview/2018/08/30/ama.html), which means that I have some experience of [participating](http://ternaus.blog/career/2020/01/08/How-I-Found-my-current-job.html) in machine learning competitions. As a result, I was very pleased to join a team of colleagues at Level 5 who wanted to create a new machine learning competition for Lyft.\n\nAfter carefully putting together a dataset and an objective, we finally put together the Lyft 3D Object Detection for Autonomous Vehicles Challenge. Over the course of this challenge, over 500 teams plugged away at achieving a top leaderboard spot. Now that the challenge has finished and the dust has settled, we would like to share with you the top solutions.\n\nI would also like to discuss what we have learned throughout the process as organizers of the competition. There are many [difficulties](https://docs.google.com/a/chalearn.org/viewer?a=v&amp;pid=sites&amp;srcid=Y2hhbGVhcm4ub3JnfHdvcmtzaG9wfGd4OjQwYzBhY2Q1N2YxNGE3YWI) to overcome when developing a machine learning challenge. From data leaks to issues with leaderboard scoring, it is common to make critical mistakes. I have provided much feedback to organizers when I was a participant in different machine learning competitions. Now it was time to get such feedback myself.\n\nThis blog post will have two parts. In the first introductory part, I will tell you how this project was started and the challenges we ran into. In the second technical part, I will describe the solutions from the top three teams in a level of depth and rigor that I would have wanted to learn from if I had been a participant.\n\n# Why does Lyft need a Machine Learning Competition?\n\nWhen I go to academic conferences I see a lot of papers that have the keywords “autonomous” and “self-driving”. For sure, all these academic works are related to the autonomous industry in some way, but for many of them, the relationship is rather weak.\n\nOne of the reasons for this gap may be that the research community does not have access to the same data as the industry. There are many academic datasets that rely solely on camera images. Meanwhile, companies, including Lyft, use a combination of camera, lidar, and radar sensors, all of which are crucial for the construction of high definition maps or the functioning of autonomy algorithms. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2Fd280f4af3e91991ad641606967cd0bbe%2FScreenshot%20from%202020-03-04%2011-28-24.png?generation=1583350172278513&amp;alt=media)\n\nPrior to 2019, the [KITTI](http://www.cvlibs.net/datasets/kitti/) dataset was the most useful dataset within the research community, that provided both camera and lidar data. KITTI was far ahead of its time and has been a staple in the self-driving industry even today, but it has its own limitations. \n\nIn 2019 the field shifted rapidly. Lyft, along with some other self-driving companies, released their own, much larger datasets. This was a great move, but it is not enough to just release the dataset. It is also necessary to promote it and engage with the community. By going these extra steps, a dataset is better able to foster collaboration within the machine learning community. This can lead to new GitHub repositories, academic papers, or blog posts. Organizing a competition is one of the most popular ways of realizing these benefits.\n\nWhen Aptiv released their [dataset](https://www.nuscenes.org/overview) in March 2019, they launched a competition as part of the WAD workshop at CVPR 2019. In addition to this, Aptiv released an [SDK](https://github.com/nutonomy/nuscenes-devkit) with useful helper functions and visualization code. For [evaluation](https://www.nuscenes.org/object-detection/?externalData=all&amp;mapData=all&amp;modalities=Any), organizers used a weighted sum of the metrics responsible for detection, orientation, velocity, and attribute estimation.\n\nLyft [released](https://medium.com/lyftlevel5/unlocking-access-to-self-driving-research-the-lyft-level-5-dataset-and-competition-d487c27b1b6c) the dataset slightly later. Since the data was structurally similar to those in the Aptiv dataset, we decided to use the same format.\n\nPreparing a good dataset is a big challenge. 90% of the success of the competition is defined by the quality of the data. I would like to thank **Robert Kesten, Muhammad Usman, John Houston, Tirth Pandya, K. Nadhamuni, Ana Ferreira, Michael Yuan, Ben Low, Ashesh Jain, Peter Ondrushka, Sami Omari, Shaswat Shah, Amruta Kulkarni, Alex Kazakova, Charlotte Tao, Lucas Platinsky, William Jiang, and Vinay Shet** for preparing and releasing the dataset.\n\n# How did we prepare for the competition?\nFrom the beginning, there was a plan to organize a challenge as a part of the competition track of NeurIPS 2019. We needed to decide who would host the competition for us. There were two options that we considered:\n1. Host ourselves: release the dataset, release the code for the metric, write a blog post about the challenge, provide prize money, and ask participants to send their predictions via email.\n2. Host through Kaggle. (I am not affiliated with Kaggle, but at the moment is the most established platform for machine competitions in the world). Kaggle would host the platform, the dataset, forum, real-time leaderboard, and notebooks in which participants may perform exploratory data analysis and even train models. There are other choices besides Kaggle, but they are less popular.\n\nBoth strategies have their own advantages, but the first one is more common. It is commonly used when organizers have a small budget or Kaggle cannot support some sort of functionality, like submitting a Docker container instead of a CSV file. Most of the challenges in competition tracks in conferences like NeurIPS, ICCV, ECCV, CVPR, ICML, KDD, MICCAI, etc follow the first path.\n\nWe decided to go with Kaggle. Although we knew we could host a quality competition ourselves, we wanted to make community engagement our highest priority.\n\nI believe that the success of the competition from the standpoint of the organizers is measured not by the titles and affiliations of the participants, nor the strength of the solutions, but rather by the number of active participants. If you only get 3-20 teams, you did not do a very good job. If the number is in the range 100-1000 you are doing well. Over 1000 teams and you’re doing extremely well! The background of the participants should not matter; whether they are from academia, industry, or are self-taught, everyone should be welcome.\n\nSuch a philosophy helps to further democratize the field, elevating the skill of the community as a whole while also encouraging the development of diverse new methods and applications.  \n\nIn my experience, the number of participants in Kaggle competitions is quite large, on the order of hundreds or thousands. The knowledge accumulated in Kaggle kernels and discussion threads became a great reference point for people investigating similar problems in the future.\n\nFollowing this way of thinking, we reached out to Kaggle and our partnership started.\n\nThere are many different tasks that one can try to solve on the Lyft dataset: 3D object detection, lidar point segmentation, 3D tracking, orientation, and velocity prediction, etc. All of them are interesting tasks.\n\nThis competition was the first ML challenge with this data type hosted on Kaggle. Its purpose was exploratory by nature, providing a testing ground for the ML community to become accustomed to this format of data. It is nice to get winning solutions that are closely related to the tasks that you face in production, but for this challenge it was secondary. We decided to go with the simplest possible task that one can have on this type of data: 3D object detection with the [3D version of the COCO mAP](https://github.com/lyft/nuscenes-devkit/blob/master/lyft_dataset_sdk/eval/detection/mAP_evaluation.py) as a metric. This turned out to be very convenient as it helped to bridge the gap between more familiar bounding box problems and this new challenge. It was also nice because it could be easily demonstrated in Python and reimplemented in the data scientist’s language of choice. We also had a limited timeline, so avoiding the potential for bugs was a great plus.\n\n# Sample model\nWe expected that most of the people joining the competition would not have hands-on experience with this type of data. Therefore we wanted to cater specifically to participants who were motivated by learning and skill growth. To me, machine Learning is an applied discipline, so it makes life much easier when there is [an example to learn from](https://towardsdatascience.com/ask-me-anything-session-with-a-kaggle-grandmaster-vladimir-i-iglovikov-942ad6a06acd). Taking someone else's end-to-end code and playing with it was always a good way to be less overwhelmed with the volume of new things you need to learn and implement correctly.\n\n[Guido Zuidhof](https://www.linkedin.com/in/guido-zuidhof-377b6947/), a [Kaggle Master](https://www.kaggle.com/gzuidhof) and colleague in the AV Research team at Lyft Level 5 prepared a [Kaggle kernel with a sample model](https://www.kaggle.com/gzuidhof/reference-model) that many of the participants used in their solutions.\n\n# Model limitations\n\nIn this challenge, I decided to maximize the flexibility given to participants with respect to model development:\n1. No limitations on the hardware.\n2. No limitations on the model size and inference time.\n3. Any external data or pre-trained models as long as they are listed in the [special thread at the discussion forum](https://www.kaggle.com/c/3d-object-detection-for-autonomous-vehicles/discussion/109361).\n\nIt is normally a good idea to enforce constraints on model development such that it is more compatible with a production environment, but I decided against this.\n\nI do not know any examples of when a winning ML competition solution made it all the way to production. I did not expect that this challenge would be any different. But this was not our goal. At the end of the day, the challenges that we face are different from the competition task. We were looking to foster innovation from the community. I was also worried that constraints may push some members of the community away.\n\n# Technical discussion\n## The dataset\n- The dataset can be downloaded from our [website](https://level5.lyft.com/dataset/) or on [Kaggle](https://www.kaggle.com/c/3d-object-detection-for-autonomous-vehicles/data).\n- It consists of a raw camera, lidar data, and HD semantic map.\n- 180 scenes, 25s each\n- 638,000 2D and 3D annotations over 18,000 objects\n- The dataset had nine classes with a large class imbalance.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2F856dd61314ae3420169e9db18a42a687%2FScreenshot%20from%202020-03-04%2011-39-23.png?generation=1583350814021061&amp;alt=media)\n\n## The problem\n\n- Metric: 3D version of the [2D detection COCO metric](http://cocodataset.org/#detection-eval): average mAP over 9 classes over thresholds [0.5, 0.55, 0.6, …, 0.9, 0.95]\n- Helper code: [Lyft SDK](https://github.com/lyft/nuscenes-devkit)\n- Duration two months: Sep 12 - Nov 12\n- Train set **40%**\n- Test set: public **30%**, private **30%**\n\n## Evaluation\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2F50e9bf683df416ec78397c72d8459d1f%2FScreenshot%20from%202020-03-04%2011-43-16.png?generation=1583351063905046&amp;alt=media)\n\nFor evaluation, we used the standard method that the Kaggle platform uses to prevent overfitting on the test set: public and private test splits.\n\nThe test set is split into two halves, the public test set, and the private test set. Participants did not know which samples in the test set belong to the public or private test sets.\n\nDuring the competition, participants were allowed to submit predictions on the whole test set twice per day. They would immediately receive their score with respect to the public test set, immortalized on the public leaderboard. This score is only a tentative measure of team performance. After the competition ends, the private leaderboard, made out of the predictions on the private part of the test set released and it is used to decide winners.\n\n# Solutions\n\n## Summary\n- Hardware:\n  - 1st place: **24** GPUs\n  - 2nd place: **17** GPUs\n  - 3rd place: **1** GPU\n- All teams used the [SECOND pytorch repo](https://github.com/traveller59/second.pytorch) in their solutions.\n- The best submission for each team was based on the ensemble of different models or variations of the same model.\n- Each team used different ensembling techniques.\n- Each team had a single model that would lead them to the top 5.\n- The strongest single model for all winning teams was based only on the lidar data.\n\nI asked winners to share their approach at the Kaggle discussion forum. If you have any questions, you may ask them there:\n\n1. [1st place solution](https://www.kaggle.com/c/3d-object-detection-for-autonomous-vehicles/discussion/122820)\n2. [2nd place solution](https://www.kaggle.com/c/3d-object-detection-for-autonomous-vehicles/discussion/123004)\n3. [3rd place solution](https://www.kaggle.com/c/3d-object-detection-for-autonomous-vehicles/discussion/117269)\n\n## First place solution\n### [Wenjing Zhang](https://www.kaggle.com/nywenjing)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2Fb44b5a31462460a061dcc7674cf05971%2FScreenshot%20from%202020-03-04%2011-49-40.png?generation=1583351445181412&amp;alt=media)\n\nWenjing is a Sr. AI Engineer at Ankobot, Singapore. He got his Ph.D. at Nanyang Technological University, where his research area was computer Graphics and 3D Geometry processing. He was a 2nd place winner at the [CVPR 2019 WAD challenge](https://sites.google.com/view/wad2019/challenge).\n\n\n\n### [Sanjay Addicam](https://www.kaggle.com/avsanjay)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2F4cbef3f1320a442cebb4d50a054e836c%2FScreenshot%20from%202020-03-04%2011-49-43.png?generation=1583351577845147&amp;alt=media)\n\nSanjay works as a CTO of the Visual Retail group at Intel. He is a Kaggle Master with 1 gold, 4 silvers, and 1 bronze medals.\n\nThis team did not use camera images or HD map data. Only lidar was used. The team did not use pre-trained models or external data.\n\nThe code was based on the [PointPillars method by nuTonomy](http://openaccess.thecvf.com/content_CVPR_2019/papers/Lang_PointPillars_Fast_Encoders_for_Object_Detection_From_Point_Clouds_CVPR_2019_paper.pdf) implemented in [SECOND](https://github.com/traveller59/second.pytorch). \n\nThe team trained seven different models with varying backbones and Voxel sizes.\n\nBackbones:\n- PIllarFeatureNet\n- PIllarFeatureNetRadius\n- PIllarFeatureNetHeight\n\nVoxel Sizes: 0.1, 0.125, 0.2, 0.25\n\n**Parameters**\n\n- Detection range: [-100,-100,-5,100,100,3]\n- No direction classifier\n- No specific post-processing, but score thresholding (0.1) and NMS\n- Adam with one-cycle policy\n- LR max:             1e-3\n- Division factor:  10\n- Weight decay:    0.01\n- Batch size:         2\n- Epochs:              30\n- Train augmentations: original, flip X, flip Y, flip XY, random rotation, random scaling, and random translation\n- Test augmentations: original, flip X, flip Y, flip XY\n\nAll 4 TTA * 7 models = 28 predictions were ensembled.\n\nThe best single model without TTA is **0.177 (5th place)**\nThe best single model with TTA is **0.197 (3rd place)**\nModel Ensemble **0.220 (1st place)**\n\nThey introduced a new method for 3D boxes fusion. Yaw angle had a 180-degree ambiguity for the IOU calculation. As a result, the direction of the predicted box is not reliable.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2Fa6324af55cfd06882aa92731cb58ed88%2FScreenshot%20from%202020-03-04%2011-56-10.png?generation=1583351839939577&amp;alt=media)\n\nIn the example above, a simple angle averaging depicted on the left would not give the desired result, while the proposed method can take into account the invariance of the evaluation metric with respect to the 180-degree rotations.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2Fc9094868164720a528b70a1b6b734886%2FScreenshot%20from%202020-03-04%2011-56-13.png?generation=1583351871230849&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2Fe58cefac22bff9ec086c9cf17ee9ee32%2FScreenshot%20from%202020-03-04%2012-54-00.png?generation=1583355290969346&amp;alt=media)\n\n### Approaches that did not work or were not attempted\n\n- Increasing the maximum number of voxels worked until some point, but decreased after this.\n- The team did not try other methods like PointRCNN or Frustrum PointNet\n\n## Second place solution\n### [Kyle Lee](https://www.kaggle.com/kylelee)\n\nKyle is a Senior AI Scientist at the Dishcraft Robotics. He got his B.SC. in Electrical and Computer Engineering at Cornell University. He has a working knowledge of ML/DL and perception concepts for object manipulation. Kyle is Kaggle Grandmaster.\n\n### General solution flow\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2F2aecec84e5e00c5399d3eeae21d53e6d%2FScreenshot%20from%202020-03-04%2012-57-10.png?generation=1583355476697920&amp;alt=media)\n\n\nThe solution is an ensemble of two approaches:\n1. Lidar based using VoxelNet/PointPillars based on SECOND, with modifications on data loading, point cloud range, voxel sizes, and maximum voxel size parameters with test-time\n2. A frustum-based approach based on [Frustum ConvNet](https://github.com/zhixinwang/frustum-convnet), where 2D object detection boxes were inferred from various re-trained [Detectron2](https://github.com/facebookresearch/detectron2) and [Tensorflow Object Detection](https://github.com/tensorflow/models/tree/master/research/object_detection) API object detection frameworks and the aligned point cloud sampled as a sequence of frustums into a fully convolutional network (FCN).\n\nThe first approach gave a public score of 0.191 or a private score of 0.188, which on its own would have been sufficient for ​2nd place​ on the leaderboard.\n\nThe second approach gave a public score of 0.171 or a private score of 0.169, which on its own would have been sufficient for​ 5th place​ on the leaderboard.\n\nSoft-NMS using the Gaussian function was used to ensemble all the above\npredictions together. The combination of the above boosted the LIDAR only approach\nby approximately +0.014 on both public/private leaderboards to 0.205 (public) / 0.202\n\nTrain time: 18 days for SECOND (simplified model: 3 days); 5 days for Frustum-ConvNet\n\nThese two approaches worked well in an ensemble for the two main reasons:\n1. Both gave strong models. (You cannot ensemble weak models and expect to get something that is really good. Garbage in, garbage out)\n2. The approaches have different strengths and weaknesses. \n\n### Pros of 2D images for F-ConvNet:\n- Denser information for smaller classes\n- Richer context/features can help distinguish confused classes\n- Leverage mature 2D object detectors\n\n### Cons of 2D images: \n- Occluded/truncated objects are unrecognizable in 2D images\n- Limited range\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2Ffe1187474323e715085869cc50128ee8%2FScreenshot%20from%202020-03-04%2013-00-57.png?generation=1583355710397955&amp;alt=media)\n\n- Point Cloud range  [-100, -100, -5, 100, 100, 3]\n- Voxel size: different values in the range 0.1x0.1 to 0.25x0.25\n- The number of voxels to be as large as possible.\n- Optimizer: Adam\n- Step-wise learning rate decay. At steps 100k, 200k learning rate was divided by 10.\n- Batch size 1, since the purpose was to maximize the number of voxels.\n- Train augmentations: original, flip X, flip Y, flip XY\n- Test time augmentations (Lidar model only): flip X, flip Y, flip XY\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2Fc82c4cbe4a0a2ed52808eabeafced527%2FScreenshot%20from%202020-03-04%2013-03-12.png?generation=1583355836007959&amp;alt=media)\n\n**Best single model (SECOND/Lidar)**:\n- Voxel size: 0.2x0.2x8\n- Private score: \n  - Without TTA: 0.170 **(top 5)**\n  - With TTA: 0.185 **(top 3)**\n- Train time (~3 days)\n- Test time (~3 hours) over 27,468 samples\n - 2.54 samples / s\n - 0.39s per sample\n\n### Approaches that did not work\n\nOversampling rare classes\n• On-road / off-road filtering\n• Adjacent timestamp object tracking/filtering\n• Rotated soft-NMS with height (3D)\n• Animals\n• PointRCNN\n\n## 3rd place solution\n\n### [Yusuke Muramatsu](https://www.kaggle.com/yukke42)\n\n### Data:\n- Pointcloud only. Camera and HD map was not used.\n- External data was not used.\n- Animals and emergency vehicle classes were not used.\n- Objects that had less than 5 points were ignored.\n### Model:\n- The combination of VoxelNet and PointPillars\n- The implementation is based on [SECOND](https://github.com/traveller59/second.pytorch)\n\n### General flow\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2F9d4ea57b7b5a84d768ac5cc274a73f2b%2FScreenshot%20from%202020-03-04%2013-07-40.png?generation=1583356106582763&amp;alt=media)\n\n\n### Training details\n- 50 epochs\n- Batch size = 4\n- Scheduler CosineAnnealingLR\n- Train augmentation: translation, scaling, rotation around the z-axis, mixup augmentation (pasting - objects to the point cloud from a library of objects created in ).\n- No Test Time Augmentation\n- Training time: about 2 to 4 days\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2F0f3cc1a59663d8a4cebbaf7497641e91%2FScreenshot%20from%202020-03-04%2013-09-15.png?generation=1583356194029533&amp;alt=media)\n\nDetection range:\n- Car, other vehicle, truck, bus =&gt; 100m x 75m\n- Pedestrian, bicycle, motorcycle =&gt; 100m x 50m\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2F7cd509331e168a34e702c5faca898ca4%2FScreenshot%20from%202020-03-04%2013-10-35.png?generation=1583356294708470&amp;alt=media)\n\nThe detection range along the z-axis is smaller by using the custom coordinate system.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2Fecf0a2f4931c4afd073ff4e1ae384753%2FScreenshot%20from%202020-03-04%2013-10-38.png?generation=1583356317234841&amp;alt=media)\n\n### Ideas that did not work:\n\n- Weighted-boxed-fusion. The score was lower than Soft NMS.\n- HD Map for object filtering.\n- Point cloud coloring using raw images or feature maps extracted from 2D detection.\n\n# Conclusion\n\nMy biggest fear entering the competition was that we would discover a critical blunder with the dataset after we had released it. Now that the competition is over, I am happy to say that everything went really well. I would say this is a success.\nIn the end, we had:\n- 660 competitors from all over the world that formed 547 teams\n- 5707 submissions\n- A dataset that was well prepared:\n  - No data leaks we found.\n  - The difference between Public and Private leaderboards is minimal.\n- An active community that created a number of pull requests to the Lyft SDK with improvements and bug fixes.\n- Plenty of positive feedback from emails and in-person thanking us for the dataset and competition.\n\n# Improvements and opportunities for next time\n\n1. Longer duration: Many people gave me feedback that it takes some time to get used to the dataset and two months that we had for the challenge is not enough. I believe 3-4 months would work better.\n2. Build a new metric to encourage data fusion: Many teams were able to work exclusively off of lidar data. Next time we should create a metric that would be improved by using more radar, lidar, and camera images while also applying constraints to inference time. \n\nI would like to give special thanks to [Christy Robertson](https://www.linkedin.com/in/christina-robertson/) who was driving the marketing part of the project and without whom it would not have any chances for success and [Erik Gaasedelen](https://www.linkedin.com/in/erikgaas/) who helped me to prepare this blog post. And last but not least, I would like to thank all the enthusiastic Kagglers who competed in the challenge. We would not have been successful without your diligence and helpfulness within the community.",
      "votes": null
    },
    {
      "id": "763909",
      "postDate": "03/05/2020 00:58:16",
      "content": "<p><a href=\"/iglovikov\">@iglovikov</a> Thanks for sharing the summary of Lyft competition. I liked the part of summary where it describes about reducing the data gap from simulations to reality. </p>\n\n<p>As mentioned for <strong>autonomous vehicles</strong> \"...Lyft, use a combination of camera, lidar, and radar sensors than only the first two types\". </p>\n\n<p><strong>A question out of curiosity</strong>: Are these the only types of data Lyft intend to release or it's the complete set of types of data?</p>\n\n<p>Please Note: I was not a participant in this competition. </p>\n\n<p>Thanks\nCheers</p>",
      "rawMarkdown": "iglovikov Thanks for sharing the summary of Lyft competition. I liked the part of summary where it describes about reducing the data gap from simulations to reality. \n\nAs mentioned for **autonomous vehicles** \"...Lyft, use a combination of camera, lidar, and radar sensors than only the first two types\". \n\n**A question out of curiosity**: Are these the only types of data Lyft intend to release or it's the complete set of types of data?\n\nPlease Note: I was not a participant in this competition. \n\nThanks\nCheers",
      "votes": null
    },
    {
      "id": "763912",
      "postDate": "03/05/2020 01:02:17",
      "content": "<p>Awesome write-up! Thank you for making it 🙏  </p>",
      "rawMarkdown": "Awesome write-up! Thank you for making it 🙏",
      "votes": null
    },
    {
      "id": "765076",
      "postDate": "03/06/2020 08:31:40",
      "content": "<p>Thanks for sharing the summary of the competition </p>",
      "rawMarkdown": "Thanks for sharing the summary of the competition",
      "votes": null
    },
    {
      "id": "766320",
      "postDate": "03/08/2020 01:52:14",
      "content": "<p>Thanks for the summary! <a href=\"/iglovikov\">@iglovikov</a>\nVery nice detailed overview. </p>",
      "rawMarkdown": "Thanks for the summary! @iglovikov\nVery nice detailed overview.",
      "votes": null
    },
    {
      "id": "771599",
      "postDate": "03/14/2020 11:31:23",
      "content": "<p>Awesome content!</p>",
      "rawMarkdown": "Awesome content!",
      "votes": null
    },
    {
      "id": "786219",
      "postDate": "03/25/2020 17:55:49",
      "content": "<p><a href=\"/iglovikov\">@iglovikov</a> thanks!</p>",
      "rawMarkdown": "iglovikov thanks!",
      "votes": null
    },
    {
      "id": "1009619",
      "postDate": "09/14/2020 06:23:48",
      "content": "<p>thank for this brief history of the competition.</p>",
      "rawMarkdown": "thank for this brief history of the competition.",
      "votes": null
    },
    {
      "id": "1093952",
      "postDate": "11/28/2020 06:55:36",
      "content": "<p>thank you for sharing</p>",
      "rawMarkdown": "thank you for sharing",
      "votes": null
    },
    {
      "id": "1827897",
      "postDate": "06/21/2022 11:38:53",
      "content": "<p>You made the same error like Tesla by changing the metric from class-avg-per-frame to frame-avg-per-class without even noticing. The models won, that ignore emergency vehicles and concentrate on the classes with lots of samples.<br>\nNot recognizing any emergency vehicle in this competition just had a minimal influence on the score. Unfortunately this costs life in the real world. Even more unfortunately that no one cared that the metric was not implemented like advertised. When i saw the Tesla report i had to think of this competition as it seems to be exactly the same error.</p>\n<p><a href=\"https://www.kaggle.com/competitions/3d-object-detection-for-autonomous-vehicles/discussion/116434#673846\" target=\"_blank\">https://www.kaggle.com/competitions/3d-object-detection-for-autonomous-vehicles/discussion/116434#673846</a><br>\n<a href=\"https://www.cbsnews.com/news/tesla-cars-crashes-emergency-vehicles/\" target=\"_blank\">https://www.cbsnews.com/news/tesla-cars-crashes-emergency-vehicles/</a></p>",
      "rawMarkdown": "You made the same error like Tesla by changing the metric from class-avg-per-frame to frame-avg-per-class without even noticing. The models won, that ignore emergency vehicles and concentrate on the classes with lots of samples.\nNot recognizing any emergency vehicle in this competition just had a minimal influence on the score. Unfortunately this costs life in the real world. Even more unfortunately that no one cared that the metric was not implemented like advertised. When i saw the Tesla report i had to think of this competition as it seems to be exactly the same error.\n\nhttps://www.kaggle.com/competitions/3d-object-detection-for-autonomous-vehicles/discussion/116434#673846\nhttps://www.cbsnews.com/news/tesla-cars-crashes-emergency-vehicles/",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1009619,
      "author_name": "venkataramananc",
      "author_url": "",
      "post_date": "09/14/2020 06:23:48",
      "content": "<p>thank for this brief history of the competition.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1093952,
      "author_name": "zilopark",
      "author_url": "",
      "post_date": "11/28/2020 06:55:36",
      "content": "<p>thank you for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1827897,
      "author_name": "marekwyborski",
      "author_url": "",
      "post_date": "06/21/2022 11:38:53",
      "content": "<p>You made the same error like Tesla by changing the metric from class-avg-per-frame to frame-avg-per-class without even noticing. The models won, that ignore emergency vehicles and concentrate on the classes with lots of samples.<br>\nNot recognizing any emergency vehicle in this competition just had a minimal influence on the score. Unfortunately this costs life in the real world. Even more unfortunately that no one cared that the metric was not implemented like advertised. When i saw the Tesla report i had to think of this competition as it seems to be exactly the same error.</p>\n<p><a href=\"https://www.kaggle.com/competitions/3d-object-detection-for-autonomous-vehicles/discussion/116434#673846\" target=\"_blank\">https://www.kaggle.com/competitions/3d-object-detection-for-autonomous-vehicles/discussion/116434#673846</a><br>\n<a href=\"https://www.cbsnews.com/news/tesla-cars-crashes-emergency-vehicles/\" target=\"_blank\">https://www.cbsnews.com/news/tesla-cars-crashes-emergency-vehicles/</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 763909,
      "author_name": "anshumoudgil",
      "author_url": "",
      "post_date": "03/05/2020 00:58:16",
      "content": "<p><a href=\"/iglovikov\">@iglovikov</a> Thanks for sharing the summary of Lyft competition. I liked the part of summary where it describes about reducing the data gap from simulations to reality. </p>\n\n<p>As mentioned for <strong>autonomous vehicles</strong> \"...Lyft, use a combination of camera, lidar, and radar sensors than only the first two types\". </p>\n\n<p><strong>A question out of curiosity</strong>: Are these the only types of data Lyft intend to release or it's the complete set of types of data?</p>\n\n<p>Please Note: I was not a participant in this competition. </p>\n\n<p>Thanks\nCheers</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 763912,
      "author_name": "blondinka",
      "author_url": "",
      "post_date": "03/05/2020 01:02:17",
      "content": "<p>Awesome write-up! Thank you for making it 🙏  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 765076,
      "author_name": "dhivyaraveendran",
      "author_url": "",
      "post_date": "03/06/2020 08:31:40",
      "content": "<p>Thanks for sharing the summary of the competition </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 766320,
      "author_name": "muhakabartay",
      "author_url": "",
      "post_date": "03/08/2020 01:52:14",
      "content": "<p>Thanks for the summary! <a href=\"/iglovikov\">@iglovikov</a>\nVery nice detailed overview. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 771599,
      "author_name": "sharanbabu",
      "author_url": "",
      "post_date": "03/14/2020 11:31:23",
      "content": "<p>Awesome content!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 786219,
      "author_name": "vitaliylyalin7000",
      "author_url": "",
      "post_date": "03/25/2020 17:55:49",
      "content": "<p><a href=\"/iglovikov\">@iglovikov</a> thanks!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "763788": "# Introduction\n\nHello! My name is [Vladimir Iglovikov](https://medium.com/lyftlevel5/employee-spotlight-vladimir-iglovikov-11b71ed02bc), and I am a Computer Vision Engineer at Lyft Level 5. Here, I work on a diverse set of projects relating to building, testing, deploying, and refining deep learning models for mapping and perception related tasks. \n\nI am [Kaggle Grandmaster](http://ternaus.blog/interview/2018/08/30/ama.html), which means that I have some experience of [participating](http://ternaus.blog/career/2020/01/08/How-I-Found-my-current-job.html) in machine learning competitions. As a result, I was very pleased to join a team of colleagues at Level 5 who wanted to create a new machine learning competition for Lyft.\n\nAfter carefully putting together a dataset and an objective, we finally put together the Lyft 3D Object Detection for Autonomous Vehicles Challenge. Over the course of this challenge, over 500 teams plugged away at achieving a top leaderboard spot. Now that the challenge has finished and the dust has settled, we would like to share with you the top solutions.\n\nI would also like to discuss what we have learned throughout the process as organizers of the competition. There are many [difficulties](https://docs.google.com/a/chalearn.org/viewer?a=v&amp;pid=sites&amp;srcid=Y2hhbGVhcm4ub3JnfHdvcmtzaG9wfGd4OjQwYzBhY2Q1N2YxNGE3YWI) to overcome when developing a machine learning challenge. From data leaks to issues with leaderboard scoring, it is common to make critical mistakes. I have provided much feedback to organizers when I was a participant in different machine learning competitions. Now it was time to get such feedback myself.\n\nThis blog post will have two parts. In the first introductory part, I will tell you how this project was started and the challenges we ran into. In the second technical part, I will describe the solutions from the top three teams in a level of depth and rigor that I would have wanted to learn from if I had been a participant.\n\n# Why does Lyft need a Machine Learning Competition?\n\nWhen I go to academic conferences I see a lot of papers that have the keywords “autonomous” and “self-driving”. For sure, all these academic works are related to the autonomous industry in some way, but for many of them, the relationship is rather weak.\n\nOne of the reasons for this gap may be that the research community does not have access to the same data as the industry. There are many academic datasets that rely solely on camera images. Meanwhile, companies, including Lyft, use a combination of camera, lidar, and radar sensors, all of which are crucial for the construction of high definition maps or the functioning of autonomy algorithms. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2Fd280f4af3e91991ad641606967cd0bbe%2FScreenshot%20from%202020-03-04%2011-28-24.png?generation=1583350172278513&amp;alt=media)\n\nPrior to 2019, the [KITTI](http://www.cvlibs.net/datasets/kitti/) dataset was the most useful dataset within the research community, that provided both camera and lidar data. KITTI was far ahead of its time and has been a staple in the self-driving industry even today, but it has its own limitations. \n\nIn 2019 the field shifted rapidly. Lyft, along with some other self-driving companies, released their own, much larger datasets. This was a great move, but it is not enough to just release the dataset. It is also necessary to promote it and engage with the community. By going these extra steps, a dataset is better able to foster collaboration within the machine learning community. This can lead to new GitHub repositories, academic papers, or blog posts. Organizing a competition is one of the most popular ways of realizing these benefits.\n\nWhen Aptiv released their [dataset](https://www.nuscenes.org/overview) in March 2019, they launched a competition as part of the WAD workshop at CVPR 2019. In addition to this, Aptiv released an [SDK](https://github.com/nutonomy/nuscenes-devkit) with useful helper functions and visualization code. For [evaluation](https://www.nuscenes.org/object-detection/?externalData=all&amp;mapData=all&amp;modalities=Any), organizers used a weighted sum of the metrics responsible for detection, orientation, velocity, and attribute estimation.\n\nLyft [released](https://medium.com/lyftlevel5/unlocking-access-to-self-driving-research-the-lyft-level-5-dataset-and-competition-d487c27b1b6c) the dataset slightly later. Since the data was structurally similar to those in the Aptiv dataset, we decided to use the same format.\n\nPreparing a good dataset is a big challenge. 90% of the success of the competition is defined by the quality of the data. I would like to thank **Robert Kesten, Muhammad Usman, John Houston, Tirth Pandya, K. Nadhamuni, Ana Ferreira, Michael Yuan, Ben Low, Ashesh Jain, Peter Ondrushka, Sami Omari, Shaswat Shah, Amruta Kulkarni, Alex Kazakova, Charlotte Tao, Lucas Platinsky, William Jiang, and Vinay Shet** for preparing and releasing the dataset.\n\n# How did we prepare for the competition?\nFrom the beginning, there was a plan to organize a challenge as a part of the competition track of NeurIPS 2019. We needed to decide who would host the competition for us. There were two options that we considered:\n1. Host ourselves: release the dataset, release the code for the metric, write a blog post about the challenge, provide prize money, and ask participants to send their predictions via email.\n2. Host through Kaggle. (I am not affiliated with Kaggle, but at the moment is the most established platform for machine competitions in the world). Kaggle would host the platform, the dataset, forum, real-time leaderboard, and notebooks in which participants may perform exploratory data analysis and even train models. There are other choices besides Kaggle, but they are less popular.\n\nBoth strategies have their own advantages, but the first one is more common. It is commonly used when organizers have a small budget or Kaggle cannot support some sort of functionality, like submitting a Docker container instead of a CSV file. Most of the challenges in competition tracks in conferences like NeurIPS, ICCV, ECCV, CVPR, ICML, KDD, MICCAI, etc follow the first path.\n\nWe decided to go with Kaggle. Although we knew we could host a quality competition ourselves, we wanted to make community engagement our highest priority.\n\nI believe that the success of the competition from the standpoint of the organizers is measured not by the titles and affiliations of the participants, nor the strength of the solutions, but rather by the number of active participants. If you only get 3-20 teams, you did not do a very good job. If the number is in the range 100-1000 you are doing well. Over 1000 teams and you’re doing extremely well! The background of the participants should not matter; whether they are from academia, industry, or are self-taught, everyone should be welcome.\n\nSuch a philosophy helps to further democratize the field, elevating the skill of the community as a whole while also encouraging the development of diverse new methods and applications.  \n\nIn my experience, the number of participants in Kaggle competitions is quite large, on the order of hundreds or thousands. The knowledge accumulated in Kaggle kernels and discussion threads became a great reference point for people investigating similar problems in the future.\n\nFollowing this way of thinking, we reached out to Kaggle and our partnership started.\n\nThere are many different tasks that one can try to solve on the Lyft dataset: 3D object detection, lidar point segmentation, 3D tracking, orientation, and velocity prediction, etc. All of them are interesting tasks.\n\nThis competition was the first ML challenge with this data type hosted on Kaggle. Its purpose was exploratory by nature, providing a testing ground for the ML community to become accustomed to this format of data. It is nice to get winning solutions that are closely related to the tasks that you face in production, but for this challenge it was secondary. We decided to go with the simplest possible task that one can have on this type of data: 3D object detection with the [3D version of the COCO mAP](https://github.com/lyft/nuscenes-devkit/blob/master/lyft_dataset_sdk/eval/detection/mAP_evaluation.py) as a metric. This turned out to be very convenient as it helped to bridge the gap between more familiar bounding box problems and this new challenge. It was also nice because it could be easily demonstrated in Python and reimplemented in the data scientist’s language of choice. We also had a limited timeline, so avoiding the potential for bugs was a great plus.\n\n# Sample model\nWe expected that most of the people joining the competition would not have hands-on experience with this type of data. Therefore we wanted to cater specifically to participants who were motivated by learning and skill growth. To me, machine Learning is an applied discipline, so it makes life much easier when there is [an example to learn from](https://towardsdatascience.com/ask-me-anything-session-with-a-kaggle-grandmaster-vladimir-i-iglovikov-942ad6a06acd). Taking someone else's end-to-end code and playing with it was always a good way to be less overwhelmed with the volume of new things you need to learn and implement correctly.\n\n[Guido Zuidhof](https://www.linkedin.com/in/guido-zuidhof-377b6947/), a [Kaggle Master](https://www.kaggle.com/gzuidhof) and colleague in the AV Research team at Lyft Level 5 prepared a [Kaggle kernel with a sample model](https://www.kaggle.com/gzuidhof/reference-model) that many of the participants used in their solutions.\n\n# Model limitations\n\nIn this challenge, I decided to maximize the flexibility given to participants with respect to model development:\n1. No limitations on the hardware.\n2. No limitations on the model size and inference time.\n3. Any external data or pre-trained models as long as they are listed in the [special thread at the discussion forum](https://www.kaggle.com/c/3d-object-detection-for-autonomous-vehicles/discussion/109361).\n\nIt is normally a good idea to enforce constraints on model development such that it is more compatible with a production environment, but I decided against this.\n\nI do not know any examples of when a winning ML competition solution made it all the way to production. I did not expect that this challenge would be any different. But this was not our goal. At the end of the day, the challenges that we face are different from the competition task. We were looking to foster innovation from the community. I was also worried that constraints may push some members of the community away.\n\n# Technical discussion\n## The dataset\n- The dataset can be downloaded from our [website](https://level5.lyft.com/dataset/) or on [Kaggle](https://www.kaggle.com/c/3d-object-detection-for-autonomous-vehicles/data).\n- It consists of a raw camera, lidar data, and HD semantic map.\n- 180 scenes, 25s each\n- 638,000 2D and 3D annotations over 18,000 objects\n- The dataset had nine classes with a large class imbalance.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2F856dd61314ae3420169e9db18a42a687%2FScreenshot%20from%202020-03-04%2011-39-23.png?generation=1583350814021061&amp;alt=media)\n\n## The problem\n\n- Metric: 3D version of the [2D detection COCO metric](http://cocodataset.org/#detection-eval): average mAP over 9 classes over thresholds [0.5, 0.55, 0.6, …, 0.9, 0.95]\n- Helper code: [Lyft SDK](https://github.com/lyft/nuscenes-devkit)\n- Duration two months: Sep 12 - Nov 12\n- Train set **40%**\n- Test set: public **30%**, private **30%**\n\n## Evaluation\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2F50e9bf683df416ec78397c72d8459d1f%2FScreenshot%20from%202020-03-04%2011-43-16.png?generation=1583351063905046&amp;alt=media)\n\nFor evaluation, we used the standard method that the Kaggle platform uses to prevent overfitting on the test set: public and private test splits.\n\nThe test set is split into two halves, the public test set, and the private test set. Participants did not know which samples in the test set belong to the public or private test sets.\n\nDuring the competition, participants were allowed to submit predictions on the whole test set twice per day. They would immediately receive their score with respect to the public test set, immortalized on the public leaderboard. This score is only a tentative measure of team performance. After the competition ends, the private leaderboard, made out of the predictions on the private part of the test set released and it is used to decide winners.\n\n# Solutions\n\n## Summary\n- Hardware:\n  - 1st place: **24** GPUs\n  - 2nd place: **17** GPUs\n  - 3rd place: **1** GPU\n- All teams used the [SECOND pytorch repo](https://github.com/traveller59/second.pytorch) in their solutions.\n- The best submission for each team was based on the ensemble of different models or variations of the same model.\n- Each team used different ensembling techniques.\n- Each team had a single model that would lead them to the top 5.\n- The strongest single model for all winning teams was based only on the lidar data.\n\nI asked winners to share their approach at the Kaggle discussion forum. If you have any questions, you may ask them there:\n\n1. [1st place solution](https://www.kaggle.com/c/3d-object-detection-for-autonomous-vehicles/discussion/122820)\n2. [2nd place solution](https://www.kaggle.com/c/3d-object-detection-for-autonomous-vehicles/discussion/123004)\n3. [3rd place solution](https://www.kaggle.com/c/3d-object-detection-for-autonomous-vehicles/discussion/117269)\n\n## First place solution\n### [Wenjing Zhang](https://www.kaggle.com/nywenjing)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2Fb44b5a31462460a061dcc7674cf05971%2FScreenshot%20from%202020-03-04%2011-49-40.png?generation=1583351445181412&amp;alt=media)\n\nWenjing is a Sr. AI Engineer at Ankobot, Singapore. He got his Ph.D. at Nanyang Technological University, where his research area was computer Graphics and 3D Geometry processing. He was a 2nd place winner at the [CVPR 2019 WAD challenge](https://sites.google.com/view/wad2019/challenge).\n\n\n\n### [Sanjay Addicam](https://www.kaggle.com/avsanjay)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2F4cbef3f1320a442cebb4d50a054e836c%2FScreenshot%20from%202020-03-04%2011-49-43.png?generation=1583351577845147&amp;alt=media)\n\nSanjay works as a CTO of the Visual Retail group at Intel. He is a Kaggle Master with 1 gold, 4 silvers, and 1 bronze medals.\n\nThis team did not use camera images or HD map data. Only lidar was used. The team did not use pre-trained models or external data.\n\nThe code was based on the [PointPillars method by nuTonomy](http://openaccess.thecvf.com/content_CVPR_2019/papers/Lang_PointPillars_Fast_Encoders_for_Object_Detection_From_Point_Clouds_CVPR_2019_paper.pdf) implemented in [SECOND](https://github.com/traveller59/second.pytorch). \n\nThe team trained seven different models with varying backbones and Voxel sizes.\n\nBackbones:\n- PIllarFeatureNet\n- PIllarFeatureNetRadius\n- PIllarFeatureNetHeight\n\nVoxel Sizes: 0.1, 0.125, 0.2, 0.25\n\n**Parameters**\n\n- Detection range: [-100,-100,-5,100,100,3]\n- No direction classifier\n- No specific post-processing, but score thresholding (0.1) and NMS\n- Adam with one-cycle policy\n- LR max:             1e-3\n- Division factor:  10\n- Weight decay:    0.01\n- Batch size:         2\n- Epochs:              30\n- Train augmentations: original, flip X, flip Y, flip XY, random rotation, random scaling, and random translation\n- Test augmentations: original, flip X, flip Y, flip XY\n\nAll 4 TTA * 7 models = 28 predictions were ensembled.\n\nThe best single model without TTA is **0.177 (5th place)**\nThe best single model with TTA is **0.197 (3rd place)**\nModel Ensemble **0.220 (1st place)**\n\nThey introduced a new method for 3D boxes fusion. Yaw angle had a 180-degree ambiguity for the IOU calculation. As a result, the direction of the predicted box is not reliable.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2Fa6324af55cfd06882aa92731cb58ed88%2FScreenshot%20from%202020-03-04%2011-56-10.png?generation=1583351839939577&amp;alt=media)\n\nIn the example above, a simple angle averaging depicted on the left would not give the desired result, while the proposed method can take into account the invariance of the evaluation metric with respect to the 180-degree rotations.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2Fc9094868164720a528b70a1b6b734886%2FScreenshot%20from%202020-03-04%2011-56-13.png?generation=1583351871230849&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2Fe58cefac22bff9ec086c9cf17ee9ee32%2FScreenshot%20from%202020-03-04%2012-54-00.png?generation=1583355290969346&amp;alt=media)\n\n### Approaches that did not work or were not attempted\n\n- Increasing the maximum number of voxels worked until some point, but decreased after this.\n- The team did not try other methods like PointRCNN or Frustrum PointNet\n\n## Second place solution\n### [Kyle Lee](https://www.kaggle.com/kylelee)\n\nKyle is a Senior AI Scientist at the Dishcraft Robotics. He got his B.SC. in Electrical and Computer Engineering at Cornell University. He has a working knowledge of ML/DL and perception concepts for object manipulation. Kyle is Kaggle Grandmaster.\n\n### General solution flow\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2F2aecec84e5e00c5399d3eeae21d53e6d%2FScreenshot%20from%202020-03-04%2012-57-10.png?generation=1583355476697920&amp;alt=media)\n\n\nThe solution is an ensemble of two approaches:\n1. Lidar based using VoxelNet/PointPillars based on SECOND, with modifications on data loading, point cloud range, voxel sizes, and maximum voxel size parameters with test-time\n2. A frustum-based approach based on [Frustum ConvNet](https://github.com/zhixinwang/frustum-convnet), where 2D object detection boxes were inferred from various re-trained [Detectron2](https://github.com/facebookresearch/detectron2) and [Tensorflow Object Detection](https://github.com/tensorflow/models/tree/master/research/object_detection) API object detection frameworks and the aligned point cloud sampled as a sequence of frustums into a fully convolutional network (FCN).\n\nThe first approach gave a public score of 0.191 or a private score of 0.188, which on its own would have been sufficient for ​2nd place​ on the leaderboard.\n\nThe second approach gave a public score of 0.171 or a private score of 0.169, which on its own would have been sufficient for​ 5th place​ on the leaderboard.\n\nSoft-NMS using the Gaussian function was used to ensemble all the above\npredictions together. The combination of the above boosted the LIDAR only approach\nby approximately +0.014 on both public/private leaderboards to 0.205 (public) / 0.202\n\nTrain time: 18 days for SECOND (simplified model: 3 days); 5 days for Frustum-ConvNet\n\nThese two approaches worked well in an ensemble for the two main reasons:\n1. Both gave strong models. (You cannot ensemble weak models and expect to get something that is really good. Garbage in, garbage out)\n2. The approaches have different strengths and weaknesses. \n\n### Pros of 2D images for F-ConvNet:\n- Denser information for smaller classes\n- Richer context/features can help distinguish confused classes\n- Leverage mature 2D object detectors\n\n### Cons of 2D images: \n- Occluded/truncated objects are unrecognizable in 2D images\n- Limited range\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2Ffe1187474323e715085869cc50128ee8%2FScreenshot%20from%202020-03-04%2013-00-57.png?generation=1583355710397955&amp;alt=media)\n\n- Point Cloud range  [-100, -100, -5, 100, 100, 3]\n- Voxel size: different values in the range 0.1x0.1 to 0.25x0.25\n- The number of voxels to be as large as possible.\n- Optimizer: Adam\n- Step-wise learning rate decay. At steps 100k, 200k learning rate was divided by 10.\n- Batch size 1, since the purpose was to maximize the number of voxels.\n- Train augmentations: original, flip X, flip Y, flip XY\n- Test time augmentations (Lidar model only): flip X, flip Y, flip XY\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2Fc82c4cbe4a0a2ed52808eabeafced527%2FScreenshot%20from%202020-03-04%2013-03-12.png?generation=1583355836007959&amp;alt=media)\n\n**Best single model (SECOND/Lidar)**:\n- Voxel size: 0.2x0.2x8\n- Private score: \n  - Without TTA: 0.170 **(top 5)**\n  - With TTA: 0.185 **(top 3)**\n- Train time (~3 days)\n- Test time (~3 hours) over 27,468 samples\n - 2.54 samples / s\n - 0.39s per sample\n\n### Approaches that did not work\n\nOversampling rare classes\n• On-road / off-road filtering\n• Adjacent timestamp object tracking/filtering\n• Rotated soft-NMS with height (3D)\n• Animals\n• PointRCNN\n\n## 3rd place solution\n\n### [Yusuke Muramatsu](https://www.kaggle.com/yukke42)\n\n### Data:\n- Pointcloud only. Camera and HD map was not used.\n- External data was not used.\n- Animals and emergency vehicle classes were not used.\n- Objects that had less than 5 points were ignored.\n### Model:\n- The combination of VoxelNet and PointPillars\n- The implementation is based on [SECOND](https://github.com/traveller59/second.pytorch)\n\n### General flow\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2F9d4ea57b7b5a84d768ac5cc274a73f2b%2FScreenshot%20from%202020-03-04%2013-07-40.png?generation=1583356106582763&amp;alt=media)\n\n\n### Training details\n- 50 epochs\n- Batch size = 4\n- Scheduler CosineAnnealingLR\n- Train augmentation: translation, scaling, rotation around the z-axis, mixup augmentation (pasting - objects to the point cloud from a library of objects created in ).\n- No Test Time Augmentation\n- Training time: about 2 to 4 days\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2F0f3cc1a59663d8a4cebbaf7497641e91%2FScreenshot%20from%202020-03-04%2013-09-15.png?generation=1583356194029533&amp;alt=media)\n\nDetection range:\n- Car, other vehicle, truck, bus =&gt; 100m x 75m\n- Pedestrian, bicycle, motorcycle =&gt; 100m x 50m\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2F7cd509331e168a34e702c5faca898ca4%2FScreenshot%20from%202020-03-04%2013-10-35.png?generation=1583356294708470&amp;alt=media)\n\nThe detection range along the z-axis is smaller by using the custom coordinate system.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F286455%2Fecf0a2f4931c4afd073ff4e1ae384753%2FScreenshot%20from%202020-03-04%2013-10-38.png?generation=1583356317234841&amp;alt=media)\n\n### Ideas that did not work:\n\n- Weighted-boxed-fusion. The score was lower than Soft NMS.\n- HD Map for object filtering.\n- Point cloud coloring using raw images or feature maps extracted from 2D detection.\n\n# Conclusion\n\nMy biggest fear entering the competition was that we would discover a critical blunder with the dataset after we had released it. Now that the competition is over, I am happy to say that everything went really well. I would say this is a success.\nIn the end, we had:\n- 660 competitors from all over the world that formed 547 teams\n- 5707 submissions\n- A dataset that was well prepared:\n  - No data leaks we found.\n  - The difference between Public and Private leaderboards is minimal.\n- An active community that created a number of pull requests to the Lyft SDK with improvements and bug fixes.\n- Plenty of positive feedback from emails and in-person thanking us for the dataset and competition.\n\n# Improvements and opportunities for next time\n\n1. Longer duration: Many people gave me feedback that it takes some time to get used to the dataset and two months that we had for the challenge is not enough. I believe 3-4 months would work better.\n2. Build a new metric to encourage data fusion: Many teams were able to work exclusively off of lidar data. Next time we should create a metric that would be improved by using more radar, lidar, and camera images while also applying constraints to inference time. \n\nI would like to give special thanks to [Christy Robertson](https://www.linkedin.com/in/christina-robertson/) who was driving the marketing part of the project and without whom it would not have any chances for success and [Erik Gaasedelen](https://www.linkedin.com/in/erikgaas/) who helped me to prepare this blog post. And last but not least, I would like to thank all the enthusiastic Kagglers who competed in the challenge. We would not have been successful without your diligence and helpfulness within the community.",
    "763909": "iglovikov Thanks for sharing the summary of Lyft competition. I liked the part of summary where it describes about reducing the data gap from simulations to reality. \n\nAs mentioned for **autonomous vehicles** \"...Lyft, use a combination of camera, lidar, and radar sensors than only the first two types\". \n\n**A question out of curiosity**: Are these the only types of data Lyft intend to release or it's the complete set of types of data?\n\nPlease Note: I was not a participant in this competition. \n\nThanks\nCheers",
    "763912": "Awesome write-up! Thank you for making it 🙏",
    "765076": "Thanks for sharing the summary of the competition",
    "766320": "Thanks for the summary! @iglovikov\nVery nice detailed overview.",
    "771599": "Awesome content!",
    "786219": "iglovikov thanks!",
    "1009619": "thank for this brief history of the competition.",
    "1093952": "thank you for sharing",
    "1827897": "You made the same error like Tesla by changing the metric from class-avg-per-frame to frame-avg-per-class without even noticing. The models won, that ignore emergency vehicles and concentrate on the classes with lots of samples.\nNot recognizing any emergency vehicle in this competition just had a minimal influence on the score. Unfortunately this costs life in the real world. Even more unfortunately that no one cared that the metric was not implemented like advertised. When i saw the Tesla report i had to think of this competition as it seems to be exactly the same error.\n\nhttps://www.kaggle.com/competitions/3d-object-detection-for-autonomous-vehicles/discussion/116434#673846\nhttps://www.cbsnews.com/news/tesla-cars-crashes-emergency-vehicles/"
  },
  "source": "meta"
}