{
  "id": 220576,
  "title": "Lyft Motion Prediction for Autonomous Vehicles Follow-up",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/220576",
  "author_name": "Luca Bergamini",
  "post_date": "2021-02-18T22:03:30.269000",
  "votes": 8,
  "comment_count": 2,
  "views": 0,
  "content": "<h1>Introduction</h1>\n<p>The Lyft 2020 competition (hosted on Kaggle) attracted more than 900 teams with a total of 14.904 submissions. After a few months since its completion, it’s time to draw some conclusions on what kagglers tried, what worked and what didn’t. The community wrote 320 posts where they shared ideas, results and suggestions. We’ve collected and filtered them to identify 5 insights that can help you hit the ground running in this challenging task. Let’s start with a short recap of what participants were asked to do and what tools Lyft provided them to start coding straight away.</p>\n<h1>The task</h1>\n<p>Motion prediction is an essential component of the autonomous stack. The task is to understand what other traffic participants will do in the short future. This can be applied to other vehicles around the AV but also vulnerable road users such as pedestrians and cyclists. For each of them, we want to predict the future location as a trajectory of 2D points. This information is crucial for the final planning of the AV trajectory, both for rule-based and deep learning based systems.</p>\n<p>However, the future is highly uncertain and a single trajectory per agent may not be enough. Even for a vehicle driving on a straight road, many different speed profiles can be chosen. All these trajectories may be in a sense “correct” (kinematically feasible and safe to take) but, in practice, only one of them will be picked when we will deploy this model. Still, it’s important to keep the possibility to generate multiple trajectories, as other models such as the planner may require them. For this reason, we highly encouraged participants to submit multiple trajectories per agent by providing them with a metric that can take into account this explicitly.</p>\n<p><a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/Gc3PtRY/Prediction-task.jpg\" alt=\"Prediction-task\"></a></p>\n<p>Prediction is a very challenging task! It may sound easy at first when you have only a bunch of agents in the scene but, as you add more of them to the mix, your model must have a “global” vision of what’s going on to correctly predict the future of them.</p>\n<h1>The data</h1>\n<p>Challenging tasks require great datasets to be solved. We provided participants with one collected by the Lyft fleet, which includes more than 1000 hours of driving and HD quality semantic map annotations. </p>\n<p><a href=\"https://ibb.co/chm2p30\"><img src=\"https://i.ibb.co/gZntX9N/Screenshot-2021-01-15-at-11-53-21.png\" alt=\"Screenshot-2021-01-15-at-11-53-21\"></a></p>\n<p>For each frame in the dataset multiple agents are available, and their annotated trajectory can be read from future frames. These annotations come from the latest version of our in-house perception and tracking system. We also provide static road geometry information, including lanes and crosswalks, as well as dynamic information such as traffic lights. This dataset is public and already available to download. Researchers can use it with no restrictions, and it’s always possible to submit results to the leaderboard and get a score, even after the end of the competition.</p>\n<h1>The tools</h1>\n<p>The motion prediction task was new to the Kaggle community, and working with such a huge dataset can look daunting at first. To encourage people from different backgrounds to take part in this challenge we decided to release <a href=\"https://github.com/lyft/l5kit\" target=\"_blank\">L5Kit</a>, a library of functionalities to work with our dataset and extract useful information. This is the very same tool we use for our internal projects, and we provide support for it as well as integrate precious feedback from the community. L5Kit provided utility functions to read and visualise data, as well as metrics to score solutions. The toolkit is particularly useful for training CNNs, as it can generate rasterized Bird-Eye-View representation from the data. These rasters can include map semantic information as well as satellite imagery.</p>\n<p><a href=\"https://ibb.co/zZ4NW5B\"><img src=\"https://i.ibb.co/qprxLnc/examples-dataset.jpg\" alt=\"examples-dataset\"></a></p>\n<p>For this year competition we also included a kernel with a CNN baseline for the task participants could use and build upon. The kernel was intended as a starting point for the competition, as the task was new, the format of the dataset unfamiliar and it was unclear how to start with it. More than 60 teams started their experiments from our kernel, and many more leveraged our pre-trained network in their solutions.</p>\n<h1>What it takes to win a competition</h1>\n<p>After reading many posts and interviewing the top scorers, we came up with a list of insights which you definitely should know when taking part into a competition about motion prediction. We report them here in no particular order.</p>\n<h2>Understand the task</h2>\n<p>Yes, it may sound trivial but you really need to invest some time to understand which world’s entities you are trying to model. Deep learning is powerful and can learn almost anything (provided the input representation is meaningful) but knowing what you are modeling can drastically change your approach. For the task of prediction, knowing that your final targets are trajectories which must be executed by an agent can give you an edge.</p>\n<p>As an example, the <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/201493\" target=\"_blank\">1st team</a> experimented with predicting only differences from a constant velocity model. Other teams tried to predict acceleration and steering instead of displacements and used a unicycle kinematic model to retrieve the trajectories, while others tried to smooth them using a Kalman filter.</p>\n<p>The 3rd team realised that the traffic light information in the rasteriser was static, and that the model might benefit from knowing the evolution of this signal through time.</p>\n<p>Again, the 6th team investigated a lot how to balance data based on the acceleration of the agents in the scene. They noticed that the different predicted modes usually matched different speed profiles and tried to come up with a scheme to sample data to be representative of those different modes.</p>\n<p>Often, these experiments have ended up in the “didn’t work” pile but the insight the team drew from them was still crucial in deciding what to do next. If you try a unicycle model and results are very close to just using displacement, you can foresee that investing your time in more refined versions is probably not going to take your score further up.</p>\n<p>As a final note, understanding the task can help you when things behave unexpectedly. This is what must have happened to the <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199636\" target=\"_blank\">9th team</a> after a lot of time spent digging and visualising trajectories. They realised that the input and target spaces we provided were misaligned by a variable angle. This issue slipped through our internal tests and was potentially affecting the whole competition. While deep learning models were still able to learn in these conditions, correctly aligning the two spaces reduced the error in half across all the teams. To our surprise, even our baseline got a performance boost by almost 2X.</p>\n<h2>Train for what you measure</h2>\n<p>Several loss functions can be used to train the same model, leading incredibly different results. If you know the metrics that will be employed to score your model, it makes sense to design a loss which is as close as possible to that metric. This way, you will directly optimise for it and your model will learn how to produce good results according to the metric itself. In the case of this competition, the metric was an extension of the well known Negative Log Likelihood to the problem of multimodal trajectory prediction.</p>\n<p>The competition’s participants understood that, and converted our scoring metric into a differentiable function almost immediately at day 1. Almost all top scorers have employed it in their final solution.</p>\n<h2>Combine efforts</h2>\n<p>Apart from the loss function, several top scorers share another common feature. Instead of just using a single model, they combine an ensemble of them to collect more and better predictions. </p>\n<p>While a single model can capture several aspects of the world (e.g. semantic information) it may fall short on others (e.g. interactions with other vehicles). Combining different models (and also different input representations) it’s crucial to get better performance on the most difficult samples. Still, it’s important to choose a mixture of diverse and strong models before thinking about how to aggregate them. Ensembling can give you those crucial points you need to rank, but you won’t get there without an accurate selection of your models. <br>\nTo prove this, here are some examples for combinations of best single models and final ensembles results for some top scorers. Ensembling does improve the final score, but the real heavy-lifters are the base models inside of it. Interestingly, the ensemble boost is very similar across teams regardless of the technique employed.</p>\n<table>\n<thead>\n<tr>\n<th>Team</th>\n<th>Best single model score</th>\n<th>Ensemble score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1st</td>\n<td>9.776</td>\n<td>9.319</td>\n</tr>\n<tr>\n<td>2nd</td>\n<td>10.496</td>\n<td>10.216</td>\n</tr>\n<tr>\n<td>4th</td>\n<td>10.846</td>\n<td>10.272</td>\n</tr>\n<tr>\n<td>9th</td>\n<td>12.285</td>\n<td>11.788</td>\n</tr>\n</tbody>\n</table>\n<p>There are several ways of combining models into ensemble, here is a (non-exhaustive) list of what did work for kagglers:</p>\n<h3>Single head over features</h3>\n<p>Many teams (such as the 1st, 21st and 22nd) relied on concatenating features of the different models together.<br>\nAfter training individual models, the individual regression heads are removed. The intermediate features are then concatenated and a new head is trained to compress knowledge from the different models into a set of final predictions. The final ensemble is then fine-tuned over the validation set to avoid overfitting the train set. The beauty of this technique is that different models can work on different input representations (e.g. different raster resolutions or raster types), while the head can focus on combining high abstraction features. All teams noticed an improvement when moving from a single model to an ensemble of this type.</p>\n<p><a href=\"https://ibb.co/BPv9cf9\"><img src=\"https://i.ibb.co/cvRjkyj/ensemble.jpg\" alt=\"ensemble\"></a></p>\n<h3>Gaussian mixture models</h3>\n<p>As part of their solution, the <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/205376\" target=\"_blank\">3rd</a> and <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199657\" target=\"_blank\">4th team</a> introduced a loss which resembles a particular case of GMM. In this setting, the task is to identify a set of final trajectories which best represent the annotated one out of a given distribution. This framework has a nice probabilistic interpretation.</p>\n<h3>Models predictions average</h3>\n<p>Although many teams tried this technique without luck, others (such as the <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199588\" target=\"_blank\">6th team</a>) succeeded in combining the proposal trajectories from the models using a mean aggregator. </p>\n<p>While other teams tried to combine the trajectories based on distance, the key intuition of the team was to compute an average based on acceleration and distance covered by the different proposals instead of displacements.</p>\n<p>One of the possible reasons why simply averaging is not enough is that even though two solutions may be perfectly sensible, their average might not. As an example, imagine a situation where an agent can either go straight or perform a lane change. Although these two modes are both plausible, their average may lead the vehicle in between the two lanes, in an unrealistic situation.</p>\n<p><a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/ZV107Yq/lane-change.jpg\" alt=\"lane-change\"></a></p>\n<h2>Train, train, train!</h2>\n<p>It may sound trivial, but DL models truly get better with more data. And for this competition, large amounts of data were available from day 1. However, with great data come great responsibilities. <br>\nHow long should you train? Should you train on the full dataset or a split? Is data augmentation still crucial with so much data? Which input representation should you try first? When is a good time to reduce the lr? Which scheduler policy should you use?</p>\n<p>The top scorers have all found early answers to these questions, so they had more time to focus on designing the model and not on the training process. Finding the global optimum is not the key factor here, as it could take several weeks to fully evaluate all the combinations. What you really want is a local optimum that gives you a strong signal on which model choices work, so you can invest heavily in them before focusing on polishing the final training details.</p>\n<h3>Training time</h3>\n<p>Two train datasets were provided to participants. The first one is around 10GB of data, while the second one is almost 10 times more and therefore has way more variance in terms of agents and scenes. All top scorers found benefits in using the full train dataset over the small one when submitting their results. Also, several teams (like the 4th and 6th) succeeded in speeding up the training time by subsampling and caching parts of the dataset, which drastically reduced the time spent iterating.</p>\n<h3>Data augmentation</h3>\n<p>Interestingly, data augmentation didn’t seem to play a crucial role in squeezing out more performance. Among the top scorers, very few teams used any type of data augmentation. As an exception, team 4th employed a cutout augmentation on the input raster. This might suggest that the data available was already enough to provide a good generalisation.<br>\nInput resolution and format<br>\nOn the other hand, all the teams widely experimented with different raster resolutions (ranging from 128 in the ensemble of the 6th place to 448 from the 1st place). Also, many teams focused on the number of history frames to include in the input. Needless to say, all these solutions employ Convolutional neural networks as backbones for the models.</p>\n<h3>Learning rate scheduling</h3>\n<p>Another widely shared “trick” on the forum was the learning rate scheduling during training. Although many teams implemented their own custom scheduling, cosine anneal was by far the most employed one among the top scorers (both 1st and 3rd used it)</p>\n<h2>Think out of the box (but don’t lose the focus)</h2>\n<p>The baseline and the toolkit we provided were heavily focused on convolutional neural networks. As such, it’s not surprising that many participants chose to use CNNs to build their initial solutions. CNNs are great and we all love them, but there is a sea of DL opportunity out there to explore!<br>\nSome brave kagglers have experimented with completely different approaches with great results, and we report here some details about their solutions:</p>\n<h3>Attention</h3>\n<p>The 3rd team successfully employed a set transformer to combine proposals into 3 candidates. In this way, each trajectory can peek at the others and learn interesting relations. Finally, 3 random vectors are used as seed to learn a weighted synthesis of the proposals and output the final candidate to be scored. <br>\nSimilarly, the 2nd team ensembled trajectories from different models using a matrix of distances between candidates as features for a set of fully connected layers. The layers’ weights can learn relations between different proposals which are then used to estimate how to combine them to obtain the final set of trajectories.<br>\nAttention (and in general transformers) is attracting a huge attention from the community as it can replace unlearned aggregators (max or avg pooling) and even convolution (see for example <a href=\"http://jalammar.github.io/illustrated-transformer/\" target=\"_blank\">here</a> for more details on the transformer architecture) </p>\n<h3>RNN</h3>\n<p>Not many teams employed RNN successfully, especially among the top scorers. As such, we were pretty surprised when we found out that the 2nd team was able to include them in their final architecture and that they were seeing noticeable benefit from using them. In their solution, features extracted from a CNN backbone are processed by a long and a short LSTM layer together with temporal information. These features are then recombined with the original ones to compute the trajectories proposal. By limiting the recurrent part of the architecture to only intermediate features, the team was able to train efficiently and avoid exploding memory allocation.</p>\n<h3>Vector representation</h3>\n<p>Rasterization is a blunt tool, and can drastically reduce the resolution of the scene around the agent. Luckily, other representations which overcome this issue exist. Two teams drew inspiration from the latest advances in self-driving vehicles and completely discarded raster-based representations in favour of vector-based ones.</p>\n<p>The <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199494\" target=\"_blank\">22nd team</a> employed <a href=\"https://github.com/charlesq34/pointnet\" target=\"_blank\">PointNet</a> (a famous architecture for point clouds detection) to process point representation of the scene (including the agent of interest and other agents around it). To our surprise, this solution is completely lane blinded, meaning it can’t rely on any static semantic information.</p>\n<p>The <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199711\" target=\"_blank\">10th team</a> chose instead <a href=\"https://blog.waymo.com/2020/05/vectornet.html\" target=\"_blank\">VectorNet</a>, which has recently been proposed for the motion prediction task. For this solution lanes and other semantic information were encoded as vectors and forwarded together with other agents. lanes and other semantic information, and is significantly faster than raster-based ones both during training and inference.</p>\n<p>However, when time is constrained, failing fast is better than chasing a rabbit down the hole for weeks. As the competition proceeded, some approaches have been deemed not worthy of further exploration, as multiple teams failed to exploit them in any meaningful way. Some of these includes:</p>\n<h3>GANs</h3>\n<p>Generative Adversarial Networks have been tried again and again by many teams. Sadly, they have all ended up adding them to the pile of “didn’t work” approaches.</p>\n<h3>3D-CNN</h3>\n<p>Some teams tried 3D convolutions, but quickly discarded them for their lack of speed and improvement over 2D solutions.</p>\n<h3>Non DL approaches</h3>\n<p>Many teams tried traditional ML and control approaches like Kalman filtering and clustering. However, considering all top solutions include some DL networks those approaches have not proven strong enough. </p>\n<h1>Conclusions</h1>\n<p>We’ve seen many great examples on how to build a winning solution for this year’s competition. We really hope all the participants enjoyed taking part in it. As hosts, we are enthusiastic about how much traction the competition got and we thoroughly enjoyed watching participants using all their tools and knowledge to climb the ladder to those top sweet places. Hopefully, after reading this post you will know a little more about how to get there in future competitions.</p>\n<p>Almost forgot, we’re <a href=\"https://www.lyft.com/careers\" target=\"_blank\">hiring</a> in the US and in Europe :)</p>",
  "messages": [
    {
      "id": 1209418,
      "postDate": "2021-02-18T22:03:30.270Z",
      "content": "<h1>Introduction</h1>\n<p>The Lyft 2020 competition (hosted on Kaggle) attracted more than 900 teams with a total of 14.904 submissions. After a few months since its completion, it’s time to draw some conclusions on what kagglers tried, what worked and what didn’t. The community wrote 320 posts where they shared ideas, results and suggestions. We’ve collected and filtered them to identify 5 insights that can help you hit the ground running in this challenging task. Let’s start with a short recap of what participants were asked to do and what tools Lyft provided them to start coding straight away.</p>\n<h1>The task</h1>\n<p>Motion prediction is an essential component of the autonomous stack. The task is to understand what other traffic participants will do in the short future. This can be applied to other vehicles around the AV but also vulnerable road users such as pedestrians and cyclists. For each of them, we want to predict the future location as a trajectory of 2D points. This information is crucial for the final planning of the AV trajectory, both for rule-based and deep learning based systems.</p>\n<p>However, the future is highly uncertain and a single trajectory per agent may not be enough. Even for a vehicle driving on a straight road, many different speed profiles can be chosen. All these trajectories may be in a sense “correct” (kinematically feasible and safe to take) but, in practice, only one of them will be picked when we will deploy this model. Still, it’s important to keep the possibility to generate multiple trajectories, as other models such as the planner may require them. For this reason, we highly encouraged participants to submit multiple trajectories per agent by providing them with a metric that can take into account this explicitly.</p>\n<p><a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/Gc3PtRY/Prediction-task.jpg\" alt=\"Prediction-task\"></a></p>\n<p>Prediction is a very challenging task! It may sound easy at first when you have only a bunch of agents in the scene but, as you add more of them to the mix, your model must have a “global” vision of what’s going on to correctly predict the future of them.</p>\n<h1>The data</h1>\n<p>Challenging tasks require great datasets to be solved. We provided participants with one collected by the Lyft fleet, which includes more than 1000 hours of driving and HD quality semantic map annotations. </p>\n<p><a href=\"https://ibb.co/chm2p30\"><img src=\"https://i.ibb.co/gZntX9N/Screenshot-2021-01-15-at-11-53-21.png\" alt=\"Screenshot-2021-01-15-at-11-53-21\"></a></p>\n<p>For each frame in the dataset multiple agents are available, and their annotated trajectory can be read from future frames. These annotations come from the latest version of our in-house perception and tracking system. We also provide static road geometry information, including lanes and crosswalks, as well as dynamic information such as traffic lights. This dataset is public and already available to download. Researchers can use it with no restrictions, and it’s always possible to submit results to the leaderboard and get a score, even after the end of the competition.</p>\n<h1>The tools</h1>\n<p>The motion prediction task was new to the Kaggle community, and working with such a huge dataset can look daunting at first. To encourage people from different backgrounds to take part in this challenge we decided to release <a href=\"https://github.com/lyft/l5kit\" target=\"_blank\">L5Kit</a>, a library of functionalities to work with our dataset and extract useful information. This is the very same tool we use for our internal projects, and we provide support for it as well as integrate precious feedback from the community. L5Kit provided utility functions to read and visualise data, as well as metrics to score solutions. The toolkit is particularly useful for training CNNs, as it can generate rasterized Bird-Eye-View representation from the data. These rasters can include map semantic information as well as satellite imagery.</p>\n<p><a href=\"https://ibb.co/zZ4NW5B\"><img src=\"https://i.ibb.co/qprxLnc/examples-dataset.jpg\" alt=\"examples-dataset\"></a></p>\n<p>For this year competition we also included a kernel with a CNN baseline for the task participants could use and build upon. The kernel was intended as a starting point for the competition, as the task was new, the format of the dataset unfamiliar and it was unclear how to start with it. More than 60 teams started their experiments from our kernel, and many more leveraged our pre-trained network in their solutions.</p>\n<h1>What it takes to win a competition</h1>\n<p>After reading many posts and interviewing the top scorers, we came up with a list of insights which you definitely should know when taking part into a competition about motion prediction. We report them here in no particular order.</p>\n<h2>Understand the task</h2>\n<p>Yes, it may sound trivial but you really need to invest some time to understand which world’s entities you are trying to model. Deep learning is powerful and can learn almost anything (provided the input representation is meaningful) but knowing what you are modeling can drastically change your approach. For the task of prediction, knowing that your final targets are trajectories which must be executed by an agent can give you an edge.</p>\n<p>As an example, the <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/201493\" target=\"_blank\">1st team</a> experimented with predicting only differences from a constant velocity model. Other teams tried to predict acceleration and steering instead of displacements and used a unicycle kinematic model to retrieve the trajectories, while others tried to smooth them using a Kalman filter.</p>\n<p>The 3rd team realised that the traffic light information in the rasteriser was static, and that the model might benefit from knowing the evolution of this signal through time.</p>\n<p>Again, the 6th team investigated a lot how to balance data based on the acceleration of the agents in the scene. They noticed that the different predicted modes usually matched different speed profiles and tried to come up with a scheme to sample data to be representative of those different modes.</p>\n<p>Often, these experiments have ended up in the “didn’t work” pile but the insight the team drew from them was still crucial in deciding what to do next. If you try a unicycle model and results are very close to just using displacement, you can foresee that investing your time in more refined versions is probably not going to take your score further up.</p>\n<p>As a final note, understanding the task can help you when things behave unexpectedly. This is what must have happened to the <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199636\" target=\"_blank\">9th team</a> after a lot of time spent digging and visualising trajectories. They realised that the input and target spaces we provided were misaligned by a variable angle. This issue slipped through our internal tests and was potentially affecting the whole competition. While deep learning models were still able to learn in these conditions, correctly aligning the two spaces reduced the error in half across all the teams. To our surprise, even our baseline got a performance boost by almost 2X.</p>\n<h2>Train for what you measure</h2>\n<p>Several loss functions can be used to train the same model, leading incredibly different results. If you know the metrics that will be employed to score your model, it makes sense to design a loss which is as close as possible to that metric. This way, you will directly optimise for it and your model will learn how to produce good results according to the metric itself. In the case of this competition, the metric was an extension of the well known Negative Log Likelihood to the problem of multimodal trajectory prediction.</p>\n<p>The competition’s participants understood that, and converted our scoring metric into a differentiable function almost immediately at day 1. Almost all top scorers have employed it in their final solution.</p>\n<h2>Combine efforts</h2>\n<p>Apart from the loss function, several top scorers share another common feature. Instead of just using a single model, they combine an ensemble of them to collect more and better predictions. </p>\n<p>While a single model can capture several aspects of the world (e.g. semantic information) it may fall short on others (e.g. interactions with other vehicles). Combining different models (and also different input representations) it’s crucial to get better performance on the most difficult samples. Still, it’s important to choose a mixture of diverse and strong models before thinking about how to aggregate them. Ensembling can give you those crucial points you need to rank, but you won’t get there without an accurate selection of your models. <br>\nTo prove this, here are some examples for combinations of best single models and final ensembles results for some top scorers. Ensembling does improve the final score, but the real heavy-lifters are the base models inside of it. Interestingly, the ensemble boost is very similar across teams regardless of the technique employed.</p>\n<table>\n<thead>\n<tr>\n<th>Team</th>\n<th>Best single model score</th>\n<th>Ensemble score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1st</td>\n<td>9.776</td>\n<td>9.319</td>\n</tr>\n<tr>\n<td>2nd</td>\n<td>10.496</td>\n<td>10.216</td>\n</tr>\n<tr>\n<td>4th</td>\n<td>10.846</td>\n<td>10.272</td>\n</tr>\n<tr>\n<td>9th</td>\n<td>12.285</td>\n<td>11.788</td>\n</tr>\n</tbody>\n</table>\n<p>There are several ways of combining models into ensemble, here is a (non-exhaustive) list of what did work for kagglers:</p>\n<h3>Single head over features</h3>\n<p>Many teams (such as the 1st, 21st and 22nd) relied on concatenating features of the different models together.<br>\nAfter training individual models, the individual regression heads are removed. The intermediate features are then concatenated and a new head is trained to compress knowledge from the different models into a set of final predictions. The final ensemble is then fine-tuned over the validation set to avoid overfitting the train set. The beauty of this technique is that different models can work on different input representations (e.g. different raster resolutions or raster types), while the head can focus on combining high abstraction features. All teams noticed an improvement when moving from a single model to an ensemble of this type.</p>\n<p><a href=\"https://ibb.co/BPv9cf9\"><img src=\"https://i.ibb.co/cvRjkyj/ensemble.jpg\" alt=\"ensemble\"></a></p>\n<h3>Gaussian mixture models</h3>\n<p>As part of their solution, the <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/205376\" target=\"_blank\">3rd</a> and <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199657\" target=\"_blank\">4th team</a> introduced a loss which resembles a particular case of GMM. In this setting, the task is to identify a set of final trajectories which best represent the annotated one out of a given distribution. This framework has a nice probabilistic interpretation.</p>\n<h3>Models predictions average</h3>\n<p>Although many teams tried this technique without luck, others (such as the <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199588\" target=\"_blank\">6th team</a>) succeeded in combining the proposal trajectories from the models using a mean aggregator. </p>\n<p>While other teams tried to combine the trajectories based on distance, the key intuition of the team was to compute an average based on acceleration and distance covered by the different proposals instead of displacements.</p>\n<p>One of the possible reasons why simply averaging is not enough is that even though two solutions may be perfectly sensible, their average might not. As an example, imagine a situation where an agent can either go straight or perform a lane change. Although these two modes are both plausible, their average may lead the vehicle in between the two lanes, in an unrealistic situation.</p>\n<p><a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/ZV107Yq/lane-change.jpg\" alt=\"lane-change\"></a></p>\n<h2>Train, train, train!</h2>\n<p>It may sound trivial, but DL models truly get better with more data. And for this competition, large amounts of data were available from day 1. However, with great data come great responsibilities. <br>\nHow long should you train? Should you train on the full dataset or a split? Is data augmentation still crucial with so much data? Which input representation should you try first? When is a good time to reduce the lr? Which scheduler policy should you use?</p>\n<p>The top scorers have all found early answers to these questions, so they had more time to focus on designing the model and not on the training process. Finding the global optimum is not the key factor here, as it could take several weeks to fully evaluate all the combinations. What you really want is a local optimum that gives you a strong signal on which model choices work, so you can invest heavily in them before focusing on polishing the final training details.</p>\n<h3>Training time</h3>\n<p>Two train datasets were provided to participants. The first one is around 10GB of data, while the second one is almost 10 times more and therefore has way more variance in terms of agents and scenes. All top scorers found benefits in using the full train dataset over the small one when submitting their results. Also, several teams (like the 4th and 6th) succeeded in speeding up the training time by subsampling and caching parts of the dataset, which drastically reduced the time spent iterating.</p>\n<h3>Data augmentation</h3>\n<p>Interestingly, data augmentation didn’t seem to play a crucial role in squeezing out more performance. Among the top scorers, very few teams used any type of data augmentation. As an exception, team 4th employed a cutout augmentation on the input raster. This might suggest that the data available was already enough to provide a good generalisation.<br>\nInput resolution and format<br>\nOn the other hand, all the teams widely experimented with different raster resolutions (ranging from 128 in the ensemble of the 6th place to 448 from the 1st place). Also, many teams focused on the number of history frames to include in the input. Needless to say, all these solutions employ Convolutional neural networks as backbones for the models.</p>\n<h3>Learning rate scheduling</h3>\n<p>Another widely shared “trick” on the forum was the learning rate scheduling during training. Although many teams implemented their own custom scheduling, cosine anneal was by far the most employed one among the top scorers (both 1st and 3rd used it)</p>\n<h2>Think out of the box (but don’t lose the focus)</h2>\n<p>The baseline and the toolkit we provided were heavily focused on convolutional neural networks. As such, it’s not surprising that many participants chose to use CNNs to build their initial solutions. CNNs are great and we all love them, but there is a sea of DL opportunity out there to explore!<br>\nSome brave kagglers have experimented with completely different approaches with great results, and we report here some details about their solutions:</p>\n<h3>Attention</h3>\n<p>The 3rd team successfully employed a set transformer to combine proposals into 3 candidates. In this way, each trajectory can peek at the others and learn interesting relations. Finally, 3 random vectors are used as seed to learn a weighted synthesis of the proposals and output the final candidate to be scored. <br>\nSimilarly, the 2nd team ensembled trajectories from different models using a matrix of distances between candidates as features for a set of fully connected layers. The layers’ weights can learn relations between different proposals which are then used to estimate how to combine them to obtain the final set of trajectories.<br>\nAttention (and in general transformers) is attracting a huge attention from the community as it can replace unlearned aggregators (max or avg pooling) and even convolution (see for example <a href=\"http://jalammar.github.io/illustrated-transformer/\" target=\"_blank\">here</a> for more details on the transformer architecture) </p>\n<h3>RNN</h3>\n<p>Not many teams employed RNN successfully, especially among the top scorers. As such, we were pretty surprised when we found out that the 2nd team was able to include them in their final architecture and that they were seeing noticeable benefit from using them. In their solution, features extracted from a CNN backbone are processed by a long and a short LSTM layer together with temporal information. These features are then recombined with the original ones to compute the trajectories proposal. By limiting the recurrent part of the architecture to only intermediate features, the team was able to train efficiently and avoid exploding memory allocation.</p>\n<h3>Vector representation</h3>\n<p>Rasterization is a blunt tool, and can drastically reduce the resolution of the scene around the agent. Luckily, other representations which overcome this issue exist. Two teams drew inspiration from the latest advances in self-driving vehicles and completely discarded raster-based representations in favour of vector-based ones.</p>\n<p>The <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199494\" target=\"_blank\">22nd team</a> employed <a href=\"https://github.com/charlesq34/pointnet\" target=\"_blank\">PointNet</a> (a famous architecture for point clouds detection) to process point representation of the scene (including the agent of interest and other agents around it). To our surprise, this solution is completely lane blinded, meaning it can’t rely on any static semantic information.</p>\n<p>The <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199711\" target=\"_blank\">10th team</a> chose instead <a href=\"https://blog.waymo.com/2020/05/vectornet.html\" target=\"_blank\">VectorNet</a>, which has recently been proposed for the motion prediction task. For this solution lanes and other semantic information were encoded as vectors and forwarded together with other agents. lanes and other semantic information, and is significantly faster than raster-based ones both during training and inference.</p>\n<p>However, when time is constrained, failing fast is better than chasing a rabbit down the hole for weeks. As the competition proceeded, some approaches have been deemed not worthy of further exploration, as multiple teams failed to exploit them in any meaningful way. Some of these includes:</p>\n<h3>GANs</h3>\n<p>Generative Adversarial Networks have been tried again and again by many teams. Sadly, they have all ended up adding them to the pile of “didn’t work” approaches.</p>\n<h3>3D-CNN</h3>\n<p>Some teams tried 3D convolutions, but quickly discarded them for their lack of speed and improvement over 2D solutions.</p>\n<h3>Non DL approaches</h3>\n<p>Many teams tried traditional ML and control approaches like Kalman filtering and clustering. However, considering all top solutions include some DL networks those approaches have not proven strong enough. </p>\n<h1>Conclusions</h1>\n<p>We’ve seen many great examples on how to build a winning solution for this year’s competition. We really hope all the participants enjoyed taking part in it. As hosts, we are enthusiastic about how much traction the competition got and we thoroughly enjoyed watching participants using all their tools and knowledge to climb the ladder to those top sweet places. Hopefully, after reading this post you will know a little more about how to get there in future competitions.</p>\n<p>Almost forgot, we’re <a href=\"https://www.lyft.com/careers\" target=\"_blank\">hiring</a> in the US and in Europe :)</p>",
      "rawMarkdown": "#Introduction\nThe Lyft 2020 competition (hosted on Kaggle) attracted more than 900 teams with a total of 14.904 submissions. After a few months since its completion, it’s time to draw some conclusions on what kagglers tried, what worked and what didn’t. The community wrote 320 posts where they shared ideas, results and suggestions. We’ve collected and filtered them to identify 5 insights that can help you hit the ground running in this challenging task. Let’s start with a short recap of what participants were asked to do and what tools Lyft provided them to start coding straight away.\n\n#The task\nMotion prediction is an essential component of the autonomous stack. The task is to understand what other traffic participants will do in the short future. This can be applied to other vehicles around the AV but also vulnerable road users such as pedestrians and cyclists. For each of them, we want to predict the future location as a trajectory of 2D points. This information is crucial for the final planning of the AV trajectory, both for rule-based and deep learning based systems.\n\nHowever, the future is highly uncertain and a single trajectory per agent may not be enough. Even for a vehicle driving on a straight road, many different speed profiles can be chosen. All these trajectories may be in a sense “correct” (kinematically feasible and safe to take) but, in practice, only one of them will be picked when we will deploy this model. Still, it’s important to keep the possibility to generate multiple trajectories, as other models such as the planner may require them. For this reason, we highly encouraged participants to submit multiple trajectories per agent by providing them with a metric that can take into account this explicitly.\n\n<a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/Gc3PtRY/Prediction-task.jpg\" alt=\"Prediction-task\" border=\"0\"></a>\n\nPrediction is a very challenging task! It may sound easy at first when you have only a bunch of agents in the scene but, as you add more of them to the mix, your model must have a “global” vision of what’s going on to correctly predict the future of them.\n#The data\nChallenging tasks require great datasets to be solved. We provided participants with one collected by the Lyft fleet, which includes more than 1000 hours of driving and HD quality semantic map annotations. \n\n<a href=\"https://ibb.co/chm2p30\"><img src=\"https://i.ibb.co/gZntX9N/Screenshot-2021-01-15-at-11-53-21.png\" alt=\"Screenshot-2021-01-15-at-11-53-21\" border=\"0\"></a>\n\nFor each frame in the dataset multiple agents are available, and their annotated trajectory can be read from future frames. These annotations come from the latest version of our in-house perception and tracking system. We also provide static road geometry information, including lanes and crosswalks, as well as dynamic information such as traffic lights. This dataset is public and already available to download. Researchers can use it with no restrictions, and it’s always possible to submit results to the leaderboard and get a score, even after the end of the competition.\n\n#The tools\nThe motion prediction task was new to the Kaggle community, and working with such a huge dataset can look daunting at first. To encourage people from different backgrounds to take part in this challenge we decided to release [L5Kit](https://github.com/lyft/l5kit), a library of functionalities to work with our dataset and extract useful information. This is the very same tool we use for our internal projects, and we provide support for it as well as integrate precious feedback from the community. L5Kit provided utility functions to read and visualise data, as well as metrics to score solutions. The toolkit is particularly useful for training CNNs, as it can generate rasterized Bird-Eye-View representation from the data. These rasters can include map semantic information as well as satellite imagery.\n\n<a href=\"https://ibb.co/zZ4NW5B\"><img src=\"https://i.ibb.co/qprxLnc/examples-dataset.jpg\" alt=\"examples-dataset\" border=\"0\"></a>\n\nFor this year competition we also included a kernel with a CNN baseline for the task participants could use and build upon. The kernel was intended as a starting point for the competition, as the task was new, the format of the dataset unfamiliar and it was unclear how to start with it. More than 60 teams started their experiments from our kernel, and many more leveraged our pre-trained network in their solutions.\n\n#What it takes to win a competition\nAfter reading many posts and interviewing the top scorers, we came up with a list of insights which you definitely should know when taking part into a competition about motion prediction. We report them here in no particular order.\n\n##Understand the task\nYes, it may sound trivial but you really need to invest some time to understand which world’s entities you are trying to model. Deep learning is powerful and can learn almost anything (provided the input representation is meaningful) but knowing what you are modeling can drastically change your approach. For the task of prediction, knowing that your final targets are trajectories which must be executed by an agent can give you an edge.\n\nAs an example, the [1st team]( https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/201493) experimented with predicting only differences from a constant velocity model. Other teams tried to predict acceleration and steering instead of displacements and used a unicycle kinematic model to retrieve the trajectories, while others tried to smooth them using a Kalman filter.\n\nThe 3rd team realised that the traffic light information in the rasteriser was static, and that the model might benefit from knowing the evolution of this signal through time.\n\nAgain, the 6th team investigated a lot how to balance data based on the acceleration of the agents in the scene. They noticed that the different predicted modes usually matched different speed profiles and tried to come up with a scheme to sample data to be representative of those different modes.\n\nOften, these experiments have ended up in the “didn’t work” pile but the insight the team drew from them was still crucial in deciding what to do next. If you try a unicycle model and results are very close to just using displacement, you can foresee that investing your time in more refined versions is probably not going to take your score further up.\n\nAs a final note, understanding the task can help you when things behave unexpectedly. This is what must have happened to the [9th team](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199636) after a lot of time spent digging and visualising trajectories. They realised that the input and target spaces we provided were misaligned by a variable angle. This issue slipped through our internal tests and was potentially affecting the whole competition. While deep learning models were still able to learn in these conditions, correctly aligning the two spaces reduced the error in half across all the teams. To our surprise, even our baseline got a performance boost by almost 2X.\n\n##Train for what you measure\nSeveral loss functions can be used to train the same model, leading incredibly different results. If you know the metrics that will be employed to score your model, it makes sense to design a loss which is as close as possible to that metric. This way, you will directly optimise for it and your model will learn how to produce good results according to the metric itself. In the case of this competition, the metric was an extension of the well known Negative Log Likelihood to the problem of multimodal trajectory prediction.\n\nThe competition’s participants understood that, and converted our scoring metric into a differentiable function almost immediately at day 1. Almost all top scorers have employed it in their final solution.\n\n##Combine efforts\nApart from the loss function, several top scorers share another common feature. Instead of just using a single model, they combine an ensemble of them to collect more and better predictions. \n\nWhile a single model can capture several aspects of the world (e.g. semantic information) it may fall short on others (e.g. interactions with other vehicles). Combining different models (and also different input representations) it’s crucial to get better performance on the most difficult samples. Still, it’s important to choose a mixture of diverse and strong models before thinking about how to aggregate them. Ensembling can give you those crucial points you need to rank, but you won’t get there without an accurate selection of your models. \nTo prove this, here are some examples for combinations of best single models and final ensembles results for some top scorers. Ensembling does improve the final score, but the real heavy-lifters are the base models inside of it. Interestingly, the ensemble boost is very similar across teams regardless of the technique employed.\n\n| Team | Best single model score | Ensemble score |\n|------|-------------------------|----------------|\n| 1st  | 9.776                   | 9.319          |\n| 2nd  | 10.496                  | 10.216         |\n| 4th  | 10.846                  | 10.272         |\n| 9th  | 12.285                  | 11.788         |\n\n\nThere are several ways of combining models into ensemble, here is a (non-exhaustive) list of what did work for kagglers:\n\n###Single head over features\nMany teams (such as the 1st, 21st and 22nd) relied on concatenating features of the different models together.\nAfter training individual models, the individual regression heads are removed. The intermediate features are then concatenated and a new head is trained to compress knowledge from the different models into a set of final predictions. The final ensemble is then fine-tuned over the validation set to avoid overfitting the train set. The beauty of this technique is that different models can work on different input representations (e.g. different raster resolutions or raster types), while the head can focus on combining high abstraction features. All teams noticed an improvement when moving from a single model to an ensemble of this type.\n\n<a href=\"https://ibb.co/BPv9cf9\"><img src=\"https://i.ibb.co/cvRjkyj/ensemble.jpg\" alt=\"ensemble\" border=\"0\"></a>\n\n###Gaussian mixture models\nAs part of their solution, the [3rd](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/205376) and [4th team](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199657) introduced a loss which resembles a particular case of GMM. In this setting, the task is to identify a set of final trajectories which best represent the annotated one out of a given distribution. This framework has a nice probabilistic interpretation.\n###Models predictions average\nAlthough many teams tried this technique without luck, others (such as the [6th team](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199588)) succeeded in combining the proposal trajectories from the models using a mean aggregator. \n\nWhile other teams tried to combine the trajectories based on distance, the key intuition of the team was to compute an average based on acceleration and distance covered by the different proposals instead of displacements.\n\nOne of the possible reasons why simply averaging is not enough is that even though two solutions may be perfectly sensible, their average might not. As an example, imagine a situation where an agent can either go straight or perform a lane change. Although these two modes are both plausible, their average may lead the vehicle in between the two lanes, in an unrealistic situation.\n\n<a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/ZV107Yq/lane-change.jpg\" alt=\"lane-change\" border=\"0\"></a>\n\n##Train, train, train!\nIt may sound trivial, but DL models truly get better with more data. And for this competition, large amounts of data were available from day 1. However, with great data come great responsibilities. \nHow long should you train? Should you train on the full dataset or a split? Is data augmentation still crucial with so much data? Which input representation should you try first? When is a good time to reduce the lr? Which scheduler policy should you use?\n\nThe top scorers have all found early answers to these questions, so they had more time to focus on designing the model and not on the training process. Finding the global optimum is not the key factor here, as it could take several weeks to fully evaluate all the combinations. What you really want is a local optimum that gives you a strong signal on which model choices work, so you can invest heavily in them before focusing on polishing the final training details.\n\n###Training time\nTwo train datasets were provided to participants. The first one is around 10GB of data, while the second one is almost 10 times more and therefore has way more variance in terms of agents and scenes. All top scorers found benefits in using the full train dataset over the small one when submitting their results. Also, several teams (like the 4th and 6th) succeeded in speeding up the training time by subsampling and caching parts of the dataset, which drastically reduced the time spent iterating.\n###Data augmentation\nInterestingly, data augmentation didn’t seem to play a crucial role in squeezing out more performance. Among the top scorers, very few teams used any type of data augmentation. As an exception, team 4th employed a cutout augmentation on the input raster. This might suggest that the data available was already enough to provide a good generalisation.\nInput resolution and format\nOn the other hand, all the teams widely experimented with different raster resolutions (ranging from 128 in the ensemble of the 6th place to 448 from the 1st place). Also, many teams focused on the number of history frames to include in the input. Needless to say, all these solutions employ Convolutional neural networks as backbones for the models.\n###Learning rate scheduling\nAnother widely shared “trick” on the forum was the learning rate scheduling during training. Although many teams implemented their own custom scheduling, cosine anneal was by far the most employed one among the top scorers (both 1st and 3rd used it)\n\n##Think out of the box (but don’t lose the focus)\nThe baseline and the toolkit we provided were heavily focused on convolutional neural networks. As such, it’s not surprising that many participants chose to use CNNs to build their initial solutions. CNNs are great and we all love them, but there is a sea of DL opportunity out there to explore!\nSome brave kagglers have experimented with completely different approaches with great results, and we report here some details about their solutions:\n\n###Attention\nThe 3rd team successfully employed a set transformer to combine proposals into 3 candidates. In this way, each trajectory can peek at the others and learn interesting relations. Finally, 3 random vectors are used as seed to learn a weighted synthesis of the proposals and output the final candidate to be scored. \nSimilarly, the 2nd team ensembled trajectories from different models using a matrix of distances between candidates as features for a set of fully connected layers. The layers’ weights can learn relations between different proposals which are then used to estimate how to combine them to obtain the final set of trajectories.\nAttention (and in general transformers) is attracting a huge attention from the community as it can replace unlearned aggregators (max or avg pooling) and even convolution (see for example [here](http://jalammar.github.io/illustrated-transformer/) for more details on the transformer architecture) \n###RNN\nNot many teams employed RNN successfully, especially among the top scorers. As such, we were pretty surprised when we found out that the 2nd team was able to include them in their final architecture and that they were seeing noticeable benefit from using them. In their solution, features extracted from a CNN backbone are processed by a long and a short LSTM layer together with temporal information. These features are then recombined with the original ones to compute the trajectories proposal. By limiting the recurrent part of the architecture to only intermediate features, the team was able to train efficiently and avoid exploding memory allocation.\n###Vector representation\nRasterization is a blunt tool, and can drastically reduce the resolution of the scene around the agent. Luckily, other representations which overcome this issue exist. Two teams drew inspiration from the latest advances in self-driving vehicles and completely discarded raster-based representations in favour of vector-based ones.\n\nThe [22nd team](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199494) employed [PointNet](https://github.com/charlesq34/pointnet) (a famous architecture for point clouds detection) to process point representation of the scene (including the agent of interest and other agents around it). To our surprise, this solution is completely lane blinded, meaning it can’t rely on any static semantic information.\n\nThe [10th team](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199711) chose instead [VectorNet](https://blog.waymo.com/2020/05/vectornet.html), which has recently been proposed for the motion prediction task. For this solution lanes and other semantic information were encoded as vectors and forwarded together with other agents. lanes and other semantic information, and is significantly faster than raster-based ones both during training and inference.\n\nHowever, when time is constrained, failing fast is better than chasing a rabbit down the hole for weeks. As the competition proceeded, some approaches have been deemed not worthy of further exploration, as multiple teams failed to exploit them in any meaningful way. Some of these includes:\n\n###GANs\nGenerative Adversarial Networks have been tried again and again by many teams. Sadly, they have all ended up adding them to the pile of “didn’t work” approaches.\n###3D-CNN\nSome teams tried 3D convolutions, but quickly discarded them for their lack of speed and improvement over 2D solutions.\n###Non DL approaches\nMany teams tried traditional ML and control approaches like Kalman filtering and clustering. However, considering all top solutions include some DL networks those approaches have not proven strong enough. \n#Conclusions\nWe’ve seen many great examples on how to build a winning solution for this year’s competition. We really hope all the participants enjoyed taking part in it. As hosts, we are enthusiastic about how much traction the competition got and we thoroughly enjoyed watching participants using all their tools and knowledge to climb the ladder to those top sweet places. Hopefully, after reading this post you will know a little more about how to get there in future competitions.\n\nAlmost forgot, we’re [hiring](https://www.lyft.com/careers) in the US and in Europe :)",
      "votes": 8
    },
    {
      "id": 1219012,
      "postDate": "2021-02-26T11:19:27.823Z",
      "content": "<p>Thanks for summarizing the whole project in one discussion forum. Really Helpful.</p>",
      "rawMarkdown": "Thanks for summarizing the whole project in one discussion forum. Really Helpful."
    },
    {
      "id": 1209419,
      "postDate": "2021-02-18T22:03:45.883Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1219012,
      "author_name": "Ramji",
      "author_url": "",
      "post_date": "2021-02-26T11:19:27.823000",
      "content": "<p>Thanks for summarizing the whole project in one discussion forum. Really Helpful.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1209419,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-02-18T22:03:45.883000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1209418": "#Introduction\nThe Lyft 2020 competition (hosted on Kaggle) attracted more than 900 teams with a total of 14.904 submissions. After a few months since its completion, it’s time to draw some conclusions on what kagglers tried, what worked and what didn’t. The community wrote 320 posts where they shared ideas, results and suggestions. We’ve collected and filtered them to identify 5 insights that can help you hit the ground running in this challenging task. Let’s start with a short recap of what participants were asked to do and what tools Lyft provided them to start coding straight away.\n\n#The task\nMotion prediction is an essential component of the autonomous stack. The task is to understand what other traffic participants will do in the short future. This can be applied to other vehicles around the AV but also vulnerable road users such as pedestrians and cyclists. For each of them, we want to predict the future location as a trajectory of 2D points. This information is crucial for the final planning of the AV trajectory, both for rule-based and deep learning based systems.\n\nHowever, the future is highly uncertain and a single trajectory per agent may not be enough. Even for a vehicle driving on a straight road, many different speed profiles can be chosen. All these trajectories may be in a sense “correct” (kinematically feasible and safe to take) but, in practice, only one of them will be picked when we will deploy this model. Still, it’s important to keep the possibility to generate multiple trajectories, as other models such as the planner may require them. For this reason, we highly encouraged participants to submit multiple trajectories per agent by providing them with a metric that can take into account this explicitly.\n\n<a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/Gc3PtRY/Prediction-task.jpg\" alt=\"Prediction-task\" border=\"0\"></a>\n\nPrediction is a very challenging task! It may sound easy at first when you have only a bunch of agents in the scene but, as you add more of them to the mix, your model must have a “global” vision of what’s going on to correctly predict the future of them.\n#The data\nChallenging tasks require great datasets to be solved. We provided participants with one collected by the Lyft fleet, which includes more than 1000 hours of driving and HD quality semantic map annotations. \n\n<a href=\"https://ibb.co/chm2p30\"><img src=\"https://i.ibb.co/gZntX9N/Screenshot-2021-01-15-at-11-53-21.png\" alt=\"Screenshot-2021-01-15-at-11-53-21\" border=\"0\"></a>\n\nFor each frame in the dataset multiple agents are available, and their annotated trajectory can be read from future frames. These annotations come from the latest version of our in-house perception and tracking system. We also provide static road geometry information, including lanes and crosswalks, as well as dynamic information such as traffic lights. This dataset is public and already available to download. Researchers can use it with no restrictions, and it’s always possible to submit results to the leaderboard and get a score, even after the end of the competition.\n\n#The tools\nThe motion prediction task was new to the Kaggle community, and working with such a huge dataset can look daunting at first. To encourage people from different backgrounds to take part in this challenge we decided to release [L5Kit](https://github.com/lyft/l5kit), a library of functionalities to work with our dataset and extract useful information. This is the very same tool we use for our internal projects, and we provide support for it as well as integrate precious feedback from the community. L5Kit provided utility functions to read and visualise data, as well as metrics to score solutions. The toolkit is particularly useful for training CNNs, as it can generate rasterized Bird-Eye-View representation from the data. These rasters can include map semantic information as well as satellite imagery.\n\n<a href=\"https://ibb.co/zZ4NW5B\"><img src=\"https://i.ibb.co/qprxLnc/examples-dataset.jpg\" alt=\"examples-dataset\" border=\"0\"></a>\n\nFor this year competition we also included a kernel with a CNN baseline for the task participants could use and build upon. The kernel was intended as a starting point for the competition, as the task was new, the format of the dataset unfamiliar and it was unclear how to start with it. More than 60 teams started their experiments from our kernel, and many more leveraged our pre-trained network in their solutions.\n\n#What it takes to win a competition\nAfter reading many posts and interviewing the top scorers, we came up with a list of insights which you definitely should know when taking part into a competition about motion prediction. We report them here in no particular order.\n\n##Understand the task\nYes, it may sound trivial but you really need to invest some time to understand which world’s entities you are trying to model. Deep learning is powerful and can learn almost anything (provided the input representation is meaningful) but knowing what you are modeling can drastically change your approach. For the task of prediction, knowing that your final targets are trajectories which must be executed by an agent can give you an edge.\n\nAs an example, the [1st team]( https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/201493) experimented with predicting only differences from a constant velocity model. Other teams tried to predict acceleration and steering instead of displacements and used a unicycle kinematic model to retrieve the trajectories, while others tried to smooth them using a Kalman filter.\n\nThe 3rd team realised that the traffic light information in the rasteriser was static, and that the model might benefit from knowing the evolution of this signal through time.\n\nAgain, the 6th team investigated a lot how to balance data based on the acceleration of the agents in the scene. They noticed that the different predicted modes usually matched different speed profiles and tried to come up with a scheme to sample data to be representative of those different modes.\n\nOften, these experiments have ended up in the “didn’t work” pile but the insight the team drew from them was still crucial in deciding what to do next. If you try a unicycle model and results are very close to just using displacement, you can foresee that investing your time in more refined versions is probably not going to take your score further up.\n\nAs a final note, understanding the task can help you when things behave unexpectedly. This is what must have happened to the [9th team](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199636) after a lot of time spent digging and visualising trajectories. They realised that the input and target spaces we provided were misaligned by a variable angle. This issue slipped through our internal tests and was potentially affecting the whole competition. While deep learning models were still able to learn in these conditions, correctly aligning the two spaces reduced the error in half across all the teams. To our surprise, even our baseline got a performance boost by almost 2X.\n\n##Train for what you measure\nSeveral loss functions can be used to train the same model, leading incredibly different results. If you know the metrics that will be employed to score your model, it makes sense to design a loss which is as close as possible to that metric. This way, you will directly optimise for it and your model will learn how to produce good results according to the metric itself. In the case of this competition, the metric was an extension of the well known Negative Log Likelihood to the problem of multimodal trajectory prediction.\n\nThe competition’s participants understood that, and converted our scoring metric into a differentiable function almost immediately at day 1. Almost all top scorers have employed it in their final solution.\n\n##Combine efforts\nApart from the loss function, several top scorers share another common feature. Instead of just using a single model, they combine an ensemble of them to collect more and better predictions. \n\nWhile a single model can capture several aspects of the world (e.g. semantic information) it may fall short on others (e.g. interactions with other vehicles). Combining different models (and also different input representations) it’s crucial to get better performance on the most difficult samples. Still, it’s important to choose a mixture of diverse and strong models before thinking about how to aggregate them. Ensembling can give you those crucial points you need to rank, but you won’t get there without an accurate selection of your models. \nTo prove this, here are some examples for combinations of best single models and final ensembles results for some top scorers. Ensembling does improve the final score, but the real heavy-lifters are the base models inside of it. Interestingly, the ensemble boost is very similar across teams regardless of the technique employed.\n\n| Team | Best single model score | Ensemble score |\n|------|-------------------------|----------------|\n| 1st  | 9.776                   | 9.319          |\n| 2nd  | 10.496                  | 10.216         |\n| 4th  | 10.846                  | 10.272         |\n| 9th  | 12.285                  | 11.788         |\n\n\nThere are several ways of combining models into ensemble, here is a (non-exhaustive) list of what did work for kagglers:\n\n###Single head over features\nMany teams (such as the 1st, 21st and 22nd) relied on concatenating features of the different models together.\nAfter training individual models, the individual regression heads are removed. The intermediate features are then concatenated and a new head is trained to compress knowledge from the different models into a set of final predictions. The final ensemble is then fine-tuned over the validation set to avoid overfitting the train set. The beauty of this technique is that different models can work on different input representations (e.g. different raster resolutions or raster types), while the head can focus on combining high abstraction features. All teams noticed an improvement when moving from a single model to an ensemble of this type.\n\n<a href=\"https://ibb.co/BPv9cf9\"><img src=\"https://i.ibb.co/cvRjkyj/ensemble.jpg\" alt=\"ensemble\" border=\"0\"></a>\n\n###Gaussian mixture models\nAs part of their solution, the [3rd](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/205376) and [4th team](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199657) introduced a loss which resembles a particular case of GMM. In this setting, the task is to identify a set of final trajectories which best represent the annotated one out of a given distribution. This framework has a nice probabilistic interpretation.\n###Models predictions average\nAlthough many teams tried this technique without luck, others (such as the [6th team](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199588)) succeeded in combining the proposal trajectories from the models using a mean aggregator. \n\nWhile other teams tried to combine the trajectories based on distance, the key intuition of the team was to compute an average based on acceleration and distance covered by the different proposals instead of displacements.\n\nOne of the possible reasons why simply averaging is not enough is that even though two solutions may be perfectly sensible, their average might not. As an example, imagine a situation where an agent can either go straight or perform a lane change. Although these two modes are both plausible, their average may lead the vehicle in between the two lanes, in an unrealistic situation.\n\n<a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/ZV107Yq/lane-change.jpg\" alt=\"lane-change\" border=\"0\"></a>\n\n##Train, train, train!\nIt may sound trivial, but DL models truly get better with more data. And for this competition, large amounts of data were available from day 1. However, with great data come great responsibilities. \nHow long should you train? Should you train on the full dataset or a split? Is data augmentation still crucial with so much data? Which input representation should you try first? When is a good time to reduce the lr? Which scheduler policy should you use?\n\nThe top scorers have all found early answers to these questions, so they had more time to focus on designing the model and not on the training process. Finding the global optimum is not the key factor here, as it could take several weeks to fully evaluate all the combinations. What you really want is a local optimum that gives you a strong signal on which model choices work, so you can invest heavily in them before focusing on polishing the final training details.\n\n###Training time\nTwo train datasets were provided to participants. The first one is around 10GB of data, while the second one is almost 10 times more and therefore has way more variance in terms of agents and scenes. All top scorers found benefits in using the full train dataset over the small one when submitting their results. Also, several teams (like the 4th and 6th) succeeded in speeding up the training time by subsampling and caching parts of the dataset, which drastically reduced the time spent iterating.\n###Data augmentation\nInterestingly, data augmentation didn’t seem to play a crucial role in squeezing out more performance. Among the top scorers, very few teams used any type of data augmentation. As an exception, team 4th employed a cutout augmentation on the input raster. This might suggest that the data available was already enough to provide a good generalisation.\nInput resolution and format\nOn the other hand, all the teams widely experimented with different raster resolutions (ranging from 128 in the ensemble of the 6th place to 448 from the 1st place). Also, many teams focused on the number of history frames to include in the input. Needless to say, all these solutions employ Convolutional neural networks as backbones for the models.\n###Learning rate scheduling\nAnother widely shared “trick” on the forum was the learning rate scheduling during training. Although many teams implemented their own custom scheduling, cosine anneal was by far the most employed one among the top scorers (both 1st and 3rd used it)\n\n##Think out of the box (but don’t lose the focus)\nThe baseline and the toolkit we provided were heavily focused on convolutional neural networks. As such, it’s not surprising that many participants chose to use CNNs to build their initial solutions. CNNs are great and we all love them, but there is a sea of DL opportunity out there to explore!\nSome brave kagglers have experimented with completely different approaches with great results, and we report here some details about their solutions:\n\n###Attention\nThe 3rd team successfully employed a set transformer to combine proposals into 3 candidates. In this way, each trajectory can peek at the others and learn interesting relations. Finally, 3 random vectors are used as seed to learn a weighted synthesis of the proposals and output the final candidate to be scored. \nSimilarly, the 2nd team ensembled trajectories from different models using a matrix of distances between candidates as features for a set of fully connected layers. The layers’ weights can learn relations between different proposals which are then used to estimate how to combine them to obtain the final set of trajectories.\nAttention (and in general transformers) is attracting a huge attention from the community as it can replace unlearned aggregators (max or avg pooling) and even convolution (see for example [here](http://jalammar.github.io/illustrated-transformer/) for more details on the transformer architecture) \n###RNN\nNot many teams employed RNN successfully, especially among the top scorers. As such, we were pretty surprised when we found out that the 2nd team was able to include them in their final architecture and that they were seeing noticeable benefit from using them. In their solution, features extracted from a CNN backbone are processed by a long and a short LSTM layer together with temporal information. These features are then recombined with the original ones to compute the trajectories proposal. By limiting the recurrent part of the architecture to only intermediate features, the team was able to train efficiently and avoid exploding memory allocation.\n###Vector representation\nRasterization is a blunt tool, and can drastically reduce the resolution of the scene around the agent. Luckily, other representations which overcome this issue exist. Two teams drew inspiration from the latest advances in self-driving vehicles and completely discarded raster-based representations in favour of vector-based ones.\n\nThe [22nd team](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199494) employed [PointNet](https://github.com/charlesq34/pointnet) (a famous architecture for point clouds detection) to process point representation of the scene (including the agent of interest and other agents around it). To our surprise, this solution is completely lane blinded, meaning it can’t rely on any static semantic information.\n\nThe [10th team](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199711) chose instead [VectorNet](https://blog.waymo.com/2020/05/vectornet.html), which has recently been proposed for the motion prediction task. For this solution lanes and other semantic information were encoded as vectors and forwarded together with other agents. lanes and other semantic information, and is significantly faster than raster-based ones both during training and inference.\n\nHowever, when time is constrained, failing fast is better than chasing a rabbit down the hole for weeks. As the competition proceeded, some approaches have been deemed not worthy of further exploration, as multiple teams failed to exploit them in any meaningful way. Some of these includes:\n\n###GANs\nGenerative Adversarial Networks have been tried again and again by many teams. Sadly, they have all ended up adding them to the pile of “didn’t work” approaches.\n###3D-CNN\nSome teams tried 3D convolutions, but quickly discarded them for their lack of speed and improvement over 2D solutions.\n###Non DL approaches\nMany teams tried traditional ML and control approaches like Kalman filtering and clustering. However, considering all top solutions include some DL networks those approaches have not proven strong enough. \n#Conclusions\nWe’ve seen many great examples on how to build a winning solution for this year’s competition. We really hope all the participants enjoyed taking part in it. As hosts, we are enthusiastic about how much traction the competition got and we thoroughly enjoyed watching participants using all their tools and knowledge to climb the ladder to those top sweet places. Hopefully, after reading this post you will know a little more about how to get there in future competitions.\n\nAlmost forgot, we’re [hiring](https://www.lyft.com/careers) in the US and in Europe :)",
    "1219012": "Thanks for summarizing the whole project in one discussion forum. Really Helpful.",
    "1209419": ""
  }
}