{
  "id": 205376,
  "title": "3rd Place Solution: Baseline + Set Transformer",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/writeups/stochastic-uplyft-3rd-place-solution-baseline-set",
  "author_name": "",
  "post_date": "2025-07-31T02:10:12.253Z",
  "votes": 27,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Update: YouTube video about our solution <a href=\"https://youtu.be/3Yz8_x38qbc\" target=\"_blank\">https://youtu.be/3Yz8_x38qbc</a></p>\n<h1>TL;DR</h1>\n<ul>\n<li>Baseline CNN regression</li>\n<li>Six 3-mode models based on: Xception41, Xception65, Xception71, EfficientNetB5</li>\n<li>One 16-mode Xception41 model</li>\n<li><a href=\"https://arxiv.org/abs/1810.00825\" target=\"_blank\">Set Transformer</a> as a second level clusterizer model for ensembling</li>\n</ul>\n<p>Our code is available on <a href=\"https://github.com/asanakoy/kaggle-lyft-motion-prediction-av\" target=\"_blank\">GitHub</a></p>\n<h1>Detailed solution</h1>\n<h2>1. Data preprocessing</h2>\n<p>The first level models used the same raster format as generated by l5kit, with next settings:</p>\n<p>raster_size=[224, 224]<br>\npixel_size=[0.5, 0.5]<br>\nego_center=[0.25, 0.5]</p>\n<p>With history frames rendered at frame offsets 0, 1, 2, 4, 8.</p>\n<p>We found the significant time with l5kit rasterizer is spent to prepare coordinates for opencv rendering due to the large number of operations on small numpy arrays, so we combined multiple stages like transform points and CV2 shift to single, numba optimized functions.</p>\n<p>This allowed us to improve performance approximately 1.6x.</p>\n<p>Another improvement was to uncompress zarr files, it’s especially useful with random access.</p>\n<p>Next improvement was to save the cached raster and all relevant information for each training sample to numpy compressed npz file. Especially with the full dataset, we saved each N-th frame for training, since the following frames are usually very similar.</p>\n<p>All optimization combined allowed to <strong>improve the CPU load during training around 6x</strong> and train multiple models simultaneously. The cached training samples for the full dataset used around 1.3TB of space. Fast nvme SSD drive is useful.</p>\n<h2>2. First level CNN models</h2>\n<p>We have tried many approaches but could not beat the baseline solution of using the imagenet pretrained CNN, avg pooling or trainable weighted pooling and fully connected layer to directly predict positions and confidences of trajectories.</p>\n<p>What made a bigger difference was the training parameters. We used SGD with a relatively high learning rate of 0.01, gradient clipping of 2 and batch size around 64-128.</p>\n<p>We used the modified CosineAnnealingWarmRestarts scheduler, starting from period of 16 epoch (200000 samples each), increasing period by 1.41421 times each cycle.</p>\n<p>We used the following models:</p>\n<ul>\n<li>Xception41, avg pool, batch size 64</li>\n<li>Xception41, avg pool, batch size 128 - similar performance to batch size of 64</li>\n<li>Xception41, learnable weighted pooling</li>\n<li>Xception41, predicts 16 modes instead of 3</li>\n<li>Xception65</li>\n<li>Xception71 - seems to be the best performing model, but was trained less comparing to Xception41 due to the shortage of time</li>\n<li>EfficientNet B5</li>\n</ul>\n<p>Initially the EfficientNet based model performed worse compared to Xception41, but with the batch size increased from 64 to 128, the performance improved significantly, with results on par with Xception.</p>\n<p>Models have been trained on the full training dataset for about 5-7 days on a single GPU (2080ti, 3090 for larger models) each.</p>\n<p>Training for longer would likely improve the score, for example Xception41 model reached following validation loss during the last 3 cycles: 11.31, 10.86, 10.37<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2000545%2Fdc9904887584210bbfa46ee33f7a9dd2%2FScreen%20Shot%202020-12-20%20at%200.20.53.png?generation=1608412893460095&amp;alt=media\" alt=\"\"></p>\n<h2>3. Ensembling (second level model)</h2>\n<p>Since we cannot simply average different models predictions, we used an approach where we find 3 trajectories that best approximate input trajectories. This can be achieved by utilizing the competition loss directly, but with input trajectories as ground truth and weighted by input confidences. This can be seen as a particular case of GMM.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2000545%2Fd3c6cac01be9fef035946f893e3a0f05%2FScreen%20Shot%202020-12-20%20at%200.25.01.png?generation=1608413124490687&amp;alt=media\" alt=\"\"><br>\nwhere n index relates to input trajectories and k index relates to 3 final output trajectories.</p>\n<p>At first we optimized this loss with a BFGS solver and this already worked pretty well, however, it was very sensitive to the initialization and tended to get stuck in local optima.</p>\n<p>We tried to use stacking on a hold-out dataset, i.e., train a 2nd level model that takes trajectories from 1st level models as input and predicts 3 completely new trajectories. We compared a bunch of different architectures from MLP to LSTM and luckily came across Set Transformer, which worked almost as well as the BFGS optimizer. Then we noticed that in the paper, they also utilize the model for the GMM task, so we tried to optimize the loss mentioned above with the model instead of stacking, i.e., train it from scratch on the whole val (or test) dataset. It consistently performed by around 0.2 better than the optimizer. We explain such a performance boost by the model's ability to leverage global statistics of the entire provided amount of training data, in contrast to the optimizer, which works sample-wise.</p>\n<p>Our Set Transformer architecture is pretty simple. As an encoder, we have a 2-layer Transformer without positional encoding. In the decoding stage, we have three so-called \"seed vectors\" that are trainable parameters. Those vectors attend to the encoded representations of the trajectories.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2000545%2F0c90849861c770da146a1c1b6ec0f32c%2FScreen%20Shot%202020-12-20%20at%200.27.11.png?generation=1608413244195430&amp;alt=media\" alt=\"\"></p>\n<p>All of the input trajectories are predictions from 3-mode models. However, we also had a 16-mode model that boosted the optimizer quite a bit, but didn't help the transformer. So what we did was to add the 16-mode model predictions to the loss, but remove from the input.</p>\n<h1>Other interesting findings and what didn't work for us</h1>\n<h2>1. Experiments with the Set Transformer</h2>\n<p>When experimenting with the Set Transformer on a hold-out train set, we found these interesting results.</p>\n<table>\n<thead>\n<tr>\n<th>Method</th>\n<th>Val nll loss</th>\n<th>Inference time on the entire val</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Optimizer</td>\n<td>12.06</td>\n<td>~ 30 min</td>\n</tr>\n<tr>\n<td>Transformer: supervised on train set</td>\n<td>12.06</td>\n<td><strong>&lt; 1 sec</strong></td>\n</tr>\n<tr>\n<td>Transformer: unsupervised on train set</td>\n<td>12.00</td>\n<td><strong>&lt; 1 sec</strong></td>\n</tr>\n<tr>\n<td>Transformer: unsupervised→supervised fine-tune on train set</td>\n<td><strong>11.98</strong></td>\n<td><strong>&lt; 1 sec</strong></td>\n</tr>\n<tr>\n<td>Transformer: unsupervised on val set</td>\n<td><strong>11.82</strong></td>\n<td>~ 30-50 min</td>\n</tr>\n</tbody>\n</table>\n<p>Supervised here denotes stacking (original NLL loss with GT trajectory as a target ), and unsupervised indicates the above-mentioned ensemble loss. </p>\n<p>In the table, we can see that unsupervised training from scratch on the validation data has the best performance. However, we can also utilize a train set to train a model that generalizes well to the validation set, outperforming the optimizer. Such an approach is speedy (as only a forward is needed to predict for a new observation) and may be more suitable for production purposes.</p>\n<h2>2. Data preprocessing not used for submitted models</h2>\n<p>We also added extra rasterizer outputs, related to traffic lights: in addition with the current rendering, we added additional 1/4th resolution planes with information about the current and previous traffic lights. We added the separate planes for known on and off traffic lights for forward, left and right directions for the current and previous moments in time. The intuition behind - to provide the model an extra information and separate unknown traffic light from the known off traffic light, has traffic light changed right now or some time ago and easier way to distinguish different signal directions. Lower resolution is sufficient to associate the signal with lanes but can be mixed to deeper model layers, where it does not have to be compressed so much with other inputs. When trained with the extra traffic light inputs, the training and validation loss dropped faster initially but converged to the same level after a few cycles. This approach may still be useful for models used outside of the small training area.</p>\n<h2>3. What we tried, with performance either on par or worse comparing to the simple baseline</h2>\n<ul>\n<li>Used kalman filter to estimate less noisy initial position, velocity, acceleration and angular rate / turn radius: Slightly faster but less stable training, the same final result.</li>\n<li>Used transformers to model interaction between agents and map, no improvement</li>\n<li>Predict the occupation map. May actually be useful for planning, but not very useful for the competition metrics. </li>\n<li>Use the predicted occupancy map of other agents as an input</li>\n<li>Predict all agents recurrently, using transformers or CNN</li>\n<li>Added positions embeddings for points on the map</li>\n<li>Different heads to predict trajectory, RNN, predict acceleration/turn, separate trajectory parameters from velocity on trajectory etc.</li>\n<li>Different number of FC (0-3) layers before the final output FC layer</li>\n</ul>\n<h1>Conlusion</h1>\n<ul>\n<li>CNN regression baseline is very hard to beat</li>\n<li>Training for longer with right parameters on the full dataset was the key</li>\n<li>l5kit can be optimized by a lot</li>\n<li>Great competition overall: lots of high quality data, nice optimizable metric, very<br>\nsupportive host</li>\n<li>Though we thought it would have been better to split train/val/test by geographical<br>\nlocation and to increase the train area.</li>\n</ul>\n<p>Thanks for reading and happy kaggling!</p>",
  "messages": [
    {
      "id": "1119284",
      "postDate": "12/19/2020 21:42:04",
      "content": "<p>Update: YouTube video about our solution <a href=\"https://youtu.be/3Yz8_x38qbc\" target=\"_blank\">https://youtu.be/3Yz8_x38qbc</a></p>\n<h1>TL;DR</h1>\n<ul>\n<li>Baseline CNN regression</li>\n<li>Six 3-mode models based on: Xception41, Xception65, Xception71, EfficientNetB5</li>\n<li>One 16-mode Xception41 model</li>\n<li><a href=\"https://arxiv.org/abs/1810.00825\" target=\"_blank\">Set Transformer</a> as a second level clusterizer model for ensembling</li>\n</ul>\n<p>Our code is available on <a href=\"https://github.com/asanakoy/kaggle-lyft-motion-prediction-av\" target=\"_blank\">GitHub</a></p>\n<h1>Detailed solution</h1>\n<h2>1. Data preprocessing</h2>\n<p>The first level models used the same raster format as generated by l5kit, with next settings:</p>\n<p>raster_size=[224, 224]<br>\npixel_size=[0.5, 0.5]<br>\nego_center=[0.25, 0.5]</p>\n<p>With history frames rendered at frame offsets 0, 1, 2, 4, 8.</p>\n<p>We found the significant time with l5kit rasterizer is spent to prepare coordinates for opencv rendering due to the large number of operations on small numpy arrays, so we combined multiple stages like transform points and CV2 shift to single, numba optimized functions.</p>\n<p>This allowed us to improve performance approximately 1.6x.</p>\n<p>Another improvement was to uncompress zarr files, it’s especially useful with random access.</p>\n<p>Next improvement was to save the cached raster and all relevant information for each training sample to numpy compressed npz file. Especially with the full dataset, we saved each N-th frame for training, since the following frames are usually very similar.</p>\n<p>All optimization combined allowed to <strong>improve the CPU load during training around 6x</strong> and train multiple models simultaneously. The cached training samples for the full dataset used around 1.3TB of space. Fast nvme SSD drive is useful.</p>\n<h2>2. First level CNN models</h2>\n<p>We have tried many approaches but could not beat the baseline solution of using the imagenet pretrained CNN, avg pooling or trainable weighted pooling and fully connected layer to directly predict positions and confidences of trajectories.</p>\n<p>What made a bigger difference was the training parameters. We used SGD with a relatively high learning rate of 0.01, gradient clipping of 2 and batch size around 64-128.</p>\n<p>We used the modified CosineAnnealingWarmRestarts scheduler, starting from period of 16 epoch (200000 samples each), increasing period by 1.41421 times each cycle.</p>\n<p>We used the following models:</p>\n<ul>\n<li>Xception41, avg pool, batch size 64</li>\n<li>Xception41, avg pool, batch size 128 - similar performance to batch size of 64</li>\n<li>Xception41, learnable weighted pooling</li>\n<li>Xception41, predicts 16 modes instead of 3</li>\n<li>Xception65</li>\n<li>Xception71 - seems to be the best performing model, but was trained less comparing to Xception41 due to the shortage of time</li>\n<li>EfficientNet B5</li>\n</ul>\n<p>Initially the EfficientNet based model performed worse compared to Xception41, but with the batch size increased from 64 to 128, the performance improved significantly, with results on par with Xception.</p>\n<p>Models have been trained on the full training dataset for about 5-7 days on a single GPU (2080ti, 3090 for larger models) each.</p>\n<p>Training for longer would likely improve the score, for example Xception41 model reached following validation loss during the last 3 cycles: 11.31, 10.86, 10.37<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2000545%2Fdc9904887584210bbfa46ee33f7a9dd2%2FScreen%20Shot%202020-12-20%20at%200.20.53.png?generation=1608412893460095&amp;alt=media\" alt=\"\"></p>\n<h2>3. Ensembling (second level model)</h2>\n<p>Since we cannot simply average different models predictions, we used an approach where we find 3 trajectories that best approximate input trajectories. This can be achieved by utilizing the competition loss directly, but with input trajectories as ground truth and weighted by input confidences. This can be seen as a particular case of GMM.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2000545%2Fd3c6cac01be9fef035946f893e3a0f05%2FScreen%20Shot%202020-12-20%20at%200.25.01.png?generation=1608413124490687&amp;alt=media\" alt=\"\"><br>\nwhere n index relates to input trajectories and k index relates to 3 final output trajectories.</p>\n<p>At first we optimized this loss with a BFGS solver and this already worked pretty well, however, it was very sensitive to the initialization and tended to get stuck in local optima.</p>\n<p>We tried to use stacking on a hold-out dataset, i.e., train a 2nd level model that takes trajectories from 1st level models as input and predicts 3 completely new trajectories. We compared a bunch of different architectures from MLP to LSTM and luckily came across Set Transformer, which worked almost as well as the BFGS optimizer. Then we noticed that in the paper, they also utilize the model for the GMM task, so we tried to optimize the loss mentioned above with the model instead of stacking, i.e., train it from scratch on the whole val (or test) dataset. It consistently performed by around 0.2 better than the optimizer. We explain such a performance boost by the model's ability to leverage global statistics of the entire provided amount of training data, in contrast to the optimizer, which works sample-wise.</p>\n<p>Our Set Transformer architecture is pretty simple. As an encoder, we have a 2-layer Transformer without positional encoding. In the decoding stage, we have three so-called \"seed vectors\" that are trainable parameters. Those vectors attend to the encoded representations of the trajectories.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2000545%2F0c90849861c770da146a1c1b6ec0f32c%2FScreen%20Shot%202020-12-20%20at%200.27.11.png?generation=1608413244195430&amp;alt=media\" alt=\"\"></p>\n<p>All of the input trajectories are predictions from 3-mode models. However, we also had a 16-mode model that boosted the optimizer quite a bit, but didn't help the transformer. So what we did was to add the 16-mode model predictions to the loss, but remove from the input.</p>\n<h1>Other interesting findings and what didn't work for us</h1>\n<h2>1. Experiments with the Set Transformer</h2>\n<p>When experimenting with the Set Transformer on a hold-out train set, we found these interesting results.</p>\n<table>\n<thead>\n<tr>\n<th>Method</th>\n<th>Val nll loss</th>\n<th>Inference time on the entire val</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Optimizer</td>\n<td>12.06</td>\n<td>~ 30 min</td>\n</tr>\n<tr>\n<td>Transformer: supervised on train set</td>\n<td>12.06</td>\n<td><strong>&lt; 1 sec</strong></td>\n</tr>\n<tr>\n<td>Transformer: unsupervised on train set</td>\n<td>12.00</td>\n<td><strong>&lt; 1 sec</strong></td>\n</tr>\n<tr>\n<td>Transformer: unsupervised→supervised fine-tune on train set</td>\n<td><strong>11.98</strong></td>\n<td><strong>&lt; 1 sec</strong></td>\n</tr>\n<tr>\n<td>Transformer: unsupervised on val set</td>\n<td><strong>11.82</strong></td>\n<td>~ 30-50 min</td>\n</tr>\n</tbody>\n</table>\n<p>Supervised here denotes stacking (original NLL loss with GT trajectory as a target ), and unsupervised indicates the above-mentioned ensemble loss. </p>\n<p>In the table, we can see that unsupervised training from scratch on the validation data has the best performance. However, we can also utilize a train set to train a model that generalizes well to the validation set, outperforming the optimizer. Such an approach is speedy (as only a forward is needed to predict for a new observation) and may be more suitable for production purposes.</p>\n<h2>2. Data preprocessing not used for submitted models</h2>\n<p>We also added extra rasterizer outputs, related to traffic lights: in addition with the current rendering, we added additional 1/4th resolution planes with information about the current and previous traffic lights. We added the separate planes for known on and off traffic lights for forward, left and right directions for the current and previous moments in time. The intuition behind - to provide the model an extra information and separate unknown traffic light from the known off traffic light, has traffic light changed right now or some time ago and easier way to distinguish different signal directions. Lower resolution is sufficient to associate the signal with lanes but can be mixed to deeper model layers, where it does not have to be compressed so much with other inputs. When trained with the extra traffic light inputs, the training and validation loss dropped faster initially but converged to the same level after a few cycles. This approach may still be useful for models used outside of the small training area.</p>\n<h2>3. What we tried, with performance either on par or worse comparing to the simple baseline</h2>\n<ul>\n<li>Used kalman filter to estimate less noisy initial position, velocity, acceleration and angular rate / turn radius: Slightly faster but less stable training, the same final result.</li>\n<li>Used transformers to model interaction between agents and map, no improvement</li>\n<li>Predict the occupation map. May actually be useful for planning, but not very useful for the competition metrics. </li>\n<li>Use the predicted occupancy map of other agents as an input</li>\n<li>Predict all agents recurrently, using transformers or CNN</li>\n<li>Added positions embeddings for points on the map</li>\n<li>Different heads to predict trajectory, RNN, predict acceleration/turn, separate trajectory parameters from velocity on trajectory etc.</li>\n<li>Different number of FC (0-3) layers before the final output FC layer</li>\n</ul>\n<h1>Conlusion</h1>\n<ul>\n<li>CNN regression baseline is very hard to beat</li>\n<li>Training for longer with right parameters on the full dataset was the key</li>\n<li>l5kit can be optimized by a lot</li>\n<li>Great competition overall: lots of high quality data, nice optimizable metric, very<br>\nsupportive host</li>\n<li>Though we thought it would have been better to split train/val/test by geographical<br>\nlocation and to increase the train area.</li>\n</ul>\n<p>Thanks for reading and happy kaggling!</p>",
      "rawMarkdown": "Update: YouTube video about our solution https://youtu.be/3Yz8_x38qbc\n\n# TL;DR\n* Baseline CNN regression\n* Six 3-mode models based on: Xception41, Xception65, Xception71, EfficientNetB5\n* One 16-mode Xception41 model\n* [Set Transformer](https://arxiv.org/abs/1810.00825) as a second level clusterizer model for ensembling\n\nOur code is available on [GitHub](https://github.com/asanakoy/kaggle-lyft-motion-prediction-av)\n\n# Detailed solution\n## 1. Data preprocessing\nThe first level models used the same raster format as generated by l5kit, with next settings:\n\nraster_size=[224, 224]\npixel_size=[0.5, 0.5]\nego_center=[0.25, 0.5]\n\nWith history frames rendered at frame offsets 0, 1, 2, 4, 8.\n\nWe found the significant time with l5kit rasterizer is spent to prepare coordinates for opencv rendering due to the large number of operations on small numpy arrays, so we combined multiple stages like transform points and CV2 shift to single, numba optimized functions.\n\nThis allowed us to improve performance approximately 1.6x.\n\nAnother improvement was to uncompress zarr files, it’s especially useful with random access.\n\nNext improvement was to save the cached raster and all relevant information for each training sample to numpy compressed npz file. Especially with the full dataset, we saved each N-th frame for training, since the following frames are usually very similar.\n\nAll optimization combined allowed to **improve the CPU load during training around 6x** and train multiple models simultaneously. The cached training samples for the full dataset used around 1.3TB of space. Fast nvme SSD drive is useful.\n\n## 2. First level CNN models\nWe have tried many approaches but could not beat the baseline solution of using the imagenet pretrained CNN, avg pooling or trainable weighted pooling and fully connected layer to directly predict positions and confidences of trajectories.\n\nWhat made a bigger difference was the training parameters. We used SGD with a relatively high learning rate of 0.01, gradient clipping of 2 and batch size around 64-128.\n\nWe used the modified CosineAnnealingWarmRestarts scheduler, starting from period of 16 epoch (200000 samples each), increasing period by 1.41421 times each cycle.\n\nWe used the following models:\n* Xception41, avg pool, batch size 64\n* Xception41, avg pool, batch size 128 - similar performance to batch size of 64\n* Xception41, learnable weighted pooling\n* Xception41, predicts 16 modes instead of 3\n* Xception65\n* Xception71 - seems to be the best performing model, but was trained less comparing to Xception41 due to the shortage of time\n* EfficientNet B5\n\nInitially the EfficientNet based model performed worse compared to Xception41, but with the batch size increased from 64 to 128, the performance improved significantly, with results on par with Xception.\n\nModels have been trained on the full training dataset for about 5-7 days on a single GPU (2080ti, 3090 for larger models) each.\n\nTraining for longer would likely improve the score, for example Xception41 model reached following validation loss during the last 3 cycles: 11.31, 10.86, 10.37\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2000545%2Fdc9904887584210bbfa46ee33f7a9dd2%2FScreen%20Shot%202020-12-20%20at%200.20.53.png?generation=1608412893460095&alt=media)\n\n## 3. Ensembling (second level model)\nSince we cannot simply average different models predictions, we used an approach where we find 3 trajectories that best approximate input trajectories. This can be achieved by utilizing the competition loss directly, but with input trajectories as ground truth and weighted by input confidences. This can be seen as a particular case of GMM.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2000545%2Fd3c6cac01be9fef035946f893e3a0f05%2FScreen%20Shot%202020-12-20%20at%200.25.01.png?generation=1608413124490687&alt=media)\nwhere n index relates to input trajectories and k index relates to 3 final output trajectories.\n\nAt first we optimized this loss with a BFGS solver and this already worked pretty well, however, it was very sensitive to the initialization and tended to get stuck in local optima.\n\nWe tried to use stacking on a hold-out dataset, i.e., train a 2nd level model that takes trajectories from 1st level models as input and predicts 3 completely new trajectories. We compared a bunch of different architectures from MLP to LSTM and luckily came across Set Transformer, which worked almost as well as the BFGS optimizer. Then we noticed that in the paper, they also utilize the model for the GMM task, so we tried to optimize the loss mentioned above with the model instead of stacking, i.e., train it from scratch on the whole val (or test) dataset. It consistently performed by around 0.2 better than the optimizer. We explain such a performance boost by the model's ability to leverage global statistics of the entire provided amount of training data, in contrast to the optimizer, which works sample-wise.\n\nOur Set Transformer architecture is pretty simple. As an encoder, we have a 2-layer Transformer without positional encoding. In the decoding stage, we have three so-called \"seed vectors\" that are trainable parameters. Those vectors attend to the encoded representations of the trajectories.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2000545%2F0c90849861c770da146a1c1b6ec0f32c%2FScreen%20Shot%202020-12-20%20at%200.27.11.png?generation=1608413244195430&alt=media)\n\nAll of the input trajectories are predictions from 3-mode models. However, we also had a 16-mode model that boosted the optimizer quite a bit, but didn't help the transformer. So what we did was to add the 16-mode model predictions to the loss, but remove from the input.\n\n# Other interesting findings and what didn't work for us\n## 1. Experiments with the Set Transformer\nWhen experimenting with the Set Transformer on a hold-out train set, we found these interesting results.\n| Method | Val nll loss | Inference time on the entire val |\n| --- | --- | --- |\n| Optimizer | 12.06 | ~ 30 min |\n| Transformer: supervised on train set | 12.06 | **< 1 sec** |\n| Transformer: unsupervised on train set | 12.00 | **< 1 sec** |\n| Transformer: unsupervised→supervised fine-tune on train set | **11.98** | **< 1 sec** |\n| Transformer: unsupervised on val set | **11.82** | ~ 30-50 min |\n\nSupervised here denotes stacking (original NLL loss with GT trajectory as a target ), and unsupervised indicates the above-mentioned ensemble loss. \n\nIn the table, we can see that unsupervised training from scratch on the validation data has the best performance. However, we can also utilize a train set to train a model that generalizes well to the validation set, outperforming the optimizer. Such an approach is speedy (as only a forward is needed to predict for a new observation) and may be more suitable for production purposes.\n\n## 2. Data preprocessing not used for submitted models\nWe also added extra rasterizer outputs, related to traffic lights: in addition with the current rendering, we added additional 1/4th resolution planes with information about the current and previous traffic lights. We added the separate planes for known on and off traffic lights for forward, left and right directions for the current and previous moments in time. The intuition behind - to provide the model an extra information and separate unknown traffic light from the known off traffic light, has traffic light changed right now or some time ago and easier way to distinguish different signal directions. Lower resolution is sufficient to associate the signal with lanes but can be mixed to deeper model layers, where it does not have to be compressed so much with other inputs. When trained with the extra traffic light inputs, the training and validation loss dropped faster initially but converged to the same level after a few cycles. This approach may still be useful for models used outside of the small training area.\n\n## 3. What we tried, with performance either on par or worse comparing to the simple baseline\n* Used kalman filter to estimate less noisy initial position, velocity, acceleration and angular rate / turn radius: Slightly faster but less stable training, the same final result.\n* Used transformers to model interaction between agents and map, no improvement\n* Predict the occupation map. May actually be useful for planning, but not very useful for the competition metrics. \n* Use the predicted occupancy map of other agents as an input\n* Predict all agents recurrently, using transformers or CNN\n* Added positions embeddings for points on the map\n* Different heads to predict trajectory, RNN, predict acceleration/turn, separate trajectory parameters from velocity on trajectory etc.\n* Different number of FC (0-3) layers before the final output FC layer\n\n# Conlusion\n* CNN regression baseline is very hard to beat\n* Training for longer with right parameters on the full dataset was the key\n* l5kit can be optimized by a lot\n* Great competition overall: lots of high quality data, nice optimizable metric, very\nsupportive host\n* Though we thought it would have been better to split train/val/test by geographical\nlocation and to increase the train area.\n\nThanks for reading and happy kaggling!",
      "votes": null
    },
    {
      "id": "1119752",
      "postDate": "12/20/2020 10:54:35",
      "content": "<p>Thanks a lot for that summary, a much better read than the news for Sunday morning.</p>\n<p>Can it be possible for you to provide the single model test csv files, I want to test your ensembling scores vs what I can get with my optimizer. Same as <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/201493#1102803\" target=\"_blank\">the check</a> that I did for the first place solution. Instead of BFGS, as you did, I used a custom fixed point procedure. It is interesting to know if it finds the global minimum more reliably compared to BFGS, and whether it outperforms the Set Transformer.</p>\n<p>What experiment tracking system has you used?</p>",
      "rawMarkdown": "Thanks a lot for that summary, a much better read than the news for Sunday morning.\n\nCan it be possible for you to provide the single model test csv files, I want to test your ensembling scores vs what I can get with my optimizer. Same as [the check](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/201493#1102803) that I did for the first place solution. Instead of BFGS, as you did, I used a custom fixed point procedure. It is interesting to know if it finds the global minimum more reliably compared to BFGS, and whether it outperforms the Set Transformer.\n\nWhat experiment tracking system has you used?",
      "votes": null
    },
    {
      "id": "1120129",
      "postDate": "12/20/2020 15:32:54",
      "content": "<p>Hi, thanks for the comment.<br>\nYou can download our test predictions in npz format from <a href=\"https://drive.google.com/drive/folders/1hM0aPGugVeAY8qUAUIWVBVcqb3KGk2rQ?usp=sharing\" target=\"_blank\">here</a>. Just ignore the sample submission csv. Our scores with these models are: public - 10.209, private - 9.404.</p>\n<p>As for the loss plot, it's Tensorboard.</p>",
      "rawMarkdown": "Hi, thanks for the comment.\nYou can download our test predictions in npz format from [here](https://drive.google.com/drive/folders/1hM0aPGugVeAY8qUAUIWVBVcqb3KGk2rQ?usp=sharing). Just ignore the sample submission csv. Our scores with these models are: public - 10.209, private - 9.404.\n\n\nAs for the loss plot, it's Tensorboard.",
      "votes": null
    },
    {
      "id": "1120535",
      "postDate": "12/20/2020 22:20:31",
      "content": "<p>Thank you. With equal weights for 7 models I got a significantly worse result. What are the models validation or LB scores, to select the weights more meaningfully?</p>\n<p><img src=\"https://imgur.com/nO7CrT8.png\" alt=\"\"></p>",
      "rawMarkdown": "Thank you. With equal weights for 7 models I got a significantly worse result. What are the models validation or LB scores, to select the weights more meaningfully?\n\n![](https://imgur.com/nO7CrT8.png)",
      "votes": null
    },
    {
      "id": "1120537",
      "postDate": "12/20/2020 22:29:02",
      "content": "<p>Not sure if we've written down the scores somewhere… You can check the LB score by yourself though :) Also note the 16-mode model. It may not fit your approach (haven't read that so not sure).</p>",
      "rawMarkdown": "Not sure if we've written down the scores somewhere... You can check the LB score by yourself though :) Also note the 16-mode model. It may not fit your approach (haven't read that so not sure).",
      "votes": null
    },
    {
      "id": "1120550",
      "postDate": "12/20/2020 22:54:40",
      "content": "<p>Well, it is too time-consuming. I think I will stop here, thanks for the help.</p>",
      "rawMarkdown": "Well, it is too time-consuming. I think I will stop here, thanks for the help.",
      "votes": null
    },
    {
      "id": "1120582",
      "postDate": "12/20/2020 23:51:14",
      "content": "<p>Thanks for sharing and congrats! It's impressive that you won 3rd place using single GPU training!</p>",
      "rawMarkdown": "Thanks for sharing and congrats! It's impressive that you won 3rd place using single GPU training!",
      "votes": null
    },
    {
      "id": "1176635",
      "postDate": "01/29/2021 17:49:57",
      "content": "<p>Hey guys,<br>\nI have created a <strong>youtube video</strong> describing our 3rd place solution in <strong>intuitive way</strong> which will be very useful for beginners or those who did not take part in this competition but want to learn about it and about our approach.<br>\nVideo link: <a href=\"https://youtu.be/3Yz8_x38qbc\" target=\"_blank\">https://youtu.be/3Yz8_x38qbc</a></p>",
      "rawMarkdown": "Hey guys,\nI have created a **youtube video** describing our 3rd place solution in **intuitive way** which will be very useful for beginners or those who did not take part in this competition but want to learn about it and about our approach.\nVideo link: https://youtu.be/3Yz8_x38qbc",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1119752,
      "author_name": "zaharch",
      "author_url": "",
      "post_date": "12/20/2020 10:54:35",
      "content": "<p>Thanks a lot for that summary, a much better read than the news for Sunday morning.</p>\n<p>Can it be possible for you to provide the single model test csv files, I want to test your ensembling scores vs what I can get with my optimizer. Same as <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/201493#1102803\" target=\"_blank\">the check</a> that I did for the first place solution. Instead of BFGS, as you did, I used a custom fixed point procedure. It is interesting to know if it finds the global minimum more reliably compared to BFGS, and whether it outperforms the Set Transformer.</p>\n<p>What experiment tracking system has you used?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1120129,
          "author_name": "aruchomu",
          "author_url": "",
          "post_date": "12/20/2020 15:32:54",
          "content": "<p>Hi, thanks for the comment.<br>\nYou can download our test predictions in npz format from <a href=\"https://drive.google.com/drive/folders/1hM0aPGugVeAY8qUAUIWVBVcqb3KGk2rQ?usp=sharing\" target=\"_blank\">here</a>. Just ignore the sample submission csv. Our scores with these models are: public - 10.209, private - 9.404.</p>\n<p>As for the loss plot, it's Tensorboard.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1120535,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "12/20/2020 22:20:31",
          "content": "<p>Thank you. With equal weights for 7 models I got a significantly worse result. What are the models validation or LB scores, to select the weights more meaningfully?</p>\n<p><img src=\"https://imgur.com/nO7CrT8.png\" alt=\"\"></p>",
          "votes": null,
          "replies": [
            {
              "id": 1120537,
              "author_name": "aruchomu",
              "author_url": "",
              "post_date": "12/20/2020 22:29:02",
              "content": "<p>Not sure if we've written down the scores somewhere… You can check the LB score by yourself though :) Also note the 16-mode model. It may not fit your approach (haven't read that so not sure).</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 1120550,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "12/20/2020 22:54:40",
          "content": "<p>Well, it is too time-consuming. I think I will stop here, thanks for the help.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1120582,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "12/20/2020 23:51:14",
      "content": "<p>Thanks for sharing and congrats! It's impressive that you won 3rd place using single GPU training!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1176635,
      "author_name": "asanakoev",
      "author_url": "",
      "post_date": "01/29/2021 17:49:57",
      "content": "<p>Hey guys,<br>\nI have created a <strong>youtube video</strong> describing our 3rd place solution in <strong>intuitive way</strong> which will be very useful for beginners or those who did not take part in this competition but want to learn about it and about our approach.<br>\nVideo link: <a href=\"https://youtu.be/3Yz8_x38qbc\" target=\"_blank\">https://youtu.be/3Yz8_x38qbc</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1119284": "Update: YouTube video about our solution https://youtu.be/3Yz8_x38qbc\n\n# TL;DR\n* Baseline CNN regression\n* Six 3-mode models based on: Xception41, Xception65, Xception71, EfficientNetB5\n* One 16-mode Xception41 model\n* [Set Transformer](https://arxiv.org/abs/1810.00825) as a second level clusterizer model for ensembling\n\nOur code is available on [GitHub](https://github.com/asanakoy/kaggle-lyft-motion-prediction-av)\n\n# Detailed solution\n## 1. Data preprocessing\nThe first level models used the same raster format as generated by l5kit, with next settings:\n\nraster_size=[224, 224]\npixel_size=[0.5, 0.5]\nego_center=[0.25, 0.5]\n\nWith history frames rendered at frame offsets 0, 1, 2, 4, 8.\n\nWe found the significant time with l5kit rasterizer is spent to prepare coordinates for opencv rendering due to the large number of operations on small numpy arrays, so we combined multiple stages like transform points and CV2 shift to single, numba optimized functions.\n\nThis allowed us to improve performance approximately 1.6x.\n\nAnother improvement was to uncompress zarr files, it’s especially useful with random access.\n\nNext improvement was to save the cached raster and all relevant information for each training sample to numpy compressed npz file. Especially with the full dataset, we saved each N-th frame for training, since the following frames are usually very similar.\n\nAll optimization combined allowed to **improve the CPU load during training around 6x** and train multiple models simultaneously. The cached training samples for the full dataset used around 1.3TB of space. Fast nvme SSD drive is useful.\n\n## 2. First level CNN models\nWe have tried many approaches but could not beat the baseline solution of using the imagenet pretrained CNN, avg pooling or trainable weighted pooling and fully connected layer to directly predict positions and confidences of trajectories.\n\nWhat made a bigger difference was the training parameters. We used SGD with a relatively high learning rate of 0.01, gradient clipping of 2 and batch size around 64-128.\n\nWe used the modified CosineAnnealingWarmRestarts scheduler, starting from period of 16 epoch (200000 samples each), increasing period by 1.41421 times each cycle.\n\nWe used the following models:\n* Xception41, avg pool, batch size 64\n* Xception41, avg pool, batch size 128 - similar performance to batch size of 64\n* Xception41, learnable weighted pooling\n* Xception41, predicts 16 modes instead of 3\n* Xception65\n* Xception71 - seems to be the best performing model, but was trained less comparing to Xception41 due to the shortage of time\n* EfficientNet B5\n\nInitially the EfficientNet based model performed worse compared to Xception41, but with the batch size increased from 64 to 128, the performance improved significantly, with results on par with Xception.\n\nModels have been trained on the full training dataset for about 5-7 days on a single GPU (2080ti, 3090 for larger models) each.\n\nTraining for longer would likely improve the score, for example Xception41 model reached following validation loss during the last 3 cycles: 11.31, 10.86, 10.37\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2000545%2Fdc9904887584210bbfa46ee33f7a9dd2%2FScreen%20Shot%202020-12-20%20at%200.20.53.png?generation=1608412893460095&alt=media)\n\n## 3. Ensembling (second level model)\nSince we cannot simply average different models predictions, we used an approach where we find 3 trajectories that best approximate input trajectories. This can be achieved by utilizing the competition loss directly, but with input trajectories as ground truth and weighted by input confidences. This can be seen as a particular case of GMM.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2000545%2Fd3c6cac01be9fef035946f893e3a0f05%2FScreen%20Shot%202020-12-20%20at%200.25.01.png?generation=1608413124490687&alt=media)\nwhere n index relates to input trajectories and k index relates to 3 final output trajectories.\n\nAt first we optimized this loss with a BFGS solver and this already worked pretty well, however, it was very sensitive to the initialization and tended to get stuck in local optima.\n\nWe tried to use stacking on a hold-out dataset, i.e., train a 2nd level model that takes trajectories from 1st level models as input and predicts 3 completely new trajectories. We compared a bunch of different architectures from MLP to LSTM and luckily came across Set Transformer, which worked almost as well as the BFGS optimizer. Then we noticed that in the paper, they also utilize the model for the GMM task, so we tried to optimize the loss mentioned above with the model instead of stacking, i.e., train it from scratch on the whole val (or test) dataset. It consistently performed by around 0.2 better than the optimizer. We explain such a performance boost by the model's ability to leverage global statistics of the entire provided amount of training data, in contrast to the optimizer, which works sample-wise.\n\nOur Set Transformer architecture is pretty simple. As an encoder, we have a 2-layer Transformer without positional encoding. In the decoding stage, we have three so-called \"seed vectors\" that are trainable parameters. Those vectors attend to the encoded representations of the trajectories.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2000545%2F0c90849861c770da146a1c1b6ec0f32c%2FScreen%20Shot%202020-12-20%20at%200.27.11.png?generation=1608413244195430&alt=media)\n\nAll of the input trajectories are predictions from 3-mode models. However, we also had a 16-mode model that boosted the optimizer quite a bit, but didn't help the transformer. So what we did was to add the 16-mode model predictions to the loss, but remove from the input.\n\n# Other interesting findings and what didn't work for us\n## 1. Experiments with the Set Transformer\nWhen experimenting with the Set Transformer on a hold-out train set, we found these interesting results.\n| Method | Val nll loss | Inference time on the entire val |\n| --- | --- | --- |\n| Optimizer | 12.06 | ~ 30 min |\n| Transformer: supervised on train set | 12.06 | **< 1 sec** |\n| Transformer: unsupervised on train set | 12.00 | **< 1 sec** |\n| Transformer: unsupervised→supervised fine-tune on train set | **11.98** | **< 1 sec** |\n| Transformer: unsupervised on val set | **11.82** | ~ 30-50 min |\n\nSupervised here denotes stacking (original NLL loss with GT trajectory as a target ), and unsupervised indicates the above-mentioned ensemble loss. \n\nIn the table, we can see that unsupervised training from scratch on the validation data has the best performance. However, we can also utilize a train set to train a model that generalizes well to the validation set, outperforming the optimizer. Such an approach is speedy (as only a forward is needed to predict for a new observation) and may be more suitable for production purposes.\n\n## 2. Data preprocessing not used for submitted models\nWe also added extra rasterizer outputs, related to traffic lights: in addition with the current rendering, we added additional 1/4th resolution planes with information about the current and previous traffic lights. We added the separate planes for known on and off traffic lights for forward, left and right directions for the current and previous moments in time. The intuition behind - to provide the model an extra information and separate unknown traffic light from the known off traffic light, has traffic light changed right now or some time ago and easier way to distinguish different signal directions. Lower resolution is sufficient to associate the signal with lanes but can be mixed to deeper model layers, where it does not have to be compressed so much with other inputs. When trained with the extra traffic light inputs, the training and validation loss dropped faster initially but converged to the same level after a few cycles. This approach may still be useful for models used outside of the small training area.\n\n## 3. What we tried, with performance either on par or worse comparing to the simple baseline\n* Used kalman filter to estimate less noisy initial position, velocity, acceleration and angular rate / turn radius: Slightly faster but less stable training, the same final result.\n* Used transformers to model interaction between agents and map, no improvement\n* Predict the occupation map. May actually be useful for planning, but not very useful for the competition metrics. \n* Use the predicted occupancy map of other agents as an input\n* Predict all agents recurrently, using transformers or CNN\n* Added positions embeddings for points on the map\n* Different heads to predict trajectory, RNN, predict acceleration/turn, separate trajectory parameters from velocity on trajectory etc.\n* Different number of FC (0-3) layers before the final output FC layer\n\n# Conlusion\n* CNN regression baseline is very hard to beat\n* Training for longer with right parameters on the full dataset was the key\n* l5kit can be optimized by a lot\n* Great competition overall: lots of high quality data, nice optimizable metric, very\nsupportive host\n* Though we thought it would have been better to split train/val/test by geographical\nlocation and to increase the train area.\n\nThanks for reading and happy kaggling!",
    "1119752": "Thanks a lot for that summary, a much better read than the news for Sunday morning.\n\nCan it be possible for you to provide the single model test csv files, I want to test your ensembling scores vs what I can get with my optimizer. Same as [the check](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/201493#1102803) that I did for the first place solution. Instead of BFGS, as you did, I used a custom fixed point procedure. It is interesting to know if it finds the global minimum more reliably compared to BFGS, and whether it outperforms the Set Transformer.\n\nWhat experiment tracking system has you used?",
    "1120129": "Hi, thanks for the comment.\nYou can download our test predictions in npz format from [here](https://drive.google.com/drive/folders/1hM0aPGugVeAY8qUAUIWVBVcqb3KGk2rQ?usp=sharing). Just ignore the sample submission csv. Our scores with these models are: public - 10.209, private - 9.404.\n\n\nAs for the loss plot, it's Tensorboard.",
    "1120535": "Thank you. With equal weights for 7 models I got a significantly worse result. What are the models validation or LB scores, to select the weights more meaningfully?\n\n![](https://imgur.com/nO7CrT8.png)",
    "1120537": "Not sure if we've written down the scores somewhere... You can check the LB score by yourself though :) Also note the 16-mode model. It may not fit your approach (haven't read that so not sure).",
    "1120550": "Well, it is too time-consuming. I think I will stop here, thanks for the help.",
    "1120582": "Thanks for sharing and congrats! It's impressive that you won 3rd place using single GPU training!",
    "1176635": "Hey guys,\nI have created a **youtube video** describing our 3rd place solution in **intuitive way** which will be very useful for beginners or those who did not take part in this competition but want to learn about it and about our approach.\nVideo link: https://youtu.be/3Yz8_x38qbc"
  },
  "source": "meta"
}