{
  "id": 201493,
  "title": "1st Place Solution & L5Kit Speedup",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/201493",
  "author_name": "Pascal Pfeiffer",
  "post_date": "2020-12-05T09:28:55.099000",
  "votes": 82,
  "comment_count": 31,
  "views": 0,
  "content": "<h1>1st Place Solution &amp; L5Kit Speedup</h1>\n<h2>Introduction</h2>\n<p>First of all, I would like to thank Lyft for this great challenge and of course for their huge dataset. I must say I was a bit overwhelmed by the sheer amount of agents in the training set, but it was a fun experience to be able to train without ever using training data twice. Well, we did, when we used pre-trained models, but more on that later.</p>\n<p>I also want to thank my Teammates <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>, <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> and <a href=\"https://www.kaggle.com/nvnnghia\" target=\"_blank\">@nvnnghia</a> for the great collaborative teamwork during the competition. We worked well together with a pipeline consisting of two repositories (version controlled custom l5kit and a training repo) and used logging in neptune.ai to keep track of the experiments. </p>\n<h2>TL;DR</h2>\n<p>Our solution is an ensemble of 4 efficientnet models trained with different rasterization parameters indivdually on train_full.zarr and stacked on validation.zarr. Key to training that many models on this huge dataset were several key improvements to the original l5kit repository to remove CPU bottleneck and enable efficient multi-GPU training using pytorch. </p>\n<h2>Improvements/Changes to the L5Kit</h2>\n<p>One of the major challenges in this competition was the slow rasterizer. We profiled the l5kit a lot and pinpointed the bottlenecks to speed it up by a factor of 4+. After that, our experiments were mostly GPU limited and multi-GPU experiments allowed us for \"fast\" iterations. Running through train_full (191_177_863+ samples) was possible within 2 days with medium sized models. </p>\n<h3>Speedups</h3>\n<p>Speed before changes with semantic view only (i5-3570K single thread, so ignore the actual time):</p>\n<pre><code>Raster-time per sample: 0.124 s\nHits         Time  Per Hit   % Time  Line Contents\n==================================================\n9   11147507.0 1238611.9     93.6    rasterizer.rasterize(history_frames, history_agents, history_tl_faces, selected_agent)\n</code></pre>\n<p>All the speedups were achieved with plain python, no C++. The main speedups came from:</p>\n<ol>\n<li>Batched transformation of boxes with modifications from <a href=\"https://github.com/lyft/l5kit/pull/167\" target=\"_blank\">https://github.com/lyft/l5kit/pull/167</a>. <br>\nIt has been said that the transformation is the bottleneck in the rasterizer as it is called a lot. We used the vectorized transformation whereever possible.</li>\n<li>In <code>box_rasterizer.py</code>, using python lists of small numpy arrays instead of one two large numpy arrays (<code>agent_images, ego_images</code>).<br>\nPython lists are surprisingly fast! <code>np.concatenate</code> on large numpy arrays is slow. Pre-allocation and writing to a single large array is even slower.<br>\nThe lists can be cast to a numpy array by <code>out_im = np.asarray(agents_images + ego_images, dtype=np.uint8)</code><br>\nThis way, boxes will stay in <code>uint8</code></li>\n<li>Moving concatenate to the GPU<br>\nAs <code>np.concatenate</code> is slow on large arrays, we do that on the GPU.<br>\nReplace <code>\"image\": image,</code> with <code>\"image_box\": image_box</code>, <code>\"image_sat\": image_sat</code>, <code>\"image_sem\": image_sem</code></li>\n<li>Speedups from <a href=\"https://github.com/lyft/l5kit/pull/140\" target=\"_blank\">https://github.com/lyft/l5kit/pull/140</a><br>\nWith the changes above, they now made a larger impact than the 7-8% stated in the PR.</li>\n</ol>\n<p>Speed after changes with semantic and satellite view (i5-3570K single thread, so ignore the actual time):</p>\n<pre><code>Raster-time per sample: 0.032 s\nHits         Time  Per Hit   % Time  Line Contents\n==================================================\n   9    2847054.0 316339.3     84.2  rasterizer.rasterize(history_frames, history_agents, history_tl_faces, selected_agent)\n</code></pre>\n<h3>Additions/Changes</h3>\n<ul>\n<li>New rasterizer which combines satellite and semantic view (<code>py_sem_sat</code>).</li>\n<li>Flag: AgentID used as a value for drawing (instead of just boolean <code>no agent = 0</code>, <code>agent = 1</code>)</li>\n<li>Flag: AgentLabelProba used as a value for drawing (instead of just boolean <code>no agent = 0</code>, <code>agent = 1</code>)</li>\n<li>Flag: Velocity used as a value for drawing (instead of just boolean <code>no agent = 0</code>, <code>agent = 1</code>)</li>\n<li>Option to draw all objects on the semantic layer, combinable with the label probas from above.</li>\n<li>Filter options for agents: <code>th_extent_ratio, th_area, th_distance_av</code> to give a larger training dataset (or harder as you want to phrase it)</li>\n<li>Multiple parallel rasterizers for ensembling models that use different raster sizes</li>\n<li>Satellite view fix for non square shapes (had some weird transformation in it before)</li>\n</ul>\n<p><img src=\"https://i.imgur.com/G3g9lEJ.png\" alt=\"rasterized_image\"></p>\n<p>We will make our L5kit code changes public in the next few days and after cleanup. Most likely with PRs to the original repo.</p>\n<h2>Journey &amp; Experiments</h2>\n<p>With all the speed improvements from our custom L5kit we were able to eliminate the previous CPU bottleneck and run experiments on an efficient multi-GPU setup. We mostly trained on 8xV100 nodes using pytorchs awesome distributed data parallel (DDP), which enables to effectivly scale a training script to multiple GPUs. This efficient training setup was key to run quite a few experiments. (See list of stuff that did not work below) In all our experiments, we used the chopped validation set for local validation as discussed <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/188695\" target=\"_blank\">here</a>. This yielded us very good prediction of the LB score (in the sub 13 range, the score was usually ~0.2 higher than local validation with a standard deviation of around 0.3). </p>\n<p>Originating from the provided baseline and the multi-mode prediction from <a href=\"https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence\" target=\"_blank\">here</a>, we experimented with raster size, pooling layers, head-modifications, usage of different backbones, learning rate scheduler, subsampling and oversampling to name a few. It was a major challenge in this competition, that experiments had to be run for at least 50 % of a full epoch of train_full.zarr or when comparing different sample sizes or learning rate scheduler even until the end to be really comparable. Subsampling the train set or stepwise decrease of the LR quickly yielded a low training and validation loss, but in the end, the performance was worse.</p>\n<p>Quite quickly, we decided to always use all data, and settled to a setup using shuffled train_full.zarr (with a fixed seed to be able to resume training) with a linear decay learning rate scheduler. To keep the training sample fixed, we also settled to an AgentDataset with <code>min_frame_history=1</code> and <code>min_frame_future=10.0</code> to best resemble the test dataset. Opposed to the many discussions about a much lower training loss, we never experienced that. Without any augmentation, the training loss is of course slightly lower, especially in the end of the epoch, but usually within 80% of the validation score. We found that larger backbones help to reach lower loss levels, but starting at the size of an EfficientNetB7, the nets seem to be overconfident with their predictions and the LB score was sometimes shaky. </p>\n<p>The first half of test is public, second half is private test. We couldn't identify any major differences in the two sets. With that knowledge it was possible to hide the true score (setting some public rows to zeros), we wonder how many teams did that and were afraid of surprises in the private LB due to this. </p>\n<h2>Best Single Model</h2>\n<p><img src=\"https://i.imgur.com/YC1WGD1.png\" alt=\"model\"></p>\n<p>Our best single model is actually one that not fully finished on the last day. It's just an EfficientNetB3 with a Linear Layer Head and dropout attached trained on an extended train_full.zarr Dataset (<code>min_frame_history=1, min_frame_future=5.0, th_distance_av=80, th_extent_ratio=1.6, th_area=0.45</code>)</p>\n<ul>\n<li>history_num_frames: 30.0</li>\n<li>raster_size: [448, 224]</li>\n<li>pixel_size: [0.1875, 0.1875]</li>\n<li>rasterizer: py_sem_sat</li>\n<li>batch_size: 64</li>\n<li>dropout: 0.3</li>\n</ul>\n<p>The model was pre-trained for 4 epochs on different lower resolution images (starting with 112, 64 and pixel size 0.75, 0.75) , so you can argue this is also an ensemble of some kind. The 5th epoch was trained with the parameters above and a very customized learning rate schedule that starts with a consine anneal and then transitions to a stepwise lr scheduler as seen in the graph below. The last evaluation score is 9.697.</p>\n<p><img src=\"https://i.imgur.com/4Q2PkvX.png\" alt=\"best_single\"></p>\n<p>The model achieves a public LB score of 9.776 and a private LB score of 9.070. With that, it would also have ranked 1st place this competition on it's own.</p>\n<p>Pretraining is VERY important! We discovered it by a mistake where we loaded an EfficientNet without pretrained weights from imagenet. It performed significantly worse by ~1.00.</p>\n<h2>Ensemble</h2>\n<p>Ensembling proved to be challenging in this competition, as traditional methods are not working well with the metric in use. We tried a lot of manual blending and form of post-processing without success. We then started building stacking models taking the raw predictions of the models as input, but this was also surprisingly not working well. At about the same time as <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> posted it <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/193908#1077899\" target=\"_blank\">here</a>, we had the idea of using the features/embeddings from each model as an input to a ensemble head. This worked much better and we could improve upon our single models specifically as we could introduce some diversity into the blend (different image size, pixel sizes, etc.) We also went with the dropout + single linear layer approach here, and all modification we tried with the head performed worse.</p>\n<p><img src=\"https://i.imgur.com/2I3M2rD.png\" alt=\"ensemble_arch\"></p>\n<p>To prevent overfitting, the ensembling was done on the chopped validation set (without using the chopped off actual validation part), in some experiments extended to validation + test set without noticeable change in the metric but with about twice the runtime. </p>\n<p>Ensembles of two similar models (even just two different checkpoints) already gave a boost of about 0.2 in the metric. Naturally, with more diverse models the ensemble performed even better and we were able to achieve our best public LB score of 9.319 and private LB score of 8.579 with an ensemble of 4 models (B3, B5, B6, B6) on two different raster sizes. </p>\n<h2>Bootstrapped Validation</h2>\n<p>We evaluated model and ensemble performance with bootstrap method on the chopped dataset. For more diverse ensembles we also got the lowest standard deviation from the bootstrap method. For our final submission it was 0.22754. The best single models had a standard deviation of ~0.27. We also identified our B7 and B8 models to have a slightly larger standard deviation in the bootstrap method of ~0.36-0.39 which may also explain their partial performance degredations on the public LB. Despite their good validation score, for that reason we excluded them in our final ensembles and final submissions. </p>\n<p><img src=\"https://i.imgur.com/l4vEAaz.png\" alt=\"bootstrap\"></p>\n<p>We plotted CV vs public LB for our submissions and found a spread of +-0.3 for most of our single models (green corridor) with just the B7 and B8 models outside of that corridor (not shown in the picture). Our ensembles are in an even smaller corridor with a spread of +-0.1 (orange corridor), proving their superior robustness and generalization capabilities over single models. </p>\n<p><img src=\"https://i.imgur.com/S6fYsBI.png\" alt=\"cv_lb\"></p>\n<p>Given this robust CV &amp; LB correlation, we were very confident that improvements in validation also lead to improvements on the leaderboard, so we did not have to sub it. So the sudden jump on public leaderboard only happened after we decided to sub our better models and was not some sudden magic we discovered.</p>\n<h2>What Didn't Work</h2>\n<ul>\n<li>Flag: AgentID used as a value for drawing (instead of just boolean <code>no agent = 0</code>, <code>agent = 1</code>)</li>\n<li>Flag: AgentLabelProba used as a value for drawing (instead of just boolean <code>no agent = 0</code>, <code>agent = 1</code>)</li>\n<li>Flag: AgentLabelProba used as a value for drawing (instead of just boolean <code>no agent = 0</code>, <code>agent = 1</code>)</li>\n<li>Option to draw all objects on the semantic layer, combinable with the label probas from above.</li>\n<li>Kalman filtering the output</li>\n<li>Any non NN prediction method</li>\n<li>Prediction of only the difference to a constant velocity model</li>\n<li>Using additional informations, such as velocity, history positions, label probas or anything else we could find in the head of the models. </li>\n<li>Heavier heads</li>\n<li>3D CNNs</li>\n<li>RNN based approaches on history frames for backbone input or outbut </li>\n<li>Using additional features from shallower layers of the backbone</li>\n<li>Ensembling with kmeans, or other analytic clustering method that we tried. Some success at scores down to around 14, but no improvements below that. </li>\n<li>Subsampling the training set. Huge speedups, but accuracy was always slightly lower</li>\n<li>TTA with modified yaw angles</li>\n<li>Unfreezing the backbone after ensembling and continue training with a second epoch</li>\n</ul>\n<h2>What We Didn't Try</h2>\n<ul>\n<li>Augmentation</li>\n<li>Multi backbone, multi rasterizer training</li>\n</ul>",
  "messages": [
    {
      "id": 1102764,
      "postDate": "2020-12-05T09:28:55.100Z",
      "content": "<h1>1st Place Solution &amp; L5Kit Speedup</h1>\n<h2>Introduction</h2>\n<p>First of all, I would like to thank Lyft for this great challenge and of course for their huge dataset. I must say I was a bit overwhelmed by the sheer amount of agents in the training set, but it was a fun experience to be able to train without ever using training data twice. Well, we did, when we used pre-trained models, but more on that later.</p>\n<p>I also want to thank my Teammates <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>, <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> and <a href=\"https://www.kaggle.com/nvnnghia\" target=\"_blank\">@nvnnghia</a> for the great collaborative teamwork during the competition. We worked well together with a pipeline consisting of two repositories (version controlled custom l5kit and a training repo) and used logging in neptune.ai to keep track of the experiments. </p>\n<h2>TL;DR</h2>\n<p>Our solution is an ensemble of 4 efficientnet models trained with different rasterization parameters indivdually on train_full.zarr and stacked on validation.zarr. Key to training that many models on this huge dataset were several key improvements to the original l5kit repository to remove CPU bottleneck and enable efficient multi-GPU training using pytorch. </p>\n<h2>Improvements/Changes to the L5Kit</h2>\n<p>One of the major challenges in this competition was the slow rasterizer. We profiled the l5kit a lot and pinpointed the bottlenecks to speed it up by a factor of 4+. After that, our experiments were mostly GPU limited and multi-GPU experiments allowed us for \"fast\" iterations. Running through train_full (191_177_863+ samples) was possible within 2 days with medium sized models. </p>\n<h3>Speedups</h3>\n<p>Speed before changes with semantic view only (i5-3570K single thread, so ignore the actual time):</p>\n<pre><code>Raster-time per sample: 0.124 s\nHits         Time  Per Hit   % Time  Line Contents\n==================================================\n9   11147507.0 1238611.9     93.6    rasterizer.rasterize(history_frames, history_agents, history_tl_faces, selected_agent)\n</code></pre>\n<p>All the speedups were achieved with plain python, no C++. The main speedups came from:</p>\n<ol>\n<li>Batched transformation of boxes with modifications from <a href=\"https://github.com/lyft/l5kit/pull/167\" target=\"_blank\">https://github.com/lyft/l5kit/pull/167</a>. <br>\nIt has been said that the transformation is the bottleneck in the rasterizer as it is called a lot. We used the vectorized transformation whereever possible.</li>\n<li>In <code>box_rasterizer.py</code>, using python lists of small numpy arrays instead of one two large numpy arrays (<code>agent_images, ego_images</code>).<br>\nPython lists are surprisingly fast! <code>np.concatenate</code> on large numpy arrays is slow. Pre-allocation and writing to a single large array is even slower.<br>\nThe lists can be cast to a numpy array by <code>out_im = np.asarray(agents_images + ego_images, dtype=np.uint8)</code><br>\nThis way, boxes will stay in <code>uint8</code></li>\n<li>Moving concatenate to the GPU<br>\nAs <code>np.concatenate</code> is slow on large arrays, we do that on the GPU.<br>\nReplace <code>\"image\": image,</code> with <code>\"image_box\": image_box</code>, <code>\"image_sat\": image_sat</code>, <code>\"image_sem\": image_sem</code></li>\n<li>Speedups from <a href=\"https://github.com/lyft/l5kit/pull/140\" target=\"_blank\">https://github.com/lyft/l5kit/pull/140</a><br>\nWith the changes above, they now made a larger impact than the 7-8% stated in the PR.</li>\n</ol>\n<p>Speed after changes with semantic and satellite view (i5-3570K single thread, so ignore the actual time):</p>\n<pre><code>Raster-time per sample: 0.032 s\nHits         Time  Per Hit   % Time  Line Contents\n==================================================\n   9    2847054.0 316339.3     84.2  rasterizer.rasterize(history_frames, history_agents, history_tl_faces, selected_agent)\n</code></pre>\n<h3>Additions/Changes</h3>\n<ul>\n<li>New rasterizer which combines satellite and semantic view (<code>py_sem_sat</code>).</li>\n<li>Flag: AgentID used as a value for drawing (instead of just boolean <code>no agent = 0</code>, <code>agent = 1</code>)</li>\n<li>Flag: AgentLabelProba used as a value for drawing (instead of just boolean <code>no agent = 0</code>, <code>agent = 1</code>)</li>\n<li>Flag: Velocity used as a value for drawing (instead of just boolean <code>no agent = 0</code>, <code>agent = 1</code>)</li>\n<li>Option to draw all objects on the semantic layer, combinable with the label probas from above.</li>\n<li>Filter options for agents: <code>th_extent_ratio, th_area, th_distance_av</code> to give a larger training dataset (or harder as you want to phrase it)</li>\n<li>Multiple parallel rasterizers for ensembling models that use different raster sizes</li>\n<li>Satellite view fix for non square shapes (had some weird transformation in it before)</li>\n</ul>\n<p><img src=\"https://i.imgur.com/G3g9lEJ.png\" alt=\"rasterized_image\"></p>\n<p>We will make our L5kit code changes public in the next few days and after cleanup. Most likely with PRs to the original repo.</p>\n<h2>Journey &amp; Experiments</h2>\n<p>With all the speed improvements from our custom L5kit we were able to eliminate the previous CPU bottleneck and run experiments on an efficient multi-GPU setup. We mostly trained on 8xV100 nodes using pytorchs awesome distributed data parallel (DDP), which enables to effectivly scale a training script to multiple GPUs. This efficient training setup was key to run quite a few experiments. (See list of stuff that did not work below) In all our experiments, we used the chopped validation set for local validation as discussed <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/188695\" target=\"_blank\">here</a>. This yielded us very good prediction of the LB score (in the sub 13 range, the score was usually ~0.2 higher than local validation with a standard deviation of around 0.3). </p>\n<p>Originating from the provided baseline and the multi-mode prediction from <a href=\"https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence\" target=\"_blank\">here</a>, we experimented with raster size, pooling layers, head-modifications, usage of different backbones, learning rate scheduler, subsampling and oversampling to name a few. It was a major challenge in this competition, that experiments had to be run for at least 50 % of a full epoch of train_full.zarr or when comparing different sample sizes or learning rate scheduler even until the end to be really comparable. Subsampling the train set or stepwise decrease of the LR quickly yielded a low training and validation loss, but in the end, the performance was worse.</p>\n<p>Quite quickly, we decided to always use all data, and settled to a setup using shuffled train_full.zarr (with a fixed seed to be able to resume training) with a linear decay learning rate scheduler. To keep the training sample fixed, we also settled to an AgentDataset with <code>min_frame_history=1</code> and <code>min_frame_future=10.0</code> to best resemble the test dataset. Opposed to the many discussions about a much lower training loss, we never experienced that. Without any augmentation, the training loss is of course slightly lower, especially in the end of the epoch, but usually within 80% of the validation score. We found that larger backbones help to reach lower loss levels, but starting at the size of an EfficientNetB7, the nets seem to be overconfident with their predictions and the LB score was sometimes shaky. </p>\n<p>The first half of test is public, second half is private test. We couldn't identify any major differences in the two sets. With that knowledge it was possible to hide the true score (setting some public rows to zeros), we wonder how many teams did that and were afraid of surprises in the private LB due to this. </p>\n<h2>Best Single Model</h2>\n<p><img src=\"https://i.imgur.com/YC1WGD1.png\" alt=\"model\"></p>\n<p>Our best single model is actually one that not fully finished on the last day. It's just an EfficientNetB3 with a Linear Layer Head and dropout attached trained on an extended train_full.zarr Dataset (<code>min_frame_history=1, min_frame_future=5.0, th_distance_av=80, th_extent_ratio=1.6, th_area=0.45</code>)</p>\n<ul>\n<li>history_num_frames: 30.0</li>\n<li>raster_size: [448, 224]</li>\n<li>pixel_size: [0.1875, 0.1875]</li>\n<li>rasterizer: py_sem_sat</li>\n<li>batch_size: 64</li>\n<li>dropout: 0.3</li>\n</ul>\n<p>The model was pre-trained for 4 epochs on different lower resolution images (starting with 112, 64 and pixel size 0.75, 0.75) , so you can argue this is also an ensemble of some kind. The 5th epoch was trained with the parameters above and a very customized learning rate schedule that starts with a consine anneal and then transitions to a stepwise lr scheduler as seen in the graph below. The last evaluation score is 9.697.</p>\n<p><img src=\"https://i.imgur.com/4Q2PkvX.png\" alt=\"best_single\"></p>\n<p>The model achieves a public LB score of 9.776 and a private LB score of 9.070. With that, it would also have ranked 1st place this competition on it's own.</p>\n<p>Pretraining is VERY important! We discovered it by a mistake where we loaded an EfficientNet without pretrained weights from imagenet. It performed significantly worse by ~1.00.</p>\n<h2>Ensemble</h2>\n<p>Ensembling proved to be challenging in this competition, as traditional methods are not working well with the metric in use. We tried a lot of manual blending and form of post-processing without success. We then started building stacking models taking the raw predictions of the models as input, but this was also surprisingly not working well. At about the same time as <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> posted it <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/193908#1077899\" target=\"_blank\">here</a>, we had the idea of using the features/embeddings from each model as an input to a ensemble head. This worked much better and we could improve upon our single models specifically as we could introduce some diversity into the blend (different image size, pixel sizes, etc.) We also went with the dropout + single linear layer approach here, and all modification we tried with the head performed worse.</p>\n<p><img src=\"https://i.imgur.com/2I3M2rD.png\" alt=\"ensemble_arch\"></p>\n<p>To prevent overfitting, the ensembling was done on the chopped validation set (without using the chopped off actual validation part), in some experiments extended to validation + test set without noticeable change in the metric but with about twice the runtime. </p>\n<p>Ensembles of two similar models (even just two different checkpoints) already gave a boost of about 0.2 in the metric. Naturally, with more diverse models the ensemble performed even better and we were able to achieve our best public LB score of 9.319 and private LB score of 8.579 with an ensemble of 4 models (B3, B5, B6, B6) on two different raster sizes. </p>\n<h2>Bootstrapped Validation</h2>\n<p>We evaluated model and ensemble performance with bootstrap method on the chopped dataset. For more diverse ensembles we also got the lowest standard deviation from the bootstrap method. For our final submission it was 0.22754. The best single models had a standard deviation of ~0.27. We also identified our B7 and B8 models to have a slightly larger standard deviation in the bootstrap method of ~0.36-0.39 which may also explain their partial performance degredations on the public LB. Despite their good validation score, for that reason we excluded them in our final ensembles and final submissions. </p>\n<p><img src=\"https://i.imgur.com/l4vEAaz.png\" alt=\"bootstrap\"></p>\n<p>We plotted CV vs public LB for our submissions and found a spread of +-0.3 for most of our single models (green corridor) with just the B7 and B8 models outside of that corridor (not shown in the picture). Our ensembles are in an even smaller corridor with a spread of +-0.1 (orange corridor), proving their superior robustness and generalization capabilities over single models. </p>\n<p><img src=\"https://i.imgur.com/S6fYsBI.png\" alt=\"cv_lb\"></p>\n<p>Given this robust CV &amp; LB correlation, we were very confident that improvements in validation also lead to improvements on the leaderboard, so we did not have to sub it. So the sudden jump on public leaderboard only happened after we decided to sub our better models and was not some sudden magic we discovered.</p>\n<h2>What Didn't Work</h2>\n<ul>\n<li>Flag: AgentID used as a value for drawing (instead of just boolean <code>no agent = 0</code>, <code>agent = 1</code>)</li>\n<li>Flag: AgentLabelProba used as a value for drawing (instead of just boolean <code>no agent = 0</code>, <code>agent = 1</code>)</li>\n<li>Flag: AgentLabelProba used as a value for drawing (instead of just boolean <code>no agent = 0</code>, <code>agent = 1</code>)</li>\n<li>Option to draw all objects on the semantic layer, combinable with the label probas from above.</li>\n<li>Kalman filtering the output</li>\n<li>Any non NN prediction method</li>\n<li>Prediction of only the difference to a constant velocity model</li>\n<li>Using additional informations, such as velocity, history positions, label probas or anything else we could find in the head of the models. </li>\n<li>Heavier heads</li>\n<li>3D CNNs</li>\n<li>RNN based approaches on history frames for backbone input or outbut </li>\n<li>Using additional features from shallower layers of the backbone</li>\n<li>Ensembling with kmeans, or other analytic clustering method that we tried. Some success at scores down to around 14, but no improvements below that. </li>\n<li>Subsampling the training set. Huge speedups, but accuracy was always slightly lower</li>\n<li>TTA with modified yaw angles</li>\n<li>Unfreezing the backbone after ensembling and continue training with a second epoch</li>\n</ul>\n<h2>What We Didn't Try</h2>\n<ul>\n<li>Augmentation</li>\n<li>Multi backbone, multi rasterizer training</li>\n</ul>",
      "rawMarkdown": "# 1st Place Solution & L5Kit Speedup\n\n## Introduction\n\nFirst of all, I would like to thank Lyft for this great challenge and of course for their huge dataset. I must say I was a bit overwhelmed by the sheer amount of agents in the training set, but it was a fun experience to be able to train without ever using training data twice. Well, we did, when we used pre-trained models, but more on that later.\n\nI also want to thank my Teammates @philippsinger, @christofhenkel and @nvnnghia for the great collaborative teamwork during the competition. We worked well together with a pipeline consisting of two repositories (version controlled custom l5kit and a training repo) and used logging in neptune.ai to keep track of the experiments. \n\n## TL;DR\n\nOur solution is an ensemble of 4 efficientnet models trained with different rasterization parameters indivdually on train_full.zarr and stacked on validation.zarr. Key to training that many models on this huge dataset were several key improvements to the original l5kit repository to remove CPU bottleneck and enable efficient multi-GPU training using pytorch. \n\n## Improvements/Changes to the L5Kit\n\nOne of the major challenges in this competition was the slow rasterizer. We profiled the l5kit a lot and pinpointed the bottlenecks to speed it up by a factor of 4+. After that, our experiments were mostly GPU limited and multi-GPU experiments allowed us for \"fast\" iterations. Running through train_full (191_177_863+ samples) was possible within 2 days with medium sized models. \n\n### Speedups\n\nSpeed before changes with semantic view only (i5-3570K single thread, so ignore the actual time):\n\n```\nRaster-time per sample: 0.124 s\nHits         Time  Per Hit   % Time  Line Contents\n==================================================\n9   11147507.0 1238611.9     93.6    rasterizer.rasterize(history_frames, history_agents, history_tl_faces, selected_agent)\n```\n\nAll the speedups were achieved with plain python, no C++. The main speedups came from:\n\n1. Batched transformation of boxes with modifications from https://github.com/lyft/l5kit/pull/167. \n   It has been said that the transformation is the bottleneck in the rasterizer as it is called a lot. We used the vectorized transformation whereever possible.\n2. In `box_rasterizer.py`, using python lists of small numpy arrays instead of one two large numpy arrays (`agent_images, ego_images`).\n   Python lists are surprisingly fast! `np.concatenate` on large numpy arrays is slow. Pre-allocation and writing to a single large array is even slower.\n   The lists can be cast to a numpy array by `out_im = np.asarray(agents_images + ego_images, dtype=np.uint8)`\n   This way, boxes will stay in `uint8`\n3. Moving concatenate to the GPU\n   As `np.concatenate` is slow on large arrays, we do that on the GPU.\n   Replace `\"image\": image,` with `\"image_box\": image_box`, `\"image_sat\": image_sat`, `\"image_sem\": image_sem`\n4. Speedups from https://github.com/lyft/l5kit/pull/140\n   With the changes above, they now made a larger impact than the 7-8% stated in the PR.\n\nSpeed after changes with semantic and satellite view (i5-3570K single thread, so ignore the actual time):\n\n```\nRaster-time per sample: 0.032 s\nHits         Time  Per Hit   % Time  Line Contents\n==================================================\n   9    2847054.0 316339.3     84.2  rasterizer.rasterize(history_frames, history_agents, history_tl_faces, selected_agent)\n```\n\n\n### Additions/Changes\n\n- New rasterizer which combines satellite and semantic view (`py_sem_sat`).\n- Flag: AgentID used as a value for drawing (instead of just boolean `no agent = 0`, `agent = 1`)\n- Flag: AgentLabelProba used as a value for drawing (instead of just boolean `no agent = 0`, `agent = 1`)\n- Flag: Velocity used as a value for drawing (instead of just boolean `no agent = 0`, `agent = 1`)\n- Option to draw all objects on the semantic layer, combinable with the label probas from above.\n- Filter options for agents: `th_extent_ratio, th_area, th_distance_av` to give a larger training dataset (or harder as you want to phrase it)\n- Multiple parallel rasterizers for ensembling models that use different raster sizes\n- Satellite view fix for non square shapes (had some weird transformation in it before)\n\n![rasterized_image](https://i.imgur.com/G3g9lEJ.png)\n\nWe will make our L5kit code changes public in the next few days and after cleanup. Most likely with PRs to the original repo.\n\n## Journey & Experiments\n\nWith all the speed improvements from our custom L5kit we were able to eliminate the previous CPU bottleneck and run experiments on an efficient multi-GPU setup. We mostly trained on 8xV100 nodes using pytorchs awesome distributed data parallel (DDP), which enables to effectivly scale a training script to multiple GPUs. This efficient training setup was key to run quite a few experiments. (See list of stuff that did not work below) In all our experiments, we used the chopped validation set for local validation as discussed [here](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/188695). This yielded us very good prediction of the LB score (in the sub 13 range, the score was usually ~0.2 higher than local validation with a standard deviation of around 0.3). \n\nOriginating from the provided baseline and the multi-mode prediction from [here](https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence), we experimented with raster size, pooling layers, head-modifications, usage of different backbones, learning rate scheduler, subsampling and oversampling to name a few. It was a major challenge in this competition, that experiments had to be run for at least 50 % of a full epoch of train_full.zarr or when comparing different sample sizes or learning rate scheduler even until the end to be really comparable. Subsampling the train set or stepwise decrease of the LR quickly yielded a low training and validation loss, but in the end, the performance was worse.\n\nQuite quickly, we decided to always use all data, and settled to a setup using shuffled train_full.zarr (with a fixed seed to be able to resume training) with a linear decay learning rate scheduler. To keep the training sample fixed, we also settled to an AgentDataset with `min_frame_history=1` and `min_frame_future=10.0` to best resemble the test dataset. Opposed to the many discussions about a much lower training loss, we never experienced that. Without any augmentation, the training loss is of course slightly lower, especially in the end of the epoch, but usually within 80% of the validation score. We found that larger backbones help to reach lower loss levels, but starting at the size of an EfficientNetB7, the nets seem to be overconfident with their predictions and the LB score was sometimes shaky. \n\nThe first half of test is public, second half is private test. We couldn't identify any major differences in the two sets. With that knowledge it was possible to hide the true score (setting some public rows to zeros), we wonder how many teams did that and were afraid of surprises in the private LB due to this. \n\n## Best Single Model\n\n![model](https://i.imgur.com/YC1WGD1.png)\n\nOur best single model is actually one that not fully finished on the last day. It's just an EfficientNetB3 with a Linear Layer Head and dropout attached trained on an extended train_full.zarr Dataset (`min_frame_history=1, min_frame_future=5.0, th_distance_av=80, th_extent_ratio=1.6, th_area=0.45`)\n\n- history_num_frames: 30.0\n- raster_size: [448, 224]\n- pixel_size: [0.1875, 0.1875]\n- rasterizer: py_sem_sat\n- batch_size: 64\n- dropout: 0.3\n\nThe model was pre-trained for 4 epochs on different lower resolution images (starting with 112, 64 and pixel size 0.75, 0.75) , so you can argue this is also an ensemble of some kind. The 5th epoch was trained with the parameters above and a very customized learning rate schedule that starts with a consine anneal and then transitions to a stepwise lr scheduler as seen in the graph below. The last evaluation score is 9.697.\n\n![best_single](https://i.imgur.com/4Q2PkvX.png)\n\nThe model achieves a public LB score of 9.776 and a private LB score of 9.070. With that, it would also have ranked 1st place this competition on it's own.\n\nPretraining is VERY important! We discovered it by a mistake where we loaded an EfficientNet without pretrained weights from imagenet. It performed significantly worse by ~1.00.\n\n## Ensemble\n\nEnsembling proved to be challenging in this competition, as traditional methods are not working well with the metric in use. We tried a lot of manual blending and form of post-processing without success. We then started building stacking models taking the raw predictions of the models as input, but this was also surprisingly not working well. At about the same time as @hengck23 posted it [here](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/193908#1077899), we had the idea of using the features/embeddings from each model as an input to a ensemble head. This worked much better and we could improve upon our single models specifically as we could introduce some diversity into the blend (different image size, pixel sizes, etc.) We also went with the dropout + single linear layer approach here, and all modification we tried with the head performed worse.\n\n![ensemble_arch](https://i.imgur.com/2I3M2rD.png)\n\nTo prevent overfitting, the ensembling was done on the chopped validation set (without using the chopped off actual validation part), in some experiments extended to validation + test set without noticeable change in the metric but with about twice the runtime. \n\nEnsembles of two similar models (even just two different checkpoints) already gave a boost of about 0.2 in the metric. Naturally, with more diverse models the ensemble performed even better and we were able to achieve our best public LB score of 9.319 and private LB score of 8.579 with an ensemble of 4 models (B3, B5, B6, B6) on two different raster sizes. \n\n## Bootstrapped Validation\n\nWe evaluated model and ensemble performance with bootstrap method on the chopped dataset. For more diverse ensembles we also got the lowest standard deviation from the bootstrap method. For our final submission it was 0.22754. The best single models had a standard deviation of ~0.27. We also identified our B7 and B8 models to have a slightly larger standard deviation in the bootstrap method of ~0.36-0.39 which may also explain their partial performance degredations on the public LB. Despite their good validation score, for that reason we excluded them in our final ensembles and final submissions. \n\n![bootstrap](https://i.imgur.com/l4vEAaz.png)\n\nWe plotted CV vs public LB for our submissions and found a spread of +-0.3 for most of our single models (green corridor) with just the B7 and B8 models outside of that corridor (not shown in the picture). Our ensembles are in an even smaller corridor with a spread of +-0.1 (orange corridor), proving their superior robustness and generalization capabilities over single models. \n\n![cv_lb](https://i.imgur.com/S6fYsBI.png)\n\nGiven this robust CV & LB correlation, we were very confident that improvements in validation also lead to improvements on the leaderboard, so we did not have to sub it. So the sudden jump on public leaderboard only happened after we decided to sub our better models and was not some sudden magic we discovered.\n\n## What Didn't Work\n\n- Flag: AgentID used as a value for drawing (instead of just boolean `no agent = 0`, `agent = 1`)\n- Flag: AgentLabelProba used as a value for drawing (instead of just boolean `no agent = 0`, `agent = 1`)\n- Flag: AgentLabelProba used as a value for drawing (instead of just boolean `no agent = 0`, `agent = 1`)\n- Option to draw all objects on the semantic layer, combinable with the label probas from above.\n- Kalman filtering the output\n- Any non NN prediction method\n- Prediction of only the difference to a constant velocity model\n- Using additional informations, such as velocity, history positions, label probas or anything else we could find in the head of the models. \n- Heavier heads\n- 3D CNNs\n- RNN based approaches on history frames for backbone input or outbut \n- Using additional features from shallower layers of the backbone\n- Ensembling with kmeans, or other analytic clustering method that we tried. Some success at scores down to around 14, but no improvements below that. \n- Subsampling the training set. Huge speedups, but accuracy was always slightly lower\n- TTA with modified yaw angles\n- Unfreezing the backbone after ensembling and continue training with a second epoch\n\n## What We Didn't Try\n\n- Augmentation\n- Multi backbone, multi rasterizer training",
      "votes": 82
    },
    {
      "id": 1102803,
      "postDate": "2020-12-05T10:33:23.437Z",
      "content": "<p>Very nice and clean work, kudos. It should be in textbooks on how to approach a DS problem correctly.</p>\n<p>I am very interested to compare your stacking vs <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199636\" target=\"_blank\">my approach</a>. Can it be possible for you to provide me with <code>submission.csv</code> files for single model predictions of your 4 models? I will make all the results public.</p>",
      "rawMarkdown": "Very nice and clean work, kudos. It should be in textbooks on how to approach a DS problem correctly.\n\nI am very interested to compare your stacking vs [my approach](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199636). Can it be possible for you to provide me with `submission.csv` files for single model predictions of your 4 models? I will make all the results public.",
      "votes": 10,
      "replies": [
        {
          "id": 1103447,
          "postDate": "2020-12-05T22:55:33.063Z",
          "content": "<p>Thank you for the praise.<br>\nI uploaded 4 submissions that are very close to the checkpoints we have used in the final ensemble. The difference should be marginal. You can find them <a href=\"https://www.kaggle.com/ilu000/lyft-submission-csvs\" target=\"_blank\">here</a><br>\nNote: first 20 rows in some submissions files are set to zero, so ignore their public LB score.</p>",
          "rawMarkdown": "Thank you for the praise.\nI uploaded 4 submissions that are very close to the checkpoints we have used in the final ensemble. The difference should be marginal. You can find them [here](https://www.kaggle.com/ilu000/lyft-submission-csvs)\nNote: first 20 rows in some submissions files are set to zero, so ignore their public LB score.",
          "votes": 4
        },
        {
          "id": 1103887,
          "postDate": "2020-12-06T12:05:33.337Z",
          "content": "<p>Thanks a lot, I got the following result<br>\n<img src=\"https://imgur.com/nlTVvUv.png\" alt=\"pic\"></p>\n<p>My analytical ensembling: 8.484, your stacking approach: 8.579<br>\nI used uniform weights 25% to each sub. I will share my ensembling code later today!</p>",
          "rawMarkdown": "Thanks a lot, I got the following result\n![pic](https://imgur.com/nlTVvUv.png)\n\nMy analytical ensembling: 8.484, your stacking approach: 8.579\nI used uniform weights 25% to each sub. I will share my ensembling code later today!",
          "votes": 8
        },
        {
          "id": 1104357,
          "postDate": "2020-12-06T20:57:22.750Z",
          "content": "<p>Awesome! Great to see an analytic method to be on par or even superior to our method. <br>\nCould you try to weight the B3 model a bit higher (it has the best score out of those). And I am very interested in your ensembling code or the specific execution. </p>",
          "rawMarkdown": "Awesome! Great to see an analytic method to be on par or even superior to our method. \nCould you try to weight the B3 model a bit higher (it has the best score out of those). And I am very interested in your ensembling code or the specific execution. ",
          "votes": 1
        },
        {
          "id": 1104383,
          "postDate": "2020-12-06T21:56:24.133Z",
          "content": "<p>I published the code here:<br>\n<a href=\"https://github.com/nosound2/lyft_mp_ensemble/blob/main/Ensembling.ipynb\" target=\"_blank\">github repo</a><br>\n(but I suspect that code is impossible to understand without explanations)</p>\n<p>And I tried to increase B3 weight, specifically used weights <code>[0.4,0.25,0.25,0.25]</code>, but it was worse 8.526. Probably too big of an increase.</p>\n<p><img src=\"https://imgur.com/cgX3YDf.png\" alt=\"\"></p>\n<p>In my final submissions I chose optimal weights based on the validation. </p>",
          "rawMarkdown": "I published the code here:\n[github repo](https://github.com/nosound2/lyft_mp_ensemble/blob/main/Ensembling.ipynb)\n(but I suspect that code is impossible to understand without explanations)\n\nAnd I tried to increase B3 weight, specifically used weights `[0.4,0.25,0.25,0.25]`, but it was worse 8.526. Probably too big of an increase.\n\n![](https://imgur.com/cgX3YDf.png)\n\nIn my final submissions I chose optimal weights based on the validation. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 1102771,
      "postDate": "2020-12-05T09:34:34.413Z",
      "content": "<p>Great team effort :D</p>",
      "rawMarkdown": "Great team effort :D",
      "votes": 10
    },
    {
      "id": 1197530,
      "postDate": "2021-02-12T08:12:18.783Z",
      "content": "<p>As we were asked to share our ensemble model, please see the snippet below:</p>\n<pre><code>import torch\nfrom torch import nn\nfrom typing import Dict\nimport timm\nfrom timm.models.layers.conv2d_same import Conv2dSame\nfrom utils.loss import pytorch_neg_multi_log_likelihood_batch\nfrom utils.poolings import GeM\n\n\nclass LMM(nn.Module):\n    def __init__(self, model_architecture, H=30, gem=False):\n        super().__init__()\n        self.H = H\n        num_history_channels = (self.H + 1) * 2\n        rgb_channels = 6\n        num_in_channels = rgb_channels + num_history_channels\n        self.future_len = 50\n        num_targets = 2 * self.future_len\n        self.num_modes = 3\n        self.num_preds = num_targets * self.num_modes\n        self.backbone = timm.create_model(model_architecture, pretrained=False)\n\n        if gem:\n            self.backbone.global_pool = GeM()\n\n        self.backbone.conv_stem = Conv2dSame(\n            num_in_channels,\n            self.backbone.conv_stem.out_channels,\n            kernel_size=self.backbone.conv_stem.kernel_size,\n            stride=self.backbone.conv_stem.stride,\n            padding=self.backbone.conv_stem.padding,\n            bias=False,\n        )\n\n        self.backbone_out_features = self.backbone.classifier.in_features\n        self.backbone.classifier = nn.Sequential(\n            nn.Identity(),\n            nn.Linear(\n                in_features=self.backbone.classifier.in_features,\n                out_features=self.backbone_out_features,\n            ),\n        )\n\n        self.lin_head = nn.Sequential(\n            nn.ReLU(),\n            nn.Linear(\n                in_features=self.backbone_out_features,\n                out_features=self.num_preds + self.num_modes,\n            ),\n        )\n\n        for param in self.parameters():\n            param.requires_grad = False\n\n    def forward(self, image_box, image_sat, image_sem, history, history_availabilities):\n        x = torch.cat(\n            (\n                image_box[:, : self.H + 1].float(),\n                image_box[:, 31 : 31 + self.H + 1].float(),\n                image_sat,\n                image_sem,\n            ),\n            1,\n        )\n        x = self.backbone(x)\n        x = self.lin_head(x)\n        return x\n\n\nclass LyftMultiModel(nn.Module):\n    def __init__(self, cfg: Dict, PARAMS, num_modes=3):\n        super().__init__()\n        self.future_len = cfg[\"model_params\"][\"future_num_frames\"]\n        num_targets = 2 * self.future_len\n        self.num_preds = num_targets * num_modes\n        self.num_modes = num_modes\n        self.PARAMS = PARAMS\n\n        self.model1 = LMM(\"tf_efficientnet_b6_ns\")\n        sd = torch.load(\"output/preds/config_ch_aws_2_checkpoint.pth\")[\"model\"]\n        self.model1.load_state_dict(sd, strict=True)\n        self.model1.lin_head = nn.Identity()\n\n        self.model3 = LMM(\"tf_efficientnet_b6_ns\")\n        sd = torch.load(\"output/preds/config_ch_aws_1_checkpoint.pth\")[\"model\"]\n        self.model3.load_state_dict(sd, strict=True)\n        backbone_out_features = self.model3.backbone_out_features\n        self.model3.lin_head = nn.Identity()\n\n        self.model4 = LMM(\"tf_efficientnet_b3_ns\", gem=True)\n        sd = torch.load(\"output/preds/philipp_config_17_ep5_checkpoint_lr5e6.pth\")[\"model\"]\n        self.model4.load_state_dict(sd, strict=True)\n        backbone_out_features_2 = self.model4.backbone_out_features\n        self.model4.lin_head = nn.Identity()\n\n        self.model5 = LMM(\"tf_efficientnet_b5_ns\", H=8)\n        sd = torch.load(\"output/preds/config_ch_ddp_16_checkpoint.pth\")[\"model\"]\n        self.model5.load_state_dict(sd)\n        backbone_out_features3 = self.model5.backbone_out_features\n        self.model5.lin_head = nn.Identity()\n\n        self.lin_head = nn.Sequential(\n            nn.ReLU(),\n            nn.Linear(\n                in_features=backbone_out_features * 2\n                + backbone_out_features_2\n                + backbone_out_features3,\n                out_features=self.num_preds + num_modes,\n            ),\n        )\n\n    def forward(\n        self,\n        image_box,\n        image_sat,\n        image_sem,\n        image_box2,\n        image_sat2,\n        image_sem2,\n        history,\n        history_availabilities,\n    ):\n\n        self.model1 = self.model1.eval()\n        self.model3 = self.model3.eval()\n        self.model4 = self.model4.eval()\n        self.model5 = self.model5.eval()\n\n        x1 = self.model1(image_box, image_sat, image_sem, history, history_availabilities)\n        x3 = self.model3(image_box, image_sat, image_sem, history, history_availabilities)\n        x4 = self.model4(image_box2, image_sat2, image_sem2, history, history_availabilities)\n        x5 = self.model5(image_box2, image_sat2, image_sem2, history, history_availabilities)\n\n        x = self.lin_head(torch.cat([x1, x3, x4, x5], dim=-1))\n\n        bs, _ = x.shape\n        pred, confidences = torch.split(x, self.num_preds, dim=1)\n        pred = pred.view(bs, self.num_modes, self.future_len, 2)\n        assert confidences.shape == (bs, self.num_modes)\n        confidences = torch.softmax(confidences, dim=1)\n        return pred, confidences\n\n\nclass LyftMultiRegressor(nn.Module):\n    \"\"\"Single mode prediction\"\"\"\n\n    def __init__(self, predictor, PARAMS):\n        super().__init__()\n        self.predictor = predictor\n        self.PARAMS = PARAMS\n\n    def forward(\n        self,\n        image_box,\n        image_sat,\n        image_sem,\n        image_box2,\n        image_sat2,\n        image_sem2,\n        targets,\n        target_availabilities,\n        history,\n        history_availabilities,\n        velocity,\n    ):\n        pred, confidences = self.predictor(\n            image_box,\n            image_sat,\n            image_sem,\n            image_box2,\n            image_sat2,\n            image_sem2,\n            history,\n            history_availabilities,\n        )\n\n        if self.PARAMS.predict_diffs:\n            pred = torch.cumsum(pred, dim=2)\n\n        loss_nll = pytorch_neg_multi_log_likelihood_batch(\n            targets, pred, confidences, target_availabilities\n        )\n\n        return loss_nll, pred, confidences\n</code></pre>",
      "rawMarkdown": "As we were asked to share our ensemble model, please see the snippet below:\n\n```\nimport torch\nfrom torch import nn\nfrom typing import Dict\nimport timm\nfrom timm.models.layers.conv2d_same import Conv2dSame\nfrom utils.loss import pytorch_neg_multi_log_likelihood_batch\nfrom utils.poolings import GeM\n\n\nclass LMM(nn.Module):\n    def __init__(self, model_architecture, H=30, gem=False):\n        super().__init__()\n        self.H = H\n        num_history_channels = (self.H + 1) * 2\n        rgb_channels = 6\n        num_in_channels = rgb_channels + num_history_channels\n        self.future_len = 50\n        num_targets = 2 * self.future_len\n        self.num_modes = 3\n        self.num_preds = num_targets * self.num_modes\n        self.backbone = timm.create_model(model_architecture, pretrained=False)\n\n        if gem:\n            self.backbone.global_pool = GeM()\n\n        self.backbone.conv_stem = Conv2dSame(\n            num_in_channels,\n            self.backbone.conv_stem.out_channels,\n            kernel_size=self.backbone.conv_stem.kernel_size,\n            stride=self.backbone.conv_stem.stride,\n            padding=self.backbone.conv_stem.padding,\n            bias=False,\n        )\n\n        self.backbone_out_features = self.backbone.classifier.in_features\n        self.backbone.classifier = nn.Sequential(\n            nn.Identity(),\n            nn.Linear(\n                in_features=self.backbone.classifier.in_features,\n                out_features=self.backbone_out_features,\n            ),\n        )\n\n        self.lin_head = nn.Sequential(\n            nn.ReLU(),\n            nn.Linear(\n                in_features=self.backbone_out_features,\n                out_features=self.num_preds + self.num_modes,\n            ),\n        )\n\n        for param in self.parameters():\n            param.requires_grad = False\n\n    def forward(self, image_box, image_sat, image_sem, history, history_availabilities):\n        x = torch.cat(\n            (\n                image_box[:, : self.H + 1].float(),\n                image_box[:, 31 : 31 + self.H + 1].float(),\n                image_sat,\n                image_sem,\n            ),\n            1,\n        )\n        x = self.backbone(x)\n        x = self.lin_head(x)\n        return x\n\n\nclass LyftMultiModel(nn.Module):\n    def __init__(self, cfg: Dict, PARAMS, num_modes=3):\n        super().__init__()\n        self.future_len = cfg[\"model_params\"][\"future_num_frames\"]\n        num_targets = 2 * self.future_len\n        self.num_preds = num_targets * num_modes\n        self.num_modes = num_modes\n        self.PARAMS = PARAMS\n\n        self.model1 = LMM(\"tf_efficientnet_b6_ns\")\n        sd = torch.load(\"output/preds/config_ch_aws_2_checkpoint.pth\")[\"model\"]\n        self.model1.load_state_dict(sd, strict=True)\n        self.model1.lin_head = nn.Identity()\n\n        self.model3 = LMM(\"tf_efficientnet_b6_ns\")\n        sd = torch.load(\"output/preds/config_ch_aws_1_checkpoint.pth\")[\"model\"]\n        self.model3.load_state_dict(sd, strict=True)\n        backbone_out_features = self.model3.backbone_out_features\n        self.model3.lin_head = nn.Identity()\n\n        self.model4 = LMM(\"tf_efficientnet_b3_ns\", gem=True)\n        sd = torch.load(\"output/preds/philipp_config_17_ep5_checkpoint_lr5e6.pth\")[\"model\"]\n        self.model4.load_state_dict(sd, strict=True)\n        backbone_out_features_2 = self.model4.backbone_out_features\n        self.model4.lin_head = nn.Identity()\n\n        self.model5 = LMM(\"tf_efficientnet_b5_ns\", H=8)\n        sd = torch.load(\"output/preds/config_ch_ddp_16_checkpoint.pth\")[\"model\"]\n        self.model5.load_state_dict(sd)\n        backbone_out_features3 = self.model5.backbone_out_features\n        self.model5.lin_head = nn.Identity()\n\n        self.lin_head = nn.Sequential(\n            nn.ReLU(),\n            nn.Linear(\n                in_features=backbone_out_features * 2\n                + backbone_out_features_2\n                + backbone_out_features3,\n                out_features=self.num_preds + num_modes,\n            ),\n        )\n\n    def forward(\n        self,\n        image_box,\n        image_sat,\n        image_sem,\n        image_box2,\n        image_sat2,\n        image_sem2,\n        history,\n        history_availabilities,\n    ):\n\n        self.model1 = self.model1.eval()\n        self.model3 = self.model3.eval()\n        self.model4 = self.model4.eval()\n        self.model5 = self.model5.eval()\n\n        x1 = self.model1(image_box, image_sat, image_sem, history, history_availabilities)\n        x3 = self.model3(image_box, image_sat, image_sem, history, history_availabilities)\n        x4 = self.model4(image_box2, image_sat2, image_sem2, history, history_availabilities)\n        x5 = self.model5(image_box2, image_sat2, image_sem2, history, history_availabilities)\n\n        x = self.lin_head(torch.cat([x1, x3, x4, x5], dim=-1))\n\n        bs, _ = x.shape\n        pred, confidences = torch.split(x, self.num_preds, dim=1)\n        pred = pred.view(bs, self.num_modes, self.future_len, 2)\n        assert confidences.shape == (bs, self.num_modes)\n        confidences = torch.softmax(confidences, dim=1)\n        return pred, confidences\n\n\nclass LyftMultiRegressor(nn.Module):\n    \"\"\"Single mode prediction\"\"\"\n\n    def __init__(self, predictor, PARAMS):\n        super().__init__()\n        self.predictor = predictor\n        self.PARAMS = PARAMS\n\n    def forward(\n        self,\n        image_box,\n        image_sat,\n        image_sem,\n        image_box2,\n        image_sat2,\n        image_sem2,\n        targets,\n        target_availabilities,\n        history,\n        history_availabilities,\n        velocity,\n    ):\n        pred, confidences = self.predictor(\n            image_box,\n            image_sat,\n            image_sem,\n            image_box2,\n            image_sat2,\n            image_sem2,\n            history,\n            history_availabilities,\n        )\n\n        if self.PARAMS.predict_diffs:\n            pred = torch.cumsum(pred, dim=2)\n\n        loss_nll = pytorch_neg_multi_log_likelihood_batch(\n            targets, pred, confidences, target_availabilities\n        )\n\n        return loss_nll, pred, confidences\n```",
      "votes": 4
    },
    {
      "id": 1103452,
      "postDate": "2020-12-05T23:15:23.503Z",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> for the write-up! and congrats for 1st place ( <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> <a href=\"https://www.kaggle.com/nvnnghia\" target=\"_blank\">@nvnnghia</a> ) !</p>\n<p>How much the satellite image affect to the score? Did you compare the score difference between <code>py_semantic</code> and your <code>py_sem_sat</code>.<br>\nI'm guessing that this part has less impact compared to other hyperparameters change, especially filter agent prob or image size/pixel size, min_frame_history to 5 etc. <br>\nWhich one was important to achieve public LB score less than 10?</p>",
      "rawMarkdown": "Thank you @ilu000 for the write-up! and congrats for 1st place ( @christofhenkel @philippsinger @nvnnghia ) !\n\nHow much the satellite image affect to the score? Did you compare the score difference between `py_semantic` and your `py_sem_sat`.\nI'm guessing that this part has less impact compared to other hyperparameters change, especially filter agent prob or image size/pixel size, min_frame_history to 5 etc. \nWhich one was important to achieve public LB score less than 10?",
      "votes": 3,
      "replies": [
        {
          "id": 1103455,
          "postDate": "2020-12-05T23:21:05.237Z",
          "content": "<p>I guess the the main difference with other teams is that you tried more bigger data, more heavy/rich input data training.<br>\nBy setting <code>min_frame_history=5</code>, or setting image size as <code>raster_size: [448, 224], pixel_size: [0.1875, 0.1875]</code>.</p>\n<p>These can be done because you improved the training speed a lot! Most of the team cannot come up training this much heavier training due to time limitation.</p>",
          "rawMarkdown": "I guess the the main difference with other teams is that you tried more bigger data, more heavy/rich input data training.\nBy setting `min_frame_history=5`, or setting image size as `raster_size: [448, 224], pixel_size: [0.1875, 0.1875]`.\n\nThese can be done because you improved the training speed a lot! Most of the team cannot come up training this much heavier training due to time limitation.",
          "votes": 2
        },
        {
          "id": 1103457,
          "postDate": "2020-12-05T23:28:26.903Z",
          "content": "<p>I am afraid we didn't do all ablation studies due to time constraints, so I can't tell the exact influence of the satellite view. It wasn't much, but with that many channels, we decided another three wouldn't hurt and may help in cases where the semantic view either fails to be drawn (happens sometimes) or in park like areas where cyclists do not drive on paved roads. </p>\n<p>Indeed, key to our solution was being able to do many experiments. We found that using even more data than whats provided with the standard settings (harder samples in this case, the train loss was significantly higher) helped to squeeze out a bit more accuracy w.r.t. the competition metric. </p>\n<p>Using more data was definitely one of the major improvements that we got. Even more improvement for a single model was gained by the pre-training on different raster-sizes. </p>",
          "rawMarkdown": "I am afraid we didn't do all ablation studies due to time constraints, so I can't tell the exact influence of the satellite view. It wasn't much, but with that many channels, we decided another three wouldn't hurt and may help in cases where the semantic view either fails to be drawn (happens sometimes) or in park like areas where cyclists do not drive on paved roads. \n\nIndeed, key to our solution was being able to do many experiments. We found that using even more data than whats provided with the standard settings (harder samples in this case, the train loss was significantly higher) helped to squeeze out a bit more accuracy w.r.t. the competition metric. \n\nUsing more data was definitely one of the major improvements that we got. Even more improvement for a single model was gained by the pre-training on different raster-sizes. ",
          "votes": 4
        },
        {
          "id": 1103481,
          "postDate": "2020-12-06T00:16:04.373Z",
          "content": "<p>I see, thank you for quick and informative reply!</p>\n<p>It seems you draw satellite image only outside of the road. We noticed that the parking area is sometimes drawn as \"road\" (e.g., parking area of Lyft office, Starbucks or fuel stands). So I'm not sure drawing only outside of the road gives enough information or not.</p>\n<p>Yeah, I noticed the possibility of using much larger dataset, but I could not execute it during the competition…</p>\n<p>\"pre-training was important\" is interesting finding! Maybe usual \"imagenet\" pretrained weight was not suitable in this task which need measure distance from semantic image, not to recognize natural images.</p>",
          "rawMarkdown": "I see, thank you for quick and informative reply!\n\nIt seems you draw satellite image only outside of the road. We noticed that the parking area is sometimes drawn as \"road\" (e.g., parking area of Lyft office, Starbucks or fuel stands). So I'm not sure drawing only outside of the road gives enough information or not.\n\nYeah, I noticed the possibility of using much larger dataset, but I could not execute it during the competition...\n\n\"pre-training was important\" is interesting finding! Maybe usual \"imagenet\" pretrained weight was not suitable in this task which need measure distance from semantic image, not to recognize natural images."
        },
        {
          "id": 1103494,
          "postDate": "2020-12-06T00:56:28.230Z",
          "content": "<p>no, we have the satellite image for each pixel. It's just 3 extra channels, but for simplicity we have shown an image with the semantic view over the satellite image in this Writeup. </p>\n<p>we have ego channels, agent channels, semantic channels, satellite channels</p>",
          "rawMarkdown": "no, we have the satellite image for each pixel. It's just 3 extra channels, but for simplicity we have shown an image with the semantic view over the satellite image in this Writeup. \n\nwe have ego channels, agent channels, semantic channels, satellite channels",
          "votes": 1
        },
        {
          "id": 1103973,
          "postDate": "2020-12-06T13:35:22.503Z",
          "content": "<p>I see, thank you for clarification!</p>\n<p>Sorry I have 1 another question, the best model was <code>history_num_frames=30</code>. Does that input have channels with 30*2 for ego/agent channels, maybe total 30 + 30 + 3 + 3 = 66 channels?</p>",
          "rawMarkdown": "I see, thank you for clarification!\n\nSorry I have 1 another question, the best model was `history_num_frames=30`. Does that input have channels with 30*2 for ego/agent channels, maybe total 30 + 30 + 3 + 3 = 66 channels?"
        },
        {
          "id": 1104367,
          "postDate": "2020-12-06T21:14:58.660Z",
          "content": "<p>Yes, 66 channels in that case</p>",
          "rawMarkdown": "Yes, 66 channels in that case",
          "votes": 1
        },
        {
          "id": 1104424,
          "postDate": "2020-12-06T23:29:33.907Z",
          "content": "<p>Wow that's very huge input! I see, thank you for answering questions :)</p>",
          "rawMarkdown": "Wow that's very huge input! I see, thank you for answering questions :)"
        },
        {
          "id": 1104851,
          "postDate": "2020-12-07T09:33:07.130Z",
          "content": "<p>I think less history frames are sufficient though, maybe 10. We just somehow started most experiments with 30 and never really questioned it until the end where we needed some speed gains for final models :)</p>",
          "rawMarkdown": "I think less history frames are sufficient though, maybe 10. We just somehow started most experiments with 30 and never really questioned it until the end where we needed some speed gains for final models :)",
          "votes": 2
        },
        {
          "id": 1105090,
          "postDate": "2020-12-07T14:11:07.550Z",
          "content": "<p>I see, then maybe agent threshold or min_frame_history=5 was the difference with other teams. Thanks for additional comment.</p>",
          "rawMarkdown": "I see, then maybe agent threshold or min_frame_history=5 was the difference with other teams. Thanks for additional comment."
        }
      ]
    },
    {
      "id": 1104365,
      "postDate": "2020-12-06T21:10:13.510Z",
      "content": "<p>That's when I read this kind of write-up that I know I still have a lot of to learn. </p>\n<p>Congratulations on the win, really impressive performance.</p>",
      "rawMarkdown": "That's when I read this kind of write-up that I know I still have a lot of to learn. \n\nCongratulations on the win, really impressive performance.",
      "votes": 4
    },
    {
      "id": 1102793,
      "postDate": "2020-12-05T10:10:28.493Z",
      "content": "<p>Amazing writeup and congratulations!</p>",
      "rawMarkdown": "Amazing writeup and congratulations!",
      "votes": 4
    },
    {
      "id": 1201855,
      "postDate": "2021-02-15T17:36:57.090Z",
      "content": "<p>A little late, <br>\ncongrats on 1st place and successful team effort.  <a href=\"https://www.kaggle.com/nvnnghia\" target=\"_blank\">@nvnnghia</a>, <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>, <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> <br>\nIt is one of the best sharing solutions in kaggle.</p>\n<p>Thanks for sharing and good write-up!</p>",
      "rawMarkdown": "A little late, \ncongrats on 1st place and successful team effort.  @nvnnghia, @philippsinger, @christofhenkel \nIt is one of the best sharing solutions in kaggle.\n\nThanks for sharing and good write-up!",
      "votes": 1
    },
    {
      "id": 1105881,
      "postDate": "2020-12-08T09:39:29.067Z",
      "content": "<p>Amazing speed up! Amazing work!<br>\nVery interesting way to do the validation. Never see a competition that has such small noise between CV and LB!<br>\nFor the normal distribution-like plot that you shown, is that the distribution of loss of all the samples in the validation set from a single model? <br>\nInteresting to see that you don't seem to have many outliers from your model. What was the largest single sample loss from your model?</p>",
      "rawMarkdown": "Amazing speed up! Amazing work!\nVery interesting way to do the validation. Never see a competition that has such small noise between CV and LB!\nFor the normal distribution-like plot that you shown, is that the distribution of loss of all the samples in the validation set from a single model? \nInteresting to see that you don't seem to have many outliers from your model. What was the largest single sample loss from your model?\n",
      "votes": 1,
      "replies": [
        {
          "id": 1105933,
          "postDate": "2020-12-08T10:46:02Z",
          "content": "<p>The distribution is from a bootstrap validation method where we randomly subsample (with multi-draw) 35k rows of the validation set. It's the histogramm over the scores of all subsamples, and not the scores of a single validation set. <br>\nWe did that, to measure the stability of the predictions (chance of shakeup).<br>\nThe largest single sample loss is ~2k as far as i remember. We all saw how skewed the score distribution is ;)</p>",
          "rawMarkdown": "The distribution is from a bootstrap validation method where we randomly subsample (with multi-draw) 35k rows of the validation set. It's the histogramm over the scores of all subsamples, and not the scores of a single validation set. \nWe did that, to measure the stability of the predictions (chance of shakeup).\nThe largest single sample loss is ~2k as far as i remember. We all saw how skewed the score distribution is ;)",
          "votes": 1
        },
        {
          "id": 1106548,
          "postDate": "2020-12-08T23:10:00.920Z",
          "content": "<p>I guess I am not very familiar with bootstrap validation. Is it like k-fold validation that you train one model on each subsample datasets? Or, it is just one model but evaluate on each subsample set?</p>\n<p>I think I saw like ~40k on my one of my ok model. ~2k is still way better than I got :)</p>",
          "rawMarkdown": "I guess I am not very familiar with bootstrap validation. Is it like k-fold validation that you train one model on each subsample datasets? Or, it is just one model but evaluate on each subsample set?\n\nI think I saw like ~40k on my one of my ok model. ~2k is still way better than I got :)"
        }
      ]
    },
    {
      "id": 1104528,
      "postDate": "2020-12-07T03:15:56.333Z",
      "content": "<p>thanks for the writeup and good work!</p>\n<p>What Didn't Work<br>\n\"Unfreezing the backbone after ensembling and continue training with a second epoch\"</p>\n<p>i experience this too. it overfits.<br>\nhowever, just unfreeze the layer 4 (last on only) of the embedding backbone is ok for low learning rate.</p>",
      "rawMarkdown": "thanks for the writeup and good work!\n\nWhat Didn't Work\n\"Unfreezing the backbone after ensembling and continue training with a second epoch\"\n\ni experience this too. it overfits.\nhowever, just unfreeze the layer 4 (last on only) of the embedding backbone is ok for low learning rate.\n",
      "votes": 1
    },
    {
      "id": 1102784,
      "postDate": "2020-12-05T09:57:51.750Z",
      "content": "<p>Congratulations, nice team effort.<br>\nThanks for sharing 👍</p>",
      "rawMarkdown": "Congratulations, nice team effort.\nThanks for sharing 👍",
      "votes": 1
    },
    {
      "id": 1103781,
      "postDate": "2020-12-06T09:12:11.140Z",
      "content": "<p>Wow. Well done to all of you - clinical execution!</p>",
      "rawMarkdown": "Wow. Well done to all of you - clinical execution!",
      "votes": 2
    },
    {
      "id": 1103709,
      "postDate": "2020-12-06T06:46:31.877Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/nvnnghia\" target=\"_blank\">@nvnnghia</a> <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> and <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> and thanks for sharing details solution! </p>",
      "rawMarkdown": "Congrats @nvnnghia @ilu000 @christofhenkel and @philippsinger and thanks for sharing details solution! ",
      "votes": 2
    },
    {
      "id": 1103213,
      "postDate": "2020-12-05T18:14:49.727Z",
      "content": "<p>Very well done, congratulations on the hard work !</p>",
      "rawMarkdown": "Very well done, congratulations on the hard work !"
    },
    {
      "id": 1102801,
      "postDate": "2020-12-05T10:31:55.070Z",
      "content": "<p>Amazing solution. is your solution improved better than others ,only by <code>py_sat_sem</code> or because of effnet or anything else. What is your Hardware anyway?</p>",
      "rawMarkdown": "Amazing solution. is your solution improved better than others ,only by `py_sat_sem` or because of effnet or anything else. What is your Hardware anyway?"
    },
    {
      "id": 1197935,
      "postDate": "2021-02-12T14:52:02.343Z",
      "rawMarkdown": "",
      "votes": -6,
      "isDeleted": true
    },
    {
      "id": 1104947,
      "postDate": "2020-12-07T11:51:15.653Z",
      "content": "<p>Inspiring work! thanks for sharing</p>",
      "rawMarkdown": "Inspiring work! thanks for sharing"
    }
  ],
  "comments": [
    {
      "id": 1102803,
      "author_name": "nosound",
      "author_url": "",
      "post_date": "2020-12-05T10:33:23.437000",
      "content": "<p>Very nice and clean work, kudos. It should be in textbooks on how to approach a DS problem correctly.</p>\n<p>I am very interested to compare your stacking vs <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199636\" target=\"_blank\">my approach</a>. Can it be possible for you to provide me with <code>submission.csv</code> files for single model predictions of your 4 models? I will make all the results public.</p>",
      "votes": 10,
      "replies": [
        {
          "id": 1103447,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-12-05T22:55:33.063000",
          "content": "<p>Thank you for the praise.<br>\nI uploaded 4 submissions that are very close to the checkpoints we have used in the final ensemble. The difference should be marginal. You can find them <a href=\"https://www.kaggle.com/ilu000/lyft-submission-csvs\" target=\"_blank\">here</a><br>\nNote: first 20 rows in some submissions files are set to zero, so ignore their public LB score.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1103887,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2020-12-06T12:05:33.337000",
          "content": "<p>Thanks a lot, I got the following result<br>\n<img src=\"https://imgur.com/nlTVvUv.png\" alt=\"pic\"></p>\n<p>My analytical ensembling: 8.484, your stacking approach: 8.579<br>\nI used uniform weights 25% to each sub. I will share my ensembling code later today!</p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 1104357,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-12-06T20:57:22.750000",
          "content": "<p>Awesome! Great to see an analytic method to be on par or even superior to our method. <br>\nCould you try to weight the B3 model a bit higher (it has the best score out of those). And I am very interested in your ensembling code or the specific execution. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1104383,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2020-12-06T21:56:24.133000",
          "content": "<p>I published the code here:<br>\n<a href=\"https://github.com/nosound2/lyft_mp_ensemble/blob/main/Ensembling.ipynb\" target=\"_blank\">github repo</a><br>\n(but I suspect that code is impossible to understand without explanations)</p>\n<p>And I tried to increase B3 weight, specifically used weights <code>[0.4,0.25,0.25,0.25]</code>, but it was worse 8.526. Probably too big of an increase.</p>\n<p><img src=\"https://imgur.com/cgX3YDf.png\" alt=\"\"></p>\n<p>In my final submissions I chose optimal weights based on the validation. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1102771,
      "author_name": "Dieter",
      "author_url": "",
      "post_date": "2020-12-05T09:34:34.413000",
      "content": "<p>Great team effort :D</p>",
      "votes": 10,
      "replies": []
    },
    {
      "id": 1197530,
      "author_name": "Pascal Pfeiffer",
      "author_url": "",
      "post_date": "2021-02-12T08:12:18.783000",
      "content": "<p>As we were asked to share our ensemble model, please see the snippet below:</p>\n<pre><code>import torch\nfrom torch import nn\nfrom typing import Dict\nimport timm\nfrom timm.models.layers.conv2d_same import Conv2dSame\nfrom utils.loss import pytorch_neg_multi_log_likelihood_batch\nfrom utils.poolings import GeM\n\n\nclass LMM(nn.Module):\n    def __init__(self, model_architecture, H=30, gem=False):\n        super().__init__()\n        self.H = H\n        num_history_channels = (self.H + 1) * 2\n        rgb_channels = 6\n        num_in_channels = rgb_channels + num_history_channels\n        self.future_len = 50\n        num_targets = 2 * self.future_len\n        self.num_modes = 3\n        self.num_preds = num_targets * self.num_modes\n        self.backbone = timm.create_model(model_architecture, pretrained=False)\n\n        if gem:\n            self.backbone.global_pool = GeM()\n\n        self.backbone.conv_stem = Conv2dSame(\n            num_in_channels,\n            self.backbone.conv_stem.out_channels,\n            kernel_size=self.backbone.conv_stem.kernel_size,\n            stride=self.backbone.conv_stem.stride,\n            padding=self.backbone.conv_stem.padding,\n            bias=False,\n        )\n\n        self.backbone_out_features = self.backbone.classifier.in_features\n        self.backbone.classifier = nn.Sequential(\n            nn.Identity(),\n            nn.Linear(\n                in_features=self.backbone.classifier.in_features,\n                out_features=self.backbone_out_features,\n            ),\n        )\n\n        self.lin_head = nn.Sequential(\n            nn.ReLU(),\n            nn.Linear(\n                in_features=self.backbone_out_features,\n                out_features=self.num_preds + self.num_modes,\n            ),\n        )\n\n        for param in self.parameters():\n            param.requires_grad = False\n\n    def forward(self, image_box, image_sat, image_sem, history, history_availabilities):\n        x = torch.cat(\n            (\n                image_box[:, : self.H + 1].float(),\n                image_box[:, 31 : 31 + self.H + 1].float(),\n                image_sat,\n                image_sem,\n            ),\n            1,\n        )\n        x = self.backbone(x)\n        x = self.lin_head(x)\n        return x\n\n\nclass LyftMultiModel(nn.Module):\n    def __init__(self, cfg: Dict, PARAMS, num_modes=3):\n        super().__init__()\n        self.future_len = cfg[\"model_params\"][\"future_num_frames\"]\n        num_targets = 2 * self.future_len\n        self.num_preds = num_targets * num_modes\n        self.num_modes = num_modes\n        self.PARAMS = PARAMS\n\n        self.model1 = LMM(\"tf_efficientnet_b6_ns\")\n        sd = torch.load(\"output/preds/config_ch_aws_2_checkpoint.pth\")[\"model\"]\n        self.model1.load_state_dict(sd, strict=True)\n        self.model1.lin_head = nn.Identity()\n\n        self.model3 = LMM(\"tf_efficientnet_b6_ns\")\n        sd = torch.load(\"output/preds/config_ch_aws_1_checkpoint.pth\")[\"model\"]\n        self.model3.load_state_dict(sd, strict=True)\n        backbone_out_features = self.model3.backbone_out_features\n        self.model3.lin_head = nn.Identity()\n\n        self.model4 = LMM(\"tf_efficientnet_b3_ns\", gem=True)\n        sd = torch.load(\"output/preds/philipp_config_17_ep5_checkpoint_lr5e6.pth\")[\"model\"]\n        self.model4.load_state_dict(sd, strict=True)\n        backbone_out_features_2 = self.model4.backbone_out_features\n        self.model4.lin_head = nn.Identity()\n\n        self.model5 = LMM(\"tf_efficientnet_b5_ns\", H=8)\n        sd = torch.load(\"output/preds/config_ch_ddp_16_checkpoint.pth\")[\"model\"]\n        self.model5.load_state_dict(sd)\n        backbone_out_features3 = self.model5.backbone_out_features\n        self.model5.lin_head = nn.Identity()\n\n        self.lin_head = nn.Sequential(\n            nn.ReLU(),\n            nn.Linear(\n                in_features=backbone_out_features * 2\n                + backbone_out_features_2\n                + backbone_out_features3,\n                out_features=self.num_preds + num_modes,\n            ),\n        )\n\n    def forward(\n        self,\n        image_box,\n        image_sat,\n        image_sem,\n        image_box2,\n        image_sat2,\n        image_sem2,\n        history,\n        history_availabilities,\n    ):\n\n        self.model1 = self.model1.eval()\n        self.model3 = self.model3.eval()\n        self.model4 = self.model4.eval()\n        self.model5 = self.model5.eval()\n\n        x1 = self.model1(image_box, image_sat, image_sem, history, history_availabilities)\n        x3 = self.model3(image_box, image_sat, image_sem, history, history_availabilities)\n        x4 = self.model4(image_box2, image_sat2, image_sem2, history, history_availabilities)\n        x5 = self.model5(image_box2, image_sat2, image_sem2, history, history_availabilities)\n\n        x = self.lin_head(torch.cat([x1, x3, x4, x5], dim=-1))\n\n        bs, _ = x.shape\n        pred, confidences = torch.split(x, self.num_preds, dim=1)\n        pred = pred.view(bs, self.num_modes, self.future_len, 2)\n        assert confidences.shape == (bs, self.num_modes)\n        confidences = torch.softmax(confidences, dim=1)\n        return pred, confidences\n\n\nclass LyftMultiRegressor(nn.Module):\n    \"\"\"Single mode prediction\"\"\"\n\n    def __init__(self, predictor, PARAMS):\n        super().__init__()\n        self.predictor = predictor\n        self.PARAMS = PARAMS\n\n    def forward(\n        self,\n        image_box,\n        image_sat,\n        image_sem,\n        image_box2,\n        image_sat2,\n        image_sem2,\n        targets,\n        target_availabilities,\n        history,\n        history_availabilities,\n        velocity,\n    ):\n        pred, confidences = self.predictor(\n            image_box,\n            image_sat,\n            image_sem,\n            image_box2,\n            image_sat2,\n            image_sem2,\n            history,\n            history_availabilities,\n        )\n\n        if self.PARAMS.predict_diffs:\n            pred = torch.cumsum(pred, dim=2)\n\n        loss_nll = pytorch_neg_multi_log_likelihood_batch(\n            targets, pred, confidences, target_availabilities\n        )\n\n        return loss_nll, pred, confidences\n</code></pre>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1103452,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2020-12-05T23:15:23.503000",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> for the write-up! and congrats for 1st place ( <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> <a href=\"https://www.kaggle.com/nvnnghia\" target=\"_blank\">@nvnnghia</a> ) !</p>\n<p>How much the satellite image affect to the score? Did you compare the score difference between <code>py_semantic</code> and your <code>py_sem_sat</code>.<br>\nI'm guessing that this part has less impact compared to other hyperparameters change, especially filter agent prob or image size/pixel size, min_frame_history to 5 etc. <br>\nWhich one was important to achieve public LB score less than 10?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1103455,
          "author_name": "corochann",
          "author_url": "",
          "post_date": "2020-12-05T23:21:05.237000",
          "content": "<p>I guess the the main difference with other teams is that you tried more bigger data, more heavy/rich input data training.<br>\nBy setting <code>min_frame_history=5</code>, or setting image size as <code>raster_size: [448, 224], pixel_size: [0.1875, 0.1875]</code>.</p>\n<p>These can be done because you improved the training speed a lot! Most of the team cannot come up training this much heavier training due to time limitation.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1103457,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-12-05T23:28:26.903000",
          "content": "<p>I am afraid we didn't do all ablation studies due to time constraints, so I can't tell the exact influence of the satellite view. It wasn't much, but with that many channels, we decided another three wouldn't hurt and may help in cases where the semantic view either fails to be drawn (happens sometimes) or in park like areas where cyclists do not drive on paved roads. </p>\n<p>Indeed, key to our solution was being able to do many experiments. We found that using even more data than whats provided with the standard settings (harder samples in this case, the train loss was significantly higher) helped to squeeze out a bit more accuracy w.r.t. the competition metric. </p>\n<p>Using more data was definitely one of the major improvements that we got. Even more improvement for a single model was gained by the pre-training on different raster-sizes. </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1103481,
          "author_name": "corochann",
          "author_url": "",
          "post_date": "2020-12-06T00:16:04.373000",
          "content": "<p>I see, thank you for quick and informative reply!</p>\n<p>It seems you draw satellite image only outside of the road. We noticed that the parking area is sometimes drawn as \"road\" (e.g., parking area of Lyft office, Starbucks or fuel stands). So I'm not sure drawing only outside of the road gives enough information or not.</p>\n<p>Yeah, I noticed the possibility of using much larger dataset, but I could not execute it during the competition…</p>\n<p>\"pre-training was important\" is interesting finding! Maybe usual \"imagenet\" pretrained weight was not suitable in this task which need measure distance from semantic image, not to recognize natural images.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1103494,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-12-06T00:56:28.230000",
          "content": "<p>no, we have the satellite image for each pixel. It's just 3 extra channels, but for simplicity we have shown an image with the semantic view over the satellite image in this Writeup. </p>\n<p>we have ego channels, agent channels, semantic channels, satellite channels</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1103973,
          "author_name": "corochann",
          "author_url": "",
          "post_date": "2020-12-06T13:35:22.503000",
          "content": "<p>I see, thank you for clarification!</p>\n<p>Sorry I have 1 another question, the best model was <code>history_num_frames=30</code>. Does that input have channels with 30*2 for ego/agent channels, maybe total 30 + 30 + 3 + 3 = 66 channels?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1104367,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-12-06T21:14:58.660000",
          "content": "<p>Yes, 66 channels in that case</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1104424,
          "author_name": "corochann",
          "author_url": "",
          "post_date": "2020-12-06T23:29:33.907000",
          "content": "<p>Wow that's very huge input! I see, thank you for answering questions :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1104851,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-12-07T09:33:07.130000",
          "content": "<p>I think less history frames are sufficient though, maybe 10. We just somehow started most experiments with 30 and never really questioned it until the end where we needed some speed gains for final models :)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1105090,
          "author_name": "corochann",
          "author_url": "",
          "post_date": "2020-12-07T14:11:07.550000",
          "content": "<p>I see, then maybe agent threshold or min_frame_history=5 was the difference with other teams. Thanks for additional comment.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1104365,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2020-12-06T21:10:13.510000",
      "content": "<p>That's when I read this kind of write-up that I know I still have a lot of to learn. </p>\n<p>Congratulations on the win, really impressive performance.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1102793,
      "author_name": "Tom Aindow",
      "author_url": "",
      "post_date": "2020-12-05T10:10:28.493000",
      "content": "<p>Amazing writeup and congratulations!</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1201855,
      "author_name": "Heroseo",
      "author_url": "",
      "post_date": "2021-02-15T17:36:57.090000",
      "content": "<p>A little late, <br>\ncongrats on 1st place and successful team effort.  <a href=\"https://www.kaggle.com/nvnnghia\" target=\"_blank\">@nvnnghia</a>, <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>, <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> <br>\nIt is one of the best sharing solutions in kaggle.</p>\n<p>Thanks for sharing and good write-up!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1105881,
      "author_name": "Louis Yang",
      "author_url": "",
      "post_date": "2020-12-08T09:39:29.067000",
      "content": "<p>Amazing speed up! Amazing work!<br>\nVery interesting way to do the validation. Never see a competition that has such small noise between CV and LB!<br>\nFor the normal distribution-like plot that you shown, is that the distribution of loss of all the samples in the validation set from a single model? <br>\nInteresting to see that you don't seem to have many outliers from your model. What was the largest single sample loss from your model?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1105933,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-12-08T10:46:02",
          "content": "<p>The distribution is from a bootstrap validation method where we randomly subsample (with multi-draw) 35k rows of the validation set. It's the histogramm over the scores of all subsamples, and not the scores of a single validation set. <br>\nWe did that, to measure the stability of the predictions (chance of shakeup).<br>\nThe largest single sample loss is ~2k as far as i remember. We all saw how skewed the score distribution is ;)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1106548,
          "author_name": "Louis Yang",
          "author_url": "",
          "post_date": "2020-12-08T23:10:00.920000",
          "content": "<p>I guess I am not very familiar with bootstrap validation. Is it like k-fold validation that you train one model on each subsample datasets? Or, it is just one model but evaluate on each subsample set?</p>\n<p>I think I saw like ~40k on my one of my ok model. ~2k is still way better than I got :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1104528,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-12-07T03:15:56.333000",
      "content": "<p>thanks for the writeup and good work!</p>\n<p>What Didn't Work<br>\n\"Unfreezing the backbone after ensembling and continue training with a second epoch\"</p>\n<p>i experience this too. it overfits.<br>\nhowever, just unfreeze the layer 4 (last on only) of the embedding backbone is ok for low learning rate.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1102784,
      "author_name": "aCode",
      "author_url": "",
      "post_date": "2020-12-05T09:57:51.750000",
      "content": "<p>Congratulations, nice team effort.<br>\nThanks for sharing 👍</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1103781,
      "author_name": "fergusoci",
      "author_url": "",
      "post_date": "2020-12-06T09:12:11.140000",
      "content": "<p>Wow. Well done to all of you - clinical execution!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1103709,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2020-12-06T06:46:31.877000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/nvnnghia\" target=\"_blank\">@nvnnghia</a> <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> and <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> and thanks for sharing details solution! </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1103213,
      "author_name": "Alin Cijov",
      "author_url": "",
      "post_date": "2020-12-05T18:14:49.727000",
      "content": "<p>Very well done, congratulations on the hard work !</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1102801,
      "author_name": "Mohammed Rizin V K",
      "author_url": "",
      "post_date": "2020-12-05T10:31:55.070000",
      "content": "<p>Amazing solution. is your solution improved better than others ,only by <code>py_sat_sem</code> or because of effnet or anything else. What is your Hardware anyway?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1197935,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-02-12T14:52:02.343000",
      "content": "",
      "votes": -6,
      "replies": []
    },
    {
      "id": 1104947,
      "author_name": "akshatp",
      "author_url": "",
      "post_date": "2020-12-07T11:51:15.653000",
      "content": "<p>Inspiring work! thanks for sharing</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1102764": "# 1st Place Solution & L5Kit Speedup\n\n## Introduction\n\nFirst of all, I would like to thank Lyft for this great challenge and of course for their huge dataset. I must say I was a bit overwhelmed by the sheer amount of agents in the training set, but it was a fun experience to be able to train without ever using training data twice. Well, we did, when we used pre-trained models, but more on that later.\n\nI also want to thank my Teammates @philippsinger, @christofhenkel and @nvnnghia for the great collaborative teamwork during the competition. We worked well together with a pipeline consisting of two repositories (version controlled custom l5kit and a training repo) and used logging in neptune.ai to keep track of the experiments. \n\n## TL;DR\n\nOur solution is an ensemble of 4 efficientnet models trained with different rasterization parameters indivdually on train_full.zarr and stacked on validation.zarr. Key to training that many models on this huge dataset were several key improvements to the original l5kit repository to remove CPU bottleneck and enable efficient multi-GPU training using pytorch. \n\n## Improvements/Changes to the L5Kit\n\nOne of the major challenges in this competition was the slow rasterizer. We profiled the l5kit a lot and pinpointed the bottlenecks to speed it up by a factor of 4+. After that, our experiments were mostly GPU limited and multi-GPU experiments allowed us for \"fast\" iterations. Running through train_full (191_177_863+ samples) was possible within 2 days with medium sized models. \n\n### Speedups\n\nSpeed before changes with semantic view only (i5-3570K single thread, so ignore the actual time):\n\n```\nRaster-time per sample: 0.124 s\nHits         Time  Per Hit   % Time  Line Contents\n==================================================\n9   11147507.0 1238611.9     93.6    rasterizer.rasterize(history_frames, history_agents, history_tl_faces, selected_agent)\n```\n\nAll the speedups were achieved with plain python, no C++. The main speedups came from:\n\n1. Batched transformation of boxes with modifications from https://github.com/lyft/l5kit/pull/167. \n   It has been said that the transformation is the bottleneck in the rasterizer as it is called a lot. We used the vectorized transformation whereever possible.\n2. In `box_rasterizer.py`, using python lists of small numpy arrays instead of one two large numpy arrays (`agent_images, ego_images`).\n   Python lists are surprisingly fast! `np.concatenate` on large numpy arrays is slow. Pre-allocation and writing to a single large array is even slower.\n   The lists can be cast to a numpy array by `out_im = np.asarray(agents_images + ego_images, dtype=np.uint8)`\n   This way, boxes will stay in `uint8`\n3. Moving concatenate to the GPU\n   As `np.concatenate` is slow on large arrays, we do that on the GPU.\n   Replace `\"image\": image,` with `\"image_box\": image_box`, `\"image_sat\": image_sat`, `\"image_sem\": image_sem`\n4. Speedups from https://github.com/lyft/l5kit/pull/140\n   With the changes above, they now made a larger impact than the 7-8% stated in the PR.\n\nSpeed after changes with semantic and satellite view (i5-3570K single thread, so ignore the actual time):\n\n```\nRaster-time per sample: 0.032 s\nHits         Time  Per Hit   % Time  Line Contents\n==================================================\n   9    2847054.0 316339.3     84.2  rasterizer.rasterize(history_frames, history_agents, history_tl_faces, selected_agent)\n```\n\n\n### Additions/Changes\n\n- New rasterizer which combines satellite and semantic view (`py_sem_sat`).\n- Flag: AgentID used as a value for drawing (instead of just boolean `no agent = 0`, `agent = 1`)\n- Flag: AgentLabelProba used as a value for drawing (instead of just boolean `no agent = 0`, `agent = 1`)\n- Flag: Velocity used as a value for drawing (instead of just boolean `no agent = 0`, `agent = 1`)\n- Option to draw all objects on the semantic layer, combinable with the label probas from above.\n- Filter options for agents: `th_extent_ratio, th_area, th_distance_av` to give a larger training dataset (or harder as you want to phrase it)\n- Multiple parallel rasterizers for ensembling models that use different raster sizes\n- Satellite view fix for non square shapes (had some weird transformation in it before)\n\n![rasterized_image](https://i.imgur.com/G3g9lEJ.png)\n\nWe will make our L5kit code changes public in the next few days and after cleanup. Most likely with PRs to the original repo.\n\n## Journey & Experiments\n\nWith all the speed improvements from our custom L5kit we were able to eliminate the previous CPU bottleneck and run experiments on an efficient multi-GPU setup. We mostly trained on 8xV100 nodes using pytorchs awesome distributed data parallel (DDP), which enables to effectivly scale a training script to multiple GPUs. This efficient training setup was key to run quite a few experiments. (See list of stuff that did not work below) In all our experiments, we used the chopped validation set for local validation as discussed [here](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/188695). This yielded us very good prediction of the LB score (in the sub 13 range, the score was usually ~0.2 higher than local validation with a standard deviation of around 0.3). \n\nOriginating from the provided baseline and the multi-mode prediction from [here](https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence), we experimented with raster size, pooling layers, head-modifications, usage of different backbones, learning rate scheduler, subsampling and oversampling to name a few. It was a major challenge in this competition, that experiments had to be run for at least 50 % of a full epoch of train_full.zarr or when comparing different sample sizes or learning rate scheduler even until the end to be really comparable. Subsampling the train set or stepwise decrease of the LR quickly yielded a low training and validation loss, but in the end, the performance was worse.\n\nQuite quickly, we decided to always use all data, and settled to a setup using shuffled train_full.zarr (with a fixed seed to be able to resume training) with a linear decay learning rate scheduler. To keep the training sample fixed, we also settled to an AgentDataset with `min_frame_history=1` and `min_frame_future=10.0` to best resemble the test dataset. Opposed to the many discussions about a much lower training loss, we never experienced that. Without any augmentation, the training loss is of course slightly lower, especially in the end of the epoch, but usually within 80% of the validation score. We found that larger backbones help to reach lower loss levels, but starting at the size of an EfficientNetB7, the nets seem to be overconfident with their predictions and the LB score was sometimes shaky. \n\nThe first half of test is public, second half is private test. We couldn't identify any major differences in the two sets. With that knowledge it was possible to hide the true score (setting some public rows to zeros), we wonder how many teams did that and were afraid of surprises in the private LB due to this. \n\n## Best Single Model\n\n![model](https://i.imgur.com/YC1WGD1.png)\n\nOur best single model is actually one that not fully finished on the last day. It's just an EfficientNetB3 with a Linear Layer Head and dropout attached trained on an extended train_full.zarr Dataset (`min_frame_history=1, min_frame_future=5.0, th_distance_av=80, th_extent_ratio=1.6, th_area=0.45`)\n\n- history_num_frames: 30.0\n- raster_size: [448, 224]\n- pixel_size: [0.1875, 0.1875]\n- rasterizer: py_sem_sat\n- batch_size: 64\n- dropout: 0.3\n\nThe model was pre-trained for 4 epochs on different lower resolution images (starting with 112, 64 and pixel size 0.75, 0.75) , so you can argue this is also an ensemble of some kind. The 5th epoch was trained with the parameters above and a very customized learning rate schedule that starts with a consine anneal and then transitions to a stepwise lr scheduler as seen in the graph below. The last evaluation score is 9.697.\n\n![best_single](https://i.imgur.com/4Q2PkvX.png)\n\nThe model achieves a public LB score of 9.776 and a private LB score of 9.070. With that, it would also have ranked 1st place this competition on it's own.\n\nPretraining is VERY important! We discovered it by a mistake where we loaded an EfficientNet without pretrained weights from imagenet. It performed significantly worse by ~1.00.\n\n## Ensemble\n\nEnsembling proved to be challenging in this competition, as traditional methods are not working well with the metric in use. We tried a lot of manual blending and form of post-processing without success. We then started building stacking models taking the raw predictions of the models as input, but this was also surprisingly not working well. At about the same time as @hengck23 posted it [here](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/193908#1077899), we had the idea of using the features/embeddings from each model as an input to a ensemble head. This worked much better and we could improve upon our single models specifically as we could introduce some diversity into the blend (different image size, pixel sizes, etc.) We also went with the dropout + single linear layer approach here, and all modification we tried with the head performed worse.\n\n![ensemble_arch](https://i.imgur.com/2I3M2rD.png)\n\nTo prevent overfitting, the ensembling was done on the chopped validation set (without using the chopped off actual validation part), in some experiments extended to validation + test set without noticeable change in the metric but with about twice the runtime. \n\nEnsembles of two similar models (even just two different checkpoints) already gave a boost of about 0.2 in the metric. Naturally, with more diverse models the ensemble performed even better and we were able to achieve our best public LB score of 9.319 and private LB score of 8.579 with an ensemble of 4 models (B3, B5, B6, B6) on two different raster sizes. \n\n## Bootstrapped Validation\n\nWe evaluated model and ensemble performance with bootstrap method on the chopped dataset. For more diverse ensembles we also got the lowest standard deviation from the bootstrap method. For our final submission it was 0.22754. The best single models had a standard deviation of ~0.27. We also identified our B7 and B8 models to have a slightly larger standard deviation in the bootstrap method of ~0.36-0.39 which may also explain their partial performance degredations on the public LB. Despite their good validation score, for that reason we excluded them in our final ensembles and final submissions. \n\n![bootstrap](https://i.imgur.com/l4vEAaz.png)\n\nWe plotted CV vs public LB for our submissions and found a spread of +-0.3 for most of our single models (green corridor) with just the B7 and B8 models outside of that corridor (not shown in the picture). Our ensembles are in an even smaller corridor with a spread of +-0.1 (orange corridor), proving their superior robustness and generalization capabilities over single models. \n\n![cv_lb](https://i.imgur.com/S6fYsBI.png)\n\nGiven this robust CV & LB correlation, we were very confident that improvements in validation also lead to improvements on the leaderboard, so we did not have to sub it. So the sudden jump on public leaderboard only happened after we decided to sub our better models and was not some sudden magic we discovered.\n\n## What Didn't Work\n\n- Flag: AgentID used as a value for drawing (instead of just boolean `no agent = 0`, `agent = 1`)\n- Flag: AgentLabelProba used as a value for drawing (instead of just boolean `no agent = 0`, `agent = 1`)\n- Flag: AgentLabelProba used as a value for drawing (instead of just boolean `no agent = 0`, `agent = 1`)\n- Option to draw all objects on the semantic layer, combinable with the label probas from above.\n- Kalman filtering the output\n- Any non NN prediction method\n- Prediction of only the difference to a constant velocity model\n- Using additional informations, such as velocity, history positions, label probas or anything else we could find in the head of the models. \n- Heavier heads\n- 3D CNNs\n- RNN based approaches on history frames for backbone input or outbut \n- Using additional features from shallower layers of the backbone\n- Ensembling with kmeans, or other analytic clustering method that we tried. Some success at scores down to around 14, but no improvements below that. \n- Subsampling the training set. Huge speedups, but accuracy was always slightly lower\n- TTA with modified yaw angles\n- Unfreezing the backbone after ensembling and continue training with a second epoch\n\n## What We Didn't Try\n\n- Augmentation\n- Multi backbone, multi rasterizer training",
    "1102803": "Very nice and clean work, kudos. It should be in textbooks on how to approach a DS problem correctly.\n\nI am very interested to compare your stacking vs [my approach](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199636). Can it be possible for you to provide me with `submission.csv` files for single model predictions of your 4 models? I will make all the results public.",
    "1102771": "Great team effort :D",
    "1197530": "As we were asked to share our ensemble model, please see the snippet below:\n\n```\nimport torch\nfrom torch import nn\nfrom typing import Dict\nimport timm\nfrom timm.models.layers.conv2d_same import Conv2dSame\nfrom utils.loss import pytorch_neg_multi_log_likelihood_batch\nfrom utils.poolings import GeM\n\n\nclass LMM(nn.Module):\n    def __init__(self, model_architecture, H=30, gem=False):\n        super().__init__()\n        self.H = H\n        num_history_channels = (self.H + 1) * 2\n        rgb_channels = 6\n        num_in_channels = rgb_channels + num_history_channels\n        self.future_len = 50\n        num_targets = 2 * self.future_len\n        self.num_modes = 3\n        self.num_preds = num_targets * self.num_modes\n        self.backbone = timm.create_model(model_architecture, pretrained=False)\n\n        if gem:\n            self.backbone.global_pool = GeM()\n\n        self.backbone.conv_stem = Conv2dSame(\n            num_in_channels,\n            self.backbone.conv_stem.out_channels,\n            kernel_size=self.backbone.conv_stem.kernel_size,\n            stride=self.backbone.conv_stem.stride,\n            padding=self.backbone.conv_stem.padding,\n            bias=False,\n        )\n\n        self.backbone_out_features = self.backbone.classifier.in_features\n        self.backbone.classifier = nn.Sequential(\n            nn.Identity(),\n            nn.Linear(\n                in_features=self.backbone.classifier.in_features,\n                out_features=self.backbone_out_features,\n            ),\n        )\n\n        self.lin_head = nn.Sequential(\n            nn.ReLU(),\n            nn.Linear(\n                in_features=self.backbone_out_features,\n                out_features=self.num_preds + self.num_modes,\n            ),\n        )\n\n        for param in self.parameters():\n            param.requires_grad = False\n\n    def forward(self, image_box, image_sat, image_sem, history, history_availabilities):\n        x = torch.cat(\n            (\n                image_box[:, : self.H + 1].float(),\n                image_box[:, 31 : 31 + self.H + 1].float(),\n                image_sat,\n                image_sem,\n            ),\n            1,\n        )\n        x = self.backbone(x)\n        x = self.lin_head(x)\n        return x\n\n\nclass LyftMultiModel(nn.Module):\n    def __init__(self, cfg: Dict, PARAMS, num_modes=3):\n        super().__init__()\n        self.future_len = cfg[\"model_params\"][\"future_num_frames\"]\n        num_targets = 2 * self.future_len\n        self.num_preds = num_targets * num_modes\n        self.num_modes = num_modes\n        self.PARAMS = PARAMS\n\n        self.model1 = LMM(\"tf_efficientnet_b6_ns\")\n        sd = torch.load(\"output/preds/config_ch_aws_2_checkpoint.pth\")[\"model\"]\n        self.model1.load_state_dict(sd, strict=True)\n        self.model1.lin_head = nn.Identity()\n\n        self.model3 = LMM(\"tf_efficientnet_b6_ns\")\n        sd = torch.load(\"output/preds/config_ch_aws_1_checkpoint.pth\")[\"model\"]\n        self.model3.load_state_dict(sd, strict=True)\n        backbone_out_features = self.model3.backbone_out_features\n        self.model3.lin_head = nn.Identity()\n\n        self.model4 = LMM(\"tf_efficientnet_b3_ns\", gem=True)\n        sd = torch.load(\"output/preds/philipp_config_17_ep5_checkpoint_lr5e6.pth\")[\"model\"]\n        self.model4.load_state_dict(sd, strict=True)\n        backbone_out_features_2 = self.model4.backbone_out_features\n        self.model4.lin_head = nn.Identity()\n\n        self.model5 = LMM(\"tf_efficientnet_b5_ns\", H=8)\n        sd = torch.load(\"output/preds/config_ch_ddp_16_checkpoint.pth\")[\"model\"]\n        self.model5.load_state_dict(sd)\n        backbone_out_features3 = self.model5.backbone_out_features\n        self.model5.lin_head = nn.Identity()\n\n        self.lin_head = nn.Sequential(\n            nn.ReLU(),\n            nn.Linear(\n                in_features=backbone_out_features * 2\n                + backbone_out_features_2\n                + backbone_out_features3,\n                out_features=self.num_preds + num_modes,\n            ),\n        )\n\n    def forward(\n        self,\n        image_box,\n        image_sat,\n        image_sem,\n        image_box2,\n        image_sat2,\n        image_sem2,\n        history,\n        history_availabilities,\n    ):\n\n        self.model1 = self.model1.eval()\n        self.model3 = self.model3.eval()\n        self.model4 = self.model4.eval()\n        self.model5 = self.model5.eval()\n\n        x1 = self.model1(image_box, image_sat, image_sem, history, history_availabilities)\n        x3 = self.model3(image_box, image_sat, image_sem, history, history_availabilities)\n        x4 = self.model4(image_box2, image_sat2, image_sem2, history, history_availabilities)\n        x5 = self.model5(image_box2, image_sat2, image_sem2, history, history_availabilities)\n\n        x = self.lin_head(torch.cat([x1, x3, x4, x5], dim=-1))\n\n        bs, _ = x.shape\n        pred, confidences = torch.split(x, self.num_preds, dim=1)\n        pred = pred.view(bs, self.num_modes, self.future_len, 2)\n        assert confidences.shape == (bs, self.num_modes)\n        confidences = torch.softmax(confidences, dim=1)\n        return pred, confidences\n\n\nclass LyftMultiRegressor(nn.Module):\n    \"\"\"Single mode prediction\"\"\"\n\n    def __init__(self, predictor, PARAMS):\n        super().__init__()\n        self.predictor = predictor\n        self.PARAMS = PARAMS\n\n    def forward(\n        self,\n        image_box,\n        image_sat,\n        image_sem,\n        image_box2,\n        image_sat2,\n        image_sem2,\n        targets,\n        target_availabilities,\n        history,\n        history_availabilities,\n        velocity,\n    ):\n        pred, confidences = self.predictor(\n            image_box,\n            image_sat,\n            image_sem,\n            image_box2,\n            image_sat2,\n            image_sem2,\n            history,\n            history_availabilities,\n        )\n\n        if self.PARAMS.predict_diffs:\n            pred = torch.cumsum(pred, dim=2)\n\n        loss_nll = pytorch_neg_multi_log_likelihood_batch(\n            targets, pred, confidences, target_availabilities\n        )\n\n        return loss_nll, pred, confidences\n```",
    "1103452": "Thank you @ilu000 for the write-up! and congrats for 1st place ( @christofhenkel @philippsinger @nvnnghia ) !\n\nHow much the satellite image affect to the score? Did you compare the score difference between `py_semantic` and your `py_sem_sat`.\nI'm guessing that this part has less impact compared to other hyperparameters change, especially filter agent prob or image size/pixel size, min_frame_history to 5 etc. \nWhich one was important to achieve public LB score less than 10?",
    "1104365": "That's when I read this kind of write-up that I know I still have a lot of to learn. \n\nCongratulations on the win, really impressive performance.",
    "1102793": "Amazing writeup and congratulations!",
    "1201855": "A little late, \ncongrats on 1st place and successful team effort.  @nvnnghia, @philippsinger, @christofhenkel \nIt is one of the best sharing solutions in kaggle.\n\nThanks for sharing and good write-up!",
    "1105881": "Amazing speed up! Amazing work!\nVery interesting way to do the validation. Never see a competition that has such small noise between CV and LB!\nFor the normal distribution-like plot that you shown, is that the distribution of loss of all the samples in the validation set from a single model? \nInteresting to see that you don't seem to have many outliers from your model. What was the largest single sample loss from your model?\n",
    "1104528": "thanks for the writeup and good work!\n\nWhat Didn't Work\n\"Unfreezing the backbone after ensembling and continue training with a second epoch\"\n\ni experience this too. it overfits.\nhowever, just unfreeze the layer 4 (last on only) of the embedding backbone is ok for low learning rate.\n",
    "1102784": "Congratulations, nice team effort.\nThanks for sharing 👍",
    "1103781": "Wow. Well done to all of you - clinical execution!",
    "1103709": "Congrats @nvnnghia @ilu000 @christofhenkel and @philippsinger and thanks for sharing details solution! ",
    "1103213": "Very well done, congratulations on the hard work !",
    "1102801": "Amazing solution. is your solution improved better than others ,only by `py_sat_sem` or because of effnet or anything else. What is your Hardware anyway?",
    "1197935": "",
    "1104947": "Inspiring work! thanks for sharing"
  }
}