{
  "id": 199657,
  "title": "4th place solution: Ensemble with GMM",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/199657",
  "author_name": "corochann",
  "post_date": "2020-11-26T16:55:14.362000",
  "votes": 58,
  "comment_count": 26,
  "views": 0,
  "content": "<p>Thank you to the organizers and congratulations to all the participants. <br>\nAlso I would like to thank my team members <a href=\"https://www.kaggle.com/zaburo\" target=\"_blank\">@zaburo</a>, <a href=\"https://www.kaggle.com/qhapaq49\" target=\"_blank\">@qhapaq49</a>, <a href=\"https://www.kaggle.com/charmq\" target=\"_blank\">@charmq</a> for our hard work, I could enjoy the competition!<br>\nThe LB was really stable in this competition due to the big amount of data, we could work on improving the model without caring about big shake-up.</p>\n<p>We have started with my baseline kernel <a href=\"https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence\" target=\"_blank\">Lyft: Training with multi-mode confidence</a>.<br>\nBelow items are substantial changes we made:</p>\n<p>[Update 2020/12/9] <strong>We have published our code:</strong></p>\n<ul>\n<li><a href=\"https://github.com/pfnet-research/kaggle-lyft-motion-prediction-4th-place-solution\" target=\"_blank\">https://github.com/pfnet-research/kaggle-lyft-motion-prediction-4th-place-solution</a></li>\n</ul>\n<h1>Short Summary</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F518134%2Fa27c649533aca3f12e23255ed2295440%2Flyft_4th_place_solution.png?generation=1606414075335190&amp;alt=media\" alt=\"\"></p>\n<p>Published baseline training pipeline was indeed already very strong. <br>\nJust modifying</p>\n<ol>\n<li>train_full.zarr</li>\n<li>l5kit==1.1.0</li>\n<li>Set min_history=0, min_future=10 in AgentDataset</li>\n<li>Cosine annealing for LR decrease until 0</li>\n</ol>\n<p>with training 1 epoch was already enough to win the prize.</p>\n<h1>1. Use train_full.zarr data</h1>\n<p>Bigger data is almost always better for deep learning model training. We used <a href=\"https://www.kaggle.com/philculliton/lyft-full-training-set\" target=\"_blank\">Lyft Full Training Set</a>.</p>\n<p>However its size is really large, containing 191M data for AgentDataset.</p>\n<p>Practically we need this modifications in order to train this big dataset in real time:</p>\n<h2>Distributed training</h2>\n<p>We implemented distributed training using <code>torch.distributed</code>.<br>\nIt usually takes about 5 days to finish 1 epoch when we use 8GPUs.</p>\n<h2>Caching some arrays into zarr beforehand to reduce on-memory usage in AgentDataset</h2>\n<p>The problem arises when we run distributed training and <code>DataLoader</code> with setting <code>num_workers</code> for multi-process data loading.</p>\n<p>In the distributed training, 8 processes run in parallel and each process invokes <code>num_workers</code> subprocess. Therefore 8 * num_workers subprocess is launched and <code>AgentDataset</code> data is copied in each subprocess.<br>\nThen Out Of Memory error occurs because <code>AgentDataset</code> internally holds <code>cumulative_sizes</code> attribute whose size is very big (<a href=\"https://github.com/lyft/l5kit/blob/1ea2f8cecfe7ad974419bf5c8519ab3a21300119/l5kit/l5kit/dataset/agent.py#L62\" target=\"_blank\">code</a>).</p>\n<p>Instead, we pre-calculated <code>track_id, scene_index, state_index</code> and saved as the zarr format. So that we can load each data from disk, and reduce on-memory usage.</p>\n<p>The Public Score was around <strong>25.742</strong> <a href=\"https://www.kaggle.com/corochann/lyft-prediction-with-multi-mode-confidence\" target=\"_blank\">kernel</a> at this stage.</p>\n<h1>2. Use l5kit==1.1.0</h1>\n<p>As mentioned in the discussion <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/186492\" target=\"_blank\">We did it all wrong</a>, image is rotated while the target value is not rotated in the previous version of l5kit==1.0.6 during the beginning of the competition.<br>\nWe updated l5kit version to 1.1.0 once it is released, which fixes this behavior.</p>\n<p>The Public LB score was jumped to <strong>15.874</strong> with this update.</p>\n<h1>3. Set min_history=0, min_future=10</h1>\n<p>As written in the “Validation Strategy” section, validation&amp;test data is made by <code>create_chopped_dataset</code> method. We noticed that this validation/test data consists of the data with A. always contains more than 10 future frames, and B sometimes it does not contain any history frame.</p>\n<p>To align the training <code>AgentDataset</code> to this test dataset behavior, we can set <strong><code>min_frame_history=0</code> and <code>min_frame_future=10</code></strong>.<br>\nI think <strong>this modification is the most important part to notice in this competition</strong>.<br>\nYou need a courage to intentionally ignore l5kit library warning ;)(<a href=\"https://github.com/lyft/l5kit/blob/1ea2f8cecfe7ad974419bf5c8519ab3a21300119/l5kit/l5kit/dataset/agent.py#L15-L17\" target=\"_blank\">code</a>).</p>\n<p>It’s very effective, the score jumped to <strong>13.059</strong>.</p>\n<h1>4. Training: with cosine annealing</h1>\n<p>Model: We trained &amp; used following models for final ensemble</p>\n<ul>\n<li>Resnet18</li>\n<li>Resnet50</li>\n<li>SEResNeXt50</li>\n<li>ecaresnet18</li>\n</ul>\n<p>However resnet18 baseline was strong, and enough to win the prize.</p>\n<p>Image size: tried (128, 128) and (224, 224). image size of 128 training proceeds faster, but image size 224 final score was slightly better.</p>\n<p>Optimizer: Adam with Cosine annealing<br>\nCosine annealing was better than Exponential decay. I think decreasing the learning rate until very close to 0 is important for final tuning.</p>\n<p>Batch size: 12 * 8 process = 96</p>\n<p>We just trained only 1 epoch to train full data. We did not downsample any of the data.</p>\n<p>Public LB score for single resnet18 model is <strong>11.341</strong>.</p>\n<h1>Augmentation</h1>\n<h2>Image augmentation</h2>\n<p>Many of the augmentation used in natural images is not appropriate for this competition task, (for example flip augmentation flips the target value as well and not realistic since right-lane, left-lane will change). We tried</p>\n<ul>\n<li>Cutout</li>\n<li>Blur</li>\n<li>Downscale<br>\nusing <a href=\"https://github.com/albumentations-team/albumentations\" target=\"_blank\">albumentations</a> library.</li>\n</ul>\n<h2>Rasterizer-level augmentation</h2>\n<p>What is different from normal image prediction is that the image is drawn by rasterizer. We can also consider applying augmentation during rasterization.</p>\n<p>I tried following augmentation by modifying <code>BoxRasterizer</code> (<a href=\"https://github.com/lyft/l5kit/blob/1ea2f8cecfe7ad974419bf5c8519ab3a21300119/l5kit/l5kit/rasterization/box_rasterizer.py\" target=\"_blank\">code</a>).</p>\n<ul>\n<li>Drop agent randomly<ul>\n<li>I assumed that the target agent’s movement does not change so much when the other agent far from the target agent exists or not. So we randomly skip drawing some of the agent boxes.</li></ul></li>\n<li>Scale extent size randomly<ul>\n<li>Even when the other agent size changes a bit, I assume that the target agent’s behavior does not change. So we scaled extent size from factor 0.9~1.1</li></ul></li>\n</ul>\n<p>I thought rasterizer-level augmentation was an interesting idea for this competition task. However we could not see big score improvement actually. Maybe the training dataset size is already big enough and its effect is not so big.</p>\n<p>By <strong>only adding cutout augmentation</strong>, we achieved to get public LB score of <strong>10.846</strong>, already enough to win the prize. </p>\n<h1>Validation Strategy</h1>\n<p>As discussed in <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/188695\" target=\"_blank\">Validation vs LB score</a>, we can run very stable validation using a chopped dataset.<br>\nHowever, it removes ground truth data, format is different from training AgentDataset and difficult to validate during training phase.</p>\n<p>What we want is <code>agents_mask_orig_bool</code> (<a href=\"https://github.com/lyft/l5kit/blob/1ea2f8cecfe7ad974419bf5c8519ab3a21300119/l5kit/l5kit/evaluation/chop_dataset.py#L68\" target=\"_blank\">code</a>).<br>\nWe saved this <code>agents_mask_orig_bool</code> and set it to the <code>agent_mask</code> argument of <code>AgentDataset</code>.<br>\nThen we could run validation during training.</p>\n<p>We sub-sampled 10000 dataset for fast validation during training, but it differs a lot from the score using a total 190327 validation dataset.<br>\nAt the end phase of the competition, we validated the trained model using a full validation dataset.</p>\n<h1>Ensemble: sample trajectory and GMM fitting</h1>\n<p>To improve the score further, how to ensemble is the key question in this task.<br>\nNo golden method is suggested in <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/180931\" target=\"_blank\">the discussion</a> and we came up with an idea to adopt the <strong>Gaussian Mixture Model</strong>.</p>\n<p>We can sample trajectory and fit the sampled points by GMM with the 3 components.<br>\nWe started from the <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.mixture.GaussianMixture.html\" target=\"_blank\">sklearn implementation</a>.<br>\nSetting <code>n_components=3</code> and <code>covariance_type=”spherical”</code> achieved the good score.<br>\nHowever sigma is fixed to 1 in this competition <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/overview/evaluation\" target=\"_blank\">metric</a>, so we also tried implementing own GMM model with fixing covariance to be 1.</p>\n<p>Ensemble by GMM was really effective, we finally achieved <strong>public LB score 10.272/private LB score 9.475</strong></p>\n<p>Thus, the ensemble pushed the public LB score from 10.846 to 10.272. But it does not change the final rank this time ;)</p>\n<p>By the way, I saw other participants used k-means clustering for ensembling coords. <br>\nI think the behavior is quite similar with using GMM, since it calls k-means clustering in the initialization of EM algorithm.</p>\n<h1>What we tried and not worked</h1>\n<p>We noticed Baseline model &amp; l5kit default rasterizer was already very strong in this competition.<br>\nWe really tried a lot, but many of the attempts failed to improve the scores. I’ll write in the reply section (since it’s already very long).</p>",
  "messages": [
    {
      "id": 1092267,
      "postDate": "2020-11-26T16:55:14.363Z",
      "content": "<p>Thank you to the organizers and congratulations to all the participants. <br>\nAlso I would like to thank my team members <a href=\"https://www.kaggle.com/zaburo\" target=\"_blank\">@zaburo</a>, <a href=\"https://www.kaggle.com/qhapaq49\" target=\"_blank\">@qhapaq49</a>, <a href=\"https://www.kaggle.com/charmq\" target=\"_blank\">@charmq</a> for our hard work, I could enjoy the competition!<br>\nThe LB was really stable in this competition due to the big amount of data, we could work on improving the model without caring about big shake-up.</p>\n<p>We have started with my baseline kernel <a href=\"https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence\" target=\"_blank\">Lyft: Training with multi-mode confidence</a>.<br>\nBelow items are substantial changes we made:</p>\n<p>[Update 2020/12/9] <strong>We have published our code:</strong></p>\n<ul>\n<li><a href=\"https://github.com/pfnet-research/kaggle-lyft-motion-prediction-4th-place-solution\" target=\"_blank\">https://github.com/pfnet-research/kaggle-lyft-motion-prediction-4th-place-solution</a></li>\n</ul>\n<h1>Short Summary</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F518134%2Fa27c649533aca3f12e23255ed2295440%2Flyft_4th_place_solution.png?generation=1606414075335190&amp;alt=media\" alt=\"\"></p>\n<p>Published baseline training pipeline was indeed already very strong. <br>\nJust modifying</p>\n<ol>\n<li>train_full.zarr</li>\n<li>l5kit==1.1.0</li>\n<li>Set min_history=0, min_future=10 in AgentDataset</li>\n<li>Cosine annealing for LR decrease until 0</li>\n</ol>\n<p>with training 1 epoch was already enough to win the prize.</p>\n<h1>1. Use train_full.zarr data</h1>\n<p>Bigger data is almost always better for deep learning model training. We used <a href=\"https://www.kaggle.com/philculliton/lyft-full-training-set\" target=\"_blank\">Lyft Full Training Set</a>.</p>\n<p>However its size is really large, containing 191M data for AgentDataset.</p>\n<p>Practically we need this modifications in order to train this big dataset in real time:</p>\n<h2>Distributed training</h2>\n<p>We implemented distributed training using <code>torch.distributed</code>.<br>\nIt usually takes about 5 days to finish 1 epoch when we use 8GPUs.</p>\n<h2>Caching some arrays into zarr beforehand to reduce on-memory usage in AgentDataset</h2>\n<p>The problem arises when we run distributed training and <code>DataLoader</code> with setting <code>num_workers</code> for multi-process data loading.</p>\n<p>In the distributed training, 8 processes run in parallel and each process invokes <code>num_workers</code> subprocess. Therefore 8 * num_workers subprocess is launched and <code>AgentDataset</code> data is copied in each subprocess.<br>\nThen Out Of Memory error occurs because <code>AgentDataset</code> internally holds <code>cumulative_sizes</code> attribute whose size is very big (<a href=\"https://github.com/lyft/l5kit/blob/1ea2f8cecfe7ad974419bf5c8519ab3a21300119/l5kit/l5kit/dataset/agent.py#L62\" target=\"_blank\">code</a>).</p>\n<p>Instead, we pre-calculated <code>track_id, scene_index, state_index</code> and saved as the zarr format. So that we can load each data from disk, and reduce on-memory usage.</p>\n<p>The Public Score was around <strong>25.742</strong> <a href=\"https://www.kaggle.com/corochann/lyft-prediction-with-multi-mode-confidence\" target=\"_blank\">kernel</a> at this stage.</p>\n<h1>2. Use l5kit==1.1.0</h1>\n<p>As mentioned in the discussion <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/186492\" target=\"_blank\">We did it all wrong</a>, image is rotated while the target value is not rotated in the previous version of l5kit==1.0.6 during the beginning of the competition.<br>\nWe updated l5kit version to 1.1.0 once it is released, which fixes this behavior.</p>\n<p>The Public LB score was jumped to <strong>15.874</strong> with this update.</p>\n<h1>3. Set min_history=0, min_future=10</h1>\n<p>As written in the “Validation Strategy” section, validation&amp;test data is made by <code>create_chopped_dataset</code> method. We noticed that this validation/test data consists of the data with A. always contains more than 10 future frames, and B sometimes it does not contain any history frame.</p>\n<p>To align the training <code>AgentDataset</code> to this test dataset behavior, we can set <strong><code>min_frame_history=0</code> and <code>min_frame_future=10</code></strong>.<br>\nI think <strong>this modification is the most important part to notice in this competition</strong>.<br>\nYou need a courage to intentionally ignore l5kit library warning ;)(<a href=\"https://github.com/lyft/l5kit/blob/1ea2f8cecfe7ad974419bf5c8519ab3a21300119/l5kit/l5kit/dataset/agent.py#L15-L17\" target=\"_blank\">code</a>).</p>\n<p>It’s very effective, the score jumped to <strong>13.059</strong>.</p>\n<h1>4. Training: with cosine annealing</h1>\n<p>Model: We trained &amp; used following models for final ensemble</p>\n<ul>\n<li>Resnet18</li>\n<li>Resnet50</li>\n<li>SEResNeXt50</li>\n<li>ecaresnet18</li>\n</ul>\n<p>However resnet18 baseline was strong, and enough to win the prize.</p>\n<p>Image size: tried (128, 128) and (224, 224). image size of 128 training proceeds faster, but image size 224 final score was slightly better.</p>\n<p>Optimizer: Adam with Cosine annealing<br>\nCosine annealing was better than Exponential decay. I think decreasing the learning rate until very close to 0 is important for final tuning.</p>\n<p>Batch size: 12 * 8 process = 96</p>\n<p>We just trained only 1 epoch to train full data. We did not downsample any of the data.</p>\n<p>Public LB score for single resnet18 model is <strong>11.341</strong>.</p>\n<h1>Augmentation</h1>\n<h2>Image augmentation</h2>\n<p>Many of the augmentation used in natural images is not appropriate for this competition task, (for example flip augmentation flips the target value as well and not realistic since right-lane, left-lane will change). We tried</p>\n<ul>\n<li>Cutout</li>\n<li>Blur</li>\n<li>Downscale<br>\nusing <a href=\"https://github.com/albumentations-team/albumentations\" target=\"_blank\">albumentations</a> library.</li>\n</ul>\n<h2>Rasterizer-level augmentation</h2>\n<p>What is different from normal image prediction is that the image is drawn by rasterizer. We can also consider applying augmentation during rasterization.</p>\n<p>I tried following augmentation by modifying <code>BoxRasterizer</code> (<a href=\"https://github.com/lyft/l5kit/blob/1ea2f8cecfe7ad974419bf5c8519ab3a21300119/l5kit/l5kit/rasterization/box_rasterizer.py\" target=\"_blank\">code</a>).</p>\n<ul>\n<li>Drop agent randomly<ul>\n<li>I assumed that the target agent’s movement does not change so much when the other agent far from the target agent exists or not. So we randomly skip drawing some of the agent boxes.</li></ul></li>\n<li>Scale extent size randomly<ul>\n<li>Even when the other agent size changes a bit, I assume that the target agent’s behavior does not change. So we scaled extent size from factor 0.9~1.1</li></ul></li>\n</ul>\n<p>I thought rasterizer-level augmentation was an interesting idea for this competition task. However we could not see big score improvement actually. Maybe the training dataset size is already big enough and its effect is not so big.</p>\n<p>By <strong>only adding cutout augmentation</strong>, we achieved to get public LB score of <strong>10.846</strong>, already enough to win the prize. </p>\n<h1>Validation Strategy</h1>\n<p>As discussed in <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/188695\" target=\"_blank\">Validation vs LB score</a>, we can run very stable validation using a chopped dataset.<br>\nHowever, it removes ground truth data, format is different from training AgentDataset and difficult to validate during training phase.</p>\n<p>What we want is <code>agents_mask_orig_bool</code> (<a href=\"https://github.com/lyft/l5kit/blob/1ea2f8cecfe7ad974419bf5c8519ab3a21300119/l5kit/l5kit/evaluation/chop_dataset.py#L68\" target=\"_blank\">code</a>).<br>\nWe saved this <code>agents_mask_orig_bool</code> and set it to the <code>agent_mask</code> argument of <code>AgentDataset</code>.<br>\nThen we could run validation during training.</p>\n<p>We sub-sampled 10000 dataset for fast validation during training, but it differs a lot from the score using a total 190327 validation dataset.<br>\nAt the end phase of the competition, we validated the trained model using a full validation dataset.</p>\n<h1>Ensemble: sample trajectory and GMM fitting</h1>\n<p>To improve the score further, how to ensemble is the key question in this task.<br>\nNo golden method is suggested in <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/180931\" target=\"_blank\">the discussion</a> and we came up with an idea to adopt the <strong>Gaussian Mixture Model</strong>.</p>\n<p>We can sample trajectory and fit the sampled points by GMM with the 3 components.<br>\nWe started from the <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.mixture.GaussianMixture.html\" target=\"_blank\">sklearn implementation</a>.<br>\nSetting <code>n_components=3</code> and <code>covariance_type=”spherical”</code> achieved the good score.<br>\nHowever sigma is fixed to 1 in this competition <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/overview/evaluation\" target=\"_blank\">metric</a>, so we also tried implementing own GMM model with fixing covariance to be 1.</p>\n<p>Ensemble by GMM was really effective, we finally achieved <strong>public LB score 10.272/private LB score 9.475</strong></p>\n<p>Thus, the ensemble pushed the public LB score from 10.846 to 10.272. But it does not change the final rank this time ;)</p>\n<p>By the way, I saw other participants used k-means clustering for ensembling coords. <br>\nI think the behavior is quite similar with using GMM, since it calls k-means clustering in the initialization of EM algorithm.</p>\n<h1>What we tried and not worked</h1>\n<p>We noticed Baseline model &amp; l5kit default rasterizer was already very strong in this competition.<br>\nWe really tried a lot, but many of the attempts failed to improve the scores. I’ll write in the reply section (since it’s already very long).</p>",
      "rawMarkdown": "Thank you to the organizers and congratulations to all the participants. \nAlso I would like to thank my team members @zaburo, @qhapaq49, @charmq for our hard work, I could enjoy the competition!\nThe LB was really stable in this competition due to the big amount of data, we could work on improving the model without caring about big shake-up.\n\nWe have started with my baseline kernel [Lyft: Training with multi-mode confidence](https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence).\nBelow items are substantial changes we made:\n\n[Update 2020/12/9] **We have published our code:**\n - https://github.com/pfnet-research/kaggle-lyft-motion-prediction-4th-place-solution\n\n# Short Summary\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F518134%2Fa27c649533aca3f12e23255ed2295440%2Flyft_4th_place_solution.png?generation=1606414075335190&alt=media)\n\nPublished baseline training pipeline was indeed already very strong. \nJust modifying\n\n1. train_full.zarr\n2. l5kit==1.1.0\n3. Set min_history=0, min_future=10 in AgentDataset\n4. Cosine annealing for LR decrease until 0\n\nwith training 1 epoch was already enough to win the prize.\n\n\n\n# 1. Use train_full.zarr data\n\nBigger data is almost always better for deep learning model training. We used [Lyft Full Training Set](https://www.kaggle.com/philculliton/lyft-full-training-set).\n\nHowever its size is really large, containing 191M data for AgentDataset.\n\nPractically we need this modifications in order to train this big dataset in real time:\n\n## Distributed training\nWe implemented distributed training using `torch.distributed`.\nIt usually takes about 5 days to finish 1 epoch when we use 8GPUs.\n\n## Caching some arrays into zarr beforehand to reduce on-memory usage in AgentDataset\nThe problem arises when we run distributed training and `DataLoader` with setting `num_workers` for multi-process data loading.\n\nIn the distributed training, 8 processes run in parallel and each process invokes `num_workers` subprocess. Therefore 8 * num_workers subprocess is launched and `AgentDataset` data is copied in each subprocess.\nThen Out Of Memory error occurs because `AgentDataset` internally holds `cumulative_sizes` attribute whose size is very big ([code](https://github.com/lyft/l5kit/blob/1ea2f8cecfe7ad974419bf5c8519ab3a21300119/l5kit/l5kit/dataset/agent.py#L62)).\n\nInstead, we pre-calculated `track_id, scene_index, state_index` and saved as the zarr format. So that we can load each data from disk, and reduce on-memory usage.\n\nThe Public Score was around **25.742** [kernel](https://www.kaggle.com/corochann/lyft-prediction-with-multi-mode-confidence) at this stage.\n\n# 2. Use l5kit==1.1.0\nAs mentioned in the discussion [We did it all wrong](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/186492), image is rotated while the target value is not rotated in the previous version of l5kit==1.0.6 during the beginning of the competition.\nWe updated l5kit version to 1.1.0 once it is released, which fixes this behavior.\n\nThe Public LB score was jumped to **15.874** with this update.\n\n# 3. Set min_history=0, min_future=10\nAs written in the “Validation Strategy” section, validation&test data is made by `create_chopped_dataset` method. We noticed that this validation/test data consists of the data with A. always contains more than 10 future frames, and B sometimes it does not contain any history frame.\n\nTo align the training `AgentDataset` to this test dataset behavior, we can set **`min_frame_history=0` and `min_frame_future=10`**.\nI think **this modification is the most important part to notice in this competition**.\nYou need a courage to intentionally ignore l5kit library warning ;)([code](https://github.com/lyft/l5kit/blob/1ea2f8cecfe7ad974419bf5c8519ab3a21300119/l5kit/l5kit/dataset/agent.py#L15-L17)).\n\nIt’s very effective, the score jumped to **13.059**.\n\n# 4. Training: with cosine annealing\n\nModel: We trained & used following models for final ensemble\n - Resnet18\n - Resnet50\n - SEResNeXt50\n - ecaresnet18\n\nHowever resnet18 baseline was strong, and enough to win the prize.\n\nImage size: tried (128, 128) and (224, 224). image size of 128 training proceeds faster, but image size 224 final score was slightly better.\n\nOptimizer: Adam with Cosine annealing\nCosine annealing was better than Exponential decay. I think decreasing the learning rate until very close to 0 is important for final tuning.\n\nBatch size: 12 * 8 process = 96\n\nWe just trained only 1 epoch to train full data. We did not downsample any of the data.\n\nPublic LB score for single resnet18 model is **11.341**.\n\n# Augmentation\n## Image augmentation\nMany of the augmentation used in natural images is not appropriate for this competition task, (for example flip augmentation flips the target value as well and not realistic since right-lane, left-lane will change). We tried\n - Cutout\n - Blur\n - Downscale\nusing [albumentations](https://github.com/albumentations-team/albumentations) library.\n\n## Rasterizer-level augmentation\nWhat is different from normal image prediction is that the image is drawn by rasterizer. We can also consider applying augmentation during rasterization.\n\nI tried following augmentation by modifying `BoxRasterizer` ([code](https://github.com/lyft/l5kit/blob/1ea2f8cecfe7ad974419bf5c8519ab3a21300119/l5kit/l5kit/rasterization/box_rasterizer.py)).\n - Drop agent randomly\n    - I assumed that the target agent’s movement does not change so much when the other agent far from the target agent exists or not. So we randomly skip drawing some of the agent boxes.\n - Scale extent size randomly\n    - Even when the other agent size changes a bit, I assume that the target agent’s behavior does not change. So we scaled extent size from factor 0.9~1.1\n\nI thought rasterizer-level augmentation was an interesting idea for this competition task. However we could not see big score improvement actually. Maybe the training dataset size is already big enough and its effect is not so big.\n\nBy **only adding cutout augmentation**, we achieved to get public LB score of **10.846**, already enough to win the prize. \n\n# Validation Strategy\n\nAs discussed in [Validation vs LB score](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/188695), we can run very stable validation using a chopped dataset.\nHowever, it removes ground truth data, format is different from training AgentDataset and difficult to validate during training phase.\n\nWhat we want is `agents_mask_orig_bool` ([code](https://github.com/lyft/l5kit/blob/1ea2f8cecfe7ad974419bf5c8519ab3a21300119/l5kit/l5kit/evaluation/chop_dataset.py#L68)).\nWe saved this `agents_mask_orig_bool` and set it to the `agent_mask` argument of `AgentDataset`.\nThen we could run validation during training.\n\nWe sub-sampled 10000 dataset for fast validation during training, but it differs a lot from the score using a total 190327 validation dataset.\nAt the end phase of the competition, we validated the trained model using a full validation dataset.\n\n# Ensemble: sample trajectory and GMM fitting\nTo improve the score further, how to ensemble is the key question in this task.\nNo golden method is suggested in [the discussion](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/180931) and we came up with an idea to adopt the **Gaussian Mixture Model**.\n\nWe can sample trajectory and fit the sampled points by GMM with the 3 components.\nWe started from the [sklearn implementation](https://scikit-learn.org/stable/modules/generated/sklearn.mixture.GaussianMixture.html).\nSetting `n_components=3` and `covariance_type=”spherical”` achieved the good score.\nHowever sigma is fixed to 1 in this competition [metric](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/overview/evaluation), so we also tried implementing own GMM model with fixing covariance to be 1.\n\nEnsemble by GMM was really effective, we finally achieved **public LB score 10.272/private LB score 9.475**\n\nThus, the ensemble pushed the public LB score from 10.846 to 10.272. But it does not change the final rank this time ;)\n\nBy the way, I saw other participants used k-means clustering for ensembling coords. \nI think the behavior is quite similar with using GMM, since it calls k-means clustering in the initialization of EM algorithm.\n\n# What we tried and not worked\nWe noticed Baseline model & l5kit default rasterizer was already very strong in this competition.\nWe really tried a lot, but many of the attempts failed to improve the scores. I’ll write in the reply section (since it’s already very long).\n",
      "votes": 57
    },
    {
      "id": 1106968,
      "postDate": "2020-12-09T09:04:51.770Z",
      "content": "<p>We have published our code:<br>\n<a href=\"https://github.com/pfnet-research/kaggle-lyft-motion-prediction-4th-place-solution\" target=\"_blank\">https://github.com/pfnet-research/kaggle-lyft-motion-prediction-4th-place-solution</a></p>",
      "rawMarkdown": "We have published our code:\nhttps://github.com/pfnet-research/kaggle-lyft-motion-prediction-4th-place-solution",
      "votes": 5,
      "replies": [
        {
          "id": 1107307,
          "postDate": "2020-12-09T15:24:16.500Z",
          "content": "<p>Great job.<br>\nThanks for sharing!</p>\n<p><a href=\"https://www.kaggle.com/corochann\" target=\"_blank\">@corochann</a> </p>",
          "rawMarkdown": "Great job.\nThanks for sharing!\n\n@corochann ",
          "votes": 1
        },
        {
          "id": 1107782,
          "postDate": "2020-12-10T00:32:09.503Z",
          "content": "<p>Thanks, please check/try our ensemble code :)</p>\n<ul>\n<li><a href=\"https://github.com/pfnet-research/kaggle-lyft-motion-prediction-4th-place-solution/tree/master/src/ensemble\" target=\"_blank\">https://github.com/pfnet-research/kaggle-lyft-motion-prediction-4th-place-solution/tree/master/src/ensemble</a></li>\n</ul>",
          "rawMarkdown": "Thanks, please check/try our ensemble code :)\n - https://github.com/pfnet-research/kaggle-lyft-motion-prediction-4th-place-solution/tree/master/src/ensemble",
          "votes": 1
        },
        {
          "id": 1107795,
          "postDate": "2020-12-10T01:03:26.083Z",
          "content": "<p>Awesome! thank you! 👍</p>",
          "rawMarkdown": "Awesome! thank you! 👍",
          "votes": 1
        },
        {
          "id": 1107809,
          "postDate": "2020-12-10T01:26:49.637Z",
          "content": "<p>😃</p>\n<p>You can find various trial implementations of our rasterizer too: (numba jit tuned version of rasterizer etc)</p>\n<ul>\n<li><a href=\"https://github.com/pfnet-research/kaggle-lyft-motion-prediction-4th-place-solution/tree/master/src/lib/rasterization\" target=\"_blank\">https://github.com/pfnet-research/kaggle-lyft-motion-prediction-4th-place-solution/tree/master/src/lib/rasterization</a></li>\n</ul>",
          "rawMarkdown": "😃\n\nYou can find various trial implementations of our rasterizer too: (numba jit tuned version of rasterizer etc)\n - https://github.com/pfnet-research/kaggle-lyft-motion-prediction-4th-place-solution/tree/master/src/lib/rasterization",
          "votes": 2
        }
      ]
    },
    {
      "id": 1096004,
      "postDate": "2020-11-30T06:27:25.403Z",
      "content": "<h1>What we tried and not worked</h1>\n<p>Sorry for the late post, these are some of the items that we tried but not worked.</p>\n<h2>Change hyper parameters</h2>\n<p>We tried several hyper parameters, especially for rasterizers. But none of them contributed to improve the model's accuracy.</p>\n<ul>\n<li><p>image size</p>\n<ul>\n<li>We tried 224x224 &amp; 128x128. The image size=128 training is faster especially because Rasterization becomes faster, and its training accuracy is almost the same until the middle of the training. However, its validation loss is a bit (only about 0.5~1.0) worse than image size = 224.</li></ul></li>\n<li><p>pixel_size</p>\n<ul>\n<li>There are many frames that the car is almost stopping now but starts in the near future. We thought that when the car starts moving, its change in the pixel is very small and CNN cannot detect it when the pixel_size is bit (resolution is rough). We tried to change <code>pixel_size</code> from default 0.5 into 0.25 or 0.15 but the accuracy becomes worse.</li></ul></li>\n<li><p>num_history</p>\n<ul>\n<li>1. Only short history predictor: Several agents have very few past history. So I thought when we train the model with only 0 past frames (i.e., input only current frame), this specific purpose model performs better for predicting the future with only 0 past frames. In the training phase, <strong>we can train this model using the agents with many past frames, by just input only a current frame.</strong> However, this model’s accuracy is worse than the default 10 history input model.</li>\n<li>2. Long history predictor: Oppositely, having more past information helps to improve the score? To check that hypothesis, we tried to input longer past history frames by setting <code>history_num_frames=7, 10</code> with <code>history_step_size=2</code> instead of default <code>history_num_frames=10</code> with <code>history_step_size=1</code>. This model’s accuracy was lower than the original model even if we only chose the validation input to have more than 20 history frames.</li></ul></li>\n</ul>\n<h2>Big, deep models</h2>\n<p>We tried <code>resnet101</code> &amp; <code>res152</code> too, but they did not work well.</p>\n<h2>Custom Rasterizer</h2>\n<p>We tried implementing our own rasterizer to add more rich information to CNN input, but all of them did not work well to improve the accuracy somehow…</p>\n<ul>\n<li><code>ChannelSemanticRasterizer</code><ul>\n<li><code>SemanticRasterizer</code> draws the semantic in RGB space, using 3 channels. We thought this is not always optimal for CNN input and tried to input 6 channels with 1. road, 2. default lane, 3. green signal lane, 4. yellow signal lane, 5. red signal lane, 6. crosswalk.</li></ul></li>\n<li><code>TLSemanticRasterizer</code>:<ul>\n<li>When we executed EDA, we thought knowing <strong>the red signal length is important</strong>. Because some frames start with the red signal as current, and start in the future when the signal changed to green.<br>\nWe input a signal length by changing the color value so that CNN can understand how long this signal color already continued (since Host car detected the signal).</li></ul></li>\n<li><code>AgentTypeBoxRasterizer</code>:<ul>\n<li>There are 4 agent types: CAR, CYCLIST, PEDESTRIAN and UNKNOWN in the original dataset. But UNKNOWN is not drawn in the original <code>BoxRasterizer</code>. Also agent type information is also dropped when drawing boxes.<br>\nWe tried to draw each agent type in different channels including UNKNOWN, to input more precise information. </li></ul></li>\n</ul>\n<p><strong>Speed up rasterizer</strong><br>\nUse numpy batch operation as much as possible, and replacing implementations  with numba jit computations. Even though it becomes faster in single process, its computation speed-up does not contribute so much for multi-process data preparation during training.</p>\n<h2>Train with Agent type</h2>\n<p>Other than trying the <code>AgentTypeBoxRasterizer</code>, we tried inputting agent type one-hot vector explicitly, but it did not contribute to improve accuracy too.</p>\n<h2>Multi-agent prediction model</h2>\n<p>The baseline kernel predicts future movement of only target agent. Instead I considered to build a model which predicts all the agent's future movement within the input image.<br>\nThe first I thought this idea speed-ups the training since it can predict multiple agent at once. However sometimes the host car detects agent with very far place, like 400 pixels far away. So we noticed it is difficult to align the image size to fixed size. And we suspended its further trial.</p>\n<h2>Yaw correction</h2>\n<p>The below figure shows the biggest error in the validation dataset. The error is extremely high when the <code>yaw</code> in the dataset was actually opposite and the model predicted the opposite way to go.<br>\nWe worked hard to check if the test dataset contains this kind of case. Indeed there seems to be some frames whose <code>yaw</code> might be opposite, however most of the time the agent is stopping in this case in the test dataset. Even though we fixed the yaw, the LB score was almost unchanged.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F518134%2Fe51c936752b5cd7bd0689f6ef9faebe9%2Fvalidation_worst_error.png?generation=1606717208803896&amp;alt=media\" alt=\"\"><br>\nThe biggest error in the validation dataset with error=43988.1 !!</p>\n<h2>Leak check</h2>\n<p>When we checked the dataset carefully, we noticed that the timestamp &amp; the map position actually overlaps within the train/validation/test dataset.<br>\nWe checked if the test dataset information was leaked from another dataset. But it seems that the timestamp is not aligned in scenes (maybe physically other host car is used to collect data, and timestamp record is not calibrated).</p>",
      "rawMarkdown": "# What we tried and not worked\n\nSorry for the late post, these are some of the items that we tried but not worked.\n\n## Change hyper parameters\n\nWe tried several hyper parameters, especially for rasterizers. But none of them contributed to improve the model's accuracy.\n - image size\n    - We tried 224x224 & 128x128. The image size=128 training is faster especially because Rasterization becomes faster, and its training accuracy is almost the same until the middle of the training. However, its validation loss is a bit (only about 0.5~1.0) worse than image size = 224.\n\n - pixel_size\n    - There are many frames that the car is almost stopping now but starts in the near future. We thought that when the car starts moving, its change in the pixel is very small and CNN cannot detect it when the pixel_size is bit (resolution is rough). We tried to change `pixel_size` from default 0.5 into 0.25 or 0.15 but the accuracy becomes worse.\n\n - num_history\n    - 1. Only short history predictor: Several agents have very few past history. So I thought when we train the model with only 0 past frames (i.e., input only current frame), this specific purpose model performs better for predicting the future with only 0 past frames. In the training phase, **we can train this model using the agents with many past frames, by just input only a current frame.** However, this model’s accuracy is worse than the default 10 history input model.\n    - 2. Long history predictor: Oppositely, having more past information helps to improve the score? To check that hypothesis, we tried to input longer past history frames by setting `history_num_frames=7, 10` with `history_step_size=2` instead of default `history_num_frames=10` with `history_step_size=1`. This model’s accuracy was lower than the original model even if we only chose the validation input to have more than 20 history frames.\n\n\n## Big, deep models\nWe tried `resnet101` & `res152` too, but they did not work well.\n\n## Custom Rasterizer\n\nWe tried implementing our own rasterizer to add more rich information to CNN input, but all of them did not work well to improve the accuracy somehow...\n - `ChannelSemanticRasterizer`\n   - `SemanticRasterizer` draws the semantic in RGB space, using 3 channels. We thought this is not always optimal for CNN input and tried to input 6 channels with 1. road, 2. default lane, 3. green signal lane, 4. yellow signal lane, 5. red signal lane, 6. crosswalk.\n - `TLSemanticRasterizer`:\n    - When we executed EDA, we thought knowing **the red signal length is important**. Because some frames start with the red signal as current, and start in the future when the signal changed to green.\nWe input a signal length by changing the color value so that CNN can understand how long this signal color already continued (since Host car detected the signal).\n - `AgentTypeBoxRasterizer`:\n   - There are 4 agent types: CAR, CYCLIST, PEDESTRIAN and UNKNOWN in the original dataset. But UNKNOWN is not drawn in the original `BoxRasterizer`. Also agent type information is also dropped when drawing boxes.\nWe tried to draw each agent type in different channels including UNKNOWN, to input more precise information. \n\n\n**Speed up rasterizer**\nUse numpy batch operation as much as possible, and replacing implementations  with numba jit computations. Even though it becomes faster in single process, its computation speed-up does not contribute so much for multi-process data preparation during training.\n\n## Train with Agent type\nOther than trying the `AgentTypeBoxRasterizer`, we tried inputting agent type one-hot vector explicitly, but it did not contribute to improve accuracy too.\n\n\n## Multi-agent prediction model\nThe baseline kernel predicts future movement of only target agent. Instead I considered to build a model which predicts all the agent's future movement within the input image.\nThe first I thought this idea speed-ups the training since it can predict multiple agent at once. However sometimes the host car detects agent with very far place, like 400 pixels far away. So we noticed it is difficult to align the image size to fixed size. And we suspended its further trial.\n\n## Yaw correction\nThe below figure shows the biggest error in the validation dataset. The error is extremely high when the `yaw` in the dataset was actually opposite and the model predicted the opposite way to go.\nWe worked hard to check if the test dataset contains this kind of case. Indeed there seems to be some frames whose `yaw` might be opposite, however most of the time the agent is stopping in this case in the test dataset. Even though we fixed the yaw, the LB score was almost unchanged.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F518134%2Fe51c936752b5cd7bd0689f6ef9faebe9%2Fvalidation_worst_error.png?generation=1606717208803896&alt=media)\nThe biggest error in the validation dataset with error=43988.1 !!\n\n\n## Leak check\nWhen we checked the dataset carefully, we noticed that the timestamp & the map position actually overlaps within the train/validation/test dataset.\nWe checked if the test dataset information was leaked from another dataset. But it seems that the timestamp is not aligned in scenes (maybe physically other host car is used to collect data, and timestamp record is not calibrated).\n\n",
      "votes": 6,
      "replies": [
        {
          "id": 1099404,
          "postDate": "2020-12-02T10:42:04.313Z",
          "content": "<p>Thanks for sharing! I think this clear list of things that don't work is way more important than those that work! <br>\nFor Lyft, I think they might want to check the \"Yaw correction\" part. Predicting a car to move in a completely opposite direction will probably be a fatal mistake in real life. Even if it only happen once.</p>",
          "rawMarkdown": "Thanks for sharing! I think this clear list of things that don't work is way more important than those that work! \nFor Lyft, I think they might want to check the \"Yaw correction\" part. Predicting a car to move in a completely opposite direction will probably be a fatal mistake in real life. Even if it only happen once.",
          "votes": 1
        },
        {
          "id": 1099529,
          "postDate": "2020-12-02T12:30:39.103Z",
          "content": "<p>Thanks for reply <a href=\"https://www.kaggle.com/louis925\" target=\"_blank\">@louis925</a>, yeah I think so too. For considering application, this \"what did not work\" is more important to consider further why.</p>\n<p>Actually \"Yaw correction\" happened many times, and I guess to correct it we need better accuracy for the previous object detection NN part where I guess this module will decide the yaw direction.<br>\nWhile the pedestrian &amp; cyclist they are more easier to decide direction from real image and indeed there is less mistake for yaw, but car is just \"box\" shape, determining its direction from image sometimes has mistake.</p>",
          "rawMarkdown": "Thanks for reply @louis925, yeah I think so too. For considering application, this \"what did not work\" is more important to consider further why.\n\nActually \"Yaw correction\" happened many times, and I guess to correct it we need better accuracy for the previous object detection NN part where I guess this module will decide the yaw direction.\nWhile the pedestrian & cyclist they are more easier to decide direction from real image and indeed there is less mistake for yaw, but car is just \"box\" shape, determining its direction from image sometimes has mistake.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1092920,
      "postDate": "2020-11-27T09:40:26.947Z",
      "content": "<p>Thank you so much for sharing. Just curious, did you use GCP or a local machine for training? Using v100 in GCP seems quite expansive.</p>",
      "rawMarkdown": "Thank you so much for sharing. Just curious, did you use GCP or a local machine for training? Using v100 in GCP seems quite expansive.",
      "votes": 3,
      "replies": [
        {
          "id": 1092923,
          "postDate": "2020-11-27T09:44:30.027Z",
          "content": "<p>We used local machine :)</p>",
          "rawMarkdown": "We used local machine :)",
          "votes": 3
        }
      ]
    },
    {
      "id": 1092700,
      "postDate": "2020-11-27T05:12:27.447Z",
      "content": "<p>Great write-up, thanks for sharing. I like your GMM ensembling.</p>",
      "rawMarkdown": "Great write-up, thanks for sharing. I like your GMM ensembling.",
      "votes": 3,
      "replies": [
        {
          "id": 1092755,
          "postDate": "2020-11-27T06:22:29.687Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> for comment. I'm looking forward to see your team's solution.</p>\n<p>I saw that your team used stacking. I also wonder if your team applied GMM ensembling, the score has increased further or not!</p>",
          "rawMarkdown": "Thank you @christofhenkel for comment. I'm looking forward to see your team's solution.\n\nI saw that your team used stacking. I also wonder if your team applied GMM ensembling, the score has increased further or not!"
        },
        {
          "id": 1092894,
          "postDate": "2020-11-27T08:52:44.610Z",
          "content": "<p>thats what I am also wondering. Might give it a try today. Thanks for posting the code </p>",
          "rawMarkdown": "thats what I am also wondering. Might give it a try today. Thanks for posting the code ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1092309,
      "postDate": "2020-11-26T17:28:17.537Z",
      "content": "<p>Thanks for sharing and well explanation.<br>\nI totally agree with you!<br>\n<a href=\"https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence\" target=\"_blank\">Lyft: Training with multi-mode confidence</a> is really strong baseline.</p>\n<p>And Congrats 4th <a href=\"https://www.kaggle.com/corochann\" target=\"_blank\">@corochann</a> and <a href=\"https://www.kaggle.com/zaburo\" target=\"_blank\">@zaburo</a>, <a href=\"https://www.kaggle.com/qhapaq49\" target=\"_blank\">@qhapaq49</a>, <a href=\"https://www.kaggle.com/charmq\" target=\"_blank\">@charmq</a>.</p>",
      "rawMarkdown": "Thanks for sharing and well explanation.\nI totally agree with you!\n[Lyft: Training with multi-mode confidence](https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence) is really strong baseline.\n\nAnd Congrats 4th @corochann and @zaburo, @qhapaq49, @charmq.",
      "votes": 3,
      "replies": [
        {
          "id": 1092326,
          "postDate": "2020-11-26T17:42:01.813Z",
          "content": "<p>Thank you! Glad to know that :) <a href=\"https://www.kaggle.com/piantic\" target=\"_blank\">@piantic</a> </p>",
          "rawMarkdown": "Thank you! Glad to know that :) @piantic "
        }
      ]
    },
    {
      "id": 1098852,
      "postDate": "2020-12-01T22:38:03.257Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/corochann\" target=\"_blank\">@corochann</a> and team. Thanks for sharing your solution.</p>",
      "rawMarkdown": "Congrats @corochann and team. Thanks for sharing your solution.",
      "votes": 1,
      "replies": [
        {
          "id": 1098932,
          "postDate": "2020-12-02T00:33:12.727Z",
          "content": "<p>Thanks, congrats <a href=\"https://www.kaggle.com/sheriytm\" target=\"_blank\">@sheriytm</a> too!</p>",
          "rawMarkdown": "Thanks, congrats @sheriytm too!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1092767,
      "postDate": "2020-11-27T06:37:08.903Z",
      "content": "<p>good work and thanks for the writeup!</p>\n<p>\"so we also tried implementing own GMM model with fixing covariance to be 1.\"<br>\ndo you have some code or pseudo-code for that? i would like to implement and try.<br>\nthanks</p>",
      "rawMarkdown": "good work and thanks for the writeup!\n\n\"so we also tried implementing own GMM model with fixing covariance to be 1.\"\ndo you have some code or pseudo-code for that? i would like to implement and try.\nthanks",
      "votes": 1,
      "replies": [
        {
          "id": 1092801,
          "postDate": "2020-11-27T07:17:34.680Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> <br>\nThanks for your comment.<br>\nWe used this code, usage is same with sklearn's GMM.</p>\n<pre><code>import numba as nb\nimport numpy as np\nfrom sklearn.mixture import GaussianMixture\n# from sklearn.mixture._gaussian_mixture import _estimate_gaussian_parameters\n\n\n@nb.jit(nb.types.Tuple(\n    (nb.float64[:], nb.float64[:, :])\n)(nb.float64[:, :], nb.float64[:, :]), nopython=True, nogil=True)\ndef _estimate_gaussian_parameters(X, resp):\n    nk = resp.sum(axis=0) + 10 * np.finfo(resp.dtype).eps\n    means = np.dot(np.ascontiguousarray(resp.T), X) / np.ascontiguousarray(np.expand_dims(nk, 1))\n    return nk, means\n\n\nclass GaussianMixtureIdentity(GaussianMixture):\n    def _initialize(self, X, resp):\n        n_samples, _ = X.shape\n        self.covariances_ = np.zeros(self.n_components)+1.0\n        self.precisions_cholesky_ = np.zeros(self.n_components)+1.0\n        weights, means = _estimate_gaussian_parameters(X, resp)\n        weights /= n_samples\n\n        self.weights_ = (weights if self.weights_init is None\n                         else self.weights_init)\n        self.means_ = means if self.means_init is None else self.means_init\n\n    def _m_step(self, X, log_resp):\n        n_samples, _ = X.shape\n        self.covariances_ = np.zeros(self.n_components)+1.0\n        self.precisions_cholesky_ = np.zeros(self.n_components)+1.0\n        self.weights_, self.means_ = _estimate_gaussian_parameters(X, np.exp(log_resp))\n        self.weights_ /= n_samples\n</code></pre>",
          "rawMarkdown": "@hengck23 \nThanks for your comment.\nWe used this code, usage is same with sklearn's GMM.\n\n```\n  \nimport numba as nb\nimport numpy as np\nfrom sklearn.mixture import GaussianMixture\n# from sklearn.mixture._gaussian_mixture import _estimate_gaussian_parameters\n\n\n@nb.jit(nb.types.Tuple(\n    (nb.float64[:], nb.float64[:, :])\n)(nb.float64[:, :], nb.float64[:, :]), nopython=True, nogil=True)\ndef _estimate_gaussian_parameters(X, resp):\n    nk = resp.sum(axis=0) + 10 * np.finfo(resp.dtype).eps\n    means = np.dot(np.ascontiguousarray(resp.T), X) / np.ascontiguousarray(np.expand_dims(nk, 1))\n    return nk, means\n\n\nclass GaussianMixtureIdentity(GaussianMixture):\n    def _initialize(self, X, resp):\n        n_samples, _ = X.shape\n        self.covariances_ = np.zeros(self.n_components)+1.0\n        self.precisions_cholesky_ = np.zeros(self.n_components)+1.0\n        weights, means = _estimate_gaussian_parameters(X, resp)\n        weights /= n_samples\n\n        self.weights_ = (weights if self.weights_init is None\n                         else self.weights_init)\n        self.means_ = means if self.means_init is None else self.means_init\n\n    def _m_step(self, X, log_resp):\n        n_samples, _ = X.shape\n        self.covariances_ = np.zeros(self.n_components)+1.0\n        self.precisions_cholesky_ = np.zeros(self.n_components)+1.0\n        self.weights_, self.means_ = _estimate_gaussian_parameters(X, np.exp(log_resp))\n        self.weights_ /= n_samples\n```",
          "votes": 6
        }
      ]
    },
    {
      "id": 1092471,
      "postDate": "2020-11-26T21:02:18.263Z",
      "content": "<p>Congrats! Wow! You set <code>min_history=0</code>! Very interesting idea! Does that increase the amount samples in AgentDataset by a lot?</p>",
      "rawMarkdown": "Congrats! Wow! You set `min_history=0`! Very interesting idea! Does that increase the amount samples in AgentDataset by a lot?",
      "votes": 1,
      "replies": [
        {
          "id": 1092600,
          "postDate": "2020-11-27T02:50:32.563Z",
          "content": "<p>Yes, the dataset becomes <code>191,177,863</code> -&gt; <code>198,474,478</code>.</p>\n<p>I guess <code>min_frame_history=0</code> increases the number, while <code>min_frame_history=10</code> reduces the number.</p>",
          "rawMarkdown": "Yes, the dataset becomes `191,177,863` -> `198,474,478`.\n\nI guess `min_frame_history=0` increases the number, while `min_frame_history=10` reduces the number.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1092322,
      "postDate": "2020-11-26T17:34:31.200Z",
      "content": "<p>Congrats! What hardware have you used to run such a massive experiment? I have trained my models for days and only covered 34M samples, and you did 191M. How many cores have you used, what was the bottleneck? Quite impressive!</p>",
      "rawMarkdown": "Congrats! What hardware have you used to run such a massive experiment? I have trained my models for days and only covered 34M samples, and you did 191M. How many cores have you used, what was the bottleneck? Quite impressive!",
      "votes": 1,
      "replies": [
        {
          "id": 1092328,
          "postDate": "2020-11-26T17:45:17.960Z",
          "content": "<p>Thank you and congrats to you too! <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> </p>\n<p>We used 8 V100 GPUs for single model training with about 32 cpu cores.<br>\nI think disk IO and rasterization process was the bottleneck compared to the normal image training.</p>",
          "rawMarkdown": "Thank you and congrats to you too! @zaharch \n\nWe used 8 V100 GPUs for single model training with about 32 cpu cores.\nI think disk IO and rasterization process was the bottleneck compared to the normal image training.",
          "votes": 3
        }
      ]
    },
    {
      "id": 1094891,
      "postDate": "2020-11-29T04:11:44.587Z",
      "content": "<p>Thanks! <a href=\"https://www.kaggle.com/prokaggler\" target=\"_blank\">@prokaggler</a> </p>",
      "rawMarkdown": "Thanks! @prokaggler "
    },
    {
      "id": 1102591,
      "postDate": "2020-12-05T05:04:27.047Z",
      "content": "<p>I love this post!</p>",
      "rawMarkdown": "I love this post!",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 1103475,
          "postDate": "2020-12-05T23:56:34.103Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/YoungseokJoung\" target=\"_blank\">@YoungseokJoung</a> !</p>",
          "rawMarkdown": "Thanks @YoungseokJoung !"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1106968,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2020-12-09T09:04:51.770000",
      "content": "<p>We have published our code:<br>\n<a href=\"https://github.com/pfnet-research/kaggle-lyft-motion-prediction-4th-place-solution\" target=\"_blank\">https://github.com/pfnet-research/kaggle-lyft-motion-prediction-4th-place-solution</a></p>",
      "votes": 5,
      "replies": [
        {
          "id": 1107307,
          "author_name": "Heroseo",
          "author_url": "",
          "post_date": "2020-12-09T15:24:16.500000",
          "content": "<p>Great job.<br>\nThanks for sharing!</p>\n<p><a href=\"https://www.kaggle.com/corochann\" target=\"_blank\">@corochann</a> </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1107782,
          "author_name": "corochann",
          "author_url": "",
          "post_date": "2020-12-10T00:32:09.503000",
          "content": "<p>Thanks, please check/try our ensemble code :)</p>\n<ul>\n<li><a href=\"https://github.com/pfnet-research/kaggle-lyft-motion-prediction-4th-place-solution/tree/master/src/ensemble\" target=\"_blank\">https://github.com/pfnet-research/kaggle-lyft-motion-prediction-4th-place-solution/tree/master/src/ensemble</a></li>\n</ul>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1107795,
          "author_name": "Dean Kang",
          "author_url": "",
          "post_date": "2020-12-10T01:03:26.083000",
          "content": "<p>Awesome! thank you! 👍</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1107809,
          "author_name": "corochann",
          "author_url": "",
          "post_date": "2020-12-10T01:26:49.637000",
          "content": "<p>😃</p>\n<p>You can find various trial implementations of our rasterizer too: (numba jit tuned version of rasterizer etc)</p>\n<ul>\n<li><a href=\"https://github.com/pfnet-research/kaggle-lyft-motion-prediction-4th-place-solution/tree/master/src/lib/rasterization\" target=\"_blank\">https://github.com/pfnet-research/kaggle-lyft-motion-prediction-4th-place-solution/tree/master/src/lib/rasterization</a></li>\n</ul>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1096004,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2020-11-30T06:27:25.403000",
      "content": "<h1>What we tried and not worked</h1>\n<p>Sorry for the late post, these are some of the items that we tried but not worked.</p>\n<h2>Change hyper parameters</h2>\n<p>We tried several hyper parameters, especially for rasterizers. But none of them contributed to improve the model's accuracy.</p>\n<ul>\n<li><p>image size</p>\n<ul>\n<li>We tried 224x224 &amp; 128x128. The image size=128 training is faster especially because Rasterization becomes faster, and its training accuracy is almost the same until the middle of the training. However, its validation loss is a bit (only about 0.5~1.0) worse than image size = 224.</li></ul></li>\n<li><p>pixel_size</p>\n<ul>\n<li>There are many frames that the car is almost stopping now but starts in the near future. We thought that when the car starts moving, its change in the pixel is very small and CNN cannot detect it when the pixel_size is bit (resolution is rough). We tried to change <code>pixel_size</code> from default 0.5 into 0.25 or 0.15 but the accuracy becomes worse.</li></ul></li>\n<li><p>num_history</p>\n<ul>\n<li>1. Only short history predictor: Several agents have very few past history. So I thought when we train the model with only 0 past frames (i.e., input only current frame), this specific purpose model performs better for predicting the future with only 0 past frames. In the training phase, <strong>we can train this model using the agents with many past frames, by just input only a current frame.</strong> However, this model’s accuracy is worse than the default 10 history input model.</li>\n<li>2. Long history predictor: Oppositely, having more past information helps to improve the score? To check that hypothesis, we tried to input longer past history frames by setting <code>history_num_frames=7, 10</code> with <code>history_step_size=2</code> instead of default <code>history_num_frames=10</code> with <code>history_step_size=1</code>. This model’s accuracy was lower than the original model even if we only chose the validation input to have more than 20 history frames.</li></ul></li>\n</ul>\n<h2>Big, deep models</h2>\n<p>We tried <code>resnet101</code> &amp; <code>res152</code> too, but they did not work well.</p>\n<h2>Custom Rasterizer</h2>\n<p>We tried implementing our own rasterizer to add more rich information to CNN input, but all of them did not work well to improve the accuracy somehow…</p>\n<ul>\n<li><code>ChannelSemanticRasterizer</code><ul>\n<li><code>SemanticRasterizer</code> draws the semantic in RGB space, using 3 channels. We thought this is not always optimal for CNN input and tried to input 6 channels with 1. road, 2. default lane, 3. green signal lane, 4. yellow signal lane, 5. red signal lane, 6. crosswalk.</li></ul></li>\n<li><code>TLSemanticRasterizer</code>:<ul>\n<li>When we executed EDA, we thought knowing <strong>the red signal length is important</strong>. Because some frames start with the red signal as current, and start in the future when the signal changed to green.<br>\nWe input a signal length by changing the color value so that CNN can understand how long this signal color already continued (since Host car detected the signal).</li></ul></li>\n<li><code>AgentTypeBoxRasterizer</code>:<ul>\n<li>There are 4 agent types: CAR, CYCLIST, PEDESTRIAN and UNKNOWN in the original dataset. But UNKNOWN is not drawn in the original <code>BoxRasterizer</code>. Also agent type information is also dropped when drawing boxes.<br>\nWe tried to draw each agent type in different channels including UNKNOWN, to input more precise information. </li></ul></li>\n</ul>\n<p><strong>Speed up rasterizer</strong><br>\nUse numpy batch operation as much as possible, and replacing implementations  with numba jit computations. Even though it becomes faster in single process, its computation speed-up does not contribute so much for multi-process data preparation during training.</p>\n<h2>Train with Agent type</h2>\n<p>Other than trying the <code>AgentTypeBoxRasterizer</code>, we tried inputting agent type one-hot vector explicitly, but it did not contribute to improve accuracy too.</p>\n<h2>Multi-agent prediction model</h2>\n<p>The baseline kernel predicts future movement of only target agent. Instead I considered to build a model which predicts all the agent's future movement within the input image.<br>\nThe first I thought this idea speed-ups the training since it can predict multiple agent at once. However sometimes the host car detects agent with very far place, like 400 pixels far away. So we noticed it is difficult to align the image size to fixed size. And we suspended its further trial.</p>\n<h2>Yaw correction</h2>\n<p>The below figure shows the biggest error in the validation dataset. The error is extremely high when the <code>yaw</code> in the dataset was actually opposite and the model predicted the opposite way to go.<br>\nWe worked hard to check if the test dataset contains this kind of case. Indeed there seems to be some frames whose <code>yaw</code> might be opposite, however most of the time the agent is stopping in this case in the test dataset. Even though we fixed the yaw, the LB score was almost unchanged.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F518134%2Fe51c936752b5cd7bd0689f6ef9faebe9%2Fvalidation_worst_error.png?generation=1606717208803896&amp;alt=media\" alt=\"\"><br>\nThe biggest error in the validation dataset with error=43988.1 !!</p>\n<h2>Leak check</h2>\n<p>When we checked the dataset carefully, we noticed that the timestamp &amp; the map position actually overlaps within the train/validation/test dataset.<br>\nWe checked if the test dataset information was leaked from another dataset. But it seems that the timestamp is not aligned in scenes (maybe physically other host car is used to collect data, and timestamp record is not calibrated).</p>",
      "votes": 6,
      "replies": [
        {
          "id": 1099404,
          "author_name": "Louis Yang",
          "author_url": "",
          "post_date": "2020-12-02T10:42:04.313000",
          "content": "<p>Thanks for sharing! I think this clear list of things that don't work is way more important than those that work! <br>\nFor Lyft, I think they might want to check the \"Yaw correction\" part. Predicting a car to move in a completely opposite direction will probably be a fatal mistake in real life. Even if it only happen once.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1099529,
          "author_name": "corochann",
          "author_url": "",
          "post_date": "2020-12-02T12:30:39.103000",
          "content": "<p>Thanks for reply <a href=\"https://www.kaggle.com/louis925\" target=\"_blank\">@louis925</a>, yeah I think so too. For considering application, this \"what did not work\" is more important to consider further why.</p>\n<p>Actually \"Yaw correction\" happened many times, and I guess to correct it we need better accuracy for the previous object detection NN part where I guess this module will decide the yaw direction.<br>\nWhile the pedestrian &amp; cyclist they are more easier to decide direction from real image and indeed there is less mistake for yaw, but car is just \"box\" shape, determining its direction from image sometimes has mistake.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1092920,
      "author_name": "Dean Kang",
      "author_url": "",
      "post_date": "2020-11-27T09:40:26.947000",
      "content": "<p>Thank you so much for sharing. Just curious, did you use GCP or a local machine for training? Using v100 in GCP seems quite expansive.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1092923,
          "author_name": "corochann",
          "author_url": "",
          "post_date": "2020-11-27T09:44:30.027000",
          "content": "<p>We used local machine :)</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1092700,
      "author_name": "Dieter",
      "author_url": "",
      "post_date": "2020-11-27T05:12:27.447000",
      "content": "<p>Great write-up, thanks for sharing. I like your GMM ensembling.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1092755,
          "author_name": "corochann",
          "author_url": "",
          "post_date": "2020-11-27T06:22:29.687000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> for comment. I'm looking forward to see your team's solution.</p>\n<p>I saw that your team used stacking. I also wonder if your team applied GMM ensembling, the score has increased further or not!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1092894,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2020-11-27T08:52:44.610000",
          "content": "<p>thats what I am also wondering. Might give it a try today. Thanks for posting the code </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1092309,
      "author_name": "Heroseo",
      "author_url": "",
      "post_date": "2020-11-26T17:28:17.537000",
      "content": "<p>Thanks for sharing and well explanation.<br>\nI totally agree with you!<br>\n<a href=\"https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence\" target=\"_blank\">Lyft: Training with multi-mode confidence</a> is really strong baseline.</p>\n<p>And Congrats 4th <a href=\"https://www.kaggle.com/corochann\" target=\"_blank\">@corochann</a> and <a href=\"https://www.kaggle.com/zaburo\" target=\"_blank\">@zaburo</a>, <a href=\"https://www.kaggle.com/qhapaq49\" target=\"_blank\">@qhapaq49</a>, <a href=\"https://www.kaggle.com/charmq\" target=\"_blank\">@charmq</a>.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1092326,
          "author_name": "corochann",
          "author_url": "",
          "post_date": "2020-11-26T17:42:01.813000",
          "content": "<p>Thank you! Glad to know that :) <a href=\"https://www.kaggle.com/piantic\" target=\"_blank\">@piantic</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1098852,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2020-12-01T22:38:03.257000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/corochann\" target=\"_blank\">@corochann</a> and team. Thanks for sharing your solution.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1098932,
          "author_name": "corochann",
          "author_url": "",
          "post_date": "2020-12-02T00:33:12.727000",
          "content": "<p>Thanks, congrats <a href=\"https://www.kaggle.com/sheriytm\" target=\"_blank\">@sheriytm</a> too!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1092767,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-11-27T06:37:08.903000",
      "content": "<p>good work and thanks for the writeup!</p>\n<p>\"so we also tried implementing own GMM model with fixing covariance to be 1.\"<br>\ndo you have some code or pseudo-code for that? i would like to implement and try.<br>\nthanks</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1092801,
          "author_name": "corochann",
          "author_url": "",
          "post_date": "2020-11-27T07:17:34.680000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> <br>\nThanks for your comment.<br>\nWe used this code, usage is same with sklearn's GMM.</p>\n<pre><code>import numba as nb\nimport numpy as np\nfrom sklearn.mixture import GaussianMixture\n# from sklearn.mixture._gaussian_mixture import _estimate_gaussian_parameters\n\n\n@nb.jit(nb.types.Tuple(\n    (nb.float64[:], nb.float64[:, :])\n)(nb.float64[:, :], nb.float64[:, :]), nopython=True, nogil=True)\ndef _estimate_gaussian_parameters(X, resp):\n    nk = resp.sum(axis=0) + 10 * np.finfo(resp.dtype).eps\n    means = np.dot(np.ascontiguousarray(resp.T), X) / np.ascontiguousarray(np.expand_dims(nk, 1))\n    return nk, means\n\n\nclass GaussianMixtureIdentity(GaussianMixture):\n    def _initialize(self, X, resp):\n        n_samples, _ = X.shape\n        self.covariances_ = np.zeros(self.n_components)+1.0\n        self.precisions_cholesky_ = np.zeros(self.n_components)+1.0\n        weights, means = _estimate_gaussian_parameters(X, resp)\n        weights /= n_samples\n\n        self.weights_ = (weights if self.weights_init is None\n                         else self.weights_init)\n        self.means_ = means if self.means_init is None else self.means_init\n\n    def _m_step(self, X, log_resp):\n        n_samples, _ = X.shape\n        self.covariances_ = np.zeros(self.n_components)+1.0\n        self.precisions_cholesky_ = np.zeros(self.n_components)+1.0\n        self.weights_, self.means_ = _estimate_gaussian_parameters(X, np.exp(log_resp))\n        self.weights_ /= n_samples\n</code></pre>",
          "votes": 6,
          "replies": []
        }
      ]
    },
    {
      "id": 1092471,
      "author_name": "Louis Yang",
      "author_url": "",
      "post_date": "2020-11-26T21:02:18.263000",
      "content": "<p>Congrats! Wow! You set <code>min_history=0</code>! Very interesting idea! Does that increase the amount samples in AgentDataset by a lot?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1092600,
          "author_name": "corochann",
          "author_url": "",
          "post_date": "2020-11-27T02:50:32.563000",
          "content": "<p>Yes, the dataset becomes <code>191,177,863</code> -&gt; <code>198,474,478</code>.</p>\n<p>I guess <code>min_frame_history=0</code> increases the number, while <code>min_frame_history=10</code> reduces the number.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1092322,
      "author_name": "nosound",
      "author_url": "",
      "post_date": "2020-11-26T17:34:31.200000",
      "content": "<p>Congrats! What hardware have you used to run such a massive experiment? I have trained my models for days and only covered 34M samples, and you did 191M. How many cores have you used, what was the bottleneck? Quite impressive!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1092328,
          "author_name": "corochann",
          "author_url": "",
          "post_date": "2020-11-26T17:45:17.960000",
          "content": "<p>Thank you and congrats to you too! <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> </p>\n<p>We used 8 V100 GPUs for single model training with about 32 cpu cores.<br>\nI think disk IO and rasterization process was the bottleneck compared to the normal image training.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1094891,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2020-11-29T04:11:44.587000",
      "content": "<p>Thanks! <a href=\"https://www.kaggle.com/prokaggler\" target=\"_blank\">@prokaggler</a> </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1102591,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-05T05:04:27.047000",
      "content": "<p>I love this post!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1103475,
          "author_name": "corochann",
          "author_url": "",
          "post_date": "2020-12-05T23:56:34.103000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/YoungseokJoung\" target=\"_blank\">@YoungseokJoung</a> !</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1092267": "Thank you to the organizers and congratulations to all the participants. \nAlso I would like to thank my team members @zaburo, @qhapaq49, @charmq for our hard work, I could enjoy the competition!\nThe LB was really stable in this competition due to the big amount of data, we could work on improving the model without caring about big shake-up.\n\nWe have started with my baseline kernel [Lyft: Training with multi-mode confidence](https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence).\nBelow items are substantial changes we made:\n\n[Update 2020/12/9] **We have published our code:**\n - https://github.com/pfnet-research/kaggle-lyft-motion-prediction-4th-place-solution\n\n# Short Summary\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F518134%2Fa27c649533aca3f12e23255ed2295440%2Flyft_4th_place_solution.png?generation=1606414075335190&alt=media)\n\nPublished baseline training pipeline was indeed already very strong. \nJust modifying\n\n1. train_full.zarr\n2. l5kit==1.1.0\n3. Set min_history=0, min_future=10 in AgentDataset\n4. Cosine annealing for LR decrease until 0\n\nwith training 1 epoch was already enough to win the prize.\n\n\n\n# 1. Use train_full.zarr data\n\nBigger data is almost always better for deep learning model training. We used [Lyft Full Training Set](https://www.kaggle.com/philculliton/lyft-full-training-set).\n\nHowever its size is really large, containing 191M data for AgentDataset.\n\nPractically we need this modifications in order to train this big dataset in real time:\n\n## Distributed training\nWe implemented distributed training using `torch.distributed`.\nIt usually takes about 5 days to finish 1 epoch when we use 8GPUs.\n\n## Caching some arrays into zarr beforehand to reduce on-memory usage in AgentDataset\nThe problem arises when we run distributed training and `DataLoader` with setting `num_workers` for multi-process data loading.\n\nIn the distributed training, 8 processes run in parallel and each process invokes `num_workers` subprocess. Therefore 8 * num_workers subprocess is launched and `AgentDataset` data is copied in each subprocess.\nThen Out Of Memory error occurs because `AgentDataset` internally holds `cumulative_sizes` attribute whose size is very big ([code](https://github.com/lyft/l5kit/blob/1ea2f8cecfe7ad974419bf5c8519ab3a21300119/l5kit/l5kit/dataset/agent.py#L62)).\n\nInstead, we pre-calculated `track_id, scene_index, state_index` and saved as the zarr format. So that we can load each data from disk, and reduce on-memory usage.\n\nThe Public Score was around **25.742** [kernel](https://www.kaggle.com/corochann/lyft-prediction-with-multi-mode-confidence) at this stage.\n\n# 2. Use l5kit==1.1.0\nAs mentioned in the discussion [We did it all wrong](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/186492), image is rotated while the target value is not rotated in the previous version of l5kit==1.0.6 during the beginning of the competition.\nWe updated l5kit version to 1.1.0 once it is released, which fixes this behavior.\n\nThe Public LB score was jumped to **15.874** with this update.\n\n# 3. Set min_history=0, min_future=10\nAs written in the “Validation Strategy” section, validation&test data is made by `create_chopped_dataset` method. We noticed that this validation/test data consists of the data with A. always contains more than 10 future frames, and B sometimes it does not contain any history frame.\n\nTo align the training `AgentDataset` to this test dataset behavior, we can set **`min_frame_history=0` and `min_frame_future=10`**.\nI think **this modification is the most important part to notice in this competition**.\nYou need a courage to intentionally ignore l5kit library warning ;)([code](https://github.com/lyft/l5kit/blob/1ea2f8cecfe7ad974419bf5c8519ab3a21300119/l5kit/l5kit/dataset/agent.py#L15-L17)).\n\nIt’s very effective, the score jumped to **13.059**.\n\n# 4. Training: with cosine annealing\n\nModel: We trained & used following models for final ensemble\n - Resnet18\n - Resnet50\n - SEResNeXt50\n - ecaresnet18\n\nHowever resnet18 baseline was strong, and enough to win the prize.\n\nImage size: tried (128, 128) and (224, 224). image size of 128 training proceeds faster, but image size 224 final score was slightly better.\n\nOptimizer: Adam with Cosine annealing\nCosine annealing was better than Exponential decay. I think decreasing the learning rate until very close to 0 is important for final tuning.\n\nBatch size: 12 * 8 process = 96\n\nWe just trained only 1 epoch to train full data. We did not downsample any of the data.\n\nPublic LB score for single resnet18 model is **11.341**.\n\n# Augmentation\n## Image augmentation\nMany of the augmentation used in natural images is not appropriate for this competition task, (for example flip augmentation flips the target value as well and not realistic since right-lane, left-lane will change). We tried\n - Cutout\n - Blur\n - Downscale\nusing [albumentations](https://github.com/albumentations-team/albumentations) library.\n\n## Rasterizer-level augmentation\nWhat is different from normal image prediction is that the image is drawn by rasterizer. We can also consider applying augmentation during rasterization.\n\nI tried following augmentation by modifying `BoxRasterizer` ([code](https://github.com/lyft/l5kit/blob/1ea2f8cecfe7ad974419bf5c8519ab3a21300119/l5kit/l5kit/rasterization/box_rasterizer.py)).\n - Drop agent randomly\n    - I assumed that the target agent’s movement does not change so much when the other agent far from the target agent exists or not. So we randomly skip drawing some of the agent boxes.\n - Scale extent size randomly\n    - Even when the other agent size changes a bit, I assume that the target agent’s behavior does not change. So we scaled extent size from factor 0.9~1.1\n\nI thought rasterizer-level augmentation was an interesting idea for this competition task. However we could not see big score improvement actually. Maybe the training dataset size is already big enough and its effect is not so big.\n\nBy **only adding cutout augmentation**, we achieved to get public LB score of **10.846**, already enough to win the prize. \n\n# Validation Strategy\n\nAs discussed in [Validation vs LB score](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/188695), we can run very stable validation using a chopped dataset.\nHowever, it removes ground truth data, format is different from training AgentDataset and difficult to validate during training phase.\n\nWhat we want is `agents_mask_orig_bool` ([code](https://github.com/lyft/l5kit/blob/1ea2f8cecfe7ad974419bf5c8519ab3a21300119/l5kit/l5kit/evaluation/chop_dataset.py#L68)).\nWe saved this `agents_mask_orig_bool` and set it to the `agent_mask` argument of `AgentDataset`.\nThen we could run validation during training.\n\nWe sub-sampled 10000 dataset for fast validation during training, but it differs a lot from the score using a total 190327 validation dataset.\nAt the end phase of the competition, we validated the trained model using a full validation dataset.\n\n# Ensemble: sample trajectory and GMM fitting\nTo improve the score further, how to ensemble is the key question in this task.\nNo golden method is suggested in [the discussion](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/180931) and we came up with an idea to adopt the **Gaussian Mixture Model**.\n\nWe can sample trajectory and fit the sampled points by GMM with the 3 components.\nWe started from the [sklearn implementation](https://scikit-learn.org/stable/modules/generated/sklearn.mixture.GaussianMixture.html).\nSetting `n_components=3` and `covariance_type=”spherical”` achieved the good score.\nHowever sigma is fixed to 1 in this competition [metric](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/overview/evaluation), so we also tried implementing own GMM model with fixing covariance to be 1.\n\nEnsemble by GMM was really effective, we finally achieved **public LB score 10.272/private LB score 9.475**\n\nThus, the ensemble pushed the public LB score from 10.846 to 10.272. But it does not change the final rank this time ;)\n\nBy the way, I saw other participants used k-means clustering for ensembling coords. \nI think the behavior is quite similar with using GMM, since it calls k-means clustering in the initialization of EM algorithm.\n\n# What we tried and not worked\nWe noticed Baseline model & l5kit default rasterizer was already very strong in this competition.\nWe really tried a lot, but many of the attempts failed to improve the scores. I’ll write in the reply section (since it’s already very long).\n",
    "1106968": "We have published our code:\nhttps://github.com/pfnet-research/kaggle-lyft-motion-prediction-4th-place-solution",
    "1096004": "# What we tried and not worked\n\nSorry for the late post, these are some of the items that we tried but not worked.\n\n## Change hyper parameters\n\nWe tried several hyper parameters, especially for rasterizers. But none of them contributed to improve the model's accuracy.\n - image size\n    - We tried 224x224 & 128x128. The image size=128 training is faster especially because Rasterization becomes faster, and its training accuracy is almost the same until the middle of the training. However, its validation loss is a bit (only about 0.5~1.0) worse than image size = 224.\n\n - pixel_size\n    - There are many frames that the car is almost stopping now but starts in the near future. We thought that when the car starts moving, its change in the pixel is very small and CNN cannot detect it when the pixel_size is bit (resolution is rough). We tried to change `pixel_size` from default 0.5 into 0.25 or 0.15 but the accuracy becomes worse.\n\n - num_history\n    - 1. Only short history predictor: Several agents have very few past history. So I thought when we train the model with only 0 past frames (i.e., input only current frame), this specific purpose model performs better for predicting the future with only 0 past frames. In the training phase, **we can train this model using the agents with many past frames, by just input only a current frame.** However, this model’s accuracy is worse than the default 10 history input model.\n    - 2. Long history predictor: Oppositely, having more past information helps to improve the score? To check that hypothesis, we tried to input longer past history frames by setting `history_num_frames=7, 10` with `history_step_size=2` instead of default `history_num_frames=10` with `history_step_size=1`. This model’s accuracy was lower than the original model even if we only chose the validation input to have more than 20 history frames.\n\n\n## Big, deep models\nWe tried `resnet101` & `res152` too, but they did not work well.\n\n## Custom Rasterizer\n\nWe tried implementing our own rasterizer to add more rich information to CNN input, but all of them did not work well to improve the accuracy somehow...\n - `ChannelSemanticRasterizer`\n   - `SemanticRasterizer` draws the semantic in RGB space, using 3 channels. We thought this is not always optimal for CNN input and tried to input 6 channels with 1. road, 2. default lane, 3. green signal lane, 4. yellow signal lane, 5. red signal lane, 6. crosswalk.\n - `TLSemanticRasterizer`:\n    - When we executed EDA, we thought knowing **the red signal length is important**. Because some frames start with the red signal as current, and start in the future when the signal changed to green.\nWe input a signal length by changing the color value so that CNN can understand how long this signal color already continued (since Host car detected the signal).\n - `AgentTypeBoxRasterizer`:\n   - There are 4 agent types: CAR, CYCLIST, PEDESTRIAN and UNKNOWN in the original dataset. But UNKNOWN is not drawn in the original `BoxRasterizer`. Also agent type information is also dropped when drawing boxes.\nWe tried to draw each agent type in different channels including UNKNOWN, to input more precise information. \n\n\n**Speed up rasterizer**\nUse numpy batch operation as much as possible, and replacing implementations  with numba jit computations. Even though it becomes faster in single process, its computation speed-up does not contribute so much for multi-process data preparation during training.\n\n## Train with Agent type\nOther than trying the `AgentTypeBoxRasterizer`, we tried inputting agent type one-hot vector explicitly, but it did not contribute to improve accuracy too.\n\n\n## Multi-agent prediction model\nThe baseline kernel predicts future movement of only target agent. Instead I considered to build a model which predicts all the agent's future movement within the input image.\nThe first I thought this idea speed-ups the training since it can predict multiple agent at once. However sometimes the host car detects agent with very far place, like 400 pixels far away. So we noticed it is difficult to align the image size to fixed size. And we suspended its further trial.\n\n## Yaw correction\nThe below figure shows the biggest error in the validation dataset. The error is extremely high when the `yaw` in the dataset was actually opposite and the model predicted the opposite way to go.\nWe worked hard to check if the test dataset contains this kind of case. Indeed there seems to be some frames whose `yaw` might be opposite, however most of the time the agent is stopping in this case in the test dataset. Even though we fixed the yaw, the LB score was almost unchanged.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F518134%2Fe51c936752b5cd7bd0689f6ef9faebe9%2Fvalidation_worst_error.png?generation=1606717208803896&alt=media)\nThe biggest error in the validation dataset with error=43988.1 !!\n\n\n## Leak check\nWhen we checked the dataset carefully, we noticed that the timestamp & the map position actually overlaps within the train/validation/test dataset.\nWe checked if the test dataset information was leaked from another dataset. But it seems that the timestamp is not aligned in scenes (maybe physically other host car is used to collect data, and timestamp record is not calibrated).\n\n",
    "1092920": "Thank you so much for sharing. Just curious, did you use GCP or a local machine for training? Using v100 in GCP seems quite expansive.",
    "1092700": "Great write-up, thanks for sharing. I like your GMM ensembling.",
    "1092309": "Thanks for sharing and well explanation.\nI totally agree with you!\n[Lyft: Training with multi-mode confidence](https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence) is really strong baseline.\n\nAnd Congrats 4th @corochann and @zaburo, @qhapaq49, @charmq.",
    "1098852": "Congrats @corochann and team. Thanks for sharing your solution.",
    "1092767": "good work and thanks for the writeup!\n\n\"so we also tried implementing own GMM model with fixing covariance to be 1.\"\ndo you have some code or pseudo-code for that? i would like to implement and try.\nthanks",
    "1092471": "Congrats! Wow! You set `min_history=0`! Very interesting idea! Does that increase the amount samples in AgentDataset by a lot?",
    "1092322": "Congrats! What hardware have you used to run such a massive experiment? I have trained my models for days and only covered 34M samples, and you did 191M. How many cores have you used, what was the bottleneck? Quite impressive!",
    "1094891": "Thanks! @prokaggler ",
    "1102591": "I love this post!"
  }
}