{
  "id": 199636,
  "title": "9th place solution",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/writeups/nosound-9th-place-solution",
  "author_name": "",
  "post_date": "2020-11-26T16:57:37.373Z",
  "votes": 45,
  "comment_count": 20,
  "views": 0,
  "content": "<p>First, it was a well organized competition. The engagement from <a href=\"https://www.kaggle.com/lucabergamini\" target=\"_blank\">@lucabergamini</a> and <a href=\"https://www.kaggle.com/iglovikov\" target=\"_blank\">@iglovikov</a> on the forum is exemplarily, and one can see how <a href=\"https://www.kaggle.com/iglovikov\" target=\"_blank\">@iglovikov</a> corrects in this competition problems that he himself has probably experienced when he was a participant, well done. Second, I think Lyft investing and organizing Kaggle competitions is not something to be taken for granted. They give the data and lots of their time, but also expose their proprietary l5kit library to the competitors. I hope it pays off to a degree with the ideas that kagglers generate, but in any case thanks to Lyft for doing this, please continue. </p>\n<p>My solution at the end is surprisingly simple, because no fancy ideas have worked well. Yet, there are several things that helped me, I describe everything below. For me, it is once again a lesson to stick to the solid basics. I remember when I started 3 months ago I was excited to try GANs, as it <a href=\"http://papers.neurips.cc/paper/8308-social-bigat-multimodal-trajectory-forecasting-using-bicycle-gan-and-graph-attention-networks.pdf\" target=\"_blank\">seemed to be the state of the art</a>. I couldn't finish further away from it.</p>\n<h1>Data</h1>\n<h3>Samples selection</h3>\n<p>The dataset parameters are <code>min_history_steps = 0</code> and <code>min_future_steps = 10</code>, which is standard and what was probably used to chop the test data. And I added another layer of filtering, - based on the state index. Stated index is a frame's count inside its scene, ranging from 0 to 250. In test, state index is always 99. If one decides to train on a chopped dataset he basically constraints himself to state index 99. But this cuts too much data, so I built a mask similar to the existing mask and require  <code>min_state_history = 30</code> and <code>min_state_future = 50</code>. I believe this change improved my scores. Without it I would get samples with state index starting from 0, which is not good.</p>\n<h3>zarr files</h3>\n<p>From the beginning I used full train zarr, and two weeks before the end I added the chopped validation and the test zarrs. That last part may sound surprising at first. Basically we have 100 frames in the test and validation, and with my states index filtering I look only at frames 30 to 50 to train on. It is a legitimate data for the training, and my motivation was to expose the models to any specifics of the test dataset, like a raining day. After all the filtering I had that much data for the training:</p>\n<ul>\n<li>140M full train</li>\n<li>1.4M test, frames 30-50</li>\n<li>1.9M validation, frames 30-50</li>\n</ul>\n<h3>Rasterization</h3>\n<ul>\n<li>12 frames <code>[0,1,2,3,4,5,8,10,13,16,20,30]</code> into the past</li>\n<li>3 channels semantic map</li>\n<li>another 3 channels of semantic map for history frame 30, gives traffic light 3s from the past</li>\n<li>future frame, when applicable (which is about <code>87%</code> of time). <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199498\" target=\"_blank\">Discussed here</a>.</li>\n<li>AV vehicle is marked with a black dot (see the above link for an example)</li>\n<li>raster size <code>[256,192]</code>, all the rest of the parameters are defaults</li>\n</ul>\n<p>All of the above modification have helped to improve the score. </p>\n<h3>Sampling</h3>\n<p>We had a lot of training data so I had a luxury to make sure that I did not attend the same sample more than once during the whole training. Basically I generated the order before the training and picked indices from it sequentially. </p>\n<h1>Model</h1>\n<p>The ensemble of 3 models, from <a href=\"https://rwightman.github.io/pytorch-image-models/results/\" target=\"_blank\">timm package</a></p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>train</th>\n<th>validation</th>\n<th>LB public</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>mixnet_m, 3 traj</td>\n<td>12.10</td>\n<td>12.62</td>\n<td>12.976</td>\n</tr>\n<tr>\n<td>mixnet_l, 3 traj</td>\n<td>11.48</td>\n<td>12.02</td>\n<td>12.285</td>\n</tr>\n<tr>\n<td>mixnet_l, 6 traj</td>\n<td>6.88</td>\n<td>7.10</td>\n<td></td>\n</tr>\n<tr>\n<td>ensembled, 3 traj</td>\n<td></td>\n<td>11.61</td>\n<td>11.788</td>\n</tr>\n</tbody>\n</table>\n<p>The approach to the model is straightforward 3 or 6 trajectories prediction, I initially got the idea from <a href=\"https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence\" target=\"_blank\">this kernel</a>, and it stuck. </p>\n<h3>Training</h3>\n<ul>\n<li>the best model <code>mixnet_l</code> was trained for <code>215 (epochs) x 64 (batch size) x 2500 (iterations)=34.4M</code> samples</li>\n<li>Adam optimizer</li>\n<li>learning rate <code>4e-4</code> decaying to <code>1e-7</code>. Most of the training on <code>1e-4</code> - <code>2e-4</code></li>\n<li>Auxiliary losses - also predicting 50 future velocities and yaws. It helped to converge faster, but not sure if it improved the accuracy. </li>\n</ul>\n<h1>Ensembling</h1>\n<p>I am proud of the ensembling procedure that I developed, however I have not tested it against the alternatives, and the benefit is only <code>0.5</code>. It seems like some people report good results with stacking, very interesting what is better. </p>\n<p>In my approach I gave weights <code>[0.8,0.3,0.5]</code> to my 3 models, which multiplied their own confidences. As a result I had 12 (3+3+6) trajectories with some probabilities assigned to them. How to ensemble those? I want to optimize the negative log likelihood metric and get three trajectories with confidences as output. If I assume that each of the 12 trajectories is a ground truth with its probability, then I have a well-defined optimization problem at hand. Still, it is a difficult non-convex optimization problem. By writing down the expression, requesting that the gradient equals to zero (local minimum requirement), I got to a fixed point formulation. Fixed point is an expression of form <code>x = f(x)</code>. From this by making several iterations I can get to an optimum. </p>\n<p>There are a couple of questions. First, does it always converge? In practice, it does always converge in 2-20 iterations. In theory it needs to be proven, I think it can be proven. Second, does it converge to a global minimum and not a local one? Given the problem structure I believe it is a global optimum, but maybe it is a global optimum only almost always. Third, does my original assumption about 12 ground truth trajectories with probabilities create a problem? It seems to me that it is a theoretically sounds approach, but it is a long discussion in itself.</p>\n<p>The advantages of the approach is that I run it in batches on GPU, very fast. The results are stable, and it accepts any number of input trajectories. Specifically, I trained the 6 trajectory model to cover a wider range of possibilities, knowing that I can ensemble them effectively. </p>\n<h1>Things that did not work, all in one big pile</h1>\n<ul>\n<li>weight decay, AdamP, AdamW, one-cycle learning rate policy</li>\n<li>freezing batch normalization layers at the last epochs of training</li>\n<li>changing agents box colors continuously based on their velocity. Bright - fast, dim - slow.</li>\n<li>GAN: train a model to differentiate between real and generated trajectories, with gradient reverse layer to propagate the gradient</li>\n<li>PCA: selecting 25 PCA components for 100 dimensional vector of the possible futures and learning to predict their coefficients</li>\n<li>predicting probability maps (like a segmentation task) and clustering those into 3 trajectories</li>\n<li>adding velocity, acceleration, yaws to the last layer of the network</li>\n<li>predicting trajectory as a difference above constant speed trajectory</li>\n<li>rasterizing output of a model and feeding it to a second level CNN model which outputs small corrections</li>\n<li>bigger raster sizes</li>\n<li>predicting heading and velocity instead of <code>x</code> and <code>y</code></li>\n<li>adding 1-2 fully connected layers on top of the model</li>\n<li>learning a \"specialized\" model, - one output trajectory for fast modes, one for turns etc.</li>\n<li>augmentations: random cutout of semantic map, removing random agents</li>\n<li>resnet18, resnet50, mobilenetv3_large_100, densenet121</li>\n<li>adding constant speed and constant acceleration models to the ensembling with low weights</li>\n<li>smoothing final trajectories with B-splines gave indecently small improvement</li>\n</ul>\n<p>Wow, I tried a lot of stuff. </p>\n<h1>Finally</h1>\n<p>Congratulations to the winners! It was a great competition, truly unique, challenging and rewarding.</p>\n<p>And with that I am happy to become Kaggle competitions grandmaster, number 197. Thank you Kaggle and kagglers for that journey!</p>",
  "messages": [
    {
      "id": "1092124",
      "postDate": "11/26/2020 15:03:29",
      "content": "<p>First, it was a well organized competition. The engagement from <a href=\"https://www.kaggle.com/lucabergamini\" target=\"_blank\">@lucabergamini</a> and <a href=\"https://www.kaggle.com/iglovikov\" target=\"_blank\">@iglovikov</a> on the forum is exemplarily, and one can see how <a href=\"https://www.kaggle.com/iglovikov\" target=\"_blank\">@iglovikov</a> corrects in this competition problems that he himself has probably experienced when he was a participant, well done. Second, I think Lyft investing and organizing Kaggle competitions is not something to be taken for granted. They give the data and lots of their time, but also expose their proprietary l5kit library to the competitors. I hope it pays off to a degree with the ideas that kagglers generate, but in any case thanks to Lyft for doing this, please continue. </p>\n<p>My solution at the end is surprisingly simple, because no fancy ideas have worked well. Yet, there are several things that helped me, I describe everything below. For me, it is once again a lesson to stick to the solid basics. I remember when I started 3 months ago I was excited to try GANs, as it <a href=\"http://papers.neurips.cc/paper/8308-social-bigat-multimodal-trajectory-forecasting-using-bicycle-gan-and-graph-attention-networks.pdf\" target=\"_blank\">seemed to be the state of the art</a>. I couldn't finish further away from it.</p>\n<h1>Data</h1>\n<h3>Samples selection</h3>\n<p>The dataset parameters are <code>min_history_steps = 0</code> and <code>min_future_steps = 10</code>, which is standard and what was probably used to chop the test data. And I added another layer of filtering, - based on the state index. Stated index is a frame's count inside its scene, ranging from 0 to 250. In test, state index is always 99. If one decides to train on a chopped dataset he basically constraints himself to state index 99. But this cuts too much data, so I built a mask similar to the existing mask and require  <code>min_state_history = 30</code> and <code>min_state_future = 50</code>. I believe this change improved my scores. Without it I would get samples with state index starting from 0, which is not good.</p>\n<h3>zarr files</h3>\n<p>From the beginning I used full train zarr, and two weeks before the end I added the chopped validation and the test zarrs. That last part may sound surprising at first. Basically we have 100 frames in the test and validation, and with my states index filtering I look only at frames 30 to 50 to train on. It is a legitimate data for the training, and my motivation was to expose the models to any specifics of the test dataset, like a raining day. After all the filtering I had that much data for the training:</p>\n<ul>\n<li>140M full train</li>\n<li>1.4M test, frames 30-50</li>\n<li>1.9M validation, frames 30-50</li>\n</ul>\n<h3>Rasterization</h3>\n<ul>\n<li>12 frames <code>[0,1,2,3,4,5,8,10,13,16,20,30]</code> into the past</li>\n<li>3 channels semantic map</li>\n<li>another 3 channels of semantic map for history frame 30, gives traffic light 3s from the past</li>\n<li>future frame, when applicable (which is about <code>87%</code> of time). <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199498\" target=\"_blank\">Discussed here</a>.</li>\n<li>AV vehicle is marked with a black dot (see the above link for an example)</li>\n<li>raster size <code>[256,192]</code>, all the rest of the parameters are defaults</li>\n</ul>\n<p>All of the above modification have helped to improve the score. </p>\n<h3>Sampling</h3>\n<p>We had a lot of training data so I had a luxury to make sure that I did not attend the same sample more than once during the whole training. Basically I generated the order before the training and picked indices from it sequentially. </p>\n<h1>Model</h1>\n<p>The ensemble of 3 models, from <a href=\"https://rwightman.github.io/pytorch-image-models/results/\" target=\"_blank\">timm package</a></p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>train</th>\n<th>validation</th>\n<th>LB public</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>mixnet_m, 3 traj</td>\n<td>12.10</td>\n<td>12.62</td>\n<td>12.976</td>\n</tr>\n<tr>\n<td>mixnet_l, 3 traj</td>\n<td>11.48</td>\n<td>12.02</td>\n<td>12.285</td>\n</tr>\n<tr>\n<td>mixnet_l, 6 traj</td>\n<td>6.88</td>\n<td>7.10</td>\n<td></td>\n</tr>\n<tr>\n<td>ensembled, 3 traj</td>\n<td></td>\n<td>11.61</td>\n<td>11.788</td>\n</tr>\n</tbody>\n</table>\n<p>The approach to the model is straightforward 3 or 6 trajectories prediction, I initially got the idea from <a href=\"https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence\" target=\"_blank\">this kernel</a>, and it stuck. </p>\n<h3>Training</h3>\n<ul>\n<li>the best model <code>mixnet_l</code> was trained for <code>215 (epochs) x 64 (batch size) x 2500 (iterations)=34.4M</code> samples</li>\n<li>Adam optimizer</li>\n<li>learning rate <code>4e-4</code> decaying to <code>1e-7</code>. Most of the training on <code>1e-4</code> - <code>2e-4</code></li>\n<li>Auxiliary losses - also predicting 50 future velocities and yaws. It helped to converge faster, but not sure if it improved the accuracy. </li>\n</ul>\n<h1>Ensembling</h1>\n<p>I am proud of the ensembling procedure that I developed, however I have not tested it against the alternatives, and the benefit is only <code>0.5</code>. It seems like some people report good results with stacking, very interesting what is better. </p>\n<p>In my approach I gave weights <code>[0.8,0.3,0.5]</code> to my 3 models, which multiplied their own confidences. As a result I had 12 (3+3+6) trajectories with some probabilities assigned to them. How to ensemble those? I want to optimize the negative log likelihood metric and get three trajectories with confidences as output. If I assume that each of the 12 trajectories is a ground truth with its probability, then I have a well-defined optimization problem at hand. Still, it is a difficult non-convex optimization problem. By writing down the expression, requesting that the gradient equals to zero (local minimum requirement), I got to a fixed point formulation. Fixed point is an expression of form <code>x = f(x)</code>. From this by making several iterations I can get to an optimum. </p>\n<p>There are a couple of questions. First, does it always converge? In practice, it does always converge in 2-20 iterations. In theory it needs to be proven, I think it can be proven. Second, does it converge to a global minimum and not a local one? Given the problem structure I believe it is a global optimum, but maybe it is a global optimum only almost always. Third, does my original assumption about 12 ground truth trajectories with probabilities create a problem? It seems to me that it is a theoretically sounds approach, but it is a long discussion in itself.</p>\n<p>The advantages of the approach is that I run it in batches on GPU, very fast. The results are stable, and it accepts any number of input trajectories. Specifically, I trained the 6 trajectory model to cover a wider range of possibilities, knowing that I can ensemble them effectively. </p>\n<h1>Things that did not work, all in one big pile</h1>\n<ul>\n<li>weight decay, AdamP, AdamW, one-cycle learning rate policy</li>\n<li>freezing batch normalization layers at the last epochs of training</li>\n<li>changing agents box colors continuously based on their velocity. Bright - fast, dim - slow.</li>\n<li>GAN: train a model to differentiate between real and generated trajectories, with gradient reverse layer to propagate the gradient</li>\n<li>PCA: selecting 25 PCA components for 100 dimensional vector of the possible futures and learning to predict their coefficients</li>\n<li>predicting probability maps (like a segmentation task) and clustering those into 3 trajectories</li>\n<li>adding velocity, acceleration, yaws to the last layer of the network</li>\n<li>predicting trajectory as a difference above constant speed trajectory</li>\n<li>rasterizing output of a model and feeding it to a second level CNN model which outputs small corrections</li>\n<li>bigger raster sizes</li>\n<li>predicting heading and velocity instead of <code>x</code> and <code>y</code></li>\n<li>adding 1-2 fully connected layers on top of the model</li>\n<li>learning a \"specialized\" model, - one output trajectory for fast modes, one for turns etc.</li>\n<li>augmentations: random cutout of semantic map, removing random agents</li>\n<li>resnet18, resnet50, mobilenetv3_large_100, densenet121</li>\n<li>adding constant speed and constant acceleration models to the ensembling with low weights</li>\n<li>smoothing final trajectories with B-splines gave indecently small improvement</li>\n</ul>\n<p>Wow, I tried a lot of stuff. </p>\n<h1>Finally</h1>\n<p>Congratulations to the winners! It was a great competition, truly unique, challenging and rewarding.</p>\n<p>And with that I am happy to become Kaggle competitions grandmaster, number 197. Thank you Kaggle and kagglers for that journey!</p>",
      "rawMarkdown": "First, it was a well organized competition. The engagement from @lucabergamini and @iglovikov on the forum is exemplarily, and one can see how @iglovikov corrects in this competition problems that he himself has probably experienced when he was a participant, well done. Second, I think Lyft investing and organizing Kaggle competitions is not something to be taken for granted. They give the data and lots of their time, but also expose their proprietary l5kit library to the competitors. I hope it pays off to a degree with the ideas that kagglers generate, but in any case thanks to Lyft for doing this, please continue. \n\nMy solution at the end is surprisingly simple, because no fancy ideas have worked well. Yet, there are several things that helped me, I describe everything below. For me, it is once again a lesson to stick to the solid basics. I remember when I started 3 months ago I was excited to try GANs, as it [seemed to be the state of the art](http://papers.neurips.cc/paper/8308-social-bigat-multimodal-trajectory-forecasting-using-bicycle-gan-and-graph-attention-networks.pdf). I couldn't finish further away from it.\n\n# Data\n\n### Samples selection\n\nThe dataset parameters are `min_history_steps = 0` and `min_future_steps = 10`, which is standard and what was probably used to chop the test data. And I added another layer of filtering, - based on the state index. Stated index is a frame's count inside its scene, ranging from 0 to 250. In test, state index is always 99. If one decides to train on a chopped dataset he basically constraints himself to state index 99. But this cuts too much data, so I built a mask similar to the existing mask and require  `min_state_history = 30` and `min_state_future = 50`. I believe this change improved my scores. Without it I would get samples with state index starting from 0, which is not good.\n\n### zarr files\n\nFrom the beginning I used full train zarr, and two weeks before the end I added the chopped validation and the test zarrs. That last part may sound surprising at first. Basically we have 100 frames in the test and validation, and with my states index filtering I look only at frames 30 to 50 to train on. It is a legitimate data for the training, and my motivation was to expose the models to any specifics of the test dataset, like a raining day. After all the filtering I had that much data for the training:\n\n- 140M full train\n- 1.4M test, frames 30-50\n- 1.9M validation, frames 30-50\n\n### Rasterization\n\n- 12 frames `[0,1,2,3,4,5,8,10,13,16,20,30]` into the past\n- 3 channels semantic map\n- another 3 channels of semantic map for history frame 30, gives traffic light 3s from the past\n- future frame, when applicable (which is about `87%` of time). [Discussed here](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199498).\n- AV vehicle is marked with a black dot (see the above link for an example)\n- raster size `[256,192]`, all the rest of the parameters are defaults\n\nAll of the above modification have helped to improve the score. \n\n### Sampling\n\nWe had a lot of training data so I had a luxury to make sure that I did not attend the same sample more than once during the whole training. Basically I generated the order before the training and picked indices from it sequentially. \n\n# Model\n\nThe ensemble of 3 models, from [timm package](https://rwightman.github.io/pytorch-image-models/results/)\n\n| Model| train | validation | LB public |\n| --- | ---|---|---|\n| mixnet_m, 3 traj | 12.10 | 12.62 | 12.976 |\n| mixnet_l, 3 traj | 11.48 | 12.02 | 12.285 |\n| mixnet_l, 6 traj | 6.88 | 7.10 |  |\n| ensembled, 3 traj |  | 11.61 | 11.788 |\n\nThe approach to the model is straightforward 3 or 6 trajectories prediction, I initially got the idea from [this kernel](https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence), and it stuck. \n\n### Training\n\n- the best model `mixnet_l` was trained for `215 (epochs) x 64 (batch size) x 2500 (iterations)=34.4M` samples\n- Adam optimizer\n- learning rate `4e-4` decaying to `1e-7`. Most of the training on `1e-4` - `2e-4`\n- Auxiliary losses - also predicting 50 future velocities and yaws. It helped to converge faster, but not sure if it improved the accuracy. \n\n# Ensembling\n\nI am proud of the ensembling procedure that I developed, however I have not tested it against the alternatives, and the benefit is only `0.5`. It seems like some people report good results with stacking, very interesting what is better. \n\nIn my approach I gave weights `[0.8,0.3,0.5]` to my 3 models, which multiplied their own confidences. As a result I had 12 (3+3+6) trajectories with some probabilities assigned to them. How to ensemble those? I want to optimize the negative log likelihood metric and get three trajectories with confidences as output. If I assume that each of the 12 trajectories is a ground truth with its probability, then I have a well-defined optimization problem at hand. Still, it is a difficult non-convex optimization problem. By writing down the expression, requesting that the gradient equals to zero (local minimum requirement), I got to a fixed point formulation. Fixed point is an expression of form `x = f(x)`. From this by making several iterations I can get to an optimum. \n\nThere are a couple of questions. First, does it always converge? In practice, it does always converge in 2-20 iterations. In theory it needs to be proven, I think it can be proven. Second, does it converge to a global minimum and not a local one? Given the problem structure I believe it is a global optimum, but maybe it is a global optimum only almost always. Third, does my original assumption about 12 ground truth trajectories with probabilities create a problem? It seems to me that it is a theoretically sounds approach, but it is a long discussion in itself.\n\nThe advantages of the approach is that I run it in batches on GPU, very fast. The results are stable, and it accepts any number of input trajectories. Specifically, I trained the 6 trajectory model to cover a wider range of possibilities, knowing that I can ensemble them effectively. \n\n# Things that did not work, all in one big pile\n\n- weight decay, AdamP, AdamW, one-cycle learning rate policy\n- freezing batch normalization layers at the last epochs of training\n- changing agents box colors continuously based on their velocity. Bright - fast, dim - slow.\n- GAN: train a model to differentiate between real and generated trajectories, with gradient reverse layer to propagate the gradient\n- PCA: selecting 25 PCA components for 100 dimensional vector of the possible futures and learning to predict their coefficients\n- predicting probability maps (like a segmentation task) and clustering those into 3 trajectories\n- adding velocity, acceleration, yaws to the last layer of the network\n- predicting trajectory as a difference above constant speed trajectory\n- rasterizing output of a model and feeding it to a second level CNN model which outputs small corrections\n- bigger raster sizes\n- predicting heading and velocity instead of `x` and `y`\n- adding 1-2 fully connected layers on top of the model\n- learning a \"specialized\" model, - one output trajectory for fast modes, one for turns etc.\n- augmentations: random cutout of semantic map, removing random agents\n- resnet18, resnet50, mobilenetv3_large_100, densenet121\n- adding constant speed and constant acceleration models to the ensembling with low weights\n- smoothing final trajectories with B-splines gave indecently small improvement\n\nWow, I tried a lot of stuff. \n\n# Finally\n\nCongratulations to the winners! It was a great competition, truly unique, challenging and rewarding.\n\nAnd with that I am happy to become Kaggle competitions grandmaster, number 197. Thank you Kaggle and kagglers for that journey!",
      "votes": null
    },
    {
      "id": "1092176",
      "postDate": "11/26/2020 15:25:50",
      "content": "<p>Thanks for sharing and good explanation!<br>\nAnd Congrats to become GM! <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> </p>",
      "rawMarkdown": "Thanks for sharing and good explanation!\nAnd Congrats to become GM! @zaharch",
      "votes": null
    },
    {
      "id": "1092182",
      "postDate": "11/26/2020 15:28:31",
      "content": "<p>Congrats on getting your silver medal</p>",
      "rawMarkdown": "Congrats on getting your silver medal",
      "votes": null
    },
    {
      "id": "1092199",
      "postDate": "11/26/2020 15:44:14",
      "content": "<p>Thank you for the explaniation, and the help during the competition.<br>\nCongrats for the Grand master rank!</p>",
      "rawMarkdown": "Thank you for the explaniation, and the help during the competition.\nCongrats for the Grand master rank!",
      "votes": null
    },
    {
      "id": "1092249",
      "postDate": "11/26/2020 16:37:03",
      "content": "<p>Thanks :) <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> </p>",
      "rawMarkdown": "Thanks :) @zaharch",
      "votes": null
    },
    {
      "id": "1092274",
      "postDate": "11/26/2020 17:01:27",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> for sharing your solution (and the things that didnt work out) and congrats on becoming a GrandMaster!</p>\n<p>I'm curious to know that how do you effectively keep track of the myriad of experiments and its effects?</p>",
      "rawMarkdown": "Thanks @zaharch for sharing your solution (and the things that didnt work out) and congrats on becoming a GrandMaster!\n\nI'm curious to know that how do you effectively keep track of the myriad of experiments and its effects?",
      "votes": null
    },
    {
      "id": "1092286",
      "postDate": "11/26/2020 17:12:11",
      "content": "<p>Nothing special, I have careful logging into csv files, with versioning of each run. Just an incremental integer, I got to version 81 at the end. For each version I write what it is about, what scores it reaches and what is my conclusion. Then move on to the next version. And I always try to make one change at a time.</p>",
      "rawMarkdown": "Nothing special, I have careful logging into csv files, with versioning of each run. Just an incremental integer, I got to version 81 at the end. For each version I write what it is about, what scores it reaches and what is my conclusion. Then move on to the next version. And I always try to make one change at a time.",
      "votes": null
    },
    {
      "id": "1092288",
      "postDate": "11/26/2020 17:13:32",
      "content": "<p>Really interesting ensembling technique! </p>\n<p>I think the fact that you were so generous in sharing your insights on the directionless trajectories at the start of the competition when you were going for solo gold to reach GM says a huge amount - that was very admirable!</p>\n<p>A well deserved gold medal…</p>",
      "rawMarkdown": "Really interesting ensembling technique! \n\nI think the fact that you were so generous in sharing your insights on the directionless trajectories at the start of the competition when you were going for solo gold to reach GM says a huge amount - that was very admirable!\n\nA well deserved gold medal...",
      "votes": null
    },
    {
      "id": "1092303",
      "postDate": "11/26/2020 17:23:54",
      "content": "<p>Thanks, fergusoci, and congratulations with your great result too. To be honest, posting the directionless trajectories topic was a safe bet. The organizers had already known the issue and worked on a fix. And it was too big of an issue to stay secret for long. Maybe maximum another week and someone else would have reported it. And additionally, I was sure that the top teams would definitely do it correctly, so mot much of an advantage to keep it secret. </p>",
      "rawMarkdown": "Thanks, fergusoci, and congratulations with your great result too. To be honest, posting the directionless trajectories topic was a safe bet. The organizers had already known the issue and worked on a fix. And it was too big of an issue to stay secret for long. Maybe maximum another week and someone else would have reported it. And additionally, I was sure that the top teams would definitely do it correctly, so mot much of an advantage to keep it secret.",
      "votes": null
    },
    {
      "id": "1092321",
      "postDate": "11/26/2020 17:34:27",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a>  and thank you indeed. It stands out the list of things you tried. At some point I'm planning to add the kinematics constraints of a bicycle model (I'm still training) which should be close to the b-spline approach you mentioned. Did you follow a specific criteria before ditching your ideas?</p>",
      "rawMarkdown": "Congrats @zaharch  and thank you indeed. It stands out the list of things you tried. At some point I'm planning to add the kinematics constraints of a bicycle model (I'm still training) which should be close to the b-spline approach you mentioned. Did you follow a specific criteria before ditching your ideas?",
      "votes": null
    },
    {
      "id": "1092396",
      "postDate": "11/26/2020 19:10:15",
      "content": "<p>Not really, just decide on a per-experiment basis what is appropriate.</p>",
      "rawMarkdown": "Not really, just decide on a per-experiment basis what is appropriate.",
      "votes": null
    },
    {
      "id": "1092569",
      "postDate": "11/27/2020 02:05:44",
      "content": "<p>Congrats on solo gold and becoming GM <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> Good job! </p>",
      "rawMarkdown": "Congrats on solo gold and becoming GM @zaharch Good job!",
      "votes": null
    },
    {
      "id": "1092677",
      "postDate": "11/27/2020 04:46:41",
      "content": "<p>Congratulations with GM! Well deserved!</p>\n<p>We used the similar approach to ensembling, tried to minimize exactly the same function, but with custom models weights.</p>\n<p>Initially we tried to use the scipy optimize with BFGS solver, but later switches to the separate model (Set Transformer based, great find by <a href=\"https://www.kaggle.com/aruchomu\" target=\"_blank\">@aruchomu</a> ), trained to minimize this function from the set of trajectory inputs, it performed better and likely found a better minimum. What is interesting, it worked better to minimize this loss (predict 3 trajectories matching the best all inputs) than to minimize the loss to known GT trajectories. When we first used the optimizer, it boosted the score by 0.8 or so, but with better model trained on the full dataset the improvement was around 0.2-0.3 with the optimizer and 0.4-0.5 with the optimizing model.</p>\n<p>We also used one of models predicting 16 modes instead of 3, it worked well in the ensemble.</p>",
      "rawMarkdown": "Congratulations with GM! Well deserved!\n\nWe used the similar approach to ensembling, tried to minimize exactly the same function, but with custom models weights.\n\nInitially we tried to use the scipy optimize with BFGS solver, but later switches to the separate model (Set Transformer based, great find by @aruchomu ), trained to minimize this function from the set of trajectory inputs, it performed better and likely found a better minimum. What is interesting, it worked better to minimize this loss (predict 3 trajectories matching the best all inputs) than to minimize the loss to known GT trajectories. When we first used the optimizer, it boosted the score by 0.8 or so, but with better model trained on the full dataset the improvement was around 0.2-0.3 with the optimizer and 0.4-0.5 with the optimizing model.\n\nWe also used one of models predicting 16 modes instead of 3, it worked well in the ensemble.",
      "votes": null
    },
    {
      "id": "1092862",
      "postDate": "11/27/2020 08:23:45",
      "content": "<p>Thanks for that description, very interesting. Is it possible that you share your implementation of the Set Transformer? I want to compare them too. And also, what did you mean by \"than to minimize the loss to known GT trajectories\".</p>",
      "rawMarkdown": "Thanks for that description, very interesting. Is it possible that you share your implementation of the Set Transformer? I want to compare them too. And also, what did you mean by \"than to minimize the loss to known GT trajectories\".",
      "votes": null
    },
    {
      "id": "1092869",
      "postDate": "11/27/2020 08:26:40",
      "content": "<p>\"adding 1-2 fully connected layers on top of the model\"</p>\n<p>i observed this too. i wonder why …</p>",
      "rawMarkdown": "\"adding 1-2 fully connected layers on top of the model\"\n\ni observed this too. i wonder why ...",
      "votes": null
    },
    {
      "id": "1092945",
      "postDate": "11/27/2020 10:17:25",
      "content": "<p>We are planning to share the source code, but may take some time to cleanup everything.</p>\n<p>With the model approach to do ensembling, we can fit the model on the out of fold data (used the folds of validation dataset for testing) to predict 3 trajectories. We tested to train this model to either directly minimize the competition loss using the known trajectories, or to train to minimize the loss of how well 3 predicted trajectories fit the N source ones (exactly the same as you have used).</p>\n<p>We found, if the model is trained to minimize the same loss as you used, it performed much better on unseen data than one trained to minimize the original competition loss to known true trajectories. It also means, since model doesn't use the known labels, we could directly fit the ensembling model on the test data, it gives slightly better results.</p>\n<p>Overall the intuition behind using the model instead of optimiser - as you mentioned it's difficult non-convex optimization problem, so I hoped the model trained on many samples could find the better minimum and less likely get stuck comparing to optimizer.</p>",
      "rawMarkdown": "We are planning to share the source code, but may take some time to cleanup everything.\n\nWith the model approach to do ensembling, we can fit the model on the out of fold data (used the folds of validation dataset for testing) to predict 3 trajectories. We tested to train this model to either directly minimize the competition loss using the known trajectories, or to train to minimize the loss of how well 3 predicted trajectories fit the N source ones (exactly the same as you have used).\n\nWe found, if the model is trained to minimize the same loss as you used, it performed much better on unseen data than one trained to minimize the original competition loss to known true trajectories. It also means, since model doesn't use the known labels, we could directly fit the ensembling model on the test data, it gives slightly better results.\n\nOverall the intuition behind using the model instead of optimiser - as you mentioned it's difficult non-convex optimization problem, so I hoped the model trained on many samples could find the better minimum and less likely get stuck comparing to optimizer.",
      "votes": null
    },
    {
      "id": "1094296",
      "postDate": "11/28/2020 13:24:40",
      "content": "<p>Very good description, thanks. And interesting insights. Regarding better results with \"my\" loss in comparison to the original true trajectories, it may be because the original loss is harder to learn. For the original loss one still needs to learn how to separate into 3 modes smartly. With \"my\" loss it is already separated into good candidates, the search field is much smaller.</p>",
      "rawMarkdown": "Very good description, thanks. And interesting insights. Regarding better results with \"my\" loss in comparison to the original true trajectories, it may be because the original loss is harder to learn. For the original loss one still needs to learn how to separate into 3 modes smartly. With \"my\" loss it is already separated into good candidates, the search field is much smaller.",
      "votes": null
    },
    {
      "id": "1094592",
      "postDate": "11/28/2020 18:40:38",
      "content": "<p>Thank you so much for sharing! Would you consider sharing the code as well? I would like to learn more from it. </p>",
      "rawMarkdown": "Thank you so much for sharing! Would you consider sharing the code as well? I would like to learn more from it.",
      "votes": null
    },
    {
      "id": "1095380",
      "postDate": "11/29/2020 14:54:02",
      "content": "<p>Actually, I think I will not share the code this time. It requires cleaning up, and I don't see too much value in it anyway. </p>",
      "rawMarkdown": "Actually, I think I will not share the code this time. It requires cleaning up, and I don't see too much value in it anyway.",
      "votes": null
    },
    {
      "id": "1098869",
      "postDate": "12/01/2020 23:06:00",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> and thanks for sharing your solution.</p>",
      "rawMarkdown": "Congrats @zaharch and thanks for sharing your solution.",
      "votes": null
    },
    {
      "id": "1103463",
      "postDate": "12/05/2020 23:41:04",
      "content": "<p>Thank you for sharing and congrats <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> <br>\ncustom past frame insertion is interesting. and it seems effective after reading 1st place solution! (<a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/201493\" target=\"_blank\">https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/201493</a>)</p>",
      "rawMarkdown": "Thank you for sharing and congrats @zaharch \ncustom past frame insertion is interesting. and it seems effective after reading 1st place solution! (https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/201493)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1092176,
      "author_name": "piantic",
      "author_url": "",
      "post_date": "11/26/2020 15:25:50",
      "content": "<p>Thanks for sharing and good explanation!<br>\nAnd Congrats to become GM! <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 1092182,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "11/26/2020 15:28:31",
          "content": "<p>Congrats on getting your silver medal</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1092249,
          "author_name": "piantic",
          "author_url": "",
          "post_date": "11/26/2020 16:37:03",
          "content": "<p>Thanks :) <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1092199,
      "author_name": "bessenyeiszilrd",
      "author_url": "",
      "post_date": "11/26/2020 15:44:14",
      "content": "<p>Thank you for the explaniation, and the help during the competition.<br>\nCongrats for the Grand master rank!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1092274,
      "author_name": "psvishnu",
      "author_url": "",
      "post_date": "11/26/2020 17:01:27",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> for sharing your solution (and the things that didnt work out) and congrats on becoming a GrandMaster!</p>\n<p>I'm curious to know that how do you effectively keep track of the myriad of experiments and its effects?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1092286,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "11/26/2020 17:12:11",
          "content": "<p>Nothing special, I have careful logging into csv files, with versioning of each run. Just an incremental integer, I got to version 81 at the end. For each version I write what it is about, what scores it reaches and what is my conclusion. Then move on to the next version. And I always try to make one change at a time.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1092288,
      "author_name": "fergusoci",
      "author_url": "",
      "post_date": "11/26/2020 17:13:32",
      "content": "<p>Really interesting ensembling technique! </p>\n<p>I think the fact that you were so generous in sharing your insights on the directionless trajectories at the start of the competition when you were going for solo gold to reach GM says a huge amount - that was very admirable!</p>\n<p>A well deserved gold medal…</p>",
      "votes": null,
      "replies": [
        {
          "id": 1092303,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "11/26/2020 17:23:54",
          "content": "<p>Thanks, fergusoci, and congratulations with your great result too. To be honest, posting the directionless trajectories topic was a safe bet. The organizers had already known the issue and worked on a fix. And it was too big of an issue to stay secret for long. Maybe maximum another week and someone else would have reported it. And additionally, I was sure that the top teams would definitely do it correctly, so mot much of an advantage to keep it secret. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1092321,
      "author_name": "ture05",
      "author_url": "",
      "post_date": "11/26/2020 17:34:27",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a>  and thank you indeed. It stands out the list of things you tried. At some point I'm planning to add the kinematics constraints of a bicycle model (I'm still training) which should be close to the b-spline approach you mentioned. Did you follow a specific criteria before ditching your ideas?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1092396,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "11/26/2020 19:10:15",
          "content": "<p>Not really, just decide on a per-experiment basis what is appropriate.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1092569,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "11/27/2020 02:05:44",
      "content": "<p>Congrats on solo gold and becoming GM <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> Good job! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1092677,
      "author_name": "dmytropoplavskiy",
      "author_url": "",
      "post_date": "11/27/2020 04:46:41",
      "content": "<p>Congratulations with GM! Well deserved!</p>\n<p>We used the similar approach to ensembling, tried to minimize exactly the same function, but with custom models weights.</p>\n<p>Initially we tried to use the scipy optimize with BFGS solver, but later switches to the separate model (Set Transformer based, great find by <a href=\"https://www.kaggle.com/aruchomu\" target=\"_blank\">@aruchomu</a> ), trained to minimize this function from the set of trajectory inputs, it performed better and likely found a better minimum. What is interesting, it worked better to minimize this loss (predict 3 trajectories matching the best all inputs) than to minimize the loss to known GT trajectories. When we first used the optimizer, it boosted the score by 0.8 or so, but with better model trained on the full dataset the improvement was around 0.2-0.3 with the optimizer and 0.4-0.5 with the optimizing model.</p>\n<p>We also used one of models predicting 16 modes instead of 3, it worked well in the ensemble.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1092862,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "11/27/2020 08:23:45",
          "content": "<p>Thanks for that description, very interesting. Is it possible that you share your implementation of the Set Transformer? I want to compare them too. And also, what did you mean by \"than to minimize the loss to known GT trajectories\".</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1092945,
          "author_name": "dmytropoplavskiy",
          "author_url": "",
          "post_date": "11/27/2020 10:17:25",
          "content": "<p>We are planning to share the source code, but may take some time to cleanup everything.</p>\n<p>With the model approach to do ensembling, we can fit the model on the out of fold data (used the folds of validation dataset for testing) to predict 3 trajectories. We tested to train this model to either directly minimize the competition loss using the known trajectories, or to train to minimize the loss of how well 3 predicted trajectories fit the N source ones (exactly the same as you have used).</p>\n<p>We found, if the model is trained to minimize the same loss as you used, it performed much better on unseen data than one trained to minimize the original competition loss to known true trajectories. It also means, since model doesn't use the known labels, we could directly fit the ensembling model on the test data, it gives slightly better results.</p>\n<p>Overall the intuition behind using the model instead of optimiser - as you mentioned it's difficult non-convex optimization problem, so I hoped the model trained on many samples could find the better minimum and less likely get stuck comparing to optimizer.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1094296,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "11/28/2020 13:24:40",
          "content": "<p>Very good description, thanks. And interesting insights. Regarding better results with \"my\" loss in comparison to the original true trajectories, it may be because the original loss is harder to learn. For the original loss one still needs to learn how to separate into 3 modes smartly. With \"my\" loss it is already separated into good candidates, the search field is much smaller.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1092869,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "11/27/2020 08:26:40",
      "content": "<p>\"adding 1-2 fully connected layers on top of the model\"</p>\n<p>i observed this too. i wonder why …</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1094592,
      "author_name": "huanvo",
      "author_url": "",
      "post_date": "11/28/2020 18:40:38",
      "content": "<p>Thank you so much for sharing! Would you consider sharing the code as well? I would like to learn more from it. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1095380,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "11/29/2020 14:54:02",
          "content": "<p>Actually, I think I will not share the code this time. It requires cleaning up, and I don't see too much value in it anyway. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1098869,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "12/01/2020 23:06:00",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> and thanks for sharing your solution.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1103463,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "12/05/2020 23:41:04",
      "content": "<p>Thank you for sharing and congrats <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> <br>\ncustom past frame insertion is interesting. and it seems effective after reading 1st place solution! (<a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/201493\" target=\"_blank\">https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/201493</a>)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1092124": "First, it was a well organized competition. The engagement from @lucabergamini and @iglovikov on the forum is exemplarily, and one can see how @iglovikov corrects in this competition problems that he himself has probably experienced when he was a participant, well done. Second, I think Lyft investing and organizing Kaggle competitions is not something to be taken for granted. They give the data and lots of their time, but also expose their proprietary l5kit library to the competitors. I hope it pays off to a degree with the ideas that kagglers generate, but in any case thanks to Lyft for doing this, please continue. \n\nMy solution at the end is surprisingly simple, because no fancy ideas have worked well. Yet, there are several things that helped me, I describe everything below. For me, it is once again a lesson to stick to the solid basics. I remember when I started 3 months ago I was excited to try GANs, as it [seemed to be the state of the art](http://papers.neurips.cc/paper/8308-social-bigat-multimodal-trajectory-forecasting-using-bicycle-gan-and-graph-attention-networks.pdf). I couldn't finish further away from it.\n\n# Data\n\n### Samples selection\n\nThe dataset parameters are `min_history_steps = 0` and `min_future_steps = 10`, which is standard and what was probably used to chop the test data. And I added another layer of filtering, - based on the state index. Stated index is a frame's count inside its scene, ranging from 0 to 250. In test, state index is always 99. If one decides to train on a chopped dataset he basically constraints himself to state index 99. But this cuts too much data, so I built a mask similar to the existing mask and require  `min_state_history = 30` and `min_state_future = 50`. I believe this change improved my scores. Without it I would get samples with state index starting from 0, which is not good.\n\n### zarr files\n\nFrom the beginning I used full train zarr, and two weeks before the end I added the chopped validation and the test zarrs. That last part may sound surprising at first. Basically we have 100 frames in the test and validation, and with my states index filtering I look only at frames 30 to 50 to train on. It is a legitimate data for the training, and my motivation was to expose the models to any specifics of the test dataset, like a raining day. After all the filtering I had that much data for the training:\n\n- 140M full train\n- 1.4M test, frames 30-50\n- 1.9M validation, frames 30-50\n\n### Rasterization\n\n- 12 frames `[0,1,2,3,4,5,8,10,13,16,20,30]` into the past\n- 3 channels semantic map\n- another 3 channels of semantic map for history frame 30, gives traffic light 3s from the past\n- future frame, when applicable (which is about `87%` of time). [Discussed here](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199498).\n- AV vehicle is marked with a black dot (see the above link for an example)\n- raster size `[256,192]`, all the rest of the parameters are defaults\n\nAll of the above modification have helped to improve the score. \n\n### Sampling\n\nWe had a lot of training data so I had a luxury to make sure that I did not attend the same sample more than once during the whole training. Basically I generated the order before the training and picked indices from it sequentially. \n\n# Model\n\nThe ensemble of 3 models, from [timm package](https://rwightman.github.io/pytorch-image-models/results/)\n\n| Model| train | validation | LB public |\n| --- | ---|---|---|\n| mixnet_m, 3 traj | 12.10 | 12.62 | 12.976 |\n| mixnet_l, 3 traj | 11.48 | 12.02 | 12.285 |\n| mixnet_l, 6 traj | 6.88 | 7.10 |  |\n| ensembled, 3 traj |  | 11.61 | 11.788 |\n\nThe approach to the model is straightforward 3 or 6 trajectories prediction, I initially got the idea from [this kernel](https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence), and it stuck. \n\n### Training\n\n- the best model `mixnet_l` was trained for `215 (epochs) x 64 (batch size) x 2500 (iterations)=34.4M` samples\n- Adam optimizer\n- learning rate `4e-4` decaying to `1e-7`. Most of the training on `1e-4` - `2e-4`\n- Auxiliary losses - also predicting 50 future velocities and yaws. It helped to converge faster, but not sure if it improved the accuracy. \n\n# Ensembling\n\nI am proud of the ensembling procedure that I developed, however I have not tested it against the alternatives, and the benefit is only `0.5`. It seems like some people report good results with stacking, very interesting what is better. \n\nIn my approach I gave weights `[0.8,0.3,0.5]` to my 3 models, which multiplied their own confidences. As a result I had 12 (3+3+6) trajectories with some probabilities assigned to them. How to ensemble those? I want to optimize the negative log likelihood metric and get three trajectories with confidences as output. If I assume that each of the 12 trajectories is a ground truth with its probability, then I have a well-defined optimization problem at hand. Still, it is a difficult non-convex optimization problem. By writing down the expression, requesting that the gradient equals to zero (local minimum requirement), I got to a fixed point formulation. Fixed point is an expression of form `x = f(x)`. From this by making several iterations I can get to an optimum. \n\nThere are a couple of questions. First, does it always converge? In practice, it does always converge in 2-20 iterations. In theory it needs to be proven, I think it can be proven. Second, does it converge to a global minimum and not a local one? Given the problem structure I believe it is a global optimum, but maybe it is a global optimum only almost always. Third, does my original assumption about 12 ground truth trajectories with probabilities create a problem? It seems to me that it is a theoretically sounds approach, but it is a long discussion in itself.\n\nThe advantages of the approach is that I run it in batches on GPU, very fast. The results are stable, and it accepts any number of input trajectories. Specifically, I trained the 6 trajectory model to cover a wider range of possibilities, knowing that I can ensemble them effectively. \n\n# Things that did not work, all in one big pile\n\n- weight decay, AdamP, AdamW, one-cycle learning rate policy\n- freezing batch normalization layers at the last epochs of training\n- changing agents box colors continuously based on their velocity. Bright - fast, dim - slow.\n- GAN: train a model to differentiate between real and generated trajectories, with gradient reverse layer to propagate the gradient\n- PCA: selecting 25 PCA components for 100 dimensional vector of the possible futures and learning to predict their coefficients\n- predicting probability maps (like a segmentation task) and clustering those into 3 trajectories\n- adding velocity, acceleration, yaws to the last layer of the network\n- predicting trajectory as a difference above constant speed trajectory\n- rasterizing output of a model and feeding it to a second level CNN model which outputs small corrections\n- bigger raster sizes\n- predicting heading and velocity instead of `x` and `y`\n- adding 1-2 fully connected layers on top of the model\n- learning a \"specialized\" model, - one output trajectory for fast modes, one for turns etc.\n- augmentations: random cutout of semantic map, removing random agents\n- resnet18, resnet50, mobilenetv3_large_100, densenet121\n- adding constant speed and constant acceleration models to the ensembling with low weights\n- smoothing final trajectories with B-splines gave indecently small improvement\n\nWow, I tried a lot of stuff. \n\n# Finally\n\nCongratulations to the winners! It was a great competition, truly unique, challenging and rewarding.\n\nAnd with that I am happy to become Kaggle competitions grandmaster, number 197. Thank you Kaggle and kagglers for that journey!",
    "1092176": "Thanks for sharing and good explanation!\nAnd Congrats to become GM! @zaharch",
    "1092182": "Congrats on getting your silver medal",
    "1092199": "Thank you for the explaniation, and the help during the competition.\nCongrats for the Grand master rank!",
    "1092249": "Thanks :) @zaharch",
    "1092274": "Thanks @zaharch for sharing your solution (and the things that didnt work out) and congrats on becoming a GrandMaster!\n\nI'm curious to know that how do you effectively keep track of the myriad of experiments and its effects?",
    "1092286": "Nothing special, I have careful logging into csv files, with versioning of each run. Just an incremental integer, I got to version 81 at the end. For each version I write what it is about, what scores it reaches and what is my conclusion. Then move on to the next version. And I always try to make one change at a time.",
    "1092288": "Really interesting ensembling technique! \n\nI think the fact that you were so generous in sharing your insights on the directionless trajectories at the start of the competition when you were going for solo gold to reach GM says a huge amount - that was very admirable!\n\nA well deserved gold medal...",
    "1092303": "Thanks, fergusoci, and congratulations with your great result too. To be honest, posting the directionless trajectories topic was a safe bet. The organizers had already known the issue and worked on a fix. And it was too big of an issue to stay secret for long. Maybe maximum another week and someone else would have reported it. And additionally, I was sure that the top teams would definitely do it correctly, so mot much of an advantage to keep it secret.",
    "1092321": "Congrats @zaharch  and thank you indeed. It stands out the list of things you tried. At some point I'm planning to add the kinematics constraints of a bicycle model (I'm still training) which should be close to the b-spline approach you mentioned. Did you follow a specific criteria before ditching your ideas?",
    "1092396": "Not really, just decide on a per-experiment basis what is appropriate.",
    "1092569": "Congrats on solo gold and becoming GM @zaharch Good job!",
    "1092677": "Congratulations with GM! Well deserved!\n\nWe used the similar approach to ensembling, tried to minimize exactly the same function, but with custom models weights.\n\nInitially we tried to use the scipy optimize with BFGS solver, but later switches to the separate model (Set Transformer based, great find by @aruchomu ), trained to minimize this function from the set of trajectory inputs, it performed better and likely found a better minimum. What is interesting, it worked better to minimize this loss (predict 3 trajectories matching the best all inputs) than to minimize the loss to known GT trajectories. When we first used the optimizer, it boosted the score by 0.8 or so, but with better model trained on the full dataset the improvement was around 0.2-0.3 with the optimizer and 0.4-0.5 with the optimizing model.\n\nWe also used one of models predicting 16 modes instead of 3, it worked well in the ensemble.",
    "1092862": "Thanks for that description, very interesting. Is it possible that you share your implementation of the Set Transformer? I want to compare them too. And also, what did you mean by \"than to minimize the loss to known GT trajectories\".",
    "1092869": "\"adding 1-2 fully connected layers on top of the model\"\n\ni observed this too. i wonder why ...",
    "1092945": "We are planning to share the source code, but may take some time to cleanup everything.\n\nWith the model approach to do ensembling, we can fit the model on the out of fold data (used the folds of validation dataset for testing) to predict 3 trajectories. We tested to train this model to either directly minimize the competition loss using the known trajectories, or to train to minimize the loss of how well 3 predicted trajectories fit the N source ones (exactly the same as you have used).\n\nWe found, if the model is trained to minimize the same loss as you used, it performed much better on unseen data than one trained to minimize the original competition loss to known true trajectories. It also means, since model doesn't use the known labels, we could directly fit the ensembling model on the test data, it gives slightly better results.\n\nOverall the intuition behind using the model instead of optimiser - as you mentioned it's difficult non-convex optimization problem, so I hoped the model trained on many samples could find the better minimum and less likely get stuck comparing to optimizer.",
    "1094296": "Very good description, thanks. And interesting insights. Regarding better results with \"my\" loss in comparison to the original true trajectories, it may be because the original loss is harder to learn. For the original loss one still needs to learn how to separate into 3 modes smartly. With \"my\" loss it is already separated into good candidates, the search field is much smaller.",
    "1094592": "Thank you so much for sharing! Would you consider sharing the code as well? I would like to learn more from it.",
    "1095380": "Actually, I think I will not share the code this time. It requires cleaning up, and I don't see too much value in it anyway.",
    "1098869": "Congrats @zaharch and thanks for sharing your solution.",
    "1103463": "Thank you for sharing and congrats @zaharch \ncustom past frame insertion is interesting. and it seems effective after reading 1st place solution! (https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/201493)"
  },
  "source": "meta"
}