{
  "id": 199649,
  "title": "7th place solution - Peter & Beluga ",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/writeups/peter-beluga-7th-place-solution-peter-beluga",
  "author_name": "",
  "post_date": "2020-11-26T16:12:29.033Z",
  "votes": 36,
  "comment_count": 5,
  "views": 0,
  "content": "<h4>Acknowledgements</h4>\n<p>Thanks for the organizers for this tough but certainly interesting challenge.<br>\nSpecial thanks for Vladimir Iglovikov and Luca Bergamini for their active forum contribution during the competition.</p>\n<p>Hats off to my team mate <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> by the time I joined him he already had optimized the hell out of l5kit and had a solid training framework.<br>\nThen he managed to boost the training speed even further by <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199583\" target=\"_blank\">rasterizing the images on GPU</a>.<br>\nWith all those improvements we were able to run dozens of experiments with different config parameters and slightly modified encodings during the last months.</p>\n<h4>Back to the future</h4>\n<p>We noticed that the training dataset and chopped validation set had slightly different feature distributions. After some digging we found that the chopped datasets (valid, test) always had availability for at least 10 future frames. It was quite the opposite than the default <code>AgentDataset</code> settings so we used that for training too. </p>\n<pre><code>AgentDataset(\n    cfg, dataset_zarr, gpu_rasterizer, agents_mask=dataset_mask,\n    min_frame_history=1, min_frame_future=10\n)\n</code></pre>\n<p>It helped both in terms of validation consistency and final score.</p>\n<h4>Poor Man's Ensembling</h4>\n<p>We did not hope that blending or any simple heuristic would help to combine different models. (I read clever tricks though and I hear that stacking works too…)</p>\n<p>The speed of the agent matters a lot and we saw that in our experiments. Intuitively for slower objects we used smaller raster size but more history.<br>\nOur final and best submission used three models based on speed (Total distance in the last 1 sec) </p>\n<ul>\n<li>[0-2] Slow model 320x220 with 3 sec history (compressed by 1.5 s) trained for 7+ days on slower examples</li>\n<li>[2-5] Medium model 320x220 with 3 sec history (compressed by 1.5 s) trained for 5+ days [1.5-12]</li>\n<li>[5+] Fast model 480x320 with 1 sec history on separate channels trained for 9+ days on [2+]</li>\n</ul>\n<h4>Things that did not work</h4>\n<ul>\n<li>We tried to use additional meta data (speed, acceleration, position, hour of the day, day of the week etc.) but it did not really help.</li>\n<li>We did not use satellite images at all. We noticed that they could have additional info (especially for pedestrians or cyclists) but it would be too slow.</li>\n<li>Different backbone. We tried a few other networks but mostly used Effnet-B2.</li>\n</ul>",
  "messages": [
    {
      "id": "1092216",
      "postDate": "11/26/2020 16:07:58",
      "content": "<h4>Acknowledgements</h4>\n<p>Thanks for the organizers for this tough but certainly interesting challenge.<br>\nSpecial thanks for Vladimir Iglovikov and Luca Bergamini for their active forum contribution during the competition.</p>\n<p>Hats off to my team mate <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> by the time I joined him he already had optimized the hell out of l5kit and had a solid training framework.<br>\nThen he managed to boost the training speed even further by <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199583\" target=\"_blank\">rasterizing the images on GPU</a>.<br>\nWith all those improvements we were able to run dozens of experiments with different config parameters and slightly modified encodings during the last months.</p>\n<h4>Back to the future</h4>\n<p>We noticed that the training dataset and chopped validation set had slightly different feature distributions. After some digging we found that the chopped datasets (valid, test) always had availability for at least 10 future frames. It was quite the opposite than the default <code>AgentDataset</code> settings so we used that for training too. </p>\n<pre><code>AgentDataset(\n    cfg, dataset_zarr, gpu_rasterizer, agents_mask=dataset_mask,\n    min_frame_history=1, min_frame_future=10\n)\n</code></pre>\n<p>It helped both in terms of validation consistency and final score.</p>\n<h4>Poor Man's Ensembling</h4>\n<p>We did not hope that blending or any simple heuristic would help to combine different models. (I read clever tricks though and I hear that stacking works too…)</p>\n<p>The speed of the agent matters a lot and we saw that in our experiments. Intuitively for slower objects we used smaller raster size but more history.<br>\nOur final and best submission used three models based on speed (Total distance in the last 1 sec) </p>\n<ul>\n<li>[0-2] Slow model 320x220 with 3 sec history (compressed by 1.5 s) trained for 7+ days on slower examples</li>\n<li>[2-5] Medium model 320x220 with 3 sec history (compressed by 1.5 s) trained for 5+ days [1.5-12]</li>\n<li>[5+] Fast model 480x320 with 1 sec history on separate channels trained for 9+ days on [2+]</li>\n</ul>\n<h4>Things that did not work</h4>\n<ul>\n<li>We tried to use additional meta data (speed, acceleration, position, hour of the day, day of the week etc.) but it did not really help.</li>\n<li>We did not use satellite images at all. We noticed that they could have additional info (especially for pedestrians or cyclists) but it would be too slow.</li>\n<li>Different backbone. We tried a few other networks but mostly used Effnet-B2.</li>\n</ul>",
      "rawMarkdown": "#### Acknowledgements \nThanks for the organizers for this tough but certainly interesting challenge.\nSpecial thanks for Vladimir Iglovikov and Luca Bergamini for their active forum contribution during the competition.\n\nHats off to my team mate @pestipeti by the time I joined him he already had optimized the hell out of l5kit and had a solid training framework.\nThen he managed to boost the training speed even further by [rasterizing the images on GPU](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199583).\nWith all those improvements we were able to run dozens of experiments with different config parameters and slightly modified encodings during the last months.\n\n#### Back to the future\nWe noticed that the training dataset and chopped validation set had slightly different feature distributions. After some digging we found that the chopped datasets (valid, test) always had availability for at least 10 future frames. It was quite the opposite than the default `AgentDataset` settings so we used that for training too. \n\n```python\nAgentDataset(\n    cfg, dataset_zarr, gpu_rasterizer, agents_mask=dataset_mask,\n    min_frame_history=1, min_frame_future=10\n)\n```\n\nIt helped both in terms of validation consistency and final score.\n\n#### Poor Man's Ensembling\nWe did not hope that blending or any simple heuristic would help to combine different models. (I read clever tricks though and I hear that stacking works too...)\n\nThe speed of the agent matters a lot and we saw that in our experiments. Intuitively for slower objects we used smaller raster size but more history.\nOur final and best submission used three models based on speed (Total distance in the last 1 sec) \n* [0-2] Slow model 320x220 with 3 sec history (compressed by 1.5 s) trained for 7+ days on slower examples\n* [2-5] Medium model 320x220 with 3 sec history (compressed by 1.5 s) trained for 5+ days [1.5-12]\n* [5+] Fast model 480x320 with 1 sec history on separate channels trained for 9+ days on [2+]\n\n#### Things that did not work\n* We tried to use additional meta data (speed, acceleration, position, hour of the day, day of the week etc.) but it did not really help.\n* We did not use satellite images at all. We noticed that they could have additional info (especially for pedestrians or cyclists) but it would be too slow.\n* Different backbone. We tried a few other networks but mostly used Effnet-B2.",
      "votes": null
    },
    {
      "id": "1092272",
      "postDate": "11/26/2020 16:57:41",
      "content": "<p>Thank you for the writeup, and congrats for the 7th place!</p>",
      "rawMarkdown": "Thank you for the writeup, and congrats for the 7th place!",
      "votes": null
    },
    {
      "id": "1092432",
      "postDate": "11/26/2020 20:10:18",
      "content": "<p>Well done! </p>\n<p>The <code>min_frame_future</code> trick is clever and more generally something I see done in a lot of Kaggle competitions, i.e. making the distributions of train and test as close as possible. </p>\n<p>Also, I better understand your comment on my discussion: these are very long training times indeed so better not forget to save the model. ;)</p>\n<p>Many thanks for sharing and congratulations again!</p>",
      "rawMarkdown": "Well done! \n\nThe `min_frame_future` trick is clever and more generally something I see done in a lot of Kaggle competitions, i.e. making the distributions of train and test as close as possible. \n\nAlso, I better understand your comment on my discussion: these are very long training times indeed so better not forget to save the model. ;)\n\nMany thanks for sharing and congratulations again!",
      "votes": null
    },
    {
      "id": "1092465",
      "postDate": "11/26/2020 20:52:27",
      "content": "<p>I wonder why adding meta data like speed, or acceleration don't work.</p>",
      "rawMarkdown": "I wonder why adding meta data like speed, or acceleration don't work.",
      "votes": null
    },
    {
      "id": "1092744",
      "postDate": "11/27/2020 06:07:47",
      "content": "<p>Thanks for sharing and congratulations <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> and <a href=\"https://www.kaggle.com/beluga\" target=\"_blank\">@beluga</a> ! </p>",
      "rawMarkdown": "Thanks for sharing and congratulations @pestipeti and @beluga !",
      "votes": null
    },
    {
      "id": "1103462",
      "postDate": "12/05/2020 23:39:09",
      "content": "<p>Thank you for sharing and congrats <a href=\"https://www.kaggle.com/gaborfodor\" target=\"_blank\">@gaborfodor</a> <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> !<br>\nI guess many of the participants were interested in GPU rasterizer work during competitioin.</p>",
      "rawMarkdown": "Thank you for sharing and congrats @gaborfodor @pestipeti !\nI guess many of the participants were interested in GPU rasterizer work during competitioin.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1092272,
      "author_name": "bessenyeiszilrd",
      "author_url": "",
      "post_date": "11/26/2020 16:57:41",
      "content": "<p>Thank you for the writeup, and congrats for the 7th place!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1092432,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "11/26/2020 20:10:18",
      "content": "<p>Well done! </p>\n<p>The <code>min_frame_future</code> trick is clever and more generally something I see done in a lot of Kaggle competitions, i.e. making the distributions of train and test as close as possible. </p>\n<p>Also, I better understand your comment on my discussion: these are very long training times indeed so better not forget to save the model. ;)</p>\n<p>Many thanks for sharing and congratulations again!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1092465,
      "author_name": "louis925",
      "author_url": "",
      "post_date": "11/26/2020 20:52:27",
      "content": "<p>I wonder why adding meta data like speed, or acceleration don't work.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1092744,
      "author_name": "ogrellier",
      "author_url": "",
      "post_date": "11/27/2020 06:07:47",
      "content": "<p>Thanks for sharing and congratulations <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> and <a href=\"https://www.kaggle.com/beluga\" target=\"_blank\">@beluga</a> ! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1103462,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "12/05/2020 23:39:09",
      "content": "<p>Thank you for sharing and congrats <a href=\"https://www.kaggle.com/gaborfodor\" target=\"_blank\">@gaborfodor</a> <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> !<br>\nI guess many of the participants were interested in GPU rasterizer work during competitioin.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1092216": "#### Acknowledgements \nThanks for the organizers for this tough but certainly interesting challenge.\nSpecial thanks for Vladimir Iglovikov and Luca Bergamini for their active forum contribution during the competition.\n\nHats off to my team mate @pestipeti by the time I joined him he already had optimized the hell out of l5kit and had a solid training framework.\nThen he managed to boost the training speed even further by [rasterizing the images on GPU](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199583).\nWith all those improvements we were able to run dozens of experiments with different config parameters and slightly modified encodings during the last months.\n\n#### Back to the future\nWe noticed that the training dataset and chopped validation set had slightly different feature distributions. After some digging we found that the chopped datasets (valid, test) always had availability for at least 10 future frames. It was quite the opposite than the default `AgentDataset` settings so we used that for training too. \n\n```python\nAgentDataset(\n    cfg, dataset_zarr, gpu_rasterizer, agents_mask=dataset_mask,\n    min_frame_history=1, min_frame_future=10\n)\n```\n\nIt helped both in terms of validation consistency and final score.\n\n#### Poor Man's Ensembling\nWe did not hope that blending or any simple heuristic would help to combine different models. (I read clever tricks though and I hear that stacking works too...)\n\nThe speed of the agent matters a lot and we saw that in our experiments. Intuitively for slower objects we used smaller raster size but more history.\nOur final and best submission used three models based on speed (Total distance in the last 1 sec) \n* [0-2] Slow model 320x220 with 3 sec history (compressed by 1.5 s) trained for 7+ days on slower examples\n* [2-5] Medium model 320x220 with 3 sec history (compressed by 1.5 s) trained for 5+ days [1.5-12]\n* [5+] Fast model 480x320 with 1 sec history on separate channels trained for 9+ days on [2+]\n\n#### Things that did not work\n* We tried to use additional meta data (speed, acceleration, position, hour of the day, day of the week etc.) but it did not really help.\n* We did not use satellite images at all. We noticed that they could have additional info (especially for pedestrians or cyclists) but it would be too slow.\n* Different backbone. We tried a few other networks but mostly used Effnet-B2.",
    "1092272": "Thank you for the writeup, and congrats for the 7th place!",
    "1092432": "Well done! \n\nThe `min_frame_future` trick is clever and more generally something I see done in a lot of Kaggle competitions, i.e. making the distributions of train and test as close as possible. \n\nAlso, I better understand your comment on my discussion: these are very long training times indeed so better not forget to save the model. ;)\n\nMany thanks for sharing and congratulations again!",
    "1092465": "I wonder why adding meta data like speed, or acceleration don't work.",
    "1092744": "Thanks for sharing and congratulations @pestipeti and @beluga !",
    "1103462": "Thank you for sharing and congrats @gaborfodor @pestipeti !\nI guess many of the participants were interested in GPU rasterizer work during competitioin."
  },
  "source": "meta"
}