{
  "id": 194128,
  "title": "[Method discussion] Multi-Head Attention for Multi-Modal Joint Vehicle Motion Forecasting",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/194128",
  "author_name": "fnands",
  "post_date": "2020-10-30T18:05:37.960000",
  "votes": 2,
  "comment_count": 0,
  "views": 0,
  "content": "<p>So today we'll have a look at <a href=\"https://arxiv.org/pdf/1910.03650.pdf\" target=\"_blank\">Multi-Head Attention for Multi-Modal Joint Vehicle Motion Forecasting</a></p>\n<p>There is a pretty <a href=\"https://www.youtube.com/watch?v=ZHvdbhOBNYU&amp;feature=youtu.be\" target=\"_blank\">nice talk here</a> by one of the authors, and there is a talk about an updated version of the <a href=\"https://slideslive.com/38923162/argoai-challenge\" target=\"_blank\">method by the same speaker here.</a> </p>\n<p><strong>Model Architecture</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F314368%2F84294bb5ed081735e5bbae14fbe58aee%2FScreenshot%20from%202020-10-30%2017-44-07.png?generation=1604076270402771&amp;alt=media\" alt=\"\"></p>\n<p>Basically the model as described in the paper takes the agent trajectories, encodes them by running a 1D CNN followed by an LSTM over each agent trajectory. <br>\nNext, MHA is used to calculate the interactions between actors, and a preliminary \"trajectory\" is predicted with an LSTM. <br>\nAnother MHA is used to then used over these intermediate trajectories to encode some time dependence, and finally a few linear layers are used to predict the output. </p>\n<p>The output is interesting as the model predicts Gaussian density function for each time point and each mode, instead of just a raw prediction, so some amount of uncertainty can be encoded. </p>\n<p><strong>Updated model</strong></p>\n<p>As you can see, this model in no way takes into account the map, and is solely relies on the information about the other agents in the scene.</p>\n<p>An updated version was used to win the Argoverse challenge, where the information about the lanes was also encoded as follows: </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F314368%2F9cbfed316717228d0b852f7fa95a3a9e%2FScreenshot%20from%202020-10-30%2018-22-32.png?generation=1604080833487414&amp;alt=media\" alt=\"\"></p>\n<p>Now the lanes are similarly encoded to the trajectories, and in the first MHA (yellow block) the agents pay attention to the lanes. <br>\nNext, another MHA is used to pay attention between agents, and finally a prediction is unrolled with an LSTM and decoded into Gaussian components as before. </p>\n<p>In principle, this provides for a reasonable efficient encoding of the scene information, needing only a trajectory per agent and line per lane. </p>\n<p>However, applying it to the Lyft lvl. 5 dataset will take a bit of creativity, as the lane lines are not given. </p>\n<p>What are your thoughts on this method? Any possible drawbacks? </p>\n<p>Right now, this does not pay attention to traffic signs, but one can imagine a few ways in which this can be added. </p>",
  "messages": [
    {
      "id": 1064978,
      "postDate": "2020-10-30T18:05:37.960Z",
      "content": "<p>So today we'll have a look at <a href=\"https://arxiv.org/pdf/1910.03650.pdf\" target=\"_blank\">Multi-Head Attention for Multi-Modal Joint Vehicle Motion Forecasting</a></p>\n<p>There is a pretty <a href=\"https://www.youtube.com/watch?v=ZHvdbhOBNYU&amp;feature=youtu.be\" target=\"_blank\">nice talk here</a> by one of the authors, and there is a talk about an updated version of the <a href=\"https://slideslive.com/38923162/argoai-challenge\" target=\"_blank\">method by the same speaker here.</a> </p>\n<p><strong>Model Architecture</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F314368%2F84294bb5ed081735e5bbae14fbe58aee%2FScreenshot%20from%202020-10-30%2017-44-07.png?generation=1604076270402771&amp;alt=media\" alt=\"\"></p>\n<p>Basically the model as described in the paper takes the agent trajectories, encodes them by running a 1D CNN followed by an LSTM over each agent trajectory. <br>\nNext, MHA is used to calculate the interactions between actors, and a preliminary \"trajectory\" is predicted with an LSTM. <br>\nAnother MHA is used to then used over these intermediate trajectories to encode some time dependence, and finally a few linear layers are used to predict the output. </p>\n<p>The output is interesting as the model predicts Gaussian density function for each time point and each mode, instead of just a raw prediction, so some amount of uncertainty can be encoded. </p>\n<p><strong>Updated model</strong></p>\n<p>As you can see, this model in no way takes into account the map, and is solely relies on the information about the other agents in the scene.</p>\n<p>An updated version was used to win the Argoverse challenge, where the information about the lanes was also encoded as follows: </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F314368%2F9cbfed316717228d0b852f7fa95a3a9e%2FScreenshot%20from%202020-10-30%2018-22-32.png?generation=1604080833487414&amp;alt=media\" alt=\"\"></p>\n<p>Now the lanes are similarly encoded to the trajectories, and in the first MHA (yellow block) the agents pay attention to the lanes. <br>\nNext, another MHA is used to pay attention between agents, and finally a prediction is unrolled with an LSTM and decoded into Gaussian components as before. </p>\n<p>In principle, this provides for a reasonable efficient encoding of the scene information, needing only a trajectory per agent and line per lane. </p>\n<p>However, applying it to the Lyft lvl. 5 dataset will take a bit of creativity, as the lane lines are not given. </p>\n<p>What are your thoughts on this method? Any possible drawbacks? </p>\n<p>Right now, this does not pay attention to traffic signs, but one can imagine a few ways in which this can be added. </p>",
      "rawMarkdown": "So today we'll have a look at [Multi-Head Attention for Multi-Modal Joint Vehicle Motion Forecasting](https://arxiv.org/pdf/1910.03650.pdf)\n\nThere is a pretty [nice talk here](https://www.youtube.com/watch?v=ZHvdbhOBNYU&feature=youtu.be) by one of the authors, and there is a talk about an updated version of the [method by the same speaker here.](https://slideslive.com/38923162/argoai-challenge) \n\n**Model Architecture**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F314368%2F84294bb5ed081735e5bbae14fbe58aee%2FScreenshot%20from%202020-10-30%2017-44-07.png?generation=1604076270402771&alt=media)\n\nBasically the model as described in the paper takes the agent trajectories, encodes them by running a 1D CNN followed by an LSTM over each agent trajectory. \nNext, MHA is used to calculate the interactions between actors, and a preliminary \"trajectory\" is predicted with an LSTM. \nAnother MHA is used to then used over these intermediate trajectories to encode some time dependence, and finally a few linear layers are used to predict the output. \n\nThe output is interesting as the model predicts Gaussian density function for each time point and each mode, instead of just a raw prediction, so some amount of uncertainty can be encoded. \n\n\n **Updated model**\n\nAs you can see, this model in no way takes into account the map, and is solely relies on the information about the other agents in the scene.\n\nAn updated version was used to win the Argoverse challenge, where the information about the lanes was also encoded as follows: \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F314368%2F9cbfed316717228d0b852f7fa95a3a9e%2FScreenshot%20from%202020-10-30%2018-22-32.png?generation=1604080833487414&alt=media)\n \n\nNow the lanes are similarly encoded to the trajectories, and in the first MHA (yellow block) the agents pay attention to the lanes. \nNext, another MHA is used to pay attention between agents, and finally a prediction is unrolled with an LSTM and decoded into Gaussian components as before. \n\n\nIn principle, this provides for a reasonable efficient encoding of the scene information, needing only a trajectory per agent and line per lane. \n\nHowever, applying it to the Lyft lvl. 5 dataset will take a bit of creativity, as the lane lines are not given. \n\n\n\nWhat are your thoughts on this method? Any possible drawbacks? \n\nRight now, this does not pay attention to traffic signs, but one can imagine a few ways in which this can be added. \n ",
      "votes": 2
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1064978": "So today we'll have a look at [Multi-Head Attention for Multi-Modal Joint Vehicle Motion Forecasting](https://arxiv.org/pdf/1910.03650.pdf)\n\nThere is a pretty [nice talk here](https://www.youtube.com/watch?v=ZHvdbhOBNYU&feature=youtu.be) by one of the authors, and there is a talk about an updated version of the [method by the same speaker here.](https://slideslive.com/38923162/argoai-challenge) \n\n**Model Architecture**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F314368%2F84294bb5ed081735e5bbae14fbe58aee%2FScreenshot%20from%202020-10-30%2017-44-07.png?generation=1604076270402771&alt=media)\n\nBasically the model as described in the paper takes the agent trajectories, encodes them by running a 1D CNN followed by an LSTM over each agent trajectory. \nNext, MHA is used to calculate the interactions between actors, and a preliminary \"trajectory\" is predicted with an LSTM. \nAnother MHA is used to then used over these intermediate trajectories to encode some time dependence, and finally a few linear layers are used to predict the output. \n\nThe output is interesting as the model predicts Gaussian density function for each time point and each mode, instead of just a raw prediction, so some amount of uncertainty can be encoded. \n\n\n **Updated model**\n\nAs you can see, this model in no way takes into account the map, and is solely relies on the information about the other agents in the scene.\n\nAn updated version was used to win the Argoverse challenge, where the information about the lanes was also encoded as follows: \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F314368%2F9cbfed316717228d0b852f7fa95a3a9e%2FScreenshot%20from%202020-10-30%2018-22-32.png?generation=1604080833487414&alt=media)\n \n\nNow the lanes are similarly encoded to the trajectories, and in the first MHA (yellow block) the agents pay attention to the lanes. \nNext, another MHA is used to pay attention between agents, and finally a prediction is unrolled with an LSTM and decoded into Gaussian components as before. \n\n\nIn principle, this provides for a reasonable efficient encoding of the scene information, needing only a trajectory per agent and line per lane. \n\nHowever, applying it to the Lyft lvl. 5 dataset will take a bit of creativity, as the lane lines are not given. \n\n\n\nWhat are your thoughts on this method? Any possible drawbacks? \n\nRight now, this does not pay attention to traffic signs, but one can imagine a few ways in which this can be added. \n "
  }
}