{
  "id": 191890,
  "title": "[Method discussion] Learning Lane Graph Representationsfor Motion Forecasting and LaneGCN",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/191890",
  "author_name": "",
  "post_date": "2020-10-19T08:45:31.316263200Z",
  "votes": 12,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Hi all, </p>\n<p>I just want to start a bit of a discussion about methods that don't use (or don't exclusively use) raster images. <br>\nI know one or two people have posted <a href=\"https://www.kaggle.com/kneroma/training-motion-prediction-with-pointnet\" target=\"_blank\">some graph/geometric learning approaches</a>, but I haven't heard a lot since. </p>\n<p>For today I want to discuss <a href=\"https://arxiv.org/pdf/2007.13732.pdf\" target=\"_blank\">this paper here</a>, by Uber ATG, where they introduce the Lane Graph Convolutional Network (LaneGCN).  </p>\n<p>There are some slides here by one of the authors: <a href=\"http://www.cs.toronto.edu/~byang/slides/LaneGCN.pdf\" target=\"_blank\">slides</a>. </p>\n<p>In any case, the method uses a directed graph representation of the map, instead of a raster, as shown below: </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F314368%2Fb4fa0f1bf431808f48b93844d4d76b55%2FScreenshot%20from%202020-10-18%2018-47-05.png?generation=1603039813540816&amp;alt=media\" alt=\"\"> </p>\n<p>Where the different colours represent left, right, successor and predecessor adjacency. </p>\n<p>They then introduce the LaneConv and dilated LaneConv operator, which shares information between the adjacent nodes. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F314368%2F737132299b2cf48c6812f3c00614a241%2FScreenshot%20from%202020-10-18%2018-47-20.png?generation=1603039977347893&amp;alt=media\" alt=\"\"></p>\n<p>Intuitively, the LaneConv operator shares information between the nodes that represent the maps. <br>\nAdditionally, the method uses spatial attention to share information between the actors (the vehicles and pedestrians etc.) and the map nodes: </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F314368%2Faf775c4e21d5b7938e5524fb33a551c4%2FScreenshot%20from%202020-10-18%2018-54-18.png?generation=1603040175073864&amp;alt=media\" alt=\"\"></p>\n<p>The idea of the above is that:</p>\n<ol>\n<li>The map nodes first share information with each other, via the LaneGCN operator. As a dilation is used to gather information about the lanes further ahead/behind the agent. </li>\n<li>Next, spatial attention is used to gather traffic information into the map nodes, i.e. the nodes pays attention to the traffic agents nearest to it. </li>\n<li>Spatial attention again gathers information about the map nodes to the actor nodes. i.e. the actor looks at which nodes are closest to it. </li>\n<li>The actors take a look (again, with attention) to what the other actors in the scene are doing. </li>\n</ol>\n<p>This information is then passed to a pretty standard prediction header to predict possible trajectories for the agent. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F314368%2F0f5a88c954a22ef8ab61555f69e3469b%2FScreenshot%20from%202020-10-19%2010-33-46.png?generation=1603096456984324&amp;alt=media\" alt=\"\"></p>\n<p>In principle, this seems to be an interesting approach that can gather rich information about lane adjacency and traffic, although there are some open questions about how to include some information, e.g. about traffic lights/stop signs and crosswalks, which are easily included in raster images. </p>\n<p>I'll make some more posts about other papers if people are interested in discussing papers like these (<a href=\"https://arxiv.org/pdf/2005.04259.pdf\" target=\"_blank\">VectorNet</a> maybe?)</p>",
  "messages": [
    {
      "id": "1053697",
      "postDate": "10/19/2020 08:45:31",
      "content": "<p>Hi all, </p>\n<p>I just want to start a bit of a discussion about methods that don't use (or don't exclusively use) raster images. <br>\nI know one or two people have posted <a href=\"https://www.kaggle.com/kneroma/training-motion-prediction-with-pointnet\" target=\"_blank\">some graph/geometric learning approaches</a>, but I haven't heard a lot since. </p>\n<p>For today I want to discuss <a href=\"https://arxiv.org/pdf/2007.13732.pdf\" target=\"_blank\">this paper here</a>, by Uber ATG, where they introduce the Lane Graph Convolutional Network (LaneGCN).  </p>\n<p>There are some slides here by one of the authors: <a href=\"http://www.cs.toronto.edu/~byang/slides/LaneGCN.pdf\" target=\"_blank\">slides</a>. </p>\n<p>In any case, the method uses a directed graph representation of the map, instead of a raster, as shown below: </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F314368%2Fb4fa0f1bf431808f48b93844d4d76b55%2FScreenshot%20from%202020-10-18%2018-47-05.png?generation=1603039813540816&amp;alt=media\" alt=\"\"> </p>\n<p>Where the different colours represent left, right, successor and predecessor adjacency. </p>\n<p>They then introduce the LaneConv and dilated LaneConv operator, which shares information between the adjacent nodes. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F314368%2F737132299b2cf48c6812f3c00614a241%2FScreenshot%20from%202020-10-18%2018-47-20.png?generation=1603039977347893&amp;alt=media\" alt=\"\"></p>\n<p>Intuitively, the LaneConv operator shares information between the nodes that represent the maps. <br>\nAdditionally, the method uses spatial attention to share information between the actors (the vehicles and pedestrians etc.) and the map nodes: </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F314368%2Faf775c4e21d5b7938e5524fb33a551c4%2FScreenshot%20from%202020-10-18%2018-54-18.png?generation=1603040175073864&amp;alt=media\" alt=\"\"></p>\n<p>The idea of the above is that:</p>\n<ol>\n<li>The map nodes first share information with each other, via the LaneGCN operator. As a dilation is used to gather information about the lanes further ahead/behind the agent. </li>\n<li>Next, spatial attention is used to gather traffic information into the map nodes, i.e. the nodes pays attention to the traffic agents nearest to it. </li>\n<li>Spatial attention again gathers information about the map nodes to the actor nodes. i.e. the actor looks at which nodes are closest to it. </li>\n<li>The actors take a look (again, with attention) to what the other actors in the scene are doing. </li>\n</ol>\n<p>This information is then passed to a pretty standard prediction header to predict possible trajectories for the agent. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F314368%2F0f5a88c954a22ef8ab61555f69e3469b%2FScreenshot%20from%202020-10-19%2010-33-46.png?generation=1603096456984324&amp;alt=media\" alt=\"\"></p>\n<p>In principle, this seems to be an interesting approach that can gather rich information about lane adjacency and traffic, although there are some open questions about how to include some information, e.g. about traffic lights/stop signs and crosswalks, which are easily included in raster images. </p>\n<p>I'll make some more posts about other papers if people are interested in discussing papers like these (<a href=\"https://arxiv.org/pdf/2005.04259.pdf\" target=\"_blank\">VectorNet</a> maybe?)</p>",
      "rawMarkdown": "Hi all, \n\nI just want to start a bit of a discussion about methods that don't use (or don't exclusively use) raster images. \nI know one or two people have posted [some graph/geometric learning approaches](https://www.kaggle.com/kneroma/training-motion-prediction-with-pointnet), but I haven't heard a lot since. \n\nFor today I want to discuss [this paper here](https://arxiv.org/pdf/2007.13732.pdf), by Uber ATG, where they introduce the Lane Graph Convolutional Network (LaneGCN).  \n\nThere are some slides here by one of the authors: [slides](http://www.cs.toronto.edu/~byang/slides/LaneGCN.pdf). \n\nIn any case, the method uses a directed graph representation of the map, instead of a raster, as shown below: \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F314368%2Fb4fa0f1bf431808f48b93844d4d76b55%2FScreenshot%20from%202020-10-18%2018-47-05.png?generation=1603039813540816&alt=media) \n\nWhere the different colours represent left, right, successor and predecessor adjacency. \n\nThey then introduce the LaneConv and dilated LaneConv operator, which shares information between the adjacent nodes. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F314368%2F737132299b2cf48c6812f3c00614a241%2FScreenshot%20from%202020-10-18%2018-47-20.png?generation=1603039977347893&alt=media)\n\nIntuitively, the LaneConv operator shares information between the nodes that represent the maps. \nAdditionally, the method uses spatial attention to share information between the actors (the vehicles and pedestrians etc.) and the map nodes: \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F314368%2Faf775c4e21d5b7938e5524fb33a551c4%2FScreenshot%20from%202020-10-18%2018-54-18.png?generation=1603040175073864&alt=media)\n\nThe idea of the above is that:\n1.  The map nodes first share information with each other, via the LaneGCN operator. As a dilation is used to gather information about the lanes further ahead/behind the agent. \n2. Next, spatial attention is used to gather traffic information into the map nodes, i.e. the nodes pays attention to the traffic agents nearest to it. \n3. Spatial attention again gathers information about the map nodes to the actor nodes. i.e. the actor looks at which nodes are closest to it. \n4. The actors take a look (again, with attention) to what the other actors in the scene are doing. \n\nThis information is then passed to a pretty standard prediction header to predict possible trajectories for the agent. \n\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F314368%2F0f5a88c954a22ef8ab61555f69e3469b%2FScreenshot%20from%202020-10-19%2010-33-46.png?generation=1603096456984324&alt=media)\n\nIn principle, this seems to be an interesting approach that can gather rich information about lane adjacency and traffic, although there are some open questions about how to include some information, e.g. about traffic lights/stop signs and crosswalks, which are easily included in raster images. \n\n\n\n\nI'll make some more posts about other papers if people are interested in discussing papers like these ([VectorNet](https://arxiv.org/pdf/2005.04259.pdf) maybe?)",
      "votes": null
    },
    {
      "id": "1053702",
      "postDate": "10/19/2020 08:52:27",
      "content": "<p>Good explanation, this is what i learn too till now. Maybe top teams already tried it. Because 99% discussing all about raster… is the main reason that ~ we don't have mapping of agents from frame to frame in a specific scene ? not sure i understand correctly. </p>\n<p>VectorNet is what i felt for this problem. Still understanding how to represent this competition data in respective format.</p>",
      "rawMarkdown": "Good explanation, this is what i learn too till now. Maybe top teams already tried it. Because 99% discussing all about raster... is the main reason that ~ we don't have mapping of agents from frame to frame in a specific scene ? not sure i understand correctly. \n\nVectorNet is what i felt for this problem. Still understanding how to represent this competition data in respective format.",
      "votes": null
    },
    {
      "id": "1053708",
      "postDate": "10/19/2020 09:00:27",
      "content": "<p>Yeah I'm busy reading the VectorNet paper, will maybe post a summary of it later this week. </p>\n<blockquote>\n  <p>we don't have mapping of agents from frame to frame in a specific scene ?</p>\n</blockquote>\n<p>I don't think I understand, we do have all the information about agents in every scene, it's maybe just a bit more work to write a different dataloader</p>",
      "rawMarkdown": "Yeah I'm busy reading the VectorNet paper, will maybe post a summary of it later this week. \n\n> we don't have mapping of agents from frame to frame in a specific scene ?\n\nI don't think I understand, we do have all the information about agents in every scene, it's maybe just a bit more work to write a different dataloader",
      "votes": null
    },
    {
      "id": "1053715",
      "postDate": "10/19/2020 09:08:31",
      "content": "<p>Yes <a href=\"https://www.kaggle.com/fnands\" target=\"_blank\">@fnands</a> -- maybe i didn't started yet. but will see if anyone already tried this experiment</p>",
      "rawMarkdown": "Yes @fnands -- maybe i didn't started yet. but will see if anyone already tried this experiment",
      "votes": null
    },
    {
      "id": "1054249",
      "postDate": "10/19/2020 19:26:47",
      "content": "<p>One question I have is hyper-parameters, there are a lot of them in this method, and not everything is mentioned in the paper.   <br>\nAs an example, what distance between map nodes was used? </p>\n<p>They often used MLP's to embed information, so how many layers do these MLPs have? One hidden layer? Two?</p>",
      "rawMarkdown": "One question I have is hyper-parameters, there are a lot of them in this method, and not everything is mentioned in the paper.   \nAs an example, what distance between map nodes was used? \n\nThey often used MLP's to embed information, so how many layers do these MLPs have? One hidden layer? Two?",
      "votes": null
    },
    {
      "id": "1054617",
      "postDate": "10/20/2020 03:46:27",
      "content": "<p>There's a little bit more info here about gcn, but free article access may be limited <br>\n<a href=\"https://towardsdatascience.com/understanding-graph-convolutional-networks-for-node-classification-a2bfdb7aba7b\" target=\"_blank\">https://towardsdatascience.com/understanding-graph-convolutional-networks-for-node-classification-a2bfdb7aba7b</a></p>\n<p>refers to SEMI-SUPERVISED CLASSIFICATION WITH GRAPH CONVOLUTIONAL NETWORKS<br>\n-Thomas N. Kipf and Max Welling<br>\n<a href=\"https://arxiv.org/pdf/1609.02907.pdf\" target=\"_blank\">https://arxiv.org/pdf/1609.02907.pdf</a></p>",
      "rawMarkdown": "There's a little bit more info here about gcn, but free article access may be limited \nhttps://towardsdatascience.com/understanding-graph-convolutional-networks-for-node-classification-a2bfdb7aba7b\n\nrefers to SEMI-SUPERVISED CLASSIFICATION WITH GRAPH CONVOLUTIONAL NETWORKS\n-Thomas N. Kipf and Max Welling\nhttps://arxiv.org/pdf/1609.02907.pdf",
      "votes": null
    },
    {
      "id": "1054850",
      "postDate": "10/20/2020 08:14:13",
      "content": "<p>Thanks, yeah I'm pretty familiar with GCNs, and that's a pretty good paper to start with, so I'm excited to see that people are trying to apply the concept to some interesting problems. </p>\n<p>I see from your notebooks you've built a lane graph, have you given much consideration to how often to \"sample\" the nodes? </p>\n<p>For now I've tried every 5m, which gives me a graph that looks like this (with the successor adj matrix as edges):</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F314368%2F145ca6d0ced2e414951179d6dee53885%2FlaneGrahp.png?generation=1603181560789971&amp;alt=media\" alt=\"\"></p>\n<p>I've done a few tests with more dense sampling, but it didn't change the results that much. </p>",
      "rawMarkdown": "Thanks, yeah I'm pretty familiar with GCNs, and that's a pretty good paper to start with, so I'm excited to see that people are trying to apply the concept to some interesting problems. \n\nI see from your notebooks you've built a lane graph, have you given much consideration to how often to \"sample\" the nodes? \n\nFor now I've tried every 5m, which gives me a graph that looks like this (with the successor adj matrix as edges):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F314368%2F145ca6d0ced2e414951179d6dee53885%2FlaneGrahp.png?generation=1603181560789971&alt=media)\n\n I've done a few tests with more dense sampling, but it didn't change the results that much.",
      "votes": null
    },
    {
      "id": "1055592",
      "postDate": "10/20/2020 23:50:29",
      "content": "<p>I think your direction estimates would be quite accurate, but how does it estimate velocity?<br>\nI have i'm hoping a few little cheats with a linear and quadratic equation I can possibly make use of, so it should probably be able to resample on the fly. </p>\n<p>Did you end up using a quadtree for localizing traffic in lanes or find another method?</p>",
      "rawMarkdown": "I think your direction estimates would be quite accurate, but how does it estimate velocity?\nI have i'm hoping a few little cheats with a linear and quadratic equation I can possibly make use of, so it should probably be able to resample on the fly. \n\nDid you end up using a quadtree for localizing traffic in lanes or find another method?",
      "votes": null
    },
    {
      "id": "1055926",
      "postDate": "10/21/2020 08:32:48",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/n3n77i\" target=\"_blank\">@n3n77i</a> , I've been thinking about it more based on the paper above. Are you trying a rule-based approach then? </p>\n<p>That could be pretty cool. </p>",
      "rawMarkdown": "Hi @n3n77i , I've been thinking about it more based on the paper above. Are you trying a rule-based approach then? \n\nThat could be pretty cool.",
      "votes": null
    },
    {
      "id": "1056616",
      "postDate": "10/21/2020 22:44:40",
      "content": "<p>I guess clustering could be considered rule-learning in a loose sense, I do have a set of estimates for future acceleration and rotation that apply based on the nearest kernel and the past data</p>\n<p>Kind of in the vein of <a href=\"https://en.wikipedia.org/wiki/Learning_classifier_system\" target=\"_blank\">https://en.wikipedia.org/wiki/Learning_classifier_system</a> with a lazy discovery heuristic</p>",
      "rawMarkdown": "I guess clustering could be considered rule-learning in a loose sense, I do have a set of estimates for future acceleration and rotation that apply based on the nearest kernel and the past data\n\nKind of in the vein of https://en.wikipedia.org/wiki/Learning_classifier_system with a lazy discovery heuristic",
      "votes": null
    },
    {
      "id": "1060169",
      "postDate": "10/25/2020 21:45:27",
      "content": "<p>Hi, fnand. I am an author of the paper and thanks for your interest. Our code will be released soon. Will update in this thread when it happens.</p>\n<p>As for the hyper-parameters, most of them just take common values and are pretty robust. We did not use a hyper-parameter for the distance between map nodes. Instead, we directly use the straight line segments of lane centerlines given by argoverse raw data as lane nodes, for the ease of reproduction. About the MLPs, usually one or two hidden layers are good.</p>",
      "rawMarkdown": "Hi, fnand. I am an author of the paper and thanks for your interest. Our code will be released soon. Will update in this thread when it happens.\n\nAs for the hyper-parameters, most of them just take common values and are pretty robust. We did not use a hyper-parameter for the distance between map nodes. Instead, we directly use the straight line segments of lane centerlines given by argoverse raw data as lane nodes, for the ease of reproduction. About the MLPs, usually one or two hidden layers are good.",
      "votes": null
    },
    {
      "id": "1060522",
      "postDate": "10/26/2020 09:19:11",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/grt123\" target=\"_blank\">@grt123</a> , thanks for the response! Yeah after looking at the Argoverse dataset I realized that it is structured in a way that creating the lane graph is a bit easier. Creating one for this dataset required quite a deep-dive into the map API, plus a couple of tricks. </p>\n<p>Can I ask you a few questions? I think I understand most of it, I just want to be sure that I understand the following correctly: </p>\n<ol>\n<li>You predict the trajectories for all agents simultaneously</li>\n<li>The input sequences and targets are all in the BEV coordinates entered around the ego vehicle, i.e. the coordinates are <strong>not</strong> shifted/rotated for each agent? So if an agent is at <code>(x, y) = (10, 40)</code>, it's trajectory will lead up to that point and the targets will originate from that point?  </li>\n<li>In the prediction header, the LRB has dim=128, but it takes as an input a concatenation of 2 feature vectors each with dim=128 (so in total dim=256), so how is this down-sampled to 128, with a extra linear layer? </li>\n</ol>",
      "rawMarkdown": "Hi @grt123 , thanks for the response! Yeah after looking at the Argoverse dataset I realized that it is structured in a way that creating the lane graph is a bit easier. Creating one for this dataset required quite a deep-dive into the map API, plus a couple of tricks. \n\nCan I ask you a few questions? I think I understand most of it, I just want to be sure that I understand the following correctly: \n\n1. You predict the trajectories for all agents simultaneously\n2. The input sequences and targets are all in the BEV coordinates entered around the ego vehicle, i.e. the coordinates are **not** shifted/rotated for each agent? So if an agent is at `(x, y) = (10, 40)`, it's trajectory will lead up to that point and the targets will originate from that point?  \n3. In the prediction header, the LRB has dim=128, but it takes as an input a concatenation of 2 feature vectors each with dim=128 (so in total dim=256), so how is this down-sampled to 128, with a extra linear layer?",
      "votes": null
    },
    {
      "id": "1061554",
      "postDate": "10/27/2020 04:50:45",
      "content": "<p>Let me first make a clarification. In argoverse, in one scene there is only one interesting actor whose trajectory is required to be predicted. This actor is called \"agent\". Other actors are not considered by the official evaluation metric.</p>\n<ol>\n<li><p>We predict the trajectories of all actors. This better trains the model than only predicting the agent.</p></li>\n<li><p>All the BEV coordinates are centered and rotated around the agent, not ego vehicle. (see page 11, implementation details). Also note that the actor input (past trajectories) are relative displacements, not absolute coordinates (see section 3.1 the second line).</p></li>\n<li><p>Yes, but with a residual block and a linear layer (see section 3.4 last two lines).</p></li>\n</ol>",
      "rawMarkdown": "Let me first make a clarification. In argoverse, in one scene there is only one interesting actor whose trajectory is required to be predicted. This actor is called \"agent\". Other actors are not considered by the official evaluation metric.\n\n1. We predict the trajectories of all actors. This better trains the model than only predicting the agent.\n\n2. All the BEV coordinates are centered and rotated around the agent, not ego vehicle. (see page 11, implementation details). Also note that the actor input (past trajectories) are relative displacements, not absolute coordinates (see section 3.1 the second line).\n\n3. Yes, but with a residual block and a linear layer (see section 3.4 last two lines).",
      "votes": null
    },
    {
      "id": "1061820",
      "postDate": "10/27/2020 10:57:40",
      "content": "<p><a href=\"https://www.kaggle.com/grt123\" target=\"_blank\">@grt123</a> , ah thanks for the clarification. </p>\n<p>I think #2 was my biggest misunderstanding. </p>\n<p>About 3 I understand the structure, I just meant that I assume the downsampling during the residual. I mean, your input is a 256 dim feature vector, and the output of your LRB is a 128 dim vector, so you have to somehow downsample the 256 feature vector before adding it as to the one with dim 128. I used a linear layer, but I just wanted to make sure what you did. </p>",
      "rawMarkdown": "grt123 , ah thanks for the clarification. \n\nI think #2 was my biggest misunderstanding. \n\n\nAbout 3 I understand the structure, I just meant that I assume the downsampling during the residual. I mean, your input is a 256 dim feature vector, and the output of your LRB is a 128 dim vector, so you have to somehow downsample the 256 feature vector before adding it as to the one with dim 128. I used a linear layer, but I just wanted to make sure what you did.",
      "votes": null
    },
    {
      "id": "1067893",
      "postDate": "11/02/2020 20:00:40",
      "content": "<p>Hi, fnand. The code is now available at: <a href=\"https://github.com/uber-research/LaneGCN\" target=\"_blank\">https://github.com/uber-research/LaneGCN</a><br>\nI think that will clarify the confusion. Thanks.</p>",
      "rawMarkdown": "Hi, fnand. The code is now available at: https://github.com/uber-research/LaneGCN\nI think that will clarify the confusion. Thanks.",
      "votes": null
    },
    {
      "id": "1067904",
      "postDate": "11/02/2020 20:14:51",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/grt123\" target=\"_blank\">@grt123</a> , thanks a lot! I'll have a look soon. </p>",
      "rawMarkdown": "Hi @grt123 , thanks a lot! I'll have a look soon.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1053702,
      "author_name": "seshurajup",
      "author_url": "",
      "post_date": "10/19/2020 08:52:27",
      "content": "<p>Good explanation, this is what i learn too till now. Maybe top teams already tried it. Because 99% discussing all about raster… is the main reason that ~ we don't have mapping of agents from frame to frame in a specific scene ? not sure i understand correctly. </p>\n<p>VectorNet is what i felt for this problem. Still understanding how to represent this competition data in respective format.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1053708,
          "author_name": "fnands",
          "author_url": "",
          "post_date": "10/19/2020 09:00:27",
          "content": "<p>Yeah I'm busy reading the VectorNet paper, will maybe post a summary of it later this week. </p>\n<blockquote>\n  <p>we don't have mapping of agents from frame to frame in a specific scene ?</p>\n</blockquote>\n<p>I don't think I understand, we do have all the information about agents in every scene, it's maybe just a bit more work to write a different dataloader</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1053715,
          "author_name": "seshurajup",
          "author_url": "",
          "post_date": "10/19/2020 09:08:31",
          "content": "<p>Yes <a href=\"https://www.kaggle.com/fnands\" target=\"_blank\">@fnands</a> -- maybe i didn't started yet. but will see if anyone already tried this experiment</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1054249,
      "author_name": "fnands",
      "author_url": "",
      "post_date": "10/19/2020 19:26:47",
      "content": "<p>One question I have is hyper-parameters, there are a lot of them in this method, and not everything is mentioned in the paper.   <br>\nAs an example, what distance between map nodes was used? </p>\n<p>They often used MLP's to embed information, so how many layers do these MLPs have? One hidden layer? Two?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1060169,
          "author_name": "grt123",
          "author_url": "",
          "post_date": "10/25/2020 21:45:27",
          "content": "<p>Hi, fnand. I am an author of the paper and thanks for your interest. Our code will be released soon. Will update in this thread when it happens.</p>\n<p>As for the hyper-parameters, most of them just take common values and are pretty robust. We did not use a hyper-parameter for the distance between map nodes. Instead, we directly use the straight line segments of lane centerlines given by argoverse raw data as lane nodes, for the ease of reproduction. About the MLPs, usually one or two hidden layers are good.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1060522,
          "author_name": "fnands",
          "author_url": "",
          "post_date": "10/26/2020 09:19:11",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/grt123\" target=\"_blank\">@grt123</a> , thanks for the response! Yeah after looking at the Argoverse dataset I realized that it is structured in a way that creating the lane graph is a bit easier. Creating one for this dataset required quite a deep-dive into the map API, plus a couple of tricks. </p>\n<p>Can I ask you a few questions? I think I understand most of it, I just want to be sure that I understand the following correctly: </p>\n<ol>\n<li>You predict the trajectories for all agents simultaneously</li>\n<li>The input sequences and targets are all in the BEV coordinates entered around the ego vehicle, i.e. the coordinates are <strong>not</strong> shifted/rotated for each agent? So if an agent is at <code>(x, y) = (10, 40)</code>, it's trajectory will lead up to that point and the targets will originate from that point?  </li>\n<li>In the prediction header, the LRB has dim=128, but it takes as an input a concatenation of 2 feature vectors each with dim=128 (so in total dim=256), so how is this down-sampled to 128, with a extra linear layer? </li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1061554,
          "author_name": "grt123",
          "author_url": "",
          "post_date": "10/27/2020 04:50:45",
          "content": "<p>Let me first make a clarification. In argoverse, in one scene there is only one interesting actor whose trajectory is required to be predicted. This actor is called \"agent\". Other actors are not considered by the official evaluation metric.</p>\n<ol>\n<li><p>We predict the trajectories of all actors. This better trains the model than only predicting the agent.</p></li>\n<li><p>All the BEV coordinates are centered and rotated around the agent, not ego vehicle. (see page 11, implementation details). Also note that the actor input (past trajectories) are relative displacements, not absolute coordinates (see section 3.1 the second line).</p></li>\n<li><p>Yes, but with a residual block and a linear layer (see section 3.4 last two lines).</p></li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1061820,
          "author_name": "fnands",
          "author_url": "",
          "post_date": "10/27/2020 10:57:40",
          "content": "<p><a href=\"https://www.kaggle.com/grt123\" target=\"_blank\">@grt123</a> , ah thanks for the clarification. </p>\n<p>I think #2 was my biggest misunderstanding. </p>\n<p>About 3 I understand the structure, I just meant that I assume the downsampling during the residual. I mean, your input is a 256 dim feature vector, and the output of your LRB is a 128 dim vector, so you have to somehow downsample the 256 feature vector before adding it as to the one with dim 128. I used a linear layer, but I just wanted to make sure what you did. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1067893,
          "author_name": "grt123",
          "author_url": "",
          "post_date": "11/02/2020 20:00:40",
          "content": "<p>Hi, fnand. The code is now available at: <a href=\"https://github.com/uber-research/LaneGCN\" target=\"_blank\">https://github.com/uber-research/LaneGCN</a><br>\nI think that will clarify the confusion. Thanks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1067904,
          "author_name": "fnands",
          "author_url": "",
          "post_date": "11/02/2020 20:14:51",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/grt123\" target=\"_blank\">@grt123</a> , thanks a lot! I'll have a look soon. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1054617,
      "author_name": "n3n77i",
      "author_url": "",
      "post_date": "10/20/2020 03:46:27",
      "content": "<p>There's a little bit more info here about gcn, but free article access may be limited <br>\n<a href=\"https://towardsdatascience.com/understanding-graph-convolutional-networks-for-node-classification-a2bfdb7aba7b\" target=\"_blank\">https://towardsdatascience.com/understanding-graph-convolutional-networks-for-node-classification-a2bfdb7aba7b</a></p>\n<p>refers to SEMI-SUPERVISED CLASSIFICATION WITH GRAPH CONVOLUTIONAL NETWORKS<br>\n-Thomas N. Kipf and Max Welling<br>\n<a href=\"https://arxiv.org/pdf/1609.02907.pdf\" target=\"_blank\">https://arxiv.org/pdf/1609.02907.pdf</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1054850,
          "author_name": "fnands",
          "author_url": "",
          "post_date": "10/20/2020 08:14:13",
          "content": "<p>Thanks, yeah I'm pretty familiar with GCNs, and that's a pretty good paper to start with, so I'm excited to see that people are trying to apply the concept to some interesting problems. </p>\n<p>I see from your notebooks you've built a lane graph, have you given much consideration to how often to \"sample\" the nodes? </p>\n<p>For now I've tried every 5m, which gives me a graph that looks like this (with the successor adj matrix as edges):</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F314368%2F145ca6d0ced2e414951179d6dee53885%2FlaneGrahp.png?generation=1603181560789971&amp;alt=media\" alt=\"\"></p>\n<p>I've done a few tests with more dense sampling, but it didn't change the results that much. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1055592,
          "author_name": "n3n77i",
          "author_url": "",
          "post_date": "10/20/2020 23:50:29",
          "content": "<p>I think your direction estimates would be quite accurate, but how does it estimate velocity?<br>\nI have i'm hoping a few little cheats with a linear and quadratic equation I can possibly make use of, so it should probably be able to resample on the fly. </p>\n<p>Did you end up using a quadtree for localizing traffic in lanes or find another method?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1055926,
          "author_name": "fnands",
          "author_url": "",
          "post_date": "10/21/2020 08:32:48",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/n3n77i\" target=\"_blank\">@n3n77i</a> , I've been thinking about it more based on the paper above. Are you trying a rule-based approach then? </p>\n<p>That could be pretty cool. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1056616,
          "author_name": "n3n77i",
          "author_url": "",
          "post_date": "10/21/2020 22:44:40",
          "content": "<p>I guess clustering could be considered rule-learning in a loose sense, I do have a set of estimates for future acceleration and rotation that apply based on the nearest kernel and the past data</p>\n<p>Kind of in the vein of <a href=\"https://en.wikipedia.org/wiki/Learning_classifier_system\" target=\"_blank\">https://en.wikipedia.org/wiki/Learning_classifier_system</a> with a lazy discovery heuristic</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1053697": "Hi all, \n\nI just want to start a bit of a discussion about methods that don't use (or don't exclusively use) raster images. \nI know one or two people have posted [some graph/geometric learning approaches](https://www.kaggle.com/kneroma/training-motion-prediction-with-pointnet), but I haven't heard a lot since. \n\nFor today I want to discuss [this paper here](https://arxiv.org/pdf/2007.13732.pdf), by Uber ATG, where they introduce the Lane Graph Convolutional Network (LaneGCN).  \n\nThere are some slides here by one of the authors: [slides](http://www.cs.toronto.edu/~byang/slides/LaneGCN.pdf). \n\nIn any case, the method uses a directed graph representation of the map, instead of a raster, as shown below: \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F314368%2Fb4fa0f1bf431808f48b93844d4d76b55%2FScreenshot%20from%202020-10-18%2018-47-05.png?generation=1603039813540816&alt=media) \n\nWhere the different colours represent left, right, successor and predecessor adjacency. \n\nThey then introduce the LaneConv and dilated LaneConv operator, which shares information between the adjacent nodes. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F314368%2F737132299b2cf48c6812f3c00614a241%2FScreenshot%20from%202020-10-18%2018-47-20.png?generation=1603039977347893&alt=media)\n\nIntuitively, the LaneConv operator shares information between the nodes that represent the maps. \nAdditionally, the method uses spatial attention to share information between the actors (the vehicles and pedestrians etc.) and the map nodes: \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F314368%2Faf775c4e21d5b7938e5524fb33a551c4%2FScreenshot%20from%202020-10-18%2018-54-18.png?generation=1603040175073864&alt=media)\n\nThe idea of the above is that:\n1.  The map nodes first share information with each other, via the LaneGCN operator. As a dilation is used to gather information about the lanes further ahead/behind the agent. \n2. Next, spatial attention is used to gather traffic information into the map nodes, i.e. the nodes pays attention to the traffic agents nearest to it. \n3. Spatial attention again gathers information about the map nodes to the actor nodes. i.e. the actor looks at which nodes are closest to it. \n4. The actors take a look (again, with attention) to what the other actors in the scene are doing. \n\nThis information is then passed to a pretty standard prediction header to predict possible trajectories for the agent. \n\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F314368%2F0f5a88c954a22ef8ab61555f69e3469b%2FScreenshot%20from%202020-10-19%2010-33-46.png?generation=1603096456984324&alt=media)\n\nIn principle, this seems to be an interesting approach that can gather rich information about lane adjacency and traffic, although there are some open questions about how to include some information, e.g. about traffic lights/stop signs and crosswalks, which are easily included in raster images. \n\n\n\n\nI'll make some more posts about other papers if people are interested in discussing papers like these ([VectorNet](https://arxiv.org/pdf/2005.04259.pdf) maybe?)",
    "1053702": "Good explanation, this is what i learn too till now. Maybe top teams already tried it. Because 99% discussing all about raster... is the main reason that ~ we don't have mapping of agents from frame to frame in a specific scene ? not sure i understand correctly. \n\nVectorNet is what i felt for this problem. Still understanding how to represent this competition data in respective format.",
    "1053708": "Yeah I'm busy reading the VectorNet paper, will maybe post a summary of it later this week. \n\n> we don't have mapping of agents from frame to frame in a specific scene ?\n\nI don't think I understand, we do have all the information about agents in every scene, it's maybe just a bit more work to write a different dataloader",
    "1053715": "Yes @fnands -- maybe i didn't started yet. but will see if anyone already tried this experiment",
    "1054249": "One question I have is hyper-parameters, there are a lot of them in this method, and not everything is mentioned in the paper.   \nAs an example, what distance between map nodes was used? \n\nThey often used MLP's to embed information, so how many layers do these MLPs have? One hidden layer? Two?",
    "1054617": "There's a little bit more info here about gcn, but free article access may be limited \nhttps://towardsdatascience.com/understanding-graph-convolutional-networks-for-node-classification-a2bfdb7aba7b\n\nrefers to SEMI-SUPERVISED CLASSIFICATION WITH GRAPH CONVOLUTIONAL NETWORKS\n-Thomas N. Kipf and Max Welling\nhttps://arxiv.org/pdf/1609.02907.pdf",
    "1054850": "Thanks, yeah I'm pretty familiar with GCNs, and that's a pretty good paper to start with, so I'm excited to see that people are trying to apply the concept to some interesting problems. \n\nI see from your notebooks you've built a lane graph, have you given much consideration to how often to \"sample\" the nodes? \n\nFor now I've tried every 5m, which gives me a graph that looks like this (with the successor adj matrix as edges):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F314368%2F145ca6d0ced2e414951179d6dee53885%2FlaneGrahp.png?generation=1603181560789971&alt=media)\n\n I've done a few tests with more dense sampling, but it didn't change the results that much.",
    "1055592": "I think your direction estimates would be quite accurate, but how does it estimate velocity?\nI have i'm hoping a few little cheats with a linear and quadratic equation I can possibly make use of, so it should probably be able to resample on the fly. \n\nDid you end up using a quadtree for localizing traffic in lanes or find another method?",
    "1055926": "Hi @n3n77i , I've been thinking about it more based on the paper above. Are you trying a rule-based approach then? \n\nThat could be pretty cool.",
    "1056616": "I guess clustering could be considered rule-learning in a loose sense, I do have a set of estimates for future acceleration and rotation that apply based on the nearest kernel and the past data\n\nKind of in the vein of https://en.wikipedia.org/wiki/Learning_classifier_system with a lazy discovery heuristic",
    "1060169": "Hi, fnand. I am an author of the paper and thanks for your interest. Our code will be released soon. Will update in this thread when it happens.\n\nAs for the hyper-parameters, most of them just take common values and are pretty robust. We did not use a hyper-parameter for the distance between map nodes. Instead, we directly use the straight line segments of lane centerlines given by argoverse raw data as lane nodes, for the ease of reproduction. About the MLPs, usually one or two hidden layers are good.",
    "1060522": "Hi @grt123 , thanks for the response! Yeah after looking at the Argoverse dataset I realized that it is structured in a way that creating the lane graph is a bit easier. Creating one for this dataset required quite a deep-dive into the map API, plus a couple of tricks. \n\nCan I ask you a few questions? I think I understand most of it, I just want to be sure that I understand the following correctly: \n\n1. You predict the trajectories for all agents simultaneously\n2. The input sequences and targets are all in the BEV coordinates entered around the ego vehicle, i.e. the coordinates are **not** shifted/rotated for each agent? So if an agent is at `(x, y) = (10, 40)`, it's trajectory will lead up to that point and the targets will originate from that point?  \n3. In the prediction header, the LRB has dim=128, but it takes as an input a concatenation of 2 feature vectors each with dim=128 (so in total dim=256), so how is this down-sampled to 128, with a extra linear layer?",
    "1061554": "Let me first make a clarification. In argoverse, in one scene there is only one interesting actor whose trajectory is required to be predicted. This actor is called \"agent\". Other actors are not considered by the official evaluation metric.\n\n1. We predict the trajectories of all actors. This better trains the model than only predicting the agent.\n\n2. All the BEV coordinates are centered and rotated around the agent, not ego vehicle. (see page 11, implementation details). Also note that the actor input (past trajectories) are relative displacements, not absolute coordinates (see section 3.1 the second line).\n\n3. Yes, but with a residual block and a linear layer (see section 3.4 last two lines).",
    "1061820": "grt123 , ah thanks for the clarification. \n\nI think #2 was my biggest misunderstanding. \n\n\nAbout 3 I understand the structure, I just meant that I assume the downsampling during the residual. I mean, your input is a 256 dim feature vector, and the output of your LRB is a 128 dim vector, so you have to somehow downsample the 256 feature vector before adding it as to the one with dim 128. I used a linear layer, but I just wanted to make sure what you did.",
    "1067893": "Hi, fnand. The code is now available at: https://github.com/uber-research/LaneGCN\nI think that will clarify the confusion. Thanks.",
    "1067904": "Hi @grt123 , thanks a lot! I'll have a look soon."
  },
  "source": "meta"
}