{
  "id": 183594,
  "title": "Motion prediction with PointNet",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/183594",
  "author_name": "",
  "post_date": "2020-09-17T10:49:03.112825300Z",
  "votes": 31,
  "comment_count": 3,
  "views": 0,
  "content": "<h1>Preditcting agent motion with PointNet architecture …</h1>\n<p>I've created two notebooks using the <strong>pointnet</strong> model to predict agents' motion :</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/kneroma/inference-motion-prediction-with-pointnet\" target=\"_blank\">Inference</a> kernel</li>\n<li><a href=\"https://www.kaggle.com/kneroma/training-motion-prediction-with-pointnet\" target=\"_blank\">Training</a> kernel</li>\n</ul>\n<h2>So why I'm using a pointnet architechture ?</h2>\n<p>Well,  I think of this competition as predicting next destination for a distribution of agents. On each frame, there are a number of <strong>UNORDERED</strong> agents. And, from  scene to scene, there is no a consistent ID to track an agent : those 2 conditions are the ones which will make a pointnet architecture very useful as pointnet  requires nor  intra-sample order, nor extra-sample order !</p>\n<h2>How do I feed the data to the pointnet model</h2>\n<p>PointNet usually requires a <strong>3D</strong> shaped array as input: (batch_size, feature_size, n_points). So, for each frame inside a scene, I read <strong>HBACKWARD=15</strong> frames from  the past and <strong>HFORWARD=50</strong> frames for the future. Using agent features from  the past, I feed an array of shape (batch_size*n_frames_per_scene,  agent_feature_sizexHBACKWARD, max_agents_count_on_frame) to the <strong>PointNet</strong> which provides me with a global <strong>contextual embedding</strong> and an <strong>agent specific embedding</strong> for each agent. Then, I concatenate the <strong>contextual embedding</strong> with each <strong>agent  embedding</strong>, the resulting array is feeded to a simple feed-forward network which outputs the <strong>300+3</strong> predictions.</p>\n<h2>Issues</h2>\n<ul>\n<li>Usually, we don't know if the agent will still in the scene during the whole HBACKWARD+HFORWARD frames. So, we have to take care to not punishing the model if the agent has gone…this is achieved by computing an <strong>agent_availability</strong> matrix which help getting rid of an agent loss function when he has gone.</li>\n<li>All the agents have not the same history length, so we have to do a zero padding for the shorter ones</li>\n</ul>",
  "messages": [
    {
      "id": "1014314",
      "postDate": "09/17/2020 10:49:03",
      "content": "<h1>Preditcting agent motion with PointNet architecture …</h1>\n<p>I've created two notebooks using the <strong>pointnet</strong> model to predict agents' motion :</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/kneroma/inference-motion-prediction-with-pointnet\" target=\"_blank\">Inference</a> kernel</li>\n<li><a href=\"https://www.kaggle.com/kneroma/training-motion-prediction-with-pointnet\" target=\"_blank\">Training</a> kernel</li>\n</ul>\n<h2>So why I'm using a pointnet architechture ?</h2>\n<p>Well,  I think of this competition as predicting next destination for a distribution of agents. On each frame, there are a number of <strong>UNORDERED</strong> agents. And, from  scene to scene, there is no a consistent ID to track an agent : those 2 conditions are the ones which will make a pointnet architecture very useful as pointnet  requires nor  intra-sample order, nor extra-sample order !</p>\n<h2>How do I feed the data to the pointnet model</h2>\n<p>PointNet usually requires a <strong>3D</strong> shaped array as input: (batch_size, feature_size, n_points). So, for each frame inside a scene, I read <strong>HBACKWARD=15</strong> frames from  the past and <strong>HFORWARD=50</strong> frames for the future. Using agent features from  the past, I feed an array of shape (batch_size*n_frames_per_scene,  agent_feature_sizexHBACKWARD, max_agents_count_on_frame) to the <strong>PointNet</strong> which provides me with a global <strong>contextual embedding</strong> and an <strong>agent specific embedding</strong> for each agent. Then, I concatenate the <strong>contextual embedding</strong> with each <strong>agent  embedding</strong>, the resulting array is feeded to a simple feed-forward network which outputs the <strong>300+3</strong> predictions.</p>\n<h2>Issues</h2>\n<ul>\n<li>Usually, we don't know if the agent will still in the scene during the whole HBACKWARD+HFORWARD frames. So, we have to take care to not punishing the model if the agent has gone…this is achieved by computing an <strong>agent_availability</strong> matrix which help getting rid of an agent loss function when he has gone.</li>\n<li>All the agents have not the same history length, so we have to do a zero padding for the shorter ones</li>\n</ul>",
      "rawMarkdown": "# Preditcting agent motion with PointNet architecture ...\n\nI've created two notebooks using the **pointnet** model to predict agents' motion :\n* [Inference](https://www.kaggle.com/kneroma/inference-motion-prediction-with-pointnet) kernel\n* [Training](https://www.kaggle.com/kneroma/training-motion-prediction-with-pointnet) kernel\n\n## So why I'm using a pointnet architechture ?\nWell,  I think of this competition as predicting next destination for a distribution of agents. On each frame, there are a number of **UNORDERED** agents. And, from  scene to scene, there is no a consistent ID to track an agent : those 2 conditions are the ones which will make a pointnet architecture very useful as pointnet  requires nor  intra-sample order, nor extra-sample order !\n\n## How do I feed the data to the pointnet model\n\nPointNet usually requires a **3D** shaped array as input: (batch_size, feature_size, n_points). So, for each frame inside a scene, I read **HBACKWARD=15** frames from  the past and **HFORWARD=50** frames for the future. Using agent features from  the past, I feed an array of shape (batch_size\\*n_frames_per_scene,  agent_feature_sizexHBACKWARD, max_agents_count_on_frame) to the **PointNet** which provides me with a global **contextual embedding** and an **agent specific embedding** for each agent. Then, I concatenate the **contextual embedding** with each **agent  embedding**, the resulting array is feeded to a simple feed-forward network which outputs the **300+3** predictions.\n\n## Issues\n* Usually, we don't know if the agent will still in the scene during the whole HBACKWARD+HFORWARD frames. So, we have to take care to not punishing the model if the agent has gone...this is achieved by computing an **agent_availability** matrix which help getting rid of an agent loss function when he has gone.\n* All the agents have not the same history length, so we have to do a zero padding for the shorter ones",
      "votes": null
    },
    {
      "id": "1018685",
      "postDate": "09/19/2020 20:47:49",
      "content": "<p>I try to add LSTM and GRU in the last layer. Result worse</p>",
      "rawMarkdown": "I try to add LSTM and GRU in the last layer. Result worse",
      "votes": null
    },
    {
      "id": "1019612",
      "postDate": "09/20/2020 14:37:12",
      "content": "<p>Great, thank you!</p>",
      "rawMarkdown": "Great, thank you!",
      "votes": null
    },
    {
      "id": "1028901",
      "postDate": "09/27/2020 10:01:12",
      "content": "<p>What you are talking about is a special case of previous work UST: <a href=\"https://arxiv.org/abs/2005.02790\" target=\"_blank\">https://arxiv.org/abs/2005.02790</a>, which is accepted by  IROS this year. Hope this helps. Stronger methods based on UST will be online very soon. Stay tuned.</p>",
      "rawMarkdown": "What you are talking about is a special case of previous work UST: https://arxiv.org/abs/2005.02790, which is accepted by  IROS this year. Hope this helps. Stronger methods based on UST will be online very soon. Stay tuned.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1018685,
      "author_name": "doanquanvietnamca",
      "author_url": "",
      "post_date": "09/19/2020 20:47:49",
      "content": "<p>I try to add LSTM and GRU in the last layer. Result worse</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1019612,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "09/20/2020 14:37:12",
      "content": "<p>Great, thank you!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1028901,
      "author_name": "winsty",
      "author_url": "",
      "post_date": "09/27/2020 10:01:12",
      "content": "<p>What you are talking about is a special case of previous work UST: <a href=\"https://arxiv.org/abs/2005.02790\" target=\"_blank\">https://arxiv.org/abs/2005.02790</a>, which is accepted by  IROS this year. Hope this helps. Stronger methods based on UST will be online very soon. Stay tuned.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1014314": "# Preditcting agent motion with PointNet architecture ...\n\nI've created two notebooks using the **pointnet** model to predict agents' motion :\n* [Inference](https://www.kaggle.com/kneroma/inference-motion-prediction-with-pointnet) kernel\n* [Training](https://www.kaggle.com/kneroma/training-motion-prediction-with-pointnet) kernel\n\n## So why I'm using a pointnet architechture ?\nWell,  I think of this competition as predicting next destination for a distribution of agents. On each frame, there are a number of **UNORDERED** agents. And, from  scene to scene, there is no a consistent ID to track an agent : those 2 conditions are the ones which will make a pointnet architecture very useful as pointnet  requires nor  intra-sample order, nor extra-sample order !\n\n## How do I feed the data to the pointnet model\n\nPointNet usually requires a **3D** shaped array as input: (batch_size, feature_size, n_points). So, for each frame inside a scene, I read **HBACKWARD=15** frames from  the past and **HFORWARD=50** frames for the future. Using agent features from  the past, I feed an array of shape (batch_size\\*n_frames_per_scene,  agent_feature_sizexHBACKWARD, max_agents_count_on_frame) to the **PointNet** which provides me with a global **contextual embedding** and an **agent specific embedding** for each agent. Then, I concatenate the **contextual embedding** with each **agent  embedding**, the resulting array is feeded to a simple feed-forward network which outputs the **300+3** predictions.\n\n## Issues\n* Usually, we don't know if the agent will still in the scene during the whole HBACKWARD+HFORWARD frames. So, we have to take care to not punishing the model if the agent has gone...this is achieved by computing an **agent_availability** matrix which help getting rid of an agent loss function when he has gone.\n* All the agents have not the same history length, so we have to do a zero padding for the shorter ones",
    "1018685": "I try to add LSTM and GRU in the last layer. Result worse",
    "1019612": "Great, thank you!",
    "1028901": "What you are talking about is a special case of previous work UST: https://arxiv.org/abs/2005.02790, which is accepted by  IROS this year. Hope this helps. Stronger methods based on UST will be online very soon. Stay tuned."
  },
  "source": "meta"
}