{
  "id": 180111,
  "title": "Summary of Behavior Prediction Paper",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/180111",
  "author_name": "",
  "post_date": "2020-09-03T22:30:20.608811Z",
  "votes": 38,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I created an outline to summarize the paper (Deep Learning-based Vehicle Behaviour Prediction for Autonomous Driving Applications: a Review) that was mentioned <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/177113\" target=\"_blank\">here</a></p>\n<p><strong>Deep Learning-based Vehicle Behaviour Prediction for Autonomous Driving Applications: a Review</strong></p>\n<p><strong>Background</strong>: Most modern literature I have seen split the problem of autonomous driving into three separate parts: <em>perception</em>, <em>planning</em>, <em>control</em>. </p>\n<p><em>Perception</em>: understanding what is going on around the vehicle that is being driven. Where are the cars, pedestrians, road lines, signs, lights, etc?  </p>\n<p><em>Planning:</em> using the information we have perceived from the world in order to create a plan for a safe path to drive through. Following lanes, not running into curbs, pedestrians or other vehicles while getting from point A to B as intended</p>\n<p><em>Control<strong>:</strong></em> Given the plan created previously apply the correct acceleration and steering in order to act out the plan. </p>\n<p>In this competition we are focusing on the planning section of the problem. Given a blend of road annotations and outputs from the lyft perception system we are trying to predict the pathing of surrounding vehicles. </p>\n<hr>\n<p><strong>Overview</strong>: One of the most important problems in self-driving vehicles is forecasting the paths of surrounding vehicles. The authors of this paper review the methods that have been explored and different ways the problem has been posed. </p>\n<p>There are a lot of acronyms used in the autonomous space. Important terms are:</p>\n<ul>\n<li><em>Target Vehicles (TVs)<strong>:</strong></em> The vehicle that we are trying to predict the path for</li>\n<li><em>Ego Vehicle (EV)<strong>:</strong></em> This is the car being driven. Our view of the world comes from this point of view. Perception comes from this car.</li>\n<li><em>Surrounding Vehicles (SVs)<strong>:</strong></em> The vehicles that are around the target vehicle and might affect its path</li>\n<li><em>Non Effective Vehicles (NVs)<strong>:</strong></em> surrounding vehicles that have no bearing on the target vehicles behavior. Can be things like parked vehicles, vehicles several lanes away, etc.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F57bc6227fa3827ea70c1571ee904cd47%2FScreen%20Shot%202020-09-03%20at%203.32.40%20PM.png?generation=1599172539573875&amp;alt=media\" alt=\"\"></p>\n<p>In the next section of the paper they define the problem formulation. The simplification from math equation to english is we are trying to predict future time steps from all target vehicles given what we can observe from the ego vehicle. This problem is very difficult to predict for all of these interacting vehicles so the problem can be simplified to a single target vehicle at a time instead of all target vehicles. In this problem we are trying to predict 50 timesteps forward with a step size of 0.1 seconds. </p>\n<p>In the section after this the various kinds of inputs, outputs and prediction methods are discussed. </p>\n<p><strong>Inputs:</strong> </p>\n<ul>\n<li><em>History of the target vehicle</em> - Ego vehicles perception of the target vehicles position across previous timesteps. This can be represented in many different ways. 3d bounding boxes, x,y coordinates, etc. Good for telling general direction of travel</li>\n<li><em>History of TV and SVs</em> - Same as previous level but with the addition of also tracking surrounding vehicles</li>\n<li><em>Simplified Bird’s Eye View</em> - This seems to be what most modern systems are using. The perception system creates a top-down of the road with its understanding of where all the vehicles are and were previously</li>\n<li><em>Raw Sensor Data</em> - Data directly from the cameras and LiDAR and other sensors. Rather than simplifying the inputs to the planning system just directly pass them as inputs to train a planning model on. Difficult because data throughput is very large and might prohibit real time usage</li>\n</ul>\n<p><strong>Outputs:</strong> </p>\n<ul>\n<li><em>Maneuver intention</em> - This is predicting the general plan rather than the exact positional future. Eg. turn right, turn left, go straight. This method is heavily limited by requiring to know all maneuvers possible. Sometimes the future of a car cannot be predefined by a simple set of maneuvers.</li>\n<li><em>Unimodal trajectory</em> - Rather than predicting a maneuver it is possible to predict a trajectory. This is predicting the future path of the target vehicle like we are doing in this competition. This gives much greater granularity and allows the model to predict outside of a fixed list of actions, but falls prey to collapsing to a single predicted path even when it is very likely there could be multiple reasonable future trajectories</li>\n<li><em>Multimodal trajectory</em> - This is an extension of the unimodal trajectory. Rather than predicting a single trajectory we can predict multiple trajectories and assign a confidence to each. Allowing us to output that it is possible a vehicle turns left here but it's more likely it continues going straight. </li>\n<li><em>Occupancy map</em> - This output differs from the others in that rather than trying to output continuous future values representing the trajectory this splits the world into a grid and then predicts occupancy of the grid cells. This makes the problem closer to a classification task instead of regression. This may in some way constrain and simplify the result, but can also run into trade-offs in terms of precision if the grid cells are not sufficiently small. </li>\n</ul>\n<p><strong>Prediction Method:</strong> </p>\n<p>There are many different combinations and permutations of these methods that are possible, but here is the general overview of the techniques that have been explored</p>\n<ul>\n<li><em>Physics based models</em><ul>\n<li>These are more conventional models that apply non-NN based methods, but seem to have been mostly surpassed by NN methods. These include things like simple constant velocity models that make assumptions that a car will continue travelling with the velocity it was last known to have. These can get more complex, but require much more hand-engineering and have been phased out. </li></ul></li>\n<li><em>Recurrent Neural Networks</em><ul>\n<li>There is a heavy temporal aspect to this problem so many people have tried applying recurrent neural networks in order to capture this relationship. </li></ul></li>\n<li><em>Convolutional Neural Networks</em><ul>\n<li>People have also attempted various different methods using CNNs. Convolutions across the time domain, and also the images generated from something like the BEV generated from the perception system. It is also possible to utilize 3D convolutions to look both across time and 2D imagery at the same time. </li></ul></li>\n<li><em>Other methods</em><ul>\n<li><em>Dense NN</em> - On state variables and other condensed representations it is possible to use simple feed forward neural networks, but these methods look fairly limited</li>\n<li><em>Graph NNs</em> - It is possible to repose the problem as a graph problem representing the relationship between the entities on the road and then apply graph convolutions</li>\n<li><em>Combinations of methods</em> - It is possible to do convolutions across the images and then RNNs across the series of images and many different combinations of these techniques across the various different inputs</li></ul></li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F42cfb426c1f511d610fb69773034b95f%2FScreen%20Shot%202020-09-03%20at%203.32.50%20PM.png?generation=1599172584952012&amp;alt=media\" alt=\"\"></p>\n<p><strong>Evaluation Metrics</strong></p>\n<ul>\n<li><em>Classification (Maneuver intention)</em><ul>\n<li>Accuracy</li>\n<li>Precision</li>\n<li>Recall</li>\n<li>F1</li>\n<li>Negative Log Likelihood</li>\n<li>Average Prediction Time - The amount of time it takes for a correct prediction of the intended class to occur. Trying to not only forecast a maneuver but also check that it predicts it early on</li></ul></li>\n<li><em>Regression (Unimodal and Multimodal trajectory)</em><ul>\n<li>Final Displacement Error - difference between final location and predicted final location</li>\n<li>MAE</li>\n<li>RMSE</li></ul></li>\n<li><em>Computation time</em><ul>\n<li>Not typically reported in papers but also highly important to make a model that is able to make predictions on reasonable hardware quickly so it can be run in realtime in a vehicle. </li></ul></li>\n</ul>",
  "messages": [
    {
      "id": "997280",
      "postDate": "09/03/2020 22:30:20",
      "content": "<p>I created an outline to summarize the paper (Deep Learning-based Vehicle Behaviour Prediction for Autonomous Driving Applications: a Review) that was mentioned <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/177113\" target=\"_blank\">here</a></p>\n<p><strong>Deep Learning-based Vehicle Behaviour Prediction for Autonomous Driving Applications: a Review</strong></p>\n<p><strong>Background</strong>: Most modern literature I have seen split the problem of autonomous driving into three separate parts: <em>perception</em>, <em>planning</em>, <em>control</em>. </p>\n<p><em>Perception</em>: understanding what is going on around the vehicle that is being driven. Where are the cars, pedestrians, road lines, signs, lights, etc?  </p>\n<p><em>Planning:</em> using the information we have perceived from the world in order to create a plan for a safe path to drive through. Following lanes, not running into curbs, pedestrians or other vehicles while getting from point A to B as intended</p>\n<p><em>Control<strong>:</strong></em> Given the plan created previously apply the correct acceleration and steering in order to act out the plan. </p>\n<p>In this competition we are focusing on the planning section of the problem. Given a blend of road annotations and outputs from the lyft perception system we are trying to predict the pathing of surrounding vehicles. </p>\n<hr>\n<p><strong>Overview</strong>: One of the most important problems in self-driving vehicles is forecasting the paths of surrounding vehicles. The authors of this paper review the methods that have been explored and different ways the problem has been posed. </p>\n<p>There are a lot of acronyms used in the autonomous space. Important terms are:</p>\n<ul>\n<li><em>Target Vehicles (TVs)<strong>:</strong></em> The vehicle that we are trying to predict the path for</li>\n<li><em>Ego Vehicle (EV)<strong>:</strong></em> This is the car being driven. Our view of the world comes from this point of view. Perception comes from this car.</li>\n<li><em>Surrounding Vehicles (SVs)<strong>:</strong></em> The vehicles that are around the target vehicle and might affect its path</li>\n<li><em>Non Effective Vehicles (NVs)<strong>:</strong></em> surrounding vehicles that have no bearing on the target vehicles behavior. Can be things like parked vehicles, vehicles several lanes away, etc.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F57bc6227fa3827ea70c1571ee904cd47%2FScreen%20Shot%202020-09-03%20at%203.32.40%20PM.png?generation=1599172539573875&amp;alt=media\" alt=\"\"></p>\n<p>In the next section of the paper they define the problem formulation. The simplification from math equation to english is we are trying to predict future time steps from all target vehicles given what we can observe from the ego vehicle. This problem is very difficult to predict for all of these interacting vehicles so the problem can be simplified to a single target vehicle at a time instead of all target vehicles. In this problem we are trying to predict 50 timesteps forward with a step size of 0.1 seconds. </p>\n<p>In the section after this the various kinds of inputs, outputs and prediction methods are discussed. </p>\n<p><strong>Inputs:</strong> </p>\n<ul>\n<li><em>History of the target vehicle</em> - Ego vehicles perception of the target vehicles position across previous timesteps. This can be represented in many different ways. 3d bounding boxes, x,y coordinates, etc. Good for telling general direction of travel</li>\n<li><em>History of TV and SVs</em> - Same as previous level but with the addition of also tracking surrounding vehicles</li>\n<li><em>Simplified Bird’s Eye View</em> - This seems to be what most modern systems are using. The perception system creates a top-down of the road with its understanding of where all the vehicles are and were previously</li>\n<li><em>Raw Sensor Data</em> - Data directly from the cameras and LiDAR and other sensors. Rather than simplifying the inputs to the planning system just directly pass them as inputs to train a planning model on. Difficult because data throughput is very large and might prohibit real time usage</li>\n</ul>\n<p><strong>Outputs:</strong> </p>\n<ul>\n<li><em>Maneuver intention</em> - This is predicting the general plan rather than the exact positional future. Eg. turn right, turn left, go straight. This method is heavily limited by requiring to know all maneuvers possible. Sometimes the future of a car cannot be predefined by a simple set of maneuvers.</li>\n<li><em>Unimodal trajectory</em> - Rather than predicting a maneuver it is possible to predict a trajectory. This is predicting the future path of the target vehicle like we are doing in this competition. This gives much greater granularity and allows the model to predict outside of a fixed list of actions, but falls prey to collapsing to a single predicted path even when it is very likely there could be multiple reasonable future trajectories</li>\n<li><em>Multimodal trajectory</em> - This is an extension of the unimodal trajectory. Rather than predicting a single trajectory we can predict multiple trajectories and assign a confidence to each. Allowing us to output that it is possible a vehicle turns left here but it's more likely it continues going straight. </li>\n<li><em>Occupancy map</em> - This output differs from the others in that rather than trying to output continuous future values representing the trajectory this splits the world into a grid and then predicts occupancy of the grid cells. This makes the problem closer to a classification task instead of regression. This may in some way constrain and simplify the result, but can also run into trade-offs in terms of precision if the grid cells are not sufficiently small. </li>\n</ul>\n<p><strong>Prediction Method:</strong> </p>\n<p>There are many different combinations and permutations of these methods that are possible, but here is the general overview of the techniques that have been explored</p>\n<ul>\n<li><em>Physics based models</em><ul>\n<li>These are more conventional models that apply non-NN based methods, but seem to have been mostly surpassed by NN methods. These include things like simple constant velocity models that make assumptions that a car will continue travelling with the velocity it was last known to have. These can get more complex, but require much more hand-engineering and have been phased out. </li></ul></li>\n<li><em>Recurrent Neural Networks</em><ul>\n<li>There is a heavy temporal aspect to this problem so many people have tried applying recurrent neural networks in order to capture this relationship. </li></ul></li>\n<li><em>Convolutional Neural Networks</em><ul>\n<li>People have also attempted various different methods using CNNs. Convolutions across the time domain, and also the images generated from something like the BEV generated from the perception system. It is also possible to utilize 3D convolutions to look both across time and 2D imagery at the same time. </li></ul></li>\n<li><em>Other methods</em><ul>\n<li><em>Dense NN</em> - On state variables and other condensed representations it is possible to use simple feed forward neural networks, but these methods look fairly limited</li>\n<li><em>Graph NNs</em> - It is possible to repose the problem as a graph problem representing the relationship between the entities on the road and then apply graph convolutions</li>\n<li><em>Combinations of methods</em> - It is possible to do convolutions across the images and then RNNs across the series of images and many different combinations of these techniques across the various different inputs</li></ul></li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F42cfb426c1f511d610fb69773034b95f%2FScreen%20Shot%202020-09-03%20at%203.32.50%20PM.png?generation=1599172584952012&amp;alt=media\" alt=\"\"></p>\n<p><strong>Evaluation Metrics</strong></p>\n<ul>\n<li><em>Classification (Maneuver intention)</em><ul>\n<li>Accuracy</li>\n<li>Precision</li>\n<li>Recall</li>\n<li>F1</li>\n<li>Negative Log Likelihood</li>\n<li>Average Prediction Time - The amount of time it takes for a correct prediction of the intended class to occur. Trying to not only forecast a maneuver but also check that it predicts it early on</li></ul></li>\n<li><em>Regression (Unimodal and Multimodal trajectory)</em><ul>\n<li>Final Displacement Error - difference between final location and predicted final location</li>\n<li>MAE</li>\n<li>RMSE</li></ul></li>\n<li><em>Computation time</em><ul>\n<li>Not typically reported in papers but also highly important to make a model that is able to make predictions on reasonable hardware quickly so it can be run in realtime in a vehicle. </li></ul></li>\n</ul>",
      "rawMarkdown": "I created an outline to summarize the paper (Deep Learning-based Vehicle Behaviour Prediction for Autonomous Driving Applications: a Review) that was mentioned [here](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/177113)\n\n\n\n**Deep Learning-based Vehicle Behaviour Prediction for Autonomous Driving Applications: a Review**\n\n**Background**: Most modern literature I have seen split the problem of autonomous driving into three separate parts: _perception_, _planning_, _control_. \n\n_Perception_: understanding what is going on around the vehicle that is being driven. Where are the cars, pedestrians, road lines, signs, lights, etc?  \n\n_Planning:_ using the information we have perceived from the world in order to create a plan for a safe path to drive through. Following lanes, not running into curbs, pedestrians or other vehicles while getting from point A to B as intended\n\n_Control**:**_ Given the plan created previously apply the correct acceleration and steering in order to act out the plan. \n\nIn this competition we are focusing on the planning section of the problem. Given a blend of road annotations and outputs from the lyft perception system we are trying to predict the pathing of surrounding vehicles. \n\n** **\n\n**Overview**: One of the most important problems in self-driving vehicles is forecasting the paths of surrounding vehicles. The authors of this paper review the methods that have been explored and different ways the problem has been posed. \n\nThere are a lot of acronyms used in the autonomous space. Important terms are:\n\n\n\n*   _Target Vehicles (TVs)**:**_ The vehicle that we are trying to predict the path for\n*   _Ego Vehicle (EV)**:**_ This is the car being driven. Our view of the world comes from this point of view. Perception comes from this car.\n*   _Surrounding Vehicles (SVs)**:**_ The vehicles that are around the target vehicle and might affect its path\n*   _Non Effective Vehicles (NVs)**:**_ surrounding vehicles that have no bearing on the target vehicles behavior. Can be things like parked vehicles, vehicles several lanes away, etc.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F57bc6227fa3827ea70c1571ee904cd47%2FScreen%20Shot%202020-09-03%20at%203.32.40%20PM.png?generation=1599172539573875&alt=media)\n\nIn the next section of the paper they define the problem formulation. The simplification from math equation to english is we are trying to predict future time steps from all target vehicles given what we can observe from the ego vehicle. This problem is very difficult to predict for all of these interacting vehicles so the problem can be simplified to a single target vehicle at a time instead of all target vehicles. In this problem we are trying to predict 50 timesteps forward with a step size of 0.1 seconds. \n\nIn the section after this the various kinds of inputs, outputs and prediction methods are discussed. \n\n**Inputs:** \n\n\n\n*   _History of the target vehicle_ - Ego vehicles perception of the target vehicles position across previous timesteps. This can be represented in many different ways. 3d bounding boxes, x,y coordinates, etc. Good for telling general direction of travel\n*   _History of TV and SVs_ - Same as previous level but with the addition of also tracking surrounding vehicles\n*   _Simplified Bird’s Eye View_ - This seems to be what most modern systems are using. The perception system creates a top-down of the road with its understanding of where all the vehicles are and were previously\n*   _Raw Sensor Data_ - Data directly from the cameras and LiDAR and other sensors. Rather than simplifying the inputs to the planning system just directly pass them as inputs to train a planning model on. Difficult because data throughput is very large and might prohibit real time usage\n\n**Outputs:** \n\n\n\n*   _Maneuver intention_ - This is predicting the general plan rather than the exact positional future. Eg. turn right, turn left, go straight. This method is heavily limited by requiring to know all maneuvers possible. Sometimes the future of a car cannot be predefined by a simple set of maneuvers.\n*   _Unimodal trajectory_ - Rather than predicting a maneuver it is possible to predict a trajectory. This is predicting the future path of the target vehicle like we are doing in this competition. This gives much greater granularity and allows the model to predict outside of a fixed list of actions, but falls prey to collapsing to a single predicted path even when it is very likely there could be multiple reasonable future trajectories\n*   _Multimodal trajectory_ - This is an extension of the unimodal trajectory. Rather than predicting a single trajectory we can predict multiple trajectories and assign a confidence to each. Allowing us to output that it is possible a vehicle turns left here but it's more likely it continues going straight. \n*   _Occupancy map_ - This output differs from the others in that rather than trying to output continuous future values representing the trajectory this splits the world into a grid and then predicts occupancy of the grid cells. This makes the problem closer to a classification task instead of regression. This may in some way constrain and simplify the result, but can also run into trade-offs in terms of precision if the grid cells are not sufficiently small. \n\n**Prediction Method:** \n\nThere are many different combinations and permutations of these methods that are possible, but here is the general overview of the techniques that have been explored\n\n\n\n*   _Physics based models_\n    *   These are more conventional models that apply non-NN based methods, but seem to have been mostly surpassed by NN methods. These include things like simple constant velocity models that make assumptions that a car will continue travelling with the velocity it was last known to have. These can get more complex, but require much more hand-engineering and have been phased out. \n*   _Recurrent Neural Networks_\n    *   There is a heavy temporal aspect to this problem so many people have tried applying recurrent neural networks in order to capture this relationship. \n*   _Convolutional Neural Networks_\n    *   People have also attempted various different methods using CNNs. Convolutions across the time domain, and also the images generated from something like the BEV generated from the perception system. It is also possible to utilize 3D convolutions to look both across time and 2D imagery at the same time. \n*   _Other methods_\n    *   _Dense NN_ - On state variables and other condensed representations it is possible to use simple feed forward neural networks, but these methods look fairly limited\n    *   _Graph NNs_ - It is possible to repose the problem as a graph problem representing the relationship between the entities on the road and then apply graph convolutions\n    *   _Combinations of methods_ - It is possible to do convolutions across the images and then RNNs across the series of images and many different combinations of these techniques across the various different inputs\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F42cfb426c1f511d610fb69773034b95f%2FScreen%20Shot%202020-09-03%20at%203.32.50%20PM.png?generation=1599172584952012&alt=media)\n\n**Evaluation Metrics**\n\n\n\n*   _Classification (Maneuver intention)_\n    *   Accuracy\n    *   Precision\n    *   Recall\n    *   F1\n    *   Negative Log Likelihood\n    *   Average Prediction Time - The amount of time it takes for a correct prediction of the intended class to occur. Trying to not only forecast a maneuver but also check that it predicts it early on\n*   _Regression (Unimodal and Multimodal trajectory)_\n    *   Final Displacement Error - difference between final location and predicted final location\n    *   MAE\n    *   RMSE\n*   _Computation time_\n    *   Not typically reported in papers but also highly important to make a model that is able to make predictions on reasonable hardware quickly so it can be run in realtime in a vehicle.",
      "votes": null
    },
    {
      "id": "999115",
      "postDate": "09/05/2020 11:03:11",
      "content": "<p>Interesting</p>",
      "rawMarkdown": "Interesting",
      "votes": null
    },
    {
      "id": "1000050",
      "postDate": "09/06/2020 08:34:51",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> for the well-organized keypoint outline👍!</p>",
      "rawMarkdown": "Thank you @ryches for the well-organized keypoint outline👍!",
      "votes": null
    },
    {
      "id": "1001110",
      "postDate": "09/07/2020 04:33:43",
      "content": "<p>Was looking for something like this. Will check it out. Thanks for sharing</p>",
      "rawMarkdown": "Was looking for something like this. Will check it out. Thanks for sharing",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1000050,
      "author_name": "hsinwenchang",
      "author_url": "",
      "post_date": "09/06/2020 08:34:51",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> for the well-organized keypoint outline👍!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 999115,
      "author_name": "hanut2702",
      "author_url": "",
      "post_date": "09/05/2020 11:03:11",
      "content": "<p>Interesting</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1001110,
      "author_name": "blessondensil294",
      "author_url": "",
      "post_date": "09/07/2020 04:33:43",
      "content": "<p>Was looking for something like this. Will check it out. Thanks for sharing</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "997280": "I created an outline to summarize the paper (Deep Learning-based Vehicle Behaviour Prediction for Autonomous Driving Applications: a Review) that was mentioned [here](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/177113)\n\n\n\n**Deep Learning-based Vehicle Behaviour Prediction for Autonomous Driving Applications: a Review**\n\n**Background**: Most modern literature I have seen split the problem of autonomous driving into three separate parts: _perception_, _planning_, _control_. \n\n_Perception_: understanding what is going on around the vehicle that is being driven. Where are the cars, pedestrians, road lines, signs, lights, etc?  \n\n_Planning:_ using the information we have perceived from the world in order to create a plan for a safe path to drive through. Following lanes, not running into curbs, pedestrians or other vehicles while getting from point A to B as intended\n\n_Control**:**_ Given the plan created previously apply the correct acceleration and steering in order to act out the plan. \n\nIn this competition we are focusing on the planning section of the problem. Given a blend of road annotations and outputs from the lyft perception system we are trying to predict the pathing of surrounding vehicles. \n\n** **\n\n**Overview**: One of the most important problems in self-driving vehicles is forecasting the paths of surrounding vehicles. The authors of this paper review the methods that have been explored and different ways the problem has been posed. \n\nThere are a lot of acronyms used in the autonomous space. Important terms are:\n\n\n\n*   _Target Vehicles (TVs)**:**_ The vehicle that we are trying to predict the path for\n*   _Ego Vehicle (EV)**:**_ This is the car being driven. Our view of the world comes from this point of view. Perception comes from this car.\n*   _Surrounding Vehicles (SVs)**:**_ The vehicles that are around the target vehicle and might affect its path\n*   _Non Effective Vehicles (NVs)**:**_ surrounding vehicles that have no bearing on the target vehicles behavior. Can be things like parked vehicles, vehicles several lanes away, etc.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F57bc6227fa3827ea70c1571ee904cd47%2FScreen%20Shot%202020-09-03%20at%203.32.40%20PM.png?generation=1599172539573875&alt=media)\n\nIn the next section of the paper they define the problem formulation. The simplification from math equation to english is we are trying to predict future time steps from all target vehicles given what we can observe from the ego vehicle. This problem is very difficult to predict for all of these interacting vehicles so the problem can be simplified to a single target vehicle at a time instead of all target vehicles. In this problem we are trying to predict 50 timesteps forward with a step size of 0.1 seconds. \n\nIn the section after this the various kinds of inputs, outputs and prediction methods are discussed. \n\n**Inputs:** \n\n\n\n*   _History of the target vehicle_ - Ego vehicles perception of the target vehicles position across previous timesteps. This can be represented in many different ways. 3d bounding boxes, x,y coordinates, etc. Good for telling general direction of travel\n*   _History of TV and SVs_ - Same as previous level but with the addition of also tracking surrounding vehicles\n*   _Simplified Bird’s Eye View_ - This seems to be what most modern systems are using. The perception system creates a top-down of the road with its understanding of where all the vehicles are and were previously\n*   _Raw Sensor Data_ - Data directly from the cameras and LiDAR and other sensors. Rather than simplifying the inputs to the planning system just directly pass them as inputs to train a planning model on. Difficult because data throughput is very large and might prohibit real time usage\n\n**Outputs:** \n\n\n\n*   _Maneuver intention_ - This is predicting the general plan rather than the exact positional future. Eg. turn right, turn left, go straight. This method is heavily limited by requiring to know all maneuvers possible. Sometimes the future of a car cannot be predefined by a simple set of maneuvers.\n*   _Unimodal trajectory_ - Rather than predicting a maneuver it is possible to predict a trajectory. This is predicting the future path of the target vehicle like we are doing in this competition. This gives much greater granularity and allows the model to predict outside of a fixed list of actions, but falls prey to collapsing to a single predicted path even when it is very likely there could be multiple reasonable future trajectories\n*   _Multimodal trajectory_ - This is an extension of the unimodal trajectory. Rather than predicting a single trajectory we can predict multiple trajectories and assign a confidence to each. Allowing us to output that it is possible a vehicle turns left here but it's more likely it continues going straight. \n*   _Occupancy map_ - This output differs from the others in that rather than trying to output continuous future values representing the trajectory this splits the world into a grid and then predicts occupancy of the grid cells. This makes the problem closer to a classification task instead of regression. This may in some way constrain and simplify the result, but can also run into trade-offs in terms of precision if the grid cells are not sufficiently small. \n\n**Prediction Method:** \n\nThere are many different combinations and permutations of these methods that are possible, but here is the general overview of the techniques that have been explored\n\n\n\n*   _Physics based models_\n    *   These are more conventional models that apply non-NN based methods, but seem to have been mostly surpassed by NN methods. These include things like simple constant velocity models that make assumptions that a car will continue travelling with the velocity it was last known to have. These can get more complex, but require much more hand-engineering and have been phased out. \n*   _Recurrent Neural Networks_\n    *   There is a heavy temporal aspect to this problem so many people have tried applying recurrent neural networks in order to capture this relationship. \n*   _Convolutional Neural Networks_\n    *   People have also attempted various different methods using CNNs. Convolutions across the time domain, and also the images generated from something like the BEV generated from the perception system. It is also possible to utilize 3D convolutions to look both across time and 2D imagery at the same time. \n*   _Other methods_\n    *   _Dense NN_ - On state variables and other condensed representations it is possible to use simple feed forward neural networks, but these methods look fairly limited\n    *   _Graph NNs_ - It is possible to repose the problem as a graph problem representing the relationship between the entities on the road and then apply graph convolutions\n    *   _Combinations of methods_ - It is possible to do convolutions across the images and then RNNs across the series of images and many different combinations of these techniques across the various different inputs\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F42cfb426c1f511d610fb69773034b95f%2FScreen%20Shot%202020-09-03%20at%203.32.50%20PM.png?generation=1599172584952012&alt=media)\n\n**Evaluation Metrics**\n\n\n\n*   _Classification (Maneuver intention)_\n    *   Accuracy\n    *   Precision\n    *   Recall\n    *   F1\n    *   Negative Log Likelihood\n    *   Average Prediction Time - The amount of time it takes for a correct prediction of the intended class to occur. Trying to not only forecast a maneuver but also check that it predicts it early on\n*   _Regression (Unimodal and Multimodal trajectory)_\n    *   Final Displacement Error - difference between final location and predicted final location\n    *   MAE\n    *   RMSE\n*   _Computation time_\n    *   Not typically reported in papers but also highly important to make a model that is able to make predictions on reasonable hardware quickly so it can be run in realtime in a vehicle.",
    "999115": "Interesting",
    "1000050": "Thank you @ryches for the well-organized keypoint outline👍!",
    "1001110": "Was looking for something like this. Will check it out. Thanks for sharing"
  },
  "source": "meta"
}