{
  "id": 179246,
  "title": "Dataset Tutorial Recording and Q&A Resources",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/179246",
  "author_name": "",
  "post_date": "2020-09-02T00:42:42.207483900Z",
  "votes": 41,
  "comment_count": 2,
  "views": 0,
  "content": "<p>A few participants asked if the dataset could be used for academic purposes.</p>\n<p>The answer is <em>yes</em>. You can, and if you will, it will make us happy.</p>\n<p>A month ago we had a <a href=\"https://vimeo.com/451293003\" target=\"_blank\">webinar</a> to help folks start building models on the dataset.</p>\n<p>We received a lot of great questions during Q&amp;A.  We did not have time to address them.</p>\n<p>I post them here as a resource for you.</p>\n<h2>Competition</h2>\n<h3>Q: Why are you focusing on prediction? Are “easier” problems like perception and planning already solved?</h3>\n<p>Perception uses segmentation and detection algorithms on the lidar, radar, and imagery data. There are many challenges there, but overall these algorithms have a lot of attention from the research community. The task for last year’s competition was: 3D object detection on the lidar and imagery data and was perception. </p>\n<p>This time we decided to explore the next part of the autonomous stack and work on prediction. The prediction problem is new and Kaggle hasn’t hosted the challenge of this type. We believe that winners will not be teams with the most hardware, but people who are the most creative. We hope that this competition will attract attention to the prediction problem, generate new ideas, and be a lot of fun for participants. </p>\n<h3>Q: I participated in the Kaggle competition hosted by Lyft last year. Does this competition have any relationship with the previous one?</h3>\n<p>The relationship is not strong. The dataset, the task, metric, the required skill set, and domain knowledge are different. </p>\n<h3>Q: Do we have access to some promotional cloud credits for people starting out in competitions?</h3>\n<p>Yes, we will be offering $300 in GCP credits to a set of participants that will beat the sample submission by organizers. </p>\n<h3>Q: Do we have to focus on predicting the action of cars? Or can we focus on the actions of humans and other agents?</h3>\n<p>Our dataset is not 100% focused on cars. It also includes pedestrians and cyclists, both in train and evaluation. However, the vast majority of agents available are vehicles. </p>\n<h3>Q: Do we need to predict the motion for multiple agents simultaneously in a given frame, or is only the ego vehicle considered?</h3>\n<p>This competition is all about other agents, so you will have to predict multiple agents in a given frame. </p>\n<h3>Q: Are we restricted to the 2D rasterization of the environment to make motion prediction decisions?</h3>\n<p>No, you are not. You have the dataset; you have a target metric. Feel free to use any method that will make your submission more robust. </p>\n<h3>Q: Are we aware of the goal/destination of the car?</h3>\n<p>Because you will be predicting other traffic agents, we are not fully aware of the destination goal of the car. However, nothing prevents you from collecting future observations for a specific agent to build the “goal” of that agent. </p>\n<h3>Q: Can the motion prediction problem be designed as a graph problem?</h3>\n<p>It is possible. If it helps with the challenge, go for it! </p>\n<h3>Q: Do you believe that the solution is in deep learning architecture? Are there any recommendations on which deep learning models will be suitable for solving this challenge?</h3>\n<p>We suspect that the solution will be based on deep learning architecture. But we will not be surprised if other methods might be better. The Kaggle community is very creative and we are very curious to see what techniques will perform well. </p>\n<h3>Q: Does the neural net have to be RESNET-50, or can we try different architectures?</h3>\n<p>No, there is no restriction for the network architecture or method that you can use in your solution. </p>\n<h3>Q: Will the competition support models that generate multiple hypotheses per agent? Or will it just support one hypothesis as seen in the example?</h3>\n<p>Yes, the submitted solutions will be scored using a multi-modal metric. However, we will put a limit on the number of concurrent hypotheses. </p>\n<h3>Q: Will you all be releasing the pre-trained baselines?</h3>\n<p>We will release a sample baseline based on the code in the <a href=\"https://github.com/lyft/l5kit/tree/master/examples/agent_motion_prediction\" target=\"_blank\">l5kit examples</a>.</p>\n<h2>Dataset</h2>\n<h3>Q: Where can I access the dataset?</h3>\n<p>You can download the dataset <a href=\"https://self-driving.lyft.com/level5/prediction/\" target=\"_blank\">here</a>. You can also view our <a href=\"https://arxiv.org/abs/2006.14480\" target=\"_blank\">paper</a>, and access the <a href=\"https://github.com/lyft/l5kit\" target=\"_blank\">L5kit</a> to help you get started.</p>\n<h3>Q: What's the main difference between the raster shown in the baseline approach and one on the visualization slide (shown @ 32:30 in the <a href=\"https://vimeo.com/451293003\" target=\"_blank\">recording</a>)?</h3>\n<p>The baseline approach employs an open-sourced version of our internal rasterizer. We’re working to close the gap between the two. Right now, the only real difference is the lane color.</p>\n<h3>Q: Can you share more about the input raster? Is the input prepared from the table entries (shown @ 16:38 in the <a href=\"https://vimeo.com/451293003\" target=\"_blank\">recording</a>)?</h3>\n<p>When we create a raster what we do is fetching an agent in a given frame in a given scene. As such, all three tables need to be queried to retrieve the final result. Moreover, because we also want to rasterize other agents from the same scene, we perform multiple queries, especially on that table. The code responsible for this part is already open-sourced and available in our L5Kit toolkit. </p>\n<h3>Q: What's the model used for training?</h3>\n<p>What's the model used for training?</p>\n<h3>Q: How do you choose which scenes to train your model?</h3>\n<p>For our deep learning models, we pick all of them. The more data the better. </p>\n<h3>Q: Object perception and tracking always have uncertainties in the object state that it outputs. Are the values for these uncertainties included in the dataset?</h3>\n<p>Yes, each detected agent has a probability assigned from perception output. Those values are stored in the dataset. </p>\n<h3>Q: Is there an approach to get information about the turn signals of the surrounding vehicles?</h3>\n<p>Currently, perception estimates the turn signal state and provides the information to the prediction of the vehicle. </p>\n<h3>Q: Does the perception stack give us the probabilities associated with each prediction?</h3>\n<p>Yes, each detected agent has a probability assigned from perception output. Those values are stored in the dataset. </p>\n<h3>Q: What are the time constraints for planning? (I.e. do you target 30Hz or more/less?)</h3>\n<p>Prediction on an AV typically targets 10 hertz with a latency of 100 milliseconds. See the Kaggle website for details regarding the competition resource and timing constraints. </p>\n<h3>Q: How many frames into the future should the model predict? What is the ideal time frame in general based on the reaction time of the ego vehicle?</h3>\n<p>For the challenge, we specifically target 50 frames in the future at 10hz. We wanted a prediction window large enough to stress the proposed solutions but small enough to fit in the input raster. </p>\n<h3>Q: What is the ratio of situations which are easy to predict (driving straight) and hard situations (collisions, jaywalking, unexpected turns)?</h3>\n<p>The dataset situations are skewed towards driving straight for the AV. However, when we consider all the agents surrounding it (the real focus of the competition), it’s difficult to draw an exact unsupervised ratio computation. </p>\n<h3>Q: Which version of Python does this use?</h3>\n<p>Python 3 (3.6 onwards) </p>\n<h3>Q: How was the ground truth data created?</h3>\n<p>The ground truth data was created using the output of our perception stack. Semantic information for the map elements has been manually annotated. </p>\n<h3>Q: I would like to reproduce the performance metrics in the baseline paper. The rasterized map has a different color for each direction of the street, but the current code on the repo does not have this detail. Can you explain?</h3>\n<p>That picture was obtained using an internal proprietary version of the semantic rasterizer. We’re working to include many of those features in our open-source version. </p>\n<h3>Q: Does the dataset provide a label for which trajectories are considered “good” or “bad” driving to make sure models learn good behavior?</h3>\n<p>No, the dataset doesn’t include any behavioural notions. </p>\n<h3>Q: Do all the agents in each frame in the dataset have trajectories?</h3>\n<p>No, agents come and go and their detection and tracking may not be perfect. We provide ways of filtering out completely unreliable agents in our toolkit. </p>\n<h3>Q: When the ego vehicle was crossing an intersection in one of the examples of the rasterized overhead view, a couple of objects morphed into an object closer to the size of a bus when they entered the intersection, and then morphed back into 2 vehicles when it crossed the intersection. Are these kinds of artifacts expected in the ground truth?</h3>\n<p>Independent of the quality of our perception stack, there will always be artifacts in the data. These situations are absolutely possible, so we and you should expect them to happen in the ground truth and learn to account for them. However, we provide tools for filtering out some of those artifacts from rasterization centers (i.e. you won’t rasterize around those faulty agents, but they will still be visible around you). This limits the noise presence in the ground truth trajectories. </p>\n<h3>Q: I am interested in using this dataset for imitation learning. In a planning setting, you normally have a route to follow. Is it possible to reliably extract this route for every scene from the lane graph protobuf provided with the dataset?</h3>\n<p>For the AV, that’s definitely possible, as the information is accurate. You can extract the translations directly from the zarr and build a route. For agents, this is still an option, but you will encounter much more variety in terms of data quality. Once you have the routes coordinates, it is easier to query the protobuf API to get lane information. </p>\n<h3>Q: What have been the limitations and failure cases in real-life scenarios and in the lab?</h3>\n<p>There are many limitations depending on the algorithm in question. For example, a pure velocity rollout may overshoot a vehicle slowing down, and undershoot on a vehicle speeding up. Incorporating the acceleration helps in many scenarios, but we still have encountered some problems — for example, a car coming to a stop ended up rolling backward with constant acceleration models. So that is an example we corrected.</p>\n<p>Basically the heuristics enter into a cat and mouse game of ever-increasing complexity to resolve edge conditions. It turns out that these models reach such specificity they can become either unintelligible or fail to generalize well enough to be useful. Particularly challenging situations revolve around occlusions, unprotected turns, parked/not parked states, stop signs, and any cases where the AV’s behavior strongly influences the agent’s behavior. </p>\n<h3>Q: Are there any high-level annotations such as pedestrian actions?</h3>\n<p>To keep things simple, no. For this competition, you will be working on the raw output of our perception stack. </p>\n<h3>Q: In frame data, we see no ego velocity/acceleration. Any reason?</h3>\n<p>We’ve been focused on predicting other agents so this hasn’t made it into the dataset yet. It should be easy to add this in for the future because these values are readily available. </p>\n<h2>SELF-DRIVING CAREERS AND RESOURCES</h2>\n<h3>Q: What internship and career opportunities are available at Lyft Level 5?</h3>\n<p>You can see our open positions <a href=\"https://www.lyft.com/careers\" target=\"_blank\">here</a>. </p>\n<h3>Q: What level of degree do I need to join the team?</h3>\n<p>Lyft does not have a degree requirement for joining the team. </p>\n<h3>Q: Can I work as an intern remotely?</h3>\n<p>Currently, Lyft team members who can conduct their work remotely are able to do so given shelter in place. However, interns normally are required to be in-office. </p>\n<h3>Q: As a student who's trying to learn more about AVs and a career in the field, how did you get your start? Any tips for a fellow student?</h3>\n<p>Check out our interview with <a href=\"https://medium.com/lyftlevel5/employee-spotlight-vladimir-iglovikov-11b71ed02bc\" target=\"_blank\">Vladimir Iglovikov</a> for his insights on this topic! </p>",
  "messages": [
    {
      "id": "994800",
      "postDate": "09/02/2020 00:42:42",
      "content": "<p>A few participants asked if the dataset could be used for academic purposes.</p>\n<p>The answer is <em>yes</em>. You can, and if you will, it will make us happy.</p>\n<p>A month ago we had a <a href=\"https://vimeo.com/451293003\" target=\"_blank\">webinar</a> to help folks start building models on the dataset.</p>\n<p>We received a lot of great questions during Q&amp;A.  We did not have time to address them.</p>\n<p>I post them here as a resource for you.</p>\n<h2>Competition</h2>\n<h3>Q: Why are you focusing on prediction? Are “easier” problems like perception and planning already solved?</h3>\n<p>Perception uses segmentation and detection algorithms on the lidar, radar, and imagery data. There are many challenges there, but overall these algorithms have a lot of attention from the research community. The task for last year’s competition was: 3D object detection on the lidar and imagery data and was perception. </p>\n<p>This time we decided to explore the next part of the autonomous stack and work on prediction. The prediction problem is new and Kaggle hasn’t hosted the challenge of this type. We believe that winners will not be teams with the most hardware, but people who are the most creative. We hope that this competition will attract attention to the prediction problem, generate new ideas, and be a lot of fun for participants. </p>\n<h3>Q: I participated in the Kaggle competition hosted by Lyft last year. Does this competition have any relationship with the previous one?</h3>\n<p>The relationship is not strong. The dataset, the task, metric, the required skill set, and domain knowledge are different. </p>\n<h3>Q: Do we have access to some promotional cloud credits for people starting out in competitions?</h3>\n<p>Yes, we will be offering $300 in GCP credits to a set of participants that will beat the sample submission by organizers. </p>\n<h3>Q: Do we have to focus on predicting the action of cars? Or can we focus on the actions of humans and other agents?</h3>\n<p>Our dataset is not 100% focused on cars. It also includes pedestrians and cyclists, both in train and evaluation. However, the vast majority of agents available are vehicles. </p>\n<h3>Q: Do we need to predict the motion for multiple agents simultaneously in a given frame, or is only the ego vehicle considered?</h3>\n<p>This competition is all about other agents, so you will have to predict multiple agents in a given frame. </p>\n<h3>Q: Are we restricted to the 2D rasterization of the environment to make motion prediction decisions?</h3>\n<p>No, you are not. You have the dataset; you have a target metric. Feel free to use any method that will make your submission more robust. </p>\n<h3>Q: Are we aware of the goal/destination of the car?</h3>\n<p>Because you will be predicting other traffic agents, we are not fully aware of the destination goal of the car. However, nothing prevents you from collecting future observations for a specific agent to build the “goal” of that agent. </p>\n<h3>Q: Can the motion prediction problem be designed as a graph problem?</h3>\n<p>It is possible. If it helps with the challenge, go for it! </p>\n<h3>Q: Do you believe that the solution is in deep learning architecture? Are there any recommendations on which deep learning models will be suitable for solving this challenge?</h3>\n<p>We suspect that the solution will be based on deep learning architecture. But we will not be surprised if other methods might be better. The Kaggle community is very creative and we are very curious to see what techniques will perform well. </p>\n<h3>Q: Does the neural net have to be RESNET-50, or can we try different architectures?</h3>\n<p>No, there is no restriction for the network architecture or method that you can use in your solution. </p>\n<h3>Q: Will the competition support models that generate multiple hypotheses per agent? Or will it just support one hypothesis as seen in the example?</h3>\n<p>Yes, the submitted solutions will be scored using a multi-modal metric. However, we will put a limit on the number of concurrent hypotheses. </p>\n<h3>Q: Will you all be releasing the pre-trained baselines?</h3>\n<p>We will release a sample baseline based on the code in the <a href=\"https://github.com/lyft/l5kit/tree/master/examples/agent_motion_prediction\" target=\"_blank\">l5kit examples</a>.</p>\n<h2>Dataset</h2>\n<h3>Q: Where can I access the dataset?</h3>\n<p>You can download the dataset <a href=\"https://self-driving.lyft.com/level5/prediction/\" target=\"_blank\">here</a>. You can also view our <a href=\"https://arxiv.org/abs/2006.14480\" target=\"_blank\">paper</a>, and access the <a href=\"https://github.com/lyft/l5kit\" target=\"_blank\">L5kit</a> to help you get started.</p>\n<h3>Q: What's the main difference between the raster shown in the baseline approach and one on the visualization slide (shown @ 32:30 in the <a href=\"https://vimeo.com/451293003\" target=\"_blank\">recording</a>)?</h3>\n<p>The baseline approach employs an open-sourced version of our internal rasterizer. We’re working to close the gap between the two. Right now, the only real difference is the lane color.</p>\n<h3>Q: Can you share more about the input raster? Is the input prepared from the table entries (shown @ 16:38 in the <a href=\"https://vimeo.com/451293003\" target=\"_blank\">recording</a>)?</h3>\n<p>When we create a raster what we do is fetching an agent in a given frame in a given scene. As such, all three tables need to be queried to retrieve the final result. Moreover, because we also want to rasterize other agents from the same scene, we perform multiple queries, especially on that table. The code responsible for this part is already open-sourced and available in our L5Kit toolkit. </p>\n<h3>Q: What's the model used for training?</h3>\n<p>What's the model used for training?</p>\n<h3>Q: How do you choose which scenes to train your model?</h3>\n<p>For our deep learning models, we pick all of them. The more data the better. </p>\n<h3>Q: Object perception and tracking always have uncertainties in the object state that it outputs. Are the values for these uncertainties included in the dataset?</h3>\n<p>Yes, each detected agent has a probability assigned from perception output. Those values are stored in the dataset. </p>\n<h3>Q: Is there an approach to get information about the turn signals of the surrounding vehicles?</h3>\n<p>Currently, perception estimates the turn signal state and provides the information to the prediction of the vehicle. </p>\n<h3>Q: Does the perception stack give us the probabilities associated with each prediction?</h3>\n<p>Yes, each detected agent has a probability assigned from perception output. Those values are stored in the dataset. </p>\n<h3>Q: What are the time constraints for planning? (I.e. do you target 30Hz or more/less?)</h3>\n<p>Prediction on an AV typically targets 10 hertz with a latency of 100 milliseconds. See the Kaggle website for details regarding the competition resource and timing constraints. </p>\n<h3>Q: How many frames into the future should the model predict? What is the ideal time frame in general based on the reaction time of the ego vehicle?</h3>\n<p>For the challenge, we specifically target 50 frames in the future at 10hz. We wanted a prediction window large enough to stress the proposed solutions but small enough to fit in the input raster. </p>\n<h3>Q: What is the ratio of situations which are easy to predict (driving straight) and hard situations (collisions, jaywalking, unexpected turns)?</h3>\n<p>The dataset situations are skewed towards driving straight for the AV. However, when we consider all the agents surrounding it (the real focus of the competition), it’s difficult to draw an exact unsupervised ratio computation. </p>\n<h3>Q: Which version of Python does this use?</h3>\n<p>Python 3 (3.6 onwards) </p>\n<h3>Q: How was the ground truth data created?</h3>\n<p>The ground truth data was created using the output of our perception stack. Semantic information for the map elements has been manually annotated. </p>\n<h3>Q: I would like to reproduce the performance metrics in the baseline paper. The rasterized map has a different color for each direction of the street, but the current code on the repo does not have this detail. Can you explain?</h3>\n<p>That picture was obtained using an internal proprietary version of the semantic rasterizer. We’re working to include many of those features in our open-source version. </p>\n<h3>Q: Does the dataset provide a label for which trajectories are considered “good” or “bad” driving to make sure models learn good behavior?</h3>\n<p>No, the dataset doesn’t include any behavioural notions. </p>\n<h3>Q: Do all the agents in each frame in the dataset have trajectories?</h3>\n<p>No, agents come and go and their detection and tracking may not be perfect. We provide ways of filtering out completely unreliable agents in our toolkit. </p>\n<h3>Q: When the ego vehicle was crossing an intersection in one of the examples of the rasterized overhead view, a couple of objects morphed into an object closer to the size of a bus when they entered the intersection, and then morphed back into 2 vehicles when it crossed the intersection. Are these kinds of artifacts expected in the ground truth?</h3>\n<p>Independent of the quality of our perception stack, there will always be artifacts in the data. These situations are absolutely possible, so we and you should expect them to happen in the ground truth and learn to account for them. However, we provide tools for filtering out some of those artifacts from rasterization centers (i.e. you won’t rasterize around those faulty agents, but they will still be visible around you). This limits the noise presence in the ground truth trajectories. </p>\n<h3>Q: I am interested in using this dataset for imitation learning. In a planning setting, you normally have a route to follow. Is it possible to reliably extract this route for every scene from the lane graph protobuf provided with the dataset?</h3>\n<p>For the AV, that’s definitely possible, as the information is accurate. You can extract the translations directly from the zarr and build a route. For agents, this is still an option, but you will encounter much more variety in terms of data quality. Once you have the routes coordinates, it is easier to query the protobuf API to get lane information. </p>\n<h3>Q: What have been the limitations and failure cases in real-life scenarios and in the lab?</h3>\n<p>There are many limitations depending on the algorithm in question. For example, a pure velocity rollout may overshoot a vehicle slowing down, and undershoot on a vehicle speeding up. Incorporating the acceleration helps in many scenarios, but we still have encountered some problems — for example, a car coming to a stop ended up rolling backward with constant acceleration models. So that is an example we corrected.</p>\n<p>Basically the heuristics enter into a cat and mouse game of ever-increasing complexity to resolve edge conditions. It turns out that these models reach such specificity they can become either unintelligible or fail to generalize well enough to be useful. Particularly challenging situations revolve around occlusions, unprotected turns, parked/not parked states, stop signs, and any cases where the AV’s behavior strongly influences the agent’s behavior. </p>\n<h3>Q: Are there any high-level annotations such as pedestrian actions?</h3>\n<p>To keep things simple, no. For this competition, you will be working on the raw output of our perception stack. </p>\n<h3>Q: In frame data, we see no ego velocity/acceleration. Any reason?</h3>\n<p>We’ve been focused on predicting other agents so this hasn’t made it into the dataset yet. It should be easy to add this in for the future because these values are readily available. </p>\n<h2>SELF-DRIVING CAREERS AND RESOURCES</h2>\n<h3>Q: What internship and career opportunities are available at Lyft Level 5?</h3>\n<p>You can see our open positions <a href=\"https://www.lyft.com/careers\" target=\"_blank\">here</a>. </p>\n<h3>Q: What level of degree do I need to join the team?</h3>\n<p>Lyft does not have a degree requirement for joining the team. </p>\n<h3>Q: Can I work as an intern remotely?</h3>\n<p>Currently, Lyft team members who can conduct their work remotely are able to do so given shelter in place. However, interns normally are required to be in-office. </p>\n<h3>Q: As a student who's trying to learn more about AVs and a career in the field, how did you get your start? Any tips for a fellow student?</h3>\n<p>Check out our interview with <a href=\"https://medium.com/lyftlevel5/employee-spotlight-vladimir-iglovikov-11b71ed02bc\" target=\"_blank\">Vladimir Iglovikov</a> for his insights on this topic! </p>",
      "rawMarkdown": "A few participants asked if the dataset could be used for academic purposes.\n\nThe answer is *yes*. You can, and if you will, it will make us happy.\n\nA month ago we had a [webinar](https://vimeo.com/451293003) to help folks start building models on the dataset.\n\nWe received a lot of great questions during Q&A.  We did not have time to address them.\n\nI post them here as a resource for you.\n\n## Competition\n\n### Q: Why are you focusing on prediction? Are “easier” problems like perception and planning already solved?\n\nPerception uses segmentation and detection algorithms on the lidar, radar, and imagery data. There are many challenges there, but overall these algorithms have a lot of attention from the research community. The task for last year’s competition was: 3D object detection on the lidar and imagery data and was perception. \n\nThis time we decided to explore the next part of the autonomous stack and work on prediction. The prediction problem is new and Kaggle hasn’t hosted the challenge of this type. We believe that winners will not be teams with the most hardware, but people who are the most creative. We hope that this competition will attract attention to the prediction problem, generate new ideas, and be a lot of fun for participants. \n\n### Q: I participated in the Kaggle competition hosted by Lyft last year. Does this competition have any relationship with the previous one? \n\nThe relationship is not strong. The dataset, the task, metric, the required skill set, and domain knowledge are different. \n\n### Q: Do we have access to some promotional cloud credits for people starting out in competitions? \n\nYes, we will be offering $300 in GCP credits to a set of participants that will beat the sample submission by organizers. \n\n### Q: Do we have to focus on predicting the action of cars? Or can we focus on the actions of humans and other agents?\n\nOur dataset is not 100% focused on cars. It also includes pedestrians and cyclists, both in train and evaluation. However, the vast majority of agents available are vehicles. \n\n### Q: Do we need to predict the motion for multiple agents simultaneously in a given frame, or is only the ego vehicle considered? \n\nThis competition is all about other agents, so you will have to predict multiple agents in a given frame. \n\n### Q: Are we restricted to the 2D rasterization of the environment to make motion prediction decisions?\n\nNo, you are not. You have the dataset; you have a target metric. Feel free to use any method that will make your submission more robust. \n\n### Q: Are we aware of the goal/destination of the car?\n\nBecause you will be predicting other traffic agents, we are not fully aware of the destination goal of the car. However, nothing prevents you from collecting future observations for a specific agent to build the “goal” of that agent. \n\n### Q: Can the motion prediction problem be designed as a graph problem?\n\nIt is possible. If it helps with the challenge, go for it! \n\n### Q: Do you believe that the solution is in deep learning architecture? Are there any recommendations on which deep learning models will be suitable for solving this challenge?\n\nWe suspect that the solution will be based on deep learning architecture. But we will not be surprised if other methods might be better. The Kaggle community is very creative and we are very curious to see what techniques will perform well. \n\n### Q: Does the neural net have to be RESNET-50, or can we try different architectures? \n\nNo, there is no restriction for the network architecture or method that you can use in your solution. \n\n### Q: Will the competition support models that generate multiple hypotheses per agent? Or will it just support one hypothesis as seen in the example? \n\nYes, the submitted solutions will be scored using a multi-modal metric. However, we will put a limit on the number of concurrent hypotheses. \n\n### Q: Will you all be releasing the pre-trained baselines?\n\nWe will release a sample baseline based on the code in the [l5kit examples](https://github.com/lyft/l5kit/tree/master/examples/agent_motion_prediction).\n\n## Dataset\n\n### Q: Where can I access the dataset?\n\nYou can download the dataset [here](https://self-driving.lyft.com/level5/prediction/). You can also view our [paper](https://arxiv.org/abs/2006.14480), and access the [L5kit](https://github.com/lyft/l5kit) to help you get started.\n\n### Q: What's the main difference between the raster shown in the baseline approach and one on the visualization slide (shown @ 32:30 in the [recording](https://vimeo.com/451293003))?\n\nThe baseline approach employs an open-sourced version of our internal rasterizer. We’re working to close the gap between the two. Right now, the only real difference is the lane color.\n\n### Q: Can you share more about the input raster? Is the input prepared from the table entries (shown @ 16:38 in the [recording](https://vimeo.com/451293003))?\n\nWhen we create a raster what we do is fetching an agent in a given frame in a given scene. As such, all three tables need to be queried to retrieve the final result. Moreover, because we also want to rasterize other agents from the same scene, we perform multiple queries, especially on that table. The code responsible for this part is already open-sourced and available in our L5Kit toolkit. \n\n### Q: What's the model used for training?\n\nWhat's the model used for training?\n\n### Q: How do you choose which scenes to train your model? \n\nFor our deep learning models, we pick all of them. The more data the better. \n\n### Q: Object perception and tracking always have uncertainties in the object state that it outputs. Are the values for these uncertainties included in the dataset?\n\nYes, each detected agent has a probability assigned from perception output. Those values are stored in the dataset. \n\n\n### Q: Is there an approach to get information about the turn signals of the surrounding vehicles?\n\nCurrently, perception estimates the turn signal state and provides the information to the prediction of the vehicle. \n\n### Q: Does the perception stack give us the probabilities associated with each prediction?\n\nYes, each detected agent has a probability assigned from perception output. Those values are stored in the dataset. \n\n### Q: What are the time constraints for planning? (I.e. do you target 30Hz or more/less?) \n\nPrediction on an AV typically targets 10 hertz with a latency of 100 milliseconds. See the Kaggle website for details regarding the competition resource and timing constraints. \n\n### Q: How many frames into the future should the model predict? What is the ideal time frame in general based on the reaction time of the ego vehicle?\n\nFor the challenge, we specifically target 50 frames in the future at 10hz. We wanted a prediction window large enough to stress the proposed solutions but small enough to fit in the input raster. \n\n### Q: What is the ratio of situations which are easy to predict (driving straight) and hard situations (collisions, jaywalking, unexpected turns)?\n\nThe dataset situations are skewed towards driving straight for the AV. However, when we consider all the agents surrounding it (the real focus of the competition), it’s difficult to draw an exact unsupervised ratio computation. \n\n### Q: Which version of Python does this use?\n\nPython 3 (3.6 onwards) \n\n### Q: How was the ground truth data created?\n\nThe ground truth data was created using the output of our perception stack. Semantic information for the map elements has been manually annotated. \n\n\n### Q: I would like to reproduce the performance metrics in the baseline paper. The rasterized map has a different color for each direction of the street, but the current code on the repo does not have this detail. Can you explain?\n\nThat picture was obtained using an internal proprietary version of the semantic rasterizer. We’re working to include many of those features in our open-source version. \n\n### Q: Does the dataset provide a label for which trajectories are considered “good” or “bad” driving to make sure models learn good behavior?\n\nNo, the dataset doesn’t include any behavioural notions. \n\n### Q: Do all the agents in each frame in the dataset have trajectories? \n\nNo, agents come and go and their detection and tracking may not be perfect. We provide ways of filtering out completely unreliable agents in our toolkit. \n\n### Q: When the ego vehicle was crossing an intersection in one of the examples of the rasterized overhead view, a couple of objects morphed into an object closer to the size of a bus when they entered the intersection, and then morphed back into 2 vehicles when it crossed the intersection. Are these kinds of artifacts expected in the ground truth?\n\nIndependent of the quality of our perception stack, there will always be artifacts in the data. These situations are absolutely possible, so we and you should expect them to happen in the ground truth and learn to account for them. However, we provide tools for filtering out some of those artifacts from rasterization centers (i.e. you won’t rasterize around those faulty agents, but they will still be visible around you). This limits the noise presence in the ground truth trajectories. \n\n### Q: I am interested in using this dataset for imitation learning. In a planning setting, you normally have a route to follow. Is it possible to reliably extract this route for every scene from the lane graph protobuf provided with the dataset?\n\nFor the AV, that’s definitely possible, as the information is accurate. You can extract the translations directly from the zarr and build a route. For agents, this is still an option, but you will encounter much more variety in terms of data quality. Once you have the routes coordinates, it is easier to query the protobuf API to get lane information. \n\n### Q: What have been the limitations and failure cases in real-life scenarios and in the lab? \n\nThere are many limitations depending on the algorithm in question. For example, a pure velocity rollout may overshoot a vehicle slowing down, and undershoot on a vehicle speeding up. Incorporating the acceleration helps in many scenarios, but we still have encountered some problems — for example, a car coming to a stop ended up rolling backward with constant acceleration models. So that is an example we corrected.\n\nBasically the heuristics enter into a cat and mouse game of ever-increasing complexity to resolve edge conditions. It turns out that these models reach such specificity they can become either unintelligible or fail to generalize well enough to be useful. Particularly challenging situations revolve around occlusions, unprotected turns, parked/not parked states, stop signs, and any cases where the AV’s behavior strongly influences the agent’s behavior. \n\n### Q: Are there any high-level annotations such as pedestrian actions?\n\nTo keep things simple, no. For this competition, you will be working on the raw output of our perception stack. \n\n### Q: In frame data, we see no ego velocity/acceleration. Any reason?\n\nWe’ve been focused on predicting other agents so this hasn’t made it into the dataset yet. It should be easy to add this in for the future because these values are readily available. \n\n## SELF-DRIVING CAREERS AND RESOURCES \n\n### Q: What internship and career opportunities are available at Lyft Level 5?\n\nYou can see our open positions [here](https://www.lyft.com/careers). \n\n### Q: What level of degree do I need to join the team?\n\nLyft does not have a degree requirement for joining the team. \n\n### Q: Can I work as an intern remotely?\n\nCurrently, Lyft team members who can conduct their work remotely are able to do so given shelter in place. However, interns normally are required to be in-office. \n\n### Q: As a student who's trying to learn more about AVs and a career in the field, how did you get your start? Any tips for a fellow student?\n\nCheck out our interview with [Vladimir Iglovikov](https://medium.com/lyftlevel5/employee-spotlight-vladimir-iglovikov-11b71ed02bc) for his insights on this topic!",
      "votes": null
    },
    {
      "id": "995730",
      "postDate": "09/02/2020 18:16:45",
      "content": "<p>For sure the winning solution will be a neural net, given the non-classicality of the loss function. But, odds are that the winning solution won't be based on computer vision models.</p>\n<p>This competition is, according to my understanding, a mix between computer vision and times series. A solution that find the right way to combine models from both of those doamains, won't only be innovative but could even score among top ranked solutions in the end of this interesting competition.</p>",
      "rawMarkdown": "For sure the winning solution will be a neural net, given the non-classicality of the loss function. But, odds are that the winning solution won't be based on computer vision models.\n\nThis competition is, according to my understanding, a mix between computer vision and times series. A solution that find the right way to combine models from both of those doamains, won't only be innovative but could even score among top ranked solutions in the end of this interesting competition.",
      "votes": null
    },
    {
      "id": "1075333",
      "postDate": "11/11/2020 15:44:26",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/iglovikov\" target=\"_blank\">@iglovikov</a>. I am playing around with the source code of l5kit trying to load data that is available but not used yet, like the velocity of agents. But I confront some confusing problems. When I am trying to rewrite the agent_sampling.py, I found that the velocity has two unfathomable value like this:<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3722352%2F63cf70ed25bd56aeab6aa325555455e7%2F1491605109394_.pic.jpg?generation=1605109414580993&amp;alt=media\" alt=\"\"></p>\n<p>Could you please give me some instructions to help me figure it out?</p>",
      "rawMarkdown": "Hi @iglovikov. I am playing around with the source code of l5kit trying to load data that is available but not used yet, like the velocity of agents. But I confront some confusing problems. When I am trying to rewrite the agent_sampling.py, I found that the velocity has two unfathomable value like this:![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3722352%2F63cf70ed25bd56aeab6aa325555455e7%2F1491605109394_.pic.jpg?generation=1605109414580993&alt=media)\n\nCould you please give me some instructions to help me figure it out?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 995730,
      "author_name": "kneroma",
      "author_url": "",
      "post_date": "09/02/2020 18:16:45",
      "content": "<p>For sure the winning solution will be a neural net, given the non-classicality of the loss function. But, odds are that the winning solution won't be based on computer vision models.</p>\n<p>This competition is, according to my understanding, a mix between computer vision and times series. A solution that find the right way to combine models from both of those doamains, won't only be innovative but could even score among top ranked solutions in the end of this interesting competition.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1075333,
      "author_name": "spicychicken38",
      "author_url": "",
      "post_date": "11/11/2020 15:44:26",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/iglovikov\" target=\"_blank\">@iglovikov</a>. I am playing around with the source code of l5kit trying to load data that is available but not used yet, like the velocity of agents. But I confront some confusing problems. When I am trying to rewrite the agent_sampling.py, I found that the velocity has two unfathomable value like this:<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3722352%2F63cf70ed25bd56aeab6aa325555455e7%2F1491605109394_.pic.jpg?generation=1605109414580993&amp;alt=media\" alt=\"\"></p>\n<p>Could you please give me some instructions to help me figure it out?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "994800": "A few participants asked if the dataset could be used for academic purposes.\n\nThe answer is *yes*. You can, and if you will, it will make us happy.\n\nA month ago we had a [webinar](https://vimeo.com/451293003) to help folks start building models on the dataset.\n\nWe received a lot of great questions during Q&A.  We did not have time to address them.\n\nI post them here as a resource for you.\n\n## Competition\n\n### Q: Why are you focusing on prediction? Are “easier” problems like perception and planning already solved?\n\nPerception uses segmentation and detection algorithms on the lidar, radar, and imagery data. There are many challenges there, but overall these algorithms have a lot of attention from the research community. The task for last year’s competition was: 3D object detection on the lidar and imagery data and was perception. \n\nThis time we decided to explore the next part of the autonomous stack and work on prediction. The prediction problem is new and Kaggle hasn’t hosted the challenge of this type. We believe that winners will not be teams with the most hardware, but people who are the most creative. We hope that this competition will attract attention to the prediction problem, generate new ideas, and be a lot of fun for participants. \n\n### Q: I participated in the Kaggle competition hosted by Lyft last year. Does this competition have any relationship with the previous one? \n\nThe relationship is not strong. The dataset, the task, metric, the required skill set, and domain knowledge are different. \n\n### Q: Do we have access to some promotional cloud credits for people starting out in competitions? \n\nYes, we will be offering $300 in GCP credits to a set of participants that will beat the sample submission by organizers. \n\n### Q: Do we have to focus on predicting the action of cars? Or can we focus on the actions of humans and other agents?\n\nOur dataset is not 100% focused on cars. It also includes pedestrians and cyclists, both in train and evaluation. However, the vast majority of agents available are vehicles. \n\n### Q: Do we need to predict the motion for multiple agents simultaneously in a given frame, or is only the ego vehicle considered? \n\nThis competition is all about other agents, so you will have to predict multiple agents in a given frame. \n\n### Q: Are we restricted to the 2D rasterization of the environment to make motion prediction decisions?\n\nNo, you are not. You have the dataset; you have a target metric. Feel free to use any method that will make your submission more robust. \n\n### Q: Are we aware of the goal/destination of the car?\n\nBecause you will be predicting other traffic agents, we are not fully aware of the destination goal of the car. However, nothing prevents you from collecting future observations for a specific agent to build the “goal” of that agent. \n\n### Q: Can the motion prediction problem be designed as a graph problem?\n\nIt is possible. If it helps with the challenge, go for it! \n\n### Q: Do you believe that the solution is in deep learning architecture? Are there any recommendations on which deep learning models will be suitable for solving this challenge?\n\nWe suspect that the solution will be based on deep learning architecture. But we will not be surprised if other methods might be better. The Kaggle community is very creative and we are very curious to see what techniques will perform well. \n\n### Q: Does the neural net have to be RESNET-50, or can we try different architectures? \n\nNo, there is no restriction for the network architecture or method that you can use in your solution. \n\n### Q: Will the competition support models that generate multiple hypotheses per agent? Or will it just support one hypothesis as seen in the example? \n\nYes, the submitted solutions will be scored using a multi-modal metric. However, we will put a limit on the number of concurrent hypotheses. \n\n### Q: Will you all be releasing the pre-trained baselines?\n\nWe will release a sample baseline based on the code in the [l5kit examples](https://github.com/lyft/l5kit/tree/master/examples/agent_motion_prediction).\n\n## Dataset\n\n### Q: Where can I access the dataset?\n\nYou can download the dataset [here](https://self-driving.lyft.com/level5/prediction/). You can also view our [paper](https://arxiv.org/abs/2006.14480), and access the [L5kit](https://github.com/lyft/l5kit) to help you get started.\n\n### Q: What's the main difference between the raster shown in the baseline approach and one on the visualization slide (shown @ 32:30 in the [recording](https://vimeo.com/451293003))?\n\nThe baseline approach employs an open-sourced version of our internal rasterizer. We’re working to close the gap between the two. Right now, the only real difference is the lane color.\n\n### Q: Can you share more about the input raster? Is the input prepared from the table entries (shown @ 16:38 in the [recording](https://vimeo.com/451293003))?\n\nWhen we create a raster what we do is fetching an agent in a given frame in a given scene. As such, all three tables need to be queried to retrieve the final result. Moreover, because we also want to rasterize other agents from the same scene, we perform multiple queries, especially on that table. The code responsible for this part is already open-sourced and available in our L5Kit toolkit. \n\n### Q: What's the model used for training?\n\nWhat's the model used for training?\n\n### Q: How do you choose which scenes to train your model? \n\nFor our deep learning models, we pick all of them. The more data the better. \n\n### Q: Object perception and tracking always have uncertainties in the object state that it outputs. Are the values for these uncertainties included in the dataset?\n\nYes, each detected agent has a probability assigned from perception output. Those values are stored in the dataset. \n\n\n### Q: Is there an approach to get information about the turn signals of the surrounding vehicles?\n\nCurrently, perception estimates the turn signal state and provides the information to the prediction of the vehicle. \n\n### Q: Does the perception stack give us the probabilities associated with each prediction?\n\nYes, each detected agent has a probability assigned from perception output. Those values are stored in the dataset. \n\n### Q: What are the time constraints for planning? (I.e. do you target 30Hz or more/less?) \n\nPrediction on an AV typically targets 10 hertz with a latency of 100 milliseconds. See the Kaggle website for details regarding the competition resource and timing constraints. \n\n### Q: How many frames into the future should the model predict? What is the ideal time frame in general based on the reaction time of the ego vehicle?\n\nFor the challenge, we specifically target 50 frames in the future at 10hz. We wanted a prediction window large enough to stress the proposed solutions but small enough to fit in the input raster. \n\n### Q: What is the ratio of situations which are easy to predict (driving straight) and hard situations (collisions, jaywalking, unexpected turns)?\n\nThe dataset situations are skewed towards driving straight for the AV. However, when we consider all the agents surrounding it (the real focus of the competition), it’s difficult to draw an exact unsupervised ratio computation. \n\n### Q: Which version of Python does this use?\n\nPython 3 (3.6 onwards) \n\n### Q: How was the ground truth data created?\n\nThe ground truth data was created using the output of our perception stack. Semantic information for the map elements has been manually annotated. \n\n\n### Q: I would like to reproduce the performance metrics in the baseline paper. The rasterized map has a different color for each direction of the street, but the current code on the repo does not have this detail. Can you explain?\n\nThat picture was obtained using an internal proprietary version of the semantic rasterizer. We’re working to include many of those features in our open-source version. \n\n### Q: Does the dataset provide a label for which trajectories are considered “good” or “bad” driving to make sure models learn good behavior?\n\nNo, the dataset doesn’t include any behavioural notions. \n\n### Q: Do all the agents in each frame in the dataset have trajectories? \n\nNo, agents come and go and their detection and tracking may not be perfect. We provide ways of filtering out completely unreliable agents in our toolkit. \n\n### Q: When the ego vehicle was crossing an intersection in one of the examples of the rasterized overhead view, a couple of objects morphed into an object closer to the size of a bus when they entered the intersection, and then morphed back into 2 vehicles when it crossed the intersection. Are these kinds of artifacts expected in the ground truth?\n\nIndependent of the quality of our perception stack, there will always be artifacts in the data. These situations are absolutely possible, so we and you should expect them to happen in the ground truth and learn to account for them. However, we provide tools for filtering out some of those artifacts from rasterization centers (i.e. you won’t rasterize around those faulty agents, but they will still be visible around you). This limits the noise presence in the ground truth trajectories. \n\n### Q: I am interested in using this dataset for imitation learning. In a planning setting, you normally have a route to follow. Is it possible to reliably extract this route for every scene from the lane graph protobuf provided with the dataset?\n\nFor the AV, that’s definitely possible, as the information is accurate. You can extract the translations directly from the zarr and build a route. For agents, this is still an option, but you will encounter much more variety in terms of data quality. Once you have the routes coordinates, it is easier to query the protobuf API to get lane information. \n\n### Q: What have been the limitations and failure cases in real-life scenarios and in the lab? \n\nThere are many limitations depending on the algorithm in question. For example, a pure velocity rollout may overshoot a vehicle slowing down, and undershoot on a vehicle speeding up. Incorporating the acceleration helps in many scenarios, but we still have encountered some problems — for example, a car coming to a stop ended up rolling backward with constant acceleration models. So that is an example we corrected.\n\nBasically the heuristics enter into a cat and mouse game of ever-increasing complexity to resolve edge conditions. It turns out that these models reach such specificity they can become either unintelligible or fail to generalize well enough to be useful. Particularly challenging situations revolve around occlusions, unprotected turns, parked/not parked states, stop signs, and any cases where the AV’s behavior strongly influences the agent’s behavior. \n\n### Q: Are there any high-level annotations such as pedestrian actions?\n\nTo keep things simple, no. For this competition, you will be working on the raw output of our perception stack. \n\n### Q: In frame data, we see no ego velocity/acceleration. Any reason?\n\nWe’ve been focused on predicting other agents so this hasn’t made it into the dataset yet. It should be easy to add this in for the future because these values are readily available. \n\n## SELF-DRIVING CAREERS AND RESOURCES \n\n### Q: What internship and career opportunities are available at Lyft Level 5?\n\nYou can see our open positions [here](https://www.lyft.com/careers). \n\n### Q: What level of degree do I need to join the team?\n\nLyft does not have a degree requirement for joining the team. \n\n### Q: Can I work as an intern remotely?\n\nCurrently, Lyft team members who can conduct their work remotely are able to do so given shelter in place. However, interns normally are required to be in-office. \n\n### Q: As a student who's trying to learn more about AVs and a career in the field, how did you get your start? Any tips for a fellow student?\n\nCheck out our interview with [Vladimir Iglovikov](https://medium.com/lyftlevel5/employee-spotlight-vladimir-iglovikov-11b71ed02bc) for his insights on this topic!",
    "995730": "For sure the winning solution will be a neural net, given the non-classicality of the loss function. But, odds are that the winning solution won't be based on computer vision models.\n\nThis competition is, according to my understanding, a mix between computer vision and times series. A solution that find the right way to combine models from both of those doamains, won't only be innovative but could even score among top ranked solutions in the end of this interesting competition.",
    "1075333": "Hi @iglovikov. I am playing around with the source code of l5kit trying to load data that is available but not used yet, like the velocity of agents. But I confront some confusing problems. When I am trying to rewrite the agent_sampling.py, I found that the velocity has two unfathomable value like this:![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3722352%2F63cf70ed25bd56aeab6aa325555455e7%2F1491605109394_.pic.jpg?generation=1605109414580993&alt=media)\n\nCould you please give me some instructions to help me figure it out?"
  },
  "source": "meta"
}