{
  "id": 178781,
  "title": "Extracting insights from Dataset paper",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/178781",
  "author_name": "Sarthak khandelwal",
  "post_date": "2020-08-31T11:28:02.508000",
  "votes": 25,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Recently I joined this competition and have gone through the research paper for the Dataset provided in this competition. Anyone would get on track much more faster if they know about the data very well and hence I decided to share the insights that I have extracted from that paper. Those who have already gone through the paper can also use this as a quick reference.</p>\n<p>Let's look at the insights below-</p>\n<h2><strong>Information about Dataset</strong></h2>\n<ul>\n<li>Dataset contains <code>1000 hours</code> of data along a fixed route. </li>\n<li>Data is being recorded along a single route and hence continous motion can be seen throughout the dataset.</li>\n<li>Used <code>20 Autonomous Vehicles</code> to prepare data (unique id has also been provided in the dataset).</li>\n<li>Dataset contains high definition semantic map with <code>15,242</code> labelled elements.</li>\n<li>Dataset contains high resolution aerial image of the <code>Palo Alto</code> area.</li>\n<li>The dataset is used for machine learning tasks such as <code>Motion Forecasting</code>, <code>Motion Planning</code> and <code>Motion Simulation</code>.</li>\n<li>This dataset contains annotations in the form of trajectories and task is to<br>\npredict the motion of the autonomous vehicle (Motion Forecasting). We are using the output from the <code>Perception Network</code> which converts the data gathered by the sensors to prepare info about the traffic agents.</li>\n<li>3 Lidars, 7 Cameras and 5 radars are attached to each autonomous vehicle to collect the data.</li>\n<li>The dataset is present in the form of n-dimensional compressed <code>zarr arrays</code>.</li>\n<li>Dataset is divided in 83-7-10% ratio.</li>\n</ul>\n<p>Dataset has three components:</p>\n<ul>\n<li>170K scenes each 25 seconds long which captures movement of self driving<br>\nvehicles and traffic participants around it.</li>\n<li>An HD semantic map that captures road rules, lane geometry and other traffic<br>\nelements.</li>\n<li>An HD aerial pic of the area.</li>\n</ul>\n<p>Each traffic participant is represented by 2.5D cuboid, yaw, yaw_rate, velocity,<br>\nacceleration and a class label.</p>\n<p>Usage of graph neural networks and leveraging Bird's Eye View (BEV) are some of the leading solutions in this domain.</p>\n<p>We can sample the dataset for various tasks separately such as for the ego vehicle (using the <code>EgoDataset</code>) or for other traffic paritcipants (using <code>AgentDataset</code>). </p>\n<p>A great image to learn about different transformations in the dataset by <a href=\"https://www.kaggle.com/corochann\" target=\"_blank\">corochann</a>- </p>\n<h2><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3449691%2Ffb20038bfc7832666b64668345b4316b%2Fl5kit_class.png?generation=1598872879049796&amp;alt=media\" alt=\"\"></h2>\n<h2><strong>Miscellanous Definitions</strong></h2>\n<p><strong>Lidar/Ladar:</strong> Method for measuring distances by illuminating target with laser light<br>\n  and calculating the reflection with a sensor.</p>\n<p><strong>Yaw</strong>: Amount of rotation of an object with respect to vertical axis.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3449691%2F153ca85b480d58b8041572fa0d96372f%2FFlight_dynamics_with_text-compressed.jpg?generation=1598868978665499&amp;alt=media\" alt=\"\"></p>\n<p><strong>HD semantic map:</strong> There are a total of <code>15,242</code> labelled traffic elements in the<br>\nsemantic map which includes <code>8,505</code> lane segments. The map was created by<br>\n<code>simultanous localization and mapping (SLAM) system</code> and was annotated by human<br>\ncurators. Information in the map could be used for planning driving strategy and<br>\nto anticipate the future movements of other traffic participants. This map is<br>\nprovided in the form of <a href=\"https://developers.google.com/protocol-buffers\" target=\"_blank\">protocol buffers</a>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3449691%2Fae237dcaf2504a6ec1483f88f973f90c%2Fsemantic_map-compressed.jpg?generation=1598867722441664&amp;alt=media\" alt=\"\"></p>\n<p><strong>HD aerial map:</strong> It surrounds the area of <code>Palo Alto</code>. It covers an area of <code>74km\nsquare</code> and is provided as <code>181 GeoTIFF tiles</code> of size <code>10560 x 10560</code> each<br>\nspanning approx 640 x 640 meters. (I think the provided size is the map size that we are rastering. Please correct if I'm wrong).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3449691%2Fd287e65706eb57c25628325910acb844%2FAerial_map-compressed.jpg?generation=1598867295682801&amp;alt=media\" alt=\"\"></p>\n<p><strong>Rasterization:</strong>  Converting electronic signals or an image described in vector<br>\ngraphics into raster images (series of pixels which when combined together form<br>\nan human interpretable image.)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3449691%2Fe75c36ec8ba0c3ac27cdd75cfa182d25%2FRaster_image.png?generation=1598867032182287&amp;alt=media\" alt=\"\"></p>\n<p>I would keep on updating this post as soon as I get new insights in competition. Suggestions are always welcomed. Thanks!</p>",
  "messages": [
    {
      "id": 992707,
      "postDate": "2020-08-31T11:28:02.507Z",
      "content": "<p>Recently I joined this competition and have gone through the research paper for the Dataset provided in this competition. Anyone would get on track much more faster if they know about the data very well and hence I decided to share the insights that I have extracted from that paper. Those who have already gone through the paper can also use this as a quick reference.</p>\n<p>Let's look at the insights below-</p>\n<h2><strong>Information about Dataset</strong></h2>\n<ul>\n<li>Dataset contains <code>1000 hours</code> of data along a fixed route. </li>\n<li>Data is being recorded along a single route and hence continous motion can be seen throughout the dataset.</li>\n<li>Used <code>20 Autonomous Vehicles</code> to prepare data (unique id has also been provided in the dataset).</li>\n<li>Dataset contains high definition semantic map with <code>15,242</code> labelled elements.</li>\n<li>Dataset contains high resolution aerial image of the <code>Palo Alto</code> area.</li>\n<li>The dataset is used for machine learning tasks such as <code>Motion Forecasting</code>, <code>Motion Planning</code> and <code>Motion Simulation</code>.</li>\n<li>This dataset contains annotations in the form of trajectories and task is to<br>\npredict the motion of the autonomous vehicle (Motion Forecasting). We are using the output from the <code>Perception Network</code> which converts the data gathered by the sensors to prepare info about the traffic agents.</li>\n<li>3 Lidars, 7 Cameras and 5 radars are attached to each autonomous vehicle to collect the data.</li>\n<li>The dataset is present in the form of n-dimensional compressed <code>zarr arrays</code>.</li>\n<li>Dataset is divided in 83-7-10% ratio.</li>\n</ul>\n<p>Dataset has three components:</p>\n<ul>\n<li>170K scenes each 25 seconds long which captures movement of self driving<br>\nvehicles and traffic participants around it.</li>\n<li>An HD semantic map that captures road rules, lane geometry and other traffic<br>\nelements.</li>\n<li>An HD aerial pic of the area.</li>\n</ul>\n<p>Each traffic participant is represented by 2.5D cuboid, yaw, yaw_rate, velocity,<br>\nacceleration and a class label.</p>\n<p>Usage of graph neural networks and leveraging Bird's Eye View (BEV) are some of the leading solutions in this domain.</p>\n<p>We can sample the dataset for various tasks separately such as for the ego vehicle (using the <code>EgoDataset</code>) or for other traffic paritcipants (using <code>AgentDataset</code>). </p>\n<p>A great image to learn about different transformations in the dataset by <a href=\"https://www.kaggle.com/corochann\" target=\"_blank\">corochann</a>- </p>\n<h2><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3449691%2Ffb20038bfc7832666b64668345b4316b%2Fl5kit_class.png?generation=1598872879049796&amp;alt=media\" alt=\"\"></h2>\n<h2><strong>Miscellanous Definitions</strong></h2>\n<p><strong>Lidar/Ladar:</strong> Method for measuring distances by illuminating target with laser light<br>\n  and calculating the reflection with a sensor.</p>\n<p><strong>Yaw</strong>: Amount of rotation of an object with respect to vertical axis.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3449691%2F153ca85b480d58b8041572fa0d96372f%2FFlight_dynamics_with_text-compressed.jpg?generation=1598868978665499&amp;alt=media\" alt=\"\"></p>\n<p><strong>HD semantic map:</strong> There are a total of <code>15,242</code> labelled traffic elements in the<br>\nsemantic map which includes <code>8,505</code> lane segments. The map was created by<br>\n<code>simultanous localization and mapping (SLAM) system</code> and was annotated by human<br>\ncurators. Information in the map could be used for planning driving strategy and<br>\nto anticipate the future movements of other traffic participants. This map is<br>\nprovided in the form of <a href=\"https://developers.google.com/protocol-buffers\" target=\"_blank\">protocol buffers</a>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3449691%2Fae237dcaf2504a6ec1483f88f973f90c%2Fsemantic_map-compressed.jpg?generation=1598867722441664&amp;alt=media\" alt=\"\"></p>\n<p><strong>HD aerial map:</strong> It surrounds the area of <code>Palo Alto</code>. It covers an area of <code>74km\nsquare</code> and is provided as <code>181 GeoTIFF tiles</code> of size <code>10560 x 10560</code> each<br>\nspanning approx 640 x 640 meters. (I think the provided size is the map size that we are rastering. Please correct if I'm wrong).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3449691%2Fd287e65706eb57c25628325910acb844%2FAerial_map-compressed.jpg?generation=1598867295682801&amp;alt=media\" alt=\"\"></p>\n<p><strong>Rasterization:</strong>  Converting electronic signals or an image described in vector<br>\ngraphics into raster images (series of pixels which when combined together form<br>\nan human interpretable image.)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3449691%2Fe75c36ec8ba0c3ac27cdd75cfa182d25%2FRaster_image.png?generation=1598867032182287&amp;alt=media\" alt=\"\"></p>\n<p>I would keep on updating this post as soon as I get new insights in competition. Suggestions are always welcomed. Thanks!</p>",
      "rawMarkdown": "Recently I joined this competition and have gone through the research paper for the Dataset provided in this competition. Anyone would get on track much more faster if they know about the data very well and hence I decided to share the insights that I have extracted from that paper. Those who have already gone through the paper can also use this as a quick reference.\n\nLet's look at the insights below-\n\n## **Information about Dataset**\n- Dataset contains `1000 hours` of data along a fixed route. \n- Data is being recorded along a single route and hence continous motion can be seen throughout the dataset.\n- Used `20 Autonomous Vehicles` to prepare data (unique id has also been provided in the dataset).\n- Dataset contains high definition semantic map with `15,242` labelled elements.\n- Dataset contains high resolution aerial image of the `Palo Alto` area.\n- The dataset is used for machine learning tasks such as `Motion Forecasting`, `Motion Planning` and `Motion Simulation`.\n- This dataset contains annotations in the form of trajectories and task is to\n  predict the motion of the autonomous vehicle (Motion Forecasting). We are using the output from the `Perception Network` which converts the data gathered by the sensors to prepare info about the traffic agents.\n- 3 Lidars, 7 Cameras and 5 radars are attached to each autonomous vehicle to collect the data.\n- The dataset is present in the form of n-dimensional compressed `zarr arrays`.\n- Dataset is divided in 83-7-10% ratio.\n\nDataset has three components:\n- 170K scenes each 25 seconds long which captures movement of self driving\n  vehicles and traffic participants around it.\n- An HD semantic map that captures road rules, lane geometry and other traffic\n  elements.\n- An HD aerial pic of the area.\n\nEach traffic participant is represented by 2.5D cuboid, yaw, yaw_rate, velocity,\nacceleration and a class label.\n\nUsage of graph neural networks and leveraging Bird's Eye View (BEV) are some of the leading solutions in this domain.\n\nWe can sample the dataset for various tasks separately such as for the ego vehicle (using the `EgoDataset`) or for other traffic paritcipants (using `AgentDataset`). \n\nA great image to learn about different transformations in the dataset by [corochann](https://www.kaggle.com/corochann)- \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3449691%2Ffb20038bfc7832666b64668345b4316b%2Fl5kit_class.png?generation=1598872879049796&alt=media)\n-------------------------------------------------------------------------------------------------------\n## **Miscellanous Definitions**\n\n**Lidar/Ladar:** Method for measuring distances by illuminating target with laser light\n  and calculating the reflection with a sensor.\n\n**Yaw**: Amount of rotation of an object with respect to vertical axis.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3449691%2F153ca85b480d58b8041572fa0d96372f%2FFlight_dynamics_with_text-compressed.jpg?generation=1598868978665499&alt=media)\n\n**HD semantic map:** There are a total of `15,242` labelled traffic elements in the\nsemantic map which includes `8,505` lane segments. The map was created by\n`simultanous localization and mapping (SLAM) system` and was annotated by human\ncurators. Information in the map could be used for planning driving strategy and\nto anticipate the future movements of other traffic participants. This map is\nprovided in the form of [protocol buffers](https://developers.google.com/protocol-buffers).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3449691%2Fae237dcaf2504a6ec1483f88f973f90c%2Fsemantic_map-compressed.jpg?generation=1598867722441664&alt=media)\n\n**HD aerial map:** It surrounds the area of `Palo Alto`. It covers an area of `74km\nsquare` and is provided as `181 GeoTIFF tiles` of size `10560 x 10560` each\nspanning approx 640 x 640 meters. (I think the provided size is the map size that we are rastering. Please correct if I'm wrong).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3449691%2Fd287e65706eb57c25628325910acb844%2FAerial_map-compressed.jpg?generation=1598867295682801&alt=media)\n\n**Rasterization:**  Converting electronic signals or an image described in vector\ngraphics into raster images (series of pixels which when combined together form\nan human interpretable image.)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3449691%2Fe75c36ec8ba0c3ac27cdd75cfa182d25%2FRaster_image.png?generation=1598867032182287&alt=media)\n\n\nI would keep on updating this post as soon as I get new insights in competition. Suggestions are always welcomed. Thanks!",
      "votes": 25
    },
    {
      "id": 992752,
      "postDate": "2020-08-31T12:20:53.087Z",
      "content": "<p>Great insights! Thanks for sharing. I had one doubt though. The \"target_positions\" in the dataset is not directly interpretable, it requires some transformation using the \"yaw\". I've observed this in l5kit's visualization functions. Is there any explanation about this? Like how these parameters are encoded in the paper?</p>",
      "rawMarkdown": "Great insights! Thanks for sharing. I had one doubt though. The \"target_positions\" in the dataset is not directly interpretable, it requires some transformation using the \"yaw\". I've observed this in l5kit's visualization functions. Is there any explanation about this? Like how these parameters are encoded in the paper?",
      "replies": [
        {
          "id": 993158,
          "postDate": "2020-08-31T17:43:34.157Z",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/kaushal2896\" target=\"_blank\">@kaushal2896</a> glad you liked my work. </p>\n<p>Basically <code>target_positions</code> describe the displacement of the vehicle (ego or any other participant) in that particular scene and we add it with the <code>centroid</code> of that vehicle to make it understand from where it should start and then move through the cordinates described in <code>target_positions</code>. We also use <code>target_yaws</code> in the draw_trajectory function which describes the rotation of the AV w.r.t vertical axis (for lane changing or taking a turn).</p>\n<p>We also used <code>world_to_image</code> attribute of the scene data as the <code>(x,y)</code> coordinates are w.r.t real world and to show it in an image format we need to convert or project that points onto the image subspace (pixels). It is pretty much similar as we do projection from a vector space into a subspace.</p>\n<p>However the above explanation is how I interpreted it. A clear information has not been mentioned AFIK. But I've found relevant areas that might be helpful-</p>\n<p><code>visualization.ipynb</code> - <br>\n\"If we want to plot the ground truth trajectory, we can convert the dataset's target_position (displacements in meters in world coordinates) into pixel coordinates in the image space, and call our utility function draw_trajectory (note that you can use this function for the predicted trajectories, as well).\"</p>\n<p><code>paper</code>-<br>\n\"It covers an area of 74km square and is provided as 181 GeoTIFF tiles of size 10560 x 10560 pixels each<br>\nspanning approx 640 x 640 meters.\"</p>",
          "rawMarkdown": "Hey @kaushal2896 glad you liked my work. \n\nBasically `target_positions` describe the displacement of the vehicle (ego or any other participant) in that particular scene and we add it with the `centroid` of that vehicle to make it understand from where it should start and then move through the cordinates described in `target_positions`. We also use `target_yaws` in the draw_trajectory function which describes the rotation of the AV w.r.t vertical axis (for lane changing or taking a turn).\n\nWe also used `world_to_image` attribute of the scene data as the `(x,y)` coordinates are w.r.t real world and to show it in an image format we need to convert or project that points onto the image subspace (pixels). It is pretty much similar as we do projection from a vector space into a subspace.\n\nHowever the above explanation is how I interpreted it. A clear information has not been mentioned AFIK. But I've found relevant areas that might be helpful-\n\n`visualization.ipynb` - \n\"If we want to plot the ground truth trajectory, we can convert the dataset's target_position (displacements in meters in world coordinates) into pixel coordinates in the image space, and call our utility function draw_trajectory (note that you can use this function for the predicted trajectories, as well).\"\n\n`paper`-\n\"It covers an area of 74km square and is provided as 181 GeoTIFF tiles of size 10560 x 10560 pixels each\nspanning approx 640 x 640 meters.\""
        }
      ]
    },
    {
      "id": 997089,
      "postDate": "2020-09-03T18:27:17.313Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 997122,
          "postDate": "2020-09-03T18:56:21.103Z",
          "content": "<p><a href=\"https://arxiv.org/pdf/2006.14480.pdf\" target=\"_blank\">https://arxiv.org/pdf/2006.14480.pdf</a></p>",
          "rawMarkdown": "https://arxiv.org/pdf/2006.14480.pdf",
          "votes": 1
        },
        {
          "id": 997168,
          "postDate": "2020-09-03T19:47:17.800Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 992752,
      "author_name": "Kaushal Shah",
      "author_url": "",
      "post_date": "2020-08-31T12:20:53.087000",
      "content": "<p>Great insights! Thanks for sharing. I had one doubt though. The \"target_positions\" in the dataset is not directly interpretable, it requires some transformation using the \"yaw\". I've observed this in l5kit's visualization functions. Is there any explanation about this? Like how these parameters are encoded in the paper?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 993158,
          "author_name": "Sarthak khandelwal",
          "author_url": "",
          "post_date": "2020-08-31T17:43:34.157000",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/kaushal2896\" target=\"_blank\">@kaushal2896</a> glad you liked my work. </p>\n<p>Basically <code>target_positions</code> describe the displacement of the vehicle (ego or any other participant) in that particular scene and we add it with the <code>centroid</code> of that vehicle to make it understand from where it should start and then move through the cordinates described in <code>target_positions</code>. We also use <code>target_yaws</code> in the draw_trajectory function which describes the rotation of the AV w.r.t vertical axis (for lane changing or taking a turn).</p>\n<p>We also used <code>world_to_image</code> attribute of the scene data as the <code>(x,y)</code> coordinates are w.r.t real world and to show it in an image format we need to convert or project that points onto the image subspace (pixels). It is pretty much similar as we do projection from a vector space into a subspace.</p>\n<p>However the above explanation is how I interpreted it. A clear information has not been mentioned AFIK. But I've found relevant areas that might be helpful-</p>\n<p><code>visualization.ipynb</code> - <br>\n\"If we want to plot the ground truth trajectory, we can convert the dataset's target_position (displacements in meters in world coordinates) into pixel coordinates in the image space, and call our utility function draw_trajectory (note that you can use this function for the predicted trajectories, as well).\"</p>\n<p><code>paper</code>-<br>\n\"It covers an area of 74km square and is provided as 181 GeoTIFF tiles of size 10560 x 10560 pixels each<br>\nspanning approx 640 x 640 meters.\"</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 997089,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-03T18:27:17.313000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 997122,
          "author_name": "sajwankit",
          "author_url": "",
          "post_date": "2020-09-03T18:56:21.103000",
          "content": "<p><a href=\"https://arxiv.org/pdf/2006.14480.pdf\" target=\"_blank\">https://arxiv.org/pdf/2006.14480.pdf</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 997168,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-09-03T19:47:17.800000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "992707": "Recently I joined this competition and have gone through the research paper for the Dataset provided in this competition. Anyone would get on track much more faster if they know about the data very well and hence I decided to share the insights that I have extracted from that paper. Those who have already gone through the paper can also use this as a quick reference.\n\nLet's look at the insights below-\n\n## **Information about Dataset**\n- Dataset contains `1000 hours` of data along a fixed route. \n- Data is being recorded along a single route and hence continous motion can be seen throughout the dataset.\n- Used `20 Autonomous Vehicles` to prepare data (unique id has also been provided in the dataset).\n- Dataset contains high definition semantic map with `15,242` labelled elements.\n- Dataset contains high resolution aerial image of the `Palo Alto` area.\n- The dataset is used for machine learning tasks such as `Motion Forecasting`, `Motion Planning` and `Motion Simulation`.\n- This dataset contains annotations in the form of trajectories and task is to\n  predict the motion of the autonomous vehicle (Motion Forecasting). We are using the output from the `Perception Network` which converts the data gathered by the sensors to prepare info about the traffic agents.\n- 3 Lidars, 7 Cameras and 5 radars are attached to each autonomous vehicle to collect the data.\n- The dataset is present in the form of n-dimensional compressed `zarr arrays`.\n- Dataset is divided in 83-7-10% ratio.\n\nDataset has three components:\n- 170K scenes each 25 seconds long which captures movement of self driving\n  vehicles and traffic participants around it.\n- An HD semantic map that captures road rules, lane geometry and other traffic\n  elements.\n- An HD aerial pic of the area.\n\nEach traffic participant is represented by 2.5D cuboid, yaw, yaw_rate, velocity,\nacceleration and a class label.\n\nUsage of graph neural networks and leveraging Bird's Eye View (BEV) are some of the leading solutions in this domain.\n\nWe can sample the dataset for various tasks separately such as for the ego vehicle (using the `EgoDataset`) or for other traffic paritcipants (using `AgentDataset`). \n\nA great image to learn about different transformations in the dataset by [corochann](https://www.kaggle.com/corochann)- \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3449691%2Ffb20038bfc7832666b64668345b4316b%2Fl5kit_class.png?generation=1598872879049796&alt=media)\n-------------------------------------------------------------------------------------------------------\n## **Miscellanous Definitions**\n\n**Lidar/Ladar:** Method for measuring distances by illuminating target with laser light\n  and calculating the reflection with a sensor.\n\n**Yaw**: Amount of rotation of an object with respect to vertical axis.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3449691%2F153ca85b480d58b8041572fa0d96372f%2FFlight_dynamics_with_text-compressed.jpg?generation=1598868978665499&alt=media)\n\n**HD semantic map:** There are a total of `15,242` labelled traffic elements in the\nsemantic map which includes `8,505` lane segments. The map was created by\n`simultanous localization and mapping (SLAM) system` and was annotated by human\ncurators. Information in the map could be used for planning driving strategy and\nto anticipate the future movements of other traffic participants. This map is\nprovided in the form of [protocol buffers](https://developers.google.com/protocol-buffers).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3449691%2Fae237dcaf2504a6ec1483f88f973f90c%2Fsemantic_map-compressed.jpg?generation=1598867722441664&alt=media)\n\n**HD aerial map:** It surrounds the area of `Palo Alto`. It covers an area of `74km\nsquare` and is provided as `181 GeoTIFF tiles` of size `10560 x 10560` each\nspanning approx 640 x 640 meters. (I think the provided size is the map size that we are rastering. Please correct if I'm wrong).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3449691%2Fd287e65706eb57c25628325910acb844%2FAerial_map-compressed.jpg?generation=1598867295682801&alt=media)\n\n**Rasterization:**  Converting electronic signals or an image described in vector\ngraphics into raster images (series of pixels which when combined together form\nan human interpretable image.)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3449691%2Fe75c36ec8ba0c3ac27cdd75cfa182d25%2FRaster_image.png?generation=1598867032182287&alt=media)\n\n\nI would keep on updating this post as soon as I get new insights in competition. Suggestions are always welcomed. Thanks!",
    "992752": "Great insights! Thanks for sharing. I had one doubt though. The \"target_positions\" in the dataset is not directly interpretable, it requires some transformation using the \"yaw\". I've observed this in l5kit's visualization functions. Is there any explanation about this? Like how these parameters are encoded in the paper?",
    "997089": ""
  }
}