{
  "id": 119051,
  "title": "Our experiments with 'semantic point clouds'",
  "url": "/competitions/3d-object-detection-for-autonomous-vehicles/discussion/119051",
  "author_name": "Simon Grest",
  "post_date": "2019-11-26T10:33:39.590000",
  "votes": 11,
  "comment_count": 2,
  "views": 0,
  "content": "<p>We (<a href=\"https://www.kaggle.com/artste\"></a><a href=\"/artste\">@artste</a> and me) added onto the reference model by running a pre-trained semantic segmentation on all the images and then 'projecting' the predicted classes onto the LIDAR point cloud.</p>\n\n<p>We posted a full piece on our work here for anyone interested in reading more: <a href=\"https://towardsdatascience.com/drawing-a-million-boxes-around-objects-on-the-roads-of-palo-alto-cd29a72ee1eb\">https://towardsdatascience.com/drawing-a-million-boxes-around-objects-on-the-roads-of-palo-alto-cd29a72ee1eb</a></p>\n\n<p>These plots illustrate how we projected the semantic classes onto the corresponding section of the point cloud:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F116844%2F7a578c86a7dc544f7ac2c2a05ecb48f0%2Fprojecting_to_point_clouds.png?generation=1574763688402625&amp;alt=media\" alt=\"\"></p>\n\n<p>Then combining them all into a 'semantic point cloud' for the sample:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F116844%2F4ede1fcc6ac0d6a6125ea546de89c5c4%2Fbanner_cropped.png?generation=1574763462396290&amp;alt=media\" alt=\"\"></p>\n\n<p>We then created a 2D representation of the 'semantic point cloud' using the Birds Eye View approach and a 5 dimensional embedding.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F116844%2F496833de03f712541f7d733940f6b960%2Fmodel_input_data.png?generation=1574763640202934&amp;alt=media\" alt=\"\"></p>\n\n<p>We also experimented with post-processing the data to try to improve the elevation and height of the predicted volumes. We combined all the LIDAR point clouds for a scene and built a terrain elevation map by taking the minimum z values pooled over space and time.</p>\n\n<p>Combined point clouds for an entire scene\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F116844%2Ff31d50c9a6c9a68791701f8eaced002d%2Fscene_point_cloud.png?generation=1574764143687336&amp;alt=media\" alt=\"\"></p>\n\n<p>Minimum pooled elevation giving terrain map\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F116844%2Fc7f5483377798be0ddf4a91d0b2c4b82%2Fminimum_elevation.png?generation=1574764200092595&amp;alt=media\" alt=\"\"></p>\n\n<p>The results on the validation set were very promising, and we were surprised that this didn't improve our score at all!</p>\n\n<p>Here's an example of the impact of using the terrain elevation map on a sample in the training set where the ego vehicle is on a hill:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F116844%2F3bc61e6f385b640f6f21ad625a4c914c%2FUntitled.png?generation=1574763952363408&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 681623,
      "postDate": "2019-11-26T10:33:39.590Z",
      "content": "<p>We (<a href=\"https://www.kaggle.com/artste\"></a><a href=\"/artste\">@artste</a> and me) added onto the reference model by running a pre-trained semantic segmentation on all the images and then 'projecting' the predicted classes onto the LIDAR point cloud.</p>\n\n<p>We posted a full piece on our work here for anyone interested in reading more: <a href=\"https://towardsdatascience.com/drawing-a-million-boxes-around-objects-on-the-roads-of-palo-alto-cd29a72ee1eb\">https://towardsdatascience.com/drawing-a-million-boxes-around-objects-on-the-roads-of-palo-alto-cd29a72ee1eb</a></p>\n\n<p>These plots illustrate how we projected the semantic classes onto the corresponding section of the point cloud:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F116844%2F7a578c86a7dc544f7ac2c2a05ecb48f0%2Fprojecting_to_point_clouds.png?generation=1574763688402625&amp;alt=media\" alt=\"\"></p>\n\n<p>Then combining them all into a 'semantic point cloud' for the sample:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F116844%2F4ede1fcc6ac0d6a6125ea546de89c5c4%2Fbanner_cropped.png?generation=1574763462396290&amp;alt=media\" alt=\"\"></p>\n\n<p>We then created a 2D representation of the 'semantic point cloud' using the Birds Eye View approach and a 5 dimensional embedding.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F116844%2F496833de03f712541f7d733940f6b960%2Fmodel_input_data.png?generation=1574763640202934&amp;alt=media\" alt=\"\"></p>\n\n<p>We also experimented with post-processing the data to try to improve the elevation and height of the predicted volumes. We combined all the LIDAR point clouds for a scene and built a terrain elevation map by taking the minimum z values pooled over space and time.</p>\n\n<p>Combined point clouds for an entire scene\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F116844%2Ff31d50c9a6c9a68791701f8eaced002d%2Fscene_point_cloud.png?generation=1574764143687336&amp;alt=media\" alt=\"\"></p>\n\n<p>Minimum pooled elevation giving terrain map\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F116844%2Fc7f5483377798be0ddf4a91d0b2c4b82%2Fminimum_elevation.png?generation=1574764200092595&amp;alt=media\" alt=\"\"></p>\n\n<p>The results on the validation set were very promising, and we were surprised that this didn't improve our score at all!</p>\n\n<p>Here's an example of the impact of using the terrain elevation map on a sample in the training set where the ego vehicle is on a hill:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F116844%2F3bc61e6f385b640f6f21ad625a4c914c%2FUntitled.png?generation=1574763952363408&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "We ([@artste ](https://www.kaggle.com/artste)and me) added onto the reference model by running a pre-trained semantic segmentation on all the images and then 'projecting' the predicted classes onto the LIDAR point cloud.\n\nWe posted a full piece on our work here for anyone interested in reading more: https://towardsdatascience.com/drawing-a-million-boxes-around-objects-on-the-roads-of-palo-alto-cd29a72ee1eb\n\nThese plots illustrate how we projected the semantic classes onto the corresponding section of the point cloud:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F116844%2F7a578c86a7dc544f7ac2c2a05ecb48f0%2Fprojecting_to_point_clouds.png?generation=1574763688402625&amp;alt=media)\n\nThen combining them all into a 'semantic point cloud' for the sample:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F116844%2F4ede1fcc6ac0d6a6125ea546de89c5c4%2Fbanner_cropped.png?generation=1574763462396290&amp;alt=media)\n\nWe then created a 2D representation of the 'semantic point cloud' using the Birds Eye View approach and a 5 dimensional embedding.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F116844%2F496833de03f712541f7d733940f6b960%2Fmodel_input_data.png?generation=1574763640202934&amp;alt=media)\n\nWe also experimented with post-processing the data to try to improve the elevation and height of the predicted volumes. We combined all the LIDAR point clouds for a scene and built a terrain elevation map by taking the minimum z values pooled over space and time.\n\nCombined point clouds for an entire scene\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F116844%2Ff31d50c9a6c9a68791701f8eaced002d%2Fscene_point_cloud.png?generation=1574764143687336&amp;alt=media)\n\nMinimum pooled elevation giving terrain map\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F116844%2Fc7f5483377798be0ddf4a91d0b2c4b82%2Fminimum_elevation.png?generation=1574764200092595&amp;alt=media)\n\nThe results on the validation set were very promising, and we were surprised that this didn't improve our score at all!\n\nHere's an example of the impact of using the terrain elevation map on a sample in the training set where the ego vehicle is on a hill:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F116844%2F3bc61e6f385b640f6f21ad625a4c914c%2FUntitled.png?generation=1574763952363408&amp;alt=media)\n\n\n\n",
      "votes": 11
    },
    {
      "id": 681762,
      "postDate": "2019-11-26T14:16:19.243Z",
      "content": "<p>Very interesting and creative, I was hoping to see some sensor fusion approaches in the solutions! </p>",
      "rawMarkdown": "Very interesting and creative, I was hoping to see some sensor fusion approaches in the solutions! ",
      "votes": 1
    },
    {
      "id": 682285,
      "postDate": "2019-11-27T07:17:39.633Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 681762,
      "author_name": "Jack Vial",
      "author_url": "",
      "post_date": "2019-11-26T14:16:19.243000",
      "content": "<p>Very interesting and creative, I was hoping to see some sensor fusion approaches in the solutions! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 682285,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-27T07:17:39.633000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "681623": "We ([@artste ](https://www.kaggle.com/artste)and me) added onto the reference model by running a pre-trained semantic segmentation on all the images and then 'projecting' the predicted classes onto the LIDAR point cloud.\n\nWe posted a full piece on our work here for anyone interested in reading more: https://towardsdatascience.com/drawing-a-million-boxes-around-objects-on-the-roads-of-palo-alto-cd29a72ee1eb\n\nThese plots illustrate how we projected the semantic classes onto the corresponding section of the point cloud:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F116844%2F7a578c86a7dc544f7ac2c2a05ecb48f0%2Fprojecting_to_point_clouds.png?generation=1574763688402625&amp;alt=media)\n\nThen combining them all into a 'semantic point cloud' for the sample:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F116844%2F4ede1fcc6ac0d6a6125ea546de89c5c4%2Fbanner_cropped.png?generation=1574763462396290&amp;alt=media)\n\nWe then created a 2D representation of the 'semantic point cloud' using the Birds Eye View approach and a 5 dimensional embedding.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F116844%2F496833de03f712541f7d733940f6b960%2Fmodel_input_data.png?generation=1574763640202934&amp;alt=media)\n\nWe also experimented with post-processing the data to try to improve the elevation and height of the predicted volumes. We combined all the LIDAR point clouds for a scene and built a terrain elevation map by taking the minimum z values pooled over space and time.\n\nCombined point clouds for an entire scene\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F116844%2Ff31d50c9a6c9a68791701f8eaced002d%2Fscene_point_cloud.png?generation=1574764143687336&amp;alt=media)\n\nMinimum pooled elevation giving terrain map\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F116844%2Fc7f5483377798be0ddf4a91d0b2c4b82%2Fminimum_elevation.png?generation=1574764200092595&amp;alt=media)\n\nThe results on the validation set were very promising, and we were surprised that this didn't improve our score at all!\n\nHere's an example of the impact of using the terrain elevation map on a sample in the training set where the ego vehicle is on a hill:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F116844%2F3bc61e6f385b640f6f21ad625a4c914c%2FUntitled.png?generation=1574763952363408&amp;alt=media)\n\n\n\n",
    "681762": "Very interesting and creative, I was hoping to see some sensor fusion approaches in the solutions! ",
    "682285": ""
  }
}