{
  "id": 113030,
  "title": "Proposal for a non GPU solution :)",
  "url": "/competitions/3d-object-detection-for-autonomous-vehicles/discussion/113030",
  "author_name": "",
  "post_date": "2019-10-16T15:51:11.285605700Z",
  "votes": 12,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I was thinking about creating and submitting another baseline to the Leaderboard, but it looks like I will not have enough time for this :(</p>\n\n<p>The idea that I had in mind should create a solution that does not require any models to be trained.</p>\n\n<ol>\n<li>Find the ground map. For example, say that the lowest 0.5 meters of Lidar points are the ground.</li>\n<li>Cut off the ground from the point cloud.</li>\n<li>Find connected components.</li>\n<li>Use PCA for each component to find 3 main axes.</li>\n<li>Fit 3D boxes to the connected components, aligning box axis with PCA axis. </li>\n<li>Project boxes on images and use pre-trained on COCO networks to assign class labels to the bounding boxes. CPU should be enough to perform the inference.</li>\n</ol>\n\n<p>Not sure how well will it work, but worth a shot, especially if you do not have GPUs at your disposal :)</p>",
  "messages": [
    {
      "id": "650697",
      "postDate": "10/16/2019 15:51:11",
      "content": "<p>I was thinking about creating and submitting another baseline to the Leaderboard, but it looks like I will not have enough time for this :(</p>\n\n<p>The idea that I had in mind should create a solution that does not require any models to be trained.</p>\n\n<ol>\n<li>Find the ground map. For example, say that the lowest 0.5 meters of Lidar points are the ground.</li>\n<li>Cut off the ground from the point cloud.</li>\n<li>Find connected components.</li>\n<li>Use PCA for each component to find 3 main axes.</li>\n<li>Fit 3D boxes to the connected components, aligning box axis with PCA axis. </li>\n<li>Project boxes on images and use pre-trained on COCO networks to assign class labels to the bounding boxes. CPU should be enough to perform the inference.</li>\n</ol>\n\n<p>Not sure how well will it work, but worth a shot, especially if you do not have GPUs at your disposal :)</p>",
      "rawMarkdown": "I was thinking about creating and submitting another baseline to the Leaderboard, but it looks like I will not have enough time for this :(\n\nThe idea that I had in mind should create a solution that does not require any models to be trained.\n\n1. Find the ground map. For example, say that the lowest 0.5 meters of Lidar points are the ground.\n2. Cut off the ground from the point cloud.\n3. Find connected components.\n4. Use PCA for each component to find 3 main axes.\n5. Fit 3D boxes to the connected components, aligning box axis with PCA axis. \n6. Project boxes on images and use pre-trained on COCO networks to assign class labels to the bounding boxes. CPU should be enough to perform the inference.\n\nNot sure how well will it work, but worth a shot, especially if you do not have GPUs at your disposal :)",
      "votes": null
    },
    {
      "id": "650703",
      "postDate": "10/16/2019 15:59:40",
      "content": "<p>There is another option. Take the network trained on KITTI or nuscences datasets and perform inference on the Lyft dataset :) </p>",
      "rawMarkdown": "There is another option. Take the network trained on KITTI or nuscences datasets and perform inference on the Lyft dataset :)",
      "votes": null
    },
    {
      "id": "650706",
      "postDate": "10/16/2019 16:00:35",
      "content": "<p>Good thinking. When autonomous vehicles will be mainstream in future it is highly likely that they would be relying on smarter yet simple software with limited computing infrastructure on board. I feel future cars with autonomous driving tech should remain cars not become a mobile data center.</p>\n\n<p>In future at some stage we would hit the limit of compute which could be done for autonomous tech like there is a limitation to how much compute can be done on a mobile phone or a laptop.</p>",
      "rawMarkdown": "Good thinking. When autonomous vehicles will be mainstream in future it is highly likely that they would be relying on smarter yet simple software with limited computing infrastructure on board. I feel future cars with autonomous driving tech should remain cars not become a mobile data center.\n\nIn future at some stage we would hit the limit of compute which could be done for autonomous tech like there is a limitation to how much compute can be done on a mobile phone or a laptop.",
      "votes": null
    },
    {
      "id": "650739",
      "postDate": "10/16/2019 16:37:56",
      "content": "<blockquote>\n  <p>Take the network trained on KITTI or nuscences datasets \n  BTW, I only found ones trained on KITTI. There should be trained on nuscenes as well, as they had competition, but I did not find those results and models with weights</p>\n</blockquote>",
      "rawMarkdown": "&gt;  Take the network trained on KITTI or nuscences datasets \nBTW, I only found ones trained on KITTI. There should be trained on nuscenes as well, as they had competition, but I did not find those results and models with weights",
      "votes": null
    },
    {
      "id": "650971",
      "postDate": "10/16/2019 22:17:04",
      "content": "<p>There are a good number of hills in Lyft (Portola / Palo Alto Hills areas maybe?) so perhaps use the MaskRCNN-based lidar segmentation baseline in the Argoverse paper.  I.e. MaskRCNN on images, use mask to segment cloud, use cloud to estimate box &amp; distance.</p>\n\n<p>I agree it would be good to have a DBScan-on-lidar baseline though.  But ground removal in Lyft isn't as easy as it is for much of Kitti.</p>\n\n<p>FWIW the pre-trained MaskRCNN in both Facebook/Detectron and tensorflow/models repos are pretty good for cars and pedestrians and CPU inference isn't too bad for either.  Probably need vision anyways to determine back/front of object, which PCA on lidar is going to confuse very easily.  Could probably train very simple 2-class regression on activations of head network to get good front/back prediction.  </p>",
      "rawMarkdown": "There are a good number of hills in Lyft (Portola / Palo Alto Hills areas maybe?) so perhaps use the MaskRCNN-based lidar segmentation baseline in the Argoverse paper.  I.e. MaskRCNN on images, use mask to segment cloud, use cloud to estimate box &amp; distance.\n\nI agree it would be good to have a DBScan-on-lidar baseline though.  But ground removal in Lyft isn't as easy as it is for much of Kitti.\n\nFWIW the pre-trained MaskRCNN in both Facebook/Detectron and tensorflow/models repos are pretty good for cars and pedestrians and CPU inference isn't too bad for either.  Probably need vision anyways to determine back/front of object, which PCA on lidar is going to confuse very easily.  Could probably train very simple 2-class regression on activations of head network to get good front/back prediction.",
      "votes": null
    },
    {
      "id": "654508",
      "postDate": "10/22/2019 00:32:17",
      "content": "<p>I was also curious to see how this approach performs. \nI tried the PointRCNN trained on Kitti for car detection and used it for inference. It got 0.02 :-D\nPerhaps finetuning it will give better results, but i don't have sufficient gpu resources to give it at try.  At times like this, i miss unlimited kaggle gpus of the old days.</p>",
      "rawMarkdown": "I was also curious to see how this approach performs. \nI tried the PointRCNN trained on Kitti for car detection and used it for inference. It got 0.02 :-D\nPerhaps finetuning it will give better results, but i don't have sufficient gpu resources to give it at try.  At times like this, i miss unlimited kaggle gpus of the old days.",
      "votes": null
    },
    {
      "id": "655877",
      "postDate": "10/23/2019 15:52:20",
      "content": "<p>Is it possible to disclose Lyft's best model or score? Curious to know. Thanks!</p>",
      "rawMarkdown": "Is it possible to disclose Lyft's best model or score? Curious to know. Thanks!",
      "votes": null
    },
    {
      "id": "655931",
      "postDate": "10/23/2019 17:29:28",
      "content": "<p>great idea!</p>",
      "rawMarkdown": "great idea!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 650703,
      "author_name": "iglovikov",
      "author_url": "",
      "post_date": "10/16/2019 15:59:40",
      "content": "<p>There is another option. Take the network trained on KITTI or nuscences datasets and perform inference on the Lyft dataset :) </p>",
      "votes": null,
      "replies": [
        {
          "id": 650739,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "10/16/2019 16:37:56",
          "content": "<blockquote>\n  <p>Take the network trained on KITTI or nuscences datasets \n  BTW, I only found ones trained on KITTI. There should be trained on nuscenes as well, as they had competition, but I did not find those results and models with weights</p>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 654508,
          "author_name": "meaninglesslives",
          "author_url": "",
          "post_date": "10/22/2019 00:32:17",
          "content": "<p>I was also curious to see how this approach performs. \nI tried the PointRCNN trained on Kitti for car detection and used it for inference. It got 0.02 :-D\nPerhaps finetuning it will give better results, but i don't have sufficient gpu resources to give it at try.  At times like this, i miss unlimited kaggle gpus of the old days.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 650706,
      "author_name": "cyberia",
      "author_url": "",
      "post_date": "10/16/2019 16:00:35",
      "content": "<p>Good thinking. When autonomous vehicles will be mainstream in future it is highly likely that they would be relying on smarter yet simple software with limited computing infrastructure on board. I feel future cars with autonomous driving tech should remain cars not become a mobile data center.</p>\n\n<p>In future at some stage we would hit the limit of compute which could be done for autonomous tech like there is a limitation to how much compute can be done on a mobile phone or a laptop.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 650971,
      "author_name": "oarphme",
      "author_url": "",
      "post_date": "10/16/2019 22:17:04",
      "content": "<p>There are a good number of hills in Lyft (Portola / Palo Alto Hills areas maybe?) so perhaps use the MaskRCNN-based lidar segmentation baseline in the Argoverse paper.  I.e. MaskRCNN on images, use mask to segment cloud, use cloud to estimate box &amp; distance.</p>\n\n<p>I agree it would be good to have a DBScan-on-lidar baseline though.  But ground removal in Lyft isn't as easy as it is for much of Kitti.</p>\n\n<p>FWIW the pre-trained MaskRCNN in both Facebook/Detectron and tensorflow/models repos are pretty good for cars and pedestrians and CPU inference isn't too bad for either.  Probably need vision anyways to determine back/front of object, which PCA on lidar is going to confuse very easily.  Could probably train very simple 2-class regression on activations of head network to get good front/back prediction.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 655877,
      "author_name": "wjshenggggg",
      "author_url": "",
      "post_date": "10/23/2019 15:52:20",
      "content": "<p>Is it possible to disclose Lyft's best model or score? Curious to know. Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 655931,
      "author_name": "krishnakatyal",
      "author_url": "",
      "post_date": "10/23/2019 17:29:28",
      "content": "<p>great idea!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "650697": "I was thinking about creating and submitting another baseline to the Leaderboard, but it looks like I will not have enough time for this :(\n\nThe idea that I had in mind should create a solution that does not require any models to be trained.\n\n1. Find the ground map. For example, say that the lowest 0.5 meters of Lidar points are the ground.\n2. Cut off the ground from the point cloud.\n3. Find connected components.\n4. Use PCA for each component to find 3 main axes.\n5. Fit 3D boxes to the connected components, aligning box axis with PCA axis. \n6. Project boxes on images and use pre-trained on COCO networks to assign class labels to the bounding boxes. CPU should be enough to perform the inference.\n\nNot sure how well will it work, but worth a shot, especially if you do not have GPUs at your disposal :)",
    "650703": "There is another option. Take the network trained on KITTI or nuscences datasets and perform inference on the Lyft dataset :)",
    "650706": "Good thinking. When autonomous vehicles will be mainstream in future it is highly likely that they would be relying on smarter yet simple software with limited computing infrastructure on board. I feel future cars with autonomous driving tech should remain cars not become a mobile data center.\n\nIn future at some stage we would hit the limit of compute which could be done for autonomous tech like there is a limitation to how much compute can be done on a mobile phone or a laptop.",
    "650739": "&gt;  Take the network trained on KITTI or nuscences datasets \nBTW, I only found ones trained on KITTI. There should be trained on nuscenes as well, as they had competition, but I did not find those results and models with weights",
    "650971": "There are a good number of hills in Lyft (Portola / Palo Alto Hills areas maybe?) so perhaps use the MaskRCNN-based lidar segmentation baseline in the Argoverse paper.  I.e. MaskRCNN on images, use mask to segment cloud, use cloud to estimate box &amp; distance.\n\nI agree it would be good to have a DBScan-on-lidar baseline though.  But ground removal in Lyft isn't as easy as it is for much of Kitti.\n\nFWIW the pre-trained MaskRCNN in both Facebook/Detectron and tensorflow/models repos are pretty good for cars and pedestrians and CPU inference isn't too bad for either.  Probably need vision anyways to determine back/front of object, which PCA on lidar is going to confuse very easily.  Could probably train very simple 2-class regression on activations of head network to get good front/back prediction.",
    "654508": "I was also curious to see how this approach performs. \nI tried the PointRCNN trained on Kitti for car detection and used it for inference. It got 0.02 :-D\nPerhaps finetuning it will give better results, but i don't have sufficient gpu resources to give it at try.  At times like this, i miss unlimited kaggle gpus of the old days.",
    "655877": "Is it possible to disclose Lyft's best model or score? Curious to know. Thanks!",
    "655931": "great idea!"
  },
  "source": "meta"
}