{
  "id": 117084,
  "title": "What models did you use?",
  "url": "/competitions/3d-object-detection-for-autonomous-vehicles/discussion/117084",
  "author_name": "Artyom Palvelev",
  "post_date": "2019-11-13T09:40:01.379000",
  "votes": 7,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Since there are no published solutions yet, let's discuss what we got here.</p>\n\n<p>According to Git, I started less than 3 weeks ago. I tried to use F-ConvNet and haven't fully made it. This model takes 2D proposals from RGB images and builds some frusta on them; then it finds the best possible box for every proposal. I have generated quite good 2D proposals using Cascade R-CNN from MMDetection, but failed to fix all the problems with F-ConvNet code.</p>\n\n<p>Here are my 2D proposals and my 3D predictions. As you can see, there are some bugs in 3D predictions. This model results in ~0.02 on LB anyway, so I think it has potential.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1859557%2Fde16142fdc98fc1c7c12a0e4d4b964c0%2F00006.jpg?generation=1573637858650833&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1859557%2F516c500f3b56e4ead706b913d0ff4c2c%2F00007.jpg?generation=1573637886063381&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1859557%2Fcef44d946c46fe4610ce65963176ac70%2F00012.jpg?generation=1573637899768471&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1859557%2Fd6e0114dfbb5a0e3c3c562dca82af08a%2F00013.jpg?generation=1573637916011425&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1859557%2Fc8733a96d30d4728ea6bed9e2ed3d33f%2Fcars.png?generation=1573637942013709&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1859557%2F5e449fe6e8687f6b1118eb7e412c327f%2Fcars2.png?generation=1573637952276608&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 672037,
      "postDate": "2019-11-13T13:14:55.820Z",
      "content": "<p>I used second.pytorch, then i generated for each 3d-box the corresponding 2d camera image (crop the 2d-bounding box from the image) like in the lyft-sdk when rendering annotations. Then i fed these images into a classification-CNN  (efficientnet-b4, BCEWithLogitsLoss) which i trained before on the corresponding images from train. The CNN was used to reclassify between [motorcycle , bicycle] ,  [truck, bus , other_vehicle] and to filter the emergency-vehicles.  I split the emergency-vehicle into 3 groups: emergency.police, emergency.ambulance and emergency.fire. Then labeled the emergency-vehicles predictions with high threshold for 3(!) times and used them to bootstrap the CNN to get a decent recognition on emergency-vehicles. I was hoping it would affect the total score more than it did. The CNN was also used to delete some false-positives in pedestrians and bicycles. Lastly i used different mask-rcnn's to find the animals and then i checked the iou of the 2d-mask-rcnn-boxes of the animals (all dogs) and the predictions from different models and epochs to find the best 3d-box. Again i was hoping for more score for the animals. So in the end what worked best for me was silent.pytorch with adjusted parameters from the nuscenes-pointpillars configuration that comes with second.pytorch and the reclassification of the corresponding 2d-box-images through a CNN. The adustments i did to second.pytorch have mainly been reducing the batch-size to 2, widening the area from 50 meter radius to about 100 meter, minimizing the voxel-size to 0.1666 and adjusting the anchors-sizes positions and reconfiguring the resamples for the different classes.</p>",
      "rawMarkdown": "I used second.pytorch, then i generated for each 3d-box the corresponding 2d camera image (crop the 2d-bounding box from the image) like in the lyft-sdk when rendering annotations. Then i fed these images into a classification-CNN  (efficientnet-b4, BCEWithLogitsLoss) which i trained before on the corresponding images from train. The CNN was used to reclassify between [motorcycle , bicycle] ,  [truck, bus , other_vehicle] and to filter the emergency-vehicles.  I split the emergency-vehicle into 3 groups: emergency.police, emergency.ambulance and emergency.fire. Then labeled the emergency-vehicles predictions with high threshold for 3(!) times and used them to bootstrap the CNN to get a decent recognition on emergency-vehicles. I was hoping it would affect the total score more than it did. The CNN was also used to delete some false-positives in pedestrians and bicycles. Lastly i used different mask-rcnn's to find the animals and then i checked the iou of the 2d-mask-rcnn-boxes of the animals (all dogs) and the predictions from different models and epochs to find the best 3d-box. Again i was hoping for more score for the animals. So in the end what worked best for me was silent.pytorch with adjusted parameters from the nuscenes-pointpillars configuration that comes with second.pytorch and the reclassification of the corresponding 2d-box-images through a CNN. The adustments i did to second.pytorch have mainly been reducing the batch-size to 2, widening the area from 50 meter radius to about 100 meter, minimizing the voxel-size to 0.1666 and adjusting the anchors-sizes positions and reconfiguring the resamples for the different classes.",
      "votes": 12,
      "replies": [
        {
          "id": 673398,
          "postDate": "2019-11-14T23:30:17.253Z",
          "content": "<p>That's really cool that you used vision as a second pass!  Did you look into the quality of the yaw / orientation estimates that you got from SECOND?  For many cars it can be very hard to tell front/back in lidar, and hence vision might have helped there.</p>",
          "rawMarkdown": "That's really cool that you used vision as a second pass!  Did you look into the quality of the yaw / orientation estimates that you got from SECOND?  For many cars it can be very hard to tell front/back in lidar, and hence vision might have helped there.",
          "votes": 1
        },
        {
          "id": 673585,
          "postDate": "2019-11-15T07:35:02.017Z",
          "content": "<p>@oarph the vision as second pass worked very good. The orientation is not important for the box, i cannot find the discussion, but for the evaluation i think it makes no difference if the direction is really the front or the back.</p>",
          "rawMarkdown": "@oarph the vision as second pass worked very good. The orientation is not important for the box, i cannot find the discussion, but for the evaluation i think it makes no difference if the direction is really the front or the back."
        }
      ]
    },
    {
      "id": 671896,
      "postDate": "2019-11-13T09:40:01.380Z",
      "content": "<p>Since there are no published solutions yet, let's discuss what we got here.</p>\n\n<p>According to Git, I started less than 3 weeks ago. I tried to use F-ConvNet and haven't fully made it. This model takes 2D proposals from RGB images and builds some frusta on them; then it finds the best possible box for every proposal. I have generated quite good 2D proposals using Cascade R-CNN from MMDetection, but failed to fix all the problems with F-ConvNet code.</p>\n\n<p>Here are my 2D proposals and my 3D predictions. As you can see, there are some bugs in 3D predictions. This model results in ~0.02 on LB anyway, so I think it has potential.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1859557%2Fde16142fdc98fc1c7c12a0e4d4b964c0%2F00006.jpg?generation=1573637858650833&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1859557%2F516c500f3b56e4ead706b913d0ff4c2c%2F00007.jpg?generation=1573637886063381&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1859557%2Fcef44d946c46fe4610ce65963176ac70%2F00012.jpg?generation=1573637899768471&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1859557%2Fd6e0114dfbb5a0e3c3c562dca82af08a%2F00013.jpg?generation=1573637916011425&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1859557%2Fc8733a96d30d4728ea6bed9e2ed3d33f%2Fcars.png?generation=1573637942013709&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1859557%2F5e449fe6e8687f6b1118eb7e412c327f%2Fcars2.png?generation=1573637952276608&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Since there are no published solutions yet, let's discuss what we got here.\n\nAccording to Git, I started less than 3 weeks ago. I tried to use F-ConvNet and haven't fully made it. This model takes 2D proposals from RGB images and builds some frusta on them; then it finds the best possible box for every proposal. I have generated quite good 2D proposals using Cascade R-CNN from MMDetection, but failed to fix all the problems with F-ConvNet code.\n\nHere are my 2D proposals and my 3D predictions. As you can see, there are some bugs in 3D predictions. This model results in ~0.02 on LB anyway, so I think it has potential.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1859557%2Fde16142fdc98fc1c7c12a0e4d4b964c0%2F00006.jpg?generation=1573637858650833&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1859557%2F516c500f3b56e4ead706b913d0ff4c2c%2F00007.jpg?generation=1573637886063381&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1859557%2Fcef44d946c46fe4610ce65963176ac70%2F00012.jpg?generation=1573637899768471&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1859557%2Fd6e0114dfbb5a0e3c3c562dca82af08a%2F00013.jpg?generation=1573637916011425&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1859557%2Fc8733a96d30d4728ea6bed9e2ed3d33f%2Fcars.png?generation=1573637942013709&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1859557%2F5e449fe6e8687f6b1118eb7e412c327f%2Fcars2.png?generation=1573637952276608&amp;alt=media)\n",
      "votes": 7
    },
    {
      "id": 673557,
      "postDate": "2019-11-15T06:25:51.923Z",
      "content": "<p>Our best submission was a EfficientNet-B4 encoder for UNet. We used BEV Image Data to predict 2D proposals.</p>",
      "rawMarkdown": "Our best submission was a EfficientNet-B4 encoder for UNet. We used BEV Image Data to predict 2D proposals.",
      "votes": 3
    },
    {
      "id": 672047,
      "postDate": "2019-11-13T13:20:44.543Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2158018%2F40caa05e579a3be884866c3a33903e74%2Fanimal.png?generation=1573651236326999&amp;alt=media\" alt=\"\"></p>\n\n<p>The green box is the mask-rcnn-box, the blue is the 3d-predicted-box and the red is the 2d-predicted-box.</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2158018%2F40caa05e579a3be884866c3a33903e74%2Fanimal.png?generation=1573651236326999&amp;alt=media)\n\nThe green box is the mask-rcnn-box, the blue is the 3d-predicted-box and the red is the 2d-predicted-box.\n",
      "votes": 3
    },
    {
      "id": 671994,
      "postDate": "2019-11-13T12:37:40.157Z",
      "content": "<p>We used UNet trained at 672x672 BEV as initial proposals + PointNet on cropped pointclouds for a better bbox fit and FP rejections. </p>\n\n<p>In our case visually boxes looked very close as well, however I think that 3D IoU is tricky when it comes to visual inspections - almost imperceptible angle deviations coupled with spatial offsets reduce the score a lot. </p>\n\n<p>We were getting around 0.8 mean IoU for most of the classes, which means we were getting 0s on higher IoU thresholds</p>",
      "rawMarkdown": "We used UNet trained at 672x672 BEV as initial proposals + PointNet on cropped pointclouds for a better bbox fit and FP rejections. \n\nIn our case visually boxes looked very close as well, however I think that 3D IoU is tricky when it comes to visual inspections - almost imperceptible angle deviations coupled with spatial offsets reduce the score a lot. \n\nWe were getting around 0.8 mean IoU for most of the classes, which means we were getting 0s on higher IoU thresholds",
      "votes": 4,
      "replies": [
        {
          "id": 672636,
          "postDate": "2019-11-14T03:12:51.010Z",
          "content": "<p><a href=\"/ruslanm\">@ruslanm</a> Thanks for sharing your ideas. I also tried to improve the resolution of BEV to 672x672 and got my best score(0.053) with segmentation. But cropping point clouds never came to my mind, which is really a loss. I wasted tons of time in debugging 3D object detection models, ComplexYOLO, SECOND, PointRCNN, but all failed to improve the score further. </p>",
          "rawMarkdown": "@ruslanm Thanks for sharing your ideas. I also tried to improve the resolution of BEV to 672x672 and got my best score(0.053) with segmentation. But cropping point clouds never came to my mind, which is really a loss. I wasted tons of time in debugging 3D object detection models, ComplexYOLO, SECOND, PointRCNN, but all failed to improve the score further. "
        }
      ]
    },
    {
      "id": 671905,
      "postDate": "2019-11-13T09:50:31.200Z",
      "content": "<p>wow, those 2d boxes look pretty great.  the 3d boxes don't look bad either, I'm surprised it only scores 0.02.  I'm really curious about the timestamps and sync of the labels in the test set.  if there's bad sync, I think you can easily have predictions that look good but otherwise don't have an IoU of at least 0.5 with the labels.  there's also of course the issue with them labeling parked cars inconsistently, which can be a big deal for scenes next to parking lots</p>",
      "rawMarkdown": "wow, those 2d boxes look pretty great.  the 3d boxes don't look bad either, I'm surprised it only scores 0.02.  I'm really curious about the timestamps and sync of the labels in the test set.  if there's bad sync, I think you can easily have predictions that look good but otherwise don't have an IoU of at least 0.5 with the labels.  there's also of course the issue with them labeling parked cars inconsistently, which can be a big deal for scenes next to parking lots",
      "votes": 2,
      "replies": [
        {
          "id": 671909,
          "postDate": "2019-11-13T10:00:31.793Z",
          "content": "<p>Cheers! 2D proposals are generated after 3 epochs, training took about 24 hours.</p>\n\n<p>I think we can mask parking lots by map. There is a function called <code>is_on_mask</code> in Lyft SDK, but it's broken. I had fun yesterday fixing it two hours before the deadline :)</p>",
          "rawMarkdown": "Cheers! 2D proposals are generated after 3 epochs, training took about 24 hours.\n\nI think we can mask parking lots by map. There is a function called `is_on_mask` in Lyft SDK, but it's broken. I had fun yesterday fixing it two hours before the deadline :)",
          "votes": 1
        },
        {
          "id": 671918,
          "postDate": "2019-11-13T10:14:27.680Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 673394,
          "postDate": "2019-11-14T23:20:53.520Z",
          "content": "<p><a href=\"/artyomp\">@artyomp</a> quick question: do you know what method you were using to project predictions into the world frame for the purpose of evaluation?  I'm curious because I think the Kaggle server might use the lidar / label timestamps, which would have very different ego poses than the cameras (since the cameras can be up to 100ms off of lidar).  Thus for any moving object (cars, pedestrians), the camera-&gt;world transform would be probably map predicted boxes to very different positions versus labels.  Probably at least a 0.5 IoU difference.</p>",
          "rawMarkdown": "@artyomp quick question: do you know what method you were using to project predictions into the world frame for the purpose of evaluation?  I'm curious because I think the Kaggle server might use the lidar / label timestamps, which would have very different ego poses than the cameras (since the cameras can be up to 100ms off of lidar).  Thus for any moving object (cars, pedestrians), the camera-&gt;world transform would be probably map predicted boxes to very different positions versus labels.  Probably at least a 0.5 IoU difference."
        }
      ]
    },
    {
      "id": 672481,
      "postDate": "2019-11-13T23:57:02.107Z",
      "content": "<p>would be great if anyone can publish their models in notebooks or github!</p>\n\n<p>i tried the second.pytorch codes but the training time was too long for my resource and time..\nUNet+PointNet is interesting and sounds like a resource-efficient solution too.</p>",
      "rawMarkdown": "would be great if anyone can publish their models in notebooks or github!\n\ni tried the second.pytorch codes but the training time was too long for my resource and time..\nUNet+PointNet is interesting and sounds like a resource-efficient solution too."
    },
    {
      "id": 671916,
      "postDate": "2019-11-13T10:13:01.207Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 672037,
      "author_name": "Marek Wyborski",
      "author_url": "",
      "post_date": "2019-11-13T13:14:55.820000",
      "content": "<p>I used second.pytorch, then i generated for each 3d-box the corresponding 2d camera image (crop the 2d-bounding box from the image) like in the lyft-sdk when rendering annotations. Then i fed these images into a classification-CNN  (efficientnet-b4, BCEWithLogitsLoss) which i trained before on the corresponding images from train. The CNN was used to reclassify between [motorcycle , bicycle] ,  [truck, bus , other_vehicle] and to filter the emergency-vehicles.  I split the emergency-vehicle into 3 groups: emergency.police, emergency.ambulance and emergency.fire. Then labeled the emergency-vehicles predictions with high threshold for 3(!) times and used them to bootstrap the CNN to get a decent recognition on emergency-vehicles. I was hoping it would affect the total score more than it did. The CNN was also used to delete some false-positives in pedestrians and bicycles. Lastly i used different mask-rcnn's to find the animals and then i checked the iou of the 2d-mask-rcnn-boxes of the animals (all dogs) and the predictions from different models and epochs to find the best 3d-box. Again i was hoping for more score for the animals. So in the end what worked best for me was silent.pytorch with adjusted parameters from the nuscenes-pointpillars configuration that comes with second.pytorch and the reclassification of the corresponding 2d-box-images through a CNN. The adustments i did to second.pytorch have mainly been reducing the batch-size to 2, widening the area from 50 meter radius to about 100 meter, minimizing the voxel-size to 0.1666 and adjusting the anchors-sizes positions and reconfiguring the resamples for the different classes.</p>",
      "votes": 12,
      "replies": [
        {
          "id": 673398,
          "author_name": "oarph",
          "author_url": "",
          "post_date": "2019-11-14T23:30:17.253000",
          "content": "<p>That's really cool that you used vision as a second pass!  Did you look into the quality of the yaw / orientation estimates that you got from SECOND?  For many cars it can be very hard to tell front/back in lidar, and hence vision might have helped there.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 673585,
          "author_name": "Marek Wyborski",
          "author_url": "",
          "post_date": "2019-11-15T07:35:02.017000",
          "content": "<p>@oarph the vision as second pass worked very good. The orientation is not important for the box, i cannot find the discussion, but for the evaluation i think it makes no difference if the direction is really the front or the back.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 673557,
      "author_name": "Ram Ramrakhya",
      "author_url": "",
      "post_date": "2019-11-15T06:25:51.923000",
      "content": "<p>Our best submission was a EfficientNet-B4 encoder for UNet. We used BEV Image Data to predict 2D proposals.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 672047,
      "author_name": "Marek Wyborski",
      "author_url": "",
      "post_date": "2019-11-13T13:20:44.543000",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2158018%2F40caa05e579a3be884866c3a33903e74%2Fanimal.png?generation=1573651236326999&amp;alt=media\" alt=\"\"></p>\n\n<p>The green box is the mask-rcnn-box, the blue is the 3d-predicted-box and the red is the 2d-predicted-box.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 671994,
      "author_name": "Ruslan Mustafin",
      "author_url": "",
      "post_date": "2019-11-13T12:37:40.157000",
      "content": "<p>We used UNet trained at 672x672 BEV as initial proposals + PointNet on cropped pointclouds for a better bbox fit and FP rejections. </p>\n\n<p>In our case visually boxes looked very close as well, however I think that 3D IoU is tricky when it comes to visual inspections - almost imperceptible angle deviations coupled with spatial offsets reduce the score a lot. </p>\n\n<p>We were getting around 0.8 mean IoU for most of the classes, which means we were getting 0s on higher IoU thresholds</p>",
      "votes": 4,
      "replies": [
        {
          "id": 672636,
          "author_name": "Bob Wang",
          "author_url": "",
          "post_date": "2019-11-14T03:12:51.010000",
          "content": "<p><a href=\"/ruslanm\">@ruslanm</a> Thanks for sharing your ideas. I also tried to improve the resolution of BEV to 672x672 and got my best score(0.053) with segmentation. But cropping point clouds never came to my mind, which is really a loss. I wasted tons of time in debugging 3D object detection models, ComplexYOLO, SECOND, PointRCNN, but all failed to improve the score further. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 671905,
      "author_name": "oarph",
      "author_url": "",
      "post_date": "2019-11-13T09:50:31.200000",
      "content": "<p>wow, those 2d boxes look pretty great.  the 3d boxes don't look bad either, I'm surprised it only scores 0.02.  I'm really curious about the timestamps and sync of the labels in the test set.  if there's bad sync, I think you can easily have predictions that look good but otherwise don't have an IoU of at least 0.5 with the labels.  there's also of course the issue with them labeling parked cars inconsistently, which can be a big deal for scenes next to parking lots</p>",
      "votes": 2,
      "replies": [
        {
          "id": 671909,
          "author_name": "Artyom Palvelev",
          "author_url": "",
          "post_date": "2019-11-13T10:00:31.793000",
          "content": "<p>Cheers! 2D proposals are generated after 3 epochs, training took about 24 hours.</p>\n\n<p>I think we can mask parking lots by map. There is a function called <code>is_on_mask</code> in Lyft SDK, but it's broken. I had fun yesterday fixing it two hours before the deadline :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 671918,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-13T10:14:27.680000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 673394,
          "author_name": "oarph",
          "author_url": "",
          "post_date": "2019-11-14T23:20:53.520000",
          "content": "<p><a href=\"/artyomp\">@artyomp</a> quick question: do you know what method you were using to project predictions into the world frame for the purpose of evaluation?  I'm curious because I think the Kaggle server might use the lidar / label timestamps, which would have very different ego poses than the cameras (since the cameras can be up to 100ms off of lidar).  Thus for any moving object (cars, pedestrians), the camera-&gt;world transform would be probably map predicted boxes to very different positions versus labels.  Probably at least a 0.5 IoU difference.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 672481,
      "author_name": "arutema47",
      "author_url": "",
      "post_date": "2019-11-13T23:57:02.107000",
      "content": "<p>would be great if anyone can publish their models in notebooks or github!</p>\n\n<p>i tried the second.pytorch codes but the training time was too long for my resource and time..\nUNet+PointNet is interesting and sounds like a resource-efficient solution too.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 671916,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-13T10:13:01.207000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "672037": "I used second.pytorch, then i generated for each 3d-box the corresponding 2d camera image (crop the 2d-bounding box from the image) like in the lyft-sdk when rendering annotations. Then i fed these images into a classification-CNN  (efficientnet-b4, BCEWithLogitsLoss) which i trained before on the corresponding images from train. The CNN was used to reclassify between [motorcycle , bicycle] ,  [truck, bus , other_vehicle] and to filter the emergency-vehicles.  I split the emergency-vehicle into 3 groups: emergency.police, emergency.ambulance and emergency.fire. Then labeled the emergency-vehicles predictions with high threshold for 3(!) times and used them to bootstrap the CNN to get a decent recognition on emergency-vehicles. I was hoping it would affect the total score more than it did. The CNN was also used to delete some false-positives in pedestrians and bicycles. Lastly i used different mask-rcnn's to find the animals and then i checked the iou of the 2d-mask-rcnn-boxes of the animals (all dogs) and the predictions from different models and epochs to find the best 3d-box. Again i was hoping for more score for the animals. So in the end what worked best for me was silent.pytorch with adjusted parameters from the nuscenes-pointpillars configuration that comes with second.pytorch and the reclassification of the corresponding 2d-box-images through a CNN. The adustments i did to second.pytorch have mainly been reducing the batch-size to 2, widening the area from 50 meter radius to about 100 meter, minimizing the voxel-size to 0.1666 and adjusting the anchors-sizes positions and reconfiguring the resamples for the different classes.",
    "671896": "Since there are no published solutions yet, let's discuss what we got here.\n\nAccording to Git, I started less than 3 weeks ago. I tried to use F-ConvNet and haven't fully made it. This model takes 2D proposals from RGB images and builds some frusta on them; then it finds the best possible box for every proposal. I have generated quite good 2D proposals using Cascade R-CNN from MMDetection, but failed to fix all the problems with F-ConvNet code.\n\nHere are my 2D proposals and my 3D predictions. As you can see, there are some bugs in 3D predictions. This model results in ~0.02 on LB anyway, so I think it has potential.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1859557%2Fde16142fdc98fc1c7c12a0e4d4b964c0%2F00006.jpg?generation=1573637858650833&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1859557%2F516c500f3b56e4ead706b913d0ff4c2c%2F00007.jpg?generation=1573637886063381&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1859557%2Fcef44d946c46fe4610ce65963176ac70%2F00012.jpg?generation=1573637899768471&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1859557%2Fd6e0114dfbb5a0e3c3c562dca82af08a%2F00013.jpg?generation=1573637916011425&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1859557%2Fc8733a96d30d4728ea6bed9e2ed3d33f%2Fcars.png?generation=1573637942013709&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1859557%2F5e449fe6e8687f6b1118eb7e412c327f%2Fcars2.png?generation=1573637952276608&amp;alt=media)\n",
    "673557": "Our best submission was a EfficientNet-B4 encoder for UNet. We used BEV Image Data to predict 2D proposals.",
    "672047": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2158018%2F40caa05e579a3be884866c3a33903e74%2Fanimal.png?generation=1573651236326999&amp;alt=media)\n\nThe green box is the mask-rcnn-box, the blue is the 3d-predicted-box and the red is the 2d-predicted-box.\n",
    "671994": "We used UNet trained at 672x672 BEV as initial proposals + PointNet on cropped pointclouds for a better bbox fit and FP rejections. \n\nIn our case visually boxes looked very close as well, however I think that 3D IoU is tricky when it comes to visual inspections - almost imperceptible angle deviations coupled with spatial offsets reduce the score a lot. \n\nWe were getting around 0.8 mean IoU for most of the classes, which means we were getting 0s on higher IoU thresholds",
    "671905": "wow, those 2d boxes look pretty great.  the 3d boxes don't look bad either, I'm surprised it only scores 0.02.  I'm really curious about the timestamps and sync of the labels in the test set.  if there's bad sync, I think you can easily have predictions that look good but otherwise don't have an IoU of at least 0.5 with the labels.  there's also of course the issue with them labeling parked cars inconsistently, which can be a big deal for scenes next to parking lots",
    "672481": "would be great if anyone can publish their models in notebooks or github!\n\ni tried the second.pytorch codes but the training time was too long for my resource and time..\nUNet+PointNet is interesting and sounds like a resource-efficient solution too.",
    "671916": ""
  }
}