{
  "id": 127145,
  "title": "(Another part of) 5th place solution",
  "url": "/competitions/pku-autonomous-driving/discussion/127145",
  "author_name": "",
  "post_date": "2020-01-22T16:52:58.119841700Z",
  "votes": 18,
  "comment_count": 6,
  "views": 0,
  "content": "<p>This is a summary of me and <a href=\"/erniechiew\">@erniechiew</a>'s approach as part of our 5th place solution. For the other part, please see [https://www.kaggle.com/c/pku-autonomous-driving/discussion/127065].</p>\n\n<h2>Quick Overview</h2>\n\n<p>Our approach is based on 2D bounding boxes: a 2D object detector (Faster R-CNN) is fine-tuned on this dataset to detect the neighboring cars. A separate network then regresses the 6D position of each car based on raw features from the bounding boxes, as well as image features from the bounding box crops.</p>\n\n<h3>Our Solution in More Detail</h3>\n\n<p>There are two key components to our solution:\n1. 2D bounding boxes\n2. 6D-pose regression</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2202881%2F0081962fbbf790fdcea88dcb36e6649f%2FPKU_car_kaggle%20(5\" alt=\"\">.png?generation=1579744111028907&amp;alt=media)</p>\n\n<h2>1. 2D Bounding Boxes</h2>\n\n<p>The “backbone” of our approach is 2D bounding boxes. Because the training data does not provide labeled bounding boxes, we first modify this kernel [https://www.kaggle.com/hypocrites/simple-eda-with-imageai-object-detection] to obtain bounding boxes for the training set. The key change we made to the kernel is that we used the provided 3D car models to obtain car dimensions, and adjusted our bounding boxes based on the car dimensions for a given car label (whereas the kernel uses a generic car dimension for all cars). We call the boxes obtained from this method as “ground truth boxes”.</p>\n\n<p>For the test set, we cannot obtain bounding boxes with this method, since the method requires the very data points we are trying to predict! We therefore resorted to 2D object detectors. First, we fine-tuned Faster-RCNN (from maskrcnn-benchmark) on the BDD100k dataset. The BDD100k tuned model is then further fine-tuned on our “ground truth boxes”. </p>\n\n<p>We found that training with larger input images resulted in noticeably better validation MAP and LB score. In the end, we trained 3 separate detector backbones with the following configurations:\n1. X-101-32x8d-FPN (input size 1373x1100)\n2. R-101-FPN (input size 2000x1602)\n3. R-50-FPN (input size 2499x2002)</p>\n\n<p>The 3 models are then used to predict bounding boxes for cars in the test images. Within each model, overlapping boxes are discarded based on an NMS threshold of 0.70, and any remaining boxes below 0.70 confidence are further discarded.</p>\n\n<p>The boxes from the three models are then ensembled using weighted boxes fusion (<a href=\"https://github.com/ZFTurbo/Weighted-Boxes-Fusion\">https://github.com/ZFTurbo/Weighted-Boxes-Fusion</a>) with equal weights. We found that by including only the boxes that are successfully fused/merged by all three models, our score improved. We call the boxes predicted by the object detectors as “predicted boxes” (as opposed to ground truth boxes).</p>\n\n<h2>2. 6D-pose regression</h2>\n\n<p>With the ground truth boxes and predicted boxes in hand, we then regress the cars’ positions using a downstream network that we simply call Q-Net (Q for Quaternion).</p>\n\n<h3>Input:</h3>\n\n<p>Q-Net consumes the following two inputs:</p>\n\n<ol>\n<li>Car bounding box crops (RGB images, resized to 128x128)</li>\n<li>10 numerical features from the bounding boxes\n<ul><li>Box width</li>\n<li>Box height</li>\n<li>Box width to height ratio</li>\n<li>Box area</li>\n<li>X-coordinate of the box center</li>\n<li>Y-coordinate of the box center</li>\n<li>“2D-distance” of box center from camera </li>\n<li>“2D-angle” between box center and camera</li>\n<li>Whether the box is close to the left boundary of the image</li>\n<li>Whether the box is close to the right boundary of the image</li></ul></li>\n</ol>\n\n<p>Both inputs are standardized appropriately.</p>\n\n<h3>Output:</h3>\n\n<p>Car translation (x, y, z)\nCar rotation (quaternion)</p>\n\n<h2>Brief Network Architecture</h2>\n\n<p>The bounding box crops are fed into a pre-trained DenseNet121 model from torchvision, while the numerical features are connected to FC layers. Both of these “input paths” are then connected to 2 “output paths”, one for predicting car rotations, and another for car translations.  </p>\n\n<p>The intuition here is that both bounding-box crops and numerical features work hand-in-hand in regressing the 6D-pose, and so information from both input components should be “communicated” or “shared” with both output paths.</p>\n\n<h2>Some Training Details</h2>\n\n<p>We trained Q-Net on the ground truth boxes, and made predictions on the predicted boxes. Initially, we thought that training with out-of-fold predicted boxes would yield better test set generalization, but this was not the case. This is most likely because in the latter case, there was a challenge to map the predicted boxes with the corresponding ground truth pose information for a given car.</p>\n\n<p>For translation we used L1-loss. For rotations, as mentioned, we converted the angles into quaternions, and used Dot Product Loss (<a href=\"https://arxiv.org/pdf/1901.09366.pdf\">https://arxiv.org/pdf/1901.09366.pdf</a>).</p>\n\n<p>We trained Q-Net on 10-folds and averaged the predictions from each fold. For translation, we simply averaged the predicted x, y, and z values respectively. For rotation, we averaged the quaternion predictions using (<a href=\"https://github.com/christophhagen/averaging-quaternions\">https://github.com/christophhagen/averaging-quaternions</a>) before converting back to Euler angles.</p>\n\n<h2>Score Summary</h2>\n\n<p>Our best model attains 0.124 on public LB and 0.116 on private LB.</p>\n\n<h2>Final Ensemble</h2>\n\n<p>Thanks to our team member <a href=\"/uiiurz1\">@uiiurz1</a>'s Centernet model, we had relatively diverse models to ensemble. However, ensembling the two models of different nature was slightly challenging. We used weighted points fusion (weighted boxes fusion but with IOU replaced with 3D-distance) to ensemble our predictions, credit to <a href=\"/uiiurz1\">@uiiurz1</a>. By ensembling, we obtained a decent boost in score: 0.131 public LB, 0.123 private LB.</p>\n\n<h2>Closing Remarks</h2>\n\n<p>We would like to congratulate all the winners, and really anyone who benefited in some way from the competition. We want to thank our teammate <a href=\"/uiiurz1\">@uiiurz1</a> for his great insights and collaborative spirit. Finally, we would also like to thank the competition host(s) and Kaggle for organizing the competition. Putting aside issues regarding the unknown competition metric and labelling methodology for cropped cars, this competition was intriguing enough to keep our minds occupied thinking about car position estimation during our drives home from work the past couple of months :)</p>",
  "messages": [
    {
      "id": "725943",
      "postDate": "01/22/2020 16:52:58",
      "content": "<p>This is a summary of me and <a href=\"/erniechiew\">@erniechiew</a>'s approach as part of our 5th place solution. For the other part, please see [https://www.kaggle.com/c/pku-autonomous-driving/discussion/127065].</p>\n\n<h2>Quick Overview</h2>\n\n<p>Our approach is based on 2D bounding boxes: a 2D object detector (Faster R-CNN) is fine-tuned on this dataset to detect the neighboring cars. A separate network then regresses the 6D position of each car based on raw features from the bounding boxes, as well as image features from the bounding box crops.</p>\n\n<h3>Our Solution in More Detail</h3>\n\n<p>There are two key components to our solution:\n1. 2D bounding boxes\n2. 6D-pose regression</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2202881%2F0081962fbbf790fdcea88dcb36e6649f%2FPKU_car_kaggle%20(5\" alt=\"\">.png?generation=1579744111028907&amp;alt=media)</p>\n\n<h2>1. 2D Bounding Boxes</h2>\n\n<p>The “backbone” of our approach is 2D bounding boxes. Because the training data does not provide labeled bounding boxes, we first modify this kernel [https://www.kaggle.com/hypocrites/simple-eda-with-imageai-object-detection] to obtain bounding boxes for the training set. The key change we made to the kernel is that we used the provided 3D car models to obtain car dimensions, and adjusted our bounding boxes based on the car dimensions for a given car label (whereas the kernel uses a generic car dimension for all cars). We call the boxes obtained from this method as “ground truth boxes”.</p>\n\n<p>For the test set, we cannot obtain bounding boxes with this method, since the method requires the very data points we are trying to predict! We therefore resorted to 2D object detectors. First, we fine-tuned Faster-RCNN (from maskrcnn-benchmark) on the BDD100k dataset. The BDD100k tuned model is then further fine-tuned on our “ground truth boxes”. </p>\n\n<p>We found that training with larger input images resulted in noticeably better validation MAP and LB score. In the end, we trained 3 separate detector backbones with the following configurations:\n1. X-101-32x8d-FPN (input size 1373x1100)\n2. R-101-FPN (input size 2000x1602)\n3. R-50-FPN (input size 2499x2002)</p>\n\n<p>The 3 models are then used to predict bounding boxes for cars in the test images. Within each model, overlapping boxes are discarded based on an NMS threshold of 0.70, and any remaining boxes below 0.70 confidence are further discarded.</p>\n\n<p>The boxes from the three models are then ensembled using weighted boxes fusion (<a href=\"https://github.com/ZFTurbo/Weighted-Boxes-Fusion\">https://github.com/ZFTurbo/Weighted-Boxes-Fusion</a>) with equal weights. We found that by including only the boxes that are successfully fused/merged by all three models, our score improved. We call the boxes predicted by the object detectors as “predicted boxes” (as opposed to ground truth boxes).</p>\n\n<h2>2. 6D-pose regression</h2>\n\n<p>With the ground truth boxes and predicted boxes in hand, we then regress the cars’ positions using a downstream network that we simply call Q-Net (Q for Quaternion).</p>\n\n<h3>Input:</h3>\n\n<p>Q-Net consumes the following two inputs:</p>\n\n<ol>\n<li>Car bounding box crops (RGB images, resized to 128x128)</li>\n<li>10 numerical features from the bounding boxes\n<ul><li>Box width</li>\n<li>Box height</li>\n<li>Box width to height ratio</li>\n<li>Box area</li>\n<li>X-coordinate of the box center</li>\n<li>Y-coordinate of the box center</li>\n<li>“2D-distance” of box center from camera </li>\n<li>“2D-angle” between box center and camera</li>\n<li>Whether the box is close to the left boundary of the image</li>\n<li>Whether the box is close to the right boundary of the image</li></ul></li>\n</ol>\n\n<p>Both inputs are standardized appropriately.</p>\n\n<h3>Output:</h3>\n\n<p>Car translation (x, y, z)\nCar rotation (quaternion)</p>\n\n<h2>Brief Network Architecture</h2>\n\n<p>The bounding box crops are fed into a pre-trained DenseNet121 model from torchvision, while the numerical features are connected to FC layers. Both of these “input paths” are then connected to 2 “output paths”, one for predicting car rotations, and another for car translations.  </p>\n\n<p>The intuition here is that both bounding-box crops and numerical features work hand-in-hand in regressing the 6D-pose, and so information from both input components should be “communicated” or “shared” with both output paths.</p>\n\n<h2>Some Training Details</h2>\n\n<p>We trained Q-Net on the ground truth boxes, and made predictions on the predicted boxes. Initially, we thought that training with out-of-fold predicted boxes would yield better test set generalization, but this was not the case. This is most likely because in the latter case, there was a challenge to map the predicted boxes with the corresponding ground truth pose information for a given car.</p>\n\n<p>For translation we used L1-loss. For rotations, as mentioned, we converted the angles into quaternions, and used Dot Product Loss (<a href=\"https://arxiv.org/pdf/1901.09366.pdf\">https://arxiv.org/pdf/1901.09366.pdf</a>).</p>\n\n<p>We trained Q-Net on 10-folds and averaged the predictions from each fold. For translation, we simply averaged the predicted x, y, and z values respectively. For rotation, we averaged the quaternion predictions using (<a href=\"https://github.com/christophhagen/averaging-quaternions\">https://github.com/christophhagen/averaging-quaternions</a>) before converting back to Euler angles.</p>\n\n<h2>Score Summary</h2>\n\n<p>Our best model attains 0.124 on public LB and 0.116 on private LB.</p>\n\n<h2>Final Ensemble</h2>\n\n<p>Thanks to our team member <a href=\"/uiiurz1\">@uiiurz1</a>'s Centernet model, we had relatively diverse models to ensemble. However, ensembling the two models of different nature was slightly challenging. We used weighted points fusion (weighted boxes fusion but with IOU replaced with 3D-distance) to ensemble our predictions, credit to <a href=\"/uiiurz1\">@uiiurz1</a>. By ensembling, we obtained a decent boost in score: 0.131 public LB, 0.123 private LB.</p>\n\n<h2>Closing Remarks</h2>\n\n<p>We would like to congratulate all the winners, and really anyone who benefited in some way from the competition. We want to thank our teammate <a href=\"/uiiurz1\">@uiiurz1</a> for his great insights and collaborative spirit. Finally, we would also like to thank the competition host(s) and Kaggle for organizing the competition. Putting aside issues regarding the unknown competition metric and labelling methodology for cropped cars, this competition was intriguing enough to keep our minds occupied thinking about car position estimation during our drives home from work the past couple of months :)</p>",
      "rawMarkdown": "This is a summary of me and @erniechiew's approach as part of our 5th place solution. For the other part, please see [https://www.kaggle.com/c/pku-autonomous-driving/discussion/127065].\n\n## Quick Overview\n\nOur approach is based on 2D bounding boxes: a 2D object detector (Faster R-CNN) is fine-tuned on this dataset to detect the neighboring cars. A separate network then regresses the 6D position of each car based on raw features from the bounding boxes, as well as image features from the bounding box crops.\n\n\n### Our Solution in More Detail\n\nThere are two key components to our solution:\n1. 2D bounding boxes\n2. 6D-pose regression\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2202881%2F0081962fbbf790fdcea88dcb36e6649f%2FPKU_car_kaggle%20(5).png?generation=1579744111028907&amp;alt=media)\n\n\n\n\n## 1. 2D Bounding Boxes\n\n\nThe “backbone” of our approach is 2D bounding boxes. Because the training data does not provide labeled bounding boxes, we first modify this kernel [https://www.kaggle.com/hypocrites/simple-eda-with-imageai-object-detection] to obtain bounding boxes for the training set. The key change we made to the kernel is that we used the provided 3D car models to obtain car dimensions, and adjusted our bounding boxes based on the car dimensions for a given car label (whereas the kernel uses a generic car dimension for all cars). We call the boxes obtained from this method as “ground truth boxes”.\n\nFor the test set, we cannot obtain bounding boxes with this method, since the method requires the very data points we are trying to predict! We therefore resorted to 2D object detectors. First, we fine-tuned Faster-RCNN (from maskrcnn-benchmark) on the BDD100k dataset. The BDD100k tuned model is then further fine-tuned on our “ground truth boxes”. \n\n\nWe found that training with larger input images resulted in noticeably better validation MAP and LB score. In the end, we trained 3 separate detector backbones with the following configurations:\n1. X-101-32x8d-FPN (input size 1373x1100)\n2. R-101-FPN (input size 2000x1602)\n3. R-50-FPN (input size 2499x2002)\n\n\n\nThe 3 models are then used to predict bounding boxes for cars in the test images. Within each model, overlapping boxes are discarded based on an NMS threshold of 0.70, and any remaining boxes below 0.70 confidence are further discarded.\n\nThe boxes from the three models are then ensembled using weighted boxes fusion (https://github.com/ZFTurbo/Weighted-Boxes-Fusion) with equal weights. We found that by including only the boxes that are successfully fused/merged by all three models, our score improved. We call the boxes predicted by the object detectors as “predicted boxes” (as opposed to ground truth boxes).\n\n\n## 2. 6D-pose regression\n\nWith the ground truth boxes and predicted boxes in hand, we then regress the cars’ positions using a downstream network that we simply call Q-Net (Q for Quaternion).\n\n### Input:\n\nQ-Net consumes the following two inputs:\n\n1. Car bounding box crops (RGB images, resized to 128x128)\n2. 10 numerical features from the bounding boxes\n- Box width\n- Box height\n- Box width to height ratio\n- Box area\n- X-coordinate of the box center\n- Y-coordinate of the box center\n- “2D-distance” of box center from camera \n- “2D-angle” between box center and camera\n- Whether the box is close to the left boundary of the image\n- Whether the box is close to the right boundary of the image\n\n\nBoth inputs are standardized appropriately.\n\n\n### Output:\n\nCar translation (x, y, z)\nCar rotation (quaternion)\n\n\n\n## Brief Network Architecture\n\nThe bounding box crops are fed into a pre-trained DenseNet121 model from torchvision, while the numerical features are connected to FC layers. Both of these “input paths” are then connected to 2 “output paths”, one for predicting car rotations, and another for car translations.  \n\nThe intuition here is that both bounding-box crops and numerical features work hand-in-hand in regressing the 6D-pose, and so information from both input components should be “communicated” or “shared” with both output paths.\n\n\n\n## Some Training Details\n\nWe trained Q-Net on the ground truth boxes, and made predictions on the predicted boxes. Initially, we thought that training with out-of-fold predicted boxes would yield better test set generalization, but this was not the case. This is most likely because in the latter case, there was a challenge to map the predicted boxes with the corresponding ground truth pose information for a given car.\n\n\nFor translation we used L1-loss. For rotations, as mentioned, we converted the angles into quaternions, and used Dot Product Loss (https://arxiv.org/pdf/1901.09366.pdf).\n\nWe trained Q-Net on 10-folds and averaged the predictions from each fold. For translation, we simply averaged the predicted x, y, and z values respectively. For rotation, we averaged the quaternion predictions using (https://github.com/christophhagen/averaging-quaternions) before converting back to Euler angles.\n\n\n\n## Score Summary\n\nOur best model attains 0.124 on public LB and 0.116 on private LB.\n\n\n## Final Ensemble\n\nThanks to our team member @uiiurz1's Centernet model, we had relatively diverse models to ensemble. However, ensembling the two models of different nature was slightly challenging. We used weighted points fusion (weighted boxes fusion but with IOU replaced with 3D-distance) to ensemble our predictions, credit to @uiiurz1. By ensembling, we obtained a decent boost in score: 0.131 public LB, 0.123 private LB.\n\n\n## Closing Remarks\n\nWe would like to congratulate all the winners, and really anyone who benefited in some way from the competition. We want to thank our teammate @uiiurz1 for his great insights and collaborative spirit. Finally, we would also like to thank the competition host(s) and Kaggle for organizing the competition. Putting aside issues regarding the unknown competition metric and labelling methodology for cropped cars, this competition was intriguing enough to keep our minds occupied thinking about car position estimation during our drives home from work the past couple of months :)",
      "votes": null
    },
    {
      "id": "725988",
      "postDate": "01/22/2020 17:34:42",
      "content": "<p>Nice Write-Up\nThanks for Sharing your Approach!! <a href=\"/css919\">@css919</a> </p>",
      "rawMarkdown": "Nice Write-Up\nThanks for Sharing your Approach!! @css919",
      "votes": null
    },
    {
      "id": "726388",
      "postDate": "01/23/2020 01:44:20",
      "content": "<p>You're welcome, thanks for reading too.</p>",
      "rawMarkdown": "You're welcome, thanks for reading too.",
      "votes": null
    },
    {
      "id": "727068",
      "postDate": "01/23/2020 12:37:41",
      "content": "<p>Thanks for haring,great job!</p>",
      "rawMarkdown": "Thanks for haring,great job!",
      "votes": null
    },
    {
      "id": "727090",
      "postDate": "01/23/2020 13:00:21",
      "content": "<p>Thanks for reading! </p>",
      "rawMarkdown": "Thanks for reading!",
      "votes": null
    },
    {
      "id": "732536",
      "postDate": "01/29/2020 23:21:47",
      "content": "<p>Congrats, thanks for sharing.</p>",
      "rawMarkdown": "Congrats, thanks for sharing.",
      "votes": null
    },
    {
      "id": "734855",
      "postDate": "02/02/2020 05:12:35",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 725988,
      "author_name": "veeralakrishna",
      "author_url": "",
      "post_date": "01/22/2020 17:34:42",
      "content": "<p>Nice Write-Up\nThanks for Sharing your Approach!! <a href=\"/css919\">@css919</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 726388,
          "author_name": "css919",
          "author_url": "",
          "post_date": "01/23/2020 01:44:20",
          "content": "<p>You're welcome, thanks for reading too.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 727068,
      "author_name": "",
      "author_url": "",
      "post_date": "01/23/2020 12:37:41",
      "content": "<p>Thanks for haring,great job!</p>",
      "votes": null,
      "replies": [
        {
          "id": 727090,
          "author_name": "css919",
          "author_url": "",
          "post_date": "01/23/2020 13:00:21",
          "content": "<p>Thanks for reading! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 732536,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "01/29/2020 23:21:47",
      "content": "<p>Congrats, thanks for sharing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 734855,
          "author_name": "css919",
          "author_url": "",
          "post_date": "02/02/2020 05:12:35",
          "content": "<p>Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "725943": "This is a summary of me and @erniechiew's approach as part of our 5th place solution. For the other part, please see [https://www.kaggle.com/c/pku-autonomous-driving/discussion/127065].\n\n## Quick Overview\n\nOur approach is based on 2D bounding boxes: a 2D object detector (Faster R-CNN) is fine-tuned on this dataset to detect the neighboring cars. A separate network then regresses the 6D position of each car based on raw features from the bounding boxes, as well as image features from the bounding box crops.\n\n\n### Our Solution in More Detail\n\nThere are two key components to our solution:\n1. 2D bounding boxes\n2. 6D-pose regression\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2202881%2F0081962fbbf790fdcea88dcb36e6649f%2FPKU_car_kaggle%20(5).png?generation=1579744111028907&amp;alt=media)\n\n\n\n\n## 1. 2D Bounding Boxes\n\n\nThe “backbone” of our approach is 2D bounding boxes. Because the training data does not provide labeled bounding boxes, we first modify this kernel [https://www.kaggle.com/hypocrites/simple-eda-with-imageai-object-detection] to obtain bounding boxes for the training set. The key change we made to the kernel is that we used the provided 3D car models to obtain car dimensions, and adjusted our bounding boxes based on the car dimensions for a given car label (whereas the kernel uses a generic car dimension for all cars). We call the boxes obtained from this method as “ground truth boxes”.\n\nFor the test set, we cannot obtain bounding boxes with this method, since the method requires the very data points we are trying to predict! We therefore resorted to 2D object detectors. First, we fine-tuned Faster-RCNN (from maskrcnn-benchmark) on the BDD100k dataset. The BDD100k tuned model is then further fine-tuned on our “ground truth boxes”. \n\n\nWe found that training with larger input images resulted in noticeably better validation MAP and LB score. In the end, we trained 3 separate detector backbones with the following configurations:\n1. X-101-32x8d-FPN (input size 1373x1100)\n2. R-101-FPN (input size 2000x1602)\n3. R-50-FPN (input size 2499x2002)\n\n\n\nThe 3 models are then used to predict bounding boxes for cars in the test images. Within each model, overlapping boxes are discarded based on an NMS threshold of 0.70, and any remaining boxes below 0.70 confidence are further discarded.\n\nThe boxes from the three models are then ensembled using weighted boxes fusion (https://github.com/ZFTurbo/Weighted-Boxes-Fusion) with equal weights. We found that by including only the boxes that are successfully fused/merged by all three models, our score improved. We call the boxes predicted by the object detectors as “predicted boxes” (as opposed to ground truth boxes).\n\n\n## 2. 6D-pose regression\n\nWith the ground truth boxes and predicted boxes in hand, we then regress the cars’ positions using a downstream network that we simply call Q-Net (Q for Quaternion).\n\n### Input:\n\nQ-Net consumes the following two inputs:\n\n1. Car bounding box crops (RGB images, resized to 128x128)\n2. 10 numerical features from the bounding boxes\n- Box width\n- Box height\n- Box width to height ratio\n- Box area\n- X-coordinate of the box center\n- Y-coordinate of the box center\n- “2D-distance” of box center from camera \n- “2D-angle” between box center and camera\n- Whether the box is close to the left boundary of the image\n- Whether the box is close to the right boundary of the image\n\n\nBoth inputs are standardized appropriately.\n\n\n### Output:\n\nCar translation (x, y, z)\nCar rotation (quaternion)\n\n\n\n## Brief Network Architecture\n\nThe bounding box crops are fed into a pre-trained DenseNet121 model from torchvision, while the numerical features are connected to FC layers. Both of these “input paths” are then connected to 2 “output paths”, one for predicting car rotations, and another for car translations.  \n\nThe intuition here is that both bounding-box crops and numerical features work hand-in-hand in regressing the 6D-pose, and so information from both input components should be “communicated” or “shared” with both output paths.\n\n\n\n## Some Training Details\n\nWe trained Q-Net on the ground truth boxes, and made predictions on the predicted boxes. Initially, we thought that training with out-of-fold predicted boxes would yield better test set generalization, but this was not the case. This is most likely because in the latter case, there was a challenge to map the predicted boxes with the corresponding ground truth pose information for a given car.\n\n\nFor translation we used L1-loss. For rotations, as mentioned, we converted the angles into quaternions, and used Dot Product Loss (https://arxiv.org/pdf/1901.09366.pdf).\n\nWe trained Q-Net on 10-folds and averaged the predictions from each fold. For translation, we simply averaged the predicted x, y, and z values respectively. For rotation, we averaged the quaternion predictions using (https://github.com/christophhagen/averaging-quaternions) before converting back to Euler angles.\n\n\n\n## Score Summary\n\nOur best model attains 0.124 on public LB and 0.116 on private LB.\n\n\n## Final Ensemble\n\nThanks to our team member @uiiurz1's Centernet model, we had relatively diverse models to ensemble. However, ensembling the two models of different nature was slightly challenging. We used weighted points fusion (weighted boxes fusion but with IOU replaced with 3D-distance) to ensemble our predictions, credit to @uiiurz1. By ensembling, we obtained a decent boost in score: 0.131 public LB, 0.123 private LB.\n\n\n## Closing Remarks\n\nWe would like to congratulate all the winners, and really anyone who benefited in some way from the competition. We want to thank our teammate @uiiurz1 for his great insights and collaborative spirit. Finally, we would also like to thank the competition host(s) and Kaggle for organizing the competition. Putting aside issues regarding the unknown competition metric and labelling methodology for cropped cars, this competition was intriguing enough to keep our minds occupied thinking about car position estimation during our drives home from work the past couple of months :)",
    "725988": "Nice Write-Up\nThanks for Sharing your Approach!! @css919",
    "726388": "You're welcome, thanks for reading too.",
    "727068": "Thanks for haring,great job!",
    "727090": "Thanks for reading!",
    "732536": "Congrats, thanks for sharing.",
    "734855": "Thanks!"
  },
  "source": "meta"
}