{
  "id": 127099,
  "title": "2nd Place solution",
  "url": "/competitions/pku-autonomous-driving/discussion/127099",
  "author_name": "stevenwudi",
  "post_date": "2020-01-22T10:07:35.673000",
  "votes": 48,
  "comment_count": 9,
  "views": 0,
  "content": "<h1>2nd Place for Kaggle_PKU_Baidu</h1>\n\n<p>Firstly congratulations to all top teams.</p>\n\n<p>Secondly I would like to congratulate to all my teammates for this collaborative team work, every member of the team is indispensable in this competition.</p>\n\n<h2>Approach</h2>\n\n<p>The overall pipeline is largely improved on previous method 6D-VNet [1].\nWe reckon we are the very few teams that didn't use CenterNet as the main network).\nThe system pipeline is as follows (the red color denotes the modules we added for this task):</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F16463%2Fbe74039a5b29b17e735b6f209c0b7e10%2Fsystem_pipeline_kaggle.PNG?generation=1579687638026915&amp;alt=media\" alt=\"\"></p>\n\n<p>The three major improvement consists of\n(1) better detector and conv backbone structure design;\n(2) post-processing with both mask and mesh information both geometrically and  via a Neural Mesh Renderer (NMR)[4] method;\n(3) novel way of ensembling of multiple models and\nweighted average of multiple predictions. </p>\n\n<p>We specify the implementation details as follows:</p>\n\n<h3>Implementation details</h3>\n\n<p>We build our framework based upon the open source project \n<a href=\"https://github.com/open-mmlab/mmdetection%5d\">MMDetection</a>. This is an excellent framework that can help us to modulised the code.\nPixel-level transform for image augmentation is called from the <a href=\"https://github.com/albumentations-team/albumentations\">Albumentations</a> library.</p>\n\n<p>The detector is a 3-stage Hybrid Task Cascade (HTC) [2] and the backbone is ImageNet pretrained High-resolution networks (HRNets) [3]. \nWe design two specific task heads for this challenge: one head taking ROIAligh feature for car class classification + quaternion regression and one head taking bounding box information \n centre location, height and width) for translation regression.\nWith this building block, we achieved private/public LB: 0.094/0.102.</p>\n\n<p>We then incorporated the training images from ApolloScape dataset and after cleaning the obviously wrong annotations, this leaves us with 6691 images and ~79,000 cars for training,\nThe kaggle dataset has around 4000 images for training,\nwe leave out 400 images randomly as validation. With tito(<a href=\"/its7171\">@its7171</a>) code for evaluation, we obtained ~0.4 mAP.\nOn public LB, we have only 0.110.\nSuch discrepancy between local validation and test mAP is a conumdrum that perplexes us until today!</p>\n\n<h3>Postprocessing</h3>\n\n<p>After visual examination, we find out the detector is working well (really well) for bounding box detection and mask segmentation (well, there are 100+ top conference paper doing the research in instance segmentation anyway). But the generated mesh from rotation and translation does not overlap quite well with the mask prediction.\nThus, we treat <code>z</code> as the oracle prediction and amend the value for <code>x</code> and <code>y</code> prediction.\nThis gives us a generous boost to 0.122/0.128 (from 0.105/0.110). </p>\n\n<h3>Model ensembles</h3>\n\n<p>Model ensemble is a necessity for kaggle top solutions: we train one model that directly regresses translation and one model regresses the <code>sigmoid</code> transformed translation.\nThe third model is trained with 0.5 flip of the image.</p>\n\n<p>Because of the speciality of the task: the network can output mask and mesh simultaneously, we merge the model by non-maximum suppression using the IoU between the predicted mesh and mask as the confident score.\nThe <code>max</code> strategy gives 3 model ensemble to 0.133/0.142.\nThe <code>average weighting</code> strategy generates even better result and is the final strategy we adopted.</p>\n\n<p>The organisors also provide the maskes that will be ignored during test evaluation.\nWe filtered out the ignore mask if predicted mesh has more than 20% overlap,\nthis will have around 0.002 mAP improvement.</p>\n\n<p>Below is the aforementioned progress we have achieved in a tabular form:</p>\n\n<p>|Method              | private LB             |  public LB|\n|:------------------: | :-------:|:-------------------------:|\n|HTC + HRNet + quaternion + translation | 0.094  | 0.102|\n|+ ApolloScape dataset      | 0.105          | 0.110|\n|+ z-&gt; x,y (postprocessing) | 0.122          | 0.128|\n|+ NMR                      | 0.127          | 0.132|\n|conf (0.1 -&gt; 0.8)          | 0.130          | 0.136|\n|+ 3 models ensemble (max)  | 0.133          | 0.142|\n|+ filter test ignore mask  | 0.136          | 0.145|\n|+ 6 models ensemble(weighted average)| 0.140 | 0.151|</p>\n\n<h2>Other bolts and nuts</h2>\n\n<h3>Visualisation using Open3D</h3>\n\n<p>We also use <a href=\"http://www.open3d.org/\">Open3d</a> to visualise the predicted validation images. The interactive 3d rendering technique allows us to examine the correctly predicted cars in the valid set. </p>\n\n<h3>Neural Mesh Renderer (NMR)</h3>\n\n<p>Neural 3D Mesh Renderer [4] is a very cool research which generates an approximate gradient for rasterization that enables the integration of rendering into neural networks.\nAfter releasing of the final private LB, we found out  using NMR actually gives a small improvement of the overall mAP. </p>\n\n<h3>What we haven't tried but think it has decent potential</h3>\n\n<ul>\n<li><p>Almost all the top winning solution adopted the CenterNet [5], it's very likely that the model ensemble with CentreNet will further boost the overall performance. We realise the universal adoptation of CenterNet in this challenge. We might be too comfortable sitting\nin our existing framework and the migration to fine-tune CenterNet seems a bit hassle which in return might ultimately causes us the top prize. </p></li>\n<li><p>Allocentric vs. Egocentric [6]. \nAllocentric representation is equivariant w.r.t. to RoI Image appearance, and is\nbetter-suited for learning. As also discussed <a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/127052\">here</a>, the modification of orientation id done by rotation matrix that moves camera center to target car center. But we are aware from the beginning that prediction of \ntranslation is far more difficult that rotation prediction, we didn't spend much effort in perfecting rotation regression.</p></li>\n</ul>\n\n<p><code>python\nyaw = 0\npitch = -np.arctan(x / z)\nroll = np.arctan(y / z)\nr = Rotation.from_euler(\"xyz\", (roll, pitch, yaw))\n</code></p>\n\n<h3>References</h3>\n\n<ul>\n<li>[1] 6D-VNet: End-To-End 6-DoF Vehicle Pose Estimation From Monocular RGB Images, Di WU et et., CVPRW2019 </li>\n<li>[2] Hybrid task cascade for instance segmentation, Chen et al., CVPR2019</li>\n<li>[3] Deep High-Resolution Representation Learning for Human Pose Estimation, Sun et al., CVPR2019</li>\n<li>[4] Neural 3D Mesh Renderer, Hiroharu Kato et al., CVPR2018</li>\n<li>[5] Objects as Points, Xingyi Zhou et al. CVPR2019</li>\n<li>[6] 3D-RCNN: Instance-level 3D Object Reconstruction via Render-and-Compare, Abhijit Kundu et al., CVPR 2018</li>\n</ul>",
  "messages": [
    {
      "id": 725635,
      "postDate": "2020-01-22T10:07:35.673Z",
      "content": "<h1>2nd Place for Kaggle_PKU_Baidu</h1>\n\n<p>Firstly congratulations to all top teams.</p>\n\n<p>Secondly I would like to congratulate to all my teammates for this collaborative team work, every member of the team is indispensable in this competition.</p>\n\n<h2>Approach</h2>\n\n<p>The overall pipeline is largely improved on previous method 6D-VNet [1].\nWe reckon we are the very few teams that didn't use CenterNet as the main network).\nThe system pipeline is as follows (the red color denotes the modules we added for this task):</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F16463%2Fbe74039a5b29b17e735b6f209c0b7e10%2Fsystem_pipeline_kaggle.PNG?generation=1579687638026915&amp;alt=media\" alt=\"\"></p>\n\n<p>The three major improvement consists of\n(1) better detector and conv backbone structure design;\n(2) post-processing with both mask and mesh information both geometrically and  via a Neural Mesh Renderer (NMR)[4] method;\n(3) novel way of ensembling of multiple models and\nweighted average of multiple predictions. </p>\n\n<p>We specify the implementation details as follows:</p>\n\n<h3>Implementation details</h3>\n\n<p>We build our framework based upon the open source project \n<a href=\"https://github.com/open-mmlab/mmdetection%5d\">MMDetection</a>. This is an excellent framework that can help us to modulised the code.\nPixel-level transform for image augmentation is called from the <a href=\"https://github.com/albumentations-team/albumentations\">Albumentations</a> library.</p>\n\n<p>The detector is a 3-stage Hybrid Task Cascade (HTC) [2] and the backbone is ImageNet pretrained High-resolution networks (HRNets) [3]. \nWe design two specific task heads for this challenge: one head taking ROIAligh feature for car class classification + quaternion regression and one head taking bounding box information \n centre location, height and width) for translation regression.\nWith this building block, we achieved private/public LB: 0.094/0.102.</p>\n\n<p>We then incorporated the training images from ApolloScape dataset and after cleaning the obviously wrong annotations, this leaves us with 6691 images and ~79,000 cars for training,\nThe kaggle dataset has around 4000 images for training,\nwe leave out 400 images randomly as validation. With tito(<a href=\"/its7171\">@its7171</a>) code for evaluation, we obtained ~0.4 mAP.\nOn public LB, we have only 0.110.\nSuch discrepancy between local validation and test mAP is a conumdrum that perplexes us until today!</p>\n\n<h3>Postprocessing</h3>\n\n<p>After visual examination, we find out the detector is working well (really well) for bounding box detection and mask segmentation (well, there are 100+ top conference paper doing the research in instance segmentation anyway). But the generated mesh from rotation and translation does not overlap quite well with the mask prediction.\nThus, we treat <code>z</code> as the oracle prediction and amend the value for <code>x</code> and <code>y</code> prediction.\nThis gives us a generous boost to 0.122/0.128 (from 0.105/0.110). </p>\n\n<h3>Model ensembles</h3>\n\n<p>Model ensemble is a necessity for kaggle top solutions: we train one model that directly regresses translation and one model regresses the <code>sigmoid</code> transformed translation.\nThe third model is trained with 0.5 flip of the image.</p>\n\n<p>Because of the speciality of the task: the network can output mask and mesh simultaneously, we merge the model by non-maximum suppression using the IoU between the predicted mesh and mask as the confident score.\nThe <code>max</code> strategy gives 3 model ensemble to 0.133/0.142.\nThe <code>average weighting</code> strategy generates even better result and is the final strategy we adopted.</p>\n\n<p>The organisors also provide the maskes that will be ignored during test evaluation.\nWe filtered out the ignore mask if predicted mesh has more than 20% overlap,\nthis will have around 0.002 mAP improvement.</p>\n\n<p>Below is the aforementioned progress we have achieved in a tabular form:</p>\n\n<p>|Method              | private LB             |  public LB|\n|:------------------: | :-------:|:-------------------------:|\n|HTC + HRNet + quaternion + translation | 0.094  | 0.102|\n|+ ApolloScape dataset      | 0.105          | 0.110|\n|+ z-&gt; x,y (postprocessing) | 0.122          | 0.128|\n|+ NMR                      | 0.127          | 0.132|\n|conf (0.1 -&gt; 0.8)          | 0.130          | 0.136|\n|+ 3 models ensemble (max)  | 0.133          | 0.142|\n|+ filter test ignore mask  | 0.136          | 0.145|\n|+ 6 models ensemble(weighted average)| 0.140 | 0.151|</p>\n\n<h2>Other bolts and nuts</h2>\n\n<h3>Visualisation using Open3D</h3>\n\n<p>We also use <a href=\"http://www.open3d.org/\">Open3d</a> to visualise the predicted validation images. The interactive 3d rendering technique allows us to examine the correctly predicted cars in the valid set. </p>\n\n<h3>Neural Mesh Renderer (NMR)</h3>\n\n<p>Neural 3D Mesh Renderer [4] is a very cool research which generates an approximate gradient for rasterization that enables the integration of rendering into neural networks.\nAfter releasing of the final private LB, we found out  using NMR actually gives a small improvement of the overall mAP. </p>\n\n<h3>What we haven't tried but think it has decent potential</h3>\n\n<ul>\n<li><p>Almost all the top winning solution adopted the CenterNet [5], it's very likely that the model ensemble with CentreNet will further boost the overall performance. We realise the universal adoptation of CenterNet in this challenge. We might be too comfortable sitting\nin our existing framework and the migration to fine-tune CenterNet seems a bit hassle which in return might ultimately causes us the top prize. </p></li>\n<li><p>Allocentric vs. Egocentric [6]. \nAllocentric representation is equivariant w.r.t. to RoI Image appearance, and is\nbetter-suited for learning. As also discussed <a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/127052\">here</a>, the modification of orientation id done by rotation matrix that moves camera center to target car center. But we are aware from the beginning that prediction of \ntranslation is far more difficult that rotation prediction, we didn't spend much effort in perfecting rotation regression.</p></li>\n</ul>\n\n<p><code>python\nyaw = 0\npitch = -np.arctan(x / z)\nroll = np.arctan(y / z)\nr = Rotation.from_euler(\"xyz\", (roll, pitch, yaw))\n</code></p>\n\n<h3>References</h3>\n\n<ul>\n<li>[1] 6D-VNet: End-To-End 6-DoF Vehicle Pose Estimation From Monocular RGB Images, Di WU et et., CVPRW2019 </li>\n<li>[2] Hybrid task cascade for instance segmentation, Chen et al., CVPR2019</li>\n<li>[3] Deep High-Resolution Representation Learning for Human Pose Estimation, Sun et al., CVPR2019</li>\n<li>[4] Neural 3D Mesh Renderer, Hiroharu Kato et al., CVPR2018</li>\n<li>[5] Objects as Points, Xingyi Zhou et al. CVPR2019</li>\n<li>[6] 3D-RCNN: Instance-level 3D Object Reconstruction via Render-and-Compare, Abhijit Kundu et al., CVPR 2018</li>\n</ul>",
      "rawMarkdown": "# 2nd Place for Kaggle_PKU_Baidu\n\nFirstly congratulations to all top teams.\n\nSecondly I would like to congratulate to all my teammates for this collaborative team work, every member of the team is indispensable in this competition.\n\n\n## Approach \n\nThe overall pipeline is largely improved on previous method 6D-VNet [[1]](#references).\nWe reckon we are the very few teams that didn't use CenterNet as the main network).\nThe system pipeline is as follows (the red color denotes the modules we added for this task):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F16463%2Fbe74039a5b29b17e735b6f209c0b7e10%2Fsystem_pipeline_kaggle.PNG?generation=1579687638026915&amp;alt=media)\n\n\nThe three major improvement consists of\n(1) better detector and conv backbone structure design;\n(2) post-processing with both mask and mesh information both geometrically and  via a Neural Mesh Renderer (NMR)[[4]](#references) method;\n(3) novel way of ensembling of multiple models and\nweighted average of multiple predictions. \n\nWe specify the implementation details as follows:\n\n\n### Implementation details\nWe build our framework based upon the open source project \n[MMDetection](https://github.com/open-mmlab/mmdetection]). This is an excellent framework that can help us to modulised the code.\nPixel-level transform for image augmentation is called from the [Albumentations](https://github.com/albumentations-team/albumentations) library.\n\nThe detector is a 3-stage Hybrid Task Cascade (HTC) [[2]](#references) and the backbone is ImageNet pretrained High-resolution networks (HRNets) [[3]](#references). \nWe design two specific task heads for this challenge: one head taking ROIAligh feature for car class classification + quaternion regression and one head taking bounding box information \n centre location, height and width) for translation regression.\nWith this building block, we achieved private/public LB: 0.094/0.102.\n\nWe then incorporated the training images from ApolloScape dataset and after cleaning the obviously wrong annotations, this leaves us with 6691 images and ~79,000 cars for training,\nThe kaggle dataset has around 4000 images for training,\nwe leave out 400 images randomly as validation. With tito(@its7171) code for evaluation, we obtained ~0.4 mAP.\nOn public LB, we have only 0.110.\nSuch discrepancy between local validation and test mAP is a conumdrum that perplexes us until today!\n\n### Postprocessing\nAfter visual examination, we find out the detector is working well (really well) for bounding box detection and mask segmentation (well, there are 100+ top conference paper doing the research in instance segmentation anyway). But the generated mesh from rotation and translation does not overlap quite well with the mask prediction.\nThus, we treat `z` as the oracle prediction and amend the value for `x` and `y` prediction.\nThis gives us a generous boost to 0.122/0.128 (from 0.105/0.110). \n\n\n### Model ensembles\nModel ensemble is a necessity for kaggle top solutions: we train one model that directly regresses translation and one model regresses the `sigmoid` transformed translation.\nThe third model is trained with 0.5 flip of the image.\n\nBecause of the speciality of the task: the network can output mask and mesh simultaneously, we merge the model by non-maximum suppression using the IoU between the predicted mesh and mask as the confident score.\nThe `max` strategy gives 3 model ensemble to 0.133/0.142.\nThe `average weighting` strategy generates even better result and is the final strategy we adopted.\n\nThe organisors also provide the maskes that will be ignored during test evaluation.\nWe filtered out the ignore mask if predicted mesh has more than 20% overlap,\nthis will have around 0.002 mAP improvement.\n\nBelow is the aforementioned progress we have achieved in a tabular form:\n\n\n\n|Method              | private LB             |  public LB|\n|:------------------: | :-------:|:-------------------------:|\n|HTC + HRNet + quaternion + translation | 0.094  | 0.102|\n|+ ApolloScape dataset      | 0.105          | 0.110|\n|+ z-&gt; x,y (postprocessing) | 0.122          | 0.128|\n|+ NMR                      | 0.127          | 0.132|\n|conf (0.1 -&gt; 0.8)          | 0.130          | 0.136|\n|+ 3 models ensemble (max)  | 0.133          | 0.142|\n|+ filter test ignore mask  | 0.136          | 0.145|\n|+ 6 models ensemble(weighted average)| 0.140 | 0.151|\n  \n  \n## Other bolts and nuts\n \n###  Visualisation using Open3D\n\nWe also use [Open3d](http://www.open3d.org/) to visualise the predicted validation images. The interactive 3d rendering technique allows us to examine the correctly predicted cars in the valid set. \n\n### Neural Mesh Renderer (NMR)\nNeural 3D Mesh Renderer [[4]](#references) is a very cool research which generates an approximate gradient for rasterization that enables the integration of rendering into neural networks.\nAfter releasing of the final private LB, we found out  using NMR actually gives a small improvement of the overall mAP. \n\n\n\n### What we haven't tried but think it has decent potential\n\n- Almost all the top winning solution adopted the CenterNet [[5]](#references), it's very likely that the model ensemble with CentreNet will further boost the overall performance. We realise the universal adoptation of CenterNet in this challenge. We might be too comfortable sitting\nin our existing framework and the migration to fine-tune CenterNet seems a bit hassle which in return might ultimately causes us the top prize. \n\n- Allocentric vs. Egocentric [[6]](#references). \nAllocentric representation is equivariant w.r.t. to RoI Image appearance, and is\nbetter-suited for learning. As also discussed [here](https://www.kaggle.com/c/pku-autonomous-driving/discussion/127052), the modification of orientation id done by rotation matrix that moves camera center to target car center. But we are aware from the beginning that prediction of \ntranslation is far more difficult that rotation prediction, we didn't spend much effort in perfecting rotation regression.\n\n```python\nyaw = 0\npitch = -np.arctan(x / z)\nroll = np.arctan(y / z)\nr = Rotation.from_euler(\"xyz\", (roll, pitch, yaw))\n```\n\n\n### References\n\n- [1] 6D-VNet: End-To-End 6-DoF Vehicle Pose Estimation From Monocular RGB Images, Di WU et et., CVPRW2019 \n- [2] Hybrid task cascade for instance segmentation, Chen et al., CVPR2019\n- [3] Deep High-Resolution Representation Learning for Human Pose Estimation, Sun et al., CVPR2019\n- [4] Neural 3D Mesh Renderer, Hiroharu Kato et al., CVPR2018\n- [5] Objects as Points, Xingyi Zhou et al. CVPR2019\n- [6] 3D-RCNN: Instance-level 3D Object Reconstruction via Render-and-Compare, Abhijit Kundu et al., CVPR 2018\n\n",
      "votes": 48
    },
    {
      "id": 736655,
      "postDate": "2020-02-04T12:33:50.970Z",
      "content": "<p>congrats </p>",
      "rawMarkdown": "congrats ",
      "votes": 1
    },
    {
      "id": 725680,
      "postDate": "2020-01-22T11:23:34.850Z",
      "content": "<p>Congrats and thank you for sharing really solid solution!</p>\n\n<blockquote>\n  <p>But we are aware from the beginning that prediction of translation is far more difficult that rotation prediction, we didn't spend much effort in perfecting rotation regression.</p>\n</blockquote>\n\n<p>I agree... I spent much time to solve this problem but the accuracy was bound by translation (depth)...</p>",
      "rawMarkdown": "Congrats and thank you for sharing really solid solution!\n\n&gt; But we are aware from the beginning that prediction of translation is far more difficult that rotation prediction, we didn't spend much effort in perfecting rotation regression.\n\nI agree... I spent much time to solve this problem but the accuracy was bound by translation (depth)...",
      "votes": 2,
      "replies": [
        {
          "id": 728956,
          "postDate": "2020-01-25T14:00:18.533Z",
          "content": "<p>We too spent way too much time in correctly converting the euler angle to quaternions...</p>",
          "rawMarkdown": "We too spent way too much time in correctly converting the euler angle to quaternions...",
          "votes": 1
        }
      ]
    },
    {
      "id": 749902,
      "postDate": "2020-02-19T01:43:19.540Z",
      "content": "<p>Did you train mmdetection on the kaggle dataset? If yes, how did you deal with the annotations that are required in that? I'm trying to test one of the kaggle images on mmdetection model but it's giving an error regarding annotations. Sorry I'm just really new at this and would appreciate guidance on this. Thanks!</p>",
      "rawMarkdown": "Did you train mmdetection on the kaggle dataset? If yes, how did you deal with the annotations that are required in that? I'm trying to test one of the kaggle images on mmdetection model but it's giving an error regarding annotations. Sorry I'm just really new at this and would appreciate guidance on this. Thanks!"
    },
    {
      "id": 732531,
      "postDate": "2020-01-29T23:14:54.113Z",
      "content": "<p>Congrats and thanks for sharing. How did you use NMR for postprocessing? What kind of postprocessing is applied?</p>",
      "rawMarkdown": "Congrats and thanks for sharing. How did you use NMR for postprocessing? What kind of postprocessing is applied?",
      "replies": [
        {
          "id": 734155,
          "postDate": "2020-02-01T02:36:57.207Z",
          "content": "<p>We combine Mask with Mesh, use silhouette as supervising signal to adjust R,T. It improves the accuracy of a single model. We plan to write a paper using the technique.\n<a href=\"/corochann\">@corochann</a> you are with preferred network, you might know the author of the NMR then~</p>",
          "rawMarkdown": "We combine Mask with Mesh, use silhouette as supervising signal to adjust R,T. It improves the accuracy of a single model. We plan to write a paper using the technique.\n@corochann you are with preferred network, you might know the author of the NMR then~"
        }
      ]
    },
    {
      "id": 725652,
      "postDate": "2020-01-22T10:44:40.260Z",
      "content": "<p>Thanks for sharing!\n1) Did you try weighted non-local neighbour embedding which you are mentioning in [1]?\n2) As I understand - you regressed full translation vector (x,y,z) but used only z from it, x,y you received from bbox? \n3) What accuracy did you get on car class prediction? Does this affect much final result?</p>",
      "rawMarkdown": "Thanks for sharing!\n1) Did you try weighted non-local neighbour embedding which you are mentioning in [1]?\n2) As I understand - you regressed full translation vector (x,y,z) but used only z from it, x,y you received from bbox? \n3) What accuracy did you get on car class prediction? Does this affect much final result?",
      "replies": [
        {
          "id": 728961,
          "postDate": "2020-01-25T14:03:55.173Z",
          "content": "<p>1) We didn't try non-local embedding or GCnet, which we should have....in the end, we are a bit contraint by the computational resources so we didn't go for it eventually.\n2) Yes, we regressed the full translation, and modify the x,y from box\n3) without augmentation, we have 95%+ accuracy on 34 car class prediction, with augmentation, its around 85% (in reality, we shouldn't use augmentation after all). The car class accuracy matters, but it is not a very difficult task...</p>",
          "rawMarkdown": "1) We didn't try non-local embedding or GCnet, which we should have....in the end, we are a bit contraint by the computational resources so we didn't go for it eventually.\n2) Yes, we regressed the full translation, and modify the x,y from box\n3) without augmentation, we have 95%+ accuracy on 34 car class prediction, with augmentation, its around 85% (in reality, we shouldn't use augmentation after all). The car class accuracy matters, but it is not a very difficult task...",
          "votes": 1
        }
      ]
    },
    {
      "id": 725653,
      "postDate": "2020-01-22T10:46:24.130Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 736655,
      "author_name": "Mohamed Tarek",
      "author_url": "",
      "post_date": "2020-02-04T12:33:50.970000",
      "content": "<p>congrats </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 725680,
      "author_name": "yu4u",
      "author_url": "",
      "post_date": "2020-01-22T11:23:34.850000",
      "content": "<p>Congrats and thank you for sharing really solid solution!</p>\n\n<blockquote>\n  <p>But we are aware from the beginning that prediction of translation is far more difficult that rotation prediction, we didn't spend much effort in perfecting rotation regression.</p>\n</blockquote>\n\n<p>I agree... I spent much time to solve this problem but the accuracy was bound by translation (depth)...</p>",
      "votes": 2,
      "replies": [
        {
          "id": 728956,
          "author_name": "stevenwudi",
          "author_url": "",
          "post_date": "2020-01-25T14:00:18.533000",
          "content": "<p>We too spent way too much time in correctly converting the euler angle to quaternions...</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 749902,
      "author_name": "aadilism",
      "author_url": "",
      "post_date": "2020-02-19T01:43:19.540000",
      "content": "<p>Did you train mmdetection on the kaggle dataset? If yes, how did you deal with the annotations that are required in that? I'm trying to test one of the kaggle images on mmdetection model but it's giving an error regarding annotations. Sorry I'm just really new at this and would appreciate guidance on this. Thanks!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 732531,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2020-01-29T23:14:54.113000",
      "content": "<p>Congrats and thanks for sharing. How did you use NMR for postprocessing? What kind of postprocessing is applied?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 734155,
          "author_name": "stevenwudi",
          "author_url": "",
          "post_date": "2020-02-01T02:36:57.207000",
          "content": "<p>We combine Mask with Mesh, use silhouette as supervising signal to adjust R,T. It improves the accuracy of a single model. We plan to write a paper using the technique.\n<a href=\"/corochann\">@corochann</a> you are with preferred network, you might know the author of the NMR then~</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 725652,
      "author_name": "Kuts Alexey",
      "author_url": "",
      "post_date": "2020-01-22T10:44:40.260000",
      "content": "<p>Thanks for sharing!\n1) Did you try weighted non-local neighbour embedding which you are mentioning in [1]?\n2) As I understand - you regressed full translation vector (x,y,z) but used only z from it, x,y you received from bbox? \n3) What accuracy did you get on car class prediction? Does this affect much final result?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 728961,
          "author_name": "stevenwudi",
          "author_url": "",
          "post_date": "2020-01-25T14:03:55.173000",
          "content": "<p>1) We didn't try non-local embedding or GCnet, which we should have....in the end, we are a bit contraint by the computational resources so we didn't go for it eventually.\n2) Yes, we regressed the full translation, and modify the x,y from box\n3) without augmentation, we have 95%+ accuracy on 34 car class prediction, with augmentation, its around 85% (in reality, we shouldn't use augmentation after all). The car class accuracy matters, but it is not a very difficult task...</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 725653,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-22T10:46:24.130000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "725635": "# 2nd Place for Kaggle_PKU_Baidu\n\nFirstly congratulations to all top teams.\n\nSecondly I would like to congratulate to all my teammates for this collaborative team work, every member of the team is indispensable in this competition.\n\n\n## Approach \n\nThe overall pipeline is largely improved on previous method 6D-VNet [[1]](#references).\nWe reckon we are the very few teams that didn't use CenterNet as the main network).\nThe system pipeline is as follows (the red color denotes the modules we added for this task):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F16463%2Fbe74039a5b29b17e735b6f209c0b7e10%2Fsystem_pipeline_kaggle.PNG?generation=1579687638026915&amp;alt=media)\n\n\nThe three major improvement consists of\n(1) better detector and conv backbone structure design;\n(2) post-processing with both mask and mesh information both geometrically and  via a Neural Mesh Renderer (NMR)[[4]](#references) method;\n(3) novel way of ensembling of multiple models and\nweighted average of multiple predictions. \n\nWe specify the implementation details as follows:\n\n\n### Implementation details\nWe build our framework based upon the open source project \n[MMDetection](https://github.com/open-mmlab/mmdetection]). This is an excellent framework that can help us to modulised the code.\nPixel-level transform for image augmentation is called from the [Albumentations](https://github.com/albumentations-team/albumentations) library.\n\nThe detector is a 3-stage Hybrid Task Cascade (HTC) [[2]](#references) and the backbone is ImageNet pretrained High-resolution networks (HRNets) [[3]](#references). \nWe design two specific task heads for this challenge: one head taking ROIAligh feature for car class classification + quaternion regression and one head taking bounding box information \n centre location, height and width) for translation regression.\nWith this building block, we achieved private/public LB: 0.094/0.102.\n\nWe then incorporated the training images from ApolloScape dataset and after cleaning the obviously wrong annotations, this leaves us with 6691 images and ~79,000 cars for training,\nThe kaggle dataset has around 4000 images for training,\nwe leave out 400 images randomly as validation. With tito(@its7171) code for evaluation, we obtained ~0.4 mAP.\nOn public LB, we have only 0.110.\nSuch discrepancy between local validation and test mAP is a conumdrum that perplexes us until today!\n\n### Postprocessing\nAfter visual examination, we find out the detector is working well (really well) for bounding box detection and mask segmentation (well, there are 100+ top conference paper doing the research in instance segmentation anyway). But the generated mesh from rotation and translation does not overlap quite well with the mask prediction.\nThus, we treat `z` as the oracle prediction and amend the value for `x` and `y` prediction.\nThis gives us a generous boost to 0.122/0.128 (from 0.105/0.110). \n\n\n### Model ensembles\nModel ensemble is a necessity for kaggle top solutions: we train one model that directly regresses translation and one model regresses the `sigmoid` transformed translation.\nThe third model is trained with 0.5 flip of the image.\n\nBecause of the speciality of the task: the network can output mask and mesh simultaneously, we merge the model by non-maximum suppression using the IoU between the predicted mesh and mask as the confident score.\nThe `max` strategy gives 3 model ensemble to 0.133/0.142.\nThe `average weighting` strategy generates even better result and is the final strategy we adopted.\n\nThe organisors also provide the maskes that will be ignored during test evaluation.\nWe filtered out the ignore mask if predicted mesh has more than 20% overlap,\nthis will have around 0.002 mAP improvement.\n\nBelow is the aforementioned progress we have achieved in a tabular form:\n\n\n\n|Method              | private LB             |  public LB|\n|:------------------: | :-------:|:-------------------------:|\n|HTC + HRNet + quaternion + translation | 0.094  | 0.102|\n|+ ApolloScape dataset      | 0.105          | 0.110|\n|+ z-&gt; x,y (postprocessing) | 0.122          | 0.128|\n|+ NMR                      | 0.127          | 0.132|\n|conf (0.1 -&gt; 0.8)          | 0.130          | 0.136|\n|+ 3 models ensemble (max)  | 0.133          | 0.142|\n|+ filter test ignore mask  | 0.136          | 0.145|\n|+ 6 models ensemble(weighted average)| 0.140 | 0.151|\n  \n  \n## Other bolts and nuts\n \n###  Visualisation using Open3D\n\nWe also use [Open3d](http://www.open3d.org/) to visualise the predicted validation images. The interactive 3d rendering technique allows us to examine the correctly predicted cars in the valid set. \n\n### Neural Mesh Renderer (NMR)\nNeural 3D Mesh Renderer [[4]](#references) is a very cool research which generates an approximate gradient for rasterization that enables the integration of rendering into neural networks.\nAfter releasing of the final private LB, we found out  using NMR actually gives a small improvement of the overall mAP. \n\n\n\n### What we haven't tried but think it has decent potential\n\n- Almost all the top winning solution adopted the CenterNet [[5]](#references), it's very likely that the model ensemble with CentreNet will further boost the overall performance. We realise the universal adoptation of CenterNet in this challenge. We might be too comfortable sitting\nin our existing framework and the migration to fine-tune CenterNet seems a bit hassle which in return might ultimately causes us the top prize. \n\n- Allocentric vs. Egocentric [[6]](#references). \nAllocentric representation is equivariant w.r.t. to RoI Image appearance, and is\nbetter-suited for learning. As also discussed [here](https://www.kaggle.com/c/pku-autonomous-driving/discussion/127052), the modification of orientation id done by rotation matrix that moves camera center to target car center. But we are aware from the beginning that prediction of \ntranslation is far more difficult that rotation prediction, we didn't spend much effort in perfecting rotation regression.\n\n```python\nyaw = 0\npitch = -np.arctan(x / z)\nroll = np.arctan(y / z)\nr = Rotation.from_euler(\"xyz\", (roll, pitch, yaw))\n```\n\n\n### References\n\n- [1] 6D-VNet: End-To-End 6-DoF Vehicle Pose Estimation From Monocular RGB Images, Di WU et et., CVPRW2019 \n- [2] Hybrid task cascade for instance segmentation, Chen et al., CVPR2019\n- [3] Deep High-Resolution Representation Learning for Human Pose Estimation, Sun et al., CVPR2019\n- [4] Neural 3D Mesh Renderer, Hiroharu Kato et al., CVPR2018\n- [5] Objects as Points, Xingyi Zhou et al. CVPR2019\n- [6] 3D-RCNN: Instance-level 3D Object Reconstruction via Render-and-Compare, Abhijit Kundu et al., CVPR 2018\n\n",
    "736655": "congrats ",
    "725680": "Congrats and thank you for sharing really solid solution!\n\n&gt; But we are aware from the beginning that prediction of translation is far more difficult that rotation prediction, we didn't spend much effort in perfecting rotation regression.\n\nI agree... I spent much time to solve this problem but the accuracy was bound by translation (depth)...",
    "749902": "Did you train mmdetection on the kaggle dataset? If yes, how did you deal with the annotations that are required in that? I'm trying to test one of the kaggle images on mmdetection model but it's giving an error regarding annotations. Sorry I'm just really new at this and would appreciate guidance on this. Thanks!",
    "732531": "Congrats and thanks for sharing. How did you use NMR for postprocessing? What kind of postprocessing is applied?",
    "725652": "Thanks for sharing!\n1) Did you try weighted non-local neighbour embedding which you are mentioning in [1]?\n2) As I understand - you regressed full translation vector (x,y,z) but used only z from it, x,y you received from bbox? \n3) What accuracy did you get on car class prediction? Does this affect much final result?",
    "725653": ""
  }
}