{
  "id": 122820,
  "title": "1st place solution (0.220 public LB)",
  "url": "/competitions/3d-object-detection-for-autonomous-vehicles/discussion/122820",
  "author_name": "Wenjing Zhang",
  "post_date": "2019-12-23T03:53:26.521000",
  "votes": 52,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hi, the post may be a little late. Thanks to organizers and congratulations to all!</p>\n\n<h1>Background</h1>\n\n<p>For the record, I want to state that I am really new to Kaggle and actually really new to machine learning. I only started to work on machine learning 9 month ago. My background is in 3D Computer Graphics so I looked into the problems more in the view of 3D than in machine learning. Pardon me if I made any silly mistake with regard to machine learning.</p>\n\n<p>Before this competition, I won 2nd place in CVPR 2019 competition WAD-Beyond Single-Frame Perception hosted by Baidu on a similar topic: 3D detection with Lidar points (<a href=\"http://wad.ai/2019/challenge.html\">http://wad.ai/2019/challenge.html</a>). The tricks we used is actually based on that competition. So we just adopt and modified the tricks and have no idea about the performance without the tricks. </p>\n\n<h1>Models</h1>\n\n<p>Our method is based on Voxelnet (<a href=\"https://github.com/traveller59/second.pytorch\">traveller59</a>) But our tricks can be applied to any other network. I gave a try with PointRCNN but got no luck. </p>\n\n<p>Basic Setting:\n- No external data\n- Data Augmentation: flip x and y, random rotation, random scaling and random translation\n- No Ground Truth Augmentation\n- Detection Range [-100,-100,-5,100,100,3]\n- No direction classifier\n- No specific post processing but score thresholding (0.1) and NMS</p>\n\n<p>Two types of voxelnet are implemented in second: FHD and PointPillars. We use both. Based on the idea that ensembling models with more difference results in better results, we use different settings to train a set of models. E.g. we use voxel size 0.1, 0.125, 0.2, 0.25 for PointPillar, 0.1 and 0.125 for FHD. Different Voxel Feature Extractors are also used for different models(e.g. PillarFeatureNet, PillarFeatureNetRadius,PillarFeatureNetRadiusHeight). RPN is modified so that the final resolution is 200 X 200 or 250 X 250. </p>\n\n<h1>Key Tricks</h1>\n\n<p>The key to our method is Test time augmentation (TTA) and model ensemble with a specific 3D box fusion method.</p>\n\n<h2><strong>Test Time Augmentation</strong></h2>\n\n<p>We transform the point cloud into several copies. Each copy is then feed into the network and get the predicted boxes. Then the predicted boxes are transformed back. For example, we can rotate the point cloud 20 degrees. After we got the predicted boxes (after NMS), we rotate the predicted boxes -20 degrees (both the center position and yaw angle are changed). We do this for many copies and fuse the results. During the competition, we only use 4 copies: original, flip x, flip y, flip x and y. We tried more copies with rotation, scaling and translation. It gives slightly better results but needs more inference time.</p>\n\n<h2><strong>Model ensemble</strong></h2>\n\n<p>We use TTA to get the predicted boxes for each model and fuse the results.</p>\n\n<p>Best Single Model (No TTA):          &gt;0.175 (I only got the score for a model with TTA score 0.193. But I  got a better model with TTA score 0.197)\nBest Single Model (TTA):    0.197</p>\n\n<p>Ensembling 7 best models: 0.220</p>\n\n<h2><strong>3D box fusion</strong></h2>\n\n<p>After we got several copies of predicted boxes (from TTA or model ensemble), we need to fuse them. We choose boxes with same label and IOU &gt; 0.6 to fuse into one. After the competition,we tried Weighted Boxes Fusion: ensembling boxes for object detection models. It actually gives better results (0.222 Public LB) but is slower. To fuse boxes, the center and size are directly weighted averaged. The key is the yaw angle. If you predict the direction and can trust it. You may use sin and cos interpolation.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3701925%2F7c611366e6f5e39e70bbbe3d0b46d12e%2Fyaw.png?generation=1577071532448853&amp;alt=media\" alt=\"\"></p>\n\n<p>However, in our implementation, we did not use direction classifier (it won’t affect the score and it can not be 100% accurate). We need to deal with one ambiguity: yaw angle is the same to yaw angle + 180 degrees. Say for one box, one prediction might give yaw angle 175 degrees or -5 degrees. <br>\nIt  does not matter since it is of the same IOU to ground truth. However, when we do box fusion, it matters. Say for the same box, another prediction gives yaw angle 5 degrees. If the first prediction is -5 degree. The average is 0 degree. It is okay. However, if the first prediction is 175 degree, the averaged boxes will be with yaw angle 90 degrees.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3701925%2F1b3f051743db56d52c52e61e943091c8%2Fyaw2.png?generation=1577071794662291&amp;alt=media\" alt=\"\"></p>\n\n<p>To deal with this problem, we use an intuitive method: average 2 * yaw instead of yaw. \nThe idea is that sin and cos interpolation of yaw angle can remove 360 degrees ambiguity. So sin and cos interpolation of 2 * yaw angle can remove 180 degrees ambiguity. I have no theoretical proof yet but intuitively feel it is correct. Actually the idea came into my head when I was telling bedtime story to my daughter. She asked me to tell the story over and over again so I just repeated it😂 !</p>\n\n<p>I uploaded our code and pretrained models. You may give a try. Have fun!</p>\n\n<p>p.s. \nFor sparse convolution, since spconv used in second might be patented, we replace it with <a href=\"https://github.com/facebookresearch/SparseConvNet\">this one</a>. It is much slower. </p>",
  "messages": [
    {
      "id": 701081,
      "postDate": "2019-12-23T03:53:26.520Z",
      "content": "<p>Hi, the post may be a little late. Thanks to organizers and congratulations to all!</p>\n\n<h1>Background</h1>\n\n<p>For the record, I want to state that I am really new to Kaggle and actually really new to machine learning. I only started to work on machine learning 9 month ago. My background is in 3D Computer Graphics so I looked into the problems more in the view of 3D than in machine learning. Pardon me if I made any silly mistake with regard to machine learning.</p>\n\n<p>Before this competition, I won 2nd place in CVPR 2019 competition WAD-Beyond Single-Frame Perception hosted by Baidu on a similar topic: 3D detection with Lidar points (<a href=\"http://wad.ai/2019/challenge.html\">http://wad.ai/2019/challenge.html</a>). The tricks we used is actually based on that competition. So we just adopt and modified the tricks and have no idea about the performance without the tricks. </p>\n\n<h1>Models</h1>\n\n<p>Our method is based on Voxelnet (<a href=\"https://github.com/traveller59/second.pytorch\">traveller59</a>) But our tricks can be applied to any other network. I gave a try with PointRCNN but got no luck. </p>\n\n<p>Basic Setting:\n- No external data\n- Data Augmentation: flip x and y, random rotation, random scaling and random translation\n- No Ground Truth Augmentation\n- Detection Range [-100,-100,-5,100,100,3]\n- No direction classifier\n- No specific post processing but score thresholding (0.1) and NMS</p>\n\n<p>Two types of voxelnet are implemented in second: FHD and PointPillars. We use both. Based on the idea that ensembling models with more difference results in better results, we use different settings to train a set of models. E.g. we use voxel size 0.1, 0.125, 0.2, 0.25 for PointPillar, 0.1 and 0.125 for FHD. Different Voxel Feature Extractors are also used for different models(e.g. PillarFeatureNet, PillarFeatureNetRadius,PillarFeatureNetRadiusHeight). RPN is modified so that the final resolution is 200 X 200 or 250 X 250. </p>\n\n<h1>Key Tricks</h1>\n\n<p>The key to our method is Test time augmentation (TTA) and model ensemble with a specific 3D box fusion method.</p>\n\n<h2><strong>Test Time Augmentation</strong></h2>\n\n<p>We transform the point cloud into several copies. Each copy is then feed into the network and get the predicted boxes. Then the predicted boxes are transformed back. For example, we can rotate the point cloud 20 degrees. After we got the predicted boxes (after NMS), we rotate the predicted boxes -20 degrees (both the center position and yaw angle are changed). We do this for many copies and fuse the results. During the competition, we only use 4 copies: original, flip x, flip y, flip x and y. We tried more copies with rotation, scaling and translation. It gives slightly better results but needs more inference time.</p>\n\n<h2><strong>Model ensemble</strong></h2>\n\n<p>We use TTA to get the predicted boxes for each model and fuse the results.</p>\n\n<p>Best Single Model (No TTA):          &gt;0.175 (I only got the score for a model with TTA score 0.193. But I  got a better model with TTA score 0.197)\nBest Single Model (TTA):    0.197</p>\n\n<p>Ensembling 7 best models: 0.220</p>\n\n<h2><strong>3D box fusion</strong></h2>\n\n<p>After we got several copies of predicted boxes (from TTA or model ensemble), we need to fuse them. We choose boxes with same label and IOU &gt; 0.6 to fuse into one. After the competition,we tried Weighted Boxes Fusion: ensembling boxes for object detection models. It actually gives better results (0.222 Public LB) but is slower. To fuse boxes, the center and size are directly weighted averaged. The key is the yaw angle. If you predict the direction and can trust it. You may use sin and cos interpolation.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3701925%2F7c611366e6f5e39e70bbbe3d0b46d12e%2Fyaw.png?generation=1577071532448853&amp;alt=media\" alt=\"\"></p>\n\n<p>However, in our implementation, we did not use direction classifier (it won’t affect the score and it can not be 100% accurate). We need to deal with one ambiguity: yaw angle is the same to yaw angle + 180 degrees. Say for one box, one prediction might give yaw angle 175 degrees or -5 degrees. <br>\nIt  does not matter since it is of the same IOU to ground truth. However, when we do box fusion, it matters. Say for the same box, another prediction gives yaw angle 5 degrees. If the first prediction is -5 degree. The average is 0 degree. It is okay. However, if the first prediction is 175 degree, the averaged boxes will be with yaw angle 90 degrees.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3701925%2F1b3f051743db56d52c52e61e943091c8%2Fyaw2.png?generation=1577071794662291&amp;alt=media\" alt=\"\"></p>\n\n<p>To deal with this problem, we use an intuitive method: average 2 * yaw instead of yaw. \nThe idea is that sin and cos interpolation of yaw angle can remove 360 degrees ambiguity. So sin and cos interpolation of 2 * yaw angle can remove 180 degrees ambiguity. I have no theoretical proof yet but intuitively feel it is correct. Actually the idea came into my head when I was telling bedtime story to my daughter. She asked me to tell the story over and over again so I just repeated it😂 !</p>\n\n<p>I uploaded our code and pretrained models. You may give a try. Have fun!</p>\n\n<p>p.s. \nFor sparse convolution, since spconv used in second might be patented, we replace it with <a href=\"https://github.com/facebookresearch/SparseConvNet\">this one</a>. It is much slower. </p>",
      "rawMarkdown": "Hi, the post may be a little late. Thanks to organizers and congratulations to all!\n\n# Background\nFor the record, I want to state that I am really new to Kaggle and actually really new to machine learning. I only started to work on machine learning 9 month ago. My background is in 3D Computer Graphics so I looked into the problems more in the view of 3D than in machine learning. Pardon me if I made any silly mistake with regard to machine learning.\n\nBefore this competition, I won 2nd place in CVPR 2019 competition WAD-Beyond Single-Frame Perception hosted by Baidu on a similar topic: 3D detection with Lidar points (http://wad.ai/2019/challenge.html). The tricks we used is actually based on that competition. So we just adopt and modified the tricks and have no idea about the performance without the tricks. \n\n# Models\nOur method is based on Voxelnet ([traveller59](https://github.com/traveller59/second.pytorch)) But our tricks can be applied to any other network. I gave a try with PointRCNN but got no luck. \n\nBasic Setting:\n- No external data\n- Data Augmentation: flip x and y, random rotation, random scaling and random translation\n- No Ground Truth Augmentation\n- Detection Range [-100,-100,-5,100,100,3]\n- No direction classifier\n- No specific post processing but score thresholding (0.1) and NMS\n\nTwo types of voxelnet are implemented in second: FHD and PointPillars. We use both. Based on the idea that ensembling models with more difference results in better results, we use different settings to train a set of models. E.g. we use voxel size 0.1, 0.125, 0.2, 0.25 for PointPillar, 0.1 and 0.125 for FHD. Different Voxel Feature Extractors are also used for different models(e.g. PillarFeatureNet, PillarFeatureNetRadius,PillarFeatureNetRadiusHeight). RPN is modified so that the final resolution is 200 X 200 or 250 X 250. \n\n# Key Tricks\nThe key to our method is Test time augmentation (TTA) and model ensemble with a specific 3D box fusion method.\n\n\n## **Test Time Augmentation**\nWe transform the point cloud into several copies. Each copy is then feed into the network and get the predicted boxes. Then the predicted boxes are transformed back. For example, we can rotate the point cloud 20 degrees. After we got the predicted boxes (after NMS), we rotate the predicted boxes -20 degrees (both the center position and yaw angle are changed). We do this for many copies and fuse the results. During the competition, we only use 4 copies: original, flip x, flip y, flip x and y. We tried more copies with rotation, scaling and translation. It gives slightly better results but needs more inference time.\n\n## **Model ensemble**\nWe use TTA to get the predicted boxes for each model and fuse the results.\n\nBest Single Model (No TTA): \t     &gt;0.175 (I only got the score for a model with TTA score 0.193. But I  got a better model with TTA score 0.197)\nBest Single Model (TTA):    0.197\n\nEnsembling 7 best models: 0.220\n\n## **3D box fusion**\n\nAfter we got several copies of predicted boxes (from TTA or model ensemble), we need to fuse them. We choose boxes with same label and IOU &gt; 0.6 to fuse into one. After the competition,we tried Weighted Boxes Fusion: ensembling boxes for object detection models. It actually gives better results (0.222 Public LB) but is slower. To fuse boxes, the center and size are directly weighted averaged. The key is the yaw angle. If you predict the direction and can trust it. You may use sin and cos interpolation.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3701925%2F7c611366e6f5e39e70bbbe3d0b46d12e%2Fyaw.png?generation=1577071532448853&amp;alt=media)\n\nHowever, in our implementation, we did not use direction classifier (it won’t affect the score and it can not be 100% accurate). We need to deal with one ambiguity: yaw angle is the same to yaw angle + 180 degrees. Say for one box, one prediction might give yaw angle 175 degrees or -5 degrees.  \nIt  does not matter since it is of the same IOU to ground truth. However, when we do box fusion, it matters. Say for the same box, another prediction gives yaw angle 5 degrees. If the first prediction is -5 degree. The average is 0 degree. It is okay. However, if the first prediction is 175 degree, the averaged boxes will be with yaw angle 90 degrees.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3701925%2F1b3f051743db56d52c52e61e943091c8%2Fyaw2.png?generation=1577071794662291&amp;alt=media)\n\n\nTo deal with this problem, we use an intuitive method: average 2 * yaw instead of yaw. \nThe idea is that sin and cos interpolation of yaw angle can remove 360 degrees ambiguity. So sin and cos interpolation of 2 * yaw angle can remove 180 degrees ambiguity. I have no theoretical proof yet but intuitively feel it is correct. Actually the idea came into my head when I was telling bedtime story to my daughter. She asked me to tell the story over and over again so I just repeated it😂 !\n\nI uploaded our code and pretrained models. You may give a try. Have fun!\n\np.s. \nFor sparse convolution, since spconv used in second might be patented, we replace it with [this one](https://github.com/facebookresearch/SparseConvNet). It is much slower. \n",
      "votes": 52
    },
    {
      "id": 763322,
      "postDate": "2020-03-04T11:18:16.763Z",
      "content": "<p>Thx for sharing!\nAnd , I think the ambiguity for this problem is removed because 2 * theta = 2 * theta' + pi, after the both sides of theta = theta' + pi multiplied by 2.😂</p>",
      "rawMarkdown": "Thx for sharing!\nAnd , I think the ambiguity for this problem is removed because 2 * theta = 2 * theta' + pi, after the both sides of theta = theta' + pi multiplied by 2.😂",
      "votes": 1
    },
    {
      "id": 701593,
      "postDate": "2019-12-23T16:36:11.300Z",
      "content": "<p>Congrats on winning the competition and thank you for sharing the codes/models!!</p>\n\n<blockquote>\n  <p>Actually the idea came into my head when I was telling bedtime story to my daughter</p>\n</blockquote>\n\n<p>Best thing on kaggle so far for me 😍 😍 😍 </p>",
      "rawMarkdown": "Congrats on winning the competition and thank you for sharing the codes/models!!\n&gt; Actually the idea came into my head when I was telling bedtime story to my daughter\n\nBest thing on kaggle so far for me 😍 😍 😍 ",
      "votes": 1
    },
    {
      "id": 1094209,
      "postDate": "2020-11-28T12:14:31.007Z",
      "content": "<p>Congrats on 1st place and writing this excellent post!👍</p>",
      "rawMarkdown": "Congrats on 1st place and writing this excellent post!👍"
    },
    {
      "id": 880119,
      "postDate": "2020-06-10T02:44:06.617Z",
      "content": "<p>Thanks so much for sharing the solution!</p>",
      "rawMarkdown": "Thanks so much for sharing the solution!"
    },
    {
      "id": 709543,
      "postDate": "2020-01-03T16:31:24.057Z",
      "content": "<p>Congrats on the winning the competition and thank you for sharing your approach. </p>",
      "rawMarkdown": "Congrats on the winning the competition and thank you for sharing your approach. "
    },
    {
      "id": 703743,
      "postDate": "2019-12-26T15:35:58.007Z",
      "content": "<p>Thank you very much for the solution write up and the code.</p>",
      "rawMarkdown": "Thank you very much for the solution write up and the code."
    },
    {
      "id": 701174,
      "postDate": "2019-12-23T07:26:06.943Z",
      "content": "<p>Thx for sharing!</p>",
      "rawMarkdown": "Thx for sharing!"
    },
    {
      "id": 701099,
      "postDate": "2019-12-23T04:46:33.263Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1289532,
      "postDate": "2021-05-01T05:50:49.590Z",
      "content": "<p>thanks for sharing!</p>",
      "rawMarkdown": "thanks for sharing!"
    }
  ],
  "comments": [
    {
      "id": 763322,
      "author_name": "NeoneoLeo",
      "author_url": "",
      "post_date": "2020-03-04T11:18:16.763000",
      "content": "<p>Thx for sharing!\nAnd , I think the ambiguity for this problem is removed because 2 * theta = 2 * theta' + pi, after the both sides of theta = theta' + pi multiplied by 2.😂</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 701593,
      "author_name": "Bibek",
      "author_url": "",
      "post_date": "2019-12-23T16:36:11.300000",
      "content": "<p>Congrats on winning the competition and thank you for sharing the codes/models!!</p>\n\n<blockquote>\n  <p>Actually the idea came into my head when I was telling bedtime story to my daughter</p>\n</blockquote>\n\n<p>Best thing on kaggle so far for me 😍 😍 😍 </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1094209,
      "author_name": "Beans",
      "author_url": "",
      "post_date": "2020-11-28T12:14:31.007000",
      "content": "<p>Congrats on 1st place and writing this excellent post!👍</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 880119,
      "author_name": "TongJin",
      "author_url": "",
      "post_date": "2020-06-10T02:44:06.617000",
      "content": "<p>Thanks so much for sharing the solution!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 709543,
      "author_name": "Jack Vial",
      "author_url": "",
      "post_date": "2020-01-03T16:31:24.057000",
      "content": "<p>Congrats on the winning the competition and thank you for sharing your approach. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 703743,
      "author_name": "Kishore M",
      "author_url": "",
      "post_date": "2019-12-26T15:35:58.007000",
      "content": "<p>Thank you very much for the solution write up and the code.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 701174,
      "author_name": "OceanWong",
      "author_url": "",
      "post_date": "2019-12-23T07:26:06.943000",
      "content": "<p>Thx for sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 701099,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-23T04:46:33.263000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1289532,
      "author_name": "Wonjun Park",
      "author_url": "",
      "post_date": "2021-05-01T05:50:49.590000",
      "content": "<p>thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "701081": "Hi, the post may be a little late. Thanks to organizers and congratulations to all!\n\n# Background\nFor the record, I want to state that I am really new to Kaggle and actually really new to machine learning. I only started to work on machine learning 9 month ago. My background is in 3D Computer Graphics so I looked into the problems more in the view of 3D than in machine learning. Pardon me if I made any silly mistake with regard to machine learning.\n\nBefore this competition, I won 2nd place in CVPR 2019 competition WAD-Beyond Single-Frame Perception hosted by Baidu on a similar topic: 3D detection with Lidar points (http://wad.ai/2019/challenge.html). The tricks we used is actually based on that competition. So we just adopt and modified the tricks and have no idea about the performance without the tricks. \n\n# Models\nOur method is based on Voxelnet ([traveller59](https://github.com/traveller59/second.pytorch)) But our tricks can be applied to any other network. I gave a try with PointRCNN but got no luck. \n\nBasic Setting:\n- No external data\n- Data Augmentation: flip x and y, random rotation, random scaling and random translation\n- No Ground Truth Augmentation\n- Detection Range [-100,-100,-5,100,100,3]\n- No direction classifier\n- No specific post processing but score thresholding (0.1) and NMS\n\nTwo types of voxelnet are implemented in second: FHD and PointPillars. We use both. Based on the idea that ensembling models with more difference results in better results, we use different settings to train a set of models. E.g. we use voxel size 0.1, 0.125, 0.2, 0.25 for PointPillar, 0.1 and 0.125 for FHD. Different Voxel Feature Extractors are also used for different models(e.g. PillarFeatureNet, PillarFeatureNetRadius,PillarFeatureNetRadiusHeight). RPN is modified so that the final resolution is 200 X 200 or 250 X 250. \n\n# Key Tricks\nThe key to our method is Test time augmentation (TTA) and model ensemble with a specific 3D box fusion method.\n\n\n## **Test Time Augmentation**\nWe transform the point cloud into several copies. Each copy is then feed into the network and get the predicted boxes. Then the predicted boxes are transformed back. For example, we can rotate the point cloud 20 degrees. After we got the predicted boxes (after NMS), we rotate the predicted boxes -20 degrees (both the center position and yaw angle are changed). We do this for many copies and fuse the results. During the competition, we only use 4 copies: original, flip x, flip y, flip x and y. We tried more copies with rotation, scaling and translation. It gives slightly better results but needs more inference time.\n\n## **Model ensemble**\nWe use TTA to get the predicted boxes for each model and fuse the results.\n\nBest Single Model (No TTA): \t     &gt;0.175 (I only got the score for a model with TTA score 0.193. But I  got a better model with TTA score 0.197)\nBest Single Model (TTA):    0.197\n\nEnsembling 7 best models: 0.220\n\n## **3D box fusion**\n\nAfter we got several copies of predicted boxes (from TTA or model ensemble), we need to fuse them. We choose boxes with same label and IOU &gt; 0.6 to fuse into one. After the competition,we tried Weighted Boxes Fusion: ensembling boxes for object detection models. It actually gives better results (0.222 Public LB) but is slower. To fuse boxes, the center and size are directly weighted averaged. The key is the yaw angle. If you predict the direction and can trust it. You may use sin and cos interpolation.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3701925%2F7c611366e6f5e39e70bbbe3d0b46d12e%2Fyaw.png?generation=1577071532448853&amp;alt=media)\n\nHowever, in our implementation, we did not use direction classifier (it won’t affect the score and it can not be 100% accurate). We need to deal with one ambiguity: yaw angle is the same to yaw angle + 180 degrees. Say for one box, one prediction might give yaw angle 175 degrees or -5 degrees.  \nIt  does not matter since it is of the same IOU to ground truth. However, when we do box fusion, it matters. Say for the same box, another prediction gives yaw angle 5 degrees. If the first prediction is -5 degree. The average is 0 degree. It is okay. However, if the first prediction is 175 degree, the averaged boxes will be with yaw angle 90 degrees.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3701925%2F1b3f051743db56d52c52e61e943091c8%2Fyaw2.png?generation=1577071794662291&amp;alt=media)\n\n\nTo deal with this problem, we use an intuitive method: average 2 * yaw instead of yaw. \nThe idea is that sin and cos interpolation of yaw angle can remove 360 degrees ambiguity. So sin and cos interpolation of 2 * yaw angle can remove 180 degrees ambiguity. I have no theoretical proof yet but intuitively feel it is correct. Actually the idea came into my head when I was telling bedtime story to my daughter. She asked me to tell the story over and over again so I just repeated it😂 !\n\nI uploaded our code and pretrained models. You may give a try. Have fun!\n\np.s. \nFor sparse convolution, since spconv used in second might be patented, we replace it with [this one](https://github.com/facebookresearch/SparseConvNet). It is much slower. \n",
    "763322": "Thx for sharing!\nAnd , I think the ambiguity for this problem is removed because 2 * theta = 2 * theta' + pi, after the both sides of theta = theta' + pi multiplied by 2.😂",
    "701593": "Congrats on winning the competition and thank you for sharing the codes/models!!\n&gt; Actually the idea came into my head when I was telling bedtime story to my daughter\n\nBest thing on kaggle so far for me 😍 😍 😍 ",
    "1094209": "Congrats on 1st place and writing this excellent post!👍",
    "880119": "Thanks so much for sharing the solution!",
    "709543": "Congrats on the winning the competition and thank you for sharing your approach. ",
    "703743": "Thank you very much for the solution write up and the code.",
    "701174": "Thx for sharing!",
    "701099": "",
    "1289532": "thanks for sharing!"
  }
}