{
  "id": 127052,
  "title": "9th Place Solution",
  "url": "/competitions/pku-autonomous-driving/writeups/yu4u-9th-place-solution",
  "author_name": "",
  "post_date": "2020-01-22T23:39:14.217Z",
  "votes": 33,
  "comment_count": 17,
  "views": 0,
  "content": "<p>I thank the host Peking University/Baidu and Kaggle team for holding this attractive competition, and congrats to all prize and medal winners.\nHere is a brief summary of my solution (under construction).</p>\n\n<h1>Approach</h1>\n\n<p>I used two independent models. The first model (Model A) detects cars and estimates their orientations. The second model (Model B) estimates depth map of each image. Using image coordinate (ix, iy) from Model A, depth (z) from Model B, and camera intrinsic, 3D coordinate (x, y, z) of each car is calculated. Estimating accurate depth is harder than car detection or orientation estimation. Thus I separated depth part to different dedicated model (Model B).</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F745525%2F0f0e8433bfc2d1baf24b657833a5ac6d%2Fpku.png?generation=1579657768348051&amp;alt=media\" alt=\"\"></p>\n\n<h1>Model A</h1>\n\n<p>Similar to CenterNet, but there are some modifications in targets and losses.</p>\n\n<h2>Targets</h2>\n\n<p>Model A detects cars, estimates their image coordinate (ix, iy)(not 3D camera coordinate (x, y, z) required for submission), and (yaw, pitch, roll).\nTherefore, the targets are (conf, dx, dy, sin(pitch), cos(pitch), yaw, roll).\nconf is confidence map for detecting car centers. (dx, dy) is relative position of car center in feature grid (0-1).\n(yaw, pitch, roll) is local orientation, not the orientation in camera coordinate system that is given as groudtruth. I used local orientation because original ground-truth orientation is hard to estimate from car appearance without context (camera position and image coordinate).</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F745525%2Fa34f2f2cd8c17f49c5c89b2db92bd5c2%2F2020-01-22%2010.32.44.png?generation=1579657792448454&amp;alt=media\" alt=\"\"></p>\n\n<p>[1] A. Mousavian et al., \"3D Bounding Box Estimation Using Deep Learning and Geometry,\" in Proc. of CVPR, 2017.</p>\n\n<p>This modification of orientation is done by rotation matrix that moves camera center to target car center:</p>\n\n<p><code>\nyaw = 0\npitch = -np.arctan(x / z)\nroll = np.arctan(y / z)\nr = Rotation.from_euler(\"xyz\", (roll, pitch, yaw))\n</code></p>\n\n<h2>Losses</h2>\n\n<p>Cross entropy is used for conf, and L2 loss is used for the other targets.</p>\n\n<h2>Architectures</h2>\n\n<p>I separately trained different models as Model A (four models) to detect different sizes of cars by grouping cars according to their depth; 0-25, 20-50, 40-80, and 70-180. For the former two groups (closer cars), efficientnet or se-resnext is used to get x32 downsampled feature map. For the latter two groups, segmentation models (efficientnet or se-resnext + FPN) are used from segmentation_models.pytorch to get finer feature maps (x16).</p>\n\n<h2>Augmentations</h2>\n\n<p>Augmentation is difficult part; it is related to how ground-truth of zoomed or flipped test images is created.</p>\n\n<ul>\n<li>flip (around principal point instead of image center: img[:, :3374] = cv2.flip(img[:, :3374], 1))</li>\n<li>rotation (around principal point; img = np.array(Image.fromarray(img).rotate(theta, center=(1686.2379, 1354.9849))) (-7 to 1 degrees)</li>\n<li>RandomBrightnessContrast</li>\n<li>Random scaling and crop</li>\n</ul>\n\n<h2>Optimizers</h2>\n\n<p>Trained for 160 epochs with Adam; LR = 0.0001 and decreased by 0.1 at epoch 100 and 140.</p>\n\n<h2>Ensemble</h2>\n\n<p>Several k-fold models are integrated at feature map level (model raw outputs are averaged), but it seems to not work (why...?)</p>\n\n<h1>Model B</h1>\n\n<h2>Input</h2>\n\n<p>For Model B, fixed area of input images are used for training and test: img[1558:, 23:3351]. Also, relative image coordinates from principal point is added to input in order to exploit the context of fixed camera position against ground (thus, input is gray image, dx, dy).</p>\n\n<h2>Targets and Losses</h2>\n\n<p>The second model predicts only the depth of cars. Actually, I trained Model B to predict (x, y, z) but used only z information.\nThus, target is 3D car coordinate (x, y, z). The loss is calculated only from the pixels of feature map that car centers exist.\nLoss function used here is MSE normalized by the distance from camera: ||pred - gt||_2 / ||gt||_2.\nThis selection comes from my assumption that the evaluation criterion about translation is relative rather than meter.</p>\n\n<h2>Architectures</h2>\n\n<p>Segmentation models (efficientnet or se-resnext + FPN) are used get  feature maps (x16).</p>\n\n<h2>Augmentations</h2>\n\n<ul>\n<li>flip (around principal point instead of image center: img[:, :3374] = cv2.flip(img[:, :3374], 1))</li>\n<li>RandomBrightnessContrast</li>\n</ul>\n\n<h2>Optimizers</h2>\n\n<p>Trained for 100 epochs with Adam; LR = 0.0001 and decreased by 0.1 at epoch 70.</p>\n\n<h1>Questions</h1>\n\n<p>After finishing the competition, several questions are still left to the participants. I really appreciate if the organizers answer to the following questions.</p>\n\n<ul>\n<li>How was ground-truth for flipped test image created? I guess simply done by x = -x, roll = -roll, pitch = -pitch and this is not accurate as discussed in <a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/123653\">https://www.kaggle.com/c/pku-autonomous-driving/discussion/123653</a></li>\n<li>How was ground-truth for zoomed test image created? I guess it is done by z = z / scale_factor. The other option is leave the ground truth as it is but I think this is less appropriate as we do not know new camera intrinsic.</li>\n<li>What was evaluation metric? Discussed in <a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/124489\">https://www.kaggle.com/c/pku-autonomous-driving/discussion/124489</a></li>\n</ul>",
  "messages": [
    {
      "id": "725317",
      "postDate": "01/22/2020 01:50:23",
      "content": "<p>I thank the host Peking University/Baidu and Kaggle team for holding this attractive competition, and congrats to all prize and medal winners.\nHere is a brief summary of my solution (under construction).</p>\n\n<h1>Approach</h1>\n\n<p>I used two independent models. The first model (Model A) detects cars and estimates their orientations. The second model (Model B) estimates depth map of each image. Using image coordinate (ix, iy) from Model A, depth (z) from Model B, and camera intrinsic, 3D coordinate (x, y, z) of each car is calculated. Estimating accurate depth is harder than car detection or orientation estimation. Thus I separated depth part to different dedicated model (Model B).</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F745525%2F0f0e8433bfc2d1baf24b657833a5ac6d%2Fpku.png?generation=1579657768348051&amp;alt=media\" alt=\"\"></p>\n\n<h1>Model A</h1>\n\n<p>Similar to CenterNet, but there are some modifications in targets and losses.</p>\n\n<h2>Targets</h2>\n\n<p>Model A detects cars, estimates their image coordinate (ix, iy)(not 3D camera coordinate (x, y, z) required for submission), and (yaw, pitch, roll).\nTherefore, the targets are (conf, dx, dy, sin(pitch), cos(pitch), yaw, roll).\nconf is confidence map for detecting car centers. (dx, dy) is relative position of car center in feature grid (0-1).\n(yaw, pitch, roll) is local orientation, not the orientation in camera coordinate system that is given as groudtruth. I used local orientation because original ground-truth orientation is hard to estimate from car appearance without context (camera position and image coordinate).</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F745525%2Fa34f2f2cd8c17f49c5c89b2db92bd5c2%2F2020-01-22%2010.32.44.png?generation=1579657792448454&amp;alt=media\" alt=\"\"></p>\n\n<p>[1] A. Mousavian et al., \"3D Bounding Box Estimation Using Deep Learning and Geometry,\" in Proc. of CVPR, 2017.</p>\n\n<p>This modification of orientation is done by rotation matrix that moves camera center to target car center:</p>\n\n<p><code>\nyaw = 0\npitch = -np.arctan(x / z)\nroll = np.arctan(y / z)\nr = Rotation.from_euler(\"xyz\", (roll, pitch, yaw))\n</code></p>\n\n<h2>Losses</h2>\n\n<p>Cross entropy is used for conf, and L2 loss is used for the other targets.</p>\n\n<h2>Architectures</h2>\n\n<p>I separately trained different models as Model A (four models) to detect different sizes of cars by grouping cars according to their depth; 0-25, 20-50, 40-80, and 70-180. For the former two groups (closer cars), efficientnet or se-resnext is used to get x32 downsampled feature map. For the latter two groups, segmentation models (efficientnet or se-resnext + FPN) are used from segmentation_models.pytorch to get finer feature maps (x16).</p>\n\n<h2>Augmentations</h2>\n\n<p>Augmentation is difficult part; it is related to how ground-truth of zoomed or flipped test images is created.</p>\n\n<ul>\n<li>flip (around principal point instead of image center: img[:, :3374] = cv2.flip(img[:, :3374], 1))</li>\n<li>rotation (around principal point; img = np.array(Image.fromarray(img).rotate(theta, center=(1686.2379, 1354.9849))) (-7 to 1 degrees)</li>\n<li>RandomBrightnessContrast</li>\n<li>Random scaling and crop</li>\n</ul>\n\n<h2>Optimizers</h2>\n\n<p>Trained for 160 epochs with Adam; LR = 0.0001 and decreased by 0.1 at epoch 100 and 140.</p>\n\n<h2>Ensemble</h2>\n\n<p>Several k-fold models are integrated at feature map level (model raw outputs are averaged), but it seems to not work (why...?)</p>\n\n<h1>Model B</h1>\n\n<h2>Input</h2>\n\n<p>For Model B, fixed area of input images are used for training and test: img[1558:, 23:3351]. Also, relative image coordinates from principal point is added to input in order to exploit the context of fixed camera position against ground (thus, input is gray image, dx, dy).</p>\n\n<h2>Targets and Losses</h2>\n\n<p>The second model predicts only the depth of cars. Actually, I trained Model B to predict (x, y, z) but used only z information.\nThus, target is 3D car coordinate (x, y, z). The loss is calculated only from the pixels of feature map that car centers exist.\nLoss function used here is MSE normalized by the distance from camera: ||pred - gt||_2 / ||gt||_2.\nThis selection comes from my assumption that the evaluation criterion about translation is relative rather than meter.</p>\n\n<h2>Architectures</h2>\n\n<p>Segmentation models (efficientnet or se-resnext + FPN) are used get  feature maps (x16).</p>\n\n<h2>Augmentations</h2>\n\n<ul>\n<li>flip (around principal point instead of image center: img[:, :3374] = cv2.flip(img[:, :3374], 1))</li>\n<li>RandomBrightnessContrast</li>\n</ul>\n\n<h2>Optimizers</h2>\n\n<p>Trained for 100 epochs with Adam; LR = 0.0001 and decreased by 0.1 at epoch 70.</p>\n\n<h1>Questions</h1>\n\n<p>After finishing the competition, several questions are still left to the participants. I really appreciate if the organizers answer to the following questions.</p>\n\n<ul>\n<li>How was ground-truth for flipped test image created? I guess simply done by x = -x, roll = -roll, pitch = -pitch and this is not accurate as discussed in <a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/123653\">https://www.kaggle.com/c/pku-autonomous-driving/discussion/123653</a></li>\n<li>How was ground-truth for zoomed test image created? I guess it is done by z = z / scale_factor. The other option is leave the ground truth as it is but I think this is less appropriate as we do not know new camera intrinsic.</li>\n<li>What was evaluation metric? Discussed in <a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/124489\">https://www.kaggle.com/c/pku-autonomous-driving/discussion/124489</a></li>\n</ul>",
      "rawMarkdown": "I thank the host Peking University/Baidu and Kaggle team for holding this attractive competition, and congrats to all prize and medal winners.\nHere is a brief summary of my solution (under construction).\n\n# Approach\nI used two independent models. The first model (Model A) detects cars and estimates their orientations. The second model (Model B) estimates depth map of each image. Using image coordinate (ix, iy) from Model A, depth (z) from Model B, and camera intrinsic, 3D coordinate (x, y, z) of each car is calculated. Estimating accurate depth is harder than car detection or orientation estimation. Thus I separated depth part to different dedicated model (Model B).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F745525%2F0f0e8433bfc2d1baf24b657833a5ac6d%2Fpku.png?generation=1579657768348051&amp;alt=media)\n\n\n# Model A\nSimilar to CenterNet, but there are some modifications in targets and losses.\n\n## Targets\nModel A detects cars, estimates their image coordinate (ix, iy)(not 3D camera coordinate (x, y, z) required for submission), and (yaw, pitch, roll).\nTherefore, the targets are (conf, dx, dy, sin(pitch), cos(pitch), yaw, roll).\nconf is confidence map for detecting car centers. (dx, dy) is relative position of car center in feature grid (0-1).\n(yaw, pitch, roll) is local orientation, not the orientation in camera coordinate system that is given as groudtruth. I used local orientation because original ground-truth orientation is hard to estimate from car appearance without context (camera position and image coordinate).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F745525%2Fa34f2f2cd8c17f49c5c89b2db92bd5c2%2F2020-01-22%2010.32.44.png?generation=1579657792448454&amp;alt=media)\n\n\n\n[1] A. Mousavian et al., \"3D Bounding Box Estimation Using Deep Learning and Geometry,\" in Proc. of CVPR, 2017.\n\nThis modification of orientation is done by rotation matrix that moves camera center to target car center:\n\n```\nyaw = 0\npitch = -np.arctan(x / z)\nroll = np.arctan(y / z)\nr = Rotation.from_euler(\"xyz\", (roll, pitch, yaw))\n```\n\n## Losses\nCross entropy is used for conf, and L2 loss is used for the other targets.\n\n## Architectures\nI separately trained different models as Model A (four models) to detect different sizes of cars by grouping cars according to their depth; 0-25, 20-50, 40-80, and 70-180. For the former two groups (closer cars), efficientnet or se-resnext is used to get x32 downsampled feature map. For the latter two groups, segmentation models (efficientnet or se-resnext + FPN) are used from segmentation_models.pytorch to get finer feature maps (x16).\n\n## Augmentations\nAugmentation is difficult part; it is related to how ground-truth of zoomed or flipped test images is created.\n\n- flip (around principal point instead of image center: img[:, :3374] = cv2.flip(img[:, :3374], 1))\n- rotation (around principal point; img = np.array(Image.fromarray(img).rotate(theta, center=(1686.2379, 1354.9849))) (-7 to 1 degrees)\n- RandomBrightnessContrast\n- Random scaling and crop\n\n## Optimizers\nTrained for 160 epochs with Adam; LR = 0.0001 and decreased by 0.1 at epoch 100 and 140.\n\n## Ensemble\nSeveral k-fold models are integrated at feature map level (model raw outputs are averaged), but it seems to not work (why...?)\n\n# Model B\n## Input\nFor Model B, fixed area of input images are used for training and test: img[1558:, 23:3351]. Also, relative image coordinates from principal point is added to input in order to exploit the context of fixed camera position against ground (thus, input is gray image, dx, dy).\n\n## Targets and Losses\nThe second model predicts only the depth of cars. Actually, I trained Model B to predict (x, y, z) but used only z information.\nThus, target is 3D car coordinate (x, y, z). The loss is calculated only from the pixels of feature map that car centers exist.\nLoss function used here is MSE normalized by the distance from camera: ||pred - gt||_2 / ||gt||_2.\nThis selection comes from my assumption that the evaluation criterion about translation is relative rather than meter.\n\n## Architectures\nSegmentation models (efficientnet or se-resnext + FPN) are used get  feature maps (x16).\n\n## Augmentations\n\n- flip (around principal point instead of image center: img[:, :3374] = cv2.flip(img[:, :3374], 1))\n- RandomBrightnessContrast\n\n## Optimizers\nTrained for 100 epochs with Adam; LR = 0.0001 and decreased by 0.1 at epoch 70.\n\n# Questions\nAfter finishing the competition, several questions are still left to the participants. I really appreciate if the organizers answer to the following questions.\n\n- How was ground-truth for flipped test image created? I guess simply done by x = -x, roll = -roll, pitch = -pitch and this is not accurate as discussed in https://www.kaggle.com/c/pku-autonomous-driving/discussion/123653\n- How was ground-truth for zoomed test image created? I guess it is done by z = z / scale_factor. The other option is leave the ground truth as it is but I think this is less appropriate as we do not know new camera intrinsic.\n- What was evaluation metric? Discussed in https://www.kaggle.com/c/pku-autonomous-driving/discussion/124489",
      "votes": null
    },
    {
      "id": "725343",
      "postDate": "01/22/2020 02:27:22",
      "content": "<p>Congratulations\nGreat Write-Up\nThanks for sharing your Approach &amp; Insights!! <a href=\"/ren4yu\">@ren4yu</a> </p>",
      "rawMarkdown": "Congratulations\nGreat Write-Up\nThanks for sharing your Approach &amp; Insights!! @ren4yu",
      "votes": null
    },
    {
      "id": "725347",
      "postDate": "01/22/2020 02:32:47",
      "content": "<p>Awesome, thanks for share</p>",
      "rawMarkdown": "Awesome, thanks for share",
      "votes": null
    },
    {
      "id": "725353",
      "postDate": "01/22/2020 02:36:31",
      "content": "<p>Interesting approach and congrats on solo gold!\nThere are still a lot of problems with this comp but I doubt the organizers will doing anything as there has been almost zero interaction with them.</p>",
      "rawMarkdown": "Interesting approach and congrats on solo gold!\nThere are still a lot of problems with this comp but I doubt the organizers will doing anything as there has been almost zero interaction with them.",
      "votes": null
    },
    {
      "id": "725362",
      "postDate": "01/22/2020 02:59:37",
      "content": "<p>can you tell me your depth average l1 loss ？</p>",
      "rawMarkdown": "can you tell me your depth average l1 loss ？",
      "votes": null
    },
    {
      "id": "725379",
      "postDate": "01/22/2020 03:26:54",
      "content": "<p>Actually, in addition to depth (z), x and y are also regressed with normalized MSE;\n((x - x')^2 + (y - y')^2 + (z - z')^2) / (x^2 + y^2 + z^2),\nwhere (x, y, z) is gt and (x', y', z') is prediction.</p>",
      "rawMarkdown": "Actually, in addition to depth (z), x and y are also regressed with normalized MSE;\n((x - x')^2 + (y - y')^2 + (z - z')^2) / (x^2 + y^2 + z^2),\nwhere (x, y, z) is gt and (x', y', z') is prediction.",
      "votes": null
    },
    {
      "id": "725387",
      "postDate": "01/22/2020 03:32:50",
      "content": "<p>Thanks for sharing this interesting and unique solution!</p>\n\n<p>About flipped test image:\nI think that flipped images in test data are all dummy.\nI made flip prediction model and predicted which images are flipped in test data.\nThen I removed predictions for flipped images from submissions and got same scores for both public and private.</p>\n\n<p>I'm guessing that images with noise is the same situation.</p>\n\n<p>Anyway it would be nice if organizers clarify this.</p>",
      "rawMarkdown": "Thanks for sharing this interesting and unique solution!\n\nAbout flipped test image:\nI think that flipped images in test data are all dummy.\nI made flip prediction model and predicted which images are flipped in test data.\nThen I removed predictions for flipped images from submissions and got same scores for both public and private.\n\nI'm guessing that images with noise is the same situation.\n\nAnyway it would be nice if organizers clarify this.",
      "votes": null
    },
    {
      "id": "725413",
      "postDate": "01/22/2020 04:18:16",
      "content": "<p>so. at the train or test pipline, you dont estimate the abs error of deep? \nfor me, at independence test set, the abs deep error is 5m more or less. the train set is 3m more or less.</p>",
      "rawMarkdown": "so. at the train or test pipline, you dont estimate the abs error of deep? \nfor me, at independence test set, the abs deep error is 5m more or less. the train set is 3m more or less.",
      "votes": null
    },
    {
      "id": "725464",
      "postDate": "01/22/2020 06:04:47",
      "content": "<p>Congratulations！\nThank you for sharing your thoughts！</p>",
      "rawMarkdown": "Congratulations！\nThank you for sharing your thoughts！",
      "votes": null
    },
    {
      "id": "725692",
      "postDate": "01/22/2020 11:55:10",
      "content": "<p>Yes, I did not check abs error because it seems to be not evaluation metric.\nAbs error should highly depend on the distance from camera (or z coordinate); distant cars are relatively difficult.</p>",
      "rawMarkdown": "Yes, I did not check abs error because it seems to be not evaluation metric.\nAbs error should highly depend on the distance from camera (or z coordinate); distant cars are relatively difficult.",
      "votes": null
    },
    {
      "id": "725694",
      "postDate": "01/22/2020 11:56:03",
      "content": "<p>I see. Really nice inspiration for checking this...</p>",
      "rawMarkdown": "I see. Really nice inspiration for checking this...",
      "votes": null
    },
    {
      "id": "725886",
      "postDate": "01/22/2020 15:18:42",
      "content": "<p>Related to the targets of Model A, I just published to calculate local orientation and search for cars with similar orientations.\n<a href=\"https://www.kaggle.com/ren4yu/show-cars-with-similar-local-orientations\">https://www.kaggle.com/ren4yu/show-cars-with-similar-local-orientations</a></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F745525%2Fecd99b3ac37448074ccf1070e15e09fb%2F2020-01-23%200.17.00.png?generation=1579706306189090&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Related to the targets of Model A, I just published to calculate local orientation and search for cars with similar orientations.\nhttps://www.kaggle.com/ren4yu/show-cars-with-similar-local-orientations\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F745525%2Fecd99b3ac37448074ccf1070e15e09fb%2F2020-01-23%200.17.00.png?generation=1579706306189090&amp;alt=media)",
      "votes": null
    },
    {
      "id": "725909",
      "postDate": "01/22/2020 15:43:37",
      "content": "<p>Thanks for sharing interesting and great solution! I have questions below:\n- How much gain did you get by the change to local orientation? In my understanding if we use centernet-like architecture, model learns the local context implicitly. And even when I use cropped image and prediction only angles, that was actually not so bad.\n- Was it also better to group cars by size and develop models for each compared to single model? Intuitively applying finer model to bigger object too is not bad idea.</p>\n\n<p>Thank you in advance! I learned a lot:)</p>",
      "rawMarkdown": "Thanks for sharing interesting and great solution! I have questions below:\n- How much gain did you get by the change to local orientation? In my understanding if we use centernet-like architecture, model learns the local context implicitly. And even when I use cropped image and prediction only angles, that was actually not so bad.\n- Was it also better to group cars by size and develop models for each compared to single model? Intuitively applying finer model to bigger object too is not bad idea.\n\nThank you in advance! I learned a lot:)",
      "votes": null
    },
    {
      "id": "729702",
      "postDate": "01/26/2020 14:42:25",
      "content": "<p>Sorry for late reply.</p>\n\n<ul>\n<li>I did not use orientation in camera coordinate throughout this competition; I believed local orientation is more robust, especially for infrequent orientations (pitch is far from 0, +-pi).</li>\n<li>The appearance sizes of cars are really different depending on the distance from camera. For the larger car models, I downsampled input images x0.5 and x0.3334 for training. Using shared feature map for these different sizes of objects are really mysterious design for me and it is somewhat surprising that the original centernet have achieved good results. I think multi-scale model would be better choice, but I used independent scale models for simplicity.</li>\n</ul>",
      "rawMarkdown": "Sorry for late reply.\n\n- I did not use orientation in camera coordinate throughout this competition; I believed local orientation is more robust, especially for infrequent orientations (pitch is far from 0, +-pi).\n- The appearance sizes of cars are really different depending on the distance from camera. For the larger car models, I downsampled input images x0.5 and x0.3334 for training. Using shared feature map for these different sizes of objects are really mysterious design for me and it is somewhat surprising that the original centernet have achieved good results. I think multi-scale model would be better choice, but I used independent scale models for simplicity.",
      "votes": null
    },
    {
      "id": "729932",
      "postDate": "01/26/2020 21:04:25",
      "content": "<p>Beautiful solution 💪 </p>",
      "rawMarkdown": "Beautiful solution 💪",
      "votes": null
    },
    {
      "id": "729946",
      "postDate": "01/26/2020 21:42:19",
      "content": "<p>I Like</p>",
      "rawMarkdown": "I Like",
      "votes": null
    },
    {
      "id": "732543",
      "postDate": "01/29/2020 23:32:01",
      "content": "<p>Congrats and thanks for sharing!</p>",
      "rawMarkdown": "Congrats and thanks for sharing!",
      "votes": null
    },
    {
      "id": "736601",
      "postDate": "02/04/2020 11:28:13",
      "content": "<p>Congrats!</p>",
      "rawMarkdown": "Congrats!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 725343,
      "author_name": "veeralakrishna",
      "author_url": "",
      "post_date": "01/22/2020 02:27:22",
      "content": "<p>Congratulations\nGreat Write-Up\nThanks for sharing your Approach &amp; Insights!! <a href=\"/ren4yu\">@ren4yu</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 725347,
      "author_name": "azamatk",
      "author_url": "",
      "post_date": "01/22/2020 02:32:47",
      "content": "<p>Awesome, thanks for share</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 725353,
      "author_name": "greatgamedota",
      "author_url": "",
      "post_date": "01/22/2020 02:36:31",
      "content": "<p>Interesting approach and congrats on solo gold!\nThere are still a lot of problems with this comp but I doubt the organizers will doing anything as there has been almost zero interaction with them.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 725362,
      "author_name": "pp2file",
      "author_url": "",
      "post_date": "01/22/2020 02:59:37",
      "content": "<p>can you tell me your depth average l1 loss ？</p>",
      "votes": null,
      "replies": [
        {
          "id": 725379,
          "author_name": "ren4yu",
          "author_url": "",
          "post_date": "01/22/2020 03:26:54",
          "content": "<p>Actually, in addition to depth (z), x and y are also regressed with normalized MSE;\n((x - x')^2 + (y - y')^2 + (z - z')^2) / (x^2 + y^2 + z^2),\nwhere (x, y, z) is gt and (x', y', z') is prediction.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 725413,
          "author_name": "pp2file",
          "author_url": "",
          "post_date": "01/22/2020 04:18:16",
          "content": "<p>so. at the train or test pipline, you dont estimate the abs error of deep? \nfor me, at independence test set, the abs deep error is 5m more or less. the train set is 3m more or less.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 725692,
          "author_name": "ren4yu",
          "author_url": "",
          "post_date": "01/22/2020 11:55:10",
          "content": "<p>Yes, I did not check abs error because it seems to be not evaluation metric.\nAbs error should highly depend on the distance from camera (or z coordinate); distant cars are relatively difficult.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 725387,
      "author_name": "its7171",
      "author_url": "",
      "post_date": "01/22/2020 03:32:50",
      "content": "<p>Thanks for sharing this interesting and unique solution!</p>\n\n<p>About flipped test image:\nI think that flipped images in test data are all dummy.\nI made flip prediction model and predicted which images are flipped in test data.\nThen I removed predictions for flipped images from submissions and got same scores for both public and private.</p>\n\n<p>I'm guessing that images with noise is the same situation.</p>\n\n<p>Anyway it would be nice if organizers clarify this.</p>",
      "votes": null,
      "replies": [
        {
          "id": 725694,
          "author_name": "ren4yu",
          "author_url": "",
          "post_date": "01/22/2020 11:56:03",
          "content": "<p>I see. Really nice inspiration for checking this...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 725464,
      "author_name": "hirokitanio",
      "author_url": "",
      "post_date": "01/22/2020 06:04:47",
      "content": "<p>Congratulations！\nThank you for sharing your thoughts！</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 725886,
      "author_name": "ren4yu",
      "author_url": "",
      "post_date": "01/22/2020 15:18:42",
      "content": "<p>Related to the targets of Model A, I just published to calculate local orientation and search for cars with similar orientations.\n<a href=\"https://www.kaggle.com/ren4yu/show-cars-with-similar-local-orientations\">https://www.kaggle.com/ren4yu/show-cars-with-similar-local-orientations</a></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F745525%2Fecd99b3ac37448074ccf1070e15e09fb%2F2020-01-23%200.17.00.png?generation=1579706306189090&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 725909,
      "author_name": "bamps53",
      "author_url": "",
      "post_date": "01/22/2020 15:43:37",
      "content": "<p>Thanks for sharing interesting and great solution! I have questions below:\n- How much gain did you get by the change to local orientation? In my understanding if we use centernet-like architecture, model learns the local context implicitly. And even when I use cropped image and prediction only angles, that was actually not so bad.\n- Was it also better to group cars by size and develop models for each compared to single model? Intuitively applying finer model to bigger object too is not bad idea.</p>\n\n<p>Thank you in advance! I learned a lot:)</p>",
      "votes": null,
      "replies": [
        {
          "id": 729702,
          "author_name": "ren4yu",
          "author_url": "",
          "post_date": "01/26/2020 14:42:25",
          "content": "<p>Sorry for late reply.</p>\n\n<ul>\n<li>I did not use orientation in camera coordinate throughout this competition; I believed local orientation is more robust, especially for infrequent orientations (pitch is far from 0, +-pi).</li>\n<li>The appearance sizes of cars are really different depending on the distance from camera. For the larger car models, I downsampled input images x0.5 and x0.3334 for training. Using shared feature map for these different sizes of objects are really mysterious design for me and it is somewhat surprising that the original centernet have achieved good results. I think multi-scale model would be better choice, but I used independent scale models for simplicity.</li>\n</ul>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 729932,
      "author_name": "alvaroibrain",
      "author_url": "",
      "post_date": "01/26/2020 21:04:25",
      "content": "<p>Beautiful solution 💪 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 729946,
      "author_name": "ijelliti",
      "author_url": "",
      "post_date": "01/26/2020 21:42:19",
      "content": "<p>I Like</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 732543,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "01/29/2020 23:32:01",
      "content": "<p>Congrats and thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 736601,
      "author_name": "",
      "author_url": "",
      "post_date": "02/04/2020 11:28:13",
      "content": "<p>Congrats!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "725317": "I thank the host Peking University/Baidu and Kaggle team for holding this attractive competition, and congrats to all prize and medal winners.\nHere is a brief summary of my solution (under construction).\n\n# Approach\nI used two independent models. The first model (Model A) detects cars and estimates their orientations. The second model (Model B) estimates depth map of each image. Using image coordinate (ix, iy) from Model A, depth (z) from Model B, and camera intrinsic, 3D coordinate (x, y, z) of each car is calculated. Estimating accurate depth is harder than car detection or orientation estimation. Thus I separated depth part to different dedicated model (Model B).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F745525%2F0f0e8433bfc2d1baf24b657833a5ac6d%2Fpku.png?generation=1579657768348051&amp;alt=media)\n\n\n# Model A\nSimilar to CenterNet, but there are some modifications in targets and losses.\n\n## Targets\nModel A detects cars, estimates their image coordinate (ix, iy)(not 3D camera coordinate (x, y, z) required for submission), and (yaw, pitch, roll).\nTherefore, the targets are (conf, dx, dy, sin(pitch), cos(pitch), yaw, roll).\nconf is confidence map for detecting car centers. (dx, dy) is relative position of car center in feature grid (0-1).\n(yaw, pitch, roll) is local orientation, not the orientation in camera coordinate system that is given as groudtruth. I used local orientation because original ground-truth orientation is hard to estimate from car appearance without context (camera position and image coordinate).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F745525%2Fa34f2f2cd8c17f49c5c89b2db92bd5c2%2F2020-01-22%2010.32.44.png?generation=1579657792448454&amp;alt=media)\n\n\n\n[1] A. Mousavian et al., \"3D Bounding Box Estimation Using Deep Learning and Geometry,\" in Proc. of CVPR, 2017.\n\nThis modification of orientation is done by rotation matrix that moves camera center to target car center:\n\n```\nyaw = 0\npitch = -np.arctan(x / z)\nroll = np.arctan(y / z)\nr = Rotation.from_euler(\"xyz\", (roll, pitch, yaw))\n```\n\n## Losses\nCross entropy is used for conf, and L2 loss is used for the other targets.\n\n## Architectures\nI separately trained different models as Model A (four models) to detect different sizes of cars by grouping cars according to their depth; 0-25, 20-50, 40-80, and 70-180. For the former two groups (closer cars), efficientnet or se-resnext is used to get x32 downsampled feature map. For the latter two groups, segmentation models (efficientnet or se-resnext + FPN) are used from segmentation_models.pytorch to get finer feature maps (x16).\n\n## Augmentations\nAugmentation is difficult part; it is related to how ground-truth of zoomed or flipped test images is created.\n\n- flip (around principal point instead of image center: img[:, :3374] = cv2.flip(img[:, :3374], 1))\n- rotation (around principal point; img = np.array(Image.fromarray(img).rotate(theta, center=(1686.2379, 1354.9849))) (-7 to 1 degrees)\n- RandomBrightnessContrast\n- Random scaling and crop\n\n## Optimizers\nTrained for 160 epochs with Adam; LR = 0.0001 and decreased by 0.1 at epoch 100 and 140.\n\n## Ensemble\nSeveral k-fold models are integrated at feature map level (model raw outputs are averaged), but it seems to not work (why...?)\n\n# Model B\n## Input\nFor Model B, fixed area of input images are used for training and test: img[1558:, 23:3351]. Also, relative image coordinates from principal point is added to input in order to exploit the context of fixed camera position against ground (thus, input is gray image, dx, dy).\n\n## Targets and Losses\nThe second model predicts only the depth of cars. Actually, I trained Model B to predict (x, y, z) but used only z information.\nThus, target is 3D car coordinate (x, y, z). The loss is calculated only from the pixels of feature map that car centers exist.\nLoss function used here is MSE normalized by the distance from camera: ||pred - gt||_2 / ||gt||_2.\nThis selection comes from my assumption that the evaluation criterion about translation is relative rather than meter.\n\n## Architectures\nSegmentation models (efficientnet or se-resnext + FPN) are used get  feature maps (x16).\n\n## Augmentations\n\n- flip (around principal point instead of image center: img[:, :3374] = cv2.flip(img[:, :3374], 1))\n- RandomBrightnessContrast\n\n## Optimizers\nTrained for 100 epochs with Adam; LR = 0.0001 and decreased by 0.1 at epoch 70.\n\n# Questions\nAfter finishing the competition, several questions are still left to the participants. I really appreciate if the organizers answer to the following questions.\n\n- How was ground-truth for flipped test image created? I guess simply done by x = -x, roll = -roll, pitch = -pitch and this is not accurate as discussed in https://www.kaggle.com/c/pku-autonomous-driving/discussion/123653\n- How was ground-truth for zoomed test image created? I guess it is done by z = z / scale_factor. The other option is leave the ground truth as it is but I think this is less appropriate as we do not know new camera intrinsic.\n- What was evaluation metric? Discussed in https://www.kaggle.com/c/pku-autonomous-driving/discussion/124489",
    "725343": "Congratulations\nGreat Write-Up\nThanks for sharing your Approach &amp; Insights!! @ren4yu",
    "725347": "Awesome, thanks for share",
    "725353": "Interesting approach and congrats on solo gold!\nThere are still a lot of problems with this comp but I doubt the organizers will doing anything as there has been almost zero interaction with them.",
    "725362": "can you tell me your depth average l1 loss ？",
    "725379": "Actually, in addition to depth (z), x and y are also regressed with normalized MSE;\n((x - x')^2 + (y - y')^2 + (z - z')^2) / (x^2 + y^2 + z^2),\nwhere (x, y, z) is gt and (x', y', z') is prediction.",
    "725387": "Thanks for sharing this interesting and unique solution!\n\nAbout flipped test image:\nI think that flipped images in test data are all dummy.\nI made flip prediction model and predicted which images are flipped in test data.\nThen I removed predictions for flipped images from submissions and got same scores for both public and private.\n\nI'm guessing that images with noise is the same situation.\n\nAnyway it would be nice if organizers clarify this.",
    "725413": "so. at the train or test pipline, you dont estimate the abs error of deep? \nfor me, at independence test set, the abs deep error is 5m more or less. the train set is 3m more or less.",
    "725464": "Congratulations！\nThank you for sharing your thoughts！",
    "725692": "Yes, I did not check abs error because it seems to be not evaluation metric.\nAbs error should highly depend on the distance from camera (or z coordinate); distant cars are relatively difficult.",
    "725694": "I see. Really nice inspiration for checking this...",
    "725886": "Related to the targets of Model A, I just published to calculate local orientation and search for cars with similar orientations.\nhttps://www.kaggle.com/ren4yu/show-cars-with-similar-local-orientations\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F745525%2Fecd99b3ac37448074ccf1070e15e09fb%2F2020-01-23%200.17.00.png?generation=1579706306189090&amp;alt=media)",
    "725909": "Thanks for sharing interesting and great solution! I have questions below:\n- How much gain did you get by the change to local orientation? In my understanding if we use centernet-like architecture, model learns the local context implicitly. And even when I use cropped image and prediction only angles, that was actually not so bad.\n- Was it also better to group cars by size and develop models for each compared to single model? Intuitively applying finer model to bigger object too is not bad idea.\n\nThank you in advance! I learned a lot:)",
    "729702": "Sorry for late reply.\n\n- I did not use orientation in camera coordinate throughout this competition; I believed local orientation is more robust, especially for infrequent orientations (pitch is far from 0, +-pi).\n- The appearance sizes of cars are really different depending on the distance from camera. For the larger car models, I downsampled input images x0.5 and x0.3334 for training. Using shared feature map for these different sizes of objects are really mysterious design for me and it is somewhat surprising that the original centernet have achieved good results. I think multi-scale model would be better choice, but I used independent scale models for simplicity.",
    "729932": "Beautiful solution 💪",
    "729946": "I Like",
    "732543": "Congrats and thanks for sharing!",
    "736601": "Congrats!"
  },
  "source": "meta"
}