{
  "id": 417191,
  "title": "3rd Place Solution - Significantly Reduced the Fluctuations caused by Randomness!",
  "url": "/competitions/image-matching-challenge-2023/writeups/3dvisualmapping-3rd-place-solution-significantly-r",
  "author_name": "",
  "post_date": "2023-06-15T02:25:33.887Z",
  "votes": 29,
  "comment_count": 6,
  "views": 0,
  "content": "<p>We are delighted to participate in this competition and would like to express gratitude to all the Kaggle staffs and Sponsors. Congratulations to all the participants. <br>\nThe team members include 陈鹏、陈建国、阮志伟 and 李伟. I would like to express my sincere gratitude to everyone for the excellent teamwork over the past month. I have thoroughly enjoyed working with all of you, and I am delighted to be a part of this team.</p>\n<h1>1 Overview</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3163774%2Fc8338666681788720689a7ead39a92ab%2F.png?generation=1686757921566293&amp;alt=media\" alt=\"pipeline\"></p>\n<h1>2 Main pipeline</h1>\n<h2>2.1 SP/SG</h2>\n<h3>2.1.1 Rotation</h3>\n<p>Rotating the image has a significant effect, since SG is lack of the rotation invariance. Therefore, for each image pair A-B, we fixed image A and rotated image B four times (0, 90, 180, 270). After performing four times SG matching, we selected the rotation angle with the most matches for the next stage. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3163774%2F33aa1271502732f84ca6420949ed19a6%2F.png?generation=1686758077155244&amp;alt=media\" alt=\"Rotation\"></p>\n<h3>2.1.2 Resize image</h3>\n<p>In the early stage of the competition, we used the original image for extracting keypoints and matching. And we noticed that the mAA score in the heritage cyprus scene was only 0.1. After experiments, we found that if we scaled the images in cyprus scene to 1920 on the longer side, the mAA score was improved to 0.6. Specifically, we didn't rescale the image size after SP keypoints extraction. Finally, in our application, we scaled the image size if the longer side was larger than 1920, otherwise kept the original image size for the following SP + SG inference.</p>\n<h3>2.1.3 SP/SG setting</h3>\n<p>We increased the NMS value of SP from 3 to 8 and set the maximum number of keypoints to 4000.</p>\n<h2>2.2 GeoVerification（RANSAC）</h2>\n<p>We used USAC_MAGSAC for geometric verification with the following configuration.<br>\ncv2.findFundamentalMat(mkpts0, mkpts1, cv2.USAC_MAGSAC, 2, 0.99999, 100000)</p>\n<h2>2.3 Setting Camera Params（Randomness）</h2>\n<p>In our experiments, we wanted to eliminate the randomness effect in our final mAA, since we observed that the same notebook could result in fluctuations of approximately 0.03 in the LB.  We found that the randomness come mainly from the Ceres (<a href=\"https://github.com/colmap/colmap/issues/404)\" target=\"_blank\">https://github.com/colmap/colmap/issues/404)</a>) optimization. <br>\nAfter analysis, the initial values of camera parameters, such as the focal length, had a significant impact on the stability of optimization results.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3163774%2F9411c508f380417b835bb253b06eae7b%2Fwall_random.png?generation=1686758232496900&amp;alt=media\" alt=\"randomness\"><br>\nSince the camera focal length information could be extracted from the image EXIF data, we provided the prior camera focal length to the camera and shared the same camera settings among images captured by the same device. Although there were still some fluctuations in the metrics, the majority of the experimental results were consistent. <br>\nThe figure below shows four submissions of the same final notebook. <strong>On the public LB, our metric fluctuates by no more than 0.004.</strong> Moreover, compared to many other teams, <strong>we do not have significant metric fluctuations between public LB and private LB.</strong> Our public LB score was amazingly close to private score. This might indicated that our method <strong>Significantly Reduced the Fluctuations caused by Randomness!</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3163774%2Ffd18321289d3745de78691742d0ca28d%2Fonline_maa.png?generation=1686758332640650&amp;alt=media\" alt=\"mAA\"></p>\n<h2>2.4 Mapper</h2>\n<p>In the COLMAP reconstruction process, we revised some default parameters in incremental mapping. After first trail, if the best model register ratio was below 1, we relaxed the mapper configuration (abs_pose_min_num_inliers=15) and re-run incremental mapping.</p>\n<h1>3 Conclusion</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3163774%2F47522f20ba1cf5b608a1806f0cf761f4%2Fconclution.png?generation=1686760502421586&amp;alt=media\" alt=\"Conclusion\"></p>\n<h1>4 Not fully tested</h1>\n<h2>4.1 Select image pairs</h2>\n<p>Our final solution took nearly 9 hours,  since we generated the image matching pairs by exhaustive method. In order to speed up our solution, we tried Efficientnet_b7, Convnextv2_huge, and DINOv2 to extract image features and generate image matching pairs by feature similarity. In offline experiments, we selected the top N/2 most similar images (N being the total number of images) for each image to form image pairs. Compared to other methods, DINOv2 performed the best. By incorporating DINOv2 into the image pairs selection, we could control the processing time to 7 hours, and achieved 0.485 in LB (compared with exhaustive 0.507).<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3163774%2Fbfdb729a111a6633b27e573788e09263%2Fdinov2.png?generation=1686760519524067&amp;alt=media\" alt=\"dinov2\"></p>\n<h2>4.2 Pixel-Perfect Structure-from-Motion</h2>\n<p>PixSFM had a good performance improvement in our local validation. However, during online testing,  it consistently times out even we run it on scenes less than 50 images.  Moreover, it took a lot time for us to install the environment and run it on kaggle. Times out Sad! </p>\n<h1>5 Ideas that did not work well</h1>\n<h2>5.1 Different detectors and matchers</h2>\n<p>We tested DKMv3, GlueStick, and SILK, but neither was able to surpass SPSG. In our experiments, we observed that DKMv3 performed better than SPSG in challenging scenes, such as those with large viewpoint differences or rotations. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3163774%2Ffae472d5266e521f18f667592eaf3454%2Fdkm.png?generation=1686758763195086&amp;alt=media\" alt=\"dkm\"><br>\nHowever, the overall metrics of DKMv3 were not as good as SPSG, which may be due to the operation of sampling dense matches we performed.</p>\n<h2>5.2 Crop image</h2>\n<p>Just like last year's winning solution, we took a crop after completing the first stage of matching. In the second stage, we performed matching again on the crop and merge all the matching results together.</p>\n<h2>5.3 Half precision</h2>\n<p>It confused our team a lot. In the kaggle notebook, we found the half-precision improved the results, but on the public LB, it returned the lower results. </p>\n<h2>5.4 Merge matches in four directions</h2>\n<p>In our methods, we rotated the image (0, 90, 180, 270) and performed four times matching. Instead of keeping all matches from four directions, we only keep the best angle matches, because keeping all matches didn't lead to any improvement in the results.</p>\n<h2>5.5 findHomography</h2>\n<p>When feature points lie on the same plane (e.g., in a wall scene) or when the camera undergoes pure rotation, the fundamental matrix degenerates. Therefore, in RANSAC, we simultaneously used the findFundamentalMat and the findHomography to calculate the inliers. However, this approach didn't lead to an improvement in the metrics.</p>\n<h2>5.6 3D Model Refinement</h2>\n<p>After the first trail of mapping, we tried to filter nosiy 3D points with large projection error or short track length, and then re-bundle adjust the model. Experically, this could help export better poses, but.. that's life.</p>",
  "messages": [
    {
      "id": "2302561",
      "postDate": "06/14/2023 16:07:50",
      "content": "<p>We are delighted to participate in this competition and would like to express gratitude to all the Kaggle staffs and Sponsors. Congratulations to all the participants. <br>\nThe team members include 陈鹏、陈建国、阮志伟 and 李伟. I would like to express my sincere gratitude to everyone for the excellent teamwork over the past month. I have thoroughly enjoyed working with all of you, and I am delighted to be a part of this team.</p>\n<h1>1 Overview</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3163774%2Fc8338666681788720689a7ead39a92ab%2F.png?generation=1686757921566293&amp;alt=media\" alt=\"pipeline\"></p>\n<h1>2 Main pipeline</h1>\n<h2>2.1 SP/SG</h2>\n<h3>2.1.1 Rotation</h3>\n<p>Rotating the image has a significant effect, since SG is lack of the rotation invariance. Therefore, for each image pair A-B, we fixed image A and rotated image B four times (0, 90, 180, 270). After performing four times SG matching, we selected the rotation angle with the most matches for the next stage. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3163774%2F33aa1271502732f84ca6420949ed19a6%2F.png?generation=1686758077155244&amp;alt=media\" alt=\"Rotation\"></p>\n<h3>2.1.2 Resize image</h3>\n<p>In the early stage of the competition, we used the original image for extracting keypoints and matching. And we noticed that the mAA score in the heritage cyprus scene was only 0.1. After experiments, we found that if we scaled the images in cyprus scene to 1920 on the longer side, the mAA score was improved to 0.6. Specifically, we didn't rescale the image size after SP keypoints extraction. Finally, in our application, we scaled the image size if the longer side was larger than 1920, otherwise kept the original image size for the following SP + SG inference.</p>\n<h3>2.1.3 SP/SG setting</h3>\n<p>We increased the NMS value of SP from 3 to 8 and set the maximum number of keypoints to 4000.</p>\n<h2>2.2 GeoVerification（RANSAC）</h2>\n<p>We used USAC_MAGSAC for geometric verification with the following configuration.<br>\ncv2.findFundamentalMat(mkpts0, mkpts1, cv2.USAC_MAGSAC, 2, 0.99999, 100000)</p>\n<h2>2.3 Setting Camera Params（Randomness）</h2>\n<p>In our experiments, we wanted to eliminate the randomness effect in our final mAA, since we observed that the same notebook could result in fluctuations of approximately 0.03 in the LB.  We found that the randomness come mainly from the Ceres (<a href=\"https://github.com/colmap/colmap/issues/404)\" target=\"_blank\">https://github.com/colmap/colmap/issues/404)</a>) optimization. <br>\nAfter analysis, the initial values of camera parameters, such as the focal length, had a significant impact on the stability of optimization results.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3163774%2F9411c508f380417b835bb253b06eae7b%2Fwall_random.png?generation=1686758232496900&amp;alt=media\" alt=\"randomness\"><br>\nSince the camera focal length information could be extracted from the image EXIF data, we provided the prior camera focal length to the camera and shared the same camera settings among images captured by the same device. Although there were still some fluctuations in the metrics, the majority of the experimental results were consistent. <br>\nThe figure below shows four submissions of the same final notebook. <strong>On the public LB, our metric fluctuates by no more than 0.004.</strong> Moreover, compared to many other teams, <strong>we do not have significant metric fluctuations between public LB and private LB.</strong> Our public LB score was amazingly close to private score. This might indicated that our method <strong>Significantly Reduced the Fluctuations caused by Randomness!</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3163774%2Ffd18321289d3745de78691742d0ca28d%2Fonline_maa.png?generation=1686758332640650&amp;alt=media\" alt=\"mAA\"></p>\n<h2>2.4 Mapper</h2>\n<p>In the COLMAP reconstruction process, we revised some default parameters in incremental mapping. After first trail, if the best model register ratio was below 1, we relaxed the mapper configuration (abs_pose_min_num_inliers=15) and re-run incremental mapping.</p>\n<h1>3 Conclusion</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3163774%2F47522f20ba1cf5b608a1806f0cf761f4%2Fconclution.png?generation=1686760502421586&amp;alt=media\" alt=\"Conclusion\"></p>\n<h1>4 Not fully tested</h1>\n<h2>4.1 Select image pairs</h2>\n<p>Our final solution took nearly 9 hours,  since we generated the image matching pairs by exhaustive method. In order to speed up our solution, we tried Efficientnet_b7, Convnextv2_huge, and DINOv2 to extract image features and generate image matching pairs by feature similarity. In offline experiments, we selected the top N/2 most similar images (N being the total number of images) for each image to form image pairs. Compared to other methods, DINOv2 performed the best. By incorporating DINOv2 into the image pairs selection, we could control the processing time to 7 hours, and achieved 0.485 in LB (compared with exhaustive 0.507).<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3163774%2Fbfdb729a111a6633b27e573788e09263%2Fdinov2.png?generation=1686760519524067&amp;alt=media\" alt=\"dinov2\"></p>\n<h2>4.2 Pixel-Perfect Structure-from-Motion</h2>\n<p>PixSFM had a good performance improvement in our local validation. However, during online testing,  it consistently times out even we run it on scenes less than 50 images.  Moreover, it took a lot time for us to install the environment and run it on kaggle. Times out Sad! </p>\n<h1>5 Ideas that did not work well</h1>\n<h2>5.1 Different detectors and matchers</h2>\n<p>We tested DKMv3, GlueStick, and SILK, but neither was able to surpass SPSG. In our experiments, we observed that DKMv3 performed better than SPSG in challenging scenes, such as those with large viewpoint differences or rotations. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3163774%2Ffae472d5266e521f18f667592eaf3454%2Fdkm.png?generation=1686758763195086&amp;alt=media\" alt=\"dkm\"><br>\nHowever, the overall metrics of DKMv3 were not as good as SPSG, which may be due to the operation of sampling dense matches we performed.</p>\n<h2>5.2 Crop image</h2>\n<p>Just like last year's winning solution, we took a crop after completing the first stage of matching. In the second stage, we performed matching again on the crop and merge all the matching results together.</p>\n<h2>5.3 Half precision</h2>\n<p>It confused our team a lot. In the kaggle notebook, we found the half-precision improved the results, but on the public LB, it returned the lower results. </p>\n<h2>5.4 Merge matches in four directions</h2>\n<p>In our methods, we rotated the image (0, 90, 180, 270) and performed four times matching. Instead of keeping all matches from four directions, we only keep the best angle matches, because keeping all matches didn't lead to any improvement in the results.</p>\n<h2>5.5 findHomography</h2>\n<p>When feature points lie on the same plane (e.g., in a wall scene) or when the camera undergoes pure rotation, the fundamental matrix degenerates. Therefore, in RANSAC, we simultaneously used the findFundamentalMat and the findHomography to calculate the inliers. However, this approach didn't lead to an improvement in the metrics.</p>\n<h2>5.6 3D Model Refinement</h2>\n<p>After the first trail of mapping, we tried to filter nosiy 3D points with large projection error or short track length, and then re-bundle adjust the model. Experically, this could help export better poses, but.. that's life.</p>",
      "rawMarkdown": "We are delighted to participate in this competition and would like to express gratitude to all the Kaggle staffs and Sponsors. Congratulations to all the participants. \nThe team members include 陈鹏、陈建国、阮志伟 and 李伟. I would like to express my sincere gratitude to everyone for the excellent teamwork over the past month. I have thoroughly enjoyed working with all of you, and I am delighted to be a part of this team.\n# 1 Overview\n![pipeline](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3163774%2Fc8338666681788720689a7ead39a92ab%2F.png?generation=1686757921566293&alt=media)\n# 2 Main pipeline\n## 2.1 SP/SG\n### 2.1.1 Rotation\nRotating the image has a significant effect, since SG is lack of the rotation invariance. Therefore, for each image pair A-B, we fixed image A and rotated image B four times (0, 90, 180, 270). After performing four times SG matching, we selected the rotation angle with the most matches for the next stage. \n![Rotation](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3163774%2F33aa1271502732f84ca6420949ed19a6%2F.png?generation=1686758077155244&alt=media)\n### 2.1.2 Resize image \nIn the early stage of the competition, we used the original image for extracting keypoints and matching. And we noticed that the mAA score in the heritage cyprus scene was only 0.1. After experiments, we found that if we scaled the images in cyprus scene to 1920 on the longer side, the mAA score was improved to 0.6. Specifically, we didn't rescale the image size after SP keypoints extraction. Finally, in our application, we scaled the image size if the longer side was larger than 1920, otherwise kept the original image size for the following SP + SG inference.\n### 2.1.3 SP/SG setting\nWe increased the NMS value of SP from 3 to 8 and set the maximum number of keypoints to 4000.\n## 2.2 GeoVerification（RANSAC）\nWe used USAC_MAGSAC for geometric verification with the following configuration.\ncv2.findFundamentalMat(mkpts0, mkpts1, cv2.USAC_MAGSAC, 2, 0.99999, 100000)\n## 2.3 Setting Camera Params（Randomness）\nIn our experiments, we wanted to eliminate the randomness effect in our final mAA, since we observed that the same notebook could result in fluctuations of approximately 0.03 in the LB.  We found that the randomness come mainly from the Ceres (https://github.com/colmap/colmap/issues/404)) optimization. \nAfter analysis, the initial values of camera parameters, such as the focal length, had a significant impact on the stability of optimization results.\n![randomness](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3163774%2F9411c508f380417b835bb253b06eae7b%2Fwall_random.png?generation=1686758232496900&alt=media)\nSince the camera focal length information could be extracted from the image EXIF data, we provided the prior camera focal length to the camera and shared the same camera settings among images captured by the same device. Although there were still some fluctuations in the metrics, the majority of the experimental results were consistent. \nThe figure below shows four submissions of the same final notebook. **On the public LB, our metric fluctuates by no more than 0.004.** Moreover, compared to many other teams, **we do not have significant metric fluctuations between public LB and private LB.** Our public LB score was amazingly close to private score. This might indicated that our method **Significantly Reduced the Fluctuations caused by Randomness!**\n![mAA](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3163774%2Ffd18321289d3745de78691742d0ca28d%2Fonline_maa.png?generation=1686758332640650&alt=media)\n## 2.4 Mapper\nIn the COLMAP reconstruction process, we revised some default parameters in incremental mapping. After first trail, if the best model register ratio was below 1, we relaxed the mapper configuration (abs_pose_min_num_inliers=15) and re-run incremental mapping.\n# 3 Conclusion\n![Conclusion](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3163774%2F47522f20ba1cf5b608a1806f0cf761f4%2Fconclution.png?generation=1686760502421586&alt=media)\n# 4 Not fully tested\n## 4.1 Select image pairs\nOur final solution took nearly 9 hours,  since we generated the image matching pairs by exhaustive method. In order to speed up our solution, we tried Efficientnet_b7, Convnextv2_huge, and DINOv2 to extract image features and generate image matching pairs by feature similarity. In offline experiments, we selected the top N/2 most similar images (N being the total number of images) for each image to form image pairs. Compared to other methods, DINOv2 performed the best. By incorporating DINOv2 into the image pairs selection, we could control the processing time to 7 hours, and achieved 0.485 in LB (compared with exhaustive 0.507).\n![dinov2](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3163774%2Fbfdb729a111a6633b27e573788e09263%2Fdinov2.png?generation=1686760519524067&alt=media)\n## 4.2 Pixel-Perfect Structure-from-Motion\nPixSFM had a good performance improvement in our local validation. However, during online testing,  it consistently times out even we run it on scenes less than 50 images.  Moreover, it took a lot time for us to install the environment and run it on kaggle. Times out Sad! \n# 5 Ideas that did not work well\n## 5.1 Different detectors and matchers\nWe tested DKMv3, GlueStick, and SILK, but neither was able to surpass SPSG. In our experiments, we observed that DKMv3 performed better than SPSG in challenging scenes, such as those with large viewpoint differences or rotations. \n![dkm](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3163774%2Ffae472d5266e521f18f667592eaf3454%2Fdkm.png?generation=1686758763195086&alt=media)\nHowever, the overall metrics of DKMv3 were not as good as SPSG, which may be due to the operation of sampling dense matches we performed.\n## 5.2 Crop image\nJust like last year's winning solution, we took a crop after completing the first stage of matching. In the second stage, we performed matching again on the crop and merge all the matching results together.\n## 5.3 Half precision\nIt confused our team a lot. In the kaggle notebook, we found the half-precision improved the results, but on the public LB, it returned the lower results. \n## 5.4 Merge matches in four directions\nIn our methods, we rotated the image (0, 90, 180, 270) and performed four times matching. Instead of keeping all matches from four directions, we only keep the best angle matches, because keeping all matches didn't lead to any improvement in the results.\n## 5.5 findHomography\nWhen feature points lie on the same plane (e.g., in a wall scene) or when the camera undergoes pure rotation, the fundamental matrix degenerates. Therefore, in RANSAC, we simultaneously used the findFundamentalMat and the findHomography to calculate the inliers. However, this approach didn't lead to an improvement in the metrics.\n## 5.6 3D Model Refinement\nAfter the first trail of mapping, we tried to filter nosiy 3D points with large projection error or short track length, and then re-bundle adjust the model. Experically, this could help export better poses, but.. that's life.",
      "votes": null
    },
    {
      "id": "2302577",
      "postDate": "06/14/2023 16:19:03",
      "content": "<p>团队介绍<br>\n滴滴出行视觉计算组以使用计算机视觉和机器学习方法来解决交通场景下的实际问题为团队目标,依托海量的数据、领先的GPU计算集群和丰富的业务场景需求，不断创新、快速迭代为大量的用户创造价值。<br>\n团队在目标检测、场景文字检测识别、图像分割、立体视觉匹配、三维重建、SLAM 等方向具有丰富的技术积累，多项技术处于业界领先水平。<br>\n实习生岗位开放招聘，欢迎加入我们。<br>\n简历投递邮箱 <a href=\"mailto:wesleyliwei@didiglobal.com\">wesleyliwei@didiglobal.com</a></p>",
      "rawMarkdown": "团队介绍\n滴滴出行视觉计算组以使用计算机视觉和机器学习方法来解决交通场景下的实际问题为团队目标,依托海量的数据、领先的GPU计算集群和丰富的业务场景需求，不断创新、快速迭代为大量的用户创造价值。\n团队在目标检测、场景文字检测识别、图像分割、立体视觉匹配、三维重建、SLAM 等方向具有丰富的技术积累，多项技术处于业界领先水平。\n实习生岗位开放招聘，欢迎加入我们。\n简历投递邮箱 wesleyliwei@didiglobal.com",
      "votes": null
    },
    {
      "id": "2302599",
      "postDate": "06/14/2023 16:36:46",
      "content": "<p>欢迎大家加入我们👋</p>",
      "rawMarkdown": "欢迎大家加入我们👋",
      "votes": null
    },
    {
      "id": "2303166",
      "postDate": "06/15/2023 06:00:37",
      "content": "<p>Congratulations on winning the 3rd place!<br>\nMay I ask how you modified the parameter <code>abs_pose_min_num_inliers=15</code>?<br>\nWhen I try to modify this parameter in pycolmap, it  report mapper_options do not have this parameter.</p>",
      "rawMarkdown": "Congratulations on winning the 3rd place!\nMay I ask how you modified the parameter `abs_pose_min_num_inliers=15`?\nWhen I try to modify this parameter in pycolmap, it  report mapper_options do not have this parameter.",
      "votes": null
    },
    {
      "id": "2303177",
      "postDate": "06/15/2023 06:14:59",
      "content": "<p>Yes, PyCOLMAP is a Python wrapper for COLMAP. PyCOLMAP does not expose many parameters of COLMAP. If you need to set these parameters, you should directly call COLMAP instead of using PyCOLMAP.</p>",
      "rawMarkdown": "Yes, PyCOLMAP is a Python wrapper for COLMAP. PyCOLMAP does not expose many parameters of COLMAP. If you need to set these parameters, you should directly call COLMAP instead of using PyCOLMAP.",
      "votes": null
    },
    {
      "id": "2303265",
      "postDate": "06/15/2023 07:37:57",
      "content": "<p>Congrats for the gold~~😃😃 Seems using rotation in image pair was significant while using SG. Thank you for sharing your ideas. </p>",
      "rawMarkdown": "Congrats for the gold~~😃😃 Seems using rotation in image pair was significant while using SG. Thank you for sharing your ideas.",
      "votes": null
    },
    {
      "id": "2303625",
      "postDate": "06/15/2023 11:44:49",
      "content": "<p>yeah, sg does not match well for roatation scenario. Some other mathers are much better, such as DKM. As experimented above, we didn't fully test the metric combining diff matchers together, as time limitation. But we found some teams used multi-thread to speed up, maybe could help.</p>",
      "rawMarkdown": "yeah, sg does not match well for roatation scenario. Some other mathers are much better, such as DKM. As experimented above, we didn't fully test the metric combining diff matchers together, as time limitation. But we found some teams used multi-thread to speed up, maybe could help.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2302577,
      "author_name": "wesleyliwei",
      "author_url": "",
      "post_date": "06/14/2023 16:19:03",
      "content": "<p>团队介绍<br>\n滴滴出行视觉计算组以使用计算机视觉和机器学习方法来解决交通场景下的实际问题为团队目标,依托海量的数据、领先的GPU计算集群和丰富的业务场景需求，不断创新、快速迭代为大量的用户创造价值。<br>\n团队在目标检测、场景文字检测识别、图像分割、立体视觉匹配、三维重建、SLAM 等方向具有丰富的技术积累，多项技术处于业界领先水平。<br>\n实习生岗位开放招聘，欢迎加入我们。<br>\n简历投递邮箱 <a href=\"mailto:wesleyliwei@didiglobal.com\">wesleyliwei@didiglobal.com</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2302599,
          "author_name": "zjgsuchenchen",
          "author_url": "",
          "post_date": "06/14/2023 16:36:46",
          "content": "<p>欢迎大家加入我们👋</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2303166,
      "author_name": "mianhtan",
      "author_url": "",
      "post_date": "06/15/2023 06:00:37",
      "content": "<p>Congratulations on winning the 3rd place!<br>\nMay I ask how you modified the parameter <code>abs_pose_min_num_inliers=15</code>?<br>\nWhen I try to modify this parameter in pycolmap, it  report mapper_options do not have this parameter.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2303177,
          "author_name": "zjgsuchenchen",
          "author_url": "",
          "post_date": "06/15/2023 06:14:59",
          "content": "<p>Yes, PyCOLMAP is a Python wrapper for COLMAP. PyCOLMAP does not expose many parameters of COLMAP. If you need to set these parameters, you should directly call COLMAP instead of using PyCOLMAP.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2303265,
      "author_name": "markrovotix",
      "author_url": "",
      "post_date": "06/15/2023 07:37:57",
      "content": "<p>Congrats for the gold~~😃😃 Seems using rotation in image pair was significant while using SG. Thank you for sharing your ideas. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2303625,
          "author_name": "wesleyliwei",
          "author_url": "",
          "post_date": "06/15/2023 11:44:49",
          "content": "<p>yeah, sg does not match well for roatation scenario. Some other mathers are much better, such as DKM. As experimented above, we didn't fully test the metric combining diff matchers together, as time limitation. But we found some teams used multi-thread to speed up, maybe could help.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2302561": "We are delighted to participate in this competition and would like to express gratitude to all the Kaggle staffs and Sponsors. Congratulations to all the participants. \nThe team members include 陈鹏、陈建国、阮志伟 and 李伟. I would like to express my sincere gratitude to everyone for the excellent teamwork over the past month. I have thoroughly enjoyed working with all of you, and I am delighted to be a part of this team.\n# 1 Overview\n![pipeline](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3163774%2Fc8338666681788720689a7ead39a92ab%2F.png?generation=1686757921566293&alt=media)\n# 2 Main pipeline\n## 2.1 SP/SG\n### 2.1.1 Rotation\nRotating the image has a significant effect, since SG is lack of the rotation invariance. Therefore, for each image pair A-B, we fixed image A and rotated image B four times (0, 90, 180, 270). After performing four times SG matching, we selected the rotation angle with the most matches for the next stage. \n![Rotation](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3163774%2F33aa1271502732f84ca6420949ed19a6%2F.png?generation=1686758077155244&alt=media)\n### 2.1.2 Resize image \nIn the early stage of the competition, we used the original image for extracting keypoints and matching. And we noticed that the mAA score in the heritage cyprus scene was only 0.1. After experiments, we found that if we scaled the images in cyprus scene to 1920 on the longer side, the mAA score was improved to 0.6. Specifically, we didn't rescale the image size after SP keypoints extraction. Finally, in our application, we scaled the image size if the longer side was larger than 1920, otherwise kept the original image size for the following SP + SG inference.\n### 2.1.3 SP/SG setting\nWe increased the NMS value of SP from 3 to 8 and set the maximum number of keypoints to 4000.\n## 2.2 GeoVerification（RANSAC）\nWe used USAC_MAGSAC for geometric verification with the following configuration.\ncv2.findFundamentalMat(mkpts0, mkpts1, cv2.USAC_MAGSAC, 2, 0.99999, 100000)\n## 2.3 Setting Camera Params（Randomness）\nIn our experiments, we wanted to eliminate the randomness effect in our final mAA, since we observed that the same notebook could result in fluctuations of approximately 0.03 in the LB.  We found that the randomness come mainly from the Ceres (https://github.com/colmap/colmap/issues/404)) optimization. \nAfter analysis, the initial values of camera parameters, such as the focal length, had a significant impact on the stability of optimization results.\n![randomness](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3163774%2F9411c508f380417b835bb253b06eae7b%2Fwall_random.png?generation=1686758232496900&alt=media)\nSince the camera focal length information could be extracted from the image EXIF data, we provided the prior camera focal length to the camera and shared the same camera settings among images captured by the same device. Although there were still some fluctuations in the metrics, the majority of the experimental results were consistent. \nThe figure below shows four submissions of the same final notebook. **On the public LB, our metric fluctuates by no more than 0.004.** Moreover, compared to many other teams, **we do not have significant metric fluctuations between public LB and private LB.** Our public LB score was amazingly close to private score. This might indicated that our method **Significantly Reduced the Fluctuations caused by Randomness!**\n![mAA](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3163774%2Ffd18321289d3745de78691742d0ca28d%2Fonline_maa.png?generation=1686758332640650&alt=media)\n## 2.4 Mapper\nIn the COLMAP reconstruction process, we revised some default parameters in incremental mapping. After first trail, if the best model register ratio was below 1, we relaxed the mapper configuration (abs_pose_min_num_inliers=15) and re-run incremental mapping.\n# 3 Conclusion\n![Conclusion](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3163774%2F47522f20ba1cf5b608a1806f0cf761f4%2Fconclution.png?generation=1686760502421586&alt=media)\n# 4 Not fully tested\n## 4.1 Select image pairs\nOur final solution took nearly 9 hours,  since we generated the image matching pairs by exhaustive method. In order to speed up our solution, we tried Efficientnet_b7, Convnextv2_huge, and DINOv2 to extract image features and generate image matching pairs by feature similarity. In offline experiments, we selected the top N/2 most similar images (N being the total number of images) for each image to form image pairs. Compared to other methods, DINOv2 performed the best. By incorporating DINOv2 into the image pairs selection, we could control the processing time to 7 hours, and achieved 0.485 in LB (compared with exhaustive 0.507).\n![dinov2](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3163774%2Fbfdb729a111a6633b27e573788e09263%2Fdinov2.png?generation=1686760519524067&alt=media)\n## 4.2 Pixel-Perfect Structure-from-Motion\nPixSFM had a good performance improvement in our local validation. However, during online testing,  it consistently times out even we run it on scenes less than 50 images.  Moreover, it took a lot time for us to install the environment and run it on kaggle. Times out Sad! \n# 5 Ideas that did not work well\n## 5.1 Different detectors and matchers\nWe tested DKMv3, GlueStick, and SILK, but neither was able to surpass SPSG. In our experiments, we observed that DKMv3 performed better than SPSG in challenging scenes, such as those with large viewpoint differences or rotations. \n![dkm](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3163774%2Ffae472d5266e521f18f667592eaf3454%2Fdkm.png?generation=1686758763195086&alt=media)\nHowever, the overall metrics of DKMv3 were not as good as SPSG, which may be due to the operation of sampling dense matches we performed.\n## 5.2 Crop image\nJust like last year's winning solution, we took a crop after completing the first stage of matching. In the second stage, we performed matching again on the crop and merge all the matching results together.\n## 5.3 Half precision\nIt confused our team a lot. In the kaggle notebook, we found the half-precision improved the results, but on the public LB, it returned the lower results. \n## 5.4 Merge matches in four directions\nIn our methods, we rotated the image (0, 90, 180, 270) and performed four times matching. Instead of keeping all matches from four directions, we only keep the best angle matches, because keeping all matches didn't lead to any improvement in the results.\n## 5.5 findHomography\nWhen feature points lie on the same plane (e.g., in a wall scene) or when the camera undergoes pure rotation, the fundamental matrix degenerates. Therefore, in RANSAC, we simultaneously used the findFundamentalMat and the findHomography to calculate the inliers. However, this approach didn't lead to an improvement in the metrics.\n## 5.6 3D Model Refinement\nAfter the first trail of mapping, we tried to filter nosiy 3D points with large projection error or short track length, and then re-bundle adjust the model. Experically, this could help export better poses, but.. that's life.",
    "2302577": "团队介绍\n滴滴出行视觉计算组以使用计算机视觉和机器学习方法来解决交通场景下的实际问题为团队目标,依托海量的数据、领先的GPU计算集群和丰富的业务场景需求，不断创新、快速迭代为大量的用户创造价值。\n团队在目标检测、场景文字检测识别、图像分割、立体视觉匹配、三维重建、SLAM 等方向具有丰富的技术积累，多项技术处于业界领先水平。\n实习生岗位开放招聘，欢迎加入我们。\n简历投递邮箱 wesleyliwei@didiglobal.com",
    "2302599": "欢迎大家加入我们👋",
    "2303166": "Congratulations on winning the 3rd place!\nMay I ask how you modified the parameter `abs_pose_min_num_inliers=15`?\nWhen I try to modify this parameter in pycolmap, it  report mapper_options do not have this parameter.",
    "2303177": "Yes, PyCOLMAP is a Python wrapper for COLMAP. PyCOLMAP does not expose many parameters of COLMAP. If you need to set these parameters, you should directly call COLMAP instead of using PyCOLMAP.",
    "2303265": "Congrats for the gold~~😃😃 Seems using rotation in image pair was significant while using SG. Thank you for sharing your ideas.",
    "2303625": "yeah, sg does not match well for roatation scenario. Some other mathers are much better, such as DKM. As experimented above, we didn't fully test the metric combining diff matchers together, as time limitation. But we found some teams used multi-thread to speed up, maybe could help."
  },
  "source": "meta"
}