{
  "id": 510338,
  "title": "3rd Place Solution: VGGSfM",
  "url": "/competitions/image-matching-challenge-2024/writeups/vggsfm-3rd-place-solution-vggsfm",
  "author_name": "",
  "post_date": "2024-06-05T23:13:05.197Z",
  "votes": 31,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I am excited to share my participation in IMC 2024 and extend my gratitude to the organizers <a href=\"https://www.kaggle.com/oldufo\" target=\"_blank\">@oldufo</a> <a href=\"https://www.kaggle.com/eduardtrulls\" target=\"_blank\">@eduardtrulls</a> , sponsors, and the Kaggle team! This was my first experience at both IMC and Kaggle, and it was truly enriching.</p>\n<p>My implementation is based on the public code of <a href=\"https://www.kaggle.com/code/nartaa/imc2024-starter/notebook\" target=\"_blank\">IMC2024 Starter</a> by <a href=\"https://www.kaggle.com/nartaa\" target=\"_blank\">@nartaa</a>. It helped a lot! This baseline uses ALIKED+LightGlue and pycolmap for reconstruction. Like others, I used the training set as the evaluation set, with 50 samples.</p>\n<p>Below is my solution for IMC 2024.</p>\n<h1>1. Introduction</h1>\n<p>The core of my solution is <a href=\"https://vggsfm.github.io/\" target=\"_blank\">VGGSfM: Visual Geometry Grounded Deep Structure From Motion</a> (<a href=\"https://vggsfm.github.io/)\" target=\"_blank\">https://vggsfm.github.io/)</a>. This work has been accepted to CVPR 2024 as a Highlight.</p>\n<p>Basically, given N input frames, VGGSfM first uses a camera predictor to estimate preliminary camera extrinsic and intrinsic parameters. After selecting a query frame and some 2D query points, it estimates 2D tracks using a track predictor. These tracks, alongside the preliminary camera data, allow us to perform triangulation to obtain 3D points, followed by a bundle adjustment to refine the camera and point combination. More details can be found in the paper. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16492520%2F0feabaadc93f98e6b15fbe9eea33503b%2Foverview.png?generation=1717614431932034&amp;alt=media\"></p>\n<h1>2. Solution</h1>\n<h2>2.1 VGGSfM across All Frames</h2>\n<p>I began by applying VGGSfM across all input frames, which proved effective on the IMC2024 evaluation set. With the input frame number of 50, its mAA is higher than the baseline by 4% (0.04). However, due to the 16 GB GPU limitation on the Kaggle server, accommodating this approach in Kaggle submission was challenging. OOM is bitter. Therefore, to accommodate the server constraints, I integrated VGGSfM into the existing pycolmap pipeline, as detailed below.</p>\n<h2>2.2 Additional Tracks</h2>\n<p>The first idea is using VGGSfM track predictor to estimate tracks, and feed them into pycolmap as additional 2D matches. For each image, I identified its N nearest frames using NetVLAD or DINO V2. Using the image as the query frame, I selected query points via Superpoint and found corresponding tracks on the N nearest frames. With N=5, this helped improving mAA by 3% on the evaluation set. However, in the public leaderboard, this only improves the score from ~0.17 to ~0.18. I suspect a higher N could yield better results, but don’t have enough submissions to try. I joined the competition quite late, after half of the competition and with ~120 submissions in total. </p>\n<p>This strategy runs track prediction once for each input frame, avoiding the quadratic complexity issue of exhaustive matching. On the Kaggle cluster, it requires 1.8 seconds per frame, in total 90 seconds for a 50-image scene. On my local A100 cluster, it takes 0.7 seconds per frame.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16492520%2F7a5fc4ec54bb1f5ce36ddad5b04ac55f%2FScreenshot%202024-06-05%20at%2020.02.12.png?generation=1717614456486524&amp;alt=media\"></p>\n<h2>2.3 SfM Track Refinement</h2>\n<p>The VGGSfM track predictor adopts a two-stage design, which first predicts coarse tracks and then refine them by local appearance. Inspired by the winner solution by ZJU3DV <a href=\"https://www.kaggle.com/xingyihezju3dv\" target=\"_blank\">@xingyihezju3dv</a> last year, I found we can refine the SfM tracks of pycolmap reconstruction. The idea is, first running the baseline with ALIKED+LightGlue (or any other methods), you can find a best reconstruction, the one with most registered images. Each valid 3D point within this reconstruction corresponds to several 2D points, and these 2D points construct a SfM track. Therefore, I extracted the image patches (31x31) for those SfM tracks, refined them by VGGSfM fine track predictor, and updated the pycolmap reconstruction with the updated 2D positions. Then, we run one more global bundle adjustment to optimize the cameras and 3D points based on the refined tracks. For example, in the Church scene, it improves the mean reprojection error from 0.64 to 0.55. </p>\n<p>Importantly, this process does not require recompiling colmap. I have provided a code segment here (<a href=\"https://www.kaggle.com/code/jianyuanv/vggsfm-to-refine-pycolmap-tracks\" target=\"_blank\">https://www.kaggle.com/code/jianyuanv/vggsfm-to-refine-pycolmap-tracks</a>) about how to do this by pycolmap and VGGSfM model. </p>\n<p>During submission, we only refine the top 4096 SfM tracks with highest reprojection errors. This will take around 150 seconds for each scene, i.e., around 30 tracks per second. In the public leaderboard, this improves our solution from ~0.18 to ~0.20.</p>\n<h2>2.4 VGGSfM to relocate missing images</h2>\n<p>We also observe that colmap usually cannot register all the images in a scene. Therefore, we try to relocate missing images by VGGSfM. For every missing image, we identified the top 5 registered images it has matches with. If an image cannot find 5 registered images, we use the nearest frames by DINO v2. </p>\n<p>Now we have a subset of 6 (1 unregistered + 5 registered) images in total. We use VGGSfM to reconstruct their camera poses. While, the poses predicted by VGGSfM stay in a different camera coordinate to the original camera coordinate of pycolmap reconstruction. We use the algorithm by umeyama to find an alignment transform for the centers of 5 registered images, and apply this alignment transform on the camera pose of the unregistered image. </p>\n<p>For each missing image, this process will take around 5 seconds. It improves the public leaderboard score from ~0.20 to ~0.21. It is worth noting that sometimes the missing image may be too hard, which may not have enough overlaps (matches) with all other frames. In this situation, bundle adjustment may fail, and we use the camera poses predicted by VGGSfM camera predictor only. To be honest I think I did something wrong for this part because the code was written in a rush. I may mess up with some camera coordinate transform, but it works not bad.</p>\n<h2>2.5 Final Solution</h2>\n<p>In addition to the methods described above, our final solution also:</p>\n<ul>\n<li><strong>a.</strong> Handles the rotation using <a href=\"https://github.com/ternaus/check_orientation\" target=\"_blank\">https://github.com/ternaus/check_orientation</a>. If the rotation prediction model predicts a score higher than 0.5 for a frame, we rotate the image by 90, 180, and 270 degrees, and pick the one with a highest number of matches.  </li>\n<li><strong>b.</strong> Uses the matches from both ALIKED+LightGlue and SP+LightGlue</li>\n<li><strong>c.</strong> Extract an area of interest for transparent images (by DBSCAN on keypoints), and run keypoint detection again on the area of interest. </li>\n</ul>\n<h2>2.6 Ideas tried but not worked obviously</h2>\n<ul>\n<li><strong>a.</strong> Histograms equalization or CLAHE.</li>\n<li><strong>b.</strong> Disambiguation methods such as the ones included in <a href=\"https://github.com/cvg/sfm-disambiguation-colmap\" target=\"_blank\">https://github.com/cvg/sfm-disambiguation-colmap</a> or <a href=\"https://github.com/RuojinCai/doppelgangers\" target=\"_blank\">Doppelgangers</a> (I probably did something wrong with Doppelgangers, but not enough time to debug it).</li>\n<li><strong>c.</strong> Homography estimation to filter out outlier matches. </li>\n<li><strong>d.</strong> Adjusting pycolmap hyper-parameters.</li>\n<li><strong>e.</strong> Test Time Augmentation, such as using different image resolution.</li>\n</ul>\n<h1>3. Learned from IMC</h1>\n<p>The three-week journey was great and highly rewarding, despite some sleep deprivation. One of the good stuffs was discovering several ways to reduce the GPU consumption of VGGSfM, which I will soon update on the VGGSfM GitHub. Additionally, it was gratifying to find that VGGSfM can effectively complement colmap when computational resources are limited. And, being familiar with Kaggle took me a lot of time. I hope to do better in the next year.</p>\n<p>(This article was written very quickly and I may further update it next week. Time to have a short break now. But feel free to drop me an email if you have any question about VGGSfM :)</p>",
  "messages": [
    {
      "id": "2857261",
      "postDate": "06/05/2024 19:15:06",
      "content": "<p>I am excited to share my participation in IMC 2024 and extend my gratitude to the organizers <a href=\"https://www.kaggle.com/oldufo\" target=\"_blank\">@oldufo</a> <a href=\"https://www.kaggle.com/eduardtrulls\" target=\"_blank\">@eduardtrulls</a> , sponsors, and the Kaggle team! This was my first experience at both IMC and Kaggle, and it was truly enriching.</p>\n<p>My implementation is based on the public code of <a href=\"https://www.kaggle.com/code/nartaa/imc2024-starter/notebook\" target=\"_blank\">IMC2024 Starter</a> by <a href=\"https://www.kaggle.com/nartaa\" target=\"_blank\">@nartaa</a>. It helped a lot! This baseline uses ALIKED+LightGlue and pycolmap for reconstruction. Like others, I used the training set as the evaluation set, with 50 samples.</p>\n<p>Below is my solution for IMC 2024.</p>\n<h1>1. Introduction</h1>\n<p>The core of my solution is <a href=\"https://vggsfm.github.io/\" target=\"_blank\">VGGSfM: Visual Geometry Grounded Deep Structure From Motion</a> (<a href=\"https://vggsfm.github.io/)\" target=\"_blank\">https://vggsfm.github.io/)</a>. This work has been accepted to CVPR 2024 as a Highlight.</p>\n<p>Basically, given N input frames, VGGSfM first uses a camera predictor to estimate preliminary camera extrinsic and intrinsic parameters. After selecting a query frame and some 2D query points, it estimates 2D tracks using a track predictor. These tracks, alongside the preliminary camera data, allow us to perform triangulation to obtain 3D points, followed by a bundle adjustment to refine the camera and point combination. More details can be found in the paper. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16492520%2F0feabaadc93f98e6b15fbe9eea33503b%2Foverview.png?generation=1717614431932034&amp;alt=media\"></p>\n<h1>2. Solution</h1>\n<h2>2.1 VGGSfM across All Frames</h2>\n<p>I began by applying VGGSfM across all input frames, which proved effective on the IMC2024 evaluation set. With the input frame number of 50, its mAA is higher than the baseline by 4% (0.04). However, due to the 16 GB GPU limitation on the Kaggle server, accommodating this approach in Kaggle submission was challenging. OOM is bitter. Therefore, to accommodate the server constraints, I integrated VGGSfM into the existing pycolmap pipeline, as detailed below.</p>\n<h2>2.2 Additional Tracks</h2>\n<p>The first idea is using VGGSfM track predictor to estimate tracks, and feed them into pycolmap as additional 2D matches. For each image, I identified its N nearest frames using NetVLAD or DINO V2. Using the image as the query frame, I selected query points via Superpoint and found corresponding tracks on the N nearest frames. With N=5, this helped improving mAA by 3% on the evaluation set. However, in the public leaderboard, this only improves the score from ~0.17 to ~0.18. I suspect a higher N could yield better results, but don’t have enough submissions to try. I joined the competition quite late, after half of the competition and with ~120 submissions in total. </p>\n<p>This strategy runs track prediction once for each input frame, avoiding the quadratic complexity issue of exhaustive matching. On the Kaggle cluster, it requires 1.8 seconds per frame, in total 90 seconds for a 50-image scene. On my local A100 cluster, it takes 0.7 seconds per frame.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16492520%2F7a5fc4ec54bb1f5ce36ddad5b04ac55f%2FScreenshot%202024-06-05%20at%2020.02.12.png?generation=1717614456486524&amp;alt=media\"></p>\n<h2>2.3 SfM Track Refinement</h2>\n<p>The VGGSfM track predictor adopts a two-stage design, which first predicts coarse tracks and then refine them by local appearance. Inspired by the winner solution by ZJU3DV <a href=\"https://www.kaggle.com/xingyihezju3dv\" target=\"_blank\">@xingyihezju3dv</a> last year, I found we can refine the SfM tracks of pycolmap reconstruction. The idea is, first running the baseline with ALIKED+LightGlue (or any other methods), you can find a best reconstruction, the one with most registered images. Each valid 3D point within this reconstruction corresponds to several 2D points, and these 2D points construct a SfM track. Therefore, I extracted the image patches (31x31) for those SfM tracks, refined them by VGGSfM fine track predictor, and updated the pycolmap reconstruction with the updated 2D positions. Then, we run one more global bundle adjustment to optimize the cameras and 3D points based on the refined tracks. For example, in the Church scene, it improves the mean reprojection error from 0.64 to 0.55. </p>\n<p>Importantly, this process does not require recompiling colmap. I have provided a code segment here (<a href=\"https://www.kaggle.com/code/jianyuanv/vggsfm-to-refine-pycolmap-tracks\" target=\"_blank\">https://www.kaggle.com/code/jianyuanv/vggsfm-to-refine-pycolmap-tracks</a>) about how to do this by pycolmap and VGGSfM model. </p>\n<p>During submission, we only refine the top 4096 SfM tracks with highest reprojection errors. This will take around 150 seconds for each scene, i.e., around 30 tracks per second. In the public leaderboard, this improves our solution from ~0.18 to ~0.20.</p>\n<h2>2.4 VGGSfM to relocate missing images</h2>\n<p>We also observe that colmap usually cannot register all the images in a scene. Therefore, we try to relocate missing images by VGGSfM. For every missing image, we identified the top 5 registered images it has matches with. If an image cannot find 5 registered images, we use the nearest frames by DINO v2. </p>\n<p>Now we have a subset of 6 (1 unregistered + 5 registered) images in total. We use VGGSfM to reconstruct their camera poses. While, the poses predicted by VGGSfM stay in a different camera coordinate to the original camera coordinate of pycolmap reconstruction. We use the algorithm by umeyama to find an alignment transform for the centers of 5 registered images, and apply this alignment transform on the camera pose of the unregistered image. </p>\n<p>For each missing image, this process will take around 5 seconds. It improves the public leaderboard score from ~0.20 to ~0.21. It is worth noting that sometimes the missing image may be too hard, which may not have enough overlaps (matches) with all other frames. In this situation, bundle adjustment may fail, and we use the camera poses predicted by VGGSfM camera predictor only. To be honest I think I did something wrong for this part because the code was written in a rush. I may mess up with some camera coordinate transform, but it works not bad.</p>\n<h2>2.5 Final Solution</h2>\n<p>In addition to the methods described above, our final solution also:</p>\n<ul>\n<li><strong>a.</strong> Handles the rotation using <a href=\"https://github.com/ternaus/check_orientation\" target=\"_blank\">https://github.com/ternaus/check_orientation</a>. If the rotation prediction model predicts a score higher than 0.5 for a frame, we rotate the image by 90, 180, and 270 degrees, and pick the one with a highest number of matches.  </li>\n<li><strong>b.</strong> Uses the matches from both ALIKED+LightGlue and SP+LightGlue</li>\n<li><strong>c.</strong> Extract an area of interest for transparent images (by DBSCAN on keypoints), and run keypoint detection again on the area of interest. </li>\n</ul>\n<h2>2.6 Ideas tried but not worked obviously</h2>\n<ul>\n<li><strong>a.</strong> Histograms equalization or CLAHE.</li>\n<li><strong>b.</strong> Disambiguation methods such as the ones included in <a href=\"https://github.com/cvg/sfm-disambiguation-colmap\" target=\"_blank\">https://github.com/cvg/sfm-disambiguation-colmap</a> or <a href=\"https://github.com/RuojinCai/doppelgangers\" target=\"_blank\">Doppelgangers</a> (I probably did something wrong with Doppelgangers, but not enough time to debug it).</li>\n<li><strong>c.</strong> Homography estimation to filter out outlier matches. </li>\n<li><strong>d.</strong> Adjusting pycolmap hyper-parameters.</li>\n<li><strong>e.</strong> Test Time Augmentation, such as using different image resolution.</li>\n</ul>\n<h1>3. Learned from IMC</h1>\n<p>The three-week journey was great and highly rewarding, despite some sleep deprivation. One of the good stuffs was discovering several ways to reduce the GPU consumption of VGGSfM, which I will soon update on the VGGSfM GitHub. Additionally, it was gratifying to find that VGGSfM can effectively complement colmap when computational resources are limited. And, being familiar with Kaggle took me a lot of time. I hope to do better in the next year.</p>\n<p>(This article was written very quickly and I may further update it next week. Time to have a short break now. But feel free to drop me an email if you have any question about VGGSfM :)</p>",
      "rawMarkdown": "I am excited to share my participation in IMC 2024 and extend my gratitude to the organizers @oldufo @eduardtrulls , sponsors, and the Kaggle team! This was my first experience at both IMC and Kaggle, and it was truly enriching.\n\nMy implementation is based on the public code of [IMC2024 Starter](https://www.kaggle.com/code/nartaa/imc2024-starter/notebook) by @nartaa. It helped a lot! This baseline uses ALIKED+LightGlue and pycolmap for reconstruction. Like others, I used the training set as the evaluation set, with 50 samples.\n\n\n\nBelow is my solution for IMC 2024.\n\n# 1. Introduction\nThe core of my solution is [VGGSfM: Visual Geometry Grounded Deep Structure From Motion](https://vggsfm.github.io/) (https://vggsfm.github.io/). This work has been accepted to CVPR 2024 as a Highlight.\n\nBasically, given N input frames, VGGSfM first uses a camera predictor to estimate preliminary camera extrinsic and intrinsic parameters. After selecting a query frame and some 2D query points, it estimates 2D tracks using a track predictor. These tracks, alongside the preliminary camera data, allow us to perform triangulation to obtain 3D points, followed by a bundle adjustment to refine the camera and point combination. More details can be found in the paper. \n\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16492520%2F0feabaadc93f98e6b15fbe9eea33503b%2Foverview.png?generation=1717614431932034&alt=media\" width=\"1100\" >\n\n\n# 2. Solution\n\n## 2.1 VGGSfM across All Frames\n\nI began by applying VGGSfM across all input frames, which proved effective on the IMC2024 evaluation set. With the input frame number of 50, its mAA is higher than the baseline by 4% (0.04). However, due to the 16 GB GPU limitation on the Kaggle server, accommodating this approach in Kaggle submission was challenging. OOM is bitter. Therefore, to accommodate the server constraints, I integrated VGGSfM into the existing pycolmap pipeline, as detailed below.\n\n## 2.2 Additional Tracks\n\nThe first idea is using VGGSfM track predictor to estimate tracks, and feed them into pycolmap as additional 2D matches. For each image, I identified its N nearest frames using NetVLAD or DINO V2. Using the image as the query frame, I selected query points via Superpoint and found corresponding tracks on the N nearest frames. With N=5, this helped improving mAA by 3% on the evaluation set. However, in the public leaderboard, this only improves the score from ~0.17 to ~0.18. I suspect a higher N could yield better results, but don’t have enough submissions to try. I joined the competition quite late, after half of the competition and with ~120 submissions in total. \n\nThis strategy runs track prediction once for each input frame, avoiding the quadratic complexity issue of exhaustive matching. On the Kaggle cluster, it requires 1.8 seconds per frame, in total 90 seconds for a 50-image scene. On my local A100 cluster, it takes 0.7 seconds per frame.\n\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16492520%2F7a5fc4ec54bb1f5ce36ddad5b04ac55f%2FScreenshot%202024-06-05%20at%2020.02.12.png?generation=1717614456486524&alt=media\" width=\"500\" >\n\n\n\n## 2.3 SfM Track Refinement \n\nThe VGGSfM track predictor adopts a two-stage design, which first predicts coarse tracks and then refine them by local appearance. Inspired by the winner solution by ZJU3DV @xingyihezju3dv last year, I found we can refine the SfM tracks of pycolmap reconstruction. The idea is, first running the baseline with ALIKED+LightGlue (or any other methods), you can find a best reconstruction, the one with most registered images. Each valid 3D point within this reconstruction corresponds to several 2D points, and these 2D points construct a SfM track. Therefore, I extracted the image patches (31x31) for those SfM tracks, refined them by VGGSfM fine track predictor, and updated the pycolmap reconstruction with the updated 2D positions. Then, we run one more global bundle adjustment to optimize the cameras and 3D points based on the refined tracks. For example, in the Church scene, it improves the mean reprojection error from 0.64 to 0.55. \n\nImportantly, this process does not require recompiling colmap. I have provided a code segment here (https://www.kaggle.com/code/jianyuanv/vggsfm-to-refine-pycolmap-tracks) about how to do this by pycolmap and VGGSfM model. \n\nDuring submission, we only refine the top 4096 SfM tracks with highest reprojection errors. This will take around 150 seconds for each scene, i.e., around 30 tracks per second. In the public leaderboard, this improves our solution from ~0.18 to ~0.20.\n\n## 2.4 VGGSfM to relocate missing images \n\nWe also observe that colmap usually cannot register all the images in a scene. Therefore, we try to relocate missing images by VGGSfM. For every missing image, we identified the top 5 registered images it has matches with. If an image cannot find 5 registered images, we use the nearest frames by DINO v2. \n\nNow we have a subset of 6 (1 unregistered + 5 registered) images in total. We use VGGSfM to reconstruct their camera poses. While, the poses predicted by VGGSfM stay in a different camera coordinate to the original camera coordinate of pycolmap reconstruction. We use the algorithm by umeyama to find an alignment transform for the centers of 5 registered images, and apply this alignment transform on the camera pose of the unregistered image. \n\nFor each missing image, this process will take around 5 seconds. It improves the public leaderboard score from ~0.20 to ~0.21. It is worth noting that sometimes the missing image may be too hard, which may not have enough overlaps (matches) with all other frames. In this situation, bundle adjustment may fail, and we use the camera poses predicted by VGGSfM camera predictor only. To be honest I think I did something wrong for this part because the code was written in a rush. I may mess up with some camera coordinate transform, but it works not bad.\n\n## 2.5 Final Solution\n\nIn addition to the methods described above, our final solution also:\n\n- **a.** Handles the rotation using https://github.com/ternaus/check_orientation. If the rotation prediction model predicts a score higher than 0.5 for a frame, we rotate the image by 90, 180, and 270 degrees, and pick the one with a highest number of matches.  \n- **b.** Uses the matches from both ALIKED+LightGlue and SP+LightGlue\n- **c.** Extract an area of interest for transparent images (by DBSCAN on keypoints), and run keypoint detection again on the area of interest. \n\n## 2.6 Ideas tried but not worked obviously \n- **a.** Histograms equalization or CLAHE.\n- **b.** Disambiguation methods such as the ones included in https://github.com/cvg/sfm-disambiguation-colmap or [Doppelgangers](https://github.com/RuojinCai/doppelgangers) (I probably did something wrong with Doppelgangers, but not enough time to debug it).\n- **c.** Homography estimation to filter out outlier matches. \n- **d.** Adjusting pycolmap hyper-parameters.\n- **e.** Test Time Augmentation, such as using different image resolution.\n\n# 3. Learned from IMC\n\nThe three-week journey was great and highly rewarding, despite some sleep deprivation. One of the good stuffs was discovering several ways to reduce the GPU consumption of VGGSfM, which I will soon update on the VGGSfM GitHub. Additionally, it was gratifying to find that VGGSfM can effectively complement colmap when computational resources are limited. And, being familiar with Kaggle took me a lot of time. I hope to do better in the next year.\n\n\n(This article was written very quickly and I may further update it next week. Time to have a short break now. But feel free to drop me an email if you have any question about VGGSfM :)",
      "votes": null
    },
    {
      "id": "2857789",
      "postDate": "06/06/2024 06:00:22",
      "content": "<p>Absolutely fantastic! It's a powerful and elegant algorithm. I'm studying your code. The only thing is, the GPU requirements are a bit high.</p>",
      "rawMarkdown": "Absolutely fantastic! It's a powerful and elegant algorithm. I'm studying your code. The only thing is, the GPU requirements are a bit high.",
      "votes": null
    },
    {
      "id": "2857808",
      "postDate": "06/06/2024 06:12:53",
      "content": "<p>An brilliant solution！I have read your paper and gained a lot. Congratulations on winning third place in this competition.</p>",
      "rawMarkdown": "An brilliant solution！I have read your paper and gained a lot. Congratulations on winning third place in this competition.",
      "votes": null
    },
    {
      "id": "2858683",
      "postDate": "06/06/2024 15:47:29",
      "content": "<p>yes, luckily I have learned something in this IMC. I am going to update the code and it should save the GPU usage by half 😄</p>",
      "rawMarkdown": "yes, luckily I have learned something in this IMC. I am going to update the code and it should save the GPU usage by half 😄",
      "votes": null
    },
    {
      "id": "2888847",
      "postDate": "06/25/2024 05:17:01",
      "content": "<p>It has been updated now</p>",
      "rawMarkdown": "It has been updated now",
      "votes": null
    },
    {
      "id": "2899892",
      "postDate": "07/01/2024 22:47:47",
      "content": "<p>well written &amp; well deserved position. congrats guys:)</p>",
      "rawMarkdown": "well written & well deserved position. congrats guys:)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2857789,
      "author_name": "sunnyykk",
      "author_url": "",
      "post_date": "06/06/2024 06:00:22",
      "content": "<p>Absolutely fantastic! It's a powerful and elegant algorithm. I'm studying your code. The only thing is, the GPU requirements are a bit high.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2858683,
          "author_name": "jianyuanv",
          "author_url": "",
          "post_date": "06/06/2024 15:47:29",
          "content": "<p>yes, luckily I have learned something in this IMC. I am going to update the code and it should save the GPU usage by half 😄</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2888847,
          "author_name": "jianyuanv",
          "author_url": "",
          "post_date": "06/25/2024 05:17:01",
          "content": "<p>It has been updated now</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2857808,
      "author_name": "ritianitian",
      "author_url": "",
      "post_date": "06/06/2024 06:12:53",
      "content": "<p>An brilliant solution！I have read your paper and gained a lot. Congratulations on winning third place in this competition.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2899892,
      "author_name": "shyamgupta196",
      "author_url": "",
      "post_date": "07/01/2024 22:47:47",
      "content": "<p>well written &amp; well deserved position. congrats guys:)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2857261": "I am excited to share my participation in IMC 2024 and extend my gratitude to the organizers @oldufo @eduardtrulls , sponsors, and the Kaggle team! This was my first experience at both IMC and Kaggle, and it was truly enriching.\n\nMy implementation is based on the public code of [IMC2024 Starter](https://www.kaggle.com/code/nartaa/imc2024-starter/notebook) by @nartaa. It helped a lot! This baseline uses ALIKED+LightGlue and pycolmap for reconstruction. Like others, I used the training set as the evaluation set, with 50 samples.\n\n\n\nBelow is my solution for IMC 2024.\n\n# 1. Introduction\nThe core of my solution is [VGGSfM: Visual Geometry Grounded Deep Structure From Motion](https://vggsfm.github.io/) (https://vggsfm.github.io/). This work has been accepted to CVPR 2024 as a Highlight.\n\nBasically, given N input frames, VGGSfM first uses a camera predictor to estimate preliminary camera extrinsic and intrinsic parameters. After selecting a query frame and some 2D query points, it estimates 2D tracks using a track predictor. These tracks, alongside the preliminary camera data, allow us to perform triangulation to obtain 3D points, followed by a bundle adjustment to refine the camera and point combination. More details can be found in the paper. \n\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16492520%2F0feabaadc93f98e6b15fbe9eea33503b%2Foverview.png?generation=1717614431932034&alt=media\" width=\"1100\" >\n\n\n# 2. Solution\n\n## 2.1 VGGSfM across All Frames\n\nI began by applying VGGSfM across all input frames, which proved effective on the IMC2024 evaluation set. With the input frame number of 50, its mAA is higher than the baseline by 4% (0.04). However, due to the 16 GB GPU limitation on the Kaggle server, accommodating this approach in Kaggle submission was challenging. OOM is bitter. Therefore, to accommodate the server constraints, I integrated VGGSfM into the existing pycolmap pipeline, as detailed below.\n\n## 2.2 Additional Tracks\n\nThe first idea is using VGGSfM track predictor to estimate tracks, and feed them into pycolmap as additional 2D matches. For each image, I identified its N nearest frames using NetVLAD or DINO V2. Using the image as the query frame, I selected query points via Superpoint and found corresponding tracks on the N nearest frames. With N=5, this helped improving mAA by 3% on the evaluation set. However, in the public leaderboard, this only improves the score from ~0.17 to ~0.18. I suspect a higher N could yield better results, but don’t have enough submissions to try. I joined the competition quite late, after half of the competition and with ~120 submissions in total. \n\nThis strategy runs track prediction once for each input frame, avoiding the quadratic complexity issue of exhaustive matching. On the Kaggle cluster, it requires 1.8 seconds per frame, in total 90 seconds for a 50-image scene. On my local A100 cluster, it takes 0.7 seconds per frame.\n\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16492520%2F7a5fc4ec54bb1f5ce36ddad5b04ac55f%2FScreenshot%202024-06-05%20at%2020.02.12.png?generation=1717614456486524&alt=media\" width=\"500\" >\n\n\n\n## 2.3 SfM Track Refinement \n\nThe VGGSfM track predictor adopts a two-stage design, which first predicts coarse tracks and then refine them by local appearance. Inspired by the winner solution by ZJU3DV @xingyihezju3dv last year, I found we can refine the SfM tracks of pycolmap reconstruction. The idea is, first running the baseline with ALIKED+LightGlue (or any other methods), you can find a best reconstruction, the one with most registered images. Each valid 3D point within this reconstruction corresponds to several 2D points, and these 2D points construct a SfM track. Therefore, I extracted the image patches (31x31) for those SfM tracks, refined them by VGGSfM fine track predictor, and updated the pycolmap reconstruction with the updated 2D positions. Then, we run one more global bundle adjustment to optimize the cameras and 3D points based on the refined tracks. For example, in the Church scene, it improves the mean reprojection error from 0.64 to 0.55. \n\nImportantly, this process does not require recompiling colmap. I have provided a code segment here (https://www.kaggle.com/code/jianyuanv/vggsfm-to-refine-pycolmap-tracks) about how to do this by pycolmap and VGGSfM model. \n\nDuring submission, we only refine the top 4096 SfM tracks with highest reprojection errors. This will take around 150 seconds for each scene, i.e., around 30 tracks per second. In the public leaderboard, this improves our solution from ~0.18 to ~0.20.\n\n## 2.4 VGGSfM to relocate missing images \n\nWe also observe that colmap usually cannot register all the images in a scene. Therefore, we try to relocate missing images by VGGSfM. For every missing image, we identified the top 5 registered images it has matches with. If an image cannot find 5 registered images, we use the nearest frames by DINO v2. \n\nNow we have a subset of 6 (1 unregistered + 5 registered) images in total. We use VGGSfM to reconstruct their camera poses. While, the poses predicted by VGGSfM stay in a different camera coordinate to the original camera coordinate of pycolmap reconstruction. We use the algorithm by umeyama to find an alignment transform for the centers of 5 registered images, and apply this alignment transform on the camera pose of the unregistered image. \n\nFor each missing image, this process will take around 5 seconds. It improves the public leaderboard score from ~0.20 to ~0.21. It is worth noting that sometimes the missing image may be too hard, which may not have enough overlaps (matches) with all other frames. In this situation, bundle adjustment may fail, and we use the camera poses predicted by VGGSfM camera predictor only. To be honest I think I did something wrong for this part because the code was written in a rush. I may mess up with some camera coordinate transform, but it works not bad.\n\n## 2.5 Final Solution\n\nIn addition to the methods described above, our final solution also:\n\n- **a.** Handles the rotation using https://github.com/ternaus/check_orientation. If the rotation prediction model predicts a score higher than 0.5 for a frame, we rotate the image by 90, 180, and 270 degrees, and pick the one with a highest number of matches.  \n- **b.** Uses the matches from both ALIKED+LightGlue and SP+LightGlue\n- **c.** Extract an area of interest for transparent images (by DBSCAN on keypoints), and run keypoint detection again on the area of interest. \n\n## 2.6 Ideas tried but not worked obviously \n- **a.** Histograms equalization or CLAHE.\n- **b.** Disambiguation methods such as the ones included in https://github.com/cvg/sfm-disambiguation-colmap or [Doppelgangers](https://github.com/RuojinCai/doppelgangers) (I probably did something wrong with Doppelgangers, but not enough time to debug it).\n- **c.** Homography estimation to filter out outlier matches. \n- **d.** Adjusting pycolmap hyper-parameters.\n- **e.** Test Time Augmentation, such as using different image resolution.\n\n# 3. Learned from IMC\n\nThe three-week journey was great and highly rewarding, despite some sleep deprivation. One of the good stuffs was discovering several ways to reduce the GPU consumption of VGGSfM, which I will soon update on the VGGSfM GitHub. Additionally, it was gratifying to find that VGGSfM can effectively complement colmap when computational resources are limited. And, being familiar with Kaggle took me a lot of time. I hope to do better in the next year.\n\n\n(This article was written very quickly and I may further update it next week. Time to have a short break now. But feel free to drop me an email if you have any question about VGGSfM :)",
    "2857789": "Absolutely fantastic! It's a powerful and elegant algorithm. I'm studying your code. The only thing is, the GPU requirements are a bit high.",
    "2857808": "An brilliant solution！I have read your paper and gained a lot. Congratulations on winning third place in this competition.",
    "2858683": "yes, luckily I have learned something in this IMC. I am going to update the code and it should save the GPU usage by half 😄",
    "2888847": "It has been updated now",
    "2899892": "well written & well deserved position. congrats guys:)"
  },
  "source": "meta"
}