{
  "id": 417407,
  "title": "1st Place Solution: Sparse + Dense matching, confidence-based merge, SfM, and then iterative refinement",
  "url": "/competitions/image-matching-challenge-2023/discussion/417407",
  "author_name": "Xingyi He",
  "post_date": "2023-06-15T15:20:12.054000",
  "votes": 84,
  "comment_count": 11,
  "views": 0,
  "content": "<h1>0. Introduction</h1>\n<p>We are delighted to be participating in the image matching challenging 2023. Thanks to the organizers, sponsors, and Kaggle staff for their efforts, and congrats to all the participants. We learn a lot from this competition and other participants.</p>\n<p>Our team members include Xingyi He, Dongli Tan, Sida Peng, Jiaming Sun, and Prof. Xiaowei Zhou. We are affiliated with the State Key Lab. of CAD&amp;CG, Zhejiang University. I would like to express my gratitude to my teammates for their hard work and dedication. </p>\n<h1>1. Overview and Motivation</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14597895%2F2b2b6c045d8a2dfa0a536090f025db02%2Fmain_fig.png?generation=1686841291013288&amp;alt=media\" alt=\"Fig.1\"><br>\nWe proposed a coarse-to-fine SfM framework to draw benefits from the recent success of detector-free matchers, while solving the multi-view inconsistency issue of detector-free matchers.<br>\nDue to the time limitation in the competition, we also incorporate the \"light-weight\" sparse feature detection and matching methods to determine image rotation and final overlap region between pairs, where the detector-free matcher will be performed upon.</p>\n<p>However, caused by the multi-view inconsistency of detector-free matchers, directly using matches for SfM will lead to a significant number of 2D and 3D points. It is hard to construct feature tracks, and the incremental mapping phase will be extremely slow.</p>\n<p>Our coarse-to-fine framework solves this issue by first quantizing matches with a confidence-guided merge approach, improving consistency while sacrificing the matching accuracy. We use the merged matches to reconstruct a coarse SfM model.<br>\nThen, we refine the coarse SfM model by a novel iterative refinement pipeline, which iterates between an attention-based multi-view matching module to refine feature tracks and a geometry refinement module to improve the reconstruction accuracy.</p>\n<h1>2. Method</h1>\n<h2>2.1 Image Pair Construction</h2>\n<p>For each image, we select k relevant images using image retrieval method.  Here we haven't found significant differences among different retrieval methods. This could potentially be attributed to the relatively small number of images or scenes in the evaluation dataset.</p>\n<h2>2.2 Matching</h2>\n<h3>2.2.1 Rotation Detection</h3>\n<p>There are some scenes within the competition datasets which contain rotated images. Since many popular learning-based matching methods can not handle this case effectively, Our approach, similar to that of many other participants, involves rotating one of the query images several times[0, π/2, π, 3π/2] and matching it with the target image, respectively. This helps to mitigate the drastic reduction in the number of matching points caused by image rotations.</p>\n<h3>2.2.2 Overlap Detection</h3>\n<p>Like last year's solution, estimating the overlap region is a commonly employed technique. We use the first round of matching to obtain the overlap region and then perform the second round of matching within them. According to the area ratio, we resize the smaller region in one image and align it with the larger region. We find a sparse matcher is capable of balancing efficiency and effectiveness.</p>\n<h3>2.2.3 Matching</h3>\n<p>We find the ensemble of multiple methods tends to outperform any individual method. Due to time constraints, we choose the combination of one sparse method (SPSG) and one dense method (LoFTR). We also find that substitute LoFTR by DKMv3 performs better in this competition.</p>\n<h2>2.3 Multi-view inconsistency problem</h2>\n<p>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14597895%2Fab3bc20584b40131a4fe28b55d7d1530%2Fnon_repeatable_problem.png?generation=1686841396817545&amp;alt=media\"> \n</p>\n<p>As shown in Fig.2, the resulting feature locations of detector-free matchers (e.g., LoFTR) in an image depend on the other image. This pair-dependent nature leads to fragmentary feature tracks when running pair-wise matching over multiple views, which makes detector-free matchers not directly applicable to existing SfM systems (e.g., COLMAP).<br>\nMoreover, as for the sparse detection and matching part, since the cropped image overlap regions are also relevant to the other image, re-detecting keypoints on the cropped images for matching also shares the same multi-view inconsistency issue.<br>\nThis issue is solved by the following coarse-to-fine SfM framework.</p>\n<h2>2.4 Coarse SfM</h2>\n<p>In this phase, we first strive for consistency by merging to reconstruct an initial coarse SfM model, which will be further refined for higher pose accuracy in the refinement phase.</p>\n<h3>2.4.1 Confidence-guided Merge</h3>\n<p>After the matching, we merge matches on each image based on confidence to improve the consistency (repeatability) of matches for SfM. For each image, we first aggregate all its matches with other images and then perform NMS with a window size of 5 to merge matches into points with the local highest confidence, as depicted in Fig.1(2). After the NMS, the number of 2D points can be significantly reduced, and the top 10000 points are selected for each image by sorting the confidence if the total point is still larger than the threshold.</p>\n<h3>2.4.2 Mapping</h3>\n<p>Based on the merged matches, we perform the coarse SfM by COLMAP. Note that the geometry verification is skipped since RANSAC is performed in the matching phase. For the reconstruction of the scene with a large number of images (~250 in this competition), we enable the parallelized bundle adjustment (PBA) in COLMAP. Specifically, since PBA uses a PCG solver, which is an inexact solution to the BA problem and unlike the exact solution of Levenberg-Marquardt (LM) solver used by default in Ceres, we enable the PBA only after a large number of images are registered (i.e., &gt;40). This is based on the intuition that the beginning of reconstruction is of critical importance, and the inexact solution of PBA may lead to a poor initialization of the scene.</p>\n\n<h2>2.5 Iterative Refinement</h2>\n<p>We proceed to refine the initial SfM model to obtain improved camera poses and point clouds. To this end, we propose an iterative refinement pipeline. Within each iteration, we first enhance the accuracy of feature tracks with a transformer-based multi-view refinement matching module.<br>\nThese refined feature tracks are then fed into a geometry refinement phase which optimizes camera poses and point clouds jointly. The geometry refinement iterates between the geometric-BA and track topology adjustment (including complete tracks, merge tracks, and filter observations). The refinement process can be performed multiple times for higher accuracy.<br>\nOur feature track refinement matching module is trained on the MegaDepth, and more details are in our paper which is soon available on arXiv.</p>\n<table>\n<thead>\n<tr>\n<th>Method</th>\n<th>score(private)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>spsg</td>\n<td>0.482</td>\n</tr>\n<tr>\n<td>spsg+LoFTR</td>\n<td>0.526</td>\n</tr>\n<tr>\n<td>spsg+LoFTR+refine</td>\n<td>0.570</td>\n</tr>\n<tr>\n<td>spsg+DKM^+refine</td>\n<td>0.594</td>\n</tr>\n</tbody>\n</table>\n<p>^only replace LoFTR in the Haiper dataset</p>\n<h1>3. Ideas tried but not worked</h1>\n<h2>3.1 Other retrieval modules</h2>\n<p>Other than NetVLad, we have also tried the Cosplace, as well as using SIFT+NN as a lightweight detector and matcher for retrieval. However, there is no noticeable improvement, even performs slightly worse than NetVLad in our framework. We think this may be because the pair construction is at the very beginning of the overall pipeline, and our framework is pretty robust to the image pair variance.</p>\n<h2>3.2 Other sparse detectors and matchers</h2>\n<p>Other than Superpoint + Superglue, we have also tried Silk + NN, which performs worse than Superpoint + Superglue. I think it may be because we did not successfully tune it to work in our framework.</p>\n<h2>3.3 Other detector-free matchers</h2>\n<p>Other than LoFTR, we also tried Matchformer and AspanFormer in our framework. We find Matcherform performs on par with LoFTR but slower, which will lead to running out of time. AspanFormer performs worse than LoFTR when used in our framework in this challenge.</p>\n<h2>3.4 Visual localization</h2>\n<p>We observe that there may image not successfully registered during mapping. Our idea is to \"focus\" on these images and regard them as a visual localization problem by trying to register them into the existing SfM model. We use a specifically trained version of LoFTR for localization, which can bring ~3% improvement on the provided training dataset. However, we did not have a spare running time quota in submission and, therefore, did not successfully evaluate visual localization in the final submission.</p>\n<h1>4. Some insights</h1>\n<h2>4.1 About the randomness</h2>\n<p>We observe that the ransac performed with matching, the ransac PnP during mapping, and the bundle adjustment multi-threading in COLMAP may contain randomness.<br>\nAfter a careful evaluation, we find the ransac randomness seed in both matching and mapping is fixed. The randomness can be dispelled by setting the number of threads to 1 in COLMAP.<br>\nTherefore, our submission can achieve exactly the same results after multiple rerunning, which helps us to evaluate the performance of our framework.</p>\n<h2>4.2 About the workload of the evaluation machine</h2>\n<p>Given that the randomness problem of our framework is fixed, we observe that the submission during the last week before the DDL is slower than (~20min) the previous submission with the same configuration.<br>\nOur final submission before the DDL using the DKM as a detector-free matcher has run out of time, which we believe may bring improvements, and we decided to choose it as one of our final submissions.<br>\nWe rerun this submission version after the DDL, and it can be successfully finished within the time limit, which achieves 59.4 finally.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14597895%2F93c7812782a9455aaeafabdffec752b9%2Ffinal_shot_2.png?generation=1686842141320946&amp;alt=media\" alt=\"\"></p>\n<h1>5. Acknowledgment</h1>\n<p>The members of our team have participated in the IMC for three consecutive years(IMC 2021, 2022, and 2023), and we are glad to see there are more and more participants in this competition, and the number of submissions achieves a new high this year. We really enjoyed the competition this year since one of the most applications of feature matching is SfM. The organizers remove the limitation of the only matching submission as in IMC2021 but limit the running time and computation resources (a machine with only 2 CPU cores and 1 GPU is provided), which makes the competition more interesting, challenging, and flexible. Thanks to the organizers, sponsors, and Kaggle staff again!</p>\n<h1>6. Suggestions</h1>\n<p>We also have some suggestions that we notice the scenes in this year's competition are mainly outdoor datasets. We think more types of scenes, such as indoor and object-level scenes with severe texture-poor regions, can be added to the competition in the future. In our recent research, we also collected a texture-poor SfM dataset which is object-centric with ground-truth annotations. We think it may be helpful for the future IMC competition, and we are glad to share it with the organizers if needed.</p>\n<p>Special thanks to the authors of the following open-source software and papers: COLMAP, SuperPoint, SuperGlue, LoFTR, DKM, HLoc, pycolmap, Cosplace, NetVlad.</p>",
  "messages": [
    {
      "id": 2303959,
      "postDate": "2023-06-15T15:20:12.053Z",
      "content": "<h1>0. Introduction</h1>\n<p>We are delighted to be participating in the image matching challenging 2023. Thanks to the organizers, sponsors, and Kaggle staff for their efforts, and congrats to all the participants. We learn a lot from this competition and other participants.</p>\n<p>Our team members include Xingyi He, Dongli Tan, Sida Peng, Jiaming Sun, and Prof. Xiaowei Zhou. We are affiliated with the State Key Lab. of CAD&amp;CG, Zhejiang University. I would like to express my gratitude to my teammates for their hard work and dedication. </p>\n<h1>1. Overview and Motivation</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14597895%2F2b2b6c045d8a2dfa0a536090f025db02%2Fmain_fig.png?generation=1686841291013288&amp;alt=media\" alt=\"Fig.1\"><br>\nWe proposed a coarse-to-fine SfM framework to draw benefits from the recent success of detector-free matchers, while solving the multi-view inconsistency issue of detector-free matchers.<br>\nDue to the time limitation in the competition, we also incorporate the \"light-weight\" sparse feature detection and matching methods to determine image rotation and final overlap region between pairs, where the detector-free matcher will be performed upon.</p>\n<p>However, caused by the multi-view inconsistency of detector-free matchers, directly using matches for SfM will lead to a significant number of 2D and 3D points. It is hard to construct feature tracks, and the incremental mapping phase will be extremely slow.</p>\n<p>Our coarse-to-fine framework solves this issue by first quantizing matches with a confidence-guided merge approach, improving consistency while sacrificing the matching accuracy. We use the merged matches to reconstruct a coarse SfM model.<br>\nThen, we refine the coarse SfM model by a novel iterative refinement pipeline, which iterates between an attention-based multi-view matching module to refine feature tracks and a geometry refinement module to improve the reconstruction accuracy.</p>\n<h1>2. Method</h1>\n<h2>2.1 Image Pair Construction</h2>\n<p>For each image, we select k relevant images using image retrieval method.  Here we haven't found significant differences among different retrieval methods. This could potentially be attributed to the relatively small number of images or scenes in the evaluation dataset.</p>\n<h2>2.2 Matching</h2>\n<h3>2.2.1 Rotation Detection</h3>\n<p>There are some scenes within the competition datasets which contain rotated images. Since many popular learning-based matching methods can not handle this case effectively, Our approach, similar to that of many other participants, involves rotating one of the query images several times[0, π/2, π, 3π/2] and matching it with the target image, respectively. This helps to mitigate the drastic reduction in the number of matching points caused by image rotations.</p>\n<h3>2.2.2 Overlap Detection</h3>\n<p>Like last year's solution, estimating the overlap region is a commonly employed technique. We use the first round of matching to obtain the overlap region and then perform the second round of matching within them. According to the area ratio, we resize the smaller region in one image and align it with the larger region. We find a sparse matcher is capable of balancing efficiency and effectiveness.</p>\n<h3>2.2.3 Matching</h3>\n<p>We find the ensemble of multiple methods tends to outperform any individual method. Due to time constraints, we choose the combination of one sparse method (SPSG) and one dense method (LoFTR). We also find that substitute LoFTR by DKMv3 performs better in this competition.</p>\n<h2>2.3 Multi-view inconsistency problem</h2>\n<p>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14597895%2Fab3bc20584b40131a4fe28b55d7d1530%2Fnon_repeatable_problem.png?generation=1686841396817545&amp;alt=media\"> \n</p>\n<p>As shown in Fig.2, the resulting feature locations of detector-free matchers (e.g., LoFTR) in an image depend on the other image. This pair-dependent nature leads to fragmentary feature tracks when running pair-wise matching over multiple views, which makes detector-free matchers not directly applicable to existing SfM systems (e.g., COLMAP).<br>\nMoreover, as for the sparse detection and matching part, since the cropped image overlap regions are also relevant to the other image, re-detecting keypoints on the cropped images for matching also shares the same multi-view inconsistency issue.<br>\nThis issue is solved by the following coarse-to-fine SfM framework.</p>\n<h2>2.4 Coarse SfM</h2>\n<p>In this phase, we first strive for consistency by merging to reconstruct an initial coarse SfM model, which will be further refined for higher pose accuracy in the refinement phase.</p>\n<h3>2.4.1 Confidence-guided Merge</h3>\n<p>After the matching, we merge matches on each image based on confidence to improve the consistency (repeatability) of matches for SfM. For each image, we first aggregate all its matches with other images and then perform NMS with a window size of 5 to merge matches into points with the local highest confidence, as depicted in Fig.1(2). After the NMS, the number of 2D points can be significantly reduced, and the top 10000 points are selected for each image by sorting the confidence if the total point is still larger than the threshold.</p>\n<h3>2.4.2 Mapping</h3>\n<p>Based on the merged matches, we perform the coarse SfM by COLMAP. Note that the geometry verification is skipped since RANSAC is performed in the matching phase. For the reconstruction of the scene with a large number of images (~250 in this competition), we enable the parallelized bundle adjustment (PBA) in COLMAP. Specifically, since PBA uses a PCG solver, which is an inexact solution to the BA problem and unlike the exact solution of Levenberg-Marquardt (LM) solver used by default in Ceres, we enable the PBA only after a large number of images are registered (i.e., &gt;40). This is based on the intuition that the beginning of reconstruction is of critical importance, and the inexact solution of PBA may lead to a poor initialization of the scene.</p>\n\n<h2>2.5 Iterative Refinement</h2>\n<p>We proceed to refine the initial SfM model to obtain improved camera poses and point clouds. To this end, we propose an iterative refinement pipeline. Within each iteration, we first enhance the accuracy of feature tracks with a transformer-based multi-view refinement matching module.<br>\nThese refined feature tracks are then fed into a geometry refinement phase which optimizes camera poses and point clouds jointly. The geometry refinement iterates between the geometric-BA and track topology adjustment (including complete tracks, merge tracks, and filter observations). The refinement process can be performed multiple times for higher accuracy.<br>\nOur feature track refinement matching module is trained on the MegaDepth, and more details are in our paper which is soon available on arXiv.</p>\n<table>\n<thead>\n<tr>\n<th>Method</th>\n<th>score(private)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>spsg</td>\n<td>0.482</td>\n</tr>\n<tr>\n<td>spsg+LoFTR</td>\n<td>0.526</td>\n</tr>\n<tr>\n<td>spsg+LoFTR+refine</td>\n<td>0.570</td>\n</tr>\n<tr>\n<td>spsg+DKM^+refine</td>\n<td>0.594</td>\n</tr>\n</tbody>\n</table>\n<p>^only replace LoFTR in the Haiper dataset</p>\n<h1>3. Ideas tried but not worked</h1>\n<h2>3.1 Other retrieval modules</h2>\n<p>Other than NetVLad, we have also tried the Cosplace, as well as using SIFT+NN as a lightweight detector and matcher for retrieval. However, there is no noticeable improvement, even performs slightly worse than NetVLad in our framework. We think this may be because the pair construction is at the very beginning of the overall pipeline, and our framework is pretty robust to the image pair variance.</p>\n<h2>3.2 Other sparse detectors and matchers</h2>\n<p>Other than Superpoint + Superglue, we have also tried Silk + NN, which performs worse than Superpoint + Superglue. I think it may be because we did not successfully tune it to work in our framework.</p>\n<h2>3.3 Other detector-free matchers</h2>\n<p>Other than LoFTR, we also tried Matchformer and AspanFormer in our framework. We find Matcherform performs on par with LoFTR but slower, which will lead to running out of time. AspanFormer performs worse than LoFTR when used in our framework in this challenge.</p>\n<h2>3.4 Visual localization</h2>\n<p>We observe that there may image not successfully registered during mapping. Our idea is to \"focus\" on these images and regard them as a visual localization problem by trying to register them into the existing SfM model. We use a specifically trained version of LoFTR for localization, which can bring ~3% improvement on the provided training dataset. However, we did not have a spare running time quota in submission and, therefore, did not successfully evaluate visual localization in the final submission.</p>\n<h1>4. Some insights</h1>\n<h2>4.1 About the randomness</h2>\n<p>We observe that the ransac performed with matching, the ransac PnP during mapping, and the bundle adjustment multi-threading in COLMAP may contain randomness.<br>\nAfter a careful evaluation, we find the ransac randomness seed in both matching and mapping is fixed. The randomness can be dispelled by setting the number of threads to 1 in COLMAP.<br>\nTherefore, our submission can achieve exactly the same results after multiple rerunning, which helps us to evaluate the performance of our framework.</p>\n<h2>4.2 About the workload of the evaluation machine</h2>\n<p>Given that the randomness problem of our framework is fixed, we observe that the submission during the last week before the DDL is slower than (~20min) the previous submission with the same configuration.<br>\nOur final submission before the DDL using the DKM as a detector-free matcher has run out of time, which we believe may bring improvements, and we decided to choose it as one of our final submissions.<br>\nWe rerun this submission version after the DDL, and it can be successfully finished within the time limit, which achieves 59.4 finally.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14597895%2F93c7812782a9455aaeafabdffec752b9%2Ffinal_shot_2.png?generation=1686842141320946&amp;alt=media\" alt=\"\"></p>\n<h1>5. Acknowledgment</h1>\n<p>The members of our team have participated in the IMC for three consecutive years(IMC 2021, 2022, and 2023), and we are glad to see there are more and more participants in this competition, and the number of submissions achieves a new high this year. We really enjoyed the competition this year since one of the most applications of feature matching is SfM. The organizers remove the limitation of the only matching submission as in IMC2021 but limit the running time and computation resources (a machine with only 2 CPU cores and 1 GPU is provided), which makes the competition more interesting, challenging, and flexible. Thanks to the organizers, sponsors, and Kaggle staff again!</p>\n<h1>6. Suggestions</h1>\n<p>We also have some suggestions that we notice the scenes in this year's competition are mainly outdoor datasets. We think more types of scenes, such as indoor and object-level scenes with severe texture-poor regions, can be added to the competition in the future. In our recent research, we also collected a texture-poor SfM dataset which is object-centric with ground-truth annotations. We think it may be helpful for the future IMC competition, and we are glad to share it with the organizers if needed.</p>\n<p>Special thanks to the authors of the following open-source software and papers: COLMAP, SuperPoint, SuperGlue, LoFTR, DKM, HLoc, pycolmap, Cosplace, NetVlad.</p>",
      "rawMarkdown": "# 0. Introduction\nWe are delighted to be participating in the image matching challenging 2023. Thanks to the organizers, sponsors, and Kaggle staff for their efforts, and congrats to all the participants. We learn a lot from this competition and other participants.\n\nOur team members include Xingyi He, Dongli Tan, Sida Peng, Jiaming Sun, and Prof. Xiaowei Zhou. We are affiliated with the State Key Lab. of CAD&CG, Zhejiang University. I would like to express my gratitude to my teammates for their hard work and dedication. \n\n# 1. Overview and Motivation\n![Fig.1](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14597895%2F2b2b6c045d8a2dfa0a536090f025db02%2Fmain_fig.png?generation=1686841291013288&alt=media)\nWe proposed a coarse-to-fine SfM framework to draw benefits from the recent success of detector-free matchers, while solving the multi-view inconsistency issue of detector-free matchers.\nDue to the time limitation in the competition, we also incorporate the \"light-weight\" sparse feature detection and matching methods to determine image rotation and final overlap region between pairs, where the detector-free matcher will be performed upon.\n\nHowever, caused by the multi-view inconsistency of detector-free matchers, directly using matches for SfM will lead to a significant number of 2D and 3D points. It is hard to construct feature tracks, and the incremental mapping phase will be extremely slow.\n\nOur coarse-to-fine framework solves this issue by first quantizing matches with a confidence-guided merge approach, improving consistency while sacrificing the matching accuracy. We use the merged matches to reconstruct a coarse SfM model.\nThen, we refine the coarse SfM model by a novel iterative refinement pipeline, which iterates between an attention-based multi-view matching module to refine feature tracks and a geometry refinement module to improve the reconstruction accuracy.\n\n# 2. Method\n## 2.1 Image Pair Construction\nFor each image, we select k relevant images using image retrieval method.  Here we haven't found significant differences among different retrieval methods. This could potentially be attributed to the relatively small number of images or scenes in the evaluation dataset.\n## 2.2 Matching\n### 2.2.1 Rotation Detection\nThere are some scenes within the competition datasets which contain rotated images. Since many popular learning-based matching methods can not handle this case effectively, Our approach, similar to that of many other participants, involves rotating one of the query images several times[0, π/2, π, 3π/2] and matching it with the target image, respectively. This helps to mitigate the drastic reduction in the number of matching points caused by image rotations.\n### 2.2.2 Overlap Detection\nLike last year's solution, estimating the overlap region is a commonly employed technique. We use the first round of matching to obtain the overlap region and then perform the second round of matching within them. According to the area ratio, we resize the smaller region in one image and align it with the larger region. We find a sparse matcher is capable of balancing efficiency and effectiveness.\n### 2.2.3 Matching\nWe find the ensemble of multiple methods tends to outperform any individual method. Due to time constraints, we choose the combination of one sparse method (SPSG) and one dense method (LoFTR). We also find that substitute LoFTR by DKMv3 performs better in this competition.\n## 2.3 Multi-view inconsistency problem\n<p align=\"center\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14597895%2Fab3bc20584b40131a4fe28b55d7d1530%2Fnon_repeatable_problem.png?generation=1686841396817545&alt=media\"  width=\"50%\" height=\"50%\"> \n</p>\nAs shown in Fig.2, the resulting feature locations of detector-free matchers (e.g., LoFTR) in an image depend on the other image. This pair-dependent nature leads to fragmentary feature tracks when running pair-wise matching over multiple views, which makes detector-free matchers not directly applicable to existing SfM systems (e.g., COLMAP).\nMoreover, as for the sparse detection and matching part, since the cropped image overlap regions are also relevant to the other image, re-detecting keypoints on the cropped images for matching also shares the same multi-view inconsistency issue.\nThis issue is solved by the following coarse-to-fine SfM framework.\n\n## 2.4 Coarse SfM\nIn this phase, we first strive for consistency by merging to reconstruct an initial coarse SfM model, which will be further refined for higher pose accuracy in the refinement phase.\n### 2.4.1 Confidence-guided Merge\nAfter the matching, we merge matches on each image based on confidence to improve the consistency (repeatability) of matches for SfM. For each image, we first aggregate all its matches with other images and then perform NMS with a window size of 5 to merge matches into points with the local highest confidence, as depicted in Fig.1(2). After the NMS, the number of 2D points can be significantly reduced, and the top 10000 points are selected for each image by sorting the confidence if the total point is still larger than the threshold.\n### 2.4.2 Mapping\nBased on the merged matches, we perform the coarse SfM by COLMAP. Note that the geometry verification is skipped since RANSAC is performed in the matching phase. For the reconstruction of the scene with a large number of images (~250 in this competition), we enable the parallelized bundle adjustment (PBA) in COLMAP. Specifically, since PBA uses a PCG solver, which is an inexact solution to the BA problem and unlike the exact solution of Levenberg-Marquardt (LM) solver used by default in Ceres, we enable the PBA only after a large number of images are registered (i.e., >40). This is based on the intuition that the beginning of reconstruction is of critical importance, and the inexact solution of PBA may lead to a poor initialization of the scene.\n\n<!-- 这里提一嘴，pba在40张图之后才启动 -->\n## 2.5 Iterative Refinement\nWe proceed to refine the initial SfM model to obtain improved camera poses and point clouds. To this end, we propose an iterative refinement pipeline. Within each iteration, we first enhance the accuracy of feature tracks with a transformer-based multi-view refinement matching module.\nThese refined feature tracks are then fed into a geometry refinement phase which optimizes camera poses and point clouds jointly. The geometry refinement iterates between the geometric-BA and track topology adjustment (including complete tracks, merge tracks, and filter observations). The refinement process can be performed multiple times for higher accuracy.\nOur feature track refinement matching module is trained on the MegaDepth, and more details are in our paper which is soon available on arXiv.\n\n| Method | score(private) |\n| --- | --- |\n| spsg  | 0.482 |\n|spsg+LoFTR| 0.526|\n|spsg+LoFTR+refine| 0.570 |\n|spsg+DKM^+refine| 0.594|\n\n^only replace LoFTR in the Haiper dataset\n\n# 3. Ideas tried but not worked\n## 3.1 Other retrieval modules\nOther than NetVLad, we have also tried the Cosplace, as well as using SIFT+NN as a lightweight detector and matcher for retrieval. However, there is no noticeable improvement, even performs slightly worse than NetVLad in our framework. We think this may be because the pair construction is at the very beginning of the overall pipeline, and our framework is pretty robust to the image pair variance.\n\n## 3.2 Other sparse detectors and matchers\nOther than Superpoint + Superglue, we have also tried Silk + NN, which performs worse than Superpoint + Superglue. I think it may be because we did not successfully tune it to work in our framework.\n\n## 3.3 Other detector-free matchers\nOther than LoFTR, we also tried Matchformer and AspanFormer in our framework. We find Matcherform performs on par with LoFTR but slower, which will lead to running out of time. AspanFormer performs worse than LoFTR when used in our framework in this challenge.\n\n## 3.4 Visual localization\nWe observe that there may image not successfully registered during mapping. Our idea is to \"focus\" on these images and regard them as a visual localization problem by trying to register them into the existing SfM model. We use a specifically trained version of LoFTR for localization, which can bring ~3% improvement on the provided training dataset. However, we did not have a spare running time quota in submission and, therefore, did not successfully evaluate visual localization in the final submission.\n\n# 4. Some insights\n## 4.1 About the randomness\nWe observe that the ransac performed with matching, the ransac PnP during mapping, and the bundle adjustment multi-threading in COLMAP may contain randomness.\nAfter a careful evaluation, we find the ransac randomness seed in both matching and mapping is fixed. The randomness can be dispelled by setting the number of threads to 1 in COLMAP.\nTherefore, our submission can achieve exactly the same results after multiple rerunning, which helps us to evaluate the performance of our framework.\n\n## 4.2 About the workload of the evaluation machine\nGiven that the randomness problem of our framework is fixed, we observe that the submission during the last week before the DDL is slower than (~20min) the previous submission with the same configuration.\nOur final submission before the DDL using the DKM as a detector-free matcher has run out of time, which we believe may bring improvements, and we decided to choose it as one of our final submissions.\nWe rerun this submission version after the DDL, and it can be successfully finished within the time limit, which achieves 59.4 finally.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14597895%2F93c7812782a9455aaeafabdffec752b9%2Ffinal_shot_2.png?generation=1686842141320946&alt=media)\n# 5. Acknowledgment\nThe members of our team have participated in the IMC for three consecutive years(IMC 2021, 2022, and 2023), and we are glad to see there are more and more participants in this competition, and the number of submissions achieves a new high this year. We really enjoyed the competition this year since one of the most applications of feature matching is SfM. The organizers remove the limitation of the only matching submission as in IMC2021 but limit the running time and computation resources (a machine with only 2 CPU cores and 1 GPU is provided), which makes the competition more interesting, challenging, and flexible. Thanks to the organizers, sponsors, and Kaggle staff again!\n\n# 6. Suggestions\nWe also have some suggestions that we notice the scenes in this year's competition are mainly outdoor datasets. We think more types of scenes, such as indoor and object-level scenes with severe texture-poor regions, can be added to the competition in the future. In our recent research, we also collected a texture-poor SfM dataset which is object-centric with ground-truth annotations. We think it may be helpful for the future IMC competition, and we are glad to share it with the organizers if needed.\n\nSpecial thanks to the authors of the following open-source software and papers: COLMAP, SuperPoint, SuperGlue, LoFTR, DKM, HLoc, pycolmap, Cosplace, NetVlad.",
      "votes": 84
    },
    {
      "id": 2304306,
      "postDate": "2023-06-15T22:49:07.657Z",
      "content": "<p>Respect well deserved, interesting approach i think if you use our modifications on NetVlad you will get above 0.6 probably<br>\nIt gave us about +0.02 on a deterministic submission<br>\nOne trick that will probably also push the score more on the private is as follows:<br>\nIf not all images registered then iterate over other models returned by colmap and take the rotation and translation from it<br>\nIt is mathematically wrong and it needs in this case an extra DLT to make it correct ( fix the relativity between models) but because of the metric it is a hack because the relative error between the unregistered images will decrease even without DLT Just iterate over the other models and register as much as possible it will give about +0.01 to the private leaderboard (Unfortunatly we didn't use it in our final submission) </p>",
      "rawMarkdown": "Respect well deserved, interesting approach i think if you use our modifications on NetVlad you will get above 0.6 probably\nIt gave us about +0.02 on a deterministic submission\nOne trick that will probably also push the score more on the private is as follows:\nIf not all images registered then iterate over other models returned by colmap and take the rotation and translation from it\nIt is mathematically wrong and it needs in this case an extra DLT to make it correct ( fix the relativity between models) but because of the metric it is a hack because the relative error between the unregistered images will decrease even without DLT Just iterate over the other models and register as much as possible it will give about +0.01 to the private leaderboard (Unfortunatly we didn't use it in our final submission) ",
      "votes": 7,
      "replies": [
        {
          "id": 2304432,
          "postDate": "2023-06-16T02:01:29.077Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 2304435,
          "postDate": "2023-06-16T02:03:37.610Z",
          "content": "<p>Thanks very much for the insight! I really like your improvement on the learning-based image retrieval  (NetVlad), which is a phase we did not elaborate on. We are glad to try your method in our framework in the future :).</p>",
          "rawMarkdown": "Thanks very much for the insight! I really like your improvement on the learning-based image retrieval  (NetVlad), which is a phase we did not elaborate on. We are glad to try your method in our framework in the future :).",
          "votes": 2,
          "replies": [
            {
              "id": 2305700,
              "postDate": "2023-06-16T21:20:34.073Z",
              "content": "<p>We would love too let us keep in touch 🙏</p>",
              "rawMarkdown": "We would love too let us keep in touch 🙏"
            }
          ]
        }
      ]
    },
    {
      "id": 2304214,
      "postDate": "2023-06-15T19:41:26.627Z",
      "content": "<p>Great work! Congrats on 1st place! Looking forward to read your paper! </p>",
      "rawMarkdown": "Great work! Congrats on 1st place! Looking forward to read your paper! ",
      "votes": 3,
      "replies": [
        {
          "id": 2305793,
          "postDate": "2023-06-16T23:08:15.780Z",
          "content": "<p><a href=\"https://www.kaggle.com/xingyihezju3dv\" target=\"_blank\">@xingyihezju3dv</a> Eager to learn more details on your approach in slides and the paper at CVPR. Great contribution, solid approach!</p>",
          "rawMarkdown": "@xingyihezju3dv Eager to learn more details on your approach in slides and the paper at CVPR. Great contribution, solid approach!"
        }
      ]
    },
    {
      "id": 2323007,
      "postDate": "2023-06-29T16:10:12.857Z",
      "content": "<p>Wow! This solution is amazing, thanks for sharing your idea.</p>",
      "rawMarkdown": "Wow! This solution is amazing, thanks for sharing your idea."
    },
    {
      "id": 2319800,
      "postDate": "2023-06-27T10:15:24.503Z",
      "content": "<p>Hi Team, Really nice work.<br>\nI wanted to have a look into the notebook, where is it? Can you please share with me ?</p>",
      "rawMarkdown": "Hi Team, Really nice work.\nI wanted to have a look into the notebook, where is it? Can you please share with me ?"
    },
    {
      "id": 2304621,
      "postDate": "2023-06-16T06:07:25.670Z",
      "content": "<p>great solution</p>",
      "rawMarkdown": "great solution"
    },
    {
      "id": 2304187,
      "postDate": "2023-06-15T18:50:30.323Z",
      "content": "<p>This is such a brilliant Solution! Thank you for posting the solution! Well deserved first position! </p>",
      "rawMarkdown": "This is such a brilliant Solution! Thank you for posting the solution! Well deserved first position! "
    },
    {
      "id": 2326904,
      "postDate": "2023-07-02T13:48:45.110Z",
      "content": "<p>thanks for sharing</p>",
      "rawMarkdown": "thanks for sharing",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2304306,
      "author_name": "ammarali32",
      "author_url": "",
      "post_date": "2023-06-15T22:49:07.657000",
      "content": "<p>Respect well deserved, interesting approach i think if you use our modifications on NetVlad you will get above 0.6 probably<br>\nIt gave us about +0.02 on a deterministic submission<br>\nOne trick that will probably also push the score more on the private is as follows:<br>\nIf not all images registered then iterate over other models returned by colmap and take the rotation and translation from it<br>\nIt is mathematically wrong and it needs in this case an extra DLT to make it correct ( fix the relativity between models) but because of the metric it is a hack because the relative error between the unregistered images will decrease even without DLT Just iterate over the other models and register as much as possible it will give about +0.01 to the private leaderboard (Unfortunatly we didn't use it in our final submission) </p>",
      "votes": 7,
      "replies": [
        {
          "id": 2304432,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-06-16T02:01:29.077000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2304435,
          "author_name": "Xingyi He",
          "author_url": "",
          "post_date": "2023-06-16T02:03:37.610000",
          "content": "<p>Thanks very much for the insight! I really like your improvement on the learning-based image retrieval  (NetVlad), which is a phase we did not elaborate on. We are glad to try your method in our framework in the future :).</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2305700,
              "author_name": "ammarali32",
              "author_url": "",
              "post_date": "2023-06-16T21:20:34.073000",
              "content": "<p>We would love too let us keep in touch 🙏</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2304214,
      "author_name": "Jaafar Mahmoud",
      "author_url": "",
      "post_date": "2023-06-15T19:41:26.627000",
      "content": "<p>Great work! Congrats on 1st place! Looking forward to read your paper! </p>",
      "votes": 3,
      "replies": [
        {
          "id": 2305793,
          "author_name": "Igor Lashkov",
          "author_url": "",
          "post_date": "2023-06-16T23:08:15.780000",
          "content": "<p><a href=\"https://www.kaggle.com/xingyihezju3dv\" target=\"_blank\">@xingyihezju3dv</a> Eager to learn more details on your approach in slides and the paper at CVPR. Great contribution, solid approach!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2323007,
      "author_name": "Hoang Pham",
      "author_url": "",
      "post_date": "2023-06-29T16:10:12.857000",
      "content": "<p>Wow! This solution is amazing, thanks for sharing your idea.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2319800,
      "author_name": "Narshima_199712",
      "author_url": "",
      "post_date": "2023-06-27T10:15:24.503000",
      "content": "<p>Hi Team, Really nice work.<br>\nI wanted to have a look into the notebook, where is it? Can you please share with me ?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2304621,
      "author_name": "michael jetson 007",
      "author_url": "",
      "post_date": "2023-06-16T06:07:25.670000",
      "content": "<p>great solution</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2304187,
      "author_name": "Mrinal Mathur",
      "author_url": "",
      "post_date": "2023-06-15T18:50:30.323000",
      "content": "<p>This is such a brilliant Solution! Thank you for posting the solution! Well deserved first position! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2326904,
      "author_name": "Yolanda Li",
      "author_url": "",
      "post_date": "2023-07-02T13:48:45.110000",
      "content": "<p>thanks for sharing</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2303959": "# 0. Introduction\nWe are delighted to be participating in the image matching challenging 2023. Thanks to the organizers, sponsors, and Kaggle staff for their efforts, and congrats to all the participants. We learn a lot from this competition and other participants.\n\nOur team members include Xingyi He, Dongli Tan, Sida Peng, Jiaming Sun, and Prof. Xiaowei Zhou. We are affiliated with the State Key Lab. of CAD&CG, Zhejiang University. I would like to express my gratitude to my teammates for their hard work and dedication. \n\n# 1. Overview and Motivation\n![Fig.1](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14597895%2F2b2b6c045d8a2dfa0a536090f025db02%2Fmain_fig.png?generation=1686841291013288&alt=media)\nWe proposed a coarse-to-fine SfM framework to draw benefits from the recent success of detector-free matchers, while solving the multi-view inconsistency issue of detector-free matchers.\nDue to the time limitation in the competition, we also incorporate the \"light-weight\" sparse feature detection and matching methods to determine image rotation and final overlap region between pairs, where the detector-free matcher will be performed upon.\n\nHowever, caused by the multi-view inconsistency of detector-free matchers, directly using matches for SfM will lead to a significant number of 2D and 3D points. It is hard to construct feature tracks, and the incremental mapping phase will be extremely slow.\n\nOur coarse-to-fine framework solves this issue by first quantizing matches with a confidence-guided merge approach, improving consistency while sacrificing the matching accuracy. We use the merged matches to reconstruct a coarse SfM model.\nThen, we refine the coarse SfM model by a novel iterative refinement pipeline, which iterates between an attention-based multi-view matching module to refine feature tracks and a geometry refinement module to improve the reconstruction accuracy.\n\n# 2. Method\n## 2.1 Image Pair Construction\nFor each image, we select k relevant images using image retrieval method.  Here we haven't found significant differences among different retrieval methods. This could potentially be attributed to the relatively small number of images or scenes in the evaluation dataset.\n## 2.2 Matching\n### 2.2.1 Rotation Detection\nThere are some scenes within the competition datasets which contain rotated images. Since many popular learning-based matching methods can not handle this case effectively, Our approach, similar to that of many other participants, involves rotating one of the query images several times[0, π/2, π, 3π/2] and matching it with the target image, respectively. This helps to mitigate the drastic reduction in the number of matching points caused by image rotations.\n### 2.2.2 Overlap Detection\nLike last year's solution, estimating the overlap region is a commonly employed technique. We use the first round of matching to obtain the overlap region and then perform the second round of matching within them. According to the area ratio, we resize the smaller region in one image and align it with the larger region. We find a sparse matcher is capable of balancing efficiency and effectiveness.\n### 2.2.3 Matching\nWe find the ensemble of multiple methods tends to outperform any individual method. Due to time constraints, we choose the combination of one sparse method (SPSG) and one dense method (LoFTR). We also find that substitute LoFTR by DKMv3 performs better in this competition.\n## 2.3 Multi-view inconsistency problem\n<p align=\"center\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14597895%2Fab3bc20584b40131a4fe28b55d7d1530%2Fnon_repeatable_problem.png?generation=1686841396817545&alt=media\"  width=\"50%\" height=\"50%\"> \n</p>\nAs shown in Fig.2, the resulting feature locations of detector-free matchers (e.g., LoFTR) in an image depend on the other image. This pair-dependent nature leads to fragmentary feature tracks when running pair-wise matching over multiple views, which makes detector-free matchers not directly applicable to existing SfM systems (e.g., COLMAP).\nMoreover, as for the sparse detection and matching part, since the cropped image overlap regions are also relevant to the other image, re-detecting keypoints on the cropped images for matching also shares the same multi-view inconsistency issue.\nThis issue is solved by the following coarse-to-fine SfM framework.\n\n## 2.4 Coarse SfM\nIn this phase, we first strive for consistency by merging to reconstruct an initial coarse SfM model, which will be further refined for higher pose accuracy in the refinement phase.\n### 2.4.1 Confidence-guided Merge\nAfter the matching, we merge matches on each image based on confidence to improve the consistency (repeatability) of matches for SfM. For each image, we first aggregate all its matches with other images and then perform NMS with a window size of 5 to merge matches into points with the local highest confidence, as depicted in Fig.1(2). After the NMS, the number of 2D points can be significantly reduced, and the top 10000 points are selected for each image by sorting the confidence if the total point is still larger than the threshold.\n### 2.4.2 Mapping\nBased on the merged matches, we perform the coarse SfM by COLMAP. Note that the geometry verification is skipped since RANSAC is performed in the matching phase. For the reconstruction of the scene with a large number of images (~250 in this competition), we enable the parallelized bundle adjustment (PBA) in COLMAP. Specifically, since PBA uses a PCG solver, which is an inexact solution to the BA problem and unlike the exact solution of Levenberg-Marquardt (LM) solver used by default in Ceres, we enable the PBA only after a large number of images are registered (i.e., >40). This is based on the intuition that the beginning of reconstruction is of critical importance, and the inexact solution of PBA may lead to a poor initialization of the scene.\n\n<!-- 这里提一嘴，pba在40张图之后才启动 -->\n## 2.5 Iterative Refinement\nWe proceed to refine the initial SfM model to obtain improved camera poses and point clouds. To this end, we propose an iterative refinement pipeline. Within each iteration, we first enhance the accuracy of feature tracks with a transformer-based multi-view refinement matching module.\nThese refined feature tracks are then fed into a geometry refinement phase which optimizes camera poses and point clouds jointly. The geometry refinement iterates between the geometric-BA and track topology adjustment (including complete tracks, merge tracks, and filter observations). The refinement process can be performed multiple times for higher accuracy.\nOur feature track refinement matching module is trained on the MegaDepth, and more details are in our paper which is soon available on arXiv.\n\n| Method | score(private) |\n| --- | --- |\n| spsg  | 0.482 |\n|spsg+LoFTR| 0.526|\n|spsg+LoFTR+refine| 0.570 |\n|spsg+DKM^+refine| 0.594|\n\n^only replace LoFTR in the Haiper dataset\n\n# 3. Ideas tried but not worked\n## 3.1 Other retrieval modules\nOther than NetVLad, we have also tried the Cosplace, as well as using SIFT+NN as a lightweight detector and matcher for retrieval. However, there is no noticeable improvement, even performs slightly worse than NetVLad in our framework. We think this may be because the pair construction is at the very beginning of the overall pipeline, and our framework is pretty robust to the image pair variance.\n\n## 3.2 Other sparse detectors and matchers\nOther than Superpoint + Superglue, we have also tried Silk + NN, which performs worse than Superpoint + Superglue. I think it may be because we did not successfully tune it to work in our framework.\n\n## 3.3 Other detector-free matchers\nOther than LoFTR, we also tried Matchformer and AspanFormer in our framework. We find Matcherform performs on par with LoFTR but slower, which will lead to running out of time. AspanFormer performs worse than LoFTR when used in our framework in this challenge.\n\n## 3.4 Visual localization\nWe observe that there may image not successfully registered during mapping. Our idea is to \"focus\" on these images and regard them as a visual localization problem by trying to register them into the existing SfM model. We use a specifically trained version of LoFTR for localization, which can bring ~3% improvement on the provided training dataset. However, we did not have a spare running time quota in submission and, therefore, did not successfully evaluate visual localization in the final submission.\n\n# 4. Some insights\n## 4.1 About the randomness\nWe observe that the ransac performed with matching, the ransac PnP during mapping, and the bundle adjustment multi-threading in COLMAP may contain randomness.\nAfter a careful evaluation, we find the ransac randomness seed in both matching and mapping is fixed. The randomness can be dispelled by setting the number of threads to 1 in COLMAP.\nTherefore, our submission can achieve exactly the same results after multiple rerunning, which helps us to evaluate the performance of our framework.\n\n## 4.2 About the workload of the evaluation machine\nGiven that the randomness problem of our framework is fixed, we observe that the submission during the last week before the DDL is slower than (~20min) the previous submission with the same configuration.\nOur final submission before the DDL using the DKM as a detector-free matcher has run out of time, which we believe may bring improvements, and we decided to choose it as one of our final submissions.\nWe rerun this submission version after the DDL, and it can be successfully finished within the time limit, which achieves 59.4 finally.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14597895%2F93c7812782a9455aaeafabdffec752b9%2Ffinal_shot_2.png?generation=1686842141320946&alt=media)\n# 5. Acknowledgment\nThe members of our team have participated in the IMC for three consecutive years(IMC 2021, 2022, and 2023), and we are glad to see there are more and more participants in this competition, and the number of submissions achieves a new high this year. We really enjoyed the competition this year since one of the most applications of feature matching is SfM. The organizers remove the limitation of the only matching submission as in IMC2021 but limit the running time and computation resources (a machine with only 2 CPU cores and 1 GPU is provided), which makes the competition more interesting, challenging, and flexible. Thanks to the organizers, sponsors, and Kaggle staff again!\n\n# 6. Suggestions\nWe also have some suggestions that we notice the scenes in this year's competition are mainly outdoor datasets. We think more types of scenes, such as indoor and object-level scenes with severe texture-poor regions, can be added to the competition in the future. In our recent research, we also collected a texture-poor SfM dataset which is object-centric with ground-truth annotations. We think it may be helpful for the future IMC competition, and we are glad to share it with the organizers if needed.\n\nSpecial thanks to the authors of the following open-source software and papers: COLMAP, SuperPoint, SuperGlue, LoFTR, DKM, HLoc, pycolmap, Cosplace, NetVlad.",
    "2304306": "Respect well deserved, interesting approach i think if you use our modifications on NetVlad you will get above 0.6 probably\nIt gave us about +0.02 on a deterministic submission\nOne trick that will probably also push the score more on the private is as follows:\nIf not all images registered then iterate over other models returned by colmap and take the rotation and translation from it\nIt is mathematically wrong and it needs in this case an extra DLT to make it correct ( fix the relativity between models) but because of the metric it is a hack because the relative error between the unregistered images will decrease even without DLT Just iterate over the other models and register as much as possible it will give about +0.01 to the private leaderboard (Unfortunatly we didn't use it in our final submission) ",
    "2304214": "Great work! Congrats on 1st place! Looking forward to read your paper! ",
    "2323007": "Wow! This solution is amazing, thanks for sharing your idea.",
    "2319800": "Hi Team, Really nice work.\nI wanted to have a look into the notebook, where is it? Can you please share with me ?",
    "2304621": "great solution",
    "2304187": "This is such a brilliant Solution! Thank you for posting the solution! Well deserved first position! ",
    "2326904": "thanks for sharing"
  }
}