{
  "id": 427143,
  "title": "(Prize Eligible) 7th place Solution using a novel matcher LightGlue",
  "url": "/competitions/image-matching-challenge-2023/writeups/rmd-3dv-prize-eligible-7th-place-solution-using-a-",
  "author_name": "",
  "post_date": "2023-07-27T15:19:26.387Z",
  "votes": 15,
  "comment_count": 1,
  "views": 0,
  "content": "<h1>Intro</h1>\n<p>Hi there, we are Alex, Andri, Felix, Deep and Philipp, 4 master students and one PhD from ETH Zurich.</p>\n<p>First of all, we would like to thank the organizers for hosting this fun and interesting challenge, we have greatly enjoyed it! We also greatly enjoy reading about the solution of other teams, who present captivating and innovative concepts through their fantastic works.</p>\n<p>We have been pushing for a prize-eligible version that could compete against solutions using SuperPoint (SP) and SuperGlue (SG). In order to achieve this, we tried out various replacements for SP such as DISK and ALIKED and moved from SuperGlue to <a href=\"https://github.com/cvg/LightGlue\" target=\"_blank\">LightGlue</a> (LG), a cheaper and more accurate local feature matcher developed at ETHZ which is released under the APACHE license. While LG provided very promising results, we were unable get a satisfactory score without using SP until the very last submission on the last day. This last submission, using an ensemble of ALIKED, DISK and SIFT, gave us enough confidence to choose it as our final submission. However, our best scoring submission would have been an ensemble using DISK, SIFT with SP which would have matched the score of 2nd place (0.562). Additionally, we had another submission that was able to match the score of 5th place on the private leaderboard but did not produce convincing train nor public scores.</p>\n<table>\n<thead>\n<tr>\n<th>Features</th>\n<th>Matchers</th>\n<th>Train</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ALIKED*</td>\n<td>LG</td>\n<td>0.763</td>\n<td>0.361</td>\n<td>0.407</td>\n</tr>\n<tr>\n<td>ALIKED+SIFT*</td>\n<td>LG+NN</td>\n<td>0.594</td>\n<td>0.434</td>\n<td>0.480</td>\n</tr>\n<tr>\n<td>DISK*</td>\n<td>LG</td>\n<td>0.761</td>\n<td>0.386</td>\n<td>0.437</td>\n</tr>\n<tr>\n<td>DISK+SIFT*</td>\n<td>LG+NN</td>\n<td><strong>0.843</strong></td>\n<td>0.438</td>\n<td>0.479</td>\n</tr>\n<tr>\n<td>ALIKED2K+DISK*</td>\n<td>LG+LG</td>\n<td>0.837</td>\n<td>0.444</td>\n<td>0.488</td>\n</tr>\n<tr>\n<td>ALIKED2K+DISK+SIFT**</td>\n<td>LG+LG+NN</td>\n<td>0.837</td>\n<td><strong>0.475</strong></td>\n<td>0.523</td>\n</tr>\n<tr>\n<td>ALIKED2K+DISK+SIFT</td>\n<td>LG(h)+LG(h)+NN</td>\n<td>0.824</td>\n<td>0.450</td>\n<td><strong>0.529</strong></td>\n</tr>\n<tr>\n<td>------------------------</td>\n<td>------------</td>\n<td>---------</td>\n<td>---------</td>\n<td>---------</td>\n</tr>\n<tr>\n<td>DISK+SP</td>\n<td>LG+LG</td>\n<td>0.876</td>\n<td>0.484</td>\n<td><strong>0.562</strong></td>\n</tr>\n<tr>\n<td>DISK+SP*</td>\n<td>LG+SG</td>\n<td>0.880</td>\n<td>0.498</td>\n<td>0.517</td>\n</tr>\n<tr>\n<td>DISK+SIFT+SP*</td>\n<td>LG+NN+LG</td>\n<td><strong>0.890</strong></td>\n<td><strong>0.511</strong></td>\n<td>0.559</td>\n</tr>\n<tr>\n<td>DISK+SIFT+SP*</td>\n<td>LG+NN+SG</td>\n<td>0.867</td>\n<td>T/o</td>\n<td>T/o</td>\n</tr>\n</tbody>\n</table>\n<p>** was our final submission, * have been submitted after the deadline, LG(h) has an increased matching threshold of LightGlue of 0.2 (default is 0.1).</p>\n<h1>LightGlue vs SuperGlue</h1>\n<p>LightGlue is an advanced matching framework developed here at ETH Zurich, which exhibits remarkable efficiency and precision. Its architecture features self- and cross-attention mechanisms, empowering it to make robust match predictions. By employing early pruning and confidence classifications, LightGlue efficiently filters out unmatchable points and terminates computations early, thus avoiding unnecessary processing. LightGlue, along with its training code, is made available under a permissive APACHE license, facilitating broader usage.</p>\n<p>In comparison to SuperGlue when combined with SuperPoint, LightGlue demonstrates superior performance in both accuracy and speed. It notably enhances the scores on the train, public, and private datasets while accomplishing these results in nearly half the time required by alternative methods.</p>\n<table>\n<thead>\n<tr>\n<th>Config</th>\n<th>Train</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>SP+SG</td>\n<td>0.643</td>\n<td>0.361</td>\n<td>0.438</td>\n</tr>\n<tr>\n<td>SP+LG</td>\n<td>0.650</td>\n<td>0.384</td>\n<td>0.461</td>\n</tr>\n</tbody>\n</table>\n<h1>Method</h1>\n<p>We developed a modular pipeline that can be called with various arguments, enabling us to try out different configurations and combine methods very easily. In our pipeline, we made heavy use of hloc, which we used as a starting point.</p>\n<h2>Image Retrieval</h2>\n<p>To avoid matching all image pairs of a scene in an exhaustive manner, we used NetVLAD to retrieve the top k images to construct our image pairs. We also tried out CosPlace but did not observe any notable improvements over NetVLAD. Depending on the configuration of each run, we either used <em>k=20</em>, <em>30</em> or <em>50</em> due to run time constraints. For our final submission, we used <em>k=30</em>.</p>\n<h2>Feature Extraction</h2>\n<p>For keypoint extraction, we combined and tried multiple alternatives. For all feature extractions, we experimented with different image sizes but finally settled on resizing the larger edge to 1600 as it provided the most robust scores:</p>\n<ul>\n<li>ALIKED: We played around with a few settings and finally chose to add it to our ensemble as it showed promising results on a few train scenes. We had to limit the number of keypoints to 2048 due to run-time limitations.</li>\n<li>DISK: DISK was the most promising replacement for SP. We tried a few different configurations and finally settled with the default using a max of 5000 keypoints.</li>\n<li>SIFT: Due to its rotation invariance and fast matching, adding sift to our ensemble turned out to boost performance, especially for heritage/dioscuri and heritage/cyprus.</li>\n<li>SP: SuperPoint was the best-performing features extractor in all our experiments, however, we did not choose it for our final submission because of its restrictive license.</li>\n</ul>\n<h2>Feature Matching</h2>\n<p>We used NN-ratio to match SIFT features. For the deep features such as DISK, ALIKED and SP, we trained LightGlue on the MegaDepth dataset.</p>\n<h2>Ensembles</h2>\n<p>The ensembles gave us the biggest boost in the score. It allowed us to run extraction and matching for different configurations and combine the matches of all configurations. This basically gives us the benefits of all used methods. The only drawback is the increased run-time and we thus had to decrease the number of retrievals. Adding SIFT was always a good option because it did not increase the run-time by much while helping to deal with rotations.</p>\n<h2>Structure-from-Motion</h2>\n<p>For the reconstruction, we used PixSfM and forced COLMAP to use shared camera parameters for some scenes.</p>\n<h3>Pixel-Perfect-SfM</h3>\n<p>We added PixSfM (after compiling a wheel for manylinux, following the build pipeline of pycolmap) as an additional refinement step to the reconstruction process. During our experiments, we noted that using PixSfM decreased the score on scenes with rotated images as the S2DNet features are not rotation invariant. We thus only used it if no rotations are found in the scene. Due to the large number of keypoints in our ensemble, we had to use the low memory configuration in all scenes, even on the very small ones.</p>\n<h3>Shared Camera Parameters</h3>\n<p>We noticed that most scenes have been taken with the same camera and therefore decided to force COLMAP to use the same camera for all images in a scene if all images have the same shape. This turned out to be especially valuable on the haiper scenes where COLMAP assigned multiple cameras.</p>\n<h2>Localizing Unregistered Images</h2>\n<p>Some images were not registered, even with a high number of matches to registered ones, possibly because the assumption of shared intrinsics was not always valid. We, therefore, introduced a post-processing step where we used the hloc toolbox to estimate the pose of unregistered images. Specifically, we checked if the camera of an unregistered image is already in the reconstruction database. If that was not the case, we would infer it from the exif data.</p>\n<h1>Other things tried</h1>\n<ul>\n<li>rotating images → We used an image orientation prediction model to correct for rotations. This worked well on the training set but reduced our score significantly upon submission.</li>\n<li>Inspired by last year's solutions, use cropping to focus matching on important regions between image pairs → Became infeasible as we would have a different set of keypoints for each pair of images used for matching.</li>\n<li>Other feature extractors and matcher such as a reimplementation of SP and dense matchers such as LoFTR, DKM → did not improve results or too slow, also unclear license for SP reimplementation.</li>\n<li>Estimated relative in-plane rotation pairwise from sift matches and then estimated the rotation for each image by propagating the rotation through the maximum spanning tree of pairwise matches. → Worked sometimes on Dioscuri but failed on other scenes.</li>\n<li>Resize for sfm did not help.</li>\n</ul>\n<h1>Acknowledgments</h1>\n<p>We would like to thank Philipp Lindenberger for his awesome guidance, tips, and support. We also want to give a huge credit to his novel matcher LightGlue. We also want to thank the <a href=\"https://cvg.ethz.ch\" target=\"_blank\">Computer Vision and Geometry Group, ETH Zurich</a> for the awesome project that started all this.</p>\n<h1>Links</h1>\n<ul>\n<li><a href=\"https://github.com/cvg/LightGlue\" target=\"_blank\">LightGlue Repo</a></li>\n<li><a href=\"https://arxiv.org/pdf/2306.13643.pdf\" target=\"_blank\">LightGlue Paper</a></li>\n<li><a href=\"https://github.com/veichta/IMC-2023\" target=\"_blank\">Solution Repo</a></li>\n<li><a href=\"https://www.kaggle.com/code/alexanderveicht/imc2023-from-repo\" target=\"_blank\">Kaggle Notebook</a></li>\n</ul>\n<h2>Per Scene Train Scores</h2>\n<h3>Heritage</h3>\n<table>\n<thead>\n<tr>\n<th>Features</th>\n<th>Matchers</th>\n<th>Cyprus</th>\n<th>Dioscuri</th>\n<th>Wall</th>\n<th>Overall</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ALIKED</td>\n<td>LG</td>\n<td>0.850</td>\n<td>0.684</td>\n<td><strong>0.967</strong></td>\n<td><strong>0.833</strong></td>\n</tr>\n<tr>\n<td>DISK</td>\n<td>LG</td>\n<td>0.314</td>\n<td>0.592</td>\n<td>0.843</td>\n<td>0.583</td>\n</tr>\n<tr>\n<td>ALIKED+SIFT</td>\n<td>LG+NN</td>\n<td>0.991</td>\n<td>0.772</td>\n<td>0.436</td>\n<td>0.733</td>\n</tr>\n<tr>\n<td>DISK+SIFT</td>\n<td>LG+NN</td>\n<td>0.993</td>\n<td>0.624</td>\n<td>0.756</td>\n<td>0.791</td>\n</tr>\n<tr>\n<td>ALIKED2K+DISK</td>\n<td>LG+LG</td>\n<td>0.792</td>\n<td>0.712</td>\n<td>0.930</td>\n<td>0.811</td>\n</tr>\n<tr>\n<td>ALIKED2K+DISK+SIFT**</td>\n<td>LG+LG+NN</td>\n<td><strong>0.993</strong></td>\n<td><strong>0.802</strong></td>\n<td>0.595</td>\n<td>0.796</td>\n</tr>\n<tr>\n<td>ALIKED2K+DISK+SIFT</td>\n<td>LG(h)+LG(h)+NN</td>\n<td><strong>0.993</strong></td>\n<td><strong>0.802</strong></td>\n<td>0.595</td>\n<td>0.796</td>\n</tr>\n</tbody>\n</table>\n<h3>Haiper</h3>\n<table>\n<thead>\n<tr>\n<th>Features</th>\n<th>Matchers</th>\n<th>bike</th>\n<th>chairs</th>\n<th>fountain</th>\n<th>Overall</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ALIKED</td>\n<td>LG</td>\n<td>0.431</td>\n<td>0.735</td>\n<td><strong>0.998</strong></td>\n<td>0.721</td>\n</tr>\n<tr>\n<td>DISK</td>\n<td>LG</td>\n<td><strong>0.926</strong></td>\n<td>0.799</td>\n<td><strong>0.998</strong></td>\n<td>0.908</td>\n</tr>\n<tr>\n<td>ALIKED+SIFT</td>\n<td>LG+NN</td>\n<td>0.579</td>\n<td>0.931</td>\n<td><strong>0.998</strong></td>\n<td>0.836</td>\n</tr>\n<tr>\n<td>DISK+SIFT</td>\n<td>LG+NN</td>\n<td>0.917</td>\n<td>0.929</td>\n<td><strong>0.998</strong></td>\n<td>0.948</td>\n</tr>\n<tr>\n<td>ALIKED2K+DISK</td>\n<td>LG+LG</td>\n<td>0.918</td>\n<td>0.812</td>\n<td><strong>0.998</strong></td>\n<td>0.909</td>\n</tr>\n<tr>\n<td>ALIKED2K+DISK+SIFT**</td>\n<td>LG+LG+NN</td>\n<td>0.922</td>\n<td>0.801</td>\n<td><strong>0.998</strong></td>\n<td>0.907</td>\n</tr>\n<tr>\n<td>ALIKED2K+DISK+SIFT</td>\n<td>LG(h)+LG(h)+NN</td>\n<td>0.920</td>\n<td>0.934</td>\n<td><strong>0.998</strong></td>\n<td><strong>0.951</strong></td>\n</tr>\n</tbody>\n</table>\n<h3>Urban</h3>\n<table>\n<thead>\n<tr>\n<th>Features</th>\n<th>Matchers</th>\n<th>kyiv-puppet-theater</th>\n<th>Overall</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ALIKED</td>\n<td>LG</td>\n<td>0.735</td>\n<td>0.735</td>\n</tr>\n<tr>\n<td>DISK</td>\n<td>LG</td>\n<td>0.793</td>\n<td>0.793</td>\n</tr>\n<tr>\n<td>ALIKED+SIFT</td>\n<td>LG+NN</td>\n<td>0.215</td>\n<td>0.215</td>\n</tr>\n<tr>\n<td>DISK+SIFT</td>\n<td>LG+NN</td>\n<td>0.789</td>\n<td>0.789</td>\n</tr>\n<tr>\n<td>ALIKED2K+DISK</td>\n<td>LG+LG</td>\n<td>0.742</td>\n<td>0.742</td>\n</tr>\n<tr>\n<td>ALIKED2K+DISK+SIFT**</td>\n<td>LG+LG+NN</td>\n<td>0.806</td>\n<td>0.806</td>\n</tr>\n<tr>\n<td>ALIKED2K+DISK+SIFT</td>\n<td>LG(h)+LG(h)+NN</td>\n<td><strong>0.824</strong></td>\n<td><strong>0.824</strong></td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": "2360205",
      "postDate": "07/26/2023 16:07:55",
      "content": "<h1>Intro</h1>\n<p>Hi there, we are Alex, Andri, Felix, Deep and Philipp, 4 master students and one PhD from ETH Zurich.</p>\n<p>First of all, we would like to thank the organizers for hosting this fun and interesting challenge, we have greatly enjoyed it! We also greatly enjoy reading about the solution of other teams, who present captivating and innovative concepts through their fantastic works.</p>\n<p>We have been pushing for a prize-eligible version that could compete against solutions using SuperPoint (SP) and SuperGlue (SG). In order to achieve this, we tried out various replacements for SP such as DISK and ALIKED and moved from SuperGlue to <a href=\"https://github.com/cvg/LightGlue\" target=\"_blank\">LightGlue</a> (LG), a cheaper and more accurate local feature matcher developed at ETHZ which is released under the APACHE license. While LG provided very promising results, we were unable get a satisfactory score without using SP until the very last submission on the last day. This last submission, using an ensemble of ALIKED, DISK and SIFT, gave us enough confidence to choose it as our final submission. However, our best scoring submission would have been an ensemble using DISK, SIFT with SP which would have matched the score of 2nd place (0.562). Additionally, we had another submission that was able to match the score of 5th place on the private leaderboard but did not produce convincing train nor public scores.</p>\n<table>\n<thead>\n<tr>\n<th>Features</th>\n<th>Matchers</th>\n<th>Train</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ALIKED*</td>\n<td>LG</td>\n<td>0.763</td>\n<td>0.361</td>\n<td>0.407</td>\n</tr>\n<tr>\n<td>ALIKED+SIFT*</td>\n<td>LG+NN</td>\n<td>0.594</td>\n<td>0.434</td>\n<td>0.480</td>\n</tr>\n<tr>\n<td>DISK*</td>\n<td>LG</td>\n<td>0.761</td>\n<td>0.386</td>\n<td>0.437</td>\n</tr>\n<tr>\n<td>DISK+SIFT*</td>\n<td>LG+NN</td>\n<td><strong>0.843</strong></td>\n<td>0.438</td>\n<td>0.479</td>\n</tr>\n<tr>\n<td>ALIKED2K+DISK*</td>\n<td>LG+LG</td>\n<td>0.837</td>\n<td>0.444</td>\n<td>0.488</td>\n</tr>\n<tr>\n<td>ALIKED2K+DISK+SIFT**</td>\n<td>LG+LG+NN</td>\n<td>0.837</td>\n<td><strong>0.475</strong></td>\n<td>0.523</td>\n</tr>\n<tr>\n<td>ALIKED2K+DISK+SIFT</td>\n<td>LG(h)+LG(h)+NN</td>\n<td>0.824</td>\n<td>0.450</td>\n<td><strong>0.529</strong></td>\n</tr>\n<tr>\n<td>------------------------</td>\n<td>------------</td>\n<td>---------</td>\n<td>---------</td>\n<td>---------</td>\n</tr>\n<tr>\n<td>DISK+SP</td>\n<td>LG+LG</td>\n<td>0.876</td>\n<td>0.484</td>\n<td><strong>0.562</strong></td>\n</tr>\n<tr>\n<td>DISK+SP*</td>\n<td>LG+SG</td>\n<td>0.880</td>\n<td>0.498</td>\n<td>0.517</td>\n</tr>\n<tr>\n<td>DISK+SIFT+SP*</td>\n<td>LG+NN+LG</td>\n<td><strong>0.890</strong></td>\n<td><strong>0.511</strong></td>\n<td>0.559</td>\n</tr>\n<tr>\n<td>DISK+SIFT+SP*</td>\n<td>LG+NN+SG</td>\n<td>0.867</td>\n<td>T/o</td>\n<td>T/o</td>\n</tr>\n</tbody>\n</table>\n<p>** was our final submission, * have been submitted after the deadline, LG(h) has an increased matching threshold of LightGlue of 0.2 (default is 0.1).</p>\n<h1>LightGlue vs SuperGlue</h1>\n<p>LightGlue is an advanced matching framework developed here at ETH Zurich, which exhibits remarkable efficiency and precision. Its architecture features self- and cross-attention mechanisms, empowering it to make robust match predictions. By employing early pruning and confidence classifications, LightGlue efficiently filters out unmatchable points and terminates computations early, thus avoiding unnecessary processing. LightGlue, along with its training code, is made available under a permissive APACHE license, facilitating broader usage.</p>\n<p>In comparison to SuperGlue when combined with SuperPoint, LightGlue demonstrates superior performance in both accuracy and speed. It notably enhances the scores on the train, public, and private datasets while accomplishing these results in nearly half the time required by alternative methods.</p>\n<table>\n<thead>\n<tr>\n<th>Config</th>\n<th>Train</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>SP+SG</td>\n<td>0.643</td>\n<td>0.361</td>\n<td>0.438</td>\n</tr>\n<tr>\n<td>SP+LG</td>\n<td>0.650</td>\n<td>0.384</td>\n<td>0.461</td>\n</tr>\n</tbody>\n</table>\n<h1>Method</h1>\n<p>We developed a modular pipeline that can be called with various arguments, enabling us to try out different configurations and combine methods very easily. In our pipeline, we made heavy use of hloc, which we used as a starting point.</p>\n<h2>Image Retrieval</h2>\n<p>To avoid matching all image pairs of a scene in an exhaustive manner, we used NetVLAD to retrieve the top k images to construct our image pairs. We also tried out CosPlace but did not observe any notable improvements over NetVLAD. Depending on the configuration of each run, we either used <em>k=20</em>, <em>30</em> or <em>50</em> due to run time constraints. For our final submission, we used <em>k=30</em>.</p>\n<h2>Feature Extraction</h2>\n<p>For keypoint extraction, we combined and tried multiple alternatives. For all feature extractions, we experimented with different image sizes but finally settled on resizing the larger edge to 1600 as it provided the most robust scores:</p>\n<ul>\n<li>ALIKED: We played around with a few settings and finally chose to add it to our ensemble as it showed promising results on a few train scenes. We had to limit the number of keypoints to 2048 due to run-time limitations.</li>\n<li>DISK: DISK was the most promising replacement for SP. We tried a few different configurations and finally settled with the default using a max of 5000 keypoints.</li>\n<li>SIFT: Due to its rotation invariance and fast matching, adding sift to our ensemble turned out to boost performance, especially for heritage/dioscuri and heritage/cyprus.</li>\n<li>SP: SuperPoint was the best-performing features extractor in all our experiments, however, we did not choose it for our final submission because of its restrictive license.</li>\n</ul>\n<h2>Feature Matching</h2>\n<p>We used NN-ratio to match SIFT features. For the deep features such as DISK, ALIKED and SP, we trained LightGlue on the MegaDepth dataset.</p>\n<h2>Ensembles</h2>\n<p>The ensembles gave us the biggest boost in the score. It allowed us to run extraction and matching for different configurations and combine the matches of all configurations. This basically gives us the benefits of all used methods. The only drawback is the increased run-time and we thus had to decrease the number of retrievals. Adding SIFT was always a good option because it did not increase the run-time by much while helping to deal with rotations.</p>\n<h2>Structure-from-Motion</h2>\n<p>For the reconstruction, we used PixSfM and forced COLMAP to use shared camera parameters for some scenes.</p>\n<h3>Pixel-Perfect-SfM</h3>\n<p>We added PixSfM (after compiling a wheel for manylinux, following the build pipeline of pycolmap) as an additional refinement step to the reconstruction process. During our experiments, we noted that using PixSfM decreased the score on scenes with rotated images as the S2DNet features are not rotation invariant. We thus only used it if no rotations are found in the scene. Due to the large number of keypoints in our ensemble, we had to use the low memory configuration in all scenes, even on the very small ones.</p>\n<h3>Shared Camera Parameters</h3>\n<p>We noticed that most scenes have been taken with the same camera and therefore decided to force COLMAP to use the same camera for all images in a scene if all images have the same shape. This turned out to be especially valuable on the haiper scenes where COLMAP assigned multiple cameras.</p>\n<h2>Localizing Unregistered Images</h2>\n<p>Some images were not registered, even with a high number of matches to registered ones, possibly because the assumption of shared intrinsics was not always valid. We, therefore, introduced a post-processing step where we used the hloc toolbox to estimate the pose of unregistered images. Specifically, we checked if the camera of an unregistered image is already in the reconstruction database. If that was not the case, we would infer it from the exif data.</p>\n<h1>Other things tried</h1>\n<ul>\n<li>rotating images → We used an image orientation prediction model to correct for rotations. This worked well on the training set but reduced our score significantly upon submission.</li>\n<li>Inspired by last year's solutions, use cropping to focus matching on important regions between image pairs → Became infeasible as we would have a different set of keypoints for each pair of images used for matching.</li>\n<li>Other feature extractors and matcher such as a reimplementation of SP and dense matchers such as LoFTR, DKM → did not improve results or too slow, also unclear license for SP reimplementation.</li>\n<li>Estimated relative in-plane rotation pairwise from sift matches and then estimated the rotation for each image by propagating the rotation through the maximum spanning tree of pairwise matches. → Worked sometimes on Dioscuri but failed on other scenes.</li>\n<li>Resize for sfm did not help.</li>\n</ul>\n<h1>Acknowledgments</h1>\n<p>We would like to thank Philipp Lindenberger for his awesome guidance, tips, and support. We also want to give a huge credit to his novel matcher LightGlue. We also want to thank the <a href=\"https://cvg.ethz.ch\" target=\"_blank\">Computer Vision and Geometry Group, ETH Zurich</a> for the awesome project that started all this.</p>\n<h1>Links</h1>\n<ul>\n<li><a href=\"https://github.com/cvg/LightGlue\" target=\"_blank\">LightGlue Repo</a></li>\n<li><a href=\"https://arxiv.org/pdf/2306.13643.pdf\" target=\"_blank\">LightGlue Paper</a></li>\n<li><a href=\"https://github.com/veichta/IMC-2023\" target=\"_blank\">Solution Repo</a></li>\n<li><a href=\"https://www.kaggle.com/code/alexanderveicht/imc2023-from-repo\" target=\"_blank\">Kaggle Notebook</a></li>\n</ul>\n<h2>Per Scene Train Scores</h2>\n<h3>Heritage</h3>\n<table>\n<thead>\n<tr>\n<th>Features</th>\n<th>Matchers</th>\n<th>Cyprus</th>\n<th>Dioscuri</th>\n<th>Wall</th>\n<th>Overall</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ALIKED</td>\n<td>LG</td>\n<td>0.850</td>\n<td>0.684</td>\n<td><strong>0.967</strong></td>\n<td><strong>0.833</strong></td>\n</tr>\n<tr>\n<td>DISK</td>\n<td>LG</td>\n<td>0.314</td>\n<td>0.592</td>\n<td>0.843</td>\n<td>0.583</td>\n</tr>\n<tr>\n<td>ALIKED+SIFT</td>\n<td>LG+NN</td>\n<td>0.991</td>\n<td>0.772</td>\n<td>0.436</td>\n<td>0.733</td>\n</tr>\n<tr>\n<td>DISK+SIFT</td>\n<td>LG+NN</td>\n<td>0.993</td>\n<td>0.624</td>\n<td>0.756</td>\n<td>0.791</td>\n</tr>\n<tr>\n<td>ALIKED2K+DISK</td>\n<td>LG+LG</td>\n<td>0.792</td>\n<td>0.712</td>\n<td>0.930</td>\n<td>0.811</td>\n</tr>\n<tr>\n<td>ALIKED2K+DISK+SIFT**</td>\n<td>LG+LG+NN</td>\n<td><strong>0.993</strong></td>\n<td><strong>0.802</strong></td>\n<td>0.595</td>\n<td>0.796</td>\n</tr>\n<tr>\n<td>ALIKED2K+DISK+SIFT</td>\n<td>LG(h)+LG(h)+NN</td>\n<td><strong>0.993</strong></td>\n<td><strong>0.802</strong></td>\n<td>0.595</td>\n<td>0.796</td>\n</tr>\n</tbody>\n</table>\n<h3>Haiper</h3>\n<table>\n<thead>\n<tr>\n<th>Features</th>\n<th>Matchers</th>\n<th>bike</th>\n<th>chairs</th>\n<th>fountain</th>\n<th>Overall</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ALIKED</td>\n<td>LG</td>\n<td>0.431</td>\n<td>0.735</td>\n<td><strong>0.998</strong></td>\n<td>0.721</td>\n</tr>\n<tr>\n<td>DISK</td>\n<td>LG</td>\n<td><strong>0.926</strong></td>\n<td>0.799</td>\n<td><strong>0.998</strong></td>\n<td>0.908</td>\n</tr>\n<tr>\n<td>ALIKED+SIFT</td>\n<td>LG+NN</td>\n<td>0.579</td>\n<td>0.931</td>\n<td><strong>0.998</strong></td>\n<td>0.836</td>\n</tr>\n<tr>\n<td>DISK+SIFT</td>\n<td>LG+NN</td>\n<td>0.917</td>\n<td>0.929</td>\n<td><strong>0.998</strong></td>\n<td>0.948</td>\n</tr>\n<tr>\n<td>ALIKED2K+DISK</td>\n<td>LG+LG</td>\n<td>0.918</td>\n<td>0.812</td>\n<td><strong>0.998</strong></td>\n<td>0.909</td>\n</tr>\n<tr>\n<td>ALIKED2K+DISK+SIFT**</td>\n<td>LG+LG+NN</td>\n<td>0.922</td>\n<td>0.801</td>\n<td><strong>0.998</strong></td>\n<td>0.907</td>\n</tr>\n<tr>\n<td>ALIKED2K+DISK+SIFT</td>\n<td>LG(h)+LG(h)+NN</td>\n<td>0.920</td>\n<td>0.934</td>\n<td><strong>0.998</strong></td>\n<td><strong>0.951</strong></td>\n</tr>\n</tbody>\n</table>\n<h3>Urban</h3>\n<table>\n<thead>\n<tr>\n<th>Features</th>\n<th>Matchers</th>\n<th>kyiv-puppet-theater</th>\n<th>Overall</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ALIKED</td>\n<td>LG</td>\n<td>0.735</td>\n<td>0.735</td>\n</tr>\n<tr>\n<td>DISK</td>\n<td>LG</td>\n<td>0.793</td>\n<td>0.793</td>\n</tr>\n<tr>\n<td>ALIKED+SIFT</td>\n<td>LG+NN</td>\n<td>0.215</td>\n<td>0.215</td>\n</tr>\n<tr>\n<td>DISK+SIFT</td>\n<td>LG+NN</td>\n<td>0.789</td>\n<td>0.789</td>\n</tr>\n<tr>\n<td>ALIKED2K+DISK</td>\n<td>LG+LG</td>\n<td>0.742</td>\n<td>0.742</td>\n</tr>\n<tr>\n<td>ALIKED2K+DISK+SIFT**</td>\n<td>LG+LG+NN</td>\n<td>0.806</td>\n<td>0.806</td>\n</tr>\n<tr>\n<td>ALIKED2K+DISK+SIFT</td>\n<td>LG(h)+LG(h)+NN</td>\n<td><strong>0.824</strong></td>\n<td><strong>0.824</strong></td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "# Intro\n\nHi there, we are Alex, Andri, Felix, Deep and Philipp, 4 master students and one PhD from ETH Zurich.\n\nFirst of all, we would like to thank the organizers for hosting this fun and interesting challenge, we have greatly enjoyed it! We also greatly enjoy reading about the solution of other teams, who present captivating and innovative concepts through their fantastic works.\n\nWe have been pushing for a prize-eligible version that could compete against solutions using SuperPoint (SP) and SuperGlue (SG). In order to achieve this, we tried out various replacements for SP such as DISK and ALIKED and moved from SuperGlue to [LightGlue](https://github.com/cvg/LightGlue) (LG), a cheaper and more accurate local feature matcher developed at ETHZ which is released under the APACHE license. While LG provided very promising results, we were unable get a satisfactory score without using SP until the very last submission on the last day. This last submission, using an ensemble of ALIKED, DISK and SIFT, gave us enough confidence to choose it as our final submission. However, our best scoring submission would have been an ensemble using DISK, SIFT with SP which would have matched the score of 2nd place (0.562). Additionally, we had another submission that was able to match the score of 5th place on the private leaderboard but did not produce convincing train nor public scores.\n\n| Features                 | Matchers       | Train     | Public    | Private   |\n| ------------------------ | -------------- | --------- | --------- | --------- |\n| ALIKED*                  | LG             | 0.763     | 0.361     | 0.407     |\n| ALIKED+SIFT*             | LG+NN          | 0.594     | 0.434     | 0.480     |\n| DISK*                    | LG             | 0.761     | 0.386     | 0.437     |\n| DISK+SIFT*               | LG+NN          | **0.843** | 0.438     | 0.479     |\n| ALIKED2K+DISK*           | LG+LG          | 0.837     | 0.444     | 0.488     |\n| ALIKED2K+DISK+SIFT**     | LG+LG+NN       | 0.837     | **0.475** | 0.523     |\n| ALIKED2K+DISK+SIFT       | LG(h)+LG(h)+NN | 0.824     | 0.450     | **0.529** |\n| ------------------------ | ------------   | --------- | --------- | --------- |\n| DISK+SP                  | LG+LG          | 0.876     | 0.484     | **0.562** |\n| DISK+SP*                 | LG+SG          | 0.880     | 0.498     | 0.517     |\n| DISK+SIFT+SP*            | LG+NN+LG       | **0.890** | **0.511** | 0.559     |\n| DISK+SIFT+SP*            | LG+NN+SG       | 0.867     | T/o       | T/o       |\n\n** was our final submission, * have been submitted after the deadline, LG(h) has an increased matching threshold of LightGlue of 0.2 (default is 0.1).\n\n# LightGlue vs SuperGlue\nLightGlue is an advanced matching framework developed here at ETH Zurich, which exhibits remarkable efficiency and precision. Its architecture features self- and cross-attention mechanisms, empowering it to make robust match predictions. By employing early pruning and confidence classifications, LightGlue efficiently filters out unmatchable points and terminates computations early, thus avoiding unnecessary processing. LightGlue, along with its training code, is made available under a permissive APACHE license, facilitating broader usage.\n\nIn comparison to SuperGlue when combined with SuperPoint, LightGlue demonstrates superior performance in both accuracy and speed. It notably enhances the scores on the train, public, and private datasets while accomplishing these results in nearly half the time required by alternative methods.\n\n| Config | Train | Public | Private |\n| ------ | ----- | ------ | ------- |\n| SP+SG  | 0.643 | 0.361  | 0.438   |\n| SP+LG  | 0.650 | 0.384  | 0.461   |\n\n# Method\nWe developed a modular pipeline that can be called with various arguments, enabling us to try out different configurations and combine methods very easily. In our pipeline, we made heavy use of hloc, which we used as a starting point.\n\n## Image Retrieval\nTo avoid matching all image pairs of a scene in an exhaustive manner, we used NetVLAD to retrieve the top k images to construct our image pairs. We also tried out CosPlace but did not observe any notable improvements over NetVLAD. Depending on the configuration of each run, we either used *k=20*, *30* or *50* due to run time constraints. For our final submission, we used *k=30*.\n\n## Feature Extraction\nFor keypoint extraction, we combined and tried multiple alternatives. For all feature extractions, we experimented with different image sizes but finally settled on resizing the larger edge to 1600 as it provided the most robust scores:\n\n- ALIKED: We played around with a few settings and finally chose to add it to our ensemble as it showed promising results on a few train scenes. We had to limit the number of keypoints to 2048 due to run-time limitations.\n- DISK: DISK was the most promising replacement for SP. We tried a few different configurations and finally settled with the default using a max of 5000 keypoints.\n- SIFT: Due to its rotation invariance and fast matching, adding sift to our ensemble turned out to boost performance, especially for heritage/dioscuri and heritage/cyprus.\n- SP: SuperPoint was the best-performing features extractor in all our experiments, however, we did not choose it for our final submission because of its restrictive license.\n\n## Feature Matching\nWe used NN-ratio to match SIFT features. For the deep features such as DISK, ALIKED and SP, we trained LightGlue on the MegaDepth dataset.\n\n## Ensembles\nThe ensembles gave us the biggest boost in the score. It allowed us to run extraction and matching for different configurations and combine the matches of all configurations. This basically gives us the benefits of all used methods. The only drawback is the increased run-time and we thus had to decrease the number of retrievals. Adding SIFT was always a good option because it did not increase the run-time by much while helping to deal with rotations.\n\n## Structure-from-Motion\nFor the reconstruction, we used PixSfM and forced COLMAP to use shared camera parameters for some scenes.\n\n### Pixel-Perfect-SfM\nWe added PixSfM (after compiling a wheel for manylinux, following the build pipeline of pycolmap) as an additional refinement step to the reconstruction process. During our experiments, we noted that using PixSfM decreased the score on scenes with rotated images as the S2DNet features are not rotation invariant. We thus only used it if no rotations are found in the scene. Due to the large number of keypoints in our ensemble, we had to use the low memory configuration in all scenes, even on the very small ones.\n\n### Shared Camera Parameters\nWe noticed that most scenes have been taken with the same camera and therefore decided to force COLMAP to use the same camera for all images in a scene if all images have the same shape. This turned out to be especially valuable on the haiper scenes where COLMAP assigned multiple cameras.\n\n## Localizing Unregistered Images\nSome images were not registered, even with a high number of matches to registered ones, possibly because the assumption of shared intrinsics was not always valid. We, therefore, introduced a post-processing step where we used the hloc toolbox to estimate the pose of unregistered images. Specifically, we checked if the camera of an unregistered image is already in the reconstruction database. If that was not the case, we would infer it from the exif data.\n\n\n# Other things tried\n\n- rotating images → We used an image orientation prediction model to correct for rotations. This worked well on the training set but reduced our score significantly upon submission.\n- Inspired by last year's solutions, use cropping to focus matching on important regions between image pairs → Became infeasible as we would have a different set of keypoints for each pair of images used for matching.\n- Other feature extractors and matcher such as a reimplementation of SP and dense matchers such as LoFTR, DKM → did not improve results or too slow, also unclear license for SP reimplementation.\n- Estimated relative in-plane rotation pairwise from sift matches and then estimated the rotation for each image by propagating the rotation through the maximum spanning tree of pairwise matches. → Worked sometimes on Dioscuri but failed on other scenes.\n- Resize for sfm did not help.\n\n# Acknowledgments\nWe would like to thank Philipp Lindenberger for his awesome guidance, tips, and support. We also want to give a huge credit to his novel matcher LightGlue. We also want to thank the [Computer Vision and Geometry Group, ETH Zurich](https://cvg.ethz.ch) for the awesome project that started all this.\n\n# Links\n- [LightGlue Repo](https://github.com/cvg/LightGlue)\n- [LightGlue Paper](https://arxiv.org/pdf/2306.13643.pdf)\n- [Solution Repo](https://github.com/veichta/IMC-2023)\n- [Kaggle Notebook](https://www.kaggle.com/code/alexanderveicht/imc2023-from-repo)\n\n## Per Scene Train Scores\n\n### Heritage\n\n| Features             | Matchers       | Cyprus    | Dioscuri  | Wall      | Overall   |\n| -------------------- | -------------- | --------- | --------- | --------- | --------- |\n| ALIKED               | LG             | 0.850     | 0.684     | **0.967** | **0.833** |\n| DISK                 | LG             | 0.314     | 0.592     | 0.843     | 0.583     |\n| ALIKED+SIFT          | LG+NN          | 0.991     | 0.772     | 0.436     | 0.733     |\n| DISK+SIFT            | LG+NN          | 0.993     | 0.624     | 0.756     | 0.791     |\n| ALIKED2K+DISK        | LG+LG          | 0.792     | 0.712     | 0.930     | 0.811     |\n| ALIKED2K+DISK+SIFT** | LG+LG+NN       | **0.993** | **0.802** | 0.595     | 0.796     |\n| ALIKED2K+DISK+SIFT   | LG(h)+LG(h)+NN | **0.993** | **0.802** | 0.595     | 0.796     |\n\n### Haiper\n\n| Features             | Matchers       | bike      | chairs | fountain  | Overall   |\n| -------------------- | -------------- | --------- | ------ | --------- | --------- |\n| ALIKED               | LG             | 0.431     | 0.735  | **0.998** | 0.721     |\n| DISK                 | LG             | **0.926** | 0.799  | **0.998** | 0.908     |\n| ALIKED+SIFT          | LG+NN          | 0.579     | 0.931  | **0.998** | 0.836     |\n| DISK+SIFT            | LG+NN          | 0.917     | 0.929  | **0.998** | 0.948     |\n| ALIKED2K+DISK        | LG+LG          | 0.918     | 0.812  | **0.998** | 0.909     |\n| ALIKED2K+DISK+SIFT** | LG+LG+NN       | 0.922     | 0.801  | **0.998** | 0.907     |\n| ALIKED2K+DISK+SIFT   | LG(h)+LG(h)+NN | 0.920     | 0.934  | **0.998** | **0.951** |\n\n### Urban\n\n| Features             | Matchers       | kyiv-puppet-theater | Overall   |\n| -------------------- | -------------- | ------------------- | --------- |\n| ALIKED               | LG             | 0.735               | 0.735     |\n| DISK                 | LG             | 0.793               | 0.793     |\n| ALIKED+SIFT          | LG+NN          | 0.215               | 0.215     |\n| DISK+SIFT            | LG+NN          | 0.789               | 0.789     |\n| ALIKED2K+DISK        | LG+LG          | 0.742               | 0.742     |\n| ALIKED2K+DISK+SIFT** | LG+LG+NN       | 0.806               | 0.806     |\n| ALIKED2K+DISK+SIFT   | LG(h)+LG(h)+NN | **0.824**           | **0.824** |",
      "votes": null
    },
    {
      "id": "2491837",
      "postDate": "10/22/2023 05:19:55",
      "content": "<p>Thank you for sharing the awesome results! LightGlue seems very valuable for the real-time high-fidelity 3D reconstruction. 👍 (also, no license problem for commercial usage.😁)</p>",
      "rawMarkdown": "Thank you for sharing the awesome results! LightGlue seems very valuable for the real-time high-fidelity 3D reconstruction. 👍 (also, no license problem for commercial usage.😁)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2491837,
      "author_name": "joon9502",
      "author_url": "",
      "post_date": "10/22/2023 05:19:55",
      "content": "<p>Thank you for sharing the awesome results! LightGlue seems very valuable for the real-time high-fidelity 3D reconstruction. 👍 (also, no license problem for commercial usage.😁)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2360205": "# Intro\n\nHi there, we are Alex, Andri, Felix, Deep and Philipp, 4 master students and one PhD from ETH Zurich.\n\nFirst of all, we would like to thank the organizers for hosting this fun and interesting challenge, we have greatly enjoyed it! We also greatly enjoy reading about the solution of other teams, who present captivating and innovative concepts through their fantastic works.\n\nWe have been pushing for a prize-eligible version that could compete against solutions using SuperPoint (SP) and SuperGlue (SG). In order to achieve this, we tried out various replacements for SP such as DISK and ALIKED and moved from SuperGlue to [LightGlue](https://github.com/cvg/LightGlue) (LG), a cheaper and more accurate local feature matcher developed at ETHZ which is released under the APACHE license. While LG provided very promising results, we were unable get a satisfactory score without using SP until the very last submission on the last day. This last submission, using an ensemble of ALIKED, DISK and SIFT, gave us enough confidence to choose it as our final submission. However, our best scoring submission would have been an ensemble using DISK, SIFT with SP which would have matched the score of 2nd place (0.562). Additionally, we had another submission that was able to match the score of 5th place on the private leaderboard but did not produce convincing train nor public scores.\n\n| Features                 | Matchers       | Train     | Public    | Private   |\n| ------------------------ | -------------- | --------- | --------- | --------- |\n| ALIKED*                  | LG             | 0.763     | 0.361     | 0.407     |\n| ALIKED+SIFT*             | LG+NN          | 0.594     | 0.434     | 0.480     |\n| DISK*                    | LG             | 0.761     | 0.386     | 0.437     |\n| DISK+SIFT*               | LG+NN          | **0.843** | 0.438     | 0.479     |\n| ALIKED2K+DISK*           | LG+LG          | 0.837     | 0.444     | 0.488     |\n| ALIKED2K+DISK+SIFT**     | LG+LG+NN       | 0.837     | **0.475** | 0.523     |\n| ALIKED2K+DISK+SIFT       | LG(h)+LG(h)+NN | 0.824     | 0.450     | **0.529** |\n| ------------------------ | ------------   | --------- | --------- | --------- |\n| DISK+SP                  | LG+LG          | 0.876     | 0.484     | **0.562** |\n| DISK+SP*                 | LG+SG          | 0.880     | 0.498     | 0.517     |\n| DISK+SIFT+SP*            | LG+NN+LG       | **0.890** | **0.511** | 0.559     |\n| DISK+SIFT+SP*            | LG+NN+SG       | 0.867     | T/o       | T/o       |\n\n** was our final submission, * have been submitted after the deadline, LG(h) has an increased matching threshold of LightGlue of 0.2 (default is 0.1).\n\n# LightGlue vs SuperGlue\nLightGlue is an advanced matching framework developed here at ETH Zurich, which exhibits remarkable efficiency and precision. Its architecture features self- and cross-attention mechanisms, empowering it to make robust match predictions. By employing early pruning and confidence classifications, LightGlue efficiently filters out unmatchable points and terminates computations early, thus avoiding unnecessary processing. LightGlue, along with its training code, is made available under a permissive APACHE license, facilitating broader usage.\n\nIn comparison to SuperGlue when combined with SuperPoint, LightGlue demonstrates superior performance in both accuracy and speed. It notably enhances the scores on the train, public, and private datasets while accomplishing these results in nearly half the time required by alternative methods.\n\n| Config | Train | Public | Private |\n| ------ | ----- | ------ | ------- |\n| SP+SG  | 0.643 | 0.361  | 0.438   |\n| SP+LG  | 0.650 | 0.384  | 0.461   |\n\n# Method\nWe developed a modular pipeline that can be called with various arguments, enabling us to try out different configurations and combine methods very easily. In our pipeline, we made heavy use of hloc, which we used as a starting point.\n\n## Image Retrieval\nTo avoid matching all image pairs of a scene in an exhaustive manner, we used NetVLAD to retrieve the top k images to construct our image pairs. We also tried out CosPlace but did not observe any notable improvements over NetVLAD. Depending on the configuration of each run, we either used *k=20*, *30* or *50* due to run time constraints. For our final submission, we used *k=30*.\n\n## Feature Extraction\nFor keypoint extraction, we combined and tried multiple alternatives. For all feature extractions, we experimented with different image sizes but finally settled on resizing the larger edge to 1600 as it provided the most robust scores:\n\n- ALIKED: We played around with a few settings and finally chose to add it to our ensemble as it showed promising results on a few train scenes. We had to limit the number of keypoints to 2048 due to run-time limitations.\n- DISK: DISK was the most promising replacement for SP. We tried a few different configurations and finally settled with the default using a max of 5000 keypoints.\n- SIFT: Due to its rotation invariance and fast matching, adding sift to our ensemble turned out to boost performance, especially for heritage/dioscuri and heritage/cyprus.\n- SP: SuperPoint was the best-performing features extractor in all our experiments, however, we did not choose it for our final submission because of its restrictive license.\n\n## Feature Matching\nWe used NN-ratio to match SIFT features. For the deep features such as DISK, ALIKED and SP, we trained LightGlue on the MegaDepth dataset.\n\n## Ensembles\nThe ensembles gave us the biggest boost in the score. It allowed us to run extraction and matching for different configurations and combine the matches of all configurations. This basically gives us the benefits of all used methods. The only drawback is the increased run-time and we thus had to decrease the number of retrievals. Adding SIFT was always a good option because it did not increase the run-time by much while helping to deal with rotations.\n\n## Structure-from-Motion\nFor the reconstruction, we used PixSfM and forced COLMAP to use shared camera parameters for some scenes.\n\n### Pixel-Perfect-SfM\nWe added PixSfM (after compiling a wheel for manylinux, following the build pipeline of pycolmap) as an additional refinement step to the reconstruction process. During our experiments, we noted that using PixSfM decreased the score on scenes with rotated images as the S2DNet features are not rotation invariant. We thus only used it if no rotations are found in the scene. Due to the large number of keypoints in our ensemble, we had to use the low memory configuration in all scenes, even on the very small ones.\n\n### Shared Camera Parameters\nWe noticed that most scenes have been taken with the same camera and therefore decided to force COLMAP to use the same camera for all images in a scene if all images have the same shape. This turned out to be especially valuable on the haiper scenes where COLMAP assigned multiple cameras.\n\n## Localizing Unregistered Images\nSome images were not registered, even with a high number of matches to registered ones, possibly because the assumption of shared intrinsics was not always valid. We, therefore, introduced a post-processing step where we used the hloc toolbox to estimate the pose of unregistered images. Specifically, we checked if the camera of an unregistered image is already in the reconstruction database. If that was not the case, we would infer it from the exif data.\n\n\n# Other things tried\n\n- rotating images → We used an image orientation prediction model to correct for rotations. This worked well on the training set but reduced our score significantly upon submission.\n- Inspired by last year's solutions, use cropping to focus matching on important regions between image pairs → Became infeasible as we would have a different set of keypoints for each pair of images used for matching.\n- Other feature extractors and matcher such as a reimplementation of SP and dense matchers such as LoFTR, DKM → did not improve results or too slow, also unclear license for SP reimplementation.\n- Estimated relative in-plane rotation pairwise from sift matches and then estimated the rotation for each image by propagating the rotation through the maximum spanning tree of pairwise matches. → Worked sometimes on Dioscuri but failed on other scenes.\n- Resize for sfm did not help.\n\n# Acknowledgments\nWe would like to thank Philipp Lindenberger for his awesome guidance, tips, and support. We also want to give a huge credit to his novel matcher LightGlue. We also want to thank the [Computer Vision and Geometry Group, ETH Zurich](https://cvg.ethz.ch) for the awesome project that started all this.\n\n# Links\n- [LightGlue Repo](https://github.com/cvg/LightGlue)\n- [LightGlue Paper](https://arxiv.org/pdf/2306.13643.pdf)\n- [Solution Repo](https://github.com/veichta/IMC-2023)\n- [Kaggle Notebook](https://www.kaggle.com/code/alexanderveicht/imc2023-from-repo)\n\n## Per Scene Train Scores\n\n### Heritage\n\n| Features             | Matchers       | Cyprus    | Dioscuri  | Wall      | Overall   |\n| -------------------- | -------------- | --------- | --------- | --------- | --------- |\n| ALIKED               | LG             | 0.850     | 0.684     | **0.967** | **0.833** |\n| DISK                 | LG             | 0.314     | 0.592     | 0.843     | 0.583     |\n| ALIKED+SIFT          | LG+NN          | 0.991     | 0.772     | 0.436     | 0.733     |\n| DISK+SIFT            | LG+NN          | 0.993     | 0.624     | 0.756     | 0.791     |\n| ALIKED2K+DISK        | LG+LG          | 0.792     | 0.712     | 0.930     | 0.811     |\n| ALIKED2K+DISK+SIFT** | LG+LG+NN       | **0.993** | **0.802** | 0.595     | 0.796     |\n| ALIKED2K+DISK+SIFT   | LG(h)+LG(h)+NN | **0.993** | **0.802** | 0.595     | 0.796     |\n\n### Haiper\n\n| Features             | Matchers       | bike      | chairs | fountain  | Overall   |\n| -------------------- | -------------- | --------- | ------ | --------- | --------- |\n| ALIKED               | LG             | 0.431     | 0.735  | **0.998** | 0.721     |\n| DISK                 | LG             | **0.926** | 0.799  | **0.998** | 0.908     |\n| ALIKED+SIFT          | LG+NN          | 0.579     | 0.931  | **0.998** | 0.836     |\n| DISK+SIFT            | LG+NN          | 0.917     | 0.929  | **0.998** | 0.948     |\n| ALIKED2K+DISK        | LG+LG          | 0.918     | 0.812  | **0.998** | 0.909     |\n| ALIKED2K+DISK+SIFT** | LG+LG+NN       | 0.922     | 0.801  | **0.998** | 0.907     |\n| ALIKED2K+DISK+SIFT   | LG(h)+LG(h)+NN | 0.920     | 0.934  | **0.998** | **0.951** |\n\n### Urban\n\n| Features             | Matchers       | kyiv-puppet-theater | Overall   |\n| -------------------- | -------------- | ------------------- | --------- |\n| ALIKED               | LG             | 0.735               | 0.735     |\n| DISK                 | LG             | 0.793               | 0.793     |\n| ALIKED+SIFT          | LG+NN          | 0.215               | 0.215     |\n| DISK+SIFT            | LG+NN          | 0.789               | 0.789     |\n| ALIKED2K+DISK        | LG+LG          | 0.742               | 0.742     |\n| ALIKED2K+DISK+SIFT** | LG+LG+NN       | 0.806               | 0.806     |\n| ALIKED2K+DISK+SIFT   | LG(h)+LG(h)+NN | **0.824**           | **0.824** |",
    "2491837": "Thank you for sharing the awesome results! LightGlue seems very valuable for the real-time high-fidelity 3D reconstruction. 👍 (also, no license problem for commercial usage.😁)"
  },
  "source": "meta"
}