{
  "id": 416873,
  "title": "2nd Place Solution for the IMC 2023 – Win Over COLMAP Randomness?!",
  "url": "/competitions/image-matching-challenge-2023/writeups/current-resistance-voltage-2nd-place-solution-for-",
  "author_name": "",
  "post_date": "2023-11-08T22:15:26.133Z",
  "votes": 49,
  "comment_count": 12,
  "views": 0,
  "content": "<h1><strong>Bonus Point</strong></h1>\n<p>We presented our solution at <a href=\"https://image-matching-workshop.github.io/\" target=\"_blank\">CVPR 2023</a>. Watch on YouTube <a href=\"https://youtu.be/9JpGjpITiDM?si=l7pGDPw4vOnZsH9H&amp;t=13519\" target=\"_blank\">here</a></p>\n<h1><strong>Intro</strong></h1>\n<p>Our team would like to deeply appreciate the Kaggle staff, Google Research, and Haiper for hosting the continuation of this exciting image matching challenge, as well as everyone here who compete and shared great discussions. My congratulations to all participants!</p>\n<p>The work we describe here is truly a joint effort of <a href=\"https://www.kaggle.com/yamsam\" target=\"_blank\">@yamsam</a>, <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a>, and <a href=\"https://www.kaggle.com/vostankovich\" target=\"_blank\">@vostankovich</a>. I’m grateful of being a part of this hardworking, cohesive, and skilled team. Thanks a lot, guys! I learned a lot from you.</p>\n<h6>We enjoyed this competition!</h6>\n<p>The best submission that was used for final scoring in private LB finished on the last day of the competition. We weren’t completely sure about this submission because it was not clear how much randomness was in it. We had a practice to re-run the same notebook code multiple times to see what scores we can get. We discussed the solution which was implemented on the last day, trusted it and it worked out. The other interesting fact is that our 2nd selected submission for final evaluation scored <strong>0.497/0.542</strong> that also allows us to take 2nd place. This selected second submission is the same as the 1st one except the “Run reconstruction multiple times” trick, that is described below. Anyway, the difference between the best submission (<strong>0.562</strong>) and the 2nd one is noticable. </p>\n<h1><strong>Overview</strong></h1>\n<p>In general, throughout our code submissions every time we fight with the randomness coming from COLMAP responsible for scene reconstruction. Our final solution is based on the use of COLMAP and pretrained SuperPoint/SuperGlue models running on different resolutions for every image in the scene. We apply a bunch of different tricks aimed at different parts of COLMAP-based pipeline in order to stabilize our solution and reach the final score.</p>\n<h1><strong>Architecture</strong></h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5065877%2F68911db7c4cc430dec05670cd196a960%2Fslide_architecture.png?generation=1687202186098466&amp;alt=media\" alt=\"architecture\"></p>\n<h1><strong>Key Takeaways:</strong></h1>\n<ul>\n<li>Initially, use all <strong>possible unique image pairs</strong> generated from the scene set. Remove the model and logic used for finding and ranking similar images in the scene. A threshold of 100 defines a minimum number of matches that we expect each image pair to have. If it is less, we discard that pair.</li>\n<li><strong>SP/SG</strong> settings: unlimited N of keypoints, the keypoint threshold is <strong>0.005</strong>, match threshold <strong>0.2</strong>, and the number of sinkhorn iteratiors is <strong>20</strong>.</li>\n<li><strong>Half precision</strong> for SP/SG helped to reduce occupied memory without sacrificing noticable accuracy. Another great performance trick is to cache keypoints/descriptors, generated by SP for each image, and, then cache SG matches for every image pair. It allowed to reduce the running time a lot.</li>\n<li><strong>TTA</strong>. Ensemble of matches extracted from images at different scales. In our local experiments, the best results are achieved with a combination of <strong>[1088, 1280, 1376]</strong>. We used np.concatenate to join matches from different models that was pretty common for the last IMC22 competition.</li>\n<li>Apply <strong>rotate detector</strong> to un-rotate images in the scene if necessary. We discovered that some scenes in the train dataset (cyprus, dioscuri) have many 90/270 rotated images. Some of the images have EXIF meta information. Unfortunately, looking into the train dataset, we did not find any specifics about the orientation the image was captured at. To address this rotation issue, we re-rotate the image to its natural orientation by employing this <a href=\"https://github.com/ternaus/check_orientation\" target=\"_blank\">solution</a>. We use it w/o any threshold and look at the number of rotations that we need to apply to an image. After applying rotation, the score for cyprus scene jumped up significantly from <strong>~0.02</strong> up to <strong>~0.55</strong>. RotNet implementation did not work out for us.</li>\n<li><strong>Set-up initial image</strong> for COLMAP reconstruction explicitly. For each image we store the number of pairs in which it seen and the number of matches these pairs produce. Then we pick the one with the highest number of pairs. If multiple images satisfy this criterion, we pick the one with the highest number of matches. It helped to boost the score.</li>\n<li>To reduce randomness in the score, we decided to do something like <strong>averaging multiple match_exhaustive</strong> calls. The idea is to run match_exhaustive for N times on the original database of matches. Then we take only those matches that that appear in 8/10 cases, other matches are neglected. It was done in a rude way with database copies write/read, etc.</li>\n<li><strong>Run reconstruction multiple times</strong> from scratch with different matchers threshold e.g. <strong>[100, 125, 75, 100]</strong>. By looking at N of registered images and number of 3D cloud points, we select the best reconstruction. This trick not only allows to find the better reconstruction by finding the better threshold for matches, but also decrease the randomness effect and acts as a countermeasure against a shake up. Due to its running time complexity, we used this strategy only for scenes having less than 40-45 images. This is the last step in our solution that helped us to boost score from <strong>0.497/0.542</strong> to <strong>0.506/0.562</strong>. We also experimented with pycolmap.incremental_mapping employing similar idea but that  scenario did not work out.</li>\n</ul>\n<h1><strong>Ideas that did not work out or not fully tested:</strong></h1>\n<p>•    <strong>TTA multi crop</strong>, no success. The idea was to split image into multiple crops and extract matches in order to find similar images in the scene and determine the best image pairs.<br>\n•    <strong>Square-sized images.</strong><br>\n•    <strong>Bigger image size</strong> (e.g., 1600) for SP/SG.<br>\n•    <strong>Manual RANSAC</strong> instead of using COLMAP internal implementation. We run experiments by disabling geometric verification, but the score was not good.<br>\n•    <strong>NMS filtering</strong> to reduce number of points by using <a href=\"https://github.com/BAILOOL/ANMS-Codes\" target=\"_blank\">ANMS</a>. <br>\n•    Filter least significant image pairs by number of outliers instead of relying on raw matches. It was quite important to look for certain number of matches for an image pair. Run experiments using SP/SG and Loftr. We got higher mAA with Loftr, but, probably, more effort needed here to make it work properly, not enough time.<br>\n•    <strong>Downscale scene images</strong> before passing them to COLMAP.<br>\n•    <strong>Pixel-Perfect Structure-from-Motion</strong>. It was a promising method to evaluate as we got a good boost locally with <a href=\"https://github.com/cvg/pixel-perfect-sfm\" target=\"_blank\">PixSfm</a>, using a single image size of 1280, and it boosted our score from <strong>0.71727</strong> to <strong>0.76253</strong>. Then, we managed to install this framework successfully and run it in Kaggle environment, but could not beat our best score at that moment. It is a heavyweight framework taking too much RAM, and we could run it only for scenes having at most ~30 images. A bit upset because we spent tons of hours compiling all this stuff.<br>\n•    <strong>Adaptive image sizes</strong>. Say, if the longest image side &gt;= 1536 for most images in the scene, we use higher image resolution for matchers ensemble e.g. [1280, 1408, 1536]. Otherwise, a default one is applied [1280,1088,1376]. Did not have enough time to test this idea. It worked locally for cyprus and wall that have big resolution. One of our last submissions implementing this idea crashed with internal error.<br>\n•    Different <strong>detectors, matchers</strong>. We tested Loftr, QuadreeAttention, AspanFormer, DKM v3, GlueStick (keypoints + lines), Patch2Pix, KeyNetAffNetHardNet, DISK, PatchNetVLAD. We also experimented with the confidence matching thresholds and number of matches, but no boost here. Eventually, SP/SG was the best choice for us. Probably, the reason why many dense-based methods did not work out for us is because of the low performance of the “repeatability” metric and high noise in the matches.<br>\n•    Different <strong>CNNs to find the most similar images</strong> in the scene and generate corresponding image pairs (NetVLAD, different pretrained timm-based backbones, CosPlace etc). We even specifically trained a model to find similar images in the scene, but no success here. Later, we gave up using this strategy at all.<br>\n•    Different <strong>keypoint/matching refinement</strong> methods (e.g., recently published <a href=\"https://github.com/TencentYoutuResearch/AdaMatcher\" target=\"_blank\">AdaMatcher</a>, Patch2Pix ), but did not have enough time. AdaMatcher seems a promising idea to try, a quote from their paper “as a refinement network for SP/SG we observe a noticeable improvement in AUC”</p>\n<h1><strong>Performance Improvements Step By Step</strong></h1>\n<table>\n<thead>\n<tr>\n<th>Method</th>\n<th>Private LB</th>\n<th>Public LB</th>\n<th>Δ Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Baseline</td>\n<td>0.382</td>\n<td>0.317</td>\n<td>–</td>\n</tr>\n<tr>\n<td>Manual RANSAC</td>\n<td>0.336</td>\n<td>0.286</td>\n<td>–0.046</td>\n</tr>\n<tr>\n<td>Image size 1024 → 1280</td>\n<td>0.407</td>\n<td>0.345</td>\n<td>+0.025</td>\n</tr>\n<tr>\n<td>Image size 1376, similarity=None</td>\n<td>0.489</td>\n<td>0.426</td>\n<td>+0.026</td>\n</tr>\n<tr>\n<td>Exhaustive matching, 8/10</td>\n<td>0.486</td>\n<td>0.441</td>\n<td>-0.003</td>\n</tr>\n<tr>\n<td>TTA, Image sizes [840, 1024, 1280]</td>\n<td>0.491</td>\n<td>0.447</td>\n<td>+0.002</td>\n</tr>\n<tr>\n<td>TTA, Image sizes [1088, 1280, 1376]</td>\n<td>0.523</td>\n<td>0.475</td>\n<td>+0.032</td>\n</tr>\n<tr>\n<td>Manual image initialization</td>\n<td>0.529</td>\n<td>0.492</td>\n<td>+0.006</td>\n</tr>\n<tr>\n<td>Rotate detection</td>\n<td>0.542</td>\n<td>0.497</td>\n<td>+0.013</td>\n</tr>\n<tr>\n<td>Multi-run reconstruction, matches thr [100, 125, 75, 100]</td>\n<td>0.562</td>\n<td>0.506</td>\n<td>+0.02</td>\n</tr>\n</tbody>\n</table>\n<h5>The final score is 0.506/0.562 in Public/Private.</h5>\n<h1><strong>Local Validation</strong></h1>\n<p>As a reference, this is one of our latest metric reports using train dataset: </p>\n<p>urban / kyiv-puppet-theater (26 images, 325 pairs) -&gt; mAA=0.921538, mAA_q=0.991077, mAA_t=0.921846<br>\nurban -&gt; mAA=0.921538</p>\n<p>heritage / cyprus (30 images, 435 pairs) -&gt; mAA=0.514713, mAA_q=0.525287, mAA_t=0.543678<br>\nheritage / wall (43 images, 903 pairs) -&gt; mAA=0.495792, mAA_q=0.875637, mAA_t=0.509302<br>\nheritage -&gt; mAA=0.505252</p>\n<p>haiper / bike (15 images, 105 pairs) -&gt; mAA=0.940952, mAA_q=0.999048, mAA_t=0.940952<br>\nhaiper / chairs (16 images, 120 pairs) -&gt; mAA=0.834167, mAA_q=0.863333, mAA_t=0.839167<br>\nhaiper / fountain (23 images, 253 pairs) -&gt; mAA=0.999605, mAA_q=1.000000, mAA_t=0.999605<br>\nhaiper -&gt; mAA=0.924908</p>\n<p><strong>Final metric -&gt; mAA=0.783900</strong></p>\n<p>Finally, we had two submissions running on the last day. One of them succeeded and allowed us to get 2nd place, but the other one did not fit the time limit, unexpectedly. Sometimes it is stressful to make a submission on the last day of the competition.</p>\n<h1><strong>Helpful Resources</strong></h1>\n<p>Special thanks to the authors of the following projects:<br>\n•    <a href=\"https://colmap.github.io/\" target=\"_blank\">COLMAP</a><br>\n•    <a href=\"https://ieeexplore.ieee.org/document/7780814\" target=\"_blank\">SuperPoint</a><br>\n•    <a href=\"https://arxiv.org/abs/1911.11763\" target=\"_blank\">SuperGlue</a><br>\n•    <a href=\"https://github.com/ternaus/check_orientation\" target=\"_blank\">check_orientation</a></p>",
  "messages": [
    {
      "id": "2300576",
      "postDate": "06/13/2023 09:08:47",
      "content": "<h1><strong>Bonus Point</strong></h1>\n<p>We presented our solution at <a href=\"https://image-matching-workshop.github.io/\" target=\"_blank\">CVPR 2023</a>. Watch on YouTube <a href=\"https://youtu.be/9JpGjpITiDM?si=l7pGDPw4vOnZsH9H&amp;t=13519\" target=\"_blank\">here</a></p>\n<h1><strong>Intro</strong></h1>\n<p>Our team would like to deeply appreciate the Kaggle staff, Google Research, and Haiper for hosting the continuation of this exciting image matching challenge, as well as everyone here who compete and shared great discussions. My congratulations to all participants!</p>\n<p>The work we describe here is truly a joint effort of <a href=\"https://www.kaggle.com/yamsam\" target=\"_blank\">@yamsam</a>, <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a>, and <a href=\"https://www.kaggle.com/vostankovich\" target=\"_blank\">@vostankovich</a>. I’m grateful of being a part of this hardworking, cohesive, and skilled team. Thanks a lot, guys! I learned a lot from you.</p>\n<h6>We enjoyed this competition!</h6>\n<p>The best submission that was used for final scoring in private LB finished on the last day of the competition. We weren’t completely sure about this submission because it was not clear how much randomness was in it. We had a practice to re-run the same notebook code multiple times to see what scores we can get. We discussed the solution which was implemented on the last day, trusted it and it worked out. The other interesting fact is that our 2nd selected submission for final evaluation scored <strong>0.497/0.542</strong> that also allows us to take 2nd place. This selected second submission is the same as the 1st one except the “Run reconstruction multiple times” trick, that is described below. Anyway, the difference between the best submission (<strong>0.562</strong>) and the 2nd one is noticable. </p>\n<h1><strong>Overview</strong></h1>\n<p>In general, throughout our code submissions every time we fight with the randomness coming from COLMAP responsible for scene reconstruction. Our final solution is based on the use of COLMAP and pretrained SuperPoint/SuperGlue models running on different resolutions for every image in the scene. We apply a bunch of different tricks aimed at different parts of COLMAP-based pipeline in order to stabilize our solution and reach the final score.</p>\n<h1><strong>Architecture</strong></h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5065877%2F68911db7c4cc430dec05670cd196a960%2Fslide_architecture.png?generation=1687202186098466&amp;alt=media\" alt=\"architecture\"></p>\n<h1><strong>Key Takeaways:</strong></h1>\n<ul>\n<li>Initially, use all <strong>possible unique image pairs</strong> generated from the scene set. Remove the model and logic used for finding and ranking similar images in the scene. A threshold of 100 defines a minimum number of matches that we expect each image pair to have. If it is less, we discard that pair.</li>\n<li><strong>SP/SG</strong> settings: unlimited N of keypoints, the keypoint threshold is <strong>0.005</strong>, match threshold <strong>0.2</strong>, and the number of sinkhorn iteratiors is <strong>20</strong>.</li>\n<li><strong>Half precision</strong> for SP/SG helped to reduce occupied memory without sacrificing noticable accuracy. Another great performance trick is to cache keypoints/descriptors, generated by SP for each image, and, then cache SG matches for every image pair. It allowed to reduce the running time a lot.</li>\n<li><strong>TTA</strong>. Ensemble of matches extracted from images at different scales. In our local experiments, the best results are achieved with a combination of <strong>[1088, 1280, 1376]</strong>. We used np.concatenate to join matches from different models that was pretty common for the last IMC22 competition.</li>\n<li>Apply <strong>rotate detector</strong> to un-rotate images in the scene if necessary. We discovered that some scenes in the train dataset (cyprus, dioscuri) have many 90/270 rotated images. Some of the images have EXIF meta information. Unfortunately, looking into the train dataset, we did not find any specifics about the orientation the image was captured at. To address this rotation issue, we re-rotate the image to its natural orientation by employing this <a href=\"https://github.com/ternaus/check_orientation\" target=\"_blank\">solution</a>. We use it w/o any threshold and look at the number of rotations that we need to apply to an image. After applying rotation, the score for cyprus scene jumped up significantly from <strong>~0.02</strong> up to <strong>~0.55</strong>. RotNet implementation did not work out for us.</li>\n<li><strong>Set-up initial image</strong> for COLMAP reconstruction explicitly. For each image we store the number of pairs in which it seen and the number of matches these pairs produce. Then we pick the one with the highest number of pairs. If multiple images satisfy this criterion, we pick the one with the highest number of matches. It helped to boost the score.</li>\n<li>To reduce randomness in the score, we decided to do something like <strong>averaging multiple match_exhaustive</strong> calls. The idea is to run match_exhaustive for N times on the original database of matches. Then we take only those matches that that appear in 8/10 cases, other matches are neglected. It was done in a rude way with database copies write/read, etc.</li>\n<li><strong>Run reconstruction multiple times</strong> from scratch with different matchers threshold e.g. <strong>[100, 125, 75, 100]</strong>. By looking at N of registered images and number of 3D cloud points, we select the best reconstruction. This trick not only allows to find the better reconstruction by finding the better threshold for matches, but also decrease the randomness effect and acts as a countermeasure against a shake up. Due to its running time complexity, we used this strategy only for scenes having less than 40-45 images. This is the last step in our solution that helped us to boost score from <strong>0.497/0.542</strong> to <strong>0.506/0.562</strong>. We also experimented with pycolmap.incremental_mapping employing similar idea but that  scenario did not work out.</li>\n</ul>\n<h1><strong>Ideas that did not work out or not fully tested:</strong></h1>\n<p>•    <strong>TTA multi crop</strong>, no success. The idea was to split image into multiple crops and extract matches in order to find similar images in the scene and determine the best image pairs.<br>\n•    <strong>Square-sized images.</strong><br>\n•    <strong>Bigger image size</strong> (e.g., 1600) for SP/SG.<br>\n•    <strong>Manual RANSAC</strong> instead of using COLMAP internal implementation. We run experiments by disabling geometric verification, but the score was not good.<br>\n•    <strong>NMS filtering</strong> to reduce number of points by using <a href=\"https://github.com/BAILOOL/ANMS-Codes\" target=\"_blank\">ANMS</a>. <br>\n•    Filter least significant image pairs by number of outliers instead of relying on raw matches. It was quite important to look for certain number of matches for an image pair. Run experiments using SP/SG and Loftr. We got higher mAA with Loftr, but, probably, more effort needed here to make it work properly, not enough time.<br>\n•    <strong>Downscale scene images</strong> before passing them to COLMAP.<br>\n•    <strong>Pixel-Perfect Structure-from-Motion</strong>. It was a promising method to evaluate as we got a good boost locally with <a href=\"https://github.com/cvg/pixel-perfect-sfm\" target=\"_blank\">PixSfm</a>, using a single image size of 1280, and it boosted our score from <strong>0.71727</strong> to <strong>0.76253</strong>. Then, we managed to install this framework successfully and run it in Kaggle environment, but could not beat our best score at that moment. It is a heavyweight framework taking too much RAM, and we could run it only for scenes having at most ~30 images. A bit upset because we spent tons of hours compiling all this stuff.<br>\n•    <strong>Adaptive image sizes</strong>. Say, if the longest image side &gt;= 1536 for most images in the scene, we use higher image resolution for matchers ensemble e.g. [1280, 1408, 1536]. Otherwise, a default one is applied [1280,1088,1376]. Did not have enough time to test this idea. It worked locally for cyprus and wall that have big resolution. One of our last submissions implementing this idea crashed with internal error.<br>\n•    Different <strong>detectors, matchers</strong>. We tested Loftr, QuadreeAttention, AspanFormer, DKM v3, GlueStick (keypoints + lines), Patch2Pix, KeyNetAffNetHardNet, DISK, PatchNetVLAD. We also experimented with the confidence matching thresholds and number of matches, but no boost here. Eventually, SP/SG was the best choice for us. Probably, the reason why many dense-based methods did not work out for us is because of the low performance of the “repeatability” metric and high noise in the matches.<br>\n•    Different <strong>CNNs to find the most similar images</strong> in the scene and generate corresponding image pairs (NetVLAD, different pretrained timm-based backbones, CosPlace etc). We even specifically trained a model to find similar images in the scene, but no success here. Later, we gave up using this strategy at all.<br>\n•    Different <strong>keypoint/matching refinement</strong> methods (e.g., recently published <a href=\"https://github.com/TencentYoutuResearch/AdaMatcher\" target=\"_blank\">AdaMatcher</a>, Patch2Pix ), but did not have enough time. AdaMatcher seems a promising idea to try, a quote from their paper “as a refinement network for SP/SG we observe a noticeable improvement in AUC”</p>\n<h1><strong>Performance Improvements Step By Step</strong></h1>\n<table>\n<thead>\n<tr>\n<th>Method</th>\n<th>Private LB</th>\n<th>Public LB</th>\n<th>Δ Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Baseline</td>\n<td>0.382</td>\n<td>0.317</td>\n<td>–</td>\n</tr>\n<tr>\n<td>Manual RANSAC</td>\n<td>0.336</td>\n<td>0.286</td>\n<td>–0.046</td>\n</tr>\n<tr>\n<td>Image size 1024 → 1280</td>\n<td>0.407</td>\n<td>0.345</td>\n<td>+0.025</td>\n</tr>\n<tr>\n<td>Image size 1376, similarity=None</td>\n<td>0.489</td>\n<td>0.426</td>\n<td>+0.026</td>\n</tr>\n<tr>\n<td>Exhaustive matching, 8/10</td>\n<td>0.486</td>\n<td>0.441</td>\n<td>-0.003</td>\n</tr>\n<tr>\n<td>TTA, Image sizes [840, 1024, 1280]</td>\n<td>0.491</td>\n<td>0.447</td>\n<td>+0.002</td>\n</tr>\n<tr>\n<td>TTA, Image sizes [1088, 1280, 1376]</td>\n<td>0.523</td>\n<td>0.475</td>\n<td>+0.032</td>\n</tr>\n<tr>\n<td>Manual image initialization</td>\n<td>0.529</td>\n<td>0.492</td>\n<td>+0.006</td>\n</tr>\n<tr>\n<td>Rotate detection</td>\n<td>0.542</td>\n<td>0.497</td>\n<td>+0.013</td>\n</tr>\n<tr>\n<td>Multi-run reconstruction, matches thr [100, 125, 75, 100]</td>\n<td>0.562</td>\n<td>0.506</td>\n<td>+0.02</td>\n</tr>\n</tbody>\n</table>\n<h5>The final score is 0.506/0.562 in Public/Private.</h5>\n<h1><strong>Local Validation</strong></h1>\n<p>As a reference, this is one of our latest metric reports using train dataset: </p>\n<p>urban / kyiv-puppet-theater (26 images, 325 pairs) -&gt; mAA=0.921538, mAA_q=0.991077, mAA_t=0.921846<br>\nurban -&gt; mAA=0.921538</p>\n<p>heritage / cyprus (30 images, 435 pairs) -&gt; mAA=0.514713, mAA_q=0.525287, mAA_t=0.543678<br>\nheritage / wall (43 images, 903 pairs) -&gt; mAA=0.495792, mAA_q=0.875637, mAA_t=0.509302<br>\nheritage -&gt; mAA=0.505252</p>\n<p>haiper / bike (15 images, 105 pairs) -&gt; mAA=0.940952, mAA_q=0.999048, mAA_t=0.940952<br>\nhaiper / chairs (16 images, 120 pairs) -&gt; mAA=0.834167, mAA_q=0.863333, mAA_t=0.839167<br>\nhaiper / fountain (23 images, 253 pairs) -&gt; mAA=0.999605, mAA_q=1.000000, mAA_t=0.999605<br>\nhaiper -&gt; mAA=0.924908</p>\n<p><strong>Final metric -&gt; mAA=0.783900</strong></p>\n<p>Finally, we had two submissions running on the last day. One of them succeeded and allowed us to get 2nd place, but the other one did not fit the time limit, unexpectedly. Sometimes it is stressful to make a submission on the last day of the competition.</p>\n<h1><strong>Helpful Resources</strong></h1>\n<p>Special thanks to the authors of the following projects:<br>\n•    <a href=\"https://colmap.github.io/\" target=\"_blank\">COLMAP</a><br>\n•    <a href=\"https://ieeexplore.ieee.org/document/7780814\" target=\"_blank\">SuperPoint</a><br>\n•    <a href=\"https://arxiv.org/abs/1911.11763\" target=\"_blank\">SuperGlue</a><br>\n•    <a href=\"https://github.com/ternaus/check_orientation\" target=\"_blank\">check_orientation</a></p>",
      "rawMarkdown": "# **Bonus Point**\nWe presented our solution at [CVPR 2023](https://image-matching-workshop.github.io/). Watch on YouTube [here](https://youtu.be/9JpGjpITiDM?si=l7pGDPw4vOnZsH9H&t=13519)\n\n# **Intro**\n\nOur team would like to deeply appreciate the Kaggle staff, Google Research, and Haiper for hosting the continuation of this exciting image matching challenge, as well as everyone here who compete and shared great discussions. My congratulations to all participants!\n\nThe work we describe here is truly a joint effort of @yamsam, @remekkinas, and @vostankovich. I’m grateful of being a part of this hardworking, cohesive, and skilled team. Thanks a lot, guys! I learned a lot from you.\n\n###### We enjoyed this competition!\n\nThe best submission that was used for final scoring in private LB finished on the last day of the competition. We weren’t completely sure about this submission because it was not clear how much randomness was in it. We had a practice to re-run the same notebook code multiple times to see what scores we can get. We discussed the solution which was implemented on the last day, trusted it and it worked out. The other interesting fact is that our 2nd selected submission for final evaluation scored **0.497/0.542** that also allows us to take 2nd place. This selected second submission is the same as the 1st one except the “Run reconstruction multiple times” trick, that is described below. Anyway, the difference between the best submission (**0.562**) and the 2nd one is noticable. \n\n# **Overview**\n\nIn general, throughout our code submissions every time we fight with the randomness coming from COLMAP responsible for scene reconstruction. Our final solution is based on the use of COLMAP and pretrained SuperPoint/SuperGlue models running on different resolutions for every image in the scene. We apply a bunch of different tricks aimed at different parts of COLMAP-based pipeline in order to stabilize our solution and reach the final score.\n\n# **Architecture**\n![architecture](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5065877%2F68911db7c4cc430dec05670cd196a960%2Fslide_architecture.png?generation=1687202186098466&alt=media)\n\n\n# **Key Takeaways:**\n\n- Initially, use all **possible unique image pairs** generated from the scene set. Remove the model and logic used for finding and ranking similar images in the scene. A threshold of 100 defines a minimum number of matches that we expect each image pair to have. If it is less, we discard that pair.\n- **SP/SG** settings: unlimited N of keypoints, the keypoint threshold is **0.005**, match threshold **0.2**, and the number of sinkhorn iteratiors is **20**.\n- **Half precision** for SP/SG helped to reduce occupied memory without sacrificing noticable accuracy. Another great performance trick is to cache keypoints/descriptors, generated by SP for each image, and, then cache SG matches for every image pair. It allowed to reduce the running time a lot.\n- **TTA**. Ensemble of matches extracted from images at different scales. In our local experiments, the best results are achieved with a combination of **[1088, 1280, 1376]**. We used np.concatenate to join matches from different models that was pretty common for the last IMC22 competition.\n- Apply **rotate detector** to un-rotate images in the scene if necessary. We discovered that some scenes in the train dataset (cyprus, dioscuri) have many 90/270 rotated images. Some of the images have EXIF meta information. Unfortunately, looking into the train dataset, we did not find any specifics about the orientation the image was captured at. To address this rotation issue, we re-rotate the image to its natural orientation by employing this [solution](https://github.com/ternaus/check_orientation ). We use it w/o any threshold and look at the number of rotations that we need to apply to an image. After applying rotation, the score for cyprus scene jumped up significantly from **~0.02** up to **~0.55**. RotNet implementation did not work out for us.\n- **Set-up initial image** for COLMAP reconstruction explicitly. For each image we store the number of pairs in which it seen and the number of matches these pairs produce. Then we pick the one with the highest number of pairs. If multiple images satisfy this criterion, we pick the one with the highest number of matches. It helped to boost the score.\n- To reduce randomness in the score, we decided to do something like **averaging multiple match_exhaustive** calls. The idea is to run match_exhaustive for N times on the original database of matches. Then we take only those matches that that appear in 8/10 cases, other matches are neglected. It was done in a rude way with database copies write/read, etc.\n- **Run reconstruction multiple times** from scratch with different matchers threshold e.g. **[100, 125, 75, 100]**. By looking at N of registered images and number of 3D cloud points, we select the best reconstruction. This trick not only allows to find the better reconstruction by finding the better threshold for matches, but also decrease the randomness effect and acts as a countermeasure against a shake up. Due to its running time complexity, we used this strategy only for scenes having less than 40-45 images. This is the last step in our solution that helped us to boost score from **0.497/0.542** to **0.506/0.562**. We also experimented with pycolmap.incremental_mapping employing similar idea but that  scenario did not work out.\n\n\n# **Ideas that did not work out or not fully tested:**\n•\t**TTA multi crop**, no success. The idea was to split image into multiple crops and extract matches in order to find similar images in the scene and determine the best image pairs.\n•\t**Square-sized images.**\n•\t**Bigger image size** (e.g., 1600) for SP/SG.\n•\t**Manual RANSAC** instead of using COLMAP internal implementation. We run experiments by disabling geometric verification, but the score was not good.\n•\t**NMS filtering** to reduce number of points by using [ANMS](https://github.com/BAILOOL/ANMS-Codes). \n•\tFilter least significant image pairs by number of outliers instead of relying on raw matches. It was quite important to look for certain number of matches for an image pair. Run experiments using SP/SG and Loftr. We got higher mAA with Loftr, but, probably, more effort needed here to make it work properly, not enough time.\n•\t**Downscale scene images** before passing them to COLMAP.\n•\t**Pixel-Perfect Structure-from-Motion**. It was a promising method to evaluate as we got a good boost locally with [PixSfm](https://github.com/cvg/pixel-perfect-sfm), using a single image size of 1280, and it boosted our score from **0.71727** to **0.76253**. Then, we managed to install this framework successfully and run it in Kaggle environment, but could not beat our best score at that moment. It is a heavyweight framework taking too much RAM, and we could run it only for scenes having at most ~30 images. A bit upset because we spent tons of hours compiling all this stuff.\n•\t**Adaptive image sizes**. Say, if the longest image side >= 1536 for most images in the scene, we use higher image resolution for matchers ensemble e.g. [1280, 1408, 1536]. Otherwise, a default one is applied [1280,1088,1376]. Did not have enough time to test this idea. It worked locally for cyprus and wall that have big resolution. One of our last submissions implementing this idea crashed with internal error.\n•\tDifferent **detectors, matchers**. We tested Loftr, QuadreeAttention, AspanFormer, DKM v3, GlueStick (keypoints + lines), Patch2Pix, KeyNetAffNetHardNet, DISK, PatchNetVLAD. We also experimented with the confidence matching thresholds and number of matches, but no boost here. Eventually, SP/SG was the best choice for us. Probably, the reason why many dense-based methods did not work out for us is because of the low performance of the “repeatability” metric and high noise in the matches.\n•\tDifferent **CNNs to find the most similar images** in the scene and generate corresponding image pairs (NetVLAD, different pretrained timm-based backbones, CosPlace etc). We even specifically trained a model to find similar images in the scene, but no success here. Later, we gave up using this strategy at all.\n•\tDifferent **keypoint/matching refinement** methods (e.g., recently published [AdaMatcher](https://github.com/TencentYoutuResearch/AdaMatcher), Patch2Pix ), but did not have enough time. AdaMatcher seems a promising idea to try, a quote from their paper “as a refinement network for SP/SG we observe a noticeable improvement in AUC”\n\n\n# **Performance Improvements Step By Step**\n| Method\t| Private LB\t| Public LB\t| Δ Private LB\n| --- | --- |--- |--- |\n|Baseline\t|0.382\t|0.317\t|–\n|Manual RANSAC\t|0.336\t|0.286\t|–0.046\n|Image size 1024 → 1280\t|0.407\t|0.345\t|+0.025\n|Image size 1376, similarity=None\t|0.489\t|0.426\t|+0.026\n|Exhaustive matching, 8/10\t|0.486\t|0.441\t|-0.003\n|TTA, Image sizes [840, 1024, 1280]\t|0.491\t|0.447\t|+0.002\n|TTA, Image sizes [1088, 1280, 1376]\t|0.523\t|0.475\t|+0.032\n|Manual image initialization\t|0.529\t|0.492\t|+0.006\n|Rotate detection\t|0.542\t|0.497\t|+0.013\n|Multi-run reconstruction, matches thr [100, 125, 75, 100]\t|0.562\t|0.506\t|+0.02\n\n##### The final score is 0.506/0.562 in Public/Private.\n\n# **Local Validation**\nAs a reference, this is one of our latest metric reports using train dataset: \n\nurban / kyiv-puppet-theater (26 images, 325 pairs) -> mAA=0.921538, mAA_q=0.991077, mAA_t=0.921846\nurban -> mAA=0.921538\n\nheritage / cyprus (30 images, 435 pairs) -> mAA=0.514713, mAA_q=0.525287, mAA_t=0.543678\nheritage / wall (43 images, 903 pairs) -> mAA=0.495792, mAA_q=0.875637, mAA_t=0.509302\nheritage -> mAA=0.505252\n\nhaiper / bike (15 images, 105 pairs) -> mAA=0.940952, mAA_q=0.999048, mAA_t=0.940952\nhaiper / chairs (16 images, 120 pairs) -> mAA=0.834167, mAA_q=0.863333, mAA_t=0.839167\nhaiper / fountain (23 images, 253 pairs) -> mAA=0.999605, mAA_q=1.000000, mAA_t=0.999605\nhaiper -> mAA=0.924908\n\n**Final metric -> mAA=0.783900**\n\nFinally, we had two submissions running on the last day. One of them succeeded and allowed us to get 2nd place, but the other one did not fit the time limit, unexpectedly. Sometimes it is stressful to make a submission on the last day of the competition.\n\n# **Helpful Resources**\nSpecial thanks to the authors of the following projects:\n•\t[COLMAP](https://colmap.github.io/)\n•\t[SuperPoint](https://ieeexplore.ieee.org/document/7780814)\n•\t[SuperGlue](https://arxiv.org/abs/1911.11763)\n•\t[check_orientation](https://github.com/ternaus/check_orientation )",
      "votes": null
    },
    {
      "id": "2300584",
      "postDate": "06/13/2023 09:20:16",
      "content": "<p>Congrats with great result!</p>\n<blockquote>\n  <p>Loftr, QuadreeAttention, AspanFormer, DKM v3, GlueStick (keypoints + lines), Patch2Pix, KeyNetAffNetHardNet, DISK, PatchNetVLAD</p>\n</blockquote>\n<p>Could you please share the numbers?</p>\n<p>The same for the global descriptor thing, please?</p>",
      "rawMarkdown": "Congrats with great result!\n\n>Loftr, QuadreeAttention, AspanFormer, DKM v3, GlueStick (keypoints + lines), Patch2Pix, KeyNetAffNetHardNet, DISK, PatchNetVLAD\n\nCould you please share the numbers?\n\nThe same for the global descriptor thing, please?",
      "votes": null
    },
    {
      "id": "2300600",
      "postDate": "06/13/2023 09:34:45",
      "content": "<p>Hey team, I just want to say a big thank you. Your work during this contest was amazing. We came up with so many ideas and things kept changing all the time.</p>\n<p>Special thanks to <a href=\"https://www.kaggle.com/vostankovich\" target=\"_blank\">@vostankovich</a>. You never gave up. I did - I quit when trying to set up pixfm, and that's not like me. I've learned from this. It's important to listen to other people's ideas and not to think that I always know best.</p>\n<p>From the start, <a href=\"https://www.kaggle.com/igorlashkov\" target=\"_blank\">@igorlashkov</a> said we should work on the rotation. I tried a few times but it didn't help me. But guess what? In the end, working on the rotation was a key part of our solution. So, big thanks to Igla for not giving up!</p>\n<p>And thank you, <a href=\"https://www.kaggle.com/oldufo\" target=\"_blank\">@oldufo</a> for talking about 3D reconstruction. I first started looking into this last year, just out of interest. The last contest made me more interested. It's a great way to learn more and get better at what we do. </p>\n<p>I'm really looking forward to next year and IMC 2024. Let's do it!</p>\n<p>One more from my side - the way I was trying to understand SfM pipeline was Meshroom. SfM was new for me so looking only in example notebooks was not enough for me. During the competition I was experimenting a lot playing with different settings to understand \"what if\".</p>\n<p><img src=\"https://i.ibb.co/dgTymVc/Meshroom.jpg\" alt=\"\"></p>",
      "rawMarkdown": "Hey team, I just want to say a big thank you. Your work during this contest was amazing. We came up with so many ideas and things kept changing all the time.\n\nSpecial thanks to @vostankovich. You never gave up. I did - I quit when trying to set up pixfm, and that's not like me. I've learned from this. It's important to listen to other people's ideas and not to think that I always know best.\n\nFrom the start, @igorlashkov said we should work on the rotation. I tried a few times but it didn't help me. But guess what? In the end, working on the rotation was a key part of our solution. So, big thanks to Igla for not giving up!\n\nAnd thank you, @oldufo for talking about 3D reconstruction. I first started looking into this last year, just out of interest. The last contest made me more interested. It's a great way to learn more and get better at what we do. \n\nI'm really looking forward to next year and IMC 2024. Let's do it!\n\n\nOne more from my side - the way I was trying to understand SfM pipeline was Meshroom. SfM was new for me so looking only in example notebooks was not enough for me. During the competition I was experimenting a lot playing with different settings to understand \"what if\".\n\n![](https://i.ibb.co/dgTymVc/Meshroom.jpg)",
      "votes": null
    },
    {
      "id": "2300614",
      "postDate": "06/13/2023 09:49:13",
      "content": "<p>The matchers were tested during start/mid of competition when we had very simple pipeline. We just tested one by one and discarded any matcher that was worse than SP/SG. All of the listed models are worse by a large margin, at least 15-20% lower than SP/SG. Some of the models also were just too slow or memory consuming along with the bad score. The closest one was LoFTR as I remember, but even ensembling it with SP/SG didn't give us any boost. We can rerun some matchers in our final pipeline for sure if needed to get the numbers.</p>\n<p>The same thing with global descriptors, we don't have numbers for them, we just discovered one time that SG can be a retrieval itself and the approach is robust enough and outperforms other models. SP/SG are really fast in half mode and with caching, so we actually completely agreed not to use any additional retrieval models for pairs search.</p>",
      "rawMarkdown": "The matchers were tested during start/mid of competition when we had very simple pipeline. We just tested one by one and discarded any matcher that was worse than SP/SG. All of the listed models are worse by a large margin, at least 15-20% lower than SP/SG. Some of the models also were just too slow or memory consuming along with the bad score. The closest one was LoFTR as I remember, but even ensembling it with SP/SG didn't give us any boost. We can rerun some matchers in our final pipeline for sure if needed to get the numbers.\n\nThe same thing with global descriptors, we don't have numbers for them, we just discovered one time that SG can be a retrieval itself and the approach is robust enough and outperforms other models. SP/SG are really fast in half mode and with caching, so we actually completely agreed not to use any additional retrieval models for pairs search.",
      "votes": null
    },
    {
      "id": "2300617",
      "postDate": "06/13/2023 09:52:24",
      "content": "<p>Congrats on 2nd place! good work!<br>\nPixSFM was really tricky! It gave boost to our validation score up to +0.1 on mAA, but unfortunately it didn't improve our LB.<br>\nYou can try it using this notebook!<br>\n<a href=\"https://www.kaggle.com/jaafarmahmoud1/pixel-perfect-sfm-on-kaggle\" target=\"_blank\">https://www.kaggle.com/jaafarmahmoud1/pixel-perfect-sfm-on-kaggle</a></p>",
      "rawMarkdown": "Congrats on 2nd place! good work!\nPixSFM was really tricky! It gave boost to our validation score up to +0.1 on mAA, but unfortunately it didn't improve our LB.\nYou can try it using this notebook!\nhttps://www.kaggle.com/jaafarmahmoud1/pixel-perfect-sfm-on-kaggle",
      "votes": null
    },
    {
      "id": "2300811",
      "postDate": "06/13/2023 13:15:35",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/igorlashkov\" target=\"_blank\">@igorlashkov</a> , congratulations for the prize and gold. I have a question</p>\n<p>How to run reconstruction multiple times? I think the reconstruction is in <code>pycolmap.incremental_mapping</code>, which consumes much time. How can you control your submission within 9 hours?</p>",
      "rawMarkdown": "Hi @igorlashkov , congratulations for the prize and gold. I have a question\n\nHow to run reconstruction multiple times? I think the reconstruction is in `pycolmap.incremental_mapping`, which consumes much time. How can you control your submission within 9 hours?",
      "votes": null
    },
    {
      "id": "2301071",
      "postDate": "06/13/2023 15:45:47",
      "content": "<p>Thanks for the detailed write up! pretty cool that sp and sg worked with small image sizes as well. In our case we used 1600 for small images and 3200 for large ones <br>\nwe have also tried iterative reconstruction but with changing the camera type it gave +0.001 boost so we ignored it especially after getting a reproducable pycolmap results, it seems that it worth further exploring<br>\nCongrats on the second place you really did a great job !!</p>",
      "rawMarkdown": "Thanks for the detailed write up! pretty cool that sp and sg worked with small image sizes as well. In our case we used 1600 for small images and 3200 for large ones \nwe have also tried iterative reconstruction but with changing the camera type it gave +0.001 boost so we ignored it especially after getting a reproducable pycolmap results, it seems that it worth further exploring\nCongrats on the second place you really did a great job !!",
      "votes": null
    },
    {
      "id": "2301144",
      "postDate": "06/13/2023 16:40:19",
      "content": "<p>Use of cache allowed us significantly increase speed of our solution. We mostly cached 3 heavy things – images resized and packed in CUDA tensor beforehand, SP keypoints, and SG matches saved in memory on a 1st run. Also, if we see that the N of matches for a certain image pair is low, we do not proceed with other image sizes, exit earlier. Special thanks to <a href=\"https://www.kaggle.com/vostankovich\" target=\"_blank\">@vostankovich</a> for making our solution blazing fast!</p>",
      "rawMarkdown": "Use of cache allowed us significantly increase speed of our solution. We mostly cached 3 heavy things – images resized and packed in CUDA tensor beforehand, SP keypoints, and SG matches saved in memory on a 1st run. Also, if we see that the N of matches for a certain image pair is low, we do not proceed with other image sizes, exit earlier. Special thanks to @vostankovich for making our solution blazing fast!",
      "votes": null
    },
    {
      "id": "2301152",
      "postDate": "06/13/2023 16:49:41",
      "content": "<p>Congratulations on the amazing result. A lot of new &amp; great stuff in this write up for me to read up on!</p>",
      "rawMarkdown": "Congratulations on the amazing result. A lot of new & great stuff in this write up for me to read up on!",
      "votes": null
    },
    {
      "id": "2301391",
      "postDate": "06/13/2023 21:24:25",
      "content": "<p>Congratulations for the excellent result and gold. It provided lots of idea to me, very helpful !!!</p>",
      "rawMarkdown": "Congratulations for the excellent result and gold. It provided lots of idea to me, very helpful !!!",
      "votes": null
    },
    {
      "id": "2301678",
      "postDate": "06/14/2023 05:05:56",
      "content": "<p>I'm glad you found the information detailed and interesting! It's impressive to see how techniques like super-resolution (SP) and style transfer (SG) can work effectively even with small image sizes. In your case, using 1600 for small images and 3200 for large ones seems like a reasonable approach. It's always fascinating to explore different strategies and see how they can contribute to the overall performance in image-related challenges.</p>",
      "rawMarkdown": "I'm glad you found the information detailed and interesting! It's impressive to see how techniques like super-resolution (SP) and style transfer (SG) can work effectively even with small image sizes. In your case, using 1600 for small images and 3200 for large ones seems like a reasonable approach. It's always fascinating to explore different strategies and see how they can contribute to the overall performance in image-related challenges.",
      "votes": null
    },
    {
      "id": "2319801",
      "postDate": "06/27/2023 10:17:10",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/igorlashkov\" target=\"_blank\">@igorlashkov</a> . Can you please share the notebook with me? Thanks</p>",
      "rawMarkdown": "Hi @igorlashkov . Can you please share the notebook with me? Thanks",
      "votes": null
    },
    {
      "id": "2613301",
      "postDate": "01/22/2024 02:04:09",
      "content": "<p><a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a>,<br>\nThank you very much to you and your team for this amazing contribution. I am trying to reimplement your solution for learning purposes. Wanted to know if the submission notebook/code is already public.</p>",
      "rawMarkdown": "remekkinas,\nThank you very much to you and your team for this amazing contribution. I am trying to reimplement your solution for learning purposes. Wanted to know if the submission notebook/code is already public.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2300584,
      "author_name": "oldufo",
      "author_url": "",
      "post_date": "06/13/2023 09:20:16",
      "content": "<p>Congrats with great result!</p>\n<blockquote>\n  <p>Loftr, QuadreeAttention, AspanFormer, DKM v3, GlueStick (keypoints + lines), Patch2Pix, KeyNetAffNetHardNet, DISK, PatchNetVLAD</p>\n</blockquote>\n<p>Could you please share the numbers?</p>\n<p>The same for the global descriptor thing, please?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2300614,
          "author_name": "vostankovich",
          "author_url": "",
          "post_date": "06/13/2023 09:49:13",
          "content": "<p>The matchers were tested during start/mid of competition when we had very simple pipeline. We just tested one by one and discarded any matcher that was worse than SP/SG. All of the listed models are worse by a large margin, at least 15-20% lower than SP/SG. Some of the models also were just too slow or memory consuming along with the bad score. The closest one was LoFTR as I remember, but even ensembling it with SP/SG didn't give us any boost. We can rerun some matchers in our final pipeline for sure if needed to get the numbers.</p>\n<p>The same thing with global descriptors, we don't have numbers for them, we just discovered one time that SG can be a retrieval itself and the approach is robust enough and outperforms other models. SP/SG are really fast in half mode and with caching, so we actually completely agreed not to use any additional retrieval models for pairs search.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2300600,
      "author_name": "remekkinas",
      "author_url": "",
      "post_date": "06/13/2023 09:34:45",
      "content": "<p>Hey team, I just want to say a big thank you. Your work during this contest was amazing. We came up with so many ideas and things kept changing all the time.</p>\n<p>Special thanks to <a href=\"https://www.kaggle.com/vostankovich\" target=\"_blank\">@vostankovich</a>. You never gave up. I did - I quit when trying to set up pixfm, and that's not like me. I've learned from this. It's important to listen to other people's ideas and not to think that I always know best.</p>\n<p>From the start, <a href=\"https://www.kaggle.com/igorlashkov\" target=\"_blank\">@igorlashkov</a> said we should work on the rotation. I tried a few times but it didn't help me. But guess what? In the end, working on the rotation was a key part of our solution. So, big thanks to Igla for not giving up!</p>\n<p>And thank you, <a href=\"https://www.kaggle.com/oldufo\" target=\"_blank\">@oldufo</a> for talking about 3D reconstruction. I first started looking into this last year, just out of interest. The last contest made me more interested. It's a great way to learn more and get better at what we do. </p>\n<p>I'm really looking forward to next year and IMC 2024. Let's do it!</p>\n<p>One more from my side - the way I was trying to understand SfM pipeline was Meshroom. SfM was new for me so looking only in example notebooks was not enough for me. During the competition I was experimenting a lot playing with different settings to understand \"what if\".</p>\n<p><img src=\"https://i.ibb.co/dgTymVc/Meshroom.jpg\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 2300617,
          "author_name": "jaafarmahmoud1",
          "author_url": "",
          "post_date": "06/13/2023 09:52:24",
          "content": "<p>Congrats on 2nd place! good work!<br>\nPixSFM was really tricky! It gave boost to our validation score up to +0.1 on mAA, but unfortunately it didn't improve our LB.<br>\nYou can try it using this notebook!<br>\n<a href=\"https://www.kaggle.com/jaafarmahmoud1/pixel-perfect-sfm-on-kaggle\" target=\"_blank\">https://www.kaggle.com/jaafarmahmoud1/pixel-perfect-sfm-on-kaggle</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2300811,
      "author_name": "forcewithme",
      "author_url": "",
      "post_date": "06/13/2023 13:15:35",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/igorlashkov\" target=\"_blank\">@igorlashkov</a> , congratulations for the prize and gold. I have a question</p>\n<p>How to run reconstruction multiple times? I think the reconstruction is in <code>pycolmap.incremental_mapping</code>, which consumes much time. How can you control your submission within 9 hours?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2301144,
          "author_name": "igorlashkov",
          "author_url": "",
          "post_date": "06/13/2023 16:40:19",
          "content": "<p>Use of cache allowed us significantly increase speed of our solution. We mostly cached 3 heavy things – images resized and packed in CUDA tensor beforehand, SP keypoints, and SG matches saved in memory on a 1st run. Also, if we see that the N of matches for a certain image pair is low, we do not proceed with other image sizes, exit earlier. Special thanks to <a href=\"https://www.kaggle.com/vostankovich\" target=\"_blank\">@vostankovich</a> for making our solution blazing fast!</p>",
          "votes": null,
          "replies": [
            {
              "id": 2319801,
              "author_name": "narshima199712",
              "author_url": "",
              "post_date": "06/27/2023 10:17:10",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/igorlashkov\" target=\"_blank\">@igorlashkov</a> . Can you please share the notebook with me? Thanks</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2301071,
      "author_name": "ammarali32",
      "author_url": "",
      "post_date": "06/13/2023 15:45:47",
      "content": "<p>Thanks for the detailed write up! pretty cool that sp and sg worked with small image sizes as well. In our case we used 1600 for small images and 3200 for large ones <br>\nwe have also tried iterative reconstruction but with changing the camera type it gave +0.001 boost so we ignored it especially after getting a reproducable pycolmap results, it seems that it worth further exploring<br>\nCongrats on the second place you really did a great job !!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2301152,
      "author_name": "anjum48",
      "author_url": "",
      "post_date": "06/13/2023 16:49:41",
      "content": "<p>Congratulations on the amazing result. A lot of new &amp; great stuff in this write up for me to read up on!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2301391,
      "author_name": "kaisenzhang",
      "author_url": "",
      "post_date": "06/13/2023 21:24:25",
      "content": "<p>Congratulations for the excellent result and gold. It provided lots of idea to me, very helpful !!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2301678,
      "author_name": "poojach7611",
      "author_url": "",
      "post_date": "06/14/2023 05:05:56",
      "content": "<p>I'm glad you found the information detailed and interesting! It's impressive to see how techniques like super-resolution (SP) and style transfer (SG) can work effectively even with small image sizes. In your case, using 1600 for small images and 3200 for large ones seems like a reasonable approach. It's always fascinating to explore different strategies and see how they can contribute to the overall performance in image-related challenges.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2613301,
      "author_name": "mukit0",
      "author_url": "",
      "post_date": "01/22/2024 02:04:09",
      "content": "<p><a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a>,<br>\nThank you very much to you and your team for this amazing contribution. I am trying to reimplement your solution for learning purposes. Wanted to know if the submission notebook/code is already public.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2300576": "# **Bonus Point**\nWe presented our solution at [CVPR 2023](https://image-matching-workshop.github.io/). Watch on YouTube [here](https://youtu.be/9JpGjpITiDM?si=l7pGDPw4vOnZsH9H&t=13519)\n\n# **Intro**\n\nOur team would like to deeply appreciate the Kaggle staff, Google Research, and Haiper for hosting the continuation of this exciting image matching challenge, as well as everyone here who compete and shared great discussions. My congratulations to all participants!\n\nThe work we describe here is truly a joint effort of @yamsam, @remekkinas, and @vostankovich. I’m grateful of being a part of this hardworking, cohesive, and skilled team. Thanks a lot, guys! I learned a lot from you.\n\n###### We enjoyed this competition!\n\nThe best submission that was used for final scoring in private LB finished on the last day of the competition. We weren’t completely sure about this submission because it was not clear how much randomness was in it. We had a practice to re-run the same notebook code multiple times to see what scores we can get. We discussed the solution which was implemented on the last day, trusted it and it worked out. The other interesting fact is that our 2nd selected submission for final evaluation scored **0.497/0.542** that also allows us to take 2nd place. This selected second submission is the same as the 1st one except the “Run reconstruction multiple times” trick, that is described below. Anyway, the difference between the best submission (**0.562**) and the 2nd one is noticable. \n\n# **Overview**\n\nIn general, throughout our code submissions every time we fight with the randomness coming from COLMAP responsible for scene reconstruction. Our final solution is based on the use of COLMAP and pretrained SuperPoint/SuperGlue models running on different resolutions for every image in the scene. We apply a bunch of different tricks aimed at different parts of COLMAP-based pipeline in order to stabilize our solution and reach the final score.\n\n# **Architecture**\n![architecture](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5065877%2F68911db7c4cc430dec05670cd196a960%2Fslide_architecture.png?generation=1687202186098466&alt=media)\n\n\n# **Key Takeaways:**\n\n- Initially, use all **possible unique image pairs** generated from the scene set. Remove the model and logic used for finding and ranking similar images in the scene. A threshold of 100 defines a minimum number of matches that we expect each image pair to have. If it is less, we discard that pair.\n- **SP/SG** settings: unlimited N of keypoints, the keypoint threshold is **0.005**, match threshold **0.2**, and the number of sinkhorn iteratiors is **20**.\n- **Half precision** for SP/SG helped to reduce occupied memory without sacrificing noticable accuracy. Another great performance trick is to cache keypoints/descriptors, generated by SP for each image, and, then cache SG matches for every image pair. It allowed to reduce the running time a lot.\n- **TTA**. Ensemble of matches extracted from images at different scales. In our local experiments, the best results are achieved with a combination of **[1088, 1280, 1376]**. We used np.concatenate to join matches from different models that was pretty common for the last IMC22 competition.\n- Apply **rotate detector** to un-rotate images in the scene if necessary. We discovered that some scenes in the train dataset (cyprus, dioscuri) have many 90/270 rotated images. Some of the images have EXIF meta information. Unfortunately, looking into the train dataset, we did not find any specifics about the orientation the image was captured at. To address this rotation issue, we re-rotate the image to its natural orientation by employing this [solution](https://github.com/ternaus/check_orientation ). We use it w/o any threshold and look at the number of rotations that we need to apply to an image. After applying rotation, the score for cyprus scene jumped up significantly from **~0.02** up to **~0.55**. RotNet implementation did not work out for us.\n- **Set-up initial image** for COLMAP reconstruction explicitly. For each image we store the number of pairs in which it seen and the number of matches these pairs produce. Then we pick the one with the highest number of pairs. If multiple images satisfy this criterion, we pick the one with the highest number of matches. It helped to boost the score.\n- To reduce randomness in the score, we decided to do something like **averaging multiple match_exhaustive** calls. The idea is to run match_exhaustive for N times on the original database of matches. Then we take only those matches that that appear in 8/10 cases, other matches are neglected. It was done in a rude way with database copies write/read, etc.\n- **Run reconstruction multiple times** from scratch with different matchers threshold e.g. **[100, 125, 75, 100]**. By looking at N of registered images and number of 3D cloud points, we select the best reconstruction. This trick not only allows to find the better reconstruction by finding the better threshold for matches, but also decrease the randomness effect and acts as a countermeasure against a shake up. Due to its running time complexity, we used this strategy only for scenes having less than 40-45 images. This is the last step in our solution that helped us to boost score from **0.497/0.542** to **0.506/0.562**. We also experimented with pycolmap.incremental_mapping employing similar idea but that  scenario did not work out.\n\n\n# **Ideas that did not work out or not fully tested:**\n•\t**TTA multi crop**, no success. The idea was to split image into multiple crops and extract matches in order to find similar images in the scene and determine the best image pairs.\n•\t**Square-sized images.**\n•\t**Bigger image size** (e.g., 1600) for SP/SG.\n•\t**Manual RANSAC** instead of using COLMAP internal implementation. We run experiments by disabling geometric verification, but the score was not good.\n•\t**NMS filtering** to reduce number of points by using [ANMS](https://github.com/BAILOOL/ANMS-Codes). \n•\tFilter least significant image pairs by number of outliers instead of relying on raw matches. It was quite important to look for certain number of matches for an image pair. Run experiments using SP/SG and Loftr. We got higher mAA with Loftr, but, probably, more effort needed here to make it work properly, not enough time.\n•\t**Downscale scene images** before passing them to COLMAP.\n•\t**Pixel-Perfect Structure-from-Motion**. It was a promising method to evaluate as we got a good boost locally with [PixSfm](https://github.com/cvg/pixel-perfect-sfm), using a single image size of 1280, and it boosted our score from **0.71727** to **0.76253**. Then, we managed to install this framework successfully and run it in Kaggle environment, but could not beat our best score at that moment. It is a heavyweight framework taking too much RAM, and we could run it only for scenes having at most ~30 images. A bit upset because we spent tons of hours compiling all this stuff.\n•\t**Adaptive image sizes**. Say, if the longest image side >= 1536 for most images in the scene, we use higher image resolution for matchers ensemble e.g. [1280, 1408, 1536]. Otherwise, a default one is applied [1280,1088,1376]. Did not have enough time to test this idea. It worked locally for cyprus and wall that have big resolution. One of our last submissions implementing this idea crashed with internal error.\n•\tDifferent **detectors, matchers**. We tested Loftr, QuadreeAttention, AspanFormer, DKM v3, GlueStick (keypoints + lines), Patch2Pix, KeyNetAffNetHardNet, DISK, PatchNetVLAD. We also experimented with the confidence matching thresholds and number of matches, but no boost here. Eventually, SP/SG was the best choice for us. Probably, the reason why many dense-based methods did not work out for us is because of the low performance of the “repeatability” metric and high noise in the matches.\n•\tDifferent **CNNs to find the most similar images** in the scene and generate corresponding image pairs (NetVLAD, different pretrained timm-based backbones, CosPlace etc). We even specifically trained a model to find similar images in the scene, but no success here. Later, we gave up using this strategy at all.\n•\tDifferent **keypoint/matching refinement** methods (e.g., recently published [AdaMatcher](https://github.com/TencentYoutuResearch/AdaMatcher), Patch2Pix ), but did not have enough time. AdaMatcher seems a promising idea to try, a quote from their paper “as a refinement network for SP/SG we observe a noticeable improvement in AUC”\n\n\n# **Performance Improvements Step By Step**\n| Method\t| Private LB\t| Public LB\t| Δ Private LB\n| --- | --- |--- |--- |\n|Baseline\t|0.382\t|0.317\t|–\n|Manual RANSAC\t|0.336\t|0.286\t|–0.046\n|Image size 1024 → 1280\t|0.407\t|0.345\t|+0.025\n|Image size 1376, similarity=None\t|0.489\t|0.426\t|+0.026\n|Exhaustive matching, 8/10\t|0.486\t|0.441\t|-0.003\n|TTA, Image sizes [840, 1024, 1280]\t|0.491\t|0.447\t|+0.002\n|TTA, Image sizes [1088, 1280, 1376]\t|0.523\t|0.475\t|+0.032\n|Manual image initialization\t|0.529\t|0.492\t|+0.006\n|Rotate detection\t|0.542\t|0.497\t|+0.013\n|Multi-run reconstruction, matches thr [100, 125, 75, 100]\t|0.562\t|0.506\t|+0.02\n\n##### The final score is 0.506/0.562 in Public/Private.\n\n# **Local Validation**\nAs a reference, this is one of our latest metric reports using train dataset: \n\nurban / kyiv-puppet-theater (26 images, 325 pairs) -> mAA=0.921538, mAA_q=0.991077, mAA_t=0.921846\nurban -> mAA=0.921538\n\nheritage / cyprus (30 images, 435 pairs) -> mAA=0.514713, mAA_q=0.525287, mAA_t=0.543678\nheritage / wall (43 images, 903 pairs) -> mAA=0.495792, mAA_q=0.875637, mAA_t=0.509302\nheritage -> mAA=0.505252\n\nhaiper / bike (15 images, 105 pairs) -> mAA=0.940952, mAA_q=0.999048, mAA_t=0.940952\nhaiper / chairs (16 images, 120 pairs) -> mAA=0.834167, mAA_q=0.863333, mAA_t=0.839167\nhaiper / fountain (23 images, 253 pairs) -> mAA=0.999605, mAA_q=1.000000, mAA_t=0.999605\nhaiper -> mAA=0.924908\n\n**Final metric -> mAA=0.783900**\n\nFinally, we had two submissions running on the last day. One of them succeeded and allowed us to get 2nd place, but the other one did not fit the time limit, unexpectedly. Sometimes it is stressful to make a submission on the last day of the competition.\n\n# **Helpful Resources**\nSpecial thanks to the authors of the following projects:\n•\t[COLMAP](https://colmap.github.io/)\n•\t[SuperPoint](https://ieeexplore.ieee.org/document/7780814)\n•\t[SuperGlue](https://arxiv.org/abs/1911.11763)\n•\t[check_orientation](https://github.com/ternaus/check_orientation )",
    "2300584": "Congrats with great result!\n\n>Loftr, QuadreeAttention, AspanFormer, DKM v3, GlueStick (keypoints + lines), Patch2Pix, KeyNetAffNetHardNet, DISK, PatchNetVLAD\n\nCould you please share the numbers?\n\nThe same for the global descriptor thing, please?",
    "2300600": "Hey team, I just want to say a big thank you. Your work during this contest was amazing. We came up with so many ideas and things kept changing all the time.\n\nSpecial thanks to @vostankovich. You never gave up. I did - I quit when trying to set up pixfm, and that's not like me. I've learned from this. It's important to listen to other people's ideas and not to think that I always know best.\n\nFrom the start, @igorlashkov said we should work on the rotation. I tried a few times but it didn't help me. But guess what? In the end, working on the rotation was a key part of our solution. So, big thanks to Igla for not giving up!\n\nAnd thank you, @oldufo for talking about 3D reconstruction. I first started looking into this last year, just out of interest. The last contest made me more interested. It's a great way to learn more and get better at what we do. \n\nI'm really looking forward to next year and IMC 2024. Let's do it!\n\n\nOne more from my side - the way I was trying to understand SfM pipeline was Meshroom. SfM was new for me so looking only in example notebooks was not enough for me. During the competition I was experimenting a lot playing with different settings to understand \"what if\".\n\n![](https://i.ibb.co/dgTymVc/Meshroom.jpg)",
    "2300614": "The matchers were tested during start/mid of competition when we had very simple pipeline. We just tested one by one and discarded any matcher that was worse than SP/SG. All of the listed models are worse by a large margin, at least 15-20% lower than SP/SG. Some of the models also were just too slow or memory consuming along with the bad score. The closest one was LoFTR as I remember, but even ensembling it with SP/SG didn't give us any boost. We can rerun some matchers in our final pipeline for sure if needed to get the numbers.\n\nThe same thing with global descriptors, we don't have numbers for them, we just discovered one time that SG can be a retrieval itself and the approach is robust enough and outperforms other models. SP/SG are really fast in half mode and with caching, so we actually completely agreed not to use any additional retrieval models for pairs search.",
    "2300617": "Congrats on 2nd place! good work!\nPixSFM was really tricky! It gave boost to our validation score up to +0.1 on mAA, but unfortunately it didn't improve our LB.\nYou can try it using this notebook!\nhttps://www.kaggle.com/jaafarmahmoud1/pixel-perfect-sfm-on-kaggle",
    "2300811": "Hi @igorlashkov , congratulations for the prize and gold. I have a question\n\nHow to run reconstruction multiple times? I think the reconstruction is in `pycolmap.incremental_mapping`, which consumes much time. How can you control your submission within 9 hours?",
    "2301071": "Thanks for the detailed write up! pretty cool that sp and sg worked with small image sizes as well. In our case we used 1600 for small images and 3200 for large ones \nwe have also tried iterative reconstruction but with changing the camera type it gave +0.001 boost so we ignored it especially after getting a reproducable pycolmap results, it seems that it worth further exploring\nCongrats on the second place you really did a great job !!",
    "2301144": "Use of cache allowed us significantly increase speed of our solution. We mostly cached 3 heavy things – images resized and packed in CUDA tensor beforehand, SP keypoints, and SG matches saved in memory on a 1st run. Also, if we see that the N of matches for a certain image pair is low, we do not proceed with other image sizes, exit earlier. Special thanks to @vostankovich for making our solution blazing fast!",
    "2301152": "Congratulations on the amazing result. A lot of new & great stuff in this write up for me to read up on!",
    "2301391": "Congratulations for the excellent result and gold. It provided lots of idea to me, very helpful !!!",
    "2301678": "I'm glad you found the information detailed and interesting! It's impressive to see how techniques like super-resolution (SP) and style transfer (SG) can work effectively even with small image sizes. In your case, using 1600 for small images and 3200 for large ones seems like a reasonable approach. It's always fascinating to explore different strategies and see how they can contribute to the overall performance in image-related challenges.",
    "2319801": "Hi @igorlashkov . Can you please share the notebook with me? Thanks",
    "2613301": "remekkinas,\nThank you very much to you and your team for this amazing contribution. I am trying to reimplement your solution for learning purposes. Wanted to know if the submission notebook/code is already public."
  },
  "source": "meta"
}