{
  "id": 582898,
  "title": "10th Place Solution",
  "url": "/competitions/image-matching-challenge-2025/writeups/team-sony-matching-10th-place-solution",
  "author_name": "",
  "post_date": "2025-06-10T22:50:38.403Z",
  "votes": 32,
  "comment_count": 1,
  "views": 0,
  "content": "<p><strong>Update:</strong> <a href=\"https://www.kaggle.com/tmyok1984\" target=\"_blank\">@tmyok1984</a> has shared the details of experiments for VGGT in <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2025/discussion/582968\" target=\"_blank\">this thread</a>, so please also check it!<br>\n(Note: VGGT is not integrated in below solution)</p>\n<hr>\n<p>First of all, I’d like to express my gratitude to the host team and all the staff members for organizing such an exciting competition and providing outstanding support throughout the challenge.</p>\n<p>This was my first time participating in the IMC, but thanks to the many insightful past IMC solutions, I managed to complete this challenge.<br>\nAlso, competing with strong participants really kept me motivated. I’m truly grateful to all of you.</p>\n<p>Our approach was strongly supported by <a href=\"https://www.kaggle.com/tmyok1984\" target=\"_blank\">@tmyok1984</a>'s leadership and the wide range of experiments &amp; discussions.<br>\nAnd also we get a great inspiration from the IMC2024 solutions by <a href=\"https://www.kaggle.com/jooott\" target=\"_blank\">@jooott</a> and <a href=\"https://www.kaggle.com/sugupoko\" target=\"_blank\">@sugupoko</a>, which influenced core ideas in our pipeline.<br>\nI really appreciate their support and teamwork during this challenging competition.</p>\n<p>In developing our solution, <strong>I especially focused on selecting image pairs with high accucary</strong> before feature matching.<br>\nThere’s almost nothing particularly special about the pipeline after the image pair selection — it’s a straightforward process using ALIKED-LightGlue for feature matching and pycolmap for camera pose estimation.<br>\nWe made a shake-up in the private leaderboard, but I think this was because our pair selection strategy fit well fortunately.</p>\n<p><br></p>\n<h2>1. Key steps</h2>\n<h3>1.1. Determine Top-k for Selecting Image Pairs (Dynamic Top-k)</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2F44480ec2a6588e2da7339330bd9ac67a%2F1.png?generation=1748952029356808&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>First, we <strong>dynamically adjust the number of top image pairs (top-k)</strong> for each dataset.</li>\n<li>In visually confusing scenes like fbk_vineyard, I noticed that incorrect image pairs are often formed, which can lead to mixing data from different clusters in the model. <br>\nTherefore, for scenes with high visual similarity, to improve the accuracy of selected image pairs, we <strong>set a very small top-k value (e.g. k = 3)</strong> to strictly select only neighboring image pairs—similar to a MST(minimum spanning tree). The similarity score was calculated based on DINOv2 patch features.</li>\n<li>In contrast, for scenes like amy_gardens, where images were taken at different times or with different cameras, there may be multiple similar images of the same place. In those cases, we used a larger top-k to keep Recall.</li>\n</ul>\n<p><strong>Note:</strong></p>\n<ul>\n<li>Regarding the above strategy, I think using a very small top-k has some risk: it can break the cluster into parts wrongly. <br>\nFor example, if there are many images from the strictly same location, they may connect only with each other, and not with nearby views.<br>\nTherefore, ideally, I think it is better to set a threshold for similarity score instead of top-k. </li>\n<li>However, setting good thresholds is hard and depends on the dataset, so in this pipeline, I accepted this risk and used this top-k approach.</li>\n<li>I also assumed that in scenes like fbk_vineyard, it’s unlikely that many images were taken from exactly the same viewpoint. (It is useless act)\nTherefore, I estimated the risk of cluster breakage to be relatively low.<ul>\n<li>But it is easy to intentionally include duplicate or nearly identical images….</li>\n<li>In such cases, additional measures may be needed—for example, if there are image pairs with too much overlap, increase the top-k value accordingly.</li></ul></li>\n</ul>\n<p><br></p>\n<h3>1.2. Get Top-k Image Pairs Ranked by Keypoint Feature-based Similarity</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2F208da39236d8b8ce2befbedec791b8b8%2F2.png?generation=1748952611810102&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Second, we select image pairs based on the top-k value computed in Step 1.\nInstead of using visual features like DINOv2 or MegaLoc, we rely on a <strong>custom similarity score based on KeyNet-AdaLAM</strong>.<ul>\n<li>I found this KeyNet-AdaLAM in <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2024/discussion/510295\" target=\"_blank\">out teammates jooott and sugupoko's pipeline in IMC2024</a>, and moreover it was originated from <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417045\" target=\"_blank\">excellent solution from the 5th-place team in IMC2023</a>.</li></ul></li>\n<li>Global descriptors could not distinguish between visually similar scenes, however, by carefully looking at small details, I found that adjacent images could be identified—like puzzle pieces that fit together.<br>\nThus, I think keypoint-based matching was considered a more rational choice for extracting accurate image pairs.</li>\n<li>KeyNet-AdaLAM was selected due to below advantages:<ul>\n<li>Robustness to scale and rotation (via AffNet &amp; OriNet)</li>\n<li>Fast exhaustive matching for all image pairs</li>\n<li>Higher precision in pair selection than other sparse matching method (e.g. ALIKED + LightGlue)</li></ul></li>\n<li>The similarity score was heuristically designed, based on the average descriptor distance of inliers and the number of inliers.<br>\nUsing this score, <strong>we could accurately extract true neighboring pairs in fbk_vineyard</strong>, as confirmed by the top-4 examples shown below.<br>\n(Since the dataset uses sequential image file names for adjacent views, the correctness of the selected pairs is evident)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2F1aadc864ec8fb39efcede0a3f217d831%2F2_1.png?generation=1748952798149857&amp;alt=media\" alt=\"\"></li>\n<li>By generating the image pairs in such a strict way, we were able to estimate the camera poses for fbk_vineyard fairly accurately.<br>\n(As for split3, there are opposite-view pairs, and it was difficult to recognize them as part of the same cluster with this pipeline…)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2F40eea1255457216c15b1f66bcfe02321%2F2_2.png?generation=1748952818025411&amp;alt=media\" alt=\"\"></li>\n</ul>\n<p><strong>Note:</strong></p>\n<ul>\n<li>To improve the accuracy of image pair matching, increasing the <code>ransac_iters</code> parameter in AdaLAM is also important (we used 384).</li>\n<li>However, higher values may cause GPU memory errors, so be careful.</li>\n</ul>\n<p><br></p>\n<h3>1.3. Extract Features by ALIKED &amp; Match by LightGlue</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2Ff545553e2ff92aaa2a4db250d2a502ac%2F3.png?generation=1748953000900906&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Next, we perform feature matching on the extracted image pairs using ALIKED + LightGlue (<code>resize_to</code>: 1024, <code>max_num_keypoints</code>: 10000).</li>\n<li>There’s not much customization in this section, but I adopt a cropping method based on DBSCAN, inspired by the <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2024/discussion/510084\" target=\"_blank\">excellent solution from the 1st-place team in IMC2024</a>.<ul>\n<li>I initially assumed that this approach might have limited impact specifically on walkthrough-type scenes like fbk_vineyard, but it still showed some positive effects in both CV and LB, so I decided to include it in our pipeline. (Especially, I remember it working well on the amy_gardens scene.)</li></ul></li>\n<li>Finally, we combine the matching results from the full images and the cropped regions, and pass them to the final step. </li>\n</ul>\n<p><br></p>\n<h3>1.4. Geometric verification &amp; Camera pose estimation (by pycolmap)</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2Fb812de83d435a69de0bc9bc20f867f20%2F4.png?generation=1748953891664264&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Finally, we remove outliers and low-confidence pairs based on MAGSAC++, save the cleaned matches to the database, and run pycolmap for camera-pose estimation.</li>\n<li>Because we already have well-curated image pairs, we don’t use match_exhaustive().</li>\n<li>Parameter tuning in pycolmap gave only small gains, but we saw some benefits by keeping <code>ba_local_max_num_iterations</code> higher.</li>\n</ul>\n<p><br><br>\nThe final metric scores for each training dataset in IMC2025 are shown below.<br>\n\"stairs\" was especially difficult to handle with the current pipeline…</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>ETs</th>\n<th>amy_gardens</th>\n<th>fbk_vineyard</th>\n<th>stairs</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>score (%)</td>\n<td>74.98</td>\n<td>40.00</td>\n<td>62.65</td>\n<td>5.26</td>\n</tr>\n<tr>\n<td>mAA (%)</td>\n<td>60.61</td>\n<td>25.00</td>\n<td>45.62</td>\n<td>2.78</td>\n</tr>\n<tr>\n<td>clusterness (%)</td>\n<td>100.00</td>\n<td>100.00</td>\n<td>100.00</td>\n<td>50.00</td>\n</tr>\n</tbody>\n</table>\n<p><br><br></p>\n<h2>2. Other Ideas</h2>\n<h3>2.1. Multi-processing</h3>\n<ul>\n<li>I referred to <a href=\"https://www.kaggle.com/tmyok1984\" target=\"_blank\">@tmyok1984</a>’s solution from last year, and applied parallel processing using two GPUs (T4 × 2).<br>\nIdeally, CPU and GPU tasks should also be processed in parallel, but in my current solution, the total runtime was acceptable without further optimization.</li>\n</ul>\n<h3>2.2 Methods That Did Not Work</h3>\n<ul>\n<li>Extractors other than ALIKED (While ALIKED+DISK+SIFT showed good results in CV, they did not work at all on the LB, so I excluded them)</li>\n<li>Rotation Handling (Although it is generally important, it did not work well in my pipeline. I guess rotation was not a critical factor in the newly added datasets this year)</li>\n<li>Refiner after Camera Pose Estimation (It didn’t match well with my pipeline. However, I believe it could still be beneficial depending on the configuration)</li>\n</ul>\n<h3>2.3 Exploration of VGGT (Not imcorporate in my pipeline)</h3>\n<ul>\n<li>The exploration of VGGT’s potential was mainly done by <a href=\"https://www.kaggle.com/tmyok1984\" target=\"_blank\">@tmyok1984</a> and <a href=\"https://www.kaggle.com/jooott\" target=\"_blank\">@jooott</a>, and we found that <strong><a href=\"https://github.com/facebookresearch/vggt\" target=\"_blank\">VGGT: Visual Geometry Grounded Transformer</a> can estimate camera poses for “stairs” scene relatively well</strong> (still not perfect).</li>\n<li>Also, additional tests by <a href=\"https://www.kaggle.com/tmyok1984\" target=\"_blank\">@tmyok1984</a> showed that <strong>VGGT also works relatively correctly for datasets with regions with sparse image coverage such as amy_gardens, and for opposite-view pairs in fbk_vineyard split3</strong>.</li>\n<li>On the other hand, absolute accuracy for VGGT's camera pose seems to be not very high, so it was difficult to make it work well in LB.</li>\n<li>Details of these experiments has shared by tmyok in <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2025/discussion/582968\" target=\"_blank\">this thread</a>, so please also check it.</li>\n</ul>\n<p><br><br>\nThanks again to all, and I'd love to join again if the challenge is held next year!</p>",
  "messages": [
    {
      "id": "3216355",
      "postDate": "06/03/2025 13:02:37",
      "content": "<p><strong>Update:</strong> <a href=\"https://www.kaggle.com/tmyok1984\" target=\"_blank\">@tmyok1984</a> has shared the details of experiments for VGGT in <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2025/discussion/582968\" target=\"_blank\">this thread</a>, so please also check it!<br>\n(Note: VGGT is not integrated in below solution)</p>\n<hr>\n<p>First of all, I’d like to express my gratitude to the host team and all the staff members for organizing such an exciting competition and providing outstanding support throughout the challenge.</p>\n<p>This was my first time participating in the IMC, but thanks to the many insightful past IMC solutions, I managed to complete this challenge.<br>\nAlso, competing with strong participants really kept me motivated. I’m truly grateful to all of you.</p>\n<p>Our approach was strongly supported by <a href=\"https://www.kaggle.com/tmyok1984\" target=\"_blank\">@tmyok1984</a>'s leadership and the wide range of experiments &amp; discussions.<br>\nAnd also we get a great inspiration from the IMC2024 solutions by <a href=\"https://www.kaggle.com/jooott\" target=\"_blank\">@jooott</a> and <a href=\"https://www.kaggle.com/sugupoko\" target=\"_blank\">@sugupoko</a>, which influenced core ideas in our pipeline.<br>\nI really appreciate their support and teamwork during this challenging competition.</p>\n<p>In developing our solution, <strong>I especially focused on selecting image pairs with high accucary</strong> before feature matching.<br>\nThere’s almost nothing particularly special about the pipeline after the image pair selection — it’s a straightforward process using ALIKED-LightGlue for feature matching and pycolmap for camera pose estimation.<br>\nWe made a shake-up in the private leaderboard, but I think this was because our pair selection strategy fit well fortunately.</p>\n<p><br></p>\n<h2>1. Key steps</h2>\n<h3>1.1. Determine Top-k for Selecting Image Pairs (Dynamic Top-k)</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2F44480ec2a6588e2da7339330bd9ac67a%2F1.png?generation=1748952029356808&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>First, we <strong>dynamically adjust the number of top image pairs (top-k)</strong> for each dataset.</li>\n<li>In visually confusing scenes like fbk_vineyard, I noticed that incorrect image pairs are often formed, which can lead to mixing data from different clusters in the model. <br>\nTherefore, for scenes with high visual similarity, to improve the accuracy of selected image pairs, we <strong>set a very small top-k value (e.g. k = 3)</strong> to strictly select only neighboring image pairs—similar to a MST(minimum spanning tree). The similarity score was calculated based on DINOv2 patch features.</li>\n<li>In contrast, for scenes like amy_gardens, where images were taken at different times or with different cameras, there may be multiple similar images of the same place. In those cases, we used a larger top-k to keep Recall.</li>\n</ul>\n<p><strong>Note:</strong></p>\n<ul>\n<li>Regarding the above strategy, I think using a very small top-k has some risk: it can break the cluster into parts wrongly. <br>\nFor example, if there are many images from the strictly same location, they may connect only with each other, and not with nearby views.<br>\nTherefore, ideally, I think it is better to set a threshold for similarity score instead of top-k. </li>\n<li>However, setting good thresholds is hard and depends on the dataset, so in this pipeline, I accepted this risk and used this top-k approach.</li>\n<li>I also assumed that in scenes like fbk_vineyard, it’s unlikely that many images were taken from exactly the same viewpoint. (It is useless act)\nTherefore, I estimated the risk of cluster breakage to be relatively low.<ul>\n<li>But it is easy to intentionally include duplicate or nearly identical images….</li>\n<li>In such cases, additional measures may be needed—for example, if there are image pairs with too much overlap, increase the top-k value accordingly.</li></ul></li>\n</ul>\n<p><br></p>\n<h3>1.2. Get Top-k Image Pairs Ranked by Keypoint Feature-based Similarity</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2F208da39236d8b8ce2befbedec791b8b8%2F2.png?generation=1748952611810102&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Second, we select image pairs based on the top-k value computed in Step 1.\nInstead of using visual features like DINOv2 or MegaLoc, we rely on a <strong>custom similarity score based on KeyNet-AdaLAM</strong>.<ul>\n<li>I found this KeyNet-AdaLAM in <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2024/discussion/510295\" target=\"_blank\">out teammates jooott and sugupoko's pipeline in IMC2024</a>, and moreover it was originated from <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417045\" target=\"_blank\">excellent solution from the 5th-place team in IMC2023</a>.</li></ul></li>\n<li>Global descriptors could not distinguish between visually similar scenes, however, by carefully looking at small details, I found that adjacent images could be identified—like puzzle pieces that fit together.<br>\nThus, I think keypoint-based matching was considered a more rational choice for extracting accurate image pairs.</li>\n<li>KeyNet-AdaLAM was selected due to below advantages:<ul>\n<li>Robustness to scale and rotation (via AffNet &amp; OriNet)</li>\n<li>Fast exhaustive matching for all image pairs</li>\n<li>Higher precision in pair selection than other sparse matching method (e.g. ALIKED + LightGlue)</li></ul></li>\n<li>The similarity score was heuristically designed, based on the average descriptor distance of inliers and the number of inliers.<br>\nUsing this score, <strong>we could accurately extract true neighboring pairs in fbk_vineyard</strong>, as confirmed by the top-4 examples shown below.<br>\n(Since the dataset uses sequential image file names for adjacent views, the correctness of the selected pairs is evident)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2F1aadc864ec8fb39efcede0a3f217d831%2F2_1.png?generation=1748952798149857&amp;alt=media\" alt=\"\"></li>\n<li>By generating the image pairs in such a strict way, we were able to estimate the camera poses for fbk_vineyard fairly accurately.<br>\n(As for split3, there are opposite-view pairs, and it was difficult to recognize them as part of the same cluster with this pipeline…)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2F40eea1255457216c15b1f66bcfe02321%2F2_2.png?generation=1748952818025411&amp;alt=media\" alt=\"\"></li>\n</ul>\n<p><strong>Note:</strong></p>\n<ul>\n<li>To improve the accuracy of image pair matching, increasing the <code>ransac_iters</code> parameter in AdaLAM is also important (we used 384).</li>\n<li>However, higher values may cause GPU memory errors, so be careful.</li>\n</ul>\n<p><br></p>\n<h3>1.3. Extract Features by ALIKED &amp; Match by LightGlue</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2Ff545553e2ff92aaa2a4db250d2a502ac%2F3.png?generation=1748953000900906&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Next, we perform feature matching on the extracted image pairs using ALIKED + LightGlue (<code>resize_to</code>: 1024, <code>max_num_keypoints</code>: 10000).</li>\n<li>There’s not much customization in this section, but I adopt a cropping method based on DBSCAN, inspired by the <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2024/discussion/510084\" target=\"_blank\">excellent solution from the 1st-place team in IMC2024</a>.<ul>\n<li>I initially assumed that this approach might have limited impact specifically on walkthrough-type scenes like fbk_vineyard, but it still showed some positive effects in both CV and LB, so I decided to include it in our pipeline. (Especially, I remember it working well on the amy_gardens scene.)</li></ul></li>\n<li>Finally, we combine the matching results from the full images and the cropped regions, and pass them to the final step. </li>\n</ul>\n<p><br></p>\n<h3>1.4. Geometric verification &amp; Camera pose estimation (by pycolmap)</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2Fb812de83d435a69de0bc9bc20f867f20%2F4.png?generation=1748953891664264&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Finally, we remove outliers and low-confidence pairs based on MAGSAC++, save the cleaned matches to the database, and run pycolmap for camera-pose estimation.</li>\n<li>Because we already have well-curated image pairs, we don’t use match_exhaustive().</li>\n<li>Parameter tuning in pycolmap gave only small gains, but we saw some benefits by keeping <code>ba_local_max_num_iterations</code> higher.</li>\n</ul>\n<p><br><br>\nThe final metric scores for each training dataset in IMC2025 are shown below.<br>\n\"stairs\" was especially difficult to handle with the current pipeline…</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>ETs</th>\n<th>amy_gardens</th>\n<th>fbk_vineyard</th>\n<th>stairs</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>score (%)</td>\n<td>74.98</td>\n<td>40.00</td>\n<td>62.65</td>\n<td>5.26</td>\n</tr>\n<tr>\n<td>mAA (%)</td>\n<td>60.61</td>\n<td>25.00</td>\n<td>45.62</td>\n<td>2.78</td>\n</tr>\n<tr>\n<td>clusterness (%)</td>\n<td>100.00</td>\n<td>100.00</td>\n<td>100.00</td>\n<td>50.00</td>\n</tr>\n</tbody>\n</table>\n<p><br><br></p>\n<h2>2. Other Ideas</h2>\n<h3>2.1. Multi-processing</h3>\n<ul>\n<li>I referred to <a href=\"https://www.kaggle.com/tmyok1984\" target=\"_blank\">@tmyok1984</a>’s solution from last year, and applied parallel processing using two GPUs (T4 × 2).<br>\nIdeally, CPU and GPU tasks should also be processed in parallel, but in my current solution, the total runtime was acceptable without further optimization.</li>\n</ul>\n<h3>2.2 Methods That Did Not Work</h3>\n<ul>\n<li>Extractors other than ALIKED (While ALIKED+DISK+SIFT showed good results in CV, they did not work at all on the LB, so I excluded them)</li>\n<li>Rotation Handling (Although it is generally important, it did not work well in my pipeline. I guess rotation was not a critical factor in the newly added datasets this year)</li>\n<li>Refiner after Camera Pose Estimation (It didn’t match well with my pipeline. However, I believe it could still be beneficial depending on the configuration)</li>\n</ul>\n<h3>2.3 Exploration of VGGT (Not imcorporate in my pipeline)</h3>\n<ul>\n<li>The exploration of VGGT’s potential was mainly done by <a href=\"https://www.kaggle.com/tmyok1984\" target=\"_blank\">@tmyok1984</a> and <a href=\"https://www.kaggle.com/jooott\" target=\"_blank\">@jooott</a>, and we found that <strong><a href=\"https://github.com/facebookresearch/vggt\" target=\"_blank\">VGGT: Visual Geometry Grounded Transformer</a> can estimate camera poses for “stairs” scene relatively well</strong> (still not perfect).</li>\n<li>Also, additional tests by <a href=\"https://www.kaggle.com/tmyok1984\" target=\"_blank\">@tmyok1984</a> showed that <strong>VGGT also works relatively correctly for datasets with regions with sparse image coverage such as amy_gardens, and for opposite-view pairs in fbk_vineyard split3</strong>.</li>\n<li>On the other hand, absolute accuracy for VGGT's camera pose seems to be not very high, so it was difficult to make it work well in LB.</li>\n<li>Details of these experiments has shared by tmyok in <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2025/discussion/582968\" target=\"_blank\">this thread</a>, so please also check it.</li>\n</ul>\n<p><br><br>\nThanks again to all, and I'd love to join again if the challenge is held next year!</p>",
      "rawMarkdown": "**Update:** @tmyok1984 has shared the details of experiments for VGGT in [this thread](https://www.kaggle.com/competitions/image-matching-challenge-2025/discussion/582968), so please also check it!\n(Note: VGGT is not integrated in below solution)\n\n-----------------------------------------------------------\nFirst of all, I’d like to express my gratitude to the host team and all the staff members for organizing such an exciting competition and providing outstanding support throughout the challenge.\n\nThis was my first time participating in the IMC, but thanks to the many insightful past IMC solutions, I managed to complete this challenge.\nAlso, competing with strong participants really kept me motivated. I’m truly grateful to all of you.\n\nOur approach was strongly supported by @tmyok1984's leadership and the wide range of experiments & discussions.\nAnd also we get a great inspiration from the IMC2024 solutions by @jooott and @sugupoko, which influenced core ideas in our pipeline.\nI really appreciate their support and teamwork during this challenging competition.\n\nIn developing our solution, **I especially focused on selecting image pairs with high accucary** before feature matching.\nThere’s almost nothing particularly special about the pipeline after the image pair selection — it’s a straightforward process using ALIKED-LightGlue for feature matching and pycolmap for camera pose estimation.\nWe made a shake-up in the private leaderboard, but I think this was because our pair selection strategy fit well fortunately.\n\n<br>\n## 1. Key steps\n\n### 1.1. Determine Top-k for Selecting Image Pairs (Dynamic Top-k)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2F44480ec2a6588e2da7339330bd9ac67a%2F1.png?generation=1748952029356808&alt=media)\n- First, we **dynamically adjust the number of top image pairs (top-k)** for each dataset.\n- In visually confusing scenes like fbk_vineyard, I noticed that incorrect image pairs are often formed, which can lead to mixing data from different clusters in the model. \nTherefore, for scenes with high visual similarity, to improve the accuracy of selected image pairs, we **set a very small top-k value (e.g. k = 3)** to strictly select only neighboring image pairs—similar to a MST(minimum spanning tree). The similarity score was calculated based on DINOv2 patch features.\n- In contrast, for scenes like amy_gardens, where images were taken at different times or with different cameras, there may be multiple similar images of the same place. In those cases, we used a larger top-k to keep Recall.\n\n**Note:**\n- Regarding the above strategy, I think using a very small top-k has some risk: it can break the cluster into parts wrongly. \nFor example, if there are many images from the strictly same location, they may connect only with each other, and not with nearby views.\nTherefore, ideally, I think it is better to set a threshold for similarity score instead of top-k. \n- However, setting good thresholds is hard and depends on the dataset, so in this pipeline, I accepted this risk and used this top-k approach.\n- I also assumed that in scenes like fbk_vineyard, it’s unlikely that many images were taken from exactly the same viewpoint. (It is useless act)\nTherefore, I estimated the risk of cluster breakage to be relatively low.\n    - But it is easy to intentionally include duplicate or nearly identical images....\n    - In such cases, additional measures may be needed—for example, if there are image pairs with too much overlap, increase the top-k value accordingly.\n\n<br>\n### 1.2. Get Top-k Image Pairs Ranked by Keypoint Feature-based Similarity\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2F208da39236d8b8ce2befbedec791b8b8%2F2.png?generation=1748952611810102&alt=media)\n- Second, we select image pairs based on the top-k value computed in Step 1.\nInstead of using visual features like DINOv2 or MegaLoc, we rely on a **custom similarity score based on KeyNet-AdaLAM**.\n    - I found this KeyNet-AdaLAM in [out teammates jooott and sugupoko's pipeline in IMC2024](https://www.kaggle.com/competitions/image-matching-challenge-2024/discussion/510295), and moreover it was originated from [excellent solution from the 5th-place team in IMC2023](https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417045).\n- Global descriptors could not distinguish between visually similar scenes, however, by carefully looking at small details, I found that adjacent images could be identified—like puzzle pieces that fit together.\nThus, I think keypoint-based matching was considered a more rational choice for extracting accurate image pairs.\n- KeyNet-AdaLAM was selected due to below advantages:\n    - Robustness to scale and rotation (via AffNet & OriNet)\n    - Fast exhaustive matching for all image pairs\n    - Higher precision in pair selection than other sparse matching method (e.g. ALIKED + LightGlue)\n- The similarity score was heuristically designed, based on the average descriptor distance of inliers and the number of inliers.\nUsing this score, **we could accurately extract true neighboring pairs in fbk_vineyard**, as confirmed by the top-4 examples shown below.\n(Since the dataset uses sequential image file names for adjacent views, the correctness of the selected pairs is evident)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2F1aadc864ec8fb39efcede0a3f217d831%2F2_1.png?generation=1748952798149857&alt=media)\n- By generating the image pairs in such a strict way, we were able to estimate the camera poses for fbk_vineyard fairly accurately.\n(As for split3, there are opposite-view pairs, and it was difficult to recognize them as part of the same cluster with this pipeline...)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2F40eea1255457216c15b1f66bcfe02321%2F2_2.png?generation=1748952818025411&alt=media)\n\n**Note:**\n- To improve the accuracy of image pair matching, increasing the `ransac_iters` parameter in AdaLAM is also important (we used 384).\n- However, higher values may cause GPU memory errors, so be careful.\n\n<br>\n### 1.3. Extract Features by ALIKED & Match by LightGlue\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2Ff545553e2ff92aaa2a4db250d2a502ac%2F3.png?generation=1748953000900906&alt=media)\n- Next, we perform feature matching on the extracted image pairs using ALIKED + LightGlue (`resize_to`: 1024, `max_num_keypoints`: 10000).\n- There’s not much customization in this section, but I adopt a cropping method based on DBSCAN, inspired by the [excellent solution from the 1st-place team in IMC2024](https://www.kaggle.com/competitions/image-matching-challenge-2024/discussion/510084).\n    - I initially assumed that this approach might have limited impact specifically on walkthrough-type scenes like fbk_vineyard, but it still showed some positive effects in both CV and LB, so I decided to include it in our pipeline. (Especially, I remember it working well on the amy_gardens scene.)\n- Finally, we combine the matching results from the full images and the cropped regions, and pass them to the final step. \n\n<br>\n### 1.4. Geometric verification & Camera pose estimation (by pycolmap)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2Fb812de83d435a69de0bc9bc20f867f20%2F4.png?generation=1748953891664264&alt=media)\n- Finally, we remove outliers and low-confidence pairs based on MAGSAC++, save the cleaned matches to the database, and run pycolmap for camera-pose estimation.\n- Because we already have well-curated image pairs, we don’t use match_exhaustive().\n- Parameter tuning in pycolmap gave only small gains, but we saw some benefits by keeping `ba_local_max_num_iterations` higher.\n\n\n<br>\nThe final metric scores for each training dataset in IMC2025 are shown below.\n\"stairs\" was especially difficult to handle with the current pipeline...\n|  | ETs | amy_gardens | fbk_vineyard | stairs |\n| --- | --- | --- | --- | --- |\n| score (%) | 74.98 | 40.00 | 62.65 | 5.26 |\n| mAA (%) | 60.61 | 25.00 | 45.62 | 2.78 |\n| clusterness (%) | 100.00 | 100.00 | 100.00 | 50.00 |\n\n\n<br><br>\n## 2. Other Ideas\n### 2.1. Multi-processing\n- I referred to @tmyok1984’s solution from last year, and applied parallel processing using two GPUs (T4 × 2).\nIdeally, CPU and GPU tasks should also be processed in parallel, but in my current solution, the total runtime was acceptable without further optimization.\n\n### 2.2 Methods That Did Not Work\n- Extractors other than ALIKED (While ALIKED+DISK+SIFT showed good results in CV, they did not work at all on the LB, so I excluded them)\n- Rotation Handling (Although it is generally important, it did not work well in my pipeline. I guess rotation was not a critical factor in the newly added datasets this year)\n- Refiner after Camera Pose Estimation (It didn’t match well with my pipeline. However, I believe it could still be beneficial depending on the configuration)\n\n### 2.3 Exploration of VGGT (Not imcorporate in my pipeline)\n- The exploration of VGGT’s potential was mainly done by @tmyok1984 and @jooott, and we found that **[VGGT: Visual Geometry Grounded Transformer](https://github.com/facebookresearch/vggt) can estimate camera poses for “stairs” scene relatively well** (still not perfect).\n- Also, additional tests by @tmyok1984 showed that **VGGT also works relatively correctly for datasets with regions with sparse image coverage such as amy_gardens, and for opposite-view pairs in fbk_vineyard split3**.\n- On the other hand, absolute accuracy for VGGT's camera pose seems to be not very high, so it was difficult to make it work well in LB.\n- Details of these experiments has shared by tmyok in [this thread](https://www.kaggle.com/competitions/image-matching-challenge-2025/discussion/582968), so please also check it.\n\n<br>\nThanks again to all, and I'd love to join again if the challenge is held next year!",
      "votes": null
    },
    {
      "id": "3242315",
      "postDate": "07/05/2025 19:11:51",
      "content": "<p>This is such an amazing solution! I am a newbie here to CV but really want to learn by recreating recipes from experts such as yourself. Is it okay if I try recreating this on my own for learning purposes and eventually publish my findings at any website (not a formal journal, just articles highlighting my experiences). Of course I fully intend to give full credits to you for the idea. </p>",
      "rawMarkdown": "This is such an amazing solution! I am a newbie here to CV but really want to learn by recreating recipes from experts such as yourself. Is it okay if I try recreating this on my own for learning purposes and eventually publish my findings at any website (not a formal journal, just articles highlighting my experiences). Of course I fully intend to give full credits to you for the idea.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3242315,
      "author_name": "mrishikesh45",
      "author_url": "",
      "post_date": "07/05/2025 19:11:51",
      "content": "<p>This is such an amazing solution! I am a newbie here to CV but really want to learn by recreating recipes from experts such as yourself. Is it okay if I try recreating this on my own for learning purposes and eventually publish my findings at any website (not a formal journal, just articles highlighting my experiences). Of course I fully intend to give full credits to you for the idea. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3216355": "**Update:** @tmyok1984 has shared the details of experiments for VGGT in [this thread](https://www.kaggle.com/competitions/image-matching-challenge-2025/discussion/582968), so please also check it!\n(Note: VGGT is not integrated in below solution)\n\n-----------------------------------------------------------\nFirst of all, I’d like to express my gratitude to the host team and all the staff members for organizing such an exciting competition and providing outstanding support throughout the challenge.\n\nThis was my first time participating in the IMC, but thanks to the many insightful past IMC solutions, I managed to complete this challenge.\nAlso, competing with strong participants really kept me motivated. I’m truly grateful to all of you.\n\nOur approach was strongly supported by @tmyok1984's leadership and the wide range of experiments & discussions.\nAnd also we get a great inspiration from the IMC2024 solutions by @jooott and @sugupoko, which influenced core ideas in our pipeline.\nI really appreciate their support and teamwork during this challenging competition.\n\nIn developing our solution, **I especially focused on selecting image pairs with high accucary** before feature matching.\nThere’s almost nothing particularly special about the pipeline after the image pair selection — it’s a straightforward process using ALIKED-LightGlue for feature matching and pycolmap for camera pose estimation.\nWe made a shake-up in the private leaderboard, but I think this was because our pair selection strategy fit well fortunately.\n\n<br>\n## 1. Key steps\n\n### 1.1. Determine Top-k for Selecting Image Pairs (Dynamic Top-k)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2F44480ec2a6588e2da7339330bd9ac67a%2F1.png?generation=1748952029356808&alt=media)\n- First, we **dynamically adjust the number of top image pairs (top-k)** for each dataset.\n- In visually confusing scenes like fbk_vineyard, I noticed that incorrect image pairs are often formed, which can lead to mixing data from different clusters in the model. \nTherefore, for scenes with high visual similarity, to improve the accuracy of selected image pairs, we **set a very small top-k value (e.g. k = 3)** to strictly select only neighboring image pairs—similar to a MST(minimum spanning tree). The similarity score was calculated based on DINOv2 patch features.\n- In contrast, for scenes like amy_gardens, where images were taken at different times or with different cameras, there may be multiple similar images of the same place. In those cases, we used a larger top-k to keep Recall.\n\n**Note:**\n- Regarding the above strategy, I think using a very small top-k has some risk: it can break the cluster into parts wrongly. \nFor example, if there are many images from the strictly same location, they may connect only with each other, and not with nearby views.\nTherefore, ideally, I think it is better to set a threshold for similarity score instead of top-k. \n- However, setting good thresholds is hard and depends on the dataset, so in this pipeline, I accepted this risk and used this top-k approach.\n- I also assumed that in scenes like fbk_vineyard, it’s unlikely that many images were taken from exactly the same viewpoint. (It is useless act)\nTherefore, I estimated the risk of cluster breakage to be relatively low.\n    - But it is easy to intentionally include duplicate or nearly identical images....\n    - In such cases, additional measures may be needed—for example, if there are image pairs with too much overlap, increase the top-k value accordingly.\n\n<br>\n### 1.2. Get Top-k Image Pairs Ranked by Keypoint Feature-based Similarity\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2F208da39236d8b8ce2befbedec791b8b8%2F2.png?generation=1748952611810102&alt=media)\n- Second, we select image pairs based on the top-k value computed in Step 1.\nInstead of using visual features like DINOv2 or MegaLoc, we rely on a **custom similarity score based on KeyNet-AdaLAM**.\n    - I found this KeyNet-AdaLAM in [out teammates jooott and sugupoko's pipeline in IMC2024](https://www.kaggle.com/competitions/image-matching-challenge-2024/discussion/510295), and moreover it was originated from [excellent solution from the 5th-place team in IMC2023](https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417045).\n- Global descriptors could not distinguish between visually similar scenes, however, by carefully looking at small details, I found that adjacent images could be identified—like puzzle pieces that fit together.\nThus, I think keypoint-based matching was considered a more rational choice for extracting accurate image pairs.\n- KeyNet-AdaLAM was selected due to below advantages:\n    - Robustness to scale and rotation (via AffNet & OriNet)\n    - Fast exhaustive matching for all image pairs\n    - Higher precision in pair selection than other sparse matching method (e.g. ALIKED + LightGlue)\n- The similarity score was heuristically designed, based on the average descriptor distance of inliers and the number of inliers.\nUsing this score, **we could accurately extract true neighboring pairs in fbk_vineyard**, as confirmed by the top-4 examples shown below.\n(Since the dataset uses sequential image file names for adjacent views, the correctness of the selected pairs is evident)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2F1aadc864ec8fb39efcede0a3f217d831%2F2_1.png?generation=1748952798149857&alt=media)\n- By generating the image pairs in such a strict way, we were able to estimate the camera poses for fbk_vineyard fairly accurately.\n(As for split3, there are opposite-view pairs, and it was difficult to recognize them as part of the same cluster with this pipeline...)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2F40eea1255457216c15b1f66bcfe02321%2F2_2.png?generation=1748952818025411&alt=media)\n\n**Note:**\n- To improve the accuracy of image pair matching, increasing the `ransac_iters` parameter in AdaLAM is also important (we used 384).\n- However, higher values may cause GPU memory errors, so be careful.\n\n<br>\n### 1.3. Extract Features by ALIKED & Match by LightGlue\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2Ff545553e2ff92aaa2a4db250d2a502ac%2F3.png?generation=1748953000900906&alt=media)\n- Next, we perform feature matching on the extracted image pairs using ALIKED + LightGlue (`resize_to`: 1024, `max_num_keypoints`: 10000).\n- There’s not much customization in this section, but I adopt a cropping method based on DBSCAN, inspired by the [excellent solution from the 1st-place team in IMC2024](https://www.kaggle.com/competitions/image-matching-challenge-2024/discussion/510084).\n    - I initially assumed that this approach might have limited impact specifically on walkthrough-type scenes like fbk_vineyard, but it still showed some positive effects in both CV and LB, so I decided to include it in our pipeline. (Especially, I remember it working well on the amy_gardens scene.)\n- Finally, we combine the matching results from the full images and the cropped regions, and pass them to the final step. \n\n<br>\n### 1.4. Geometric verification & Camera pose estimation (by pycolmap)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2Fb812de83d435a69de0bc9bc20f867f20%2F4.png?generation=1748953891664264&alt=media)\n- Finally, we remove outliers and low-confidence pairs based on MAGSAC++, save the cleaned matches to the database, and run pycolmap for camera-pose estimation.\n- Because we already have well-curated image pairs, we don’t use match_exhaustive().\n- Parameter tuning in pycolmap gave only small gains, but we saw some benefits by keeping `ba_local_max_num_iterations` higher.\n\n\n<br>\nThe final metric scores for each training dataset in IMC2025 are shown below.\n\"stairs\" was especially difficult to handle with the current pipeline...\n|  | ETs | amy_gardens | fbk_vineyard | stairs |\n| --- | --- | --- | --- | --- |\n| score (%) | 74.98 | 40.00 | 62.65 | 5.26 |\n| mAA (%) | 60.61 | 25.00 | 45.62 | 2.78 |\n| clusterness (%) | 100.00 | 100.00 | 100.00 | 50.00 |\n\n\n<br><br>\n## 2. Other Ideas\n### 2.1. Multi-processing\n- I referred to @tmyok1984’s solution from last year, and applied parallel processing using two GPUs (T4 × 2).\nIdeally, CPU and GPU tasks should also be processed in parallel, but in my current solution, the total runtime was acceptable without further optimization.\n\n### 2.2 Methods That Did Not Work\n- Extractors other than ALIKED (While ALIKED+DISK+SIFT showed good results in CV, they did not work at all on the LB, so I excluded them)\n- Rotation Handling (Although it is generally important, it did not work well in my pipeline. I guess rotation was not a critical factor in the newly added datasets this year)\n- Refiner after Camera Pose Estimation (It didn’t match well with my pipeline. However, I believe it could still be beneficial depending on the configuration)\n\n### 2.3 Exploration of VGGT (Not imcorporate in my pipeline)\n- The exploration of VGGT’s potential was mainly done by @tmyok1984 and @jooott, and we found that **[VGGT: Visual Geometry Grounded Transformer](https://github.com/facebookresearch/vggt) can estimate camera poses for “stairs” scene relatively well** (still not perfect).\n- Also, additional tests by @tmyok1984 showed that **VGGT also works relatively correctly for datasets with regions with sparse image coverage such as amy_gardens, and for opposite-view pairs in fbk_vineyard split3**.\n- On the other hand, absolute accuracy for VGGT's camera pose seems to be not very high, so it was difficult to make it work well in LB.\n- Details of these experiments has shared by tmyok in [this thread](https://www.kaggle.com/competitions/image-matching-challenge-2025/discussion/582968), so please also check it.\n\n<br>\nThanks again to all, and I'd love to join again if the challenge is held next year!",
    "3242315": "This is such an amazing solution! I am a newbie here to CV but really want to learn by recreating recipes from experts such as yourself. Is it okay if I try recreating this on my own for learning purposes and eventually publish my findings at any website (not a formal journal, just articles highlighting my experiences). Of course I fully intend to give full credits to you for the idea."
  },
  "source": "meta"
}