{
  "id": 583464,
  "title": "9th Place Solution: DINOv2-Optimized Filtering",
  "url": "/competitions/image-matching-challenge-2025/discussion/583464",
  "author_name": "ymg_aq",
  "post_date": "2025-06-07T04:50:19.601000",
  "votes": 23,
  "comment_count": 0,
  "views": 0,
  "content": "<p>First and foremost, I would like to extend my sincere gratitude to the organizers of the 2025 Image Matching Challenge for hosting such an inspiring and meticulously run competition.</p>\n<h1>1. Overview</h1>\n<p>My method is designed around a two-stage philosophy: a “lenient first stage” that tries not to miss any true positives, followed by “subsequent stages” that prune false matches. Because the competition’s metric is the harmonic mean of clustering accuracy and pose accuracy, sacrificing either side drastically lowers the final score. Consequently, the first stage must keep as many potentially correct pairs as possible, while later stages must suppress error propagation to achieve high-precision 3-D reconstruction.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4717152%2F3310941dcd9cd34d82d2ecee34b73077%2F9th_solution_pipeline.png?generation=1749269384826324&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Two-stage design: a recall-oriented first stage keeps image pairs, followed by a precision-oriented stage that removes false matches.</li>\n<li>Similarity filtering: retain pairs with DINOv2 cosine ≦ 0.15 or within the top 50 similarities, but switch to exhaustive matching when the dataset has fewer than 100 images.</li>\n<li>Multi-resolution matching: run ALIKED + LightGlue at 840 px → 1280 px → 2048 px, use DBSCAN to focus on dense regions, then apply RANSAC for final outlier removal.</li>\n</ul>\n<h1>2. Solution Details</h1>\n<h2>2-1. Similarity Filtering (DINOv2 or exhaustive matching)</h2>\n<p>In preprocessing and image-pair generation, global feature vectors are extracted with a pretrained DINOv2 model, and cosine similarities are computed for every image pair. Pairs that fall below the similarity threshold of 0.15 are first filtered out, and the top 50 most similar pairs are then selected. If the number of images is less than 100, an adaptive mechanism switches to exhaustive matching for computational efficiency.</p>\n<h2>2-2. Multi-resolution keypoint detection and matching (ALIKED + LightGlue)</h2>\n<p>The pipeline then performs multi-resolution keypoint detection and matching, re-implementing <a href=\"https://www.kaggle.com/code/vostankovich/imc2024-1st-place-solution#Merge-csv\" target=\"_blank\">the first-place solution from IMC 2024</a>. In the initial phase, images are resized to 840 pixels, and rotation-aware matching is executed. Up to 1024 keypoints are extracted with the ALIKED detector, LightGlue finds correspondences, and rotations of 0°, 90°, 180°, and 270° are compensated. In the second phase the resolution is raised to 1280 pixels, up to 8192 keypoints are detected, and a stricter filtering threshold of 0.2 is applied. The third phase processes images at 2048 pixels to obtain the most precise matches.</p>\n<p>After matching, DBSCAN clustering identifies high-density regions of correspondences. Based on this analysis, rectangular areas where important features concentrate are calculated; these become crop regions that efficiently restrict the target area for subsequent high-precision matching. Detailed analysis is carried out on these regions at 1280 pixels and 2048 pixels, and finally the results from all resolutions and regions are integrated. In the outlier-removal and optimization stage, RANSAC detects and eliminates mismatched correspondences, and a final filter removes pairs whose match score falls below a threshold, thereby completing the keypoint detection and matching process.</p>\n<h2>2-3. 3-D reconstruction and pose estimation (COLMAP)</h2>\n<p>The features, matches, and fundamental matrices generated by the pipeline are converted into a COLMAP database and imported to provide the foundation for 3-D reconstruction. Incremental reconstruction begins with the pair having the largest number of matches. Bundle adjustment continuously optimizes the model while new images are added step by step, and at the same time the 3-D point cloud is updated. Several reconstructions are attempted under different initial conditions to find the best result. Independent scene clusters are detected and processed separately, and the final camera poses are obtained. Cluster labels are assigned in a straightforward manner according to the camera model.</p>\n<h2>2-4. Acceleration</h2>\n<p>To balance the load during parallel processing on 2x T4 GPUs, the dataset is divided into two groups so that the sum of the squares of the image counts is nearly equal. Correcting the imbalance that sometimes occurred in the original implementation shortens the processing time by up to about 15% in certain pipelines.</p>\n<h1>3. Scores</h1>\n<h3>Train Dataset</h3>\n<table>\n<thead>\n<tr>\n<th>Dataset</th>\n<th>Score</th>\n<th>mAA</th>\n<th>Clusterness</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>imc2023_haiper</td>\n<td>68.08%</td>\n<td>73.33%</td>\n<td>63.53%</td>\n</tr>\n<tr>\n<td>imc2023_heritage</td>\n<td>87.56%</td>\n<td>77.88%</td>\n<td>100.00%</td>\n</tr>\n<tr>\n<td>imc2023_theather_imc2024_church</td>\n<td>68.24%</td>\n<td>51.79%</td>\n<td>100.00%</td>\n</tr>\n<tr>\n<td>imc2024_dioscuri_baalshamin</td>\n<td>91.73%</td>\n<td>84.72%</td>\n<td>100.00%</td>\n</tr>\n<tr>\n<td>imc2024_lizard_pond</td>\n<td>72.88%</td>\n<td>57.33%</td>\n<td>100.00%</td>\n</tr>\n<tr>\n<td>pt_brandenburg_british_buckingham</td>\n<td>67.70%</td>\n<td>78.59%</td>\n<td>59.47%</td>\n</tr>\n<tr>\n<td>pt_piazzasanmarco_grandplace</td>\n<td>87.40%</td>\n<td>77.62%</td>\n<td>100.00%</td>\n</tr>\n<tr>\n<td>pt_sacrecoeur_trevi_tajmahal</td>\n<td>93.33%</td>\n<td>87.50%</td>\n<td>100.00%</td>\n</tr>\n<tr>\n<td>pt_stpeters_stpauls</td>\n<td>61.35%</td>\n<td>79.38%</td>\n<td>50.00%</td>\n</tr>\n<tr>\n<td>amy_gardens</td>\n<td>22.12%</td>\n<td>12.44%</td>\n<td>100.00%</td>\n</tr>\n<tr>\n<td>fbk_vineyard</td>\n<td>42.37%</td>\n<td>39.77%</td>\n<td>45.32%</td>\n</tr>\n<tr>\n<td>ETs</td>\n<td>64.94%</td>\n<td>48.08%</td>\n<td>100.00%</td>\n</tr>\n<tr>\n<td>stairs</td>\n<td>0.00%</td>\n<td>0.00%</td>\n<td>71.43%</td>\n</tr>\n<tr>\n<td><strong>Average over all datasets</strong></td>\n<td>63.67%</td>\n<td>59.11%</td>\n<td>83.83%</td>\n</tr>\n</tbody>\n</table>\n<h3>Submissions</h3>\n<p>Scores by differences in first-stage filtering.</p>\n<table>\n<thead>\n<tr>\n<th>Method</th>\n<th>Private</th>\n<th>Public</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>DINOv2 (similarity &lt; 0.15, top-k = 50, #images ≧ 100)</strong></td>\n<td><strong>44.69</strong></td>\n<td><strong>41.08</strong></td>\n</tr>\n<tr>\n<td>DINOv2 (similarity &lt; 0.3, top-k = 20, #images ≧ 20)</td>\n<td>42.93</td>\n<td>39.87</td>\n</tr>\n<tr>\n<td>Exhaustive matching</td>\n<td>41.74</td>\n<td>36.37</td>\n</tr>\n</tbody>\n</table>\n<h1>4. Attempts that did not work</h1>\n<ul>\n<li>Adaptive threshold tuning for DINOv2 was attempted using Bayesian optimization, but each evaluation step was computationally heavy and the parameter space could not converge within the time limit. </li>\n<li>A binary classifier that identified clusters using edge weights in the closed graph—such as the number of inlier matches and the reprojection error—was trained independently of COLMAP’s camera-model clustering, but precision fell sharply when recall was increased, leaving F1 unimproved, so the idea was abandoned. </li>\n<li>Experiments that used <a href=\"https://github.com/nianticlabs/mickey\" target=\"_blank\">MicKey</a> features led to an overly dense LightGlue match graph, which became over-connected; accumulated loop-closure errors caused the final poses to diverge. </li>\n<li>Applying <a href=\"https://github.com/zju3dv/LoFTR\" target=\"_blank\">LoFTR</a> to all pairs and then again after filtering greatly increased the number of high-resolution pairs, but incorrect correspondences retained from the coarse pass over-fit and converged incorrectly in scenes with low parallax.</li>\n</ul>",
  "messages": [
    {
      "id": 3219028,
      "postDate": "2025-06-07T04:50:19.600Z",
      "content": "<p>First and foremost, I would like to extend my sincere gratitude to the organizers of the 2025 Image Matching Challenge for hosting such an inspiring and meticulously run competition.</p>\n<h1>1. Overview</h1>\n<p>My method is designed around a two-stage philosophy: a “lenient first stage” that tries not to miss any true positives, followed by “subsequent stages” that prune false matches. Because the competition’s metric is the harmonic mean of clustering accuracy and pose accuracy, sacrificing either side drastically lowers the final score. Consequently, the first stage must keep as many potentially correct pairs as possible, while later stages must suppress error propagation to achieve high-precision 3-D reconstruction.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4717152%2F3310941dcd9cd34d82d2ecee34b73077%2F9th_solution_pipeline.png?generation=1749269384826324&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Two-stage design: a recall-oriented first stage keeps image pairs, followed by a precision-oriented stage that removes false matches.</li>\n<li>Similarity filtering: retain pairs with DINOv2 cosine ≦ 0.15 or within the top 50 similarities, but switch to exhaustive matching when the dataset has fewer than 100 images.</li>\n<li>Multi-resolution matching: run ALIKED + LightGlue at 840 px → 1280 px → 2048 px, use DBSCAN to focus on dense regions, then apply RANSAC for final outlier removal.</li>\n</ul>\n<h1>2. Solution Details</h1>\n<h2>2-1. Similarity Filtering (DINOv2 or exhaustive matching)</h2>\n<p>In preprocessing and image-pair generation, global feature vectors are extracted with a pretrained DINOv2 model, and cosine similarities are computed for every image pair. Pairs that fall below the similarity threshold of 0.15 are first filtered out, and the top 50 most similar pairs are then selected. If the number of images is less than 100, an adaptive mechanism switches to exhaustive matching for computational efficiency.</p>\n<h2>2-2. Multi-resolution keypoint detection and matching (ALIKED + LightGlue)</h2>\n<p>The pipeline then performs multi-resolution keypoint detection and matching, re-implementing <a href=\"https://www.kaggle.com/code/vostankovich/imc2024-1st-place-solution#Merge-csv\" target=\"_blank\">the first-place solution from IMC 2024</a>. In the initial phase, images are resized to 840 pixels, and rotation-aware matching is executed. Up to 1024 keypoints are extracted with the ALIKED detector, LightGlue finds correspondences, and rotations of 0°, 90°, 180°, and 270° are compensated. In the second phase the resolution is raised to 1280 pixels, up to 8192 keypoints are detected, and a stricter filtering threshold of 0.2 is applied. The third phase processes images at 2048 pixels to obtain the most precise matches.</p>\n<p>After matching, DBSCAN clustering identifies high-density regions of correspondences. Based on this analysis, rectangular areas where important features concentrate are calculated; these become crop regions that efficiently restrict the target area for subsequent high-precision matching. Detailed analysis is carried out on these regions at 1280 pixels and 2048 pixels, and finally the results from all resolutions and regions are integrated. In the outlier-removal and optimization stage, RANSAC detects and eliminates mismatched correspondences, and a final filter removes pairs whose match score falls below a threshold, thereby completing the keypoint detection and matching process.</p>\n<h2>2-3. 3-D reconstruction and pose estimation (COLMAP)</h2>\n<p>The features, matches, and fundamental matrices generated by the pipeline are converted into a COLMAP database and imported to provide the foundation for 3-D reconstruction. Incremental reconstruction begins with the pair having the largest number of matches. Bundle adjustment continuously optimizes the model while new images are added step by step, and at the same time the 3-D point cloud is updated. Several reconstructions are attempted under different initial conditions to find the best result. Independent scene clusters are detected and processed separately, and the final camera poses are obtained. Cluster labels are assigned in a straightforward manner according to the camera model.</p>\n<h2>2-4. Acceleration</h2>\n<p>To balance the load during parallel processing on 2x T4 GPUs, the dataset is divided into two groups so that the sum of the squares of the image counts is nearly equal. Correcting the imbalance that sometimes occurred in the original implementation shortens the processing time by up to about 15% in certain pipelines.</p>\n<h1>3. Scores</h1>\n<h3>Train Dataset</h3>\n<table>\n<thead>\n<tr>\n<th>Dataset</th>\n<th>Score</th>\n<th>mAA</th>\n<th>Clusterness</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>imc2023_haiper</td>\n<td>68.08%</td>\n<td>73.33%</td>\n<td>63.53%</td>\n</tr>\n<tr>\n<td>imc2023_heritage</td>\n<td>87.56%</td>\n<td>77.88%</td>\n<td>100.00%</td>\n</tr>\n<tr>\n<td>imc2023_theather_imc2024_church</td>\n<td>68.24%</td>\n<td>51.79%</td>\n<td>100.00%</td>\n</tr>\n<tr>\n<td>imc2024_dioscuri_baalshamin</td>\n<td>91.73%</td>\n<td>84.72%</td>\n<td>100.00%</td>\n</tr>\n<tr>\n<td>imc2024_lizard_pond</td>\n<td>72.88%</td>\n<td>57.33%</td>\n<td>100.00%</td>\n</tr>\n<tr>\n<td>pt_brandenburg_british_buckingham</td>\n<td>67.70%</td>\n<td>78.59%</td>\n<td>59.47%</td>\n</tr>\n<tr>\n<td>pt_piazzasanmarco_grandplace</td>\n<td>87.40%</td>\n<td>77.62%</td>\n<td>100.00%</td>\n</tr>\n<tr>\n<td>pt_sacrecoeur_trevi_tajmahal</td>\n<td>93.33%</td>\n<td>87.50%</td>\n<td>100.00%</td>\n</tr>\n<tr>\n<td>pt_stpeters_stpauls</td>\n<td>61.35%</td>\n<td>79.38%</td>\n<td>50.00%</td>\n</tr>\n<tr>\n<td>amy_gardens</td>\n<td>22.12%</td>\n<td>12.44%</td>\n<td>100.00%</td>\n</tr>\n<tr>\n<td>fbk_vineyard</td>\n<td>42.37%</td>\n<td>39.77%</td>\n<td>45.32%</td>\n</tr>\n<tr>\n<td>ETs</td>\n<td>64.94%</td>\n<td>48.08%</td>\n<td>100.00%</td>\n</tr>\n<tr>\n<td>stairs</td>\n<td>0.00%</td>\n<td>0.00%</td>\n<td>71.43%</td>\n</tr>\n<tr>\n<td><strong>Average over all datasets</strong></td>\n<td>63.67%</td>\n<td>59.11%</td>\n<td>83.83%</td>\n</tr>\n</tbody>\n</table>\n<h3>Submissions</h3>\n<p>Scores by differences in first-stage filtering.</p>\n<table>\n<thead>\n<tr>\n<th>Method</th>\n<th>Private</th>\n<th>Public</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>DINOv2 (similarity &lt; 0.15, top-k = 50, #images ≧ 100)</strong></td>\n<td><strong>44.69</strong></td>\n<td><strong>41.08</strong></td>\n</tr>\n<tr>\n<td>DINOv2 (similarity &lt; 0.3, top-k = 20, #images ≧ 20)</td>\n<td>42.93</td>\n<td>39.87</td>\n</tr>\n<tr>\n<td>Exhaustive matching</td>\n<td>41.74</td>\n<td>36.37</td>\n</tr>\n</tbody>\n</table>\n<h1>4. Attempts that did not work</h1>\n<ul>\n<li>Adaptive threshold tuning for DINOv2 was attempted using Bayesian optimization, but each evaluation step was computationally heavy and the parameter space could not converge within the time limit. </li>\n<li>A binary classifier that identified clusters using edge weights in the closed graph—such as the number of inlier matches and the reprojection error—was trained independently of COLMAP’s camera-model clustering, but precision fell sharply when recall was increased, leaving F1 unimproved, so the idea was abandoned. </li>\n<li>Experiments that used <a href=\"https://github.com/nianticlabs/mickey\" target=\"_blank\">MicKey</a> features led to an overly dense LightGlue match graph, which became over-connected; accumulated loop-closure errors caused the final poses to diverge. </li>\n<li>Applying <a href=\"https://github.com/zju3dv/LoFTR\" target=\"_blank\">LoFTR</a> to all pairs and then again after filtering greatly increased the number of high-resolution pairs, but incorrect correspondences retained from the coarse pass over-fit and converged incorrectly in scenes with low parallax.</li>\n</ul>",
      "rawMarkdown": "First and foremost, I would like to extend my sincere gratitude to the organizers of the 2025 Image Matching Challenge for hosting such an inspiring and meticulously run competition.\n\n# 1. Overview\nMy method is designed around a two-stage philosophy: a “lenient first stage” that tries not to miss any true positives, followed by “subsequent stages” that prune false matches. Because the competition’s metric is the harmonic mean of clustering accuracy and pose accuracy, sacrificing either side drastically lowers the final score. Consequently, the first stage must keep as many potentially correct pairs as possible, while later stages must suppress error propagation to achieve high-precision 3-D reconstruction.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4717152%2F3310941dcd9cd34d82d2ecee34b73077%2F9th_solution_pipeline.png?generation=1749269384826324&alt=media)\n\n- Two-stage design: a recall-oriented first stage keeps image pairs, followed by a precision-oriented stage that removes false matches.\n- Similarity filtering: retain pairs with DINOv2 cosine ≦ 0.15 or within the top 50 similarities, but switch to exhaustive matching when the dataset has fewer than 100 images.\n- Multi-resolution matching: run ALIKED + LightGlue at 840 px → 1280 px → 2048 px, use DBSCAN to focus on dense regions, then apply RANSAC for final outlier removal.\n\n# 2. Solution Details\n\n## 2-1. Similarity Filtering (DINOv2 or exhaustive matching)\nIn preprocessing and image-pair generation, global feature vectors are extracted with a pretrained DINOv2 model, and cosine similarities are computed for every image pair. Pairs that fall below the similarity threshold of 0.15 are first filtered out, and the top 50 most similar pairs are then selected. If the number of images is less than 100, an adaptive mechanism switches to exhaustive matching for computational efficiency.\n\n## 2-2. Multi-resolution keypoint detection and matching (ALIKED + LightGlue)\nThe pipeline then performs multi-resolution keypoint detection and matching, re-implementing [the first-place solution from IMC 2024](https://www.kaggle.com/code/vostankovich/imc2024-1st-place-solution#Merge-csv). In the initial phase, images are resized to 840 pixels, and rotation-aware matching is executed. Up to 1024 keypoints are extracted with the ALIKED detector, LightGlue finds correspondences, and rotations of 0°, 90°, 180°, and 270° are compensated. In the second phase the resolution is raised to 1280 pixels, up to 8192 keypoints are detected, and a stricter filtering threshold of 0.2 is applied. The third phase processes images at 2048 pixels to obtain the most precise matches.\n\nAfter matching, DBSCAN clustering identifies high-density regions of correspondences. Based on this analysis, rectangular areas where important features concentrate are calculated; these become crop regions that efficiently restrict the target area for subsequent high-precision matching. Detailed analysis is carried out on these regions at 1280 pixels and 2048 pixels, and finally the results from all resolutions and regions are integrated. In the outlier-removal and optimization stage, RANSAC detects and eliminates mismatched correspondences, and a final filter removes pairs whose match score falls below a threshold, thereby completing the keypoint detection and matching process.\n\n## 2-3. 3-D reconstruction and pose estimation (COLMAP)\nThe features, matches, and fundamental matrices generated by the pipeline are converted into a COLMAP database and imported to provide the foundation for 3-D reconstruction. Incremental reconstruction begins with the pair having the largest number of matches. Bundle adjustment continuously optimizes the model while new images are added step by step, and at the same time the 3-D point cloud is updated. Several reconstructions are attempted under different initial conditions to find the best result. Independent scene clusters are detected and processed separately, and the final camera poses are obtained. Cluster labels are assigned in a straightforward manner according to the camera model.\n\n## 2-4. Acceleration\nTo balance the load during parallel processing on 2x T4 GPUs, the dataset is divided into two groups so that the sum of the squares of the image counts is nearly equal. Correcting the imbalance that sometimes occurred in the original implementation shortens the processing time by up to about 15% in certain pipelines.\n\n# 3. Scores\n### Train Dataset\n| Dataset                                     | Score  | mAA    | Clusterness |\n|---------------------------------------------|--------|--------|-------------|\n| imc2023_haiper                              | 68.08% | 73.33% | 63.53%      |\n| imc2023_heritage                            | 87.56% | 77.88% | 100.00%     |\n| imc2023_theather_imc2024_church             | 68.24% | 51.79% | 100.00%     |\n| imc2024_dioscuri_baalshamin                 | 91.73% | 84.72% | 100.00%     |\n| imc2024_lizard_pond                         | 72.88% | 57.33% | 100.00%     |\n| pt_brandenburg_british_buckingham           | 67.70% | 78.59% | 59.47%      |\n| pt_piazzasanmarco_grandplace                | 87.40% | 77.62% | 100.00%     |\n| pt_sacrecoeur_trevi_tajmahal                | 93.33% | 87.50% | 100.00%     |\n| pt_stpeters_stpauls                         | 61.35% | 79.38% | 50.00%      |\n| amy_gardens                                 | 22.12% | 12.44% | 100.00%     |\n| fbk_vineyard                                | 42.37% | 39.77% | 45.32%      |\n| ETs                                         | 64.94% | 48.08% | 100.00%     |\n| stairs                                      | 0.00%  | 0.00%  | 71.43%      |\n| **Average over all datasets**               | 63.67% | 59.11% | 83.83%      |\n\n### Submissions\nScores by differences in first-stage filtering.\n\n| Method                                                | Private | Public |\n|-------------------------------------------------------|--------:|-------:|\n| **DINOv2 (similarity < 0.15, top-k = 50, #images ≧ 100)** |  **44.69** |  **41.08** |\n| DINOv2 (similarity < 0.3, top-k = 20, #images ≧ 20)   |  42.93 |  39.87 |\n| Exhaustive matching                                 |  41.74 |  36.37 |\n\n# 4. Attempts that did not work\n- Adaptive threshold tuning for DINOv2 was attempted using Bayesian optimization, but each evaluation step was computationally heavy and the parameter space could not converge within the time limit. \n- A binary classifier that identified clusters using edge weights in the closed graph—such as the number of inlier matches and the reprojection error—was trained independently of COLMAP’s camera-model clustering, but precision fell sharply when recall was increased, leaving F1 unimproved, so the idea was abandoned. \n- Experiments that used [MicKey](https://github.com/nianticlabs/mickey) features led to an overly dense LightGlue match graph, which became over-connected; accumulated loop-closure errors caused the final poses to diverge. \n- Applying [LoFTR](https://github.com/zju3dv/LoFTR) to all pairs and then again after filtering greatly increased the number of high-resolution pairs, but incorrect correspondences retained from the coarse pass over-fit and converged incorrectly in scenes with low parallax.",
      "votes": 23
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3219028": "First and foremost, I would like to extend my sincere gratitude to the organizers of the 2025 Image Matching Challenge for hosting such an inspiring and meticulously run competition.\n\n# 1. Overview\nMy method is designed around a two-stage philosophy: a “lenient first stage” that tries not to miss any true positives, followed by “subsequent stages” that prune false matches. Because the competition’s metric is the harmonic mean of clustering accuracy and pose accuracy, sacrificing either side drastically lowers the final score. Consequently, the first stage must keep as many potentially correct pairs as possible, while later stages must suppress error propagation to achieve high-precision 3-D reconstruction.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4717152%2F3310941dcd9cd34d82d2ecee34b73077%2F9th_solution_pipeline.png?generation=1749269384826324&alt=media)\n\n- Two-stage design: a recall-oriented first stage keeps image pairs, followed by a precision-oriented stage that removes false matches.\n- Similarity filtering: retain pairs with DINOv2 cosine ≦ 0.15 or within the top 50 similarities, but switch to exhaustive matching when the dataset has fewer than 100 images.\n- Multi-resolution matching: run ALIKED + LightGlue at 840 px → 1280 px → 2048 px, use DBSCAN to focus on dense regions, then apply RANSAC for final outlier removal.\n\n# 2. Solution Details\n\n## 2-1. Similarity Filtering (DINOv2 or exhaustive matching)\nIn preprocessing and image-pair generation, global feature vectors are extracted with a pretrained DINOv2 model, and cosine similarities are computed for every image pair. Pairs that fall below the similarity threshold of 0.15 are first filtered out, and the top 50 most similar pairs are then selected. If the number of images is less than 100, an adaptive mechanism switches to exhaustive matching for computational efficiency.\n\n## 2-2. Multi-resolution keypoint detection and matching (ALIKED + LightGlue)\nThe pipeline then performs multi-resolution keypoint detection and matching, re-implementing [the first-place solution from IMC 2024](https://www.kaggle.com/code/vostankovich/imc2024-1st-place-solution#Merge-csv). In the initial phase, images are resized to 840 pixels, and rotation-aware matching is executed. Up to 1024 keypoints are extracted with the ALIKED detector, LightGlue finds correspondences, and rotations of 0°, 90°, 180°, and 270° are compensated. In the second phase the resolution is raised to 1280 pixels, up to 8192 keypoints are detected, and a stricter filtering threshold of 0.2 is applied. The third phase processes images at 2048 pixels to obtain the most precise matches.\n\nAfter matching, DBSCAN clustering identifies high-density regions of correspondences. Based on this analysis, rectangular areas where important features concentrate are calculated; these become crop regions that efficiently restrict the target area for subsequent high-precision matching. Detailed analysis is carried out on these regions at 1280 pixels and 2048 pixels, and finally the results from all resolutions and regions are integrated. In the outlier-removal and optimization stage, RANSAC detects and eliminates mismatched correspondences, and a final filter removes pairs whose match score falls below a threshold, thereby completing the keypoint detection and matching process.\n\n## 2-3. 3-D reconstruction and pose estimation (COLMAP)\nThe features, matches, and fundamental matrices generated by the pipeline are converted into a COLMAP database and imported to provide the foundation for 3-D reconstruction. Incremental reconstruction begins with the pair having the largest number of matches. Bundle adjustment continuously optimizes the model while new images are added step by step, and at the same time the 3-D point cloud is updated. Several reconstructions are attempted under different initial conditions to find the best result. Independent scene clusters are detected and processed separately, and the final camera poses are obtained. Cluster labels are assigned in a straightforward manner according to the camera model.\n\n## 2-4. Acceleration\nTo balance the load during parallel processing on 2x T4 GPUs, the dataset is divided into two groups so that the sum of the squares of the image counts is nearly equal. Correcting the imbalance that sometimes occurred in the original implementation shortens the processing time by up to about 15% in certain pipelines.\n\n# 3. Scores\n### Train Dataset\n| Dataset                                     | Score  | mAA    | Clusterness |\n|---------------------------------------------|--------|--------|-------------|\n| imc2023_haiper                              | 68.08% | 73.33% | 63.53%      |\n| imc2023_heritage                            | 87.56% | 77.88% | 100.00%     |\n| imc2023_theather_imc2024_church             | 68.24% | 51.79% | 100.00%     |\n| imc2024_dioscuri_baalshamin                 | 91.73% | 84.72% | 100.00%     |\n| imc2024_lizard_pond                         | 72.88% | 57.33% | 100.00%     |\n| pt_brandenburg_british_buckingham           | 67.70% | 78.59% | 59.47%      |\n| pt_piazzasanmarco_grandplace                | 87.40% | 77.62% | 100.00%     |\n| pt_sacrecoeur_trevi_tajmahal                | 93.33% | 87.50% | 100.00%     |\n| pt_stpeters_stpauls                         | 61.35% | 79.38% | 50.00%      |\n| amy_gardens                                 | 22.12% | 12.44% | 100.00%     |\n| fbk_vineyard                                | 42.37% | 39.77% | 45.32%      |\n| ETs                                         | 64.94% | 48.08% | 100.00%     |\n| stairs                                      | 0.00%  | 0.00%  | 71.43%      |\n| **Average over all datasets**               | 63.67% | 59.11% | 83.83%      |\n\n### Submissions\nScores by differences in first-stage filtering.\n\n| Method                                                | Private | Public |\n|-------------------------------------------------------|--------:|-------:|\n| **DINOv2 (similarity < 0.15, top-k = 50, #images ≧ 100)** |  **44.69** |  **41.08** |\n| DINOv2 (similarity < 0.3, top-k = 20, #images ≧ 20)   |  42.93 |  39.87 |\n| Exhaustive matching                                 |  41.74 |  36.37 |\n\n# 4. Attempts that did not work\n- Adaptive threshold tuning for DINOv2 was attempted using Bayesian optimization, but each evaluation step was computationally heavy and the parameter space could not converge within the time limit. \n- A binary classifier that identified clusters using edge weights in the closed graph—such as the number of inlier matches and the reprojection error—was trained independently of COLMAP’s camera-model clustering, but precision fell sharply when recall was increased, leaving F1 unimproved, so the idea was abandoned. \n- Experiments that used [MicKey](https://github.com/nianticlabs/mickey) features led to an overly dense LightGlue match graph, which became over-connected; accumulated loop-closure errors caused the final poses to diverge. \n- Applying [LoFTR](https://github.com/zju3dv/LoFTR) to all pairs and then again after filtering greatly increased the number of high-resolution pairs, but incorrect correspondences retained from the coarse pass over-fit and converged incorrectly in scenes with low parallax."
  }
}