{
  "id": 583185,
  "title": "12th Place Solution - Pushing Limits of COLMAP",
  "url": "/competitions/image-matching-challenge-2025/discussion/583185",
  "author_name": "Igor Lashkov",
  "post_date": "2025-06-05T09:08:09.967000",
  "votes": 26,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Congratulations to all the participants – <strong>nearly 1,000 teams</strong> this year again! <br><br>\nThe unexpected changes in the competition due to the metric bug caused considerable confusion and disrupted our plans at a critical moment. <br><br>\nNevertheless, …, we'd like to share a small piece of our contribution with everyone.</p>\n<h2>Overview</h2>\n<p>We simultaneously developed <strong>two distinct submission pipelines</strong> using our previous <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2024/discussion/510084\" target=\"_blank\">IMC2024 winner solution</a>. The first approach relied solely on external matchers and COLMAP for reconstruction. In the second approach — aligned with the strategies of top-ranking teams — we utilized a clustering method based on DINOv2, followed by scene reconstruction using COLMAP.</p>\n<h2>Pipeline 1. COLMAP only</h2>\n<p>For image pair selection, we decided not to use any image retrieval method. Instead, we relied exclusively on matches generated by ALIKED/LightGlue (with image resolutions of 1280 and 2048) and DISK (resolution 1536), along with the clustering capabilities provided by COLMAP. Specifically, we applied a threshold of 30  per individual detector and 80 matches across all detectors in the ensemble for a given image. Similar to previous IMCs, to get more accurate matches we performed rotation of the images to the natural orientation. The two-view geometry was computed from image point correspondences using RANSAC, rather than COLMAP's internal implementation. COLMAP is capable of generating multiple reconstruction models, each representing a distinct cluster.</p>\n<p>Changes to incremental mapping of COLMAP:</p>\n<pre><code> = {\n    : 3,\n    : 10,\n}\n</code></pre>\n<p>In this case, there is no need to provide an image to initialize the reconstruction.</p>\n<h3>Key Takeaways:</h3>\n<ul>\n<li><strong>Gradually filter unique image pairs</strong> within the scene based on the number of matches</li>\n<li><strong>ALIKED+LightGlue finetuned settings</strong> with unlimited N of keypoints produced by ALIKED n16 and the LG parameters providing accurate results.</li>\n<li><strong>Cache keypoints &amp; descriptors</strong>, generated by ALIKED for each image. It helped to reduce the running time a lot.</li>\n<li><strong>Multi-GPU acceleration. Mixed precision with GPU T4x2</strong> hardware helps to reduce the time of image matching stage significantly. We employ both GPUs to perform SfM in parallel.</li>\n<li><strong>Matches TTA.</strong> An ensemble of matches extracted from high-resolution images of different scales. In LB experiments, the best results were obtained with a combination of 1280, 1536, and 2048 for both original and cropped images.</li>\n</ul>\n<h2>Pipeline 2. Clustering</h2>\n<p>We experimented with clustering the dataset images in advance using global feature descriptors.<br>\nCompared to the approach of matching all images and clustering based on COLMAP’s output (which we ended up using for our submission), this method achieves nearly the same score in about 3/4 of the processing time.</p>\n<h3>Overview of Processing Steps</h3>\n<ol>\n<li><p>Feature extraction and construction of a rotation-robust similarity matrix  </p>\n<ul>\n<li><p>Obtain an embedding vector for each image using a feature-extraction model such as DINOv2.</p></li>\n<li><p>Extract features for four rotated versions of each image (0°, 90°, 180°, 270°) and compute the pair-wise similarities across these orientations.</p></li>\n<li><p>For every image pair, keep the highest similarity score and its corresponding rotation, yielding a rotation-robust similarity matrix.</p></li></ul></li>\n<li><p>1st stage clustering (coarse clustering)  </p>\n<ul>\n<li><p>Using the similarity matrix, find the top N nearest neighbors for each image and connect pairs whose similarity exceeds a chosen threshold.</p></li>\n<li><p>Treat the resulting adjacency matrix as a graph and extract connected components to form the initial clusters.</p></li></ul></li>\n<li><p>2nd Stage Cluster Merging (Small-Cluster Consolidation)</p>\n<ul>\n<li><p>For clusters that remain small, search for a merge target by counting the number of matched image pairs with other clusters.</p></li>\n<li><p>ALIKED+LG matching for the top K pairs with high similarity between clusters, merging clusters with a certain number of matches.</p></li></ul></li>\n<li><p>Outlier Handling</p>\n<ul>\n<li>Clusters whose size is still extremely small after merging are re-assigned to an “outliers” cluster.</li></ul></li>\n</ol>\n<h3>Some insights</h3>\n<ul>\n<li>We experimented with models like <strong>DINOv2</strong> and <strong>SigLIP-v2</strong>, but <strong>MegaLoc</strong> showed the best performance on the private test set as a feature extractor.</li>\n<li>Graph-based clustering methods were more effective than HDBSCAN in terms of both CV and LB scores. Our hypothesis is that even within the same scene, the presence of many “chain-like” or “bridging” images reduces high-density regions. As a result, density-based clustering methods like HDBSCAN may not be well-suited for this type of data.</li>\n</ul>\n<h3>Score &amp; Processing Time</h3>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>Public LB</th>\n<th>Private LB</th>\n<th>Processing Time</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>COLMAP base clustering</td>\n<td>41.86</td>\n<td>44.56</td>\n<td>8h45m</td>\n</tr>\n<tr>\n<td>pre-clustering by MegaLoc</td>\n<td>40.99</td>\n<td>44.28</td>\n<td>6h40m</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>All of the results were obtained using the same matching setup: ALIKED1280 + ALIKED2048 + DISK1536, which showed the best performance on the private test set.</li>\n</ul>\n<h2>Ideas that did not work out:</h2>\n<ul>\n<li><strong>TTA multi-crop</strong> for image pair filtering.</li>\n<li><strong>Different detectors, matchers.</strong> We tested SP/SG, LoFTR, OmniGlue, GIM. Eventually, ALIKED, LightGlue and DISK were the best LB choice for us.</li>\n<li><strong>VGGT.</strong> Ran out of memory.</li>\n<li><strong>Newest COLMAP version 3.11.1.</strong> Compared to 0.6.1, the score dropped.</li>\n<li><strong>GLOMAP</strong>: Compared to COLMAP, GLOMAP raised a time out exception surprisingly.</li>\n</ul>\n<p>In conclusion, the <strong>solely use of COLMAP</strong> in your approach can bring you remarkably close to a gold medal. <br><br>\nSee you in the next competitions!</p>",
  "messages": [
    {
      "id": 3217657,
      "postDate": "2025-06-05T09:08:09.967Z",
      "content": "<p>Congratulations to all the participants – <strong>nearly 1,000 teams</strong> this year again! <br><br>\nThe unexpected changes in the competition due to the metric bug caused considerable confusion and disrupted our plans at a critical moment. <br><br>\nNevertheless, …, we'd like to share a small piece of our contribution with everyone.</p>\n<h2>Overview</h2>\n<p>We simultaneously developed <strong>two distinct submission pipelines</strong> using our previous <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2024/discussion/510084\" target=\"_blank\">IMC2024 winner solution</a>. The first approach relied solely on external matchers and COLMAP for reconstruction. In the second approach — aligned with the strategies of top-ranking teams — we utilized a clustering method based on DINOv2, followed by scene reconstruction using COLMAP.</p>\n<h2>Pipeline 1. COLMAP only</h2>\n<p>For image pair selection, we decided not to use any image retrieval method. Instead, we relied exclusively on matches generated by ALIKED/LightGlue (with image resolutions of 1280 and 2048) and DISK (resolution 1536), along with the clustering capabilities provided by COLMAP. Specifically, we applied a threshold of 30  per individual detector and 80 matches across all detectors in the ensemble for a given image. Similar to previous IMCs, to get more accurate matches we performed rotation of the images to the natural orientation. The two-view geometry was computed from image point correspondences using RANSAC, rather than COLMAP's internal implementation. COLMAP is capable of generating multiple reconstruction models, each representing a distinct cluster.</p>\n<p>Changes to incremental mapping of COLMAP:</p>\n<pre><code> = {\n    : 3,\n    : 10,\n}\n</code></pre>\n<p>In this case, there is no need to provide an image to initialize the reconstruction.</p>\n<h3>Key Takeaways:</h3>\n<ul>\n<li><strong>Gradually filter unique image pairs</strong> within the scene based on the number of matches</li>\n<li><strong>ALIKED+LightGlue finetuned settings</strong> with unlimited N of keypoints produced by ALIKED n16 and the LG parameters providing accurate results.</li>\n<li><strong>Cache keypoints &amp; descriptors</strong>, generated by ALIKED for each image. It helped to reduce the running time a lot.</li>\n<li><strong>Multi-GPU acceleration. Mixed precision with GPU T4x2</strong> hardware helps to reduce the time of image matching stage significantly. We employ both GPUs to perform SfM in parallel.</li>\n<li><strong>Matches TTA.</strong> An ensemble of matches extracted from high-resolution images of different scales. In LB experiments, the best results were obtained with a combination of 1280, 1536, and 2048 for both original and cropped images.</li>\n</ul>\n<h2>Pipeline 2. Clustering</h2>\n<p>We experimented with clustering the dataset images in advance using global feature descriptors.<br>\nCompared to the approach of matching all images and clustering based on COLMAP’s output (which we ended up using for our submission), this method achieves nearly the same score in about 3/4 of the processing time.</p>\n<h3>Overview of Processing Steps</h3>\n<ol>\n<li><p>Feature extraction and construction of a rotation-robust similarity matrix  </p>\n<ul>\n<li><p>Obtain an embedding vector for each image using a feature-extraction model such as DINOv2.</p></li>\n<li><p>Extract features for four rotated versions of each image (0°, 90°, 180°, 270°) and compute the pair-wise similarities across these orientations.</p></li>\n<li><p>For every image pair, keep the highest similarity score and its corresponding rotation, yielding a rotation-robust similarity matrix.</p></li></ul></li>\n<li><p>1st stage clustering (coarse clustering)  </p>\n<ul>\n<li><p>Using the similarity matrix, find the top N nearest neighbors for each image and connect pairs whose similarity exceeds a chosen threshold.</p></li>\n<li><p>Treat the resulting adjacency matrix as a graph and extract connected components to form the initial clusters.</p></li></ul></li>\n<li><p>2nd Stage Cluster Merging (Small-Cluster Consolidation)</p>\n<ul>\n<li><p>For clusters that remain small, search for a merge target by counting the number of matched image pairs with other clusters.</p></li>\n<li><p>ALIKED+LG matching for the top K pairs with high similarity between clusters, merging clusters with a certain number of matches.</p></li></ul></li>\n<li><p>Outlier Handling</p>\n<ul>\n<li>Clusters whose size is still extremely small after merging are re-assigned to an “outliers” cluster.</li></ul></li>\n</ol>\n<h3>Some insights</h3>\n<ul>\n<li>We experimented with models like <strong>DINOv2</strong> and <strong>SigLIP-v2</strong>, but <strong>MegaLoc</strong> showed the best performance on the private test set as a feature extractor.</li>\n<li>Graph-based clustering methods were more effective than HDBSCAN in terms of both CV and LB scores. Our hypothesis is that even within the same scene, the presence of many “chain-like” or “bridging” images reduces high-density regions. As a result, density-based clustering methods like HDBSCAN may not be well-suited for this type of data.</li>\n</ul>\n<h3>Score &amp; Processing Time</h3>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>Public LB</th>\n<th>Private LB</th>\n<th>Processing Time</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>COLMAP base clustering</td>\n<td>41.86</td>\n<td>44.56</td>\n<td>8h45m</td>\n</tr>\n<tr>\n<td>pre-clustering by MegaLoc</td>\n<td>40.99</td>\n<td>44.28</td>\n<td>6h40m</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>All of the results were obtained using the same matching setup: ALIKED1280 + ALIKED2048 + DISK1536, which showed the best performance on the private test set.</li>\n</ul>\n<h2>Ideas that did not work out:</h2>\n<ul>\n<li><strong>TTA multi-crop</strong> for image pair filtering.</li>\n<li><strong>Different detectors, matchers.</strong> We tested SP/SG, LoFTR, OmniGlue, GIM. Eventually, ALIKED, LightGlue and DISK were the best LB choice for us.</li>\n<li><strong>VGGT.</strong> Ran out of memory.</li>\n<li><strong>Newest COLMAP version 3.11.1.</strong> Compared to 0.6.1, the score dropped.</li>\n<li><strong>GLOMAP</strong>: Compared to COLMAP, GLOMAP raised a time out exception surprisingly.</li>\n</ul>\n<p>In conclusion, the <strong>solely use of COLMAP</strong> in your approach can bring you remarkably close to a gold medal. <br><br>\nSee you in the next competitions!</p>",
      "rawMarkdown": "Congratulations to all the participants – **nearly 1,000 teams** this year again! <br />\nThe unexpected changes in the competition due to the metric bug caused considerable confusion and disrupted our plans at a critical moment. <br />\nNevertheless, ..., we'd like to share a small piece of our contribution with everyone.\n\n## Overview\n\nWe simultaneously developed **two distinct submission pipelines** using our previous [IMC2024 winner solution](https://www.kaggle.com/competitions/image-matching-challenge-2024/discussion/510084). The first approach relied solely on external matchers and COLMAP for reconstruction. In the second approach — aligned with the strategies of top-ranking teams — we utilized a clustering method based on DINOv2, followed by scene reconstruction using COLMAP.\n\n## Pipeline 1. COLMAP only\nFor image pair selection, we decided not to use any image retrieval method. Instead, we relied exclusively on matches generated by ALIKED/LightGlue (with image resolutions of 1280 and 2048) and DISK (resolution 1536), along with the clustering capabilities provided by COLMAP. Specifically, we applied a threshold of 30  per individual detector and 80 matches across all detectors in the ensemble for a given image. Similar to previous IMCs, to get more accurate matches we performed rotation of the images to the natural orientation. The two-view geometry was computed from image point correspondences using RANSAC, rather than COLMAP's internal implementation. COLMAP is capable of generating multiple reconstruction models, each representing a distinct cluster.\n\nChanges to incremental mapping of COLMAP:\n```\ncolmap_mapper_options = {\n    \"min_model_size\": 3,\n    \"max_num_models\": 10,\n}\n```\nIn this case, there is no need to provide an image to initialize the reconstruction.\n\n### Key Takeaways:\n- **Gradually filter unique image pairs** within the scene based on the number of matches\n- **ALIKED+LightGlue finetuned settings** with unlimited N of keypoints produced by ALIKED n16 and the LG parameters providing accurate results.\n- **Cache keypoints & descriptors**, generated by ALIKED for each image. It helped to reduce the running time a lot.\n- **Multi-GPU acceleration. Mixed precision with GPU T4x2** hardware helps to reduce the time of image matching stage significantly. We employ both GPUs to perform SfM in parallel.\n- **Matches TTA.** An ensemble of matches extracted from high-resolution images of different scales. In LB experiments, the best results were obtained with a combination of 1280, 1536, and 2048 for both original and cropped images.\n\n\n## Pipeline 2. Clustering\nWe experimented with clustering the dataset images in advance using global feature descriptors.\nCompared to the approach of matching all images and clustering based on COLMAP’s output (which we ended up using for our submission), this method achieves nearly the same score in about 3/4 of the processing time.\n\n### Overview of Processing Steps\n1. Feature extraction and construction of a rotation-robust similarity matrix  \n    - Obtain an embedding vector for each image using a feature-extraction model such as DINOv2.\n\n    - Extract features for four rotated versions of each image (0°, 90°, 180°, 270°) and compute the pair-wise similarities across these orientations.\n\n    - For every image pair, keep the highest similarity score and its corresponding rotation, yielding a rotation-robust similarity matrix.\n\n2. 1st stage clustering (coarse clustering)  \n    - Using the similarity matrix, find the top N nearest neighbors for each image and connect pairs whose similarity exceeds a chosen threshold.\n\n    - Treat the resulting adjacency matrix as a graph and extract connected components to form the initial clusters.\n\n3. 2nd Stage Cluster Merging (Small-Cluster Consolidation)\n    - For clusters that remain small, search for a merge target by counting the number of matched image pairs with other clusters.\n\n    - ALIKED+LG matching for the top K pairs with high similarity between clusters, merging clusters with a certain number of matches.\n\n4. Outlier Handling\n    - Clusters whose size is still extremely small after merging are re-assigned to an “outliers” cluster.\n\n### Some insights\n- We experimented with models like **DINOv2** and **SigLIP-v2**, but **MegaLoc** showed the best performance on the private test set as a feature extractor.\n- Graph-based clustering methods were more effective than HDBSCAN in terms of both CV and LB scores. Our hypothesis is that even within the same scene, the presence of many “chain-like” or “bridging” images reduces high-density regions. As a result, density-based clustering methods like HDBSCAN may not be well-suited for this type of data.\n\n### Score & Processing Time\n|                           | Public LB | Private LB | Processing Time | \n| ------------------------- | --------- | ---------- | --------------- | \n| COLMAP base clustering    | 41.86     | 44.56      | 8h45m           | \n| pre-clustering by MegaLoc | 40.99     | 44.28      | 6h40m           | \n\n- All of the results were obtained using the same matching setup: ALIKED1280 + ALIKED2048 + DISK1536, which showed the best performance on the private test set.\n\n\n## Ideas that did not work out:\n- **TTA multi-crop** for image pair filtering.\n- **Different detectors, matchers.** We tested SP/SG, LoFTR, OmniGlue, GIM. Eventually, ALIKED, LightGlue and DISK were the best LB choice for us.\n- **VGGT.** Ran out of memory.\n- **Newest COLMAP version 3.11.1.** Compared to 0.6.1, the score dropped.\n- **GLOMAP**: Compared to COLMAP, GLOMAP raised a time out exception surprisingly.\n\nIn conclusion, the **solely use of COLMAP** in your approach can bring you remarkably close to a gold medal. <br />\nSee you in the next competitions!",
      "votes": 26
    },
    {
      "id": 3218404,
      "postDate": "2025-06-06T06:24:46.150Z",
      "content": "<p>Great writeup! I'm sorry to see you missed the gold</p>",
      "rawMarkdown": "Great writeup! I'm sorry to see you missed the gold"
    },
    {
      "id": 3223159,
      "postDate": "2025-06-12T23:59:52.653Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3218404,
      "author_name": "Zacchaeus",
      "author_url": "",
      "post_date": "2025-06-06T06:24:46.150000",
      "content": "<p>Great writeup! I'm sorry to see you missed the gold</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3223159,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-06-12T23:59:52.653000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3217657": "Congratulations to all the participants – **nearly 1,000 teams** this year again! <br />\nThe unexpected changes in the competition due to the metric bug caused considerable confusion and disrupted our plans at a critical moment. <br />\nNevertheless, ..., we'd like to share a small piece of our contribution with everyone.\n\n## Overview\n\nWe simultaneously developed **two distinct submission pipelines** using our previous [IMC2024 winner solution](https://www.kaggle.com/competitions/image-matching-challenge-2024/discussion/510084). The first approach relied solely on external matchers and COLMAP for reconstruction. In the second approach — aligned with the strategies of top-ranking teams — we utilized a clustering method based on DINOv2, followed by scene reconstruction using COLMAP.\n\n## Pipeline 1. COLMAP only\nFor image pair selection, we decided not to use any image retrieval method. Instead, we relied exclusively on matches generated by ALIKED/LightGlue (with image resolutions of 1280 and 2048) and DISK (resolution 1536), along with the clustering capabilities provided by COLMAP. Specifically, we applied a threshold of 30  per individual detector and 80 matches across all detectors in the ensemble for a given image. Similar to previous IMCs, to get more accurate matches we performed rotation of the images to the natural orientation. The two-view geometry was computed from image point correspondences using RANSAC, rather than COLMAP's internal implementation. COLMAP is capable of generating multiple reconstruction models, each representing a distinct cluster.\n\nChanges to incremental mapping of COLMAP:\n```\ncolmap_mapper_options = {\n    \"min_model_size\": 3,\n    \"max_num_models\": 10,\n}\n```\nIn this case, there is no need to provide an image to initialize the reconstruction.\n\n### Key Takeaways:\n- **Gradually filter unique image pairs** within the scene based on the number of matches\n- **ALIKED+LightGlue finetuned settings** with unlimited N of keypoints produced by ALIKED n16 and the LG parameters providing accurate results.\n- **Cache keypoints & descriptors**, generated by ALIKED for each image. It helped to reduce the running time a lot.\n- **Multi-GPU acceleration. Mixed precision with GPU T4x2** hardware helps to reduce the time of image matching stage significantly. We employ both GPUs to perform SfM in parallel.\n- **Matches TTA.** An ensemble of matches extracted from high-resolution images of different scales. In LB experiments, the best results were obtained with a combination of 1280, 1536, and 2048 for both original and cropped images.\n\n\n## Pipeline 2. Clustering\nWe experimented with clustering the dataset images in advance using global feature descriptors.\nCompared to the approach of matching all images and clustering based on COLMAP’s output (which we ended up using for our submission), this method achieves nearly the same score in about 3/4 of the processing time.\n\n### Overview of Processing Steps\n1. Feature extraction and construction of a rotation-robust similarity matrix  \n    - Obtain an embedding vector for each image using a feature-extraction model such as DINOv2.\n\n    - Extract features for four rotated versions of each image (0°, 90°, 180°, 270°) and compute the pair-wise similarities across these orientations.\n\n    - For every image pair, keep the highest similarity score and its corresponding rotation, yielding a rotation-robust similarity matrix.\n\n2. 1st stage clustering (coarse clustering)  \n    - Using the similarity matrix, find the top N nearest neighbors for each image and connect pairs whose similarity exceeds a chosen threshold.\n\n    - Treat the resulting adjacency matrix as a graph and extract connected components to form the initial clusters.\n\n3. 2nd Stage Cluster Merging (Small-Cluster Consolidation)\n    - For clusters that remain small, search for a merge target by counting the number of matched image pairs with other clusters.\n\n    - ALIKED+LG matching for the top K pairs with high similarity between clusters, merging clusters with a certain number of matches.\n\n4. Outlier Handling\n    - Clusters whose size is still extremely small after merging are re-assigned to an “outliers” cluster.\n\n### Some insights\n- We experimented with models like **DINOv2** and **SigLIP-v2**, but **MegaLoc** showed the best performance on the private test set as a feature extractor.\n- Graph-based clustering methods were more effective than HDBSCAN in terms of both CV and LB scores. Our hypothesis is that even within the same scene, the presence of many “chain-like” or “bridging” images reduces high-density regions. As a result, density-based clustering methods like HDBSCAN may not be well-suited for this type of data.\n\n### Score & Processing Time\n|                           | Public LB | Private LB | Processing Time | \n| ------------------------- | --------- | ---------- | --------------- | \n| COLMAP base clustering    | 41.86     | 44.56      | 8h45m           | \n| pre-clustering by MegaLoc | 40.99     | 44.28      | 6h40m           | \n\n- All of the results were obtained using the same matching setup: ALIKED1280 + ALIKED2048 + DISK1536, which showed the best performance on the private test set.\n\n\n## Ideas that did not work out:\n- **TTA multi-crop** for image pair filtering.\n- **Different detectors, matchers.** We tested SP/SG, LoFTR, OmniGlue, GIM. Eventually, ALIKED, LightGlue and DISK were the best LB choice for us.\n- **VGGT.** Ran out of memory.\n- **Newest COLMAP version 3.11.1.** Compared to 0.6.1, the score dropped.\n- **GLOMAP**: Compared to COLMAP, GLOMAP raised a time out exception surprisingly.\n\nIn conclusion, the **solely use of COLMAP** in your approach can bring you remarkably close to a gold medal. <br />\nSee you in the next competitions!",
    "3218404": "Great writeup! I'm sorry to see you missed the gold",
    "3223159": ""
  }
}