{
  "id": 583977,
  "title": "14th place solution",
  "url": "/competitions/image-matching-challenge-2025/discussion/583977",
  "author_name": "Saif daoud",
  "post_date": "2025-06-10T20:40:54.666000",
  "votes": 14,
  "comment_count": 0,
  "views": 0,
  "content": "<p>First, I would like to thank the organizers and Kaggle team for this exciting competition. <br>\nMy teammate <a href=\"https://www.kaggle.com/khlifimohamed\" target=\"_blank\">@khlifimohamed</a> and I are proud to have finished in 14th place. While this was our official ranking, we also have other submissions that weren't selected and could have placed us in the gold medal tier. Regardless of the outcome, the experience has been immensely rewarding.</p>\n<p><strong>1. Overview:</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8034884%2F1bf87faf1fa218693cc53300ae062d38%2FImage_matching_pipeline.png?generation=1749574535884613&amp;alt=media\" alt=\"\"></p>\n<p><strong>2. Global Feature Extraction + Clustering:</strong></p>\n<p>We use <strong>DINOv2</strong> to extract global image embeddings, which are then projected via <strong>t-SNE</strong> for dimensionality reduction. These low-dimensional features are clustered using <strong>HDBSCAN</strong>, a density-based algorithm that groups images based on visual similarity and also detects outliers. To ensure robust clustering, we defined a <strong>confidence score 𝑠</strong> as the harmonic mean between HDBSCAN clustering confidence and the silouhette score. Cluster assignments are accepted only if 𝑠 &gt; 0.6.</p>\n<p><strong>3. Keypoints Detection + Exhaustive Pair Matching:</strong></p>\n<p>This part of our solution is largely based on the 2024 first-place solution, which we found to be highly effective and well-optimized. For local feature extraction and matching, we used <strong>ALIKED</strong> to detect keypoints and <strong>LightGlue</strong> for robust matching. As for their parameters, we leveraged two resolutions: <strong>1280</strong> and <strong>1088</strong> for the full image, and <strong>1280</strong> for cropped regions. The cropping is guided by DBSCAN, which identifies dense regions of matched keypoints in order to focus on relevant areas. To keep runtime practical (~6 hours), we limited the number of keypoints to 8192 per image. For geometric verification, we also adopted the custom RANSAC implementation from last year solution.</p>\n<p><strong>4. Reconstruction:</strong></p>\n<p>We used <strong>PyCOLMAP v3.11.1</strong> for incremental 3D reconstruction. <br>\nTo improve reconstruction stability and reduce randomness, we fixed the first image pair used for initialization. The <strong>first image</strong> is the most matched one, and the <strong>second image</strong> is the one with the most keypoints matched to the first. This initialization consistently increased our reconstruction quality and score. However, fixing the initial pair introduces a risk of divergence in the reconstruction process. To overcome this issue, we implemented a fallback strategy: if PyCOLMAP failed to converge, we iteratively tested alternative second images (e.g., second-best, third-best), and selected the first converging reconstruction.</p>\n<p><strong>4.1. Via Clustrering:</strong></p>\n<p>When the confidence score 𝑠 &gt; 0.6 (as described in Section 2), <br>\nwe leverage the clustering method to guide reconstruction. However, clustering can be error-prone in some scenes. Therefore, we adopted the following approach:</p>\n<ul>\n<li>If s &gt; 0.85 : we run the reconstruction on each predicted cluster.</li>\n<li>If 0.6 &lt; s ≤ 0.85: we use clustering <strong>only for pair initialization</strong>, and we run reconstructions on all the images, followed by a postprocessing to remove the overlap between the obtained reconstructions.</li>\n</ul>\n<p><strong>4.2. Iterative approach:</strong></p>\n<p>This approach is used when the confidence score 𝑠 &lt; 0.6 or if there was some issues when reconstructing via clustering. We run multiple rounds of reconstruction over all images, removing only those whose all paired images have already been successfully registered. The iterations continues until there are no remain images. Then, we apply a postprocessing step to merge reconstructions and eliminate overlaps.</p>\n<p><strong>Our code:</strong> <a href=\"https://github.com/saif-daoud/IMC-2025-14th-place-solution\" target=\"_blank\">https://github.com/saif-daoud/IMC-2025-14th-place-solution</a><br>\n<strong>Kaggle kernel:</strong> <a href=\"https://www.kaggle.com/code/saifdaoud2/pycolmap-imc-2025/notebook?scriptVersionId=243122336\" target=\"_blank\">https://www.kaggle.com/code/saifdaoud2/pycolmap-imc-2025/notebook?scriptVersionId=243122336</a></p>",
  "messages": [
    {
      "id": 3221361,
      "postDate": "2025-06-10T20:40:54.667Z",
      "content": "<p>First, I would like to thank the organizers and Kaggle team for this exciting competition. <br>\nMy teammate <a href=\"https://www.kaggle.com/khlifimohamed\" target=\"_blank\">@khlifimohamed</a> and I are proud to have finished in 14th place. While this was our official ranking, we also have other submissions that weren't selected and could have placed us in the gold medal tier. Regardless of the outcome, the experience has been immensely rewarding.</p>\n<p><strong>1. Overview:</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8034884%2F1bf87faf1fa218693cc53300ae062d38%2FImage_matching_pipeline.png?generation=1749574535884613&amp;alt=media\" alt=\"\"></p>\n<p><strong>2. Global Feature Extraction + Clustering:</strong></p>\n<p>We use <strong>DINOv2</strong> to extract global image embeddings, which are then projected via <strong>t-SNE</strong> for dimensionality reduction. These low-dimensional features are clustered using <strong>HDBSCAN</strong>, a density-based algorithm that groups images based on visual similarity and also detects outliers. To ensure robust clustering, we defined a <strong>confidence score 𝑠</strong> as the harmonic mean between HDBSCAN clustering confidence and the silouhette score. Cluster assignments are accepted only if 𝑠 &gt; 0.6.</p>\n<p><strong>3. Keypoints Detection + Exhaustive Pair Matching:</strong></p>\n<p>This part of our solution is largely based on the 2024 first-place solution, which we found to be highly effective and well-optimized. For local feature extraction and matching, we used <strong>ALIKED</strong> to detect keypoints and <strong>LightGlue</strong> for robust matching. As for their parameters, we leveraged two resolutions: <strong>1280</strong> and <strong>1088</strong> for the full image, and <strong>1280</strong> for cropped regions. The cropping is guided by DBSCAN, which identifies dense regions of matched keypoints in order to focus on relevant areas. To keep runtime practical (~6 hours), we limited the number of keypoints to 8192 per image. For geometric verification, we also adopted the custom RANSAC implementation from last year solution.</p>\n<p><strong>4. Reconstruction:</strong></p>\n<p>We used <strong>PyCOLMAP v3.11.1</strong> for incremental 3D reconstruction. <br>\nTo improve reconstruction stability and reduce randomness, we fixed the first image pair used for initialization. The <strong>first image</strong> is the most matched one, and the <strong>second image</strong> is the one with the most keypoints matched to the first. This initialization consistently increased our reconstruction quality and score. However, fixing the initial pair introduces a risk of divergence in the reconstruction process. To overcome this issue, we implemented a fallback strategy: if PyCOLMAP failed to converge, we iteratively tested alternative second images (e.g., second-best, third-best), and selected the first converging reconstruction.</p>\n<p><strong>4.1. Via Clustrering:</strong></p>\n<p>When the confidence score 𝑠 &gt; 0.6 (as described in Section 2), <br>\nwe leverage the clustering method to guide reconstruction. However, clustering can be error-prone in some scenes. Therefore, we adopted the following approach:</p>\n<ul>\n<li>If s &gt; 0.85 : we run the reconstruction on each predicted cluster.</li>\n<li>If 0.6 &lt; s ≤ 0.85: we use clustering <strong>only for pair initialization</strong>, and we run reconstructions on all the images, followed by a postprocessing to remove the overlap between the obtained reconstructions.</li>\n</ul>\n<p><strong>4.2. Iterative approach:</strong></p>\n<p>This approach is used when the confidence score 𝑠 &lt; 0.6 or if there was some issues when reconstructing via clustering. We run multiple rounds of reconstruction over all images, removing only those whose all paired images have already been successfully registered. The iterations continues until there are no remain images. Then, we apply a postprocessing step to merge reconstructions and eliminate overlaps.</p>\n<p><strong>Our code:</strong> <a href=\"https://github.com/saif-daoud/IMC-2025-14th-place-solution\" target=\"_blank\">https://github.com/saif-daoud/IMC-2025-14th-place-solution</a><br>\n<strong>Kaggle kernel:</strong> <a href=\"https://www.kaggle.com/code/saifdaoud2/pycolmap-imc-2025/notebook?scriptVersionId=243122336\" target=\"_blank\">https://www.kaggle.com/code/saifdaoud2/pycolmap-imc-2025/notebook?scriptVersionId=243122336</a></p>",
      "rawMarkdown": "First, I would like to thank the organizers and Kaggle team for this exciting competition. \nMy teammate @khlifimohamed and I are proud to have finished in 14th place. While this was our official ranking, we also have other submissions that weren't selected and could have placed us in the gold medal tier. Regardless of the outcome, the experience has been immensely rewarding.\n\n\n**1. Overview:**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8034884%2F1bf87faf1fa218693cc53300ae062d38%2FImage_matching_pipeline.png?generation=1749574535884613&alt=media)\n\n\n**2. Global Feature Extraction + Clustering:**\n\nWe use **DINOv2** to extract global image embeddings, which are then projected via **t-SNE** for dimensionality reduction. These low-dimensional features are clustered using **HDBSCAN**, a density-based algorithm that groups images based on visual similarity and also detects outliers. To ensure robust clustering, we defined a **confidence score 𝑠** as the harmonic mean between HDBSCAN clustering confidence and the silouhette score. Cluster assignments are accepted only if 𝑠 > 0.6.\n\n\n**3. Keypoints Detection + Exhaustive Pair Matching:**\n\nThis part of our solution is largely based on the 2024 first-place solution, which we found to be highly effective and well-optimized. For local feature extraction and matching, we used **ALIKED** to detect keypoints and **LightGlue** for robust matching. As for their parameters, we leveraged two resolutions: **1280** and **1088** for the full image, and **1280** for cropped regions. The cropping is guided by DBSCAN, which identifies dense regions of matched keypoints in order to focus on relevant areas. To keep runtime practical (~6 hours), we limited the number of keypoints to 8192 per image. For geometric verification, we also adopted the custom RANSAC implementation from last year solution.\n\n\n**4. Reconstruction:**\n\nWe used **PyCOLMAP v3.11.1** for incremental 3D reconstruction. \nTo improve reconstruction stability and reduce randomness, we fixed the first image pair used for initialization. The **first image** is the most matched one, and the **second image** is the one with the most keypoints matched to the first. This initialization consistently increased our reconstruction quality and score. However, fixing the initial pair introduces a risk of divergence in the reconstruction process. To overcome this issue, we implemented a fallback strategy: if PyCOLMAP failed to converge, we iteratively tested alternative second images (e.g., second-best, third-best), and selected the first converging reconstruction.\n\n\n**4.1. Via Clustrering:**\n\nWhen the confidence score 𝑠 > 0.6 (as described in Section 2), \nwe leverage the clustering method to guide reconstruction. However, clustering can be error-prone in some scenes. Therefore, we adopted the following approach:\n- If s > 0.85 : we run the reconstruction on each predicted cluster.\n- If 0.6 < s ≤ 0.85: we use clustering **only for pair initialization**, and we run reconstructions on all the images, followed by a postprocessing to remove the overlap between the obtained reconstructions.\n\n\n**4.2. Iterative approach:**\n\nThis approach is used when the confidence score 𝑠 < 0.6 or if there was some issues when reconstructing via clustering. We run multiple rounds of reconstruction over all images, removing only those whose all paired images have already been successfully registered. The iterations continues until there are no remain images. Then, we apply a postprocessing step to merge reconstructions and eliminate overlaps.\n\n\n**Our code:** https://github.com/saif-daoud/IMC-2025-14th-place-solution\n**Kaggle kernel:** https://www.kaggle.com/code/saifdaoud2/pycolmap-imc-2025/notebook?scriptVersionId=243122336",
      "votes": 14
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3221361": "First, I would like to thank the organizers and Kaggle team for this exciting competition. \nMy teammate @khlifimohamed and I are proud to have finished in 14th place. While this was our official ranking, we also have other submissions that weren't selected and could have placed us in the gold medal tier. Regardless of the outcome, the experience has been immensely rewarding.\n\n\n**1. Overview:**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8034884%2F1bf87faf1fa218693cc53300ae062d38%2FImage_matching_pipeline.png?generation=1749574535884613&alt=media)\n\n\n**2. Global Feature Extraction + Clustering:**\n\nWe use **DINOv2** to extract global image embeddings, which are then projected via **t-SNE** for dimensionality reduction. These low-dimensional features are clustered using **HDBSCAN**, a density-based algorithm that groups images based on visual similarity and also detects outliers. To ensure robust clustering, we defined a **confidence score 𝑠** as the harmonic mean between HDBSCAN clustering confidence and the silouhette score. Cluster assignments are accepted only if 𝑠 > 0.6.\n\n\n**3. Keypoints Detection + Exhaustive Pair Matching:**\n\nThis part of our solution is largely based on the 2024 first-place solution, which we found to be highly effective and well-optimized. For local feature extraction and matching, we used **ALIKED** to detect keypoints and **LightGlue** for robust matching. As for their parameters, we leveraged two resolutions: **1280** and **1088** for the full image, and **1280** for cropped regions. The cropping is guided by DBSCAN, which identifies dense regions of matched keypoints in order to focus on relevant areas. To keep runtime practical (~6 hours), we limited the number of keypoints to 8192 per image. For geometric verification, we also adopted the custom RANSAC implementation from last year solution.\n\n\n**4. Reconstruction:**\n\nWe used **PyCOLMAP v3.11.1** for incremental 3D reconstruction. \nTo improve reconstruction stability and reduce randomness, we fixed the first image pair used for initialization. The **first image** is the most matched one, and the **second image** is the one with the most keypoints matched to the first. This initialization consistently increased our reconstruction quality and score. However, fixing the initial pair introduces a risk of divergence in the reconstruction process. To overcome this issue, we implemented a fallback strategy: if PyCOLMAP failed to converge, we iteratively tested alternative second images (e.g., second-best, third-best), and selected the first converging reconstruction.\n\n\n**4.1. Via Clustrering:**\n\nWhen the confidence score 𝑠 > 0.6 (as described in Section 2), \nwe leverage the clustering method to guide reconstruction. However, clustering can be error-prone in some scenes. Therefore, we adopted the following approach:\n- If s > 0.85 : we run the reconstruction on each predicted cluster.\n- If 0.6 < s ≤ 0.85: we use clustering **only for pair initialization**, and we run reconstructions on all the images, followed by a postprocessing to remove the overlap between the obtained reconstructions.\n\n\n**4.2. Iterative approach:**\n\nThis approach is used when the confidence score 𝑠 < 0.6 or if there was some issues when reconstructing via clustering. We run multiple rounds of reconstruction over all images, removing only those whose all paired images have already been successfully registered. The iterations continues until there are no remain images. Then, we apply a postprocessing step to merge reconstructions and eliminate overlaps.\n\n\n**Our code:** https://github.com/saif-daoud/IMC-2025-14th-place-solution\n**Kaggle kernel:** https://www.kaggle.com/code/saifdaoud2/pycolmap-imc-2025/notebook?scriptVersionId=243122336"
  }
}