{
  "id": 583184,
  "title": "7th Place Solution",
  "url": "/competitions/image-matching-challenge-2025/writeups/kaqln-7th-place-solution",
  "author_name": "",
  "post_date": "2025-06-05T09:06:43.879209600Z",
  "votes": 25,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Thanks to all the organizers and participants of the Image Matching Challenge. This competition has given us the opportunity to catch up on many 3D vision related studies and to better understand the implementation of COLMAP.</p>\n<p>Our solution is a customized version of <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2024/discussion/510084\" target=\"_blank\">the IMC'24 winning solution</a>. Our new experiments, which were not part of the IMC'24 solution<strong>,</strong> largely ended in failure. Key features are as follows:</p>\n<ul>\n<li>Matching and rotation correction for all image pairs</li>\n<li>Clustering after matching</li>\n</ul>\n<p>The overall pipeline is as follows:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6388%2F25551dac1337fe90cfbc5233fba0b314%2Foverview.png?generation=1749113468303722&amp;alt=media\" alt=\"\"></p>\n<h2>Clustering using similarity graph</h2>\n<p>To extract clusters from the dataset, we created a similarity graph based on inlier counts. Edges below a threshold were removed from this graph, and the resulting connected components were treated as clusters. Isolated points and small connected components were considered outliers.</p>\n<h2>Failed Attempts</h2>\n<ul>\n<li>I considered adopting MASt3R-SfM. However, when I ran and evaluated it on ETs and stairs in my local env, my experiments did not yield the expected improvements. In these experiments, I only used MASt3R's semi-dense matches and evaluated the camera poses output by MASt3R's fastNN after converting them to the competition format. While the visualized results initially appeared to estimate camera poses accurately, they didn't seem to achieve the accuracy required to surpass the evaluation metric threshold for this competition.</li>\n<li>VGGSfM and VGGT require a significant amount of VRAM. Therefore, I adopted an approach that only uses track refinement, as the 2nd place solution of IMC'24 employed. Due to changes in the <code>pycolmap</code> API, I had to modify a lot of code. It's possible there were bugs, but at least in my experiments, I couldn't achieve any improvements.</li>\n</ul>",
  "messages": [
    {
      "id": "3217656",
      "postDate": "06/05/2025 09:06:43",
      "content": "<p>Thanks to all the organizers and participants of the Image Matching Challenge. This competition has given us the opportunity to catch up on many 3D vision related studies and to better understand the implementation of COLMAP.</p>\n<p>Our solution is a customized version of <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2024/discussion/510084\" target=\"_blank\">the IMC'24 winning solution</a>. Our new experiments, which were not part of the IMC'24 solution<strong>,</strong> largely ended in failure. Key features are as follows:</p>\n<ul>\n<li>Matching and rotation correction for all image pairs</li>\n<li>Clustering after matching</li>\n</ul>\n<p>The overall pipeline is as follows:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6388%2F25551dac1337fe90cfbc5233fba0b314%2Foverview.png?generation=1749113468303722&amp;alt=media\" alt=\"\"></p>\n<h2>Clustering using similarity graph</h2>\n<p>To extract clusters from the dataset, we created a similarity graph based on inlier counts. Edges below a threshold were removed from this graph, and the resulting connected components were treated as clusters. Isolated points and small connected components were considered outliers.</p>\n<h2>Failed Attempts</h2>\n<ul>\n<li>I considered adopting MASt3R-SfM. However, when I ran and evaluated it on ETs and stairs in my local env, my experiments did not yield the expected improvements. In these experiments, I only used MASt3R's semi-dense matches and evaluated the camera poses output by MASt3R's fastNN after converting them to the competition format. While the visualized results initially appeared to estimate camera poses accurately, they didn't seem to achieve the accuracy required to surpass the evaluation metric threshold for this competition.</li>\n<li>VGGSfM and VGGT require a significant amount of VRAM. Therefore, I adopted an approach that only uses track refinement, as the 2nd place solution of IMC'24 employed. Due to changes in the <code>pycolmap</code> API, I had to modify a lot of code. It's possible there were bugs, but at least in my experiments, I couldn't achieve any improvements.</li>\n</ul>",
      "rawMarkdown": "Thanks to all the organizers and participants of the Image Matching Challenge. This competition has given us the opportunity to catch up on many 3D vision related studies and to better understand the implementation of COLMAP.\n\nOur solution is a customized version of [the IMC'24 winning solution](https://www.kaggle.com/competitions/image-matching-challenge-2024/discussion/510084). Our new experiments, which were not part of the IMC'24 solution**,** largely ended in failure. Key features are as follows:\n\n- Matching and rotation correction for all image pairs\n- Clustering after matching\n\nThe overall pipeline is as follows:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6388%2F25551dac1337fe90cfbc5233fba0b314%2Foverview.png?generation=1749113468303722&alt=media)\n\n## Clustering using similarity graph\n\nTo extract clusters from the dataset, we created a similarity graph based on inlier counts. Edges below a threshold were removed from this graph, and the resulting connected components were treated as clusters. Isolated points and small connected components were considered outliers.\n\n## Failed Attempts\n\n- I considered adopting MASt3R-SfM. However, when I ran and evaluated it on ETs and stairs in my local env, my experiments did not yield the expected improvements. In these experiments, I only used MASt3R's semi-dense matches and evaluated the camera poses output by MASt3R's fastNN after converting them to the competition format. While the visualized results initially appeared to estimate camera poses accurately, they didn't seem to achieve the accuracy required to surpass the evaluation metric threshold for this competition.\n- VGGSfM and VGGT require a significant amount of VRAM. Therefore, I adopted an approach that only uses track refinement, as the 2nd place solution of IMC'24 employed. Due to changes in the `pycolmap` API, I had to modify a lot of code. It's possible there were bugs, but at least in my experiments, I couldn't achieve any improvements.",
      "votes": null
    },
    {
      "id": "3217742",
      "postDate": "06/05/2025 11:18:53",
      "content": "<p>Congratulations!<br>\nWere there any parameters that had a significant impact on the algorithm?<br>\nFor example, the number of keypoints output by aliked, the threshold, colmap’s min_model_size, and so on.</p>",
      "rawMarkdown": "Congratulations!\nWere there any parameters that had a significant impact on the algorithm?\nFor example, the number of keypoints output by aliked, the threshold, colmap’s min_model_size, and so on.",
      "votes": null
    },
    {
      "id": "3217816",
      "postDate": "06/05/2025 12:33:37",
      "content": "<p>for colmap :<br>\nmin_model_size : 3 , max_num_models : 5  <br>\nand the other params can be found here </p>",
      "rawMarkdown": "for colmap :\nmin_model_size : 3 , max_num_models : 5  \nand the other params can be found here",
      "votes": null
    },
    {
      "id": "3217845",
      "postDate": "06/05/2025 12:56:51",
      "content": "<p>Most parameters were not tuned from the IMC'24 configurations.<br>\nChanges to the image size had the most significant impact.<br>\nWhile one might easily expect performance to improve with larger image sizes, I observed that the performance gains for the image size used by ALIKED in the earlier part of the pipeline converged more quickly.<br>\nConsequently, in my final configuration, the image size for ALIKED, which takes cropped images as input, is the largest at 1792.</p>",
      "rawMarkdown": "Most parameters were not tuned from the IMC'24 configurations.\nChanges to the image size had the most significant impact.\nWhile one might easily expect performance to improve with larger image sizes, I observed that the performance gains for the image size used by ALIKED in the earlier part of the pipeline converged more quickly.\nConsequently, in my final configuration, the image size for ALIKED, which takes cropped images as input, is the largest at 1792.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3217742,
      "author_name": "sugupoko",
      "author_url": "",
      "post_date": "06/05/2025 11:18:53",
      "content": "<p>Congratulations!<br>\nWere there any parameters that had a significant impact on the algorithm?<br>\nFor example, the number of keypoints output by aliked, the threshold, colmap’s min_model_size, and so on.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3217816,
          "author_name": "justforfun44",
          "author_url": "",
          "post_date": "06/05/2025 12:33:37",
          "content": "<p>for colmap :<br>\nmin_model_size : 3 , max_num_models : 5  <br>\nand the other params can be found here </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3217845,
          "author_name": "confirm",
          "author_url": "",
          "post_date": "06/05/2025 12:56:51",
          "content": "<p>Most parameters were not tuned from the IMC'24 configurations.<br>\nChanges to the image size had the most significant impact.<br>\nWhile one might easily expect performance to improve with larger image sizes, I observed that the performance gains for the image size used by ALIKED in the earlier part of the pipeline converged more quickly.<br>\nConsequently, in my final configuration, the image size for ALIKED, which takes cropped images as input, is the largest at 1792.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3217656": "Thanks to all the organizers and participants of the Image Matching Challenge. This competition has given us the opportunity to catch up on many 3D vision related studies and to better understand the implementation of COLMAP.\n\nOur solution is a customized version of [the IMC'24 winning solution](https://www.kaggle.com/competitions/image-matching-challenge-2024/discussion/510084). Our new experiments, which were not part of the IMC'24 solution**,** largely ended in failure. Key features are as follows:\n\n- Matching and rotation correction for all image pairs\n- Clustering after matching\n\nThe overall pipeline is as follows:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6388%2F25551dac1337fe90cfbc5233fba0b314%2Foverview.png?generation=1749113468303722&alt=media)\n\n## Clustering using similarity graph\n\nTo extract clusters from the dataset, we created a similarity graph based on inlier counts. Edges below a threshold were removed from this graph, and the resulting connected components were treated as clusters. Isolated points and small connected components were considered outliers.\n\n## Failed Attempts\n\n- I considered adopting MASt3R-SfM. However, when I ran and evaluated it on ETs and stairs in my local env, my experiments did not yield the expected improvements. In these experiments, I only used MASt3R's semi-dense matches and evaluated the camera poses output by MASt3R's fastNN after converting them to the competition format. While the visualized results initially appeared to estimate camera poses accurately, they didn't seem to achieve the accuracy required to surpass the evaluation metric threshold for this competition.\n- VGGSfM and VGGT require a significant amount of VRAM. Therefore, I adopted an approach that only uses track refinement, as the 2nd place solution of IMC'24 employed. Due to changes in the `pycolmap` API, I had to modify a lot of code. It's possible there were bugs, but at least in my experiments, I couldn't achieve any improvements.",
    "3217742": "Congratulations!\nWere there any parameters that had a significant impact on the algorithm?\nFor example, the number of keypoints output by aliked, the threshold, colmap’s min_model_size, and so on.",
    "3217816": "for colmap :\nmin_model_size : 3 , max_num_models : 5  \nand the other params can be found here",
    "3217845": "Most parameters were not tuned from the IMC'24 configurations.\nChanges to the image size had the most significant impact.\nWhile one might easily expect performance to improve with larger image sizes, I observed that the performance gains for the image size used by ALIKED in the earlier part of the pipeline converged more quickly.\nConsequently, in my final configuration, the image size for ALIKED, which takes cropped images as input, is the largest at 1792."
  },
  "source": "meta"
}