{
  "id": 583711,
  "title": "5th Place Solution : MASt3R is All You Need",
  "url": "/competitions/image-matching-challenge-2025/writeups/sayan-paul-5th-place-solution-mast3r-is-all-you-ne",
  "author_name": "",
  "post_date": "2025-06-08T18:49:26.433970Z",
  "votes": 18,
  "comment_count": 1,
  "views": 0,
  "content": "<p>First and foremost, I'd like to thank the organizers and Kaggle team for hosting such an exciting and challenging competition. Congratulations to all the top participants! While I’m not new to the field of 3D vision, this was my first time properly participating in the Image Matching Challenge. One of the main challenges I faced was joining the competition quite late - around two weeks before the deadline. As a result, a significant portion of my time went into setting up Kaggle notebooks, leaving me with limited opportunities to iterate and experiment with different approaches and their variants. That said, let's dive into the simple yet effective solution that worked for me.</p>\n<h4>Overview</h4>\n<p>From recent personal experiments, I’ve observed that foundation models like DUSt3R, MASt3R, and VGGT offer superior matching performance compared to traditional detector-descriptor-matcher pipelines (e.g., ALIKED or SuperPoint + LightGlue). Based on this, I built a straightforward pipeline using MASt3R and evaluated it on the IMC-2025 dataset. As expected, it performed well in terms of matching accuracy. However, a major challenge was the high computational cost of the MASt3R matcher, which led to frequent notebook timeouts during inference. I addressed this by implementing an efficient image-pair shortlisting strategy, tuning its hyperparameters, and applying a few engineering optimizations to reduce runtime without significantly impacting performance.<br>\nThe pipeline that gave the best result on Public LB :-</p>\n<ol>\n<li>Image-Pair Similarity based Shortlisting using <a href=\"https://arxiv.org/abs/2409.19152\" target=\"_blank\">MASt3R-ASMK</a></li>\n<li><a href=\"https://arxiv.org/abs/2406.09756\" target=\"_blank\">MASt3R</a> Semi-Dense Matching on the shortlisted image pairs</li>\n<li>COLMAP based Verification, Reconstruction and Clustering<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2474539%2F21b88c6ebf846a383c84751f92005605%2Fpipeline_overview.png?generation=1749406349577280&amp;alt=media\" alt=\"PipelIne Overview Block Diagram\"></li>\n</ol>\n<h4>Solution Details</h4>\n<h5>1. Image-Pair Similarity based Shortlisting</h5>\n<p> I tried 3 methods to compute a similarity matrix between all the images of a dataset in a fast way and shortlist the pairs using certain thresholds, for the more expensive MASt3R Matcher to run on them. </p>\n<p>(i) <a href=\"https://arxiv.org/abs/2304.07193\" target=\"_blank\"><strong>DINO-v2</strong></a> : Similarity matrix was computed by normalizing the (1 - distance_matrix), where L2-distance was used.<br>\n(ii) <a href=\"https://arxiv.org/abs/2409.19152\" target=\"_blank\"><strong>MASt3R-ASMK</strong></a> : Used the official implementation to compute the similarity matrix.<br>\n(iii) <a href=\"https://arxiv.org/abs/2404.19174\" target=\"_blank\"><strong>XFeat local-feature-aggregation</strong></a> : XFeat is a local keypoint feature extractor which is very fast but comparable to Superpoint, etc in terms of accuracy. Created a custom function to compute the \"mean cosine-similarity\" of top-k nearest-neighbor matching keypoint features between 2 images, as the image similarity.</p>\n<p>\nThe similarity matrix from each method was passed to a function with the hyper-parameters topk_min, topk_max, topk_percentile and sim_thres were used to shortlist the candidate set of pairs for image matching.\n</p>\n<ul>\n<li><strong>sim_thres</strong> : Hard threshold on similarity score </li>\n<li><strong>topk_percentile</strong> : Having a fixed top-k threshold restricts adapting the no. of similar images when the dataset size varies. Defining it as a percentile makes it adaptive.</li>\n<li><strong>topk_min</strong> and <strong>topk_max</strong> : The minimum and maximum no. of image pairs allowed for matching, these thresholds limit the pairs in case top-k percentile pairs are too low or high in number.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2474539%2F66e8cbd723372cee58c08306655820ff%2Ftable1.png?generation=1749407386792438&amp;alt=media\" alt=\"table-1\"></li>\n</ul>\n<p>As you can observe from the above table that MASt3R-ASMK performed better than the others, so it was chosen. The hyper-parameters topk_min, topk_max, topk_percentile and sim_thres were tuned separately for each method using the local validation dataset (IMC-2025-train). I submitted some of the hyper-parameters combinations to the public LB to verify the best choice. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2474539%2Fc5415dfc4506d097c523861b6ea4ca53%2Ftable2.png?generation=1749407519100892&amp;alt=media\" alt=\"table-2\"></p>\n<p> Though MASt3R Matcher has high precision and is able to discard false positives by predicting very less matches, it still helps to pass only relevant image-pairs. Firstly, because the matching and reconstruction needs to be completed within the 9h time budget. Secondly, more pairs doesn't necessarily mean better accuracy (refer to Table-2 row-4). </p>\n<p> Another thing that I wanted to try is the ensemble of the image-shortlisting methods but due to lack of time and attempts, couldn't test it. </p>\n<h5>2. Image Matching</h5>\n<p>I used the MASt3R model's feature extraction <code>(image_size = 512)</code> and semi-dense matching using Fast-Reciprocal-NN from the official implementation <code>(subsample = 8, pixel_tol = 5)</code> . I tried to tune the min_conf_thres (Minimum Confidence Threshold) parameter for the matches a little bit, keeping all other params fixed. <code>(MASt3R-ASMK: topk_min - 10, topk_max - 30, topk_percentile - 0.3, sim_thres - 0.001)</code> .<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2474539%2F90c8fd042909465f98e3b99377b7b3b2%2Ftable3.png?generation=1749407724111672&amp;alt=media\" alt=\"table-3\"></p>\n<h5>3. COLMAP based Verification, Reconstruction, and Clustering</h5>\n<p> Once the matches are estimated using MASt3R, they are exported to COLMAP format and final image pair indexes are dumped. Then pyCOLMAP API's geometric verification is used to verify the matches. </p>\n<p> After that, pyCOLMAP's Incremental Mapping pipeline is used to generate the list of reconstructions or models. Due to MASt3R's superior matching precision, COLMAP is able to easily cluster images into separate sparse models. The following parameter values were used for the mapping. </p>\n<pre><code> = pycolmap.IncrementalPipelineOptions()\n = \n = \n</code></pre>\n<p> I have also tried pre-clustering the images using either MASt3R-ASMK or DINO-v2, followed by MASt3R matching in each cluster and then generating single reconstruction model by either COLMAP (Incremental-SFM) or GLOMAP (Global-SFM). But this approach performed worser than the automatic clustering using COLMAP, on local validation set, so discarded. </p>\n<h5>4. Engineering Tips and Tricks</h5>\n<ul>\n<li>Parallel Processing of Individual Datasets on the 2 x T4 GPUs (weighted distribution based on dataset size).</li>\n<li>Build cuRoPE module of CroCo (MASt3R) using CUDA.</li>\n<li>Isolate modules into subprocesses which can randomly crash (like for e.g. pyCOLMAP) and retry execution for max_retries</li>\n</ul>",
  "messages": [
    {
      "id": "3220085",
      "postDate": "06/08/2025 18:49:26",
      "content": "<p>First and foremost, I'd like to thank the organizers and Kaggle team for hosting such an exciting and challenging competition. Congratulations to all the top participants! While I’m not new to the field of 3D vision, this was my first time properly participating in the Image Matching Challenge. One of the main challenges I faced was joining the competition quite late - around two weeks before the deadline. As a result, a significant portion of my time went into setting up Kaggle notebooks, leaving me with limited opportunities to iterate and experiment with different approaches and their variants. That said, let's dive into the simple yet effective solution that worked for me.</p>\n<h4>Overview</h4>\n<p>From recent personal experiments, I’ve observed that foundation models like DUSt3R, MASt3R, and VGGT offer superior matching performance compared to traditional detector-descriptor-matcher pipelines (e.g., ALIKED or SuperPoint + LightGlue). Based on this, I built a straightforward pipeline using MASt3R and evaluated it on the IMC-2025 dataset. As expected, it performed well in terms of matching accuracy. However, a major challenge was the high computational cost of the MASt3R matcher, which led to frequent notebook timeouts during inference. I addressed this by implementing an efficient image-pair shortlisting strategy, tuning its hyperparameters, and applying a few engineering optimizations to reduce runtime without significantly impacting performance.<br>\nThe pipeline that gave the best result on Public LB :-</p>\n<ol>\n<li>Image-Pair Similarity based Shortlisting using <a href=\"https://arxiv.org/abs/2409.19152\" target=\"_blank\">MASt3R-ASMK</a></li>\n<li><a href=\"https://arxiv.org/abs/2406.09756\" target=\"_blank\">MASt3R</a> Semi-Dense Matching on the shortlisted image pairs</li>\n<li>COLMAP based Verification, Reconstruction and Clustering<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2474539%2F21b88c6ebf846a383c84751f92005605%2Fpipeline_overview.png?generation=1749406349577280&amp;alt=media\" alt=\"PipelIne Overview Block Diagram\"></li>\n</ol>\n<h4>Solution Details</h4>\n<h5>1. Image-Pair Similarity based Shortlisting</h5>\n<p> I tried 3 methods to compute a similarity matrix between all the images of a dataset in a fast way and shortlist the pairs using certain thresholds, for the more expensive MASt3R Matcher to run on them. </p>\n<p>(i) <a href=\"https://arxiv.org/abs/2304.07193\" target=\"_blank\"><strong>DINO-v2</strong></a> : Similarity matrix was computed by normalizing the (1 - distance_matrix), where L2-distance was used.<br>\n(ii) <a href=\"https://arxiv.org/abs/2409.19152\" target=\"_blank\"><strong>MASt3R-ASMK</strong></a> : Used the official implementation to compute the similarity matrix.<br>\n(iii) <a href=\"https://arxiv.org/abs/2404.19174\" target=\"_blank\"><strong>XFeat local-feature-aggregation</strong></a> : XFeat is a local keypoint feature extractor which is very fast but comparable to Superpoint, etc in terms of accuracy. Created a custom function to compute the \"mean cosine-similarity\" of top-k nearest-neighbor matching keypoint features between 2 images, as the image similarity.</p>\n<p>\nThe similarity matrix from each method was passed to a function with the hyper-parameters topk_min, topk_max, topk_percentile and sim_thres were used to shortlist the candidate set of pairs for image matching.\n</p>\n<ul>\n<li><strong>sim_thres</strong> : Hard threshold on similarity score </li>\n<li><strong>topk_percentile</strong> : Having a fixed top-k threshold restricts adapting the no. of similar images when the dataset size varies. Defining it as a percentile makes it adaptive.</li>\n<li><strong>topk_min</strong> and <strong>topk_max</strong> : The minimum and maximum no. of image pairs allowed for matching, these thresholds limit the pairs in case top-k percentile pairs are too low or high in number.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2474539%2F66e8cbd723372cee58c08306655820ff%2Ftable1.png?generation=1749407386792438&amp;alt=media\" alt=\"table-1\"></li>\n</ul>\n<p>As you can observe from the above table that MASt3R-ASMK performed better than the others, so it was chosen. The hyper-parameters topk_min, topk_max, topk_percentile and sim_thres were tuned separately for each method using the local validation dataset (IMC-2025-train). I submitted some of the hyper-parameters combinations to the public LB to verify the best choice. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2474539%2Fc5415dfc4506d097c523861b6ea4ca53%2Ftable2.png?generation=1749407519100892&amp;alt=media\" alt=\"table-2\"></p>\n<p> Though MASt3R Matcher has high precision and is able to discard false positives by predicting very less matches, it still helps to pass only relevant image-pairs. Firstly, because the matching and reconstruction needs to be completed within the 9h time budget. Secondly, more pairs doesn't necessarily mean better accuracy (refer to Table-2 row-4). </p>\n<p> Another thing that I wanted to try is the ensemble of the image-shortlisting methods but due to lack of time and attempts, couldn't test it. </p>\n<h5>2. Image Matching</h5>\n<p>I used the MASt3R model's feature extraction <code>(image_size = 512)</code> and semi-dense matching using Fast-Reciprocal-NN from the official implementation <code>(subsample = 8, pixel_tol = 5)</code> . I tried to tune the min_conf_thres (Minimum Confidence Threshold) parameter for the matches a little bit, keeping all other params fixed. <code>(MASt3R-ASMK: topk_min - 10, topk_max - 30, topk_percentile - 0.3, sim_thres - 0.001)</code> .<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2474539%2F90c8fd042909465f98e3b99377b7b3b2%2Ftable3.png?generation=1749407724111672&amp;alt=media\" alt=\"table-3\"></p>\n<h5>3. COLMAP based Verification, Reconstruction, and Clustering</h5>\n<p> Once the matches are estimated using MASt3R, they are exported to COLMAP format and final image pair indexes are dumped. Then pyCOLMAP API's geometric verification is used to verify the matches. </p>\n<p> After that, pyCOLMAP's Incremental Mapping pipeline is used to generate the list of reconstructions or models. Due to MASt3R's superior matching precision, COLMAP is able to easily cluster images into separate sparse models. The following parameter values were used for the mapping. </p>\n<pre><code> = pycolmap.IncrementalPipelineOptions()\n = \n = \n</code></pre>\n<p> I have also tried pre-clustering the images using either MASt3R-ASMK or DINO-v2, followed by MASt3R matching in each cluster and then generating single reconstruction model by either COLMAP (Incremental-SFM) or GLOMAP (Global-SFM). But this approach performed worser than the automatic clustering using COLMAP, on local validation set, so discarded. </p>\n<h5>4. Engineering Tips and Tricks</h5>\n<ul>\n<li>Parallel Processing of Individual Datasets on the 2 x T4 GPUs (weighted distribution based on dataset size).</li>\n<li>Build cuRoPE module of CroCo (MASt3R) using CUDA.</li>\n<li>Isolate modules into subprocesses which can randomly crash (like for e.g. pyCOLMAP) and retry execution for max_retries</li>\n</ul>",
      "rawMarkdown": "First and foremost, I'd like to thank the organizers and Kaggle team for hosting such an exciting and challenging competition. Congratulations to all the top participants! While I’m not new to the field of 3D vision, this was my first time properly participating in the Image Matching Challenge. One of the main challenges I faced was joining the competition quite late - around two weeks before the deadline. As a result, a significant portion of my time went into setting up Kaggle notebooks, leaving me with limited opportunities to iterate and experiment with different approaches and their variants. That said, let's dive into the simple yet effective solution that worked for me.\n\n#### Overview\n\nFrom recent personal experiments, I’ve observed that foundation models like DUSt3R, MASt3R, and VGGT offer superior matching performance compared to traditional detector-descriptor-matcher pipelines (e.g., ALIKED or SuperPoint + LightGlue). Based on this, I built a straightforward pipeline using MASt3R and evaluated it on the IMC-2025 dataset. As expected, it performed well in terms of matching accuracy. However, a major challenge was the high computational cost of the MASt3R matcher, which led to frequent notebook timeouts during inference. I addressed this by implementing an efficient image-pair shortlisting strategy, tuning its hyperparameters, and applying a few engineering optimizations to reduce runtime without significantly impacting performance.\n\nThe pipeline that gave the best result on Public LB :-\n1. Image-Pair Similarity based Shortlisting using [MASt3R-ASMK](https://arxiv.org/abs/2409.19152)\n2. [MASt3R](https://arxiv.org/abs/2406.09756) Semi-Dense Matching on the shortlisted image pairs\n3. COLMAP based Verification, Reconstruction and Clustering\n\n![PipelIne Overview Block Diagram](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2474539%2F21b88c6ebf846a383c84751f92005605%2Fpipeline_overview.png?generation=1749406349577280&alt=media)\n\n\n#### Solution Details\n\n##### 1. Image-Pair Similarity based Shortlisting \n\n<p> I tried 3 methods to compute a similarity matrix between all the images of a dataset in a fast way and shortlist the pairs using certain thresholds, for the more expensive MASt3R Matcher to run on them. </p>\n\n(i) [**DINO-v2**](https://arxiv.org/abs/2304.07193) : Similarity matrix was computed by normalizing the (1 - distance_matrix), where L2-distance was used.\n(ii) [**MASt3R-ASMK**](https://arxiv.org/abs/2409.19152) : Used the official implementation to compute the similarity matrix.\n(iii) [**XFeat local-feature-aggregation**](https://arxiv.org/abs/2404.19174) : XFeat is a local keypoint feature extractor which is very fast but comparable to Superpoint, etc in terms of accuracy. Created a custom function to compute the \"mean cosine-similarity\" of top-k nearest-neighbor matching keypoint features between 2 images, as the image similarity.\n<p>\nThe similarity matrix from each method was passed to a function with the hyper-parameters topk_min, topk_max, topk_percentile and sim_thres were used to shortlist the candidate set of pairs for image matching.\n</p>\n- **sim_thres** : Hard threshold on similarity score \n- **topk_percentile** : Having a fixed top-k threshold restricts adapting the no. of similar images when the dataset size varies. Defining it as a percentile makes it adaptive.\n- **topk_min** and **topk_max** : The minimum and maximum no. of image pairs allowed for matching, these thresholds limit the pairs in case top-k percentile pairs are too low or high in number.\n\n![table-1](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2474539%2F66e8cbd723372cee58c08306655820ff%2Ftable1.png?generation=1749407386792438&alt=media)\n\n<p>As you can observe from the above table that MASt3R-ASMK performed better than the others, so it was chosen. The hyper-parameters topk_min, topk_max, topk_percentile and sim_thres were tuned separately for each method using the local validation dataset (IMC-2025-train). I submitted some of the hyper-parameters combinations to the public LB to verify the best choice. </p>\n\n![table-2](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2474539%2Fc5415dfc4506d097c523861b6ea4ca53%2Ftable2.png?generation=1749407519100892&alt=media)\n\n<p> Though MASt3R Matcher has high precision and is able to discard false positives by predicting very less matches, it still helps to pass only relevant image-pairs. Firstly, because the matching and reconstruction needs to be completed within the 9h time budget. Secondly, more pairs doesn't necessarily mean better accuracy (refer to Table-2 row-4). </p>\n\n<p> Another thing that I wanted to try is the ensemble of the image-shortlisting methods but due to lack of time and attempts, couldn't test it. </p>\n\n##### 2. Image Matching\n\nI used the MASt3R model's feature extraction `(image_size = 512)` and semi-dense matching using Fast-Reciprocal-NN from the official implementation `(subsample = 8, pixel_tol = 5)` . I tried to tune the min_conf_thres (Minimum Confidence Threshold) parameter for the matches a little bit, keeping all other params fixed. `(MASt3R-ASMK: topk_min - 10, topk_max - 30, topk_percentile - 0.3, sim_thres - 0.001)` .\n\n![table-3](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2474539%2F90c8fd042909465f98e3b99377b7b3b2%2Ftable3.png?generation=1749407724111672&alt=media)\n\n##### 3. COLMAP based Verification, Reconstruction, and Clustering\n\n<p> Once the matches are estimated using MASt3R, they are exported to COLMAP format and final image pair indexes are dumped. Then pyCOLMAP API's geometric verification is used to verify the matches. </p>\n\n<p> After that, pyCOLMAP's Incremental Mapping pipeline is used to generate the list of reconstructions or models. Due to MASt3R's superior matching precision, COLMAP is able to easily cluster images into separate sparse models. The following parameter values were used for the mapping. </p>\n\n```\nmapper_options = pycolmap.IncrementalPipelineOptions()\nmapper_options.min_model_size = 3\nmapper_options.max_num_models = 25\n```\n\n<p> I have also tried pre-clustering the images using either MASt3R-ASMK or DINO-v2, followed by MASt3R matching in each cluster and then generating single reconstruction model by either COLMAP (Incremental-SFM) or GLOMAP (Global-SFM). But this approach performed worser than the automatic clustering using COLMAP, on local validation set, so discarded. </p>\n\n##### 4. Engineering Tips and Tricks\n\n- Parallel Processing of Individual Datasets on the 2 x T4 GPUs (weighted distribution based on dataset size).\n- Build cuRoPE module of CroCo (MASt3R) using CUDA.\n- Isolate modules into subprocesses which can randomly crash (like for e.g. pyCOLMAP) and retry execution for max_retries",
      "votes": null
    },
    {
      "id": "3220787",
      "postDate": "06/10/2025 01:32:06",
      "content": "<p>Congratulations on the gold medal!<br>\nI was really impressed by your simple yet powerful solution leveraging MASt3R.</p>\n<p>Sorry if this is a basic question, but I’d like to ask something about the matching part of MASt3R.<br>\nWhen you mentioned \"Fast-Reciprocal-NN from the official implementation,\" I believe you were referring to this <a href=\"https://github.com/naver/mast3r/blob/fb8dfe783043764f6ad78dd021bfcd88bc1d73a2/mast3r/fast_nn.py#L109\" target=\"_blank\">implementation</a>.<br>\nHow can I change the min_conf_thres parameter?</p>\n<p>I wanted to try it myself, but I couldn’t find a corresponding argument in the function definition, so I wasn’t sure how to set it.<br>\nI’d appreciate it if you could let me know.</p>",
      "rawMarkdown": "Congratulations on the gold medal!\nI was really impressed by your simple yet powerful solution leveraging MASt3R.\n\nSorry if this is a basic question, but I’d like to ask something about the matching part of MASt3R.\nWhen you mentioned \"Fast-Reciprocal-NN from the official implementation,\" I believe you were referring to this [implementation](https://github.com/naver/mast3r/blob/fb8dfe783043764f6ad78dd021bfcd88bc1d73a2/mast3r/fast_nn.py#L109).\nHow can I change the min_conf_thres parameter?\n\nI wanted to try it myself, but I couldn’t find a corresponding argument in the function definition, so I wasn’t sure how to set it.\nI’d appreciate it if you could let me know.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3220787,
      "author_name": "kashiwaba",
      "author_url": "",
      "post_date": "06/10/2025 01:32:06",
      "content": "<p>Congratulations on the gold medal!<br>\nI was really impressed by your simple yet powerful solution leveraging MASt3R.</p>\n<p>Sorry if this is a basic question, but I’d like to ask something about the matching part of MASt3R.<br>\nWhen you mentioned \"Fast-Reciprocal-NN from the official implementation,\" I believe you were referring to this <a href=\"https://github.com/naver/mast3r/blob/fb8dfe783043764f6ad78dd021bfcd88bc1d73a2/mast3r/fast_nn.py#L109\" target=\"_blank\">implementation</a>.<br>\nHow can I change the min_conf_thres parameter?</p>\n<p>I wanted to try it myself, but I couldn’t find a corresponding argument in the function definition, so I wasn’t sure how to set it.<br>\nI’d appreciate it if you could let me know.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3220085": "First and foremost, I'd like to thank the organizers and Kaggle team for hosting such an exciting and challenging competition. Congratulations to all the top participants! While I’m not new to the field of 3D vision, this was my first time properly participating in the Image Matching Challenge. One of the main challenges I faced was joining the competition quite late - around two weeks before the deadline. As a result, a significant portion of my time went into setting up Kaggle notebooks, leaving me with limited opportunities to iterate and experiment with different approaches and their variants. That said, let's dive into the simple yet effective solution that worked for me.\n\n#### Overview\n\nFrom recent personal experiments, I’ve observed that foundation models like DUSt3R, MASt3R, and VGGT offer superior matching performance compared to traditional detector-descriptor-matcher pipelines (e.g., ALIKED or SuperPoint + LightGlue). Based on this, I built a straightforward pipeline using MASt3R and evaluated it on the IMC-2025 dataset. As expected, it performed well in terms of matching accuracy. However, a major challenge was the high computational cost of the MASt3R matcher, which led to frequent notebook timeouts during inference. I addressed this by implementing an efficient image-pair shortlisting strategy, tuning its hyperparameters, and applying a few engineering optimizations to reduce runtime without significantly impacting performance.\n\nThe pipeline that gave the best result on Public LB :-\n1. Image-Pair Similarity based Shortlisting using [MASt3R-ASMK](https://arxiv.org/abs/2409.19152)\n2. [MASt3R](https://arxiv.org/abs/2406.09756) Semi-Dense Matching on the shortlisted image pairs\n3. COLMAP based Verification, Reconstruction and Clustering\n\n![PipelIne Overview Block Diagram](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2474539%2F21b88c6ebf846a383c84751f92005605%2Fpipeline_overview.png?generation=1749406349577280&alt=media)\n\n\n#### Solution Details\n\n##### 1. Image-Pair Similarity based Shortlisting \n\n<p> I tried 3 methods to compute a similarity matrix between all the images of a dataset in a fast way and shortlist the pairs using certain thresholds, for the more expensive MASt3R Matcher to run on them. </p>\n\n(i) [**DINO-v2**](https://arxiv.org/abs/2304.07193) : Similarity matrix was computed by normalizing the (1 - distance_matrix), where L2-distance was used.\n(ii) [**MASt3R-ASMK**](https://arxiv.org/abs/2409.19152) : Used the official implementation to compute the similarity matrix.\n(iii) [**XFeat local-feature-aggregation**](https://arxiv.org/abs/2404.19174) : XFeat is a local keypoint feature extractor which is very fast but comparable to Superpoint, etc in terms of accuracy. Created a custom function to compute the \"mean cosine-similarity\" of top-k nearest-neighbor matching keypoint features between 2 images, as the image similarity.\n<p>\nThe similarity matrix from each method was passed to a function with the hyper-parameters topk_min, topk_max, topk_percentile and sim_thres were used to shortlist the candidate set of pairs for image matching.\n</p>\n- **sim_thres** : Hard threshold on similarity score \n- **topk_percentile** : Having a fixed top-k threshold restricts adapting the no. of similar images when the dataset size varies. Defining it as a percentile makes it adaptive.\n- **topk_min** and **topk_max** : The minimum and maximum no. of image pairs allowed for matching, these thresholds limit the pairs in case top-k percentile pairs are too low or high in number.\n\n![table-1](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2474539%2F66e8cbd723372cee58c08306655820ff%2Ftable1.png?generation=1749407386792438&alt=media)\n\n<p>As you can observe from the above table that MASt3R-ASMK performed better than the others, so it was chosen. The hyper-parameters topk_min, topk_max, topk_percentile and sim_thres were tuned separately for each method using the local validation dataset (IMC-2025-train). I submitted some of the hyper-parameters combinations to the public LB to verify the best choice. </p>\n\n![table-2](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2474539%2Fc5415dfc4506d097c523861b6ea4ca53%2Ftable2.png?generation=1749407519100892&alt=media)\n\n<p> Though MASt3R Matcher has high precision and is able to discard false positives by predicting very less matches, it still helps to pass only relevant image-pairs. Firstly, because the matching and reconstruction needs to be completed within the 9h time budget. Secondly, more pairs doesn't necessarily mean better accuracy (refer to Table-2 row-4). </p>\n\n<p> Another thing that I wanted to try is the ensemble of the image-shortlisting methods but due to lack of time and attempts, couldn't test it. </p>\n\n##### 2. Image Matching\n\nI used the MASt3R model's feature extraction `(image_size = 512)` and semi-dense matching using Fast-Reciprocal-NN from the official implementation `(subsample = 8, pixel_tol = 5)` . I tried to tune the min_conf_thres (Minimum Confidence Threshold) parameter for the matches a little bit, keeping all other params fixed. `(MASt3R-ASMK: topk_min - 10, topk_max - 30, topk_percentile - 0.3, sim_thres - 0.001)` .\n\n![table-3](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2474539%2F90c8fd042909465f98e3b99377b7b3b2%2Ftable3.png?generation=1749407724111672&alt=media)\n\n##### 3. COLMAP based Verification, Reconstruction, and Clustering\n\n<p> Once the matches are estimated using MASt3R, they are exported to COLMAP format and final image pair indexes are dumped. Then pyCOLMAP API's geometric verification is used to verify the matches. </p>\n\n<p> After that, pyCOLMAP's Incremental Mapping pipeline is used to generate the list of reconstructions or models. Due to MASt3R's superior matching precision, COLMAP is able to easily cluster images into separate sparse models. The following parameter values were used for the mapping. </p>\n\n```\nmapper_options = pycolmap.IncrementalPipelineOptions()\nmapper_options.min_model_size = 3\nmapper_options.max_num_models = 25\n```\n\n<p> I have also tried pre-clustering the images using either MASt3R-ASMK or DINO-v2, followed by MASt3R matching in each cluster and then generating single reconstruction model by either COLMAP (Incremental-SFM) or GLOMAP (Global-SFM). But this approach performed worser than the automatic clustering using COLMAP, on local validation set, so discarded. </p>\n\n##### 4. Engineering Tips and Tricks\n\n- Parallel Processing of Individual Datasets on the 2 x T4 GPUs (weighted distribution based on dataset size).\n- Build cuRoPE module of CroCo (MASt3R) using CUDA.\n- Isolate modules into subprocesses which can randomly crash (like for e.g. pyCOLMAP) and retry execution for max_retries",
    "3220787": "Congratulations on the gold medal!\nI was really impressed by your simple yet powerful solution leveraging MASt3R.\n\nSorry if this is a basic question, but I’d like to ask something about the matching part of MASt3R.\nWhen you mentioned \"Fast-Reciprocal-NN from the official implementation,\" I believe you were referring to this [implementation](https://github.com/naver/mast3r/blob/fb8dfe783043764f6ad78dd021bfcd88bc1d73a2/mast3r/fast_nn.py#L109).\nHow can I change the min_conf_thres parameter?\n\nI wanted to try it myself, but I couldn’t find a corresponding argument in the function definition, so I wasn’t sure how to set it.\nI’d appreciate it if you could let me know."
  },
  "source": "meta"
}