{
  "id": 583058,
  "title": "1st Place Solution",
  "url": "/competitions/image-matching-challenge-2025/writeups/ns64-1st-place-solution",
  "author_name": "",
  "post_date": "2025-06-15T13:29:28.973Z",
  "votes": 66,
  "comment_count": 18,
  "views": 0,
  "content": "<p>Thank you to the organizers and Kaggle team for this exciting competition. Congratulations to all the participants. Although I first participated in IMC in 2022, I'm glad that this time I achieved my best results ever.</p>\n<p>This year, I focused on utilizing recent 3D geometric foundation models such as MASt3R and VGGT. I was surprised at their potential. I respect the authors who developed such great models.</p>\n<h2>1. Overview</h2>\n<p>I developed a simple MASt3R-based pipeline.</p>\n<p>As far as I've tried, image matching using MASt3R's local feature head appears to achieve significantly better results than other detector-based methods such as ALIKED+LG on IMC25.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F563483%2Fefd3ae4c091bb1f01f9f9d43a1469d8b%2Ffig1.png?generation=1749042274433083&amp;alt=media\" alt=\"overview\"></p>\n<ul>\n<li>Pre-clustering using coarse MASt3R matching on the top-k neighbor images (not shown in the figure).</li>\n<li>Image pair extraction with the combination of multiple shortlists derived from image retrieval.</li>\n<li>Utilizing MASt3R (<a href=\"https://github.com/naver/mast3r\" target=\"_blank\">https://github.com/naver/mast3r</a>) as a matcher. In addition to MASt3R semi-dense matches,<br>\nkeypoints extracted by other keypoint detectors are also matched using MASt3R.</li>\n<li>Generating reconstructions using a general COLMAP pipeline (based on <a href=\"https://www.kaggle.com/code/eduardtrulls/imc25-submission\" target=\"_blank\">https://www.kaggle.com/code/eduardtrulls/imc25-submission</a>)</li>\n</ul>\n<h2>2. Solution</h2>\n<p>I observed that MASt3R-based matching achieves higher precision (i.e., fewer mismatches) than other methods, so it can match images in messy scenes where detector-based methods typically fail.<br>\nI think the following points are probably useful for robust matching:</p>\n<ul>\n<li>3D geometric features (From DUSt3R)</li>\n<li>MASt3R is trained not only with the MegaDepth dataset but also with several object-centric datasets</li>\n</ul>\n<p>By simply using the MASt3R model as a semi-dense detector-free matcher and integrating it into a general COLMAP pipeline, I achieved a score of 42-45 on the PublicLB. Furthermore, I found that increasing the number of image pairs led to a higher score, reaching approximately 50. Therefore, I thought that how image pairs could be increased within limited computational time might be an important consideration.</p>\n<h3>2.1. Clustering</h3>\n<p>I developed a pre-clustering approach based on MASt3R matches, but this approach was not adopted in the final version. It turned out that, since image matching primarily uses MASt3R, there wasn't a significant difference whether I used pre-clustering or multiple reconstructions from COLMAP.</p>\n<p>The pre-clustering (which was not used) approach is outlined below:</p>\n<ol>\n<li>Initialize the cluster label for each image to -1</li>\n<li>Extract N images from the scene using farthest point sampling, and assign each of them a unique cluster label (0, 1, …, N-1).</li>\n<li>For each of the N seed images (from step 2), run 1-vs-all matching with MASt3R against other unclustered images. If an image is matched, assign the seed image's cluster label to the matched image. Otherwise, assign a new cluster label to it.</li>\n<li>Use the newly matched images as the next queries, and repeat 1-vs-k matching iteratively until all images are assigned a cluster label</li>\n<li>Treat small clusters as \"outliers\"</li>\n</ol>\n<h3>2.2. Shortlist</h3>\n<p>Since the MASt3R matcher is computationally more intensive than detector-based methods, a shortlist of image pairs for matching is still important.</p>\n<p>I generated candidate pairs for matching within a scene by taking the union of neighbors from multiple image retrieval results. I used the following four global features for image retrieval.</p>\n<ul>\n<li>MASt3R-ASMK (From MASt3R-SfM <a href=\"https://arxiv.org/abs/2409.19152\" target=\"_blank\">https://arxiv.org/abs/2409.19152</a>)</li>\n<li>MASt3R-SPoC (From MASt3R-SfM <a href=\"https://arxiv.org/abs/2409.19152\" target=\"_blank\">https://arxiv.org/abs/2409.19152</a>)</li>\n<li>DINOv2</li>\n<li>ISC (<a href=\"https://arxiv.org/abs/2112.04323\" target=\"_blank\">https://arxiv.org/abs/2112.04323</a>)</li>\n</ul>\n<p>Although MASt3R-ASMK could extract pairs with sufficient coverage, adding the other models slightly improved the score.</p>\n<table>\n<thead>\n<tr>\n<th>Features</th>\n<th>Parameters</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>MASt3R-ASMK</td>\n<td>n=10, k=25 (See MASt3R-SfM paper)</td>\n</tr>\n<tr>\n<td>MASt3R-SPoC</td>\n<td>topk=10</td>\n</tr>\n<tr>\n<td>DINOv2</td>\n<td>topk=10</td>\n</tr>\n<tr>\n<td>ISC</td>\n<td>topk=10</td>\n</tr>\n</tbody>\n</table>\n<p>Note that using only MASt3R-ASMK with heavier settings (e.g., MASt3R-ASMK(n=10, k=60)) can also achieve a high score, but the proposed method is faster while achieving a similar score.</p>\n<h3>2.3. Pairwise matching</h3>\n<p>First, matches for an image pair are computed by MASt3R. This implementation is based on <code>fast_reciprocal_NNs()</code> from the code in the official repository. I used default parameters of subsample=8 and pixel_tol=5.</p>\n<p>In addition to the semi-dense matches, keypoints extracted by other keypoint detectors are also fed to the MASt3R matcher (Inspired by MP-SfM <a href=\"https://arxiv.org/abs/2504.20040)\" target=\"_blank\">https://arxiv.org/abs/2504.20040)</a>. These additional keypoints might regionally overlap with points subsampled by MASt3R, but this approach improved the score compared to using only MASt3R matches.</p>\n<p>I used ALIKED and SuperPoint as the additional keypoint detectors. While I also tried SIFT, GIMSuperPoint, and DaD, I ultimately adopted the combination of ALIKED and SuperPoint because this combination yielded the best LB score.</p>\n<p>The detailed configurations are as follows:</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Parameters</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>MASt3R</td>\n<td>size=512, threshold=1.001</td>\n</tr>\n<tr>\n<td>ALIKED detector</td>\n<td>size=1280, max_keypoints=4096</td>\n</tr>\n<tr>\n<td>SuperPoint detector</td>\n<td>size=1600, max_keypoints=4096, threshold=0.0005</td>\n</tr>\n</tbody>\n</table>\n<h3>2.4. Engineering tips</h3>\n<p>In addition, I applied the following techniques to make the pipeline faster:</p>\n<ul>\n<li>Build the curope (RoPE2D) module with CUDA.</li>\n<li>Replace attention implementations used in mast3r/dust3r/croco with <code>torch.nn.functional.scaled_dot_product_attention</code>.<br>\n(to enable flash attention)</li>\n<li>Fix <code>use_amp</code> args in the MASt3R inference function to be used correctly.</li>\n<li>Use the T4 x 2 environment in Kaggle, and run the submission pipeline in parallel over scene subsets, split by dataset in <code>submission.csv</code>.</li>\n</ul>\n<p>According to Speedy MASt3R (<a href=\"https://arxiv.org/abs/2503.10017)\" target=\"_blank\">https://arxiv.org/abs/2503.10017)</a>, TensorRT can further accelerate the MASt3R model. However, I wasn't able to convert the model.</p>\n<h3>2.5. Local/Public/Private Score</h3>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>amy_gardens</th>\n<th>fbk_vineyard</th>\n<th>ETs</th>\n<th>stairs</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Best submission</td>\n<td>40.65</td>\n<td>72.09</td>\n<td>59.46</td>\n<td>15.89</td>\n<td>52.64</td>\n<td>56.00</td>\n</tr>\n<tr>\n<td>w/ Pre-clustering</td>\n<td>37.02</td>\n<td>47.25</td>\n<td>59.46</td>\n<td>20.35</td>\n<td>50.22</td>\n<td>50.93</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>The scores of <code>amy_gardens</code> and <code>fbk_vineyard</code> seemed to be unstable.</li>\n<li><code>stairs</code> was difficult (As a side note, VGGT achieved the highest score of <code>25.07</code> in my local experiments)</li>\n</ul>\n<h2>3. What did not work</h2>\n<ul>\n<li>GLOMAP: I tried GLOMAP instead of COLMAP with the pre-clustering approach, but it didn't improve the score.</li>\n<li>Coarse-to-Fine matching: Probably, there was a bug in my implementation.</li>\n<li>Using monocular depth estimation: I tried to filter mismatched pairs using depth information.</li>\n<li>VGGT: The tracking head of VGGT with pre-extracted keypoints worked well on local tests, but it resulted in a <code>TimeoutError</code> in submission.</li>\n</ul>\n<h2>Code</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/ns6464/imc2025-1st-place-solution\" target=\"_blank\">Notebook</a></li>\n<li><a href=\"https://github.com/ns-rokuyon/kaggle-image-matching-challenge-2025\" target=\"_blank\">Github</a></li>\n</ul>",
  "messages": [
    {
      "id": "3217069",
      "postDate": "06/04/2025 13:46:15",
      "content": "<p>Thank you to the organizers and Kaggle team for this exciting competition. Congratulations to all the participants. Although I first participated in IMC in 2022, I'm glad that this time I achieved my best results ever.</p>\n<p>This year, I focused on utilizing recent 3D geometric foundation models such as MASt3R and VGGT. I was surprised at their potential. I respect the authors who developed such great models.</p>\n<h2>1. Overview</h2>\n<p>I developed a simple MASt3R-based pipeline.</p>\n<p>As far as I've tried, image matching using MASt3R's local feature head appears to achieve significantly better results than other detector-based methods such as ALIKED+LG on IMC25.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F563483%2Fefd3ae4c091bb1f01f9f9d43a1469d8b%2Ffig1.png?generation=1749042274433083&amp;alt=media\" alt=\"overview\"></p>\n<ul>\n<li>Pre-clustering using coarse MASt3R matching on the top-k neighbor images (not shown in the figure).</li>\n<li>Image pair extraction with the combination of multiple shortlists derived from image retrieval.</li>\n<li>Utilizing MASt3R (<a href=\"https://github.com/naver/mast3r\" target=\"_blank\">https://github.com/naver/mast3r</a>) as a matcher. In addition to MASt3R semi-dense matches,<br>\nkeypoints extracted by other keypoint detectors are also matched using MASt3R.</li>\n<li>Generating reconstructions using a general COLMAP pipeline (based on <a href=\"https://www.kaggle.com/code/eduardtrulls/imc25-submission\" target=\"_blank\">https://www.kaggle.com/code/eduardtrulls/imc25-submission</a>)</li>\n</ul>\n<h2>2. Solution</h2>\n<p>I observed that MASt3R-based matching achieves higher precision (i.e., fewer mismatches) than other methods, so it can match images in messy scenes where detector-based methods typically fail.<br>\nI think the following points are probably useful for robust matching:</p>\n<ul>\n<li>3D geometric features (From DUSt3R)</li>\n<li>MASt3R is trained not only with the MegaDepth dataset but also with several object-centric datasets</li>\n</ul>\n<p>By simply using the MASt3R model as a semi-dense detector-free matcher and integrating it into a general COLMAP pipeline, I achieved a score of 42-45 on the PublicLB. Furthermore, I found that increasing the number of image pairs led to a higher score, reaching approximately 50. Therefore, I thought that how image pairs could be increased within limited computational time might be an important consideration.</p>\n<h3>2.1. Clustering</h3>\n<p>I developed a pre-clustering approach based on MASt3R matches, but this approach was not adopted in the final version. It turned out that, since image matching primarily uses MASt3R, there wasn't a significant difference whether I used pre-clustering or multiple reconstructions from COLMAP.</p>\n<p>The pre-clustering (which was not used) approach is outlined below:</p>\n<ol>\n<li>Initialize the cluster label for each image to -1</li>\n<li>Extract N images from the scene using farthest point sampling, and assign each of them a unique cluster label (0, 1, …, N-1).</li>\n<li>For each of the N seed images (from step 2), run 1-vs-all matching with MASt3R against other unclustered images. If an image is matched, assign the seed image's cluster label to the matched image. Otherwise, assign a new cluster label to it.</li>\n<li>Use the newly matched images as the next queries, and repeat 1-vs-k matching iteratively until all images are assigned a cluster label</li>\n<li>Treat small clusters as \"outliers\"</li>\n</ol>\n<h3>2.2. Shortlist</h3>\n<p>Since the MASt3R matcher is computationally more intensive than detector-based methods, a shortlist of image pairs for matching is still important.</p>\n<p>I generated candidate pairs for matching within a scene by taking the union of neighbors from multiple image retrieval results. I used the following four global features for image retrieval.</p>\n<ul>\n<li>MASt3R-ASMK (From MASt3R-SfM <a href=\"https://arxiv.org/abs/2409.19152\" target=\"_blank\">https://arxiv.org/abs/2409.19152</a>)</li>\n<li>MASt3R-SPoC (From MASt3R-SfM <a href=\"https://arxiv.org/abs/2409.19152\" target=\"_blank\">https://arxiv.org/abs/2409.19152</a>)</li>\n<li>DINOv2</li>\n<li>ISC (<a href=\"https://arxiv.org/abs/2112.04323\" target=\"_blank\">https://arxiv.org/abs/2112.04323</a>)</li>\n</ul>\n<p>Although MASt3R-ASMK could extract pairs with sufficient coverage, adding the other models slightly improved the score.</p>\n<table>\n<thead>\n<tr>\n<th>Features</th>\n<th>Parameters</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>MASt3R-ASMK</td>\n<td>n=10, k=25 (See MASt3R-SfM paper)</td>\n</tr>\n<tr>\n<td>MASt3R-SPoC</td>\n<td>topk=10</td>\n</tr>\n<tr>\n<td>DINOv2</td>\n<td>topk=10</td>\n</tr>\n<tr>\n<td>ISC</td>\n<td>topk=10</td>\n</tr>\n</tbody>\n</table>\n<p>Note that using only MASt3R-ASMK with heavier settings (e.g., MASt3R-ASMK(n=10, k=60)) can also achieve a high score, but the proposed method is faster while achieving a similar score.</p>\n<h3>2.3. Pairwise matching</h3>\n<p>First, matches for an image pair are computed by MASt3R. This implementation is based on <code>fast_reciprocal_NNs()</code> from the code in the official repository. I used default parameters of subsample=8 and pixel_tol=5.</p>\n<p>In addition to the semi-dense matches, keypoints extracted by other keypoint detectors are also fed to the MASt3R matcher (Inspired by MP-SfM <a href=\"https://arxiv.org/abs/2504.20040)\" target=\"_blank\">https://arxiv.org/abs/2504.20040)</a>. These additional keypoints might regionally overlap with points subsampled by MASt3R, but this approach improved the score compared to using only MASt3R matches.</p>\n<p>I used ALIKED and SuperPoint as the additional keypoint detectors. While I also tried SIFT, GIMSuperPoint, and DaD, I ultimately adopted the combination of ALIKED and SuperPoint because this combination yielded the best LB score.</p>\n<p>The detailed configurations are as follows:</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Parameters</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>MASt3R</td>\n<td>size=512, threshold=1.001</td>\n</tr>\n<tr>\n<td>ALIKED detector</td>\n<td>size=1280, max_keypoints=4096</td>\n</tr>\n<tr>\n<td>SuperPoint detector</td>\n<td>size=1600, max_keypoints=4096, threshold=0.0005</td>\n</tr>\n</tbody>\n</table>\n<h3>2.4. Engineering tips</h3>\n<p>In addition, I applied the following techniques to make the pipeline faster:</p>\n<ul>\n<li>Build the curope (RoPE2D) module with CUDA.</li>\n<li>Replace attention implementations used in mast3r/dust3r/croco with <code>torch.nn.functional.scaled_dot_product_attention</code>.<br>\n(to enable flash attention)</li>\n<li>Fix <code>use_amp</code> args in the MASt3R inference function to be used correctly.</li>\n<li>Use the T4 x 2 environment in Kaggle, and run the submission pipeline in parallel over scene subsets, split by dataset in <code>submission.csv</code>.</li>\n</ul>\n<p>According to Speedy MASt3R (<a href=\"https://arxiv.org/abs/2503.10017)\" target=\"_blank\">https://arxiv.org/abs/2503.10017)</a>, TensorRT can further accelerate the MASt3R model. However, I wasn't able to convert the model.</p>\n<h3>2.5. Local/Public/Private Score</h3>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>amy_gardens</th>\n<th>fbk_vineyard</th>\n<th>ETs</th>\n<th>stairs</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Best submission</td>\n<td>40.65</td>\n<td>72.09</td>\n<td>59.46</td>\n<td>15.89</td>\n<td>52.64</td>\n<td>56.00</td>\n</tr>\n<tr>\n<td>w/ Pre-clustering</td>\n<td>37.02</td>\n<td>47.25</td>\n<td>59.46</td>\n<td>20.35</td>\n<td>50.22</td>\n<td>50.93</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>The scores of <code>amy_gardens</code> and <code>fbk_vineyard</code> seemed to be unstable.</li>\n<li><code>stairs</code> was difficult (As a side note, VGGT achieved the highest score of <code>25.07</code> in my local experiments)</li>\n</ul>\n<h2>3. What did not work</h2>\n<ul>\n<li>GLOMAP: I tried GLOMAP instead of COLMAP with the pre-clustering approach, but it didn't improve the score.</li>\n<li>Coarse-to-Fine matching: Probably, there was a bug in my implementation.</li>\n<li>Using monocular depth estimation: I tried to filter mismatched pairs using depth information.</li>\n<li>VGGT: The tracking head of VGGT with pre-extracted keypoints worked well on local tests, but it resulted in a <code>TimeoutError</code> in submission.</li>\n</ul>\n<h2>Code</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/ns6464/imc2025-1st-place-solution\" target=\"_blank\">Notebook</a></li>\n<li><a href=\"https://github.com/ns-rokuyon/kaggle-image-matching-challenge-2025\" target=\"_blank\">Github</a></li>\n</ul>",
      "rawMarkdown": "Thank you to the organizers and Kaggle team for this exciting competition. Congratulations to all the participants. Although I first participated in IMC in 2022, I'm glad that this time I achieved my best results ever.\n\nThis year, I focused on utilizing recent 3D geometric foundation models such as MASt3R and VGGT. I was surprised at their potential. I respect the authors who developed such great models.\n\n\n## 1. Overview\n\nI developed a simple MASt3R-based pipeline.\n\nAs far as I've tried, image matching using MASt3R's local feature head appears to achieve significantly better results than other detector-based methods such as ALIKED+LG on IMC25.\n\n![overview](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F563483%2Fefd3ae4c091bb1f01f9f9d43a1469d8b%2Ffig1.png?generation=1749042274433083&alt=media)\n\n- Pre-clustering using coarse MASt3R matching on the top-k neighbor images (not shown in the figure).\n- Image pair extraction with the combination of multiple shortlists derived from image retrieval.\n- Utilizing MASt3R (https://github.com/naver/mast3r) as a matcher. In addition to MASt3R semi-dense matches,\n  keypoints extracted by other keypoint detectors are also matched using MASt3R.\n- Generating reconstructions using a general COLMAP pipeline (based on https://www.kaggle.com/code/eduardtrulls/imc25-submission)\n\n\n## 2. Solution\n\nI observed that MASt3R-based matching achieves higher precision (i.e., fewer mismatches) than other methods, so it can match images in messy scenes where detector-based methods typically fail.\nI think the following points are probably useful for robust matching:\n\n- 3D geometric features (From DUSt3R)\n- MASt3R is trained not only with the MegaDepth dataset but also with several object-centric datasets\n\nBy simply using the MASt3R model as a semi-dense detector-free matcher and integrating it into a general COLMAP pipeline, I achieved a score of 42-45 on the PublicLB. Furthermore, I found that increasing the number of image pairs led to a higher score, reaching approximately 50. Therefore, I thought that how image pairs could be increased within limited computational time might be an important consideration.\n\n\n### 2.1. Clustering\n\nI developed a pre-clustering approach based on MASt3R matches, but this approach was not adopted in the final version. It turned out that, since image matching primarily uses MASt3R, there wasn't a significant difference whether I used pre-clustering or multiple reconstructions from COLMAP.\n\nThe pre-clustering (which was not used) approach is outlined below:\n\n1. Initialize the cluster label for each image to -1\n2. Extract N images from the scene using farthest point sampling, and assign each of them a unique cluster label (0, 1, ..., N-1).\n3. For each of the N seed images (from step 2), run 1-vs-all matching with MASt3R against other unclustered images. If an image is matched, assign the seed image's cluster label to the matched image. Otherwise, assign a new cluster label to it.\n4. Use the newly matched images as the next queries, and repeat 1-vs-k matching iteratively until all images are assigned a cluster label\n5. Treat small clusters as \"outliers\"\n\n\n### 2.2. Shortlist\n\nSince the MASt3R matcher is computationally more intensive than detector-based methods, a shortlist of image pairs for matching is still important.\n\nI generated candidate pairs for matching within a scene by taking the union of neighbors from multiple image retrieval results. I used the following four global features for image retrieval.\n\n- MASt3R-ASMK (From MASt3R-SfM https://arxiv.org/abs/2409.19152)\n- MASt3R-SPoC (From MASt3R-SfM https://arxiv.org/abs/2409.19152)\n- DINOv2\n- ISC (https://arxiv.org/abs/2112.04323)\n\nAlthough MASt3R-ASMK could extract pairs with sufficient coverage, adding the other models slightly improved the score.\n\n| Features    | Parameters                        |\n|:------------|:----------------------------------|\n| MASt3R-ASMK | n=10, k=25 (See MASt3R-SfM paper) |\n| MASt3R-SPoC | topk=10                           |\n| DINOv2      | topk=10                           |\n| ISC         | topk=10                           |\n\nNote that using only MASt3R-ASMK with heavier settings (e.g., MASt3R-ASMK(n=10, k=60)) can also achieve a high score, but the proposed method is faster while achieving a similar score.\n\n\n### 2.3. Pairwise matching\n\nFirst, matches for an image pair are computed by MASt3R. This implementation is based on `fast_reciprocal_NNs()` from the code in the official repository. I used default parameters of subsample=8 and pixel_tol=5.\n\nIn addition to the semi-dense matches, keypoints extracted by other keypoint detectors are also fed to the MASt3R matcher (Inspired by MP-SfM https://arxiv.org/abs/2504.20040). These additional keypoints might regionally overlap with points subsampled by MASt3R, but this approach improved the score compared to using only MASt3R matches.\n\nI used ALIKED and SuperPoint as the additional keypoint detectors. While I also tried SIFT, GIMSuperPoint, and DaD, I ultimately adopted the combination of ALIKED and SuperPoint because this combination yielded the best LB score.\n\nThe detailed configurations are as follows:\n\n| Model                   | Parameters                                      |\n|:------------------------|:------------------------------------------------|\n| MASt3R                  | size=512, threshold=1.001                       |\n| ALIKED detector         | size=1280, max_keypoints=4096                   |\n| SuperPoint detector     | size=1600, max_keypoints=4096, threshold=0.0005 |\n\n\n### 2.4. Engineering tips\n\nIn addition, I applied the following techniques to make the pipeline faster:\n\n- Build the curope (RoPE2D) module with CUDA.\n- Replace attention implementations used in mast3r/dust3r/croco with `torch.nn.functional.scaled_dot_product_attention`.\n  (to enable flash attention)\n- Fix `use_amp` args in the MASt3R inference function to be used correctly.\n- Use the T4 x 2 environment in Kaggle, and run the submission pipeline in parallel over scene subsets, split by dataset in `submission.csv`.\n\nAccording to Speedy MASt3R (https://arxiv.org/abs/2503.10017), TensorRT can further accelerate the MASt3R model. However, I wasn't able to convert the model.\n\n\n### 2.5. Local/Public/Private Score\n\n|                   | amy_gardens | fbk_vineyard | ETs    | stairs  | Public  | Private  |\n|:------------------|:------------|:-------------|:-------|:--------|:--------|:---------|\n| Best submission   | 40.65       | 72.09        | 59.46  | 15.89   | 52.64   | 56.00    |\n| w/ Pre-clustering | 37.02       | 47.25        | 59.46  | 20.35   | 50.22   | 50.93    |\n\n- The scores of `amy_gardens` and `fbk_vineyard` seemed to be unstable.\n- `stairs` was difficult (As a side note, VGGT achieved the highest score of `25.07` in my local experiments)\n\n\n## 3. What did not work\n\n- GLOMAP: I tried GLOMAP instead of COLMAP with the pre-clustering approach, but it didn't improve the score.\n- Coarse-to-Fine matching: Probably, there was a bug in my implementation.\n- Using monocular depth estimation: I tried to filter mismatched pairs using depth information.\n- VGGT: The tracking head of VGGT with pre-extracted keypoints worked well on local tests, but it resulted in a `TimeoutError` in submission.\n\n\n## Code\n\n- [Notebook] (https://www.kaggle.com/code/ns6464/imc2025-1st-place-solution)\n- [Github](https://github.com/ns-rokuyon/kaggle-image-matching-challenge-2025)",
      "votes": null
    },
    {
      "id": "3217236",
      "postDate": "06/04/2025 17:31:57",
      "content": "<p>Impressive! Truly worthy of the winning solution.<br>\nDo you have any plans to share the code?</p>",
      "rawMarkdown": "Impressive! Truly worthy of the winning solution.\nDo you have any plans to share the code?",
      "votes": null
    },
    {
      "id": "3217361",
      "postDate": "06/05/2025 00:00:05",
      "content": "<p>Thank you!<br>\nYes, I plan to share the code later.</p>",
      "rawMarkdown": "Thank you!\nYes, I plan to share the code later.",
      "votes": null
    },
    {
      "id": "3217558",
      "postDate": "06/05/2025 06:16:57",
      "content": "<p>Congratulations on winning 1st place! I'm very impressed with your approach.<br>\nAm I correct in understanding that you didn't perform any processing to correct image rotation? In my solution, my PrivateLB dropped from 45.58 to 38.67 when I didn't apply rotation correction using image matching. If you truly didn't perform any rotation correction, I'm just amazed by MASt3R's high matching capability.</p>",
      "rawMarkdown": "Congratulations on winning 1st place! I'm very impressed with your approach.\nAm I correct in understanding that you didn't perform any processing to correct image rotation? In my solution, my PrivateLB dropped from 45.58 to 38.67 when I didn't apply rotation correction using image matching. If you truly didn't perform any rotation correction, I'm just amazed by MASt3R's high matching capability.",
      "votes": null
    },
    {
      "id": "3217625",
      "postDate": "06/05/2025 08:06:14",
      "content": "<p>Thank you for amazing writeup!</p>\n<blockquote>\n  <p>While I also tried SIFT, GIMSuperPoint, and DaD, I ultimately adopted the combination of ALIKED and SuperPoint because this combination yielded the best LB score.</p>\n</blockquote>\n<p>Would you mind please to share those results with other detectors? </p>",
      "rawMarkdown": "Thank you for amazing writeup!\n\n>While I also tried SIFT, GIMSuperPoint, and DaD, I ultimately adopted the combination of ALIKED and SuperPoint because this combination yielded the best LB score.\n\nWould you mind please to share those results with other detectors?",
      "votes": null
    },
    {
      "id": "3217965",
      "postDate": "06/05/2025 15:45:15",
      "content": "<p>Thank you! congrats on 7th place!</p>\n<blockquote>\n  <p>Am I correct in understanding that you didn't perform any processing to correct image rotation?</p>\n</blockquote>\n<p>Yes, you're correct. I didn't use any rotation as preprocessing. I thought rotation wasn't particularly important for the IMC25 dataset after my first submission using MASt3R achieved a good score.</p>",
      "rawMarkdown": "Thank you! congrats on 7th place!\n\n> Am I correct in understanding that you didn't perform any processing to correct image rotation?\n\nYes, you're correct. I didn't use any rotation as preprocessing. I thought rotation wasn't particularly important for the IMC25 dataset after my first submission using MASt3R achieved a good score.",
      "votes": null
    },
    {
      "id": "3217968",
      "postDate": "06/05/2025 15:49:45",
      "content": "<p>Thank you for reading!</p>\n<p>Here are some other results. Unfortunately, there are not ideal ablation studies due to the daily submission limit.</p>\n<table>\n<thead>\n<tr>\n<th>Sparse Detectors</th>\n<th>w/ Pre-clustering</th>\n<th>Shortlist</th>\n<th>Verification</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ALIKED(4k)+DaD(4k)</td>\n<td>No</td>\n<td>ASMK(k=30) SPoC(k=10) DINOv2(k=10) ISC(k=10)</td>\n<td>RANSAC(COLMAP)</td>\n<td>52.41</td>\n<td>55.31</td>\n</tr>\n<tr>\n<td>ALIKED(4k)+SIFT(4k)</td>\n<td>Yes</td>\n<td>ASMK(k=30) SPoC(k=10) DINOv2(k=10) ISC(k=10)</td>\n<td>MAGSAC(OpenCV)</td>\n<td>50.66</td>\n<td>52.10</td>\n</tr>\n<tr>\n<td>SIFT(4k)</td>\n<td>Yes</td>\n<td>ASMK(k=30) SPoC(k=10) DINOv2(k=10) ISC(k=10)</td>\n<td>MAGSAC(OpenCV)</td>\n<td>50.50</td>\n<td>51.82</td>\n</tr>\n<tr>\n<td>ALIKED(2k)+SIFT(2k)+SP(2k)+DaD(2k)</td>\n<td>Yes</td>\n<td>ASMK(k=30) SPoC(k=10) DINOv2(k=10) ISC(k=10)</td>\n<td>MAGSAC(OpenCV)</td>\n<td>50.33</td>\n<td>50.00</td>\n</tr>\n<tr>\n<td>GIMSP(4k)+ALIKED(4k)</td>\n<td>No</td>\n<td>ASMK(k=60)</td>\n<td>MAGSAC(OpenCV)</td>\n<td>49.25</td>\n<td>53.83</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "Thank you for reading!\n\nHere are some other results. Unfortunately, there are not ideal ablation studies due to the daily submission limit.\n\n| Sparse Detectors                   | w/ Pre-clustering | Shortlist                                    | Verification   | Public  | Private  |\n|:-----------------------------------|:------------------|:---------------------------------------------|:---------------|:--------|:---------|\n| ALIKED(4k)+DaD(4k)                 | No                | ASMK(k=30) SPoC(k=10) DINOv2(k=10) ISC(k=10) | RANSAC(COLMAP) | 52.41   | 55.31    |\n| ALIKED(4k)+SIFT(4k)                | Yes               | ASMK(k=30) SPoC(k=10) DINOv2(k=10) ISC(k=10) | MAGSAC(OpenCV) | 50.66   | 52.10    |\n| SIFT(4k)                           | Yes               | ASMK(k=30) SPoC(k=10) DINOv2(k=10) ISC(k=10) | MAGSAC(OpenCV) | 50.50   | 51.82    |\n| ALIKED(2k)+SIFT(2k)+SP(2k)+DaD(2k) | Yes               | ASMK(k=30) SPoC(k=10) DINOv2(k=10) ISC(k=10) | MAGSAC(OpenCV) | 50.33   | 50.00    |\n| GIMSP(4k)+ALIKED(4k)               | No                | ASMK(k=60)                                   | MAGSAC(OpenCV) | 49.25   | 53.83    |",
      "votes": null
    },
    {
      "id": "3218171",
      "postDate": "06/05/2025 23:39:55",
      "content": "<p>Congrats on winning 1st place and thanks for sharing your solution.<br>\nI learned a lot from this especially how you used MASt3R's for image matching and added keypoints from ALIKED and SuperPoint to improve results. Your engineering tips were also very helpful.<br>\ni'm new to 3D matching, so it was great to see how you handled performance and scoring.<br>\nJust wondering, How did you decide which keypoint  detectors to combime? Did some combinations not work well?<br>\nThanks again and great job!</p>",
      "rawMarkdown": "Congrats on winning 1st place and thanks for sharing your solution.\nI learned a lot from this especially how you used MASt3R's for image matching and added keypoints from ALIKED and SuperPoint to improve results. Your engineering tips were also very helpful.\ni'm new to 3D matching, so it was great to see how you handled performance and scoring.\nJust wondering, How did you decide which keypoint  detectors to combime? Did some combinations not work well?\nThanks again and great job!",
      "votes": null
    },
    {
      "id": "3218631",
      "postDate": "06/06/2025 13:17:02",
      "content": "<p>Congratulations on 1st place!<br>\nI was truly impressed by your simple yet innovative solution leveraging MASt3r.</p>\n<p>In relation to this question, if you happen to know: what was the score without adding keypoints from ALIKED and SuperPoint?<br>\nIn other words, how much did incorporating keypoints from ALIKED and SuperPoint actually contribute to the final score?</p>\n<p>Also, I’m really looking forward to the release of your source code.<br>\nI’m especially interested in how you implemented passing keypoints detected by other detectors into MASt3r.</p>",
      "rawMarkdown": "Congratulations on 1st place!\nI was truly impressed by your simple yet innovative solution leveraging MASt3r.\n\nIn relation to this question, if you happen to know: what was the score without adding keypoints from ALIKED and SuperPoint?\nIn other words, how much did incorporating keypoints from ALIKED and SuperPoint actually contribute to the final score?\n\nAlso, I’m really looking forward to the release of your source code.\nI’m especially interested in how you implemented passing keypoints detected by other detectors into MASt3r.",
      "votes": null
    },
    {
      "id": "3219108",
      "postDate": "06/07/2025 07:12:38",
      "content": "<p>Congratulations！I'm very appreciate for the impressive solution, learnt a lot! 👍</p>",
      "rawMarkdown": "Congratulations！I'm very appreciate for the impressive solution, learnt a lot! 👍",
      "votes": null
    },
    {
      "id": "3219292",
      "postDate": "06/07/2025 12:50:18",
      "content": "<p>Thank you!<br>\nActually, the contribution of the additional keypoints was relatively small, leading to a score increase of only approximately 0.5 to 1.0 points.</p>\n<blockquote>\n  <p>I’m especially interested in how you implemented passing keypoints detected by other detectors into MASt3r.</p>\n</blockquote>\n<p>I used MP-SfM implementations below. <a href=\"https://github.com/cvg/mpsfm/blob/main/mpsfm/extraction/pairwise/models/mast3r.py#L112\" target=\"_blank\">https://github.com/cvg/mpsfm/blob/main/mpsfm/extraction/pairwise/models/mast3r.py#L112</a></p>",
      "rawMarkdown": "Thank you!\nActually, the contribution of the additional keypoints was relatively small, leading to a score increase of only approximately 0.5 to 1.0 points.\n\n> I’m especially interested in how you implemented passing keypoints detected by other detectors into MASt3r.\n\nI used MP-SfM implementations below. https://github.com/cvg/mpsfm/blob/main/mpsfm/extraction/pairwise/models/mast3r.py#L112",
      "votes": null
    },
    {
      "id": "3219298",
      "postDate": "06/07/2025 12:58:12",
      "content": "<p>Thank you!<br>\nI selected the keypoint detectors based roughly on PublicLB score.<br>\nYou can see several combination results here: <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2025/discussion/583058#3217968\" target=\"_blank\">https://www.kaggle.com/competitions/image-matching-challenge-2025/discussion/583058#3217968</a></p>",
      "rawMarkdown": "Thank you!\nI selected the keypoint detectors based roughly on PublicLB score.\nYou can see several combination results here: https://www.kaggle.com/competitions/image-matching-challenge-2025/discussion/583058#3217968",
      "votes": null
    },
    {
      "id": "3220131",
      "postDate": "06/08/2025 20:52:54",
      "content": "<p>Thank you so much for the quick reply and for sharing the detailed results link!</p>\n<p>The table you posted comparing different keypoint detector combinations is really helpful. It’s interesting to see how ALIKED + SuperPoint gave the best balance of accuracy and speed.</p>\n<p>I’ll explore the combinations further — really appreciate the transparency in sharing all this!</p>",
      "rawMarkdown": "Thank you so much for the quick reply and for sharing the detailed results link!\n\nThe table you posted comparing different keypoint detector combinations is really helpful. It’s interesting to see how ALIKED + SuperPoint gave the best balance of accuracy and speed.\n\nI’ll explore the combinations further — really appreciate the transparency in sharing all this!",
      "votes": null
    },
    {
      "id": "3220547",
      "postDate": "06/09/2025 13:44:33",
      "content": "<p>Impressive! Thanks for sharing.</p>",
      "rawMarkdown": "Impressive! Thanks for sharing.",
      "votes": null
    },
    {
      "id": "3221750",
      "postDate": "06/11/2025 12:01:46",
      "content": "<p>Congrats and thanks for the write-up.</p>\n<p>Would this pipeline also work using only ALIKED/SuperPoint keypoints + the MASt3R Matcher (ie not using the dense MASt3R matches)?</p>",
      "rawMarkdown": "Congrats and thanks for the write-up.\n\nWould this pipeline also work using only ALIKED/SuperPoint keypoints + the MASt3R Matcher (ie not using the dense MASt3R matches)?",
      "votes": null
    },
    {
      "id": "3222200",
      "postDate": "06/12/2025 02:27:19",
      "content": "<p>Congrats and thanks for the write-up.</p>",
      "rawMarkdown": "Congrats and thanks for the write-up.",
      "votes": null
    },
    {
      "id": "3222715",
      "postDate": "06/12/2025 12:29:08",
      "content": "<p>great work done</p>",
      "rawMarkdown": "great work done",
      "votes": null
    },
    {
      "id": "3222761",
      "postDate": "06/12/2025 13:32:52",
      "content": "<p>Thank you!<br>\nUsing only ALIKED/SuperPoint keypoints with the MASt3R Matcher performed worse than also using the dense MASt3R matches.<br>\nHere are the results (Please note that these results are for the ALIKED+DaD combination):</p>\n<table>\n<thead>\n<tr>\n<th>MASt3R matches</th>\n<th>Sparse keypoints</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Dense+Sparse</td>\n<td>ALIKED+DaD</td>\n<td>52.41</td>\n<td>55.31</td>\n</tr>\n<tr>\n<td>Sparse</td>\n<td>ALIKED+DaD</td>\n<td>48.59</td>\n<td>53.79</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "Thank you!\nUsing only ALIKED/SuperPoint keypoints with the MASt3R Matcher performed worse than also using the dense MASt3R matches.\nHere are the results (Please note that these results are for the ALIKED+DaD combination):\n\n| MASt3R matches |  Sparse keypoints  |  Public  | Private  |\n|:---|:------------------|:--------|:---------|\n| Dense+Sparse | ALIKED+DaD  | 52.41  |  55.31  |\n| Sparse | ALIKED+DaD  | 48.59  | 53.79 |",
      "votes": null
    },
    {
      "id": "3236026",
      "postDate": "06/29/2025 22:44:57",
      "content": "<p>Congratulations on winning 1st place! I'm very impressed with your approach.This is a very Great way to sort libraries and use them</p>",
      "rawMarkdown": "Congratulations on winning 1st place! I'm very impressed with your approach.This is a very Great way to sort libraries and use them",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3217236,
      "author_name": "tmyok1984",
      "author_url": "",
      "post_date": "06/04/2025 17:31:57",
      "content": "<p>Impressive! Truly worthy of the winning solution.<br>\nDo you have any plans to share the code?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3217361,
          "author_name": "ns6464",
          "author_url": "",
          "post_date": "06/05/2025 00:00:05",
          "content": "<p>Thank you!<br>\nYes, I plan to share the code later.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3219108,
              "author_name": "",
              "author_url": "",
              "post_date": "06/07/2025 07:12:38",
              "content": "<p>Congratulations！I'm very appreciate for the impressive solution, learnt a lot! 👍</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3217558,
      "author_name": "confirm",
      "author_url": "",
      "post_date": "06/05/2025 06:16:57",
      "content": "<p>Congratulations on winning 1st place! I'm very impressed with your approach.<br>\nAm I correct in understanding that you didn't perform any processing to correct image rotation? In my solution, my PrivateLB dropped from 45.58 to 38.67 when I didn't apply rotation correction using image matching. If you truly didn't perform any rotation correction, I'm just amazed by MASt3R's high matching capability.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3217965,
          "author_name": "ns6464",
          "author_url": "",
          "post_date": "06/05/2025 15:45:15",
          "content": "<p>Thank you! congrats on 7th place!</p>\n<blockquote>\n  <p>Am I correct in understanding that you didn't perform any processing to correct image rotation?</p>\n</blockquote>\n<p>Yes, you're correct. I didn't use any rotation as preprocessing. I thought rotation wasn't particularly important for the IMC25 dataset after my first submission using MASt3R achieved a good score.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3217625,
      "author_name": "oldufo",
      "author_url": "",
      "post_date": "06/05/2025 08:06:14",
      "content": "<p>Thank you for amazing writeup!</p>\n<blockquote>\n  <p>While I also tried SIFT, GIMSuperPoint, and DaD, I ultimately adopted the combination of ALIKED and SuperPoint because this combination yielded the best LB score.</p>\n</blockquote>\n<p>Would you mind please to share those results with other detectors? </p>",
      "votes": null,
      "replies": [
        {
          "id": 3217968,
          "author_name": "ns6464",
          "author_url": "",
          "post_date": "06/05/2025 15:49:45",
          "content": "<p>Thank you for reading!</p>\n<p>Here are some other results. Unfortunately, there are not ideal ablation studies due to the daily submission limit.</p>\n<table>\n<thead>\n<tr>\n<th>Sparse Detectors</th>\n<th>w/ Pre-clustering</th>\n<th>Shortlist</th>\n<th>Verification</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ALIKED(4k)+DaD(4k)</td>\n<td>No</td>\n<td>ASMK(k=30) SPoC(k=10) DINOv2(k=10) ISC(k=10)</td>\n<td>RANSAC(COLMAP)</td>\n<td>52.41</td>\n<td>55.31</td>\n</tr>\n<tr>\n<td>ALIKED(4k)+SIFT(4k)</td>\n<td>Yes</td>\n<td>ASMK(k=30) SPoC(k=10) DINOv2(k=10) ISC(k=10)</td>\n<td>MAGSAC(OpenCV)</td>\n<td>50.66</td>\n<td>52.10</td>\n</tr>\n<tr>\n<td>SIFT(4k)</td>\n<td>Yes</td>\n<td>ASMK(k=30) SPoC(k=10) DINOv2(k=10) ISC(k=10)</td>\n<td>MAGSAC(OpenCV)</td>\n<td>50.50</td>\n<td>51.82</td>\n</tr>\n<tr>\n<td>ALIKED(2k)+SIFT(2k)+SP(2k)+DaD(2k)</td>\n<td>Yes</td>\n<td>ASMK(k=30) SPoC(k=10) DINOv2(k=10) ISC(k=10)</td>\n<td>MAGSAC(OpenCV)</td>\n<td>50.33</td>\n<td>50.00</td>\n</tr>\n<tr>\n<td>GIMSP(4k)+ALIKED(4k)</td>\n<td>No</td>\n<td>ASMK(k=60)</td>\n<td>MAGSAC(OpenCV)</td>\n<td>49.25</td>\n<td>53.83</td>\n</tr>\n</tbody>\n</table>",
          "votes": null,
          "replies": [
            {
              "id": 3218631,
              "author_name": "kashiwaba",
              "author_url": "",
              "post_date": "06/06/2025 13:17:02",
              "content": "<p>Congratulations on 1st place!<br>\nI was truly impressed by your simple yet innovative solution leveraging MASt3r.</p>\n<p>In relation to this question, if you happen to know: what was the score without adding keypoints from ALIKED and SuperPoint?<br>\nIn other words, how much did incorporating keypoints from ALIKED and SuperPoint actually contribute to the final score?</p>\n<p>Also, I’m really looking forward to the release of your source code.<br>\nI’m especially interested in how you implemented passing keypoints detected by other detectors into MASt3r.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3219292,
                  "author_name": "ns6464",
                  "author_url": "",
                  "post_date": "06/07/2025 12:50:18",
                  "content": "<p>Thank you!<br>\nActually, the contribution of the additional keypoints was relatively small, leading to a score increase of only approximately 0.5 to 1.0 points.</p>\n<blockquote>\n  <p>I’m especially interested in how you implemented passing keypoints detected by other detectors into MASt3r.</p>\n</blockquote>\n<p>I used MP-SfM implementations below. <a href=\"https://github.com/cvg/mpsfm/blob/main/mpsfm/extraction/pairwise/models/mast3r.py#L112\" target=\"_blank\">https://github.com/cvg/mpsfm/blob/main/mpsfm/extraction/pairwise/models/mast3r.py#L112</a></p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3218171,
      "author_name": "mhamzashafqat",
      "author_url": "",
      "post_date": "06/05/2025 23:39:55",
      "content": "<p>Congrats on winning 1st place and thanks for sharing your solution.<br>\nI learned a lot from this especially how you used MASt3R's for image matching and added keypoints from ALIKED and SuperPoint to improve results. Your engineering tips were also very helpful.<br>\ni'm new to 3D matching, so it was great to see how you handled performance and scoring.<br>\nJust wondering, How did you decide which keypoint  detectors to combime? Did some combinations not work well?<br>\nThanks again and great job!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3219298,
          "author_name": "ns6464",
          "author_url": "",
          "post_date": "06/07/2025 12:58:12",
          "content": "<p>Thank you!<br>\nI selected the keypoint detectors based roughly on PublicLB score.<br>\nYou can see several combination results here: <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2025/discussion/583058#3217968\" target=\"_blank\">https://www.kaggle.com/competitions/image-matching-challenge-2025/discussion/583058#3217968</a></p>",
          "votes": null,
          "replies": [
            {
              "id": 3220131,
              "author_name": "mhamzashafqat",
              "author_url": "",
              "post_date": "06/08/2025 20:52:54",
              "content": "<p>Thank you so much for the quick reply and for sharing the detailed results link!</p>\n<p>The table you posted comparing different keypoint detector combinations is really helpful. It’s interesting to see how ALIKED + SuperPoint gave the best balance of accuracy and speed.</p>\n<p>I’ll explore the combinations further — really appreciate the transparency in sharing all this!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3220547,
      "author_name": "aniketpotabatti",
      "author_url": "",
      "post_date": "06/09/2025 13:44:33",
      "content": "<p>Impressive! Thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3221750,
      "author_name": "danielkorth",
      "author_url": "",
      "post_date": "06/11/2025 12:01:46",
      "content": "<p>Congrats and thanks for the write-up.</p>\n<p>Would this pipeline also work using only ALIKED/SuperPoint keypoints + the MASt3R Matcher (ie not using the dense MASt3R matches)?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3222761,
          "author_name": "ns6464",
          "author_url": "",
          "post_date": "06/12/2025 13:32:52",
          "content": "<p>Thank you!<br>\nUsing only ALIKED/SuperPoint keypoints with the MASt3R Matcher performed worse than also using the dense MASt3R matches.<br>\nHere are the results (Please note that these results are for the ALIKED+DaD combination):</p>\n<table>\n<thead>\n<tr>\n<th>MASt3R matches</th>\n<th>Sparse keypoints</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Dense+Sparse</td>\n<td>ALIKED+DaD</td>\n<td>52.41</td>\n<td>55.31</td>\n</tr>\n<tr>\n<td>Sparse</td>\n<td>ALIKED+DaD</td>\n<td>48.59</td>\n<td>53.79</td>\n</tr>\n</tbody>\n</table>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3222200,
      "author_name": "dlresearch",
      "author_url": "",
      "post_date": "06/12/2025 02:27:19",
      "content": "<p>Congrats and thanks for the write-up.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3222715,
      "author_name": "adityapratapsingh5",
      "author_url": "",
      "post_date": "06/12/2025 12:29:08",
      "content": "<p>great work done</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3236026,
      "author_name": "abubakaridrees",
      "author_url": "",
      "post_date": "06/29/2025 22:44:57",
      "content": "<p>Congratulations on winning 1st place! I'm very impressed with your approach.This is a very Great way to sort libraries and use them</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3217069": "Thank you to the organizers and Kaggle team for this exciting competition. Congratulations to all the participants. Although I first participated in IMC in 2022, I'm glad that this time I achieved my best results ever.\n\nThis year, I focused on utilizing recent 3D geometric foundation models such as MASt3R and VGGT. I was surprised at their potential. I respect the authors who developed such great models.\n\n\n## 1. Overview\n\nI developed a simple MASt3R-based pipeline.\n\nAs far as I've tried, image matching using MASt3R's local feature head appears to achieve significantly better results than other detector-based methods such as ALIKED+LG on IMC25.\n\n![overview](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F563483%2Fefd3ae4c091bb1f01f9f9d43a1469d8b%2Ffig1.png?generation=1749042274433083&alt=media)\n\n- Pre-clustering using coarse MASt3R matching on the top-k neighbor images (not shown in the figure).\n- Image pair extraction with the combination of multiple shortlists derived from image retrieval.\n- Utilizing MASt3R (https://github.com/naver/mast3r) as a matcher. In addition to MASt3R semi-dense matches,\n  keypoints extracted by other keypoint detectors are also matched using MASt3R.\n- Generating reconstructions using a general COLMAP pipeline (based on https://www.kaggle.com/code/eduardtrulls/imc25-submission)\n\n\n## 2. Solution\n\nI observed that MASt3R-based matching achieves higher precision (i.e., fewer mismatches) than other methods, so it can match images in messy scenes where detector-based methods typically fail.\nI think the following points are probably useful for robust matching:\n\n- 3D geometric features (From DUSt3R)\n- MASt3R is trained not only with the MegaDepth dataset but also with several object-centric datasets\n\nBy simply using the MASt3R model as a semi-dense detector-free matcher and integrating it into a general COLMAP pipeline, I achieved a score of 42-45 on the PublicLB. Furthermore, I found that increasing the number of image pairs led to a higher score, reaching approximately 50. Therefore, I thought that how image pairs could be increased within limited computational time might be an important consideration.\n\n\n### 2.1. Clustering\n\nI developed a pre-clustering approach based on MASt3R matches, but this approach was not adopted in the final version. It turned out that, since image matching primarily uses MASt3R, there wasn't a significant difference whether I used pre-clustering or multiple reconstructions from COLMAP.\n\nThe pre-clustering (which was not used) approach is outlined below:\n\n1. Initialize the cluster label for each image to -1\n2. Extract N images from the scene using farthest point sampling, and assign each of them a unique cluster label (0, 1, ..., N-1).\n3. For each of the N seed images (from step 2), run 1-vs-all matching with MASt3R against other unclustered images. If an image is matched, assign the seed image's cluster label to the matched image. Otherwise, assign a new cluster label to it.\n4. Use the newly matched images as the next queries, and repeat 1-vs-k matching iteratively until all images are assigned a cluster label\n5. Treat small clusters as \"outliers\"\n\n\n### 2.2. Shortlist\n\nSince the MASt3R matcher is computationally more intensive than detector-based methods, a shortlist of image pairs for matching is still important.\n\nI generated candidate pairs for matching within a scene by taking the union of neighbors from multiple image retrieval results. I used the following four global features for image retrieval.\n\n- MASt3R-ASMK (From MASt3R-SfM https://arxiv.org/abs/2409.19152)\n- MASt3R-SPoC (From MASt3R-SfM https://arxiv.org/abs/2409.19152)\n- DINOv2\n- ISC (https://arxiv.org/abs/2112.04323)\n\nAlthough MASt3R-ASMK could extract pairs with sufficient coverage, adding the other models slightly improved the score.\n\n| Features    | Parameters                        |\n|:------------|:----------------------------------|\n| MASt3R-ASMK | n=10, k=25 (See MASt3R-SfM paper) |\n| MASt3R-SPoC | topk=10                           |\n| DINOv2      | topk=10                           |\n| ISC         | topk=10                           |\n\nNote that using only MASt3R-ASMK with heavier settings (e.g., MASt3R-ASMK(n=10, k=60)) can also achieve a high score, but the proposed method is faster while achieving a similar score.\n\n\n### 2.3. Pairwise matching\n\nFirst, matches for an image pair are computed by MASt3R. This implementation is based on `fast_reciprocal_NNs()` from the code in the official repository. I used default parameters of subsample=8 and pixel_tol=5.\n\nIn addition to the semi-dense matches, keypoints extracted by other keypoint detectors are also fed to the MASt3R matcher (Inspired by MP-SfM https://arxiv.org/abs/2504.20040). These additional keypoints might regionally overlap with points subsampled by MASt3R, but this approach improved the score compared to using only MASt3R matches.\n\nI used ALIKED and SuperPoint as the additional keypoint detectors. While I also tried SIFT, GIMSuperPoint, and DaD, I ultimately adopted the combination of ALIKED and SuperPoint because this combination yielded the best LB score.\n\nThe detailed configurations are as follows:\n\n| Model                   | Parameters                                      |\n|:------------------------|:------------------------------------------------|\n| MASt3R                  | size=512, threshold=1.001                       |\n| ALIKED detector         | size=1280, max_keypoints=4096                   |\n| SuperPoint detector     | size=1600, max_keypoints=4096, threshold=0.0005 |\n\n\n### 2.4. Engineering tips\n\nIn addition, I applied the following techniques to make the pipeline faster:\n\n- Build the curope (RoPE2D) module with CUDA.\n- Replace attention implementations used in mast3r/dust3r/croco with `torch.nn.functional.scaled_dot_product_attention`.\n  (to enable flash attention)\n- Fix `use_amp` args in the MASt3R inference function to be used correctly.\n- Use the T4 x 2 environment in Kaggle, and run the submission pipeline in parallel over scene subsets, split by dataset in `submission.csv`.\n\nAccording to Speedy MASt3R (https://arxiv.org/abs/2503.10017), TensorRT can further accelerate the MASt3R model. However, I wasn't able to convert the model.\n\n\n### 2.5. Local/Public/Private Score\n\n|                   | amy_gardens | fbk_vineyard | ETs    | stairs  | Public  | Private  |\n|:------------------|:------------|:-------------|:-------|:--------|:--------|:---------|\n| Best submission   | 40.65       | 72.09        | 59.46  | 15.89   | 52.64   | 56.00    |\n| w/ Pre-clustering | 37.02       | 47.25        | 59.46  | 20.35   | 50.22   | 50.93    |\n\n- The scores of `amy_gardens` and `fbk_vineyard` seemed to be unstable.\n- `stairs` was difficult (As a side note, VGGT achieved the highest score of `25.07` in my local experiments)\n\n\n## 3. What did not work\n\n- GLOMAP: I tried GLOMAP instead of COLMAP with the pre-clustering approach, but it didn't improve the score.\n- Coarse-to-Fine matching: Probably, there was a bug in my implementation.\n- Using monocular depth estimation: I tried to filter mismatched pairs using depth information.\n- VGGT: The tracking head of VGGT with pre-extracted keypoints worked well on local tests, but it resulted in a `TimeoutError` in submission.\n\n\n## Code\n\n- [Notebook] (https://www.kaggle.com/code/ns6464/imc2025-1st-place-solution)\n- [Github](https://github.com/ns-rokuyon/kaggle-image-matching-challenge-2025)",
    "3217236": "Impressive! Truly worthy of the winning solution.\nDo you have any plans to share the code?",
    "3217361": "Thank you!\nYes, I plan to share the code later.",
    "3217558": "Congratulations on winning 1st place! I'm very impressed with your approach.\nAm I correct in understanding that you didn't perform any processing to correct image rotation? In my solution, my PrivateLB dropped from 45.58 to 38.67 when I didn't apply rotation correction using image matching. If you truly didn't perform any rotation correction, I'm just amazed by MASt3R's high matching capability.",
    "3217625": "Thank you for amazing writeup!\n\n>While I also tried SIFT, GIMSuperPoint, and DaD, I ultimately adopted the combination of ALIKED and SuperPoint because this combination yielded the best LB score.\n\nWould you mind please to share those results with other detectors?",
    "3217965": "Thank you! congrats on 7th place!\n\n> Am I correct in understanding that you didn't perform any processing to correct image rotation?\n\nYes, you're correct. I didn't use any rotation as preprocessing. I thought rotation wasn't particularly important for the IMC25 dataset after my first submission using MASt3R achieved a good score.",
    "3217968": "Thank you for reading!\n\nHere are some other results. Unfortunately, there are not ideal ablation studies due to the daily submission limit.\n\n| Sparse Detectors                   | w/ Pre-clustering | Shortlist                                    | Verification   | Public  | Private  |\n|:-----------------------------------|:------------------|:---------------------------------------------|:---------------|:--------|:---------|\n| ALIKED(4k)+DaD(4k)                 | No                | ASMK(k=30) SPoC(k=10) DINOv2(k=10) ISC(k=10) | RANSAC(COLMAP) | 52.41   | 55.31    |\n| ALIKED(4k)+SIFT(4k)                | Yes               | ASMK(k=30) SPoC(k=10) DINOv2(k=10) ISC(k=10) | MAGSAC(OpenCV) | 50.66   | 52.10    |\n| SIFT(4k)                           | Yes               | ASMK(k=30) SPoC(k=10) DINOv2(k=10) ISC(k=10) | MAGSAC(OpenCV) | 50.50   | 51.82    |\n| ALIKED(2k)+SIFT(2k)+SP(2k)+DaD(2k) | Yes               | ASMK(k=30) SPoC(k=10) DINOv2(k=10) ISC(k=10) | MAGSAC(OpenCV) | 50.33   | 50.00    |\n| GIMSP(4k)+ALIKED(4k)               | No                | ASMK(k=60)                                   | MAGSAC(OpenCV) | 49.25   | 53.83    |",
    "3218171": "Congrats on winning 1st place and thanks for sharing your solution.\nI learned a lot from this especially how you used MASt3R's for image matching and added keypoints from ALIKED and SuperPoint to improve results. Your engineering tips were also very helpful.\ni'm new to 3D matching, so it was great to see how you handled performance and scoring.\nJust wondering, How did you decide which keypoint  detectors to combime? Did some combinations not work well?\nThanks again and great job!",
    "3218631": "Congratulations on 1st place!\nI was truly impressed by your simple yet innovative solution leveraging MASt3r.\n\nIn relation to this question, if you happen to know: what was the score without adding keypoints from ALIKED and SuperPoint?\nIn other words, how much did incorporating keypoints from ALIKED and SuperPoint actually contribute to the final score?\n\nAlso, I’m really looking forward to the release of your source code.\nI’m especially interested in how you implemented passing keypoints detected by other detectors into MASt3r.",
    "3219108": "Congratulations！I'm very appreciate for the impressive solution, learnt a lot! 👍",
    "3219292": "Thank you!\nActually, the contribution of the additional keypoints was relatively small, leading to a score increase of only approximately 0.5 to 1.0 points.\n\n> I’m especially interested in how you implemented passing keypoints detected by other detectors into MASt3r.\n\nI used MP-SfM implementations below. https://github.com/cvg/mpsfm/blob/main/mpsfm/extraction/pairwise/models/mast3r.py#L112",
    "3219298": "Thank you!\nI selected the keypoint detectors based roughly on PublicLB score.\nYou can see several combination results here: https://www.kaggle.com/competitions/image-matching-challenge-2025/discussion/583058#3217968",
    "3220131": "Thank you so much for the quick reply and for sharing the detailed results link!\n\nThe table you posted comparing different keypoint detector combinations is really helpful. It’s interesting to see how ALIKED + SuperPoint gave the best balance of accuracy and speed.\n\nI’ll explore the combinations further — really appreciate the transparency in sharing all this!",
    "3220547": "Impressive! Thanks for sharing.",
    "3221750": "Congrats and thanks for the write-up.\n\nWould this pipeline also work using only ALIKED/SuperPoint keypoints + the MASt3R Matcher (ie not using the dense MASt3R matches)?",
    "3222200": "Congrats and thanks for the write-up.",
    "3222715": "great work done",
    "3222761": "Thank you!\nUsing only ALIKED/SuperPoint keypoints with the MASt3R Matcher performed worse than also using the dense MASt3R matches.\nHere are the results (Please note that these results are for the ALIKED+DaD combination):\n\n| MASt3R matches |  Sparse keypoints  |  Public  | Private  |\n|:---|:------------------|:--------|:---------|\n| Dense+Sparse | ALIKED+DaD  | 52.41  |  55.31  |\n| Sparse | ALIKED+DaD  | 48.59  | 53.79 |",
    "3236026": "Congratulations on winning 1st place! I'm very impressed with your approach.This is a very Great way to sort libraries and use them"
  },
  "source": "meta"
}