{
  "id": 583097,
  "title": "11th place solution",
  "url": "/competitions/image-matching-challenge-2025/discussion/583097",
  "author_name": "motono0223",
  "post_date": "2025-06-04T16:58:30.891000",
  "votes": 33,
  "comment_count": 2,
  "views": 0,
  "content": "<h1>Overview</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8163878%2Fa1a8238d8943872b2cd95d87778af334%2FScreenshot%202025-06-05%2001.46.42.png?generation=1749055688789666&amp;alt=media\" alt=\"\"></p>\n<h1>Notebook</h1>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/motono0223/imc2025-11th-place-solution\" target=\"_blank\">https://www.kaggle.com/code/motono0223/imc2025-11th-place-solution</a><ul>\n<li>private lb = 43.34</li>\n<li>public lb = 43.53</li></ul></li>\n</ul>\n<h1>Descriptions</h1>\n<h2>Create image pairs</h2>\n<ul>\n<li>global feature extractor: DinoV2/Base</li>\n<li>Topk = 150</li>\n</ul>\n<h2>Rerank image pairs</h2>\n<ul>\n<li><p>Due to the large number of image pairs, it is necessary to reduce the processing time for image matching.</p></li>\n<li><p>To address this, image pairs are re-ranked and filtered using the ALIKED model (local features) in this process.</p></li>\n<li><p>Specifically, the procedure is as follows:</p>\n<ul>\n<li>For each image, local descriptors are extracted using the ALIKED model.<ul>\n<li>The image size is set to 2048 pixels, with 2048 keypoints and a detection threshold of 0.01.</li>\n<li>All descriptors are normalized with norm=1.</li></ul></li>\n<li>A Euclidean distance matrix is computed between the local descriptors of each image pair.<ul>\n<li>If image1 has N descriptors and image2 has M descriptors, an NxM distance matrix is calculated using the GPU.</li></ul></li>\n<li>For each row in the distance matrix, the minimum value is extracted (resulting in N minimum values).</li>\n<li>If these minimum values are below the threshold of 1.0, they are counted as valid matches.<br>\nImage pairs with 15 or fewer valid matches are discarded.</li></ul></li>\n<li><p>The key point is that since LightGlue (GNN) is computationally expensive, a fast simulation of image matching is performed using the distance matrix as a lightweight approximation.</p></li>\n</ul>\n<h2>Image Matching</h2>\n<ul>\n<li>Local feature extractor: Aliked (n16)<ul>\n<li>resize_to: 1536 pixel</li>\n<li>max_num_keypoints: 8192 points</li>\n<li>detection_threshold: 0.01</li></ul></li>\n<li>Matcher: LighGlue</li>\n<li>post process<ul>\n<li>the number of matches &gt; 15 matches</li></ul></li>\n</ul>\n<h2>Filter image pairs</h2>\n<ul>\n<li><p>This step filters image pairs whose matches are not critical for SfM. The pruning accelerates the reconstruction and gives COLMAP’s incremental mapping more opportunities to reconstruct small clusters.</p></li>\n<li><p>Image-matching results can be expressed as a network in which images are nodes and match counts are edges; the number of edges connected to each node can then be counted.</p></li>\n<li><p>Non-essential pairs are removed by capping the number of edges per image node (threshold = 6). If a node has more than six edges, those with the fewest matches are discarded until only six remain.</p>\n<ul>\n<li>In COLMAP’s incremental mapping, an initial pair is chosen and additional images are registered one by one; the initial pair is selected statistically from the matches stored in the database.[1]</li>\n<li>Match results sometimes concentrate on a handful of images, causing the incremental mapping output to favor certain clusters.</li>\n<li>Pruning pairs connected to images with many matches increases the likelihood that initial pairs will be formed inside smaller clusters.</li></ul></li>\n</ul>\n<h2>Rough clustering</h2>\n<ul>\n<li>A lightweight clustering routine is implemented with NetworkX.</li>\n<li>For every image pair, the number of matches is read from <code>matches.h5</code>; edges with 150 or more matches are added to <code>G = nx.Graph()</code>.</li>\n<li>Clusters are then obtained via <code>nx.weakly_connected_components(G)</code>.</li>\n</ul>\n<h2>colmap (exhaustive matching)</h2>\n<ul>\n<li>Computation is performed with <strong>pycolmap</strong>’s <code>exhaustive_matching</code>.</li>\n<li>All settings remain at their defaults (pycolmap 0.6.1).</li>\n</ul>\n<h2>colmap (incremental mapping)</h2>\n<ul>\n<li>For each clustered image set, <code>incremental_mapping</code> in <strong>pycolmap</strong> is executed four times.</li>\n<li>Parameters:</li>\n</ul>\n<pre><code>mapper_options = pycolmap.IncrementalPipelineOptions()\nmapper_options.multiple_models   = \nmapper_options.min_model_size    = \nmapper_options.max_num_models = \nmapper_options.init_num_trials   = \n</code></pre>\n<h2>Merge maps</h2>\n<ul>\n<li>For every cluster, maps are merged.</li>\n<li>Maps are selected for merging when:<ul>\n<li>Six or more images are shared between the maps, and</li>\n<li>Each map’s <code>mean_reproj_error</code> is ≤ 2.0.</li></ul></li>\n<li>Merging is carried out with <strong>COLMAP</strong>’s native APIs (not pycolmap):<ul>\n<li><code>model_merger</code> combines the models.</li>\n<li><code>bundle_adjuster</code> optimizes camera parameters.<ul>\n<li>--BundleAdjustment.refine_focal_length = 1</li>\n<li>--BundleAdjustment.refine_principal_point = 1</li>\n<li>--BundleAdjustment.refine_extra_params = 1</li></ul></li></ul></li>\n</ul>\n<p>[1] pycolmap.incremental mapping()<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8163878%2F19d0e899b13ff451ce99c550dea94422%2FScreenshot%202025-06-05%2001.35.24.png?generation=1749055849424519&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 3217205,
      "postDate": "2025-06-04T16:58:30.890Z",
      "content": "<h1>Overview</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8163878%2Fa1a8238d8943872b2cd95d87778af334%2FScreenshot%202025-06-05%2001.46.42.png?generation=1749055688789666&amp;alt=media\" alt=\"\"></p>\n<h1>Notebook</h1>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/motono0223/imc2025-11th-place-solution\" target=\"_blank\">https://www.kaggle.com/code/motono0223/imc2025-11th-place-solution</a><ul>\n<li>private lb = 43.34</li>\n<li>public lb = 43.53</li></ul></li>\n</ul>\n<h1>Descriptions</h1>\n<h2>Create image pairs</h2>\n<ul>\n<li>global feature extractor: DinoV2/Base</li>\n<li>Topk = 150</li>\n</ul>\n<h2>Rerank image pairs</h2>\n<ul>\n<li><p>Due to the large number of image pairs, it is necessary to reduce the processing time for image matching.</p></li>\n<li><p>To address this, image pairs are re-ranked and filtered using the ALIKED model (local features) in this process.</p></li>\n<li><p>Specifically, the procedure is as follows:</p>\n<ul>\n<li>For each image, local descriptors are extracted using the ALIKED model.<ul>\n<li>The image size is set to 2048 pixels, with 2048 keypoints and a detection threshold of 0.01.</li>\n<li>All descriptors are normalized with norm=1.</li></ul></li>\n<li>A Euclidean distance matrix is computed between the local descriptors of each image pair.<ul>\n<li>If image1 has N descriptors and image2 has M descriptors, an NxM distance matrix is calculated using the GPU.</li></ul></li>\n<li>For each row in the distance matrix, the minimum value is extracted (resulting in N minimum values).</li>\n<li>If these minimum values are below the threshold of 1.0, they are counted as valid matches.<br>\nImage pairs with 15 or fewer valid matches are discarded.</li></ul></li>\n<li><p>The key point is that since LightGlue (GNN) is computationally expensive, a fast simulation of image matching is performed using the distance matrix as a lightweight approximation.</p></li>\n</ul>\n<h2>Image Matching</h2>\n<ul>\n<li>Local feature extractor: Aliked (n16)<ul>\n<li>resize_to: 1536 pixel</li>\n<li>max_num_keypoints: 8192 points</li>\n<li>detection_threshold: 0.01</li></ul></li>\n<li>Matcher: LighGlue</li>\n<li>post process<ul>\n<li>the number of matches &gt; 15 matches</li></ul></li>\n</ul>\n<h2>Filter image pairs</h2>\n<ul>\n<li><p>This step filters image pairs whose matches are not critical for SfM. The pruning accelerates the reconstruction and gives COLMAP’s incremental mapping more opportunities to reconstruct small clusters.</p></li>\n<li><p>Image-matching results can be expressed as a network in which images are nodes and match counts are edges; the number of edges connected to each node can then be counted.</p></li>\n<li><p>Non-essential pairs are removed by capping the number of edges per image node (threshold = 6). If a node has more than six edges, those with the fewest matches are discarded until only six remain.</p>\n<ul>\n<li>In COLMAP’s incremental mapping, an initial pair is chosen and additional images are registered one by one; the initial pair is selected statistically from the matches stored in the database.[1]</li>\n<li>Match results sometimes concentrate on a handful of images, causing the incremental mapping output to favor certain clusters.</li>\n<li>Pruning pairs connected to images with many matches increases the likelihood that initial pairs will be formed inside smaller clusters.</li></ul></li>\n</ul>\n<h2>Rough clustering</h2>\n<ul>\n<li>A lightweight clustering routine is implemented with NetworkX.</li>\n<li>For every image pair, the number of matches is read from <code>matches.h5</code>; edges with 150 or more matches are added to <code>G = nx.Graph()</code>.</li>\n<li>Clusters are then obtained via <code>nx.weakly_connected_components(G)</code>.</li>\n</ul>\n<h2>colmap (exhaustive matching)</h2>\n<ul>\n<li>Computation is performed with <strong>pycolmap</strong>’s <code>exhaustive_matching</code>.</li>\n<li>All settings remain at their defaults (pycolmap 0.6.1).</li>\n</ul>\n<h2>colmap (incremental mapping)</h2>\n<ul>\n<li>For each clustered image set, <code>incremental_mapping</code> in <strong>pycolmap</strong> is executed four times.</li>\n<li>Parameters:</li>\n</ul>\n<pre><code>mapper_options = pycolmap.IncrementalPipelineOptions()\nmapper_options.multiple_models   = \nmapper_options.min_model_size    = \nmapper_options.max_num_models = \nmapper_options.init_num_trials   = \n</code></pre>\n<h2>Merge maps</h2>\n<ul>\n<li>For every cluster, maps are merged.</li>\n<li>Maps are selected for merging when:<ul>\n<li>Six or more images are shared between the maps, and</li>\n<li>Each map’s <code>mean_reproj_error</code> is ≤ 2.0.</li></ul></li>\n<li>Merging is carried out with <strong>COLMAP</strong>’s native APIs (not pycolmap):<ul>\n<li><code>model_merger</code> combines the models.</li>\n<li><code>bundle_adjuster</code> optimizes camera parameters.<ul>\n<li>--BundleAdjustment.refine_focal_length = 1</li>\n<li>--BundleAdjustment.refine_principal_point = 1</li>\n<li>--BundleAdjustment.refine_extra_params = 1</li></ul></li></ul></li>\n</ul>\n<p>[1] pycolmap.incremental mapping()<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8163878%2F19d0e899b13ff451ce99c550dea94422%2FScreenshot%202025-06-05%2001.35.24.png?generation=1749055849424519&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "# Overview\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8163878%2Fa1a8238d8943872b2cd95d87778af334%2FScreenshot%202025-06-05%2001.46.42.png?generation=1749055688789666&alt=media)\n\n# Notebook\n- https://www.kaggle.com/code/motono0223/imc2025-11th-place-solution\n    - private lb = 43.34\n    - public lb = 43.53\n\n# Descriptions\n\n## Create image pairs\n- global feature extractor: DinoV2/Base\n- Topk = 150\n\n## Rerank image pairs\n* Due to the large number of image pairs, it is necessary to reduce the processing time for image matching.\n* To address this, image pairs are re-ranked and filtered using the ALIKED model (local features) in this process.\n* Specifically, the procedure is as follows:\n    * For each image, local descriptors are extracted using the ALIKED model.\n        * The image size is set to 2048 pixels, with 2048 keypoints and a detection threshold of 0.01.\n        * All descriptors are normalized with norm=1.\n    * A Euclidean distance matrix is computed between the local descriptors of each image pair.\n        * If image1 has N descriptors and image2 has M descriptors, an NxM distance matrix is calculated using the GPU.\n    * For each row in the distance matrix, the minimum value is extracted (resulting in N minimum values).\n    * If these minimum values are below the threshold of 1.0, they are counted as valid matches.\n       Image pairs with 15 or fewer valid matches are discarded.\n\n* The key point is that since LightGlue (GNN) is computationally expensive, a fast simulation of image matching is performed using the distance matrix as a lightweight approximation.\n\n## Image Matching\n* Local feature extractor: Aliked (n16)\n   * resize_to: 1536 pixel\n   * max_num_keypoints: 8192 points\n   * detection_threshold: 0.01\n* Matcher: LighGlue\n* post process\n   * the number of matches > 15 matches\n\n## Filter image pairs\n* This step filters image pairs whose matches are not critical for SfM. The pruning accelerates the reconstruction and gives COLMAP’s incremental mapping more opportunities to reconstruct small clusters.\n* Image-matching results can be expressed as a network in which images are nodes and match counts are edges; the number of edges connected to each node can then be counted.\n* Non-essential pairs are removed by capping the number of edges per image node (threshold = 6). If a node has more than six edges, those with the fewest matches are discarded until only six remain.\n\n  * In COLMAP’s incremental mapping, an initial pair is chosen and additional images are registered one by one; the initial pair is selected statistically from the matches stored in the database.[1]\n  * Match results sometimes concentrate on a handful of images, causing the incremental mapping output to favor certain clusters.\n  * Pruning pairs connected to images with many matches increases the likelihood that initial pairs will be formed inside smaller clusters.\n\n## Rough clustering\n* A lightweight clustering routine is implemented with NetworkX.\n* For every image pair, the number of matches is read from `matches.h5`; edges with 150 or more matches are added to `G = nx.Graph()`.\n* Clusters are then obtained via `nx.weakly_connected_components(G)`.\n\n## colmap (exhaustive matching)\n* Computation is performed with **pycolmap**’s `exhaustive_matching`.\n* All settings remain at their defaults (pycolmap 0.6.1).\n\n## colmap (incremental mapping)\n* For each clustered image set, `incremental_mapping` in **pycolmap** is executed four times.\n* Parameters:\n\n```python\nmapper_options = pycolmap.IncrementalPipelineOptions()\nmapper_options.multiple_models   = True\nmapper_options.min_model_size    = 2\nmapper_options.max_num_models = 15\nmapper_options.init_num_trials   = 1000\n```\n\n## Merge maps\n* For every cluster, maps are merged.\n* Maps are selected for merging when:\n  * Six or more images are shared between the maps, and\n  * Each map’s `mean_reproj_error` is ≤ 2.0.\n* Merging is carried out with **COLMAP**’s native APIs (not pycolmap):\n  * `model_merger` combines the models.\n  * `bundle_adjuster` optimizes camera parameters.\n       - --BundleAdjustment.refine_focal_length = 1\n       - --BundleAdjustment.refine_principal_point = 1\n       - --BundleAdjustment.refine_extra_params = 1\n\n\n[1] pycolmap.incremental mapping()\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8163878%2F19d0e899b13ff451ce99c550dea94422%2FScreenshot%202025-06-05%2001.35.24.png?generation=1749055849424519&alt=media)\n\n",
      "votes": 33
    },
    {
      "id": 3217219,
      "postDate": "2025-06-04T17:11:58.237Z",
      "content": "<p>Really appreciate you sharing this! To be honest, I’ve been curious about your solution ever since I saw you hit 100 on the ETs — that was incredible. I tried tuning for a long time but couldn’t crack it. Looking forward to learning from you!</p>",
      "rawMarkdown": "Really appreciate you sharing this! To be honest, I’ve been curious about your solution ever since I saw you hit 100 on the ETs — that was incredible. I tried tuning for a long time but couldn’t crack it. Looking forward to learning from you!"
    },
    {
      "id": 3217458,
      "postDate": "2025-06-05T03:17:23.273Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3217219,
      "author_name": "Helen",
      "author_url": "",
      "post_date": "2025-06-04T17:11:58.237000",
      "content": "<p>Really appreciate you sharing this! To be honest, I’ve been curious about your solution ever since I saw you hit 100 on the ETs — that was incredible. I tried tuning for a long time but couldn’t crack it. Looking forward to learning from you!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3217458,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-06-05T03:17:23.273000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3217205": "# Overview\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8163878%2Fa1a8238d8943872b2cd95d87778af334%2FScreenshot%202025-06-05%2001.46.42.png?generation=1749055688789666&alt=media)\n\n# Notebook\n- https://www.kaggle.com/code/motono0223/imc2025-11th-place-solution\n    - private lb = 43.34\n    - public lb = 43.53\n\n# Descriptions\n\n## Create image pairs\n- global feature extractor: DinoV2/Base\n- Topk = 150\n\n## Rerank image pairs\n* Due to the large number of image pairs, it is necessary to reduce the processing time for image matching.\n* To address this, image pairs are re-ranked and filtered using the ALIKED model (local features) in this process.\n* Specifically, the procedure is as follows:\n    * For each image, local descriptors are extracted using the ALIKED model.\n        * The image size is set to 2048 pixels, with 2048 keypoints and a detection threshold of 0.01.\n        * All descriptors are normalized with norm=1.\n    * A Euclidean distance matrix is computed between the local descriptors of each image pair.\n        * If image1 has N descriptors and image2 has M descriptors, an NxM distance matrix is calculated using the GPU.\n    * For each row in the distance matrix, the minimum value is extracted (resulting in N minimum values).\n    * If these minimum values are below the threshold of 1.0, they are counted as valid matches.\n       Image pairs with 15 or fewer valid matches are discarded.\n\n* The key point is that since LightGlue (GNN) is computationally expensive, a fast simulation of image matching is performed using the distance matrix as a lightweight approximation.\n\n## Image Matching\n* Local feature extractor: Aliked (n16)\n   * resize_to: 1536 pixel\n   * max_num_keypoints: 8192 points\n   * detection_threshold: 0.01\n* Matcher: LighGlue\n* post process\n   * the number of matches > 15 matches\n\n## Filter image pairs\n* This step filters image pairs whose matches are not critical for SfM. The pruning accelerates the reconstruction and gives COLMAP’s incremental mapping more opportunities to reconstruct small clusters.\n* Image-matching results can be expressed as a network in which images are nodes and match counts are edges; the number of edges connected to each node can then be counted.\n* Non-essential pairs are removed by capping the number of edges per image node (threshold = 6). If a node has more than six edges, those with the fewest matches are discarded until only six remain.\n\n  * In COLMAP’s incremental mapping, an initial pair is chosen and additional images are registered one by one; the initial pair is selected statistically from the matches stored in the database.[1]\n  * Match results sometimes concentrate on a handful of images, causing the incremental mapping output to favor certain clusters.\n  * Pruning pairs connected to images with many matches increases the likelihood that initial pairs will be formed inside smaller clusters.\n\n## Rough clustering\n* A lightweight clustering routine is implemented with NetworkX.\n* For every image pair, the number of matches is read from `matches.h5`; edges with 150 or more matches are added to `G = nx.Graph()`.\n* Clusters are then obtained via `nx.weakly_connected_components(G)`.\n\n## colmap (exhaustive matching)\n* Computation is performed with **pycolmap**’s `exhaustive_matching`.\n* All settings remain at their defaults (pycolmap 0.6.1).\n\n## colmap (incremental mapping)\n* For each clustered image set, `incremental_mapping` in **pycolmap** is executed four times.\n* Parameters:\n\n```python\nmapper_options = pycolmap.IncrementalPipelineOptions()\nmapper_options.multiple_models   = True\nmapper_options.min_model_size    = 2\nmapper_options.max_num_models = 15\nmapper_options.init_num_trials   = 1000\n```\n\n## Merge maps\n* For every cluster, maps are merged.\n* Maps are selected for merging when:\n  * Six or more images are shared between the maps, and\n  * Each map’s `mean_reproj_error` is ≤ 2.0.\n* Merging is carried out with **COLMAP**’s native APIs (not pycolmap):\n  * `model_merger` combines the models.\n  * `bundle_adjuster` optimizes camera parameters.\n       - --BundleAdjustment.refine_focal_length = 1\n       - --BundleAdjustment.refine_principal_point = 1\n       - --BundleAdjustment.refine_extra_params = 1\n\n\n[1] pycolmap.incremental mapping()\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8163878%2F19d0e899b13ff451ce99c550dea94422%2FScreenshot%202025-06-05%2001.35.24.png?generation=1749055849424519&alt=media)\n\n",
    "3217219": "Really appreciate you sharing this! To be honest, I’ve been curious about your solution ever since I saw you hit 100 on the ETs — that was incredible. I tried tuning for a long time but couldn’t crack it. Looking forward to learning from you!",
    "3217458": ""
  }
}