{
  "id": 583401,
  "title": "3rd Place Solution",
  "url": "/competitions/image-matching-challenge-2025/writeups/roni-heka-3rd-place-solution",
  "author_name": "",
  "post_date": "2025-06-07T00:30:52.347Z",
  "votes": 25,
  "comment_count": 9,
  "views": 0,
  "content": "<p>First and foremost, I would like to express my sincere gratitude to the organizers for hosting the Image Matching Challenge again this year. I am always inspired by the new themes you introduce each year, and I truly enjoy the opportunity to learn and experiment with new techniques.</p>\n<h2>Overview</h2>\n<p>My solution consists of a pre-matching stage with rotation augmentation, followed by multiple matching rounds using tiled images. The core ideas of rotation augmentation and image tiling were also used in <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/416918\" target=\"_blank\">my approach at IMC2023</a>.</p>\n<p>The overall pipeline is summarized in this diagram:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9249230%2Ffbfdc7bd888157ce842fdacc4567cbf6%2F20250606_solution.pptx%20(1).png?generation=1749222583567967&amp;alt=media\" alt=\"\"></p>\n<p>Like many other teams, I utilized a T4x2 GPU setup for parallel processing.</p>\n<h2>Solution Details</h2>\n<h3>2.1. Pre-matching</h3>\n<p>I leveraged the baseline pipeline, using DINOv2 to generate a list of image pairs with high similarity. I found the optimal parameters to be a minimum of 50-60 pairs per image (min_pair) and a similarity threshold (sim_th) of 0.2.</p>\n<p>For each pair, I performed matching using ALIKED + LightGlue (with longside set to 960px and max_keypoints to 4096). Pairs with more than 40 inlier matches were considered valid. Since some datasets contained rotated images, I implemented a retry mechanism with rotation augmentation for pairs that initially failed to match.</p>\n<h3>2.2. Matching with Tiled Images</h3>\n<p>For the valid pairs from the pre-matching stage, I uniformly tiled each image into four quadrants. I then performed matching on these tiles, again using ALIKED + LightGlue (max_keypoints=4096). The optimal longside setting for the tiles was either 1024 or 1216.</p>\n<p>A key part of my strategy was to create a set of five images for matching: the four tiles plus the original image downscaled by half. This resulted in 25 matching combinations per original image pair (5x5), which allows for more robust matching between images with significant differences in scale and perspective. While this may seem computationally expensive, the low computational cost of LightGlue made it feasible within the given time constraints. <br>\n<em>UPDATE</em>: I've posted the implemention of the pipeline in the <em>comments</em> section below.</p>\n<p>A reperesentative scores by the training data is summarized here. I couldn't find an effective method for the stairs dataset in the end…</p>\n<table>\n<thead>\n<tr>\n<th>heritage</th>\n<th>theather_church</th>\n<th>dioscuri_baalshamin</th>\n<th>lizard_pond</th>\n<th>piazzasanmarco</th>\n<th>amy_gardens</th>\n<th>fbk_vineyard</th>\n<th>ETs</th>\n<th>stairs</th>\n<th>Average</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>88.13</td>\n<td>61.39</td>\n<td>91.73</td>\n<td>76.16</td>\n<td>60.25</td>\n<td>23.91</td>\n<td>37.88</td>\n<td>66.67</td>\n<td>4.3</td>\n<td>56.71</td>\n<td>48.97</td>\n<td>49.05</td>\n</tr>\n</tbody>\n</table>\n<p>Actually, I have not fully verified why this specific tiling strategy was so effective. However, I believe one reason is that it helps mitigate the issue of keypoints being concentrated in only one part of the image.</p>\n<p>To improve time efficiency, I cached the results of the ALIKED feature extraction for reuse. After matching, I restored the keypoint coordinates to their original image space, used RANSAC to filter outliers, and then combined these keypoints with those from the pre-matching stage. The RANSAC parameters were taken from <a href=\"https://www.kaggle.com/code/vostankovich/imc2024-1st-place-solution\" target=\"_blank\">the 2024 1st place team's solution,</a> and I would like to extend my thanks to them.</p>\n<h3>2.3. Other Points</h3>\n<p>Registering pycolmap results: My pycolmap processing was based on the official baseline. However, I paid an attention to the order in which reconstruction results from incremental_mapping are registered. My understanding for this competition is that a larger number of registered images in the largest cluster of a scene generally leads to a better score (assuming the estimated camera poses are correct). Since images can be part of multiple clusters, it's crucial to prevent the results from larger, more robust clusters from being overwritten by those from smaller clusters.<br>\nTo achieve this, I sorted the reconstructed maps by the number of registered images in ascending order before registering the final poses. This modification seemed to yield a small boost on the leaderboard.</p>\n<p>Like this:</p>\n<pre><code>map_info_list = []\nfor map_idx, cur_map in maps.items():\n    num_registered_images = cur_map.num_registered_images()\n    map_info_list.append({\n        : map_idx,\n        : cur_map,\n        : num_registered_images\n    })\n\nsorted_map_info_list = sorted(map_info_list, key=lambda x: x[])\nfor map_info in sorted_map_info_list:\n    \n</code></pre>\n<p>Parameter Tuning: The parameters for ALIKED and LightGlue, as well as the input image sizes, had a significant impact on accuracy. Careful tuning was essential.</p>\n<h2>What Didn't Work in My Case</h2>\n<ul>\n<li>Using SIFT instead of ALIKED: While SIFT was faster, it resulted in a lower score on most datasets compared to ALIKED. I didn't have a chance to try an ensemble of ALIKED and SIFT.</li>\n<li>Scene Pre-Clustering: I attempted to pre-cluster images within datasets that contained multiple scenes.<br>\nI tried clustering based on a distance matrix derived from the number of keypoint matches from the pre-matching stage, but I couldn't find a reliable method to accurately segment all scenes.<br>\nEven when scenes were correctly segmented, it did not lead to a significant score improvement (for this competition's metric) unless the subsequent matching step could find enough keypoints for a robust reconstruction.</li>\n<li>While pairwise matching-based clustering had low accuracy, I found VGGT to be powerful from a clustering perspective, as it seems to find context among a group of dozens of images. I experimented with using SIFT feature points as query_points and identifying pairs based on the number of points exceeding vis_score and conf_score thresholds for a target image. This method achieved a higher true positive rate for pair identification than the pairwise approach in section 2.1 and successfully clustered many datasets.<br>\nHowever, in terms of the resulting pose accuracy and final score, it did not show a significant advantage in my setup, possibly due to the resizing.<br>\nThe use of VGGT has been deeply explored and discussed by <a href=\"https://www.kaggle.com/tmyok1984\" target=\"_blank\">@tmyok1984</a> san in <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2025/discussion/582968\" target=\"_blank\">this discussion post.</a> I'm learning a great deal from reading the post.</li>\n</ul>",
  "messages": [
    {
      "id": "3218701",
      "postDate": "06/06/2025 15:36:04",
      "content": "<p>First and foremost, I would like to express my sincere gratitude to the organizers for hosting the Image Matching Challenge again this year. I am always inspired by the new themes you introduce each year, and I truly enjoy the opportunity to learn and experiment with new techniques.</p>\n<h2>Overview</h2>\n<p>My solution consists of a pre-matching stage with rotation augmentation, followed by multiple matching rounds using tiled images. The core ideas of rotation augmentation and image tiling were also used in <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/416918\" target=\"_blank\">my approach at IMC2023</a>.</p>\n<p>The overall pipeline is summarized in this diagram:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9249230%2Ffbfdc7bd888157ce842fdacc4567cbf6%2F20250606_solution.pptx%20(1).png?generation=1749222583567967&amp;alt=media\" alt=\"\"></p>\n<p>Like many other teams, I utilized a T4x2 GPU setup for parallel processing.</p>\n<h2>Solution Details</h2>\n<h3>2.1. Pre-matching</h3>\n<p>I leveraged the baseline pipeline, using DINOv2 to generate a list of image pairs with high similarity. I found the optimal parameters to be a minimum of 50-60 pairs per image (min_pair) and a similarity threshold (sim_th) of 0.2.</p>\n<p>For each pair, I performed matching using ALIKED + LightGlue (with longside set to 960px and max_keypoints to 4096). Pairs with more than 40 inlier matches were considered valid. Since some datasets contained rotated images, I implemented a retry mechanism with rotation augmentation for pairs that initially failed to match.</p>\n<h3>2.2. Matching with Tiled Images</h3>\n<p>For the valid pairs from the pre-matching stage, I uniformly tiled each image into four quadrants. I then performed matching on these tiles, again using ALIKED + LightGlue (max_keypoints=4096). The optimal longside setting for the tiles was either 1024 or 1216.</p>\n<p>A key part of my strategy was to create a set of five images for matching: the four tiles plus the original image downscaled by half. This resulted in 25 matching combinations per original image pair (5x5), which allows for more robust matching between images with significant differences in scale and perspective. While this may seem computationally expensive, the low computational cost of LightGlue made it feasible within the given time constraints. <br>\n<em>UPDATE</em>: I've posted the implemention of the pipeline in the <em>comments</em> section below.</p>\n<p>A reperesentative scores by the training data is summarized here. I couldn't find an effective method for the stairs dataset in the end…</p>\n<table>\n<thead>\n<tr>\n<th>heritage</th>\n<th>theather_church</th>\n<th>dioscuri_baalshamin</th>\n<th>lizard_pond</th>\n<th>piazzasanmarco</th>\n<th>amy_gardens</th>\n<th>fbk_vineyard</th>\n<th>ETs</th>\n<th>stairs</th>\n<th>Average</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>88.13</td>\n<td>61.39</td>\n<td>91.73</td>\n<td>76.16</td>\n<td>60.25</td>\n<td>23.91</td>\n<td>37.88</td>\n<td>66.67</td>\n<td>4.3</td>\n<td>56.71</td>\n<td>48.97</td>\n<td>49.05</td>\n</tr>\n</tbody>\n</table>\n<p>Actually, I have not fully verified why this specific tiling strategy was so effective. However, I believe one reason is that it helps mitigate the issue of keypoints being concentrated in only one part of the image.</p>\n<p>To improve time efficiency, I cached the results of the ALIKED feature extraction for reuse. After matching, I restored the keypoint coordinates to their original image space, used RANSAC to filter outliers, and then combined these keypoints with those from the pre-matching stage. The RANSAC parameters were taken from <a href=\"https://www.kaggle.com/code/vostankovich/imc2024-1st-place-solution\" target=\"_blank\">the 2024 1st place team's solution,</a> and I would like to extend my thanks to them.</p>\n<h3>2.3. Other Points</h3>\n<p>Registering pycolmap results: My pycolmap processing was based on the official baseline. However, I paid an attention to the order in which reconstruction results from incremental_mapping are registered. My understanding for this competition is that a larger number of registered images in the largest cluster of a scene generally leads to a better score (assuming the estimated camera poses are correct). Since images can be part of multiple clusters, it's crucial to prevent the results from larger, more robust clusters from being overwritten by those from smaller clusters.<br>\nTo achieve this, I sorted the reconstructed maps by the number of registered images in ascending order before registering the final poses. This modification seemed to yield a small boost on the leaderboard.</p>\n<p>Like this:</p>\n<pre><code>map_info_list = []\nfor map_idx, cur_map in maps.items():\n    num_registered_images = cur_map.num_registered_images()\n    map_info_list.append({\n        : map_idx,\n        : cur_map,\n        : num_registered_images\n    })\n\nsorted_map_info_list = sorted(map_info_list, key=lambda x: x[])\nfor map_info in sorted_map_info_list:\n    \n</code></pre>\n<p>Parameter Tuning: The parameters for ALIKED and LightGlue, as well as the input image sizes, had a significant impact on accuracy. Careful tuning was essential.</p>\n<h2>What Didn't Work in My Case</h2>\n<ul>\n<li>Using SIFT instead of ALIKED: While SIFT was faster, it resulted in a lower score on most datasets compared to ALIKED. I didn't have a chance to try an ensemble of ALIKED and SIFT.</li>\n<li>Scene Pre-Clustering: I attempted to pre-cluster images within datasets that contained multiple scenes.<br>\nI tried clustering based on a distance matrix derived from the number of keypoint matches from the pre-matching stage, but I couldn't find a reliable method to accurately segment all scenes.<br>\nEven when scenes were correctly segmented, it did not lead to a significant score improvement (for this competition's metric) unless the subsequent matching step could find enough keypoints for a robust reconstruction.</li>\n<li>While pairwise matching-based clustering had low accuracy, I found VGGT to be powerful from a clustering perspective, as it seems to find context among a group of dozens of images. I experimented with using SIFT feature points as query_points and identifying pairs based on the number of points exceeding vis_score and conf_score thresholds for a target image. This method achieved a higher true positive rate for pair identification than the pairwise approach in section 2.1 and successfully clustered many datasets.<br>\nHowever, in terms of the resulting pose accuracy and final score, it did not show a significant advantage in my setup, possibly due to the resizing.<br>\nThe use of VGGT has been deeply explored and discussed by <a href=\"https://www.kaggle.com/tmyok1984\" target=\"_blank\">@tmyok1984</a> san in <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2025/discussion/582968\" target=\"_blank\">this discussion post.</a> I'm learning a great deal from reading the post.</li>\n</ul>",
      "rawMarkdown": "First and foremost, I would like to express my sincere gratitude to the organizers for hosting the Image Matching Challenge again this year. I am always inspired by the new themes you introduce each year, and I truly enjoy the opportunity to learn and experiment with new techniques.\n\n## Overview\nMy solution consists of a pre-matching stage with rotation augmentation, followed by multiple matching rounds using tiled images. The core ideas of rotation augmentation and image tiling were also used in [my approach at IMC2023](https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/416918).\n\nThe overall pipeline is summarized in this diagram:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9249230%2Ffbfdc7bd888157ce842fdacc4567cbf6%2F20250606_solution.pptx%20(1).png?generation=1749222583567967&alt=media)\n\nLike many other teams, I utilized a T4x2 GPU setup for parallel processing.\n\n## Solution Details\n### 2.1. Pre-matching\nI leveraged the baseline pipeline, using DINOv2 to generate a list of image pairs with high similarity. I found the optimal parameters to be a minimum of 50-60 pairs per image (min_pair) and a similarity threshold (sim_th) of 0.2.\n\nFor each pair, I performed matching using ALIKED + LightGlue (with longside set to 960px and max_keypoints to 4096). Pairs with more than 40 inlier matches were considered valid. Since some datasets contained rotated images, I implemented a retry mechanism with rotation augmentation for pairs that initially failed to match.\n\n### 2.2. Matching with Tiled Images\nFor the valid pairs from the pre-matching stage, I uniformly tiled each image into four quadrants. I then performed matching on these tiles, again using ALIKED + LightGlue (max_keypoints=4096). The optimal longside setting for the tiles was either 1024 or 1216.\n\nA key part of my strategy was to create a set of five images for matching: the four tiles plus the original image downscaled by half. This resulted in 25 matching combinations per original image pair (5x5), which allows for more robust matching between images with significant differences in scale and perspective. While this may seem computationally expensive, the low computational cost of LightGlue made it feasible within the given time constraints. \n*UPDATE*: I've posted the implemention of the pipeline in the *comments* section below.\n\nA reperesentative scores by the training data is summarized here. I couldn't find an effective method for the stairs dataset in the end...\n|heritage\t|theather_church\t|dioscuri_baalshamin\t|lizard_pond\t|piazzasanmarco\t|amy_gardens\t|fbk_vineyard\t|ETs\t|stairs\t|Average\t|Public\t|Private|\n|:--|:--|:--|:--|:--|:--|:--|:--|:--|:--|:--|:--|\n|88.13\t|61.39\t|91.73\t|76.16\t|60.25\t|23.91\t|37.88\t|66.67\t|4.3\t|56.71\t|48.97\t|49.05|\n\n\nActually, I have not fully verified why this specific tiling strategy was so effective. However, I believe one reason is that it helps mitigate the issue of keypoints being concentrated in only one part of the image.\n\nTo improve time efficiency, I cached the results of the ALIKED feature extraction for reuse. After matching, I restored the keypoint coordinates to their original image space, used RANSAC to filter outliers, and then combined these keypoints with those from the pre-matching stage. The RANSAC parameters were taken from [the 2024 1st place team's solution,](https://www.kaggle.com/code/vostankovich/imc2024-1st-place-solution) and I would like to extend my thanks to them.\n\n### 2.3. Other Points\nRegistering pycolmap results: My pycolmap processing was based on the official baseline. However, I paid an attention to the order in which reconstruction results from incremental_mapping are registered. My understanding for this competition is that a larger number of registered images in the largest cluster of a scene generally leads to a better score (assuming the estimated camera poses are correct). Since images can be part of multiple clusters, it's crucial to prevent the results from larger, more robust clusters from being overwritten by those from smaller clusters.\nTo achieve this, I sorted the reconstructed maps by the number of registered images in ascending order before registering the final poses. This modification seemed to yield a small boost on the leaderboard.\n\nLike this:\n``` \nmap_info_list = []\nfor map_idx, cur_map in maps.items():\n    num_registered_images = cur_map.num_registered_images()\n    map_info_list.append({\n        \"original_map_index\": map_idx,\n        \"map_object\": cur_map,\n        \"num_reg_images\": num_registered_images\n    })\n# Sort by the number of registered images in ascending order\nsorted_map_info_list = sorted(map_info_list, key=lambda x: x[\"num_reg_images\"])\nfor map_info in sorted_map_info_list:\n    # (Register poses...)\n``` \n\nParameter Tuning: The parameters for ALIKED and LightGlue, as well as the input image sizes, had a significant impact on accuracy. Careful tuning was essential.\n\n## What Didn't Work in My Case\n- Using SIFT instead of ALIKED: While SIFT was faster, it resulted in a lower score on most datasets compared to ALIKED. I didn't have a chance to try an ensemble of ALIKED and SIFT.\n- Scene Pre-Clustering: I attempted to pre-cluster images within datasets that contained multiple scenes.\nI tried clustering based on a distance matrix derived from the number of keypoint matches from the pre-matching stage, but I couldn't find a reliable method to accurately segment all scenes.\nEven when scenes were correctly segmented, it did not lead to a significant score improvement (for this competition's metric) unless the subsequent matching step could find enough keypoints for a robust reconstruction.\n- While pairwise matching-based clustering had low accuracy, I found VGGT to be powerful from a clustering perspective, as it seems to find context among a group of dozens of images. I experimented with using SIFT feature points as query_points and identifying pairs based on the number of points exceeding vis_score and conf_score thresholds for a target image. This method achieved a higher true positive rate for pair identification than the pairwise approach in section 2.1 and successfully clustered many datasets.\nHowever, in terms of the resulting pose accuracy and final score, it did not show a significant advantage in my setup, possibly due to the resizing.\nThe use of VGGT has been deeply explored and discussed by @tmyok1984 san in [this discussion post.](https://www.kaggle.com/competitions/image-matching-challenge-2025/discussion/582968) I'm learning a great deal from reading the post.",
      "votes": null
    },
    {
      "id": "3218778",
      "postDate": "06/06/2025 17:33:11",
      "content": "<p>Thanks for sharing your solution.  <br>\nI first learned about the tiled images approach from your solution two years ago,  <br>\nand I'm surprised to see it still worked effectively for this year's challenge.</p>\n<p>I'm curious why you decided to use this strategy again,  <br>\nand what kind of evaluation scores you got on the validation data.</p>\n<p>Do you have a table with numbers for <code>amy_gardens</code> or <code>fbk_vineyard</code> when using the tiling strategy?</p>",
      "rawMarkdown": "Thanks for sharing your solution.  \nI first learned about the tiled images approach from your solution two years ago,  \nand I'm surprised to see it still worked effectively for this year's challenge.\n\nI'm curious why you decided to use this strategy again,  \nand what kind of evaluation scores you got on the validation data.\n\nDo you have a table with numbers for `amy_gardens` or `fbk_vineyard` when using the tiling strategy?",
      "votes": null
    },
    {
      "id": "3218782",
      "postDate": "06/06/2025 17:41:39",
      "content": "<blockquote>\n  <p>While this may seem computationally expensive, the low computational cost of LightGlue made it feasible within the given time constraints.</p>\n</blockquote>\n<p>I understand that LightGlue has a relatively low computational cost, but performing 25 matching combinations per original image pair still seems computationally intensive.<br>\nCould you share any implementation strategies or optimizations you used to make this feasible?<br>\nIt would be great to learn from your approach — do you have any plans to share the implementation code?</p>",
      "rawMarkdown": "> While this may seem computationally expensive, the low computational cost of LightGlue made it feasible within the given time constraints.\n\nI understand that LightGlue has a relatively low computational cost, but performing 25 matching combinations per original image pair still seems computationally intensive.\nCould you share any implementation strategies or optimizations you used to make this feasible?\nIt would be great to learn from your approach — do you have any plans to share the implementation code?",
      "votes": null
    },
    {
      "id": "3218788",
      "postDate": "06/06/2025 17:46:39",
      "content": "<p>Thanks for mentioning my post about VGGT!<br>\nIt definitely had great potential, but I couldn’t quite make it work — ended up not boosting the LB at all 😅</p>",
      "rawMarkdown": "Thanks for mentioning my post about VGGT!\nIt definitely had great potential, but I couldn’t quite make it work — ended up not boosting the LB at all 😅",
      "votes": null
    },
    {
      "id": "3218927",
      "postDate": "06/06/2025 23:38:35",
      "content": "<p>Thank you for your comments!<br>\nTo keep the execution time under the 9-hour limit, I had to constrain min_pairs and sim_th as described above. Another necessary optimization was setting the LightGlue confidence threshold for pre-matching to 0.2, which helped reduce false positives among the selected pairs.</p>\n<p>I'll post the code here that takes a pair of image tensors performs the tiling and matching, returns the keypoint coordinates and ALIKED caches.<br>\nThe pipeline is inspired by <a href=\"https://www.kaggle.com/code/chankhavu/loftr-superglue-dkm-with-inspiration\" target=\"_blank\">this work</a>.</p>\n<pre><code> (torch.nn.Module):\n     ():\n        ().__init__()\n        .device=device\n        .extractor = ALIKED(weights=, **extractor_cfg).to(.device).()\n        .lightglue = LightGlue(**lg_cfg).to(.device)\n        .ttas = (())\n        .tta2id = {k: i  i, k  (.ttas)}\n        .tta_combination = [[i, j]  i  ()  j  ()]\n\n     ():\n        pred = {}        \n        cache, key1, key2, quad = cache_args \n        img_key_list = [key1,key2]\n        data[] = {: data[]}\n        data[] = {: data[]}\n        keypoints_dict, descriptors_dict = {}, {}\n         i, img_key  ([, ]):\n            keypoints_list, descriptors_list = [], []\n             i == :\n                _quad = \n            :\n                _quad = quad\n                cache[img_key_list[i]][_quad][]:\n                no_cache = \n            :\n                no_cache = \n            \n             j, img  (data[img_key][]):\n                 no_cache:\n                    img = img.unsqueeze()\n                    cache[img_key_list[i]][_quad][j][] = .extractor.extract(img, resize=)\n                pred = cache[img_key_list[i]][_quad][j][]\n                keypoints_list.append(pred[])\n                descriptors_list.append(pred[])\n            keypoints_dict[img_key] = keypoints_list\n            descriptors_dict[img_key] = descriptors_list\n        \n        group_pred_list = []\n         tta_group  .tta_combination:\n            group_idx = .tta2id[tta_group[]], .tta2id[tta_group[]]\n            i0, i1 = group_idx[], group_idx[]\n            data[][], data[][]= keypoints_dict[][i0], descriptors_dict[][i0]\n            data[][], data[][]= keypoints_dict[][i1], descriptors_dict[][i1]\n            group_pred = .lightglue(data)\n            group_pred.update({:data[][],\n                                :data[][]})\n            group_pred_list.append(group_pred)\n         group_pred_list, cache\n\n :\n     ():\n        .device = device\n        .extractor_cfg = extractor_cfg\n        .lg_cfg = lg_cfg\n        ._lightglue_matcher = LightGlueCustomMatching_sep(\n            device=.device, extractor_cfg=.extractor_cfg, lg_cfg=.lg_cfg\n            ).().to(device,dtype)\n        .conf_thresh = conf_th\n\n     ():\n        \n        img = img.clone()\n         long_side   :\n            scale = long_side / (img.shape[], img.shape[])\n            w = (img.shape[] * scale)\n            h = (img.shape[] * scale)\n            img = torch.nn.functional.interpolate(img, size=(h, w), mode=, align_corners=)\n        :\n            scale = \n         img, scale\n\n     ():\n        \n        h, w = image.shape[], image.shape[]\n         h %  != :\n            h = h - \n         w %  != :\n            w = w - \n        image = image[:, :, :h, :w]\n         [image[:, :, :h//, :w//], \n            image[:, :, :h//, w//:], \n            image[:, :, h//:, :w//], \n            image[:, :, h//:, w//:],\n            transforms.functional.resize(image, size=(h//,w//))]\n\n     ():\n        \n         quadrant == :\n            coords[:, ] += w//\n         quadrant == :\n            coords[:, ] += h//\n         quadrant == :\n            coords[:, ] += w//\n            coords[:, ] += h//\n         quadrant == :\n            coords = [[y*, x*]  y, x  coords]            \n         coords\n\n     ():\n         torch.no_grad():\n            img_ts0, scale0 = .prep_img(img_ts0, input_longside)\n            img_ts1, scale1 = .prep_img(img_ts1, input_longside)\n            img_parts0 = .split_image(img_ts0) \n            img_parts1 = .split_image(img_ts1)\n            cat_mkpts0, cat_mkpts1 = [], []\n            pred, cache = ._lightglue_matcher.forward_flat(\n                data={\n                    : torch.cat(img_parts0),\n                    : torch.cat(img_parts1),\n                },\n            cache_args=cache_args)\n\n            mkpts0, mkpts1= [], []\n             idx, [i0,i1]  (.tta_combination):\n                group_pred = pred[idx]\n                pred_aug={}\n                use_keys = [, , , ]\n                 k  use_keys:\n                    v = group_pred[k]\n                     (v, torch.Tensor):\n                        pred_aug[k] = v[].detach().cpu().numpy().squeeze()\n                    :\n                        pred_aug[k] = v               \n                kpts0, kpts1 = pred_aug[], pred_aug[]\n                matches = pred_aug[]\n                valid = matches &gt; -\n                \n                 (kpts0[valid]) &gt; :\n                    mkpts0 = .reconstruct_coords(kpts0[valid], i0, img_ts0.shape[], img_ts0.shape[])\n                    mkpts1 = .reconstruct_coords(kpts1[matches[valid]], i1, img_ts1.shape[], img_ts1.shape[])\n                    cat_mkpts0.append(mkpts0)\n                    cat_mkpts1.append(mkpts1)\n\n             (cat_mkpts0) &gt; :\n                cat_mkpts0 = np.concatenate(cat_mkpts0)\n                cat_mkpts1 = np.concatenate(cat_mkpts1)\n            :\n                _, key1, key2, _ = cache_args\n                ()\n                 np.empty((, )), np.empty((, )), cache\n            \n            :\n                _, inliers = cv2.findFundamentalMat(cat_mkpts0, cat_mkpts1, cv2.USAC_MAGSAC, ransacReprojThreshold=, confidence=, maxIters=)\n                inliers = inliers.ravel() &gt; \n                cat_mkpts0 = cat_mkpts0[inliers]\n                cat_mkpts1 = cat_mkpts1[inliers]\n             Exception:\n                _, key1, key2, _ = cache_args\n                ()\n                 np.empty((, )), np.empty((, )), cache\n\n            mask0 = (cat_mkpts0[:, ] &gt;= ) &amp; (cat_mkpts0[:, ] &lt; img_ts0.shape[]) &amp; (cat_mkpts0[:, ] &gt;= ) &amp; (cat_mkpts0[:, ] &lt; img_ts0.shape[])\n            mask1 = (cat_mkpts1[:, ] &gt;= ) &amp; (cat_mkpts1[:, ] &lt; img_ts1.shape[]) &amp; (cat_mkpts1[:, ] &gt;= ) &amp; (cat_mkpts1[:, ] &lt; img_ts1.shape[])\n             cat_mkpts0[mask0 &amp; mask1] / scale0, cat_mkpts1[mask0 &amp; mask1] / scale1, cache\n</code></pre>",
      "rawMarkdown": "Thank you for your comments!\nTo keep the execution time under the 9-hour limit, I had to constrain min_pairs and sim_th as described above. Another necessary optimization was setting the LightGlue confidence threshold for pre-matching to 0.2, which helped reduce false positives among the selected pairs.\n\nI'll post the code here that takes a pair of image tensors performs the tiling and matching, returns the keypoint coordinates and ALIKED caches.\nThe pipeline is inspired by [this work](https://www.kaggle.com/code/chankhavu/loftr-superglue-dkm-with-inspiration).\n\n```\nclass LightGlueCustomMatching_sep(torch.nn.Module):\n    def __init__(self, device=None, extractor_cfg=None, lg_cfg=None):\n        super().__init__()\n        self.device=device\n        self.extractor = ALIKED(weights=f\"/kaggle/input/imc24lightglue/weights/aliked-n16.pth\", **extractor_cfg).to(self.device).eval()\n        self.lightglue = LightGlue(**lg_cfg).to(self.device)\n        self.ttas = list(range(5))\n        self.tta2id = {k: i for i, k in enumerate(self.ttas)}\n        self.tta_combination = [[i, j] for i in range(5) for j in range(5)]\n\n    def forward_flat(self, data, cache_args):\n        pred = {}        \n        cache, key1, key2, quad = cache_args # quad: Rotation times of image1. 0:No rotation, 1:90deg, 2:180deg, 3:270deg\n        img_key_list = [key1,key2]\n        data[\"image0\"] = {\"image\": data[\"image0\"]}\n        data[\"image1\"] = {\"image\": data[\"image1\"]}\n        keypoints_dict, descriptors_dict = {}, {}\n        for i, img_key in enumerate(['image0', 'image1']):\n            keypoints_list, descriptors_list = [], []\n            if i == 0:\n                _quad = 0\n            else:\n                _quad = quad\n            if \"pred\" not in cache[img_key_list[i]][_quad][0]:\n                no_cache = True\n            else:\n                no_cache = False\n            # Get ALIKED descriptors\n            for j, img in enumerate(data[img_key][\"image\"]):\n                if no_cache:\n                    img = img.unsqueeze(0)\n                    cache[img_key_list[i]][_quad][j][\"pred\"] = self.extractor.extract(img, resize=None)\n                pred = cache[img_key_list[i]][_quad][j][\"pred\"]\n                keypoints_list.append(pred['keypoints'])\n                descriptors_list.append(pred['descriptors'])\n            keypoints_dict[img_key] = keypoints_list\n            descriptors_dict[img_key] = descriptors_list\n        # Prepare data for LightGlue and run matching one by one\n        group_pred_list = []\n        for tta_group in self.tta_combination:\n            group_idx = self.tta2id[tta_group[0]], self.tta2id[tta_group[1]]\n            i0, i1 = group_idx[0], group_idx[1]\n            data[\"image0\"][\"keypoints\"], data[\"image0\"][\"descriptors\"]= keypoints_dict['image0'][i0], descriptors_dict['image0'][i0]\n            data[\"image1\"][\"keypoints\"], data[\"image1\"][\"descriptors\"]= keypoints_dict['image1'][i1], descriptors_dict['image1'][i1]\n            group_pred = self.lightglue(data)\n            group_pred.update({\"keypoints0\":data['image0'][\"keypoints\"],\n                                \"keypoints1\":data['image1'][\"keypoints\"]})\n            group_pred_list.append(group_pred)\n        return group_pred_list, cache\n    \nclass LightGlueMatcherPipeline_sep:\n    def __init__(self, device=None, conf_th=None, extractor_cfg=None, lg_cfg=None):\n        self.device = device\n        self.extractor_cfg = extractor_cfg\n        self.lg_cfg = lg_cfg\n        self._lightglue_matcher = LightGlueCustomMatching_sep(\n            device=self.device, extractor_cfg=self.extractor_cfg, lg_cfg=self.lg_cfg\n            ).eval().to(device,dtype)\n        self.conf_thresh = conf_th\n    \n    def prep_img(self, img, long_side=None):\n        \"\"\"Resize the tensor image to a specified long side.\"\"\"\n        img = img.clone()\n        if long_side is not None:\n            scale = long_side / max(img.shape[2], img.shape[3])\n            w = int(img.shape[3] * scale)\n            h = int(img.shape[2] * scale)\n            img = torch.nn.functional.interpolate(img, size=(h, w), mode='bilinear', align_corners=False)\n        else:\n            scale = 1.0\n        return img, scale\n    \n    def split_image(self, image):\n        \"\"\"Split the image into 4 quadrants and return them along with a resized version.\"\"\"\n        h, w = image.shape[2], image.shape[3]\n        if h % 2 != 0:\n            h = h - 1\n        if w % 2 != 0:\n            w = w - 1\n        image = image[:, :, :h, :w]\n        return [image[:, :, :h//2, :w//2], \n            image[:, :, :h//2, w//2:], \n            image[:, :, h//2:, :w//2], \n            image[:, :, h//2:, w//2:],\n            transforms.functional.resize(image, size=(h//2,w//2))]\n\n    def reconstruct_coords(self, coords, quadrant, w, h):\n        \"\"\"Reconstruct coordinates based on the separation quadrant.\"\"\"\n        if quadrant == 1:\n            coords[:, 0] += w//2\n        elif quadrant == 2:\n            coords[:, 1] += h//2\n        elif quadrant == 3:\n            coords[:, 0] += w//2\n            coords[:, 1] += h//2\n        elif quadrant == 4:\n            coords = [[y*2, x*2] for y, x in coords]            \n        return coords\n    \n    def __call__(self, img_ts0, img_ts1, cache_args, input_longside=None):\n        with torch.no_grad():\n            img_ts0, scale0 = self.prep_img(img_ts0, input_longside)\n            img_ts1, scale1 = self.prep_img(img_ts1, input_longside)\n            img_parts0 = self.split_image(img_ts0) \n            img_parts1 = self.split_image(img_ts1)\n            cat_mkpts0, cat_mkpts1 = [], []\n            pred, cache = self._lightglue_matcher.forward_flat(\n                data={\n                    \"image0\": torch.cat(img_parts0),\n                    \"image1\": torch.cat(img_parts1),\n                },\n            cache_args=cache_args)\n\n            mkpts0, mkpts1= [], []\n            for idx, [i0,i1] in enumerate(self.tta_combination):\n                group_pred = pred[idx]\n                pred_aug={}\n                use_keys = [\"keypoints0\", \"keypoints1\", \"matches0\", \"matching_scores0\"]\n                for k in use_keys:\n                    v = group_pred[k]\n                    if isinstance(v, torch.Tensor):\n                        pred_aug[k] = v[0].detach().cpu().numpy().squeeze()\n                    else:\n                        pred_aug[k] = v               \n                kpts0, kpts1 = pred_aug[\"keypoints0\"], pred_aug[\"keypoints1\"]\n                matches = pred_aug[\"matches0\"]\n                valid = matches > -1\n                # Recover keypoint coordinates based on the quadrant\n                if len(kpts0[valid]) > 0:\n                    mkpts0 = self.reconstruct_coords(kpts0[valid], i0, img_ts0.shape[3], img_ts0.shape[2])\n                    mkpts1 = self.reconstruct_coords(kpts1[matches[valid]], i1, img_ts1.shape[3], img_ts1.shape[2])\n                    cat_mkpts0.append(mkpts0)\n                    cat_mkpts1.append(mkpts1)\n\n            if len(cat_mkpts0) > 0:\n                cat_mkpts0 = np.concatenate(cat_mkpts0)\n                cat_mkpts1 = np.concatenate(cat_mkpts1)\n            else:\n                _, key1, key2, _ = cache_args\n                print(f\"No matches at {key1} vs. {key2}\")\n                return np.empty((0, 2)), np.empty((0, 2)), cache\n            # RANSAC\n            try:\n                _, inliers = cv2.findFundamentalMat(cat_mkpts0, cat_mkpts1, cv2.USAC_MAGSAC, ransacReprojThreshold=5, confidence=0.9999, maxIters=50000)\n                inliers = inliers.ravel() > 0\n                cat_mkpts0 = cat_mkpts0[inliers]\n                cat_mkpts1 = cat_mkpts1[inliers]\n            except Exception:\n                _, key1, key2, _ = cache_args\n                print(f\"Error in findFundamentalMat: {key1}-{key2}\")\n                return np.empty((0, 2)), np.empty((0, 2)), cache\n\n            mask0 = (cat_mkpts0[:, 0] >= 0) & (cat_mkpts0[:, 0] < img_ts0.shape[3]) & (cat_mkpts0[:, 1] >= 0) & (cat_mkpts0[:, 1] < img_ts0.shape[2])\n            mask1 = (cat_mkpts1[:, 0] >= 0) & (cat_mkpts1[:, 0] < img_ts1.shape[3]) & (cat_mkpts1[:, 1] >= 0) & (cat_mkpts1[:, 1] < img_ts1.shape[2])\n            return cat_mkpts0[mask0 & mask1] / scale0, cat_mkpts1[mask0 & mask1] / scale1, cache\n```",
      "votes": null
    },
    {
      "id": "3218933",
      "postDate": "06/06/2025 23:47:44",
      "content": "<p>Congratulations on 3rd place! Your approach with rotation augmentation and the 5-image tiling strategy is truly innovative for robust matching. <br>\nCould you elaborate on why you chose a min_pair of 50-60 and a sim_th of 0.2 for your pre-matching stage?</p>",
      "rawMarkdown": "Congratulations on 3rd place! Your approach with rotation augmentation and the 5-image tiling strategy is truly innovative for robust matching. \nCould you elaborate on why you chose a min_pair of 50-60 and a sim_th of 0.2 for your pre-matching stage?",
      "votes": null
    },
    {
      "id": "3218936",
      "postDate": "06/07/2025 00:06:04",
      "content": "<p>Thanks for the congratulations and for the great question!<br>\nThe main reasons were:</p>\n<ul>\n<li>The distribution of similarity scores obtained by DINOv2 changes significantly from one dataset to another. On a dataset where the average similarity is high, a fixed threshold could generate a massive number of candidate pairs. This would spend too much time on pre-matching at a specific dataset. This is why the sim_th is set to low.</li>\n<li>The subsequent matching process on the tiled images is time-consuming. Through experimentation, we found that increasing min_pair above the 50-60 range pushed our total runtime over the limit, which would cause a timeout.</li>\n</ul>",
      "rawMarkdown": "Thanks for the congratulations and for the great question!\nThe main reasons were:\n- The distribution of similarity scores obtained by DINOv2 changes significantly from one dataset to another. On a dataset where the average similarity is high, a fixed threshold could generate a massive number of candidate pairs. This would spend too much time on pre-matching at a specific dataset. This is why the sim_th is set to low.\n- The subsequent matching process on the tiled images is time-consuming. Through experimentation, we found that increasing min_pair above the 50-60 range pushed our total runtime over the limit, which would cause a timeout.",
      "votes": null
    },
    {
      "id": "3218943",
      "postDate": "06/07/2025 00:40:45",
      "content": "<p>I did experiment with a variety of other strategies for this year's challenge. I tried several combinations of multi-scale image sets and also tested cropping. However, for my particular pipeline, the tiling strategy still was found to be the most effective. </p>\n<p>I have updated the main text to include the evaluation results on the training data.<br>\nLater, I will provide an update with comparative data against other methods :)</p>",
      "rawMarkdown": "I did experiment with a variety of other strategies for this year's challenge. I tried several combinations of multi-scale image sets and also tested cropping. However, for my particular pipeline, the tiling strategy still was found to be the most effective. \n\nI have updated the main text to include the evaluation results on the training data.\nLater, I will provide an update with comparative data against other methods :)",
      "votes": null
    },
    {
      "id": "3220952",
      "postDate": "06/10/2025 07:33:41",
      "content": "<p>Hi! Congrats again — super impressive work! Would you mind sharing the ALIKED and LightGlue parameters you used?<br>\nI followed the setup you posted and built a similar version. After several runs and tuning attempts, I feel like you’ve really nailed down some excellent parameter choices. Would love to learn from what worked best for you!</p>",
      "rawMarkdown": "Hi! Congrats again — super impressive work! Would you mind sharing the ALIKED and LightGlue parameters you used?\nI followed the setup you posted and built a similar version. After several runs and tuning attempts, I feel like you’ve really nailed down some excellent parameter choices. Would love to learn from what worked best for you!",
      "votes": null
    },
    {
      "id": "3355686",
      "postDate": "11/30/2025 17:34:54",
      "content": "<p>Congratulations on 3rd place ! Your solution is indeed a very smart attempt. I am a beginner in this competition. I was wondering if you or anyone else would be kind enough to review my solution . I tried to replicate your solution but either of the two things keep happening :\n1) A very low score like 14 or 15 \n2) Submission scoring error : because the timelimit ended and an entire submission csv file could not be generated . Therefore, illegal submission.\nGithub : <a href=\"https://github.com/MujtabaJunaid/imc-2025-naive-attempt/blob/main/3rd_position_copy_attempt.py\" target=\"_blank\">https://github.com/MujtabaJunaid/imc-2025-naive-attempt/blob/main/3rd_position_copy_attempt.py</a>\nI would be very thankful for any advice suggestions.\nThank you and my apologies for any inconvenience.</p>",
      "rawMarkdown": "Congratulations on 3rd place ! Your solution is indeed a very smart attempt. I am a beginner in this competition. I was wondering if you or anyone else would be kind enough to review my solution . I tried to replicate your solution but either of the two things keep happening :\n1) A very low score like 14 or 15 \n2) Submission scoring error : because the timelimit ended and an entire submission csv file could not be generated . Therefore, illegal submission.\nGithub : https://github.com/MujtabaJunaid/imc-2025-naive-attempt/blob/main/3rd_position_copy_attempt.py\nI would be very thankful for any advice suggestions.\nThank you and my apologies for any inconvenience.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3218778,
      "author_name": "tmyok1984",
      "author_url": "",
      "post_date": "06/06/2025 17:33:11",
      "content": "<p>Thanks for sharing your solution.  <br>\nI first learned about the tiled images approach from your solution two years ago,  <br>\nand I'm surprised to see it still worked effectively for this year's challenge.</p>\n<p>I'm curious why you decided to use this strategy again,  <br>\nand what kind of evaluation scores you got on the validation data.</p>\n<p>Do you have a table with numbers for <code>amy_gardens</code> or <code>fbk_vineyard</code> when using the tiling strategy?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3218788,
          "author_name": "tmyok1984",
          "author_url": "",
          "post_date": "06/06/2025 17:46:39",
          "content": "<p>Thanks for mentioning my post about VGGT!<br>\nIt definitely had great potential, but I couldn’t quite make it work — ended up not boosting the LB at all 😅</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3218943,
          "author_name": "roniheka",
          "author_url": "",
          "post_date": "06/07/2025 00:40:45",
          "content": "<p>I did experiment with a variety of other strategies for this year's challenge. I tried several combinations of multi-scale image sets and also tested cropping. However, for my particular pipeline, the tiling strategy still was found to be the most effective. </p>\n<p>I have updated the main text to include the evaluation results on the training data.<br>\nLater, I will provide an update with comparative data against other methods :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3218782,
      "author_name": "tmyok1984",
      "author_url": "",
      "post_date": "06/06/2025 17:41:39",
      "content": "<blockquote>\n  <p>While this may seem computationally expensive, the low computational cost of LightGlue made it feasible within the given time constraints.</p>\n</blockquote>\n<p>I understand that LightGlue has a relatively low computational cost, but performing 25 matching combinations per original image pair still seems computationally intensive.<br>\nCould you share any implementation strategies or optimizations you used to make this feasible?<br>\nIt would be great to learn from your approach — do you have any plans to share the implementation code?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3218927,
          "author_name": "roniheka",
          "author_url": "",
          "post_date": "06/06/2025 23:38:35",
          "content": "<p>Thank you for your comments!<br>\nTo keep the execution time under the 9-hour limit, I had to constrain min_pairs and sim_th as described above. Another necessary optimization was setting the LightGlue confidence threshold for pre-matching to 0.2, which helped reduce false positives among the selected pairs.</p>\n<p>I'll post the code here that takes a pair of image tensors performs the tiling and matching, returns the keypoint coordinates and ALIKED caches.<br>\nThe pipeline is inspired by <a href=\"https://www.kaggle.com/code/chankhavu/loftr-superglue-dkm-with-inspiration\" target=\"_blank\">this work</a>.</p>\n<pre><code> (torch.nn.Module):\n     ():\n        ().__init__()\n        .device=device\n        .extractor = ALIKED(weights=, **extractor_cfg).to(.device).()\n        .lightglue = LightGlue(**lg_cfg).to(.device)\n        .ttas = (())\n        .tta2id = {k: i  i, k  (.ttas)}\n        .tta_combination = [[i, j]  i  ()  j  ()]\n\n     ():\n        pred = {}        \n        cache, key1, key2, quad = cache_args \n        img_key_list = [key1,key2]\n        data[] = {: data[]}\n        data[] = {: data[]}\n        keypoints_dict, descriptors_dict = {}, {}\n         i, img_key  ([, ]):\n            keypoints_list, descriptors_list = [], []\n             i == :\n                _quad = \n            :\n                _quad = quad\n                cache[img_key_list[i]][_quad][]:\n                no_cache = \n            :\n                no_cache = \n            \n             j, img  (data[img_key][]):\n                 no_cache:\n                    img = img.unsqueeze()\n                    cache[img_key_list[i]][_quad][j][] = .extractor.extract(img, resize=)\n                pred = cache[img_key_list[i]][_quad][j][]\n                keypoints_list.append(pred[])\n                descriptors_list.append(pred[])\n            keypoints_dict[img_key] = keypoints_list\n            descriptors_dict[img_key] = descriptors_list\n        \n        group_pred_list = []\n         tta_group  .tta_combination:\n            group_idx = .tta2id[tta_group[]], .tta2id[tta_group[]]\n            i0, i1 = group_idx[], group_idx[]\n            data[][], data[][]= keypoints_dict[][i0], descriptors_dict[][i0]\n            data[][], data[][]= keypoints_dict[][i1], descriptors_dict[][i1]\n            group_pred = .lightglue(data)\n            group_pred.update({:data[][],\n                                :data[][]})\n            group_pred_list.append(group_pred)\n         group_pred_list, cache\n\n :\n     ():\n        .device = device\n        .extractor_cfg = extractor_cfg\n        .lg_cfg = lg_cfg\n        ._lightglue_matcher = LightGlueCustomMatching_sep(\n            device=.device, extractor_cfg=.extractor_cfg, lg_cfg=.lg_cfg\n            ).().to(device,dtype)\n        .conf_thresh = conf_th\n\n     ():\n        \n        img = img.clone()\n         long_side   :\n            scale = long_side / (img.shape[], img.shape[])\n            w = (img.shape[] * scale)\n            h = (img.shape[] * scale)\n            img = torch.nn.functional.interpolate(img, size=(h, w), mode=, align_corners=)\n        :\n            scale = \n         img, scale\n\n     ():\n        \n        h, w = image.shape[], image.shape[]\n         h %  != :\n            h = h - \n         w %  != :\n            w = w - \n        image = image[:, :, :h, :w]\n         [image[:, :, :h//, :w//], \n            image[:, :, :h//, w//:], \n            image[:, :, h//:, :w//], \n            image[:, :, h//:, w//:],\n            transforms.functional.resize(image, size=(h//,w//))]\n\n     ():\n        \n         quadrant == :\n            coords[:, ] += w//\n         quadrant == :\n            coords[:, ] += h//\n         quadrant == :\n            coords[:, ] += w//\n            coords[:, ] += h//\n         quadrant == :\n            coords = [[y*, x*]  y, x  coords]            \n         coords\n\n     ():\n         torch.no_grad():\n            img_ts0, scale0 = .prep_img(img_ts0, input_longside)\n            img_ts1, scale1 = .prep_img(img_ts1, input_longside)\n            img_parts0 = .split_image(img_ts0) \n            img_parts1 = .split_image(img_ts1)\n            cat_mkpts0, cat_mkpts1 = [], []\n            pred, cache = ._lightglue_matcher.forward_flat(\n                data={\n                    : torch.cat(img_parts0),\n                    : torch.cat(img_parts1),\n                },\n            cache_args=cache_args)\n\n            mkpts0, mkpts1= [], []\n             idx, [i0,i1]  (.tta_combination):\n                group_pred = pred[idx]\n                pred_aug={}\n                use_keys = [, , , ]\n                 k  use_keys:\n                    v = group_pred[k]\n                     (v, torch.Tensor):\n                        pred_aug[k] = v[].detach().cpu().numpy().squeeze()\n                    :\n                        pred_aug[k] = v               \n                kpts0, kpts1 = pred_aug[], pred_aug[]\n                matches = pred_aug[]\n                valid = matches &gt; -\n                \n                 (kpts0[valid]) &gt; :\n                    mkpts0 = .reconstruct_coords(kpts0[valid], i0, img_ts0.shape[], img_ts0.shape[])\n                    mkpts1 = .reconstruct_coords(kpts1[matches[valid]], i1, img_ts1.shape[], img_ts1.shape[])\n                    cat_mkpts0.append(mkpts0)\n                    cat_mkpts1.append(mkpts1)\n\n             (cat_mkpts0) &gt; :\n                cat_mkpts0 = np.concatenate(cat_mkpts0)\n                cat_mkpts1 = np.concatenate(cat_mkpts1)\n            :\n                _, key1, key2, _ = cache_args\n                ()\n                 np.empty((, )), np.empty((, )), cache\n            \n            :\n                _, inliers = cv2.findFundamentalMat(cat_mkpts0, cat_mkpts1, cv2.USAC_MAGSAC, ransacReprojThreshold=, confidence=, maxIters=)\n                inliers = inliers.ravel() &gt; \n                cat_mkpts0 = cat_mkpts0[inliers]\n                cat_mkpts1 = cat_mkpts1[inliers]\n             Exception:\n                _, key1, key2, _ = cache_args\n                ()\n                 np.empty((, )), np.empty((, )), cache\n\n            mask0 = (cat_mkpts0[:, ] &gt;= ) &amp; (cat_mkpts0[:, ] &lt; img_ts0.shape[]) &amp; (cat_mkpts0[:, ] &gt;= ) &amp; (cat_mkpts0[:, ] &lt; img_ts0.shape[])\n            mask1 = (cat_mkpts1[:, ] &gt;= ) &amp; (cat_mkpts1[:, ] &lt; img_ts1.shape[]) &amp; (cat_mkpts1[:, ] &gt;= ) &amp; (cat_mkpts1[:, ] &lt; img_ts1.shape[])\n             cat_mkpts0[mask0 &amp; mask1] / scale0, cat_mkpts1[mask0 &amp; mask1] / scale1, cache\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3218933,
      "author_name": "tyyuki",
      "author_url": "",
      "post_date": "06/06/2025 23:47:44",
      "content": "<p>Congratulations on 3rd place! Your approach with rotation augmentation and the 5-image tiling strategy is truly innovative for robust matching. <br>\nCould you elaborate on why you chose a min_pair of 50-60 and a sim_th of 0.2 for your pre-matching stage?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3218936,
          "author_name": "roniheka",
          "author_url": "",
          "post_date": "06/07/2025 00:06:04",
          "content": "<p>Thanks for the congratulations and for the great question!<br>\nThe main reasons were:</p>\n<ul>\n<li>The distribution of similarity scores obtained by DINOv2 changes significantly from one dataset to another. On a dataset where the average similarity is high, a fixed threshold could generate a massive number of candidate pairs. This would spend too much time on pre-matching at a specific dataset. This is why the sim_th is set to low.</li>\n<li>The subsequent matching process on the tiled images is time-consuming. Through experimentation, we found that increasing min_pair above the 50-60 range pushed our total runtime over the limit, which would cause a timeout.</li>\n</ul>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3220952,
      "author_name": "yuanlin08",
      "author_url": "",
      "post_date": "06/10/2025 07:33:41",
      "content": "<p>Hi! Congrats again — super impressive work! Would you mind sharing the ALIKED and LightGlue parameters you used?<br>\nI followed the setup you posted and built a similar version. After several runs and tuning attempts, I feel like you’ve really nailed down some excellent parameter choices. Would love to learn from what worked best for you!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3355686,
      "author_name": "mujtabajunaid",
      "author_url": "",
      "post_date": "11/30/2025 17:34:54",
      "content": "<p>Congratulations on 3rd place ! Your solution is indeed a very smart attempt. I am a beginner in this competition. I was wondering if you or anyone else would be kind enough to review my solution . I tried to replicate your solution but either of the two things keep happening :\n1) A very low score like 14 or 15 \n2) Submission scoring error : because the timelimit ended and an entire submission csv file could not be generated . Therefore, illegal submission.\nGithub : <a href=\"https://github.com/MujtabaJunaid/imc-2025-naive-attempt/blob/main/3rd_position_copy_attempt.py\" target=\"_blank\">https://github.com/MujtabaJunaid/imc-2025-naive-attempt/blob/main/3rd_position_copy_attempt.py</a>\nI would be very thankful for any advice suggestions.\nThank you and my apologies for any inconvenience.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3218701": "First and foremost, I would like to express my sincere gratitude to the organizers for hosting the Image Matching Challenge again this year. I am always inspired by the new themes you introduce each year, and I truly enjoy the opportunity to learn and experiment with new techniques.\n\n## Overview\nMy solution consists of a pre-matching stage with rotation augmentation, followed by multiple matching rounds using tiled images. The core ideas of rotation augmentation and image tiling were also used in [my approach at IMC2023](https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/416918).\n\nThe overall pipeline is summarized in this diagram:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9249230%2Ffbfdc7bd888157ce842fdacc4567cbf6%2F20250606_solution.pptx%20(1).png?generation=1749222583567967&alt=media)\n\nLike many other teams, I utilized a T4x2 GPU setup for parallel processing.\n\n## Solution Details\n### 2.1. Pre-matching\nI leveraged the baseline pipeline, using DINOv2 to generate a list of image pairs with high similarity. I found the optimal parameters to be a minimum of 50-60 pairs per image (min_pair) and a similarity threshold (sim_th) of 0.2.\n\nFor each pair, I performed matching using ALIKED + LightGlue (with longside set to 960px and max_keypoints to 4096). Pairs with more than 40 inlier matches were considered valid. Since some datasets contained rotated images, I implemented a retry mechanism with rotation augmentation for pairs that initially failed to match.\n\n### 2.2. Matching with Tiled Images\nFor the valid pairs from the pre-matching stage, I uniformly tiled each image into four quadrants. I then performed matching on these tiles, again using ALIKED + LightGlue (max_keypoints=4096). The optimal longside setting for the tiles was either 1024 or 1216.\n\nA key part of my strategy was to create a set of five images for matching: the four tiles plus the original image downscaled by half. This resulted in 25 matching combinations per original image pair (5x5), which allows for more robust matching between images with significant differences in scale and perspective. While this may seem computationally expensive, the low computational cost of LightGlue made it feasible within the given time constraints. \n*UPDATE*: I've posted the implemention of the pipeline in the *comments* section below.\n\nA reperesentative scores by the training data is summarized here. I couldn't find an effective method for the stairs dataset in the end...\n|heritage\t|theather_church\t|dioscuri_baalshamin\t|lizard_pond\t|piazzasanmarco\t|amy_gardens\t|fbk_vineyard\t|ETs\t|stairs\t|Average\t|Public\t|Private|\n|:--|:--|:--|:--|:--|:--|:--|:--|:--|:--|:--|:--|\n|88.13\t|61.39\t|91.73\t|76.16\t|60.25\t|23.91\t|37.88\t|66.67\t|4.3\t|56.71\t|48.97\t|49.05|\n\n\nActually, I have not fully verified why this specific tiling strategy was so effective. However, I believe one reason is that it helps mitigate the issue of keypoints being concentrated in only one part of the image.\n\nTo improve time efficiency, I cached the results of the ALIKED feature extraction for reuse. After matching, I restored the keypoint coordinates to their original image space, used RANSAC to filter outliers, and then combined these keypoints with those from the pre-matching stage. The RANSAC parameters were taken from [the 2024 1st place team's solution,](https://www.kaggle.com/code/vostankovich/imc2024-1st-place-solution) and I would like to extend my thanks to them.\n\n### 2.3. Other Points\nRegistering pycolmap results: My pycolmap processing was based on the official baseline. However, I paid an attention to the order in which reconstruction results from incremental_mapping are registered. My understanding for this competition is that a larger number of registered images in the largest cluster of a scene generally leads to a better score (assuming the estimated camera poses are correct). Since images can be part of multiple clusters, it's crucial to prevent the results from larger, more robust clusters from being overwritten by those from smaller clusters.\nTo achieve this, I sorted the reconstructed maps by the number of registered images in ascending order before registering the final poses. This modification seemed to yield a small boost on the leaderboard.\n\nLike this:\n``` \nmap_info_list = []\nfor map_idx, cur_map in maps.items():\n    num_registered_images = cur_map.num_registered_images()\n    map_info_list.append({\n        \"original_map_index\": map_idx,\n        \"map_object\": cur_map,\n        \"num_reg_images\": num_registered_images\n    })\n# Sort by the number of registered images in ascending order\nsorted_map_info_list = sorted(map_info_list, key=lambda x: x[\"num_reg_images\"])\nfor map_info in sorted_map_info_list:\n    # (Register poses...)\n``` \n\nParameter Tuning: The parameters for ALIKED and LightGlue, as well as the input image sizes, had a significant impact on accuracy. Careful tuning was essential.\n\n## What Didn't Work in My Case\n- Using SIFT instead of ALIKED: While SIFT was faster, it resulted in a lower score on most datasets compared to ALIKED. I didn't have a chance to try an ensemble of ALIKED and SIFT.\n- Scene Pre-Clustering: I attempted to pre-cluster images within datasets that contained multiple scenes.\nI tried clustering based on a distance matrix derived from the number of keypoint matches from the pre-matching stage, but I couldn't find a reliable method to accurately segment all scenes.\nEven when scenes were correctly segmented, it did not lead to a significant score improvement (for this competition's metric) unless the subsequent matching step could find enough keypoints for a robust reconstruction.\n- While pairwise matching-based clustering had low accuracy, I found VGGT to be powerful from a clustering perspective, as it seems to find context among a group of dozens of images. I experimented with using SIFT feature points as query_points and identifying pairs based on the number of points exceeding vis_score and conf_score thresholds for a target image. This method achieved a higher true positive rate for pair identification than the pairwise approach in section 2.1 and successfully clustered many datasets.\nHowever, in terms of the resulting pose accuracy and final score, it did not show a significant advantage in my setup, possibly due to the resizing.\nThe use of VGGT has been deeply explored and discussed by @tmyok1984 san in [this discussion post.](https://www.kaggle.com/competitions/image-matching-challenge-2025/discussion/582968) I'm learning a great deal from reading the post.",
    "3218778": "Thanks for sharing your solution.  \nI first learned about the tiled images approach from your solution two years ago,  \nand I'm surprised to see it still worked effectively for this year's challenge.\n\nI'm curious why you decided to use this strategy again,  \nand what kind of evaluation scores you got on the validation data.\n\nDo you have a table with numbers for `amy_gardens` or `fbk_vineyard` when using the tiling strategy?",
    "3218782": "> While this may seem computationally expensive, the low computational cost of LightGlue made it feasible within the given time constraints.\n\nI understand that LightGlue has a relatively low computational cost, but performing 25 matching combinations per original image pair still seems computationally intensive.\nCould you share any implementation strategies or optimizations you used to make this feasible?\nIt would be great to learn from your approach — do you have any plans to share the implementation code?",
    "3218788": "Thanks for mentioning my post about VGGT!\nIt definitely had great potential, but I couldn’t quite make it work — ended up not boosting the LB at all 😅",
    "3218927": "Thank you for your comments!\nTo keep the execution time under the 9-hour limit, I had to constrain min_pairs and sim_th as described above. Another necessary optimization was setting the LightGlue confidence threshold for pre-matching to 0.2, which helped reduce false positives among the selected pairs.\n\nI'll post the code here that takes a pair of image tensors performs the tiling and matching, returns the keypoint coordinates and ALIKED caches.\nThe pipeline is inspired by [this work](https://www.kaggle.com/code/chankhavu/loftr-superglue-dkm-with-inspiration).\n\n```\nclass LightGlueCustomMatching_sep(torch.nn.Module):\n    def __init__(self, device=None, extractor_cfg=None, lg_cfg=None):\n        super().__init__()\n        self.device=device\n        self.extractor = ALIKED(weights=f\"/kaggle/input/imc24lightglue/weights/aliked-n16.pth\", **extractor_cfg).to(self.device).eval()\n        self.lightglue = LightGlue(**lg_cfg).to(self.device)\n        self.ttas = list(range(5))\n        self.tta2id = {k: i for i, k in enumerate(self.ttas)}\n        self.tta_combination = [[i, j] for i in range(5) for j in range(5)]\n\n    def forward_flat(self, data, cache_args):\n        pred = {}        \n        cache, key1, key2, quad = cache_args # quad: Rotation times of image1. 0:No rotation, 1:90deg, 2:180deg, 3:270deg\n        img_key_list = [key1,key2]\n        data[\"image0\"] = {\"image\": data[\"image0\"]}\n        data[\"image1\"] = {\"image\": data[\"image1\"]}\n        keypoints_dict, descriptors_dict = {}, {}\n        for i, img_key in enumerate(['image0', 'image1']):\n            keypoints_list, descriptors_list = [], []\n            if i == 0:\n                _quad = 0\n            else:\n                _quad = quad\n            if \"pred\" not in cache[img_key_list[i]][_quad][0]:\n                no_cache = True\n            else:\n                no_cache = False\n            # Get ALIKED descriptors\n            for j, img in enumerate(data[img_key][\"image\"]):\n                if no_cache:\n                    img = img.unsqueeze(0)\n                    cache[img_key_list[i]][_quad][j][\"pred\"] = self.extractor.extract(img, resize=None)\n                pred = cache[img_key_list[i]][_quad][j][\"pred\"]\n                keypoints_list.append(pred['keypoints'])\n                descriptors_list.append(pred['descriptors'])\n            keypoints_dict[img_key] = keypoints_list\n            descriptors_dict[img_key] = descriptors_list\n        # Prepare data for LightGlue and run matching one by one\n        group_pred_list = []\n        for tta_group in self.tta_combination:\n            group_idx = self.tta2id[tta_group[0]], self.tta2id[tta_group[1]]\n            i0, i1 = group_idx[0], group_idx[1]\n            data[\"image0\"][\"keypoints\"], data[\"image0\"][\"descriptors\"]= keypoints_dict['image0'][i0], descriptors_dict['image0'][i0]\n            data[\"image1\"][\"keypoints\"], data[\"image1\"][\"descriptors\"]= keypoints_dict['image1'][i1], descriptors_dict['image1'][i1]\n            group_pred = self.lightglue(data)\n            group_pred.update({\"keypoints0\":data['image0'][\"keypoints\"],\n                                \"keypoints1\":data['image1'][\"keypoints\"]})\n            group_pred_list.append(group_pred)\n        return group_pred_list, cache\n    \nclass LightGlueMatcherPipeline_sep:\n    def __init__(self, device=None, conf_th=None, extractor_cfg=None, lg_cfg=None):\n        self.device = device\n        self.extractor_cfg = extractor_cfg\n        self.lg_cfg = lg_cfg\n        self._lightglue_matcher = LightGlueCustomMatching_sep(\n            device=self.device, extractor_cfg=self.extractor_cfg, lg_cfg=self.lg_cfg\n            ).eval().to(device,dtype)\n        self.conf_thresh = conf_th\n    \n    def prep_img(self, img, long_side=None):\n        \"\"\"Resize the tensor image to a specified long side.\"\"\"\n        img = img.clone()\n        if long_side is not None:\n            scale = long_side / max(img.shape[2], img.shape[3])\n            w = int(img.shape[3] * scale)\n            h = int(img.shape[2] * scale)\n            img = torch.nn.functional.interpolate(img, size=(h, w), mode='bilinear', align_corners=False)\n        else:\n            scale = 1.0\n        return img, scale\n    \n    def split_image(self, image):\n        \"\"\"Split the image into 4 quadrants and return them along with a resized version.\"\"\"\n        h, w = image.shape[2], image.shape[3]\n        if h % 2 != 0:\n            h = h - 1\n        if w % 2 != 0:\n            w = w - 1\n        image = image[:, :, :h, :w]\n        return [image[:, :, :h//2, :w//2], \n            image[:, :, :h//2, w//2:], \n            image[:, :, h//2:, :w//2], \n            image[:, :, h//2:, w//2:],\n            transforms.functional.resize(image, size=(h//2,w//2))]\n\n    def reconstruct_coords(self, coords, quadrant, w, h):\n        \"\"\"Reconstruct coordinates based on the separation quadrant.\"\"\"\n        if quadrant == 1:\n            coords[:, 0] += w//2\n        elif quadrant == 2:\n            coords[:, 1] += h//2\n        elif quadrant == 3:\n            coords[:, 0] += w//2\n            coords[:, 1] += h//2\n        elif quadrant == 4:\n            coords = [[y*2, x*2] for y, x in coords]            \n        return coords\n    \n    def __call__(self, img_ts0, img_ts1, cache_args, input_longside=None):\n        with torch.no_grad():\n            img_ts0, scale0 = self.prep_img(img_ts0, input_longside)\n            img_ts1, scale1 = self.prep_img(img_ts1, input_longside)\n            img_parts0 = self.split_image(img_ts0) \n            img_parts1 = self.split_image(img_ts1)\n            cat_mkpts0, cat_mkpts1 = [], []\n            pred, cache = self._lightglue_matcher.forward_flat(\n                data={\n                    \"image0\": torch.cat(img_parts0),\n                    \"image1\": torch.cat(img_parts1),\n                },\n            cache_args=cache_args)\n\n            mkpts0, mkpts1= [], []\n            for idx, [i0,i1] in enumerate(self.tta_combination):\n                group_pred = pred[idx]\n                pred_aug={}\n                use_keys = [\"keypoints0\", \"keypoints1\", \"matches0\", \"matching_scores0\"]\n                for k in use_keys:\n                    v = group_pred[k]\n                    if isinstance(v, torch.Tensor):\n                        pred_aug[k] = v[0].detach().cpu().numpy().squeeze()\n                    else:\n                        pred_aug[k] = v               \n                kpts0, kpts1 = pred_aug[\"keypoints0\"], pred_aug[\"keypoints1\"]\n                matches = pred_aug[\"matches0\"]\n                valid = matches > -1\n                # Recover keypoint coordinates based on the quadrant\n                if len(kpts0[valid]) > 0:\n                    mkpts0 = self.reconstruct_coords(kpts0[valid], i0, img_ts0.shape[3], img_ts0.shape[2])\n                    mkpts1 = self.reconstruct_coords(kpts1[matches[valid]], i1, img_ts1.shape[3], img_ts1.shape[2])\n                    cat_mkpts0.append(mkpts0)\n                    cat_mkpts1.append(mkpts1)\n\n            if len(cat_mkpts0) > 0:\n                cat_mkpts0 = np.concatenate(cat_mkpts0)\n                cat_mkpts1 = np.concatenate(cat_mkpts1)\n            else:\n                _, key1, key2, _ = cache_args\n                print(f\"No matches at {key1} vs. {key2}\")\n                return np.empty((0, 2)), np.empty((0, 2)), cache\n            # RANSAC\n            try:\n                _, inliers = cv2.findFundamentalMat(cat_mkpts0, cat_mkpts1, cv2.USAC_MAGSAC, ransacReprojThreshold=5, confidence=0.9999, maxIters=50000)\n                inliers = inliers.ravel() > 0\n                cat_mkpts0 = cat_mkpts0[inliers]\n                cat_mkpts1 = cat_mkpts1[inliers]\n            except Exception:\n                _, key1, key2, _ = cache_args\n                print(f\"Error in findFundamentalMat: {key1}-{key2}\")\n                return np.empty((0, 2)), np.empty((0, 2)), cache\n\n            mask0 = (cat_mkpts0[:, 0] >= 0) & (cat_mkpts0[:, 0] < img_ts0.shape[3]) & (cat_mkpts0[:, 1] >= 0) & (cat_mkpts0[:, 1] < img_ts0.shape[2])\n            mask1 = (cat_mkpts1[:, 0] >= 0) & (cat_mkpts1[:, 0] < img_ts1.shape[3]) & (cat_mkpts1[:, 1] >= 0) & (cat_mkpts1[:, 1] < img_ts1.shape[2])\n            return cat_mkpts0[mask0 & mask1] / scale0, cat_mkpts1[mask0 & mask1] / scale1, cache\n```",
    "3218933": "Congratulations on 3rd place! Your approach with rotation augmentation and the 5-image tiling strategy is truly innovative for robust matching. \nCould you elaborate on why you chose a min_pair of 50-60 and a sim_th of 0.2 for your pre-matching stage?",
    "3218936": "Thanks for the congratulations and for the great question!\nThe main reasons were:\n- The distribution of similarity scores obtained by DINOv2 changes significantly from one dataset to another. On a dataset where the average similarity is high, a fixed threshold could generate a massive number of candidate pairs. This would spend too much time on pre-matching at a specific dataset. This is why the sim_th is set to low.\n- The subsequent matching process on the tiled images is time-consuming. Through experimentation, we found that increasing min_pair above the 50-60 range pushed our total runtime over the limit, which would cause a timeout.",
    "3218943": "I did experiment with a variety of other strategies for this year's challenge. I tried several combinations of multi-scale image sets and also tested cropping. However, for my particular pipeline, the tiling strategy still was found to be the most effective. \n\nI have updated the main text to include the evaluation results on the training data.\nLater, I will provide an update with comparative data against other methods :)",
    "3220952": "Hi! Congrats again — super impressive work! Would you mind sharing the ALIKED and LightGlue parameters you used?\nI followed the setup you posted and built a similar version. After several runs and tuning attempts, I feel like you’ve really nailed down some excellent parameter choices. Would love to learn from what worked best for you!",
    "3355686": "Congratulations on 3rd place ! Your solution is indeed a very smart attempt. I am a beginner in this competition. I was wondering if you or anyone else would be kind enough to review my solution . I tried to replicate your solution but either of the two things keep happening :\n1) A very low score like 14 or 15 \n2) Submission scoring error : because the timelimit ended and an entire submission csv file could not be generated . Therefore, illegal submission.\nGithub : https://github.com/MujtabaJunaid/imc-2025-naive-attempt/blob/main/3rd_position_copy_attempt.py\nI would be very thankful for any advice suggestions.\nThank you and my apologies for any inconvenience."
  },
  "source": "meta"
}