{
  "id": 583575,
  "title": "80th Place Solution: Two-Stage ALIKED+LightGlue Pipeline and Adaptive OmniGlue-Enhancement Pipeline",
  "url": "/competitions/image-matching-challenge-2025/discussion/583575",
  "author_name": "kappa",
  "post_date": "2025-06-07T22:47:11.739000",
  "votes": 9,
  "comment_count": 0,
  "views": 0,
  "content": "<h1>80th Place Solution: Two-Stage ALIKED+LightGlue Pipeline and Adaptive OmniGlue-Enhancement Pipeline</h1>\n<p>We thank the IMC 2025 organisers for hosting this challenging competition and the Kaggle community for the valuable discussions and insights. Our implementation builds upon excellent prior work in the image matching community.</p>\n<p>We have noticed that no OmniGlue-enhanced/based solution is mentioned in the discussion. Although our solution did not perform as well as the top solutions, it might be worth sharing our solution and findings with the community.</p>\n<h2>Overview</h2>\n<p>We, <a href=\"https://www.kaggle.com/kazumax0720\" target=\"_blank\">@kazumax0720</a> and I, have prepared two pipelines for the final submission: 1) DINOv2 then Two-stage ALIKED+LightGlue and 2) 1)+budget  constraint OmniGlue matching pipeline.</p>\n<p>Our approach uses <a href=\"https://www.kaggle.com/octaviograu\" target=\"_blank\">@octaviograu</a>'s <a href=\"https://www.kaggle.com/code/octaviograu/baseline-dinov2-aliked-lightglue\" target=\"_blank\">public notebook</a> as a baseline and added a few components such as image rotation, second-stage matching and OmniGlue matching.</p>\n<h2>Method</h2>\n<h3>1) DINOv2 then Two-stage ALIKED+LightGlue</h3>\n<h4>Stage 1: Image Retrieval with DINOv2</h4>\n<ul>\n<li>Extract global features using a DINOv2 vision transformer</li>\n<li>Perform similarity-based image pairing to reduce computational complexity</li>\n<li>Apply adaptive thresholding to balance recall against computational efficiency</li>\n</ul>\n<h4>Stage 2: Two-Stage Feature Matching</h4>\n<ul>\n<li>Primary matching: ALIKED keypoint detection + LightGlue matching</li>\n<li>Secondary matching: Crop matching with ALIKED + LightGlue</li>\n<li>Geometric verification: MAGSAC-based outlier rejection with adaptive thresholds</li>\n</ul>\n<p>The below diagram  demonstrates the DINOv2 then Two-stage ALIKED+LightGlue:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6398606%2F56f15dd4a67b30f2abac24f712611321%2FIMC25-solution.png?generation=1749334903365639&amp;alt=media\" alt=\"\"></p>\n<h3>2) DINOv2 then Two-stage ALIKED+LightGlue + OmniGlue Geometric failure trigger</h3>\n<p>OmniGlue is computationally more expensive than ALIKED+LightGlue. Thus, building upon the core pipeline above, we have developed an adaptive OmniGlue enhancement pipeline.</p>\n<p>The diagram below demonstrates the OmniGlue-enhanced pipeline:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6398606%2F9f7526a6268d18b62fd096519c38eac2%2FIMC25-solution-OmniGlue.png?generation=1749334934055549&amp;alt=media\" alt=\"\"></p>\n<h4>GPU Optimisation</h4>\n<p>The original OmniGlue implementation had significant performance bottlenecks that made it impractical for large-scale competition use, especially CPU-bound DINO feature extraction. The original implementation ran DINO features on CPU, creating a major computational bottleneck. Therefore, we have moved the entire DINO feature extraction pipeline to GPU:</p>\n<pre><code> :\n     ():\n        .device = torch.device(device  torch.cuda.is_available()  )\n        .use_mixed_precision = use_mixed_precision\n\n        \n        .model = dino.vit_base()\n        state_dict = torch.load(cpt_path, map_location=, weights_only=)\n        .model.load_state_dict(state_dict)\n        .model = .model.to(.device)\n        .model.()\n</code></pre>\n<p>By implementing this and other minor tricks, we could shorten the image pair processing time from 10 seconds to 1 - 2 seconds per image pair, which made it possible for us to use OmniGlue partially in our pipeline.</p>\n<h4>Geometric Distribution Trigger</h4>\n<p>OmniGlue matching is triggered by \"poor\" geometric distribution of ALIKED+LightGlue matches as even with GPU-optimisation of OmniGlue, it is still computationally too expensive to use OmniGlue for all the image pairs (after seeing other top solutions, it probably was possible to limit the number of image pairs to process). Because of this constraint, we also limit the trigger of OmniGlue up to 20% of the dataset.</p>\n<p>The \"poor\" geometric distribution trigger is rather a simple diagnosis combination of the initial matches distribution. Here are three criteria for triggering, and if two criteria are satisfied, the OmniGlue matching would be triggered. The following are the criteria:</p>\n<p>1) Insufficient matches: check if the number of matches is simply less than threshold=6<br>\n2) Spatial Coverage:</p>\n<pre><code>\nx_range = np.(x_coords) - np.(x_coords)\ny_range = np.(y_coords) - np.(y_coords)\n\n\nx_std = np.std(x_coords)\ny_std = np.std(y_coords)\n\n\nmin_range_threshold = (, x_std * geometric_multiplier, y_std * geometric_multiplier)\n\n\npoor_indicators.append(x_range &lt; min_range_threshold)\npoor_indicators.append(y_range &lt; min_range_threshold)\n</code></pre>\n<p>3) Grid-based clustering:</p>\n<p>We divide the image into a 4x4 grid and check the occupancy. If the occupancy is less than the threshold of 0.5, it is categorised as a poor geometrical distribution</p>\n<pre><code>\nn_grid =   \nx_bins = np.linspace(np.(x_coords), np.(x_coords), n_grid + )\ny_bins = np.linspace(np.(y_coords), np.(y_coords), n_grid + )\n\n\ngrid_counts = np.zeros((n_grid, n_grid))\n i  ((x_coords)):\n    x_idx = (n_grid - , (, np.digitize(x_coords[i], x_bins) - ))\n    y_idx = (n_grid - , (, np.digitize(y_coords[i], y_bins) - ))\n    grid_counts[x_idx, y_idx] += \n\n\nnon_empty_cells = np.(grid_counts &gt; )\ntotal_cells = n_grid * n_grid\npoor_indicators.append(non_empty_cells &lt; total_cells / )\n</code></pre>\n<p>When OmniGlue is triggered, 60 - 70% of them successfully add new matches to the original matches.</p>\n<h2>Final Submission Results</h2>\n<p>The table below demonstrates the private and public leaderboard scores of our final submission solutions, along with local cross-validation scores. Although the local CV scores suggest the OmniGlue-enhanced pipeline performs better than the Two-stage ALIKED+LightGlue matching pipeline, the leaderboard scores of the Two-stage ALIKED+LightGlue pipeline are better than the OmniGlue-enhanced pipeline.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>Private LB</th>\n<th>Public LB</th>\n<th>ETs</th>\n<th>amy_gardens</th>\n<th>fbk_vineyard</th>\n<th>stairs</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Two-stage ALIKED+LightGlue<br>(80th place)</td>\n<td>34.18</td>\n<td>34.86</td>\n<td>37.35</td>\n<td>40.16</td>\n<td>25.35</td>\n<td>6.25</td>\n</tr>\n<tr>\n<td>+OmniGlue Geometric failure<br> trigger 20%</td>\n<td>32.57</td>\n<td>32.20</td>\n<td>37.55</td>\n<td>39.51</td>\n<td>30.48</td>\n<td>9.09</td>\n</tr>\n</tbody>\n</table>\n<h2>Discussion and Analysis</h2>\n<h3>Score Robustness Analysis</h3>\n<p>We have noticed randomness in the OmniGlue-enhanced pipelines leaderboard score. After the competition ended, we have submitted the same solutions 3 times and summarised the statistics in the table below. The table below lists 3 pipelines' results - both of our final submission and a random 15% of a dataset OmniGlue matching + fully triggered OmniGlue matching for the \"stairs\" dataset.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>Two-stage<br>ALIKED+LightGlue</th>\n<th>+OmniGlue Geometric<br>failure trigger 20%</th>\n<th>+15% random OmniGlue<br>+ Full stairs dataset</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Mean Private LB</td>\n<td>33.98</td>\n<td>34.9</td>\n<td>35.94</td>\n</tr>\n<tr>\n<td>Mean Public LB</td>\n<td>34.86</td>\n<td>33.42</td>\n<td>32.46</td>\n</tr>\n<tr>\n<td>Private STD</td>\n<td>0.84</td>\n<td>2.89</td>\n<td>1.67</td>\n</tr>\n<tr>\n<td>Public STD</td>\n<td>0.46</td>\n<td>2.84</td>\n<td>1.26</td>\n</tr>\n<tr>\n<td>Private Max</td>\n<td>34.59</td>\n<td>37.07</td>\n<td>37.83</td>\n</tr>\n<tr>\n<td>Public Max</td>\n<td>34.33</td>\n<td>35.09</td>\n<td>33.43</td>\n</tr>\n<tr>\n<td>Private Min</td>\n<td>33.02</td>\n<td>31.62</td>\n<td>34.68</td>\n</tr>\n<tr>\n<td>Public Min</td>\n<td>33.44</td>\n<td>30.14</td>\n<td>31.04</td>\n</tr>\n</tbody>\n</table>\n<p>The table demonstrates that the Two-stage ALIKED+LightGlue scores are relatively stable (standard deviation of 0.46 in Public LB). On the other hand, both OmniGlue-enhanced pipelines' standard deviations are larger (2.84 and 1.26 in Public LB, respectively). This is discussed in the Score Variance Investigation section.</p>\n<h3>Geometric OmniGlue Trigger Analysis</h3>\n<p>Our geometric failure trigger was designed with the intuition that poorly distributed matches indicate potential improvement through OmniGlue. However, experimental results and some intuitions suggest this hypothesis requires refinement:</p>\n<h4>Potential Issues with Geometric Triggering</h4>\n<ul>\n<li>Object-centric datasets: In scenes like \"ETs\" with single prominent objects, matches naturally concentrate around facial/body features</li>\n<li>Scene coherence: Some image pairs may legitimately lack geometric diversity</li>\n<li>Inter-cluster confusion: Triggering on geometric failure may attempt to match fundamentally different scenes</li>\n</ul>\n<p>This is an example of Inter-cluster confusion, where we should have discarded the pair instead of triggering OmniGlue. Triggering OmniGlue for such a pair is not only useless for the reconstruction but also wastes computational resources.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6398606%2Ffc34b66d2161e847e154a98f3fbb81fc%2FOmniGlue-ETs.png?generation=1749334981140001&amp;alt=media\" alt=\"\"></p>\n<h4>Evidence for Alternative Strategies</h4>\n<ul>\n<li>Random 15% triggering achieved better average performance than geometric triggering</li>\n<li>Suggests OmniGlue's advantage may be in overall matching quality rather than \"rescue\" scenarios</li>\n</ul>\n<p>This is where OmniGlue performs better than ALIKED+LightGlue matching.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6398606%2F023a75a9d82278c792aa907801a7e695%2FOmniGlue-staris.png?generation=1749335012683604&amp;alt=media\" alt=\"\"></p>\n<h3>Score Variance Investigation</h3>\n<p>The significant variance in OmniGlue-enhanced solutions raises important questions about evaluation stability.</p>\n<p>Hypothesised Contributing Factors:</p>\n<ul>\n<li>Threshold sensitivity: Small changes in match quality affecting clustering decisions</li>\n<li>Computational precision: GPU vs CPU differences in floating-point operations</li>\n</ul>\n<h3>Ablation Study</h3>\n<p>The table below shows the incremental development of our pipeline. The reported scores are the highest scores of the same pipeline.</p>\n<table>\n<thead>\n<tr>\n<th>Pipeline</th>\n<th>Private LB</th>\n<th>Public LB</th>\n<th>Remarks</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Baseline DINOv2 ALIKED+LG</td>\n<td>30.34</td>\n<td>32.38</td>\n<td>Public notebook</td>\n</tr>\n<tr>\n<td>+Two-stage matching</td>\n<td>34.56</td>\n<td>33.45</td>\n<td></td>\n</tr>\n<tr>\n<td>+Image Rotation</td>\n<td>34.18</td>\n<td>34.86</td>\n<td>Final core pipeline (80th place)</td>\n</tr>\n<tr>\n<td>+OmniGlue (geometric trigger)</td>\n<td>37.07</td>\n<td>35.09</td>\n<td>Another submitted solution</td>\n</tr>\n<tr>\n<td>+OmniGlue for stairs + 15% random</td>\n<td>37.83</td>\n<td>33.43</td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<h3>Key takeaways</h3>\n<ul>\n<li>GPU-optimised OmniGlue implementation</li>\n<li>Grid-based match distribution: how we used it in the pipeline might not be the best method, however this way of analysis for the match distribution would be useful</li>\n</ul>",
  "messages": [
    {
      "id": 3219542,
      "postDate": "2025-06-07T22:47:11.740Z",
      "content": "<h1>80th Place Solution: Two-Stage ALIKED+LightGlue Pipeline and Adaptive OmniGlue-Enhancement Pipeline</h1>\n<p>We thank the IMC 2025 organisers for hosting this challenging competition and the Kaggle community for the valuable discussions and insights. Our implementation builds upon excellent prior work in the image matching community.</p>\n<p>We have noticed that no OmniGlue-enhanced/based solution is mentioned in the discussion. Although our solution did not perform as well as the top solutions, it might be worth sharing our solution and findings with the community.</p>\n<h2>Overview</h2>\n<p>We, <a href=\"https://www.kaggle.com/kazumax0720\" target=\"_blank\">@kazumax0720</a> and I, have prepared two pipelines for the final submission: 1) DINOv2 then Two-stage ALIKED+LightGlue and 2) 1)+budget  constraint OmniGlue matching pipeline.</p>\n<p>Our approach uses <a href=\"https://www.kaggle.com/octaviograu\" target=\"_blank\">@octaviograu</a>'s <a href=\"https://www.kaggle.com/code/octaviograu/baseline-dinov2-aliked-lightglue\" target=\"_blank\">public notebook</a> as a baseline and added a few components such as image rotation, second-stage matching and OmniGlue matching.</p>\n<h2>Method</h2>\n<h3>1) DINOv2 then Two-stage ALIKED+LightGlue</h3>\n<h4>Stage 1: Image Retrieval with DINOv2</h4>\n<ul>\n<li>Extract global features using a DINOv2 vision transformer</li>\n<li>Perform similarity-based image pairing to reduce computational complexity</li>\n<li>Apply adaptive thresholding to balance recall against computational efficiency</li>\n</ul>\n<h4>Stage 2: Two-Stage Feature Matching</h4>\n<ul>\n<li>Primary matching: ALIKED keypoint detection + LightGlue matching</li>\n<li>Secondary matching: Crop matching with ALIKED + LightGlue</li>\n<li>Geometric verification: MAGSAC-based outlier rejection with adaptive thresholds</li>\n</ul>\n<p>The below diagram  demonstrates the DINOv2 then Two-stage ALIKED+LightGlue:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6398606%2F56f15dd4a67b30f2abac24f712611321%2FIMC25-solution.png?generation=1749334903365639&amp;alt=media\" alt=\"\"></p>\n<h3>2) DINOv2 then Two-stage ALIKED+LightGlue + OmniGlue Geometric failure trigger</h3>\n<p>OmniGlue is computationally more expensive than ALIKED+LightGlue. Thus, building upon the core pipeline above, we have developed an adaptive OmniGlue enhancement pipeline.</p>\n<p>The diagram below demonstrates the OmniGlue-enhanced pipeline:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6398606%2F9f7526a6268d18b62fd096519c38eac2%2FIMC25-solution-OmniGlue.png?generation=1749334934055549&amp;alt=media\" alt=\"\"></p>\n<h4>GPU Optimisation</h4>\n<p>The original OmniGlue implementation had significant performance bottlenecks that made it impractical for large-scale competition use, especially CPU-bound DINO feature extraction. The original implementation ran DINO features on CPU, creating a major computational bottleneck. Therefore, we have moved the entire DINO feature extraction pipeline to GPU:</p>\n<pre><code> :\n     ():\n        .device = torch.device(device  torch.cuda.is_available()  )\n        .use_mixed_precision = use_mixed_precision\n\n        \n        .model = dino.vit_base()\n        state_dict = torch.load(cpt_path, map_location=, weights_only=)\n        .model.load_state_dict(state_dict)\n        .model = .model.to(.device)\n        .model.()\n</code></pre>\n<p>By implementing this and other minor tricks, we could shorten the image pair processing time from 10 seconds to 1 - 2 seconds per image pair, which made it possible for us to use OmniGlue partially in our pipeline.</p>\n<h4>Geometric Distribution Trigger</h4>\n<p>OmniGlue matching is triggered by \"poor\" geometric distribution of ALIKED+LightGlue matches as even with GPU-optimisation of OmniGlue, it is still computationally too expensive to use OmniGlue for all the image pairs (after seeing other top solutions, it probably was possible to limit the number of image pairs to process). Because of this constraint, we also limit the trigger of OmniGlue up to 20% of the dataset.</p>\n<p>The \"poor\" geometric distribution trigger is rather a simple diagnosis combination of the initial matches distribution. Here are three criteria for triggering, and if two criteria are satisfied, the OmniGlue matching would be triggered. The following are the criteria:</p>\n<p>1) Insufficient matches: check if the number of matches is simply less than threshold=6<br>\n2) Spatial Coverage:</p>\n<pre><code>\nx_range = np.(x_coords) - np.(x_coords)\ny_range = np.(y_coords) - np.(y_coords)\n\n\nx_std = np.std(x_coords)\ny_std = np.std(y_coords)\n\n\nmin_range_threshold = (, x_std * geometric_multiplier, y_std * geometric_multiplier)\n\n\npoor_indicators.append(x_range &lt; min_range_threshold)\npoor_indicators.append(y_range &lt; min_range_threshold)\n</code></pre>\n<p>3) Grid-based clustering:</p>\n<p>We divide the image into a 4x4 grid and check the occupancy. If the occupancy is less than the threshold of 0.5, it is categorised as a poor geometrical distribution</p>\n<pre><code>\nn_grid =   \nx_bins = np.linspace(np.(x_coords), np.(x_coords), n_grid + )\ny_bins = np.linspace(np.(y_coords), np.(y_coords), n_grid + )\n\n\ngrid_counts = np.zeros((n_grid, n_grid))\n i  ((x_coords)):\n    x_idx = (n_grid - , (, np.digitize(x_coords[i], x_bins) - ))\n    y_idx = (n_grid - , (, np.digitize(y_coords[i], y_bins) - ))\n    grid_counts[x_idx, y_idx] += \n\n\nnon_empty_cells = np.(grid_counts &gt; )\ntotal_cells = n_grid * n_grid\npoor_indicators.append(non_empty_cells &lt; total_cells / )\n</code></pre>\n<p>When OmniGlue is triggered, 60 - 70% of them successfully add new matches to the original matches.</p>\n<h2>Final Submission Results</h2>\n<p>The table below demonstrates the private and public leaderboard scores of our final submission solutions, along with local cross-validation scores. Although the local CV scores suggest the OmniGlue-enhanced pipeline performs better than the Two-stage ALIKED+LightGlue matching pipeline, the leaderboard scores of the Two-stage ALIKED+LightGlue pipeline are better than the OmniGlue-enhanced pipeline.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>Private LB</th>\n<th>Public LB</th>\n<th>ETs</th>\n<th>amy_gardens</th>\n<th>fbk_vineyard</th>\n<th>stairs</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Two-stage ALIKED+LightGlue<br>(80th place)</td>\n<td>34.18</td>\n<td>34.86</td>\n<td>37.35</td>\n<td>40.16</td>\n<td>25.35</td>\n<td>6.25</td>\n</tr>\n<tr>\n<td>+OmniGlue Geometric failure<br> trigger 20%</td>\n<td>32.57</td>\n<td>32.20</td>\n<td>37.55</td>\n<td>39.51</td>\n<td>30.48</td>\n<td>9.09</td>\n</tr>\n</tbody>\n</table>\n<h2>Discussion and Analysis</h2>\n<h3>Score Robustness Analysis</h3>\n<p>We have noticed randomness in the OmniGlue-enhanced pipelines leaderboard score. After the competition ended, we have submitted the same solutions 3 times and summarised the statistics in the table below. The table below lists 3 pipelines' results - both of our final submission and a random 15% of a dataset OmniGlue matching + fully triggered OmniGlue matching for the \"stairs\" dataset.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>Two-stage<br>ALIKED+LightGlue</th>\n<th>+OmniGlue Geometric<br>failure trigger 20%</th>\n<th>+15% random OmniGlue<br>+ Full stairs dataset</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Mean Private LB</td>\n<td>33.98</td>\n<td>34.9</td>\n<td>35.94</td>\n</tr>\n<tr>\n<td>Mean Public LB</td>\n<td>34.86</td>\n<td>33.42</td>\n<td>32.46</td>\n</tr>\n<tr>\n<td>Private STD</td>\n<td>0.84</td>\n<td>2.89</td>\n<td>1.67</td>\n</tr>\n<tr>\n<td>Public STD</td>\n<td>0.46</td>\n<td>2.84</td>\n<td>1.26</td>\n</tr>\n<tr>\n<td>Private Max</td>\n<td>34.59</td>\n<td>37.07</td>\n<td>37.83</td>\n</tr>\n<tr>\n<td>Public Max</td>\n<td>34.33</td>\n<td>35.09</td>\n<td>33.43</td>\n</tr>\n<tr>\n<td>Private Min</td>\n<td>33.02</td>\n<td>31.62</td>\n<td>34.68</td>\n</tr>\n<tr>\n<td>Public Min</td>\n<td>33.44</td>\n<td>30.14</td>\n<td>31.04</td>\n</tr>\n</tbody>\n</table>\n<p>The table demonstrates that the Two-stage ALIKED+LightGlue scores are relatively stable (standard deviation of 0.46 in Public LB). On the other hand, both OmniGlue-enhanced pipelines' standard deviations are larger (2.84 and 1.26 in Public LB, respectively). This is discussed in the Score Variance Investigation section.</p>\n<h3>Geometric OmniGlue Trigger Analysis</h3>\n<p>Our geometric failure trigger was designed with the intuition that poorly distributed matches indicate potential improvement through OmniGlue. However, experimental results and some intuitions suggest this hypothesis requires refinement:</p>\n<h4>Potential Issues with Geometric Triggering</h4>\n<ul>\n<li>Object-centric datasets: In scenes like \"ETs\" with single prominent objects, matches naturally concentrate around facial/body features</li>\n<li>Scene coherence: Some image pairs may legitimately lack geometric diversity</li>\n<li>Inter-cluster confusion: Triggering on geometric failure may attempt to match fundamentally different scenes</li>\n</ul>\n<p>This is an example of Inter-cluster confusion, where we should have discarded the pair instead of triggering OmniGlue. Triggering OmniGlue for such a pair is not only useless for the reconstruction but also wastes computational resources.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6398606%2Ffc34b66d2161e847e154a98f3fbb81fc%2FOmniGlue-ETs.png?generation=1749334981140001&amp;alt=media\" alt=\"\"></p>\n<h4>Evidence for Alternative Strategies</h4>\n<ul>\n<li>Random 15% triggering achieved better average performance than geometric triggering</li>\n<li>Suggests OmniGlue's advantage may be in overall matching quality rather than \"rescue\" scenarios</li>\n</ul>\n<p>This is where OmniGlue performs better than ALIKED+LightGlue matching.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6398606%2F023a75a9d82278c792aa907801a7e695%2FOmniGlue-staris.png?generation=1749335012683604&amp;alt=media\" alt=\"\"></p>\n<h3>Score Variance Investigation</h3>\n<p>The significant variance in OmniGlue-enhanced solutions raises important questions about evaluation stability.</p>\n<p>Hypothesised Contributing Factors:</p>\n<ul>\n<li>Threshold sensitivity: Small changes in match quality affecting clustering decisions</li>\n<li>Computational precision: GPU vs CPU differences in floating-point operations</li>\n</ul>\n<h3>Ablation Study</h3>\n<p>The table below shows the incremental development of our pipeline. The reported scores are the highest scores of the same pipeline.</p>\n<table>\n<thead>\n<tr>\n<th>Pipeline</th>\n<th>Private LB</th>\n<th>Public LB</th>\n<th>Remarks</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Baseline DINOv2 ALIKED+LG</td>\n<td>30.34</td>\n<td>32.38</td>\n<td>Public notebook</td>\n</tr>\n<tr>\n<td>+Two-stage matching</td>\n<td>34.56</td>\n<td>33.45</td>\n<td></td>\n</tr>\n<tr>\n<td>+Image Rotation</td>\n<td>34.18</td>\n<td>34.86</td>\n<td>Final core pipeline (80th place)</td>\n</tr>\n<tr>\n<td>+OmniGlue (geometric trigger)</td>\n<td>37.07</td>\n<td>35.09</td>\n<td>Another submitted solution</td>\n</tr>\n<tr>\n<td>+OmniGlue for stairs + 15% random</td>\n<td>37.83</td>\n<td>33.43</td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<h3>Key takeaways</h3>\n<ul>\n<li>GPU-optimised OmniGlue implementation</li>\n<li>Grid-based match distribution: how we used it in the pipeline might not be the best method, however this way of analysis for the match distribution would be useful</li>\n</ul>",
      "rawMarkdown": "# 80th Place Solution: Two-Stage ALIKED+LightGlue Pipeline and Adaptive OmniGlue-Enhancement Pipeline\n\nWe thank the IMC 2025 organisers for hosting this challenging competition and the Kaggle community for the valuable discussions and insights. Our implementation builds upon excellent prior work in the image matching community.\n\nWe have noticed that no OmniGlue-enhanced/based solution is mentioned in the discussion. Although our solution did not perform as well as the top solutions, it might be worth sharing our solution and findings with the community.\n\n## Overview\n\nWe, @kazumax0720 and I, have prepared two pipelines for the final submission: 1) DINOv2 then Two-stage ALIKED+LightGlue and 2) 1)+budget  constraint OmniGlue matching pipeline.\n\nOur approach uses @octaviograu's [public notebook](https://www.kaggle.com/code/octaviograu/baseline-dinov2-aliked-lightglue) as a baseline and added a few components such as image rotation, second-stage matching and OmniGlue matching.\n\n## Method\n\n### 1) DINOv2 then Two-stage ALIKED+LightGlue\n\n#### Stage 1: Image Retrieval with DINOv2\n\n- Extract global features using a DINOv2 vision transformer\n- Perform similarity-based image pairing to reduce computational complexity\n- Apply adaptive thresholding to balance recall against computational efficiency\n\n#### Stage 2: Two-Stage Feature Matching\n\n- Primary matching: ALIKED keypoint detection + LightGlue matching\n- Secondary matching: Crop matching with ALIKED + LightGlue\n- Geometric verification: MAGSAC-based outlier rejection with adaptive thresholds\n\nThe below diagram  demonstrates the DINOv2 then Two-stage ALIKED+LightGlue:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6398606%2F56f15dd4a67b30f2abac24f712611321%2FIMC25-solution.png?generation=1749334903365639&alt=media)\n\n### 2) DINOv2 then Two-stage ALIKED+LightGlue + OmniGlue Geometric failure trigger\n\nOmniGlue is computationally more expensive than ALIKED+LightGlue. Thus, building upon the core pipeline above, we have developed an adaptive OmniGlue enhancement pipeline.\n\nThe diagram below demonstrates the OmniGlue-enhanced pipeline:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6398606%2F9f7526a6268d18b62fd096519c38eac2%2FIMC25-solution-OmniGlue.png?generation=1749334934055549&alt=media)\n\n#### GPU Optimisation\n\nThe original OmniGlue implementation had significant performance bottlenecks that made it impractical for large-scale competition use, especially CPU-bound DINO feature extraction. The original implementation ran DINO features on CPU, creating a major computational bottleneck. Therefore, we have moved the entire DINO feature extraction pipeline to GPU:\n\n```python\nclass OptimizedDINOExtract:\n    def __init__(self, cpt_path, device='cuda', use_mixed_precision=True):\n        self.device = torch.device(device if torch.cuda.is_available() else 'cpu')\n        self.use_mixed_precision = use_mixed_precision\n        \n        # Load DINO model directly on GPU\n        self.model = dino.vit_base()\n        state_dict = torch.load(cpt_path, map_location='cpu', weights_only=False)\n        self.model.load_state_dict(state_dict)\n        self.model = self.model.to(self.device)\n        self.model.eval()\n```\n\nBy implementing this and other minor tricks, we could shorten the image pair processing time from 10 seconds to 1 - 2 seconds per image pair, which made it possible for us to use OmniGlue partially in our pipeline.\n\n#### Geometric Distribution Trigger\n\nOmniGlue matching is triggered by \"poor\" geometric distribution of ALIKED+LightGlue matches as even with GPU-optimisation of OmniGlue, it is still computationally too expensive to use OmniGlue for all the image pairs (after seeing other top solutions, it probably was possible to limit the number of image pairs to process). Because of this constraint, we also limit the trigger of OmniGlue up to 20% of the dataset.\n\nThe \"poor\" geometric distribution trigger is rather a simple diagnosis combination of the initial matches distribution. Here are three criteria for triggering, and if two criteria are satisfied, the OmniGlue matching would be triggered. The following are the criteria:\n\n1) Insufficient matches: check if the number of matches is simply less than threshold=6\n2) Spatial Coverage:\n\n```python\n# Extract coordinate ranges\nx_range = np.max(x_coords) - np.min(x_coords)\ny_range = np.max(y_coords) - np.min(y_coords)\n\n# Calculate statistical spread\nx_std = np.std(x_coords)\ny_std = np.std(y_coords)\n\n# Adaptive threshold based on data distribution\nmin_range_threshold = max(100, x_std * geometric_multiplier, y_std * geometric_multiplier)\n\n# Evaluate coverage\npoor_indicators.append(x_range < min_range_threshold)\npoor_indicators.append(y_range < min_range_threshold)\n```\n\n3) Grid-based clustering:\n\nWe divide the image into a 4x4 grid and check the occupancy. If the occupancy is less than the threshold of 0.5, it is categorised as a poor geometrical distribution\n\n```python\n# Create a spatial grid\nn_grid = 4  # 4×4 grid = 16 cells\nx_bins = np.linspace(np.min(x_coords), np.max(x_coords), n_grid + 1)\ny_bins = np.linspace(np.min(y_coords), np.max(y_coords), n_grid + 1)\n\n# Count matches per grid cell\ngrid_counts = np.zeros((n_grid, n_grid))\nfor i in range(len(x_coords)):\n    x_idx = min(n_grid - 1, max(0, np.digitize(x_coords[i], x_bins) - 1))\n    y_idx = min(n_grid - 1, max(0, np.digitize(y_coords[i], y_bins) - 1))\n    grid_counts[x_idx, y_idx] += 1\n\n# Evaluate spatial uniformity\nnon_empty_cells = np.sum(grid_counts > 0)\ntotal_cells = n_grid * n_grid\npoor_indicators.append(non_empty_cells < total_cells / 2)\n```\n\nWhen OmniGlue is triggered, 60 - 70% of them successfully add new matches to the original matches.\n\n## Final Submission Results\n\nThe table below demonstrates the private and public leaderboard scores of our final submission solutions, along with local cross-validation scores. Although the local CV scores suggest the OmniGlue-enhanced pipeline performs better than the Two-stage ALIKED+LightGlue matching pipeline, the leaderboard scores of the Two-stage ALIKED+LightGlue pipeline are better than the OmniGlue-enhanced pipeline.\n\n|                                             \t| Private LB \t| Public LB \t| ETs   \t| amy_gardens \t| fbk_vineyard \t| stairs \t|\n|---------------------------------------------\t|------------\t|-----------\t|-------\t|-------------\t|--------------\t|--------\t|\n| Two-stage ALIKED+LightGlue<br>(80th place)  \t| 34.18      \t| 34.86     \t| 37.35 \t| 40.16       \t| 25.35        \t| 6.25   \t|\n| +OmniGlue Geometric failure<br> trigger 20% \t| 32.57      \t| 32.20     \t| 37.55 \t| 39.51       \t| 30.48        \t| 9.09   \t|\n\n\n## Discussion and Analysis\n\n### Score Robustness Analysis\n\nWe have noticed randomness in the OmniGlue-enhanced pipelines leaderboard score. After the competition ended, we have submitted the same solutions 3 times and summarised the statistics in the table below. The table below lists 3 pipelines' results - both of our final submission and a random 15% of a dataset OmniGlue matching + fully triggered OmniGlue matching for the \"stairs\" dataset.\n\n|                 \t| Two-stage<br>ALIKED+LightGlue \t| +OmniGlue Geometric<br>failure trigger 20% \t| +15% random OmniGlue<br>+ Full stairs dataset \t|\n|-----------------\t|:-----------------------------:\t|:------------------------------------------:\t|:---------------------------------------------:\t|\n| Mean Private LB \t|             33.98             \t|                    34.9                    \t|                     35.94                     \t|\n| Mean Public LB  \t|             34.86             \t|                    33.42                   \t|                     32.46                     \t|\n| Private STD     \t|              0.84             \t|                    2.89                    \t|                      1.67                     \t|\n| Public STD      \t|              0.46             \t|                    2.84                    \t|                      1.26                     \t|\n| Private Max     \t|             34.59             \t|                    37.07                   \t|                     37.83                     \t|\n| Public Max      \t|             34.33             \t|                    35.09                   \t|                     33.43                     \t|\n| Private Min     \t|             33.02             \t|                    31.62                   \t|                     34.68                     \t|\n| Public Min      \t|             33.44             \t|                    30.14                   \t|                     31.04                     \t|\n\n\nThe table demonstrates that the Two-stage ALIKED+LightGlue scores are relatively stable (standard deviation of 0.46 in Public LB). On the other hand, both OmniGlue-enhanced pipelines' standard deviations are larger (2.84 and 1.26 in Public LB, respectively). This is discussed in the Score Variance Investigation section.\n\n### Geometric OmniGlue Trigger Analysis\n\nOur geometric failure trigger was designed with the intuition that poorly distributed matches indicate potential improvement through OmniGlue. However, experimental results and some intuitions suggest this hypothesis requires refinement:\n\n#### Potential Issues with Geometric Triggering\n\n- Object-centric datasets: In scenes like \"ETs\" with single prominent objects, matches naturally concentrate around facial/body features\n- Scene coherence: Some image pairs may legitimately lack geometric diversity\n- Inter-cluster confusion: Triggering on geometric failure may attempt to match fundamentally different scenes\n\nThis is an example of Inter-cluster confusion, where we should have discarded the pair instead of triggering OmniGlue. Triggering OmniGlue for such a pair is not only useless for the reconstruction but also wastes computational resources.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6398606%2Ffc34b66d2161e847e154a98f3fbb81fc%2FOmniGlue-ETs.png?generation=1749334981140001&alt=media)\n\n#### Evidence for Alternative Strategies\n\n- Random 15% triggering achieved better average performance than geometric triggering\n- Suggests OmniGlue's advantage may be in overall matching quality rather than \"rescue\" scenarios\n\nThis is where OmniGlue performs better than ALIKED+LightGlue matching.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6398606%2F023a75a9d82278c792aa907801a7e695%2FOmniGlue-staris.png?generation=1749335012683604&alt=media)\n\n### Score Variance Investigation\n\nThe significant variance in OmniGlue-enhanced solutions raises important questions about evaluation stability.\n\nHypothesised Contributing Factors:\n\n- Threshold sensitivity: Small changes in match quality affecting clustering decisions\n- Computational precision: GPU vs CPU differences in floating-point operations\n\n### Ablation Study\n\nThe table below shows the incremental development of our pipeline. The reported scores are the highest scores of the same pipeline.\n\n| Pipeline                          \t| Private LB \t| Public LB \t| Remarks                          \t|\n|-----------------------------------\t|------------\t|-----------\t|----------------------------------\t|\n| Baseline DINOv2 ALIKED+LG         \t|    30.34   \t|   32.38   \t| Public notebook                  \t|\n| +Two-stage matching               \t|    34.56   \t|   33.45   \t|                                  \t|\n| +Image Rotation                   \t|    34.18   \t|   34.86   \t| Final core pipeline (80th place) \t|\n| +OmniGlue (geometric trigger)     \t|    37.07   \t|   35.09   \t| Another submitted solution       \t|\n| +OmniGlue for stairs + 15% random \t|    37.83   \t|   33.43   \t|                                  \t|\n\n### Key takeaways\n\n- GPU-optimised OmniGlue implementation\n- Grid-based match distribution: how we used it in the pipeline might not be the best method, however this way of analysis for the match distribution would be useful",
      "votes": 9
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3219542": "# 80th Place Solution: Two-Stage ALIKED+LightGlue Pipeline and Adaptive OmniGlue-Enhancement Pipeline\n\nWe thank the IMC 2025 organisers for hosting this challenging competition and the Kaggle community for the valuable discussions and insights. Our implementation builds upon excellent prior work in the image matching community.\n\nWe have noticed that no OmniGlue-enhanced/based solution is mentioned in the discussion. Although our solution did not perform as well as the top solutions, it might be worth sharing our solution and findings with the community.\n\n## Overview\n\nWe, @kazumax0720 and I, have prepared two pipelines for the final submission: 1) DINOv2 then Two-stage ALIKED+LightGlue and 2) 1)+budget  constraint OmniGlue matching pipeline.\n\nOur approach uses @octaviograu's [public notebook](https://www.kaggle.com/code/octaviograu/baseline-dinov2-aliked-lightglue) as a baseline and added a few components such as image rotation, second-stage matching and OmniGlue matching.\n\n## Method\n\n### 1) DINOv2 then Two-stage ALIKED+LightGlue\n\n#### Stage 1: Image Retrieval with DINOv2\n\n- Extract global features using a DINOv2 vision transformer\n- Perform similarity-based image pairing to reduce computational complexity\n- Apply adaptive thresholding to balance recall against computational efficiency\n\n#### Stage 2: Two-Stage Feature Matching\n\n- Primary matching: ALIKED keypoint detection + LightGlue matching\n- Secondary matching: Crop matching with ALIKED + LightGlue\n- Geometric verification: MAGSAC-based outlier rejection with adaptive thresholds\n\nThe below diagram  demonstrates the DINOv2 then Two-stage ALIKED+LightGlue:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6398606%2F56f15dd4a67b30f2abac24f712611321%2FIMC25-solution.png?generation=1749334903365639&alt=media)\n\n### 2) DINOv2 then Two-stage ALIKED+LightGlue + OmniGlue Geometric failure trigger\n\nOmniGlue is computationally more expensive than ALIKED+LightGlue. Thus, building upon the core pipeline above, we have developed an adaptive OmniGlue enhancement pipeline.\n\nThe diagram below demonstrates the OmniGlue-enhanced pipeline:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6398606%2F9f7526a6268d18b62fd096519c38eac2%2FIMC25-solution-OmniGlue.png?generation=1749334934055549&alt=media)\n\n#### GPU Optimisation\n\nThe original OmniGlue implementation had significant performance bottlenecks that made it impractical for large-scale competition use, especially CPU-bound DINO feature extraction. The original implementation ran DINO features on CPU, creating a major computational bottleneck. Therefore, we have moved the entire DINO feature extraction pipeline to GPU:\n\n```python\nclass OptimizedDINOExtract:\n    def __init__(self, cpt_path, device='cuda', use_mixed_precision=True):\n        self.device = torch.device(device if torch.cuda.is_available() else 'cpu')\n        self.use_mixed_precision = use_mixed_precision\n        \n        # Load DINO model directly on GPU\n        self.model = dino.vit_base()\n        state_dict = torch.load(cpt_path, map_location='cpu', weights_only=False)\n        self.model.load_state_dict(state_dict)\n        self.model = self.model.to(self.device)\n        self.model.eval()\n```\n\nBy implementing this and other minor tricks, we could shorten the image pair processing time from 10 seconds to 1 - 2 seconds per image pair, which made it possible for us to use OmniGlue partially in our pipeline.\n\n#### Geometric Distribution Trigger\n\nOmniGlue matching is triggered by \"poor\" geometric distribution of ALIKED+LightGlue matches as even with GPU-optimisation of OmniGlue, it is still computationally too expensive to use OmniGlue for all the image pairs (after seeing other top solutions, it probably was possible to limit the number of image pairs to process). Because of this constraint, we also limit the trigger of OmniGlue up to 20% of the dataset.\n\nThe \"poor\" geometric distribution trigger is rather a simple diagnosis combination of the initial matches distribution. Here are three criteria for triggering, and if two criteria are satisfied, the OmniGlue matching would be triggered. The following are the criteria:\n\n1) Insufficient matches: check if the number of matches is simply less than threshold=6\n2) Spatial Coverage:\n\n```python\n# Extract coordinate ranges\nx_range = np.max(x_coords) - np.min(x_coords)\ny_range = np.max(y_coords) - np.min(y_coords)\n\n# Calculate statistical spread\nx_std = np.std(x_coords)\ny_std = np.std(y_coords)\n\n# Adaptive threshold based on data distribution\nmin_range_threshold = max(100, x_std * geometric_multiplier, y_std * geometric_multiplier)\n\n# Evaluate coverage\npoor_indicators.append(x_range < min_range_threshold)\npoor_indicators.append(y_range < min_range_threshold)\n```\n\n3) Grid-based clustering:\n\nWe divide the image into a 4x4 grid and check the occupancy. If the occupancy is less than the threshold of 0.5, it is categorised as a poor geometrical distribution\n\n```python\n# Create a spatial grid\nn_grid = 4  # 4×4 grid = 16 cells\nx_bins = np.linspace(np.min(x_coords), np.max(x_coords), n_grid + 1)\ny_bins = np.linspace(np.min(y_coords), np.max(y_coords), n_grid + 1)\n\n# Count matches per grid cell\ngrid_counts = np.zeros((n_grid, n_grid))\nfor i in range(len(x_coords)):\n    x_idx = min(n_grid - 1, max(0, np.digitize(x_coords[i], x_bins) - 1))\n    y_idx = min(n_grid - 1, max(0, np.digitize(y_coords[i], y_bins) - 1))\n    grid_counts[x_idx, y_idx] += 1\n\n# Evaluate spatial uniformity\nnon_empty_cells = np.sum(grid_counts > 0)\ntotal_cells = n_grid * n_grid\npoor_indicators.append(non_empty_cells < total_cells / 2)\n```\n\nWhen OmniGlue is triggered, 60 - 70% of them successfully add new matches to the original matches.\n\n## Final Submission Results\n\nThe table below demonstrates the private and public leaderboard scores of our final submission solutions, along with local cross-validation scores. Although the local CV scores suggest the OmniGlue-enhanced pipeline performs better than the Two-stage ALIKED+LightGlue matching pipeline, the leaderboard scores of the Two-stage ALIKED+LightGlue pipeline are better than the OmniGlue-enhanced pipeline.\n\n|                                             \t| Private LB \t| Public LB \t| ETs   \t| amy_gardens \t| fbk_vineyard \t| stairs \t|\n|---------------------------------------------\t|------------\t|-----------\t|-------\t|-------------\t|--------------\t|--------\t|\n| Two-stage ALIKED+LightGlue<br>(80th place)  \t| 34.18      \t| 34.86     \t| 37.35 \t| 40.16       \t| 25.35        \t| 6.25   \t|\n| +OmniGlue Geometric failure<br> trigger 20% \t| 32.57      \t| 32.20     \t| 37.55 \t| 39.51       \t| 30.48        \t| 9.09   \t|\n\n\n## Discussion and Analysis\n\n### Score Robustness Analysis\n\nWe have noticed randomness in the OmniGlue-enhanced pipelines leaderboard score. After the competition ended, we have submitted the same solutions 3 times and summarised the statistics in the table below. The table below lists 3 pipelines' results - both of our final submission and a random 15% of a dataset OmniGlue matching + fully triggered OmniGlue matching for the \"stairs\" dataset.\n\n|                 \t| Two-stage<br>ALIKED+LightGlue \t| +OmniGlue Geometric<br>failure trigger 20% \t| +15% random OmniGlue<br>+ Full stairs dataset \t|\n|-----------------\t|:-----------------------------:\t|:------------------------------------------:\t|:---------------------------------------------:\t|\n| Mean Private LB \t|             33.98             \t|                    34.9                    \t|                     35.94                     \t|\n| Mean Public LB  \t|             34.86             \t|                    33.42                   \t|                     32.46                     \t|\n| Private STD     \t|              0.84             \t|                    2.89                    \t|                      1.67                     \t|\n| Public STD      \t|              0.46             \t|                    2.84                    \t|                      1.26                     \t|\n| Private Max     \t|             34.59             \t|                    37.07                   \t|                     37.83                     \t|\n| Public Max      \t|             34.33             \t|                    35.09                   \t|                     33.43                     \t|\n| Private Min     \t|             33.02             \t|                    31.62                   \t|                     34.68                     \t|\n| Public Min      \t|             33.44             \t|                    30.14                   \t|                     31.04                     \t|\n\n\nThe table demonstrates that the Two-stage ALIKED+LightGlue scores are relatively stable (standard deviation of 0.46 in Public LB). On the other hand, both OmniGlue-enhanced pipelines' standard deviations are larger (2.84 and 1.26 in Public LB, respectively). This is discussed in the Score Variance Investigation section.\n\n### Geometric OmniGlue Trigger Analysis\n\nOur geometric failure trigger was designed with the intuition that poorly distributed matches indicate potential improvement through OmniGlue. However, experimental results and some intuitions suggest this hypothesis requires refinement:\n\n#### Potential Issues with Geometric Triggering\n\n- Object-centric datasets: In scenes like \"ETs\" with single prominent objects, matches naturally concentrate around facial/body features\n- Scene coherence: Some image pairs may legitimately lack geometric diversity\n- Inter-cluster confusion: Triggering on geometric failure may attempt to match fundamentally different scenes\n\nThis is an example of Inter-cluster confusion, where we should have discarded the pair instead of triggering OmniGlue. Triggering OmniGlue for such a pair is not only useless for the reconstruction but also wastes computational resources.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6398606%2Ffc34b66d2161e847e154a98f3fbb81fc%2FOmniGlue-ETs.png?generation=1749334981140001&alt=media)\n\n#### Evidence for Alternative Strategies\n\n- Random 15% triggering achieved better average performance than geometric triggering\n- Suggests OmniGlue's advantage may be in overall matching quality rather than \"rescue\" scenarios\n\nThis is where OmniGlue performs better than ALIKED+LightGlue matching.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6398606%2F023a75a9d82278c792aa907801a7e695%2FOmniGlue-staris.png?generation=1749335012683604&alt=media)\n\n### Score Variance Investigation\n\nThe significant variance in OmniGlue-enhanced solutions raises important questions about evaluation stability.\n\nHypothesised Contributing Factors:\n\n- Threshold sensitivity: Small changes in match quality affecting clustering decisions\n- Computational precision: GPU vs CPU differences in floating-point operations\n\n### Ablation Study\n\nThe table below shows the incremental development of our pipeline. The reported scores are the highest scores of the same pipeline.\n\n| Pipeline                          \t| Private LB \t| Public LB \t| Remarks                          \t|\n|-----------------------------------\t|------------\t|-----------\t|----------------------------------\t|\n| Baseline DINOv2 ALIKED+LG         \t|    30.34   \t|   32.38   \t| Public notebook                  \t|\n| +Two-stage matching               \t|    34.56   \t|   33.45   \t|                                  \t|\n| +Image Rotation                   \t|    34.18   \t|   34.86   \t| Final core pipeline (80th place) \t|\n| +OmniGlue (geometric trigger)     \t|    37.07   \t|   35.09   \t| Another submitted solution       \t|\n| +OmniGlue for stairs + 15% random \t|    37.83   \t|   33.43   \t|                                  \t|\n\n### Key takeaways\n\n- GPU-optimised OmniGlue implementation\n- Grid-based match distribution: how we used it in the pipeline might not be the best method, however this way of analysis for the match distribution would be useful"
  }
}