{
  "id": 582968,
  "title": "Memory-Efficient VGGT Tracking with Pose Refinement",
  "url": "/competitions/image-matching-challenge-2025/discussion/582968",
  "author_name": "tmyok",
  "post_date": "2025-06-04T00:10:14.303000",
  "votes": 28,
  "comment_count": 0,
  "views": 0,
  "content": "<p><strong>Update:</strong> Code is now available<br>\n<a href=\"https://www.kaggle.com/code/tmyok1984/imc2025-exp66-vggt-with-refiner?scriptVersionId=243529440\" target=\"_blank\">Notebook</a>  <br>\n<a href=\"https://www.kaggle.com/datasets/tmyok1984/imc2025-exp66-vggt-with-incremental-model-refiner\" target=\"_blank\">Scripts</a></p>\n<hr>\n<p>We introduced our team's solution&nbsp;<a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2025/discussion/582898\" target=\"_blank\">here</a>. In this section, I will focus on the VGGT-based approach that I worked on.</p>\n<p><strong>E2E 3D Geometric Foundation Models</strong>&nbsp;have been attracting attention as potential game-changers in the field of image matching. Among them, I chose to apply&nbsp;<strong>VGGT</strong>, developed by <a href=\"https://www.kaggle.com/jianyuanv\" target=\"_blank\">@jianyuanv</a>, to this IMC challenge. However, applying VGGT to IMC2025 posed three key challenges:</p>\n<ul>\n<li>It does not assume inputs containing multiple unrelated scenes</li>\n<li>It requires large amounts of VRAM</li>\n<li>Its camera pose estimation accuracy is relatively low</li>\n</ul>\n<p>To address these issues, we implemented the following improvements:</p>\n<ul>\n<li>Robust image matching based on global features using the&nbsp;<strong>VGGT tracker</strong></li>\n<li>Chunk-wise processing of VGGT tracker to reduce memory usage</li>\n<li>Camera pose refinement using high-precision correspondences from&nbsp;<strong>ALIKED + LightGlue</strong></li>\n</ul>\n<h2>Overview</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2F5c181b25c0664a66e0d351bfac0c6105%2Foverview.svg?generation=1748995502557044&amp;alt=media\" alt=\"\"></p>\n<p><strong>Note</strong>: To meet the 9-hour time constraint imposed by Code Requirements, the following additional optimizations were made:</p>\n<ul>\n<li>Pre-filtering of candidate image pairs using&nbsp;<strong>keynet-adalam</strong>&nbsp;<a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2025/discussion/582898\" target=\"_blank\">details</a></li>\n<li>VGGT processing was limited to dataset with ≤128 images</li>\n</ul>\n<h2>VGGT Tracker</h2>\n<p>VGGT provides not only camera pose estimation, but also a&nbsp;<strong>tracking</strong>&nbsp;function. Through experiments, we observed that it captures global context more effectively than conventional keypoint-based methods.</p>\n<p>For example, in the&nbsp;fbk_<em>vineyard</em>&nbsp;dataset, the correspondences stopped appropriately at scene transition boundaries.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2F012feab94341ecf48e0eaedf334048e9%2FVGGT-tracker.png?generation=1748995541504832&amp;alt=media\" alt=\"\"></p>\n<h2>Memory Reduction via Chunked Processing</h2>\n<p>In the original implementation, all target images were input at once for a single query image, resulting in high memory usage. To address this, I split the target images into 1/N chunks and processed them sequentially with the query image. This significantly reduced VRAM consumption.</p>\n<h2>Camera Pose Refinement</h2>\n<p>While the matching from the VGGT tracker captured the overall scene structure, its absolute accuracy was limited. This is partly due to VGGT resizing images to 518, which leads to reduced spatial precision in keypoint localization.</p>\n<p>To overcome this, I refined the coarse camera poses estimated by VGGT using high-precision correspondences obtained from&nbsp;ALIKED + LightGlue. The refinement was done using the&nbsp;<code>incremental_model_refiner</code>, which was also used in the 1st place solution of IMC2023.</p>\n<p><a href=\"https://github.com/zju3dv/DetectorFreeSfM/blob/main/src/sfm_runner/sfm_model_geometry_refiner.py\" target=\"_blank\">https://github.com/zju3dv/DetectorFreeSfM/blob/main/src/sfm_runner/sfm_model_geometry_refiner.py</a></p>\n<h2>Experiments</h2>\n<p>We evaluated the effectiveness of our method using validation data provided by the organizers.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2F2bb53df3cdff4aae5da03ecd08118c69%2Fresults.png?generation=1748995584825997&amp;alt=media\" alt=\"\"></p>\n<table>\n<thead>\n<tr>\n<th>Method</th>\n<th>Public LB</th>\n<th>Private LB</th>\n<th>Val (amy_gardens)</th>\n<th>Val (fbk_vineyard)</th>\n<th>Val (ETs)</th>\n<th>Val (stairs)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>LightGlue with refiner</td>\n<td>39.91</td>\n<td>42.33</td>\n<td>27.19</td>\n<td>45.91</td>\n<td>61.33</td>\n<td>6.25</td>\n</tr>\n<tr>\n<td>VGGT with refiner</td>\n<td>38.55</td>\n<td>41.02</td>\n<td>44.88</td>\n<td>63.26</td>\n<td>66.67</td>\n<td>11.77</td>\n</tr>\n<tr>\n<td>VGGT w/o refiner</td>\n<td>31.04</td>\n<td>36.65</td>\n<td>39.34</td>\n<td>66.23</td>\n<td>17.54</td>\n<td>14.29</td>\n</tr>\n</tbody>\n</table>\n<p>Below are some qualitative results on selected scenes:</p>\n<p><strong>amy_gardens</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2Fc47def199d10a5760c9137fafea5d1af%2Famy_gardens.png?generation=1748995608568829&amp;alt=media\" alt=\"\"></p>\n<p><strong>fbk_vineyard</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2F6cf484b03ac0275c9fce24190cdd5e1c%2Fvineyard1.png?generation=1748995649687449&amp;alt=media\" alt=\"\"></p>\n<p><strong>ETs</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2Fc21f6edecaf0ded471376ff0f816d827%2FET1.png?generation=1748995665949055&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2F29d19e09d6db3ef6e622e356d15c981b%2FET2.png?generation=1748995676428148&amp;alt=media\" alt=\"\"></p>\n<p>Although the proposed method performed well on the validation data, it underperformed on the leaderboard compared to conventional LightGlue-based approaches. As a result, the final submission was based on the baseline LightGlue-based method.</p>",
  "messages": [
    {
      "id": 3216675,
      "postDate": "2025-06-04T00:10:14.303Z",
      "content": "<p><strong>Update:</strong> Code is now available<br>\n<a href=\"https://www.kaggle.com/code/tmyok1984/imc2025-exp66-vggt-with-refiner?scriptVersionId=243529440\" target=\"_blank\">Notebook</a>  <br>\n<a href=\"https://www.kaggle.com/datasets/tmyok1984/imc2025-exp66-vggt-with-incremental-model-refiner\" target=\"_blank\">Scripts</a></p>\n<hr>\n<p>We introduced our team's solution&nbsp;<a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2025/discussion/582898\" target=\"_blank\">here</a>. In this section, I will focus on the VGGT-based approach that I worked on.</p>\n<p><strong>E2E 3D Geometric Foundation Models</strong>&nbsp;have been attracting attention as potential game-changers in the field of image matching. Among them, I chose to apply&nbsp;<strong>VGGT</strong>, developed by <a href=\"https://www.kaggle.com/jianyuanv\" target=\"_blank\">@jianyuanv</a>, to this IMC challenge. However, applying VGGT to IMC2025 posed three key challenges:</p>\n<ul>\n<li>It does not assume inputs containing multiple unrelated scenes</li>\n<li>It requires large amounts of VRAM</li>\n<li>Its camera pose estimation accuracy is relatively low</li>\n</ul>\n<p>To address these issues, we implemented the following improvements:</p>\n<ul>\n<li>Robust image matching based on global features using the&nbsp;<strong>VGGT tracker</strong></li>\n<li>Chunk-wise processing of VGGT tracker to reduce memory usage</li>\n<li>Camera pose refinement using high-precision correspondences from&nbsp;<strong>ALIKED + LightGlue</strong></li>\n</ul>\n<h2>Overview</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2F5c181b25c0664a66e0d351bfac0c6105%2Foverview.svg?generation=1748995502557044&amp;alt=media\" alt=\"\"></p>\n<p><strong>Note</strong>: To meet the 9-hour time constraint imposed by Code Requirements, the following additional optimizations were made:</p>\n<ul>\n<li>Pre-filtering of candidate image pairs using&nbsp;<strong>keynet-adalam</strong>&nbsp;<a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2025/discussion/582898\" target=\"_blank\">details</a></li>\n<li>VGGT processing was limited to dataset with ≤128 images</li>\n</ul>\n<h2>VGGT Tracker</h2>\n<p>VGGT provides not only camera pose estimation, but also a&nbsp;<strong>tracking</strong>&nbsp;function. Through experiments, we observed that it captures global context more effectively than conventional keypoint-based methods.</p>\n<p>For example, in the&nbsp;fbk_<em>vineyard</em>&nbsp;dataset, the correspondences stopped appropriately at scene transition boundaries.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2F012feab94341ecf48e0eaedf334048e9%2FVGGT-tracker.png?generation=1748995541504832&amp;alt=media\" alt=\"\"></p>\n<h2>Memory Reduction via Chunked Processing</h2>\n<p>In the original implementation, all target images were input at once for a single query image, resulting in high memory usage. To address this, I split the target images into 1/N chunks and processed them sequentially with the query image. This significantly reduced VRAM consumption.</p>\n<h2>Camera Pose Refinement</h2>\n<p>While the matching from the VGGT tracker captured the overall scene structure, its absolute accuracy was limited. This is partly due to VGGT resizing images to 518, which leads to reduced spatial precision in keypoint localization.</p>\n<p>To overcome this, I refined the coarse camera poses estimated by VGGT using high-precision correspondences obtained from&nbsp;ALIKED + LightGlue. The refinement was done using the&nbsp;<code>incremental_model_refiner</code>, which was also used in the 1st place solution of IMC2023.</p>\n<p><a href=\"https://github.com/zju3dv/DetectorFreeSfM/blob/main/src/sfm_runner/sfm_model_geometry_refiner.py\" target=\"_blank\">https://github.com/zju3dv/DetectorFreeSfM/blob/main/src/sfm_runner/sfm_model_geometry_refiner.py</a></p>\n<h2>Experiments</h2>\n<p>We evaluated the effectiveness of our method using validation data provided by the organizers.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2F2bb53df3cdff4aae5da03ecd08118c69%2Fresults.png?generation=1748995584825997&amp;alt=media\" alt=\"\"></p>\n<table>\n<thead>\n<tr>\n<th>Method</th>\n<th>Public LB</th>\n<th>Private LB</th>\n<th>Val (amy_gardens)</th>\n<th>Val (fbk_vineyard)</th>\n<th>Val (ETs)</th>\n<th>Val (stairs)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>LightGlue with refiner</td>\n<td>39.91</td>\n<td>42.33</td>\n<td>27.19</td>\n<td>45.91</td>\n<td>61.33</td>\n<td>6.25</td>\n</tr>\n<tr>\n<td>VGGT with refiner</td>\n<td>38.55</td>\n<td>41.02</td>\n<td>44.88</td>\n<td>63.26</td>\n<td>66.67</td>\n<td>11.77</td>\n</tr>\n<tr>\n<td>VGGT w/o refiner</td>\n<td>31.04</td>\n<td>36.65</td>\n<td>39.34</td>\n<td>66.23</td>\n<td>17.54</td>\n<td>14.29</td>\n</tr>\n</tbody>\n</table>\n<p>Below are some qualitative results on selected scenes:</p>\n<p><strong>amy_gardens</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2Fc47def199d10a5760c9137fafea5d1af%2Famy_gardens.png?generation=1748995608568829&amp;alt=media\" alt=\"\"></p>\n<p><strong>fbk_vineyard</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2F6cf484b03ac0275c9fce24190cdd5e1c%2Fvineyard1.png?generation=1748995649687449&amp;alt=media\" alt=\"\"></p>\n<p><strong>ETs</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2Fc21f6edecaf0ded471376ff0f816d827%2FET1.png?generation=1748995665949055&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2F29d19e09d6db3ef6e622e356d15c981b%2FET2.png?generation=1748995676428148&amp;alt=media\" alt=\"\"></p>\n<p>Although the proposed method performed well on the validation data, it underperformed on the leaderboard compared to conventional LightGlue-based approaches. As a result, the final submission was based on the baseline LightGlue-based method.</p>",
      "rawMarkdown": "**Update:** Code is now available\n[Notebook](https://www.kaggle.com/code/tmyok1984/imc2025-exp66-vggt-with-refiner?scriptVersionId=243529440)  \n[Scripts](https://www.kaggle.com/datasets/tmyok1984/imc2025-exp66-vggt-with-incremental-model-refiner)\n\n----\n\nWe introduced our team's solution [here](https://www.kaggle.com/competitions/image-matching-challenge-2025/discussion/582898). In this section, I will focus on the VGGT-based approach that I worked on.\n\n**E2E 3D Geometric Foundation Models** have been attracting attention as potential game-changers in the field of image matching. Among them, I chose to apply **VGGT**, developed by @jianyuanv, to this IMC challenge. However, applying VGGT to IMC2025 posed three key challenges:\n\n- It does not assume inputs containing multiple unrelated scenes\n- It requires large amounts of VRAM\n- Its camera pose estimation accuracy is relatively low\n\nTo address these issues, we implemented the following improvements:\n\n- Robust image matching based on global features using the **VGGT tracker**\n- Chunk-wise processing of VGGT tracker to reduce memory usage\n- Camera pose refinement using high-precision correspondences from **ALIKED + LightGlue**\n\n## Overview\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2F5c181b25c0664a66e0d351bfac0c6105%2Foverview.svg?generation=1748995502557044&alt=media)\n\n**Note**: To meet the 9-hour time constraint imposed by Code Requirements, the following additional optimizations were made:\n\n- Pre-filtering of candidate image pairs using **keynet-adalam** [details](https://www.kaggle.com/competitions/image-matching-challenge-2025/discussion/582898)\n- VGGT processing was limited to dataset with ≤128 images\n\n## VGGT Tracker\n\nVGGT provides not only camera pose estimation, but also a **tracking** function. Through experiments, we observed that it captures global context more effectively than conventional keypoint-based methods.\n\nFor example, in the fbk_*vineyard* dataset, the correspondences stopped appropriately at scene transition boundaries.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2F012feab94341ecf48e0eaedf334048e9%2FVGGT-tracker.png?generation=1748995541504832&alt=media)\n\n## Memory Reduction via Chunked Processing\n\nIn the original implementation, all target images were input at once for a single query image, resulting in high memory usage. To address this, I split the target images into 1/N chunks and processed them sequentially with the query image. This significantly reduced VRAM consumption.\n\n## Camera Pose Refinement\n\nWhile the matching from the VGGT tracker captured the overall scene structure, its absolute accuracy was limited. This is partly due to VGGT resizing images to 518, which leads to reduced spatial precision in keypoint localization.\n\nTo overcome this, I refined the coarse camera poses estimated by VGGT using high-precision correspondences obtained from ALIKED + LightGlue. The refinement was done using the `incremental_model_refiner`, which was also used in the 1st place solution of IMC2023.\n\nhttps://github.com/zju3dv/DetectorFreeSfM/blob/main/src/sfm_runner/sfm_model_geometry_refiner.py\n\n## Experiments\n\nWe evaluated the effectiveness of our method using validation data provided by the organizers.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2F2bb53df3cdff4aae5da03ecd08118c69%2Fresults.png?generation=1748995584825997&alt=media)\n\n| Method                 | Public LB | Private LB | Val (amy_gardens) | Val (fbk_vineyard) | Val (ETs) | Val (stairs) |\n|------------------------|-----------|------------|--------------------|---------------------|-----------|---------------|\n| LightGlue with refiner | 39.91     | 42.33      | 27.19              | 45.91               | 61.33     | 6.25          |\n| VGGT with refiner      | 38.55     | 41.02      | 44.88              | 63.26               | 66.67     | 11.77         |\n| VGGT w/o refiner       | 31.04     | 36.65      | 39.34              | 66.23               | 17.54     | 14.29         |\n\n\nBelow are some qualitative results on selected scenes:\n\n**amy_gardens**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2Fc47def199d10a5760c9137fafea5d1af%2Famy_gardens.png?generation=1748995608568829&alt=media)\n\n**fbk_vineyard**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2F6cf484b03ac0275c9fce24190cdd5e1c%2Fvineyard1.png?generation=1748995649687449&alt=media)\n\n**ETs**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2Fc21f6edecaf0ded471376ff0f816d827%2FET1.png?generation=1748995665949055&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2F29d19e09d6db3ef6e622e356d15c981b%2FET2.png?generation=1748995676428148&alt=media)\n\nAlthough the proposed method performed well on the validation data, it underperformed on the leaderboard compared to conventional LightGlue-based approaches. As a result, the final submission was based on the baseline LightGlue-based method.",
      "votes": 28
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3216675": "**Update:** Code is now available\n[Notebook](https://www.kaggle.com/code/tmyok1984/imc2025-exp66-vggt-with-refiner?scriptVersionId=243529440)  \n[Scripts](https://www.kaggle.com/datasets/tmyok1984/imc2025-exp66-vggt-with-incremental-model-refiner)\n\n----\n\nWe introduced our team's solution [here](https://www.kaggle.com/competitions/image-matching-challenge-2025/discussion/582898). In this section, I will focus on the VGGT-based approach that I worked on.\n\n**E2E 3D Geometric Foundation Models** have been attracting attention as potential game-changers in the field of image matching. Among them, I chose to apply **VGGT**, developed by @jianyuanv, to this IMC challenge. However, applying VGGT to IMC2025 posed three key challenges:\n\n- It does not assume inputs containing multiple unrelated scenes\n- It requires large amounts of VRAM\n- Its camera pose estimation accuracy is relatively low\n\nTo address these issues, we implemented the following improvements:\n\n- Robust image matching based on global features using the **VGGT tracker**\n- Chunk-wise processing of VGGT tracker to reduce memory usage\n- Camera pose refinement using high-precision correspondences from **ALIKED + LightGlue**\n\n## Overview\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2F5c181b25c0664a66e0d351bfac0c6105%2Foverview.svg?generation=1748995502557044&alt=media)\n\n**Note**: To meet the 9-hour time constraint imposed by Code Requirements, the following additional optimizations were made:\n\n- Pre-filtering of candidate image pairs using **keynet-adalam** [details](https://www.kaggle.com/competitions/image-matching-challenge-2025/discussion/582898)\n- VGGT processing was limited to dataset with ≤128 images\n\n## VGGT Tracker\n\nVGGT provides not only camera pose estimation, but also a **tracking** function. Through experiments, we observed that it captures global context more effectively than conventional keypoint-based methods.\n\nFor example, in the fbk_*vineyard* dataset, the correspondences stopped appropriately at scene transition boundaries.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2F012feab94341ecf48e0eaedf334048e9%2FVGGT-tracker.png?generation=1748995541504832&alt=media)\n\n## Memory Reduction via Chunked Processing\n\nIn the original implementation, all target images were input at once for a single query image, resulting in high memory usage. To address this, I split the target images into 1/N chunks and processed them sequentially with the query image. This significantly reduced VRAM consumption.\n\n## Camera Pose Refinement\n\nWhile the matching from the VGGT tracker captured the overall scene structure, its absolute accuracy was limited. This is partly due to VGGT resizing images to 518, which leads to reduced spatial precision in keypoint localization.\n\nTo overcome this, I refined the coarse camera poses estimated by VGGT using high-precision correspondences obtained from ALIKED + LightGlue. The refinement was done using the `incremental_model_refiner`, which was also used in the 1st place solution of IMC2023.\n\nhttps://github.com/zju3dv/DetectorFreeSfM/blob/main/src/sfm_runner/sfm_model_geometry_refiner.py\n\n## Experiments\n\nWe evaluated the effectiveness of our method using validation data provided by the organizers.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2F2bb53df3cdff4aae5da03ecd08118c69%2Fresults.png?generation=1748995584825997&alt=media)\n\n| Method                 | Public LB | Private LB | Val (amy_gardens) | Val (fbk_vineyard) | Val (ETs) | Val (stairs) |\n|------------------------|-----------|------------|--------------------|---------------------|-----------|---------------|\n| LightGlue with refiner | 39.91     | 42.33      | 27.19              | 45.91               | 61.33     | 6.25          |\n| VGGT with refiner      | 38.55     | 41.02      | 44.88              | 63.26               | 66.67     | 11.77         |\n| VGGT w/o refiner       | 31.04     | 36.65      | 39.34              | 66.23               | 17.54     | 14.29         |\n\n\nBelow are some qualitative results on selected scenes:\n\n**amy_gardens**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2Fc47def199d10a5760c9137fafea5d1af%2Famy_gardens.png?generation=1748995608568829&alt=media)\n\n**fbk_vineyard**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2F6cf484b03ac0275c9fce24190cdd5e1c%2Fvineyard1.png?generation=1748995649687449&alt=media)\n\n**ETs**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2Fc21f6edecaf0ded471376ff0f816d827%2FET1.png?generation=1748995665949055&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2F29d19e09d6db3ef6e622e356d15c981b%2FET2.png?generation=1748995676428148&alt=media)\n\nAlthough the proposed method performed well on the validation data, it underperformed on the leaderboard compared to conventional LightGlue-based approaches. As a result, the final submission was based on the baseline LightGlue-based method."
  }
}