{
  "id": 510499,
  "title": "2nd Place Solution: MST-Aided SfM & Transparent Scene Solution [Prize Eligible]",
  "url": "/competitions/image-matching-challenge-2024/writeups/aloc-2nd-place-solution-mst-aided-sfm-transparent-",
  "author_name": "",
  "post_date": "2024-07-01T12:58:22.590Z",
  "votes": 44,
  "comment_count": 7,
  "views": 0,
  "content": "<p>We would like to express our gratitude to the Kaggle community and the organizers from Czech Technical University in Prague for their contributions to this competition. We also appreciate enthusiastic discussions from all participants. Congratulations to all the participating teams!<br>\nThe work described here is actually a joint effort by <a href=\"https://www.kaggle.com/sunnyykk\" target=\"_blank\">@sunnyykk</a>, <a href=\"https://www.kaggle.com/gdchenhao\" target=\"_blank\">@gdchenhao</a>, <a href=\"https://www.kaggle.com/mayunchaoamap\" target=\"_blank\">@mayunchaoamap</a> and <a href=\"https://www.kaggle.com/wangshengyi96\" target=\"_blank\">@wangshengyi96</a>. We especially thank our mentor Zhang Tao for his strong support and guidance.<br>\nI'm very glad to have been part of such an excellent team participating in IMC-2024. Throughout the competition, we have gained a lot and learned a lot.<br>\nIt is an honor to share our solution with you now.</p>\n<h2>1. Overview</h2>\n<p>Undoubtedly, the biggest difference between this year's competition and previous competitions is the appearence of transparent and reflective scenes. After many trials, we found it challenging to develop a general solution that could handle both transparent and conventional scenes simultaneously. Therefore, we designed different methods to tackle the two types of scenes.<br>\n<strong>Conventional Scenes</strong>: We designed an iterative optimization SfM scheme based on the Minimum Spanning Tree (MST). We use the coarse model reconstructed from the most concise data association as a skeleton and iteratively add redundant associations to optimize the accuracy of the coarse model.<br>\n<strong>Transparent and Reflective Scenes</strong>: We assume that the camera captures a transparent object in a circumferential manner and focus on calculating the shooting sequence of the images. Then we place the images in the corresponding positions and orientations. Notably, we designed a new global descriptor that can effectively capture detailed information in these scenes, enabling accurate camera pose estimation.<br>\nMoreover, our code is based on the open-source solution of the 7th team <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/427143\" target=\"_blank\">RMD-3DV</a> from IMC-2023. This robust baseline, which combines technologies like feature ensemble, pixsfm, and reloc with colmap, allowed us to build upon their framework without starting from scratch. We appreciate their generous contribution.</p>\n<h2>2. Method</h2>\n<p>This is the pipeline of our solution:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15672041%2F42d7d2c74bfa064e7f18da3a4a575c8c%2F2024-07-01%2020.54.59.png?generation=1719838528246968&amp;alt=media\"></p>\n<h3>2.1 Preprocessing</h3>\n<p>We start by performing rotation detection on the images and determine whether the scene is transparent or not.</p>\n<ul>\n<li><strong>Rotation Detection</strong>:&nbsp;We use a rotation detection model to predict and correct the image rotation. However, recognizing that the model's predictions are not always accurate, we retain the original rotations if less than 10% of the images are predicted rotated.</li>\n<li><strong>Shared Camera Intrinsics</strong>:&nbsp;If all the image dimensions are identical, we set all cameras to share the same internal parameters, occasionally bringing a 0.01 improvement in results.</li>\n<li><strong>Transparency Detection</strong>:&nbsp;We calculate the average difference between different images to determine whether the scene is transparent or not for separate handling.</li>\n</ul>\n<h3>2.2 Global Features</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15672041%2Fca78e1978e7ceb711ec0f41a64e1e27b%2F2024-07-01%2020.56.14.png?generation=1719838610536552&amp;alt=media\"></p>\n<p>We designed a stronger global feature descriptor that yields more reliable image pairs during the image retrieval phase.<br>\nSpecifically, we developed a global descriptor combining point and patch features. The basic approach involves extracting point features (ALIKED) and patch features (DINO) from an image, then establishing a one-to-one correspondence based on their spatial relationships. Using clustering and the VLAD algorithm, we generate global descriptors. This approach allows the clustering algorithm to achieve unsupervised learning of scene features, and incorporating DINO further elevates the learning potential.<br>\nOur method outperforms NetVLAD, AnyLoc, DINO (GAP or GMP), and SALAD on VPR-related datasets. While the image retrieval in the SfM pipeline typically results in many candidate matches (30+), which provides robustness to high recall rates, distinguishing retrieval capabilities becomes less pronounced. Compared to other configurations and excluding transparent scenes, the scores of NetVLAD and our global features are as follows:</p>\n<table>\n<thead>\n<tr>\n<th>Global Descriptor</th>\n<th>Private</th>\n<th>Public</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>NetVLAD</td>\n<td>0.241</td>\n<td>0.230</td>\n</tr>\n<tr>\n<td><strong>Ours</strong></td>\n<td><strong>0.245</strong></td>\n<td><strong>0.247</strong></td>\n</tr>\n</tbody>\n</table>\n<h3>2.3 Local Features</h3>\n<p>We utilized three types of local features and use the ensemble of their matching results:</p>\n<ul>\n<li><strong>Dedode v2 + Dual Softmax</strong></li>\n<li><strong>DISK + LightGlue</strong></li>\n<li><strong>SIFT + Nearest Neighbor</strong></li>\n</ul>\n<p>The v2 version of the Dedode detector produces richer and more evenly distributed feature points. We selected the pre-trained G-upright as the descriptor and dual softmax as the matcher.</p>\n<table>\n<thead>\n<tr>\n<th>Local Feature Configuration</th>\n<th>Private</th>\n<th>Public</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ALIKED + DISK + SIFT</td>\n<td><strong>0.184</strong></td>\n<td>0.169</td>\n</tr>\n<tr>\n<td>SuperPoint + DISK + SIFT</td>\n<td>0.178</td>\n<td>0.172</td>\n</tr>\n<tr>\n<td>Dedode v1 + DISK + SIFT</td>\n<td>0.179</td>\n<td>0.177</td>\n</tr>\n<tr>\n<td><strong>Dedode v2 + DISK + SIFT</strong></td>\n<td><strong>0.184</strong></td>\n<td><strong>0.185</strong></td>\n</tr>\n</tbody>\n</table>\n<h3>2.4 MST-Aided Coarse-to-Fine SfM Solution</h3>\n<p>In the data association phase of SfM, extensive feature matching is typically undertaken to enhance the robustness and accuracy of SfM. However, increasing data associations in scenes with repetitive textures can result in more incorrect matches. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15672041%2Fd7468879e081409467ad74f5fc895c83%2F2024-06-06%2020.17.12.png?generation=1717676240666535&amp;alt=media\"><br>\nThis issue was evident in the reconstruction result of church scene in the train set.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15672041%2F06b5b48d023cc71d8183039d0902750b%2F2024-06-06%2020.02.28.png?generation=1717675369117964&amp;alt=media\"> <br>\nTo address this, we proposed a coarse-to-fine SfM solution based on the Minimum Spanning Tree (MST)</p>\n<p></p><ul><br>\n<li><strong>Stage 1</strong>:&nbsp;We construct a similarity graph where vertices represent images, and edges represent similarity. By computing the MST, we obtain a globally optimal data association linking all image nodes, which is used for the first SfM. Experiments showed this method removes a large amount of incorrect associations, significantly improving coarse-grained accuracy but somewhat losing fine-grained accuracy.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15672041%2Ff2258dbcc3d0707a98f7db089871a659%2F1.png?generation=1717675415406880&amp;alt=media\"></li><br>\n<li><strong>Stage 2</strong>:&nbsp;We use full data associations and the coarse model from Stage 1 providing initial camera pose priors for geometric verification. This filters out incorrect feature matches in the full data association, leading to the final model. The data redundancy maintains the coarse-grained advantages of Stage 1 while compensating for its fine-grained accuracy losses.<br><p></p>\n<table>\n<thead>\n<tr>\n<th>SfM Solution</th>\n<th>Private</th>\n<th>Public</th>\n<th><br></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Direct SfM</td>\n<td>0.240</td>\n<td>0.253</td>\n<td><br></td>\n</tr>\n<tr>\n<td>MST-Aided SfM (Stage 1 Only)</td>\n<td>0.216</td>\n<td>0.218</td>\n<td><br></td>\n</tr>\n<tr>\n<td><strong>MST-Aided SfM (Stage 1 + 2)</strong></td>\n<td><strong>0.258</strong></td>\n<td><strong>0.268</strong></td>\n<td></td></tr></tbody></table></li>\n\n\n</ul>\n\n\n\n\n\n<h3>2.5 Post-Processing</h3>\n\n\n\n\n\n\n<p>Following the experiences from previous competitions, we employed pixsfm to optimize the SfM model. Additionally, we deployed an HLoc-based relocalization module to process unregistered images, which typically resulted in a 0~0.01 score improvement.</p>\n<h3>2.6 Transparent Scenes</h3>\n<p>For transparent scenes, we tried to use various local features, including ALIKED, DISK, LoFTR, and DKMV3, but unfortunately none delivered satisfactory results. By observing the training set, we assumed circumferential camera capture of transparent targets. Using similarity graphs from global features, we calculate the min-cost path to determine the image shooting sequence, arranging all cameras in a closed loop to calculate their positions and orientations.<br>\nInterestingly, visualizing the global features of a cylinder also revealed a highly ordered 2D plan layout, aiding significantly in recovering the image capture sequence. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15672041%2F8625f3098ae89ad3ebfb6f7202344376%2F2024-06-06%2020.37.07.png?generation=1717677454710544&amp;alt=media\"><br>\nThe solution for transparent scenes improved scores by approximately <strong>0.06</strong> eventually.</p>\n<table>\n<thead>\n<tr>\n<th>Solution</th>\n<th>Private</th>\n<th>Public</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Without Transparent Solution</td>\n<td>0.184</td>\n<td>0.185</td>\n</tr>\n<tr>\n<td><strong>With Transparent Solution</strong></td>\n<td><strong>0.242</strong></td>\n<td><strong>0.249</strong></td>\n</tr>\n</tbody>\n</table>\n<h2>3. Other Ideas</h2>\n<h3>3.1 Methods that did not work</h3>\n<ul>\n<li><strong>Enhanced First Phase of MST-Aided SfM</strong>:&nbsp;By adding more redundant edges (e.g., top-k, alternative key paths) to MST, aiming to balance recall and precision in the first reconstruction phase. While private score reached a high of <strong>0.263</strong>, the public score was only <strong>0.249</strong>, so this was not selected.</li>\n<li><strong>Lightglue Matcher for Dedode</strong>:&nbsp;We tested Kornia's lightglue matcher for dedode (b/g), but it performed worse than dual softmax.</li>\n<li><strong>Day-Night Challenge</strong>:&nbsp;To tackle day-night challenges in pond and lizard datasets, we attempted brightness-based day/night classification, normalized feature distribution, but encountered unknown \"Throw Exception\" obstacles.</li>\n<li><strong>Dense Matchers</strong>:&nbsp;We tested dense matchers like LoFTR, EfficientLoFTR, and DKMV3 in conventional scenes but saw no score improvement, likely due to our unoptimized parameters (matching threshold, grid size, etc.).</li>\n<li><strong>Neural Network-Based SfM</strong>:&nbsp;We explored neural network-based SfM methods like Dust3R for transparent scenes, but results were not pretty well.</li>\n<li><strong>Dense Optical Flow</strong>:&nbsp;Using dense optical flow to restore the image capture sequence in transparent scenes also worked but was less effective than our proposed global feature.</li>\n</ul>\n<h3>3.2 Methods that we haven't tried</h3>\n<ul>\n<li><strong>TTA</strong>:&nbsp;We didn't try TTA methods due to runtime considerations.</li>\n<li><strong>Multi-Scale Matching</strong>:&nbsp;We didn't detect covisible regions or perform multi-scale local feature matching.<br>\n﻿</li>\n</ul>",
  "messages": [
    {
      "id": "2858332",
      "postDate": "06/06/2024 12:33:00",
      "content": "<p>We would like to express our gratitude to the Kaggle community and the organizers from Czech Technical University in Prague for their contributions to this competition. We also appreciate enthusiastic discussions from all participants. Congratulations to all the participating teams!<br>\nThe work described here is actually a joint effort by <a href=\"https://www.kaggle.com/sunnyykk\" target=\"_blank\">@sunnyykk</a>, <a href=\"https://www.kaggle.com/gdchenhao\" target=\"_blank\">@gdchenhao</a>, <a href=\"https://www.kaggle.com/mayunchaoamap\" target=\"_blank\">@mayunchaoamap</a> and <a href=\"https://www.kaggle.com/wangshengyi96\" target=\"_blank\">@wangshengyi96</a>. We especially thank our mentor Zhang Tao for his strong support and guidance.<br>\nI'm very glad to have been part of such an excellent team participating in IMC-2024. Throughout the competition, we have gained a lot and learned a lot.<br>\nIt is an honor to share our solution with you now.</p>\n<h2>1. Overview</h2>\n<p>Undoubtedly, the biggest difference between this year's competition and previous competitions is the appearence of transparent and reflective scenes. After many trials, we found it challenging to develop a general solution that could handle both transparent and conventional scenes simultaneously. Therefore, we designed different methods to tackle the two types of scenes.<br>\n<strong>Conventional Scenes</strong>: We designed an iterative optimization SfM scheme based on the Minimum Spanning Tree (MST). We use the coarse model reconstructed from the most concise data association as a skeleton and iteratively add redundant associations to optimize the accuracy of the coarse model.<br>\n<strong>Transparent and Reflective Scenes</strong>: We assume that the camera captures a transparent object in a circumferential manner and focus on calculating the shooting sequence of the images. Then we place the images in the corresponding positions and orientations. Notably, we designed a new global descriptor that can effectively capture detailed information in these scenes, enabling accurate camera pose estimation.<br>\nMoreover, our code is based on the open-source solution of the 7th team <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/427143\" target=\"_blank\">RMD-3DV</a> from IMC-2023. This robust baseline, which combines technologies like feature ensemble, pixsfm, and reloc with colmap, allowed us to build upon their framework without starting from scratch. We appreciate their generous contribution.</p>\n<h2>2. Method</h2>\n<p>This is the pipeline of our solution:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15672041%2F42d7d2c74bfa064e7f18da3a4a575c8c%2F2024-07-01%2020.54.59.png?generation=1719838528246968&amp;alt=media\"></p>\n<h3>2.1 Preprocessing</h3>\n<p>We start by performing rotation detection on the images and determine whether the scene is transparent or not.</p>\n<ul>\n<li><strong>Rotation Detection</strong>:&nbsp;We use a rotation detection model to predict and correct the image rotation. However, recognizing that the model's predictions are not always accurate, we retain the original rotations if less than 10% of the images are predicted rotated.</li>\n<li><strong>Shared Camera Intrinsics</strong>:&nbsp;If all the image dimensions are identical, we set all cameras to share the same internal parameters, occasionally bringing a 0.01 improvement in results.</li>\n<li><strong>Transparency Detection</strong>:&nbsp;We calculate the average difference between different images to determine whether the scene is transparent or not for separate handling.</li>\n</ul>\n<h3>2.2 Global Features</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15672041%2Fca78e1978e7ceb711ec0f41a64e1e27b%2F2024-07-01%2020.56.14.png?generation=1719838610536552&amp;alt=media\"></p>\n<p>We designed a stronger global feature descriptor that yields more reliable image pairs during the image retrieval phase.<br>\nSpecifically, we developed a global descriptor combining point and patch features. The basic approach involves extracting point features (ALIKED) and patch features (DINO) from an image, then establishing a one-to-one correspondence based on their spatial relationships. Using clustering and the VLAD algorithm, we generate global descriptors. This approach allows the clustering algorithm to achieve unsupervised learning of scene features, and incorporating DINO further elevates the learning potential.<br>\nOur method outperforms NetVLAD, AnyLoc, DINO (GAP or GMP), and SALAD on VPR-related datasets. While the image retrieval in the SfM pipeline typically results in many candidate matches (30+), which provides robustness to high recall rates, distinguishing retrieval capabilities becomes less pronounced. Compared to other configurations and excluding transparent scenes, the scores of NetVLAD and our global features are as follows:</p>\n<table>\n<thead>\n<tr>\n<th>Global Descriptor</th>\n<th>Private</th>\n<th>Public</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>NetVLAD</td>\n<td>0.241</td>\n<td>0.230</td>\n</tr>\n<tr>\n<td><strong>Ours</strong></td>\n<td><strong>0.245</strong></td>\n<td><strong>0.247</strong></td>\n</tr>\n</tbody>\n</table>\n<h3>2.3 Local Features</h3>\n<p>We utilized three types of local features and use the ensemble of their matching results:</p>\n<ul>\n<li><strong>Dedode v2 + Dual Softmax</strong></li>\n<li><strong>DISK + LightGlue</strong></li>\n<li><strong>SIFT + Nearest Neighbor</strong></li>\n</ul>\n<p>The v2 version of the Dedode detector produces richer and more evenly distributed feature points. We selected the pre-trained G-upright as the descriptor and dual softmax as the matcher.</p>\n<table>\n<thead>\n<tr>\n<th>Local Feature Configuration</th>\n<th>Private</th>\n<th>Public</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ALIKED + DISK + SIFT</td>\n<td><strong>0.184</strong></td>\n<td>0.169</td>\n</tr>\n<tr>\n<td>SuperPoint + DISK + SIFT</td>\n<td>0.178</td>\n<td>0.172</td>\n</tr>\n<tr>\n<td>Dedode v1 + DISK + SIFT</td>\n<td>0.179</td>\n<td>0.177</td>\n</tr>\n<tr>\n<td><strong>Dedode v2 + DISK + SIFT</strong></td>\n<td><strong>0.184</strong></td>\n<td><strong>0.185</strong></td>\n</tr>\n</tbody>\n</table>\n<h3>2.4 MST-Aided Coarse-to-Fine SfM Solution</h3>\n<p>In the data association phase of SfM, extensive feature matching is typically undertaken to enhance the robustness and accuracy of SfM. However, increasing data associations in scenes with repetitive textures can result in more incorrect matches. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15672041%2Fd7468879e081409467ad74f5fc895c83%2F2024-06-06%2020.17.12.png?generation=1717676240666535&amp;alt=media\"><br>\nThis issue was evident in the reconstruction result of church scene in the train set.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15672041%2F06b5b48d023cc71d8183039d0902750b%2F2024-06-06%2020.02.28.png?generation=1717675369117964&amp;alt=media\"> <br>\nTo address this, we proposed a coarse-to-fine SfM solution based on the Minimum Spanning Tree (MST)</p>\n<p></p><ul><br>\n<li><strong>Stage 1</strong>:&nbsp;We construct a similarity graph where vertices represent images, and edges represent similarity. By computing the MST, we obtain a globally optimal data association linking all image nodes, which is used for the first SfM. Experiments showed this method removes a large amount of incorrect associations, significantly improving coarse-grained accuracy but somewhat losing fine-grained accuracy.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15672041%2Ff2258dbcc3d0707a98f7db089871a659%2F1.png?generation=1717675415406880&amp;alt=media\"></li><br>\n<li><strong>Stage 2</strong>:&nbsp;We use full data associations and the coarse model from Stage 1 providing initial camera pose priors for geometric verification. This filters out incorrect feature matches in the full data association, leading to the final model. The data redundancy maintains the coarse-grained advantages of Stage 1 while compensating for its fine-grained accuracy losses.<br><p></p>\n<table>\n<thead>\n<tr>\n<th>SfM Solution</th>\n<th>Private</th>\n<th>Public</th>\n<th><br></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Direct SfM</td>\n<td>0.240</td>\n<td>0.253</td>\n<td><br></td>\n</tr>\n<tr>\n<td>MST-Aided SfM (Stage 1 Only)</td>\n<td>0.216</td>\n<td>0.218</td>\n<td><br></td>\n</tr>\n<tr>\n<td><strong>MST-Aided SfM (Stage 1 + 2)</strong></td>\n<td><strong>0.258</strong></td>\n<td><strong>0.268</strong></td>\n<td></td></tr></tbody></table></li>\n\n\n</ul>\n\n\n\n\n\n<h3>2.5 Post-Processing</h3>\n\n\n\n\n\n\n<p>Following the experiences from previous competitions, we employed pixsfm to optimize the SfM model. Additionally, we deployed an HLoc-based relocalization module to process unregistered images, which typically resulted in a 0~0.01 score improvement.</p>\n<h3>2.6 Transparent Scenes</h3>\n<p>For transparent scenes, we tried to use various local features, including ALIKED, DISK, LoFTR, and DKMV3, but unfortunately none delivered satisfactory results. By observing the training set, we assumed circumferential camera capture of transparent targets. Using similarity graphs from global features, we calculate the min-cost path to determine the image shooting sequence, arranging all cameras in a closed loop to calculate their positions and orientations.<br>\nInterestingly, visualizing the global features of a cylinder also revealed a highly ordered 2D plan layout, aiding significantly in recovering the image capture sequence. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15672041%2F8625f3098ae89ad3ebfb6f7202344376%2F2024-06-06%2020.37.07.png?generation=1717677454710544&amp;alt=media\"><br>\nThe solution for transparent scenes improved scores by approximately <strong>0.06</strong> eventually.</p>\n<table>\n<thead>\n<tr>\n<th>Solution</th>\n<th>Private</th>\n<th>Public</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Without Transparent Solution</td>\n<td>0.184</td>\n<td>0.185</td>\n</tr>\n<tr>\n<td><strong>With Transparent Solution</strong></td>\n<td><strong>0.242</strong></td>\n<td><strong>0.249</strong></td>\n</tr>\n</tbody>\n</table>\n<h2>3. Other Ideas</h2>\n<h3>3.1 Methods that did not work</h3>\n<ul>\n<li><strong>Enhanced First Phase of MST-Aided SfM</strong>:&nbsp;By adding more redundant edges (e.g., top-k, alternative key paths) to MST, aiming to balance recall and precision in the first reconstruction phase. While private score reached a high of <strong>0.263</strong>, the public score was only <strong>0.249</strong>, so this was not selected.</li>\n<li><strong>Lightglue Matcher for Dedode</strong>:&nbsp;We tested Kornia's lightglue matcher for dedode (b/g), but it performed worse than dual softmax.</li>\n<li><strong>Day-Night Challenge</strong>:&nbsp;To tackle day-night challenges in pond and lizard datasets, we attempted brightness-based day/night classification, normalized feature distribution, but encountered unknown \"Throw Exception\" obstacles.</li>\n<li><strong>Dense Matchers</strong>:&nbsp;We tested dense matchers like LoFTR, EfficientLoFTR, and DKMV3 in conventional scenes but saw no score improvement, likely due to our unoptimized parameters (matching threshold, grid size, etc.).</li>\n<li><strong>Neural Network-Based SfM</strong>:&nbsp;We explored neural network-based SfM methods like Dust3R for transparent scenes, but results were not pretty well.</li>\n<li><strong>Dense Optical Flow</strong>:&nbsp;Using dense optical flow to restore the image capture sequence in transparent scenes also worked but was less effective than our proposed global feature.</li>\n</ul>\n<h3>3.2 Methods that we haven't tried</h3>\n<ul>\n<li><strong>TTA</strong>:&nbsp;We didn't try TTA methods due to runtime considerations.</li>\n<li><strong>Multi-Scale Matching</strong>:&nbsp;We didn't detect covisible regions or perform multi-scale local feature matching.<br>\n﻿</li>\n</ul>",
      "rawMarkdown": "We would like to express our gratitude to the Kaggle community and the organizers from Czech Technical University in Prague for their contributions to this competition. We also appreciate enthusiastic discussions from all participants. Congratulations to all the participating teams!\nThe work described here is actually a joint effort by @sunnyykk, @gdchenhao, @mayunchaoamap and @wangshengyi96. We especially thank our mentor Zhang Tao for his strong support and guidance.\nI'm very glad to have been part of such an excellent team participating in IMC-2024. Throughout the competition, we have gained a lot and learned a lot.\nIt is an honor to share our solution with you now.\n\n## 1. Overview\nUndoubtedly, the biggest difference between this year's competition and previous competitions is the appearence of transparent and reflective scenes. After many trials, we found it challenging to develop a general solution that could handle both transparent and conventional scenes simultaneously. Therefore, we designed different methods to tackle the two types of scenes.\n**Conventional Scenes**: We designed an iterative optimization SfM scheme based on the Minimum Spanning Tree (MST). We use the coarse model reconstructed from the most concise data association as a skeleton and iteratively add redundant associations to optimize the accuracy of the coarse model.\n**Transparent and Reflective Scenes**: We assume that the camera captures a transparent object in a circumferential manner and focus on calculating the shooting sequence of the images. Then we place the images in the corresponding positions and orientations. Notably, we designed a new global descriptor that can effectively capture detailed information in these scenes, enabling accurate camera pose estimation.\nMoreover, our code is based on the open-source solution of the 7th team [RMD-3DV](https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/427143) from IMC-2023. This robust baseline, which combines technologies like feature ensemble, pixsfm, and reloc with colmap, allowed us to build upon their framework without starting from scratch. We appreciate their generous contribution.\n\n## 2. Method\nThis is the pipeline of our solution:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15672041%2F42d7d2c74bfa064e7f18da3a4a575c8c%2F2024-07-01%2020.54.59.png?generation=1719838528246968&alt=media)\n### 2.1 Preprocessing\nWe start by performing rotation detection on the images and determine whether the scene is transparent or not.\n- **Rotation Detection**: We use a rotation detection model to predict and correct the image rotation. However, recognizing that the model's predictions are not always accurate, we retain the original rotations if less than 10% of the images are predicted rotated.\n- **Shared Camera Intrinsics**: If all the image dimensions are identical, we set all cameras to share the same internal parameters, occasionally bringing a 0.01 improvement in results.\n- **Transparency Detection**: We calculate the average difference between different images to determine whether the scene is transparent or not for separate handling.\n\n   \n### 2.2 Global Features\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15672041%2Fca78e1978e7ceb711ec0f41a64e1e27b%2F2024-07-01%2020.56.14.png?generation=1719838610536552&alt=media =600x300)\n\n\nWe designed a stronger global feature descriptor that yields more reliable image pairs during the image retrieval phase.\nSpecifically, we developed a global descriptor combining point and patch features. The basic approach involves extracting point features (ALIKED) and patch features (DINO) from an image, then establishing a one-to-one correspondence based on their spatial relationships. Using clustering and the VLAD algorithm, we generate global descriptors. This approach allows the clustering algorithm to achieve unsupervised learning of scene features, and incorporating DINO further elevates the learning potential.\nOur method outperforms NetVLAD, AnyLoc, DINO (GAP or GMP), and SALAD on VPR-related datasets. While the image retrieval in the SfM pipeline typically results in many candidate matches (30+), which provides robustness to high recall rates, distinguishing retrieval capabilities becomes less pronounced. Compared to other configurations and excluding transparent scenes, the scores of NetVLAD and our global features are as follows:\n\n| Global Descriptor        | Private     | Public  |\n| ------------- |:-------------:| -----:|\n| NetVLAD  | 0.241 | 0.230 |\n| **Ours**      | **0.245**    |  **0.247** |\n\n### 2.3 Local Features\nWe utilized three types of local features and use the ensemble of their matching results:\n- **Dedode v2 + Dual Softmax**\n- **DISK + LightGlue**\n- **SIFT + Nearest Neighbor**\n\nThe v2 version of the Dedode detector produces richer and more evenly distributed feature points. We selected the pre-trained G-upright as the descriptor and dual softmax as the matcher.\n\n| Local Feature Configuration        | Private     | Public  |\n| ------------- |:-------------:| -----:|\n| ALIKED + DISK + SIFT  | **0.184** | 0.169 |\n| SuperPoint + DISK + SIFT | 0.178    |  0.172 |\n| Dedode v1 + DISK + SIFT | 0.179   |  0.177 |\n| **Dedode v2 + DISK + SIFT** | **0.184**   | **0.185** |\n\n### 2.4 MST-Aided Coarse-to-Fine SfM Solution\nIn the data association phase of SfM, extensive feature matching is typically undertaken to enhance the robustness and accuracy of SfM. However, increasing data associations in scenes with repetitive textures can result in more incorrect matches. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15672041%2Fd7468879e081409467ad74f5fc895c83%2F2024-06-06%2020.17.12.png?generation=1717676240666535&alt=media =400x265)\nThis issue was evident in the reconstruction result of church scene in the train set.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15672041%2F06b5b48d023cc71d8183039d0902750b%2F2024-06-06%2020.02.28.png?generation=1717675369117964&alt=media =730x200) \nTo address this, we proposed a coarse-to-fine SfM solution based on the Minimum Spanning Tree (MST)\n- **Stage 1**: We construct a similarity graph where vertices represent images, and edges represent similarity. By computing the MST, we obtain a globally optimal data association linking all image nodes, which is used for the first SfM. Experiments showed this method removes a large amount of incorrect associations, significantly improving coarse-grained accuracy but somewhat losing fine-grained accuracy.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15672041%2Ff2258dbcc3d0707a98f7db089871a659%2F1.png?generation=1717675415406880&alt=media =550x500)\n- **Stage 2**: We use full data associations and the coarse model from Stage 1 providing initial camera pose priors for geometric verification. This filters out incorrect feature matches in the full data association, leading to the final model. The data redundancy maintains the coarse-grained advantages of Stage 1 while compensating for its fine-grained accuracy losses.\n| SfM Solution        | Private     | Public  |\n| ------------- |:-------------:| -----:|\n| Direct SfM  | 0.240 | 0.253 |\n| MST-Aided SfM (Stage 1 Only)      | 0.216    | 0.218 |\n| **MST-Aided SfM (Stage 1 + 2)**   | **0.258**  | **0.268** |\n\n### 2.5 Post-Processing\nFollowing the experiences from previous competitions, we employed pixsfm to optimize the SfM model. Additionally, we deployed an HLoc-based relocalization module to process unregistered images, which typically resulted in a 0~0.01 score improvement.\n\n### 2.6 Transparent Scenes\nFor transparent scenes, we tried to use various local features, including ALIKED, DISK, LoFTR, and DKMV3, but unfortunately none delivered satisfactory results. By observing the training set, we assumed circumferential camera capture of transparent targets. Using similarity graphs from global features, we calculate the min-cost path to determine the image shooting sequence, arranging all cameras in a closed loop to calculate their positions and orientations.\nInterestingly, visualizing the global features of a cylinder also revealed a highly ordered 2D plan layout, aiding significantly in recovering the image capture sequence. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15672041%2F8625f3098ae89ad3ebfb6f7202344376%2F2024-06-06%2020.37.07.png?generation=1717677454710544&alt=media =700x280)\nThe solution for transparent scenes improved scores by approximately **0.06** eventually.\n| Solution       | Private     | Public  |\n| ------------- |:-------------:| -----:|\n| Without Transparent Solution  | 0.184 | 0.185 |\n| **With Transparent Solution**   | **0.242**  | **0.249** |\n\n## 3. Other Ideas\n### 3.1 Methods that did not work\n- **Enhanced First Phase of MST-Aided SfM**: By adding more redundant edges (e.g., top-k, alternative key paths) to MST, aiming to balance recall and precision in the first reconstruction phase. While private score reached a high of **0.263**, the public score was only **0.249**, so this was not selected.\n- **Lightglue Matcher for Dedode**: We tested Kornia's lightglue matcher for dedode (b/g), but it performed worse than dual softmax.\n- **Day-Night Challenge**: To tackle day-night challenges in pond and lizard datasets, we attempted brightness-based day/night classification, normalized feature distribution, but encountered unknown \"Throw Exception\" obstacles.\n- **Dense Matchers**: We tested dense matchers like LoFTR, EfficientLoFTR, and DKMV3 in conventional scenes but saw no score improvement, likely due to our unoptimized parameters (matching threshold, grid size, etc.).\n- **Neural Network-Based SfM**: We explored neural network-based SfM methods like Dust3R for transparent scenes, but results were not pretty well.\n- **Dense Optical Flow**: Using dense optical flow to restore the image capture sequence in transparent scenes also worked but was less effective than our proposed global feature.\n\n### 3.2 Methods that we haven't tried\n- **TTA**: We didn't try TTA methods due to runtime considerations.\n- **Multi-Scale Matching**: We didn't detect covisible regions or perform multi-scale local feature matching.\n﻿",
      "votes": null
    },
    {
      "id": "2860339",
      "postDate": "06/07/2024 15:01:35",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": null
    },
    {
      "id": "2862990",
      "postDate": "06/09/2024 07:09:48",
      "content": "<p>Thank you for sharing the solution.<br>\nIt was a very interesting solution.</p>\n<p>I have a question about the Global Features.<br>\nWhen it says \"establishing a one-to-one correspondence based on their spatial relationships,\" does this mean to concatenate the feature of a keypoint and the feature of the patch where the keypoint is located?</p>\n<p>For example, in the attached image, is the final feature corresponding to keypoint K6 a concatenation of the feature of keypoint K6 (feat_K6) and the feature of patch P24 (feat_P24) where keypoint K6 is located?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3823496%2F4e145c21dd6f25d6b4dd8ac16ea2981a%2F2nd_place_question.jpg?generation=1717916605082142&amp;alt=media\"></p>",
      "rawMarkdown": "Thank you for sharing the solution.\nIt was a very interesting solution.\n\nI have a question about the Global Features.\nWhen it says \"establishing a one-to-one correspondence based on their spatial relationships,\" does this mean to concatenate the feature of a keypoint and the feature of the patch where the keypoint is located?\n\nFor example, in the attached image, is the final feature corresponding to keypoint K6 a concatenation of the feature of keypoint K6 (feat_K6) and the feature of patch P24 (feat_P24) where keypoint K6 is located?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3823496%2F4e145c21dd6f25d6b4dd8ac16ea2981a%2F2nd_place_question.jpg?generation=1717916605082142&alt=media)",
      "votes": null
    },
    {
      "id": "2865834",
      "postDate": "06/11/2024 03:09:50",
      "content": "<p>Congratulations with the 1st place!<br>\nThat's right! Exactly. <br>\nExtract the patch feature where each point is located and concatenate it to the point feature. <br>\nWe believe this is a direct way to enhance the semantics of the point, giving it better discriminative power for scene semantics, thereby better serving downstream tasks like constructing MST and restoring the order of transparent objects.</p>",
      "rawMarkdown": "Congratulations with the 1st place!\nThat's right! Exactly. \nExtract the patch feature where each point is located and concatenate it to the point feature. \nWe believe this is a direct way to enhance the semantics of the point, giving it better discriminative power for scene semantics, thereby better serving downstream tasks like constructing MST and restoring the order of transparent objects.",
      "votes": null
    },
    {
      "id": "2866658",
      "postDate": "06/11/2024 12:46:19",
      "content": "<p>Thank you for your response.<br>\nI think it's a simple yet revolutionary and excellent global feature that performs well in downstream tasks for all scenes, including transparent ones. I'm very impressed.</p>\n<p>Also, congratulations again on your 2nd place finish!</p>",
      "rawMarkdown": "Thank you for your response.\nI think it's a simple yet revolutionary and excellent global feature that performs well in downstream tasks for all scenes, including transparent ones. I'm very impressed.\n\nAlso, congratulations again on your 2nd place finish!",
      "votes": null
    },
    {
      "id": "2872144",
      "postDate": "06/14/2024 15:47:28",
      "content": "<p>Thank you for sharing the solution 👏</p>",
      "rawMarkdown": "Thank you for sharing the solution 👏",
      "votes": null
    },
    {
      "id": "2927238",
      "postDate": "07/18/2024 11:21:52",
      "content": "<p>Congratulation! Thanks for sharing the great solution especially the MST-Aided Coarse-to-Fine SfM pipeline.</p>\n<p>But I still have some questions:</p>\n<ol>\n<li>what is the relationship between your similarity graph and MST? </li>\n<li>how you construct the similarity graph and MST? </li>\n</ol>\n<p>My understanding is that the edge with the largest weight in the similarity graph is selected to construct the MST. In other words, the MST is the shortest path that connects the vertices in the similarity graph with the largest sum of weights (similarity)?</p>\n<p>Looking forward to your answer, the more detailed the better, thank you very much.</p>",
      "rawMarkdown": "Congratulation! Thanks for sharing the great solution especially the MST-Aided Coarse-to-Fine SfM pipeline.\n\nBut I still have some questions:\n1. what is the relationship between your similarity graph and MST? \n2. how you construct the similarity graph and MST? \n\nMy understanding is that the edge with the largest weight in the similarity graph is selected to construct the MST. In other words, the MST is the shortest path that connects the vertices in the similarity graph with the largest sum of weights (similarity)?\n\nLooking forward to your answer, the more detailed the better, thank you very much.",
      "votes": null
    },
    {
      "id": "3000069",
      "postDate": "09/27/2024 09:21:17",
      "content": "<p>Your understanding is roughly correct, but MST is a tree like structure. For our solution, MST is the spanning tree for the similarity graph with minimal cost (highest similarity).</p>",
      "rawMarkdown": "Your understanding is roughly correct, but MST is a tree like structure. For our solution, MST is the spanning tree for the similarity graph with minimal cost (highest similarity).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2860339,
      "author_name": "bmcwrap",
      "author_url": "",
      "post_date": "06/07/2024 15:01:35",
      "content": "<p>Thanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2862990,
      "author_name": "kashiwaba",
      "author_url": "",
      "post_date": "06/09/2024 07:09:48",
      "content": "<p>Thank you for sharing the solution.<br>\nIt was a very interesting solution.</p>\n<p>I have a question about the Global Features.<br>\nWhen it says \"establishing a one-to-one correspondence based on their spatial relationships,\" does this mean to concatenate the feature of a keypoint and the feature of the patch where the keypoint is located?</p>\n<p>For example, in the attached image, is the final feature corresponding to keypoint K6 a concatenation of the feature of keypoint K6 (feat_K6) and the feature of patch P24 (feat_P24) where keypoint K6 is located?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3823496%2F4e145c21dd6f25d6b4dd8ac16ea2981a%2F2nd_place_question.jpg?generation=1717916605082142&amp;alt=media\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 2865834,
          "author_name": "sunnyykk",
          "author_url": "",
          "post_date": "06/11/2024 03:09:50",
          "content": "<p>Congratulations with the 1st place!<br>\nThat's right! Exactly. <br>\nExtract the patch feature where each point is located and concatenate it to the point feature. <br>\nWe believe this is a direct way to enhance the semantics of the point, giving it better discriminative power for scene semantics, thereby better serving downstream tasks like constructing MST and restoring the order of transparent objects.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2866658,
              "author_name": "kashiwaba",
              "author_url": "",
              "post_date": "06/11/2024 12:46:19",
              "content": "<p>Thank you for your response.<br>\nI think it's a simple yet revolutionary and excellent global feature that performs well in downstream tasks for all scenes, including transparent ones. I'm very impressed.</p>\n<p>Also, congratulations again on your 2nd place finish!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2872144,
      "author_name": "metinmekiabullrahman",
      "author_url": "",
      "post_date": "06/14/2024 15:47:28",
      "content": "<p>Thank you for sharing the solution 👏</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2927238,
      "author_name": "vanjing",
      "author_url": "",
      "post_date": "07/18/2024 11:21:52",
      "content": "<p>Congratulation! Thanks for sharing the great solution especially the MST-Aided Coarse-to-Fine SfM pipeline.</p>\n<p>But I still have some questions:</p>\n<ol>\n<li>what is the relationship between your similarity graph and MST? </li>\n<li>how you construct the similarity graph and MST? </li>\n</ol>\n<p>My understanding is that the edge with the largest weight in the similarity graph is selected to construct the MST. In other words, the MST is the shortest path that connects the vertices in the similarity graph with the largest sum of weights (similarity)?</p>\n<p>Looking forward to your answer, the more detailed the better, thank you very much.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3000069,
          "author_name": "wangshengyi96",
          "author_url": "",
          "post_date": "09/27/2024 09:21:17",
          "content": "<p>Your understanding is roughly correct, but MST is a tree like structure. For our solution, MST is the spanning tree for the similarity graph with minimal cost (highest similarity).</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2858332": "We would like to express our gratitude to the Kaggle community and the organizers from Czech Technical University in Prague for their contributions to this competition. We also appreciate enthusiastic discussions from all participants. Congratulations to all the participating teams!\nThe work described here is actually a joint effort by @sunnyykk, @gdchenhao, @mayunchaoamap and @wangshengyi96. We especially thank our mentor Zhang Tao for his strong support and guidance.\nI'm very glad to have been part of such an excellent team participating in IMC-2024. Throughout the competition, we have gained a lot and learned a lot.\nIt is an honor to share our solution with you now.\n\n## 1. Overview\nUndoubtedly, the biggest difference between this year's competition and previous competitions is the appearence of transparent and reflective scenes. After many trials, we found it challenging to develop a general solution that could handle both transparent and conventional scenes simultaneously. Therefore, we designed different methods to tackle the two types of scenes.\n**Conventional Scenes**: We designed an iterative optimization SfM scheme based on the Minimum Spanning Tree (MST). We use the coarse model reconstructed from the most concise data association as a skeleton and iteratively add redundant associations to optimize the accuracy of the coarse model.\n**Transparent and Reflective Scenes**: We assume that the camera captures a transparent object in a circumferential manner and focus on calculating the shooting sequence of the images. Then we place the images in the corresponding positions and orientations. Notably, we designed a new global descriptor that can effectively capture detailed information in these scenes, enabling accurate camera pose estimation.\nMoreover, our code is based on the open-source solution of the 7th team [RMD-3DV](https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/427143) from IMC-2023. This robust baseline, which combines technologies like feature ensemble, pixsfm, and reloc with colmap, allowed us to build upon their framework without starting from scratch. We appreciate their generous contribution.\n\n## 2. Method\nThis is the pipeline of our solution:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15672041%2F42d7d2c74bfa064e7f18da3a4a575c8c%2F2024-07-01%2020.54.59.png?generation=1719838528246968&alt=media)\n### 2.1 Preprocessing\nWe start by performing rotation detection on the images and determine whether the scene is transparent or not.\n- **Rotation Detection**: We use a rotation detection model to predict and correct the image rotation. However, recognizing that the model's predictions are not always accurate, we retain the original rotations if less than 10% of the images are predicted rotated.\n- **Shared Camera Intrinsics**: If all the image dimensions are identical, we set all cameras to share the same internal parameters, occasionally bringing a 0.01 improvement in results.\n- **Transparency Detection**: We calculate the average difference between different images to determine whether the scene is transparent or not for separate handling.\n\n   \n### 2.2 Global Features\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15672041%2Fca78e1978e7ceb711ec0f41a64e1e27b%2F2024-07-01%2020.56.14.png?generation=1719838610536552&alt=media =600x300)\n\n\nWe designed a stronger global feature descriptor that yields more reliable image pairs during the image retrieval phase.\nSpecifically, we developed a global descriptor combining point and patch features. The basic approach involves extracting point features (ALIKED) and patch features (DINO) from an image, then establishing a one-to-one correspondence based on their spatial relationships. Using clustering and the VLAD algorithm, we generate global descriptors. This approach allows the clustering algorithm to achieve unsupervised learning of scene features, and incorporating DINO further elevates the learning potential.\nOur method outperforms NetVLAD, AnyLoc, DINO (GAP or GMP), and SALAD on VPR-related datasets. While the image retrieval in the SfM pipeline typically results in many candidate matches (30+), which provides robustness to high recall rates, distinguishing retrieval capabilities becomes less pronounced. Compared to other configurations and excluding transparent scenes, the scores of NetVLAD and our global features are as follows:\n\n| Global Descriptor        | Private     | Public  |\n| ------------- |:-------------:| -----:|\n| NetVLAD  | 0.241 | 0.230 |\n| **Ours**      | **0.245**    |  **0.247** |\n\n### 2.3 Local Features\nWe utilized three types of local features and use the ensemble of their matching results:\n- **Dedode v2 + Dual Softmax**\n- **DISK + LightGlue**\n- **SIFT + Nearest Neighbor**\n\nThe v2 version of the Dedode detector produces richer and more evenly distributed feature points. We selected the pre-trained G-upright as the descriptor and dual softmax as the matcher.\n\n| Local Feature Configuration        | Private     | Public  |\n| ------------- |:-------------:| -----:|\n| ALIKED + DISK + SIFT  | **0.184** | 0.169 |\n| SuperPoint + DISK + SIFT | 0.178    |  0.172 |\n| Dedode v1 + DISK + SIFT | 0.179   |  0.177 |\n| **Dedode v2 + DISK + SIFT** | **0.184**   | **0.185** |\n\n### 2.4 MST-Aided Coarse-to-Fine SfM Solution\nIn the data association phase of SfM, extensive feature matching is typically undertaken to enhance the robustness and accuracy of SfM. However, increasing data associations in scenes with repetitive textures can result in more incorrect matches. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15672041%2Fd7468879e081409467ad74f5fc895c83%2F2024-06-06%2020.17.12.png?generation=1717676240666535&alt=media =400x265)\nThis issue was evident in the reconstruction result of church scene in the train set.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15672041%2F06b5b48d023cc71d8183039d0902750b%2F2024-06-06%2020.02.28.png?generation=1717675369117964&alt=media =730x200) \nTo address this, we proposed a coarse-to-fine SfM solution based on the Minimum Spanning Tree (MST)\n- **Stage 1**: We construct a similarity graph where vertices represent images, and edges represent similarity. By computing the MST, we obtain a globally optimal data association linking all image nodes, which is used for the first SfM. Experiments showed this method removes a large amount of incorrect associations, significantly improving coarse-grained accuracy but somewhat losing fine-grained accuracy.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15672041%2Ff2258dbcc3d0707a98f7db089871a659%2F1.png?generation=1717675415406880&alt=media =550x500)\n- **Stage 2**: We use full data associations and the coarse model from Stage 1 providing initial camera pose priors for geometric verification. This filters out incorrect feature matches in the full data association, leading to the final model. The data redundancy maintains the coarse-grained advantages of Stage 1 while compensating for its fine-grained accuracy losses.\n| SfM Solution        | Private     | Public  |\n| ------------- |:-------------:| -----:|\n| Direct SfM  | 0.240 | 0.253 |\n| MST-Aided SfM (Stage 1 Only)      | 0.216    | 0.218 |\n| **MST-Aided SfM (Stage 1 + 2)**   | **0.258**  | **0.268** |\n\n### 2.5 Post-Processing\nFollowing the experiences from previous competitions, we employed pixsfm to optimize the SfM model. Additionally, we deployed an HLoc-based relocalization module to process unregistered images, which typically resulted in a 0~0.01 score improvement.\n\n### 2.6 Transparent Scenes\nFor transparent scenes, we tried to use various local features, including ALIKED, DISK, LoFTR, and DKMV3, but unfortunately none delivered satisfactory results. By observing the training set, we assumed circumferential camera capture of transparent targets. Using similarity graphs from global features, we calculate the min-cost path to determine the image shooting sequence, arranging all cameras in a closed loop to calculate their positions and orientations.\nInterestingly, visualizing the global features of a cylinder also revealed a highly ordered 2D plan layout, aiding significantly in recovering the image capture sequence. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15672041%2F8625f3098ae89ad3ebfb6f7202344376%2F2024-06-06%2020.37.07.png?generation=1717677454710544&alt=media =700x280)\nThe solution for transparent scenes improved scores by approximately **0.06** eventually.\n| Solution       | Private     | Public  |\n| ------------- |:-------------:| -----:|\n| Without Transparent Solution  | 0.184 | 0.185 |\n| **With Transparent Solution**   | **0.242**  | **0.249** |\n\n## 3. Other Ideas\n### 3.1 Methods that did not work\n- **Enhanced First Phase of MST-Aided SfM**: By adding more redundant edges (e.g., top-k, alternative key paths) to MST, aiming to balance recall and precision in the first reconstruction phase. While private score reached a high of **0.263**, the public score was only **0.249**, so this was not selected.\n- **Lightglue Matcher for Dedode**: We tested Kornia's lightglue matcher for dedode (b/g), but it performed worse than dual softmax.\n- **Day-Night Challenge**: To tackle day-night challenges in pond and lizard datasets, we attempted brightness-based day/night classification, normalized feature distribution, but encountered unknown \"Throw Exception\" obstacles.\n- **Dense Matchers**: We tested dense matchers like LoFTR, EfficientLoFTR, and DKMV3 in conventional scenes but saw no score improvement, likely due to our unoptimized parameters (matching threshold, grid size, etc.).\n- **Neural Network-Based SfM**: We explored neural network-based SfM methods like Dust3R for transparent scenes, but results were not pretty well.\n- **Dense Optical Flow**: Using dense optical flow to restore the image capture sequence in transparent scenes also worked but was less effective than our proposed global feature.\n\n### 3.2 Methods that we haven't tried\n- **TTA**: We didn't try TTA methods due to runtime considerations.\n- **Multi-Scale Matching**: We didn't detect covisible regions or perform multi-scale local feature matching.\n﻿",
    "2860339": "Thanks for sharing",
    "2862990": "Thank you for sharing the solution.\nIt was a very interesting solution.\n\nI have a question about the Global Features.\nWhen it says \"establishing a one-to-one correspondence based on their spatial relationships,\" does this mean to concatenate the feature of a keypoint and the feature of the patch where the keypoint is located?\n\nFor example, in the attached image, is the final feature corresponding to keypoint K6 a concatenation of the feature of keypoint K6 (feat_K6) and the feature of patch P24 (feat_P24) where keypoint K6 is located?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3823496%2F4e145c21dd6f25d6b4dd8ac16ea2981a%2F2nd_place_question.jpg?generation=1717916605082142&alt=media)",
    "2865834": "Congratulations with the 1st place!\nThat's right! Exactly. \nExtract the patch feature where each point is located and concatenate it to the point feature. \nWe believe this is a direct way to enhance the semantics of the point, giving it better discriminative power for scene semantics, thereby better serving downstream tasks like constructing MST and restoring the order of transparent objects.",
    "2866658": "Thank you for your response.\nI think it's a simple yet revolutionary and excellent global feature that performs well in downstream tasks for all scenes, including transparent ones. I'm very impressed.\n\nAlso, congratulations again on your 2nd place finish!",
    "2872144": "Thank you for sharing the solution 👏",
    "2927238": "Congratulation! Thanks for sharing the great solution especially the MST-Aided Coarse-to-Fine SfM pipeline.\n\nBut I still have some questions:\n1. what is the relationship between your similarity graph and MST? \n2. how you construct the similarity graph and MST? \n\nMy understanding is that the edge with the largest weight in the similarity graph is selected to construct the MST. In other words, the MST is the shortest path that connects the vertices in the similarity graph with the largest sum of weights (similarity)?\n\nLooking forward to your answer, the more detailed the better, thank you very much.",
    "3000069": "Your understanding is roughly correct, but MST is a tree like structure. For our solution, MST is the spanning tree for the similarity graph with minimal cost (highest similarity)."
  },
  "source": "meta"
}