{
  "id": 511291,
  "title": "6th Place Solution: Detector-Free SfM & Transparent Scene Trick",
  "url": "/competitions/image-matching-challenge-2024/discussion/511291",
  "author_name": "Hao Yu",
  "post_date": "2024-06-10T06:57:11.613000",
  "votes": 17,
  "comment_count": 0,
  "views": 0,
  "content": "<h1>1. Intro</h1>\n<p>We are delighted to participate in the Image Matching Challenge 2024. We would like to express our gratitude to the organizers, sponsors, and the staff of Kaggle for their efforts in making this competition possible. We also thank all the participants for their valuable suggestions and assistance.<br>\nOur team consists of Hao Yu, Xingyi He, Dongli Tan, Sida Peng, and Xiaowei Zhou. We are affiliated with the State Key Laboratory of CAD&amp;CG at Zhejiang University. I’m very grateful for the hard work and dedication of our teammates in this competition.</p>\n<h1>2. Overview</h1>\n<p>Our final solution involves using <a href=\"https://github.com/zju3dv/DetectorFreeSfM\" target=\"_blank\">Detector-free Structure from Motion (DFSfM)</a> for general scenes and an image order recovery strategy for transparent scenes. For general scenes, DFSfM continues the winning strategy from <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417407\" target=\"_blank\">our IMC 2023 solution</a>, and for transparent scenes, we identified the patterns in the camera trajectories for transparent scenes in IMC 2024. By uniformly sampling the camera center positions on a circular camera trajectory and then recovering the order of each image in relation to the camera's position on the trajectory, we were able to estimate the 3D poses of the images.</p>\n<h1>3. Method</h1>\n<h2>3.1 Pipeline for general scenes</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17083104%2F7741132428fce484496fd4a511cd54f7%2Fpipeline.png?generation=1718000339918824&amp;alt=media\"><br>\nSimilar to IMC 2023, we utilized the DFSfM, a coarse-to-fine Structure from Motion (SfM) framework. <br>\nThe original intention behind the design of DFSfM was to address the multi-view inconsistency issues caused by the detector-free matcher LoFTR. We found that dense matchers like DKM and RoMa still have this issue and that they generate a large number of dense 2D points, significantly increasing the computational load. <br>\nTherefore, it is necessary to merge matches for each view using a confidence-guided merging method, sacrificing some matching accuracy to improve consistency, and then refine the tracks after SfM. We used the merged matches to reconstruct a coarse SfM model. <br>\nSubsequently, we refined the rough SfM model through a novel iterative refinement pipeline that iterates between an attention-based multi-view matching module and a geometric refinement module to enhance reconstruction accuracy. <br>\nDue to the time constraints of the competition, we also employed a \"lightweight\" sparse feature detection and matching method to determine the image rotation and the final overlap areas between pairs of images, where the dense matcher (DKM, RoMa) will be executed.</p>\n<h3>3.1.1 Construct Image Pairs</h3>\n<p>Pairs from Retrieval (NetVLAD): retrieval involves using an image retrieval method to select k relevant images for each image. Here, we did not observe significant differences between different retrieval methods. </p>\n<h3>3.1.2 Matching</h3>\n<h4>3.1.2.1 Roation Detection (See 2.2.1 in <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417407\" target=\"_blank\">our IMC 2023 solution</a>)</h4>\n<h4>3.1.2.2 Overlap Detection (See 2.2.2 in <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417407\" target=\"_blank\">our IMC 2023 solution</a>)</h4>\n<h4>3.1.2.3 Sparse Matching + Dense Matching</h4>\n<p>For scenes without highr resolution images, we used Superpoint + Superglue for sparse matching, and RoMa for dense matching instead. Due to the original RoMa being too time-consuming, we replaced the feature extraction backbone of RoMa with vit-b and retrained a model. The results showed that our retrained RoMa could achieve a speed close to DKMv3, and the local evaluation indicated a significant improvement in performance compared to DKMv3.<br>\nFor scenes with high-resolution images, we found that both the original RoMa and our retrained RoMa did not perform well, so we adopted DKMv3 as the dense matching model.</p>\n<h3>3.1.3 Multi-view inconsistency problem for dense matching(See 2.3 in <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417407\" target=\"_blank\">our IMC 2023 solution</a>）</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17083104%2Fa1d2d7df1db86361357a5d302c90e796%2Fimc24_non_repeatable_problem.jpg?generation=1718010371154180&amp;alt=media\"></p>\n<p>Last year, when we used the semi-dense matching model LoFTR, the coarse-to-fine matching process of LoFTR resulted in the simultaneous existence of grid-level points and pixel-level points, which caused a multi-view inconsistency problem. This year, we adopted dense matching models such as DKM and RoMa. Although the matches generated are all at the pixel-level, since DKM and RoMa sample matches from flow, they also have the multi-view inconsistency problem. Therefore, we also addressed this issue through a coarse-to-fine architecture.</p>\n<h3>3.1.4 Coarse SfM</h3>\n<h4>3.1.4.1 Confidence-guided Merge(See 2.4.1 in our <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417407\" target=\"_blank\">IMC 2023 solution</a>）</h4>\n<h4>3.1.4.2 Mapping twice</h4>\n<p>Based on the merged matches, we perform the coarse Structure from Motion (SfM) using COLMAP. We drew upon the experience from IMC 2023 (thanks to <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417191\" target=\"_blank\">3rd Place Solution - Significantly Reduced the Fluctuations caused by Randomness!</a> ) and found that after the initial COLMAP reconstruction, by relaxing the parameters of COLMAP and running it again, and then selecting the model with a greater number of registered images, this brought us  improvement.</p>\n<h3>3.1.5 Iterative Refinement (See 2.4.1 in our <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417407\" target=\"_blank\">IMC 2023 solution</a>）</h3>\n<h3>3.1.6 Results</h3>\n<p>Here are our results for LB and local validation score.</p>\n<table>\n<thead>\n<tr>\n<th><strong>Method</strong></th>\n<th><strong>Private LB</strong></th>\n<th><strong>Public LB</strong></th>\n<th><strong>Val (avg.)</strong></th>\n<th><strong>pond</strong></th>\n<th><strong>church</strong></th>\n<th><strong>dioscuri</strong></th>\n<th><strong>lizard</strong></th>\n<th><strong>multi-temple</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>DFSfM(SuperPoint + SuperGlue + DKMv3)</td>\n<td>0.164</td>\n<td>0.171</td>\n<td>0.399</td>\n<td>0.394</td>\n<td>0.195</td>\n<td>0.480</td>\n<td>0.495</td>\n<td>0.428</td>\n</tr>\n<tr>\n<td>DFSfM(SuperPoint + SuperGlue +Our Retrained RoMa)</td>\n<td><strong>0.167</strong></td>\n<td><strong>0.175</strong></td>\n<td><strong>0.470</strong></td>\n<td>0.487</td>\n<td>0.190</td>\n<td>0.542</td>\n<td>0.743</td>\n<td>0.390</td>\n</tr>\n</tbody>\n</table>\n<p>Due to the large number of images in \"pond\" and \"lizard\", we sampled around 100 images for these two scenes.</p>\n<h2><strong>3.2 Pipeline for transparent scenes</strong></h2>\n<p>We made extensive efforts and found that matching combined with COLMAP completely fails for transparent scenes. By observing the characteristics of the camera center trajectory in local transparent scenes, and according to the IMC 2024 evaluation metrics (the trajectory of the camera center, rather than specific rotation and translation), we designed a unique processing strategy for transparent scenes: generating a circular camera trajectory and recovering the order of the images.</p>\n<h3><strong>3.2.1 Validate the idea on local transparent scene</strong></h3>\n<p>The images provided locally include the order of each image within the scene in their names, so we generated a circular trajectory and uniformly sampled the camera center coordinates according to the image order. By doing this, we found that mAA for the cylinder and cup in our local evaluation reached over 0.9.</p>\n<table>\n<thead>\n<tr>\n<th><strong>scene</strong></th>\n<th><strong>score</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>cylinder</td>\n<td>0.919</td>\n</tr>\n<tr>\n<td>cup</td>\n<td>0.995</td>\n</tr>\n</tbody>\n</table>\n<h3><strong>3.2.2 How can we distinguish these transparent scenes on Kaggle ？</strong></h3>\n<p>Based on the characteristics of the content in transparent scene images to segment (foreground and background). We used the <a href=\"https://github.com/YangtaoWANG95/TokenCut\" target=\"_blank\">tokencut</a> segmentation model to perform foreground segmentation on each image. If the foreground area segmented from all images is roughly consistent, it indicates that the camera trajectory for this scene is approximately circular, and we mark this scene as a transparent scene.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17083104%2F7bb37389e98f83cd45c21eb30cfa9152%2Fsegment.png?generation=1718000535898820&amp;alt=media\"></p>\n<h3><strong>3.2.3  How can we recover the image order on Kaggle?</strong></h3>\n<p>After distinguishing the transparent scenes, how to restore the image order on Kaggle is a challenge because the image names in the Kaggle dataset are garbled, so we can't directly sample the camera center of each image on the generated circular trajectory as we did locally with known order. Therefore, a specialized method is needed to restore the image order and place each image in its correct position within the scene.</p>\n<p>We designed a strategy based on the <em>Image Similarity Matrix + TSP</em> algorithm to restore the image order. For instance, considering a scenario with images labeled 1, 2, and 3, which are situated on a circular path, the objective of the optimization task is to maximize the similarity among the image pairs: 1 and 2, 2 and 3, as well as 3 and 1. Thus, this problem can be formulated as a Traveling Salesman Problem (TSP), which is about finding the shortest possible route that visits a set of cities and returns to the origin. (The <em>Image Distance Matrix</em> is equal to <em>1 - Image Similarity Matrix</em> ).</p>\n<p>After constructing the optimization problem, the key step is to estimate a similarity matrix that is closest to the ground-truth. We tried two methods.</p>\n<p>(1)We used a retrieval method (NetVLAD) to calculate the global feature similarity of images and build the similarity matrix.</p>\n<p>(2)We employed a Sparse Matching + RANSAC approach, using the number of matches generated between two images to represent their similarity.</p>\n<p>We validated the recovery effect locally and also ran tests on Kaggle based on previous best approach (DFSfM with our retrained RoMa model)<br>\nLocal Transparent Scenes:</p>\n<table>\n<thead>\n<tr>\n<th>Method\\Scene</th>\n<th>cup</th>\n<th>cylinder</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Random Recovery</td>\n<td>0.081</td>\n<td>0.066</td>\n</tr>\n<tr>\n<td>NetVLAD</td>\n<td><strong>0.101</strong></td>\n<td>0.242</td>\n</tr>\n<tr>\n<td>SP + SG</td>\n<td>0.020</td>\n<td>0.112</td>\n</tr>\n<tr>\n<td>SIFT + NN</td>\n<td>0.030</td>\n<td>0.606</td>\n</tr>\n<tr>\n<td>ALIKED + NN</td>\n<td>0.056</td>\n<td><strong>0.919</strong></td>\n</tr>\n</tbody>\n</table>\n<p>Kaggle Submissions:</p>\n<table>\n<thead>\n<tr>\n<th>Method</th>\n<th>Private LB</th>\n<th>Public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>DFSfM(RoMa)</td>\n<td>0.167</td>\n<td>0.175</td>\n</tr>\n<tr>\n<td>DFSfM(RoMa) + NetVLAD</td>\n<td>0.165</td>\n<td>0.171</td>\n</tr>\n<tr>\n<td>DFSfM(RoMa) + SP + SG</td>\n<td>0.189</td>\n<td>0.199</td>\n</tr>\n<tr>\n<td>DFSfM(RoMa) + SIFT + NN</td>\n<td><strong>0.201</strong></td>\n<td>0.207</td>\n</tr>\n<tr>\n<td>DFSfM(RoMa)+ ALIKED + NN</td>\n<td>0.191</td>\n<td><strong>0.225</strong></td>\n</tr>\n</tbody>\n</table>\n<p>In the end, we chose the DFSfM (RoMa + ALIKED + NN) approach, which achieved the highest score of <strong>0.225</strong> on the Public LB, with a Private LB score of <strong>0.191</strong>. Regrettably, although the DFSfM (RoMa + SIFT + NN) approach only scored <strong>0.207</strong> on the public LB, it reached the highest score of <strong>0.201</strong> on the Private LB.</p>\n<h1><strong>4. Ideas tried but not worked</strong></h1>\n<h2><strong>4.1 For general scenes</strong></h2>\n<h3><strong>4.1.1 Other image retrieval methods</strong></h3>\n<ul>\n<li>Anyloc</li>\n<li>Eigenplaces</li>\n<li>Rotation detection before image retrieval</li>\n</ul>\n<p>We attempted to use the latest retrieval methods such as Anyloc and Eigenplaces, and found that the results did not show significant differences. At the same time, we also tried to add image rotation detection before retrieval, and found that the results did not improve.</p>\n<h3><strong>4.1.1 Other image matching methods</strong></h3>\n<ul>\n<li>DeDoDe + LG</li>\n<li>DISK + LG</li>\n<li>ALIKED + LG</li>\n</ul>\n<p>We also tried some other sparse matching methods, such as ALIKED + LightGlue, DISK + LightGlue, DeDoDe + LightGlue, and found that the results did not significantly improve. Both the local and public leaderboard results were slightly worse than those of SuperPoint + SuperGlue. We believe the possible reason might be that we were unable to adjust the parameters correctly.</p>\n<h2><strong>4.2 For transparent scenes</strong></h2>\n<ul>\n<li>DUSt3R</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17083104%2F17e9beac384fcdede5a7e4ff4026862e%2Fdustr.png?generation=1718000589088088&amp;alt=media\"></p>\n<p>We attempted to handle the special case of transparent scenes using DUSt3R and found that, like other matching methods, it was unable to effectively process transparent scenes.</p>\n<h1><strong>5. Acknowledgment:</strong></h1>\n<p>Once again, I would like to express my gratitude to the organizers for their contributions to this competition, and I appreciate the hard work and dedication of my teammates and all the participants.</p>",
  "messages": [
    {
      "id": 2864482,
      "postDate": "2024-06-10T06:57:11.613Z",
      "content": "<h1>1. Intro</h1>\n<p>We are delighted to participate in the Image Matching Challenge 2024. We would like to express our gratitude to the organizers, sponsors, and the staff of Kaggle for their efforts in making this competition possible. We also thank all the participants for their valuable suggestions and assistance.<br>\nOur team consists of Hao Yu, Xingyi He, Dongli Tan, Sida Peng, and Xiaowei Zhou. We are affiliated with the State Key Laboratory of CAD&amp;CG at Zhejiang University. I’m very grateful for the hard work and dedication of our teammates in this competition.</p>\n<h1>2. Overview</h1>\n<p>Our final solution involves using <a href=\"https://github.com/zju3dv/DetectorFreeSfM\" target=\"_blank\">Detector-free Structure from Motion (DFSfM)</a> for general scenes and an image order recovery strategy for transparent scenes. For general scenes, DFSfM continues the winning strategy from <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417407\" target=\"_blank\">our IMC 2023 solution</a>, and for transparent scenes, we identified the patterns in the camera trajectories for transparent scenes in IMC 2024. By uniformly sampling the camera center positions on a circular camera trajectory and then recovering the order of each image in relation to the camera's position on the trajectory, we were able to estimate the 3D poses of the images.</p>\n<h1>3. Method</h1>\n<h2>3.1 Pipeline for general scenes</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17083104%2F7741132428fce484496fd4a511cd54f7%2Fpipeline.png?generation=1718000339918824&amp;alt=media\"><br>\nSimilar to IMC 2023, we utilized the DFSfM, a coarse-to-fine Structure from Motion (SfM) framework. <br>\nThe original intention behind the design of DFSfM was to address the multi-view inconsistency issues caused by the detector-free matcher LoFTR. We found that dense matchers like DKM and RoMa still have this issue and that they generate a large number of dense 2D points, significantly increasing the computational load. <br>\nTherefore, it is necessary to merge matches for each view using a confidence-guided merging method, sacrificing some matching accuracy to improve consistency, and then refine the tracks after SfM. We used the merged matches to reconstruct a coarse SfM model. <br>\nSubsequently, we refined the rough SfM model through a novel iterative refinement pipeline that iterates between an attention-based multi-view matching module and a geometric refinement module to enhance reconstruction accuracy. <br>\nDue to the time constraints of the competition, we also employed a \"lightweight\" sparse feature detection and matching method to determine the image rotation and the final overlap areas between pairs of images, where the dense matcher (DKM, RoMa) will be executed.</p>\n<h3>3.1.1 Construct Image Pairs</h3>\n<p>Pairs from Retrieval (NetVLAD): retrieval involves using an image retrieval method to select k relevant images for each image. Here, we did not observe significant differences between different retrieval methods. </p>\n<h3>3.1.2 Matching</h3>\n<h4>3.1.2.1 Roation Detection (See 2.2.1 in <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417407\" target=\"_blank\">our IMC 2023 solution</a>)</h4>\n<h4>3.1.2.2 Overlap Detection (See 2.2.2 in <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417407\" target=\"_blank\">our IMC 2023 solution</a>)</h4>\n<h4>3.1.2.3 Sparse Matching + Dense Matching</h4>\n<p>For scenes without highr resolution images, we used Superpoint + Superglue for sparse matching, and RoMa for dense matching instead. Due to the original RoMa being too time-consuming, we replaced the feature extraction backbone of RoMa with vit-b and retrained a model. The results showed that our retrained RoMa could achieve a speed close to DKMv3, and the local evaluation indicated a significant improvement in performance compared to DKMv3.<br>\nFor scenes with high-resolution images, we found that both the original RoMa and our retrained RoMa did not perform well, so we adopted DKMv3 as the dense matching model.</p>\n<h3>3.1.3 Multi-view inconsistency problem for dense matching(See 2.3 in <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417407\" target=\"_blank\">our IMC 2023 solution</a>）</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17083104%2Fa1d2d7df1db86361357a5d302c90e796%2Fimc24_non_repeatable_problem.jpg?generation=1718010371154180&amp;alt=media\"></p>\n<p>Last year, when we used the semi-dense matching model LoFTR, the coarse-to-fine matching process of LoFTR resulted in the simultaneous existence of grid-level points and pixel-level points, which caused a multi-view inconsistency problem. This year, we adopted dense matching models such as DKM and RoMa. Although the matches generated are all at the pixel-level, since DKM and RoMa sample matches from flow, they also have the multi-view inconsistency problem. Therefore, we also addressed this issue through a coarse-to-fine architecture.</p>\n<h3>3.1.4 Coarse SfM</h3>\n<h4>3.1.4.1 Confidence-guided Merge(See 2.4.1 in our <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417407\" target=\"_blank\">IMC 2023 solution</a>）</h4>\n<h4>3.1.4.2 Mapping twice</h4>\n<p>Based on the merged matches, we perform the coarse Structure from Motion (SfM) using COLMAP. We drew upon the experience from IMC 2023 (thanks to <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417191\" target=\"_blank\">3rd Place Solution - Significantly Reduced the Fluctuations caused by Randomness!</a> ) and found that after the initial COLMAP reconstruction, by relaxing the parameters of COLMAP and running it again, and then selecting the model with a greater number of registered images, this brought us  improvement.</p>\n<h3>3.1.5 Iterative Refinement (See 2.4.1 in our <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417407\" target=\"_blank\">IMC 2023 solution</a>）</h3>\n<h3>3.1.6 Results</h3>\n<p>Here are our results for LB and local validation score.</p>\n<table>\n<thead>\n<tr>\n<th><strong>Method</strong></th>\n<th><strong>Private LB</strong></th>\n<th><strong>Public LB</strong></th>\n<th><strong>Val (avg.)</strong></th>\n<th><strong>pond</strong></th>\n<th><strong>church</strong></th>\n<th><strong>dioscuri</strong></th>\n<th><strong>lizard</strong></th>\n<th><strong>multi-temple</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>DFSfM(SuperPoint + SuperGlue + DKMv3)</td>\n<td>0.164</td>\n<td>0.171</td>\n<td>0.399</td>\n<td>0.394</td>\n<td>0.195</td>\n<td>0.480</td>\n<td>0.495</td>\n<td>0.428</td>\n</tr>\n<tr>\n<td>DFSfM(SuperPoint + SuperGlue +Our Retrained RoMa)</td>\n<td><strong>0.167</strong></td>\n<td><strong>0.175</strong></td>\n<td><strong>0.470</strong></td>\n<td>0.487</td>\n<td>0.190</td>\n<td>0.542</td>\n<td>0.743</td>\n<td>0.390</td>\n</tr>\n</tbody>\n</table>\n<p>Due to the large number of images in \"pond\" and \"lizard\", we sampled around 100 images for these two scenes.</p>\n<h2><strong>3.2 Pipeline for transparent scenes</strong></h2>\n<p>We made extensive efforts and found that matching combined with COLMAP completely fails for transparent scenes. By observing the characteristics of the camera center trajectory in local transparent scenes, and according to the IMC 2024 evaluation metrics (the trajectory of the camera center, rather than specific rotation and translation), we designed a unique processing strategy for transparent scenes: generating a circular camera trajectory and recovering the order of the images.</p>\n<h3><strong>3.2.1 Validate the idea on local transparent scene</strong></h3>\n<p>The images provided locally include the order of each image within the scene in their names, so we generated a circular trajectory and uniformly sampled the camera center coordinates according to the image order. By doing this, we found that mAA for the cylinder and cup in our local evaluation reached over 0.9.</p>\n<table>\n<thead>\n<tr>\n<th><strong>scene</strong></th>\n<th><strong>score</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>cylinder</td>\n<td>0.919</td>\n</tr>\n<tr>\n<td>cup</td>\n<td>0.995</td>\n</tr>\n</tbody>\n</table>\n<h3><strong>3.2.2 How can we distinguish these transparent scenes on Kaggle ？</strong></h3>\n<p>Based on the characteristics of the content in transparent scene images to segment (foreground and background). We used the <a href=\"https://github.com/YangtaoWANG95/TokenCut\" target=\"_blank\">tokencut</a> segmentation model to perform foreground segmentation on each image. If the foreground area segmented from all images is roughly consistent, it indicates that the camera trajectory for this scene is approximately circular, and we mark this scene as a transparent scene.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17083104%2F7bb37389e98f83cd45c21eb30cfa9152%2Fsegment.png?generation=1718000535898820&amp;alt=media\"></p>\n<h3><strong>3.2.3  How can we recover the image order on Kaggle?</strong></h3>\n<p>After distinguishing the transparent scenes, how to restore the image order on Kaggle is a challenge because the image names in the Kaggle dataset are garbled, so we can't directly sample the camera center of each image on the generated circular trajectory as we did locally with known order. Therefore, a specialized method is needed to restore the image order and place each image in its correct position within the scene.</p>\n<p>We designed a strategy based on the <em>Image Similarity Matrix + TSP</em> algorithm to restore the image order. For instance, considering a scenario with images labeled 1, 2, and 3, which are situated on a circular path, the objective of the optimization task is to maximize the similarity among the image pairs: 1 and 2, 2 and 3, as well as 3 and 1. Thus, this problem can be formulated as a Traveling Salesman Problem (TSP), which is about finding the shortest possible route that visits a set of cities and returns to the origin. (The <em>Image Distance Matrix</em> is equal to <em>1 - Image Similarity Matrix</em> ).</p>\n<p>After constructing the optimization problem, the key step is to estimate a similarity matrix that is closest to the ground-truth. We tried two methods.</p>\n<p>(1)We used a retrieval method (NetVLAD) to calculate the global feature similarity of images and build the similarity matrix.</p>\n<p>(2)We employed a Sparse Matching + RANSAC approach, using the number of matches generated between two images to represent their similarity.</p>\n<p>We validated the recovery effect locally and also ran tests on Kaggle based on previous best approach (DFSfM with our retrained RoMa model)<br>\nLocal Transparent Scenes:</p>\n<table>\n<thead>\n<tr>\n<th>Method\\Scene</th>\n<th>cup</th>\n<th>cylinder</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Random Recovery</td>\n<td>0.081</td>\n<td>0.066</td>\n</tr>\n<tr>\n<td>NetVLAD</td>\n<td><strong>0.101</strong></td>\n<td>0.242</td>\n</tr>\n<tr>\n<td>SP + SG</td>\n<td>0.020</td>\n<td>0.112</td>\n</tr>\n<tr>\n<td>SIFT + NN</td>\n<td>0.030</td>\n<td>0.606</td>\n</tr>\n<tr>\n<td>ALIKED + NN</td>\n<td>0.056</td>\n<td><strong>0.919</strong></td>\n</tr>\n</tbody>\n</table>\n<p>Kaggle Submissions:</p>\n<table>\n<thead>\n<tr>\n<th>Method</th>\n<th>Private LB</th>\n<th>Public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>DFSfM(RoMa)</td>\n<td>0.167</td>\n<td>0.175</td>\n</tr>\n<tr>\n<td>DFSfM(RoMa) + NetVLAD</td>\n<td>0.165</td>\n<td>0.171</td>\n</tr>\n<tr>\n<td>DFSfM(RoMa) + SP + SG</td>\n<td>0.189</td>\n<td>0.199</td>\n</tr>\n<tr>\n<td>DFSfM(RoMa) + SIFT + NN</td>\n<td><strong>0.201</strong></td>\n<td>0.207</td>\n</tr>\n<tr>\n<td>DFSfM(RoMa)+ ALIKED + NN</td>\n<td>0.191</td>\n<td><strong>0.225</strong></td>\n</tr>\n</tbody>\n</table>\n<p>In the end, we chose the DFSfM (RoMa + ALIKED + NN) approach, which achieved the highest score of <strong>0.225</strong> on the Public LB, with a Private LB score of <strong>0.191</strong>. Regrettably, although the DFSfM (RoMa + SIFT + NN) approach only scored <strong>0.207</strong> on the public LB, it reached the highest score of <strong>0.201</strong> on the Private LB.</p>\n<h1><strong>4. Ideas tried but not worked</strong></h1>\n<h2><strong>4.1 For general scenes</strong></h2>\n<h3><strong>4.1.1 Other image retrieval methods</strong></h3>\n<ul>\n<li>Anyloc</li>\n<li>Eigenplaces</li>\n<li>Rotation detection before image retrieval</li>\n</ul>\n<p>We attempted to use the latest retrieval methods such as Anyloc and Eigenplaces, and found that the results did not show significant differences. At the same time, we also tried to add image rotation detection before retrieval, and found that the results did not improve.</p>\n<h3><strong>4.1.1 Other image matching methods</strong></h3>\n<ul>\n<li>DeDoDe + LG</li>\n<li>DISK + LG</li>\n<li>ALIKED + LG</li>\n</ul>\n<p>We also tried some other sparse matching methods, such as ALIKED + LightGlue, DISK + LightGlue, DeDoDe + LightGlue, and found that the results did not significantly improve. Both the local and public leaderboard results were slightly worse than those of SuperPoint + SuperGlue. We believe the possible reason might be that we were unable to adjust the parameters correctly.</p>\n<h2><strong>4.2 For transparent scenes</strong></h2>\n<ul>\n<li>DUSt3R</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17083104%2F17e9beac384fcdede5a7e4ff4026862e%2Fdustr.png?generation=1718000589088088&amp;alt=media\"></p>\n<p>We attempted to handle the special case of transparent scenes using DUSt3R and found that, like other matching methods, it was unable to effectively process transparent scenes.</p>\n<h1><strong>5. Acknowledgment:</strong></h1>\n<p>Once again, I would like to express my gratitude to the organizers for their contributions to this competition, and I appreciate the hard work and dedication of my teammates and all the participants.</p>",
      "rawMarkdown": "# 1. Intro\nWe are delighted to participate in the Image Matching Challenge 2024. We would like to express our gratitude to the organizers, sponsors, and the staff of Kaggle for their efforts in making this competition possible. We also thank all the participants for their valuable suggestions and assistance.\nOur team consists of Hao Yu, Xingyi He, Dongli Tan, Sida Peng, and Xiaowei Zhou. We are affiliated with the State Key Laboratory of CAD&CG at Zhejiang University. I’m very grateful for the hard work and dedication of our teammates in this competition.\n# 2. Overview\nOur final solution involves using [Detector-free Structure from Motion (DFSfM)](https://github.com/zju3dv/DetectorFreeSfM) for general scenes and an image order recovery strategy for transparent scenes. For general scenes, DFSfM continues the winning strategy from [our IMC 2023 solution](https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417407), and for transparent scenes, we identified the patterns in the camera trajectories for transparent scenes in IMC 2024. By uniformly sampling the camera center positions on a circular camera trajectory and then recovering the order of each image in relation to the camera's position on the trajectory, we were able to estimate the 3D poses of the images.\n# 3. Method\n## 3.1 Pipeline for general scenes\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17083104%2F7741132428fce484496fd4a511cd54f7%2Fpipeline.png?generation=1718000339918824&alt=media)\nSimilar to IMC 2023, we utilized the DFSfM, a coarse-to-fine Structure from Motion (SfM) framework. \nThe original intention behind the design of DFSfM was to address the multi-view inconsistency issues caused by the detector-free matcher LoFTR. We found that dense matchers like DKM and RoMa still have this issue and that they generate a large number of dense 2D points, significantly increasing the computational load. \nTherefore, it is necessary to merge matches for each view using a confidence-guided merging method, sacrificing some matching accuracy to improve consistency, and then refine the tracks after SfM. We used the merged matches to reconstruct a coarse SfM model. \nSubsequently, we refined the rough SfM model through a novel iterative refinement pipeline that iterates between an attention-based multi-view matching module and a geometric refinement module to enhance reconstruction accuracy. \nDue to the time constraints of the competition, we also employed a \"lightweight\" sparse feature detection and matching method to determine the image rotation and the final overlap areas between pairs of images, where the dense matcher (DKM, RoMa) will be executed.\n### 3.1.1 Construct Image Pairs\nPairs from Retrieval (NetVLAD): retrieval involves using an image retrieval method to select k relevant images for each image. Here, we did not observe significant differences between different retrieval methods. \n### 3.1.2 Matching\n#### 3.1.2.1 Roation Detection (See 2.2.1 in [our IMC 2023 solution](https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417407))\n#### 3.1.2.2 Overlap Detection (See 2.2.2 in [our IMC 2023 solution](https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417407))\n#### 3.1.2.3 Sparse Matching + Dense Matching\nFor scenes without highr resolution images, we used Superpoint + Superglue for sparse matching, and RoMa for dense matching instead. Due to the original RoMa being too time-consuming, we replaced the feature extraction backbone of RoMa with vit-b and retrained a model. The results showed that our retrained RoMa could achieve a speed close to DKMv3, and the local evaluation indicated a significant improvement in performance compared to DKMv3.\nFor scenes with high-resolution images, we found that both the original RoMa and our retrained RoMa did not perform well, so we adopted DKMv3 as the dense matching model.\n### 3.1.3 Multi-view inconsistency problem for dense matching(See 2.3 in [our IMC 2023 solution](https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417407)）\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17083104%2Fa1d2d7df1db86361357a5d302c90e796%2Fimc24_non_repeatable_problem.jpg?generation=1718010371154180&alt=media\" style=\"zoom: 15%;\" />\n\nLast year, when we used the semi-dense matching model LoFTR, the coarse-to-fine matching process of LoFTR resulted in the simultaneous existence of grid-level points and pixel-level points, which caused a multi-view inconsistency problem. This year, we adopted dense matching models such as DKM and RoMa. Although the matches generated are all at the pixel-level, since DKM and RoMa sample matches from flow, they also have the multi-view inconsistency problem. Therefore, we also addressed this issue through a coarse-to-fine architecture.\n\n### 3.1.4 Coarse SfM\n\n#### 3.1.4.1 Confidence-guided Merge(See 2.4.1 in our [IMC 2023 solution](https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417407)）\n\n#### 3.1.4.2 Mapping twice\n\nBased on the merged matches, we perform the coarse Structure from Motion (SfM) using COLMAP. We drew upon the experience from IMC 2023 (thanks to [3rd Place Solution - Significantly Reduced the Fluctuations caused by Randomness!](https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417191) ) and found that after the initial COLMAP reconstruction, by relaxing the parameters of COLMAP and running it again, and then selecting the model with a greater number of registered images, this brought us  improvement.\n\n### 3.1.5 Iterative Refinement (See 2.4.1 in our [IMC 2023 solution](https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417407)）\n\n### 3.1.6 Results\n\nHere are our results for LB and local validation score.\n\n| **Method**                                        | **Private LB** | **Public LB** | **Val (avg.)** | **pond** | **church** | **dioscuri** | **lizard** | **multi-temple** |\n| ------------------------------------------------- | -------------- | ------------- | -------------- | -------- | ---------- | ------------ | ---------- | :--------------- |\n| DFSfM(SuperPoint + SuperGlue + DKMv3)             | 0.164          | 0.171         | 0.399      | 0.394    | 0.195      | 0.480        | 0.495      | 0.428            |\n| DFSfM(SuperPoint + SuperGlue +Our Retrained RoMa) |  **0.167**         | **0.175**         | **0.470**      | 0.487    | 0.190      | 0.542        | 0.743      | 0.390            |\n\nDue to the large number of images in \"pond\" and \"lizard\", we sampled around 100 images for these two scenes.\n\n## **3.2 Pipeline for transparent scenes**\n\nWe made extensive efforts and found that matching combined with COLMAP completely fails for transparent scenes. By observing the characteristics of the camera center trajectory in local transparent scenes, and according to the IMC 2024 evaluation metrics (the trajectory of the camera center, rather than specific rotation and translation), we designed a unique processing strategy for transparent scenes: generating a circular camera trajectory and recovering the order of the images.\n\n### **3.2.1 Validate the idea on local transparent scene**\n\nThe images provided locally include the order of each image within the scene in their names, so we generated a circular trajectory and uniformly sampled the camera center coordinates according to the image order. By doing this, we found that mAA for the cylinder and cup in our local evaluation reached over 0.9.\n\n| **scene** | **score** |\n| --------- | --------- |\n| cylinder  | 0.919     |\n| cup       | 0.995     |\n\n### **3.2.2 How can we distinguish these transparent scenes on Kaggle ？**\n\nBased on the characteristics of the content in transparent scene images to segment (foreground and background). We used the [tokencut](https://github.com/YangtaoWANG95/TokenCut) segmentation model to perform foreground segmentation on each image. If the foreground area segmented from all images is roughly consistent, it indicates that the camera trajectory for this scene is approximately circular, and we mark this scene as a transparent scene.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17083104%2F7bb37389e98f83cd45c21eb30cfa9152%2Fsegment.png?generation=1718000535898820&alt=media)\n\n### **3.2.3  How can we recover the image order on Kaggle?**\n\nAfter distinguishing the transparent scenes, how to restore the image order on Kaggle is a challenge because the image names in the Kaggle dataset are garbled, so we can't directly sample the camera center of each image on the generated circular trajectory as we did locally with known order. Therefore, a specialized method is needed to restore the image order and place each image in its correct position within the scene.\n\nWe designed a strategy based on the *Image Similarity Matrix + TSP* algorithm to restore the image order. For instance, considering a scenario with images labeled 1, 2, and 3, which are situated on a circular path, the objective of the optimization task is to maximize the similarity among the image pairs: 1 and 2, 2 and 3, as well as 3 and 1. Thus, this problem can be formulated as a Traveling Salesman Problem (TSP), which is about finding the shortest possible route that visits a set of cities and returns to the origin. (The *Image Distance Matrix* is equal to *1 - Image Similarity Matrix* ).\n\nAfter constructing the optimization problem, the key step is to estimate a similarity matrix that is closest to the ground-truth. We tried two methods.\n\n(1)We used a retrieval method (NetVLAD) to calculate the global feature similarity of images and build the similarity matrix.\n\n(2)We employed a Sparse Matching + RANSAC approach, using the number of matches generated between two images to represent their similarity.\n\nWe validated the recovery effect locally and also ran tests on Kaggle based on previous best approach (DFSfM with our retrained RoMa model)\nLocal Transparent Scenes:\n| Method\\Scene    | cup   | cylinder |\n| --------------- | ----- | -------- |\n| Random Recovery | 0.081 | 0.066    |\n| NetVLAD         | **0.101** | 0.242    |\n| SP + SG         | 0.020 | 0.112    |\n| SIFT + NN       | 0.030 | 0.606    |\n| ALIKED + NN     | 0.056 | **0.919**    |\n\nKaggle Submissions:\n| Method                   | Private LB | Public LB |\n| ------------------------ | ---------- | --------- |\n| DFSfM(RoMa)              | 0.167      | 0.175     |\n| DFSfM(RoMa) + NetVLAD    | 0.165      | 0.171     |\n| DFSfM(RoMa) + SP + SG    | 0.189      | 0.199     |\n| DFSfM(RoMa) + SIFT + NN  | **0.201**      | 0.207     |\n| DFSfM(RoMa)+ ALIKED + NN | 0.191      | **0.225**     |\n\nIn the end, we chose the DFSfM (RoMa + ALIKED + NN) approach, which achieved the highest score of **0.225** on the Public LB, with a Private LB score of **0.191**. Regrettably, although the DFSfM (RoMa + SIFT + NN) approach only scored **0.207** on the public LB, it reached the highest score of **0.201** on the Private LB.\n\n# **4. Ideas tried but not worked**\n\n## **4.1 For general scenes**\n\n### **4.1.1 Other image retrieval methods**\n\n- Anyloc\n- Eigenplaces\n- Rotation detection before image retrieval\n\nWe attempted to use the latest retrieval methods such as Anyloc and Eigenplaces, and found that the results did not show significant differences. At the same time, we also tried to add image rotation detection before retrieval, and found that the results did not improve.\n\n### **4.1.1 Other image matching methods**\n\n- DeDoDe + LG\n- DISK + LG\n- ALIKED + LG\n\nWe also tried some other sparse matching methods, such as ALIKED + LightGlue, DISK + LightGlue, DeDoDe + LightGlue, and found that the results did not significantly improve. Both the local and public leaderboard results were slightly worse than those of SuperPoint + SuperGlue. We believe the possible reason might be that we were unable to adjust the parameters correctly.\n\n## **4.2 For transparent scenes**\n\n- DUSt3R\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17083104%2F17e9beac384fcdede5a7e4ff4026862e%2Fdustr.png?generation=1718000589088088&alt=media)\n\nWe attempted to handle the special case of transparent scenes using DUSt3R and found that, like other matching methods, it was unable to effectively process transparent scenes.\n\n# **5. Acknowledgment:**\n\nOnce again, I would like to express my gratitude to the organizers for their contributions to this competition, and I appreciate the hard work and dedication of my teammates and all the participants.",
      "votes": 17
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2864482": "# 1. Intro\nWe are delighted to participate in the Image Matching Challenge 2024. We would like to express our gratitude to the organizers, sponsors, and the staff of Kaggle for their efforts in making this competition possible. We also thank all the participants for their valuable suggestions and assistance.\nOur team consists of Hao Yu, Xingyi He, Dongli Tan, Sida Peng, and Xiaowei Zhou. We are affiliated with the State Key Laboratory of CAD&CG at Zhejiang University. I’m very grateful for the hard work and dedication of our teammates in this competition.\n# 2. Overview\nOur final solution involves using [Detector-free Structure from Motion (DFSfM)](https://github.com/zju3dv/DetectorFreeSfM) for general scenes and an image order recovery strategy for transparent scenes. For general scenes, DFSfM continues the winning strategy from [our IMC 2023 solution](https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417407), and for transparent scenes, we identified the patterns in the camera trajectories for transparent scenes in IMC 2024. By uniformly sampling the camera center positions on a circular camera trajectory and then recovering the order of each image in relation to the camera's position on the trajectory, we were able to estimate the 3D poses of the images.\n# 3. Method\n## 3.1 Pipeline for general scenes\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17083104%2F7741132428fce484496fd4a511cd54f7%2Fpipeline.png?generation=1718000339918824&alt=media)\nSimilar to IMC 2023, we utilized the DFSfM, a coarse-to-fine Structure from Motion (SfM) framework. \nThe original intention behind the design of DFSfM was to address the multi-view inconsistency issues caused by the detector-free matcher LoFTR. We found that dense matchers like DKM and RoMa still have this issue and that they generate a large number of dense 2D points, significantly increasing the computational load. \nTherefore, it is necessary to merge matches for each view using a confidence-guided merging method, sacrificing some matching accuracy to improve consistency, and then refine the tracks after SfM. We used the merged matches to reconstruct a coarse SfM model. \nSubsequently, we refined the rough SfM model through a novel iterative refinement pipeline that iterates between an attention-based multi-view matching module and a geometric refinement module to enhance reconstruction accuracy. \nDue to the time constraints of the competition, we also employed a \"lightweight\" sparse feature detection and matching method to determine the image rotation and the final overlap areas between pairs of images, where the dense matcher (DKM, RoMa) will be executed.\n### 3.1.1 Construct Image Pairs\nPairs from Retrieval (NetVLAD): retrieval involves using an image retrieval method to select k relevant images for each image. Here, we did not observe significant differences between different retrieval methods. \n### 3.1.2 Matching\n#### 3.1.2.1 Roation Detection (See 2.2.1 in [our IMC 2023 solution](https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417407))\n#### 3.1.2.2 Overlap Detection (See 2.2.2 in [our IMC 2023 solution](https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417407))\n#### 3.1.2.3 Sparse Matching + Dense Matching\nFor scenes without highr resolution images, we used Superpoint + Superglue for sparse matching, and RoMa for dense matching instead. Due to the original RoMa being too time-consuming, we replaced the feature extraction backbone of RoMa with vit-b and retrained a model. The results showed that our retrained RoMa could achieve a speed close to DKMv3, and the local evaluation indicated a significant improvement in performance compared to DKMv3.\nFor scenes with high-resolution images, we found that both the original RoMa and our retrained RoMa did not perform well, so we adopted DKMv3 as the dense matching model.\n### 3.1.3 Multi-view inconsistency problem for dense matching(See 2.3 in [our IMC 2023 solution](https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417407)）\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17083104%2Fa1d2d7df1db86361357a5d302c90e796%2Fimc24_non_repeatable_problem.jpg?generation=1718010371154180&alt=media\" style=\"zoom: 15%;\" />\n\nLast year, when we used the semi-dense matching model LoFTR, the coarse-to-fine matching process of LoFTR resulted in the simultaneous existence of grid-level points and pixel-level points, which caused a multi-view inconsistency problem. This year, we adopted dense matching models such as DKM and RoMa. Although the matches generated are all at the pixel-level, since DKM and RoMa sample matches from flow, they also have the multi-view inconsistency problem. Therefore, we also addressed this issue through a coarse-to-fine architecture.\n\n### 3.1.4 Coarse SfM\n\n#### 3.1.4.1 Confidence-guided Merge(See 2.4.1 in our [IMC 2023 solution](https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417407)）\n\n#### 3.1.4.2 Mapping twice\n\nBased on the merged matches, we perform the coarse Structure from Motion (SfM) using COLMAP. We drew upon the experience from IMC 2023 (thanks to [3rd Place Solution - Significantly Reduced the Fluctuations caused by Randomness!](https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417191) ) and found that after the initial COLMAP reconstruction, by relaxing the parameters of COLMAP and running it again, and then selecting the model with a greater number of registered images, this brought us  improvement.\n\n### 3.1.5 Iterative Refinement (See 2.4.1 in our [IMC 2023 solution](https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417407)）\n\n### 3.1.6 Results\n\nHere are our results for LB and local validation score.\n\n| **Method**                                        | **Private LB** | **Public LB** | **Val (avg.)** | **pond** | **church** | **dioscuri** | **lizard** | **multi-temple** |\n| ------------------------------------------------- | -------------- | ------------- | -------------- | -------- | ---------- | ------------ | ---------- | :--------------- |\n| DFSfM(SuperPoint + SuperGlue + DKMv3)             | 0.164          | 0.171         | 0.399      | 0.394    | 0.195      | 0.480        | 0.495      | 0.428            |\n| DFSfM(SuperPoint + SuperGlue +Our Retrained RoMa) |  **0.167**         | **0.175**         | **0.470**      | 0.487    | 0.190      | 0.542        | 0.743      | 0.390            |\n\nDue to the large number of images in \"pond\" and \"lizard\", we sampled around 100 images for these two scenes.\n\n## **3.2 Pipeline for transparent scenes**\n\nWe made extensive efforts and found that matching combined with COLMAP completely fails for transparent scenes. By observing the characteristics of the camera center trajectory in local transparent scenes, and according to the IMC 2024 evaluation metrics (the trajectory of the camera center, rather than specific rotation and translation), we designed a unique processing strategy for transparent scenes: generating a circular camera trajectory and recovering the order of the images.\n\n### **3.2.1 Validate the idea on local transparent scene**\n\nThe images provided locally include the order of each image within the scene in their names, so we generated a circular trajectory and uniformly sampled the camera center coordinates according to the image order. By doing this, we found that mAA for the cylinder and cup in our local evaluation reached over 0.9.\n\n| **scene** | **score** |\n| --------- | --------- |\n| cylinder  | 0.919     |\n| cup       | 0.995     |\n\n### **3.2.2 How can we distinguish these transparent scenes on Kaggle ？**\n\nBased on the characteristics of the content in transparent scene images to segment (foreground and background). We used the [tokencut](https://github.com/YangtaoWANG95/TokenCut) segmentation model to perform foreground segmentation on each image. If the foreground area segmented from all images is roughly consistent, it indicates that the camera trajectory for this scene is approximately circular, and we mark this scene as a transparent scene.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17083104%2F7bb37389e98f83cd45c21eb30cfa9152%2Fsegment.png?generation=1718000535898820&alt=media)\n\n### **3.2.3  How can we recover the image order on Kaggle?**\n\nAfter distinguishing the transparent scenes, how to restore the image order on Kaggle is a challenge because the image names in the Kaggle dataset are garbled, so we can't directly sample the camera center of each image on the generated circular trajectory as we did locally with known order. Therefore, a specialized method is needed to restore the image order and place each image in its correct position within the scene.\n\nWe designed a strategy based on the *Image Similarity Matrix + TSP* algorithm to restore the image order. For instance, considering a scenario with images labeled 1, 2, and 3, which are situated on a circular path, the objective of the optimization task is to maximize the similarity among the image pairs: 1 and 2, 2 and 3, as well as 3 and 1. Thus, this problem can be formulated as a Traveling Salesman Problem (TSP), which is about finding the shortest possible route that visits a set of cities and returns to the origin. (The *Image Distance Matrix* is equal to *1 - Image Similarity Matrix* ).\n\nAfter constructing the optimization problem, the key step is to estimate a similarity matrix that is closest to the ground-truth. We tried two methods.\n\n(1)We used a retrieval method (NetVLAD) to calculate the global feature similarity of images and build the similarity matrix.\n\n(2)We employed a Sparse Matching + RANSAC approach, using the number of matches generated between two images to represent their similarity.\n\nWe validated the recovery effect locally and also ran tests on Kaggle based on previous best approach (DFSfM with our retrained RoMa model)\nLocal Transparent Scenes:\n| Method\\Scene    | cup   | cylinder |\n| --------------- | ----- | -------- |\n| Random Recovery | 0.081 | 0.066    |\n| NetVLAD         | **0.101** | 0.242    |\n| SP + SG         | 0.020 | 0.112    |\n| SIFT + NN       | 0.030 | 0.606    |\n| ALIKED + NN     | 0.056 | **0.919**    |\n\nKaggle Submissions:\n| Method                   | Private LB | Public LB |\n| ------------------------ | ---------- | --------- |\n| DFSfM(RoMa)              | 0.167      | 0.175     |\n| DFSfM(RoMa) + NetVLAD    | 0.165      | 0.171     |\n| DFSfM(RoMa) + SP + SG    | 0.189      | 0.199     |\n| DFSfM(RoMa) + SIFT + NN  | **0.201**      | 0.207     |\n| DFSfM(RoMa)+ ALIKED + NN | 0.191      | **0.225**     |\n\nIn the end, we chose the DFSfM (RoMa + ALIKED + NN) approach, which achieved the highest score of **0.225** on the Public LB, with a Private LB score of **0.191**. Regrettably, although the DFSfM (RoMa + SIFT + NN) approach only scored **0.207** on the public LB, it reached the highest score of **0.201** on the Private LB.\n\n# **4. Ideas tried but not worked**\n\n## **4.1 For general scenes**\n\n### **4.1.1 Other image retrieval methods**\n\n- Anyloc\n- Eigenplaces\n- Rotation detection before image retrieval\n\nWe attempted to use the latest retrieval methods such as Anyloc and Eigenplaces, and found that the results did not show significant differences. At the same time, we also tried to add image rotation detection before retrieval, and found that the results did not improve.\n\n### **4.1.1 Other image matching methods**\n\n- DeDoDe + LG\n- DISK + LG\n- ALIKED + LG\n\nWe also tried some other sparse matching methods, such as ALIKED + LightGlue, DISK + LightGlue, DeDoDe + LightGlue, and found that the results did not significantly improve. Both the local and public leaderboard results were slightly worse than those of SuperPoint + SuperGlue. We believe the possible reason might be that we were unable to adjust the parameters correctly.\n\n## **4.2 For transparent scenes**\n\n- DUSt3R\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17083104%2F17e9beac384fcdede5a7e4ff4026862e%2Fdustr.png?generation=1718000589088088&alt=media)\n\nWe attempted to handle the special case of transparent scenes using DUSt3R and found that, like other matching methods, it was unable to effectively process transparent scenes.\n\n# **5. Acknowledgment:**\n\nOnce again, I would like to express my gratitude to the organizers for their contributions to this competition, and I appreciate the hard work and dedication of my teammates and all the participants."
  }
}