{
  "id": 510611,
  "title": "4th Place Solution: ALIKED+LightGlue is all you need [Prize Eligible]",
  "url": "/competitions/image-matching-challenge-2024/discussion/510611",
  "author_name": "tmyok",
  "post_date": "2024-06-06T21:03:13.763000",
  "votes": 46,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I am delighted to participate in the Image Matching Challenge again, two years after my last participation in 2022. I would like to extend my gratitude to the host and the Kaggle team for organizing the competition. I also pay my respects to all the competitors who competed against each other and completed the challenge together.</p>\n<p>This year's competition features scenes with challenging image quality compared to the previous IMC2022. It is worthwhile to evaluate the performance of the latest machine learning techniques. Additionally, I have learned that some of the distributed data includes scenes that have already been lost, such as the Temple of Baalshamin. This highlights the importance of digital archiving using 3D reconstruction technology and further underscores the social significance of this competition.</p>\n<h2>1. Specific Challenges Encountered</h2>\n<p>In explaining my solution, I would like to outline the key challenges encountered.</p>\n<h3>1.1 Rotated Images</h3>\n<p>In addition to conditions related to the shooting environment, there was a possibility that the host intentionally rotated some of the images. With the EXIF information removed, I had to rely solely on the images themselves to address this issue.</p>\n<h3>1.2 Transparent Scenes</h3>\n<p>The baseline code revealed that it could barely handle transparent objects. The images lacked texture and created reflections and specularities.</p>\n<h2>2. Solution</h2>\n<h3>2.1 Overview</h3>\n<p>Like other teams, my solution processes transparent and non-transparent scenes separately. After processing them separately, 3D reconstruction with COLMAP is performed using the matching results obtained from each.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2F159bf28bda8dbb0c83c394c2ec0f0e86%2Fsolution.jpg?generation=1717707107726771&amp;alt=media\" alt=\"overview\"></p>\n<h3>2.2 Non-Transparent Scenes</h3>\n<h4>2.2.1 Keypoint Detection</h4>\n<p>Keypoints were generated and cached while rotating the images by 90 degrees at a time. ALIKED-n16 was used, and keypoints for each rotation angle were retained.</p>\n<h4>2.2.2 Matching Stage</h4>\n<p>LightGlue was used for evaluating the matches. For a fixed key1, keypoints from key2 were evaluated in four patterns, and the combination with the highest number of matches was adopted. Referring to the IMC2023 2nd place solution (<a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/416873\" target=\"_blank\">link</a>), matches threshold was evaluated in two patterns: 100 and 125.</p>\n<h3>2.3 Transparent Scenes</h3>\n<h4>2.3.1 Foreground Segmentation - \"bottle\" class is all you need -</h4>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2Ffe4b9b8bdebbbd178b9b63388ce971e4%2Fcylinder_matching_initial.png?generation=1717707155993809&amp;alt=media\" alt=\"cylinder_matching_initial\"></p>\n<p>By closely observing failure cases in transparent scenes, it was noticed that many keypoints appeared in background areas, leading to failures in camera pose estimation. To suppress keypoints in background areas, I investigated foreground extraction methods. Using the DINOv2 Segmenter (<a href=\"https://github.com/facebookresearch/dinov2/blob/main/notebooks/semantic_segmentation.ipynb\" target=\"_blank\">link</a>), I discovered that the VOC2012 model assigned class5 to the foreground. According to the <a href=\"http://host.robots.ox.ac.uk/pascal/VOC/voc2012/segexamples/index.html\" target=\"_blank\">link</a>, class5 is assigned to \"bottle\". By treating this class as \"transparent\", I hypothesized that I could achieve high-precision segmentation.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2F6dd026c50f87131ec407026c2d513bf9%2Fcylinder_segmented.png?generation=1717707179881772&amp;alt=media\" alt=\"cylinder_segmented\"></p>\n<h4>2.3.2 Keypoint Detection with Original Scale</h4>\n<p>The images in transparent scenes were relatively large and all of the same size, so I decided to detect keypoints at the original scale without resizing. Considering VRAM and processing time, keypoints were detected in 1024x1024 grid units. Combining this with the DINOv2 Segmenter, keypoints were detected only in the foreground area.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2F8b95ff498022a69cc46e32d6134a49e3%2Fcylinder_keypoints_initial.png?generation=1717707437917224&amp;alt=media\" alt=\"keypoints initial\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2Fff16f90a5b4ac647de3cbe3349b1cd41%2Fcylinder_keypoints_result.png?generation=1717707455484348&amp;alt=media\" alt=\"keypoints proposed\"></p>\n<h4>2.3.3 Feature Matching</h4>\n<p>Given the shooting conditions this time, I determined that there was no need to search extensively for matches. Therefore, corresponding points were searched only between corresponding grids during keypoint detection. This significantly reduced the search range.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2F206ad587a4bff93ef5eb97ce7367a2ac%2Fcylinder_matching_result.png?generation=1717707266905858&amp;alt=media\" alt=\"cylinder_matching_proposed\"></p>\n<h3>2.4 Other Tricks</h3>\n<h4>2.4.1 Get Pairs Exhaustive</h4>\n<p>Since the number of images per scene was not large in IMC2024, exhaustive matching for all pairs was performed instead of searching for pairs with DINOv2 or EfficientNet. This reduced the risk of missing matches due to low embedding-based similarity.</p>\n<h4>2.4.2 Use <em>ALL</em> Images</h4>\n<blockquote>\n  <p>I think it's natural for the score to increase if you increase the number of images, as it makes it easier to triangulate.</p>\n</blockquote>\n<p>Inspired by Camaro's comment (<a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2024/discussion/494789#2768517\" target=\"_blank\">link</a>), I wondered if using images beyond those in submission.csv could facilitate easier 3D reconstruction. A simple LB probing revealed that there were images other than those listed in submission.csv in the test data's images folder. Using up to 100 image sets for validation significantly improved the validation score, so it was expected to be effective on the LB as well.</p>\n<h2>3. Results</h2>\n<p>After the competition ended, I conducted several late submissions to evaluate how much each additional technique contributed to improving the leaderboard (LB) score. For local validation, I used a subset of approximately 50 images generated with the following notebook.</p>\n<p><a href=\"https://www.kaggle.com/code/tmyok1984/imc2024-validation\" target=\"_blank\">https://www.kaggle.com/code/tmyok1984/imc2024-validation</a></p>\n<table>\n<thead>\n<tr>\n<th>No,</th>\n<th>Transparent trick</th>\n<th>Exhaustive matching</th>\n<th>Use all images</th>\n<th>Private LB</th>\n<th>Public LB</th>\n<th>Val (avg.)</th>\n<th>church</th>\n<th>dioscuri</th>\n<th>lizard</th>\n<th>temple-baalshamin</th>\n<th>pond</th>\n<th>glass_cup</th>\n<th>glass_cylinder</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td></td>\n<td></td>\n<td></td>\n<td>0.149</td>\n<td>0.136</td>\n<td>0.26</td>\n<td>0.24</td>\n<td>0.47</td>\n<td>0.54</td>\n<td>0.42</td>\n<td>0.10</td>\n<td>0.02</td>\n<td>0.03</td>\n</tr>\n<tr>\n<td>2</td>\n<td>✓</td>\n<td></td>\n<td></td>\n<td>0.184</td>\n<td>0.171</td>\n<td>0.32</td>\n<td>0.24</td>\n<td>0.47</td>\n<td>0.54</td>\n<td>0.40</td>\n<td>0.09</td>\n<td>0.02</td>\n<td>0.47</td>\n</tr>\n<tr>\n<td>3</td>\n<td>✓</td>\n<td>✓</td>\n<td></td>\n<td>0.186</td>\n<td>0.176</td>\n<td>0.34</td>\n<td>0.24</td>\n<td>0.52</td>\n<td>0.51</td>\n<td>0.42</td>\n<td>0.17</td>\n<td>0.02</td>\n<td>0.47</td>\n</tr>\n<tr>\n<td>4</td>\n<td>✓</td>\n<td>✓</td>\n<td>✓</td>\n<td>0.197</td>\n<td>0.194</td>\n<td>0.43</td>\n<td>0.31</td>\n<td>0.56</td>\n<td>0.79</td>\n<td>0.41</td>\n<td>0.46</td>\n<td>0.02</td>\n<td>0.47</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>Baseline (1) : Private LB=0.149 (around 300th place), Public LB=0.136 (around 300th place)</li>\n<li>Add Transparent trick (1+2) : Private LB=0.184 (around 8th place), Public LB=0.171 (around 50th place)</li>\n<li>Add Exhaustive matching (1+2+3) : Private LB=0.186 (around 7th place), Public LB=0.176 (around 30th place)</li>\n<li>Add ALL images (1+2+3+4) : Private LB=0.197 (4th place), Public LB=0.194 (around 7th place)</li>\n</ul>\n<p>This ablation study revealed the following points:</p>\n<ol>\n<li><p>Effect of Transparent trick: Adding the trick for transparent scenes significantly improved the Private LB score from 0.149 to 0.184. This demonstrated that handling transparent scenes greatly contributes to overall performance improvement.</p></li>\n<li><p>Effect of Exhaustive matching: Adding the exhaustive matching method for all pairs further improved the score slightly. This indicated that the risk of missing matches due to low embedding-based similarity was effectively reduced.</p></li>\n<li><p>Effect of Using All Images: Using images beyond those in submission.csv resulted in the most significant score improvement. This confirmed that utilizing additional image information greatly enhances the accuracy of 3D reconstruction.</p></li>\n</ol>\n<p>Overall, handling transparent scenes and using all available images were shown to be key factors for ranking high on the leaderboard.</p>\n<h2>4. Implementation</h2>\n<p>This section shares techniques related to implementation.</p>\n<h3>4.1 Multiple Process / Multiple GPUs</h3>\n<p>As highlighted in previous solutions, IMC is characterized by the substantial computational cost of both CPU and GPU processing. Running the CPU and GPU in parallel can potentially double the processing speed. Additionally, using GPUs with two T4 cards allows for parallelizing GPU processing, thereby doubling the processing capacity.</p>\n<p>Although the importance of these aspects has been mentioned in past solutions, it is very rare for reference code to be made public. As part of my contribution to the community, I have shared the code I used this time. The base part of the implementation can be used beyond IMC, so please refer to it.</p>\n<p><a href=\"https://www.kaggle.com/code/tmyok1984/imc2024-exp556\" target=\"_blank\">https://www.kaggle.com/code/tmyok1984/imc2024-exp556</a></p>\n<h3>4.2 Utility Script</h3>\n<p>When using packages not available in the Kaggle environment, offline installation is necessary. However, performing offline package installation in the submission notebook wastes submission time. To solve this problem, I always use the utility script feature. By adding a utility script notebook with pre-installed packages, necessary packages can be installed beforehand, making the submission notebook more efficient. The utility script I used this time is below.</p>\n<p><a href=\"https://www.kaggle.com/code/tmyok1984/imc2024-install-once\" target=\"_blank\">https://www.kaggle.com/code/tmyok1984/imc2024-install-once</a></p>\n<p>The following link provides detailed information on how to create utility scripts, so please refer to it.</p>\n<p><a href=\"https://www.kaggle.com/code/kononenko/pip-install-once\" target=\"_blank\">https://www.kaggle.com/code/kononenko/pip-install-once</a></p>\n<h2>5. What Did Not Work (For Me)</h2>\n<ul>\n<li>Symmetries-and-repeats: I spent most of the competition period on symmetries-and-repeats. Ultimately, I implemented a combination of a  <a href=\"https://github.com/RuojinCai/doppelgangers\" target=\"_blank\">Doppelgangers classifier</a> and 3D geometrical verification. Although I adopted it in the final submission and it worked well locally in some cases, the LB consistently worsened during the ablation study with the late submission. I experienced the Concorde effect firsthand.</li>\n<li>Overlap region：In our 10th place solution for IMC2022, we used DKM-ROI (<a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2022/discussion/328903\" target=\"_blank\">link1</a>, <a href=\"https://www.kaggle.com/code/tmyok1984/imc2022-ensemble-with-dkm-roi\" target=\"_blank\">link2</a>). This time, there was an excellent publicly available notebook (<a href=\"https://www.kaggle.com/code/nartaa/imc24-overlap-detection\" target=\"_blank\">link</a>), so I implemented it as well but did not achieve good results. While high matching accuracy was required in IMC2022, in IMC2024, high robustness was required, so its effectiveness was limited. Specifically, if the initial value for finding overlap was incorrect, the subsequent steps would break down significantly. In contrast, the 1st place solution  (<a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2024/discussion/510084\" target=\"_blank\">link</a>) successfully estimated the ROI using multiple viewpoints robustly and achieved excellent results.</li>\n<li>Non-Maximum Suppression：This also strongly boosted our LB in IMC2022. However, ALIKED was strong enough this time that even when LoFTR, DeDoDe, and RoMa were ensembled, no improvement was observed in my local validation.</li>\n<li>COLMAP hyperparameters：Using the help(pycolmap) command, many configurable parameters could be seen. Various parameters were changed, but no significant effect was obtained.</li>\n</ul>",
  "messages": [
    {
      "id": 2859197,
      "postDate": "2024-06-06T21:03:13.763Z",
      "content": "<p>I am delighted to participate in the Image Matching Challenge again, two years after my last participation in 2022. I would like to extend my gratitude to the host and the Kaggle team for organizing the competition. I also pay my respects to all the competitors who competed against each other and completed the challenge together.</p>\n<p>This year's competition features scenes with challenging image quality compared to the previous IMC2022. It is worthwhile to evaluate the performance of the latest machine learning techniques. Additionally, I have learned that some of the distributed data includes scenes that have already been lost, such as the Temple of Baalshamin. This highlights the importance of digital archiving using 3D reconstruction technology and further underscores the social significance of this competition.</p>\n<h2>1. Specific Challenges Encountered</h2>\n<p>In explaining my solution, I would like to outline the key challenges encountered.</p>\n<h3>1.1 Rotated Images</h3>\n<p>In addition to conditions related to the shooting environment, there was a possibility that the host intentionally rotated some of the images. With the EXIF information removed, I had to rely solely on the images themselves to address this issue.</p>\n<h3>1.2 Transparent Scenes</h3>\n<p>The baseline code revealed that it could barely handle transparent objects. The images lacked texture and created reflections and specularities.</p>\n<h2>2. Solution</h2>\n<h3>2.1 Overview</h3>\n<p>Like other teams, my solution processes transparent and non-transparent scenes separately. After processing them separately, 3D reconstruction with COLMAP is performed using the matching results obtained from each.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2F159bf28bda8dbb0c83c394c2ec0f0e86%2Fsolution.jpg?generation=1717707107726771&amp;alt=media\" alt=\"overview\"></p>\n<h3>2.2 Non-Transparent Scenes</h3>\n<h4>2.2.1 Keypoint Detection</h4>\n<p>Keypoints were generated and cached while rotating the images by 90 degrees at a time. ALIKED-n16 was used, and keypoints for each rotation angle were retained.</p>\n<h4>2.2.2 Matching Stage</h4>\n<p>LightGlue was used for evaluating the matches. For a fixed key1, keypoints from key2 were evaluated in four patterns, and the combination with the highest number of matches was adopted. Referring to the IMC2023 2nd place solution (<a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/416873\" target=\"_blank\">link</a>), matches threshold was evaluated in two patterns: 100 and 125.</p>\n<h3>2.3 Transparent Scenes</h3>\n<h4>2.3.1 Foreground Segmentation - \"bottle\" class is all you need -</h4>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2Ffe4b9b8bdebbbd178b9b63388ce971e4%2Fcylinder_matching_initial.png?generation=1717707155993809&amp;alt=media\" alt=\"cylinder_matching_initial\"></p>\n<p>By closely observing failure cases in transparent scenes, it was noticed that many keypoints appeared in background areas, leading to failures in camera pose estimation. To suppress keypoints in background areas, I investigated foreground extraction methods. Using the DINOv2 Segmenter (<a href=\"https://github.com/facebookresearch/dinov2/blob/main/notebooks/semantic_segmentation.ipynb\" target=\"_blank\">link</a>), I discovered that the VOC2012 model assigned class5 to the foreground. According to the <a href=\"http://host.robots.ox.ac.uk/pascal/VOC/voc2012/segexamples/index.html\" target=\"_blank\">link</a>, class5 is assigned to \"bottle\". By treating this class as \"transparent\", I hypothesized that I could achieve high-precision segmentation.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2F6dd026c50f87131ec407026c2d513bf9%2Fcylinder_segmented.png?generation=1717707179881772&amp;alt=media\" alt=\"cylinder_segmented\"></p>\n<h4>2.3.2 Keypoint Detection with Original Scale</h4>\n<p>The images in transparent scenes were relatively large and all of the same size, so I decided to detect keypoints at the original scale without resizing. Considering VRAM and processing time, keypoints were detected in 1024x1024 grid units. Combining this with the DINOv2 Segmenter, keypoints were detected only in the foreground area.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2F8b95ff498022a69cc46e32d6134a49e3%2Fcylinder_keypoints_initial.png?generation=1717707437917224&amp;alt=media\" alt=\"keypoints initial\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2Fff16f90a5b4ac647de3cbe3349b1cd41%2Fcylinder_keypoints_result.png?generation=1717707455484348&amp;alt=media\" alt=\"keypoints proposed\"></p>\n<h4>2.3.3 Feature Matching</h4>\n<p>Given the shooting conditions this time, I determined that there was no need to search extensively for matches. Therefore, corresponding points were searched only between corresponding grids during keypoint detection. This significantly reduced the search range.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2F206ad587a4bff93ef5eb97ce7367a2ac%2Fcylinder_matching_result.png?generation=1717707266905858&amp;alt=media\" alt=\"cylinder_matching_proposed\"></p>\n<h3>2.4 Other Tricks</h3>\n<h4>2.4.1 Get Pairs Exhaustive</h4>\n<p>Since the number of images per scene was not large in IMC2024, exhaustive matching for all pairs was performed instead of searching for pairs with DINOv2 or EfficientNet. This reduced the risk of missing matches due to low embedding-based similarity.</p>\n<h4>2.4.2 Use <em>ALL</em> Images</h4>\n<blockquote>\n  <p>I think it's natural for the score to increase if you increase the number of images, as it makes it easier to triangulate.</p>\n</blockquote>\n<p>Inspired by Camaro's comment (<a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2024/discussion/494789#2768517\" target=\"_blank\">link</a>), I wondered if using images beyond those in submission.csv could facilitate easier 3D reconstruction. A simple LB probing revealed that there were images other than those listed in submission.csv in the test data's images folder. Using up to 100 image sets for validation significantly improved the validation score, so it was expected to be effective on the LB as well.</p>\n<h2>3. Results</h2>\n<p>After the competition ended, I conducted several late submissions to evaluate how much each additional technique contributed to improving the leaderboard (LB) score. For local validation, I used a subset of approximately 50 images generated with the following notebook.</p>\n<p><a href=\"https://www.kaggle.com/code/tmyok1984/imc2024-validation\" target=\"_blank\">https://www.kaggle.com/code/tmyok1984/imc2024-validation</a></p>\n<table>\n<thead>\n<tr>\n<th>No,</th>\n<th>Transparent trick</th>\n<th>Exhaustive matching</th>\n<th>Use all images</th>\n<th>Private LB</th>\n<th>Public LB</th>\n<th>Val (avg.)</th>\n<th>church</th>\n<th>dioscuri</th>\n<th>lizard</th>\n<th>temple-baalshamin</th>\n<th>pond</th>\n<th>glass_cup</th>\n<th>glass_cylinder</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td></td>\n<td></td>\n<td></td>\n<td>0.149</td>\n<td>0.136</td>\n<td>0.26</td>\n<td>0.24</td>\n<td>0.47</td>\n<td>0.54</td>\n<td>0.42</td>\n<td>0.10</td>\n<td>0.02</td>\n<td>0.03</td>\n</tr>\n<tr>\n<td>2</td>\n<td>✓</td>\n<td></td>\n<td></td>\n<td>0.184</td>\n<td>0.171</td>\n<td>0.32</td>\n<td>0.24</td>\n<td>0.47</td>\n<td>0.54</td>\n<td>0.40</td>\n<td>0.09</td>\n<td>0.02</td>\n<td>0.47</td>\n</tr>\n<tr>\n<td>3</td>\n<td>✓</td>\n<td>✓</td>\n<td></td>\n<td>0.186</td>\n<td>0.176</td>\n<td>0.34</td>\n<td>0.24</td>\n<td>0.52</td>\n<td>0.51</td>\n<td>0.42</td>\n<td>0.17</td>\n<td>0.02</td>\n<td>0.47</td>\n</tr>\n<tr>\n<td>4</td>\n<td>✓</td>\n<td>✓</td>\n<td>✓</td>\n<td>0.197</td>\n<td>0.194</td>\n<td>0.43</td>\n<td>0.31</td>\n<td>0.56</td>\n<td>0.79</td>\n<td>0.41</td>\n<td>0.46</td>\n<td>0.02</td>\n<td>0.47</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>Baseline (1) : Private LB=0.149 (around 300th place), Public LB=0.136 (around 300th place)</li>\n<li>Add Transparent trick (1+2) : Private LB=0.184 (around 8th place), Public LB=0.171 (around 50th place)</li>\n<li>Add Exhaustive matching (1+2+3) : Private LB=0.186 (around 7th place), Public LB=0.176 (around 30th place)</li>\n<li>Add ALL images (1+2+3+4) : Private LB=0.197 (4th place), Public LB=0.194 (around 7th place)</li>\n</ul>\n<p>This ablation study revealed the following points:</p>\n<ol>\n<li><p>Effect of Transparent trick: Adding the trick for transparent scenes significantly improved the Private LB score from 0.149 to 0.184. This demonstrated that handling transparent scenes greatly contributes to overall performance improvement.</p></li>\n<li><p>Effect of Exhaustive matching: Adding the exhaustive matching method for all pairs further improved the score slightly. This indicated that the risk of missing matches due to low embedding-based similarity was effectively reduced.</p></li>\n<li><p>Effect of Using All Images: Using images beyond those in submission.csv resulted in the most significant score improvement. This confirmed that utilizing additional image information greatly enhances the accuracy of 3D reconstruction.</p></li>\n</ol>\n<p>Overall, handling transparent scenes and using all available images were shown to be key factors for ranking high on the leaderboard.</p>\n<h2>4. Implementation</h2>\n<p>This section shares techniques related to implementation.</p>\n<h3>4.1 Multiple Process / Multiple GPUs</h3>\n<p>As highlighted in previous solutions, IMC is characterized by the substantial computational cost of both CPU and GPU processing. Running the CPU and GPU in parallel can potentially double the processing speed. Additionally, using GPUs with two T4 cards allows for parallelizing GPU processing, thereby doubling the processing capacity.</p>\n<p>Although the importance of these aspects has been mentioned in past solutions, it is very rare for reference code to be made public. As part of my contribution to the community, I have shared the code I used this time. The base part of the implementation can be used beyond IMC, so please refer to it.</p>\n<p><a href=\"https://www.kaggle.com/code/tmyok1984/imc2024-exp556\" target=\"_blank\">https://www.kaggle.com/code/tmyok1984/imc2024-exp556</a></p>\n<h3>4.2 Utility Script</h3>\n<p>When using packages not available in the Kaggle environment, offline installation is necessary. However, performing offline package installation in the submission notebook wastes submission time. To solve this problem, I always use the utility script feature. By adding a utility script notebook with pre-installed packages, necessary packages can be installed beforehand, making the submission notebook more efficient. The utility script I used this time is below.</p>\n<p><a href=\"https://www.kaggle.com/code/tmyok1984/imc2024-install-once\" target=\"_blank\">https://www.kaggle.com/code/tmyok1984/imc2024-install-once</a></p>\n<p>The following link provides detailed information on how to create utility scripts, so please refer to it.</p>\n<p><a href=\"https://www.kaggle.com/code/kononenko/pip-install-once\" target=\"_blank\">https://www.kaggle.com/code/kononenko/pip-install-once</a></p>\n<h2>5. What Did Not Work (For Me)</h2>\n<ul>\n<li>Symmetries-and-repeats: I spent most of the competition period on symmetries-and-repeats. Ultimately, I implemented a combination of a  <a href=\"https://github.com/RuojinCai/doppelgangers\" target=\"_blank\">Doppelgangers classifier</a> and 3D geometrical verification. Although I adopted it in the final submission and it worked well locally in some cases, the LB consistently worsened during the ablation study with the late submission. I experienced the Concorde effect firsthand.</li>\n<li>Overlap region：In our 10th place solution for IMC2022, we used DKM-ROI (<a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2022/discussion/328903\" target=\"_blank\">link1</a>, <a href=\"https://www.kaggle.com/code/tmyok1984/imc2022-ensemble-with-dkm-roi\" target=\"_blank\">link2</a>). This time, there was an excellent publicly available notebook (<a href=\"https://www.kaggle.com/code/nartaa/imc24-overlap-detection\" target=\"_blank\">link</a>), so I implemented it as well but did not achieve good results. While high matching accuracy was required in IMC2022, in IMC2024, high robustness was required, so its effectiveness was limited. Specifically, if the initial value for finding overlap was incorrect, the subsequent steps would break down significantly. In contrast, the 1st place solution  (<a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2024/discussion/510084\" target=\"_blank\">link</a>) successfully estimated the ROI using multiple viewpoints robustly and achieved excellent results.</li>\n<li>Non-Maximum Suppression：This also strongly boosted our LB in IMC2022. However, ALIKED was strong enough this time that even when LoFTR, DeDoDe, and RoMa were ensembled, no improvement was observed in my local validation.</li>\n<li>COLMAP hyperparameters：Using the help(pycolmap) command, many configurable parameters could be seen. Various parameters were changed, but no significant effect was obtained.</li>\n</ul>",
      "rawMarkdown": "I am delighted to participate in the Image Matching Challenge again, two years after my last participation in 2022. I would like to extend my gratitude to the host and the Kaggle team for organizing the competition. I also pay my respects to all the competitors who competed against each other and completed the challenge together.\n\nThis year's competition features scenes with challenging image quality compared to the previous IMC2022. It is worthwhile to evaluate the performance of the latest machine learning techniques. Additionally, I have learned that some of the distributed data includes scenes that have already been lost, such as the Temple of Baalshamin. This highlights the importance of digital archiving using 3D reconstruction technology and further underscores the social significance of this competition.\n\n## 1. Specific Challenges Encountered\n\nIn explaining my solution, I would like to outline the key challenges encountered.\n\n### 1.1 Rotated Images\nIn addition to conditions related to the shooting environment, there was a possibility that the host intentionally rotated some of the images. With the EXIF information removed, I had to rely solely on the images themselves to address this issue.\n\n### 1.2 Transparent Scenes\nThe baseline code revealed that it could barely handle transparent objects. The images lacked texture and created reflections and specularities.\n\n## 2. Solution\n\n### 2.1 Overview\n\nLike other teams, my solution processes transparent and non-transparent scenes separately. After processing them separately, 3D reconstruction with COLMAP is performed using the matching results obtained from each.\n\n![overview](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2F159bf28bda8dbb0c83c394c2ec0f0e86%2Fsolution.jpg?generation=1717707107726771&alt=media)\n\n### 2.2 Non-Transparent Scenes\n\n#### 2.2.1 Keypoint Detection\n\nKeypoints were generated and cached while rotating the images by 90 degrees at a time. ALIKED-n16 was used, and keypoints for each rotation angle were retained.\n\n#### 2.2.2 Matching Stage\nLightGlue was used for evaluating the matches. For a fixed key1, keypoints from key2 were evaluated in four patterns, and the combination with the highest number of matches was adopted. Referring to the IMC2023 2nd place solution ([link](https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/416873)), matches threshold was evaluated in two patterns: 100 and 125.\n\n### 2.3 Transparent Scenes\n#### 2.3.1 Foreground Segmentation - \"bottle\" class is all you need -\n\n![cylinder_matching_initial](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2Ffe4b9b8bdebbbd178b9b63388ce971e4%2Fcylinder_matching_initial.png?generation=1717707155993809&alt=media)\n\nBy closely observing failure cases in transparent scenes, it was noticed that many keypoints appeared in background areas, leading to failures in camera pose estimation. To suppress keypoints in background areas, I investigated foreground extraction methods. Using the DINOv2 Segmenter ([link](https://github.com/facebookresearch/dinov2/blob/main/notebooks/semantic_segmentation.ipynb)), I discovered that the VOC2012 model assigned class5 to the foreground. According to the [link](http://host.robots.ox.ac.uk/pascal/VOC/voc2012/segexamples/index.html), class5 is assigned to \"bottle\". By treating this class as \"transparent\", I hypothesized that I could achieve high-precision segmentation.\n\n![cylinder_segmented](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2F6dd026c50f87131ec407026c2d513bf9%2Fcylinder_segmented.png?generation=1717707179881772&alt=media)\n\n#### 2.3.2 Keypoint Detection with Original Scale\nThe images in transparent scenes were relatively large and all of the same size, so I decided to detect keypoints at the original scale without resizing. Considering VRAM and processing time, keypoints were detected in 1024x1024 grid units. Combining this with the DINOv2 Segmenter, keypoints were detected only in the foreground area.\n\n![keypoints initial](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2F8b95ff498022a69cc46e32d6134a49e3%2Fcylinder_keypoints_initial.png?generation=1717707437917224&alt=media)\n\n![keypoints proposed](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2Fff16f90a5b4ac647de3cbe3349b1cd41%2Fcylinder_keypoints_result.png?generation=1717707455484348&alt=media)\n\n#### 2.3.3 Feature Matching\nGiven the shooting conditions this time, I determined that there was no need to search extensively for matches. Therefore, corresponding points were searched only between corresponding grids during keypoint detection. This significantly reduced the search range.\n\n![cylinder_matching_proposed](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2F206ad587a4bff93ef5eb97ce7367a2ac%2Fcylinder_matching_result.png?generation=1717707266905858&alt=media)\n\n### 2.4 Other Tricks\n#### 2.4.1 Get Pairs Exhaustive\nSince the number of images per scene was not large in IMC2024, exhaustive matching for all pairs was performed instead of searching for pairs with DINOv2 or EfficientNet. This reduced the risk of missing matches due to low embedding-based similarity.\n\n#### 2.4.2 Use *ALL* Images\n\n> I think it's natural for the score to increase if you increase the number of images, as it makes it easier to triangulate.\n\nInspired by Camaro's comment ([link](https://www.kaggle.com/competitions/image-matching-challenge-2024/discussion/494789#2768517)), I wondered if using images beyond those in submission.csv could facilitate easier 3D reconstruction. A simple LB probing revealed that there were images other than those listed in submission.csv in the test data's images folder. Using up to 100 image sets for validation significantly improved the validation score, so it was expected to be effective on the LB as well.\n\n## 3. Results\n\nAfter the competition ended, I conducted several late submissions to evaluate how much each additional technique contributed to improving the leaderboard (LB) score. For local validation, I used a subset of approximately 50 images generated with the following notebook.\n\n[https://www.kaggle.com/code/tmyok1984/imc2024-validation](https://www.kaggle.com/code/tmyok1984/imc2024-validation)\n\n| No, | Transparent trick | Exhaustive matching | Use all images | Private LB | Public LB | Val (avg.) | church | dioscuri | lizard | temple-baalshamin | pond | glass_cup | glass_cylinder |\n| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |\n| 1 |  |  |  | 0.149 | 0.136 | 0.26  | 0.24  | 0.47  | 0.54  | 0.42  | 0.10  | 0.02  | 0.03  |\n| 2 | ✓ |  |  | 0.184 | 0.171 | 0.32  | 0.24  | 0.47  | 0.54  | 0.40  | 0.09  | 0.02  | 0.47  |\n| 3 | ✓ | ✓ |  | 0.186 | 0.176 | 0.34  | 0.24  | 0.52  | 0.51  | 0.42  | 0.17  | 0.02  | 0.47  |\n| 4 | ✓ | ✓ | ✓ | 0.197 | 0.194 | 0.43 | 0.31 | 0.56 | 0.79 | 0.41 | 0.46 | 0.02 | 0.47 |\n\n- Baseline (1) : Private LB=0.149 (around 300th place), Public LB=0.136 (around 300th place)\n- Add Transparent trick (1+2) : Private LB=0.184 (around 8th place), Public LB=0.171 (around 50th place)\n- Add Exhaustive matching (1+2+3) : Private LB=0.186 (around 7th place), Public LB=0.176 (around 30th place)\n- Add ALL images (1+2+3+4) : Private LB=0.197 (4th place), Public LB=0.194 (around 7th place)\n\nThis ablation study revealed the following points:\n\n1. Effect of Transparent trick: Adding the trick for transparent scenes significantly improved the Private LB score from 0.149 to 0.184. This demonstrated that handling transparent scenes greatly contributes to overall performance improvement.\n\n2. Effect of Exhaustive matching: Adding the exhaustive matching method for all pairs further improved the score slightly. This indicated that the risk of missing matches due to low embedding-based similarity was effectively reduced.\n\n3. Effect of Using All Images: Using images beyond those in submission.csv resulted in the most significant score improvement. This confirmed that utilizing additional image information greatly enhances the accuracy of 3D reconstruction.\n\nOverall, handling transparent scenes and using all available images were shown to be key factors for ranking high on the leaderboard.\n\n## 4. Implementation\n\nThis section shares techniques related to implementation.\n\n### 4.1 Multiple Process / Multiple GPUs\nAs highlighted in previous solutions, IMC is characterized by the substantial computational cost of both CPU and GPU processing. Running the CPU and GPU in parallel can potentially double the processing speed. Additionally, using GPUs with two T4 cards allows for parallelizing GPU processing, thereby doubling the processing capacity.\n\nAlthough the importance of these aspects has been mentioned in past solutions, it is very rare for reference code to be made public. As part of my contribution to the community, I have shared the code I used this time. The base part of the implementation can be used beyond IMC, so please refer to it.\n\n[https://www.kaggle.com/code/tmyok1984/imc2024-exp556](https://www.kaggle.com/code/tmyok1984/imc2024-exp556)\n\n### 4.2 Utility Script\n\nWhen using packages not available in the Kaggle environment, offline installation is necessary. However, performing offline package installation in the submission notebook wastes submission time. To solve this problem, I always use the utility script feature. By adding a utility script notebook with pre-installed packages, necessary packages can be installed beforehand, making the submission notebook more efficient. The utility script I used this time is below.\n\n[https://www.kaggle.com/code/tmyok1984/imc2024-install-once](https://www.kaggle.com/code/tmyok1984/imc2024-install-once)\n\nThe following link provides detailed information on how to create utility scripts, so please refer to it.\n\n[https://www.kaggle.com/code/kononenko/pip-install-once](https://www.kaggle.com/code/kononenko/pip-install-once)\n\n## 5. What Did Not Work (For Me)\n- Symmetries-and-repeats: I spent most of the competition period on symmetries-and-repeats. Ultimately, I implemented a combination of a  [Doppelgangers classifier](https://github.com/RuojinCai/doppelgangers) and 3D geometrical verification. Although I adopted it in the final submission and it worked well locally in some cases, the LB consistently worsened during the ablation study with the late submission. I experienced the Concorde effect firsthand.\n- Overlap region：In our 10th place solution for IMC2022, we used DKM-ROI ([link1](https://www.kaggle.com/competitions/image-matching-challenge-2022/discussion/328903), [link2](https://www.kaggle.com/code/tmyok1984/imc2022-ensemble-with-dkm-roi)). This time, there was an excellent publicly available notebook ([link](https://www.kaggle.com/code/nartaa/imc24-overlap-detection)), so I implemented it as well but did not achieve good results. While high matching accuracy was required in IMC2022, in IMC2024, high robustness was required, so its effectiveness was limited. Specifically, if the initial value for finding overlap was incorrect, the subsequent steps would break down significantly. In contrast, the 1st place solution  ([link](https://www.kaggle.com/competitions/image-matching-challenge-2024/discussion/510084)) successfully estimated the ROI using multiple viewpoints robustly and achieved excellent results.\n- Non-Maximum Suppression：This also strongly boosted our LB in IMC2022. However, ALIKED was strong enough this time that even when LoFTR, DeDoDe, and RoMa were ensembled, no improvement was observed in my local validation.\n- COLMAP hyperparameters：Using the help(pycolmap) command, many configurable parameters could be seen. Various parameters were changed, but no significant effect was obtained.",
      "votes": 46
    },
    {
      "id": 2863385,
      "postDate": "2024-06-09T11:23:12.363Z",
      "content": "<p>I have published my code in GitHub: <a href=\"https://github.com/tmyok/kaggle-image-matching-challenge-2024\" target=\"_blank\">https://github.com/tmyok/kaggle-image-matching-challenge-2024</a></p>",
      "rawMarkdown": "I have published my code in GitHub: https://github.com/tmyok/kaggle-image-matching-challenge-2024",
      "votes": 4
    },
    {
      "id": 2864972,
      "postDate": "2024-06-10T13:01:03.537Z",
      "content": "<p>Thank you for sharing the solution.<br>\nI found it very interesting.</p>\n<p>I have a question about the transparent scene section. For transparent scenes, were all pairs matched exhaustively in the same way as for non-transparent scenes?</p>",
      "rawMarkdown": "Thank you for sharing the solution.\nI found it very interesting.\n\nI have a question about the transparent scene section. For transparent scenes, were all pairs matched exhaustively in the same way as for non-transparent scenes?",
      "votes": 2,
      "replies": [
        {
          "id": 2865290,
          "postDate": "2024-06-10T16:32:39.617Z",
          "content": "<p>Yes, that's correct. For transparent scenes, all pairs were matched exhaustively in the same way as for non-transparent scenes. Please refer to the following link for implementation details:<br>\n<a href=\"https://www.kaggle.com/code/tmyok1984/imc2024-exp556#Feature-matching\" target=\"_blank\">https://www.kaggle.com/code/tmyok1984/imc2024-exp556#Feature-matching</a></p>\n<p>By the way, congratulations on your victory. I learned a lot from reading your team's solution.</p>",
          "rawMarkdown": "Yes, that's correct. For transparent scenes, all pairs were matched exhaustively in the same way as for non-transparent scenes. Please refer to the following link for implementation details:\nhttps://www.kaggle.com/code/tmyok1984/imc2024-exp556#Feature-matching\n\nBy the way, congratulations on your victory. I learned a lot from reading your team's solution.",
          "votes": 2,
          "replies": [
            {
              "id": 2865696,
              "postDate": "2024-06-10T23:31:03.587Z",
              "content": "<p>Thank you for telling me.<br>\nThat's interesting because I was also initially inferring transparent scenes using a similar method to yours, but I couldn't get valid results without some method of estimating consecutive image pairs.<br>\nMaybe a grid division that I hadn't thought of was more effective. I'll check the shared source code for details.</p>\n<p>Also, congratulations again on the solo gold!</p>",
              "rawMarkdown": "Thank you for telling me.\nThat's interesting because I was also initially inferring transparent scenes using a similar method to yours, but I couldn't get valid results without some method of estimating consecutive image pairs.\nMaybe a grid division that I hadn't thought of was more effective. I'll check the shared source code for details.\n\nAlso, congratulations again on the solo gold!",
              "votes": 1
            },
            {
              "id": 2865703,
              "postDate": "2024-06-10T23:54:47.033Z",
              "content": "<p>I was satisfied with my own method and didn't work on further improving the accuracy for transparent scenes. I learned a lot from your team's TSP-based method.</p>\n<p>Congrats on becoming a competitions grandmaster!</p>",
              "rawMarkdown": "I was satisfied with my own method and didn't work on further improving the accuracy for transparent scenes. I learned a lot from your team's TSP-based method.\n\nCongrats on becoming a competitions grandmaster!",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2860391,
      "postDate": "2024-06-07T15:31:24.550Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2859306,
      "postDate": "2024-06-07T00:27:15.287Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true,
      "replies": [
        {
          "id": 2865309,
          "postDate": "2024-06-10T16:39:50.220Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2863385,
      "author_name": "tmyok",
      "author_url": "",
      "post_date": "2024-06-09T11:23:12.363000",
      "content": "<p>I have published my code in GitHub: <a href=\"https://github.com/tmyok/kaggle-image-matching-challenge-2024\" target=\"_blank\">https://github.com/tmyok/kaggle-image-matching-challenge-2024</a></p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2864972,
      "author_name": "YumeNeko",
      "author_url": "",
      "post_date": "2024-06-10T13:01:03.537000",
      "content": "<p>Thank you for sharing the solution.<br>\nI found it very interesting.</p>\n<p>I have a question about the transparent scene section. For transparent scenes, were all pairs matched exhaustively in the same way as for non-transparent scenes?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2865290,
          "author_name": "tmyok",
          "author_url": "",
          "post_date": "2024-06-10T16:32:39.617000",
          "content": "<p>Yes, that's correct. For transparent scenes, all pairs were matched exhaustively in the same way as for non-transparent scenes. Please refer to the following link for implementation details:<br>\n<a href=\"https://www.kaggle.com/code/tmyok1984/imc2024-exp556#Feature-matching\" target=\"_blank\">https://www.kaggle.com/code/tmyok1984/imc2024-exp556#Feature-matching</a></p>\n<p>By the way, congratulations on your victory. I learned a lot from reading your team's solution.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2865696,
              "author_name": "YumeNeko",
              "author_url": "",
              "post_date": "2024-06-10T23:31:03.587000",
              "content": "<p>Thank you for telling me.<br>\nThat's interesting because I was also initially inferring transparent scenes using a similar method to yours, but I couldn't get valid results without some method of estimating consecutive image pairs.<br>\nMaybe a grid division that I hadn't thought of was more effective. I'll check the shared source code for details.</p>\n<p>Also, congratulations again on the solo gold!</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2865703,
              "author_name": "tmyok",
              "author_url": "",
              "post_date": "2024-06-10T23:54:47.033000",
              "content": "<p>I was satisfied with my own method and didn't work on further improving the accuracy for transparent scenes. I learned a lot from your team's TSP-based method.</p>\n<p>Congrats on becoming a competitions grandmaster!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2860391,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-06-07T15:31:24.550000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2859306,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-06-07T00:27:15.287000",
      "content": "",
      "votes": -1,
      "replies": [
        {
          "id": 2865309,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-06-10T16:39:50.220000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2859197": "I am delighted to participate in the Image Matching Challenge again, two years after my last participation in 2022. I would like to extend my gratitude to the host and the Kaggle team for organizing the competition. I also pay my respects to all the competitors who competed against each other and completed the challenge together.\n\nThis year's competition features scenes with challenging image quality compared to the previous IMC2022. It is worthwhile to evaluate the performance of the latest machine learning techniques. Additionally, I have learned that some of the distributed data includes scenes that have already been lost, such as the Temple of Baalshamin. This highlights the importance of digital archiving using 3D reconstruction technology and further underscores the social significance of this competition.\n\n## 1. Specific Challenges Encountered\n\nIn explaining my solution, I would like to outline the key challenges encountered.\n\n### 1.1 Rotated Images\nIn addition to conditions related to the shooting environment, there was a possibility that the host intentionally rotated some of the images. With the EXIF information removed, I had to rely solely on the images themselves to address this issue.\n\n### 1.2 Transparent Scenes\nThe baseline code revealed that it could barely handle transparent objects. The images lacked texture and created reflections and specularities.\n\n## 2. Solution\n\n### 2.1 Overview\n\nLike other teams, my solution processes transparent and non-transparent scenes separately. After processing them separately, 3D reconstruction with COLMAP is performed using the matching results obtained from each.\n\n![overview](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2F159bf28bda8dbb0c83c394c2ec0f0e86%2Fsolution.jpg?generation=1717707107726771&alt=media)\n\n### 2.2 Non-Transparent Scenes\n\n#### 2.2.1 Keypoint Detection\n\nKeypoints were generated and cached while rotating the images by 90 degrees at a time. ALIKED-n16 was used, and keypoints for each rotation angle were retained.\n\n#### 2.2.2 Matching Stage\nLightGlue was used for evaluating the matches. For a fixed key1, keypoints from key2 were evaluated in four patterns, and the combination with the highest number of matches was adopted. Referring to the IMC2023 2nd place solution ([link](https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/416873)), matches threshold was evaluated in two patterns: 100 and 125.\n\n### 2.3 Transparent Scenes\n#### 2.3.1 Foreground Segmentation - \"bottle\" class is all you need -\n\n![cylinder_matching_initial](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2Ffe4b9b8bdebbbd178b9b63388ce971e4%2Fcylinder_matching_initial.png?generation=1717707155993809&alt=media)\n\nBy closely observing failure cases in transparent scenes, it was noticed that many keypoints appeared in background areas, leading to failures in camera pose estimation. To suppress keypoints in background areas, I investigated foreground extraction methods. Using the DINOv2 Segmenter ([link](https://github.com/facebookresearch/dinov2/blob/main/notebooks/semantic_segmentation.ipynb)), I discovered that the VOC2012 model assigned class5 to the foreground. According to the [link](http://host.robots.ox.ac.uk/pascal/VOC/voc2012/segexamples/index.html), class5 is assigned to \"bottle\". By treating this class as \"transparent\", I hypothesized that I could achieve high-precision segmentation.\n\n![cylinder_segmented](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2F6dd026c50f87131ec407026c2d513bf9%2Fcylinder_segmented.png?generation=1717707179881772&alt=media)\n\n#### 2.3.2 Keypoint Detection with Original Scale\nThe images in transparent scenes were relatively large and all of the same size, so I decided to detect keypoints at the original scale without resizing. Considering VRAM and processing time, keypoints were detected in 1024x1024 grid units. Combining this with the DINOv2 Segmenter, keypoints were detected only in the foreground area.\n\n![keypoints initial](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2F8b95ff498022a69cc46e32d6134a49e3%2Fcylinder_keypoints_initial.png?generation=1717707437917224&alt=media)\n\n![keypoints proposed](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2Fff16f90a5b4ac647de3cbe3349b1cd41%2Fcylinder_keypoints_result.png?generation=1717707455484348&alt=media)\n\n#### 2.3.3 Feature Matching\nGiven the shooting conditions this time, I determined that there was no need to search extensively for matches. Therefore, corresponding points were searched only between corresponding grids during keypoint detection. This significantly reduced the search range.\n\n![cylinder_matching_proposed](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611870%2F206ad587a4bff93ef5eb97ce7367a2ac%2Fcylinder_matching_result.png?generation=1717707266905858&alt=media)\n\n### 2.4 Other Tricks\n#### 2.4.1 Get Pairs Exhaustive\nSince the number of images per scene was not large in IMC2024, exhaustive matching for all pairs was performed instead of searching for pairs with DINOv2 or EfficientNet. This reduced the risk of missing matches due to low embedding-based similarity.\n\n#### 2.4.2 Use *ALL* Images\n\n> I think it's natural for the score to increase if you increase the number of images, as it makes it easier to triangulate.\n\nInspired by Camaro's comment ([link](https://www.kaggle.com/competitions/image-matching-challenge-2024/discussion/494789#2768517)), I wondered if using images beyond those in submission.csv could facilitate easier 3D reconstruction. A simple LB probing revealed that there were images other than those listed in submission.csv in the test data's images folder. Using up to 100 image sets for validation significantly improved the validation score, so it was expected to be effective on the LB as well.\n\n## 3. Results\n\nAfter the competition ended, I conducted several late submissions to evaluate how much each additional technique contributed to improving the leaderboard (LB) score. For local validation, I used a subset of approximately 50 images generated with the following notebook.\n\n[https://www.kaggle.com/code/tmyok1984/imc2024-validation](https://www.kaggle.com/code/tmyok1984/imc2024-validation)\n\n| No, | Transparent trick | Exhaustive matching | Use all images | Private LB | Public LB | Val (avg.) | church | dioscuri | lizard | temple-baalshamin | pond | glass_cup | glass_cylinder |\n| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |\n| 1 |  |  |  | 0.149 | 0.136 | 0.26  | 0.24  | 0.47  | 0.54  | 0.42  | 0.10  | 0.02  | 0.03  |\n| 2 | ✓ |  |  | 0.184 | 0.171 | 0.32  | 0.24  | 0.47  | 0.54  | 0.40  | 0.09  | 0.02  | 0.47  |\n| 3 | ✓ | ✓ |  | 0.186 | 0.176 | 0.34  | 0.24  | 0.52  | 0.51  | 0.42  | 0.17  | 0.02  | 0.47  |\n| 4 | ✓ | ✓ | ✓ | 0.197 | 0.194 | 0.43 | 0.31 | 0.56 | 0.79 | 0.41 | 0.46 | 0.02 | 0.47 |\n\n- Baseline (1) : Private LB=0.149 (around 300th place), Public LB=0.136 (around 300th place)\n- Add Transparent trick (1+2) : Private LB=0.184 (around 8th place), Public LB=0.171 (around 50th place)\n- Add Exhaustive matching (1+2+3) : Private LB=0.186 (around 7th place), Public LB=0.176 (around 30th place)\n- Add ALL images (1+2+3+4) : Private LB=0.197 (4th place), Public LB=0.194 (around 7th place)\n\nThis ablation study revealed the following points:\n\n1. Effect of Transparent trick: Adding the trick for transparent scenes significantly improved the Private LB score from 0.149 to 0.184. This demonstrated that handling transparent scenes greatly contributes to overall performance improvement.\n\n2. Effect of Exhaustive matching: Adding the exhaustive matching method for all pairs further improved the score slightly. This indicated that the risk of missing matches due to low embedding-based similarity was effectively reduced.\n\n3. Effect of Using All Images: Using images beyond those in submission.csv resulted in the most significant score improvement. This confirmed that utilizing additional image information greatly enhances the accuracy of 3D reconstruction.\n\nOverall, handling transparent scenes and using all available images were shown to be key factors for ranking high on the leaderboard.\n\n## 4. Implementation\n\nThis section shares techniques related to implementation.\n\n### 4.1 Multiple Process / Multiple GPUs\nAs highlighted in previous solutions, IMC is characterized by the substantial computational cost of both CPU and GPU processing. Running the CPU and GPU in parallel can potentially double the processing speed. Additionally, using GPUs with two T4 cards allows for parallelizing GPU processing, thereby doubling the processing capacity.\n\nAlthough the importance of these aspects has been mentioned in past solutions, it is very rare for reference code to be made public. As part of my contribution to the community, I have shared the code I used this time. The base part of the implementation can be used beyond IMC, so please refer to it.\n\n[https://www.kaggle.com/code/tmyok1984/imc2024-exp556](https://www.kaggle.com/code/tmyok1984/imc2024-exp556)\n\n### 4.2 Utility Script\n\nWhen using packages not available in the Kaggle environment, offline installation is necessary. However, performing offline package installation in the submission notebook wastes submission time. To solve this problem, I always use the utility script feature. By adding a utility script notebook with pre-installed packages, necessary packages can be installed beforehand, making the submission notebook more efficient. The utility script I used this time is below.\n\n[https://www.kaggle.com/code/tmyok1984/imc2024-install-once](https://www.kaggle.com/code/tmyok1984/imc2024-install-once)\n\nThe following link provides detailed information on how to create utility scripts, so please refer to it.\n\n[https://www.kaggle.com/code/kononenko/pip-install-once](https://www.kaggle.com/code/kononenko/pip-install-once)\n\n## 5. What Did Not Work (For Me)\n- Symmetries-and-repeats: I spent most of the competition period on symmetries-and-repeats. Ultimately, I implemented a combination of a  [Doppelgangers classifier](https://github.com/RuojinCai/doppelgangers) and 3D geometrical verification. Although I adopted it in the final submission and it worked well locally in some cases, the LB consistently worsened during the ablation study with the late submission. I experienced the Concorde effect firsthand.\n- Overlap region：In our 10th place solution for IMC2022, we used DKM-ROI ([link1](https://www.kaggle.com/competitions/image-matching-challenge-2022/discussion/328903), [link2](https://www.kaggle.com/code/tmyok1984/imc2022-ensemble-with-dkm-roi)). This time, there was an excellent publicly available notebook ([link](https://www.kaggle.com/code/nartaa/imc24-overlap-detection)), so I implemented it as well but did not achieve good results. While high matching accuracy was required in IMC2022, in IMC2024, high robustness was required, so its effectiveness was limited. Specifically, if the initial value for finding overlap was incorrect, the subsequent steps would break down significantly. In contrast, the 1st place solution  ([link](https://www.kaggle.com/competitions/image-matching-challenge-2024/discussion/510084)) successfully estimated the ROI using multiple viewpoints robustly and achieved excellent results.\n- Non-Maximum Suppression：This also strongly boosted our LB in IMC2022. However, ALIKED was strong enough this time that even when LoFTR, DeDoDe, and RoMa were ensembled, no improvement was observed in my local validation.\n- COLMAP hyperparameters：Using the help(pycolmap) command, many configurable parameters could be seen. Various parameters were changed, but no significant effect was obtained.",
    "2863385": "I have published my code in GitHub: https://github.com/tmyok/kaggle-image-matching-challenge-2024",
    "2864972": "Thank you for sharing the solution.\nI found it very interesting.\n\nI have a question about the transparent scene section. For transparent scenes, were all pairs matched exhaustively in the same way as for non-transparent scenes?",
    "2860391": "",
    "2859306": ""
  }
}