{
  "id": 515089,
  "title": "[10th] Place Solution for the  Image Matching Challenge 2024 - Hexathlon",
  "url": "/competitions/image-matching-challenge-2024/discussion/515089",
  "author_name": "A.S.TENSHU",
  "post_date": "2024-06-26T21:29:27.736000",
  "votes": 5,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I am thrilled to share my experience participating in the Image Matching Challenge 2024 - Hexathlon. First of all, I wholeheartedly express my appreciation to the Kaggle platform, the organizers <a href=\"https://www.kaggle.com/oldufo\" target=\"_blank\">@oldufo</a> and <a href=\"https://www.kaggle.com/eduardtrulls\" target=\"_blank\">@eduardtrulls</a>, and the sponsors. Secondly, I would like to thank all the kaggler. Throughout the competition, I have found the insightful posts and code in the discussion and code area to be incredibly beneficial for my learning and progress. I am inspired and would now like to share my solution with the community, hoping to offer some assistance and inspiration to other participants. \nBelow is my solution for Image Matching Challenge 2024 - Hexathlon. </p>\n<h2>Background</h2>\n<p>The objective of this contest is to create detailed 3D maps from collections of images captured in a variety of contexts and settings. Participants are tasked with crafting a model capable of producing precise spatial depictions, irrespective of the origin of the images—be they aerial photos from drones, shots within thick woodlands, scenes from the dark of night, or any of the six distinct problem types. \n<a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2024/overview\" target=\"_blank\">https://www.kaggle.com/competitions/image-matching-challenge-2024/overview</a>\n<a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2024/data\" target=\"_blank\">https://www.kaggle.com/competitions/image-matching-challenge-2024/data</a> </p>\n<h2>Method</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18363154%2Fbd4437067acd08a6f58ac81a0525ace4%2F1.jpg?generation=1719433837098548&amp;alt=media\" alt=\"\"></p>\n<h2>Overview</h2>\n<p>My processing flow consists of five parts, Image Rotation Detection, Image Matching, Image Keypoint Extraction, Image Keypoint Matching, and Image Keypoint Fusion. I tried multiple sets of image matching hyperparameters, image matching models, and image keypoint extraction models locally. Considering the limitations of online computational resources and reasoning time, I finally chose Aliked and Affnet+hardnet for keypoint extraction, lightglue and adalam for keypoint matching. For the matching of transparent images, which is the difficult part of the competition, the number of key point matches is increased by cropping the image, thus improving the score.</p>\n<h2>Dataset</h2>\n<p>train dataset </p>\n<table>\n<thead>\n<tr>\n<th>scene</th>\n<th>total count</th>\n<th>size</th>\n<th>max count</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>pond</td>\n<td>1117</td>\n<td>Width 576, Height 1024</td>\n<td>877</td>\n</tr>\n<tr>\n<td>lizard</td>\n<td>711</td>\n<td>Width 580, Height 1024</td>\n<td>284</td>\n</tr>\n<tr>\n<td>church</td>\n<td>110</td>\n<td>Width 768, Height 1024</td>\n<td>92</td>\n</tr>\n<tr>\n<td>dioscuri</td>\n<td>70</td>\n<td>Width 1024, Height 768</td>\n<td>27</td>\n</tr>\n<tr>\n<td>multi-temporal-temple-baalshamin</td>\n<td>68</td>\n<td>Width 1920, Height 1440</td>\n<td>10</td>\n</tr>\n<tr>\n<td>transp_obj_glass_cup</td>\n<td>36</td>\n<td>Width 4608, Height 3288</td>\n<td>36</td>\n</tr>\n<tr>\n<td>transp_obj_glass_cylinder</td>\n<td>36</td>\n<td>Width 6048, Height 4032</td>\n<td>36</td>\n</tr>\n</tbody>\n</table>\n<p><br>\nObserving the training set images based on <a href=\"https://www.kaggle.com/code/moritake04/eda-imc2024-preview-all-images\" target=\"_blank\">EDA</a>, it was found that there were a lot of rotated images in the dioscuri scene, my considerations were to perform rotation detection and use the results of the rotation detection for correction as well as to consider the use of a feature extractor that is not sensitive to rotation. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18363154%2F585b309c91efda9908241fa11501c244%2F2.jpg?generation=1719434974727758&amp;alt=media\" alt=\"\"> </p>\n<h2>Image Retrieval</h2>\n<p>In order to match all pairs of images of a scene in an exhaustive way, the hyperparameters of the open source baseline are tuned. </p>\n<h2>Feature Extraction</h2>\n<p>As mentioned in the dataset analysis section, one of the first things I noticed in the dataset was that certain scenes contained a large number of rotated images, and I tried to solve this problem in two ways: </p>\n<ol>\n<li>use rotation invariant feature matchers (AffNet/HardNet). </li>\n<li>A lightweight orientation detector is used to detect the rotation angle and rotate the image pair accordingly so that the two images have similar orientations. My consideration is that the rotation also affects the matching between the images because the rotation changes the distribution of the pixel points on the x, y axis. </li>\n</ol>\n<p>For keypoint extraction, the combination of ALIKED+LightGlue and AffNet was finally used because ALIKED was a very slow feature extractor in the pre-competition experiments, yet had a better score performance compared to other feature extractors (DISK,SIFT). Therefore, the pre-competition attempts were parameter tuned for ALIKED, including num_features,min_matches,resize_to etc. \nAfter reading the winning solutions of the 2023 competition, I learned that feature extractors such as keynet, affnet, etc. appeared in many of the winning solutions, and I implemented my own affhardnet from the 2023 open source code for feature extraction. Unfortunately, affnet, while rotationally robust, is also a slower feature extractor, which is a considerable challenge for possible model fusion. \nThanks to <a href=\"https://www.kaggle.com/code/motono0223/imc-2024-multi-models-pipeline\" target=\"_blank\">https://www.kaggle.com/code/motono0223/imc-2024-multi-models-pipeline</a> for the parallelization of the image matching and COLMAP processes, as well as Aliked's speed optimization, which made model fusion possible. \nKey feature extraction for transparent images is a important problem, I have a relatively simple treatment in this piece, for transparent images, I found that performing center cropping can improve the matching effect to some extent. </p>\n<h2>Feature Matching</h2>\n<p>Use AdaLAM to match all possible pairs and merge the results with lightglue matches. AdaLAM has a number of parameters that can be tuned, such as  force_seed_mnn, search_expansion, ransac_iters, but due to the inherent stochastic nature, I don't think the gain from tuning these hyperparameters is measurable. </p>\n<h2>Incremental Mapper</h2>\n<p>Based on the same reasoning, I used COLMAP's incremental mapper for the reconstruction, using almost the default parameters except that min_model_size was set to 3. </p>\n<h2>Ensembles</h2>\n<p>Based on past experience, model fusion can lead to significant enhancements, especially with different models. My initial expectation was to have the Aliked+lighteglue matching results fused with Affnet using a different feature extractor, as most of the IMC2023 winners did, and strangely when I chose the three-model fusion, the commit timeout expired. In the end, I was left with the option of submitting the results of the two-model fusion. Referring to the model parameters shared by the IMC2023 contestants, simple parameter adjustments were made and used as the final submission. </p>\n<h2>What didn't work</h2>\n<p>When selecting different HardNet weights, there is no significant change in local validation. Using different image embeddings, there is no notable variation in local validation. </p>\n<h2>Sources</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/motono0223/imc-2024-multi-models-pipeline\" target=\"_blank\">https://www.kaggle.com/code/motono0223/imc-2024-multi-models-pipeline</a></li>\n<li><a href=\"https://www.kaggle.com/code/moritake04/eda-imc2024-preview-all-images\" target=\"_blank\">https://www.kaggle.com/code/moritake04/eda-imc2024-preview-all-images</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417407\" target=\"_blank\">https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417407</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/416873\" target=\"_blank\">https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/416873</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/416918\" target=\"_blank\">https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/416918</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/416816\" target=\"_blank\">https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/416816</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417045\" target=\"_blank\">https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417045</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417002\" target=\"_blank\">https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417002</a></li>\n</ul>\n<h2>Submission</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18363154%2F642076cb9d86bfbc8a5744874a9ea8f9%2F3.png?generation=1719436726030718&amp;alt=media\" alt=\"\"></p>\n<h2>Code</h2>\n<p>I cleared the winning code (removing unnecessary comments and functions) and the link to the cleared winning code is as follows, Note that the code is inherently random, and the results will vary each time the code is run.</p>\n<p>Cleaned Winner Code\n<a href=\"https://www.kaggle.com/code/kirvk013/fork-of-imc-2024-multi-models-pipeline-523c69\" target=\"_blank\">https://www.kaggle.com/code/kirvk013/fork-of-imc-2024-multi-models-pipeline-523c69</a>\nLocal Test Code\n<a href=\"https://www.kaggle.com/code/kirvk013/fork-of-local-cv?scriptVersionId=185634850\" target=\"_blank\">https://www.kaggle.com/code/kirvk013/fork-of-local-cv?scriptVersionId=185634850</a></p>",
  "messages": [
    {
      "id": 2891786,
      "postDate": "2024-06-26T21:29:27.737Z",
      "content": "<p>I am thrilled to share my experience participating in the Image Matching Challenge 2024 - Hexathlon. First of all, I wholeheartedly express my appreciation to the Kaggle platform, the organizers <a href=\"https://www.kaggle.com/oldufo\" target=\"_blank\">@oldufo</a> and <a href=\"https://www.kaggle.com/eduardtrulls\" target=\"_blank\">@eduardtrulls</a>, and the sponsors. Secondly, I would like to thank all the kaggler. Throughout the competition, I have found the insightful posts and code in the discussion and code area to be incredibly beneficial for my learning and progress. I am inspired and would now like to share my solution with the community, hoping to offer some assistance and inspiration to other participants. \nBelow is my solution for Image Matching Challenge 2024 - Hexathlon. </p>\n<h2>Background</h2>\n<p>The objective of this contest is to create detailed 3D maps from collections of images captured in a variety of contexts and settings. Participants are tasked with crafting a model capable of producing precise spatial depictions, irrespective of the origin of the images—be they aerial photos from drones, shots within thick woodlands, scenes from the dark of night, or any of the six distinct problem types. \n<a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2024/overview\" target=\"_blank\">https://www.kaggle.com/competitions/image-matching-challenge-2024/overview</a>\n<a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2024/data\" target=\"_blank\">https://www.kaggle.com/competitions/image-matching-challenge-2024/data</a> </p>\n<h2>Method</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18363154%2Fbd4437067acd08a6f58ac81a0525ace4%2F1.jpg?generation=1719433837098548&amp;alt=media\" alt=\"\"></p>\n<h2>Overview</h2>\n<p>My processing flow consists of five parts, Image Rotation Detection, Image Matching, Image Keypoint Extraction, Image Keypoint Matching, and Image Keypoint Fusion. I tried multiple sets of image matching hyperparameters, image matching models, and image keypoint extraction models locally. Considering the limitations of online computational resources and reasoning time, I finally chose Aliked and Affnet+hardnet for keypoint extraction, lightglue and adalam for keypoint matching. For the matching of transparent images, which is the difficult part of the competition, the number of key point matches is increased by cropping the image, thus improving the score.</p>\n<h2>Dataset</h2>\n<p>train dataset </p>\n<table>\n<thead>\n<tr>\n<th>scene</th>\n<th>total count</th>\n<th>size</th>\n<th>max count</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>pond</td>\n<td>1117</td>\n<td>Width 576, Height 1024</td>\n<td>877</td>\n</tr>\n<tr>\n<td>lizard</td>\n<td>711</td>\n<td>Width 580, Height 1024</td>\n<td>284</td>\n</tr>\n<tr>\n<td>church</td>\n<td>110</td>\n<td>Width 768, Height 1024</td>\n<td>92</td>\n</tr>\n<tr>\n<td>dioscuri</td>\n<td>70</td>\n<td>Width 1024, Height 768</td>\n<td>27</td>\n</tr>\n<tr>\n<td>multi-temporal-temple-baalshamin</td>\n<td>68</td>\n<td>Width 1920, Height 1440</td>\n<td>10</td>\n</tr>\n<tr>\n<td>transp_obj_glass_cup</td>\n<td>36</td>\n<td>Width 4608, Height 3288</td>\n<td>36</td>\n</tr>\n<tr>\n<td>transp_obj_glass_cylinder</td>\n<td>36</td>\n<td>Width 6048, Height 4032</td>\n<td>36</td>\n</tr>\n</tbody>\n</table>\n<p><br>\nObserving the training set images based on <a href=\"https://www.kaggle.com/code/moritake04/eda-imc2024-preview-all-images\" target=\"_blank\">EDA</a>, it was found that there were a lot of rotated images in the dioscuri scene, my considerations were to perform rotation detection and use the results of the rotation detection for correction as well as to consider the use of a feature extractor that is not sensitive to rotation. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18363154%2F585b309c91efda9908241fa11501c244%2F2.jpg?generation=1719434974727758&amp;alt=media\" alt=\"\"> </p>\n<h2>Image Retrieval</h2>\n<p>In order to match all pairs of images of a scene in an exhaustive way, the hyperparameters of the open source baseline are tuned. </p>\n<h2>Feature Extraction</h2>\n<p>As mentioned in the dataset analysis section, one of the first things I noticed in the dataset was that certain scenes contained a large number of rotated images, and I tried to solve this problem in two ways: </p>\n<ol>\n<li>use rotation invariant feature matchers (AffNet/HardNet). </li>\n<li>A lightweight orientation detector is used to detect the rotation angle and rotate the image pair accordingly so that the two images have similar orientations. My consideration is that the rotation also affects the matching between the images because the rotation changes the distribution of the pixel points on the x, y axis. </li>\n</ol>\n<p>For keypoint extraction, the combination of ALIKED+LightGlue and AffNet was finally used because ALIKED was a very slow feature extractor in the pre-competition experiments, yet had a better score performance compared to other feature extractors (DISK,SIFT). Therefore, the pre-competition attempts were parameter tuned for ALIKED, including num_features,min_matches,resize_to etc. \nAfter reading the winning solutions of the 2023 competition, I learned that feature extractors such as keynet, affnet, etc. appeared in many of the winning solutions, and I implemented my own affhardnet from the 2023 open source code for feature extraction. Unfortunately, affnet, while rotationally robust, is also a slower feature extractor, which is a considerable challenge for possible model fusion. \nThanks to <a href=\"https://www.kaggle.com/code/motono0223/imc-2024-multi-models-pipeline\" target=\"_blank\">https://www.kaggle.com/code/motono0223/imc-2024-multi-models-pipeline</a> for the parallelization of the image matching and COLMAP processes, as well as Aliked's speed optimization, which made model fusion possible. \nKey feature extraction for transparent images is a important problem, I have a relatively simple treatment in this piece, for transparent images, I found that performing center cropping can improve the matching effect to some extent. </p>\n<h2>Feature Matching</h2>\n<p>Use AdaLAM to match all possible pairs and merge the results with lightglue matches. AdaLAM has a number of parameters that can be tuned, such as  force_seed_mnn, search_expansion, ransac_iters, but due to the inherent stochastic nature, I don't think the gain from tuning these hyperparameters is measurable. </p>\n<h2>Incremental Mapper</h2>\n<p>Based on the same reasoning, I used COLMAP's incremental mapper for the reconstruction, using almost the default parameters except that min_model_size was set to 3. </p>\n<h2>Ensembles</h2>\n<p>Based on past experience, model fusion can lead to significant enhancements, especially with different models. My initial expectation was to have the Aliked+lighteglue matching results fused with Affnet using a different feature extractor, as most of the IMC2023 winners did, and strangely when I chose the three-model fusion, the commit timeout expired. In the end, I was left with the option of submitting the results of the two-model fusion. Referring to the model parameters shared by the IMC2023 contestants, simple parameter adjustments were made and used as the final submission. </p>\n<h2>What didn't work</h2>\n<p>When selecting different HardNet weights, there is no significant change in local validation. Using different image embeddings, there is no notable variation in local validation. </p>\n<h2>Sources</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/motono0223/imc-2024-multi-models-pipeline\" target=\"_blank\">https://www.kaggle.com/code/motono0223/imc-2024-multi-models-pipeline</a></li>\n<li><a href=\"https://www.kaggle.com/code/moritake04/eda-imc2024-preview-all-images\" target=\"_blank\">https://www.kaggle.com/code/moritake04/eda-imc2024-preview-all-images</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417407\" target=\"_blank\">https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417407</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/416873\" target=\"_blank\">https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/416873</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/416918\" target=\"_blank\">https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/416918</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/416816\" target=\"_blank\">https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/416816</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417045\" target=\"_blank\">https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417045</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417002\" target=\"_blank\">https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417002</a></li>\n</ul>\n<h2>Submission</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18363154%2F642076cb9d86bfbc8a5744874a9ea8f9%2F3.png?generation=1719436726030718&amp;alt=media\" alt=\"\"></p>\n<h2>Code</h2>\n<p>I cleared the winning code (removing unnecessary comments and functions) and the link to the cleared winning code is as follows, Note that the code is inherently random, and the results will vary each time the code is run.</p>\n<p>Cleaned Winner Code\n<a href=\"https://www.kaggle.com/code/kirvk013/fork-of-imc-2024-multi-models-pipeline-523c69\" target=\"_blank\">https://www.kaggle.com/code/kirvk013/fork-of-imc-2024-multi-models-pipeline-523c69</a>\nLocal Test Code\n<a href=\"https://www.kaggle.com/code/kirvk013/fork-of-local-cv?scriptVersionId=185634850\" target=\"_blank\">https://www.kaggle.com/code/kirvk013/fork-of-local-cv?scriptVersionId=185634850</a></p>",
      "rawMarkdown": "I am thrilled to share my experience participating in the Image Matching Challenge 2024 - Hexathlon. First of all, I wholeheartedly express my appreciation to the Kaggle platform, the organizers @oldufo and @eduardtrulls, and the sponsors. Secondly, I would like to thank all the kaggler. Throughout the competition, I have found the insightful posts and code in the discussion and code area to be incredibly beneficial for my learning and progress. I am inspired and would now like to share my solution with the community, hoping to offer some assistance and inspiration to other participants. \nBelow is my solution for Image Matching Challenge 2024 - Hexathlon. \n## Background \nThe objective of this contest is to create detailed 3D maps from collections of images captured in a variety of contexts and settings. Participants are tasked with crafting a model capable of producing precise spatial depictions, irrespective of the origin of the images—be they aerial photos from drones, shots within thick woodlands, scenes from the dark of night, or any of the six distinct problem types. \n[https://www.kaggle.com/competitions/image-matching-challenge-2024/overview](https://www.kaggle.com/competitions/image-matching-challenge-2024/overview)\n[https://www.kaggle.com/competitions/image-matching-challenge-2024/data](https://www.kaggle.com/competitions/image-matching-challenge-2024/data) \n## Method \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18363154%2Fbd4437067acd08a6f58ac81a0525ace4%2F1.jpg?generation=1719433837098548&alt=media)\n## Overview \n\nMy processing flow consists of five parts, Image Rotation Detection, Image Matching, Image Keypoint Extraction, Image Keypoint Matching, and Image Keypoint Fusion. I tried multiple sets of image matching hyperparameters, image matching models, and image keypoint extraction models locally. Considering the limitations of online computational resources and reasoning time, I finally chose Aliked and Affnet+hardnet for keypoint extraction, lightglue and adalam for keypoint matching. For the matching of transparent images, which is the difficult part of the competition, the number of key point matches is increased by cropping the image, thus improving the score.\n\n## Dataset \ntrain dataset \n| scene | total count | size | max count |\n| --- | --- | --- | --- |\n| pond | 1117 | Width 576, Height 1024 | 877 |\n| lizard | 711 | Width 580, Height 1024 | 284 |\n| church | 110 | Width 768, Height 1024 | 92 |\n| dioscuri | 70 | Width 1024, Height 768 | 27 |\n| multi-temporal-temple-baalshamin | 68 | Width 1920, Height 1440 | 10 |\n| transp_obj_glass_cup | 36 | Width 4608, Height 3288 | 36 |\n| transp_obj_glass_cylinder | 36 | Width 6048, Height 4032 | 36 |\n\n <br/>\nObserving the training set images based on [EDA](https://www.kaggle.com/code/moritake04/eda-imc2024-preview-all-images), it was found that there were a lot of rotated images in the dioscuri scene, my considerations were to perform rotation detection and use the results of the rotation detection for correction as well as to consider the use of a feature extractor that is not sensitive to rotation. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18363154%2F585b309c91efda9908241fa11501c244%2F2.jpg?generation=1719434974727758&alt=media) \n## Image Retrieval\nIn order to match all pairs of images of a scene in an exhaustive way, the hyperparameters of the open source baseline are tuned. \n## Feature Extraction \nAs mentioned in the dataset analysis section, one of the first things I noticed in the dataset was that certain scenes contained a large number of rotated images, and I tried to solve this problem in two ways: \n1. use rotation invariant feature matchers (AffNet/HardNet). \n2. A lightweight orientation detector is used to detect the rotation angle and rotate the image pair accordingly so that the two images have similar orientations. My consideration is that the rotation also affects the matching between the images because the rotation changes the distribution of the pixel points on the x, y axis. \n \nFor keypoint extraction, the combination of ALIKED+LightGlue and AffNet was finally used because ALIKED was a very slow feature extractor in the pre-competition experiments, yet had a better score performance compared to other feature extractors (DISK,SIFT). Therefore, the pre-competition attempts were parameter tuned for ALIKED, including num_features,min_matches,resize_to etc. \nAfter reading the winning solutions of the 2023 competition, I learned that feature extractors such as keynet, affnet, etc. appeared in many of the winning solutions, and I implemented my own affhardnet from the 2023 open source code for feature extraction. Unfortunately, affnet, while rotationally robust, is also a slower feature extractor, which is a considerable challenge for possible model fusion. \nThanks to [https://www.kaggle.com/code/motono0223/imc-2024-multi-models-pipeline](https://www.kaggle.com/code/motono0223/imc-2024-multi-models-pipeline) for the parallelization of the image matching and COLMAP processes, as well as Aliked's speed optimization, which made model fusion possible. \nKey feature extraction for transparent images is a important problem, I have a relatively simple treatment in this piece, for transparent images, I found that performing center cropping can improve the matching effect to some extent. \n## Feature Matching \nUse AdaLAM to match all possible pairs and merge the results with lightglue matches. AdaLAM has a number of parameters that can be tuned, such as  force_seed_mnn, search_expansion, ransac_iters, but due to the inherent stochastic nature, I don't think the gain from tuning these hyperparameters is measurable. \n## Incremental Mapper \nBased on the same reasoning, I used COLMAP's incremental mapper for the reconstruction, using almost the default parameters except that min_model_size was set to 3. \n## Ensembles \nBased on past experience, model fusion can lead to significant enhancements, especially with different models. My initial expectation was to have the Aliked+lighteglue matching results fused with Affnet using a different feature extractor, as most of the IMC2023 winners did, and strangely when I chose the three-model fusion, the commit timeout expired. In the end, I was left with the option of submitting the results of the two-model fusion. Referring to the model parameters shared by the IMC2023 contestants, simple parameter adjustments were made and used as the final submission. \n## What didn't work\nWhen selecting different HardNet weights, there is no significant change in local validation. Using different image embeddings, there is no notable variation in local validation. \n## Sources \n- https://www.kaggle.com/code/motono0223/imc-2024-multi-models-pipeline\n- https://www.kaggle.com/code/moritake04/eda-imc2024-preview-all-images\n- https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417407\n- https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/416873\n- https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/416918\n- https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/416816\n- https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417045\n- https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417002\n\n##Submission\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18363154%2F642076cb9d86bfbc8a5744874a9ea8f9%2F3.png?generation=1719436726030718&alt=media)\n##Code\nI cleared the winning code (removing unnecessary comments and functions) and the link to the cleared winning code is as follows, Note that the code is inherently random, and the results will vary each time the code is run.\n\nCleaned Winner Code\nhttps://www.kaggle.com/code/kirvk013/fork-of-imc-2024-multi-models-pipeline-523c69\nLocal Test Code\nhttps://www.kaggle.com/code/kirvk013/fork-of-local-cv?scriptVersionId=185634850\n\n",
      "votes": 5
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2891786": "I am thrilled to share my experience participating in the Image Matching Challenge 2024 - Hexathlon. First of all, I wholeheartedly express my appreciation to the Kaggle platform, the organizers @oldufo and @eduardtrulls, and the sponsors. Secondly, I would like to thank all the kaggler. Throughout the competition, I have found the insightful posts and code in the discussion and code area to be incredibly beneficial for my learning and progress. I am inspired and would now like to share my solution with the community, hoping to offer some assistance and inspiration to other participants. \nBelow is my solution for Image Matching Challenge 2024 - Hexathlon. \n## Background \nThe objective of this contest is to create detailed 3D maps from collections of images captured in a variety of contexts and settings. Participants are tasked with crafting a model capable of producing precise spatial depictions, irrespective of the origin of the images—be they aerial photos from drones, shots within thick woodlands, scenes from the dark of night, or any of the six distinct problem types. \n[https://www.kaggle.com/competitions/image-matching-challenge-2024/overview](https://www.kaggle.com/competitions/image-matching-challenge-2024/overview)\n[https://www.kaggle.com/competitions/image-matching-challenge-2024/data](https://www.kaggle.com/competitions/image-matching-challenge-2024/data) \n## Method \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18363154%2Fbd4437067acd08a6f58ac81a0525ace4%2F1.jpg?generation=1719433837098548&alt=media)\n## Overview \n\nMy processing flow consists of five parts, Image Rotation Detection, Image Matching, Image Keypoint Extraction, Image Keypoint Matching, and Image Keypoint Fusion. I tried multiple sets of image matching hyperparameters, image matching models, and image keypoint extraction models locally. Considering the limitations of online computational resources and reasoning time, I finally chose Aliked and Affnet+hardnet for keypoint extraction, lightglue and adalam for keypoint matching. For the matching of transparent images, which is the difficult part of the competition, the number of key point matches is increased by cropping the image, thus improving the score.\n\n## Dataset \ntrain dataset \n| scene | total count | size | max count |\n| --- | --- | --- | --- |\n| pond | 1117 | Width 576, Height 1024 | 877 |\n| lizard | 711 | Width 580, Height 1024 | 284 |\n| church | 110 | Width 768, Height 1024 | 92 |\n| dioscuri | 70 | Width 1024, Height 768 | 27 |\n| multi-temporal-temple-baalshamin | 68 | Width 1920, Height 1440 | 10 |\n| transp_obj_glass_cup | 36 | Width 4608, Height 3288 | 36 |\n| transp_obj_glass_cylinder | 36 | Width 6048, Height 4032 | 36 |\n\n <br/>\nObserving the training set images based on [EDA](https://www.kaggle.com/code/moritake04/eda-imc2024-preview-all-images), it was found that there were a lot of rotated images in the dioscuri scene, my considerations were to perform rotation detection and use the results of the rotation detection for correction as well as to consider the use of a feature extractor that is not sensitive to rotation. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18363154%2F585b309c91efda9908241fa11501c244%2F2.jpg?generation=1719434974727758&alt=media) \n## Image Retrieval\nIn order to match all pairs of images of a scene in an exhaustive way, the hyperparameters of the open source baseline are tuned. \n## Feature Extraction \nAs mentioned in the dataset analysis section, one of the first things I noticed in the dataset was that certain scenes contained a large number of rotated images, and I tried to solve this problem in two ways: \n1. use rotation invariant feature matchers (AffNet/HardNet). \n2. A lightweight orientation detector is used to detect the rotation angle and rotate the image pair accordingly so that the two images have similar orientations. My consideration is that the rotation also affects the matching between the images because the rotation changes the distribution of the pixel points on the x, y axis. \n \nFor keypoint extraction, the combination of ALIKED+LightGlue and AffNet was finally used because ALIKED was a very slow feature extractor in the pre-competition experiments, yet had a better score performance compared to other feature extractors (DISK,SIFT). Therefore, the pre-competition attempts were parameter tuned for ALIKED, including num_features,min_matches,resize_to etc. \nAfter reading the winning solutions of the 2023 competition, I learned that feature extractors such as keynet, affnet, etc. appeared in many of the winning solutions, and I implemented my own affhardnet from the 2023 open source code for feature extraction. Unfortunately, affnet, while rotationally robust, is also a slower feature extractor, which is a considerable challenge for possible model fusion. \nThanks to [https://www.kaggle.com/code/motono0223/imc-2024-multi-models-pipeline](https://www.kaggle.com/code/motono0223/imc-2024-multi-models-pipeline) for the parallelization of the image matching and COLMAP processes, as well as Aliked's speed optimization, which made model fusion possible. \nKey feature extraction for transparent images is a important problem, I have a relatively simple treatment in this piece, for transparent images, I found that performing center cropping can improve the matching effect to some extent. \n## Feature Matching \nUse AdaLAM to match all possible pairs and merge the results with lightglue matches. AdaLAM has a number of parameters that can be tuned, such as  force_seed_mnn, search_expansion, ransac_iters, but due to the inherent stochastic nature, I don't think the gain from tuning these hyperparameters is measurable. \n## Incremental Mapper \nBased on the same reasoning, I used COLMAP's incremental mapper for the reconstruction, using almost the default parameters except that min_model_size was set to 3. \n## Ensembles \nBased on past experience, model fusion can lead to significant enhancements, especially with different models. My initial expectation was to have the Aliked+lighteglue matching results fused with Affnet using a different feature extractor, as most of the IMC2023 winners did, and strangely when I chose the three-model fusion, the commit timeout expired. In the end, I was left with the option of submitting the results of the two-model fusion. Referring to the model parameters shared by the IMC2023 contestants, simple parameter adjustments were made and used as the final submission. \n## What didn't work\nWhen selecting different HardNet weights, there is no significant change in local validation. Using different image embeddings, there is no notable variation in local validation. \n## Sources \n- https://www.kaggle.com/code/motono0223/imc-2024-multi-models-pipeline\n- https://www.kaggle.com/code/moritake04/eda-imc2024-preview-all-images\n- https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417407\n- https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/416873\n- https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/416918\n- https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/416816\n- https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417045\n- https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417002\n\n##Submission\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18363154%2F642076cb9d86bfbc8a5744874a9ea8f9%2F3.png?generation=1719436726030718&alt=media)\n##Code\nI cleared the winning code (removing unnecessary comments and functions) and the link to the cleared winning code is as follows, Note that the code is inherently random, and the results will vary each time the code is run.\n\nCleaned Winner Code\nhttps://www.kaggle.com/code/kirvk013/fork-of-imc-2024-multi-models-pipeline-523c69\nLocal Test Code\nhttps://www.kaggle.com/code/kirvk013/fork-of-local-cv?scriptVersionId=185634850\n\n"
  }
}