{
  "id": 417126,
  "title": "34th place solution",
  "url": "/competitions/image-matching-challenge-2023/discussion/417126",
  "author_name": "Junghoon Sung",
  "post_date": "2023-06-14T10:04:22.581000",
  "votes": 10,
  "comment_count": 0,
  "views": 0,
  "content": "<h1>34th place solution</h1>\n<p>Thank you to the organizers of the IMC 2023 Challenge and the Kaggle officials for their efforts. This challenge has been very helpful.</p>\n<h1>Overview</h1>\n<p>Our code started with the submission-example baseline provided by the host. The pipeline has the following architecture:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8362662%2F6c55c1a4b581a462acef473aac6f8bff%2FMicrosoftTeams-image%20(7).png?generation=1686732896225750&amp;alt=media\" alt=\"\"></p>\n<h1>keep point</h1>\n<h3>Get image pair shortlist</h3>\n<p>we started looking for a model provided by the timm library in 'get_image_pair_shortlist'. We determined that this problem was similar to a classification problem, so we used the model 'tf_efficientnet_b8'. It was experimentally better than the default 'tf_efficientnet_b7'. And we conducted experiments by changing the sim_th parameter. We found that the performance was the best when sim_th was set to 0.6.</p>\n<h3>Keypoint detect / matching</h3>\n<h4>Model select</h4>\n<p>We only used 'KeyNetAffNetHardNet'. We spent a lot of time in the model selection process. We experimented with model ensembling and parameter tuning for 'LoFTR', 'DISK', 'KeyNetAffNetHardNet', 'KeyNetAffNetSoSNet', 'DKM', 'Silk', and 'RootSIFT'. However, in our personal experiments, 'KeyNetAffNetHardNet' performed the best.</p>\n<h4>Ensemble for multi-resolution</h4>\n<p>We examined datasets/scenes with various resolutions and conducted experiments with various resolutions. Among them, the optimal resolutions were 1696 and 1536. We also experimented with them as single-resolutions, but the performance was not good. We improved the performance by ensembling the results of applying 1696 resolution and 1536 resolution.</p>\n<h4>Find the optimal parameters</h4>\n<p>As we used 'KeyNetAffNetHardNet', we conducted experiments by adjusting numerous related parameters.<br>\nThe final submitted parameters are as follows.</p>\n<h5>detect_features</h5>\n<p><strong><em><em>num_feats = 20000</em></em></strong><br>\n<strong><em><em>matching_alg = 'adalam' # smnn, adalam</em></em></strong><br>\n<strong><em><em>min_matches = 10</em></em></strong></p>\n<h5>matche_features</h5>\n<p><strong><em><em>ransac_iters = 128</em></em></strong><br>\n<strong><em><em>search_expansion = 16</em></em></strong></p>\n<h1>The final score:</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8362662%2Fea0805dba29a4ee979b9804bed6c9af2%2F15361696keynet.png?generation=1686734928396262&amp;alt=media\" alt=\"\"></p>\n<p><strong><em><em>Public LB: 0.451</em></em></strong></p>\n<p><strong><em><em>Private LB: 0.471</em></em></strong></p>\n<h1>Other things I tried</h1>\n<p>The pipeline has the following architecture:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8362662%2F829cc787a3904f3701fab10edffffbfe%2FMicrosoftTeams-image%20(6).png?generation=1686732818842259&amp;alt=media\" alt=\"\"></p>\n<h1>keep point</h1>\n<h3>Get image pair shortlist</h3>\n<p>we used the top-ranked model 'convnext_large_mlp.clip_laion2b_soup_ft_in12k_in1k_384'. It was experimentally better than the efficientnet_b8.</p>\n<h3>Keypoint detect / matching</h3>\n<h4>High-resolution input image</h4>\n<p>First, we examined datasets/scenes with various resolutions. Among them, we confirmed that Cyprus and Wall have ultra-high-resolution images. We determined that if we reduce the image size too much during the resize stage, a lot of information would be lost in these ultra-high-resolution images. Therefore, we checked the GPU memory provided by Kaggle and resized the images to a manageable resolution. (We selected 1920, 1696, and 1848 depending on the experiment.)</p>\n<h4>Model select</h4>\n<p>We only used 'KeyNetAffNetHardNet'.</p>\n<h4>Ensemble for multi-resolution / single-resolution</h4>\n<p>Due to various factors, we designed the pipeline to distinguish architecture and applied multi-resolution for datasets/scenes containing input images with a resolution of 4K or higher to perform more keypoint detection and matches. On the other hand, we applied single-resolution for datasets/scenes with input images of FHD or lower because forcing upsampling of input images causes noise. Since applying multi-resolution can detect many noisy keypoints, we applied single-resolution.</p>\n<p>Additionally, when we previously applied multi-resolution for all datasets/scenes and submitted, we were able to avoid run-timeout issues.</p>\n<h3>Colmap / reconstruction</h3>\n<p>When performing submission with the above options, run-timeout often occurred. We thought it was due to the resolution being too high to resize. However, we thought it would be better to improve other parts such as feature matching accuracy and reconstruction for better results.<br>\nThis idea was inspired by this discussion (<a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/413551)\" target=\"_blank\">https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/413551)</a>. For the train set, I believe that registering new images has a greater impact on performance than accuracy through GBA, so I tried to reduce the execution time by reducing the number of GBA iterations. Therefore, I set the options by doubling or tripling the GBA-related options.<br>\n<code>Mapper.ba_global_images_ratio = 1.1 * 3 (default: 1.1)</code><br>\n<code>Mapper.ba_global_points_ratio = 1.1 * 3 (default: 1.1)</code><br>\n<code>Mapper.ba_global_images_freq = 500 * 3 (default: 500)</code>  <br>\n<code>Mapper.ba_global_points_freq = 250000 * 3 (default: 250000)</code></p>\n<h1>The final score:</h1>\n<p>urban / kyiv-puppet-theater (26 images, 325 pairs) -&gt; mAA=0.896615, mAA_q=0.954154, mAA_t=0.896615</p>\n<p>urban -&gt; mAA=0.896615</p>\n<p>heritage / dioscuri (174 images, 15051 pairs) -&gt; mAA=0.609727, mAA_q=0.692964, mAA_t=0.614750</p>\n<p>heritage / cyprus (30 images, 435 pairs) -&gt; mAA=0.926207, mAA_q=0.976092, mAA_t=0.931494</p>\n<p>heritage / wall (43 images, 903 pairs) -&gt; mAA=0.767331, mAA_q=0.954153, mAA_t=0.771539</p>\n<p>heritage -&gt; mAA=0.767755</p>\n<p>haiper / bike (15 images, 105 pairs) -&gt; mAA=0.929524, mAA_q=1.000000, mAA_t=0.929524</p>\n<p>haiper / chairs (16 images, 120 pairs) -&gt; mAA=0.980000, mAA_q=1.000000, mAA_t=0.980000</p>\n<p>haiper / fountain (23 images, 253 pairs) -&gt; mAA=1.000000, mAA_q=1.000000, mAA_t=1.000000</p>\n<p>haiper -&gt; mAA=0.969841</p>\n<p><strong>Final metric -&gt; mAA=0.878070</strong></p>\n<p><strong>Public LB: 0.448</strong></p>\n<p><strong>Private LB: 0.461</strong></p>\n<h1>Sad story</h1>\n<p>I won't say much. I didn't select the right submission, and ended up receiving a low score for the funny situation. Haha!!!<br>\nIt's okay. This is also part of my skill.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8362662%2Fab00053f1ee12803e39d35388cd6b6de%2FUntitled%20(1).png?generation=1686728295436104&amp;alt=media\" alt=\"\"></p>\n<p>The difference between these two is as follows:</p>\n<p><strong><em><em>mAA : 0.504</em></em></strong></p>\n<p>get_image_pair_shorlist model : <strong><em><em>convnext_large_mlp.clip_laion2b_soup_ft_in12k_in1k_384</em></em></strong></p>\n<p><strong><em><em>mAA : 0.471</em></em></strong></p>\n<p>get_image_pair_shorlist model : <strong><em><em>tf_efficientnet_b8</em></em></strong></p>",
  "messages": [
    {
      "id": 2302060,
      "postDate": "2023-06-14T10:04:22.580Z",
      "content": "<h1>34th place solution</h1>\n<p>Thank you to the organizers of the IMC 2023 Challenge and the Kaggle officials for their efforts. This challenge has been very helpful.</p>\n<h1>Overview</h1>\n<p>Our code started with the submission-example baseline provided by the host. The pipeline has the following architecture:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8362662%2F6c55c1a4b581a462acef473aac6f8bff%2FMicrosoftTeams-image%20(7).png?generation=1686732896225750&amp;alt=media\" alt=\"\"></p>\n<h1>keep point</h1>\n<h3>Get image pair shortlist</h3>\n<p>we started looking for a model provided by the timm library in 'get_image_pair_shortlist'. We determined that this problem was similar to a classification problem, so we used the model 'tf_efficientnet_b8'. It was experimentally better than the default 'tf_efficientnet_b7'. And we conducted experiments by changing the sim_th parameter. We found that the performance was the best when sim_th was set to 0.6.</p>\n<h3>Keypoint detect / matching</h3>\n<h4>Model select</h4>\n<p>We only used 'KeyNetAffNetHardNet'. We spent a lot of time in the model selection process. We experimented with model ensembling and parameter tuning for 'LoFTR', 'DISK', 'KeyNetAffNetHardNet', 'KeyNetAffNetSoSNet', 'DKM', 'Silk', and 'RootSIFT'. However, in our personal experiments, 'KeyNetAffNetHardNet' performed the best.</p>\n<h4>Ensemble for multi-resolution</h4>\n<p>We examined datasets/scenes with various resolutions and conducted experiments with various resolutions. Among them, the optimal resolutions were 1696 and 1536. We also experimented with them as single-resolutions, but the performance was not good. We improved the performance by ensembling the results of applying 1696 resolution and 1536 resolution.</p>\n<h4>Find the optimal parameters</h4>\n<p>As we used 'KeyNetAffNetHardNet', we conducted experiments by adjusting numerous related parameters.<br>\nThe final submitted parameters are as follows.</p>\n<h5>detect_features</h5>\n<p><strong><em><em>num_feats = 20000</em></em></strong><br>\n<strong><em><em>matching_alg = 'adalam' # smnn, adalam</em></em></strong><br>\n<strong><em><em>min_matches = 10</em></em></strong></p>\n<h5>matche_features</h5>\n<p><strong><em><em>ransac_iters = 128</em></em></strong><br>\n<strong><em><em>search_expansion = 16</em></em></strong></p>\n<h1>The final score:</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8362662%2Fea0805dba29a4ee979b9804bed6c9af2%2F15361696keynet.png?generation=1686734928396262&amp;alt=media\" alt=\"\"></p>\n<p><strong><em><em>Public LB: 0.451</em></em></strong></p>\n<p><strong><em><em>Private LB: 0.471</em></em></strong></p>\n<h1>Other things I tried</h1>\n<p>The pipeline has the following architecture:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8362662%2F829cc787a3904f3701fab10edffffbfe%2FMicrosoftTeams-image%20(6).png?generation=1686732818842259&amp;alt=media\" alt=\"\"></p>\n<h1>keep point</h1>\n<h3>Get image pair shortlist</h3>\n<p>we used the top-ranked model 'convnext_large_mlp.clip_laion2b_soup_ft_in12k_in1k_384'. It was experimentally better than the efficientnet_b8.</p>\n<h3>Keypoint detect / matching</h3>\n<h4>High-resolution input image</h4>\n<p>First, we examined datasets/scenes with various resolutions. Among them, we confirmed that Cyprus and Wall have ultra-high-resolution images. We determined that if we reduce the image size too much during the resize stage, a lot of information would be lost in these ultra-high-resolution images. Therefore, we checked the GPU memory provided by Kaggle and resized the images to a manageable resolution. (We selected 1920, 1696, and 1848 depending on the experiment.)</p>\n<h4>Model select</h4>\n<p>We only used 'KeyNetAffNetHardNet'.</p>\n<h4>Ensemble for multi-resolution / single-resolution</h4>\n<p>Due to various factors, we designed the pipeline to distinguish architecture and applied multi-resolution for datasets/scenes containing input images with a resolution of 4K or higher to perform more keypoint detection and matches. On the other hand, we applied single-resolution for datasets/scenes with input images of FHD or lower because forcing upsampling of input images causes noise. Since applying multi-resolution can detect many noisy keypoints, we applied single-resolution.</p>\n<p>Additionally, when we previously applied multi-resolution for all datasets/scenes and submitted, we were able to avoid run-timeout issues.</p>\n<h3>Colmap / reconstruction</h3>\n<p>When performing submission with the above options, run-timeout often occurred. We thought it was due to the resolution being too high to resize. However, we thought it would be better to improve other parts such as feature matching accuracy and reconstruction for better results.<br>\nThis idea was inspired by this discussion (<a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/413551)\" target=\"_blank\">https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/413551)</a>. For the train set, I believe that registering new images has a greater impact on performance than accuracy through GBA, so I tried to reduce the execution time by reducing the number of GBA iterations. Therefore, I set the options by doubling or tripling the GBA-related options.<br>\n<code>Mapper.ba_global_images_ratio = 1.1 * 3 (default: 1.1)</code><br>\n<code>Mapper.ba_global_points_ratio = 1.1 * 3 (default: 1.1)</code><br>\n<code>Mapper.ba_global_images_freq = 500 * 3 (default: 500)</code>  <br>\n<code>Mapper.ba_global_points_freq = 250000 * 3 (default: 250000)</code></p>\n<h1>The final score:</h1>\n<p>urban / kyiv-puppet-theater (26 images, 325 pairs) -&gt; mAA=0.896615, mAA_q=0.954154, mAA_t=0.896615</p>\n<p>urban -&gt; mAA=0.896615</p>\n<p>heritage / dioscuri (174 images, 15051 pairs) -&gt; mAA=0.609727, mAA_q=0.692964, mAA_t=0.614750</p>\n<p>heritage / cyprus (30 images, 435 pairs) -&gt; mAA=0.926207, mAA_q=0.976092, mAA_t=0.931494</p>\n<p>heritage / wall (43 images, 903 pairs) -&gt; mAA=0.767331, mAA_q=0.954153, mAA_t=0.771539</p>\n<p>heritage -&gt; mAA=0.767755</p>\n<p>haiper / bike (15 images, 105 pairs) -&gt; mAA=0.929524, mAA_q=1.000000, mAA_t=0.929524</p>\n<p>haiper / chairs (16 images, 120 pairs) -&gt; mAA=0.980000, mAA_q=1.000000, mAA_t=0.980000</p>\n<p>haiper / fountain (23 images, 253 pairs) -&gt; mAA=1.000000, mAA_q=1.000000, mAA_t=1.000000</p>\n<p>haiper -&gt; mAA=0.969841</p>\n<p><strong>Final metric -&gt; mAA=0.878070</strong></p>\n<p><strong>Public LB: 0.448</strong></p>\n<p><strong>Private LB: 0.461</strong></p>\n<h1>Sad story</h1>\n<p>I won't say much. I didn't select the right submission, and ended up receiving a low score for the funny situation. Haha!!!<br>\nIt's okay. This is also part of my skill.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8362662%2Fab00053f1ee12803e39d35388cd6b6de%2FUntitled%20(1).png?generation=1686728295436104&amp;alt=media\" alt=\"\"></p>\n<p>The difference between these two is as follows:</p>\n<p><strong><em><em>mAA : 0.504</em></em></strong></p>\n<p>get_image_pair_shorlist model : <strong><em><em>convnext_large_mlp.clip_laion2b_soup_ft_in12k_in1k_384</em></em></strong></p>\n<p><strong><em><em>mAA : 0.471</em></em></strong></p>\n<p>get_image_pair_shorlist model : <strong><em><em>tf_efficientnet_b8</em></em></strong></p>",
      "rawMarkdown": "# 34th place solution\nThank you to the organizers of the IMC 2023 Challenge and the Kaggle officials for their efforts. This challenge has been very helpful.\n\n# Overview\nOur code started with the submission-example baseline provided by the host. The pipeline has the following architecture:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8362662%2F6c55c1a4b581a462acef473aac6f8bff%2FMicrosoftTeams-image%20(7).png?generation=1686732896225750&alt=media)\n\n# keep point\n### Get image pair shortlist\nwe started looking for a model provided by the timm library in 'get_image_pair_shortlist'. We determined that this problem was similar to a classification problem, so we used the model 'tf_efficientnet_b8'. It was experimentally better than the default 'tf_efficientnet_b7'. And we conducted experiments by changing the sim_th parameter. We found that the performance was the best when sim_th was set to 0.6.\n\n### Keypoint detect / matching\n#### Model select\nWe only used 'KeyNetAffNetHardNet'. We spent a lot of time in the model selection process. We experimented with model ensembling and parameter tuning for 'LoFTR', 'DISK', 'KeyNetAffNetHardNet', 'KeyNetAffNetSoSNet', 'DKM', 'Silk', and 'RootSIFT'. However, in our personal experiments, 'KeyNetAffNetHardNet' performed the best.\n#### Ensemble for multi-resolution \nWe examined datasets/scenes with various resolutions and conducted experiments with various resolutions. Among them, the optimal resolutions were 1696 and 1536. We also experimented with them as single-resolutions, but the performance was not good. We improved the performance by ensembling the results of applying 1696 resolution and 1536 resolution.\n####  Find the optimal parameters\nAs we used 'KeyNetAffNetHardNet', we conducted experiments by adjusting numerous related parameters.\nThe final submitted parameters are as follows.\n\n##### detect_features \n****num_feats = 20000****\n****matching_alg = 'adalam' # smnn, adalam****\n****min_matches = 10****\n##### matche_features\n****ransac_iters = 128****\n****search_expansion = 16****\n\n# The final score:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8362662%2Fea0805dba29a4ee979b9804bed6c9af2%2F15361696keynet.png?generation=1686734928396262&alt=media)\n\n****Public LB: 0.451****\n\n****Private LB: 0.471****\n\n# Other things I tried\nThe pipeline has the following architecture:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8362662%2F829cc787a3904f3701fab10edffffbfe%2FMicrosoftTeams-image%20(6).png?generation=1686732818842259&alt=media)\n\n# keep point\n### Get image pair shortlist\nwe used the top-ranked model 'convnext_large_mlp.clip_laion2b_soup_ft_in12k_in1k_384'. It was experimentally better than the efficientnet_b8.\n\n### Keypoint detect / matching\n#### High-resolution input image\nFirst, we examined datasets/scenes with various resolutions. Among them, we confirmed that Cyprus and Wall have ultra-high-resolution images. We determined that if we reduce the image size too much during the resize stage, a lot of information would be lost in these ultra-high-resolution images. Therefore, we checked the GPU memory provided by Kaggle and resized the images to a manageable resolution. (We selected 1920, 1696, and 1848 depending on the experiment.)\n#### Model select\nWe only used 'KeyNetAffNetHardNet'.\n#### Ensemble for multi-resolution / single-resolution\nDue to various factors, we designed the pipeline to distinguish architecture and applied multi-resolution for datasets/scenes containing input images with a resolution of 4K or higher to perform more keypoint detection and matches. On the other hand, we applied single-resolution for datasets/scenes with input images of FHD or lower because forcing upsampling of input images causes noise. Since applying multi-resolution can detect many noisy keypoints, we applied single-resolution.\n\nAdditionally, when we previously applied multi-resolution for all datasets/scenes and submitted, we were able to avoid run-timeout issues.\n\n### Colmap / reconstruction\nWhen performing submission with the above options, run-timeout often occurred. We thought it was due to the resolution being too high to resize. However, we thought it would be better to improve other parts such as feature matching accuracy and reconstruction for better results.\nThis idea was inspired by this discussion (https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/413551). For the train set, I believe that registering new images has a greater impact on performance than accuracy through GBA, so I tried to reduce the execution time by reducing the number of GBA iterations. Therefore, I set the options by doubling or tripling the GBA-related options.\n```Mapper.ba_global_images_ratio = 1.1 * 3 (default: 1.1)```\n```Mapper.ba_global_points_ratio = 1.1 * 3 (default: 1.1)```\n```Mapper.ba_global_images_freq = 500 * 3 (default: 500)```  \n```Mapper.ba_global_points_freq = 250000 * 3 (default: 250000)```\n\n# The final score:\nurban / kyiv-puppet-theater (26 images, 325 pairs) -> mAA=0.896615, mAA_q=0.954154, mAA_t=0.896615\n\nurban -> mAA=0.896615\n\nheritage / dioscuri (174 images, 15051 pairs) -> mAA=0.609727, mAA_q=0.692964, mAA_t=0.614750\n\nheritage / cyprus (30 images, 435 pairs) -> mAA=0.926207, mAA_q=0.976092, mAA_t=0.931494\n\nheritage / wall (43 images, 903 pairs) -> mAA=0.767331, mAA_q=0.954153, mAA_t=0.771539\n\nheritage -> mAA=0.767755\n\nhaiper / bike (15 images, 105 pairs) -> mAA=0.929524, mAA_q=1.000000, mAA_t=0.929524\n\nhaiper / chairs (16 images, 120 pairs) -> mAA=0.980000, mAA_q=1.000000, mAA_t=0.980000\n\nhaiper / fountain (23 images, 253 pairs) -> mAA=1.000000, mAA_q=1.000000, mAA_t=1.000000\n\nhaiper -> mAA=0.969841\n\n**Final metric -> mAA=0.878070**\n\n**Public LB: 0.448**\n\n**Private LB: 0.461**\n\n# Sad story\nI won't say much. I didn't select the right submission, and ended up receiving a low score for the funny situation. Haha!!!\nIt's okay. This is also part of my skill.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8362662%2Fab00053f1ee12803e39d35388cd6b6de%2FUntitled%20(1).png?generation=1686728295436104&alt=media)\n\nThe difference between these two is as follows:\n\n****mAA : 0.504****\n\nget_image_pair_shorlist model : ****convnext_large_mlp.clip_laion2b_soup_ft_in12k_in1k_384****\n\n****mAA : 0.471****\n\nget_image_pair_shorlist model : ****tf_efficientnet_b8****",
      "votes": 10
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2302060": "# 34th place solution\nThank you to the organizers of the IMC 2023 Challenge and the Kaggle officials for their efforts. This challenge has been very helpful.\n\n# Overview\nOur code started with the submission-example baseline provided by the host. The pipeline has the following architecture:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8362662%2F6c55c1a4b581a462acef473aac6f8bff%2FMicrosoftTeams-image%20(7).png?generation=1686732896225750&alt=media)\n\n# keep point\n### Get image pair shortlist\nwe started looking for a model provided by the timm library in 'get_image_pair_shortlist'. We determined that this problem was similar to a classification problem, so we used the model 'tf_efficientnet_b8'. It was experimentally better than the default 'tf_efficientnet_b7'. And we conducted experiments by changing the sim_th parameter. We found that the performance was the best when sim_th was set to 0.6.\n\n### Keypoint detect / matching\n#### Model select\nWe only used 'KeyNetAffNetHardNet'. We spent a lot of time in the model selection process. We experimented with model ensembling and parameter tuning for 'LoFTR', 'DISK', 'KeyNetAffNetHardNet', 'KeyNetAffNetSoSNet', 'DKM', 'Silk', and 'RootSIFT'. However, in our personal experiments, 'KeyNetAffNetHardNet' performed the best.\n#### Ensemble for multi-resolution \nWe examined datasets/scenes with various resolutions and conducted experiments with various resolutions. Among them, the optimal resolutions were 1696 and 1536. We also experimented with them as single-resolutions, but the performance was not good. We improved the performance by ensembling the results of applying 1696 resolution and 1536 resolution.\n####  Find the optimal parameters\nAs we used 'KeyNetAffNetHardNet', we conducted experiments by adjusting numerous related parameters.\nThe final submitted parameters are as follows.\n\n##### detect_features \n****num_feats = 20000****\n****matching_alg = 'adalam' # smnn, adalam****\n****min_matches = 10****\n##### matche_features\n****ransac_iters = 128****\n****search_expansion = 16****\n\n# The final score:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8362662%2Fea0805dba29a4ee979b9804bed6c9af2%2F15361696keynet.png?generation=1686734928396262&alt=media)\n\n****Public LB: 0.451****\n\n****Private LB: 0.471****\n\n# Other things I tried\nThe pipeline has the following architecture:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8362662%2F829cc787a3904f3701fab10edffffbfe%2FMicrosoftTeams-image%20(6).png?generation=1686732818842259&alt=media)\n\n# keep point\n### Get image pair shortlist\nwe used the top-ranked model 'convnext_large_mlp.clip_laion2b_soup_ft_in12k_in1k_384'. It was experimentally better than the efficientnet_b8.\n\n### Keypoint detect / matching\n#### High-resolution input image\nFirst, we examined datasets/scenes with various resolutions. Among them, we confirmed that Cyprus and Wall have ultra-high-resolution images. We determined that if we reduce the image size too much during the resize stage, a lot of information would be lost in these ultra-high-resolution images. Therefore, we checked the GPU memory provided by Kaggle and resized the images to a manageable resolution. (We selected 1920, 1696, and 1848 depending on the experiment.)\n#### Model select\nWe only used 'KeyNetAffNetHardNet'.\n#### Ensemble for multi-resolution / single-resolution\nDue to various factors, we designed the pipeline to distinguish architecture and applied multi-resolution for datasets/scenes containing input images with a resolution of 4K or higher to perform more keypoint detection and matches. On the other hand, we applied single-resolution for datasets/scenes with input images of FHD or lower because forcing upsampling of input images causes noise. Since applying multi-resolution can detect many noisy keypoints, we applied single-resolution.\n\nAdditionally, when we previously applied multi-resolution for all datasets/scenes and submitted, we were able to avoid run-timeout issues.\n\n### Colmap / reconstruction\nWhen performing submission with the above options, run-timeout often occurred. We thought it was due to the resolution being too high to resize. However, we thought it would be better to improve other parts such as feature matching accuracy and reconstruction for better results.\nThis idea was inspired by this discussion (https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/413551). For the train set, I believe that registering new images has a greater impact on performance than accuracy through GBA, so I tried to reduce the execution time by reducing the number of GBA iterations. Therefore, I set the options by doubling or tripling the GBA-related options.\n```Mapper.ba_global_images_ratio = 1.1 * 3 (default: 1.1)```\n```Mapper.ba_global_points_ratio = 1.1 * 3 (default: 1.1)```\n```Mapper.ba_global_images_freq = 500 * 3 (default: 500)```  \n```Mapper.ba_global_points_freq = 250000 * 3 (default: 250000)```\n\n# The final score:\nurban / kyiv-puppet-theater (26 images, 325 pairs) -> mAA=0.896615, mAA_q=0.954154, mAA_t=0.896615\n\nurban -> mAA=0.896615\n\nheritage / dioscuri (174 images, 15051 pairs) -> mAA=0.609727, mAA_q=0.692964, mAA_t=0.614750\n\nheritage / cyprus (30 images, 435 pairs) -> mAA=0.926207, mAA_q=0.976092, mAA_t=0.931494\n\nheritage / wall (43 images, 903 pairs) -> mAA=0.767331, mAA_q=0.954153, mAA_t=0.771539\n\nheritage -> mAA=0.767755\n\nhaiper / bike (15 images, 105 pairs) -> mAA=0.929524, mAA_q=1.000000, mAA_t=0.929524\n\nhaiper / chairs (16 images, 120 pairs) -> mAA=0.980000, mAA_q=1.000000, mAA_t=0.980000\n\nhaiper / fountain (23 images, 253 pairs) -> mAA=1.000000, mAA_q=1.000000, mAA_t=1.000000\n\nhaiper -> mAA=0.969841\n\n**Final metric -> mAA=0.878070**\n\n**Public LB: 0.448**\n\n**Private LB: 0.461**\n\n# Sad story\nI won't say much. I didn't select the right submission, and ended up receiving a low score for the funny situation. Haha!!!\nIt's okay. This is also part of my skill.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8362662%2Fab00053f1ee12803e39d35388cd6b6de%2FUntitled%20(1).png?generation=1686728295436104&alt=media)\n\nThe difference between these two is as follows:\n\n****mAA : 0.504****\n\nget_image_pair_shorlist model : ****convnext_large_mlp.clip_laion2b_soup_ft_in12k_in1k_384****\n\n****mAA : 0.471****\n\nget_image_pair_shorlist model : ****tf_efficientnet_b8****"
  }
}