{
  "id": 417186,
  "title": "30th place solution",
  "url": "/competitions/image-matching-challenge-2023/discussion/417186",
  "author_name": "Khoa Ngo",
  "post_date": "2023-06-14T15:49:04.179000",
  "votes": 11,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Congratulations to everyone for the journey we had throughout the competition. Also, thanks to the organizers for bringing the Image Matching Challenge to Kaggle again.<br>\nI will give a brief of my solution to this challenge.</p>\n<h1>Architecture</h1>\n<p>I came to the competition with very limited time and only a little experience from IMC 2022. So, it is likely the first time I walked through the flow of 3D reconstruction.<br>\nI strictly followed the host pipeline, which I split into 3 modules:</p>\n<ul>\n<li>Global descriptors</li>\n<li>Local descriptors (matching)</li>\n<li>Reconstruction</li>\n</ul>\n<p>My main work was focused on improving their efficiency separately.</p>\n<h1>Global descriptors</h1>\n<p>From my point of view, well-trained models on a <strong>landmark</strong> dataset could provide better descriptors than ImageNet pre-trained backbones.<br>\nAs a result, I utilized some of the models that I had trained to compete in the <a href=\"https://www.kaggle.com/competitions/landmark-recognition-2021\" target=\"_blank\">Google Landmark 2021</a>, and then concatenate them to a global descriptor:</p>\n<pre><code>EfficientNetV2-M \nEfficientNetV2-L  \n.                  |--&gt;[concat]--&gt; [fc ]\nResNeSt-       /  \nResNeSt-      /\n</code></pre>\n<h1>Local descriptors</h1>\n<p>This year, competitors are required to perform matching in a strict time interval.<br>\nI focused on <strong>detector-based</strong> (2-stage) methods only, because I thought I could save time on the points detector part (for example, to match <em>(image_i, image_j)</em> and <em>(image_i, image_k)</em>, semi-dense and dense methods will have to \"extract\" <em>image_i</em> two times). When I read other top team solutions, it seemed to be a wrong decision I had made, since such an amount of good matching models are omitted 😭. However, here is the list of methods I tried:<br>\n<strong>Detector</strong>: SuperPoint, SuperPoint + FeatureBooster, KeyNetAffNetHardNet, DISK, SiLK, ALIKE.<br>\n<strong>Matcher</strong>: SuperGlue, GlueStick, SGMNet, AdaLAM.<br>\nWith the detector, I found that <strong>SuperPoint</strong> gave superior results than others.<br>\nWith the matcher, <strong>SuperGlue</strong> showed the best performance in accuracy and efficiency. <strong>GlueStick</strong> is quite good but slower. <strong>SGMNet</strong> is quite fast but lower. I then ensemble keypoints and matches from their predictions and filter out duplicates.</p>\n<h1>Reconstruction</h1>\n<p>I didn't think I could improve much on this, so I only played with <strong>colmap parameters</strong> a bit to find out a (maybe) better combination than the default one.<br>\nSome parameters I changed:</p>\n<pre><code>max_num_trials\nb_images_freq\nb_max_num_iterations\nb_points_freq\nb_max_num_iterations\ninit_num_trials\nmax_num_models\nmin_model_size\n</code></pre>\n<p>I could save a little time with a \"lighter\" combination of parameters while still keeping the accuracy.</p>\n<h1>Final thoughts</h1>\n<p>I guess mine is quite a simple solution, but still give me a silver :D.<br>\nHowever, the knowledge I gained from the competition may be the best I could achieve.<br>\nThank you for your reading and happy Kaggling!</p>",
  "messages": [
    {
      "id": 2302535,
      "postDate": "2023-06-14T15:49:04.180Z",
      "content": "<p>Congratulations to everyone for the journey we had throughout the competition. Also, thanks to the organizers for bringing the Image Matching Challenge to Kaggle again.<br>\nI will give a brief of my solution to this challenge.</p>\n<h1>Architecture</h1>\n<p>I came to the competition with very limited time and only a little experience from IMC 2022. So, it is likely the first time I walked through the flow of 3D reconstruction.<br>\nI strictly followed the host pipeline, which I split into 3 modules:</p>\n<ul>\n<li>Global descriptors</li>\n<li>Local descriptors (matching)</li>\n<li>Reconstruction</li>\n</ul>\n<p>My main work was focused on improving their efficiency separately.</p>\n<h1>Global descriptors</h1>\n<p>From my point of view, well-trained models on a <strong>landmark</strong> dataset could provide better descriptors than ImageNet pre-trained backbones.<br>\nAs a result, I utilized some of the models that I had trained to compete in the <a href=\"https://www.kaggle.com/competitions/landmark-recognition-2021\" target=\"_blank\">Google Landmark 2021</a>, and then concatenate them to a global descriptor:</p>\n<pre><code>EfficientNetV2-M \nEfficientNetV2-L  \n.                  |--&gt;[concat]--&gt; [fc ]\nResNeSt-       /  \nResNeSt-      /\n</code></pre>\n<h1>Local descriptors</h1>\n<p>This year, competitors are required to perform matching in a strict time interval.<br>\nI focused on <strong>detector-based</strong> (2-stage) methods only, because I thought I could save time on the points detector part (for example, to match <em>(image_i, image_j)</em> and <em>(image_i, image_k)</em>, semi-dense and dense methods will have to \"extract\" <em>image_i</em> two times). When I read other top team solutions, it seemed to be a wrong decision I had made, since such an amount of good matching models are omitted 😭. However, here is the list of methods I tried:<br>\n<strong>Detector</strong>: SuperPoint, SuperPoint + FeatureBooster, KeyNetAffNetHardNet, DISK, SiLK, ALIKE.<br>\n<strong>Matcher</strong>: SuperGlue, GlueStick, SGMNet, AdaLAM.<br>\nWith the detector, I found that <strong>SuperPoint</strong> gave superior results than others.<br>\nWith the matcher, <strong>SuperGlue</strong> showed the best performance in accuracy and efficiency. <strong>GlueStick</strong> is quite good but slower. <strong>SGMNet</strong> is quite fast but lower. I then ensemble keypoints and matches from their predictions and filter out duplicates.</p>\n<h1>Reconstruction</h1>\n<p>I didn't think I could improve much on this, so I only played with <strong>colmap parameters</strong> a bit to find out a (maybe) better combination than the default one.<br>\nSome parameters I changed:</p>\n<pre><code>max_num_trials\nb_images_freq\nb_max_num_iterations\nb_points_freq\nb_max_num_iterations\ninit_num_trials\nmax_num_models\nmin_model_size\n</code></pre>\n<p>I could save a little time with a \"lighter\" combination of parameters while still keeping the accuracy.</p>\n<h1>Final thoughts</h1>\n<p>I guess mine is quite a simple solution, but still give me a silver :D.<br>\nHowever, the knowledge I gained from the competition may be the best I could achieve.<br>\nThank you for your reading and happy Kaggling!</p>",
      "rawMarkdown": "Congratulations to everyone for the journey we had throughout the competition. Also, thanks to the organizers for bringing the Image Matching Challenge to Kaggle again.\nI will give a brief of my solution to this challenge.\n# Architecture\nI came to the competition with very limited time and only a little experience from IMC 2022. So, it is likely the first time I walked through the flow of 3D reconstruction.\nI strictly followed the host pipeline, which I split into 3 modules:\n- Global descriptors\n- Local descriptors (matching)\n- Reconstruction\n\nMy main work was focused on improving their efficiency separately.\n# Global descriptors\nFrom my point of view, well-trained models on a **landmark** dataset could provide better descriptors than ImageNet pre-trained backbones.\nAs a result, I utilized some of the models that I had trained to compete in the [Google Landmark 2021](https://www.kaggle.com/competitions/landmark-recognition-2021), and then concatenate them to a global descriptor:\n```\n\nEfficientNetV2-M \\\nEfficientNetV2-L  \\\n.                  |-->[concat]--> [fc 2048]\nResNeSt-200       /  \nResNeSt-269      /\n```\n# Local descriptors\nThis year, competitors are required to perform matching in a strict time interval.\nI focused on **detector-based** (2-stage) methods only, because I thought I could save time on the points detector part (for example, to match *(image_i, image_j)* and *(image_i, image_k)*, semi-dense and dense methods will have to \"extract\" *image_i* two times). When I read other top team solutions, it seemed to be a wrong decision I had made, since such an amount of good matching models are omitted 😭. However, here is the list of methods I tried:\n**Detector**: SuperPoint, SuperPoint + FeatureBooster, KeyNetAffNetHardNet, DISK, SiLK, ALIKE.\n**Matcher**: SuperGlue, GlueStick, SGMNet, AdaLAM.\nWith the detector, I found that **SuperPoint** gave superior results than others.\nWith the matcher, **SuperGlue** showed the best performance in accuracy and efficiency. **GlueStick** is quite good but slower. **SGMNet** is quite fast but lower. I then ensemble keypoints and matches from their predictions and filter out duplicates.\n# Reconstruction\nI didn't think I could improve much on this, so I only played with **colmap parameters** a bit to find out a (maybe) better combination than the default one.\nSome parameters I changed:\n```\nmax_num_trials\nba_global_images_freq\nba_global_max_num_iterations\nba_global_points_freq\nba_local_max_num_iterations\ninit_num_trials\nmax_num_models\nmin_model_size\n```\nI could save a little time with a \"lighter\" combination of parameters while still keeping the accuracy.\n\n# Final thoughts\nI guess mine is quite a simple solution, but still give me a silver :D.\nHowever, the knowledge I gained from the competition may be the best I could achieve.\nThank you for your reading and happy Kaggling!\n",
      "votes": 10
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2302535": "Congratulations to everyone for the journey we had throughout the competition. Also, thanks to the organizers for bringing the Image Matching Challenge to Kaggle again.\nI will give a brief of my solution to this challenge.\n# Architecture\nI came to the competition with very limited time and only a little experience from IMC 2022. So, it is likely the first time I walked through the flow of 3D reconstruction.\nI strictly followed the host pipeline, which I split into 3 modules:\n- Global descriptors\n- Local descriptors (matching)\n- Reconstruction\n\nMy main work was focused on improving their efficiency separately.\n# Global descriptors\nFrom my point of view, well-trained models on a **landmark** dataset could provide better descriptors than ImageNet pre-trained backbones.\nAs a result, I utilized some of the models that I had trained to compete in the [Google Landmark 2021](https://www.kaggle.com/competitions/landmark-recognition-2021), and then concatenate them to a global descriptor:\n```\n\nEfficientNetV2-M \\\nEfficientNetV2-L  \\\n.                  |-->[concat]--> [fc 2048]\nResNeSt-200       /  \nResNeSt-269      /\n```\n# Local descriptors\nThis year, competitors are required to perform matching in a strict time interval.\nI focused on **detector-based** (2-stage) methods only, because I thought I could save time on the points detector part (for example, to match *(image_i, image_j)* and *(image_i, image_k)*, semi-dense and dense methods will have to \"extract\" *image_i* two times). When I read other top team solutions, it seemed to be a wrong decision I had made, since such an amount of good matching models are omitted 😭. However, here is the list of methods I tried:\n**Detector**: SuperPoint, SuperPoint + FeatureBooster, KeyNetAffNetHardNet, DISK, SiLK, ALIKE.\n**Matcher**: SuperGlue, GlueStick, SGMNet, AdaLAM.\nWith the detector, I found that **SuperPoint** gave superior results than others.\nWith the matcher, **SuperGlue** showed the best performance in accuracy and efficiency. **GlueStick** is quite good but slower. **SGMNet** is quite fast but lower. I then ensemble keypoints and matches from their predictions and filter out duplicates.\n# Reconstruction\nI didn't think I could improve much on this, so I only played with **colmap parameters** a bit to find out a (maybe) better combination than the default one.\nSome parameters I changed:\n```\nmax_num_trials\nba_global_images_freq\nba_global_max_num_iterations\nba_global_points_freq\nba_local_max_num_iterations\ninit_num_trials\nmax_num_models\nmin_model_size\n```\nI could save a little time with a \"lighter\" combination of parameters while still keeping the accuracy.\n\n# Final thoughts\nI guess mine is quite a simple solution, but still give me a silver :D.\nHowever, the knowledge I gained from the competition may be the best I could achieve.\nThank you for your reading and happy Kaggling!\n"
  }
}