{
  "id": 417045,
  "title": "6th Place Solution -  [ --- ]AffNetHardNet8 + AdaLAM",
  "url": "/competitions/image-matching-challenge-2023/writeups/maxchen303-6th-place-solution-affnethardnet8-adala",
  "author_name": "",
  "post_date": "2023-06-22T19:47:47.037Z",
  "votes": 29,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I would like to express my thanks to the competition organizers and Kaggle staff for hosting this amazing competition. I would also like to thank the competition hosts (@oldufo , <a href=\"https://www.kaggle.com/eduardtrulls\" target=\"_blank\">@eduardtrulls</a>) for providing helpful materials and a great example submission, which allowed me to quickly get up to speed with the competition.<br>\nParticipating in this competition and IMC2022 has taught me a lot about image matching, Kornia, and SfM. </p>\n<h1>1. Overview</h1>\n<p>My final solution was based on the submission example, but I pushed the limits of KeyNetAffNetHardNet + AdaLAM matcher with Colmap. </p>\n<ul>\n<li>I implemented four local feature detectors that are similar to KeynetAffnetHardnet (<a href=\"https://kornia.readthedocs.io/en/latest/feature.html#kornia.feature.GFTTAffNetHardNet\" target=\"_blank\">Example from Kornia</a>), and each detector extracted local features for all the images (with limited size) in a scene. See more info about other non-learning based keypoint detectors <a href=\"https://kornia.readthedocs.io/en/latest/feature.html\" target=\"_blank\">here</a>. </li>\n<li>Using HardNet8 rather than HardNet although I didn't see a performance difference.</li>\n<li>AdaLAM matcher was used to match all the pairs and get the matches for the scene. Only pair with an average matching distance &lt; 0.5 were kept. 'force_seed_mnn' was set to True and 'ransac_iters' was increased to 256 compared with the submission example.</li>\n<li>The matches from the four detectors were then combined.</li>\n<li>I used USAC_MAGSAC to get the fundamental matrix for each pair and only kept the inlier matches. The results are written to the database as two-view geometry. <code>incremental_mapping</code> is then performed without <code>match_exhaustive</code>.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3964695%2Fbb53c3f1e7ebad9394a673a421b47d11%2Fdetector.png?generation=1686695228576401&amp;alt=media\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3964695%2F051e9413d0ce9cd8357723ba7efab1be%2Ffull_pipeline.png?generation=1686695522836099&amp;alt=media\"></p>\n<h2>1.1. How I reached this solution</h2>\n<p>After running the submission example a few times, I thought that the key to finding a good solution was to identify a good shortlist of image pairs and apply similar solutions from IMC2022. I noticed that the KeyNetAffNetHardNet solution only needed to run the slower local feature detection once for each image, and the matching using AdaLAM matcher was fast enough to match each pair. Therefore, I decided to use KeyNetAffNetHardNet and the matching distance of AdaLAM to find the matching/pairing shortlist. However, with some optimizations, it ended up being the final solution in the last week.</p>\n<h2>1.2. KeyNetAffNetHardNet</h2>\n<p>I was surprised by the performance of KeyNetAffNetHardNet after using AdaLAM to match all possible pairs. With an increased number of features to 8000 and a maximum longer edge of 1600, I was able to achieve a score of 0.414/0.457 (Public/Private).</p>\n<p>During my experimentation, I discovered that the matching distance can be used to determine whether two images have overlapping areas or not. See my <a href=\"https://www.kaggle.com/code/maxchen303/imc2023-test-notes\" target=\"_blank\">Test Notebook</a> for some experiments using KeyNetAffNetHardNet + Adalam for pairing.<br>\nDisabling the Upright option (enabling OriNet) can make the Adalam matching more robust in handling rotated images.</p>\n<h1>2. Implementation details with Colmap</h1>\n<h2>2.1. Using focal length</h2>\n<p>In the submission example, there is a section that extracts the focal length from image exif. However, the \"FocalLengthIn35mmFilm\" property may not be found using the <code>image.get_exif()</code> method. In some cases, the focal length information may exist in the <code>exif_ifd</code> (as described in this <a href=\"https://stackoverflow.com/questions/68033479/how-to-show-all-the-metadata-about-images\" target=\"_blank\">Stack Overflow post</a>). <br>\nTo extract the focal length from the <code>exif_ifd</code>, I used the following code:</p>\n<pre><code>exif = image.getexif()\nexif_ifd = exif.get_ifd()\nexif.update(exif_ifd)\n</code></pre>\n<p>If the focal length is found in the exif, I also set the \"prior_focal_length\" flag to true when adding the camera to the database. This improved the mAA for some scenes (mAA of urban / kyiv-puppet-theater 0.764 -&gt; 0.812).</p>\n<p>According to the  <a href=\"https://colmap.github.io/tutorial.html#database-management\" target=\"_blank\">Colmap tutorial</a>:</p>\n<blockquote>\n  <p>By setting the prior_focal_length flag to 0 or 1, you can give a hint whether the reconstruction algorithm should trust the focal length value.</p>\n</blockquote>\n<h2>2.2. Handling Randomness: Bypass match_exhaustive()</h2>\n<p></p>\n<p></p>\n<h2>2.3. Multiprocessing</h2>\n<p>I discovered that adding matches from other AffNetHardNet detectors could increase the overall mAA score. To fit more similar detectors into the process, I used multiprocessing to ensure that the second CPU core was fully utilized. The final solution takes about 7.5 hours to run using four detectors with edges no longer than 1600.</p>\n<p>To optimize the process, I sorted all the scenes by the number of images before processing, from greater to smaller. For each scene, there were two stages: matching to generate the database and reconstruction from the database. The reconstruction of different scenes is allowed to run in parallel with the matchings. Ideally, the pipeline should look like the following:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3964695%2Fc9cbe4d2f30c6874e597dd5949c2b62e%2Fmultiprocessing.png?generation=1686696885412256&amp;alt=media\"></p>\n<h1>3. Other things I tried</h1>\n<ul>\n<li>After discovering that KeyNetAffNetHardNet was good at finding matching/pairing shortlists, I spent a lot of time exploring pair-wise matching such as LoFTR, SE-LoFTR, and DKMv3. However, these methods were slow even just running on the selected pairs and did not improve the mAAs in my implementations.</li>\n<li>Use AdaLAM matcher with other keypoint detectors such as DISK, SiLK, and ALIKE: I also used OriNet and AffNet to convert the keypoints from these detectors to Lafs and get the HardNet8 descriptors. I then tried matching on the native descriptors, HardNet8 descriptors, or concatenated descriptors. This approach seemed to perform better than matching on the native descriptors without Lafs. If I had more time, I would like to explore this direction further.</li>\n<li>Running local feature detection without resizing by splitting large images into smaller images that could fit into the GPU memory. However, this approach was extremely slow and did not improve performance. I found that the mAAs did not improve beyond a certain resolution. Allowing the longer edge to be 1600 or 2048 yielded similar scores.</li>\n<li>Tuning Colmap mapping options: Too many options and difficult to evaluate the outcomes.</li>\n</ul>\n<h1>4. Local Validation Score</h1>\n<pre><code> / kyiv-puppet-theater ( images,  pairs) -&gt; mAA=., mAA_q=., mAA_t=.\n -&gt; mAA=.\n\n / dioscuri ( images,  pairs) -&gt; mAA=., mAA_q=., mAA_t=.\n / cyprus ( images,  pairs) -&gt; mAA=., mAA_q=., mAA_t=.\n / wall ( images,  pairs) -&gt; mAA=., mAA_q=., mAA_t=.\n -&gt; mAA=.\n\n / bike ( images,  pairs) -&gt; mAA=., mAA_q=., mAA_t=.\n / chairs ( images,  pairs) -&gt; mAA=., mAA_q=., mAA_t=.\n / fountain ( images,  pairs) -&gt; mAA=., mAA_q=., mAA_t=.\n -&gt; mAA=.\n\n metric -&gt; mAA=. (t: . sec.)\n</code></pre>",
  "messages": [
    {
      "id": "2301458",
      "postDate": "06/14/2023 00:38:58",
      "content": "<p>I would like to express my thanks to the competition organizers and Kaggle staff for hosting this amazing competition. I would also like to thank the competition hosts (@oldufo , <a href=\"https://www.kaggle.com/eduardtrulls\" target=\"_blank\">@eduardtrulls</a>) for providing helpful materials and a great example submission, which allowed me to quickly get up to speed with the competition.<br>\nParticipating in this competition and IMC2022 has taught me a lot about image matching, Kornia, and SfM. </p>\n<h1>1. Overview</h1>\n<p>My final solution was based on the submission example, but I pushed the limits of KeyNetAffNetHardNet + AdaLAM matcher with Colmap. </p>\n<ul>\n<li>I implemented four local feature detectors that are similar to KeynetAffnetHardnet (<a href=\"https://kornia.readthedocs.io/en/latest/feature.html#kornia.feature.GFTTAffNetHardNet\" target=\"_blank\">Example from Kornia</a>), and each detector extracted local features for all the images (with limited size) in a scene. See more info about other non-learning based keypoint detectors <a href=\"https://kornia.readthedocs.io/en/latest/feature.html\" target=\"_blank\">here</a>. </li>\n<li>Using HardNet8 rather than HardNet although I didn't see a performance difference.</li>\n<li>AdaLAM matcher was used to match all the pairs and get the matches for the scene. Only pair with an average matching distance &lt; 0.5 were kept. 'force_seed_mnn' was set to True and 'ransac_iters' was increased to 256 compared with the submission example.</li>\n<li>The matches from the four detectors were then combined.</li>\n<li>I used USAC_MAGSAC to get the fundamental matrix for each pair and only kept the inlier matches. The results are written to the database as two-view geometry. <code>incremental_mapping</code> is then performed without <code>match_exhaustive</code>.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3964695%2Fbb53c3f1e7ebad9394a673a421b47d11%2Fdetector.png?generation=1686695228576401&amp;alt=media\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3964695%2F051e9413d0ce9cd8357723ba7efab1be%2Ffull_pipeline.png?generation=1686695522836099&amp;alt=media\"></p>\n<h2>1.1. How I reached this solution</h2>\n<p>After running the submission example a few times, I thought that the key to finding a good solution was to identify a good shortlist of image pairs and apply similar solutions from IMC2022. I noticed that the KeyNetAffNetHardNet solution only needed to run the slower local feature detection once for each image, and the matching using AdaLAM matcher was fast enough to match each pair. Therefore, I decided to use KeyNetAffNetHardNet and the matching distance of AdaLAM to find the matching/pairing shortlist. However, with some optimizations, it ended up being the final solution in the last week.</p>\n<h2>1.2. KeyNetAffNetHardNet</h2>\n<p>I was surprised by the performance of KeyNetAffNetHardNet after using AdaLAM to match all possible pairs. With an increased number of features to 8000 and a maximum longer edge of 1600, I was able to achieve a score of 0.414/0.457 (Public/Private).</p>\n<p>During my experimentation, I discovered that the matching distance can be used to determine whether two images have overlapping areas or not. See my <a href=\"https://www.kaggle.com/code/maxchen303/imc2023-test-notes\" target=\"_blank\">Test Notebook</a> for some experiments using KeyNetAffNetHardNet + Adalam for pairing.<br>\nDisabling the Upright option (enabling OriNet) can make the Adalam matching more robust in handling rotated images.</p>\n<h1>2. Implementation details with Colmap</h1>\n<h2>2.1. Using focal length</h2>\n<p>In the submission example, there is a section that extracts the focal length from image exif. However, the \"FocalLengthIn35mmFilm\" property may not be found using the <code>image.get_exif()</code> method. In some cases, the focal length information may exist in the <code>exif_ifd</code> (as described in this <a href=\"https://stackoverflow.com/questions/68033479/how-to-show-all-the-metadata-about-images\" target=\"_blank\">Stack Overflow post</a>). <br>\nTo extract the focal length from the <code>exif_ifd</code>, I used the following code:</p>\n<pre><code>exif = image.getexif()\nexif_ifd = exif.get_ifd()\nexif.update(exif_ifd)\n</code></pre>\n<p>If the focal length is found in the exif, I also set the \"prior_focal_length\" flag to true when adding the camera to the database. This improved the mAA for some scenes (mAA of urban / kyiv-puppet-theater 0.764 -&gt; 0.812).</p>\n<p>According to the  <a href=\"https://colmap.github.io/tutorial.html#database-management\" target=\"_blank\">Colmap tutorial</a>:</p>\n<blockquote>\n  <p>By setting the prior_focal_length flag to 0 or 1, you can give a hint whether the reconstruction algorithm should trust the focal length value.</p>\n</blockquote>\n<h2>2.2. Handling Randomness: Bypass match_exhaustive()</h2>\n<p></p>\n<p></p>\n<h2>2.3. Multiprocessing</h2>\n<p>I discovered that adding matches from other AffNetHardNet detectors could increase the overall mAA score. To fit more similar detectors into the process, I used multiprocessing to ensure that the second CPU core was fully utilized. The final solution takes about 7.5 hours to run using four detectors with edges no longer than 1600.</p>\n<p>To optimize the process, I sorted all the scenes by the number of images before processing, from greater to smaller. For each scene, there were two stages: matching to generate the database and reconstruction from the database. The reconstruction of different scenes is allowed to run in parallel with the matchings. Ideally, the pipeline should look like the following:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3964695%2Fc9cbe4d2f30c6874e597dd5949c2b62e%2Fmultiprocessing.png?generation=1686696885412256&amp;alt=media\"></p>\n<h1>3. Other things I tried</h1>\n<ul>\n<li>After discovering that KeyNetAffNetHardNet was good at finding matching/pairing shortlists, I spent a lot of time exploring pair-wise matching such as LoFTR, SE-LoFTR, and DKMv3. However, these methods were slow even just running on the selected pairs and did not improve the mAAs in my implementations.</li>\n<li>Use AdaLAM matcher with other keypoint detectors such as DISK, SiLK, and ALIKE: I also used OriNet and AffNet to convert the keypoints from these detectors to Lafs and get the HardNet8 descriptors. I then tried matching on the native descriptors, HardNet8 descriptors, or concatenated descriptors. This approach seemed to perform better than matching on the native descriptors without Lafs. If I had more time, I would like to explore this direction further.</li>\n<li>Running local feature detection without resizing by splitting large images into smaller images that could fit into the GPU memory. However, this approach was extremely slow and did not improve performance. I found that the mAAs did not improve beyond a certain resolution. Allowing the longer edge to be 1600 or 2048 yielded similar scores.</li>\n<li>Tuning Colmap mapping options: Too many options and difficult to evaluate the outcomes.</li>\n</ul>\n<h1>4. Local Validation Score</h1>\n<pre><code> / kyiv-puppet-theater ( images,  pairs) -&gt; mAA=., mAA_q=., mAA_t=.\n -&gt; mAA=.\n\n / dioscuri ( images,  pairs) -&gt; mAA=., mAA_q=., mAA_t=.\n / cyprus ( images,  pairs) -&gt; mAA=., mAA_q=., mAA_t=.\n / wall ( images,  pairs) -&gt; mAA=., mAA_q=., mAA_t=.\n -&gt; mAA=.\n\n / bike ( images,  pairs) -&gt; mAA=., mAA_q=., mAA_t=.\n / chairs ( images,  pairs) -&gt; mAA=., mAA_q=., mAA_t=.\n / fountain ( images,  pairs) -&gt; mAA=., mAA_q=., mAA_t=.\n -&gt; mAA=.\n\n metric -&gt; mAA=. (t: . sec.)\n</code></pre>",
      "rawMarkdown": "I would like to express my thanks to the competition organizers and Kaggle staff for hosting this amazing competition. I would also like to thank the competition hosts (@oldufo , @eduardtrulls) for providing helpful materials and a great example submission, which allowed me to quickly get up to speed with the competition.\nParticipating in this competition and IMC2022 has taught me a lot about image matching, Kornia, and SfM. \n\n# 1. Overview\nMy final solution was based on the submission example, but I pushed the limits of KeyNetAffNetHardNet + AdaLAM matcher with Colmap. \n- I implemented four local feature detectors that are similar to KeynetAffnetHardnet ([Example from Kornia](https://kornia.readthedocs.io/en/latest/feature.html#kornia.feature.GFTTAffNetHardNet)), and each detector extracted local features for all the images (with limited size) in a scene. See more info about other non-learning based keypoint detectors [here](https://kornia.readthedocs.io/en/latest/feature.html). \n- Using HardNet8 rather than HardNet although I didn't see a performance difference.\n- AdaLAM matcher was used to match all the pairs and get the matches for the scene. Only pair with an average matching distance < 0.5 were kept. 'force_seed_mnn' was set to True and 'ransac_iters' was increased to 256 compared with the submission example.\n- The matches from the four detectors were then combined.\n- I used USAC_MAGSAC to get the fundamental matrix for each pair and only kept the inlier matches. The results are written to the database as two-view geometry. `incremental_mapping` is then performed without `match_exhaustive`.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3964695%2Fbb53c3f1e7ebad9394a673a421b47d11%2Fdetector.png?generation=1686695228576401&alt=media\" width=\"480\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3964695%2F051e9413d0ce9cd8357723ba7efab1be%2Ffull_pipeline.png?generation=1686695522836099&alt=media\" width=\"640\">\n\n## 1.1. How I reached this solution\nAfter running the submission example a few times, I thought that the key to finding a good solution was to identify a good shortlist of image pairs and apply similar solutions from IMC2022. I noticed that the KeyNetAffNetHardNet solution only needed to run the slower local feature detection once for each image, and the matching using AdaLAM matcher was fast enough to match each pair. Therefore, I decided to use KeyNetAffNetHardNet and the matching distance of AdaLAM to find the matching/pairing shortlist. However, with some optimizations, it ended up being the final solution in the last week.\n\n## 1.2. KeyNetAffNetHardNet\nI was surprised by the performance of KeyNetAffNetHardNet after using AdaLAM to match all possible pairs. With an increased number of features to 8000 and a maximum longer edge of 1600, I was able to achieve a score of 0.414/0.457 (Public/Private).\n\nDuring my experimentation, I discovered that the matching distance can be used to determine whether two images have overlapping areas or not. See my [Test Notebook](https://www.kaggle.com/code/maxchen303/imc2023-test-notes) for some experiments using KeyNetAffNetHardNet + Adalam for pairing.\nDisabling the Upright option (enabling OriNet) can make the Adalam matching more robust in handling rotated images.\n\n# 2. Implementation details with Colmap\n## 2.1. Using focal length\nIn the submission example, there is a section that extracts the focal length from image exif. However, the \"FocalLengthIn35mmFilm\" property may not be found using the `image.get_exif()` method. In some cases, the focal length information may exist in the `exif_ifd` (as described in this [Stack Overflow post](https://stackoverflow.com/questions/68033479/how-to-show-all-the-metadata-about-images)). \nTo extract the focal length from the `exif_ifd`, I used the following code:\n\n```python\nexif = image.getexif()\nexif_ifd = exif.get_ifd(0x8769)\nexif.update(exif_ifd)\n```\nIf the focal length is found in the exif, I also set the \"prior_focal_length\" flag to true when adding the camera to the database. This improved the mAA for some scenes (mAA of urban / kyiv-puppet-theater 0.764 -> 0.812).\n\nAccording to the  [Colmap tutorial](https://colmap.github.io/tutorial.html#database-management):\n>By setting the prior_focal_length flag to 0 or 1, you can give a hint whether the reconstruction algorithm should trust the focal length value.\n\n## 2.2. Handling Randomness: Bypass match_exhaustive()\n~~ After testing my solution many times on the training dataset, I noticed that the evaluation results were random even when using the same code. I found that the source of this randomness was the `match_exhaustive()` function, which runs RANSAC on the matches and generates [two-view geometry](https://github.com/colmap/colmap/blob/dev/src/colmap/estimators/two_view_geometry.h) for each image pair in the database. ~~\n\n~~ To address this issue, I decided to bypass the `match_exhaustive` function and use USAC_MAGSAC from OpenCV to find the fundamental matrix for each pair. I then wrote the two_view_geometry directly into the database. With this change, I was able to get deterministic mAAs when running the same code. By tuning the USAC_MAGSAC parameters, I was able to ensure the best score for reconstructing the same matches.  ~~\n\n## 2.3. Multiprocessing\nI discovered that adding matches from other AffNetHardNet detectors could increase the overall mAA score. To fit more similar detectors into the process, I used multiprocessing to ensure that the second CPU core was fully utilized. The final solution takes about 7.5 hours to run using four detectors with edges no longer than 1600.\n\nTo optimize the process, I sorted all the scenes by the number of images before processing, from greater to smaller. For each scene, there were two stages: matching to generate the database and reconstruction from the database. The reconstruction of different scenes is allowed to run in parallel with the matchings. Ideally, the pipeline should look like the following:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3964695%2Fc9cbe4d2f30c6874e597dd5949c2b62e%2Fmultiprocessing.png?generation=1686696885412256&alt=media\" width=\"640\">\n\n# 3. Other things I tried\n- After discovering that KeyNetAffNetHardNet was good at finding matching/pairing shortlists, I spent a lot of time exploring pair-wise matching such as LoFTR, SE-LoFTR, and DKMv3. However, these methods were slow even just running on the selected pairs and did not improve the mAAs in my implementations.\n- Use AdaLAM matcher with other keypoint detectors such as DISK, SiLK, and ALIKE: I also used OriNet and AffNet to convert the keypoints from these detectors to Lafs and get the HardNet8 descriptors. I then tried matching on the native descriptors, HardNet8 descriptors, or concatenated descriptors. This approach seemed to perform better than matching on the native descriptors without Lafs. If I had more time, I would like to explore this direction further.\n- Running local feature detection without resizing by splitting large images into smaller images that could fit into the GPU memory. However, this approach was extremely slow and did not improve performance. I found that the mAAs did not improve beyond a certain resolution. Allowing the longer edge to be 1600 or 2048 yielded similar scores.\n- Tuning Colmap mapping options: Too many options and difficult to evaluate the outcomes.\n\n# 4. Local Validation Score \n```\nurban / kyiv-puppet-theater (26 images, 325 pairs) -> mAA=0.904923, mAA_q=0.931077, mAA_t=0.914154\nurban -> mAA=0.904923\n\nheritage / dioscuri (174 images, 15051 pairs) -> mAA=0.906996, mAA_q=0.982559, mAA_t=0.910856\nheritage / cyprus (30 images, 435 pairs) -> mAA=0.855172, mAA_q=0.868966, mAA_t=0.865977\nheritage / wall (43 images, 903 pairs) -> mAA=0.484496, mAA_q=0.820377, mAA_t=0.499336\nheritage -> mAA=0.748888\n\nhaiper / bike (15 images, 105 pairs) -> mAA=0.921905, mAA_q=0.999048, mAA_t=0.921905\nhaiper / chairs (16 images, 120 pairs) -> mAA=0.979167, mAA_q=0.999167, mAA_t=0.979167\nhaiper / fountain (23 images, 253 pairs) -> mAA=0.999605, mAA_q=1.000000, mAA_t=0.999605\nhaiper -> mAA=0.966892\n\nFinal metric -> mAA=0.873568 (t: 1.7704436779022217 sec.)\n```",
      "votes": null
    },
    {
      "id": "2301620",
      "postDate": "06/14/2023 04:23:47",
      "content": "<p>Thank you for sharing your brilliant ideas and congrats to getting a solo gold! The tricks to use camera focal length, cache the keypoints and use MAGSAC directly are really impressive. </p>",
      "rawMarkdown": "Thank you for sharing your brilliant ideas and congrats to getting a solo gold! The tricks to use camera focal length, cache the keypoints and use MAGSAC directly are really impressive.",
      "votes": null
    },
    {
      "id": "2301947",
      "postDate": "06/14/2023 08:50:18",
      "content": "<p>Wow, man! I never thought that kornia local features could be pushed to this limit!</p>",
      "rawMarkdown": "Wow, man! I never thought that kornia local features could be pushed to this limit!",
      "votes": null
    },
    {
      "id": "2304141",
      "postDate": "06/15/2023 17:57:19",
      "content": "<p>I have a question though, theoretically how does including the focal length in the Colmap parameter improve the accuracy of calculating the camera poses?</p>",
      "rawMarkdown": "I have a question though, theoretically how does including the focal length in the Colmap parameter improve the accuracy of calculating the camera poses?",
      "votes": null
    },
    {
      "id": "2304241",
      "postDate": "06/15/2023 20:35:48",
      "content": "<p>I am actually new to SfM but I found something from the Colmap source code that may answer your question:</p>\n<ol>\n<li>When the \"estimate_focal_length\" option is True,  the absolute pose estimator will quadratically sample some <a href=\"https://github.com/colmap/colmap/blob/43de802cfb3ed2bd155150e7e5e3e8c8dd5aaa3e/src/estimators/pose.h#LL56C1-L56C1\" target=\"_blank\">(30 by default)</a> scale factors and select the best one to correct the camera model <a href=\"https://github.com/colmap/colmap/blob/43de802cfb3ed2bd155150e7e5e3e8c8dd5aaa3e/src/estimators/pose.cc#L88\" target=\"_blank\">(Source Code)</a>. I think if the given initial focal length is off by too much and the number of samples is not enough, the \"best\" estimated camera model won't be good enough for the <a href=\"https://github.com/colmap/colmap/blob/43de802cfb3ed2bd155150e7e5e3e8c8dd5aaa3e/src/estimators/pose.cc#L50\" target=\"_blank\">absolute pose estimation</a> and refinement.</li>\n<li>Setting \"prior_focal_length\" flag to 1  will disable the sampling (<em>estimate_focal_length</em>) mentioned above and only use the given camera model in absolute pose estimation <a href=\"https://github.com/colmap/colmap/blob/43de802cfb3ed2bd155150e7e5e3e8c8dd5aaa3e/src/sfm/incremental_mapper.cc#L430\" target=\"_blank\">(Source Code)</a>. This won't affect the refinement options.</li>\n<li>Images from a camera with a prior focal length will have a higher priority to be selected as the initial pair. <a href=\"https://github.com/colmap/colmap/blob/43de802cfb3ed2bd155150e7e5e3e8c8dd5aaa3e/src/sfm/incremental_mapper.cc#L845\" target=\"_blank\">Source Code</a>.</li>\n</ol>",
      "rawMarkdown": "I am actually new to SfM but I found something from the Colmap source code that may answer your question:\n1. When the \"estimate_focal_length\" option is True,  the absolute pose estimator will quadratically sample some [(30 by default)](https://github.com/colmap/colmap/blob/43de802cfb3ed2bd155150e7e5e3e8c8dd5aaa3e/src/estimators/pose.h#LL56C1-L56C1) scale factors and select the best one to correct the camera model [(Source Code)](https://github.com/colmap/colmap/blob/43de802cfb3ed2bd155150e7e5e3e8c8dd5aaa3e/src/estimators/pose.cc#L88). I think if the given initial focal length is off by too much and the number of samples is not enough, the \"best\" estimated camera model won't be good enough for the [absolute pose estimation](https://github.com/colmap/colmap/blob/43de802cfb3ed2bd155150e7e5e3e8c8dd5aaa3e/src/estimators/pose.cc#L50) and refinement.\n2. Setting \"prior_focal_length\" flag to 1  will disable the sampling (*estimate_focal_length*) mentioned above and only use the given camera model in absolute pose estimation [(Source Code)](https://github.com/colmap/colmap/blob/43de802cfb3ed2bd155150e7e5e3e8c8dd5aaa3e/src/sfm/incremental_mapper.cc#L430). This won't affect the refinement options.\n3. Images from a camera with a prior focal length will have a higher priority to be selected as the initial pair. [Source Code] (https://github.com/colmap/colmap/blob/43de802cfb3ed2bd155150e7e5e3e8c8dd5aaa3e/src/sfm/incremental_mapper.cc#L845).",
      "votes": null
    },
    {
      "id": "2304887",
      "postDate": "06/16/2023 09:42:55",
      "content": "<p>Some Images contain Exif data which you can get the true value of the focal length… Adding a prior in optimization with the correct value of the focal length would help into making the optimization easier to converge and more constrained to local minimum… <br>\nBut here one need to be careful as the Exif data might be manipulated as in the case of Urban-Kyiv, which the true values of the Exif are not correct due to scale manipulation of the images.</p>",
      "rawMarkdown": "Some Images contain Exif data which you can get the true value of the focal length... Adding a prior in optimization with the correct value of the focal length would help into making the optimization easier to converge and more constrained to local minimum... \nBut here one need to be careful as the Exif data might be manipulated as in the case of Urban-Kyiv, which the true values of the Exif are not correct due to scale manipulation of the images.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2301620,
      "author_name": "anonymousyuxiang",
      "author_url": "",
      "post_date": "06/14/2023 04:23:47",
      "content": "<p>Thank you for sharing your brilliant ideas and congrats to getting a solo gold! The tricks to use camera focal length, cache the keypoints and use MAGSAC directly are really impressive. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2301947,
      "author_name": "oldufo",
      "author_url": "",
      "post_date": "06/14/2023 08:50:18",
      "content": "<p>Wow, man! I never thought that kornia local features could be pushed to this limit!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2304141,
      "author_name": "anonymousyuxiang",
      "author_url": "",
      "post_date": "06/15/2023 17:57:19",
      "content": "<p>I have a question though, theoretically how does including the focal length in the Colmap parameter improve the accuracy of calculating the camera poses?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2304241,
          "author_name": "maxchen303",
          "author_url": "",
          "post_date": "06/15/2023 20:35:48",
          "content": "<p>I am actually new to SfM but I found something from the Colmap source code that may answer your question:</p>\n<ol>\n<li>When the \"estimate_focal_length\" option is True,  the absolute pose estimator will quadratically sample some <a href=\"https://github.com/colmap/colmap/blob/43de802cfb3ed2bd155150e7e5e3e8c8dd5aaa3e/src/estimators/pose.h#LL56C1-L56C1\" target=\"_blank\">(30 by default)</a> scale factors and select the best one to correct the camera model <a href=\"https://github.com/colmap/colmap/blob/43de802cfb3ed2bd155150e7e5e3e8c8dd5aaa3e/src/estimators/pose.cc#L88\" target=\"_blank\">(Source Code)</a>. I think if the given initial focal length is off by too much and the number of samples is not enough, the \"best\" estimated camera model won't be good enough for the <a href=\"https://github.com/colmap/colmap/blob/43de802cfb3ed2bd155150e7e5e3e8c8dd5aaa3e/src/estimators/pose.cc#L50\" target=\"_blank\">absolute pose estimation</a> and refinement.</li>\n<li>Setting \"prior_focal_length\" flag to 1  will disable the sampling (<em>estimate_focal_length</em>) mentioned above and only use the given camera model in absolute pose estimation <a href=\"https://github.com/colmap/colmap/blob/43de802cfb3ed2bd155150e7e5e3e8c8dd5aaa3e/src/sfm/incremental_mapper.cc#L430\" target=\"_blank\">(Source Code)</a>. This won't affect the refinement options.</li>\n<li>Images from a camera with a prior focal length will have a higher priority to be selected as the initial pair. <a href=\"https://github.com/colmap/colmap/blob/43de802cfb3ed2bd155150e7e5e3e8c8dd5aaa3e/src/sfm/incremental_mapper.cc#L845\" target=\"_blank\">Source Code</a>.</li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2304887,
          "author_name": "jaafarmahmoud1",
          "author_url": "",
          "post_date": "06/16/2023 09:42:55",
          "content": "<p>Some Images contain Exif data which you can get the true value of the focal length… Adding a prior in optimization with the correct value of the focal length would help into making the optimization easier to converge and more constrained to local minimum… <br>\nBut here one need to be careful as the Exif data might be manipulated as in the case of Urban-Kyiv, which the true values of the Exif are not correct due to scale manipulation of the images.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2301458": "I would like to express my thanks to the competition organizers and Kaggle staff for hosting this amazing competition. I would also like to thank the competition hosts (@oldufo , @eduardtrulls) for providing helpful materials and a great example submission, which allowed me to quickly get up to speed with the competition.\nParticipating in this competition and IMC2022 has taught me a lot about image matching, Kornia, and SfM. \n\n# 1. Overview\nMy final solution was based on the submission example, but I pushed the limits of KeyNetAffNetHardNet + AdaLAM matcher with Colmap. \n- I implemented four local feature detectors that are similar to KeynetAffnetHardnet ([Example from Kornia](https://kornia.readthedocs.io/en/latest/feature.html#kornia.feature.GFTTAffNetHardNet)), and each detector extracted local features for all the images (with limited size) in a scene. See more info about other non-learning based keypoint detectors [here](https://kornia.readthedocs.io/en/latest/feature.html). \n- Using HardNet8 rather than HardNet although I didn't see a performance difference.\n- AdaLAM matcher was used to match all the pairs and get the matches for the scene. Only pair with an average matching distance < 0.5 were kept. 'force_seed_mnn' was set to True and 'ransac_iters' was increased to 256 compared with the submission example.\n- The matches from the four detectors were then combined.\n- I used USAC_MAGSAC to get the fundamental matrix for each pair and only kept the inlier matches. The results are written to the database as two-view geometry. `incremental_mapping` is then performed without `match_exhaustive`.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3964695%2Fbb53c3f1e7ebad9394a673a421b47d11%2Fdetector.png?generation=1686695228576401&alt=media\" width=\"480\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3964695%2F051e9413d0ce9cd8357723ba7efab1be%2Ffull_pipeline.png?generation=1686695522836099&alt=media\" width=\"640\">\n\n## 1.1. How I reached this solution\nAfter running the submission example a few times, I thought that the key to finding a good solution was to identify a good shortlist of image pairs and apply similar solutions from IMC2022. I noticed that the KeyNetAffNetHardNet solution only needed to run the slower local feature detection once for each image, and the matching using AdaLAM matcher was fast enough to match each pair. Therefore, I decided to use KeyNetAffNetHardNet and the matching distance of AdaLAM to find the matching/pairing shortlist. However, with some optimizations, it ended up being the final solution in the last week.\n\n## 1.2. KeyNetAffNetHardNet\nI was surprised by the performance of KeyNetAffNetHardNet after using AdaLAM to match all possible pairs. With an increased number of features to 8000 and a maximum longer edge of 1600, I was able to achieve a score of 0.414/0.457 (Public/Private).\n\nDuring my experimentation, I discovered that the matching distance can be used to determine whether two images have overlapping areas or not. See my [Test Notebook](https://www.kaggle.com/code/maxchen303/imc2023-test-notes) for some experiments using KeyNetAffNetHardNet + Adalam for pairing.\nDisabling the Upright option (enabling OriNet) can make the Adalam matching more robust in handling rotated images.\n\n# 2. Implementation details with Colmap\n## 2.1. Using focal length\nIn the submission example, there is a section that extracts the focal length from image exif. However, the \"FocalLengthIn35mmFilm\" property may not be found using the `image.get_exif()` method. In some cases, the focal length information may exist in the `exif_ifd` (as described in this [Stack Overflow post](https://stackoverflow.com/questions/68033479/how-to-show-all-the-metadata-about-images)). \nTo extract the focal length from the `exif_ifd`, I used the following code:\n\n```python\nexif = image.getexif()\nexif_ifd = exif.get_ifd(0x8769)\nexif.update(exif_ifd)\n```\nIf the focal length is found in the exif, I also set the \"prior_focal_length\" flag to true when adding the camera to the database. This improved the mAA for some scenes (mAA of urban / kyiv-puppet-theater 0.764 -> 0.812).\n\nAccording to the  [Colmap tutorial](https://colmap.github.io/tutorial.html#database-management):\n>By setting the prior_focal_length flag to 0 or 1, you can give a hint whether the reconstruction algorithm should trust the focal length value.\n\n## 2.2. Handling Randomness: Bypass match_exhaustive()\n~~ After testing my solution many times on the training dataset, I noticed that the evaluation results were random even when using the same code. I found that the source of this randomness was the `match_exhaustive()` function, which runs RANSAC on the matches and generates [two-view geometry](https://github.com/colmap/colmap/blob/dev/src/colmap/estimators/two_view_geometry.h) for each image pair in the database. ~~\n\n~~ To address this issue, I decided to bypass the `match_exhaustive` function and use USAC_MAGSAC from OpenCV to find the fundamental matrix for each pair. I then wrote the two_view_geometry directly into the database. With this change, I was able to get deterministic mAAs when running the same code. By tuning the USAC_MAGSAC parameters, I was able to ensure the best score for reconstructing the same matches.  ~~\n\n## 2.3. Multiprocessing\nI discovered that adding matches from other AffNetHardNet detectors could increase the overall mAA score. To fit more similar detectors into the process, I used multiprocessing to ensure that the second CPU core was fully utilized. The final solution takes about 7.5 hours to run using four detectors with edges no longer than 1600.\n\nTo optimize the process, I sorted all the scenes by the number of images before processing, from greater to smaller. For each scene, there were two stages: matching to generate the database and reconstruction from the database. The reconstruction of different scenes is allowed to run in parallel with the matchings. Ideally, the pipeline should look like the following:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3964695%2Fc9cbe4d2f30c6874e597dd5949c2b62e%2Fmultiprocessing.png?generation=1686696885412256&alt=media\" width=\"640\">\n\n# 3. Other things I tried\n- After discovering that KeyNetAffNetHardNet was good at finding matching/pairing shortlists, I spent a lot of time exploring pair-wise matching such as LoFTR, SE-LoFTR, and DKMv3. However, these methods were slow even just running on the selected pairs and did not improve the mAAs in my implementations.\n- Use AdaLAM matcher with other keypoint detectors such as DISK, SiLK, and ALIKE: I also used OriNet and AffNet to convert the keypoints from these detectors to Lafs and get the HardNet8 descriptors. I then tried matching on the native descriptors, HardNet8 descriptors, or concatenated descriptors. This approach seemed to perform better than matching on the native descriptors without Lafs. If I had more time, I would like to explore this direction further.\n- Running local feature detection without resizing by splitting large images into smaller images that could fit into the GPU memory. However, this approach was extremely slow and did not improve performance. I found that the mAAs did not improve beyond a certain resolution. Allowing the longer edge to be 1600 or 2048 yielded similar scores.\n- Tuning Colmap mapping options: Too many options and difficult to evaluate the outcomes.\n\n# 4. Local Validation Score \n```\nurban / kyiv-puppet-theater (26 images, 325 pairs) -> mAA=0.904923, mAA_q=0.931077, mAA_t=0.914154\nurban -> mAA=0.904923\n\nheritage / dioscuri (174 images, 15051 pairs) -> mAA=0.906996, mAA_q=0.982559, mAA_t=0.910856\nheritage / cyprus (30 images, 435 pairs) -> mAA=0.855172, mAA_q=0.868966, mAA_t=0.865977\nheritage / wall (43 images, 903 pairs) -> mAA=0.484496, mAA_q=0.820377, mAA_t=0.499336\nheritage -> mAA=0.748888\n\nhaiper / bike (15 images, 105 pairs) -> mAA=0.921905, mAA_q=0.999048, mAA_t=0.921905\nhaiper / chairs (16 images, 120 pairs) -> mAA=0.979167, mAA_q=0.999167, mAA_t=0.979167\nhaiper / fountain (23 images, 253 pairs) -> mAA=0.999605, mAA_q=1.000000, mAA_t=0.999605\nhaiper -> mAA=0.966892\n\nFinal metric -> mAA=0.873568 (t: 1.7704436779022217 sec.)\n```",
    "2301620": "Thank you for sharing your brilliant ideas and congrats to getting a solo gold! The tricks to use camera focal length, cache the keypoints and use MAGSAC directly are really impressive.",
    "2301947": "Wow, man! I never thought that kornia local features could be pushed to this limit!",
    "2304141": "I have a question though, theoretically how does including the focal length in the Colmap parameter improve the accuracy of calculating the camera poses?",
    "2304241": "I am actually new to SfM but I found something from the Colmap source code that may answer your question:\n1. When the \"estimate_focal_length\" option is True,  the absolute pose estimator will quadratically sample some [(30 by default)](https://github.com/colmap/colmap/blob/43de802cfb3ed2bd155150e7e5e3e8c8dd5aaa3e/src/estimators/pose.h#LL56C1-L56C1) scale factors and select the best one to correct the camera model [(Source Code)](https://github.com/colmap/colmap/blob/43de802cfb3ed2bd155150e7e5e3e8c8dd5aaa3e/src/estimators/pose.cc#L88). I think if the given initial focal length is off by too much and the number of samples is not enough, the \"best\" estimated camera model won't be good enough for the [absolute pose estimation](https://github.com/colmap/colmap/blob/43de802cfb3ed2bd155150e7e5e3e8c8dd5aaa3e/src/estimators/pose.cc#L50) and refinement.\n2. Setting \"prior_focal_length\" flag to 1  will disable the sampling (*estimate_focal_length*) mentioned above and only use the given camera model in absolute pose estimation [(Source Code)](https://github.com/colmap/colmap/blob/43de802cfb3ed2bd155150e7e5e3e8c8dd5aaa3e/src/sfm/incremental_mapper.cc#L430). This won't affect the refinement options.\n3. Images from a camera with a prior focal length will have a higher priority to be selected as the initial pair. [Source Code] (https://github.com/colmap/colmap/blob/43de802cfb3ed2bd155150e7e5e3e8c8dd5aaa3e/src/sfm/incremental_mapper.cc#L845).",
    "2304887": "Some Images contain Exif data which you can get the true value of the focal length... Adding a prior in optimization with the correct value of the focal length would help into making the optimization easier to converge and more constrained to local minimum... \nBut here one need to be careful as the Exif data might be manipulated as in the case of Urban-Kyiv, which the true values of the Exif are not correct due to scale manipulation of the images."
  },
  "source": "meta"
}