{
  "id": 417002,
  "title": "16th Place Solution",
  "url": "/competitions/image-matching-challenge-2023/discussion/417002",
  "author_name": "Yuxiang Huang",
  "post_date": "2023-06-13T19:44:36.580000",
  "votes": 20,
  "comment_count": 0,
  "views": 0,
  "content": "<p>First, we would like to say thank you to the organizers and Kaggle staff for setting up this challenge, it has been an amazing experience for us. </p>\n<p><strong>Our solution</strong><br>\nFor our final submission, we used the an emsemble of SuperPoint, KeyNet/AffNet/HardNet and SOSNet as our feature detectors/descriptors, and used SuperGlue and Adalam for feature matching. We used Colmap for reconstruction and camera localization. We also used <a href=\"https://github.com/cvg/Hierarchical-Localization\" target=\"_blank\">hloc</a> to speed up the pipeline and make it more scalable for validation and testing. </p>\n<p><strong>Image Retrieval</strong><br>\nWe used <a href=\"https://openaccess.thecvf.com/content_cvpr_2016/papers/Arandjelovic_NetVLAD_CNN_Architecture_CVPR_2016_paper.pdf\" target=\"_blank\">NetVLAD </a>as implemented in <a href=\"https://github.com/cvg/Hierarchical-Localization\" target=\"_blank\">hloc</a> as the global feature descriptor for image retrieval</p>\n<p><strong>Feature Matching</strong><br>\nThe first thing we noticed in the dataset was that some scenes contain a lot of rotated images, and we tried to tackle this problem with 2 approaches: <br>\n1) use rotation invariant feature matchers (e.g., KeyNet/AffNet/HardNet, SOSNet). <br>\n2) use a <a href=\"https://github.com/ternaus/check_orientation\" target=\"_blank\">lightweight orientation detector</a> to detect the rotation angles and rotate the image pairs accordingly so that both images have a similar orientation (for simplicity, we only set the rotation angles to 90, 180 and 270 degrees). </p>\n<p>We proceeded with both approaches and found out that both approaches achieve a similar improvement on the Heritage dataset, however, by ensembling more feature matchers, we observe some extra improvements on Urban and Haiper datasets, so we finally took this approach, and this ensemble achieved the best results for us within the time limit of 9h: <strong>SuperGlue + KeyNet/AffNet/HardNet (with Adalam) + SOSNet (with Adalam)</strong>. Using orientation compensation on the ensembled model does not bring any extra improvements. </p>\n<p><strong>Things that did not work</strong>:<br>\n1) We first tried <a href=\"https://github.com/pidahbus/deep-image-orientation-angle-detection\" target=\"_blank\">this SOTA orientation detector</a>, however it consumes too much memory and could not be integrated into our pipeline on Kaggle<br>\n2) We found <strong>KeyNet/AffNet/HardNet + Adalam</strong> in the baseline the <strong>best single feature matcher</strong> without any preprocessing -- We could achieve 0.455/0.433 (equal to 47th place) by only tuning its parameters and resizing the input images to 1600, however, when we integrated them into our pipeline using hloc, its performance dropped significantly to 0.334/0.277 (locally as well, mainly on urban), we tried to investigate but still do not know why. <br>\n3) We experimented on a lot of recent feature matchers and ensembles, including DKMv3, DISK, LoFTR, SiLK, DAC, and they either do not perform as well or are too slow when integrated into the pipeline. In general, we found that end-to-end dense matchers not well suited for this multiview challenge despite their success in last year's two view challenge, their speed is too slow and the scores they achieve are also not as good. Here are some local validation results:</p>\n<ol>\n<li>SiLK (on ~800x600):<br>\nurban: 0.125<br>\nhaiper: 0.165</li>\n<li>DKMv3 (on ~800x600 and it's still very slow):<br>\nheritage: 0.185<br>\nhaiper: 0.510</li>\n<li>DISK (on ~1600x1200):<br>\nurban: 0.461<br>\nheritage: 0.292 (0.452 with rotation compensation)<br>\nhaiper: 0.433</li>\n<li>SOSNet with Adalam (on ~1600x1200):<br>\nurban: 0.031<br>\nheritage: 0.460 (same with rotation compensation)<br>\nhaiper: 0.653</li>\n<li>Sift / Rootsift with Adalam (on ~1600x1200):<br>\nurban: 0.02<br>\nheritage: 0.396<br>\nhaiper: 0.635</li>\n<li>DAC: the results are very bad</li>\n</ol>\n<p><strong>Reconstruction</strong><br>\nAfter merging all the match points from the ensemble, we apply <a href=\"https://github.com/colmap/pycolmap/blob/743a4ac305183f96d2a4cfce7c7f6418b31b8598/pipeline/match_features.cc#L76\" target=\"_blank\">geometric verification</a> in Colmap before reconstructing the model, which speeds up the reconstruction. <br>\n<strong>Things that did not work</strong>: <br>\n1) We tried using Pixel-Perfect SFM, we set it up locally and it gave descent results visually comparable to our pipeline, but since we could not get it up running on Kaggle we did not proceed further. <br>\n2) We tried using MAGSAC++ to replace the default RANSAC function Colmap uses to remove bad matching points before reconstructing the model, but we did not see a significant difference in the final scores. </p>",
  "messages": [
    {
      "id": 2301329,
      "postDate": "2023-06-13T19:44:36.580Z",
      "content": "<p>First, we would like to say thank you to the organizers and Kaggle staff for setting up this challenge, it has been an amazing experience for us. </p>\n<p><strong>Our solution</strong><br>\nFor our final submission, we used the an emsemble of SuperPoint, KeyNet/AffNet/HardNet and SOSNet as our feature detectors/descriptors, and used SuperGlue and Adalam for feature matching. We used Colmap for reconstruction and camera localization. We also used <a href=\"https://github.com/cvg/Hierarchical-Localization\" target=\"_blank\">hloc</a> to speed up the pipeline and make it more scalable for validation and testing. </p>\n<p><strong>Image Retrieval</strong><br>\nWe used <a href=\"https://openaccess.thecvf.com/content_cvpr_2016/papers/Arandjelovic_NetVLAD_CNN_Architecture_CVPR_2016_paper.pdf\" target=\"_blank\">NetVLAD </a>as implemented in <a href=\"https://github.com/cvg/Hierarchical-Localization\" target=\"_blank\">hloc</a> as the global feature descriptor for image retrieval</p>\n<p><strong>Feature Matching</strong><br>\nThe first thing we noticed in the dataset was that some scenes contain a lot of rotated images, and we tried to tackle this problem with 2 approaches: <br>\n1) use rotation invariant feature matchers (e.g., KeyNet/AffNet/HardNet, SOSNet). <br>\n2) use a <a href=\"https://github.com/ternaus/check_orientation\" target=\"_blank\">lightweight orientation detector</a> to detect the rotation angles and rotate the image pairs accordingly so that both images have a similar orientation (for simplicity, we only set the rotation angles to 90, 180 and 270 degrees). </p>\n<p>We proceeded with both approaches and found out that both approaches achieve a similar improvement on the Heritage dataset, however, by ensembling more feature matchers, we observe some extra improvements on Urban and Haiper datasets, so we finally took this approach, and this ensemble achieved the best results for us within the time limit of 9h: <strong>SuperGlue + KeyNet/AffNet/HardNet (with Adalam) + SOSNet (with Adalam)</strong>. Using orientation compensation on the ensembled model does not bring any extra improvements. </p>\n<p><strong>Things that did not work</strong>:<br>\n1) We first tried <a href=\"https://github.com/pidahbus/deep-image-orientation-angle-detection\" target=\"_blank\">this SOTA orientation detector</a>, however it consumes too much memory and could not be integrated into our pipeline on Kaggle<br>\n2) We found <strong>KeyNet/AffNet/HardNet + Adalam</strong> in the baseline the <strong>best single feature matcher</strong> without any preprocessing -- We could achieve 0.455/0.433 (equal to 47th place) by only tuning its parameters and resizing the input images to 1600, however, when we integrated them into our pipeline using hloc, its performance dropped significantly to 0.334/0.277 (locally as well, mainly on urban), we tried to investigate but still do not know why. <br>\n3) We experimented on a lot of recent feature matchers and ensembles, including DKMv3, DISK, LoFTR, SiLK, DAC, and they either do not perform as well or are too slow when integrated into the pipeline. In general, we found that end-to-end dense matchers not well suited for this multiview challenge despite their success in last year's two view challenge, their speed is too slow and the scores they achieve are also not as good. Here are some local validation results:</p>\n<ol>\n<li>SiLK (on ~800x600):<br>\nurban: 0.125<br>\nhaiper: 0.165</li>\n<li>DKMv3 (on ~800x600 and it's still very slow):<br>\nheritage: 0.185<br>\nhaiper: 0.510</li>\n<li>DISK (on ~1600x1200):<br>\nurban: 0.461<br>\nheritage: 0.292 (0.452 with rotation compensation)<br>\nhaiper: 0.433</li>\n<li>SOSNet with Adalam (on ~1600x1200):<br>\nurban: 0.031<br>\nheritage: 0.460 (same with rotation compensation)<br>\nhaiper: 0.653</li>\n<li>Sift / Rootsift with Adalam (on ~1600x1200):<br>\nurban: 0.02<br>\nheritage: 0.396<br>\nhaiper: 0.635</li>\n<li>DAC: the results are very bad</li>\n</ol>\n<p><strong>Reconstruction</strong><br>\nAfter merging all the match points from the ensemble, we apply <a href=\"https://github.com/colmap/pycolmap/blob/743a4ac305183f96d2a4cfce7c7f6418b31b8598/pipeline/match_features.cc#L76\" target=\"_blank\">geometric verification</a> in Colmap before reconstructing the model, which speeds up the reconstruction. <br>\n<strong>Things that did not work</strong>: <br>\n1) We tried using Pixel-Perfect SFM, we set it up locally and it gave descent results visually comparable to our pipeline, but since we could not get it up running on Kaggle we did not proceed further. <br>\n2) We tried using MAGSAC++ to replace the default RANSAC function Colmap uses to remove bad matching points before reconstructing the model, but we did not see a significant difference in the final scores. </p>",
      "rawMarkdown": "First, we would like to say thank you to the organizers and Kaggle staff for setting up this challenge, it has been an amazing experience for us. \n\n**Our solution**\nFor our final submission, we used the an emsemble of SuperPoint, KeyNet/AffNet/HardNet and SOSNet as our feature detectors/descriptors, and used SuperGlue and Adalam for feature matching. We used Colmap for reconstruction and camera localization. We also used [hloc](https://github.com/cvg/Hierarchical-Localization) to speed up the pipeline and make it more scalable for validation and testing. \n\n**Image Retrieval**\nWe used [NetVLAD ](https://openaccess.thecvf.com/content_cvpr_2016/papers/Arandjelovic_NetVLAD_CNN_Architecture_CVPR_2016_paper.pdf)as implemented in [hloc](https://github.com/cvg/Hierarchical-Localization) as the global feature descriptor for image retrieval\n\n**Feature Matching**\nThe first thing we noticed in the dataset was that some scenes contain a lot of rotated images, and we tried to tackle this problem with 2 approaches: \n1) use rotation invariant feature matchers (e.g., KeyNet/AffNet/HardNet, SOSNet). \n2) use a [lightweight orientation detector](https://github.com/ternaus/check_orientation) to detect the rotation angles and rotate the image pairs accordingly so that both images have a similar orientation (for simplicity, we only set the rotation angles to 90, 180 and 270 degrees). \n\nWe proceeded with both approaches and found out that both approaches achieve a similar improvement on the Heritage dataset, however, by ensembling more feature matchers, we observe some extra improvements on Urban and Haiper datasets, so we finally took this approach, and this ensemble achieved the best results for us within the time limit of 9h: **SuperGlue + KeyNet/AffNet/HardNet (with Adalam) + SOSNet (with Adalam)**. Using orientation compensation on the ensembled model does not bring any extra improvements. \n\n**Things that did not work**:\n1) We first tried [this SOTA orientation detector](https://github.com/pidahbus/deep-image-orientation-angle-detection), however it consumes too much memory and could not be integrated into our pipeline on Kaggle\n2) We found **KeyNet/AffNet/HardNet + Adalam** in the baseline the **best single feature matcher** without any preprocessing -- We could achieve 0.455/0.433 (equal to 47th place) by only tuning its parameters and resizing the input images to 1600, however, when we integrated them into our pipeline using hloc, its performance dropped significantly to 0.334/0.277 (locally as well, mainly on urban), we tried to investigate but still do not know why. \n3) We experimented on a lot of recent feature matchers and ensembles, including DKMv3, DISK, LoFTR, SiLK, DAC, and they either do not perform as well or are too slow when integrated into the pipeline. In general, we found that end-to-end dense matchers not well suited for this multiview challenge despite their success in last year's two view challenge, their speed is too slow and the scores they achieve are also not as good. Here are some local validation results:\n1. SiLK (on ~800x600):\nurban: 0.125\nhaiper: 0.165\n2. DKMv3 (on ~800x600 and it's still very slow):\nheritage: 0.185\nhaiper: 0.510\n3. DISK (on ~1600x1200):\nurban: 0.461\nheritage: 0.292 (0.452 with rotation compensation)\nhaiper: 0.433\n4. SOSNet with Adalam (on ~1600x1200):\nurban: 0.031\nheritage: 0.460 (same with rotation compensation)\nhaiper: 0.653\n5. Sift / Rootsift with Adalam (on ~1600x1200):\nurban: 0.02\nheritage: 0.396\nhaiper: 0.635\n5. DAC: the results are very bad\n\n**Reconstruction**\nAfter merging all the match points from the ensemble, we apply [geometric verification](https://github.com/colmap/pycolmap/blob/743a4ac305183f96d2a4cfce7c7f6418b31b8598/pipeline/match_features.cc#L76) in Colmap before reconstructing the model, which speeds up the reconstruction. \n**Things that did not work**: \n1) We tried using Pixel-Perfect SFM, we set it up locally and it gave descent results visually comparable to our pipeline, but since we could not get it up running on Kaggle we did not proceed further. \n2) We tried using MAGSAC++ to replace the default RANSAC function Colmap uses to remove bad matching points before reconstructing the model, but we did not see a significant difference in the final scores. \n\n\n",
      "votes": 20
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2301329": "First, we would like to say thank you to the organizers and Kaggle staff for setting up this challenge, it has been an amazing experience for us. \n\n**Our solution**\nFor our final submission, we used the an emsemble of SuperPoint, KeyNet/AffNet/HardNet and SOSNet as our feature detectors/descriptors, and used SuperGlue and Adalam for feature matching. We used Colmap for reconstruction and camera localization. We also used [hloc](https://github.com/cvg/Hierarchical-Localization) to speed up the pipeline and make it more scalable for validation and testing. \n\n**Image Retrieval**\nWe used [NetVLAD ](https://openaccess.thecvf.com/content_cvpr_2016/papers/Arandjelovic_NetVLAD_CNN_Architecture_CVPR_2016_paper.pdf)as implemented in [hloc](https://github.com/cvg/Hierarchical-Localization) as the global feature descriptor for image retrieval\n\n**Feature Matching**\nThe first thing we noticed in the dataset was that some scenes contain a lot of rotated images, and we tried to tackle this problem with 2 approaches: \n1) use rotation invariant feature matchers (e.g., KeyNet/AffNet/HardNet, SOSNet). \n2) use a [lightweight orientation detector](https://github.com/ternaus/check_orientation) to detect the rotation angles and rotate the image pairs accordingly so that both images have a similar orientation (for simplicity, we only set the rotation angles to 90, 180 and 270 degrees). \n\nWe proceeded with both approaches and found out that both approaches achieve a similar improvement on the Heritage dataset, however, by ensembling more feature matchers, we observe some extra improvements on Urban and Haiper datasets, so we finally took this approach, and this ensemble achieved the best results for us within the time limit of 9h: **SuperGlue + KeyNet/AffNet/HardNet (with Adalam) + SOSNet (with Adalam)**. Using orientation compensation on the ensembled model does not bring any extra improvements. \n\n**Things that did not work**:\n1) We first tried [this SOTA orientation detector](https://github.com/pidahbus/deep-image-orientation-angle-detection), however it consumes too much memory and could not be integrated into our pipeline on Kaggle\n2) We found **KeyNet/AffNet/HardNet + Adalam** in the baseline the **best single feature matcher** without any preprocessing -- We could achieve 0.455/0.433 (equal to 47th place) by only tuning its parameters and resizing the input images to 1600, however, when we integrated them into our pipeline using hloc, its performance dropped significantly to 0.334/0.277 (locally as well, mainly on urban), we tried to investigate but still do not know why. \n3) We experimented on a lot of recent feature matchers and ensembles, including DKMv3, DISK, LoFTR, SiLK, DAC, and they either do not perform as well or are too slow when integrated into the pipeline. In general, we found that end-to-end dense matchers not well suited for this multiview challenge despite their success in last year's two view challenge, their speed is too slow and the scores they achieve are also not as good. Here are some local validation results:\n1. SiLK (on ~800x600):\nurban: 0.125\nhaiper: 0.165\n2. DKMv3 (on ~800x600 and it's still very slow):\nheritage: 0.185\nhaiper: 0.510\n3. DISK (on ~1600x1200):\nurban: 0.461\nheritage: 0.292 (0.452 with rotation compensation)\nhaiper: 0.433\n4. SOSNet with Adalam (on ~1600x1200):\nurban: 0.031\nheritage: 0.460 (same with rotation compensation)\nhaiper: 0.653\n5. Sift / Rootsift with Adalam (on ~1600x1200):\nurban: 0.02\nheritage: 0.396\nhaiper: 0.635\n5. DAC: the results are very bad\n\n**Reconstruction**\nAfter merging all the match points from the ensemble, we apply [geometric verification](https://github.com/colmap/pycolmap/blob/743a4ac305183f96d2a4cfce7c7f6418b31b8598/pipeline/match_features.cc#L76) in Colmap before reconstructing the model, which speeds up the reconstruction. \n**Things that did not work**: \n1) We tried using Pixel-Perfect SFM, we set it up locally and it gave descent results visually comparable to our pipeline, but since we could not get it up running on Kaggle we did not proceed further. \n2) We tried using MAGSAC++ to replace the default RANSAC function Colmap uses to remove bad matching points before reconstructing the model, but we did not see a significant difference in the final scores. \n\n\n"
  }
}