{
  "id": 418139,
  "title": "26th Place Solution: SP&SG + Rotation",
  "url": "/competitions/image-matching-challenge-2023/discussion/418139",
  "author_name": "HaoHong",
  "post_date": "2023-06-19T01:07:18.179000",
  "votes": 3,
  "comment_count": 0,
  "views": 0,
  "content": "<p>First of all, I would like to thank the organizers and kaggle staff for organizing this wonderful competition :)</p>\n<h1>Overview</h1>\n<p>First, we select the image pair by using EfficientNet, just like the baseline. We use SuperPoint&amp;SuperGlue for feature matching, and use cv2.MAGSAC to remove some obvious error points. Then feed the points into pycolmap for reconstruction.</p>\n<h1>SP&amp;SG</h1>\n<p>We use SP&amp;SG for feature matching at 3 different resolutions. We found that SP&amp;SG didn’t perform well in resolution that higer than 1600, perhaps it's due to the training data. So fianlly we set the resolution as 1500 * 1000, 1024 * 1024 and 1000 * 750, respectively represent image with 3:2, 1:1, 4:3 aspect ratio. SuperGlue’s sinkhorn iteration is set to 5, more iteration time didn’t improve the result but require lot of time.<br>\nThen we use MAGSAC and set the reprojection error as 3 to remove some obvious error points. Although it did not show a significant improvement in the pubilc LB, it seems to help slightly in the private LB.</p>\n<pre><code> , inliers = cv2.findFundamentalMat(mkpts0_superglue, mkpts1_superglue, cv2.USAC_MAGSAC, , ., ) \n</code></pre>\n<h1>Other feature matching method we have try</h1>\n<p>We have try to esemble <code>LoFTR, DKM, GlueStick</code> into the pileline, but they didn’t improve the result and sometime may even reduces the reconstruction accuracy, even when we use MAGSAC to filter the matching points with a reprojection error less than 0.5 pixel. I don’t know if there is some mistake with the way we use, but I have to say that SP&amp;SG is quite strong.<br>\nWe even try to use optical flow to filter the matching result ,but optical flow didn’t act well in wide baseline image pair.</p>\n<h1>Some tricks</h1>\n<ol>\n<li>Rotation<br>\nThe pretrain checkpoint of SP&amp;SG we used don’t have the robust rotational invariance like SIFT, so we make a simple judgment that if matching points are less than 20, we just rotate the image by 90° and -90°. With the rotate operation, our score improve from 0.36 to 0.39 in the public LB.</li>\n<li>Swap pair<br>\nWe found that SP&amp;SG may act different when selecting different image as the reference image, so I swap the pair when it didn’t match enough points. It slightly improve the score (about 0.01) in the LB.</li>\n<li>Make it more efficient<br>\nIn order to keep the runtime below 9 hours, we make a simple judgment that if SP&amp;SG gets less than 20 points in one resolution then stop the matching of this image pair. Thought we try to use LoFTR and DKM to deal with those pair, it didn’t help much. With this little trick we can finish the test data in 6h and didn’t result in any drop in the score, this make it possible for us to incorporate other matching algorithms.</li>\n</ol>",
  "messages": [
    {
      "id": 2308459,
      "postDate": "2023-06-19T01:07:18.180Z",
      "content": "<p>First of all, I would like to thank the organizers and kaggle staff for organizing this wonderful competition :)</p>\n<h1>Overview</h1>\n<p>First, we select the image pair by using EfficientNet, just like the baseline. We use SuperPoint&amp;SuperGlue for feature matching, and use cv2.MAGSAC to remove some obvious error points. Then feed the points into pycolmap for reconstruction.</p>\n<h1>SP&amp;SG</h1>\n<p>We use SP&amp;SG for feature matching at 3 different resolutions. We found that SP&amp;SG didn’t perform well in resolution that higer than 1600, perhaps it's due to the training data. So fianlly we set the resolution as 1500 * 1000, 1024 * 1024 and 1000 * 750, respectively represent image with 3:2, 1:1, 4:3 aspect ratio. SuperGlue’s sinkhorn iteration is set to 5, more iteration time didn’t improve the result but require lot of time.<br>\nThen we use MAGSAC and set the reprojection error as 3 to remove some obvious error points. Although it did not show a significant improvement in the pubilc LB, it seems to help slightly in the private LB.</p>\n<pre><code> , inliers = cv2.findFundamentalMat(mkpts0_superglue, mkpts1_superglue, cv2.USAC_MAGSAC, , ., ) \n</code></pre>\n<h1>Other feature matching method we have try</h1>\n<p>We have try to esemble <code>LoFTR, DKM, GlueStick</code> into the pileline, but they didn’t improve the result and sometime may even reduces the reconstruction accuracy, even when we use MAGSAC to filter the matching points with a reprojection error less than 0.5 pixel. I don’t know if there is some mistake with the way we use, but I have to say that SP&amp;SG is quite strong.<br>\nWe even try to use optical flow to filter the matching result ,but optical flow didn’t act well in wide baseline image pair.</p>\n<h1>Some tricks</h1>\n<ol>\n<li>Rotation<br>\nThe pretrain checkpoint of SP&amp;SG we used don’t have the robust rotational invariance like SIFT, so we make a simple judgment that if matching points are less than 20, we just rotate the image by 90° and -90°. With the rotate operation, our score improve from 0.36 to 0.39 in the public LB.</li>\n<li>Swap pair<br>\nWe found that SP&amp;SG may act different when selecting different image as the reference image, so I swap the pair when it didn’t match enough points. It slightly improve the score (about 0.01) in the LB.</li>\n<li>Make it more efficient<br>\nIn order to keep the runtime below 9 hours, we make a simple judgment that if SP&amp;SG gets less than 20 points in one resolution then stop the matching of this image pair. Thought we try to use LoFTR and DKM to deal with those pair, it didn’t help much. With this little trick we can finish the test data in 6h and didn’t result in any drop in the score, this make it possible for us to incorporate other matching algorithms.</li>\n</ol>",
      "rawMarkdown": "First of all, I would like to thank the organizers and kaggle staff for organizing this wonderful competition :)\n# Overview\nFirst, we select the image pair by using EfficientNet, just like the baseline. We use SuperPoint&SuperGlue for feature matching, and use cv2.MAGSAC to remove some obvious error points. Then feed the points into pycolmap for reconstruction.\n# SP&SG\nWe use SP&SG for feature matching at 3 different resolutions. We found that SP&SG didn’t perform well in resolution that higer than 1600, perhaps it's due to the training data. So fianlly we set the resolution as 1500 * 1000, 1024 * 1024 and 1000 * 750, respectively represent image with 3:2, 1:1, 4:3 aspect ratio. SuperGlue’s sinkhorn iteration is set to 5, more iteration time didn’t improve the result but require lot of time.\nThen we use MAGSAC and set the reprojection error as 3 to remove some obvious error points. Although it did not show a significant improvement in the pubilc LB, it seems to help slightly in the private LB.\n```\n Fm, inliers = cv2.findFundamentalMat(mkpts0_superglue, mkpts1_superglue, cv2.USAC_MAGSAC, 3, 0.999, 100000) \n```\n# Other feature matching method we have try\nWe have try to esemble `LoFTR, DKM, GlueStick` into the pileline, but they didn’t improve the result and sometime may even reduces the reconstruction accuracy, even when we use MAGSAC to filter the matching points with a reprojection error less than 0.5 pixel. I don’t know if there is some mistake with the way we use, but I have to say that SP&SG is quite strong.\nWe even try to use optical flow to filter the matching result ,but optical flow didn’t act well in wide baseline image pair.\n# Some tricks\n1. Rotation\nThe pretrain checkpoint of SP&SG we used don’t have the robust rotational invariance like SIFT, so we make a simple judgment that if matching points are less than 20, we just rotate the image by 90° and -90°. With the rotate operation, our score improve from 0.36 to 0.39 in the public LB.\n2. Swap pair\nWe found that SP&SG may act different when selecting different image as the reference image, so I swap the pair when it didn’t match enough points. It slightly improve the score (about 0.01) in the LB.\n3. Make it more efficient\nIn order to keep the runtime below 9 hours, we make a simple judgment that if SP&SG gets less than 20 points in one resolution then stop the matching of this image pair. Thought we try to use LoFTR and DKM to deal with those pair, it didn’t help much. With this little trick we can finish the test data in 6h and didn’t result in any drop in the score, this make it possible for us to incorporate other matching algorithms.",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2308459": "First of all, I would like to thank the organizers and kaggle staff for organizing this wonderful competition :)\n# Overview\nFirst, we select the image pair by using EfficientNet, just like the baseline. We use SuperPoint&SuperGlue for feature matching, and use cv2.MAGSAC to remove some obvious error points. Then feed the points into pycolmap for reconstruction.\n# SP&SG\nWe use SP&SG for feature matching at 3 different resolutions. We found that SP&SG didn’t perform well in resolution that higer than 1600, perhaps it's due to the training data. So fianlly we set the resolution as 1500 * 1000, 1024 * 1024 and 1000 * 750, respectively represent image with 3:2, 1:1, 4:3 aspect ratio. SuperGlue’s sinkhorn iteration is set to 5, more iteration time didn’t improve the result but require lot of time.\nThen we use MAGSAC and set the reprojection error as 3 to remove some obvious error points. Although it did not show a significant improvement in the pubilc LB, it seems to help slightly in the private LB.\n```\n Fm, inliers = cv2.findFundamentalMat(mkpts0_superglue, mkpts1_superglue, cv2.USAC_MAGSAC, 3, 0.999, 100000) \n```\n# Other feature matching method we have try\nWe have try to esemble `LoFTR, DKM, GlueStick` into the pileline, but they didn’t improve the result and sometime may even reduces the reconstruction accuracy, even when we use MAGSAC to filter the matching points with a reprojection error less than 0.5 pixel. I don’t know if there is some mistake with the way we use, but I have to say that SP&SG is quite strong.\nWe even try to use optical flow to filter the matching result ,but optical flow didn’t act well in wide baseline image pair.\n# Some tricks\n1. Rotation\nThe pretrain checkpoint of SP&SG we used don’t have the robust rotational invariance like SIFT, so we make a simple judgment that if matching points are less than 20, we just rotate the image by 90° and -90°. With the rotate operation, our score improve from 0.36 to 0.39 in the public LB.\n2. Swap pair\nWe found that SP&SG may act different when selecting different image as the reference image, so I swap the pair when it didn’t match enough points. It slightly improve the score (about 0.01) in the LB.\n3. Make it more efficient\nIn order to keep the runtime below 9 hours, we make a simple judgment that if SP&SG gets less than 20 points in one resolution then stop the matching of this image pair. Thought we try to use LoFTR and DKM to deal with those pair, it didn’t help much. With this little trick we can finish the test data in 6h and didn’t result in any drop in the score, this make it possible for us to incorporate other matching algorithms."
  }
}