{
  "id": 416777,
  "title": "46th solution",
  "url": "/competitions/image-matching-challenge-2023/writeups/kyoukuntaro-46th-solution",
  "author_name": "",
  "post_date": "2023-06-13T01:48:13.898835800Z",
  "votes": 22,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Our method consists of three simple parts: keypoint matching, structure from motion, and post-processing. I will briefly explain each of them with a focus on the differences from the baseline.</p>\n<h2>Keypoint Detect and Matching</h2>\n<p>We adopted a method that performs keypoint extraction and matching separately, rather than an end-to-end matching method that can share 3D model points across many images. Ultimately, we only used KeyNetAffNetHardNet, but if we had more time, we would have liked to ensemble it with SuperPoint-based methods.<br>\nBy making simple changes listed below, we can improve the score.</p>\n<ul>\n<li>increasing the number of keypoint (2048→2048*4).</li>\n<li>extracting keypoints from different resized images.</li>\n<li>using algorithm adalam and Orinet written in the codes.</li>\n<li>using both the Fundamental matrix and the Homography matrix and merging the two results<br>\nfor narrowing down the matching based on geometric characteristics.</li>\n<li>setting an upper limit on the number of feature point matches per pair of images for feature point matching to reduce the computational complexity of 3D reconstruction.</li>\n</ul>\n<h2>Structure from Motion</h2>\n<p>We tried multiple minimum matching numbers and used the model that estimated the largest number of images that could be estimated as the final estimation result. However, if the proportion of images that could be estimated exceeded the threshold when trying the final matching number from the largest one, we did not try any minimum matching below it.</p>\n<h2>Post-Processing</h2>\n<p>Images that could not be estimated by colmap were estimated using cv2.solvePnPRansac.</p>",
  "messages": [
    {
      "id": "2300012",
      "postDate": "06/13/2023 01:48:13",
      "content": "<p>Our method consists of three simple parts: keypoint matching, structure from motion, and post-processing. I will briefly explain each of them with a focus on the differences from the baseline.</p>\n<h2>Keypoint Detect and Matching</h2>\n<p>We adopted a method that performs keypoint extraction and matching separately, rather than an end-to-end matching method that can share 3D model points across many images. Ultimately, we only used KeyNetAffNetHardNet, but if we had more time, we would have liked to ensemble it with SuperPoint-based methods.<br>\nBy making simple changes listed below, we can improve the score.</p>\n<ul>\n<li>increasing the number of keypoint (2048→2048*4).</li>\n<li>extracting keypoints from different resized images.</li>\n<li>using algorithm adalam and Orinet written in the codes.</li>\n<li>using both the Fundamental matrix and the Homography matrix and merging the two results<br>\nfor narrowing down the matching based on geometric characteristics.</li>\n<li>setting an upper limit on the number of feature point matches per pair of images for feature point matching to reduce the computational complexity of 3D reconstruction.</li>\n</ul>\n<h2>Structure from Motion</h2>\n<p>We tried multiple minimum matching numbers and used the model that estimated the largest number of images that could be estimated as the final estimation result. However, if the proportion of images that could be estimated exceeded the threshold when trying the final matching number from the largest one, we did not try any minimum matching below it.</p>\n<h2>Post-Processing</h2>\n<p>Images that could not be estimated by colmap were estimated using cv2.solvePnPRansac.</p>",
      "rawMarkdown": "Our method consists of three simple parts: keypoint matching, structure from motion, and post-processing. I will briefly explain each of them with a focus on the differences from the baseline.\n## Keypoint Detect and Matching\nWe adopted a method that performs keypoint extraction and matching separately, rather than an end-to-end matching method that can share 3D model points across many images. Ultimately, we only used KeyNetAffNetHardNet, but if we had more time, we would have liked to ensemble it with SuperPoint-based methods.\nBy making simple changes listed below, we can improve the score.\n- increasing the number of keypoint (2048→2048*4).\n- extracting keypoints from different resized images.\n- using algorithm adalam and Orinet written in the codes.\n- using both the Fundamental matrix and the Homography matrix and merging the two results\nfor narrowing down the matching based on geometric characteristics.\n- setting an upper limit on the number of feature point matches per pair of images for feature point matching to reduce the computational complexity of 3D reconstruction.\n\n## Structure from Motion\nWe tried multiple minimum matching numbers and used the model that estimated the largest number of images that could be estimated as the final estimation result. However, if the proportion of images that could be estimated exceeded the threshold when trying the final matching number from the largest one, we did not try any minimum matching below it.\n\n## Post-Processing\nImages that could not be estimated by colmap were estimated using cv2.solvePnPRansac.",
      "votes": null
    },
    {
      "id": "2300057",
      "postDate": "06/13/2023 02:39:09",
      "content": "<p>thank you for your sharing！</p>",
      "rawMarkdown": "thank you for your sharing！",
      "votes": null
    },
    {
      "id": "2300059",
      "postDate": "06/13/2023 02:43:43",
      "content": "<p>You can also share your solution on this topic  <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/416737\" target=\"_blank\">https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/416737</a></p>",
      "rawMarkdown": "You can also share your solution on this topic  https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/416737",
      "votes": null
    },
    {
      "id": "2300129",
      "postDate": "06/13/2023 03:36:00",
      "content": "<p>Thanks for sharing! You have some really good ideas, may I ask you a few questions? <br>\n1) How did you merge the results from fundamental matrix and homography?<br>\n2) For your post-processing, from my understanding, cv2.solvePnPRansac can only outputs camera poses given the correspondences between the 3D reconstructed points and their corresponding 2D points on the image, how did you find the 3D-2D correspondances? <br>\n3) Did these 2 steps help boosting your score? </p>",
      "rawMarkdown": "Thanks for sharing! You have some really good ideas, may I ask you a few questions? \n1) How did you merge the results from fundamental matrix and homography?\n2) For your post-processing, from my understanding, cv2.solvePnPRansac can only outputs camera poses given the correspondences between the 3D reconstructed points and their corresponding 2D points on the image, how did you find the 3D-2D correspondances? \n3) Did these 2 steps help boosting your score?",
      "votes": null
    },
    {
      "id": "2300156",
      "postDate": "06/13/2023 04:05:47",
      "content": "<p>Thank you for your question.<br>\n1)  We narrowed each of the matches using the two methods and simply merged them. Of course, we treated the matching that covered the two methods as one.<br>\n2) This was accomplished by using matches.h5 and images.txt and points3d.txt to correspond. It was a bit tedious, but the file contained the necessary information.<br>\n3) Both methods improved the score by about 0.01. Method 2 had a smaller contribution than I expected.</p>",
      "rawMarkdown": "Thank you for your question.\n1)  We narrowed each of the matches using the two methods and simply merged them. Of course, we treated the matching that covered the two methods as one.\n2) This was accomplished by using matches.h5 and images.txt and points3d.txt to correspond. It was a bit tedious, but the file contained the necessary information.\n3) Both methods improved the score by about 0.01. Method 2 had a smaller contribution than I expected.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2300057,
      "author_name": "zhouqun92",
      "author_url": "",
      "post_date": "06/13/2023 02:39:09",
      "content": "<p>thank you for your sharing！</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2300059,
      "author_name": "zhouqun92",
      "author_url": "",
      "post_date": "06/13/2023 02:43:43",
      "content": "<p>You can also share your solution on this topic  <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/416737\" target=\"_blank\">https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/416737</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2300129,
      "author_name": "anonymousyuxiang",
      "author_url": "",
      "post_date": "06/13/2023 03:36:00",
      "content": "<p>Thanks for sharing! You have some really good ideas, may I ask you a few questions? <br>\n1) How did you merge the results from fundamental matrix and homography?<br>\n2) For your post-processing, from my understanding, cv2.solvePnPRansac can only outputs camera poses given the correspondences between the 3D reconstructed points and their corresponding 2D points on the image, how did you find the 3D-2D correspondances? <br>\n3) Did these 2 steps help boosting your score? </p>",
      "votes": null,
      "replies": [
        {
          "id": 2300156,
          "author_name": "yawata",
          "author_url": "",
          "post_date": "06/13/2023 04:05:47",
          "content": "<p>Thank you for your question.<br>\n1)  We narrowed each of the matches using the two methods and simply merged them. Of course, we treated the matching that covered the two methods as one.<br>\n2) This was accomplished by using matches.h5 and images.txt and points3d.txt to correspond. It was a bit tedious, but the file contained the necessary information.<br>\n3) Both methods improved the score by about 0.01. Method 2 had a smaller contribution than I expected.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2300012": "Our method consists of three simple parts: keypoint matching, structure from motion, and post-processing. I will briefly explain each of them with a focus on the differences from the baseline.\n## Keypoint Detect and Matching\nWe adopted a method that performs keypoint extraction and matching separately, rather than an end-to-end matching method that can share 3D model points across many images. Ultimately, we only used KeyNetAffNetHardNet, but if we had more time, we would have liked to ensemble it with SuperPoint-based methods.\nBy making simple changes listed below, we can improve the score.\n- increasing the number of keypoint (2048→2048*4).\n- extracting keypoints from different resized images.\n- using algorithm adalam and Orinet written in the codes.\n- using both the Fundamental matrix and the Homography matrix and merging the two results\nfor narrowing down the matching based on geometric characteristics.\n- setting an upper limit on the number of feature point matches per pair of images for feature point matching to reduce the computational complexity of 3D reconstruction.\n\n## Structure from Motion\nWe tried multiple minimum matching numbers and used the model that estimated the largest number of images that could be estimated as the final estimation result. However, if the proportion of images that could be estimated exceeded the threshold when trying the final matching number from the largest one, we did not try any minimum matching below it.\n\n## Post-Processing\nImages that could not be estimated by colmap were estimated using cv2.solvePnPRansac.",
    "2300057": "thank you for your sharing！",
    "2300059": "You can also share your solution on this topic  https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/416737",
    "2300129": "Thanks for sharing! You have some really good ideas, may I ask you a few questions? \n1) How did you merge the results from fundamental matrix and homography?\n2) For your post-processing, from my understanding, cv2.solvePnPRansac can only outputs camera poses given the correspondences between the 3D reconstructed points and their corresponding 2D points on the image, how did you find the 3D-2D correspondances? \n3) Did these 2 steps help boosting your score?",
    "2300156": "Thank you for your question.\n1)  We narrowed each of the matches using the two methods and simply merged them. Of course, we treated the matching that covered the two methods as one.\n2) This was accomplished by using matches.h5 and images.txt and points3d.txt to correspond. It was a bit tedious, but the file contained the necessary information.\n3) Both methods improved the score by about 0.01. Method 2 had a smaller contribution than I expected."
  },
  "source": "meta"
}