{
  "id": 487622,
  "title": "Is 3D model reconstruction strictly necessary?",
  "url": "/competitions/image-matching-challenge-2024/discussion/487622",
  "author_name": "Serhii Hrynko",
  "post_date": "2024-03-29T19:21:19.905000",
  "votes": 4,
  "comment_count": 3,
  "views": 0,
  "content": "<p>It looks like submission data as well as the competition metric does not include any parts of the actual 3d model.<br>\nInstead we are asked to predict relative positions (read distances) between the cameras. Even orientations don't seem to affect final score.</p>\n<p>E.g. submission with translation vectors equal to center coordinates and all identity rotation matrices seems to yield same score as submission with correct relative orientations.</p>\n<p>The question is: can reconstruction of 3d model, followed by referencing images to it, followed by estimation of camera parameters, produce any better results than directly predicting distances between cameras based on initial set of images?</p>\n<p>I'm personally betting on direct camera position/distance prediction, especially since given images are not specifically prepared for 3d model reconstruction thus making any results derived from such reconstruction extremely unreliable.</p>",
  "messages": [
    {
      "id": 2722824,
      "postDate": "2024-03-29T19:21:19.907Z",
      "content": "<p>It looks like submission data as well as the competition metric does not include any parts of the actual 3d model.<br>\nInstead we are asked to predict relative positions (read distances) between the cameras. Even orientations don't seem to affect final score.</p>\n<p>E.g. submission with translation vectors equal to center coordinates and all identity rotation matrices seems to yield same score as submission with correct relative orientations.</p>\n<p>The question is: can reconstruction of 3d model, followed by referencing images to it, followed by estimation of camera parameters, produce any better results than directly predicting distances between cameras based on initial set of images?</p>\n<p>I'm personally betting on direct camera position/distance prediction, especially since given images are not specifically prepared for 3d model reconstruction thus making any results derived from such reconstruction extremely unreliable.</p>",
      "rawMarkdown": "It looks like submission data as well as the competition metric does not include any parts of the actual 3d model.\nInstead we are asked to predict relative positions (read distances) between the cameras. Even orientations don't seem to affect final score.\n\nE.g. submission with translation vectors equal to center coordinates and all identity rotation matrices seems to yield same score as submission with correct relative orientations.\n\nThe question is: can reconstruction of 3d model, followed by referencing images to it, followed by estimation of camera parameters, produce any better results than directly predicting distances between cameras based on initial set of images?\n\nI'm personally betting on direct camera position/distance prediction, especially since given images are not specifically prepared for 3d model reconstruction thus making any results derived from such reconstruction extremely unreliable.",
      "votes": 4
    },
    {
      "id": 2722874,
      "postDate": "2024-03-29T20:17:31.440Z",
      "content": "<p>E.g. here is one article I found discussing this topic <a href=\"https://andrewjkramer.net/camera-pose-estimation-using-convolutional-neural-networks/\" target=\"_blank\">https://andrewjkramer.net/camera-pose-estimation-using-convolutional-neural-networks/</a></p>\n<blockquote>\n  <p>Feature matching methods were, in general, more accurate. However, there are certain cases in which the neural network performs better. For higher changes in position and orientation the CNN can slightly outperform feature based methods. Also, the CNN can perform significantly better in cases where it is difficult to find and match features. For example, when the images do not contain a sufficient number of features.</p>\n</blockquote>",
      "rawMarkdown": "E.g. here is one article I found discussing this topic https://andrewjkramer.net/camera-pose-estimation-using-convolutional-neural-networks/\n\n> Feature matching methods were, in general, more accurate. However, there are certain cases in which the neural network performs better. For higher changes in position and orientation the CNN can slightly outperform feature based methods. Also, the CNN can perform significantly better in cases where it is difficult to find and match features. For example, when the images do not contain a sufficient number of features.",
      "votes": 1
    },
    {
      "id": 2722842,
      "postDate": "2024-03-29T19:41:05.070Z",
      "content": "<p>Absolutely. Existing SfM methods give you both 3D model and a point cloud as a byproduct. But if you can provide directly the camera poses, that is enough.</p>",
      "rawMarkdown": "Absolutely. Existing SfM methods give you both 3D model and a point cloud as a byproduct. But if you can provide directly the camera poses, that is enough.\n",
      "votes": 2
    },
    {
      "id": 2722839,
      "postDate": "2024-03-29T19:31:46.827Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2722874,
      "author_name": "Serhii Hrynko",
      "author_url": "",
      "post_date": "2024-03-29T20:17:31.440000",
      "content": "<p>E.g. here is one article I found discussing this topic <a href=\"https://andrewjkramer.net/camera-pose-estimation-using-convolutional-neural-networks/\" target=\"_blank\">https://andrewjkramer.net/camera-pose-estimation-using-convolutional-neural-networks/</a></p>\n<blockquote>\n  <p>Feature matching methods were, in general, more accurate. However, there are certain cases in which the neural network performs better. For higher changes in position and orientation the CNN can slightly outperform feature based methods. Also, the CNN can perform significantly better in cases where it is difficult to find and match features. For example, when the images do not contain a sufficient number of features.</p>\n</blockquote>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2722842,
      "author_name": "old-ufo",
      "author_url": "",
      "post_date": "2024-03-29T19:41:05.070000",
      "content": "<p>Absolutely. Existing SfM methods give you both 3D model and a point cloud as a byproduct. But if you can provide directly the camera poses, that is enough.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2722839,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-03-29T19:31:46.827000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2722824": "It looks like submission data as well as the competition metric does not include any parts of the actual 3d model.\nInstead we are asked to predict relative positions (read distances) between the cameras. Even orientations don't seem to affect final score.\n\nE.g. submission with translation vectors equal to center coordinates and all identity rotation matrices seems to yield same score as submission with correct relative orientations.\n\nThe question is: can reconstruction of 3d model, followed by referencing images to it, followed by estimation of camera parameters, produce any better results than directly predicting distances between cameras based on initial set of images?\n\nI'm personally betting on direct camera position/distance prediction, especially since given images are not specifically prepared for 3d model reconstruction thus making any results derived from such reconstruction extremely unreliable.",
    "2722874": "E.g. here is one article I found discussing this topic https://andrewjkramer.net/camera-pose-estimation-using-convolutional-neural-networks/\n\n> Feature matching methods were, in general, more accurate. However, there are certain cases in which the neural network performs better. For higher changes in position and orientation the CNN can slightly outperform feature based methods. Also, the CNN can perform significantly better in cases where it is difficult to find and match features. For example, when the images do not contain a sufficient number of features.",
    "2722842": "Absolutely. Existing SfM methods give you both 3D model and a point cloud as a byproduct. But if you can provide directly the camera poses, that is enough.\n",
    "2722839": ""
  }
}