{
  "id": 414175,
  "title": "Can we train two models, one each for rotation matrix and translation vector?",
  "url": "/competitions/image-matching-challenge-2023/discussion/414175",
  "author_name": "",
  "post_date": "2023-05-31T16:42:06.574088800Z",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I observed that the entries of the rotation matrix are within the interval [-1,1] and the entries of the translation vector are unbounded.</p>\n<p>So, i was hoping to train two different models for prediction rotation matrix and translation vector respectively. I feel it will be convenient even to choose different loss functions.</p>\n<p>Can we train two different models for this competition?</p>",
  "messages": [
    {
      "id": "2282567",
      "postDate": "05/31/2023 16:42:06",
      "content": "<p>I observed that the entries of the rotation matrix are within the interval [-1,1] and the entries of the translation vector are unbounded.</p>\n<p>So, i was hoping to train two different models for prediction rotation matrix and translation vector respectively. I feel it will be convenient even to choose different loss functions.</p>\n<p>Can we train two different models for this competition?</p>",
      "rawMarkdown": "I observed that the entries of the rotation matrix are within the interval [-1,1] and the entries of the translation vector are unbounded.\n\nSo, i was hoping to train two different models for prediction rotation matrix and translation vector respectively. I feel it will be convenient even to choose different loss functions.\n\n\nCan we train two different models for this competition?",
      "votes": null
    },
    {
      "id": "2284145",
      "postDate": "06/01/2023 18:53:01",
      "content": "<p>You can try, but this is not the usual way of doing structure from motion. There is a reliable pipeline as provided by the organizers to calculate the camera poses in general cases, and it has been mathematically proven and optimized since a long time ago. It would be much more efficient to just try training/using DL models to identify and match feature points in pairs of images and use the geometric relationship of those feature points to calculate the relative camera poses. </p>\n<p>Training a DL model to calculate the rotation and translation matrices separately is not a good idea, one thing you need to make sure is that the relative poses between each pairs of images are consistent (you can't make the model output one R and one T matrix from a single image), and that alone is not easy to do. I suggest you go with the well established traditional approach. </p>",
      "rawMarkdown": "You can try, but this is not the usual way of doing structure from motion. There is a reliable pipeline as provided by the organizers to calculate the camera poses in general cases, and it has been mathematically proven and optimized since a long time ago. It would be much more efficient to just try training/using DL models to identify and match feature points in pairs of images and use the geometric relationship of those feature points to calculate the relative camera poses. \n\nTraining a DL model to calculate the rotation and translation matrices separately is not a good idea, one thing you need to make sure is that the relative poses between each pairs of images are consistent (you can't make the model output one R and one T matrix from a single image), and that alone is not easy to do. I suggest you go with the well established traditional approach.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2284145,
      "author_name": "anonymousyuxiang",
      "author_url": "",
      "post_date": "06/01/2023 18:53:01",
      "content": "<p>You can try, but this is not the usual way of doing structure from motion. There is a reliable pipeline as provided by the organizers to calculate the camera poses in general cases, and it has been mathematically proven and optimized since a long time ago. It would be much more efficient to just try training/using DL models to identify and match feature points in pairs of images and use the geometric relationship of those feature points to calculate the relative camera poses. </p>\n<p>Training a DL model to calculate the rotation and translation matrices separately is not a good idea, one thing you need to make sure is that the relative poses between each pairs of images are consistent (you can't make the model output one R and one T matrix from a single image), and that alone is not easy to do. I suggest you go with the well established traditional approach. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2282567": "I observed that the entries of the rotation matrix are within the interval [-1,1] and the entries of the translation vector are unbounded.\n\nSo, i was hoping to train two different models for prediction rotation matrix and translation vector respectively. I feel it will be convenient even to choose different loss functions.\n\n\nCan we train two different models for this competition?",
    "2284145": "You can try, but this is not the usual way of doing structure from motion. There is a reliable pipeline as provided by the organizers to calculate the camera poses in general cases, and it has been mathematically proven and optimized since a long time ago. It would be much more efficient to just try training/using DL models to identify and match feature points in pairs of images and use the geometric relationship of those feature points to calculate the relative camera poses. \n\nTraining a DL model to calculate the rotation and translation matrices separately is not a good idea, one thing you need to make sure is that the relative poses between each pairs of images are consistent (you can't make the model output one R and one T matrix from a single image), and that alone is not easy to do. I suggest you go with the well established traditional approach."
  },
  "source": "meta"
}