{
  "id": 571771,
  "title": "Clarification on Rotation/Translation Reference and Ground Truth Availability",
  "url": "/competitions/image-matching-challenge-2025/discussion/571771",
  "author_name": "",
  "post_date": "2025-04-05T13:47:27.242336100Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi, I’ve been reading through the dataset description and past discussions, and I have a few important questions that I haven't been able to resolve yet. I hope these can help clarify things for others as well😄</p>\n<h1>1. What is the reference frame for the rotation matrix and translation vector?</h1>\n<p>In the submission, we are expected to provide a rotation matrix <code>R</code> and a translation vector <code>T</code> for each image.<br>\nHowever, I haven't found a clear explanation of <strong>what these are defined relative to</strong>. <br>\nIs the world coordinate system arbitrary per cluster? <br>\nAre we free to define it as long as the camera poses are consistent within the scene?</p>\n<h1>2. The train_labels.csv file; \"Sample values are random\" — what does that mean exactly?</h1>\n<p>I found the information for <code>rotation_matrix</code> and <code>translation_vector</code>, <strong>\"Sample values are random\"</strong>, in Dataset Description.<br>\nDoes it mean they in <code>train_labels.csv</code> are not actual ground truth values, but rather random placeholders?<br>\nIf so, we cannot use them to understand what the coordinate system of each scene looks like.</p>\n<h1>3. If ground truth R and T values are not available, how can we evaluate our models during development?</h1>\n<p>Since we don’t know what the poses should be, and the provided R/T in training are random, does that mean we must rely solely on the leaderboard score to judge whether our method is improving?<br>\nHave previous participants used proxy tasks like clustering evaluation to guide development?</p>\n<hr>\n<p>I’m concerned that without access to at least some reliable R/T values, we cannot understand how the poses are defined, nor compare outputs to anything during development.<br>\nIf there's any information or best practice on this that I may have missed, I’d really appreciate it if someone could clarify. Thank you!</p>",
  "messages": [
    {
      "id": "3171274",
      "postDate": "04/05/2025 13:47:27",
      "content": "<p>Hi, I’ve been reading through the dataset description and past discussions, and I have a few important questions that I haven't been able to resolve yet. I hope these can help clarify things for others as well😄</p>\n<h1>1. What is the reference frame for the rotation matrix and translation vector?</h1>\n<p>In the submission, we are expected to provide a rotation matrix <code>R</code> and a translation vector <code>T</code> for each image.<br>\nHowever, I haven't found a clear explanation of <strong>what these are defined relative to</strong>. <br>\nIs the world coordinate system arbitrary per cluster? <br>\nAre we free to define it as long as the camera poses are consistent within the scene?</p>\n<h1>2. The train_labels.csv file; \"Sample values are random\" — what does that mean exactly?</h1>\n<p>I found the information for <code>rotation_matrix</code> and <code>translation_vector</code>, <strong>\"Sample values are random\"</strong>, in Dataset Description.<br>\nDoes it mean they in <code>train_labels.csv</code> are not actual ground truth values, but rather random placeholders?<br>\nIf so, we cannot use them to understand what the coordinate system of each scene looks like.</p>\n<h1>3. If ground truth R and T values are not available, how can we evaluate our models during development?</h1>\n<p>Since we don’t know what the poses should be, and the provided R/T in training are random, does that mean we must rely solely on the leaderboard score to judge whether our method is improving?<br>\nHave previous participants used proxy tasks like clustering evaluation to guide development?</p>\n<hr>\n<p>I’m concerned that without access to at least some reliable R/T values, we cannot understand how the poses are defined, nor compare outputs to anything during development.<br>\nIf there's any information or best practice on this that I may have missed, I’d really appreciate it if someone could clarify. Thank you!</p>",
      "rawMarkdown": "Hi, I’ve been reading through the dataset description and past discussions, and I have a few important questions that I haven't been able to resolve yet. I hope these can help clarify things for others as well😄\n\n# 1. What is the reference frame for the rotation matrix and translation vector?\nIn the submission, we are expected to provide a rotation matrix `R` and a translation vector `T` for each image.\nHowever, I haven't found a clear explanation of **what these are defined relative to**. \nIs the world coordinate system arbitrary per cluster? \nAre we free to define it as long as the camera poses are consistent within the scene?\n\n# 2. The train_labels.csv file; \"Sample values are random\" — what does that mean exactly?\nI found the information for `rotation_matrix` and `translation_vector`, **\"Sample values are random\"**, in Dataset Description.\nDoes it mean they in `train_labels.csv` are not actual ground truth values, but rather random placeholders?\nIf so, we cannot use them to understand what the coordinate system of each scene looks like.\n\n# 3. If ground truth R and T values are not available, how can we evaluate our models during development?\nSince we don’t know what the poses should be, and the provided R/T in training are random, does that mean we must rely solely on the leaderboard score to judge whether our method is improving?\nHave previous participants used proxy tasks like clustering evaluation to guide development?\n\n---\nI’m concerned that without access to at least some reliable R/T values, we cannot understand how the poses are defined, nor compare outputs to anything during development.\nIf there's any information or best practice on this that I may have missed, I’d really appreciate it if someone could clarify. Thank you!",
      "votes": null
    },
    {
      "id": "3171301",
      "postDate": "04/05/2025 14:35:56",
      "content": "<p>Hi,<br>\nyou can set any consistent reference system for each scene, the metric is designed to maximize the final score; the code of the full metric is available inside the util scripts provided. </p>\n<p>For further details on this and your other doubts please refers to the metric description and the discussion section of the past edition (<a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2024)\" target=\"_blank\">https://www.kaggle.com/competitions/image-matching-challenge-2024)</a>. </p>",
      "rawMarkdown": "Hi,\nyou can set any consistent reference system for each scene, the metric is designed to maximize the final score; the code of the full metric is available inside the util scripts provided. \n\nFor further details on this and your other doubts please refers to the metric description and the discussion section of the past edition (https://www.kaggle.com/competitions/image-matching-challenge-2024).",
      "votes": null
    },
    {
      "id": "3171686",
      "postDate": "04/06/2025 00:54:32",
      "content": "<p>Thanks for your reply! <br>\nNow, I understand what you said and the concept of IMC2025. <br>\nAnd, I found similar question in IMC2024. <br>\n<a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2024/discussion/487622\" target=\"_blank\">https://www.kaggle.com/competitions/image-matching-challenge-2024/discussion/487622</a></p>",
      "rawMarkdown": "Thanks for your reply! \nNow, I understand what you said and the concept of IMC2025. \nAnd, I found similar question in IMC2024. \nhttps://www.kaggle.com/competitions/image-matching-challenge-2024/discussion/487622",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3171301,
      "author_name": "fabiobellavia",
      "author_url": "",
      "post_date": "04/05/2025 14:35:56",
      "content": "<p>Hi,<br>\nyou can set any consistent reference system for each scene, the metric is designed to maximize the final score; the code of the full metric is available inside the util scripts provided. </p>\n<p>For further details on this and your other doubts please refers to the metric description and the discussion section of the past edition (<a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2024)\" target=\"_blank\">https://www.kaggle.com/competitions/image-matching-challenge-2024)</a>. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3171686,
          "author_name": "thoth000",
          "author_url": "",
          "post_date": "04/06/2025 00:54:32",
          "content": "<p>Thanks for your reply! <br>\nNow, I understand what you said and the concept of IMC2025. <br>\nAnd, I found similar question in IMC2024. <br>\n<a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2024/discussion/487622\" target=\"_blank\">https://www.kaggle.com/competitions/image-matching-challenge-2024/discussion/487622</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3171274": "Hi, I’ve been reading through the dataset description and past discussions, and I have a few important questions that I haven't been able to resolve yet. I hope these can help clarify things for others as well😄\n\n# 1. What is the reference frame for the rotation matrix and translation vector?\nIn the submission, we are expected to provide a rotation matrix `R` and a translation vector `T` for each image.\nHowever, I haven't found a clear explanation of **what these are defined relative to**. \nIs the world coordinate system arbitrary per cluster? \nAre we free to define it as long as the camera poses are consistent within the scene?\n\n# 2. The train_labels.csv file; \"Sample values are random\" — what does that mean exactly?\nI found the information for `rotation_matrix` and `translation_vector`, **\"Sample values are random\"**, in Dataset Description.\nDoes it mean they in `train_labels.csv` are not actual ground truth values, but rather random placeholders?\nIf so, we cannot use them to understand what the coordinate system of each scene looks like.\n\n# 3. If ground truth R and T values are not available, how can we evaluate our models during development?\nSince we don’t know what the poses should be, and the provided R/T in training are random, does that mean we must rely solely on the leaderboard score to judge whether our method is improving?\nHave previous participants used proxy tasks like clustering evaluation to guide development?\n\n---\nI’m concerned that without access to at least some reliable R/T values, we cannot understand how the poses are defined, nor compare outputs to anything during development.\nIf there's any information or best practice on this that I may have missed, I’d really appreciate it if someone could clarify. Thank you!",
    "3171301": "Hi,\nyou can set any consistent reference system for each scene, the metric is designed to maximize the final score; the code of the full metric is available inside the util scripts provided. \n\nFor further details on this and your other doubts please refers to the metric description and the discussion section of the past edition (https://www.kaggle.com/competitions/image-matching-challenge-2024).",
    "3171686": "Thanks for your reply! \nNow, I understand what you said and the concept of IMC2025. \nAnd, I found similar question in IMC2024. \nhttps://www.kaggle.com/competitions/image-matching-challenge-2024/discussion/487622"
  },
  "source": "meta"
}