{
  "id": 402361,
  "title": "A question about evaluation methodology",
  "url": "/competitions/image-matching-challenge-2023/discussion/402361",
  "author_name": "",
  "post_date": "2023-04-18T04:02:22.224740800Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I am trying to understand how evaluation work? As I could understand:</p>\n<ol>\n<li>The submission contains camera pose estimation for each image. So the output is one transformation matrix per image, splitted to rotation and translation matrix.</li>\n<li>The camera pose estimation is obtained from the set of transformation matrix aligning each pair of the images in the scene</li>\n</ol>\n<p>This is forward process<br>\nNow, the evaluation is based on the accuracy of the p.2, if I understand it correctly. This is intermediate step and its results are not included in submission. So, in order to use them in the evaluation, they are <strong>recreated during the evaluation</strong> by one of the possible methodologies. What is this methology doesn't matter to competitors. Do I understand it correctly, so far?</p>",
  "messages": [
    {
      "id": "2225281",
      "postDate": "04/18/2023 04:02:22",
      "content": "<p>I am trying to understand how evaluation work? As I could understand:</p>\n<ol>\n<li>The submission contains camera pose estimation for each image. So the output is one transformation matrix per image, splitted to rotation and translation matrix.</li>\n<li>The camera pose estimation is obtained from the set of transformation matrix aligning each pair of the images in the scene</li>\n</ol>\n<p>This is forward process<br>\nNow, the evaluation is based on the accuracy of the p.2, if I understand it correctly. This is intermediate step and its results are not included in submission. So, in order to use them in the evaluation, they are <strong>recreated during the evaluation</strong> by one of the possible methodologies. What is this methology doesn't matter to competitors. Do I understand it correctly, so far?</p>",
      "rawMarkdown": "I am trying to understand how evaluation work? As I could understand:\n1. The submission contains camera pose estimation for each image. So the output is one transformation matrix per image, splitted to rotation and translation matrix.\n2. The camera pose estimation is obtained from the set of transformation matrix aligning each pair of the images in the scene\n\nThis is forward process\nNow, the evaluation is based on the accuracy of the p.2, if I understand it correctly. This is intermediate step and its results are not included in submission. So, in order to use them in the evaluation, they are **recreated during the evaluation** by one of the possible methodologies. What is this methology doesn't matter to competitors. Do I understand it correctly, so far?",
      "votes": null
    },
    {
      "id": "2225483",
      "postDate": "04/18/2023 07:44:05",
      "content": "<p>Sorry, I'm having trouble parsing your post. Evaluating the quality of a 3D reconstruction is not trivial. Reconstructions will typically have a different scale (one point cloud being \"larger\" than the other). Not all images may be registered. We can't match ground truth/predicted point clouds because we don't have the predictions, and they may have very different properties (your solution may not even have point clouds, we ask only for poses). There are different ways to do this, some more costly than others.</p>\n<p>We chose a very simple metric: we evaluate pairwise pose accuracy, for every possible pair of images. Maybe it's easier to understand if you look at <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2022/overview\" target=\"_blank\">last year's</a> metric, which is quite similar. This year's metric is <a href=\"https://www.kaggle.com/code/eduardtrulls/imc2023-evaluation\" target=\"_blank\">here</a>: you may want to try running it on one of the small scenes (or just perturb the ground truth) and trivially output the metric for every pair of images.</p>",
      "rawMarkdown": "Sorry, I'm having trouble parsing your post. Evaluating the quality of a 3D reconstruction is not trivial. Reconstructions will typically have a different scale (one point cloud being \"larger\" than the other). Not all images may be registered. We can't match ground truth/predicted point clouds because we don't have the predictions, and they may have very different properties (your solution may not even have point clouds, we ask only for poses). There are different ways to do this, some more costly than others.\n\nWe chose a very simple metric: we evaluate pairwise pose accuracy, for every possible pair of images. Maybe it's easier to understand if you look at [last year's](https://www.kaggle.com/competitions/image-matching-challenge-2022/overview) metric, which is quite similar. This year's metric is [here](https://www.kaggle.com/code/eduardtrulls/imc2023-evaluation): you may want to try running it on one of the small scenes (or just perturb the ground truth) and trivially output the metric for every pair of images.",
      "votes": null
    },
    {
      "id": "2225540",
      "postDate": "04/18/2023 08:32:27",
      "content": "<p><a href=\"https://www.kaggle.com/eduardtrulls\" target=\"_blank\">@eduardtrulls</a>, thank you for your response.<br>\nSo as I could understand from the code, your metrics works something like that. <strong>For each pair of images</strong> belonging to the same scene:</p>\n<ol>\n<li>You take camera poses from ground truth and calculate some metric that I will refer as <em>pose difference</em> (you also could pre-calculate this difference only once and save it to file).</li>\n<li>Now you calculate the <em>pose difference</em> for the submitted estimated poses for the same pair of images</li>\n<li>Finally, you calculate the error between two <em>pose differences</em> (GT and estimated).</li>\n</ol>\n<p>I intentionally skipped math details of calculations, in order to keep strict to the concept.<br>\nFinally you combine all errors between pose differences and get a final score. Is that correct?</p>",
      "rawMarkdown": "eduardtrulls, thank you for your response.\nSo as I could understand from the code, your metrics works something like that. **For each pair of images** belonging to the same scene:\n1.  You take camera poses from ground truth and calculate some metric that I will refer as *pose difference* (you also could pre-calculate this difference only once and save it to file).\n2. Now you calculate the *pose difference* for the submitted estimated poses for the same pair of images\n3. Finally, you calculate the error between two *pose differences* (GT and estimated).\n\nI intentionally skipped math details of calculations, in order to keep strict to the concept.\nFinally you combine all errors between pose differences and get a final score. Is that correct?",
      "votes": null
    },
    {
      "id": "2225832",
      "postDate": "04/18/2023 13:13:18",
      "content": "<p>Yes, that is correct </p>",
      "rawMarkdown": "Yes, that is correct",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2225483,
      "author_name": "eduardtrulls",
      "author_url": "",
      "post_date": "04/18/2023 07:44:05",
      "content": "<p>Sorry, I'm having trouble parsing your post. Evaluating the quality of a 3D reconstruction is not trivial. Reconstructions will typically have a different scale (one point cloud being \"larger\" than the other). Not all images may be registered. We can't match ground truth/predicted point clouds because we don't have the predictions, and they may have very different properties (your solution may not even have point clouds, we ask only for poses). There are different ways to do this, some more costly than others.</p>\n<p>We chose a very simple metric: we evaluate pairwise pose accuracy, for every possible pair of images. Maybe it's easier to understand if you look at <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2022/overview\" target=\"_blank\">last year's</a> metric, which is quite similar. This year's metric is <a href=\"https://www.kaggle.com/code/eduardtrulls/imc2023-evaluation\" target=\"_blank\">here</a>: you may want to try running it on one of the small scenes (or just perturb the ground truth) and trivially output the metric for every pair of images.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2225540,
          "author_name": "yurikreinin",
          "author_url": "",
          "post_date": "04/18/2023 08:32:27",
          "content": "<p><a href=\"https://www.kaggle.com/eduardtrulls\" target=\"_blank\">@eduardtrulls</a>, thank you for your response.<br>\nSo as I could understand from the code, your metrics works something like that. <strong>For each pair of images</strong> belonging to the same scene:</p>\n<ol>\n<li>You take camera poses from ground truth and calculate some metric that I will refer as <em>pose difference</em> (you also could pre-calculate this difference only once and save it to file).</li>\n<li>Now you calculate the <em>pose difference</em> for the submitted estimated poses for the same pair of images</li>\n<li>Finally, you calculate the error between two <em>pose differences</em> (GT and estimated).</li>\n</ol>\n<p>I intentionally skipped math details of calculations, in order to keep strict to the concept.<br>\nFinally you combine all errors between pose differences and get a final score. Is that correct?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2225832,
              "author_name": "oldufo",
              "author_url": "",
              "post_date": "04/18/2023 13:13:18",
              "content": "<p>Yes, that is correct </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2225281": "I am trying to understand how evaluation work? As I could understand:\n1. The submission contains camera pose estimation for each image. So the output is one transformation matrix per image, splitted to rotation and translation matrix.\n2. The camera pose estimation is obtained from the set of transformation matrix aligning each pair of the images in the scene\n\nThis is forward process\nNow, the evaluation is based on the accuracy of the p.2, if I understand it correctly. This is intermediate step and its results are not included in submission. So, in order to use them in the evaluation, they are **recreated during the evaluation** by one of the possible methodologies. What is this methology doesn't matter to competitors. Do I understand it correctly, so far?",
    "2225483": "Sorry, I'm having trouble parsing your post. Evaluating the quality of a 3D reconstruction is not trivial. Reconstructions will typically have a different scale (one point cloud being \"larger\" than the other). Not all images may be registered. We can't match ground truth/predicted point clouds because we don't have the predictions, and they may have very different properties (your solution may not even have point clouds, we ask only for poses). There are different ways to do this, some more costly than others.\n\nWe chose a very simple metric: we evaluate pairwise pose accuracy, for every possible pair of images. Maybe it's easier to understand if you look at [last year's](https://www.kaggle.com/competitions/image-matching-challenge-2022/overview) metric, which is quite similar. This year's metric is [here](https://www.kaggle.com/code/eduardtrulls/imc2023-evaluation): you may want to try running it on one of the small scenes (or just perturb the ground truth) and trivially output the metric for every pair of images.",
    "2225540": "eduardtrulls, thank you for your response.\nSo as I could understand from the code, your metrics works something like that. **For each pair of images** belonging to the same scene:\n1.  You take camera poses from ground truth and calculate some metric that I will refer as *pose difference* (you also could pre-calculate this difference only once and save it to file).\n2. Now you calculate the *pose difference* for the submitted estimated poses for the same pair of images\n3. Finally, you calculate the error between two *pose differences* (GT and estimated).\n\nI intentionally skipped math details of calculations, in order to keep strict to the concept.\nFinally you combine all errors between pose differences and get a final score. Is that correct?",
    "2225832": "Yes, that is correct"
  },
  "source": "meta"
}