{
  "id": 497071,
  "title": "Training Data mismatch ",
  "url": "/competitions/image-matching-challenge-2024/discussion/497071",
  "author_name": "",
  "post_date": "2024-04-23T14:03:05.001541200Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>In the training dataset, for each scene, there is an sfm folder which contains the file <strong>images.txt</strong>. In this file, the following data is present:</p>\n<blockquote>\n  <p>Image list with two lines of data per image:<br>\n  IMAGE_ID, QW, QX, QY, QZ, TX, TY, TZ, CAMERA_ID, NAME<br>\n  POINTS2D[] as (X, Y, POINT3D_ID)</p>\n</blockquote>\n<p>From this, I understand that QW, QX, QY, and QZ represent the rotation in quaternion (which I transformed into an R matrix), and TX, TY, and TZ represent the translation. However, in the file <strong>train_labels.csv,</strong> the rotation and translation matrices do not match with the aforementioned ones.</p>\n<p>Furthermore, in the same file, there is a calibration matrix. So, for these training data, do we know the intrinsics, right? Because I read in another discussion that we don't know them. <br>\n<a href=\"https://www.kaggle.com/oldufo\" target=\"_blank\">@oldufo</a> <a href=\"https://www.kaggle.com/fabiobellavia\" target=\"_blank\">@fabiobellavia</a> </p>",
  "messages": [
    {
      "id": "2769697",
      "postDate": "04/23/2024 14:03:05",
      "content": "<p>In the training dataset, for each scene, there is an sfm folder which contains the file <strong>images.txt</strong>. In this file, the following data is present:</p>\n<blockquote>\n  <p>Image list with two lines of data per image:<br>\n  IMAGE_ID, QW, QX, QY, QZ, TX, TY, TZ, CAMERA_ID, NAME<br>\n  POINTS2D[] as (X, Y, POINT3D_ID)</p>\n</blockquote>\n<p>From this, I understand that QW, QX, QY, and QZ represent the rotation in quaternion (which I transformed into an R matrix), and TX, TY, and TZ represent the translation. However, in the file <strong>train_labels.csv,</strong> the rotation and translation matrices do not match with the aforementioned ones.</p>\n<p>Furthermore, in the same file, there is a calibration matrix. So, for these training data, do we know the intrinsics, right? Because I read in another discussion that we don't know them. <br>\n<a href=\"https://www.kaggle.com/oldufo\" target=\"_blank\">@oldufo</a> <a href=\"https://www.kaggle.com/fabiobellavia\" target=\"_blank\">@fabiobellavia</a> </p>",
      "rawMarkdown": "In the training dataset, for each scene, there is an sfm folder which contains the file **images.txt**. In this file, the following data is present:\n>Image list with two lines of data per image:\nIMAGE_ID, QW, QX, QY, QZ, TX, TY, TZ, CAMERA_ID, NAME\nPOINTS2D[] as (X, Y, POINT3D_ID)\n\nFrom this, I understand that QW, QX, QY, and QZ represent the rotation in quaternion (which I transformed into an R matrix), and TX, TY, and TZ represent the translation. However, in the file **train_labels.csv,** the rotation and translation matrices do not match with the aforementioned ones.\n\nFurthermore, in the same file, there is a calibration matrix. So, for these training data, do we know the intrinsics, right? Because I read in another discussion that we don't know them. \n@oldufo @fabiobellavia",
      "votes": null
    },
    {
      "id": "2771510",
      "postDate": "04/24/2024 09:12:31",
      "content": "<blockquote>\n  <p>From this, I understand that QW, QX, QY, and QZ represent the rotation in quaternion (which I transformed into an R matrix), and TX, TY, and TZ represent the translation. However, in the file train_labels.csv, the rotation and translation matrices do not match with the aforementioned ones.</p>\n</blockquote>\n<p>That is not true. If you read the reconstruction with colmap like:</p>\n<pre><code>import pycolmap\nrec = pycolmap()\nimg=rec\nprint (())\n</code></pre>\n<p>You will get:</p>\n<pre><code>array([[- ,  ,  ],\n       [-,   , -],\n       [-, -,  ]])\n</code></pre>\n<p>And for the train_labels.csv:</p>\n<pre><code>import pandas as pd\ndf =pd()\nchurch_df = df==]\nimg111=church_df==]\n)\n</code></pre>\n<p>You will get </p>\n<pre><code>['-0.;0.734;0.1;-0.484;0.65;-0.653;-0.91;-0.395;0.56']\n</code></pre>\n<p>Which is exactly the same as rotation matrix. <br>\nRegarding the translation, though, the translation in colmap format is different from the <code>train_labels.csv</code> by some constant. The colmap translation is scale-less, whereas <code>train_labels.csv</code> gives you things in meters. <br>\nHowever, given that evaluation metric aligns the provided poses to the GT poses, this scale factor does not matter.</p>\n<blockquote>\n  <p>Furthermore, in the same file, there is a calibration matrix. So, for these training data, do we know the intrinsics, right? Because I read in another discussion that we don't know them.</p>\n</blockquote>\n<p>You know everything for the training data. For the test data you have only images. </p>",
      "rawMarkdown": ">From this, I understand that QW, QX, QY, and QZ represent the rotation in quaternion (which I transformed into an R matrix), and TX, TY, and TZ represent the translation. However, in the file train_labels.csv, the rotation and translation matrices do not match with the aforementioned ones.\n\nThat is not true. If you read the reconstruction with colmap like:\n\n```python3\nimport pycolmap\nrec = pycolmap.Reconstruction('path_to_church')\nimg=rec.images[111]\nprint (img.cam_from_world.rotation.matrix())\n```\n\nYou will get:\n\n```\narray([[-0.0180495 ,  0.40600763,  0.91369142],\n       [-0.47979086,  0.7982316 , -0.36417996],\n       [-0.87719721, -0.44495406,  0.18039107]])\n```\n\nAnd for the train_labels.csv:\n```python3\nimport pandas as pd\ndf =pd.read_csv('train_labels.csv')\nchurch_df = df[df['dataset']=='church']\nimg111=church_df[church_df['image_name']=='00111.png']\nprint(img111.rotation_matrix.to_numpy())\n```\n\nYou will get \n```\n['-0.018049501107115562;0.40600763053442734;0.913691424638321;-0.47979085925197484;0.7982316005790165;-0.36417995993095653;-0.8771972109440591;-0.44495406030834395;0.1803910677585856']\n```\n\nWhich is exactly the same as rotation matrix. \nRegarding the translation, though, the translation in colmap format is different from the `train_labels.csv` by some constant. The colmap translation is scale-less, whereas `train_labels.csv` gives you things in meters. \nHowever, given that evaluation metric aligns the provided poses to the GT poses, this scale factor does not matter.\n\n\n>Furthermore, in the same file, there is a calibration matrix. So, for these training data, do we know the intrinsics, right? Because I read in another discussion that we don't know them.\n\nYou know everything for the training data. For the test data you have only images.",
      "votes": null
    },
    {
      "id": "2772069",
      "postDate": "04/24/2024 14:11:42",
      "content": "<p>Yes, you are right. I made a mistake in my calculations. Thank you for your answer! What are the units for the focal length and the principal points?</p>",
      "rawMarkdown": "Yes, you are right. I made a mistake in my calculations. Thank you for your answer! What are the units for the focal length and the principal points?",
      "votes": null
    },
    {
      "id": "2773585",
      "postDate": "04/24/2024 19:42:36",
      "content": "<p>They are pixels.</p>",
      "rawMarkdown": "They are pixels.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2771510,
      "author_name": "oldufo",
      "author_url": "",
      "post_date": "04/24/2024 09:12:31",
      "content": "<blockquote>\n  <p>From this, I understand that QW, QX, QY, and QZ represent the rotation in quaternion (which I transformed into an R matrix), and TX, TY, and TZ represent the translation. However, in the file train_labels.csv, the rotation and translation matrices do not match with the aforementioned ones.</p>\n</blockquote>\n<p>That is not true. If you read the reconstruction with colmap like:</p>\n<pre><code>import pycolmap\nrec = pycolmap()\nimg=rec\nprint (())\n</code></pre>\n<p>You will get:</p>\n<pre><code>array([[- ,  ,  ],\n       [-,   , -],\n       [-, -,  ]])\n</code></pre>\n<p>And for the train_labels.csv:</p>\n<pre><code>import pandas as pd\ndf =pd()\nchurch_df = df==]\nimg111=church_df==]\n)\n</code></pre>\n<p>You will get </p>\n<pre><code>['-0.;0.734;0.1;-0.484;0.65;-0.653;-0.91;-0.395;0.56']\n</code></pre>\n<p>Which is exactly the same as rotation matrix. <br>\nRegarding the translation, though, the translation in colmap format is different from the <code>train_labels.csv</code> by some constant. The colmap translation is scale-less, whereas <code>train_labels.csv</code> gives you things in meters. <br>\nHowever, given that evaluation metric aligns the provided poses to the GT poses, this scale factor does not matter.</p>\n<blockquote>\n  <p>Furthermore, in the same file, there is a calibration matrix. So, for these training data, do we know the intrinsics, right? Because I read in another discussion that we don't know them.</p>\n</blockquote>\n<p>You know everything for the training data. For the test data you have only images. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2772069,
          "author_name": "silavassiliou",
          "author_url": "",
          "post_date": "04/24/2024 14:11:42",
          "content": "<p>Yes, you are right. I made a mistake in my calculations. Thank you for your answer! What are the units for the focal length and the principal points?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2773585,
              "author_name": "oldufo",
              "author_url": "",
              "post_date": "04/24/2024 19:42:36",
              "content": "<p>They are pixels.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2769697": "In the training dataset, for each scene, there is an sfm folder which contains the file **images.txt**. In this file, the following data is present:\n>Image list with two lines of data per image:\nIMAGE_ID, QW, QX, QY, QZ, TX, TY, TZ, CAMERA_ID, NAME\nPOINTS2D[] as (X, Y, POINT3D_ID)\n\nFrom this, I understand that QW, QX, QY, and QZ represent the rotation in quaternion (which I transformed into an R matrix), and TX, TY, and TZ represent the translation. However, in the file **train_labels.csv,** the rotation and translation matrices do not match with the aforementioned ones.\n\nFurthermore, in the same file, there is a calibration matrix. So, for these training data, do we know the intrinsics, right? Because I read in another discussion that we don't know them. \n@oldufo @fabiobellavia",
    "2771510": ">From this, I understand that QW, QX, QY, and QZ represent the rotation in quaternion (which I transformed into an R matrix), and TX, TY, and TZ represent the translation. However, in the file train_labels.csv, the rotation and translation matrices do not match with the aforementioned ones.\n\nThat is not true. If you read the reconstruction with colmap like:\n\n```python3\nimport pycolmap\nrec = pycolmap.Reconstruction('path_to_church')\nimg=rec.images[111]\nprint (img.cam_from_world.rotation.matrix())\n```\n\nYou will get:\n\n```\narray([[-0.0180495 ,  0.40600763,  0.91369142],\n       [-0.47979086,  0.7982316 , -0.36417996],\n       [-0.87719721, -0.44495406,  0.18039107]])\n```\n\nAnd for the train_labels.csv:\n```python3\nimport pandas as pd\ndf =pd.read_csv('train_labels.csv')\nchurch_df = df[df['dataset']=='church']\nimg111=church_df[church_df['image_name']=='00111.png']\nprint(img111.rotation_matrix.to_numpy())\n```\n\nYou will get \n```\n['-0.018049501107115562;0.40600763053442734;0.913691424638321;-0.47979085925197484;0.7982316005790165;-0.36417995993095653;-0.8771972109440591;-0.44495406030834395;0.1803910677585856']\n```\n\nWhich is exactly the same as rotation matrix. \nRegarding the translation, though, the translation in colmap format is different from the `train_labels.csv` by some constant. The colmap translation is scale-less, whereas `train_labels.csv` gives you things in meters. \nHowever, given that evaluation metric aligns the provided poses to the GT poses, this scale factor does not matter.\n\n\n>Furthermore, in the same file, there is a calibration matrix. So, for these training data, do we know the intrinsics, right? Because I read in another discussion that we don't know them.\n\nYou know everything for the training data. For the test data you have only images.",
    "2772069": "Yes, you are right. I made a mistake in my calculations. Thank you for your answer! What are the units for the focal length and the principal points?",
    "2773585": "They are pixels."
  },
  "source": "meta"
}