{
  "id": 573183,
  "title": "All the THEORY you need to start Image Matching",
  "url": "/competitions/image-matching-challenge-2025/discussion/573183",
  "author_name": "Dmytro Buhai",
  "post_date": "2025-04-14T05:49:17.749000",
  "votes": 57,
  "comment_count": 6,
  "views": 0,
  "content": "<p></p><h2>📂 Dataset Structure</h2><p></p>\n<p>Each dataset consists of multiple scenes. A scene is a group of related images that belong together (e.g., different viewpoints of the same object). Some images in a dataset may be <strong>outliers</strong>, meaning they don’t belong to any scene (e.g., random photos taken in the area).</p>\n<p>In the <strong>test data</strong>, all images are mixed in one folder. Your task is:</p>\n<ol>\n<li>Cluster the images into groups that represent scenes  </li>\n<li>Assign each image to a cluster or label it as an <code>outlier</code>  </li>\n<li>Reconstruct the camera pose (R, T) for each image (except outliers)</li>\n</ol>\n<hr>\n<p></p><h1>🧠 Camera Pose Estimation and mAA</h1><p></p>\n<ul>\n<li>A rotation matrix (3×3) - where the camera is pointing  </li>\n<li>A translation vector (3D) - where the camera is located  </li>\n</ul>\n<p><b>\"Ground truth\"</b> is the true known pose of the camera when the image was taken. In training data, this is given to you. In test data, you try to predict it.</p>\n<p>Camera center <b>C</b> in world coordinates is calculated as:  <br>\n<strong>C = −R · T</strong><br>\nThis transforms the camera pose into a 3D point in space.</p>\n<p>Camera centres of one scene:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F25747348%2F61b7f20998538db19346d6d34c961920%2F2025-04-14%20083034.jpg?generation=1744608643370046&amp;alt=media\" alt=\"\"></p>\n<hr>\n<h3><b>Structure from Motion (SfM)</b> is a pipeline that reconstructs:</h3>\n<ul>\n<li>Camera poses for each image  </li>\n<li>A sparse 3D point cloud of the scene</li>\n</ul>\n<p>Your predicted camera centers may be rotated differently, shifted (translated), and scaled.  <br>\nThis is normal in SfM - reconstructions can be geometrically correct but misaligned.</p>\n<p>So, to compare your predicted camera centers <b>C</b> to the ground truth centers <b>Cg</b>, a similarity transformation <b>T</b> (i.e. scale, rotation and translation altogether) is applied. This aligns your 3D camera predictions to the ground truth up to scale and orientation.  <br>\n(Some of your predicted camera poses may be wrong - far from ground truth - and we want to ignore bad data and only use good matches to compute the best alignment.)</p>\n<p>To compute a similarity transform (rotation, translation, and scale) between two sets of 3D points, you need at least 3 point correspondences that are not collinear.</p>\n<hr>\n<p>A <strong>triplet</strong> is a group of three matching camera centers. Each one has:</p>\n<ul>\n<li>A predicted 3D position <b>Ci</b>  </li>\n<li>A known ground-truth position <b>Cgi</b>  </li>\n</ul>\n<p>So a triplet is:  <br>\n<code>{(C1, Cg1), (C2, Cg2), (C3, Cg3)}</code></p>\n<p><b>RANSAC</b> = Random Sample Consensus - Algorithm to randomly test many triplets, ignore outliers, and find the best transformation.</p>\n<p>The <code>thresholds</code> column lists multiple numeric values per scene. These represent distance or error thresholds that are used to evaluate how accurate predicted 3D camera poses are (compared to ground truth). Each threshold is a maximum allowable error (in meters) to consider a pose prediction as correct.</p>\n<hr>\n<p>So use <b>Horn’s method</b> (algorithm used to compute the best similarity transformation between two sets of 3D points) to compute a similarity transformation.  <br>\nFind scale <b>s</b>, rotation <b>R</b>, translation <b>t</b> so that:</p>\n<pre><code> = s ⋅ R ⋅ Ci + t</code></pre>\n<p>This is our first estimate <b>T′</b>.  <br>\nThen apply <b>T′</b> to all predicted camera centers <b>Ci</b> and for each one count how many of your predictions become <b>registered cameras</b> (i.e., within error threshold of ground truth):</p>\n<pre><code> = (‖ T(Ci) − Cgi ‖ &lt; threshold)</code></pre>\n<p>Where <code>‖ · ‖</code> denotes Euclidean (L2) distance between two 3D points.</p>\n<p>Try many random triplets and get many candidate transformations <b>T′</b>.  <br>\nFor each one, count how many registered cameras it produces and save the best one.  <br>\nTake the best transformation <b>T′</b>, and re-estimate it using all the inliers (not just the original triplet).</p>\n<hr>\n<p></p><h2>📊 mAA = mean Average Accuracy</h2><p></p>\n<p>It measures how accurately your predicted camera centers match the ground truth, after alignment with a similarity transformation.</p>\n<p>As discussed before, you use triplets and Horn’s method inside RANSAC.  <br>\nYou find the best transformation:</p>\n<pre><code> = s ⋅ R ⋅ Ci + t</code></pre>\n<p>This aligns your predicted camera centers to the same coordinate frame as the ground-truth.  <br>\nThen apply this transformation - for each predicted camera center <b>Ci</b>, compute:  <br>\n<code>Ci_aligned = T(Ci)</code></p>\n<p>Now for each threshold <b>t_j</b>, check if the difference between <b>Ci_aligned</b> and ground-truth center <b>Cgi</b> is less than <b>t_j</b>:</p>\n<pre><code> = (‖ T(Ci) − Cgi ‖ &lt; threshold)</code></pre>\n<p>Then compute registration ratio per threshold for scene:  <br>\nFor each <b>N' = N − 3</b> images of the scene and each threshold <b>t_j</b>, compute the percentage of registered cameras:</p>\n<pre><code>rj = ( / ') * (registered)</code></pre>\n<p>That gives you the <b>Average Accuracy (AA)</b> for that scene.  <br>\nThen find average across all thresholds:</p>\n<pre><code>  ( / k) * sum(rj)</code></pre>\n<p>Where:</p>\n<ul>\n<li><b>k</b> = number of thresholds</li>\n</ul>\n<p>Once you compute <b>mAA per scene</b>, you calculate final mAA for all scenes by averaging across scenes.</p>\n<hr>\n<p></p><h2>🧠 Clustering and Final Evaluation Score</h2><p></p>\n<p>Each scene from the ground-truth data is denoted as <code>S&lt;sub&gt;ki&lt;/sub&gt;</code>, and your predicted clusters are <code>C&lt;sub&gt;kj&lt;/sub&gt;</code>.</p>\n<p>The evaluation will:</p>\n<ul>\n<li>Compare your clusters with the true scenes  </li>\n<li>Check how well your clustering matches the original grouping  </li>\n<li>Measure how accurate your pose predictions are (as explained earlier with mAA)</li>\n</ul>\n<p>Each ground-truth scene <code>S&lt;sub&gt;ki&lt;/sub&gt;</code> is matched to the user-submitted cluster <code>C&lt;sub&gt;kj&lt;/sub&gt;</code> that gives the <b>highest mAA</b> after alignment.  <br>\nAll images labeled as outliers are <b>excluded</b> from this matching step.</p>\n<hr>\n<h3>📊 Clustering Score (Precision)</h3>\n<p>After the best cluster <code>C&lt;sub&gt;kji&lt;/sub&gt;</code> is selected for each scene <code>S&lt;sub&gt;ki&lt;/sub&gt;</code>, compute the <b>clustering score</b> as:</p>\n<pre><code>clustering_score = ||||</code></pre>\n<p>Where:</p>\n<ul>\n<li><code>|S ∩ C|</code> is the number of images that correctly overlap between your predicted cluster and the real scene  </li>\n<li><code>|C|</code> is the total number of images in your predicted cluster</li>\n</ul>\n<p>This measures <b>how pure your cluster is</b> - i.e., out of everything you said belongs to this scene, how many actually do?</p>\n<hr>\n<h3>📈 Pose Score (mAA)</h3>\n<p>As explained earlier, for each image:</p>\n<ul>\n<li>Apply the best alignment transform T to your predicted camera centers  </li>\n<li>Check if each transformed pose is within the threshold of the ground-truth center  </li>\n<li>Count the number of registered cameras at each threshold  </li>\n<li>Compute registration ratios across thresholds  </li>\n<li>Average across thresholds → mAA per scene  </li>\n<li>Average over scenes → mAA for dataset</li>\n</ul>\n<p>This mAA represents <b>recall</b> - how many real scene images you found and aligned correctly.</p>\n<hr>\n<h3>🔗 Final Score Per Dataset</h3>\n<p>Now both mAA (recall-like) and clustering score (precision-like) are combined using the <b>harmonic mean</b>, similar to the F1-score in classification:</p>\n<pre><code> =  * (mAA * clustering_score) / (mAA + clustering_score)</code></pre>\n<p>This ensures that:</p>\n<ul>\n<li>If either clustering or pose accuracy is bad, your score is penalized  </li>\n<li>You must do <b>both clustering and reconstruction well</b> to score high</li>\n</ul>\n<hr>\n<h3>📊 Final Challenge Score</h3>\n<p>You will repeat this evaluation for each dataset <code>D&lt;sub&gt;k&lt;/sub&gt;</code>.  <br>\nThe final score for your submission is:</p>\n<pre><code></code></pre>\n<p>i.e., the average harmonic mean across all datasets.</p>\n<hr>\n<p></p><h1>🧩 Scene Clustering &amp; Pose Estimation (SfM)</h1><p></p>\n<hr>\n<h4></h4><h3>📌 <b>1) Scene Clustering</b></h3>\n<hr>\n<h5></h5><h4>📘 <b>1.1) Unsupervised Clustering on Global Descriptors</b></h4>\n<p>Extract global embeddings using models like <code>CLIP</code>, <code>DINOv2</code>, etc., then cluster them using algorithms like <b>DBSCAN</b>, <b>KMeans</b>, <b>Agglomerative Clustering</b>…</p>\n<p>📌 <strong>Image features</strong> are compact vector representations of the visual content of the image.</p>\n<hr>\n<h5></h5><h4>📦 Global Descriptors</h4>\n<p>Global descriptors are single vectors that represent the entire image - its objects, textures, and layout.</p>\n<p>✅ <strong>Use cases</strong>:</p>\n<ul>\n<li>Image clustering  </li>\n<li>Image retrieval  </li>\n<li>Scene classification  </li>\n</ul>\n<p><br></p>\n<p><strong>DINOv2</strong> is part of the <strong>DINO</strong> (Distillation with No Labels) family of self-supervised vision transformers:</p>\n<ul>\n<li>Learns by matching feature vectors of different augmented views of the same image.</li>\n<li>Ensures different images are mapped to distant feature vectors.</li>\n<li>Requires no manual labels (contrastive learning).</li>\n</ul>\n<hr>\n<h4>🔧 How Global Descriptor Clustering Works</h4>\n<ul>\n<li>Extract a single embedding per image (e.g., using DINOv2).</li>\n<li>Compute pairwise distances (typically cosine distance) between all embeddings.</li>\n<li>Cluster embeddings using:<ul>\n<li><strong>DBSCAN</strong> - density-based clustering without needing the number of clusters.</li>\n<li><strong>Agglomerative Clustering</strong> - hierarchical grouping.</li>\n<li><strong>KMeans</strong> - if the number of clusters is approximately known.</li>\n<li><strong>Graph-based Clustering</strong> - (see below 👇).</li></ul></li>\n<li>The result: each cluster represents one predicted scene.</li>\n</ul>\n<hr>\n<h4></h4><h4>📘 <b> Graph-Based Clustering</b></h4>\n<p>✅ In our pipeline:</p>\n<ul>\n<li>Compute <strong>pairwise cosine distances</strong> between image embeddings extracted by <strong>DINOv2</strong>.</li>\n<li>Build a <strong>graph</strong> where:<ul>\n<li>Each node represents an image.</li>\n<li>An edge is created between two images if their distance is below a threshold (e.g., <code>0.3</code>).</li></ul></li>\n<li><strong>Connected components</strong> of this graph are treated as <strong>predicted scenes</strong>.</li>\n<li>Images with no edges (isolated nodes) are labeled as <strong>outliers</strong>.</li>\n</ul>\n<p>\nThis approach is faster and more scalable than DBSCAN for large datasets, and does not require knowing the number of clusters beforehand.\n</p>\n<p>📌 <strong>Note</strong>:  <br>\nI do not use ALIKED or LightGlue at this stage.  <br>\nThey are used later in the pipeline for <strong>local feature extraction and matching</strong> - when reconstructing 3D camera poses.</p>\n<hr>\n<h5></h5><h4>🔍 Local Keypoints &amp; Descriptors (General Theory)</h4>\n<p>While  my solution clusters images using global descriptors, another common approach in image matching is to use <strong>local features</strong>:</p>\n<p><strong>Keypoints</strong> are distinctive image regions (corners, blobs, edges) that:</p>\n<ul>\n<li>Are stable under changes in viewpoint and lighting.</li>\n<li>Can be reliably detected and matched across different images.</li>\n</ul>\n<p>🧠 <strong>Common keypoint detectors</strong>:</p>\n<ul>\n<li><strong>Hand-crafted</strong>: SIFT, ORB, AKAZE</li>\n<li><strong>Learned</strong>: SuperPoint, D2Net, R2D2</li>\n</ul>\n<p><strong>Local Descriptors</strong>:  <br>\nSmall feature vectors extracted around keypoints describing their local visual neighborhood.  <br>\nUsed for matching keypoints between different images.</p>\n<hr>\n<h4>⚡ Important Note</h4>\n<p>In our case:</p>\n<ul>\n<li><strong>Global features (DINOv2)</strong> are used for initial scene clustering via <strong>graph-based connected components</strong>.</li>\n<li><strong>Local features (ALIKED)</strong> and <strong>feature matching (LightGlue)</strong> are only used afterward for precise camera pose reconstruction (<strong>Structure-from-Motion</strong>, SfM).</li>\n</ul>\n<hr>\n<h4></h4><h3>📌 <b>2) Estimating Camera Poses (SfM)</b></h3>\n<hr>\n<h4>📐 What is Structure-from-Motion (SfM)?</h4>\n<p>\n<b>Structure-from-Motion (SfM)</b> is a computer vision pipeline that reconstructs:\n</p><ul>\n<li>📸 The pose of each camera (Rotation matrix <b>R</b> and Translation vector <b>T</b>).</li>\n<li>🧱 A sparse 3D point cloud of the scene (estimated 3D locations of keypoints).</li>\n</ul>\n\nGiven only a set of 2D images - without depth, without camera poses - SfM **recovers** both the camera trajectories and the scene structure.\n<p></p>\n<hr>\n<h4>📐 Simplified SfM Pipeline (General)</h4>\n<ol>\n<li><b>Pairwise Matching</b> - Find correspondences between keypoints detected in different images.</li>\n<li><b>Relative Pose Estimation</b> - Estimate the relative rotation (R) and translation (T) between pairs of images using 2D matches (via Essential Matrix decomposition).</li>\n<li><b>Triangulation</b> - Estimate 3D coordinates of points by triangulating matched keypoints from multiple views.</li>\n<li><b>Incremental SfM</b> - Add images one-by-one by solving PnP (Perspective-n-Point) problems.</li>\n<li><b>Bundle Adjustment</b> - Optimize all camera poses and 3D points jointly to minimize reprojection errors.</li>\n</ol>\n<hr>\n<h4>ALIKED, LightGlue, pycolmap</h4>\n<p>\n</p><ul>\n<li><b>Local feature detection</b>:<strong>ALIKED</strong> - a fast, learned local feature detector and descriptor extractor based on deep learning.</li>\n<li><b>Feature matching</b>: <strong>LightGlue</strong> - an attention-based neural matcher that finds reliable correspondences even under strong viewpoint and illumination changes.</li>\n<li><b>3D reconstruction and camera pose estimation</b>: <strong>pycolmap</strong> - a Python wrapper over <strong>COLMAP</strong>, a powerful Structure-from-Motion software, to reconstruct the camera trajectories and sparse 3D point cloud automatically.</li>\n</ul>\n\n<p></p>\n<hr>\n<p><strong>Triangulation</strong>    - Estimating 3D coordinates of a point from two or more images of it (2D → 3D)  <br>\n<strong>PnP (Perspective-n-Point)</strong> - Given 2D-3D correspondences, estimate the camera pose (R, T)  <br>\n<strong>Incremental SfM</strong> - Build the scene step by step: triangulate → PnP → repeat</p>\n<h4>🔧 How It Works in Practice</h4>\n<ol>\n<li><b>Extract ALIKED local features</b> for each image - keypoints + descriptors are saved into .h5 files.</li>\n<li><b>Match features using LightGlue</b> - for selected image pairs (shortlisted by global similarity or exhaustively).</li>\n<li><b>Import features and matches into a COLMAP database</b> using <code>h5_to_db</code> utilities.</li>\n<li><b>Run exhaustive matching and geometric verification</b> inside COLMAP (pycolmap API).</li>\n<li><b>Perform incremental reconstruction</b> (triangulation + pose estimation + bundle adjustment).</li>\n<li><b>Output:</b> For each image: \n<ul>\n<li>Rotation matrix <code>R</code></li>\n<li>Translation vector <code>T</code></li>\n</ul>\n</li>\n<li>If an image cannot be reconstructed (e.g., insufficient matches) - assign <code>NaN</code> to its pose fields.</li>\n</ol>\n<hr>\n<h4>📦 Tools Used</h4>\n<ul>\n<li><b>ALIKED</b> - lightweight keypoint extractor based on deep learning, efficient for fast and robust feature detection and description.</li>\n<li><b>LightGlue</b> - transformer-based feature matcher optimized for large viewpoint, scale, and illumination changes. Learns how to attend to correct matches.</li>\n<li><b>COLMAP (via pycolmap)</b> - one of the most powerful SfM engines, supporting incremental mapping, exhaustive matching, triangulation, PnP pose estimation, and bundle adjustment.</li>\n</ul>\n<hr>\n<h2></h2>\n<p>\nUsing COLMAP, pycolmap, or another SfM engine, we perform the following:\n</p><ul>\n<li>Input: All images assigned to the same cluster.</li>\n<li>Process:\n    <ul>\n    <li>Extract ALIKED features.</li>\n    <li>Match using LightGlue.</li>\n    <li>Import data into COLMAP database.</li>\n    <li>Run incremental reconstruction via pycolmap.</li>\n    </ul>\n</li>\n<li>Output: For each image, recover:\n    <ul>\n    <li>Rotation matrix <code>R</code> (3x3)</li>\n    <li>Translation vector <code>T</code> (3x1)</li>\n    </ul>\n</li>\n<li>If reconstruction fails for some images: assign NaN values to their pose fields.</li>\n</ul>\n<p></p>\n<hr>",
  "messages": [
    {
      "id": 3178394,
      "postDate": "2025-04-14T05:49:17.750Z",
      "content": "<p></p><h2>📂 Dataset Structure</h2><p></p>\n<p>Each dataset consists of multiple scenes. A scene is a group of related images that belong together (e.g., different viewpoints of the same object). Some images in a dataset may be <strong>outliers</strong>, meaning they don’t belong to any scene (e.g., random photos taken in the area).</p>\n<p>In the <strong>test data</strong>, all images are mixed in one folder. Your task is:</p>\n<ol>\n<li>Cluster the images into groups that represent scenes  </li>\n<li>Assign each image to a cluster or label it as an <code>outlier</code>  </li>\n<li>Reconstruct the camera pose (R, T) for each image (except outliers)</li>\n</ol>\n<hr>\n<p></p><h1>🧠 Camera Pose Estimation and mAA</h1><p></p>\n<ul>\n<li>A rotation matrix (3×3) - where the camera is pointing  </li>\n<li>A translation vector (3D) - where the camera is located  </li>\n</ul>\n<p><b>\"Ground truth\"</b> is the true known pose of the camera when the image was taken. In training data, this is given to you. In test data, you try to predict it.</p>\n<p>Camera center <b>C</b> in world coordinates is calculated as:  <br>\n<strong>C = −R · T</strong><br>\nThis transforms the camera pose into a 3D point in space.</p>\n<p>Camera centres of one scene:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F25747348%2F61b7f20998538db19346d6d34c961920%2F2025-04-14%20083034.jpg?generation=1744608643370046&amp;alt=media\" alt=\"\"></p>\n<hr>\n<h3><b>Structure from Motion (SfM)</b> is a pipeline that reconstructs:</h3>\n<ul>\n<li>Camera poses for each image  </li>\n<li>A sparse 3D point cloud of the scene</li>\n</ul>\n<p>Your predicted camera centers may be rotated differently, shifted (translated), and scaled.  <br>\nThis is normal in SfM - reconstructions can be geometrically correct but misaligned.</p>\n<p>So, to compare your predicted camera centers <b>C</b> to the ground truth centers <b>Cg</b>, a similarity transformation <b>T</b> (i.e. scale, rotation and translation altogether) is applied. This aligns your 3D camera predictions to the ground truth up to scale and orientation.  <br>\n(Some of your predicted camera poses may be wrong - far from ground truth - and we want to ignore bad data and only use good matches to compute the best alignment.)</p>\n<p>To compute a similarity transform (rotation, translation, and scale) between two sets of 3D points, you need at least 3 point correspondences that are not collinear.</p>\n<hr>\n<p>A <strong>triplet</strong> is a group of three matching camera centers. Each one has:</p>\n<ul>\n<li>A predicted 3D position <b>Ci</b>  </li>\n<li>A known ground-truth position <b>Cgi</b>  </li>\n</ul>\n<p>So a triplet is:  <br>\n<code>{(C1, Cg1), (C2, Cg2), (C3, Cg3)}</code></p>\n<p><b>RANSAC</b> = Random Sample Consensus - Algorithm to randomly test many triplets, ignore outliers, and find the best transformation.</p>\n<p>The <code>thresholds</code> column lists multiple numeric values per scene. These represent distance or error thresholds that are used to evaluate how accurate predicted 3D camera poses are (compared to ground truth). Each threshold is a maximum allowable error (in meters) to consider a pose prediction as correct.</p>\n<hr>\n<p>So use <b>Horn’s method</b> (algorithm used to compute the best similarity transformation between two sets of 3D points) to compute a similarity transformation.  <br>\nFind scale <b>s</b>, rotation <b>R</b>, translation <b>t</b> so that:</p>\n<pre><code> = s ⋅ R ⋅ Ci + t</code></pre>\n<p>This is our first estimate <b>T′</b>.  <br>\nThen apply <b>T′</b> to all predicted camera centers <b>Ci</b> and for each one count how many of your predictions become <b>registered cameras</b> (i.e., within error threshold of ground truth):</p>\n<pre><code> = (‖ T(Ci) − Cgi ‖ &lt; threshold)</code></pre>\n<p>Where <code>‖ · ‖</code> denotes Euclidean (L2) distance between two 3D points.</p>\n<p>Try many random triplets and get many candidate transformations <b>T′</b>.  <br>\nFor each one, count how many registered cameras it produces and save the best one.  <br>\nTake the best transformation <b>T′</b>, and re-estimate it using all the inliers (not just the original triplet).</p>\n<hr>\n<p></p><h2>📊 mAA = mean Average Accuracy</h2><p></p>\n<p>It measures how accurately your predicted camera centers match the ground truth, after alignment with a similarity transformation.</p>\n<p>As discussed before, you use triplets and Horn’s method inside RANSAC.  <br>\nYou find the best transformation:</p>\n<pre><code> = s ⋅ R ⋅ Ci + t</code></pre>\n<p>This aligns your predicted camera centers to the same coordinate frame as the ground-truth.  <br>\nThen apply this transformation - for each predicted camera center <b>Ci</b>, compute:  <br>\n<code>Ci_aligned = T(Ci)</code></p>\n<p>Now for each threshold <b>t_j</b>, check if the difference between <b>Ci_aligned</b> and ground-truth center <b>Cgi</b> is less than <b>t_j</b>:</p>\n<pre><code> = (‖ T(Ci) − Cgi ‖ &lt; threshold)</code></pre>\n<p>Then compute registration ratio per threshold for scene:  <br>\nFor each <b>N' = N − 3</b> images of the scene and each threshold <b>t_j</b>, compute the percentage of registered cameras:</p>\n<pre><code>rj = ( / ') * (registered)</code></pre>\n<p>That gives you the <b>Average Accuracy (AA)</b> for that scene.  <br>\nThen find average across all thresholds:</p>\n<pre><code>  ( / k) * sum(rj)</code></pre>\n<p>Where:</p>\n<ul>\n<li><b>k</b> = number of thresholds</li>\n</ul>\n<p>Once you compute <b>mAA per scene</b>, you calculate final mAA for all scenes by averaging across scenes.</p>\n<hr>\n<p></p><h2>🧠 Clustering and Final Evaluation Score</h2><p></p>\n<p>Each scene from the ground-truth data is denoted as <code>S&lt;sub&gt;ki&lt;/sub&gt;</code>, and your predicted clusters are <code>C&lt;sub&gt;kj&lt;/sub&gt;</code>.</p>\n<p>The evaluation will:</p>\n<ul>\n<li>Compare your clusters with the true scenes  </li>\n<li>Check how well your clustering matches the original grouping  </li>\n<li>Measure how accurate your pose predictions are (as explained earlier with mAA)</li>\n</ul>\n<p>Each ground-truth scene <code>S&lt;sub&gt;ki&lt;/sub&gt;</code> is matched to the user-submitted cluster <code>C&lt;sub&gt;kj&lt;/sub&gt;</code> that gives the <b>highest mAA</b> after alignment.  <br>\nAll images labeled as outliers are <b>excluded</b> from this matching step.</p>\n<hr>\n<h3>📊 Clustering Score (Precision)</h3>\n<p>After the best cluster <code>C&lt;sub&gt;kji&lt;/sub&gt;</code> is selected for each scene <code>S&lt;sub&gt;ki&lt;/sub&gt;</code>, compute the <b>clustering score</b> as:</p>\n<pre><code>clustering_score = ||||</code></pre>\n<p>Where:</p>\n<ul>\n<li><code>|S ∩ C|</code> is the number of images that correctly overlap between your predicted cluster and the real scene  </li>\n<li><code>|C|</code> is the total number of images in your predicted cluster</li>\n</ul>\n<p>This measures <b>how pure your cluster is</b> - i.e., out of everything you said belongs to this scene, how many actually do?</p>\n<hr>\n<h3>📈 Pose Score (mAA)</h3>\n<p>As explained earlier, for each image:</p>\n<ul>\n<li>Apply the best alignment transform T to your predicted camera centers  </li>\n<li>Check if each transformed pose is within the threshold of the ground-truth center  </li>\n<li>Count the number of registered cameras at each threshold  </li>\n<li>Compute registration ratios across thresholds  </li>\n<li>Average across thresholds → mAA per scene  </li>\n<li>Average over scenes → mAA for dataset</li>\n</ul>\n<p>This mAA represents <b>recall</b> - how many real scene images you found and aligned correctly.</p>\n<hr>\n<h3>🔗 Final Score Per Dataset</h3>\n<p>Now both mAA (recall-like) and clustering score (precision-like) are combined using the <b>harmonic mean</b>, similar to the F1-score in classification:</p>\n<pre><code> =  * (mAA * clustering_score) / (mAA + clustering_score)</code></pre>\n<p>This ensures that:</p>\n<ul>\n<li>If either clustering or pose accuracy is bad, your score is penalized  </li>\n<li>You must do <b>both clustering and reconstruction well</b> to score high</li>\n</ul>\n<hr>\n<h3>📊 Final Challenge Score</h3>\n<p>You will repeat this evaluation for each dataset <code>D&lt;sub&gt;k&lt;/sub&gt;</code>.  <br>\nThe final score for your submission is:</p>\n<pre><code></code></pre>\n<p>i.e., the average harmonic mean across all datasets.</p>\n<hr>\n<p></p><h1>🧩 Scene Clustering &amp; Pose Estimation (SfM)</h1><p></p>\n<hr>\n<h4></h4><h3>📌 <b>1) Scene Clustering</b></h3>\n<hr>\n<h5></h5><h4>📘 <b>1.1) Unsupervised Clustering on Global Descriptors</b></h4>\n<p>Extract global embeddings using models like <code>CLIP</code>, <code>DINOv2</code>, etc., then cluster them using algorithms like <b>DBSCAN</b>, <b>KMeans</b>, <b>Agglomerative Clustering</b>…</p>\n<p>📌 <strong>Image features</strong> are compact vector representations of the visual content of the image.</p>\n<hr>\n<h5></h5><h4>📦 Global Descriptors</h4>\n<p>Global descriptors are single vectors that represent the entire image - its objects, textures, and layout.</p>\n<p>✅ <strong>Use cases</strong>:</p>\n<ul>\n<li>Image clustering  </li>\n<li>Image retrieval  </li>\n<li>Scene classification  </li>\n</ul>\n<p><br></p>\n<p><strong>DINOv2</strong> is part of the <strong>DINO</strong> (Distillation with No Labels) family of self-supervised vision transformers:</p>\n<ul>\n<li>Learns by matching feature vectors of different augmented views of the same image.</li>\n<li>Ensures different images are mapped to distant feature vectors.</li>\n<li>Requires no manual labels (contrastive learning).</li>\n</ul>\n<hr>\n<h4>🔧 How Global Descriptor Clustering Works</h4>\n<ul>\n<li>Extract a single embedding per image (e.g., using DINOv2).</li>\n<li>Compute pairwise distances (typically cosine distance) between all embeddings.</li>\n<li>Cluster embeddings using:<ul>\n<li><strong>DBSCAN</strong> - density-based clustering without needing the number of clusters.</li>\n<li><strong>Agglomerative Clustering</strong> - hierarchical grouping.</li>\n<li><strong>KMeans</strong> - if the number of clusters is approximately known.</li>\n<li><strong>Graph-based Clustering</strong> - (see below 👇).</li></ul></li>\n<li>The result: each cluster represents one predicted scene.</li>\n</ul>\n<hr>\n<h4></h4><h4>📘 <b> Graph-Based Clustering</b></h4>\n<p>✅ In our pipeline:</p>\n<ul>\n<li>Compute <strong>pairwise cosine distances</strong> between image embeddings extracted by <strong>DINOv2</strong>.</li>\n<li>Build a <strong>graph</strong> where:<ul>\n<li>Each node represents an image.</li>\n<li>An edge is created between two images if their distance is below a threshold (e.g., <code>0.3</code>).</li></ul></li>\n<li><strong>Connected components</strong> of this graph are treated as <strong>predicted scenes</strong>.</li>\n<li>Images with no edges (isolated nodes) are labeled as <strong>outliers</strong>.</li>\n</ul>\n<p>\nThis approach is faster and more scalable than DBSCAN for large datasets, and does not require knowing the number of clusters beforehand.\n</p>\n<p>📌 <strong>Note</strong>:  <br>\nI do not use ALIKED or LightGlue at this stage.  <br>\nThey are used later in the pipeline for <strong>local feature extraction and matching</strong> - when reconstructing 3D camera poses.</p>\n<hr>\n<h5></h5><h4>🔍 Local Keypoints &amp; Descriptors (General Theory)</h4>\n<p>While  my solution clusters images using global descriptors, another common approach in image matching is to use <strong>local features</strong>:</p>\n<p><strong>Keypoints</strong> are distinctive image regions (corners, blobs, edges) that:</p>\n<ul>\n<li>Are stable under changes in viewpoint and lighting.</li>\n<li>Can be reliably detected and matched across different images.</li>\n</ul>\n<p>🧠 <strong>Common keypoint detectors</strong>:</p>\n<ul>\n<li><strong>Hand-crafted</strong>: SIFT, ORB, AKAZE</li>\n<li><strong>Learned</strong>: SuperPoint, D2Net, R2D2</li>\n</ul>\n<p><strong>Local Descriptors</strong>:  <br>\nSmall feature vectors extracted around keypoints describing their local visual neighborhood.  <br>\nUsed for matching keypoints between different images.</p>\n<hr>\n<h4>⚡ Important Note</h4>\n<p>In our case:</p>\n<ul>\n<li><strong>Global features (DINOv2)</strong> are used for initial scene clustering via <strong>graph-based connected components</strong>.</li>\n<li><strong>Local features (ALIKED)</strong> and <strong>feature matching (LightGlue)</strong> are only used afterward for precise camera pose reconstruction (<strong>Structure-from-Motion</strong>, SfM).</li>\n</ul>\n<hr>\n<h4></h4><h3>📌 <b>2) Estimating Camera Poses (SfM)</b></h3>\n<hr>\n<h4>📐 What is Structure-from-Motion (SfM)?</h4>\n<p>\n<b>Structure-from-Motion (SfM)</b> is a computer vision pipeline that reconstructs:\n</p><ul>\n<li>📸 The pose of each camera (Rotation matrix <b>R</b> and Translation vector <b>T</b>).</li>\n<li>🧱 A sparse 3D point cloud of the scene (estimated 3D locations of keypoints).</li>\n</ul>\n\nGiven only a set of 2D images - without depth, without camera poses - SfM **recovers** both the camera trajectories and the scene structure.\n<p></p>\n<hr>\n<h4>📐 Simplified SfM Pipeline (General)</h4>\n<ol>\n<li><b>Pairwise Matching</b> - Find correspondences between keypoints detected in different images.</li>\n<li><b>Relative Pose Estimation</b> - Estimate the relative rotation (R) and translation (T) between pairs of images using 2D matches (via Essential Matrix decomposition).</li>\n<li><b>Triangulation</b> - Estimate 3D coordinates of points by triangulating matched keypoints from multiple views.</li>\n<li><b>Incremental SfM</b> - Add images one-by-one by solving PnP (Perspective-n-Point) problems.</li>\n<li><b>Bundle Adjustment</b> - Optimize all camera poses and 3D points jointly to minimize reprojection errors.</li>\n</ol>\n<hr>\n<h4>ALIKED, LightGlue, pycolmap</h4>\n<p>\n</p><ul>\n<li><b>Local feature detection</b>:<strong>ALIKED</strong> - a fast, learned local feature detector and descriptor extractor based on deep learning.</li>\n<li><b>Feature matching</b>: <strong>LightGlue</strong> - an attention-based neural matcher that finds reliable correspondences even under strong viewpoint and illumination changes.</li>\n<li><b>3D reconstruction and camera pose estimation</b>: <strong>pycolmap</strong> - a Python wrapper over <strong>COLMAP</strong>, a powerful Structure-from-Motion software, to reconstruct the camera trajectories and sparse 3D point cloud automatically.</li>\n</ul>\n\n<p></p>\n<hr>\n<p><strong>Triangulation</strong>    - Estimating 3D coordinates of a point from two or more images of it (2D → 3D)  <br>\n<strong>PnP (Perspective-n-Point)</strong> - Given 2D-3D correspondences, estimate the camera pose (R, T)  <br>\n<strong>Incremental SfM</strong> - Build the scene step by step: triangulate → PnP → repeat</p>\n<h4>🔧 How It Works in Practice</h4>\n<ol>\n<li><b>Extract ALIKED local features</b> for each image - keypoints + descriptors are saved into .h5 files.</li>\n<li><b>Match features using LightGlue</b> - for selected image pairs (shortlisted by global similarity or exhaustively).</li>\n<li><b>Import features and matches into a COLMAP database</b> using <code>h5_to_db</code> utilities.</li>\n<li><b>Run exhaustive matching and geometric verification</b> inside COLMAP (pycolmap API).</li>\n<li><b>Perform incremental reconstruction</b> (triangulation + pose estimation + bundle adjustment).</li>\n<li><b>Output:</b> For each image: \n<ul>\n<li>Rotation matrix <code>R</code></li>\n<li>Translation vector <code>T</code></li>\n</ul>\n</li>\n<li>If an image cannot be reconstructed (e.g., insufficient matches) - assign <code>NaN</code> to its pose fields.</li>\n</ol>\n<hr>\n<h4>📦 Tools Used</h4>\n<ul>\n<li><b>ALIKED</b> - lightweight keypoint extractor based on deep learning, efficient for fast and robust feature detection and description.</li>\n<li><b>LightGlue</b> - transformer-based feature matcher optimized for large viewpoint, scale, and illumination changes. Learns how to attend to correct matches.</li>\n<li><b>COLMAP (via pycolmap)</b> - one of the most powerful SfM engines, supporting incremental mapping, exhaustive matching, triangulation, PnP pose estimation, and bundle adjustment.</li>\n</ul>\n<hr>\n<h2></h2>\n<p>\nUsing COLMAP, pycolmap, or another SfM engine, we perform the following:\n</p><ul>\n<li>Input: All images assigned to the same cluster.</li>\n<li>Process:\n    <ul>\n    <li>Extract ALIKED features.</li>\n    <li>Match using LightGlue.</li>\n    <li>Import data into COLMAP database.</li>\n    <li>Run incremental reconstruction via pycolmap.</li>\n    </ul>\n</li>\n<li>Output: For each image, recover:\n    <ul>\n    <li>Rotation matrix <code>R</code> (3x3)</li>\n    <li>Translation vector <code>T</code> (3x1)</li>\n    </ul>\n</li>\n<li>If reconstruction fails for some images: assign NaN values to their pose fields.</li>\n</ul>\n<p></p>\n<hr>",
      "rawMarkdown": "<center><h2 style=\"color:#1f77b4; \">📂 Dataset Structure</h2></center>\n\nEach dataset consists of multiple scenes. A scene is a group of related images that belong together (e.g., different viewpoints of the same object). Some images in a dataset may be <strong>outliers</strong>, meaning they don’t belong to any scene (e.g., random photos taken in the area).\n\nIn the <strong>test data</strong>, all images are mixed in one folder. Your task is:\n1. Cluster the images into groups that represent scenes  \n2. Assign each image to a cluster or label it as an <code>outlier</code>  \n3. Reconstruct the camera pose (R, T) for each image (except outliers)\n\n---\n\n<center><h1 style=\"color:#1f77b4;\">🧠 Camera Pose Estimation and mAA</h1></center>\n\n- A rotation matrix (3×3) - where the camera is pointing  \n- A translation vector (3D) - where the camera is located  \n\n<b>\"Ground truth\"</b> is the true known pose of the camera when the image was taken. In training data, this is given to you. In test data, you try to predict it.\n\nCamera center <b>C</b> in world coordinates is calculated as:  \n**C = −R<sup>T</sup> · T**\nThis transforms the camera pose into a 3D point in space.\n\nCamera centres of one scene:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F25747348%2F61b7f20998538db19346d6d34c961920%2F2025-04-14%20083034.jpg?generation=1744608643370046&alt=media)\n\n---\n\n### <b>Structure from Motion (SfM)</b> is a pipeline that reconstructs:\n- Camera poses for each image  \n- A sparse 3D point cloud of the scene\n\nYour predicted camera centers may be rotated differently, shifted (translated), and scaled.  \nThis is normal in SfM - reconstructions can be geometrically correct but misaligned.\n\nSo, to compare your predicted camera centers <b>C</b> to the ground truth centers <b>Cg</b>, a similarity transformation <b>T</b> (i.e. scale, rotation and translation altogether) is applied. This aligns your 3D camera predictions to the ground truth up to scale and orientation.  \n(Some of your predicted camera poses may be wrong - far from ground truth - and we want to ignore bad data and only use good matches to compute the best alignment.)\n\nTo compute a similarity transform (rotation, translation, and scale) between two sets of 3D points, you need at least 3 point correspondences that are not collinear.\n\n---\n\nA <strong>triplet</strong> is a group of three matching camera centers. Each one has:\n- A predicted 3D position <b>Ci</b>  \n- A known ground-truth position <b>Cgi</b>  \n\nSo a triplet is:  \n<code>{(C1, Cg1), (C2, Cg2), (C3, Cg3)}</code>\n\n<b>RANSAC</b> = Random Sample Consensus - Algorithm to randomly test many triplets, ignore outliers, and find the best transformation.\n\nThe <code>thresholds</code> column lists multiple numeric values per scene. These represent distance or error thresholds that are used to evaluate how accurate predicted 3D camera poses are (compared to ground truth). Each threshold is a maximum allowable error (in meters) to consider a pose prediction as correct.\n\n---\n\nSo use <b>Horn’s method</b> (algorithm used to compute the best similarity transformation between two sets of 3D points) to compute a similarity transformation.  \nFind scale <b>s</b>, rotation <b>R</b>, translation <b>t</b> so that:\n\n<pre><code>T(Ci) = s ⋅ R ⋅ Ci + t</code></pre>\n\nThis is our first estimate <b>T′</b>.  \nThen apply <b>T′</b> to all predicted camera centers <b>Ci</b> and for each one count how many of your predictions become <b>registered cameras</b> (i.e., within error threshold of ground truth):\n\n<pre><code>is_registered = (‖ T(Ci) − Cgi ‖ < threshold)</code></pre>\n\nWhere <code>‖ · ‖</code> denotes Euclidean (L2) distance between two 3D points.\n\nTry many random triplets and get many candidate transformations <b>T′</b>.  \nFor each one, count how many registered cameras it produces and save the best one.  \nTake the best transformation <b>T′</b>, and re-estimate it using all the inliers (not just the original triplet).\n\n---\n\n<center><h2 style=\"color:#2ca02c;\">📊 mAA = mean Average Accuracy</h2></center>\n\nIt measures how accurately your predicted camera centers match the ground truth, after alignment with a similarity transformation.\n\nAs discussed before, you use triplets and Horn’s method inside RANSAC.  \nYou find the best transformation:\n\n<pre><code>T(Ci) = s ⋅ R ⋅ Ci + t</code></pre>\n\nThis aligns your predicted camera centers to the same coordinate frame as the ground-truth.  \nThen apply this transformation - for each predicted camera center <b>Ci</b>, compute:  \n<code>Ci_aligned = T(Ci)</code>\n\nNow for each threshold <b>t_j</b>, check if the difference between <b>Ci_aligned</b> and ground-truth center <b>Cgi</b> is less than <b>t_j</b>:\n\n<pre><code>is_registered = (‖ T(Ci) − Cgi ‖ < threshold)</code></pre>\n\nThen compute registration ratio per threshold for scene:  \nFor each <b>N' = N − 3</b> images of the scene and each threshold <b>t_j</b>, compute the percentage of registered cameras:\n\n<pre><code>rj = (1 / N') * sum(registered)</code></pre>\n\nThat gives you the <b>Average Accuracy (AA)</b> for that scene.  \nThen find average across all thresholds:\n\n<pre><code>mAA = (1 / k) * sum(rj)</code></pre>\n\nWhere:\n- <b>k</b> = number of thresholds\n\nOnce you compute <b>mAA per scene</b>, you calculate final mAA for all scenes by averaging across scenes.\n\n---\n\n<center><h2 style=\"color:#1f77b4;\">🧠 Clustering and Final Evaluation Score</h2></center>\n\nEach scene from the ground-truth data is denoted as <code>S<sub>ki</sub></code>, and your predicted clusters are <code>C<sub>kj</sub></code>.\n\nThe evaluation will:\n- Compare your clusters with the true scenes  \n- Check how well your clustering matches the original grouping  \n- Measure how accurate your pose predictions are (as explained earlier with mAA)\n\nEach ground-truth scene <code>S<sub>ki</sub></code> is matched to the user-submitted cluster <code>C<sub>kj</sub></code> that gives the <b>highest mAA</b> after alignment.  \nAll images labeled as outliers are <b>excluded</b> from this matching step.\n\n---\n\n<h3 style=\"color:#2ca02c;\">📊 Clustering Score (Precision)</h3>\n\nAfter the best cluster <code>C<sub>kji</sub></code> is selected for each scene <code>S<sub>ki</sub></code>, compute the <b>clustering score</b> as:\n\n<pre><code>clustering_score = |S ∩ C| / |C|</code></pre>\n\nWhere:\n- <code>|S ∩ C|</code> is the number of images that correctly overlap between your predicted cluster and the real scene  \n- <code>|C|</code> is the total number of images in your predicted cluster\n\nThis measures <b>how pure your cluster is</b> - i.e., out of everything you said belongs to this scene, how many actually do?\n\n---\n\n<h3 style=\"color:#2ca02c;\">📈 Pose Score (mAA)</h3>\n\nAs explained earlier, for each image:\n- Apply the best alignment transform T to your predicted camera centers  \n- Check if each transformed pose is within the threshold of the ground-truth center  \n- Count the number of registered cameras at each threshold  \n- Compute registration ratios across thresholds  \n- Average across thresholds → mAA per scene  \n- Average over scenes → mAA for dataset\n\nThis mAA represents <b>recall</b> - how many real scene images you found and aligned correctly.\n\n---\n\n<h3 style=\"color:#2ca02c;\">🔗 Final Score Per Dataset</h3>\n\nNow both mAA (recall-like) and clustering score (precision-like) are combined using the <b>harmonic mean</b>, similar to the F1-score in classification:\n\n<pre><code>final_score_per_dataset = 2 * (mAA * clustering_score) / (mAA + clustering_score)</code></pre>\n\nThis ensures that:\n- If either clustering or pose accuracy is bad, your score is penalized  \n- You must do <b>both clustering and reconstruction well</b> to score high\n\n---\n\n<h3 style=\"color:#ff7f0e;\">📊 Final Challenge Score</h3>\n\nYou will repeat this evaluation for each dataset <code>D<sub>k</sub></code>.  \nThe final score for your submission is:\n\n<pre><code>mean(final_score_per_dataset)</code></pre>\n\ni.e., the average harmonic mean across all datasets.\n\n---\n<center><h1 style=\"color:#1f77b4;\">🧩 Scene Clustering & Pose Estimation (SfM)</h1></center>\n\n---\n\n#### <h3 style=\"color:#2ca02c;\">📌 <b>1) Scene Clustering</b></h3>\n\n---\n\n##### <h4 style=\"color:#2ca02c;\">📘 <b>1.1) Unsupervised Clustering on Global Descriptors</b></h4>\n\nExtract global embeddings using models like <code>CLIP</code>, <code>DINOv2</code>, etc., then cluster them using algorithms like <b>DBSCAN</b>, <b>KMeans</b>, <b>Agglomerative Clustering</b>...\n\n📌 **Image features** are compact vector representations of the visual content of the image.\n\n---\n\n##### <h4 style=\"color:#1f77b4;\">📦 Global Descriptors</h4>\n\nGlobal descriptors are single vectors that represent the entire image - its objects, textures, and layout.\n\n✅ **Use cases**:\n- Image clustering  \n- Image retrieval  \n- Scene classification  \n\n<br>\n\n**DINOv2** is part of the **DINO** (Distillation with No Labels) family of self-supervised vision transformers:\n- Learns by matching feature vectors of different augmented views of the same image.\n- Ensures different images are mapped to distant feature vectors.\n- Requires no manual labels (contrastive learning).\n\n---\n\n<h4 style=\"color:#1f77b4;\">🔧 How Global Descriptor Clustering Works</h4>\n\n- Extract a single embedding per image (e.g., using DINOv2).\n- Compute pairwise distances (typically cosine distance) between all embeddings.\n- Cluster embeddings using:\n  - **DBSCAN** - density-based clustering without needing the number of clusters.\n  - **Agglomerative Clustering** - hierarchical grouping.\n  - **KMeans** - if the number of clusters is approximately known.\n  - **Graph-based Clustering** - (see below 👇).\n- The result: each cluster represents one predicted scene.\n\n---\n\n#### <h4 style=\"color:#2ca02c;\">📘 <b> Graph-Based Clustering</b></h4>\n\n✅ In our pipeline:\n- Compute **pairwise cosine distances** between image embeddings extracted by **DINOv2**.\n- Build a **graph** where:\n  - Each node represents an image.\n  - An edge is created between two images if their distance is below a threshold (e.g., <code>0.3</code>).\n- **Connected components** of this graph are treated as **predicted scenes**.\n- Images with no edges (isolated nodes) are labeled as **outliers**.\n\n<p style=\"font-size:16px;\">\nThis approach is faster and more scalable than DBSCAN for large datasets, and does not require knowing the number of clusters beforehand.\n</p>\n\n📌 **Note**:  \nI do not use ALIKED or LightGlue at this stage.  \nThey are used later in the pipeline for **local feature extraction and matching** - when reconstructing 3D camera poses.\n\n---\n\n##### <h4 style=\"color:#1f77b4;\">🔍 Local Keypoints & Descriptors (General Theory)</h4>\n\nWhile  my solution clusters images using global descriptors, another common approach in image matching is to use **local features**:\n\n**Keypoints** are distinctive image regions (corners, blobs, edges) that:\n- Are stable under changes in viewpoint and lighting.\n- Can be reliably detected and matched across different images.\n\n🧠 **Common keypoint detectors**:\n- **Hand-crafted**: SIFT, ORB, AKAZE\n- **Learned**: SuperPoint, D2Net, R2D2\n\n**Local Descriptors**:  \nSmall feature vectors extracted around keypoints describing their local visual neighborhood.  \nUsed for matching keypoints between different images.\n\n---\n\n<h4 style=\"color:rgb(255, 11, 214);\">⚡ Important Note</h4>\n\nIn our case:\n- **Global features (DINOv2)** are used for initial scene clustering via **graph-based connected components**.\n- **Local features (ALIKED)** and **feature matching (LightGlue)** are only used afterward for precise camera pose reconstruction (**Structure-from-Motion**, SfM).\n\n\n---\n#### <h3 style=\"color:#2ca02c;\">📌 <b>2) Estimating Camera Poses (SfM)</b></h3>\n\n---\n\n<h4 style=\"color:#1f77b4;\">📐 What is Structure-from-Motion (SfM)?</h4>\n\n<p style=\"font-size:16px;\">\n<b>Structure-from-Motion (SfM)</b> is a computer vision pipeline that reconstructs:\n<ul>\n<li>📸 The pose of each camera (Rotation matrix <b>R</b> and Translation vector <b>T</b>).</li>\n<li>🧱 A sparse 3D point cloud of the scene (estimated 3D locations of keypoints).</li>\n</ul>\n\nGiven only a set of 2D images - without depth, without camera poses - SfM **recovers** both the camera trajectories and the scene structure.\n</p>\n\n---\n\n<h4 style=\"color:#1f77b4;\">📐 Simplified SfM Pipeline (General)</h4>\n\n<ol style=\"font-size:16px;\">\n<li><b>Pairwise Matching</b> - Find correspondences between keypoints detected in different images.</li>\n<li><b>Relative Pose Estimation</b> - Estimate the relative rotation (R) and translation (T) between pairs of images using 2D matches (via Essential Matrix decomposition).</li>\n<li><b>Triangulation</b> - Estimate 3D coordinates of points by triangulating matched keypoints from multiple views.</li>\n<li><b>Incremental SfM</b> - Add images one-by-one by solving PnP (Perspective-n-Point) problems.</li>\n<li><b>Bundle Adjustment</b> - Optimize all camera poses and 3D points jointly to minimize reprojection errors.</li>\n</ol>\n\n---\n\n<h4 style=\"color:#2ca02c;\">ALIKED, LightGlue, pycolmap</h4>\n\n<p style=\"font-size:16px;\">\n<ul>\n<li><b>Local feature detection</b>:<strong>ALIKED</strong> - a fast, learned local feature detector and descriptor extractor based on deep learning.</li>\n<li><b>Feature matching</b>: <strong>LightGlue</strong> - an attention-based neural matcher that finds reliable correspondences even under strong viewpoint and illumination changes.</li>\n<li><b>3D reconstruction and camera pose estimation</b>: <strong>pycolmap</strong> - a Python wrapper over <strong>COLMAP</strong>, a powerful Structure-from-Motion software, to reconstruct the camera trajectories and sparse 3D point cloud automatically.</li>\n</ul>\n\n</p>\n\n---\n**Triangulation**\t- Estimating 3D coordinates of a point from two or more images of it (2D → 3D)  \n**PnP (Perspective-n-Point)** - Given 2D-3D correspondences, estimate the camera pose (R, T)  \n**Incremental SfM** - Build the scene step by step: triangulate → PnP → repeat\n\n<h4 style=\"color:#1f77b4;\">🔧 How It Works in Practice</h4>\n\n<ol style=\"font-size:16px;\">\n<li><b>Extract ALIKED local features</b> for each image - keypoints + descriptors are saved into .h5 files.</li>\n<li><b>Match features using LightGlue</b> - for selected image pairs (shortlisted by global similarity or exhaustively).</li>\n<li><b>Import features and matches into a COLMAP database</b> using <code>h5_to_db</code> utilities.</li>\n<li><b>Run exhaustive matching and geometric verification</b> inside COLMAP (pycolmap API).</li>\n<li><b>Perform incremental reconstruction</b> (triangulation + pose estimation + bundle adjustment).</li>\n<li><b>Output:</b> For each image: \n<ul>\n<li>Rotation matrix <code>R</code></li>\n<li>Translation vector <code>T</code></li>\n</ul>\n</li>\n<li>If an image cannot be reconstructed (e.g., insufficient matches) - assign <code>NaN</code> to its pose fields.</li>\n</ol>\n\n---\n\n<h4 style=\"color:#1f77b4;\">📦 Tools Used</h4>\n\n<ul style=\"font-size:16px;\">\n<li><b>ALIKED</b> - lightweight keypoint extractor based on deep learning, efficient for fast and robust feature detection and description.</li>\n<li><b>LightGlue</b> - transformer-based feature matcher optimized for large viewpoint, scale, and illumination changes. Learns how to attend to correct matches.</li>\n<li><b>COLMAP (via pycolmap)</b> - one of the most powerful SfM engines, supporting incremental mapping, exhaustive matching, triangulation, PnP pose estimation, and bundle adjustment.</li>\n</ul>\n\n---\n\n<h2 style=\"color:rgb(255, 11, 214);\"><span style=\"font-size: 30px;\">📸 Reconstruct Camera Poses (SfM)</span></h2>\n\n<p style=\"font-size:16px;\">\nUsing COLMAP, pycolmap, or another SfM engine, we perform the following:\n<ul>\n<li>Input: All images assigned to the same cluster.</li>\n<li>Process:\n    <ul>\n    <li>Extract ALIKED features.</li>\n    <li>Match using LightGlue.</li>\n    <li>Import data into COLMAP database.</li>\n    <li>Run incremental reconstruction via pycolmap.</li>\n    </ul>\n</li>\n<li>Output: For each image, recover:\n    <ul>\n    <li>Rotation matrix <code>R</code> (3x3)</li>\n    <li>Translation vector <code>T</code> (3x1)</li>\n    </ul>\n</li>\n<li>If reconstruction fails for some images: assign NaN values to their pose fields.</li>\n</ul>\n</p>\n\n---",
      "votes": 57
    },
    {
      "id": 3204068,
      "postDate": "2025-05-17T18:17:30.153Z",
      "content": "<p>Great stuff! And all in very details, just what I needed to get started! Thanks <a href=\"https://www.kaggle.com/dmytrobuhai\" target=\"_blank\">@dmytrobuhai</a> !</p>",
      "rawMarkdown": "Great stuff! And all in very details, just what I needed to get started! Thanks @dmytrobuhai !",
      "votes": 1
    },
    {
      "id": 3203991,
      "postDate": "2025-05-17T16:35:55.340Z",
      "content": "<p>Thank you a lot for the extremely useful and concise explanations! It is my first day during which I am deeply diving into the competition details and this topic definitely helped me to quickly grasp main theoretical concepts. 🔥</p>",
      "rawMarkdown": "Thank you a lot for the extremely useful and concise explanations! It is my first day during which I am deeply diving into the competition details and this topic definitely helped me to quickly grasp main theoretical concepts. 🔥",
      "votes": 1
    },
    {
      "id": 3186803,
      "postDate": "2025-04-25T07:32:38.420Z",
      "content": "<p>I have some confusion about the calculation formula for the Camera center C in world coordinates. Could you please explain how it is derived? I would be very grateful if you could reply.</p>",
      "rawMarkdown": "I have some confusion about the calculation formula for the Camera center C in world coordinates. Could you please explain how it is derived? I would be very grateful if you could reply.",
      "votes": 1,
      "replies": [
        {
          "id": 3187230,
          "postDate": "2025-04-25T17:32:22.147Z",
          "content": "<p>hello, this formula: </p>\n<pre>X_cam = R ⋅ X_world + T\n</pre>\n<p>transforms a point from the world coordinate system (global 3D space) into the camera coordinate system.</p>\n<blockquote>\n  <p>Think of it this way: the world is fixed, you have a camera somewhere in that world, looking in a specific direction. You want to understand how a 3D point in the world appears from the perspective of the camera.</p>\n</blockquote>\n<p>That means take the point in world space, rotate it into the camera's view direction, and then translate it so that the camera becomes the new origin</p>\n<hr>\n<p>If the camera is not aligned with the world axes, then we need to rotate the world to the camera’s orientation. Each row (or column, depending on convention) of <strong>R</strong> is a unit vector representing one axis of the rotated frame.</p>\n<p>Firstly, let's clarify that <strong>vectors are orthonormal</strong> if they satisfy two conditions:</p>\n<ol>\n<li>Vectors are <strong>perpendicular</strong> to each other  </li>\n<li>Each vector has <strong>unit length</strong></li>\n</ol>\n<p>A <strong>rotation matrix <code>R</code></strong> is an <strong>orthonormal linear transformation</strong> that:</p>\n<p>a. <strong>Preserves lengths</strong>:  <br>\n   </p>\n<pre>||R x|| = ||x||</pre>\n<p>b. <strong>Preserves angles between vectors</strong>:  <br>\n   </p>\n<pre>⟨R x, R y⟩ = ⟨x, y⟩</pre>\n<p>c. <strong>Is orthogonal</strong>:  <br>\n   </p>\n<pre>R^T = R⁻¹</pre>\n<p>For any orthogonal matrix:  <br>\n   </p>\n<pre>R^T ⋅ R = I</pre>\n<p><br>\n   So multiplying by the inverse matrix on the right gives:  <br>\n   </p>\n<pre>R^T = R⁻¹</pre>\n<p>The transpose undoes the transformation - it rotates back.</p>\n<p>d. <strong>Has determinant</strong>:  <br>\n   </p>\n<pre>det(R) = +1</pre>\n<p><br>\n   (i.e., it's a <strong>proper rotation</strong>, not a reflection)</p>\n<hr>\n<p>Let <code>X_world</code> be a 3D point in the world frame:  </p>\n<pre>X_world = [x, y, z]^T</pre>\n<p>And let:  </p>\n<pre>R = [r1 r2 r3]</pre>\n<p><br>\nwhere <code>r1</code>, <code>r2</code>, <code>r3</code> are the <strong>unit vectors</strong> of the new basis (camera axes).</p>\n<p>Then the rotated point becomes:</p>\n<pre>X_cam = R ⋅ X_world = x ⋅ r1 + y ⋅ r2 + z ⋅ r3\n</pre>\n<p>This expresses the vector <code>X_world</code> in the new frame defined by <code>R</code>.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F25747348%2F9adb1e1db134ce9caec75bda1361cd27%2F1_rkaT6sY7OHDworLvSiWv4w.jpg?generation=1745602417430547&amp;alt=media\" alt=\"\"></p>\n<h2><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F25747348%2F4c52edbc3d78990c5f8e1d2481f073d2%2FRotation-and-translation-parameters-of-the-camera.png?generation=1745602430125479&amp;alt=media\" alt=\"\"></h2>\n<p>After rotation, we <strong>shift the entire coordinate frame to the camera's origin</strong>.</p>\n<p><code>T</code> is the <strong>position of the world origin in the camera's coordinate system</strong>. </p>\n<p>So:</p>\n<pre>+T translates the rotated point to the camera’s local coordinate frame\n</pre>\n<p>This <strong>moves the origin</strong> so that the camera becomes the new reference point.</p>\n<hr>\n<h3>Inverse Transformation</h3>\n<p>To go back (camera → world), you need to <strong>undo the transformation</strong>:</p>\n<p>Start with:</p>\n<pre>X_cam = R ⋅ X_world + T\n</pre>\n<p>Subtract <code>T</code> and multiply by <code>R⁻¹</code> on the left:</p>\n<pre>R⁻¹ ⋅ (X_cam - T)= R⁻¹ ⋅ R ⋅ X_word \n</pre>\n<p>We discuss early that  R^T = R⁻¹ , so</p>\n<pre>R^T ⋅ (X_cam - T)= I ⋅ X_word (where I is unit matrix)\n</pre>\n<p>Than:</p>\n<pre>X_world = R^T ⋅ (X_cam − T)\n</pre>\n<p>Now apply this <strong>inverse transformation</strong> to the <strong>camera's own center</strong> in its local frame:</p>\n<pre>C_cam = [0, 0, 0]^T   (camera origin in its own frame)\nC_world = R^T ⋅ (C_cam − T) = R^T ⋅ (0 − T) = − R^T ⋅ T\n</pre>\n<blockquote>\n  <p>Note: <code>T</code> in superscript is the <strong>transpose</strong>, not exponentiation.</p>\n</blockquote>\n<p>This formula gives the <strong>camera center <code>C</code> in world coordinates</strong> - exactly what we use during evaluation.</p>",
          "rawMarkdown": "hello, this formula: \n\n<pre>\nX_cam = R ⋅ X_world + T\n</pre>\n\ntransforms a point from the world coordinate system (global 3D space) into the camera coordinate system.\n\n> Think of it this way: the world is fixed, you have a camera somewhere in that world, looking in a specific direction. You want to understand how a 3D point in the world appears from the perspective of the camera.\n\nThat means take the point in world space, rotate it into the camera's view direction, and then translate it so that the camera becomes the new origin\n\n---\n\nIf the camera is not aligned with the world axes, then we need to rotate the world to the camera’s orientation. Each row (or column, depending on convention) of **R** is a unit vector representing one axis of the rotated frame.\n\nFirstly, let's clarify that **vectors are orthonormal** if they satisfy two conditions:\n\n1. Vectors are **perpendicular** to each other  \n2. Each vector has **unit length**\n\nA **rotation matrix `R`** is an **orthonormal linear transformation** that:\n\na. **Preserves lengths**:  \n   <pre>||R x|| = ||x||</pre>\n\nb. **Preserves angles between vectors**:  \n   <pre>⟨R x, R y⟩ = ⟨x, y⟩</pre>\n\nc. **Is orthogonal**:  \n   <pre>R^T = R⁻¹</pre>  \n\n   For any orthogonal matrix:  \n   <pre>R^T ⋅ R = I</pre>  \n   So multiplying by the inverse matrix on the right gives:  \n   <pre>R^T = R⁻¹</pre>\n  The transpose undoes the transformation - it rotates back.\n\nd. **Has determinant**:  \n   <pre>det(R) = +1</pre>  \n   (i.e., it's a **proper rotation**, not a reflection)\n\n---\n\nLet `X_world` be a 3D point in the world frame:  \n<pre>X_world = [x, y, z]^T</pre>\n\nAnd let:  \n<pre>R = [r1 r2 r3]</pre>  \nwhere `r1`, `r2`, `r3` are the **unit vectors** of the new basis (camera axes).\n\nThen the rotated point becomes:\n\n<pre>\nX_cam = R ⋅ X_world = x ⋅ r1 + y ⋅ r2 + z ⋅ r3\n</pre>\n\nThis expresses the vector `X_world` in the new frame defined by `R`.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F25747348%2F9adb1e1db134ce9caec75bda1361cd27%2F1_rkaT6sY7OHDworLvSiWv4w.jpg?generation=1745602417430547&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F25747348%2F4c52edbc3d78990c5f8e1d2481f073d2%2FRotation-and-translation-parameters-of-the-camera.png?generation=1745602430125479&alt=media)\n---\n\nAfter rotation, we **shift the entire coordinate frame to the camera's origin**.\n\n`T` is the **position of the world origin in the camera's coordinate system**. \n\nSo:\n\n<pre>\n+T translates the rotated point to the camera’s local coordinate frame\n</pre>\n\nThis **moves the origin** so that the camera becomes the new reference point.\n\n---\n\n### Inverse Transformation\n\nTo go back (camera → world), you need to **undo the transformation**:\n\nStart with:\n\n<pre>\nX_cam = R ⋅ X_world + T\n</pre>\n\nSubtract `T` and multiply by `R⁻¹` on the left:\n\n<pre>\nR⁻¹ ⋅ (X_cam - T)= R⁻¹ ⋅ R ⋅ X_word \n</pre>\nWe discuss early that  R^T = R⁻¹ , so\n<pre>\nR^T ⋅ (X_cam - T)= I ⋅ X_word (where I is unit matrix)\n</pre>\nThan:\n<pre>\nX_world = R^T ⋅ (X_cam − T)\n</pre>\n\nNow apply this **inverse transformation** to the **camera's own center** in its local frame:\n\n<pre>\nC_cam = [0, 0, 0]^T   (camera origin in its own frame)\nC_world = R^T ⋅ (C_cam − T) = R^T ⋅ (0 − T) = − R^T ⋅ T\n</pre>\n\n> Note: `T` in superscript is the **transpose**, not exponentiation.\n\nThis formula gives the **camera center `C` in world coordinates** - exactly what we use during evaluation.\n\n",
          "votes": 10,
          "replies": [
            {
              "id": 3187537,
              "postDate": "2025-04-26T07:01:57.993Z",
              "content": "<p>Thank you so much for your detailed answer! It really helped me understand everything. I truly appreciate your kindness and patience!</p>",
              "rawMarkdown": "Thank you so much for your detailed answer! It really helped me understand everything. I truly appreciate your kindness and patience!",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 3178395,
      "postDate": "2025-04-14T05:50:12.840Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3204068,
      "author_name": "Diganta",
      "author_url": "",
      "post_date": "2025-05-17T18:17:30.153000",
      "content": "<p>Great stuff! And all in very details, just what I needed to get started! Thanks <a href=\"https://www.kaggle.com/dmytrobuhai\" target=\"_blank\">@dmytrobuhai</a> !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3203991,
      "author_name": "Vyacheslav Efimov",
      "author_url": "",
      "post_date": "2025-05-17T16:35:55.340000",
      "content": "<p>Thank you a lot for the extremely useful and concise explanations! It is my first day during which I am deeply diving into the competition details and this topic definitely helped me to quickly grasp main theoretical concepts. 🔥</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3186803,
      "author_name": "stardustlove",
      "author_url": "",
      "post_date": "2025-04-25T07:32:38.420000",
      "content": "<p>I have some confusion about the calculation formula for the Camera center C in world coordinates. Could you please explain how it is derived? I would be very grateful if you could reply.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3187230,
          "author_name": "Dmytro Buhai",
          "author_url": "",
          "post_date": "2025-04-25T17:32:22.147000",
          "content": "<p>hello, this formula: </p>\n<pre>X_cam = R ⋅ X_world + T\n</pre>\n<p>transforms a point from the world coordinate system (global 3D space) into the camera coordinate system.</p>\n<blockquote>\n  <p>Think of it this way: the world is fixed, you have a camera somewhere in that world, looking in a specific direction. You want to understand how a 3D point in the world appears from the perspective of the camera.</p>\n</blockquote>\n<p>That means take the point in world space, rotate it into the camera's view direction, and then translate it so that the camera becomes the new origin</p>\n<hr>\n<p>If the camera is not aligned with the world axes, then we need to rotate the world to the camera’s orientation. Each row (or column, depending on convention) of <strong>R</strong> is a unit vector representing one axis of the rotated frame.</p>\n<p>Firstly, let's clarify that <strong>vectors are orthonormal</strong> if they satisfy two conditions:</p>\n<ol>\n<li>Vectors are <strong>perpendicular</strong> to each other  </li>\n<li>Each vector has <strong>unit length</strong></li>\n</ol>\n<p>A <strong>rotation matrix <code>R</code></strong> is an <strong>orthonormal linear transformation</strong> that:</p>\n<p>a. <strong>Preserves lengths</strong>:  <br>\n   </p>\n<pre>||R x|| = ||x||</pre>\n<p>b. <strong>Preserves angles between vectors</strong>:  <br>\n   </p>\n<pre>⟨R x, R y⟩ = ⟨x, y⟩</pre>\n<p>c. <strong>Is orthogonal</strong>:  <br>\n   </p>\n<pre>R^T = R⁻¹</pre>\n<p>For any orthogonal matrix:  <br>\n   </p>\n<pre>R^T ⋅ R = I</pre>\n<p><br>\n   So multiplying by the inverse matrix on the right gives:  <br>\n   </p>\n<pre>R^T = R⁻¹</pre>\n<p>The transpose undoes the transformation - it rotates back.</p>\n<p>d. <strong>Has determinant</strong>:  <br>\n   </p>\n<pre>det(R) = +1</pre>\n<p><br>\n   (i.e., it's a <strong>proper rotation</strong>, not a reflection)</p>\n<hr>\n<p>Let <code>X_world</code> be a 3D point in the world frame:  </p>\n<pre>X_world = [x, y, z]^T</pre>\n<p>And let:  </p>\n<pre>R = [r1 r2 r3]</pre>\n<p><br>\nwhere <code>r1</code>, <code>r2</code>, <code>r3</code> are the <strong>unit vectors</strong> of the new basis (camera axes).</p>\n<p>Then the rotated point becomes:</p>\n<pre>X_cam = R ⋅ X_world = x ⋅ r1 + y ⋅ r2 + z ⋅ r3\n</pre>\n<p>This expresses the vector <code>X_world</code> in the new frame defined by <code>R</code>.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F25747348%2F9adb1e1db134ce9caec75bda1361cd27%2F1_rkaT6sY7OHDworLvSiWv4w.jpg?generation=1745602417430547&amp;alt=media\" alt=\"\"></p>\n<h2><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F25747348%2F4c52edbc3d78990c5f8e1d2481f073d2%2FRotation-and-translation-parameters-of-the-camera.png?generation=1745602430125479&amp;alt=media\" alt=\"\"></h2>\n<p>After rotation, we <strong>shift the entire coordinate frame to the camera's origin</strong>.</p>\n<p><code>T</code> is the <strong>position of the world origin in the camera's coordinate system</strong>. </p>\n<p>So:</p>\n<pre>+T translates the rotated point to the camera’s local coordinate frame\n</pre>\n<p>This <strong>moves the origin</strong> so that the camera becomes the new reference point.</p>\n<hr>\n<h3>Inverse Transformation</h3>\n<p>To go back (camera → world), you need to <strong>undo the transformation</strong>:</p>\n<p>Start with:</p>\n<pre>X_cam = R ⋅ X_world + T\n</pre>\n<p>Subtract <code>T</code> and multiply by <code>R⁻¹</code> on the left:</p>\n<pre>R⁻¹ ⋅ (X_cam - T)= R⁻¹ ⋅ R ⋅ X_word \n</pre>\n<p>We discuss early that  R^T = R⁻¹ , so</p>\n<pre>R^T ⋅ (X_cam - T)= I ⋅ X_word (where I is unit matrix)\n</pre>\n<p>Than:</p>\n<pre>X_world = R^T ⋅ (X_cam − T)\n</pre>\n<p>Now apply this <strong>inverse transformation</strong> to the <strong>camera's own center</strong> in its local frame:</p>\n<pre>C_cam = [0, 0, 0]^T   (camera origin in its own frame)\nC_world = R^T ⋅ (C_cam − T) = R^T ⋅ (0 − T) = − R^T ⋅ T\n</pre>\n<blockquote>\n  <p>Note: <code>T</code> in superscript is the <strong>transpose</strong>, not exponentiation.</p>\n</blockquote>\n<p>This formula gives the <strong>camera center <code>C</code> in world coordinates</strong> - exactly what we use during evaluation.</p>",
          "votes": 10,
          "replies": [
            {
              "id": 3187537,
              "author_name": "stardustlove",
              "author_url": "",
              "post_date": "2025-04-26T07:01:57.993000",
              "content": "<p>Thank you so much for your detailed answer! It really helped me understand everything. I truly appreciate your kindness and patience!</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3178395,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-04-14T05:50:12.840000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3178394": "<center><h2 style=\"color:#1f77b4; \">📂 Dataset Structure</h2></center>\n\nEach dataset consists of multiple scenes. A scene is a group of related images that belong together (e.g., different viewpoints of the same object). Some images in a dataset may be <strong>outliers</strong>, meaning they don’t belong to any scene (e.g., random photos taken in the area).\n\nIn the <strong>test data</strong>, all images are mixed in one folder. Your task is:\n1. Cluster the images into groups that represent scenes  \n2. Assign each image to a cluster or label it as an <code>outlier</code>  \n3. Reconstruct the camera pose (R, T) for each image (except outliers)\n\n---\n\n<center><h1 style=\"color:#1f77b4;\">🧠 Camera Pose Estimation and mAA</h1></center>\n\n- A rotation matrix (3×3) - where the camera is pointing  \n- A translation vector (3D) - where the camera is located  \n\n<b>\"Ground truth\"</b> is the true known pose of the camera when the image was taken. In training data, this is given to you. In test data, you try to predict it.\n\nCamera center <b>C</b> in world coordinates is calculated as:  \n**C = −R<sup>T</sup> · T**\nThis transforms the camera pose into a 3D point in space.\n\nCamera centres of one scene:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F25747348%2F61b7f20998538db19346d6d34c961920%2F2025-04-14%20083034.jpg?generation=1744608643370046&alt=media)\n\n---\n\n### <b>Structure from Motion (SfM)</b> is a pipeline that reconstructs:\n- Camera poses for each image  \n- A sparse 3D point cloud of the scene\n\nYour predicted camera centers may be rotated differently, shifted (translated), and scaled.  \nThis is normal in SfM - reconstructions can be geometrically correct but misaligned.\n\nSo, to compare your predicted camera centers <b>C</b> to the ground truth centers <b>Cg</b>, a similarity transformation <b>T</b> (i.e. scale, rotation and translation altogether) is applied. This aligns your 3D camera predictions to the ground truth up to scale and orientation.  \n(Some of your predicted camera poses may be wrong - far from ground truth - and we want to ignore bad data and only use good matches to compute the best alignment.)\n\nTo compute a similarity transform (rotation, translation, and scale) between two sets of 3D points, you need at least 3 point correspondences that are not collinear.\n\n---\n\nA <strong>triplet</strong> is a group of three matching camera centers. Each one has:\n- A predicted 3D position <b>Ci</b>  \n- A known ground-truth position <b>Cgi</b>  \n\nSo a triplet is:  \n<code>{(C1, Cg1), (C2, Cg2), (C3, Cg3)}</code>\n\n<b>RANSAC</b> = Random Sample Consensus - Algorithm to randomly test many triplets, ignore outliers, and find the best transformation.\n\nThe <code>thresholds</code> column lists multiple numeric values per scene. These represent distance or error thresholds that are used to evaluate how accurate predicted 3D camera poses are (compared to ground truth). Each threshold is a maximum allowable error (in meters) to consider a pose prediction as correct.\n\n---\n\nSo use <b>Horn’s method</b> (algorithm used to compute the best similarity transformation between two sets of 3D points) to compute a similarity transformation.  \nFind scale <b>s</b>, rotation <b>R</b>, translation <b>t</b> so that:\n\n<pre><code>T(Ci) = s ⋅ R ⋅ Ci + t</code></pre>\n\nThis is our first estimate <b>T′</b>.  \nThen apply <b>T′</b> to all predicted camera centers <b>Ci</b> and for each one count how many of your predictions become <b>registered cameras</b> (i.e., within error threshold of ground truth):\n\n<pre><code>is_registered = (‖ T(Ci) − Cgi ‖ < threshold)</code></pre>\n\nWhere <code>‖ · ‖</code> denotes Euclidean (L2) distance between two 3D points.\n\nTry many random triplets and get many candidate transformations <b>T′</b>.  \nFor each one, count how many registered cameras it produces and save the best one.  \nTake the best transformation <b>T′</b>, and re-estimate it using all the inliers (not just the original triplet).\n\n---\n\n<center><h2 style=\"color:#2ca02c;\">📊 mAA = mean Average Accuracy</h2></center>\n\nIt measures how accurately your predicted camera centers match the ground truth, after alignment with a similarity transformation.\n\nAs discussed before, you use triplets and Horn’s method inside RANSAC.  \nYou find the best transformation:\n\n<pre><code>T(Ci) = s ⋅ R ⋅ Ci + t</code></pre>\n\nThis aligns your predicted camera centers to the same coordinate frame as the ground-truth.  \nThen apply this transformation - for each predicted camera center <b>Ci</b>, compute:  \n<code>Ci_aligned = T(Ci)</code>\n\nNow for each threshold <b>t_j</b>, check if the difference between <b>Ci_aligned</b> and ground-truth center <b>Cgi</b> is less than <b>t_j</b>:\n\n<pre><code>is_registered = (‖ T(Ci) − Cgi ‖ < threshold)</code></pre>\n\nThen compute registration ratio per threshold for scene:  \nFor each <b>N' = N − 3</b> images of the scene and each threshold <b>t_j</b>, compute the percentage of registered cameras:\n\n<pre><code>rj = (1 / N') * sum(registered)</code></pre>\n\nThat gives you the <b>Average Accuracy (AA)</b> for that scene.  \nThen find average across all thresholds:\n\n<pre><code>mAA = (1 / k) * sum(rj)</code></pre>\n\nWhere:\n- <b>k</b> = number of thresholds\n\nOnce you compute <b>mAA per scene</b>, you calculate final mAA for all scenes by averaging across scenes.\n\n---\n\n<center><h2 style=\"color:#1f77b4;\">🧠 Clustering and Final Evaluation Score</h2></center>\n\nEach scene from the ground-truth data is denoted as <code>S<sub>ki</sub></code>, and your predicted clusters are <code>C<sub>kj</sub></code>.\n\nThe evaluation will:\n- Compare your clusters with the true scenes  \n- Check how well your clustering matches the original grouping  \n- Measure how accurate your pose predictions are (as explained earlier with mAA)\n\nEach ground-truth scene <code>S<sub>ki</sub></code> is matched to the user-submitted cluster <code>C<sub>kj</sub></code> that gives the <b>highest mAA</b> after alignment.  \nAll images labeled as outliers are <b>excluded</b> from this matching step.\n\n---\n\n<h3 style=\"color:#2ca02c;\">📊 Clustering Score (Precision)</h3>\n\nAfter the best cluster <code>C<sub>kji</sub></code> is selected for each scene <code>S<sub>ki</sub></code>, compute the <b>clustering score</b> as:\n\n<pre><code>clustering_score = |S ∩ C| / |C|</code></pre>\n\nWhere:\n- <code>|S ∩ C|</code> is the number of images that correctly overlap between your predicted cluster and the real scene  \n- <code>|C|</code> is the total number of images in your predicted cluster\n\nThis measures <b>how pure your cluster is</b> - i.e., out of everything you said belongs to this scene, how many actually do?\n\n---\n\n<h3 style=\"color:#2ca02c;\">📈 Pose Score (mAA)</h3>\n\nAs explained earlier, for each image:\n- Apply the best alignment transform T to your predicted camera centers  \n- Check if each transformed pose is within the threshold of the ground-truth center  \n- Count the number of registered cameras at each threshold  \n- Compute registration ratios across thresholds  \n- Average across thresholds → mAA per scene  \n- Average over scenes → mAA for dataset\n\nThis mAA represents <b>recall</b> - how many real scene images you found and aligned correctly.\n\n---\n\n<h3 style=\"color:#2ca02c;\">🔗 Final Score Per Dataset</h3>\n\nNow both mAA (recall-like) and clustering score (precision-like) are combined using the <b>harmonic mean</b>, similar to the F1-score in classification:\n\n<pre><code>final_score_per_dataset = 2 * (mAA * clustering_score) / (mAA + clustering_score)</code></pre>\n\nThis ensures that:\n- If either clustering or pose accuracy is bad, your score is penalized  \n- You must do <b>both clustering and reconstruction well</b> to score high\n\n---\n\n<h3 style=\"color:#ff7f0e;\">📊 Final Challenge Score</h3>\n\nYou will repeat this evaluation for each dataset <code>D<sub>k</sub></code>.  \nThe final score for your submission is:\n\n<pre><code>mean(final_score_per_dataset)</code></pre>\n\ni.e., the average harmonic mean across all datasets.\n\n---\n<center><h1 style=\"color:#1f77b4;\">🧩 Scene Clustering & Pose Estimation (SfM)</h1></center>\n\n---\n\n#### <h3 style=\"color:#2ca02c;\">📌 <b>1) Scene Clustering</b></h3>\n\n---\n\n##### <h4 style=\"color:#2ca02c;\">📘 <b>1.1) Unsupervised Clustering on Global Descriptors</b></h4>\n\nExtract global embeddings using models like <code>CLIP</code>, <code>DINOv2</code>, etc., then cluster them using algorithms like <b>DBSCAN</b>, <b>KMeans</b>, <b>Agglomerative Clustering</b>...\n\n📌 **Image features** are compact vector representations of the visual content of the image.\n\n---\n\n##### <h4 style=\"color:#1f77b4;\">📦 Global Descriptors</h4>\n\nGlobal descriptors are single vectors that represent the entire image - its objects, textures, and layout.\n\n✅ **Use cases**:\n- Image clustering  \n- Image retrieval  \n- Scene classification  \n\n<br>\n\n**DINOv2** is part of the **DINO** (Distillation with No Labels) family of self-supervised vision transformers:\n- Learns by matching feature vectors of different augmented views of the same image.\n- Ensures different images are mapped to distant feature vectors.\n- Requires no manual labels (contrastive learning).\n\n---\n\n<h4 style=\"color:#1f77b4;\">🔧 How Global Descriptor Clustering Works</h4>\n\n- Extract a single embedding per image (e.g., using DINOv2).\n- Compute pairwise distances (typically cosine distance) between all embeddings.\n- Cluster embeddings using:\n  - **DBSCAN** - density-based clustering without needing the number of clusters.\n  - **Agglomerative Clustering** - hierarchical grouping.\n  - **KMeans** - if the number of clusters is approximately known.\n  - **Graph-based Clustering** - (see below 👇).\n- The result: each cluster represents one predicted scene.\n\n---\n\n#### <h4 style=\"color:#2ca02c;\">📘 <b> Graph-Based Clustering</b></h4>\n\n✅ In our pipeline:\n- Compute **pairwise cosine distances** between image embeddings extracted by **DINOv2**.\n- Build a **graph** where:\n  - Each node represents an image.\n  - An edge is created between two images if their distance is below a threshold (e.g., <code>0.3</code>).\n- **Connected components** of this graph are treated as **predicted scenes**.\n- Images with no edges (isolated nodes) are labeled as **outliers**.\n\n<p style=\"font-size:16px;\">\nThis approach is faster and more scalable than DBSCAN for large datasets, and does not require knowing the number of clusters beforehand.\n</p>\n\n📌 **Note**:  \nI do not use ALIKED or LightGlue at this stage.  \nThey are used later in the pipeline for **local feature extraction and matching** - when reconstructing 3D camera poses.\n\n---\n\n##### <h4 style=\"color:#1f77b4;\">🔍 Local Keypoints & Descriptors (General Theory)</h4>\n\nWhile  my solution clusters images using global descriptors, another common approach in image matching is to use **local features**:\n\n**Keypoints** are distinctive image regions (corners, blobs, edges) that:\n- Are stable under changes in viewpoint and lighting.\n- Can be reliably detected and matched across different images.\n\n🧠 **Common keypoint detectors**:\n- **Hand-crafted**: SIFT, ORB, AKAZE\n- **Learned**: SuperPoint, D2Net, R2D2\n\n**Local Descriptors**:  \nSmall feature vectors extracted around keypoints describing their local visual neighborhood.  \nUsed for matching keypoints between different images.\n\n---\n\n<h4 style=\"color:rgb(255, 11, 214);\">⚡ Important Note</h4>\n\nIn our case:\n- **Global features (DINOv2)** are used for initial scene clustering via **graph-based connected components**.\n- **Local features (ALIKED)** and **feature matching (LightGlue)** are only used afterward for precise camera pose reconstruction (**Structure-from-Motion**, SfM).\n\n\n---\n#### <h3 style=\"color:#2ca02c;\">📌 <b>2) Estimating Camera Poses (SfM)</b></h3>\n\n---\n\n<h4 style=\"color:#1f77b4;\">📐 What is Structure-from-Motion (SfM)?</h4>\n\n<p style=\"font-size:16px;\">\n<b>Structure-from-Motion (SfM)</b> is a computer vision pipeline that reconstructs:\n<ul>\n<li>📸 The pose of each camera (Rotation matrix <b>R</b> and Translation vector <b>T</b>).</li>\n<li>🧱 A sparse 3D point cloud of the scene (estimated 3D locations of keypoints).</li>\n</ul>\n\nGiven only a set of 2D images - without depth, without camera poses - SfM **recovers** both the camera trajectories and the scene structure.\n</p>\n\n---\n\n<h4 style=\"color:#1f77b4;\">📐 Simplified SfM Pipeline (General)</h4>\n\n<ol style=\"font-size:16px;\">\n<li><b>Pairwise Matching</b> - Find correspondences between keypoints detected in different images.</li>\n<li><b>Relative Pose Estimation</b> - Estimate the relative rotation (R) and translation (T) between pairs of images using 2D matches (via Essential Matrix decomposition).</li>\n<li><b>Triangulation</b> - Estimate 3D coordinates of points by triangulating matched keypoints from multiple views.</li>\n<li><b>Incremental SfM</b> - Add images one-by-one by solving PnP (Perspective-n-Point) problems.</li>\n<li><b>Bundle Adjustment</b> - Optimize all camera poses and 3D points jointly to minimize reprojection errors.</li>\n</ol>\n\n---\n\n<h4 style=\"color:#2ca02c;\">ALIKED, LightGlue, pycolmap</h4>\n\n<p style=\"font-size:16px;\">\n<ul>\n<li><b>Local feature detection</b>:<strong>ALIKED</strong> - a fast, learned local feature detector and descriptor extractor based on deep learning.</li>\n<li><b>Feature matching</b>: <strong>LightGlue</strong> - an attention-based neural matcher that finds reliable correspondences even under strong viewpoint and illumination changes.</li>\n<li><b>3D reconstruction and camera pose estimation</b>: <strong>pycolmap</strong> - a Python wrapper over <strong>COLMAP</strong>, a powerful Structure-from-Motion software, to reconstruct the camera trajectories and sparse 3D point cloud automatically.</li>\n</ul>\n\n</p>\n\n---\n**Triangulation**\t- Estimating 3D coordinates of a point from two or more images of it (2D → 3D)  \n**PnP (Perspective-n-Point)** - Given 2D-3D correspondences, estimate the camera pose (R, T)  \n**Incremental SfM** - Build the scene step by step: triangulate → PnP → repeat\n\n<h4 style=\"color:#1f77b4;\">🔧 How It Works in Practice</h4>\n\n<ol style=\"font-size:16px;\">\n<li><b>Extract ALIKED local features</b> for each image - keypoints + descriptors are saved into .h5 files.</li>\n<li><b>Match features using LightGlue</b> - for selected image pairs (shortlisted by global similarity or exhaustively).</li>\n<li><b>Import features and matches into a COLMAP database</b> using <code>h5_to_db</code> utilities.</li>\n<li><b>Run exhaustive matching and geometric verification</b> inside COLMAP (pycolmap API).</li>\n<li><b>Perform incremental reconstruction</b> (triangulation + pose estimation + bundle adjustment).</li>\n<li><b>Output:</b> For each image: \n<ul>\n<li>Rotation matrix <code>R</code></li>\n<li>Translation vector <code>T</code></li>\n</ul>\n</li>\n<li>If an image cannot be reconstructed (e.g., insufficient matches) - assign <code>NaN</code> to its pose fields.</li>\n</ol>\n\n---\n\n<h4 style=\"color:#1f77b4;\">📦 Tools Used</h4>\n\n<ul style=\"font-size:16px;\">\n<li><b>ALIKED</b> - lightweight keypoint extractor based on deep learning, efficient for fast and robust feature detection and description.</li>\n<li><b>LightGlue</b> - transformer-based feature matcher optimized for large viewpoint, scale, and illumination changes. Learns how to attend to correct matches.</li>\n<li><b>COLMAP (via pycolmap)</b> - one of the most powerful SfM engines, supporting incremental mapping, exhaustive matching, triangulation, PnP pose estimation, and bundle adjustment.</li>\n</ul>\n\n---\n\n<h2 style=\"color:rgb(255, 11, 214);\"><span style=\"font-size: 30px;\">📸 Reconstruct Camera Poses (SfM)</span></h2>\n\n<p style=\"font-size:16px;\">\nUsing COLMAP, pycolmap, or another SfM engine, we perform the following:\n<ul>\n<li>Input: All images assigned to the same cluster.</li>\n<li>Process:\n    <ul>\n    <li>Extract ALIKED features.</li>\n    <li>Match using LightGlue.</li>\n    <li>Import data into COLMAP database.</li>\n    <li>Run incremental reconstruction via pycolmap.</li>\n    </ul>\n</li>\n<li>Output: For each image, recover:\n    <ul>\n    <li>Rotation matrix <code>R</code> (3x3)</li>\n    <li>Translation vector <code>T</code> (3x1)</li>\n    </ul>\n</li>\n<li>If reconstruction fails for some images: assign NaN values to their pose fields.</li>\n</ul>\n</p>\n\n---",
    "3204068": "Great stuff! And all in very details, just what I needed to get started! Thanks @dmytrobuhai !",
    "3203991": "Thank you a lot for the extremely useful and concise explanations! It is my first day during which I am deeply diving into the competition details and this topic definitely helped me to quickly grasp main theoretical concepts. 🔥",
    "3186803": "I have some confusion about the calculation formula for the Camera center C in world coordinates. Could you please explain how it is derived? I would be very grateful if you could reply.",
    "3178395": ""
  }
}