{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":91498,"databundleVersionId":11655853,"sourceType":"competition"},{"sourceId":11296095,"sourceType":"datasetVersion","datasetId":7063493}],"dockerImageVersionId":30919,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### Step 1.1: Import Libraries\n\n#### Explanation\nWe’ll begin by importing the necessary libraries for exploring the dataset. These libraries will help us navigate the file system, load images, handle numerical data, work with CSV files, and visualize images.\n- `os`: For navigating the file system (e.g., listing files in directories).\n- `cv2` (OpenCV): For loading and processing images (e.g., reading PNG files).\n- `numpy`: For numerical operations (e.g., manipulating image arrays).\n- `pandas`: For loading and manipulating CSV files (e.g., `sample_submission.csv`).\n- `matplotlib.pyplot`: For visualizing images in a grid.\n- `pathlib.Path`: For handling file paths in a cross-platform way.","metadata":{"execution":{"iopub.status.busy":"2025-04-04T11:24:38.148820Z","iopub.execute_input":"2025-04-04T11:24:38.149222Z","execution_failed":"2025-04-04T11:24:46.429Z"}}},{"cell_type":"code","source":"# Step 1.1: Import Libraries\nimport os\nimport cv2\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nfrom pathlib import Path","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-06T15:12:37.139017Z","iopub.execute_input":"2025-04-06T15:12:37.139379Z","iopub.status.idle":"2025-04-06T15:12:37.143344Z","shell.execute_reply.started":"2025-04-06T15:12:37.139345Z","shell.execute_reply":"2025-04-06T15:12:37.142497Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Step 1.2: Define Paths\n\n#### Explanation\nWe’ll define the paths to the dataset files and directories to ensure we can access them consistently throughout the notebook. We’ll also add a print statement to confirm the paths are set correctly.\n\n- `data_path`: The root directory of the competition dataset.\n- `train_path`: Path to the `train` directory containing training images.\n- `test_path`: Path to the `test` directory containing test images.\n- `sample_submission_path`, `train_labels_path`, `train_thresholds_path`: Paths to the CSV files.\n- A print statement will display each path to verify they are correct.","metadata":{}},{"cell_type":"code","source":"# Step 1.2: Define Paths\ndata_path = Path('/kaggle/input/image-matching-challenge-2025')\ntrain_path = data_path / 'train'\ntest_path = data_path / 'test'\nsample_submission_path = data_path / 'sample_submission.csv'\ntrain_labels_path = data_path / 'train_labels.csv'\ntrain_thresholds_path = data_path / 'train_thresholds.csv'\n\n# Print paths to confirm\nprint(\"Data Path:\", data_path)\nprint(\"Train Path:\", train_path)\nprint(\"Test Path:\", test_path)\nprint(\"Sample Submission Path:\", sample_submission_path)\nprint(\"Train Labels Path:\", train_labels_path)\nprint(\"Train Thresholds Path:\", train_thresholds_path)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-06T15:12:37.184552Z","iopub.execute_input":"2025-04-06T15:12:37.184786Z","iopub.status.idle":"2025-04-06T15:12:37.192641Z","shell.execute_reply.started":"2025-04-06T15:12:37.184766Z","shell.execute_reply":"2025-04-06T15:12:37.191158Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Step 1.3: Explore the Directory Structure\n\nIn this step, we aim to understand the organization of the image data by exploring the directory structure of the train and test directories. This is a crucial initial step because it helps us confirm the dataset’s layout, identify the number of images, and understand how scenes or datasets are grouped. We’ll use a method to recursively traverse the directories, listing each subdirectory, the number of files it contains, and a few sample filenames. This will give us insights into the diversity of scenes, the scale of the dataset, and any potential challenges, such as the presence of outliers or varying numbers of images per scene. By printing this information, we can verify that the data is accessible and structured as expected, setting the stage for further exploration of the images and CSV files.","metadata":{}},{"cell_type":"code","source":"# Step 1.3: Explore the Directory Structure\nprint(\"Train directory contents:\")\nfor root, dirs, files in os.walk(train_path):\n    print(f\"Directory: {root}\")\n    print(f\"Number of files: {len(files)}\")\n    if len(files) > 0:\n        print(f\"Sample files: {files[:3]}\")\n    print()\n\nprint(\"Test directory contents:\")\nfor root, dirs, files in os.walk(test_path):\n    print(f\"Directory: {root}\")\n    print(f\"Number of files: {len(files)}\")\n    if len(files) > 0:\n        print(f\"Sample files: {files[:3]}\")\n    print()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-06T15:12:37.193813Z","iopub.execute_input":"2025-04-06T15:12:37.194131Z","iopub.status.idle":"2025-04-06T15:12:38.763743Z","shell.execute_reply.started":"2025-04-06T15:12:37.194077Z","shell.execute_reply":"2025-04-06T15:12:38.762922Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Step 1.4: Load and Display CSV Files\n\nNow that we’ve explored the directory structure, we need to examine the CSV files provided in the dataset: sample_submission.csv, train_labels.csv, and train_thresholds.csv. These files contain critical information about the task, such as the expected submission format, ground truth data for training, and thresholds for processing matches. We’ll load each CSV file into a DataFrame and display the first few rows to understand their structure, including column names and data types. Additionally, we’ll print the shape of each DataFrame to get a sense of the data volume. This step is essential for confirming the task requirements (e.g., what we need to predict) and understanding the ground truth available for training, which will guide our modeling approach.","metadata":{}},{"cell_type":"code","source":"# Step 1.4: Load and Display CSV Files\n# Sample Submission\nprint(\"Sample Submission:\")\nsample_submission = pd.read_csv(sample_submission_path)\nprint(sample_submission.head())\nprint(f\"Shape: {sample_submission.shape}\")\nprint()\n\n# Train Labels\nprint(\"Train Labels:\")\ntrain_labels = pd.read_csv(train_labels_path)\nprint(train_labels.head())\nprint(f\"Shape: {train_labels.shape}\")\nprint()\n\n# Train Thresholds\nprint(\"Train Thresholds:\")\ntrain_thresholds = pd.read_csv(train_thresholds_path)\nprint(train_thresholds.head())\nprint(f\"Shape: {train_thresholds.shape}\")\nprint()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-06T15:12:38.765038Z","iopub.execute_input":"2025-04-06T15:12:38.765328Z","iopub.status.idle":"2025-04-06T15:12:38.798537Z","shell.execute_reply.started":"2025-04-06T15:12:38.765305Z","shell.execute_reply":"2025-04-06T15:12:38.797773Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Step 1.5: Visualize Sample Images\n\nHaving explored the directory structure and CSV files, we now need to visually inspect the images in the training set to understand their characteristics, such as scene types, lighting conditions, and image dimensions. This step is crucial for planning preprocessing steps, such as resizing or augmentation, and for getting a sense of the visual challenges (e.g., occlusions, varying angles). We’ll select a few sample images from the training set, load them, and display them in a grid. We’ll also print their file paths and dimensions to confirm they load correctly and to check for consistency in image sizes. Since the filenames indicate a .png extension, we’ll ensure we’re loading PNG files, and we’ll add debugging to handle any loading issues.","metadata":{}},{"cell_type":"code","source":"# Step 1.5: Visualize Sample Images from Train\n# Get a few sample images (use .png since filenames have .png extension)\ntrain_images = list(train_path.glob('**/*.png'))[:4]\nprint(\"Sample image paths:\")\nfor img_path in train_images:\n    print(img_path)\n\n# Load and visualize images\nfig, axes = plt.subplots(2, 2, figsize=(10, 10))\naxes = axes.ravel()\n\nfor idx, img_path in enumerate(train_images):\n    # Load image\n    img = cv2.imread(str(img_path))\n    if img is None:\n        print(f\"Failed to load image: {img_path}\")\n        continue\n    # Convert BGR to RGB\n    img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)\n    # Display image\n    axes[idx].imshow(img)\n    axes[idx].set_title(f\"Image {idx+1}: {img_path.name}\")\n    axes[idx].axis('off')\n    print(f\"Image {idx+1} shape: {img.shape}\")\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-06T15:12:38.799583Z","iopub.execute_input":"2025-04-06T15:12:38.799878Z","iopub.status.idle":"2025-04-06T15:12:39.799766Z","shell.execute_reply.started":"2025-04-06T15:12:38.799847Z","shell.execute_reply":"2025-04-06T15:12:39.798264Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Step 1.6: Analyze the Explored Data\n\n## Explanation\n\n\nWith the directory structure, CSV files, and sample images explored, we need to analyze the data to confirm the task requirements and understand the dataset’s characteristics. We’ll summarize the findings from the train and test directories, the CSV files (sample_submission.csv, train_labels.csv, train_thresholds.csv), and the visualized images. This analysis will help us identify the key tasks (e.g., clustering, image matching, 3D reconstruction), assess the diversity of scenes, and note any challenges (e.g., varying image sizes, outliers). We’ll print a summary of the total number of images, the structure of the CSV files, and observations about the images to consolidate our understanding before defining the problem in the next step.","metadata":{}},{"cell_type":"code","source":"# Step 1.6: Analyze the Explored Data\n# Summarize the directory structure\ntrain_dirs = [d for d in os.listdir(train_path) if os.path.isdir(os.path.join(train_path, d))]\ntest_dirs = [d for d in os.listdir(test_path) if os.path.isdir(os.path.join(test_path, d))]\n\ntrain_image_count = sum(len(files) for root, dirs, files in os.walk(train_path) if files)\ntest_image_count = sum(len(files) for root, dirs, files in os.walk(test_path) if files)\n\nprint(\"Summary of Directory Structure:\")\nprint(f\"Number of train datasets: {len(train_dirs)}\")\nprint(f\"Train datasets: {train_dirs}\")\nprint(f\"Total train images: {train_image_count}\")\nprint(f\"Number of test datasets: {len(test_dirs)}\")\nprint(f\"Test datasets: {test_dirs}\")\nprint(f\"Total test images: {test_image_count}\")\nprint()\n\n# Summarize CSV files\nprint(\"Summary of CSV Files:\")\nprint(\"Sample Submission Columns:\", list(sample_submission.columns))\nprint(\"Sample Submission Shape:\", sample_submission.shape)\nprint(\"Train Labels Columns:\", list(train_labels.columns))\nprint(\"Train Labels Shape:\", train_labels.shape)\nprint(\"Train Thresholds Columns:\", list(train_thresholds.columns))\nprint(\"Train Thresholds Shape:\", train_thresholds.shape)\nprint()\n\n# Summarize image characteristics\nprint(\"Image Characteristics:\")\nprint(\"Sample image paths (from previous step):\")\nfor img_path in train_images:\n    print(img_path)\nprint(\"Sample image shapes (from previous step):\")\nprint(\"Image 1 shape: (1024, 576, 3)\")\nprint(\"Image 2 shape: (768, 1024, 3)\")\nprint(\"Image 3 shape: (1024, 576, 3)\")\nprint(\"Image 4 shape: (768, 1024, 3)\")\nprint(\"Observation: Images are from the 'amy_gardens' dataset, depicting a garden with peach trees, supported by poles and wires, under cloudy skies.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-06T15:12:39.800753Z","iopub.execute_input":"2025-04-06T15:12:39.801023Z","iopub.status.idle":"2025-04-06T15:12:39.832840Z","shell.execute_reply.started":"2025-04-06T15:12:39.801000Z","shell.execute_reply":"2025-04-06T15:12:39.831813Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Step 2.1: Define the Problem Technically\n\nWith the dataset explored, we need to define the problem technically by identifying the key tasks, evaluation metrics, and challenges based on the data. The sample_submission.csv indicates that we must predict the scene (cluster) and camera pose (rotation matrix and translation vector) for each test image, suggesting a two-part task: clustering images into scenes and performing 3D reconstruction. The train_labels.csv provides ground truth poses for training images, which we can use to validate our pipeline, but we’ll need to generate keypoint matches between images ourselves. The train_thresholds.csv offers thresholds for filtering matches, which will be useful during image matching. The images vary in size, and some datasets include outliers, which we’ll need to handle. We’ll summarize the tasks, hypothesize the evaluation metric, and outline our approach, printing these details to ensure clarity before proceeding to implementation.","metadata":{}},{"cell_type":"code","source":"# Step 2.1: Define the Problem Technically\n# Define the tasks\nprint(\"Task Definition:\")\nprint(\"1. Clustering: Assign each test image to a scene (e.g., 'cluster0' in sample_submission.csv).\")\nprint(\"   - Training data is already clustered (e.g., 'fountain' in imc2023_haiper).\")\nprint(\"   - Test data requires clustering, handling outliers (e.g., 'outliers_out_et003.png').\")\nprint(\"2. 3D Reconstruction: For each image, predict its camera pose (rotation_matrix, translation_vector).\")\nprint(\"   - Requires image matching to find keypoint correspondences between images in the same scene.\")\nprint(\"   - Use matches to estimate camera poses and reconstruct the 3D scene.\")\nprint()\n\n# Hypothesize the evaluation metric\nprint(\"Evaluation Metric (Hypothesized):\")\nprint(\"- Clustering: Likely evaluated using Adjusted Rand Index (ARI) or purity, comparing predicted clusters to ground truth.\")\nprint(\"- 3D Reconstruction: Likely evaluated using reprojection error or pose accuracy (error in rotation and translation).\")\nprint(\"Note: train_thresholds.csv suggests image matching performance may be evaluated at different thresholds (e.g., mAP).\")\nprint()\n\n# Outline the approach\nprint(\"Approach Outline:\")\nprint(\"1. Clustering:\")\nprint(\"   - Extract image embeddings using a pre-trained CNN (e.g., ResNet).\")\nprint(\"   - Cluster images using HDBSCAN or K-means to group them into scenes.\")\nprint(\"2. 3D Reconstruction:\")\nprint(\"   - Image Matching: Use LoFTR or SuperGlue to find keypoint matches between image pairs.\")\nprint(\"   - Structure from Motion (SfM): Use COLMAP to estimate camera poses and reconstruct the 3D scene.\")\nprint(\"   - Post-Processing: Filter matches using thresholds from train_thresholds.csv.\")\nprint()\n\n# Identify challenges\nprint(\"Challenges:\")\nprint(\"- Varying image sizes (e.g., 1024x576, 768x1024) require resizing for consistency.\")\nprint(\"- Outliers in datasets (e.g., 'outliers_out_et003.png') need robust clustering.\")\nprint(\"- Diverse scenes (landmarks, heritage, natural) may require domain-specific handling.\")\nprint(\"- Hidden test set may include new scenes, requiring generalization.\")\nprint(\"- Notebook must run within Kaggle's time limits (typically 20 minutes).\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-06T15:12:39.833790Z","iopub.execute_input":"2025-04-06T15:12:39.834049Z","iopub.status.idle":"2025-04-06T15:12:39.845677Z","shell.execute_reply.started":"2025-04-06T15:12:39.834026Z","shell.execute_reply":"2025-04-06T15:12:39.844614Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Step 2.2: Extract Image Embeddings for Clustering\n\nTo cluster the test images into scenes, we first need to extract feature embeddings that capture the visual content of each image. We’ll use a pre-trained convolutional neural network (CNN), specifically ResNet50, which is available in the torchvision library and has been pre-trained on ImageNet. This model will generate a fixed-size embedding for each image, which we can then use for clustering. We’ll load the test images, preprocess them (resize and normalize as required by ResNet50), pass them through the model to get embeddings, and store the embeddings along with their corresponding image paths. We’ll print the number of images processed and the shape of the embeddings to confirm the extraction is successful. This step is foundational for clustering, as the quality of the embeddings will directly impact the clustering performance.","metadata":{}},{"cell_type":"code","source":"# Step 2.2: Extract Image Embeddings for Clustering\nimport torch\nimport torchvision.models as models\nimport torchvision.transforms as transforms\nfrom PIL import Image\n\n# Load pre-trained ResNet50 model\nmodel = models.resnet50(pretrained=True)\nmodel.eval()  # Set to evaluation mode\nmodel = model.cuda() if torch.cuda.is_available() else model  # Use GPU if available\n\n# Remove the final fully connected layer to get embeddings\nmodel = torch.nn.Sequential(*list(model.children())[:-1])\n\n# Define image preprocessing\npreprocess = transforms.Compose([\n    transforms.Resize((224, 224)),  # Resize to 224x224 as required by ResNet\n    transforms.ToTensor(),  # Convert to tensor\n    transforms.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225])  # Normalize for ImageNet\n])\n\n# Get all test images\ntest_images = list(test_path.glob('**/*.png'))\nprint(f\"Total test images to process: {len(test_images)}\")\n\n# Extract embeddings\nembeddings = []\nimage_paths = []\n\nfor img_path in test_images:\n    # Load and preprocess image\n    img = Image.open(img_path).convert('RGB')\n    img_tensor = preprocess(img).unsqueeze(0)  # Add batch dimension\n    if torch.cuda.is_available():\n        img_tensor = img_tensor.cuda()\n    \n    # Get embedding\n    with torch.no_grad():\n        embedding = model(img_tensor)\n    embedding = embedding.cpu().numpy().flatten()  # Flatten to 1D array\n    embeddings.append(embedding)\n    image_paths.append(img_path)\n\n# Convert to numpy array\nembeddings = np.array(embeddings)\nprint(f\"Embeddings shape: {embeddings.shape}\")\nprint(f\"Number of images processed: {len(image_paths)}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-06T15:12:39.846775Z","iopub.execute_input":"2025-04-06T15:12:39.847058Z","iopub.status.idle":"2025-04-06T15:12:43.906520Z","shell.execute_reply.started":"2025-04-06T15:12:39.847034Z","shell.execute_reply":"2025-04-06T15:12:43.905673Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Step 2.3: Install HDBSCAN and Cluster Test Images\n## Explanation\n\n\nBefore we can cluster the test images using HDBSCAN, we need to install the hdbscan library, as it’s not available in the default Kaggle environment. We’ll use pip to install it directly within the notebook. Once installed, we’ll proceed with clustering the test images using the embeddings extracted in the previous step. HDBSCAN is a robust clustering algorithm that can handle outliers, which is important given the presence of images labeled as outliers in the dataset. We’ll cluster the images, assign labels, and map them to scene names (e.g., cluster0, cluster1), printing the number of clusters, the number of images per cluster, and the cluster assignments for a few images to verify the results. This will prepare us for image matching within each cluster.","metadata":{}},{"cell_type":"code","source":"# Step 2.3: Install HDBSCAN and Cluster Test Images\n# Install hdbscan\n!pip install hdbscan\n\nimport hdbscan\n\n# Perform clustering with HDBSCAN\nclusterer = hdbscan.HDBSCAN(min_cluster_size=5, min_samples=3)\ncluster_labels = clusterer.fit_predict(embeddings)\n\n# Analyze clustering results\nunique_labels = np.unique(cluster_labels)\nnum_clusters = len(unique_labels) - (1 if -1 in unique_labels else 0)  # Exclude noise label (-1)\nprint(f\"Number of clusters found: {num_clusters}\")\nprint(f\"Cluster labels: {unique_labels}\")\n\n# Count images per cluster\nfor label in unique_labels:\n    if label == -1:\n        print(f\"Number of noise points (label -1): {np.sum(cluster_labels == label)}\")\n    else:\n        print(f\"Number of images in cluster {label}: {np.sum(cluster_labels == label)}\")\n\n# Map cluster labels to scene names (e.g., cluster0, cluster1, ...)\nscene_mapping = {label: f\"cluster{label}\" for label in unique_labels if label != -1}\nscene_mapping[-1] = \"noise\"  # Label for outliers\n\n# Assign scene names to images\nimage_scenes = [scene_mapping[label] for label in cluster_labels]\n\n# Print cluster assignments for the first few images\nprint(\"\\nCluster assignments for the first 5 images:\")\nfor i in range(min(5, len(image_paths))):\n    print(f\"Image: {image_paths[i].name}, Scene: {image_scenes[i]}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-06T15:12:43.908819Z","iopub.execute_input":"2025-04-06T15:12:43.909072Z","iopub.status.idle":"2025-04-06T15:12:47.304972Z","shell.execute_reply.started":"2025-04-06T15:12:43.909051Z","shell.execute_reply":"2025-04-06T15:12:47.304028Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Step 2.4: Visualize Clusters Using Dimensionality Reduction\n\nTo better understand the clustering results, we’ll visualize the embeddings in a 2D space by reducing their dimensionality from 2048 dimensions to 2 dimensions using UMAP (Uniform Manifold Approximation and Projection), a popular technique for visualizing high-dimensional data. We’ll then create a scatter plot where each point represents an image, colored by its cluster label, with noise points (label -1) marked distinctly. This visualization will help us assess whether the clusters are well-separated and if the noise points are appropriately identified. We’ll also print a summary of the clustering results to confirm the number of clusters and images per cluster, ensuring we have a clear understanding before proceeding to image matching within each cluster.","metadata":{}},{"cell_type":"code","source":"# Step 2.4: Visualize Clusters Using Dimensionality Reduction\n# Install umap-learn\n!pip install umap-learn\n\nimport umap\nimport seaborn as sns\n\n# Reduce dimensionality of embeddings to 2D using UMAP\nreducer = umap.UMAP(n_components=2, random_state=42)\nembeddings_2d = reducer.fit_transform(embeddings)\n\n# Create a scatter plot of the 2D embeddings\nplt.figure(figsize=(10, 8))\nsns.scatterplot(x=embeddings_2d[:, 0], y=embeddings_2d[:, 1], hue=cluster_labels, palette=\"deep\", style=cluster_labels, size=cluster_labels, sizes=(50, 200))\nplt.title(\"2D Visualization of Test Image Clusters (UMAP)\")\nplt.xlabel(\"UMAP Component 1\")\nplt.ylabel(\"UMAP Component 2\")\nplt.legend(title=\"Cluster Label\")\nplt.show()\n\n# Print summary of clustering results\nprint(\"Clustering Summary:\")\nprint(f\"Number of clusters found: {num_clusters}\")\nprint(f\"Total images: {len(cluster_labels)}\")\nfor label in unique_labels:\n    if label == -1:\n        print(f\"Number of noise points (label -1): {np.sum(cluster_labels == label)}\")\n    else:\n        print(f\"Number of images in cluster {label}: {np.sum(cluster_labels == label)}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-06T15:12:47.307373Z","iopub.execute_input":"2025-04-06T15:12:47.307637Z","iopub.status.idle":"2025-04-06T15:13:24.691492Z","shell.execute_reply.started":"2025-04-06T15:12:47.307614Z","shell.execute_reply":"2025-04-06T15:13:24.690428Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Step 3.1: Correct Image Matching with LoFTR and Visualize Matches\n\nThe error occurs because LoFTR expects a single-channel grayscale image (shape [batch_size, 1, height, width]), but the input tensor has 2 channels. This might be due to an issue with the batch dimension or an incorrect assumption about the input shape after grayscale conversion. We’ll fix this by explicitly ensuring the input images are loaded, converted to grayscale, and batched correctly for LoFTR. We’ll select two images from cluster 0, perform keypoint matching, and visualize the matches by drawing lines between corresponding keypoints in a combined image. We’ll also print the shape of the input tensors to debug the issue and confirm the number of matches found. This visualization will help us verify that the matching process is working correctly before proceeding to 3D reconstruction.","metadata":{}},{"cell_type":"code","source":"# Step 3.1: Correct Image Matching with LoFTR and Visualize Matches\nimport kornia\nfrom kornia.feature import LoFTR\nimport kornia as K\nimport kornia.feature as KF\n\n# Select images from the largest cluster (cluster 0)\ncluster_0_indices = [i for i, label in enumerate(cluster_labels) if label == 0]\nif len(cluster_0_indices) < 2:\n    print(\"Not enough images in cluster 0 to perform matching.\")\nelse:\n    # Select the first two images from cluster 0\n    img1_path = image_paths[cluster_0_indices[0]]\n    img2_path = image_paths[cluster_0_indices[1]]\n    \n    # Load images using PIL to ensure correct RGB loading\n    img1 = Image.open(img1_path).convert('RGB')\n    img2 = Image.open(img2_path).convert('RGB')\n    \n    # Preprocess images: resize and convert to tensor\n    preprocess = transforms.Compose([\n        transforms.Resize((640, 640)),  # Resize to a smaller size for faster processing\n        transforms.ToTensor(),  # Convert to tensor\n    ])\n    \n    img1_tensor = preprocess(img1)\n    img2_tensor = preprocess(img2)\n    \n    # Convert to grayscale for LoFTR\n    img1_gray = K.color.rgb_to_grayscale(img1_tensor)  # Should be [1, H, W]\n    img2_gray = K.color.rgb_to_grayscale(img2_tensor)  # Should be [1, H, W]\n    \n    # Add batch dimension explicitly\n    img1_gray = img1_gray.unsqueeze(0)  # Shape: [1, 1, H, W]\n    img2_gray = img2_gray.unsqueeze(0)  # Shape: [1, 1, H, W]\n    \n    # Print shapes to debug\n    print(f\"Image 1 grayscale shape: {img1_gray.shape}\")\n    print(f\"Image 2 grayscale shape: {img2_gray.shape}\")\n    \n    # Move to GPU if available\n    device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')\n    img1_gray = img1_gray.to(device)\n    img2_gray = img2_gray.to(device)\n    \n    # Initialize LoFTR\n    matcher = LoFTR(pretrained='outdoor').to(device)\n    matcher.eval()\n    \n    # Prepare input for LoFTR\n    input_dict = {\n        \"image0\": img1_gray,  # Shape: [1, 1, H, W]\n        \"image1\": img2_gray   # Shape: [1, 1, H, W]\n    }\n    \n    # Perform matching\n    with torch.no_grad():\n        correspondences = matcher(input_dict)\n    \n    # Extract keypoints and matches\n    mkpts0 = correspondences['keypoints0'].cpu().numpy()\n    mkpts1 = correspondences['keypoints1'].cpu().numpy()\n    print(f\"Number of matches found: {len(mkpts0)}\")\n    \n    # Load images for visualization using OpenCV\n    img1_np = cv2.imread(str(img1_path))\n    img2_np = cv2.imread(str(img2_path))\n    img1_np = cv2.cvtColor(img1_np, cv2.COLOR_BGR2RGB)\n    img2_np = cv2.cvtColor(img2_np, cv2.COLOR_BGR2RGB)\n    \n    # Resize images for visualization to match the input size to LoFTR\n    img1_np = cv2.resize(img1_np, (640, 640))\n    img2_np = cv2.resize(img2_np, (640, 640))\n    \n    # Create a combined image to draw matches\n    h1, w1 = img1_np.shape[:2]\n    h2, w2 = img2_np.shape[:2]\n    combined_img = np.zeros((max(h1, h2), w1 + w2, 3), dtype=np.uint8)\n    combined_img[:h1, :w1] = img1_np\n    combined_img[:h2, w1:w1+w2] = img2_np\n    \n    # Draw matches\n    for (x1, y1), (x2, y2) in zip(mkpts0, mkpts1):\n        x2_shifted = x2 + w1  # Shift x-coordinate for the second image\n        cv2.circle(combined_img, (int(x1), int(y1)), 5, (0, 255, 0), 2)\n        cv2.circle(combined_img, (int(x2_shifted), int(y2)), 5, (0, 255, 0), 2)\n        cv2.line(combined_img, (int(x1), int(y1)), (int(x2_shifted), int(y2)), (255, 0, 0), 1)\n    \n    # Display the result\n    plt.figure(figsize=(15, 5))\n    plt.imshow(combined_img)\n    plt.title(f\"Keypoint Matches Between {img1_path.name} and {img2_path.name}\")\n    plt.axis('off')\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-06T15:13:24.692666Z","iopub.execute_input":"2025-04-06T15:13:24.693225Z","iopub.status.idle":"2025-04-06T15:13:30.967183Z","shell.execute_reply.started":"2025-04-06T15:13:24.693197Z","shell.execute_reply":"2025-04-06T15:13:30.965705Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Step 3.2: Adjust SfM Options for PyCOLMAP and Visualize 3D Point Cloud\n\nThe error occurred because the init_min_num_inliers attribute is not available in the current pycolmap IncrementalPipelineOptions. We’ll inspect the available attributes of IncrementalPipelineOptions to find the correct options for controlling the SfM pipeline. We’ll set reasonable defaults for the options, such as min_num_matches, and proceed with the COLMAP pipeline for cluster 0, running feature extraction, matching, and SfM to reconstruct the 3D scene. After reconstruction, we’ll extract the camera poses, print the number of images registered and 3D points reconstructed, and visualize the 3D point cloud to confirm the quality of the reconstruction. This visualization will help us assess the scene structure before processing other clusters.","metadata":{}},{"cell_type":"code","source":"# Install pycolmap if not already installed\n!pip install pycolmap\n\n# Step 3.2: Adjust SfM Options for PyCOLMAP and Visualize 3D Point Cloud\nimport pycolmap\nimport numpy as np\nfrom pathlib import Path\nfrom PIL import Image\nimport matplotlib.pyplot as plt\n\n# Create a directory for COLMAP workspace\ncolmap_dir = Path('/kaggle/working/colmap')\ncolmap_dir.mkdir(exist_ok=True)\nimage_dir = colmap_dir / 'images'\nimage_dir.mkdir(exist_ok=True)\n\n# Copy images from cluster 0 to the COLMAP image directory\ncluster_0_paths = [image_paths[i] for i in cluster_0_indices]\nfor i, img_path in enumerate(cluster_0_paths):\n    img = Image.open(img_path).convert('RGB')\n    img.save(image_dir / f\"image_{i}.png\")\n\n# Set up paths for COLMAP\ndatabase_path = colmap_dir / 'database.db'\noutput_path = colmap_dir / 'sparse'\noutput_path.mkdir(exist_ok=True)\n\n# Create IncrementalPipelineOptions object and set attributes\noptions = pycolmap.IncrementalPipelineOptions()\n\n# Inspect available attributes (for debugging)\nprint(\"Available attributes in IncrementalPipelineOptions:\")\nprint(dir(options))\n\n# Set available options (adjust based on available attributes)\noptions.min_num_matches = 15\noptions.ba_refine_focal_length = True\noptions.ba_refine_principal_point = False\noptions.ba_refine_extra_params = False\n# Instead of init_min_num_inliers, we can use min_model_size or similar if available\nif hasattr(options, 'min_model_size'):\n    options.min_model_size = 10  # Minimum number of inliers for initial model\nelse:\n    print(\"min_model_size not available, proceeding with default options.\")\n\n# Run the full COLMAP pipeline: feature extraction, matching, and SfM\ntry:\n    # Extract features\n    pycolmap.extract_features(\n        database_path=str(database_path),\n        image_path=str(image_dir),\n        camera_mode=pycolmap.CameraMode.AUTO\n    )\n    \n    # Perform exhaustive matching\n    pycolmap.match_exhaustive(str(database_path))\n    \n    # Run incremental mapping (SfM)\n    reconstructions = pycolmap.incremental_mapping(\n        database_path=str(database_path),\n        image_path=str(image_dir),\n        output_path=str(output_path),\n        options=options\n    )\n    \n    # Since incremental_mapping returns a dict of reconstructions, get the first one\n    if reconstructions:\n        reconstruction = list(reconstructions.values())[0]  # Take the first reconstruction\n        print(f\"Number of images registered: {reconstruction.num_reg_images()}\")\n        print(f\"Number of 3D points reconstructed: {reconstruction.num_points3D()}\")\n        \n        # Extract 3D points for visualization\n        points3d = []\n        for point3d_id in reconstruction.points3D:\n            point3d = reconstruction.points3D[point3d_id]\n            points3d.append(point3d.xyz)\n        points3d = np.array(points3d)\n        \n        # Visualize the 3D point cloud\n        fig = plt.figure(figsize=(10, 8))\n        ax = fig.add_subplot(111, projection='3d')\n        ax.scatter(points3d[:, 0], points3d[:, 1], points3d[:, 2], s=1, c='b', marker='o')\n        ax.set_title(\"3D Point Cloud of Cluster 0\")\n        ax.set_xlabel(\"X\")\n        ax.set_ylabel(\"Y\")\n        ax.set_zlabel(\"Z\")\n        plt.show()\n        \n        # Store camera poses for submission\n        camera_poses = {}\n        image_names = [f\"image_{i}.png\" for i in range(len(cluster_0_paths))]\n        for img_id in reconstruction.images:\n            img = reconstruction.images[img_id]\n            rotation_matrix = img.qvec2rotmat()\n            translation_vector = img.tvec\n            img_name = image_names[img_id-1]\n            camera_poses[img_name] = (rotation_matrix.flatten(), translation_vector)\n    else:\n        print(\"Reconstruction failed: No reconstructions returned.\")\nexcept Exception as e:\n    print(f\"Reconstruction failed: {e}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-06T15:13:30.968128Z","iopub.execute_input":"2025-04-06T15:13:30.968415Z","iopub.status.idle":"2025-04-06T15:16:21.860586Z","shell.execute_reply.started":"2025-04-06T15:13:30.968390Z","shell.execute_reply":"2025-04-06T15:16:21.859517Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Step 3.3: Use COLMAP’s High-Level Pipeline with Relaxed Constraints and Visualize 3D Reconstruction\n\nThe error occurred because the add_camera method is not available in the current pycolmap API, indicating that manual database management is not supported. We’ll revert to using pycolmap’s high-level pipeline functions (extract_features, match_exhaustive, and incremental_mapping), which handle camera calibration, feature extraction, and matching internally. Our previous attempts only registered 2 out of 51 images in cluster 0, likely due to strict constraints in the SfM pipeline. We’ll relax these constraints further by lowering min_num_matches and min_model_size, and increasing init_num_trials to allow more attempts at initializing the reconstruction. After reconstructing the scene, we’ll extract the camera poses using the correct method (rotmat() for the rotation matrix and tvec for the translation vector), print the number of images registered and 3D points reconstructed, and visualize the 3D point cloud along with the camera positions in a 3D plot to assess the scene structure.","metadata":{}},{"cell_type":"code","source":"# Step 3.3: Use COLMAP’s High-Level Pipeline with Relaxed Constraints and Visualize 3D Reconstruction\nimport pycolmap\n\n# Create a directory for COLMAP workspace\ncolmap_dir = Path('/kaggle/working/colmap')\ncolmap_dir.mkdir(exist_ok=True)\nimage_dir = colmap_dir / 'images'\nimage_dir.mkdir(exist_ok=True)\n\n# Copy images from cluster 0 to the COLMAP image directory\ncluster_0_paths = [image_paths[i] for i in cluster_0_indices]\nfor i, img_path in enumerate(cluster_0_paths):\n    img = Image.open(img_path).convert('RGB')\n    img.save(image_dir / f\"image_{i}.png\")\n\n# Set up paths for COLMAP\ndatabase_path = colmap_dir / 'database.db'\noutput_path = colmap_dir / 'sparse'\noutput_path.mkdir(exist_ok=True)\n\n# Create IncrementalPipelineOptions object and set attributes\noptions = pycolmap.IncrementalPipelineOptions()\noptions.min_num_matches = 5    # Lowered to allow more matches\noptions.min_model_size = 3     # Lowered to allow smaller initial models\noptions.init_num_trials = 2000  # Increased to allow more attempts at initialization\noptions.ba_refine_focal_length = True\noptions.ba_refine_principal_point = False\noptions.ba_refine_extra_params = False\n\n# Run the full COLMAP pipeline: feature extraction, matching, and SfM\ntry:\n    # Extract features\n    pycolmap.extract_features(\n        database_path=str(database_path),\n        image_path=str(image_dir),\n        camera_mode=pycolmap.CameraMode.AUTO\n    )\n    \n    # Perform exhaustive matching\n    pycolmap.match_exhaustive(str(database_path))\n    \n    # Run incremental mapping (SfM)\n    reconstructions = pycolmap.incremental_mapping(\n        database_path=str(database_path),\n        image_path=str(image_dir),\n        output_path=str(output_path),\n        options=options\n    )\n    \n    # Since incremental_mapping returns a dict of reconstructions, get the first one\n    if reconstructions:\n        reconstruction = list(reconstructions.values())[0]  # Take the first reconstruction\n        print(f\"Number of images registered: {reconstruction.num_reg_images()}\")\n        print(f\"Number of 3D points reconstructed: {reconstruction.num_points3D()}\")\n        \n        # Extract 3D points for visualization\n        points3d = []\n        for point3d_id in reconstruction.points3D:\n            point3d = reconstruction.points3D[point3d_id]\n            points3d.append(point3d.xyz)\n        points3d = np.array(points3d)\n        \n        # Extract camera positions for visualization\n        camera_positions = []\n        for img_id in reconstruction.images:\n            img = reconstruction.images[img_id]\n            # Compute camera position as -R^T * t\n            rotation = img.rotmat()  # Use rotmat() to get the rotation matrix\n            translation = img.tvec\n            position = -np.dot(rotation.T, translation)\n            camera_positions.append(position)\n        camera_positions = np.array(camera_positions)\n        \n        # Visualize the 3D point cloud and camera positions\n        fig = plt.figure(figsize=(10, 8))\n        ax = fig.add_subplot(111, projection='3d')\n        # Plot 3D points\n        ax.scatter(points3d[:, 0], points3d[:, 1], points3d[:, 2], s=1, c='b', marker='o', label='3D Points')\n        # Plot camera positions\n        ax.scatter(camera_positions[:, 0], camera_positions[:, 1], camera_positions[:, 2], s=50, c='r', marker='^', label='Cameras')\n        ax.set_title(\"3D Point Cloud and Camera Positions of Cluster 0\")\n        ax.set_xlabel(\"X\")\n        ax.set_ylabel(\"Y\")\n        ax.set_zlabel(\"Z\")\n        ax.legend()\n        plt.show()\n        \n        # Store camera poses for submission\n        camera_poses = {}\n        image_names = [f\"image_{i}.png\" for i in range(len(cluster_0_paths))]\n        for img_id in reconstruction.images:\n            img = reconstruction.images[img_id]\n            rotation_matrix = img.rotmat()  # Use rotmat() to get the rotation matrix\n            translation_vector = img.tvec\n            img_name = image_names[img_id-1]\n            camera_poses[img_name] = (rotation_matrix.flatten(), translation_vector)\n    else:\n        print(\"Reconstruction failed: No reconstructions returned.\")\nexcept Exception as e:\n    print(f\"Reconstruction failed: {e}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-06T15:16:21.861717Z","iopub.execute_input":"2025-04-06T15:16:21.861971Z","iopub.status.idle":"2025-04-06T15:17:03.481749Z","shell.execute_reply.started":"2025-04-06T15:16:21.861947Z","shell.execute_reply":"2025-04-06T15:17:03.480764Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Step 3.4: Adjust Feature Extraction, Fix Camera Pose Extraction, and Visualize 3D Reconstruction\n\nThe error occurred because the rotmat() method is not available in the current pycolmap API. We’ll fix this by accessing the rotation matrix using the rotation attribute of the Image object and the translation vector via the tvec attribute. The low number of registered images (2 out of 51) indicates that COLMAP’s default feature extraction and matching settings are not sufficient for this dataset. We’ll adjust the feature extraction settings by increasing the number of features extracted (peak_threshold and max_num_features) to improve matching. We’ll also keep the relaxed SfM constraints (min_num_matches, min_model_size, and init_num_trials) to encourage more images to be registered. After reconstructing the scene, we’ll extract the camera poses, print the number of images registered and 3D points reconstructed, and visualize the 3D point cloud along with the camera positions in a 3D plot to assess the scene structure.","metadata":{}},{"cell_type":"code","source":"# Step 3.4: Fix Feature Extraction Argument and Visualize 3D Reconstruction\nimport pycolmap\n\n# Create a directory for COLMAP workspace\ncolmap_dir = Path('/kaggle/working/colmap')\ncolmap_dir.mkdir(exist_ok=True)\nimage_dir = colmap_dir / 'images'\nimage_dir.mkdir(exist_ok=True)\n\n# Copy images from cluster 0 to the COLMAP image directory\ncluster_0_paths = [image_paths[i] for i in cluster_0_indices]\nfor i, img_path in enumerate(cluster_0_paths):\n    img = Image.open(img_path).convert('RGB')\n    img.save(image_dir / f\"image_{i}.png\")\n\n# Set up paths for COLMAP\ndatabase_path = colmap_dir / 'database.db'\noutput_path = colmap_dir / 'sparse'\noutput_path.mkdir(exist_ok=True)\n\n# Create SiftExtractionOptions to adjust feature extraction\nsift_options = pycolmap.SiftExtractionOptions()\nsift_options.peak_threshold = 0.01  # Lower threshold to extract more features\nsift_options.max_num_features = 8192  # Increase the maximum number of features\n\n# Create IncrementalPipelineOptions object and set attributes\noptions = pycolmap.IncrementalPipelineOptions()\noptions.min_num_matches = 5    # Lowered to allow more matches\noptions.min_model_size = 3     # Lowered to allow smaller initial models\noptions.init_num_trials = 2000  # Increased to allow more attempts at initialization\noptions.ba_refine_focal_length = True\noptions.ba_refine_principal_point = False\noptions.ba_refine_extra_params = False\n\n# Run the full COLMAP pipeline: feature extraction, matching, and SfM\ntry:\n    # Extract features with adjusted options\n    pycolmap.extract_features(\n        database_path=str(database_path),\n        image_path=str(image_dir),\n        camera_mode=pycolmap.CameraMode.AUTO,\n        sift_options=sift_options  # Fixed argument name\n    )\n    \n    # Perform exhaustive matching\n    pycolmap.match_exhaustive(str(database_path))\n    \n    # Run incremental mapping (SfM)\n    reconstructions = pycolmap.incremental_mapping(\n        database_path=str(database_path),\n        image_path=str(image_dir),\n        output_path=str(output_path),\n        options=options\n    )\n    \n    # Since incremental_mapping returns a dict of reconstructions, get the first one\n    if reconstructions:\n        reconstruction = list(reconstructions.values())[0]  # Take the first reconstruction\n        print(f\"Number of images registered: {reconstruction.num_reg_images()}\")\n        print(f\"Number of 3D points reconstructed: {reconstruction.num_points3D()}\")\n        \n        # Extract 3D points for visualization\n        points3d = []\n        for point3d_id in reconstruction.points3D:\n            point3d = reconstruction.points3D[point3d_id]\n            points3d.append(point3d.xyz)\n        points3d = np.array(points3d)\n        \n        # Extract camera positions for visualization\n        camera_positions = []\n        for img_id in reconstruction.images:\n            img = reconstruction.images[img_id]\n            # Compute camera position as -R^T * t\n            rotation = img.rotation  # Use rotation attribute to get the rotation matrix\n            translation = img.tvec\n            position = -np.dot(rotation.T, translation)\n            camera_positions.append(position)\n        camera_positions = np.array(camera_positions)\n        \n        # Visualize the 3D point cloud and camera positions\n        fig = plt.figure(figsize=(10, 8))\n        ax = fig.add_subplot(111, projection='3d')\n        # Plot 3D points\n        ax.scatter(points3d[:, 0], points3d[:, 1], points3d[:, 2], s=1, c='b', marker='o', label='3D Points')\n        # Plot camera positions\n        ax.scatter(camera_positions[:, 0], camera_positions[:, 1], camera_positions[:, 2], s=50, c='r', marker='^', label='Cameras')\n        ax.set_title(\"3D Point Cloud and Camera Positions of Cluster 0\")\n        ax.set_xlabel(\"X\")\n        ax.set_ylabel(\"Y\")\n        ax.set_zlabel(\"Z\")\n        ax.legend()\n        plt.show()\n        \n        # Store camera poses for submission\n        camera_poses = {}\n        image_names = [f\"image_{i}.png\" for i in range(len(cluster_0_paths))]\n        for img_id in reconstruction.images:\n            img = reconstruction.images[img_id]\n            rotation_matrix = img.rotation  # Use rotation attribute to get the rotation matrix\n            translation_vector = img.tvec\n            img_name = image_names[img_id-1]\n            camera_poses[img_name] = (rotation_matrix.flatten(), translation_vector)\n    else:\n        print(\"Reconstruction failed: No reconstructions returned.\")\nexcept Exception as e:\n    print(f\"Reconstruction failed: {e}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-06T15:17:03.482549Z","iopub.execute_input":"2025-04-06T15:17:03.482777Z","iopub.status.idle":"2025-04-06T15:17:45.339766Z","shell.execute_reply.started":"2025-04-06T15:17:03.482753Z","shell.execute_reply":"2025-04-06T15:17:45.338801Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Step 3.5: Fix Matching Argument, Relax Matching Constraints, and Visualize 3D Reconstruction\n\nThe error occurred because the match_exhaustive function expects the SIFT matching options to be passed as sift_options, not sift_matching_options. We’ll fix this by renaming the argument. From the available attributes in SiftMatchingOptions, we’ll lower max_ratio to 0.9 (from the default 0.8) and increase max_distance to 0.8 (from the default 0.7) to allow more matches. We’ll also use TwoViewGeometryOptions to relax the geometric verification constraints by lowering the min_num_inliers for a pair to be considered valid. We’ll continue working with the first 10 images of cluster 0 to speed up debugging. After reconstructing the scene, we’ll extract the camera poses using the rotation_matrix() method for the rotation matrix and tvec for the translation vector, print the number of images registered and 3D points reconstructed, and visualize the 3D point cloud along with the camera positions in a 3D plot to assess the scene structure.","metadata":{}},{"cell_type":"code","source":"import pycolmap\nimport numpy as np\nimport matplotlib.pyplot as plt\nfrom PIL import Image\nfrom pathlib import Path\n\n# Define function to convert quaternion to rotation matrix\ndef qvec_to_rotmat(qvec):\n    \"\"\"Convert quaternion to rotation matrix.\"\"\"\n    q0, q1, q2, q3 = qvec\n    return np.array([\n        [1 - 2*q2**2 - 2*q3**2, 2*q1*q2 - 2*q0*q3,     2*q1*q3 + 2*q0*q2],\n        [2*q1*q2 + 2*q0*q3,     1 - 2*q1**2 - 2*q3**2, 2*q2*q3 - 2*q0*q1],\n        [2*q1*q3 - 2*q0*q2,     2*q2*q3 + 2*q0*q1,     1 - 2*q1**2 - 2*q2**2]\n    ])\n\n# Step 3.5: Fix Matching Argument, Relax Matching Constraints, and Visualize 3D Reconstruction\ncolmap_dir = Path('/kaggle/working/colmap')\ncolmap_dir.mkdir(exist_ok=True)\nimage_dir = colmap_dir / 'images'\nimage_dir.mkdir(exist_ok=True)\n\n# Copy a subset of images from cluster 0 to the COLMAP image directory (to speed up debugging)\ncluster_0_paths = [image_paths[i] for i in cluster_0_indices[:10]]  # Use only the first 10 images\nfor i, img_path in enumerate(cluster_0_paths):\n    img = Image.open(img_path).convert('RGB')\n    img.save(image_dir / f\"image_{i}.png\")\n\n# Set up paths for COLMAP\ndatabase_path = colmap_dir / 'database.db'\noutput_path = colmap_dir / 'sparse'\noutput_path.mkdir(exist_ok=True)\n\n# Create SiftExtractionOptions to adjust feature extraction\nsift_extraction_options = pycolmap.SiftExtractionOptions()\nsift_extraction_options.peak_threshold = 0.01  # Lower threshold to extract more features\nsift_extraction_options.max_num_features = 8192  # Increase the maximum number of features\n\n# Create SiftMatchingOptions and set attributes to relax matching constraints\nsift_options = pycolmap.SiftMatchingOptions()\nsift_options.max_ratio = 0.9  # Relaxed to allow more matches (default is 0.8)\nsift_options.max_distance = 0.8  # Relaxed to allow more matches (default is 0.7)\nsift_options.cross_check = True  # Keep cross-check enabled for robustness\n\n# Create TwoViewGeometryOptions to relax geometric verification\nverification_options = pycolmap.TwoViewGeometryOptions()\nverification_options.min_num_inliers = 5  # Lowered to allow more pairs to be considered valid\n\n# Create IncrementalPipelineOptions object and set attributes\noptions = pycolmap.IncrementalPipelineOptions()\noptions.min_num_matches = 2    # Further lowered to allow more matches\noptions.min_model_size = 2     # Further lowered to allow smaller initial models\noptions.init_num_trials = 10000  # Further increased to allow more attempts at initialization\noptions.ba_refine_focal_length = True\noptions.ba_refine_principal_point = False\noptions.ba_refine_extra_params = False\n\n# Run the full COLMAP pipeline: feature extraction, matching, and SfM\ntry:\n    # Extract features with adjusted options\n    pycolmap.extract_features(\n        database_path=str(database_path),\n        image_path=str(image_dir),\n        camera_mode=pycolmap.CameraMode.AUTO,\n        sift_options=sift_extraction_options\n    )\n    \n    # Perform exhaustive matching with adjusted options\n    pycolmap.match_exhaustive(\n        database_path=str(database_path),\n        sift_options=sift_options,\n        verification_options=verification_options\n    )\n    \n    # Run incremental mapping (SfM)\n    reconstructions = pycolmap.incremental_mapping(\n        database_path=str(database_path),\n        image_path=str(image_dir),\n        output_path=str(output_path),\n        options=options\n    )\n    \n    # Since incremental_mapping returns a dict of reconstructions, get the first one\n    if reconstructions:\n        reconstruction = list(reconstructions.values())[0]  # Take the first reconstruction\n        print(f\"Number of images registered: {reconstruction.num_reg_images()}\")\n        print(f\"Number of 3D points reconstructed: {reconstruction.num_points3D()}\")\n        \n        # Extract 3D points for visualization\n        points3d = []\n        for point3d_id in reconstruction.points3D:\n            point3d = reconstruction.points3D[point3d_id]\n            points3d.append(point3d.xyz)\n        points3d = np.array(points3d)\n        \n        # Extract camera positions for visualization\n        camera_positions = []\n        for img_id in reconstruction.images:\n            img = reconstruction.images[img_id]\n            # Compute camera position as -R^T * t\n            rotation = qvec_to_rotmat(img.qvec)  # Use qvec_to_rotmat() to get the rotation matrix\n            translation = img.tvec\n            position = -np.dot(rotation.T, translation)\n            camera_positions.append(position)\n        camera_positions = np.array(camera_positions)\n        \n        # Visualize the 3D point cloud and camera positions\n        fig = plt.figure(figsize=(10, 8))\n        ax = fig.add_subplot(111, projection='3d')\n        # Plot 3D points\n        ax.scatter(points3d[:, 0], points3d[:, 1], points3d[:, 2], s=1, c='b', marker='o', label='3D Points')\n        # Plot camera positions\n        ax.scatter(camera_positions[:, 0], camera_positions[:, 1], camera_positions[:, 2], s=50, c='r', marker='^', label='Cameras')\n        ax.set_title(\"3D Point Cloud and Camera Positions of Cluster 0 (First 10 Images)\")\n        ax.set_xlabel(\"X\")\n        ax.set_ylabel(\"Y\")\n        ax.set_zlabel(\"Z\")\n        ax.legend()\n        plt.show()\n        \n        # Store camera poses for submission\n        camera_poses = {}\n        image_names = [f\"image_{i}.png\" for i in range(len(cluster_0_paths))]\n        for img_id in reconstruction.images:\n            img = reconstruction.images[img_id]\n            rotation_matrix = qvec_to_rotmat(img.qvec)  # Use qvec_to_rotmat() to get the rotation matrix\n            translation_vector = img.tvec\n            img_name = image_names[img_id-1]\n            camera_poses[img_name] = (rotation_matrix.flatten(), translation_vector)\n    else:\n        print(\"Reconstruction failed: No reconstructions returned.\")\nexcept Exception as e:\n    print(f\"Reconstruction failed: {e}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-06T15:17:45.340593Z","iopub.execute_input":"2025-04-06T15:17:45.340851Z","iopub.status.idle":"2025-04-06T15:18:50.241297Z","shell.execute_reply.started":"2025-04-06T15:17:45.340833Z","shell.execute_reply":"2025-04-06T15:18:50.240232Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Step 3.6: Debug Image Attributes, Fix Camera Pose Extraction, and Visualize 3D Reconstruction\n\nThe error occurred because the rotation_matrix() method is not available in the current pycolmap API. To fix this, we’ll print the available attributes of the Image object to identify the correct way to access the rotation matrix and translation vector. Based on the pycolmap API, the rotation is likely stored as a quaternion in the qvec attribute, which we can convert to a rotation matrix using a utility function, and the translation vector should be accessible via the tvec attribute. We’ll continue working with the first 10 images of cluster 0 to keep the runs fast for debugging. After reconstructing the scene, we’ll extract the camera poses, print the number of images registered and 3D points reconstructed, and visualize the 3D point cloud along with the camera positions in a 3D plot to assess the scene structure. Once we resolve the camera pose extraction, we’ll revisit the issue of low registration (only 2 out of 10 images).","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport matplotlib.pyplot as plt\nfrom pathlib import Path\nfrom PIL import Image\nimport pycolmap\n\n# === Setup COLMAP workspace ===\ncolmap_dir = Path('/kaggle/working/colmap')\ncolmap_dir.mkdir(exist_ok=True)\nimage_dir = colmap_dir / 'images'\nimage_dir.mkdir(exist_ok=True)\n\n# === Select and copy a subset of cluster 0 images ===\ncluster_0_paths = [image_paths[i] for i in cluster_0_indices[:10]]  # Use only first 10 images\nfor i, img_path in enumerate(cluster_0_paths):\n    img = Image.open(img_path).convert('RGB')\n    img.save(image_dir / f\"image_{i}.png\")\n\n# === Define database and output paths ===\ndatabase_path = colmap_dir / 'database.db'\noutput_path = colmap_dir / 'sparse'\noutput_path.mkdir(exist_ok=True)\n\n# === Feature extraction options ===\nsift_extraction_options = pycolmap.SiftExtractionOptions()\nsift_extraction_options.peak_threshold = 0.01\nsift_extraction_options.max_num_features = 8192\n\n# === Feature matching options ===\nsift_options = pycolmap.SiftMatchingOptions()\nsift_options.max_ratio = 0.9\nsift_options.max_distance = 0.8\nsift_options.cross_check = True\n\n# === Geometric verification options ===\nverification_options = pycolmap.TwoViewGeometryOptions()\nverification_options.min_num_inliers = 5\n\n# === Incremental pipeline options ===\noptions = pycolmap.IncrementalPipelineOptions()\noptions.min_num_matches = 2\noptions.min_model_size = 2\noptions.init_num_trials = 10000\noptions.ba_refine_focal_length = True\noptions.ba_refine_principal_point = False\noptions.ba_refine_extra_params = False\n\n# === Run the COLMAP SfM pipeline ===\ntry:\n    # Step 1: Extract features\n    pycolmap.extract_features(\n        database_path=str(database_path),\n        image_path=str(image_dir),\n        camera_mode=pycolmap.CameraMode.AUTO,\n        sift_options=sift_extraction_options\n    )\n\n    # Step 2: Match features\n    pycolmap.match_exhaustive(\n        database_path=str(database_path),\n        sift_options=sift_options,\n        verification_options=verification_options\n    )\n\n    # Step 3: Run SfM (incremental mapping)\n    reconstructions = pycolmap.incremental_mapping(\n        database_path=str(database_path),\n        image_path=str(image_dir),\n        output_path=str(output_path),\n        options=options\n    )\n\n    if reconstructions:\n        reconstruction = list(reconstructions.values())[0]\n        print(f\"Number of images registered: {reconstruction.num_reg_images()}\")\n        print(f\"Number of 3D points reconstructed: {reconstruction.num_points3D()}\")\n\n        # === Debug one camera pose ===\n        for img_id in reconstruction.images:\n            img = reconstruction.images[img_id]\n            cam_from_world = img.cam_from_world\n            rotation = cam_from_world.rotation\n            print(f\"Available attributes in Rotation3d object:\")\n            print(dir(rotation))\n            translation = cam_from_world.translation\n            print(f\"Type of translation: {type(translation)}\")\n            break  # Only debug the first image\n\n        # === Extract 3D points for visualization ===\n        points3d = []\n        for point3d_id in reconstruction.points3D:\n            point3d = reconstruction.points3D[point3d_id]\n            points3d.append(point3d.xyz)\n        points3d = np.array(points3d)\n\n        # === Extract camera positions ===\n        camera_positions = []\n        for img_id in reconstruction.images:\n            img = reconstruction.images[img_id]\n            cam_from_world = img.cam_from_world\n            rotation = cam_from_world.rotation.matrix()\n            translation = cam_from_world.translation\n            position = -np.dot(rotation.T, translation)\n            camera_positions.append(position)\n        camera_positions = np.array(camera_positions)\n\n        # === Visualize the 3D point cloud and camera positions ===\n        fig = plt.figure(figsize=(10, 8))\n        ax = fig.add_subplot(111, projection='3d')\n        ax.scatter(points3d[:, 0], points3d[:, 1], points3d[:, 2], s=1, c='b', label='3D Points')\n        ax.scatter(camera_positions[:, 0], camera_positions[:, 1], camera_positions[:, 2], s=50, c='r', marker='^', label='Cameras')\n        ax.set_title(\"3D Point Cloud and Camera Positions (Cluster 0 - First 10 Images)\")\n        ax.set_xlabel(\"X\")\n        ax.set_ylabel(\"Y\")\n        ax.set_zlabel(\"Z\")\n        ax.legend()\n        plt.show()\n\n        # === Store camera poses ===\n        camera_poses = {}\n        for img_id in reconstruction.images:\n            img = reconstruction.images[img_id]\n            cam_from_world = img.cam_from_world\n            rotation_matrix = cam_from_world.rotation.matrix()\n            translation_vector = cam_from_world.translation\n            img_name = img.name  # ✅ Use image name from pycolmap\n            camera_poses[img_name] = (rotation_matrix.flatten(), translation_vector)\n\n    else:\n        print(\"Reconstruction failed: No reconstructions returned.\")\n\nexcept Exception as e:\n    print(f\"Reconstruction failed: {e}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-06T15:18:50.242200Z","iopub.execute_input":"2025-04-06T15:18:50.242503Z","iopub.status.idle":"2025-04-06T15:19:55.936435Z","shell.execute_reply.started":"2025-04-06T15:18:50.242470Z","shell.execute_reply":"2025-04-06T15:19:55.935516Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Step 3.7: Visualize Images and Switch to Sequential Matching to Improve Registration\n\nThe SfM pipeline is now working correctly for camera pose extraction and visualization, but only 2 out of 10 images are being registered. To understand why, we’ll first visualize the first 10 images of cluster 0 to check for overlap, image quality, and potential issues like blur or lighting variations. Then, we’ll modify the pipeline to use sequential matching instead of exhaustive matching, assuming the images might be ordered in a sequence (e.g., taken as a video or panorama). Sequential matching matches consecutive pairs of images, which can be more effective if the images have temporal or spatial continuity. We’ll also lower the min_num_inliers in TwoViewGeometryOptions to 3 to allow more pairs to be considered valid. After reconstructing the scene, we’ll print the number of images registered and 3D points reconstructed, and visualize the 3D point cloud along with the camera positions in a 3D plot to assess the scene structure.","metadata":{}},{"cell_type":"code","source":"# Step 3.7: Fix Image Visualization and Rerun SfM with Sequential Matching\nimport pycolmap\nimport numpy as np\nimport matplotlib.pyplot as plt\nfrom pathlib import Path\nfrom PIL import Image\n\n# === Setup COLMAP workspace ===\ncolmap_dir = Path('/kaggle/working/colmap')\ncolmap_dir.mkdir(exist_ok=True)\nimage_dir = colmap_dir / 'images'\nimage_dir.mkdir(exist_ok=True)\n\n# === Select and copy a subset of cluster 0 images ===\ncluster_0_paths = [image_paths[i] for i in cluster_0_indices[:10]]  # Use only first 10 images\nfor i, img_path in enumerate(cluster_0_paths):\n    img = Image.open(img_path).convert('RGB')\n    saved_path = image_dir / f\"image_{i}.png\"\n    img.save(saved_path)\n    print(f\"Saved image {i} to: {saved_path}\")\n\n# === Visualize the first 10 images ===\nfig, axes = plt.subplots(2, 5, figsize=(20, 8))\naxes = axes.flatten()\nfor i in range(10):\n    img_path = image_dir / f\"image_{i}.png\"\n    try:\n        img = Image.open(img_path).convert('RGB')\n        img_array = np.array(img)  # Convert PIL image to NumPy array for matplotlib\n        axes[i].imshow(img_array)\n        axes[i].set_title(f\"Image {i}\")\n        axes[i].axis('off')\n    except Exception as e:\n        print(f\"Failed to load image {i}: {e}\")\n        axes[i].set_title(f\"Image {i} (Failed)\")\n        axes[i].axis('off')\nplt.tight_layout()\nplt.show()\n\n# === Define database and output paths ===\ndatabase_path = colmap_dir / 'database.db'\noutput_path = colmap_dir / 'sparse'\noutput_path.mkdir(exist_ok=True)\n\n# === Feature extraction options ===\nsift_extraction_options = pycolmap.SiftExtractionOptions()\nsift_extraction_options.peak_threshold = 0.01\nsift_extraction_options.max_num_features = 8192\n\n# === Feature matching options ===\nsift_options = pycolmap.SiftMatchingOptions()\nsift_options.max_ratio = 0.9\nsift_options.max_distance = 0.8\nsift_options.cross_check = True\n\n# === Geometric verification options ===\nverification_options = pycolmap.TwoViewGeometryOptions()\nverification_options.min_num_inliers = 3  # Further lowered to allow more pairs\n\n# === Incremental pipeline options ===\noptions = pycolmap.IncrementalPipelineOptions()\noptions.min_num_matches = 2\noptions.min_model_size = 2\noptions.init_num_trials = 10000\noptions.ba_refine_focal_length = True\noptions.ba_refine_principal_point = False\noptions.ba_refine_extra_params = False\n\n# === Run the COLMAP SfM pipeline ===\ntry:\n    # Step 1: Extract features\n    pycolmap.extract_features(\n        database_path=str(database_path),\n        image_path=str(image_dir),\n        camera_mode=pycolmap.CameraMode.AUTO,\n        sift_options=sift_extraction_options\n    )\n\n    # Step 2: Match features using sequential matching\n    pycolmap.match_sequential(\n        database_path=str(database_path),\n        sift_options=sift_options,\n        verification_options=verification_options\n    )\n\n    # Step 3: Run SfM (incremental mapping)\n    reconstructions = pycolmap.incremental_mapping(\n        database_path=str(database_path),\n        image_path=str(image_dir),\n        output_path=str(output_path),\n        options=options\n    )\n\n    if reconstructions:\n        reconstruction = list(reconstructions.values())[0]\n        print(f\"Number of images registered: {reconstruction.num_reg_images()}\")\n        print(f\"Number of 3D points reconstructed: {reconstruction.num_points3D()}\")\n\n        # === Extract 3D points for visualization ===\n        points3d = []\n        for point3d_id in reconstruction.points3D:\n            point3d = reconstruction.points3D[point3d_id]\n            points3d.append(point3d.xyz)\n        points3d = np.array(points3d)\n\n        # === Extract camera positions ===\n        camera_positions = []\n        for img_id in reconstruction.images:\n            img = reconstruction.images[img_id]\n            cam_from_world = img.cam_from_world\n            rotation = cam_from_world.rotation.matrix()\n            translation = cam_from_world.translation\n            position = -np.dot(rotation.T, translation)\n            camera_positions.append(position)\n        camera_positions = np.array(camera_positions)\n\n        # === Visualize the 3D point cloud and camera positions ===\n        fig = plt.figure(figsize=(10, 8))\n        ax = fig.add_subplot(111, projection='3d')\n        ax.scatter(points3d[:, 0], points3d[:, 1], points3d[:, 2], s=1, c='b', label='3D Points')\n        ax.scatter(camera_positions[:, 0], camera_positions[:, 1], camera_positions[:, 2], s=50, c='r', marker='^', label='Cameras')\n        ax.set_title(\"3D Point Cloud and Camera Positions (Cluster 0 - First 10 Images)\")\n        ax.set_xlabel(\"X\")\n        ax.set_ylabel(\"Y\")\n        ax.set_zlabel(\"Z\")\n        ax.legend()\n        plt.show()\n\n        # === Store camera poses ===\n        camera_poses = {}\n        for img_id in reconstruction.images:\n            img = reconstruction.images[img_id]\n            cam_from_world = img.cam_from_world\n            rotation_matrix = cam_from_world.rotation.matrix()\n            translation_vector = cam_from_world.translation\n            img_name = img.name\n            camera_poses[img_name] = (rotation_matrix.flatten(), translation_vector)\n\n    else:\n        print(\"Reconstruction failed: No reconstructions returned.\")\n\nexcept Exception as e:\n    print(f\"Reconstruction failed: {e}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-06T15:19:55.937311Z","iopub.execute_input":"2025-04-06T15:19:55.937559Z","iopub.status.idle":"2025-04-06T15:20:51.659287Z","shell.execute_reply.started":"2025-04-06T15:19:55.937541Z","shell.execute_reply":"2025-04-06T15:20:51.658311Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Step 3.7: Preprocess Images, Use Exhaustive Matching, and Relax Constraints to Improve Registration\n\n## Explanation\n\nThe SfM pipeline is only registering 2 out of 10 images due to limited overlap, low texture, and lighting variations. To address this, we’ll:\n\n1. **Switch to exhaustive matching:** Since the images are not strictly sequential, we’ll use match_exhaustive to try all possible pairs, which should increase the chances of finding matches.\n2. **Relax constraints further:** Lower the min_num_inliers in TwoViewGeometryOptions to 2 and reduce min_num_matches and min_model_size in\n3. IncrementalPipelineOptions to 1 to allow more images to be registered. After reconstructing the scene, we’ll print the number of images registered and 3D points reconstructed, and visualize the 3D point cloud along with the camera positions in a 3D plot to assess the scene structure.","metadata":{}},{"cell_type":"code","source":"!pip install torch torchvision\n!pip install opencv-python","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-06T15:20:51.660263Z","iopub.execute_input":"2025-04-06T15:20:51.660531Z","iopub.status.idle":"2025-04-06T15:20:58.553364Z","shell.execute_reply.started":"2025-04-06T15:20:51.660509Z","shell.execute_reply":"2025-04-06T15:20:58.552192Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"!git clone https://github.com/magicleap/SuperGluePretrainedNetwork.git\n!wget https://github.com/magicleap/SuperGluePretrainedNetwork/raw/master/models/weights/superglue_indoor.pth -P SuperGluePretrainedNetwork/models/weights/","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-06T15:20:58.554602Z","iopub.execute_input":"2025-04-06T15:20:58.554951Z","iopub.status.idle":"2025-04-06T15:21:09.490212Z","shell.execute_reply.started":"2025-04-06T15:20:58.554915Z","shell.execute_reply":"2025-04-06T15:21:09.488978Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import sys\nsys.path.append('SuperGluePretrainedNetwork')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-06T15:21:09.491440Z","iopub.execute_input":"2025-04-06T15:21:09.491796Z","iopub.status.idle":"2025-04-06T15:21:09.497159Z","shell.execute_reply.started":"2025-04-06T15:21:09.491759Z","shell.execute_reply":"2025-04-06T15:21:09.495565Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Step 3.9: Use SuperPoint and SuperGlue with COLMAP Command-Line Integration (Retry)\nimport pycolmap\nimport numpy as np\nimport matplotlib.pyplot as plt\nfrom pathlib import Path\nfrom PIL import Image\nimport cv2\nimport time\nimport os\nimport torch\nimport sys\nimport subprocess\nimport warnings\nwarnings.filterwarnings(\"ignore\", category=FutureWarning)\n\nsys.path.append('SuperGluePretrainedNetwork')\nfrom models.matching import Matching\nfrom models.utils import (frame2tensor, make_matching_plot)\n\n# === Install COLMAP in Kaggle environment ===\nprint(\"Installing COLMAP dependencies...\")\ntry:\n    # Install required packages\n    subprocess.run(\"apt-get update\", shell=True, check=True)\n    subprocess.run(\"apt-get install -y build-essential cmake libboost-all-dev libeigen3-dev libceres-dev libfreeimage-dev libmetis-dev libgoogle-glog-dev libgflags-dev libsqlite3-dev libglew-dev qtbase5-dev libqt5opengl5-dev libcgal-dev\", shell=True, check=True)\n    \n    # Download and install COLMAP\n    print(\"Downloading and installing COLMAP...\")\n    subprocess.run(\"git clone https://github.com/colmap/colmap.git /kaggle/working/colmap-repo\", shell=True, check=True)\n    subprocess.run(\"mkdir -p /kaggle/working/colmap-repo/build\", shell=True, check=True)\n    subprocess.run(\"cmake .. -DCMAKE_BUILD_TYPE=Release\", shell=True, check=True, cwd=\"/kaggle/working/colmap-repo/build\")\n    subprocess.run(\"make -j$(nproc)\", shell=True, check=True, cwd=\"/kaggle/working/colmap-repo/build\")\n    subprocess.run(\"make install\", shell=True, check=True, cwd=\"/kaggle/working/colmap-repo/build\")\n    \n    # Verify installation\n    result = subprocess.run(\"colmap --version\", shell=True, capture_output=True, text=True)\n    print(\"COLMAP version:\", result.stdout)\n    if result.returncode != 0:\n        raise Exception(\"COLMAP installation failed.\")\nexcept Exception as e:\n    print(f\"Error installing COLMAP: {e}\")\n    print(\"Falling back to default COLMAP installation method...\")\n    subprocess.run(\"apt-get install -y colmap\", shell=True, check=True)\n    result = subprocess.run(\"colmap --version\", shell=True, capture_output=True, text=True)\n    print(\"COLMAP version:\", result.stdout)\n\n# === Set device ===\ndevice = torch.device('cuda' if torch.cuda.is_available() else 'cpu')\nprint(f\"Using device: {device}\")\n\n# === Define image paths ===\nstairs_dir = Path('/kaggle/input/image-matching-challenge-2025/train/stairs')\nimage_paths = [str(stairs_dir / img) for img in os.listdir(stairs_dir) if img.endswith(('.png', '.jpg', '.jpeg'))]\nprint(f\"Total images in stairs directory: {len(image_paths)}\")\nprint(\"First few image paths:\")\nfor path in image_paths[:5]:\n    print(path)\n\n# === Define cluster_0_indices (select first 5 images) ===\ncluster_0_indices = list(range(min(5, len(image_paths))))\nprint(f\"Cluster 0 indices: {cluster_0_indices}\")\n\n# === Setup COLMAP workspace ===\ncolmap_dir = Path('/kaggle/working/colmap')\ncolmap_dir.mkdir(exist_ok=True)\nimage_dir = colmap_dir / 'images'\nimage_dir.mkdir(exist_ok=True)\nfeatures_dir = colmap_dir / 'features'\nfeatures_dir.mkdir(exist_ok=True)\nmatches_dir = colmap_dir / 'matches'\nmatches_dir.mkdir(exist_ok=True)\n\n# === Select and preprocess a subset of cluster 0 images ===\ncluster_0_paths = [image_paths[i] for i in cluster_0_indices]\nprint(\"Paths in cluster_0_paths:\")\nfor path in cluster_0_paths:\n    print(path)\n\nfor i, img_path in enumerate(cluster_0_paths):\n    img = Image.open(img_path).convert('RGB')\n    img_np = np.array(img)\n    r, g, b = cv2.split(img_np)\n    clahe = cv2.createCLAHE(clipLimit=2.0, tileGridSize=(8, 8))\n    r_clahe = clahe.apply(r)\n    g_clahe = clahe.apply(g)\n    b_clahe = clahe.apply(b)\n    img_clahe = cv2.merge([r_clahe, g_clahe, b_clahe])\n    img_pil = Image.fromarray(img_clahe)\n    img_pil.save(image_dir / f\"image_{i}.png\")\n\n# === Visualize the preprocessed images ===\nfig, axes = plt.subplots(1, 5, figsize=(20, 4))\naxes = axes.flatten()\nfor i, img_path in enumerate(cluster_0_paths):\n    img = Image.open(image_dir / f\"image_{i}.png\")\n    axes[i].imshow(img)\n    axes[i].set_title(f\"Preprocessed Image {i}\")\n    axes[i].axis('off')\nplt.tight_layout()\nplt.show()\nplt.close('all')\n\n# === Initialize SuperGlue ===\nconfig = {\n    'superpoint': {\n        'nms_radius': 4,\n        'keypoint_threshold': 0.005,\n        'max_keypoints': 1024\n    },\n    'superglue': {\n        'weights': 'indoor',\n        'sinkhorn_iterations': 20,\n        'match_threshold': 0.2,\n    }\n}\nmatching = Matching(config).eval().to(device)\n\n# === Define database and output paths ===\ndatabase_path = colmap_dir / 'database.db'\noutput_path = colmap_dir / 'sparse'\noutput_path.mkdir(exist_ok=True)\n\n# === Step 1: Initialize the COLMAP database with images ===\nsift_extraction_options = pycolmap.SiftExtractionOptions()\nsift_extraction_options.peak_threshold = 0.001\nsift_extraction_options.edge_threshold = 20\nsift_extraction_options.max_num_features = 16384\n\nstart_time = time.time()\npycolmap.extract_features(\n    database_path=str(database_path),\n    image_path=str(image_dir),\n    camera_mode=pycolmap.CameraMode.AUTO,\n    sift_options=sift_extraction_options\n)\nprint(f\"Feature extraction (initialization) took {time.time() - start_time:.2f} seconds.\")\n\n# === Step 2: Run SuperPoint and SuperGlue for feature detection and matching ===\nkeypoints_dict = {}\ndescriptors_dict = {}\nmatches_dict = {}\n\nstart_time = time.time()\nimage_pairs = [(i, j) for i in range(len(cluster_0_paths)) for j in range(i + 1, len(cluster_0_paths))]\nfor idx0, idx1 in image_pairs:\n    img0 = cv2.imread(str(image_dir / f\"image_{idx0}.png\"), cv2.IMREAD_GRAYSCALE)\n    img1 = cv2.imread(str(image_dir / f\"image_{idx1}.png\"), cv2.IMREAD_GRAYSCALE)\n    \n    img0 = cv2.resize(img0, (640, 480))\n    img1 = cv2.resize(img1, (640, 480))\n    \n    inp0 = frame2tensor(img0, device)\n    inp1 = frame2tensor(img1, device)\n    \n    pred = matching({'image0': inp0, 'image1': inp1})\n    pred = {k: v[0].detach().cpu().numpy() for k, v in pred.items()}\n    kpts0, kpts1 = pred['keypoints0'], pred['keypoints1']\n    matches, conf = pred['matches0'], pred['matching_scores0']\n    \n    valid = matches > -1\n    mkpts0 = kpts0[valid]\n    mkpts1 = kpts1[matches[valid]]\n    match_conf = conf[valid]\n    \n    img0_name = f\"image_{idx0}.png\"\n    img1_name = f\"image_{idx1}.png\"\n    if img0_name not in keypoints_dict:\n        keypoints_dict[img0_name] = kpts0\n        descriptors_dict[img0_name] = np.zeros((len(kpts0), 128), dtype=np.float32)\n    if img1_name not in keypoints_dict:\n        keypoints_dict[img1_name] = kpts1\n        descriptors_dict[img1_name] = np.zeros((len(kpts1), 128), dtype=np.float32)\n    \n    matches_dict[(img0_name, img1_name)] = (mkpts0, mkpts1, match_conf)\n    \n    if idx0 == 0 and idx1 == 1:\n        color = np.random.randint(0, 255, (len(mkpts0), 3)) / 255.0\n        text = [\n            f'SuperGlue',\n            f'Keypoints: {len(kpts0)}:{len(kpts1)}',\n            f'Matches: {len(mkpts0)}'\n        ]\n        plt.figure(figsize=(12, 6))\n        make_matching_plot(img0, img1, kpts0, kpts1, mkpts0, mkpts1, color, text, path=None, show_keypoints=True)\n        fig = plt.gcf()\n        fig.canvas.draw()\n        plot_img = np.frombuffer(fig.canvas.tostring_rgb(), dtype=np.uint8)\n        plot_img = plot_img.reshape(fig.canvas.get_width_height()[::-1] + (3,))\n        plt.close(fig)\n        plt.figure(figsize=(12, 6))\n        plt.imshow(plot_img)\n        plt.axis('off')\n        plt.show()\n        plt.close('all')\n\nprint(f\"SuperGlue matching took {time.time() - start_time:.2f} seconds.\")\n\n# === Step 3: Save SuperPoint keypoints to text files ===\nfor img_name in keypoints_dict:\n    keypoints = keypoints_dict[img_name]\n    descriptors = descriptors_dict[img_name]\n    keypoints_colmap = np.zeros((len(keypoints), 4), dtype=np.float32)\n    keypoints_colmap[:, :2] = keypoints\n    keypoints_colmap[:, 2] = 1.0\n    keypoints_colmap[:, 3] = 0.0\n    \n    feature_file = features_dir / f\"{img_name}.txt\"\n    with open(feature_file, 'w') as f:\n        f.write(f\"{len(keypoints)} 128\\n\")\n        for kp, desc in zip(keypoints_colmap, descriptors):\n            f.write(f\"{kp[0]} {kp[1]} {kp[2]} {kp[3]} {' '.join(map(str, desc))}\\n\")\n\n# === Step 4: Save SuperGlue matches to a text file ===\nmatches_file = matches_dir / \"matches.txt\"\nwith open(matches_file, 'w') as f:\n    for (img0_name, img1_name), (mkpts0, mkpts1, match_conf) in matches_dict.items():\n        kpts0 = keypoints_dict[img0_name]\n        kpts1 = keypoints_dict[img1_name]\n        matches = []\n        for i, (mkpt0, mkpt1) in enumerate(zip(mkpts0, mkpts1)):\n            idx0 = np.where((kpts0 == mkpt0).all(axis=1))[0][0]\n            idx1 = np.where((kpts1 == mkpt1).all(axis=1))[0][0]\n            matches.append((idx0, idx1))\n        if matches:\n            f.write(f\"{img0_name} {img1_name}\\n\")\n            for idx0, idx1 in matches:\n                f.write(f\"{idx0} {idx1}\\n\")\n            f.write(\"\\n\")\n\n# === Step 5: Import features and matches using COLMAP command-line tools ===\nstart_time = time.time()\ntry:\n    subprocess.run([\n        \"colmap\", \"feature_importer\",\n        \"--database_path\", str(database_path),\n        \"--image_path\", str(image_dir),\n        \"--import_path\", str(features_dir)\n    ], check=True)\n    \n    subprocess.run([\n        \"colmap\", \"matches_importer\",\n        \"--database_path\", str(database_path),\n        \"--match_list_path\", str(matches_file),\n        \"--match_type\", \"pairs\"\n    ], check=True)\n    print(f\"Importing features and matches took {time.time() - start_time:.2f} seconds.\")\nexcept Exception as e:\n    print(f\"Error importing features and matches: {e}\")\n    print(\"Falling back to default COLMAP matching in the next step...\")\n\n# === Step 6: Incremental mapping ===\noptions = pycolmap.IncrementalPipelineOptions()\noptions.min_num_matches = 1\noptions.min_model_size = 2\noptions.init_num_trials = 5000\noptions.ba_refine_focal_length = True\noptions.ba_refine_principal_point = False\noptions.ba_refine_extra_params = False\n\nstart_time = time.time()\nreconstructions = pycolmap.incremental_mapping(\n    database_path=str(database_path),\n    image_path=str(image_dir),\n    output_path=str(output_path),\n    options=options\n)\nprint(f\"Incremental mapping took {time.time() - start_time:.2f} seconds.\")\n\nif reconstructions:\n    reconstruction = list(reconstructions.values())[0]\n    print(f\"Number of images registered: {reconstruction.num_reg_images()}\")\n    print(f\"Number of 3D points reconstructed: {reconstruction.num_points3D()}\")\n\n    points3d = []\n    for point3d_id in reconstruction.points3D:\n        point3d = reconstruction.points3D[point3d_id]\n        points3d.append(point3d.xyz)\n    points3d = np.array(points3d)\n\n    camera_positions = []\n    for img_id in reconstruction.images:\n        img = reconstruction.images[img_id]\n        cam_from_world = img.cam_from_world\n        rotation = cam_from_world.rotation.matrix()\n        translation = cam_from_world.translation\n        position = -np.dot(rotation.T, translation)\n        camera_positions.append(position)\n    camera_positions = np.array(camera_positions)\n\n    fig = plt.figure(figsize=(10, 8))\n    ax = fig.add_subplot(111, projection='3d')\n    ax.scatter(points3d[:, 0], points3d[:, 1], points3d[:, 2], s=1, c='b', label='3D Points')\n    ax.scatter(camera_positions[:, 0], camera_positions[:, 1], camera_positions[:, 2], s=50, c='r', marker='^', label='Cameras')\n    ax.set_title(\"3D Point Cloud and Camera Positions (Cluster 0 - First 5 Images)\")\n    ax.set_xlabel(\"X\")\n    ax.set_ylabel(\"Y\")\n    ax.set_zlabel(\"Z\")\n    ax.legend()\n    plt.show()\n    plt.close('all')\nelse:\n    print(\"Reconstruction failed: No reconstructions returned.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-06T15:21:09.498054Z","iopub.execute_input":"2025-04-06T15:21:09.498343Z","iopub.status.idle":"2025-04-06T15:22:29.654179Z","shell.execute_reply.started":"2025-04-06T15:21:09.498324Z","shell.execute_reply":"2025-04-06T15:22:29.653086Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 1st","metadata":{}},{"cell_type":"code","source":"# Step 3.10 (Simplified): Custom SfM Pipeline with SuperGlue Matches (5 Images, Single Cluster, No Bundle Adjustment)\nimport numpy as np\nimport matplotlib.pyplot as plt\nfrom pathlib import Path\nfrom PIL import Image\nimport cv2\nimport time\nimport os\nimport torch\nimport sys\nimport warnings\nwarnings.filterwarnings(\"ignore\", category=FutureWarning)\n\nsys.path.append('SuperGluePretrainedNetwork')\nfrom models.matching import Matching\nfrom models.utils import (frame2tensor, make_matching_plot)\n\n# === Set device ===\ndevice = torch.device('cuda' if torch.cuda.is_available() else 'cpu')\nprint(f\"Using device: {device}\")\n\n# === Define image paths ===\nstairs_dir = Path('/kaggle/input/image-matching-challenge-2025/train/stairs')\nimage_paths = [str(stairs_dir / img) for img in os.listdir(stairs_dir) if img.endswith(('.png', '.jpg', '.jpeg'))]\nprint(f\"Total images in stairs directory: {len(image_paths)}\")\nprint(\"First few image paths:\")\nfor path in image_paths[:5]:\n    print(path)\n\n# === Select first 5 images ===\nnum_images = 5\nselected_indices = list(range(min(num_images, len(image_paths))))\nprint(f\"Selected indices: {selected_indices}\")\n\n# === Define a single cluster ===\nclusters = [selected_indices]  # Single cluster with 5 images\nprint(f\"Clusters: {clusters}\")\n\n# === Setup workspace ===\nwork_dir = Path('/kaggle/working/custom_sfm')\nwork_dir.mkdir(exist_ok=True)\nimage_dir = work_dir / 'images'\nimage_dir.mkdir(exist_ok=True)\n\n# === Copy all selected images ===\nselected_paths = [image_paths[i] for i in selected_indices]\nfor i, img_path in enumerate(selected_paths):\n    img = Image.open(img_path).convert('RGB')\n    img.save(image_dir / f\"image_{i}.png\")\n\n# === Initialize SuperGlue ===\nconfig = {\n    'superpoint': {\n        'nms_radius': 4,\n        'keypoint_threshold': 0.003,\n        'max_keypoints': 2048  # Reduced to speed up processing\n    },\n    'superglue': {\n        'weights': 'indoor',\n        'sinkhorn_iterations': 20,\n        'match_threshold': 0.05  # Increased to reduce matches\n    }\n}\nmatching = Matching(config).eval().to(device)\n\n# === Function to run SfM on a single cluster ===\ndef run_sfm_on_cluster(cluster_indices, image_dir, matching, device):\n    keypoints_dict = {}\n    matches_dict = {}\n    image_sizes = {}\n    \n    print(f\"Processing cluster with indices: {cluster_indices}\")\n    \n    # Run SuperPoint and SuperGlue\n    start_time = time.time()\n    image_pairs = [(i, j) for i in range(len(cluster_indices)) for j in range(i + 1, len(cluster_indices))]\n    edges = []\n    for idx0, idx1 in image_pairs:\n        global_idx0, global_idx1 = cluster_indices[idx0], cluster_indices[idx1]\n        img0 = cv2.imread(str(image_dir / f\"image_{global_idx0}.png\"), cv2.IMREAD_GRAYSCALE)\n        img1 = cv2.imread(str(image_dir / f\"image_{global_idx1}.png\"), cv2.IMREAD_GRAYSCALE)\n        \n        orig_size0 = img0.shape[::-1]\n        orig_size1 = img1.shape[::-1]\n        \n        img0 = cv2.resize(img0, (640, 480))\n        img1 = cv2.resize(img1, (640, 480))\n        \n        inp0 = frame2tensor(img0, device)\n        inp1 = frame2tensor(img1, device)\n        \n        pred = matching({'image0': inp0, 'image1': inp1})\n        pred = {k: v[0].detach().cpu().numpy() for k, v in pred.items()}\n        kpts0, kpts1 = pred['keypoints0'], pred['keypoints1']\n        matches, conf = pred['matches0'], pred['matching_scores0']\n        \n        valid = matches > -1\n        mkpts0 = kpts0[valid]\n        mkpts1 = kpts1[matches[valid]]\n        \n        scale0 = (orig_size0[0] / 640, orig_size0[1] / 480)\n        scale1 = (orig_size1[0] / 640, orig_size1[1] / 480)\n        kpts0[:, 0] *= scale0[0]\n        kpts0[:, 1] *= scale0[1]\n        kpts1[:, 0] *= scale1[0]\n        kpts1[:, 1] *= scale1[1]\n        mkpts0[:, 0] *= scale0[0]\n        mkpts0[:, 1] *= scale0[1]\n        mkpts1[:, 0] *= scale1[0]\n        mkpts1[:, 1] *= scale1[1]\n        \n        img0_name = f\"image_{global_idx0}.png\"\n        img1_name = f\"image_{global_idx1}.png\"\n        if img0_name not in keypoints_dict:\n            keypoints_dict[img0_name] = kpts0\n            image_sizes[img0_name] = orig_size0\n        if img1_name not in keypoints_dict:\n            keypoints_dict[img1_name] = kpts1\n            image_sizes[img1_name] = orig_size1\n        \n        matches_dict[(img0_name, img1_name)] = (mkpts0, mkpts1)\n        \n        edges.append((len(mkpts0), idx0, idx1))\n        \n        print(f\"Matches between {img0_name} and {img1_name}: {len(mkpts0)}\")\n    \n    print(f\"SuperGlue matching for cluster took {time.time() - start_time:.2f} seconds.\")\n    \n    # Sort edges by number of matches (descending)\n    edges.sort(key=lambda x: -x[0])\n    parent = list(range(len(cluster_indices)))\n    rank = [0] * len(cluster_indices)\n    \n    def find(x):\n        if parent[x] != x:\n            parent[x] = find(parent[x])\n        return parent[x]\n    \n    def union(x, y):\n        px, py = find(x), find(y)\n        if px == py:\n            return\n        if rank[px] < rank[py]:\n            px, py = py, px\n        parent[py] = px\n        if rank[px] == rank[py]:\n            rank[px] += 1\n    \n    selected_pairs = []\n    for num_matches, idx0, idx1 in edges:\n        if num_matches < 8:\n            continue\n        if find(idx0) != find(idx1):\n            union(idx0, idx1)\n            selected_pairs.append((idx0, idx1))\n    \n    if not selected_pairs:\n        print(\"No pairs with sufficient matches found in cluster.\")\n        return np.array([]), np.array([]), np.array([]), {}\n    \n    # Initialize poses with the first selected pair\n    idx0, idx1 = selected_pairs[0]\n    global_idx0, global_idx1 = cluster_indices[idx0], cluster_indices[idx1]\n    img0_name = f\"image_{global_idx0}.png\"\n    img1_name = f\"image_{global_idx1}.png\"\n    \n    mkpts0, mkpts1 = matches_dict[(img0_name, img1_name)]\n    mkpts0 = mkpts0.astype(np.float32)\n    mkpts1 = mkpts1.astype(np.float32)\n    \n    # Estimate focal length\n    image_width, image_height = image_sizes[img0_name]\n    focal_length = max(image_width, image_height) * 1.2\n    cx, cy = image_width / 2, image_height / 2\n    K = np.array([\n        [focal_length, 0, cx],\n        [0, focal_length, cy],\n        [0, 0, 1]\n    ], dtype=np.float32)\n    \n    # Estimate essential matrix\n    E, mask = cv2.findEssentialMat(mkpts0, mkpts1, K, method=cv2.RANSAC, prob=0.999, threshold=1.0)\n    if E is None or E.shape != (3, 3):\n        print(f\"Failed to estimate essential matrix for initial pair {img0_name} and {img1_name}\")\n        return np.array([]), np.array([]), np.array([]), {}\n    \n    # Recover pose\n    _, R, t, mask = cv2.recoverPose(E, mkpts0, mkpts1, K)\n    if R is None or t is None:\n        print(f\"Failed to recover pose for initial pair {img0_name} and {img1_name}\")\n        return np.array([]), np.array([]), np.array([]), {}\n    \n    R = R.astype(np.float32)\n    t = t.astype(np.float32).flatten()\n    \n    # Initialize poses\n    poses = {img0_name: (np.eye(3, dtype=np.float32), np.zeros(3, dtype=np.float32))}\n    poses[img1_name] = (R, t)\n    registered_indices = {idx0, idx1}\n    \n    # Process remaining pairs\n    start_time = time.time()\n    for idx0, idx1 in selected_pairs[1:]:\n        global_idx0, global_idx1 = cluster_indices[idx0], cluster_indices[idx1]\n        img0_name = f\"image_{global_idx0}.png\"\n        img1_name = f\"image_{global_idx1}.png\"\n        \n        if img0_name in poses and img1_name in poses:\n            continue\n        elif img0_name not in poses and img1_name not in poses:\n            continue\n        \n        if img1_name in poses and img0_name not in poses:\n            img0_name, img1_name = img1_name, img0_name\n            idx0, idx1 = idx1, idx0\n        \n        key1 = (img0_name, img1_name)\n        key2 = (img1_name, img0_name)\n        if key1 in matches_dict:\n            mkpts0, mkpts1 = matches_dict[key1]\n        elif key2 in matches_dict:\n            mkpts1, mkpts0 = matches_dict[key2]\n        else:\n            print(f\"Matches not found for pair ({img0_name}, {img1_name})\")\n            continue\n        \n        mkpts0 = mkpts0.astype(np.float32)\n        mkpts1 = mkpts1.astype(np.float32)\n        \n        E, mask = cv2.findEssentialMat(mkpts0, mkpts1, K, method=cv2.RANSAC, prob=0.999, threshold=1.0)\n        if E is None or E.shape != (3, 3):\n            print(f\"Failed to estimate essential matrix between {img0_name} and {img1_name}\")\n            continue\n        \n        _, R, t, mask = cv2.recoverPose(E, mkpts0, mkpts1, K)\n        if R is None or t is None:\n            print(f\"Failed to recover pose between {img0_name} and {img1_name}\")\n            continue\n        \n        R = R.astype(np.float32)\n        t = t.astype(np.float32).flatten()\n        \n        R0, t0 = poses[img0_name]\n        R = R0 @ R\n        t = R0 @ t + t0\n        \n        poses[img1_name] = (R, t)\n        registered_indices.add(idx1)\n    \n    print(f\"Pose estimation took {time.time() - start_time:.2f} seconds.\")\n    \n    # Triangulate 3D points\n    start_time = time.time()\n    points3d = []\n    colors = []\n    for idx0, idx1 in image_pairs:\n        global_idx0, global_idx1 = cluster_indices[idx0], cluster_indices[idx1]\n        img0_name = f\"image_{global_idx0}.png\"\n        img1_name = f\"image_{global_idx1}.png\"\n        \n        if img0_name not in poses or img1_name not in poses:\n            continue\n        \n        key1 = (img0_name, img1_name)\n        key2 = (img1_name, img0_name)\n        if key1 in matches_dict:\n            mkpts0, mkpts1 = matches_dict[key1]\n        elif key2 in matches_dict:\n            mkpts1, mkpts0 = matches_dict[key2]\n        else:\n            print(f\"Matches not found for pair ({img0_name}, {img1_name}) during triangulation\")\n            continue\n        \n        if len(mkpts0) != len(mkpts1):\n            print(f\"Mismatch in number of points between {img0_name} and {img1_name}: {len(mkpts0)} vs {len(mkpts1)}\")\n            continue\n        \n        mkpts0 = mkpts0.astype(np.float32)\n        mkpts1 = mkpts1.astype(np.float32)\n        \n        R0, t0 = poses[img0_name]\n        R1, t1 = poses[img1_name]\n        \n        P0 = K @ np.hstack((R0, t0.reshape(3, 1)))\n        P1 = K @ np.hstack((R1, t1.reshape(3, 1)))\n        P0 = P0.astype(np.float32)\n        P1 = P1.astype(np.float32)\n        \n        pts0 = mkpts0.T.astype(np.float32)\n        pts1 = mkpts1.T.astype(np.float32)\n        \n        if pts0.shape[1] == 0 or pts1.shape[1] == 0:\n            print(f\"Skipping triangulation due to zero matches between {img0_name} and {img1_name}\")\n            continue\n        \n        points4d = cv2.triangulatePoints(P0, P1, pts0, pts1)\n        points3d_h = points4d[:3] / points4d[3]\n        points3d.extend(points3d_h.T)\n        \n        img0 = cv2.imread(str(image_dir / img0_name))\n        img0 = cv2.resize(img0, image_sizes[img0_name])\n        for pt in mkpts0:\n            x, y = int(pt[0]), int(pt[1])\n            if 0 <= x < img0.shape[1] and 0 <= y < img0.shape[0]:\n                colors.append(img0[y, x] / 255.0)\n            else:\n                colors.append([0, 0, 0])\n    \n    points3d = np.array(points3d)\n    colors = np.array(colors)\n    \n    # Remove invalid points\n    valid = np.all(np.isfinite(points3d), axis=1) & (np.abs(points3d) < 1e5).all(axis=1)\n    points3d = points3d[valid]\n    colors = colors[valid]\n    \n    print(f\"Triangulation took {time.time() - start_time:.2f} seconds.\")\n    \n    # Camera positions\n    camera_positions = []\n    for img_name in poses:\n        R, t = poses[img_name]\n        pos = -R.T @ t\n        camera_positions.append(pos)\n    camera_positions = np.array(camera_positions)\n    \n    # Visualize\n    fig = plt.figure(figsize=(10, 8))\n    ax = fig.add_subplot(111, projection='3d')\n    ax.scatter(points3d[:, 0], points3d[:, 1], points3d[:, 2], c=colors, s=1, label='3D Points')\n    ax.scatter(camera_positions[:, 0], camera_positions[:, 1], camera_positions[:, 2], s=50, c='r', marker='^', label='Cameras')\n    ax.set_title(f\"3D Point Cloud and Camera Positions (Cluster {cluster_indices[0]}-{cluster_indices[-1]})\")\n    ax.set_xlabel(\"X\")\n    ax.set_ylabel(\"Y\")\n    ax.set_zlabel(\"Z\")\n    ax.legend()\n    plt.show()\n    plt.close('all')\n    \n    print(f\"Cluster {cluster_indices[0]}-{cluster_indices[-1]}: {len(points3d)} 3D points, {len(poses)} cameras\")\n    \n    return points3d, colors, camera_positions, poses\n\n# === Process the single cluster ===\nall_points3d = []\nall_colors = []\nall_camera_positions = []\nall_poses = {}\n\nstart_time = time.time()\nfor cluster_idx, cluster_indices in enumerate(clusters):\n    points3d, colors, camera_positions, poses = run_sfm_on_cluster(cluster_indices, image_dir, matching, device)\n    all_points3d.append(points3d)\n    all_colors.append(colors)\n    all_camera_positions.append(camera_positions)\n    all_poses.update(poses)\n\n# Since there's only one cluster, no merging is needed\npoints3d = all_points3d[0]\ncolors = all_colors[0]\ncamera_positions = all_camera_positions[0]\n\nprint(f\"Total runtime: {time.time() - start_time:.2f} seconds\")\nprint(f\"Total number of 3D points reconstructed: {len(points3d)}\")\nprint(f\"Total number of cameras estimated: {len(camera_positions)}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-06T15:22:29.655445Z","iopub.execute_input":"2025-04-06T15:22:29.655806Z","iopub.status.idle":"2025-04-06T15:22:36.230054Z","shell.execute_reply.started":"2025-04-06T15:22:29.655770Z","shell.execute_reply":"2025-04-06T15:22:36.229040Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 2nd","metadata":{}},{"cell_type":"code","source":"# Step 3.10 (Step 2): Custom SfM Pipeline with SuperGlue Matches (10 Images, 2 Clusters, No Bundle Adjustment)\nimport numpy as np\nimport matplotlib.pyplot as plt\nfrom pathlib import Path\nfrom PIL import Image\nimport cv2\nimport time\nimport os\nimport torch\nimport sys\nimport warnings\nwarnings.filterwarnings(\"ignore\", category=FutureWarning)\n\nsys.path.append('SuperGluePretrainedNetwork')\nfrom models.matching import Matching\nfrom models.utils import (frame2tensor, make_matching_plot)\n\n# === Set device ===\ndevice = torch.device('cuda' if torch.cuda.is_available() else 'cpu')\nprint(f\"Using device: {device}\")\n\n# === Define image paths ===\nstairs_dir = Path('/kaggle/input/image-matching-challenge-2025/train/stairs')\nimage_paths = [str(stairs_dir / img) for img in os.listdir(stairs_dir) if img.endswith(('.png', '.jpg', '.jpeg'))]\nprint(f\"Total images in stairs directory: {len(image_paths)}\")\nprint(\"First few image paths:\")\nfor path in image_paths[:5]:\n    print(path)\n\n# === Select first 10 images ===\nnum_images = 10\nselected_indices = list(range(min(num_images, len(image_paths))))\nprint(f\"Selected indices: {selected_indices}\")\n\n# === Define clusters (5 images per cluster, with overlap of 2 images) ===\ncluster_size = 5\noverlap = 2\nclusters = []\nfor i in range(0, len(selected_indices), cluster_size - overlap):\n    cluster_indices = selected_indices[i:i + cluster_size]\n    if len(cluster_indices) >= 2:\n        clusters.append(cluster_indices)\nprint(f\"Clusters: {clusters}\")\n\n# === Setup workspace ===\nwork_dir = Path('/kaggle/working/custom_sfm')\nwork_dir.mkdir(exist_ok=True)\nimage_dir = work_dir / 'images'\nimage_dir.mkdir(exist_ok=True)\n\n# === Copy all selected images ===\nselected_paths = [image_paths[i] for i in selected_indices]\nfor i, img_path in enumerate(selected_paths):\n    img = Image.open(img_path).convert('RGB')\n    img.save(image_dir / f\"image_{i}.png\")\n\n# === Initialize SuperGlue ===\nconfig = {\n    'superpoint': {\n        'nms_radius': 4,\n        'keypoint_threshold': 0.003,\n        'max_keypoints': 2048\n    },\n    'superglue': {\n        'weights': 'indoor',\n        'sinkhorn_iterations': 20,\n        'match_threshold': 0.05\n    }\n}\nmatching = Matching(config).eval().to(device)\n\n# === Function to run SfM on a single cluster ===\ndef run_sfm_on_cluster(cluster_indices, image_dir, matching, device, cluster_idx, clusters):\n    keypoints_dict = {}\n    matches_dict = {}\n    image_sizes = {}\n    \n    print(f\"Processing cluster with indices: {cluster_indices}\")\n    \n    # Run SuperPoint and SuperGlue\n    start_time = time.time()\n    image_pairs = [(i, j) for i in range(len(cluster_indices)) for j in range(i + 1, len(cluster_indices))]\n    edges = []\n    for idx0, idx1 in image_pairs:\n        global_idx0, global_idx1 = cluster_indices[idx0], cluster_indices[idx1]\n        img0 = cv2.imread(str(image_dir / f\"image_{global_idx0}.png\"), cv2.IMREAD_GRAYSCALE)\n        img1 = cv2.imread(str(image_dir / f\"image_{global_idx1}.png\"), cv2.IMREAD_GRAYSCALE)\n        \n        orig_size0 = img0.shape[::-1]\n        orig_size1 = img1.shape[::-1]\n        \n        img0 = cv2.resize(img0, (640, 480))\n        img1 = cv2.resize(img1, (640, 480))\n        \n        inp0 = frame2tensor(img0, device)\n        inp1 = frame2tensor(img1, device)\n        \n        pred = matching({'image0': inp0, 'image1': inp1})\n        pred = {k: v[0].detach().cpu().numpy() for k, v in pred.items()}\n        kpts0, kpts1 = pred['keypoints0'], pred['keypoints1']\n        matches, conf = pred['matches0'], pred['matching_scores0']\n        \n        valid = matches > -1\n        mkpts0 = kpts0[valid]\n        mkpts1 = kpts1[matches[valid]]\n        \n        scale0 = (orig_size0[0] / 640, orig_size0[1] / 480)\n        scale1 = (orig_size1[0] / 640, orig_size1[1] / 480)\n        kpts0[:, 0] *= scale0[0]\n        kpts0[:, 1] *= scale0[1]\n        kpts1[:, 0] *= scale1[0]\n        kpts1[:, 1] *= scale1[1]\n        mkpts0[:, 0] *= scale0[0]\n        mkpts0[:, 1] *= scale0[1]\n        mkpts1[:, 0] *= scale1[0]\n        mkpts1[:, 1] *= scale1[1]\n        \n        img0_name = f\"image_{global_idx0}.png\"\n        img1_name = f\"image_{global_idx1}.png\"\n        if img0_name not in keypoints_dict:\n            keypoints_dict[img0_name] = kpts0\n            image_sizes[img0_name] = orig_size0\n        if img1_name not in keypoints_dict:\n            keypoints_dict[img1_name] = kpts1\n            image_sizes[img1_name] = orig_size1\n        \n        matches_dict[(img0_name, img1_name)] = (mkpts0, mkpts1)\n        \n        priority = 0\n        if cluster_idx < len(clusters) - 1:\n            overlap_indices = clusters[cluster_idx][-overlap:]\n            if global_idx0 in overlap_indices or global_idx1 in overlap_indices:\n                priority = 1\n        \n        edges.append((len(mkpts0), idx0, idx1, priority))\n        \n        print(f\"Matches between {img0_name} and {img1_name}: {len(mkpts0)}\")\n    \n    print(f\"SuperGlue matching for cluster took {time.time() - start_time:.2f} seconds.\")\n    \n    # Sort edges\n    edges.sort(key=lambda x: (-x[3], -x[0]))\n    parent = list(range(len(cluster_indices)))\n    rank = [0] * len(cluster_indices)\n    \n    def find(x):\n        if parent[x] != x:\n            parent[x] = find(parent[x])\n        return parent[x]\n    \n    def union(x, y):\n        px, py = find(x), find(y)\n        if px == py:\n            return\n        if rank[px] < rank[py]:\n            px, py = py, px\n        parent[py] = px\n        if rank[px] == rank[py]:\n            rank[px] += 1\n    \n    selected_pairs = []\n    for num_matches, idx0, idx1, _ in edges:\n        if num_matches < 8:\n            continue\n        if find(idx0) != find(idx1):\n            union(idx0, idx1)\n            selected_pairs.append((idx0, idx1))\n    \n    if cluster_idx < len(clusters) - 1:\n        overlap_indices = clusters[cluster_idx][-overlap:]\n        for overlap_idx in overlap_indices:\n            overlap_local_idx = cluster_indices.index(overlap_idx)\n            if overlap_local_idx not in {idx0 for idx0, _ in selected_pairs} and overlap_local_idx not in {idx1 for _, idx1 in selected_pairs}:\n                best_pair = None\n                best_num_matches = 0\n                for num_matches, idx0, idx1, _ in edges:\n                    if num_matches < 8:\n                        continue\n                    if idx0 == overlap_local_idx or idx1 == overlap_local_idx:\n                        if num_matches > best_num_matches:\n                            best_num_matches = num_matches\n                            best_pair = (idx0, idx1)\n                if best_pair:\n                    idx0, idx1 = best_pair\n                    if find(idx0) != find(idx1):\n                        union(idx0, idx1)\n                        selected_pairs.append((idx0, idx1))\n    \n    if not selected_pairs:\n        print(\"No pairs with sufficient matches found in cluster.\")\n        return np.array([]), np.array([]), np.array([]), {}\n    \n    # Initialize poses\n    idx0, idx1 = selected_pairs[0]\n    global_idx0, global_idx1 = cluster_indices[idx0], cluster_indices[idx1]\n    img0_name = f\"image_{global_idx0}.png\"\n    img1_name = f\"image_{global_idx1}.png\"\n    \n    mkpts0, mkpts1 = matches_dict[(img0_name, img1_name)]\n    mkpts0 = mkpts0.astype(np.float32)\n    mkpts1 = mkpts1.astype(np.float32)\n    \n    image_width, image_height = image_sizes[img0_name]\n    focal_length = max(image_width, image_height) * 1.2\n    cx, cy = image_width / 2, image_height / 2\n    K = np.array([\n        [focal_length, 0, cx],\n        [0, focal_length, cy],\n        [0, 0, 1]\n    ], dtype=np.float32)\n    \n    E, mask = cv2.findEssentialMat(mkpts0, mkpts1, K, method=cv2.RANSAC, prob=0.999, threshold=1.0)\n    if E is None or E.shape != (3, 3):\n        print(f\"Failed to estimate essential matrix for initial pair {img0_name} and {img1_name}\")\n        return np.array([]), np.array([]), np.array([]), {}\n    \n    _, R, t, mask = cv2.recoverPose(E, mkpts0, mkpts1, K)\n    if R is None or t is None:\n        print(f\"Failed to recover pose for initial pair {img0_name} and {img1_name}\")\n        return np.array([]), np.array([]), np.array([]), {}\n    \n    R = R.astype(np.float32)\n    t = t.astype(np.float32).flatten()\n    \n    poses = {img0_name: (np.eye(3, dtype=np.float32), np.zeros(3, dtype=np.float32))}\n    poses[img1_name] = (R, t)\n    registered_indices = {idx0, idx1}\n    \n    # Process remaining pairs\n    start_time = time.time()\n    for idx0, idx1 in selected_pairs[1:]:\n        global_idx0, global_idx1 = cluster_indices[idx0], cluster_indices[idx1]\n        img0_name = f\"image_{global_idx0}.png\"\n        img1_name = f\"image_{global_idx1}.png\"\n        \n        if img0_name in poses and img1_name in poses:\n            continue\n        elif img0_name not in poses and img1_name not in poses:\n            continue\n        \n        if img1_name in poses and img0_name not in poses:\n            img0_name, img1_name = img1_name, img0_name\n            idx0, idx1 = idx1, idx0\n        \n        key1 = (img0_name, img1_name)\n        key2 = (img1_name, img0_name)\n        if key1 in matches_dict:\n            mkpts0, mkpts1 = matches_dict[key1]\n        elif key2 in matches_dict:\n            mkpts1, mkpts0 = matches_dict[key2]\n        else:\n            print(f\"Matches not found for pair ({img0_name}, {img1_name})\")\n            continue\n        \n        mkpts0 = mkpts0.astype(np.float32)\n        mkpts1 = mkpts1.astype(np.float32)\n        \n        E, mask = cv2.findEssentialMat(mkpts0, mkpts1, K, method=cv2.RANSAC, prob=0.999, threshold=1.0)\n        if E is None or E.shape != (3, 3):\n            print(f\"Failed to estimate essential matrix between {img0_name} and {img1_name}\")\n            continue\n        \n        _, R, t, mask = cv2.recoverPose(E, mkpts0, mkpts1, K)\n        if R is None or t is None:\n            print(f\"Failed to recover pose between {img0_name} and {img1_name}\")\n            continue\n        \n        R = R.astype(np.float32)\n        t = t.astype(np.float32).flatten()\n        \n        R0, t0 = poses[img0_name]\n        R = R0 @ R\n        t = R0 @ t + t0\n        \n        poses[img1_name] = (R, t)\n        registered_indices.add(idx1)\n    \n    print(f\"Pose estimation took {time.time() - start_time:.2f} seconds.\")\n    \n    # Triangulate 3D points\n    start_time = time.time()\n    points3d = []\n    colors = []\n    for idx0, idx1 in image_pairs:\n        global_idx0, global_idx1 = cluster_indices[idx0], cluster_indices[idx1]\n        img0_name = f\"image_{global_idx0}.png\"\n        img1_name = f\"image_{global_idx1}.png\"\n        \n        if img0_name not in poses or img1_name not in poses:\n            continue\n        \n        key1 = (img0_name, img1_name)\n        key2 = (img1_name, img0_name)\n        if key1 in matches_dict:\n            mkpts0, mkpts1 = matches_dict[key1]\n        elif key2 in matches_dict:\n            mkpts1, mkpts0 = matches_dict[key2]\n        else:\n            print(f\"Matches not found for pair ({img0_name}, {img1_name}) during triangulation\")\n            continue\n        \n        if len(mkpts0) != len(mkpts1):\n            print(f\"Mismatch in number of points between {img0_name} and {img1_name}: {len(mkpts0)} vs {len(mkpts1)}\")\n            continue\n        \n        mkpts0 = mkpts0.astype(np.float32)\n        mkpts1 = mkpts1.astype(np.float32)\n        \n        R0, t0 = poses[img0_name]\n        R1, t1 = poses[img1_name]\n        \n        P0 = K @ np.hstack((R0, t0.reshape(3, 1)))\n        P1 = K @ np.hstack((R1, t1.reshape(3, 1)))\n        P0 = P0.astype(np.float32)\n        P1 = P1.astype(np.float32)\n        \n        pts0 = mkpts0.T.astype(np.float32)\n        pts1 = mkpts1.T.astype(np.float32)\n        \n        if pts0.shape[1] == 0 or pts1.shape[1] == 0:\n            print(f\"Skipping triangulation due to zero matches between {img0_name} and {img1_name}\")\n            continue\n        \n        points4d = cv2.triangulatePoints(P0, P1, pts0, pts1)\n        points3d_h = points4d[:3] / points4d[3]\n        points3d.extend(points3d_h.T)\n        \n        img0 = cv2.imread(str(image_dir / img0_name))\n        img0 = cv2.resize(img0, image_sizes[img0_name])\n        for pt in mkpts0:\n            x, y = int(pt[0]), int(pt[1])\n            if 0 <= x < img0.shape[1] and 0 <= y < img0.shape[0]:\n                colors.append(img0[y, x] / 255.0)\n            else:\n                colors.append([0, 0, 0])\n    \n    points3d = np.array(points3d)\n    colors = np.array(colors)\n    \n    valid = np.all(np.isfinite(points3d), axis=1) & (np.abs(points3d) < 1e5).all(axis=1)\n    points3d = points3d[valid]\n    colors = colors[valid]\n    \n    print(f\"Triangulation took {time.time() - start_time:.2f} seconds.\")\n    \n    # Camera positions\n    camera_positions = []\n    for img_name in poses:\n        R, t = poses[img_name]\n        pos = -R.T @ t\n        camera_positions.append(pos)\n    camera_positions = np.array(camera_positions)\n    \n    # Visualize\n    fig = plt.figure(figsize=(10, 8))\n    ax = fig.add_subplot(111, projection='3d')\n    ax.scatter(points3d[:, 0], points3d[:, 1], points3d[:, 2], c=colors, s=1, label='3D Points')\n    ax.scatter(camera_positions[:, 0], camera_positions[:, 1], camera_positions[:, 2], s=50, c='r', marker='^', label='Cameras')\n    ax.set_title(f\"3D Point Cloud and Camera Positions (Cluster {cluster_indices[0]}-{cluster_indices[-1]})\")\n    ax.set_xlabel(\"X\")\n    ax.set_ylabel(\"Y\")\n    ax.set_zlabel(\"Z\")\n    ax.legend()\n    plt.show()\n    plt.close('all')\n    \n    print(f\"Cluster {cluster_indices[0]}-{cluster_indices[-1]}: {len(points3d)} 3D points, {len(poses)} cameras\")\n    \n    return points3d, colors, camera_positions, poses\n\n# === Process each cluster ===\nall_points3d = []\nall_colors = []\nall_camera_positions = []\nall_poses = {}\n\nstart_time = time.time()\nfor cluster_idx, cluster_indices in enumerate(clusters):\n    points3d, colors, camera_positions, poses = run_sfm_on_cluster(cluster_indices, image_dir, matching, device, cluster_idx, clusters)\n    all_points3d.append(points3d)\n    all_colors.append(colors)\n    all_camera_positions.append(camera_positions)\n    all_poses.update(poses)\n\n# Merge reconstructions\nif not all_points3d or len(all_points3d[0]) == 0:\n    print(\"No points reconstructed in the first cluster. Cannot proceed with merging.\")\nelse:\n    merged_points3d = all_points3d[0]\n    merged_colors = all_colors[0]\n    merged_camera_positions = all_camera_positions[0]\n\n    for i in range(1, len(clusters)):\n        overlap_image = f\"image_{clusters[i-1][-1]}.png\"\n        if overlap_image not in all_poses:\n            print(f\"Cannot align clusters {i-1} and {i} using overlapping image {overlap_image}.\")\n            continue\n        \n        R0_prev, t0_prev = all_poses[overlap_image]\n        R0_curr, t0_curr = all_poses[overlap_image]\n        \n        R_align = R0_prev @ np.linalg.inv(R0_curr)\n        t_align = t0_prev - R_align @ t0_curr\n        \n        points3d_curr = all_points3d[i]\n        if len(points3d_curr) == 0:\n            print(f\"Cluster {i} has no points to merge\")\n            continue\n        points3d_curr = (R_align @ points3d_curr.T).T + t_align\n        all_points3d[i] = points3d_curr\n        merged_points3d = np.vstack((merged_points3d, points3d_curr))\n        merged_colors = np.vstack((merged_colors, all_colors[i]))\n        \n        camera_positions_curr = all_camera_positions[i]\n        if len(camera_positions_curr) == 0:\n            print(f\"Cluster {i} has no camera positions to merge\")\n            continue\n        camera_positions_curr = (R_align @ camera_positions_curr.T).T + t_align\n        all_camera_positions[i] = camera_positions_curr\n        merged_camera_positions = np.vstack((merged_camera_positions, camera_positions_curr))\n        \n        for img_name in list(all_poses.keys()):\n            if img_name in all_poses:\n                R, t = all_poses[img_name]\n                R = R_align @ R\n                t = R_align @ t + t_align\n                all_poses[img_name] = (R, t)\n\n    # Remove duplicates in camera positions\n    unique_camera_positions = []\n    seen_images = set()\n    for i, pos in enumerate(merged_camera_positions):\n        img_idx = i % len(selected_indices)\n        img_name = f\"image_{img_idx}.png\"\n        if img_name not in seen_images:\n            unique_camera_positions.append(pos)\n            seen_images.add(img_name)\n    unique_camera_positions = np.array(unique_camera_positions)\n\n    # Visualize merged reconstruction\n    fig = plt.figure(figsize=(10, 8))\n    ax = fig.add_subplot(111, projection='3d')\n    ax.scatter(merged_points3d[:, 0], merged_points3d[:, 1], merged_points3d[:, 2], c=merged_colors, s=1, label='3D Points')\n    ax.scatter(unique_camera_positions[:, 0], unique_camera_positions[:, 1], unique_camera_positions[:, 2], s=50, c='r', marker='^', label='Cameras')\n    ax.set_title(\"Merged 3D Point Cloud and Camera Positions (First 10 Images)\")\n    ax.set_xlabel(\"X\")\n    ax.set_ylabel(\"Y\")\n    ax.set_zlabel(\"Z\")\n    ax.legend()\n    plt.show()\n    plt.close('all')\n\n    print(f\"Total runtime: {time.time() - start_time:.2f} seconds\")\n    print(f\"Total number of 3D points reconstructed: {len(merged_points3d)}\")\n    print(f\"Total number of cameras estimated: {len(unique_camera_positions)}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-06T15:22:36.231142Z","iopub.execute_input":"2025-04-06T15:22:36.231460Z","iopub.status.idle":"2025-04-06T15:22:50.528092Z","shell.execute_reply.started":"2025-04-06T15:22:36.231422Z","shell.execute_reply":"2025-04-06T15:22:50.527210Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 3rd","metadata":{}},{"cell_type":"code","source":"# Step 3.10 (Step 3): Custom SfM Pipeline with SuperGlue Matches and COLMAP Bundle Adjustment (10 Images)\nimport pycolmap\nimport numpy as np\nimport matplotlib.pyplot as plt\nfrom pathlib import Path\nfrom PIL import Image\nimport cv2\nimport time\nimport os\nimport torch\nimport sys\nimport warnings\nwarnings.filterwarnings(\"ignore\", category=FutureWarning)\n\nsys.path.append('SuperGluePretrainedNetwork')\nfrom models.matching import Matching\nfrom models.utils import (frame2tensor, make_matching_plot)\n\n# === Set device ===\ndevice = torch.device('cuda' if torch.cuda.is_available() else 'cpu')\nprint(f\"Using device: {device}\")\n\n# === Define image paths ===\nstairs_dir = Path('/kaggle/input/image-matching-challenge-2025/train/stairs')\nimage_paths = [str(stairs_dir / img) for img in os.listdir(stairs_dir) if img.endswith(('.png', '.jpg', '.jpeg'))]\nprint(f\"Total images in stairs directory: {len(image_paths)}\")\nprint(\"First few image paths:\")\nfor path in image_paths[:5]:\n    print(path)\n\n# === Select first 10 images ===\nnum_images = 10\nselected_indices = list(range(min(num_images, len(image_paths))))\nprint(f\"Selected indices: {selected_indices}\")\n\n# === Setup workspace ===\nwork_dir = Path('/kaggle/working/colmap')\nwork_dir.mkdir(exist_ok=True)\nimage_dir = work_dir / 'images'\nimage_dir.mkdir(exist_ok=True)\n\n# === Copy all selected images ===\nselected_paths = [image_paths[i] for i in selected_indices]\nfor i, img_path in enumerate(selected_paths):\n    img = Image.open(img_path).convert('RGB')\n    img.save(image_dir / f\"image_{i}.png\")\n\n# === Initialize SuperGlue ===\nconfig = {\n    'superpoint': {\n        'nms_radius': 4,\n        'keypoint_threshold': 0.003,\n        'max_keypoints': 2048\n    },\n    'superglue': {\n        'weights': 'indoor',\n        'sinkhorn_iterations': 20,\n        'match_threshold': 0.05\n    }\n}\nmatching = Matching(config).eval().to(device)\n\n# === Run SuperPoint and SuperGlue ===\nkeypoints_dict = {}\nmatches_dict = {}\nimage_sizes = {}\n\nstart_time = time.time()\nimage_pairs = [(i, j) for i in range(len(selected_indices)) for j in range(i + 1, len(selected_indices))]\nfor idx0, idx1 in image_pairs:\n    img0 = cv2.imread(str(image_dir / f\"image_{idx0}.png\"), cv2.IMREAD_GRAYSCALE)\n    img1 = cv2.imread(str(image_dir / f\"image_{idx1}.png\"), cv2.IMREAD_GRAYSCALE)\n    \n    orig_size0 = img0.shape[::-1]\n    orig_size1 = img1.shape[::-1]\n    \n    img0 = cv2.resize(img0, (640, 480))\n    img1 = cv2.resize(img1, (640, 480))\n    \n    inp0 = frame2tensor(img0, device)\n    inp1 = frame2tensor(img1, device)\n    \n    pred = matching({'image0': inp0, 'image1': inp1})\n    pred = {k: v[0].detach().cpu().numpy() for k, v in pred.items()}\n    kpts0, kpts1 = pred['keypoints0'], pred['keypoints1']\n    matches, conf = pred['matches0'], pred['matching_scores0']\n    \n    valid = matches > -1\n    mkpts0 = kpts0[valid]\n    mkpts1 = kpts1[matches[valid]]\n    \n    scale0 = (orig_size0[0] / 640, orig_size0[1] / 480)\n    scale1 = (orig_size1[0] / 640, orig_size1[1] / 480)\n    kpts0[:, 0] *= scale0[0]\n    kpts0[:, 1] *= scale0[1]\n    kpts1[:, 0] *= scale1[0]\n    kpts1[:, 1] *= scale1[1]\n    mkpts0[:, 0] *= scale0[0]\n    mkpts0[:, 1] *= scale0[1]\n    mkpts1[:, 0] *= scale1[0]\n    mkpts1[:, 1] *= scale1[1]\n    \n    img0_name = f\"image_{idx0}.png\"\n    img1_name = f\"image_{idx1}.png\"\n    if img0_name not in keypoints_dict:\n        keypoints_dict[img0_name] = kpts0\n        image_sizes[img0_name] = orig_size0\n    if img1_name not in keypoints_dict:\n        keypoints_dict[img1_name] = kpts1\n        image_sizes[img1_name] = orig_size1\n    \n    matches_dict[(img0_name, img1_name)] = (mkpts0, mkpts1)\n    \n    print(f\"Matches between {img0_name} and {img1_name}: {len(mkpts0)}\")\n\nprint(f\"SuperGlue matching took {time.time() - start_time:.2f} seconds.\")\n\n# === Setup COLMAP database ===\ndatabase_path = work_dir / 'database.db'\noutput_path = work_dir / 'sparse'\noutput_path.mkdir(exist_ok=True)\n\n# === Import features and matches into COLMAP ===\ntry:\n    db = pycolmap.Database(str(database_path))\n    \n    camera_ids = {}\n    for img_name in keypoints_dict:\n        width, height = image_sizes[img_name]\n        focal_length = max(width, height) * 1.2\n        camera = pycolmap.Camera(\n            model=\"SIMPLE_PINHOLE\",\n            width=width,\n            height=height,\n            params=[focal_length, width / 2, height / 2]\n        )\n        camera_id = db.add_camera(camera)\n        camera_ids[img_name] = camera_id\n    \n    image_ids = {}\n    for img_name in keypoints_dict:\n        image_id = db.add_image(img_name, camera_ids[img_name])\n        image_ids[img_name] = image_id\n    \n    for img_name, kpts in keypoints_dict.items():\n        image_id = image_ids[img_name]\n        keypoints = np.zeros((len(kpts), 4), dtype=np.float32)\n        keypoints[:, :2] = kpts\n        keypoints[:, 2] = 1.0  # Dummy sigma\n        db.add_keypoints(image_id, keypoints)\n    \n    for (img0_name, img1_name), (mkpts0, mkpts1) in matches_dict.items():\n        image_id0 = image_ids[img0_name]\n        image_id1 = image_ids[img1_name]\n        kpts0 = keypoints_dict[img0_name]\n        kpts1 = keypoints_dict[img1_name]\n        \n        matches = []\n        for i, (pt0, pt1) in enumerate(zip(mkpts0, mkpts1)):\n            idx0 = np.where((kpts0 == pt0).all(axis=1))[0]\n            idx1 = np.where((kpts1 == pt1).all(axis=1))[0]\n            if len(idx0) == 1 and len(idx1) == 1:\n                matches.append((idx0[0], idx1[0]))\n        matches = np.array(matches, dtype=np.uint32)\n        if len(matches) > 0:\n            db.add_matches(image_id0, image_id1, matches)\n    \n    db.commit()\n    db.close()\n    \n    # === Run COLMAP SfM ===\n    options = pycolmap.IncrementalPipelineOptions()\n    options.min_num_matches = 8\n    options.min_model_size = 2\n    options.init_num_trials = 50000\n    \n    start_time = time.time()\n    reconstructions = pycolmap.incremental_mapping(\n        database_path=str(database_path),\n        image_path=str(image_dir),\n        output_path=str(output_path),\n        options=options\n    )\n    print(f\"Incremental mapping (with bundle adjustment) took {time.time() - start_time:.2f} seconds.\")\n    \n    if reconstructions:\n        reconstruction = list(reconstructions.values())[0]\n        print(f\"Number of images registered: {reconstruction.num_reg_images()}\")\n        print(f\"Number of 3D points reconstructed: {reconstruction.num_points3D()}\")\n        \n        points3d = []\n        colors = []\n        for point3d_id in reconstruction.points3D:\n            point3d = reconstruction.points3D[point3d_id]\n            points3d.append(point3d.xyz)\n            colors.append(point3d.color / 255.0)\n        points3d = np.array(points3d)\n        colors = np.array(colors)\n        \n        camera_positions = []\n        for img_id in reconstruction.images:\n            img = reconstruction.images[img_id]\n            cam_from_world = img.cam_from_world\n            rotation = cam_from_world.rotation.matrix()\n            translation = cam_from_world.translation\n            position = -np.dot(rotation.T, translation)\n            camera_positions.append(position)\n        camera_positions = np.array(camera_positions)\n        \n        fig = plt.figure(figsize=(10, 8))\n        ax = fig.add_subplot(111, projection='3d')\n        ax.scatter(points3d[:, 0], points3d[:, 1], points3d[:, 2], c=colors, s=1, label='3D Points')\n        ax.scatter(camera_positions[:, 0], camera_positions[:, 1], camera_positions[:, 2], s=50, c='r', marker='^', label='Cameras')\n        ax.set_title(\"3D Point Cloud and Camera Positions (First 10 Images)\")\n        ax.set_xlabel(\"X\")\n        ax.set_ylabel(\"Y\")\n        ax.set_zlabel(\"Z\")\n        ax.legend()\n        plt.show()\n        plt.close('all')\n    else:\n        print(\"Reconstruction failed: No reconstructions returned.\")\n\nexcept AttributeError as e:\n    print(f\"pycolmap does not support importing features and matches: {e}\")\n    print(\"Please proceed without bundle adjustment or use a different environment with COLMAP support.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-06T15:22:50.532025Z","iopub.execute_input":"2025-04-06T15:22:50.532285Z","iopub.status.idle":"2025-04-06T15:23:03.708230Z","shell.execute_reply.started":"2025-04-06T15:22:50.532264Z","shell.execute_reply":"2025-04-06T15:23:03.707190Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 4th","metadata":{}},{"cell_type":"code","source":"# Step 3.10 (Step 3): Custom SfM Pipeline with SuperGlue Matches and COLMAP Bundle Adjustment (10 Images)\nimport pycolmap\nimport numpy as np\nimport matplotlib.pyplot as plt\nfrom pathlib import Path\nfrom PIL import Image\nimport cv2\nimport time\nimport os\nimport torch\nimport sys\nimport warnings\nwarnings.filterwarnings(\"ignore\", category=FutureWarning)\n\nsys.path.append('SuperGluePretrainedNetwork')\nfrom models.matching import Matching\nfrom models.utils import (frame2tensor, make_matching_plot)\n\n# === Set device ===\ndevice = torch.device('cuda' if torch.cuda.is_available() else 'cpu')\nprint(f\"Using device: {device}\")\n\n# === Define image paths ===\nstairs_dir = Path('/kaggle/input/image-matching-challenge-2025/train/stairs')\nimage_paths = [str(stairs_dir / img) for img in os.listdir(stairs_dir) if img.endswith(('.png', '.jpg', '.jpeg'))]\nprint(f\"Total images in stairs directory: {len(image_paths)}\")\nprint(\"First few image paths:\")\nfor path in image_paths[:5]:\n    print(path)\n\n# === Select first 10 images ===\nnum_images = 10\nselected_indices = list(range(min(num_images, len(image_paths))))\nprint(f\"Selected indices: {selected_indices}\")\n\n# === Setup workspace ===\nwork_dir = Path('/kaggle/working/colmap')\nwork_dir.mkdir(exist_ok=True)\nimage_dir = work_dir / 'images'\nimage_dir.mkdir(exist_ok=True)\n\n# === Copy all selected images ===\nselected_paths = [image_paths[i] for i in selected_indices]\nfor i, img_path in enumerate(selected_paths):\n    img = Image.open(img_path).convert('RGB')\n    img.save(image_dir / f\"image_{i}.png\")\n\n# === Initialize SuperGlue ===\nconfig = {\n    'superpoint': {\n        'nms_radius': 4,\n        'keypoint_threshold': 0.003,\n        'max_keypoints': 2048\n    },\n    'superglue': {\n        'weights': 'indoor',\n        'sinkhorn_iterations': 20,\n        'match_threshold': 0.05\n    }\n}\nmatching = Matching(config).eval().to(device)\n\n# === Run SuperPoint and SuperGlue ===\nkeypoints_dict = {}\nmatches_dict = {}\nimage_sizes = {}\n\nstart_time = time.time()\nimage_pairs = [(i, j) for i in range(len(selected_indices)) for j in range(i + 1, len(selected_indices))]\nfor idx0, idx1 in image_pairs:\n    img0 = cv2.imread(str(image_dir / f\"image_{idx0}.png\"), cv2.IMREAD_GRAYSCALE)\n    img1 = cv2.imread(str(image_dir / f\"image_{idx1}.png\"), cv2.IMREAD_GRAYSCALE)\n    \n    orig_size0 = img0.shape[::-1]\n    orig_size1 = img1.shape[::-1]\n    \n    img0 = cv2.resize(img0, (640, 480))\n    img1 = cv2.resize(img1, (640, 480))\n    \n    inp0 = frame2tensor(img0, device)\n    inp1 = frame2tensor(img1, device)\n    \n    pred = matching({'image0': inp0, 'image1': inp1})\n    pred = {k: v[0].detach().cpu().numpy() for k, v in pred.items()}\n    kpts0, kpts1 = pred['keypoints0'], pred['keypoints1']\n    matches, conf = pred['matches0'], pred['matching_scores0']\n    \n    valid = matches > -1\n    mkpts0 = kpts0[valid]\n    mkpts1 = kpts1[matches[valid]]\n    \n    scale0 = (orig_size0[0] / 640, orig_size0[1] / 480)\n    scale1 = (orig_size1[0] / 640, orig_size1[1] / 480)\n    kpts0[:, 0] *= scale0[0]\n    kpts0[:, 1] *= scale0[1]\n    kpts1[:, 0] *= scale1[0]\n    kpts1[:, 1] *= scale1[1]\n    mkpts0[:, 0] *= scale0[0]\n    mkpts0[:, 1] *= scale0[1]\n    mkpts1[:, 0] *= scale1[0]\n    mkpts1[:, 1] *= scale1[1]\n    \n    img0_name = f\"image_{idx0}.png\"\n    img1_name = f\"image_{idx1}.png\"\n    if img0_name not in keypoints_dict:\n        keypoints_dict[img0_name] = kpts0\n        image_sizes[img0_name] = orig_size0\n    if img1_name not in keypoints_dict:\n        keypoints_dict[img1_name] = kpts1\n        image_sizes[img1_name] = orig_size1\n    \n    matches_dict[(img0_name, img1_name)] = (mkpts0, mkpts1)\n    \n    print(f\"Matches between {img0_name} and {img1_name}: {len(mkpts0)}\")\n\nprint(f\"SuperGlue matching took {time.time() - start_time:.2f} seconds.\")\n\n# === Setup COLMAP database ===\ndatabase_path = work_dir / 'database.db'\noutput_path = work_dir / 'sparse'\noutput_path.mkdir(exist_ok=True)\n\n# === Import features and matches into COLMAP ===\ntry:\n    db = pycolmap.Database(str(database_path))\n    \n    camera_ids = {}\n    for img_name in keypoints_dict:\n        width, height = image_sizes[img_name]\n        focal_length = max(width, height) * 1.2\n        camera = pycolmap.Camera(\n            model=\"SIMPLE_PINHOLE\",\n            width=width,\n            height=height,\n            params=[focal_length, width / 2, height / 2]\n        )\n        camera_id = db.add_camera(camera)\n        camera_ids[img_name] = camera_id\n    \n    image_ids = {}\n    for img_name in keypoints_dict:\n        image_id = db.add_image(img_name, camera_ids[img_name])\n        image_ids[img_name] = image_id\n    \n    for img_name, kpts in keypoints_dict.items():\n        image_id = image_ids[img_name]\n        keypoints = np.zeros((len(kpts), 4), dtype=np.float32)\n        keypoints[:, :2] = kpts\n        keypoints[:, 2] = 1.0  # Dummy sigma\n        db.add_keypoints(image_id, keypoints)\n    \n    for (img0_name, img1_name), (mkpts0, mkpts1) in matches_dict.items():\n        image_id0 = image_ids[img0_name]\n        image_id1 = image_ids[img1_name]\n        kpts0 = keypoints_dict[img0_name]\n        kpts1 = keypoints_dict[img1_name]\n        \n        matches = []\n        for i, (pt0, pt1) in enumerate(zip(mkpts0, mkpts1)):\n            idx0 = np.where((kpts0 == pt0).all(axis=1))[0]\n            idx1 = np.where((kpts1 == pt1).all(axis=1))[0]\n            if len(idx0) == 1 and len(idx1) == 1:\n                matches.append((idx0[0], idx1[0]))\n        matches = np.array(matches, dtype=np.uint32)\n        if len(matches) > 0:\n            db.add_matches(image_id0, image_id1, matches)\n    \n    db.commit()\n    db.close()\n    \n    # === Run COLMAP SfM ===\n    options = pycolmap.IncrementalPipelineOptions()\n    options.min_num_matches = 8\n    options.min_model_size = 2\n    options.init_num_trials = 50000\n    \n    start_time = time.time()\n    reconstructions = pycolmap.incremental_mapping(\n        database_path=str(database_path),\n        image_path=str(image_dir),\n        output_path=str(output_path),\n        options=options\n    )\n    print(f\"Incremental mapping (with bundle adjustment) took {time.time() - start_time:.2f} seconds.\")\n    \n    if reconstructions:\n        reconstruction = list(reconstructions.values())[0]\n        print(f\"Number of images registered: {reconstruction.num_reg_images()}\")\n        print(f\"Number of 3D points reconstructed: {reconstruction.num_points3D()}\")\n        \n        points3d = []\n        colors = []\n        for point3d_id in reconstruction.points3D:\n            point3d = reconstruction.points3D[point3d_id]\n            points3d.append(point3d.xyz)\n            colors.append(point3d.color / 255.0)\n        points3d = np.array(points3d)\n        colors = np.array(colors)\n        \n        camera_positions = []\n        for img_id in reconstruction.images:\n            img = reconstruction.images[img_id]\n            cam_from_world = img.cam_from_world\n            rotation = cam_from_world.rotation.matrix()\n            translation = cam_from_world.translation\n            position = -np.dot(rotation.T, translation)\n            camera_positions.append(position)\n        camera_positions = np.array(camera_positions)\n        \n        fig = plt.figure(figsize=(10, 8))\n        ax = fig.add_subplot(111, projection='3d')\n        ax.scatter(points3d[:, 0], points3d[:, 1], points3d[:, 2], c=colors, s=1, label='3D Points')\n        ax.scatter(camera_positions[:, 0], camera_positions[:, 1], camera_positions[:, 2], s=50, c='r', marker='^', label='Cameras')\n        ax.set_title(\"3D Point Cloud and Camera Positions (First 10 Images)\")\n        ax.set_xlabel(\"X\")\n        ax.set_ylabel(\"Y\")\n        ax.set_zlabel(\"Z\")\n        ax.legend()\n        plt.show()\n        plt.close('all')\n    else:\n        print(\"Reconstruction failed: No reconstructions returned.\")\n\nexcept AttributeError as e:\n    print(f\"pycolmap does not support importing features and matches: {e}\")\n    print(\"Please proceed without bundle adjustment or use a different environment with COLMAP support.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-06T15:23:03.709921Z","iopub.execute_input":"2025-04-06T15:23:03.710327Z","iopub.status.idle":"2025-04-06T15:23:16.887907Z","shell.execute_reply.started":"2025-04-06T15:23:03.710287Z","shell.execute_reply":"2025-04-06T15:23:16.886891Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 5th","metadata":{}},{"cell_type":"code","source":"# Step 5: Custom SfM Pipeline for Image Matching Challenge 2025 (All Images, Format Output for Submission)\nimport numpy as np\nimport matplotlib.pyplot as plt\nfrom pathlib import Path\nfrom PIL import Image\nimport cv2\nimport time\nimport os\nimport torch\nimport sys\nimport warnings\nfrom scipy.optimize import least_squares\nwarnings.filterwarnings(\"ignore\", category=FutureWarning)\n\nsys.path.append('SuperGluePretrainedNetwork')\nfrom models.matching import Matching\nfrom models.utils import (frame2tensor, make_matching_plot)\n\n# === Set device ===\ndevice = torch.device('cuda' if torch.cuda.is_available() else 'cpu')\nprint(f\"Using device: {device}\")\n\n# === Define image paths ===\nstairs_dir = Path('/kaggle/input/image-matching-challenge-2025/train/stairs')\nimage_paths = [str(stairs_dir / img) for img in os.listdir(stairs_dir) if img.endswith(('.png', '.jpg', '.jpeg'))]\nprint(f\"Total images in stairs directory: {len(image_paths)}\")\nprint(\"First few image paths:\")\nfor path in image_paths[:5]:\n    print(path)\n\n# === Select all images ===\nnum_images = len(image_paths)\nselected_indices = list(range(num_images))\nprint(f\"Selected indices: {selected_indices}\")\n\n# === Define clusters (5 images per cluster, with overlap of 2 images) ===\ncluster_size = 5\noverlap = 2\nclusters = []\nfor i in range(0, len(selected_indices), cluster_size - overlap):\n    cluster_indices = selected_indices[i:i + cluster_size]\n    if len(cluster_indices) >= 2:\n        clusters.append(cluster_indices)\nprint(f\"Clusters: {clusters}\")\n\n# === Setup workspace ===\nwork_dir = Path('/kaggle/working/custom_sfm')\nwork_dir.mkdir(exist_ok=True)\nimage_dir = work_dir / 'images'\nimage_dir.mkdir(exist_ok=True)\noutput_dir = work_dir / 'output'\noutput_dir.mkdir(exist_ok=True)\n\n# === Copy all selected images ===\nselected_paths = [image_paths[i] for i in selected_indices]\nfor i, img_path in enumerate(selected_paths):\n    img = Image.open(img_path).convert('RGB')\n    img.save(image_dir / f\"image_{i}.png\")\n\n# === Initialize SuperGlue ===\nconfig = {\n    'superpoint': {\n        'nms_radius': 4,\n        'keypoint_threshold': 0.003,\n        'max_keypoints': 2048\n    },\n    'superglue': {\n        'weights': 'indoor',\n        'sinkhorn_iterations': 20,\n        'match_threshold': 0.03\n    }\n}\nmatching = Matching(config).eval().to(device)\n\n# === Function to run SfM on a single cluster ===\ndef run_sfm_on_cluster(cluster_indices, image_dir, matching, device, cluster_idx, clusters):\n    keypoints_dict = {}\n    matches_dict = {}\n    image_sizes = {}\n    point_to_images = []  # Track which images see each 3D point\n    \n    print(f\"Processing cluster with indices: {cluster_indices}\")\n    \n    # Run SuperPoint and SuperGlue\n    start_time = time.time()\n    image_pairs = [(i, j) for i in range(len(cluster_indices)) for j in range(i + 1, len(cluster_indices))]\n    edges = []\n    for idx0, idx1 in image_pairs:\n        global_idx0, global_idx1 = cluster_indices[idx0], cluster_indices[idx1]\n        img0 = cv2.imread(str(image_dir / f\"image_{global_idx0}.png\"), cv2.IMREAD_GRAYSCALE)\n        img1 = cv2.imread(str(image_dir / f\"image_{global_idx1}.png\"), cv2.IMREAD_GRAYSCALE)\n        \n        orig_size0 = img0.shape[::-1]\n        orig_size1 = img1.shape[::-1]\n        \n        img0 = cv2.resize(img0, (640, 480))\n        img1 = cv2.resize(img1, (640, 480))\n        \n        inp0 = frame2tensor(img0, device)\n        inp1 = frame2tensor(img1, device)\n        \n        pred = matching({'image0': inp0, 'image1': inp1})\n        pred = {k: v[0].detach().cpu().numpy() for k, v in pred.items()}\n        kpts0, kpts1 = pred['keypoints0'], pred['keypoints1']\n        matches, conf = pred['matches0'], pred['matching_scores0']\n        \n        valid = matches > -1\n        mkpts0 = kpts0[valid]\n        mkpts1 = kpts1[matches[valid]]\n        \n        scale0 = (orig_size0[0] / 640, orig_size0[1] / 480)\n        scale1 = (orig_size1[0] / 640, orig_size1[1] / 480)\n        kpts0[:, 0] *= scale0[0]\n        kpts0[:, 1] *= scale0[1]\n        kpts1[:, 0] *= scale1[0]\n        kpts1[:, 1] *= scale1[1]\n        mkpts0[:, 0] *= scale0[0]\n        mkpts0[:, 1] *= scale0[1]\n        mkpts1[:, 0] *= scale1[0]\n        mkpts1[:, 1] *= scale1[1]\n        \n        img0_name = f\"image_{global_idx0}.png\"\n        img1_name = f\"image_{global_idx1}.png\"\n        if img0_name not in keypoints_dict:\n            keypoints_dict[img0_name] = kpts0\n            image_sizes[img0_name] = orig_size0\n        if img1_name not in keypoints_dict:\n            keypoints_dict[img1_name] = kpts1\n            image_sizes[img1_name] = orig_size1\n        \n        matches_dict[(img0_name, img1_name)] = (mkpts0, mkpts1)\n        \n        priority = 0\n        if cluster_idx < len(clusters) - 1:\n            overlap_indices = clusters[cluster_idx][-overlap:]\n            if global_idx0 in overlap_indices or global_idx1 in overlap_indices:\n                priority = 1\n        \n        edges.append((len(mkpts0), idx0, idx1, priority))\n        \n        print(f\"Matches between {img0_name} and {img1_name}: {len(mkpts0)}\")\n    \n    print(f\"SuperGlue matching for cluster took {time.time() - start_time:.2f} seconds.\")\n    \n    # Sort edges\n    edges.sort(key=lambda x: (-x[3], -x[0]))\n    parent = list(range(len(cluster_indices)))\n    rank = [0] * len(cluster_indices)\n    \n    def find(x):\n        if parent[x] != x:\n            parent[x] = find(parent[x])\n        return parent[x]\n    \n    def union(x, y):\n        px, py = find(x), find(y)\n        if px == py:\n            return\n        if rank[px] < rank[py]:\n            px, py = py, px\n        parent[py] = px\n        if rank[px] == rank[py]:\n            rank[px] += 1\n    \n    selected_pairs = []\n    for num_matches, idx0, idx1, _ in edges:\n        if num_matches < 8:\n            continue\n        if find(idx0) != find(idx1):\n            union(idx0, idx1)\n            selected_pairs.append((idx0, idx1))\n    \n    if cluster_idx < len(clusters) - 1:\n        overlap_indices = clusters[cluster_idx][-overlap:]\n        for overlap_idx in overlap_indices:\n            overlap_local_idx = cluster_indices.index(overlap_idx)\n            if overlap_local_idx not in {idx0 for idx0, _ in selected_pairs} and overlap_local_idx not in {idx1 for _, idx1 in selected_pairs}:\n                best_pair = None\n                best_num_matches = 0\n                for num_matches, idx0, idx1, _ in edges:\n                    if num_matches < 8:\n                        continue\n                    if idx0 == overlap_local_idx or idx1 == overlap_local_idx:\n                        if num_matches > best_num_matches:\n                            best_num_matches = num_matches\n                            best_pair = (idx0, idx1)\n                if best_pair:\n                    idx0, idx1 = best_pair\n                    if find(idx0) != find(idx1):\n                        union(idx0, idx1)\n                        selected_pairs.append((idx0, idx1))\n    \n    if not selected_pairs:\n        print(\"No pairs with sufficient matches found in cluster.\")\n        return np.array([]), np.array([]), np.array([]), {}, [], {}\n    \n    # Initialize poses\n    idx0, idx1 = selected_pairs[0]\n    global_idx0, global_idx1 = cluster_indices[idx0], cluster_indices[idx1]\n    img0_name = f\"image_{global_idx0}.png\"\n    img1_name = f\"image_{global_idx1}.png\"\n    \n    mkpts0, mkpts1 = matches_dict[(img0_name, img1_name)]\n    mkpts0 = mkpts0.astype(np.float32)\n    mkpts1 = mkpts1.astype(np.float32)\n    \n    image_width, image_height = image_sizes[img0_name]\n    focal_length = max(image_width, image_height) * 1.2\n    cx, cy = image_width / 2, image_height / 2\n    K = np.array([\n        [focal_length, 0, cx],\n        [0, focal_length, cy],\n        [0, 0, 1]\n    ], dtype=np.float32)\n    \n    E, mask = cv2.findEssentialMat(mkpts0, mkpts1, K, method=cv2.RANSAC, prob=0.999, threshold=1.0)\n    if E is None or E.shape != (3, 3):\n        print(f\"Failed to estimate essential matrix for initial pair {img0_name} and {img1_name}\")\n        return np.array([]), np.array([]), np.array([]), {}, [], {}\n    \n    _, R, t, mask = cv2.recoverPose(E, mkpts0, mkpts1, K)\n    if R is None or t is None:\n        print(f\"Failed to recover pose for initial pair {img0_name} and {img1_name}\")\n        return np.array([]), np.array([]), np.array([]), {}, [], {}\n    \n    R = R.astype(np.float32)\n    t = t.astype(np.float32).flatten()\n    \n    poses = {img0_name: (np.eye(3, dtype=np.float32), np.zeros(3, dtype=np.float32))}\n    poses[img1_name] = (R, t)\n    registered_indices = {idx0, idx1}\n    \n    # Process remaining pairs\n    start_time = time.time()\n    for idx0, idx1 in selected_pairs[1:]:\n        global_idx0, global_idx1 = cluster_indices[idx0], cluster_indices[idx1]\n        img0_name = f\"image_{global_idx0}.png\"\n        img1_name = f\"image_{global_idx1}.png\"\n        \n        if img0_name in poses and img1_name in poses:\n            continue\n        elif img0_name not in poses and img1_name not in poses:\n            continue\n        \n        if img1_name in poses and img0_name not in poses:\n            img0_name, img1_name = img1_name, img0_name\n            idx0, idx1 = idx1, idx0\n        \n        key1 = (img0_name, img1_name)\n        key2 = (img1_name, img0_name)\n        if key1 in matches_dict:\n            mkpts0, mkpts1 = matches_dict[key1]\n        elif key2 in matches_dict:\n            mkpts1, mkpts0 = matches_dict[key2]\n        else:\n            print(f\"Matches not found for pair ({img0_name}, {img1_name})\")\n            continue\n        \n        mkpts0 = mkpts0.astype(np.float32)\n        mkpts1 = mkpts1.astype(np.float32)\n        \n        E, mask = cv2.findEssentialMat(mkpts0, mkpts1, K, method=cv2.RANSAC, prob=0.999, threshold=1.0)\n        if E is None or E.shape != (3, 3):\n            print(f\"Failed to estimate essential matrix between {img0_name} and {img1_name}\")\n            continue\n        \n        _, R, t, mask = cv2.recoverPose(E, mkpts0, mkpts1, K)\n        if R is None or t is None:\n            print(f\"Failed to recover pose between {img0_name} and {img1_name}\")\n            continue\n        \n        R = R.astype(np.float32)\n        t = t.astype(np.float32).flatten()\n        \n        R0, t0 = poses[img0_name]\n        R = R0 @ R\n        t = R0 @ t + t0\n        \n        poses[img1_name] = (R, t)\n        registered_indices.add(idx1)\n    \n    print(f\"Pose estimation took {time.time() - start_time:.2f} seconds.\")\n    \n    # Triangulate 3D points\n    start_time = time.time()\n    points3d = []\n    colors = []\n    for idx0, idx1 in image_pairs:\n        global_idx0, global_idx1 = cluster_indices[idx0], cluster_indices[idx1]\n        img0_name = f\"image_{global_idx0}.png\"\n        img1_name = f\"image_{global_idx1}.png\"\n        \n        if img0_name not in poses or img1_name not in poses:\n            continue\n        \n        key1 = (img0_name, img1_name)\n        key2 = (img1_name, img0_name)\n        if key1 in matches_dict:\n            mkpts0, mkpts1 = matches_dict[key1]\n        elif key2 in matches_dict:\n            mkpts1, mkpts0 = matches_dict[key2]\n        else:\n            print(f\"Matches not found for pair ({img0_name}, {img1_name}) during triangulation\")\n            continue\n        \n        if len(mkpts0) != len(mkpts1):\n            print(f\"Mismatch in number of points between {img0_name} and {img1_name}: {len(mkpts0)} vs {len(mkpts1)}\")\n            continue\n        \n        mkpts0 = mkpts0.astype(np.float32)\n        mkpts1 = mkpts1.astype(np.float32)\n        \n        R0, t0 = poses[img0_name]\n        R1, t1 = poses[img1_name]\n        \n        P0 = K @ np.hstack((R0, t0.reshape(3, 1)))\n        P1 = K @ np.hstack((R1, t1.reshape(3, 1)))\n        P0 = P0.astype(np.float32)\n        P1 = P1.astype(np.float32)\n        \n        pts0 = mkpts0.T.astype(np.float32)\n        pts1 = mkpts1.T.astype(np.float32)\n        \n        if pts0.shape[1] == 0 or pts1.shape[1] == 0:\n            print(f\"Skipping triangulation due to zero matches between {img0_name} and {img1_name}\")\n            continue\n        \n        points4d = cv2.triangulatePoints(P0, P1, pts0, pts1)\n        points3d_h = points4d[:3] / points4d[3]\n        points3d.extend(points3d_h.T)\n        \n        for _ in range(pts0.shape[1]):\n            point_to_images.append([(img0_name, pts0[:, _]), (img1_name, pts1[:, _])])\n        \n        img0 = cv2.imread(str(image_dir / img0_name))\n        img0 = cv2.resize(img0, image_sizes[img0_name])\n        for pt in mkpts0:\n            x, y = int(pt[0]), int(pt[1])\n            if 0 <= x < img0.shape[1] and 0 <= y < img0.shape[0]:\n                colors.append(img0[y, x] / 255.0)\n            else:\n                colors.append([0, 0, 0])\n    \n    points3d = np.array(points3d)\n    colors = np.array(colors)\n    \n    valid = np.all(np.isfinite(points3d), axis=1) & (np.abs(points3d) < 1e5).all(axis=1)\n    points3d = points3d[valid]\n    colors = colors[valid]\n    point_to_images = [pt for pt, v in zip(point_to_images, valid) if v]\n    \n    print(f\"Triangulation took {time.time() - start_time:.2f} seconds.\")\n    \n    # Lightweight Bundle Adjustment\n    start_time = time.time()\n    def project(points3d, R, t, K):\n        points3d = points3d.T\n        points = R @ points3d + t.reshape(3, 1)\n        points = K @ points\n        points = points[:2] / points[2]\n        return points.T\n    \n    def reprojection_error(params, points3d, observations, K, img_names):\n        num_cameras = len(img_names)\n        num_points = len(points3d)\n        \n        # Unpack parameters (only optimize translations and a single focal length)\n        translations = params[:num_cameras * 3].reshape(num_cameras, 3)\n        focal_length = params[-1]\n        K_opt = K.copy()\n        K_opt[0, 0] = K_opt[1, 1] = focal_length\n        \n        errors = []\n        for i, img_name in enumerate(img_names):\n            R, _ = poses[img_name]\n            t = translations[i]\n            for j, obs in enumerate(observations):\n                for img_obs, pt in obs:\n                    if img_obs == img_name:\n                        proj = project(points3d[j:j+1], R, t, K_opt)\n                        errors.append(proj[0] - pt)\n        \n        return np.concatenate(errors)\n    \n    # Pack parameters for optimization (only optimize translations and focal length)\n    img_names = list(poses.keys())\n    translations = np.array([t for _, t in poses.values()])\n    params = np.hstack((translations.ravel(), focal_length))\n    \n    # Run bundle adjustment with limited iterations\n    result = least_squares(\n        reprojection_error,\n        params,\n        args=(points3d, point_to_images, K, img_names),\n        max_nfev=10,\n        ftol=1e-4,\n        xtol=1e-4\n    )\n    optimized_params = result.x\n    \n    # Unpack optimized parameters\n    num_cameras = len(img_names)\n    translations = optimized_params[:num_cameras * 3].reshape(num_cameras, 3)\n    for i, img_name in enumerate(img_names):\n        R, _ = poses[img_name]\n        poses[img_name] = (R, translations[i])\n    \n    print(f\"Bundle adjustment took {time.time() - start_time:.2f} seconds.\")\n    \n    # Camera positions\n    camera_positions = []\n    for img_name in poses:\n        R, t = poses[img_name]\n        pos = -R.T @ t\n        camera_positions.append(pos)\n    camera_positions = np.array(camera_positions)\n    \n    # Visualize\n    fig = plt.figure(figsize=(10, 8))\n    ax = fig.add_subplot(111, projection='3d')\n    ax.scatter(points3d[:, 0], points3d[:, 1], points3d[:, 2], c=colors, s=1, label='3D Points')\n    ax.scatter(camera_positions[:, 0], camera_positions[:, 1], camera_positions[:, 2], s=50, c='r', marker='^', label='Cameras')\n    ax.set_title(f\"3D Point Cloud and Camera Positions (Cluster {cluster_indices[0]}-{cluster_indices[-1]})\")\n    ax.set_xlabel(\"X\")\n    ax.set_ylabel(\"Y\")\n    ax.set_zlabel(\"Z\")\n    ax.legend()\n    plt.show()\n    plt.close('all')\n    \n    print(f\"Cluster {cluster_indices[0]}-{cluster_indices[-1]}: {len(points3d)} 3D points, {len(poses)} cameras\")\n    \n    return points3d, colors, camera_positions, poses, point_to_images, K\n\n# === Process each cluster ===\nall_points3d = []\nall_colors = []\nall_camera_positions = []\nall_poses = {}\nall_point_to_images = []\nall_Ks = {}\n\nstart_time = time.time()\nfor cluster_idx, cluster_indices in enumerate(clusters):\n    points3d, colors, camera_positions, poses, point_to_images, K = run_sfm_on_cluster(cluster_indices, image_dir, matching, device, cluster_idx, clusters)\n    all_points3d.append(points3d)\n    all_colors.append(colors)\n    all_camera_positions.append(camera_positions)\n    all_poses.update(poses)\n    all_point_to_images.append(point_to_images)\n    all_Ks.update({img_name: K for img_name in poses})\n\n# Merge reconstructions\nif not all_points3d or len(all_points3d[0]) == 0:\n    print(\"No points reconstructed in the first cluster. Cannot proceed with merging.\")\nelse:\n    merged_points3d = all_points3d[0]\n    merged_colors = all_colors[0]\n    merged_camera_positions = all_camera_positions[0]\n    merged_point_to_images = all_point_to_images[0]\n\n    for i in range(1, len(clusters)):\n        overlap_image = f\"image_{clusters[i-1][-1]}.png\"\n        if overlap_image not in all_poses:\n            print(f\"Cannot align clusters {i-1} and {i} using overlapping image {overlap_image}.\")\n            continue\n        \n        R0_prev, t0_prev = all_poses[overlap_image]\n        R0_curr, t0_curr = all_poses[overlap_image]\n        \n        R_align = R0_prev @ np.linalg.inv(R0_curr)\n        t_align = t0_prev - R_align @ t0_curr\n        \n        points3d_curr = all_points3d[i]\n        if len(points3d_curr) == 0:\n            print(f\"Cluster {i} has no points to merge\")\n            continue\n        points3d_curr = (R_align @ points3d_curr.T).T + t_align\n        all_points3d[i] = points3d_curr\n        merged_points3d = np.vstack((merged_points3d, points3d_curr))\n        merged_colors = np.vstack((merged_colors, all_colors[i]))\n        merged_point_to_images.extend(all_point_to_images[i])\n        \n        camera_positions_curr = all_camera_positions[i]\n        if len(camera_positions_curr) == 0:\n            print(f\"Cluster {i} has no camera positions to merge\")\n            continue\n        camera_positions_curr = (R_align @ camera_positions_curr.T).T + t_align\n        all_camera_positions[i] = camera_positions_curr\n        merged_camera_positions = np.vstack((merged_camera_positions, camera_positions_curr))\n        \n        for img_name in list(all_poses.keys()):\n            if img_name in all_poses:\n                R, t = all_poses[img_name]\n                R = R_align @ R\n                t = R_align @ t + t_align\n                all_poses[img_name] = (R, t)\n\n    # Remove duplicates in camera positions\n    unique_camera_positions = []\n    seen_images = set()\n    for i, pos in enumerate(merged_camera_positions):\n        img_idx = i % len(selected_indices)\n        img_name = f\"image_{img_idx}.png\"\n        if img_name not in seen_images:\n            unique_camera_positions.append(pos)\n            seen_images.add(img_name)\n    unique_camera_positions = np.array(unique_camera_positions)\n\n    # Visualize merged reconstruction\n    fig = plt.figure(figsize=(10, 8))\n    ax = fig.add_subplot(111, projection='3d')\n    ax.scatter(merged_points3d[:, 0], merged_points3d[:, 1], merged_points3d[:, 2], c=merged_colors, s=1, label='3D Points')\n    ax.scatter(unique_camera_positions[:, 0], unique_camera_positions[:, 1], unique_camera_positions[:, 2], s=50, c='r', marker='^', label='Cameras')\n    ax.set_title(\"Merged 3D Point Cloud and Camera Positions (All Images)\")\n    ax.set_xlabel(\"X\")\n    ax.set_ylabel(\"Y\")\n    ax.set_zlabel(\"Z\")\n    ax.legend()\n    plt.show()\n    plt.close('all')\n\n    print(f\"Total runtime: {time.time() - start_time:.2f} seconds\")\n    print(f\"Total number of 3D points reconstructed: {len(merged_points3d)}\")\n    print(f\"Total number of cameras estimated: {len(unique_camera_positions)}\")\n\n    # === Save output in COLMAP format ===\n    # cameras.txt\n    with open(output_dir / 'cameras.txt', 'w') as f:\n        f.write(\"# Camera list with one line of data per camera:\\n\")\n        f.write(\"#   CAMERA_ID, MODEL, WIDTH, HEIGHT, PARAMS[]\\n\")\n        f.write(\"# Number of cameras: {}\\n\".format(len(all_poses)))\n        for i, img_name in enumerate(sorted(all_poses.keys())):\n            width, height = all_Ks[img_name][0, 2] * 2, all_Ks[img_name][1, 2] * 2\n            focal_length = all_Ks[img_name][0, 0]\n            cx, cy = all_Ks[img_name][0, 2], all_Ks[img_name][1, 2]\n            f.write(f\"{i+1} SIMPLE_PINHOLE {int(width)} {int(height)} {focal_length} {cx} {cy}\\n\")\n\n    # images.txt\n    with open(output_dir / 'images.txt', 'w') as f:\n        f.write(\"# Image list with two lines of data per image:\\n\")\n        f.write(\"#   IMAGE_ID, QW, QX, QY, QZ, TX, TY, TZ, CAMERA_ID, NAME\\n\")\n        f.write(\"#   POINTS2D[] as (X, Y, POINT3D_ID)\\n\")\n        f.write(\"# Number of images: {}\\n\".format(len(all_poses)))\n        for i, img_name in enumerate(sorted(all_poses.keys())):\n            R, t = all_poses[img_name]\n            # Convert rotation matrix to quaternion\n            w = np.sqrt(1.0 + R[0, 0] + R[1, 1] + R[2, 2]) / 2.0\n            if w < 1e-10:\n                w = 1e-10\n            x = (R[2, 1] - R[1, 2]) / (4 * w)\n            y = (R[0, 2] - R[2, 0]) / (4 * w)\n            z = (R[1, 0] - R[0, 1]) / (4 * w)\n            f.write(f\"{i+1} {w} {x} {y} {z} {t[0]} {t[1]} {t[2]} {i+1} {img_name}\\n\")\n            # Add 2D-3D correspondences (simplified, assuming we track them)\n            f.write(\"\\n\")  # We'll leave POINTS2D empty for now; can be added if needed\n\n    # points3D.txt\n    with open(output_dir / 'points3D.txt', 'w') as f:\n        f.write(\"# 3D point list with one line of data per point:\\n\")\n        f.write(\"#   POINT3D_ID, X, Y, Z, R, G, B, ERROR, TRACK[] as (IMAGE_ID, POINT2D_IDX)\\n\")\n        f.write(\"# Number of points: {}\\n\".format(len(merged_points3d)))\n        for i, (point, color) in enumerate(zip(merged_points3d, merged_colors)):\n            rgb = (color * 255).astype(int)\n            f.write(f\"{i+1} {point[0]} {point[1]} {point[2]} {rgb[0]} {rgb[1]} {rgb[2]} 1.0\\n\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-06T15:23:16.888990Z","iopub.execute_input":"2025-04-06T15:23:16.889255Z","iopub.status.idle":"2025-04-06T15:24:47.490291Z","shell.execute_reply.started":"2025-04-06T15:23:16.889234Z","shell.execute_reply":"2025-04-06T15:24:47.489341Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 6th","metadata":{}},{"cell_type":"code","source":"# Step 6: Process All Scenes in Test Set and Generate Submission for Image Matching Challenge 2025\nimport numpy as np\nfrom pathlib import Path\nimport cv2\nimport time\nimport torch\nimport sys\nimport warnings\nfrom scipy.optimize import least_squares\nwarnings.filterwarnings(\"ignore\", category=FutureWarning)\n\nsys.path.append('SuperGluePretrainedNetwork')\nfrom models.matching import Matching\nfrom models.utils import (frame2tensor, make_matching_plot)\n\n# === Function to save point cloud as PLY ===\ndef save_ply(points3d, output_path):\n    header = f\"\"\"ply\nformat ascii 1.0\nelement vertex {len(points3d)}\nproperty float x\nproperty float y\nproperty float z\nproperty uchar red\nproperty uchar green\nproperty uchar blue\nend_header\n\"\"\"\n    colors = np.random.randint(0, 255, size=(len(points3d), 3))\n    with open(output_path, 'w') as f:\n        f.write(header)\n        for point, color in zip(points3d, colors):\n            f.write(f\"{point[0]} {point[1]} {point[2]} {color[0]} {color[1]} {color[2]}\\n\")\n\n# === Set device ===\ndevice = torch.device('cuda' if torch.cuda.is_available() else 'cpu')\nprint(f\"Using device: {device}\")\n\n# === Initialize SuperGlue ===\nconfig = {\n    'superpoint': {\n        'nms_radius': 4,\n        'keypoint_threshold': 0.003,\n        'max_keypoints': 8192  # Increased\n    },\n    'superglue': {\n        'weights': 'indoor',  # Test 'outdoor' if needed\n        'sinkhorn_iterations': 20,\n        'match_threshold': 0.01\n    }\n}\nmatching = Matching(config).eval().to(device)\nprint(\"Loaded SuperPoint model\")\nprint(\"Loaded SuperGlue model (\\\"indoor\\\" weights)\")\n\n# === Function to run SfM on a single cluster ===\ndef run_sfm_on_cluster(cluster_indices, image_dir, image_paths, matching, device, cluster_idx, clusters):\n    keypoints_dict = {}\n    matches_dict = {}\n    image_sizes = {}\n    point_to_images = []\n    \n    print(f\"Processing cluster with indices: {cluster_indices}\")\n    \n    start_time = time.time()\n    image_pairs = [(i, j) for i in range(len(cluster_indices)) for j in range(i + 1, len(cluster_indices))]\n    edges = []\n    for idx0, idx1 in image_pairs:\n        global_idx0, global_idx1 = cluster_indices[idx0], cluster_indices[idx1]\n        img0 = cv2.imread(image_paths[global_idx0], cv2.IMREAD_GRAYSCALE)\n        img1 = cv2.imread(image_paths[global_idx1], cv2.IMREAD_GRAYSCALE)\n        \n        # Apply histogram equalization\n        img0 = cv2.equalizeHist(img0)\n        img1 = cv2.equalizeHist(img1)\n        \n        orig_size0 = img0.shape[::-1]\n        orig_size1 = img1.shape[::-1]\n        \n        img0 = cv2.resize(img0, (640, 480))\n        img1 = cv2.resize(img1, (640, 480))\n        \n        inp0 = frame2tensor(img0, device)\n        inp1 = frame2tensor(img1, device)\n        \n        pred = matching({'image0': inp0, 'image1': inp1})\n        pred = {k: v[0].detach().cpu().numpy() for k, v in pred.items()}\n        kpts0, kpts1 = pred['keypoints0'], pred['keypoints1']\n        matches, conf = pred['matches0'], pred['matching_scores0']\n        \n        valid = matches > -1\n        mkpts0 = kpts0[valid]\n        mkpts1 = kpts1[matches[valid]]\n        \n        scale0 = (orig_size0[0] / 640, orig_size0[1] / 480)\n        scale1 = (orig_size1[0] / 640, orig_size1[1] / 480)\n        kpts0[:, 0] *= scale0[0]\n        kpts0[:, 1] *= scale0[1]\n        kpts1[:, 0] *= scale1[0]\n        kpts1[:, 1] *= scale1[1]\n        mkpts0[:, 0] *= scale0[0]\n        mkpts0[:, 1] *= scale0[1]\n        mkpts1[:, 0] *= scale1[0]\n        mkpts1[:, 1] *= scale1[1]\n        \n        img0_name = Path(image_paths[global_idx0]).name\n        img1_name = Path(image_paths[global_idx1]).name\n        if img0_name not in keypoints_dict:\n            keypoints_dict[img0_name] = kpts0\n            image_sizes[img0_name] = orig_size0\n        if img1_name not in keypoints_dict:\n            keypoints_dict[img1_name] = kpts1\n            image_sizes[img1_name] = orig_size1\n        \n        matches_dict[(img0_name, img1_name)] = (mkpts0, mkpts1)\n        \n        priority = 0\n        if cluster_idx < len(clusters) - 1:\n            overlap_indices = clusters[cluster_idx][-overlap:]\n            if global_idx0 in overlap_indices or global_idx1 in overlap_indices:\n                priority = 1\n        \n        edges.append((len(mkpts0), idx0, idx1, priority))\n        \n        print(f\"Matches between {img0_name} and {img1_name}: {len(mkpts0)}\")\n    \n    print(f\"SuperGlue matching for cluster took {time.time() - start_time:.2f} seconds.\")\n    \n    edges.sort(key=lambda x: (-x[3], -x[0]))\n    parent = list(range(len(cluster_indices)))\n    rank = [0] * len(cluster_indices)\n    \n    def find(x):\n        if parent[x] != x:\n            parent[x] = find(parent[x])\n        return parent[x]\n    \n    def union(x, y):\n        px, py = find(x), find(y)\n        if px == py:\n            return\n        if rank[px] < rank[py]:\n            px, py = py, px\n        parent[py] = px\n        if rank[px] == rank[py]:\n            rank[px] += 1\n    \n    selected_pairs = []\n    for num_matches, idx0, idx1, _ in edges:\n        if num_matches < 8:\n            continue\n        if find(idx0) != find(idx1):\n            union(idx0, idx1)\n            selected_pairs.append((idx0, idx1))\n    \n    if cluster_idx < len(clusters) - 1:\n        overlap_indices = clusters[cluster_idx][-overlap:]\n        for overlap_idx in overlap_indices:\n            overlap_local_idx = cluster_indices.index(overlap_idx)\n            if overlap_local_idx not in {idx0 for idx0, _ in selected_pairs} and overlap_local_idx not in {idx1 for _, idx1 in selected_pairs}:\n                best_pair = None\n                best_num_matches = 0\n                for num_matches, idx0, idx1, _ in edges:\n                    if num_matches < 8:\n                        continue\n                    if idx0 == overlap_local_idx or idx1 == overlap_local_idx:\n                        if num_matches > best_num_matches:\n                            best_num_matches = num_matches\n                            best_pair = (idx0, idx1)\n                if best_pair:\n                    idx0, idx1 = best_pair\n                    if find(idx0) != find(idx1):\n                        union(idx0, idx1)\n                        selected_pairs.append((idx0, idx1))\n    \n    if not selected_pairs:\n        print(\"No pairs with sufficient matches found in cluster.\")\n        return {}, [], {}\n    \n    idx0, idx1 = selected_pairs[0]\n    global_idx0, global_idx1 = cluster_indices[idx0], cluster_indices[idx1]\n    img0_name = Path(image_paths[global_idx0]).name\n    img1_name = Path(image_paths[global_idx1]).name\n    \n    mkpts0, mkpts1 = matches_dict[(img0_name, img1_name)]\n    mkpts0 = mkpts0.astype(np.float32)\n    mkpts1 = mkpts1.astype(np.float32)\n    \n    image_width, image_height = image_sizes[img0_name]\n    focal_length = max(image_width, image_height) * 1.2\n    cx, cy = image_width / 2, image_height / 2\n    K = np.array([\n        [focal_length, 0, cx],\n        [0, focal_length, cy],\n        [0, 0, 1]\n    ], dtype=np.float32)\n    \n    E, mask = cv2.findEssentialMat(mkpts0, mkpts1, K, method=cv2.RANSAC, prob=0.999, threshold=1.0)\n    if E is None or E.shape != (3, 3):\n        print(f\"Failed to estimate essential matrix for initial pair {img0_name} and {img1_name}\")\n        return {}, [], {}\n    \n    _, R, t, mask = cv2.recoverPose(E, mkpts0, mkpts1, K)\n    if R is None or t is None:\n        print(f\"Failed to recover pose for initial pair {img0_name} and {img1_name}\")\n        return {}, [], {}\n    \n    R = R.astype(np.float32)\n    t = t.astype(np.float32).flatten()\n    \n    poses = {img0_name: (np.eye(3, dtype=np.float32), np.zeros(3, dtype=np.float32))}\n    poses[img1_name] = (R, t)\n    registered_indices = {idx0, idx1}\n    \n    start_time = time.time()\n    for idx0, idx1 in selected_pairs[1:]:\n        global_idx0, global_idx1 = cluster_indices[idx0], cluster_indices[idx1]\n        img0_name = Path(image_paths[global_idx0]).name\n        img1_name = Path(image_paths[global_idx1]).name\n        \n        if img0_name in poses and img1_name in poses:\n            continue\n        elif img0_name not in poses and img1_name not in poses:\n            continue\n        \n        if img1_name in poses and img0_name not in poses:\n            img0_name, img1_name = img1_name, img0_name\n            idx0, idx1 = idx1, idx0\n        \n        key1 = (img0_name, img1_name)\n        key2 = (img1_name, img0_name)\n        if key1 in matches_dict:\n            mkpts0, mkpts1 = matches_dict[key1]\n        elif key2 in matches_dict:\n            mkpts1, mkpts0 = matches_dict[key2]\n        else:\n            print(f\"Matches not found for pair ({img0_name}, {img1_name})\")\n            continue\n        \n        mkpts0 = mkpts0.astype(np.float32)\n        mkpts1 = mkpts1.astype(np.float32)\n        \n        E, mask = cv2.findEssentialMat(mkpts0, mkpts1, K, method=cv2.RANSAC, prob=0.999, threshold=1.0)\n        if E is None or E.shape != (3, 3):\n            print(f\"Failed to estimate essential matrix between {img0_name} and {img1_name}\")\n            continue\n        \n        _, R, t, mask = cv2.recoverPose(E, mkpts0, mkpts1, K)\n        if R is None or t is None:\n            print(f\"Failed to recover pose between {img0_name} and {img1_name}\")\n            continue\n        \n        R = R.astype(np.float32)\n        t = t.astype(np.float32).flatten()\n        \n        R0, t0 = poses[img0_name]\n        R = R0 @ R\n        t = R0 @ t + t0\n        \n        poses[img1_name] = (R, t)\n        registered_indices.add(idx1)\n    \n    print(f\"Pose estimation took {time.time() - start_time:.2f} seconds.\")\n    \n    start_time = time.time()\n    points3d = []\n    for idx0, idx1 in image_pairs:\n        global_idx0, global_idx1 = cluster_indices[idx0], cluster_indices[idx1]\n        img0_name = Path(image_paths[global_idx0]).name\n        img1_name = Path(image_paths[global_idx1]).name\n        \n        if img0_name not in poses or img1_name not in poses:\n            continue\n        \n        key1 = (img0_name, img1_name)\n        key2 = (img1_name, img0_name)\n        if key1 in matches_dict:\n            mkpts0, mkpts1 = matches_dict[key1]\n        elif key2 in matches_dict:\n            mkpts1, mkpts0 = matches_dict[key2]\n        else:\n            print(f\"Matches not found for pair ({img0_name}, {img1_name}) during triangulation\")\n            continue\n        \n        if len(mkpts0) != len(mkpts1):\n            print(f\"Mismatch in number of points between {img0_name} and {img1_name}: {len(mkpts0)} vs {len(mkpts1)}\")\n            continue\n        \n        mkpts0 = mkpts0.astype(np.float32)\n        mkpts1 = mkpts1.astype(np.float32)\n        \n        R0, t0 = poses[img0_name]\n        R1, t1 = poses[img1_name]\n        \n        P0 = K @ np.hstack((R0, t0.reshape(3, 1)))\n        P1 = K @ np.hstack((R1, t1.reshape(3, 1)))\n        P0 = P0.astype(np.float32)\n        P1 = P1.astype(np.float32)\n        \n        pts0 = mkpts0.T.astype(np.float32)\n        pts1 = mkpts1.T.astype(np.float32)\n        \n        if pts0.shape[1] == 0 or pts1.shape[1] == 0:\n            print(f\"Skipping triangulation due to zero matches between {img0_name} and {img1_name}\")\n            continue\n        \n        points4d = cv2.triangulatePoints(P0, P1, pts0, pts1)\n        points3d_h = points4d[:3] / points4d[3]\n        points3d.extend(points3d_h.T)\n        \n        for _ in range(pts0.shape[1]):\n            point_to_images.append([(img0_name, pts0[:, _]), (img1_name, pts1[:, _])])\n    \n    points3d = np.array(points3d)\n    \n    valid = np.all(np.isfinite(points3d), axis=1) & (np.abs(points3d) < 1e4).all(axis=1)\n    points3d = points3d[valid]\n    point_to_images = [pt for pt, v in zip(point_to_images, valid) if v]\n    \n    print(f\"Triangulation took {time.time() - start_time:.2f} seconds.\")\n    \n    start_time = time.time()\n    def project(points3d, R, t, K):\n        points3d = points3d.T\n        points = R @ points3d + t.reshape(3, 1)\n        points = K @ points\n        points = points[:2] / points[2]\n        return points.T\n    \n    def reprojection_error(params, points3d, observations, K, img_names):\n        num_cameras = len(img_names)\n        num_points = len(points3d)\n        \n        translations = params[:num_cameras * 3].reshape(num_cameras, 3)\n        focal_length = params[-1]\n        K_opt = K.copy()\n        K_opt[0, 0] = K_opt[1, 1] = focal_length\n        \n        errors = []\n        for i, img_name in enumerate(img_names):\n            R, _ = poses[img_name]\n            t = translations[i]\n            for j, obs in enumerate(observations):\n                for img_obs, pt in obs:\n                    if img_obs == img_name:\n                        proj = project(points3d[j:j+1], R, t, K_opt)\n                        errors.append(proj[0] - pt)\n        \n        return np.concatenate(errors)\n    \n    img_names = list(poses.keys())\n    translations = np.array([t for _, t in poses.values()])\n    params = np.hstack((translations.ravel(), focal_length))\n    \n    result = least_squares(\n        reprojection_error,\n        params,\n        args=(points3d, point_to_images, K, img_names),\n        max_nfev=50,\n        ftol=1e-4,\n        xtol=1e-4\n    )\n    optimized_params = result.x\n    \n    num_cameras = len(img_names)\n    translations = optimized_params[:num_cameras * 3].reshape(num_cameras, 3)\n    for i, img_name in enumerate(img_names):\n        R, _ = poses[img_name]\n        poses[img_name] = (R, translations[i])\n    \n    print(f\"Bundle adjustment took {time.time() - start_time:.2f} seconds.\")\n    \n    return poses, points3d, K\n\n# === Process all scenes in the test set ===\ntest_dir = Path('/kaggle/input/image-matching-challenge-2025/test')\nsubmission_data = []\n\nfor scene_dir in test_dir.iterdir():\n    if not scene_dir.is_dir():\n        continue\n    scene_name = scene_dir.name\n    print(f\"\\nProcessing scene: {scene_name}\")\n    \n    image_paths = [str(img) for img in scene_dir.iterdir() if img.suffix in ('.png', '.jpg', '.jpeg')]\n    image_paths.sort()\n    print(f\"Number of images in scene: {len(image_paths)}\")\n    \n    num_images = len(image_paths)\n    selected_indices = list(range(num_images))\n    \n    cluster_size = 4\n    overlap = 2\n    clusters = []\n    for i in range(0, len(selected_indices), cluster_size - overlap):\n        cluster_indices = selected_indices[i:i + cluster_size]\n        if len(cluster_indices) >= 2:\n            clusters.append(cluster_indices)\n    print(f\"Clusters: {clusters}\")\n    \n    all_poses = {}\n    all_points3d = []\n    all_Ks = {}\n    \n    start_time = time.time()\n    for cluster_idx, cluster_indices in enumerate(clusters):\n        poses, points3d, K = run_sfm_on_cluster(cluster_indices, None, image_paths, matching, device, cluster_idx, clusters)\n        all_poses.update(poses)\n        all_points3d.append(points3d)\n        all_Ks.update({img_name: K for img_name in poses})\n    \n    if not all_points3d or len(all_points3d[0]) == 0:\n        print(f\"No points reconstructed in the first cluster for scene {scene_name}. Skipping.\")\n        continue\n    \n    merged_points3d = all_points3d[0]\n    for i in range(1, len(clusters)):\n        overlap_image = Path(image_paths[clusters[i-1][-1]]).name\n        if overlap_image not in all_poses:\n            print(f\"Cannot align clusters {i-1} and {i} using overlapping image {overlap_image}.\")\n            continue\n        \n        R0_prev, t0_prev = all_poses[overlap_image]\n        R0_curr, t0_curr = all_poses[overlap_image]\n        \n        R_align = R0_prev @ np.linalg.inv(R0_curr)\n        t_align = t0_prev - R_align @ t0_curr\n        \n        points3d_curr = all_points3d[i]\n        if len(points3d_curr) == 0:\n            print(f\"Cluster {i} has no points to merge\")\n            continue\n        points3d_curr = (R_align @ points3d_curr.T).T + t_align\n        all_points3d[i] = points3d_curr\n        merged_points3d = np.vstack((merged_points3d, points3d_curr))\n        \n        for img_name in list(all_poses.keys()):\n            if img_name in all_poses:\n                R, t = all_poses[img_name]\n                R = R_align @ R\n                t = R_align @ t + t_align\n                all_poses[img_name] = (R, t)\n    \n    save_ply(merged_points3d, f\"/kaggle/working/{scene_name}_point_cloud.ply\")\n    print(f\"Point cloud saved at: /kaggle/working/{scene_name}_point_cloud.ply\")\n    \n    print(f\"Scene {scene_name} processed in {time.time() - start_time:.2f} seconds\")\n    print(f\"Total number of 3D points reconstructed: {len(merged_points3d)}\")\n    print(f\"Total number of cameras estimated: {len(all_poses)}\")\n    \n    for img_name in all_poses:\n        R, t = all_poses[img_name]\n        w = np.sqrt(1.0 + R[0, 0] + R[1, 1] + R[2, 2]) / 2.0\n        if w < 1e-10:\n            w = 1e-10\n        x = (R[2, 1] - R[1, 2]) / (4 * w)\n        y = (R[0, 2] - R[2, 0]) / (4 * w)\n        z = (R[1, 0] - R[0, 1]) / (4 * w)\n        image_id = f\"{scene_name}/{img_name}\"\n        submission_data.append((image_id, w, x, y, z, t[0], t[1], t[2]))\n\n# === Ensure all images are included in submission ===\nall_image_paths = []\nfor scene_dir in test_dir.iterdir():\n    if not scene_dir.is_dir():\n        continue\n    scene_name = scene_dir.name\n    image_paths = [str(img) for img in scene_dir.iterdir() if img.suffix in ('.png', '.jpg', '.jpeg')]\n    image_paths.sort()\n    for img_path in image_paths:\n        img_name = Path(img_path).name\n        image_id = f\"{scene_name}/{img_name}\"\n        all_image_paths.append(image_id)\n\nfor image_id in all_image_paths:\n    if not any(data[0] == image_id for data in submission_data):\n        submission_data.append((image_id, 1.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0))\n\n# === Generate submission file ===\nsubmission_file = Path('/kaggle/working/submission.csv')\nwith open(submission_file, 'w') as f:\n    f.write(\"image_path,rotation_w,rotation_x,rotation_y,rotation_z,translation_x,translation_y,translation_z\\n\")\n    for data in submission_data:\n        image_id, w, x, y, z, tx, ty, tz = data\n        f.write(f\"{image_id},{w},{x},{y},{z},{tx},{ty},{tz}\\n\")\n\nprint(f\"Submission file generated at: {submission_file}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-06T15:24:47.491233Z","iopub.execute_input":"2025-04-06T15:24:47.491560Z","iopub.status.idle":"2025-04-06T15:26:27.071888Z","shell.execute_reply.started":"2025-04-06T15:24:47.491516Z","shell.execute_reply":"2025-04-06T15:26:27.070903Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Corrected SfM Pipeline¶","metadata":{}},{"cell_type":"code","source":"# === Process all scenes in the test set ===\ntest_dir = Path('/kaggle/input/image-matching-challenge-2025/test')\nsubmission_data = []\n\nprint(f\"Test directory: {test_dir}\")\nprint(f\"Contents of test directory: {list(test_dir.iterdir())}\")\n\nfor scene_dir in test_dir.iterdir():\n    if not scene_dir.is_dir():\n        print(f\"Skipping non-directory: {scene_dir}\")\n        continue\n    scene_name = scene_dir.name\n    print(f\"\\nProcessing scene: {scene_name}\")\n\n    image_paths = [str(img) for img in scene_dir.iterdir() if img.suffix in ('.png', '.jpg', '.jpeg')]\n    image_paths.sort()\n    print(f\"Number of images in scene: {len(image_paths)}\")\n\n    num_images = len(image_paths)\n    selected_indices = list(range(num_images))\n\n    cluster_size = 4\n    overlap = 2\n    clusters = []\n    for i in range(0, len(selected_indices), cluster_size - overlap):\n        cluster_indices = selected_indices[i:i + cluster_size]\n        if len(cluster_indices) >= 2:\n            clusters.append(cluster_indices)\n    print(f\"Clusters: {clusters}\")\n\n    all_poses = {}\n    all_points3d = []\n    all_Ks = {}\n\n    start_time = time.time()\n    for cluster_idx, cluster_indices in enumerate(clusters):\n        poses, points3d, K = run_sfm_on_cluster(cluster_indices, None, image_paths, matching, device, cluster_idx, clusters)\n        print(f\"Cluster {cluster_idx}: Number of poses estimated: {len(poses)}\")\n        all_poses.update(poses)\n        all_points3d.append(points3d)\n        all_Ks.update({img_name: K for img_name in poses})\n\n    print(f\"Total poses after processing all clusters: {len(all_poses)}\")\n\n    if not all_points3d or len(all_points3d[0]) == 0:\n        print(f\"No points reconstructed in the first cluster for scene {scene_name}. Skipping.\")\n        continue\n\n    merged_points3d = all_points3d[0]\n    for i in range(1, len(clusters)):\n        overlap_image = Path(image_paths[clusters[i-1][-1]]).name\n        if overlap_image not in all_poses:\n            print(f\"Cannot align clusters {i-1} and {i} using overlapping image {overlap_image}.\")\n            continue\n\n        R0_prev, t0_prev = all_poses[overlap_image]\n        R0_curr, t0_curr = all_poses[overlap_image]\n\n        R_align = R0_prev @ np.linalg.inv(R0_curr)\n        t_align = t0_prev - R_align @ t0_curr\n\n        points3d_curr = all_points3d[i]\n        if len(points3d_curr) == 0:\n            print(f\"Cluster {i} has no points to merge\")\n            continue\n        points3d_curr = (R_align @ points3d_curr.T).T + t_align\n        all_points3d[i] = points3d_curr\n        merged_points3d = np.vstack((merged_points3d, points3d_curr))\n\n        for img_name in list(all_poses.keys()):\n            if img_name in all_poses:\n                R, t = all_poses[img_name]\n                R = R_align @ R\n                t = R_align @ t + t_align\n                all_poses[img_name] = (R, t)\n\n    save_ply(merged_points3d, f\"/kaggle/working/{scene_name}_point_cloud.ply\")\n    print(f\"Point cloud saved at: /kaggle/working/{scene_name}_point_cloud.ply\")\n\n    print(f\"Scene {scene_name} processed in {time.time() - start_time:.2f} seconds\")\n    print(f\"Total number of 3D points reconstructed: {len(merged_points3d)}\")\n    print(f\"Total number of cameras estimated: {len(all_poses)}\")\n\n    print(f\"Adding poses to submission_data for scene {scene_name}\")\n    for img_name in all_poses:\n        R, t = all_poses[img_name]\n        w = np.sqrt(1.0 + R[0, 0] + R[1, 1] + R[2, 2]) / 2.0\n        if w < 1e-10:\n            w = 1e-10\n        x = (R[2, 1] - R[1, 2]) / (4 * w)\n        y = (R[0, 2] - R[2, 0]) / (4 * w)\n        z = (R[1, 0] - R[0, 1]) / (4 * w)\n        image_id = f\"{scene_name}_{img_name}_public\"\n        submission_data.append((image_id, w, x, y, z, t[0], t[1], t[2]))\n        print(f\"Added to submission_data: {image_id}\")\n\nprint(f\"Total entries in submission_data after processing all scenes: {len(submission_data)}\")\n\n# === Load the sample submission file to get all expected image paths ===\nsample_submission_path = '/kaggle/input/image-matching-challenge-2025/sample_submission.csv'\nprint(f\"Loading sample submission file from: {sample_submission_path}\")\nsample_submission_df = pd.read_csv(sample_submission_path)\nprint(f\"Total images in sample submission: {len(sample_submission_df)}\")\nprint(f\"Columns in sample submission: {sample_submission_df.columns.tolist()}\")\nprint(f\"First few rows of sample submission:\\n{sample_submission_df.head()}\")\n\n# === Debug: Print sample entries in submission_data ===\nprint(\"Sample entries in submission_data:\")\nfor entry in submission_data[:5]:\n    print(entry)\n\n# === Convert submission_data to a DataFrame ===\n# The image_id in submission_data is already in the correct format (e.g., ETs_another_et_another_et001.png_public)\nsubmission_df = pd.DataFrame(\n    submission_data,\n    columns=['image_path', 'rotation_w', 'rotation_x', 'rotation_y', 'rotation_z', 'translation_x', 'translation_y', 'translation_z']\n)\nprint(f\"submission_data has {len(submission_df)} entries before merging\")\n\n# === Merge with sample submission to ensure all images are included ===\n# The image_path column in submission_df matches the image_id column in sample_submission_df\nmerged_df = sample_submission_df[['image_id']].merge(\n    submission_df,\n    left_on='image_id',\n    right_on='image_path',\n    how='left'\n)\n\n# Drop the redundant image_path column after merging\nmerged_df = merged_df.drop(columns=['image_path'])\n\n# Fill missing values with default poses\ndefault_pose = {'rotation_w': 1.0, 'rotation_x': 0.0, 'rotation_y': 0.0, 'rotation_z': 0.0,\n                'translation_x': 0.0, 'translation_y': 0.0, 'translation_z': 0.0}\nmerged_df.fillna(default_pose, inplace=True)\n\nprint(f\"After merging, submission has {len(merged_df)} entries\")\nprint(f\"Sample rows from merged submission:\\n{merged_df.head()}\")\n\n# === Generate submission file ===\n# Ensure the submission file has the expected column names (Kaggle expects 'image_path')\nmerged_df = merged_df.rename(columns={'image_id': 'image_path'})\nsubmission_file = Path('/kaggle/working/submission.csv')\nmerged_df.to_csv(submission_file, index=False)\nprint(f\"Submission file generated at: {submission_file}\")\n\n# Verify the submission file\nsubmission_df = pd.read_csv(submission_file)\nprint(f\"Number of rows in submission.csv: {len(submission_df)}\")\nprint(f\"First few rows of submission.csv:\\n{submission_df.head()}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-06T15:26:27.072822Z","iopub.execute_input":"2025-04-06T15:26:27.073147Z","iopub.status.idle":"2025-04-06T15:28:07.166526Z","shell.execute_reply.started":"2025-04-06T15:26:27.073120Z","shell.execute_reply":"2025-04-06T15:28:07.165390Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Below is the complete, submission-ready code. I'll include comments to explain each section and the improvements made.¶\n","metadata":{}},{"cell_type":"code","source":"# === Import Libraries ===\nimport numpy as np\nfrom pathlib import Path\nimport cv2\nimport time\nimport torch\nimport sys\nimport warnings\nimport pandas as pd\nfrom scipy.optimize import least_squares\nwarnings.filterwarnings(\"ignore\", category=FutureWarning)\n\n# Add SuperGlue path\nsys.path.append('SuperGluePretrainedNetwork')\nfrom models.matching import Matching\nfrom models.utils import frame2tensor\n\n# === Function to Save Point Cloud as PLY ===\ndef save_ply(points3d, output_path):\n    header = f\"\"\"ply\nformat ascii 1.0\nelement vertex {len(points3d)}\nproperty float x\nproperty float y\nproperty float z\nproperty uchar red\nproperty uchar green\nproperty uchar blue\nend_header\n\"\"\"\n    colors = np.random.randint(0, 255, size=(len(points3d), 3))\n    with open(output_path, 'w') as f:\n        f.write(header)\n        for point, color in zip(points3d, colors):\n            f.write(f\"{point[0]} {point[1]} {point[2]} {color[0]} {color[1]} {color[2]}\\n\")\n\n# === Function to Initialize SuperGlue with Scene-Specific Weights ===\ndef initialize_superglue(scene_name):\n    print(f\"Initializing SuperGlue for scene: {scene_name}\")\n    start_time = time.time()\n    config = {\n        'superpoint': {\n            'nms_radius': 4,\n            'keypoint_threshold': 0.005,\n            'max_keypoints': 4096\n        },\n        'superglue': {\n            'sinkhorn_iterations': 30 if scene_name.lower() == 'stairs' else 20,  # Increased for stairs\n            'match_threshold': 0.03 if scene_name.lower() == 'stairs' else 0.02   # Increased for stairs\n        }\n    }\n    indoor_scenes = ['ETs', 'imc2023_theather_imc2024_church', 'imc2024_dioscuri_baalshamin',\n                     'pt_brandenburg_british_buckingham', 'pt_piazzasanmarco_grandplace',\n                     'pt_sacrecoeur_trevi_tajmahal', 'pt_stpeters_stpauls']\n    if any(scene.lower() in scene_name.lower() for scene in indoor_scenes):\n        config['superglue']['weights'] = 'indoor'\n        print(f\"Using 'indoor' weights for scene: {scene_name}\")\n    else:\n        config['superglue']['weights'] = 'outdoor'\n        print(f\"Using 'outdoor' weights for scene: {scene_name}\")\n    device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')\n    matching = Matching(config).eval().to(device)\n    print(f\"SuperGlue initialized in {time.time() - start_time:.2f} seconds\")\n    return matching, device\n\n# === Function to Recover Pose with Fallback ===\ndef recover_pose_with_fallback(mkpts0, mkpts1, K, img0_name, img1_name):\n    # First attempt: Essential matrix\n    E, mask = cv2.findEssentialMat(mkpts0, mkpts1, K, method=cv2.RANSAC, prob=0.999, threshold=0.5, maxIters=2000)\n    if E is None or E.shape != (3, 3):\n        print(f\"Essential matrix estimation failed for pair {img0_name} and {img1_name}. Trying homography.\")\n    else:\n        inliers, R, t, mask = cv2.recoverPose(E, mkpts0, mkpts1, K)\n        if inliers >= 10:  # Require at least 10 inliers\n            return R, t, inliers\n    \n    # Fallback: Homography (assuming planar scene, common in stairs)\n    H, mask = cv2.findHomography(mkpts0, mkpts1, method=cv2.RANSAC, ransacReprojThreshold=0.5, confidence=0.999, maxIters=2000)\n    if H is None or H.shape != (3, 3):\n        print(f\"Homography estimation failed for pair {img0_name} and {img1_name}. Giving up.\")\n        return None, None, 0\n    \n    # Decompose homography to get possible poses\n    num, Rs, ts, normals = cv2.decomposeHomographyMat(H, K)\n    best_R, best_t, best_inliers = None, None, 0\n    for i in range(num):\n        R, t = Rs[i], ts[i]\n        points3d = cv2.triangulatePoints(K @ np.hstack((np.eye(3), np.zeros((3, 1)))), K @ np.hstack((R, t)), mkpts0.T, mkpts1.T)\n        points3d = points3d[:3] / points3d[3]\n        points = K @ (R @ points3d + t.reshape(3, 1))\n        points = points[:2] / points[2]\n        errors = np.linalg.norm(points.T - mkpts1, axis=1)\n        inliers = np.sum(errors < 1.0)\n        if inliers > best_inliers:\n            best_inliers = inliers\n            best_R, best_t = R, t\n    \n    if best_inliers >= 10:\n        return best_R, best_t, best_inliers\n    else:\n        print(f\"Insufficient inliers from homography decomposition for pair {img0_name} and {img1_name}.\")\n        return None, None, 0\n\n# === Function to Run SfM on a Single Cluster ===\ndef run_sfm_on_cluster(args):\n    cluster_indices, image_dir, image_paths, scene_name, cluster_idx, clusters, overlap = args\n    print(f\"Starting cluster {cluster_idx} for scene {scene_name}\")\n    \n    try:\n        # Initialize SuperGlue and device\n        matching, device = initialize_superglue(scene_name)\n        \n        keypoints_dict = {}\n        matches_dict = {}\n        image_sizes = {}\n        point_to_images = []\n        \n        print(f\"Processing cluster {cluster_idx} with indices: {cluster_indices}\")\n        \n        start_time = time.time()\n        image_pairs = [(i, j) for i in range(len(cluster_indices)) for j in range(i + 1, len(cluster_indices))]\n        edges = []\n        for idx0, idx1 in image_pairs:\n            global_idx0, global_idx1 = cluster_indices[idx0], cluster_indices[idx1]\n            img0 = cv2.imread(image_paths[global_idx0], cv2.IMREAD_GRAYSCALE)\n            img1 = cv2.imread(image_paths[global_idx1], cv2.IMREAD_GRAYSCALE)\n            \n            if img0 is None or img1 is None:\n                print(f\"Failed to load images: {image_paths[global_idx0]} or {image_paths[global_idx1]}\")\n                continue\n            \n            img0 = cv2.equalizeHist(img0)\n            img1 = cv2.equalizeHist(img1)\n            \n            orig_size0 = img0.shape[::-1]\n            orig_size1 = img1.shape[::-1]\n            \n            img0 = cv2.resize(img0, (640, 480))\n            img1 = cv2.resize(img1, (640, 480))\n            \n            inp0 = frame2tensor(img0, device)\n            inp1 = frame2tensor(img1, device)\n            \n            pred = matching({'image0': inp0, 'image1': inp1})\n            pred = {k: v[0].detach().cpu().numpy() for k, v in pred.items()}\n            kpts0, kpts1 = pred['keypoints0'], pred['keypoints1']\n            matches, conf = pred['matches0'], pred['matching_scores0']\n            \n            valid = matches > -1\n            mkpts0 = kpts0[valid]\n            mkpts1 = kpts1[matches[valid]]\n            \n            scale0 = (orig_size0[0] / 640, orig_size0[1] / 480)\n            scale1 = (orig_size1[0] / 640, orig_size1[1] / 480)\n            kpts0[:, 0] *= scale0[0]\n            kpts0[:, 1] *= scale0[1]\n            kpts1[:, 0] *= scale1[0]\n            kpts1[:, 1] *= scale1[1]\n            mkpts0[:, 0] *= scale0[0]\n            mkpts0[:, 1] *= scale0[1]\n            mkpts1[:, 0] *= scale1[0]\n            mkpts1[:, 1] *= scale1[1]\n            \n            img0_name = Path(image_paths[global_idx0]).name\n            img1_name = Path(image_paths[global_idx1]).name\n            if img0_name not in keypoints_dict:\n                keypoints_dict[img0_name] = kpts0\n                image_sizes[img0_name] = orig_size0\n            if img1_name not in keypoints_dict:\n                keypoints_dict[img1_name] = kpts1\n                image_sizes[img1_name] = orig_size1\n            \n            matches_dict[(img0_name, img1_name)] = (mkpts0, mkpts1)\n            \n            priority = 0\n            if cluster_idx < len(clusters) - 1:\n                overlap_indices = clusters[cluster_idx][-overlap:]\n                if global_idx0 in overlap_indices or global_idx1 in overlap_indices:\n                    priority = 1\n            \n            edges.append((len(mkpts0), idx0, idx1, priority))\n            \n            print(f\"Matches between {img0_name} and {img1_name}: {len(mkpts0)}\")\n        \n        print(f\"SuperGlue matching for cluster took {time.time() - start_time:.2f} seconds.\")\n        \n        # Sort edges by priority and number of matches\n        edges.sort(key=lambda x: (-x[3], -x[0]))\n        parent = list(range(len(cluster_indices)))\n        rank = [0] * len(cluster_indices)\n        \n        def find(x):\n            if parent[x] != x:\n                parent[x] = find(parent[x])\n            return parent[x]\n        \n        def union(x, y):\n            px, py = find(x), find(y)\n            if px == py:\n                return\n            if rank[px] < rank[py]:\n                px, py = py, px\n            parent[py] = px\n            if rank[px] == rank[py]:\n                rank[px] += 1\n        \n        selected_pairs = []\n        for num_matches, idx0, idx1, _ in edges:\n            if num_matches < 6:\n                continue\n            if find(idx0) != find(idx1):\n                union(idx0, idx1)\n                selected_pairs.append((idx0, idx1))\n        \n        # Ensure overlap images are included\n        if cluster_idx < len(clusters) - 1:\n            overlap_indices = clusters[cluster_idx][-overlap:]\n            for overlap_idx in overlap_indices:\n                overlap_local_idx = cluster_indices.index(overlap_idx)\n                if overlap_local_idx not in {idx0 for idx0, _ in selected_pairs} and overlap_local_idx not in {idx1 for _, idx1 in selected_pairs}:\n                    best_pair = None\n                    best_num_matches = 0\n                    for num_matches, idx0, idx1, _ in edges:\n                        if num_matches < 6:\n                            continue\n                        if idx0 == overlap_local_idx or idx1 == overlap_local_idx:\n                            if num_matches > best_num_matches:\n                                best_num_matches = num_matches\n                                best_pair = (idx0, idx1)\n                    if best_pair:\n                        idx0, idx1 = best_pair\n                        if find(idx0) != find(idx1):\n                            union(idx0, idx1)\n                            selected_pairs.append((idx0, idx1))\n        \n        # Fallback: If no pairs are selected, try pairwise processing\n        if not selected_pairs:\n            print(\"No pairs with sufficient matches found in cluster. Trying pairwise processing...\")\n            selected_pairs = []\n            for i in range(len(cluster_indices) - 1):\n                idx0, idx1 = i, i + 1\n                global_idx0, global_idx1 = cluster_indices[idx0], cluster_indices[idx1]\n                img0_name = Path(image_paths[global_idx0]).name\n                img1_name = Path(image_paths[global_idx1]).name\n                if (img0_name, img1_name) in matches_dict:\n                    mkpts0, _ = matches_dict[(img0_name, img1_name)]\n                    if len(mkpts0) >= 6:\n                        selected_pairs.append((idx0, idx1))\n        \n        if not selected_pairs:\n            print(\"Pairwise processing also failed. Skipping cluster.\")\n            return {}, [], {}, []\n        \n        # Initialize poses with the first pair\n        idx0, idx1 = selected_pairs[0]\n        global_idx0, global_idx1 = cluster_indices[idx0], cluster_indices[idx1]\n        img0_name = Path(image_paths[global_idx0]).name\n        img1_name = Path(image_paths[global_idx1]).name\n        \n        if (img0_name, img1_name) not in matches_dict:\n            print(f\"Initial pair ({img0_name}, {img1_name}) not in matches_dict. Skipping cluster.\")\n            return {}, [], {}, []\n        \n        mkpts0, mkpts1 = matches_dict[(img0_name, img1_name)]\n        mkpts0 = mkpts0.astype(np.float32)\n        mkpts1 = mkpts1.astype(np.float32)\n        \n        image_width, image_height = image_sizes[img0_name]\n        focal_length = max(image_width, image_height) * 1.0\n        cx, cy = image_width / 2, image_height / 2\n        K = np.array([\n            [focal_length, 0, cx],\n            [0, focal_length, cy],\n            [0, 0, 1]\n        ], dtype=np.float32)\n        \n        R, t, inliers = recover_pose_with_fallback(mkpts0, mkpts1, K, img0_name, img1_name)\n        if R is None or t is None:\n            print(f\"Failed to recover pose for initial pair {img0_name} and {img1_name}. Skipping cluster.\")\n            return {}, [], {}, []\n        \n        R = R.astype(np.float32)\n        t = t.astype(np.float32).flatten()\n        \n        poses = {img0_name: (np.eye(3, dtype=np.float32), np.zeros(3, dtype=np.float32))}\n        poses[img1_name] = (R, t)\n        registered_indices = {idx0, idx1}\n        \n        # Pose estimation for remaining pairs\n        start_time = time.time()\n        for idx0, idx1 in selected_pairs[1:]:\n            global_idx0, global_idx1 = cluster_indices[idx0], cluster_indices[idx1]\n            img0_name = Path(image_paths[global_idx0]).name\n            img1_name = Path(image_paths[global_idx1]).name\n            \n            # Skip if both images are already registered\n            if img0_name in poses and img1_name in poses:\n                continue\n            # Skip if neither image is registered (cannot anchor the pair)\n            if img0_name not in poses and img1_name not in poses:\n                continue\n            \n            # Ensure img0_name is the registered image\n            if img1_name in poses and img0_name not in poses:\n                img0_name, img1_name = img1_name, img0_name\n                idx0, idx1 = idx1, idx0\n            \n            # Check if matches exist for this pair\n            key1 = (img0_name, img1_name)\n            key2 = (img1_name, img0_name)\n            if key1 in matches_dict:\n                mkpts0, mkpts1 = matches_dict[key1]\n            elif key2 in matches_dict:\n                mkpts1, mkpts0 = matches_dict[key2]\n            else:\n                print(f\"Matches not found for pair ({img0_name}, {img1_name}) in matches_dict. Skipping pair.\")\n                continue\n            \n            # Ensure we have enough matches\n            if len(mkpts0) < 6:\n                print(f\"Insufficient matches ({len(mkpts0)}) for pair ({img0_name}, {img1_name}). Skipping pair.\")\n                continue\n            \n            mkpts0 = mkpts0.astype(np.float32)\n            mkpts1 = mkpts1.astype(np.float32)\n            \n            R, t, inliers = recover_pose_with_fallback(mkpts0, mkpts1, K, img0_name, img1_name)\n            if R is None or t is None:\n                print(f\"Failed to recover pose between {img0_name} and {img1_name}. Skipping pair.\")\n                continue\n            \n            R = R.astype(np.float32)\n            t = t.astype(np.float32).flatten()\n            \n            R0, t0 = poses[img0_name]\n            R = R0 @ R\n            t = R0 @ t + t0\n            \n            poses[img1_name] = (R, t)\n            registered_indices.add(idx1)\n        \n        # Fallback: Register remaining images pairwise\n        remaining_indices = set(range(len(cluster_indices))) - registered_indices\n        for idx in sorted(remaining_indices):\n            global_idx = cluster_indices[idx]\n            img_name = Path(image_paths[global_idx]).name\n            best_pair = None\n            best_num_matches = 0\n            for num_matches, idx0, idx1, _ in edges:\n                if num_matches < 6:\n                    continue\n                if idx0 == idx or idx1 == idx:\n                    other_idx = idx1 if idx0 == idx else idx0\n                    if other_idx in registered_indices:\n                        if num_matches > best_num_matches:\n                            best_num_matches = num_matches\n                            best_pair = (idx0, idx1)\n            if best_pair:\n                idx0, idx1 = best_pair\n                global_idx0, global_idx1 = cluster_indices[idx0], cluster_indices[idx1]\n                img0_name = Path(image_paths[global_idx0]).name\n                img1_name = Path(image_paths[global_idx1]).name\n                \n                if img0_name in poses and img1_name in poses:\n                    continue\n                elif img0_name not in poses and img1_name not in poses:\n                    continue\n                \n                if img1_name in poses and img0_name not in poses:\n                    img0_name, img1_name = img1_name, img0_name\n                    idx0, idx1 = idx1, idx0\n                \n                key1 = (img0_name, img1_name)\n                key2 = (img1_name, img0_name)\n                if key1 in matches_dict:\n                    mkpts0, mkpts1 = matches_dict[key1]\n                elif key2 in matches_dict:\n                    mkpts1, mkpts0 = matches_dict[key2]\n                else:\n                    print(f\"Fallback: Matches not found for pair ({img0_name}, {img1_name}) in matches_dict. Skipping pair.\")\n                    continue\n                \n                if len(mkpts0) < 6:\n                    print(f\"Fallback: Insufficient matches ({len(mkpts0)}) for pair ({img0_name}, {img1_name}). Skipping pair.\")\n                    continue\n                \n                mkpts0 = mkpts0.astype(np.float32)\n                mkpts1 = mkpts1.astype(np.float32)\n                \n                R, t, inliers = recover_pose_with_fallback(mkpts0, mkpts1, K, img0_name, img1_name)\n                if R is None or t is None:\n                    print(f\"Fallback: Failed to recover pose between {img0_name} and {img1_name}. Skipping pair.\")\n                    continue\n                \n                R = R.astype(np.float32)\n                t = t.astype(np.float32).flatten()\n                \n                R0, t0 = poses[img0_name]\n                R = R0 @ R\n                t = R0 @ t + t0\n                \n                poses[img1_name] = (R, t)\n                registered_indices.add(idx1)\n        \n        print(f\"Pose estimation took {time.time() - start_time:.2f} seconds.\")\n        \n        # Triangulation\n        start_time = time.time()\n        points3d = []\n        cluster_point_to_images = []\n        for idx0, idx1 in image_pairs:\n            global_idx0, global_idx1 = cluster_indices[idx0], cluster_indices[idx1]\n            img0_name = Path(image_paths[global_idx0]).name\n            img1_name = Path(image_paths[global_idx1]).name\n            \n            if img0_name not in poses or img1_name not in poses:\n                continue\n            \n            key1 = (img0_name, img1_name)\n            key2 = (img1_name, img0_name)\n            if key1 in matches_dict:\n                mkpts0, mkpts1 = matches_dict[key1]\n            elif key2 in matches_dict:\n                mkpts1, mkpts0 = matches_dict[key2]\n            else:\n                print(f\"Matches not found for pair ({img0_name}, {img1_name}) during triangulation\")\n                continue\n            \n            if len(mkpts0) != len(mkpts1):\n                print(f\"Mismatch in number of points between {img0_name} and {img1_name}: {len(mkpts0)} vs {len(mkpts1)}\")\n                continue\n            \n            mkpts0 = mkpts0.astype(np.float32)\n            mkpts1 = mkpts1.astype(np.float32)\n            \n            R0, t0 = poses[img0_name]\n            R1, t1 = poses[img1_name]\n            \n            P0 = K @ np.hstack((R0, t0.reshape(3, 1)))\n            P1 = K @ np.hstack((R1, t1.reshape(3, 1)))\n            P0 = P0.astype(np.float32)\n            P1 = P1.astype(np.float32)\n            \n            pts0 = mkpts0.T.astype(np.float32)\n            pts1 = mkpts1.T.astype(np.float32)\n            \n            if pts0.shape[1] == 0 or pts1.shape[1] == 0:\n                print(f\"Skipping triangulation due to zero matches between {img0_name} and {img1_name}\")\n                continue\n            \n            points4d = cv2.triangulatePoints(P0, P1, pts0, pts1)\n            points3d_h = points4d[:3] / points4d[3]\n            points3d.extend(points3d_h.T)\n            \n            for _ in range(pts0.shape[1]):\n                cluster_point_to_images.append([(img0_name, pts0[:, _]), (img1_name, pts1[:, _])])\n        \n        points3d = np.array(points3d)\n        \n        valid = np.all(np.isfinite(points3d), axis=1) & (np.abs(points3d) < 1e4).all(axis=1)\n        points3d = points3d[valid]\n        cluster_point_to_images = [pt for pt, v in zip(cluster_point_to_images, valid) if v]\n        \n        print(f\"Triangulation took {time.time() - start_time:.2f} seconds.\")\n        \n        # Local Bundle Adjustment\n        if len(poses) < 3 or len(cluster_point_to_images) == 0:\n            print(\"Skipping bundle adjustment due to insufficient poses or observations.\")\n        else:\n            start_time = time.time()\n            def project(points3d, R, t, K):\n                points3d = points3d.T\n                points = R @ points3d + t.reshape(3, 1)\n                points = K @ points\n                points = points[:2] / points[2]\n                return points.T\n            \n            def reprojection_error(params, points3d, observations, K, img_names):\n                num_cameras = len(img_names)\n                num_points = len(points3d)\n                \n                translations = params[:num_cameras * 3].reshape(num_cameras, 3)\n                focal_length = params[-1]\n                K_opt = K.copy()\n                K_opt[0, 0] = K_opt[1, 1] = focal_length\n                \n                errors = []\n                for i, img_name in enumerate(img_names):\n                    R, _ = poses[img_name]\n                    t = translations[i]\n                    for j, obs in enumerate(observations):\n                        for img_obs, pt in obs:\n                            if img_obs == img_name:\n                                proj = project(points3d[j:j+1], R, t, K_opt)\n                                errors.append(proj[0] - pt)\n                \n                if len(errors) == 0:\n                    print(\"No reprojection errors computed. Returning zeros.\")\n                    return np.zeros(num_cameras * 3 + 1)\n                return np.concatenate(errors)\n            \n            img_names = list(poses.keys())\n            translations = np.array([t for _, t in poses.values()])\n            params = np.hstack((translations.ravel(), focal_length))\n            \n            result = least_squares(\n                reprojection_error,\n                params,\n                args=(points3d, cluster_point_to_images, K, img_names),\n                max_nfev=20,  # Reduced for faster processing\n                ftol=1e-5,\n                xtol=1e-5\n            )\n            optimized_params = result.x\n            \n            num_cameras = len(img_names)\n            translations = optimized_params[:num_cameras * 3].reshape(num_cameras, 3)\n            for i, img_name in enumerate(img_names):\n                R, _ = poses[img_name]\n                poses[img_name] = (R, translations[i])\n            \n            print(f\"Bundle adjustment took {time.time() - start_time:.2f} seconds.\")\n        \n        return poses, points3d, K, cluster_point_to_images\n    \n    except Exception as e:\n        print(f\"Error in cluster {cluster_idx} for scene {scene_name}: {str(e)}\")\n        return {}, [], {}, []\n\n# === Process All Scenes in the Test Set ===\ntest_dir = Path('/kaggle/input/image-matching-challenge-2025/test')\nsubmission_data = []\n\nprint(f\"Test directory: {test_dir}\")\nprint(f\"Contents of test directory: {list(test_dir.iterdir())}\")\n\n# Load sample submission to get all expected scenes\nsample_submission_path = '/kaggle/input/image-matching-challenge-2025/sample_submission.csv'\nsample_submission_df = pd.read_csv(sample_submission_path)\nprint(f\"Total images in sample submission: {len(sample_submission_df)}\")\nunique_scenes = sample_submission_df['dataset'].unique()\nprint(f\"Unique scenes in sample submission: {unique_scenes}\")\n\n# Iterate over all scenes in the sample submission\nfor scene_name in unique_scenes:\n    scene_dir = test_dir / scene_name\n    if not scene_dir.is_dir():\n        print(f\"Scene directory {scene_dir} does not exist, skipping (will be available during evaluation).\")\n        continue\n    print(f\"\\nProcessing scene: {scene_name}\")\n\n    image_paths = [str(img) for img in scene_dir.iterdir() if img.suffix in ('.png', '.jpg', '.jpeg')]\n    image_paths.sort()\n    print(f\"Number of images in scene: {len(image_paths)}\")\n\n    num_images = len(image_paths)\n    selected_indices = list(range(num_images))\n\n    # Adjust cluster size for stairs\n    cluster_size = 8 if scene_name.lower() == 'stairs' else 10  # Reduced for stairs\n    overlap = 2 if scene_name.lower() == 'stairs' else 5        # Reduced overlap for stairs\n    clusters = []\n    for i in range(0, len(selected_indices), cluster_size - overlap):\n        cluster_indices = selected_indices[i:i + cluster_size]\n        if len(cluster_indices) >= 2:\n            clusters.append(cluster_indices)\n    print(f\"Clusters: {clusters}\")\n\n    all_poses = {}\n    all_points3d = []\n    all_Ks = {}\n    all_observations = []\n\n    start_time = time.time()\n    print(f\"Using sequential processing for scene {scene_name}\")\n    for cluster_idx, cluster_indices in enumerate(clusters):\n        result = run_sfm_on_cluster((cluster_indices, None, image_paths, scene_name, cluster_idx, clusters, overlap))\n        poses, points3d, K, cluster_point_to_images = result\n        all_poses.update(poses)\n        if len(points3d) > 0:\n            all_points3d.append(points3d)\n        all_Ks.update({img_name: K for img_name in poses})\n        all_observations.extend(cluster_point_to_images)\n\n    print(f\"Total poses after processing all clusters: {len(all_poses)}\")\n\n    if not all_points3d or len(all_points3d[0]) == 0:\n        print(f\"No points reconstructed in the first cluster for scene {scene_name}. Skipping.\")\n        continue\n\n    # Merge point clouds and align clusters\n    merged_points3d = all_points3d[0]\n    for i in range(1, len(all_points3d)):\n        if i >= len(clusters):\n            continue\n        overlap_images = [Path(image_paths[idx]).name for idx in clusters[i-1][-overlap:]]\n        alignment_found = False\n        for overlap_image in overlap_images:\n            if overlap_image in all_poses:\n                R0_prev, t0_prev = all_poses[overlap_image]\n                R0_curr, t0_curr = all_poses[overlap_image]\n\n                R_align = R0_prev @ np.linalg.inv(R0_curr)\n                t_align = t0_prev - R_align @ t0_curr\n\n                points3d_curr = all_points3d[i]\n                if len(points3d_curr) == 0:\n                    print(f\"Cluster {i} has no points to merge\")\n                    continue\n                points3d_curr = (R_align @ points3d_curr.T).T + t_align\n                all_points3d[i] = points3d_curr\n                merged_points3d = np.vstack((merged_points3d, points3d_curr))\n\n                for img_name in list(all_poses.keys()):\n                    if img_name in all_poses:\n                        R, t = all_poses[img_name]\n                        R = R_align @ R\n                        t = R_align @ t + t_align\n                        all_poses[img_name] = (R, t)\n                alignment_found = True\n                break\n        if not alignment_found:\n            print(f\"Cannot align clusters {i-1} and {i} due to missing overlap poses.\")\n\n    # Global bundle adjustment (simplified)\n    if len(all_poses) >= 3 and len(all_observations) > 0:\n        start_time = time.time()\n        def project(points3d, R, t, K):\n            points3d = points3d.T\n            points = R @ points3d + t.reshape(3, 1)\n            points = K @ points\n            points = points[:2] / points[2]\n            return points.T\n        \n        def reprojection_error(params, points3d, observations, K, img_names):\n            num_cameras = len(img_names)\n            num_points = len(points3d)\n            \n            translations = params[:num_cameras * 3].reshape(num_cameras, 3)\n            focal_length = params[-1]\n            K_opt = K.copy()\n            K_opt[0, 0] = K_opt[1, 1] = focal_length\n            \n            errors = []\n            for i, img_name in enumerate(img_names):\n                R, _ = all_poses[img_name]\n                t = translations[i]\n                for j, obs in enumerate(observations):\n                    for img_obs, pt in obs:\n                        if img_obs == img_name:\n                            proj = project(points3d[j:j+1], R, t, K_opt)\n                            errors.append(proj[0] - pt)\n            \n            if len(errors) == 0:\n                print(\"No reprojection errors computed in global BA. Returning zeros.\")\n                return np.zeros(num_cameras * 3 + 1)\n            return np.concatenate(errors)\n        \n        img_names = list(all_poses.keys())\n        translations = np.array([t for _, t in all_poses.values()])\n        params = np.hstack((translations.ravel(), focal_length))\n        \n        result = least_squares(\n            reprojection_error,\n            params,\n            args=(merged_points3d, all_observations, K, img_names),\n            max_nfev=20,  # Reduced for faster processing\n            ftol=1e-5,\n            xtol=1e-5\n        )\n        optimized_params = result.x\n        \n        num_cameras = len(img_names)\n        translations = optimized_params[:num_cameras * 3].reshape(num_cameras, 3)\n        for i, img_name in enumerate(img_names):\n            R, _ = all_poses[img_name]\n            all_poses[img_name] = (R, translations[i])\n        \n        print(f\"Global bundle adjustment took {time.time() - start_time:.2f} seconds.\")\n\n    save_ply(merged_points3d, f\"/kaggle/working/{scene_name}_point_cloud.ply\")\n    print(f\"Point cloud saved at: /kaggle/working/{scene_name}_point_cloud.ply\")\n\n    print(f\"Scene {scene_name} processed in {time.time() - start_time:.2f} seconds\")\n    print(f\"Total number of 3D points reconstructed: {len(merged_points3d)}\")\n    print(f\"Total number of cameras estimated: {len(all_poses)}\")\n\n    print(f\"Adding poses to submission_data for scene {scene_name}\")\n    for img_name in all_poses:\n        R, t = all_poses[img_name]\n        w = np.sqrt(1.0 + R[0, 0] + R[1, 1] + R[2, 2]) / 2.0\n        if w < 1e-10:\n            w = 1e-10\n        x = (R[2, 1] - R[1, 2]) / (4 * w)\n        y = (R[0, 2] - R[2, 0]) / (4 * w)\n        z = (R[1, 0] - R[0, 1]) / (4 * w)\n        image_id = f\"{scene_name}/{img_name}\"\n        submission_data.append((image_id, w, x, y, z, t[0], t[1], t[2]))\n        print(f\"Added to submission_data: {image_id}\")\n\nprint(f\"Total entries in submission_data after processing all scenes: {len(submission_data)}\")\n\n# === Merge with Sample Submission ===\nprint(f\"Loading sample submission file from: {sample_submission_path}\")\nprint(f\"Total images in sample submission: {len(sample_submission_df)}\")\nprint(f\"Columns in sample submission: {sample_submission_df.columns.tolist()}\")\nprint(f\"First few rows of sample submission:\\n{sample_submission_df.head()}\")\n\nsubmission_df = pd.DataFrame(\n    submission_data,\n    columns=['image_path', 'rotation_w', 'rotation_x', 'rotation_y', 'rotation_z', 'translation_x', 'translation_y', 'translation_z']\n)\nprint(f\"submission_data has {len(submission_df)} entries before merging\")\n\nsample_submission_df['image_path'] = sample_submission_df['dataset'] + '/' + sample_submission_df['image']\nmerged_df = sample_submission_df[['image_path']].merge(\n    submission_df,\n    on='image_path',\n    how='left'\n)\n\ndefault_pose = {\n    'rotation_w': 1.0, 'rotation_x': 0.0, 'rotation_y': 0.0, 'rotation_z': 0.0,\n    'translation_x': 0.0, 'translation_y': 0.0, 'translation_z': 0.0\n}\nmerged_df.fillna(default_pose, inplace=True)\n\nprint(f\"After merging, submission has {len(merged_df)} entries\")\nprint(f\"Sample rows from merged submission:\\n{merged_df.head()}\")\n\n# === Generate Submission File ===\nsubmission_file = Path('/kaggle/working/submission.csv')\nmerged_df.to_csv(submission_file, index=False)\nprint(f\"Submission file generated at: {submission_file}\")\n\n# Verify the submission file\nsubmission_df = pd.read_csv(submission_file)\nprint(f\"Number of rows in submission.csv: {len(submission_df)}\")\nprint(f\"First few rows of submission.csv:\\n{submission_df.head()}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-06T15:28:07.167874Z","iopub.execute_input":"2025-04-06T15:28:07.168229Z","iopub.status.idle":"2025-04-06T15:49:22.036146Z","shell.execute_reply.started":"2025-04-06T15:28:07.168196Z","shell.execute_reply":"2025-04-06T15:49:22.034950Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\n\n# Read the submission.csv file\n# If you're on your local machine, replace the path with the location of your downloaded file\n# If you're in a Kaggle notebook, upload the file or read it from /kaggle/working/\ndf = pd.read_csv('/kaggle/working/submission.csv')  # Adjust the path as needed\n\n# Validate the file\nprint(f\"Number of rows: {len(df)}\")\nprint(f\"Columns: {df.columns.tolist()}\")\n\n# Expected values\nexpected_rows = 1945\nexpected_columns = ['image_path', 'rotation_w', 'rotation_x', 'rotation_y', 'rotation_z', \n                    'translation_x', 'translation_y', 'translation_z']\n\n# Check row count\nassert len(df) == expected_rows, f\"Expected {expected_rows} rows, but got {len(df)}\"\n\n# Check columns\nassert list(df.columns) == expected_columns, f\"Expected columns {expected_columns}, but got {df.columns.tolist()}\"\n\n# Check for NaN or infinity\nnumeric_cols = ['rotation_w', 'rotation_x', 'rotation_y', 'rotation_z', \n                'translation_x', 'translation_y', 'translation_z']\nassert not df[numeric_cols].isnull().any().any(), \"Submission contains NaN values\"\nassert np.isfinite(df[numeric_cols]).all().all(), \"Submission contains non-finite values (inf or -inf)\"\n\n# Check quaternion normalization\nquaternion_cols = ['rotation_w', 'rotation_x', 'rotation_y', 'rotation_z']\nquaternion_norm = (df[quaternion_cols]**2).sum(axis=1)\nassert np.allclose(quaternion_norm, 1.0, atol=1e-6), \"Quaternions are not normalized\"\n\n# Check image_path format\nfor path in df['image_path']:\n    assert isinstance(path, str) and len(path.split('/')) == 2, f\"Invalid image_path format: {path}\"\n\nprint(\"Submission file validation passed!\")\n\n# Re-save the file with proper encoding and line endings\ndf.to_csv('/kaggle/working/submission_clean.csv', index=False, encoding='utf-8', lineterminator='\\n')\nprint(\"Re-saved submission as submission_clean.csv\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\n\n# Rename submission_clean.csv to submission.csv\nos.rename('/kaggle/working/submission_clean.csv', '/kaggle/working/submission.csv')\nprint(\"Renamed submission_clean.csv to submission.csv\")\n\n# Confirm the file exists\nprint(\"Files in /kaggle/working/:\")\nfor file in os.listdir('/kaggle/working'):\n    print(f\"- {file}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-06T15:49:22.067752Z","iopub.execute_input":"2025-04-06T15:49:22.067977Z","iopub.status.idle":"2025-04-06T15:49:22.074455Z","shell.execute_reply.started":"2025-04-06T15:49:22.067957Z","shell.execute_reply":"2025-04-06T15:49:22.073418Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import shutil\n\n# Clean up /kaggle/working/ except for submission.csv\nfiles_to_keep = ['submission.csv']\nfor file in os.listdir('/kaggle/working'):\n    if file not in files_to_keep:\n        file_path = os.path.join('/kaggle/working', file)\n        if os.path.isfile(file_path):\n            os.remove(file_path)\n            print(f\"Removed {file_path}\")\n        elif os.path.isdir(file_path):\n            shutil.rmtree(file_path)\n            print(f\"Removed directory {file_path}\")\n\n# Confirm remaining files\nprint(\"Remaining files in /kaggle/working/:\")\nfor file in os.listdir('/kaggle/working'):\n    print(f\"- {file}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-06T15:49:22.075350Z","iopub.execute_input":"2025-04-06T15:49:22.075959Z","iopub.status.idle":"2025-04-06T15:49:22.236624Z","shell.execute_reply.started":"2025-04-06T15:49:22.075924Z","shell.execute_reply":"2025-04-06T15:49:22.235714Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\nimport shutil\n\n# Create the weights directory if it doesn't exist\nweights_dir = '/kaggle/working/SuperGluePretrainedNetwork/models/weights/'\nos.makedirs(weights_dir, exist_ok=True)\n\n# Path to the weights dataset (adjust based on your dataset name)\nweights_dataset_path = '/kaggle/input/superglue-weights/'\n\n# List of weight files\nweight_files = ['superpoint_v1.pth', 'superglue_indoor.pth', 'superglue_outdoor.pth']\n\n# Copy weights from the dataset to the expected directory\nfor weight_file in weight_files:\n    src_path = os.path.join(weights_dataset_path, weight_file)\n    dst_path = os.path.join(weights_dir, weight_file)\n    if os.path.exists(src_path):\n        shutil.copy(src_path, dst_path)\n        print(f\"Copied {weight_file} to {dst_path}\")\n    else:\n        raise FileNotFoundError(f\"Weight file {src_path} not found in dataset\")\n\n# Verify that the files exist\nfor weight_file in weight_files:\n    weight_path = os.path.join(weights_dir, weight_file)\n    if os.path.exists(weight_path):\n        print(f\"Found {weight_path}\")\n    else:\n        raise FileNotFoundError(f\"Failed to find {weight_path}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-06T15:49:22.237408Z","iopub.execute_input":"2025-04-06T15:49:22.237629Z","iopub.status.idle":"2025-04-06T15:49:23.587720Z","shell.execute_reply.started":"2025-04-06T15:49:22.237611Z","shell.execute_reply":"2025-04-06T15:49:23.586719Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Solution: Implement Clustering and Outlier Detection","metadata":{}},{"cell_type":"code","source":"!git clone https://github.com/magicleap/SuperGluePretrainedNetwork.git","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-06T15:49:23.588692Z","iopub.execute_input":"2025-04-06T15:49:23.588955Z","iopub.status.idle":"2025-04-06T15:49:23.764971Z","shell.execute_reply.started":"2025-04-06T15:49:23.588932Z","shell.execute_reply":"2025-04-06T15:49:23.763876Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import sys\nimport os\n\n# Path to the SuperGluePretrainedNetwork dataset\nsuperglue_dataset_path = '/kaggle/input/supergluepretrainednetwork/SuperGluePretrainedNetwork'\nsuperglue_working_path = '/kaggle/working/SuperGluePretrainedNetwork'\n\nif os.path.exists(superglue_dataset_path):\n    sys.path.append(superglue_dataset_path)\n    print(f\"Added {superglue_dataset_path} to sys.path\")\nelif os.path.exists(superglue_working_path):\n    sys.path.append(superglue_working_path)\n    print(f\"Added {superglue_working_path} to sys.path\")\nelse:\n    raise FileNotFoundError(\"SuperGluePretrainedNetwork directory not found in /kaggle/input/ or /kaggle/working/\")\n\nfrom models.matching import Matching\nfrom models.utils import frame2tensor","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Copy SuperGluePretrainedNetwork to /kaggle/working/\nif not os.path.exists(superglue_working_path):\n    shutil.copytree(superglue_dataset_path, superglue_working_path)\n    print(f\"Copied SuperGluePretrainedNetwork to {superglue_working_path}\")\nelse:\n    print(f\"{superglue_working_path} already exists, skipping copy\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-06T15:57:29.481059Z","iopub.execute_input":"2025-04-06T15:57:29.481458Z","iopub.status.idle":"2025-04-06T15:57:29.486672Z","shell.execute_reply.started":"2025-04-06T15:57:29.481424Z","shell.execute_reply":"2025-04-06T15:57:29.485668Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# === Import Libraries ===\nimport numpy as np\nfrom pathlib import Path\nimport cv2\nimport time\nimport torch\nimport sys\nimport warnings\nimport pandas as pd\nfrom scipy.optimize import least_squares\nimport os\nimport shutil\nfrom sklearn.cluster import DBSCAN\nwarnings.filterwarnings(\"ignore\", category=FutureWarning)\n\n# === Set Up SuperGluePretrainedNetwork ===\n# Path to the SuperGluePretrainedNetwork dataset\nsuperglue_dataset_path = '/kaggle/input/supergluepretrainednetwork/SuperGluePretrainedNetwork'\nsuperglue_working_path = '/kaggle/working/SuperGluePretrainedNetwork'\n\n# Add SuperGluePretrainedNetwork to sys.path\nif os.path.exists(superglue_dataset_path):\n    sys.path.append(superglue_dataset_path)\n    print(f\"Added {superglue_dataset_path} to sys.path\")\nelif os.path.exists(superglue_working_path):\n    sys.path.append(superglue_working_path)\n    print(f\"Added {superglue_working_path} to sys.path\")\nelse:\n    raise FileNotFoundError(\"SuperGluePretrainedNetwork directory not found in /kaggle/input/ or /kaggle/working/\")\n\n# Copy SuperGluePretrainedNetwork to /kaggle/working/\nif not os.path.exists(superglue_working_path):\n    shutil.copytree(superglue_dataset_path, superglue_working_path)\n    print(f\"Copied SuperGluePretrainedNetwork to {superglue_working_path}\")\nelse:\n    print(f\"{superglue_working_path} already exists, skipping copy\")\n\n# Import SuperGlue modules\nfrom models.matching import Matching\nfrom models.utils import frame2tensor\n\n# === Set Up SuperGlue Weights ===\n# Create the weights directory\nweights_dir = '/kaggle/working/SuperGluePretrainedNetwork/models/weights/'\nos.makedirs(weights_dir, exist_ok=True)\n\n# Path to the weights dataset (adjust based on your dataset name)\nweights_dataset_path = '/kaggle/input/superglue-weights/'\n\n# List of weight files\nweight_files = ['superpoint_v1.pth', 'superglue_indoor.pth', 'superglue_outdoor.pth']\n\n# Copy weights from the dataset to the expected directory\nfor weight_file in weight_files:\n    src_path = os.path.join(weights_dataset_path, weight_file)\n    dst_path = os.path.join(weights_dir, weight_file)\n    if os.path.exists(src_path):\n        shutil.copy(src_path, dst_path)\n        print(f\"Copied {weight_file} to {dst_path}\")\n    else:\n        raise FileNotFoundError(f\"Weight file {src_path} not found in dataset\")\n\n# Verify that the files exist\nfor weight_file in weight_files:\n    weight_path = os.path.join(weights_dir, weight_file)\n    if os.path.exists(weight_path):\n        print(f\"Found {weight_path}\")\n    else:\n        raise FileNotFoundError(f\"Failed to find {weight_path}\")\n\n# Rest of your code (save_ply, initialize_superglue, cluster_images, etc.) remains the same","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-06T15:57:32.183945Z","iopub.execute_input":"2025-04-06T15:57:32.184296Z","iopub.status.idle":"2025-04-06T15:57:32.386219Z","shell.execute_reply.started":"2025-04-06T15:57:32.184268Z","shell.execute_reply":"2025-04-06T15:57:32.385195Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Phase 1: Setup and Initial Functions\nimport numpy as np\nimport cv2\nimport pandas as pd\nfrom pathlib import Path\nimport time\nfrom scipy.spatial.transform import Rotation\nfrom scipy.optimize import least_squares\nfrom concurrent.futures import ProcessPoolExecutor\nimport multiprocessing\nimport os\n\n# Utility functions\ndef rotation_matrix_to_quaternion(R):\n    \"\"\"Convert a 3x3 rotation matrix to a quaternion [w, x, y, z].\"\"\"\n    if not np.all(np.isfinite(R)) or R.shape != (3, 3):\n        return np.array([1.0, 0.0, 0.0, 0.0], dtype=np.float64)\n    rot = Rotation.from_matrix(R)\n    q = rot.as_quat()  # Returns [x, y, z, w]\n    return np.array([q[3], q[0], q[1], q[2]], dtype=np.float64)  # Reorder to [w, x, y, z]\n\ndef quaternion_to_rotation_matrix(q):\n    \"\"\"Convert a quaternion [w, x, y, z] to a 3x3 rotation matrix.\"\"\"\n    if not np.all(np.isfinite(q)):\n        return np.eye(3, dtype=np.float64)\n    rot = Rotation.from_quat([q[1], q[2], q[3], q[0]])  # Reorder [x, y, z, w]\n    return rot.as_matrix()\n\ndef project(points3d, R, t, K):\n    \"\"\"Project 3D points to 2D using camera intrinsics and extrinsics.\"\"\"\n    if not (np.all(np.isfinite(points3d)) and np.all(np.isfinite(R)) and np.all(np.isfinite(t)) and np.all(np.isfinite(K))):\n        return np.zeros((points3d.shape[0], 2), dtype=np.float32)\n    points3d = points3d.T\n    P = K @ np.hstack((R, t.reshape(3, 1)))\n    points2d = P @ np.vstack((points3d, np.ones((1, points3d.shape[1]))))\n    points2d = points2d[:2] / (points2d[2] + 1e-8)\n    return points2d.T\n\ndef validate_pose(R, t, img_name=\"Unknown\"):\n    \"\"\"Validate that R and t have the correct shapes and are finite.\"\"\"\n    if R.shape != (3, 3):\n        print(f\"Invalid rotation matrix shape for {img_name}: {R.shape}\")\n        return False\n    if t.shape != (3,):\n        print(f\"Invalid translation vector shape for {img_name}: {t.shape}\")\n        return False\n    if not (np.all(np.isfinite(R)) and np.all(np.isfinite(t))):\n        print(f\"Non-finite values in pose for {img_name}: R={R}, t={t}\")\n        return False\n    return True\n\ndef normalize_translation(t, img_name=\"Unknown\"):\n    \"\"\"Ensure translation vector is a 1D array of shape (3,).\"\"\"\n    if t is None:\n        print(f\"Translation vector is None for {img_name}. Using default.\")\n        return np.array([0.0, 0.0, 0.0], dtype=np.float32)\n    t = np.array(t, dtype=np.float32).flatten()\n    if t.size != 3:\n        print(f\"Cannot normalize translation vector for {img_name}: size={t.size}\")\n        return np.array([0.0, 0.0, 0.0], dtype=np.float32)\n    return t\n\n# Load sample submission\nsample_submission = pd.read_csv('/kaggle/input/image-matching-challenge-2025/sample_submission.csv')\nprint(\"Columns in sample_submission.csv:\", sample_submission.columns.tolist())\n\nimage_path_column = 'image'\nif image_path_column not in sample_submission.columns:\n    raise ValueError(\"Expected 'image' column in sample_submission.csv\")\n\n# Function to infer the correct scene\ndef infer_scene_from_image(dataset, image_name):\n    if dataset == 'ETs':\n        if image_name.startswith('et_et') or image_name.startswith('another_et_another_et') or image_name.startswith('outliers_out_et'):\n            return 'et'\n    elif dataset == 'amy_gardens':\n        if image_name.startswith('peach_'):\n            return 'peach'\n    elif dataset == 'stairs':\n        if image_name.startswith('stairs_split_1_'):\n            return 'stairs_split_1'\n        elif image_name.startswith('stairs_split_2_'):\n            return 'stairs_split_2'\n    elif dataset == 'pt_stpeters_stpauls':\n        if image_name.startswith('st_peters_square_'):\n            return 'st_peters_square'\n        return 'pt_stpeters_stpauls'\n    elif dataset == 'imc2024_dioscuri_baalshamin':\n        return 'dioscuri_baalshamin'\n    elif dataset == 'imc2024_lizard_pond':\n        return 'lizard_pond'\n    elif dataset == 'imc2023_haiper':\n        return 'haiper'\n    elif dataset == 'imc2023_heritage':\n        return 'heritage'\n    elif dataset == 'imc2023_theather_imc2024_church':\n        return 'theather_imc2024_church'\n    elif dataset == 'fbk_vineyard':\n        return 'fbk_vineyard'\n    elif dataset == 'pt_brandenburg_british_buckingham':\n        return 'brandenburg_british_buckingham'\n    elif dataset == 'pt_piazzasanmarco_grandplace':\n        return 'piazzasanmarco_grandplace'\n    elif dataset == 'pt_sacrecoeur_trevi_tajmahal':\n        return 'sacrecoeur_trevi_tajmahal'\n    return dataset\n\n# Add inferred scene column\nsample_submission['inferred_scene'] = sample_submission.apply(\n    lambda row: infer_scene_from_image(row['dataset'], row['image']), axis=1\n)\n\n# Group images by dataset and inferred scene\ngrouped_images = sample_submission.groupby(['dataset', 'inferred_scene'])\nprint(\"Grouped dataset and inferred scene combinations:\", list(grouped_images.groups.keys()))\n\n# Initialize summary tracking\nsummary = {\n    'total_groups': len(grouped_images),\n    'successful_groups': 0,\n    'failed_groups': [],\n    'total_images': len(sample_submission),\n    'computed_poses': 0,\n    'sequential_poses': 0,\n    'images_not_found': 0,\n    'high_reprojection_errors': 0\n}\n\n# Phase 2: Feature Matching (Using SIFT)\ndef cluster_images(image_paths, scene_name, min_matches_threshold=10):\n    \"\"\"Cluster images using a sliding window approach for efficiency.\"\"\"\n    print(f\"Clustering images for scene: {scene_name}\")\n    start_time = time.time()\n    num_images = len(image_paths)\n    if num_images < 2:\n        print(\"Fewer than 2 images. Using sequential clustering.\")\n        return [list(range(num_images))], []\n\n    window_size = 7\n    edges = []\n\n    # Create SIFT and FLANN matcher for clustering\n    sift = cv2.SIFT_create(\n        nfeatures=5000,  # Increased to get more features\n        contrastThreshold=0.01,\n        edgeThreshold=15\n    )\n    FLANN_INDEX_KDTREE = 1\n    index_params = dict(algorithm=FLANN_INDEX_KDTREE, trees=5)\n    search_params = dict(checks=50)\n    flann = cv2.FlannBasedMatcher(index_params, search_params)\n\n    for i in range(num_images):\n        for j in range(i + 1, min(i + window_size + 1, num_images)):\n            img0 = cv2.imread(image_paths[i], cv2.IMREAD_GRAYSCALE)\n            img1 = cv2.imread(image_paths[j], cv2.IMREAD_GRAYSCALE)\n            if img0 is None or img1 is None:\n                print(f\"Failed to load images: {image_paths[i]} or {image_paths[j]}\")\n                continue\n            img0 = cv2.equalizeHist(img0)\n            img1 = cv2.equalizeHist(img1)\n            img0 = cv2.resize(img0, (320, 240))\n            img1 = cv2.resize(img1, (320, 240))\n            kp0, des0 = sift.detectAndCompute(img0, None)\n            kp1, des1 = sift.detectAndCompute(img1, None)\n            if des0 is None or des1 is None:\n                continue\n            matches = flann.knnMatch(des0, des1, k=2)\n            good_matches = []\n            for m, n in matches:\n                if m.distance < 0.6 * n.distance:\n                    good_matches.append(m)\n            num_matches = len(good_matches)\n            if num_matches >= min_matches_threshold:\n                edges.append((num_matches, i, j))\n\n    if not edges:\n        print(\"No matches found between any image pairs. Using sequential clustering.\")\n        return [list(range(num_images))], []\n\n    # Build a graph and find connected components\n    from collections import defaultdict\n    graph = defaultdict(list)\n    for _, i, j in edges:\n        graph[i].append(j)\n        graph[j].append(i)\n\n    def find_component(start, graph, visited):\n        component = []\n        stack = [start]\n        while stack:\n            node = stack.pop()\n            if node not in visited:\n                visited.add(node)\n                component.append(node)\n                stack.extend(n for n in graph[node] if n not in visited)\n        return component\n\n    visited = set()\n    clusters = []\n    for i in range(num_images):\n        if i not in visited:\n            component = find_component(i, graph, visited)\n            if len(component) >= 2:\n                clusters.append(sorted(component))\n\n    noise_indices = [i for i in range(num_images) if i not in visited]\n\n    if not clusters:\n        print(\"No clusters found. Using sequential clustering.\")\n        clusters = [list(range(num_images))]\n\n    print(f\"Clustering took {time.time() - start_time:.2f} seconds\")\n    return clusters, noise_indices\n\n# Phase 3: Pose Estimation\ndef recover_pose_with_fallback(mkpts0, mkpts1, K, img0_name, img1_name):\n    \"\"\"Recover pose between two images with robust fallbacks.\"\"\"\n    if len(mkpts0) < 5:\n        print(f\"Too few matches ({len(mkpts0)}) between {img0_name} and {img1_name}. Cannot estimate pose.\")\n        return None, None, 0\n\n    max_iters = 5000\n    E, mask = cv2.findEssentialMat(\n        mkpts0, mkpts1, K, method=cv2.RANSAC, prob=0.999, threshold=0.5, maxIters=max_iters\n    )\n    if E is None or E.shape != (3, 3):\n        print(f\"Essential matrix estimation failed for pair {img0_name} and {img1_name}. Trying homography.\")\n    else:\n        inliers, R, t, mask = cv2.recoverPose(E, mkpts0, mkpts1, K, mask=mask)\n        if inliers >= 5:\n            P0 = K @ np.hstack((np.eye(3), np.zeros((3, 1))))\n            P1 = K @ np.hstack((R, t.reshape(3, 1)))\n            points4d = cv2.triangulatePoints(P0, P1, mkpts0.T, mkpts1.T)\n            points3d = points4d[:3] / (points4d[3] + 1e-8)\n            points3d_cam2 = R @ points3d + t.reshape(3, 1)\n            in_front = (points3d[2, :] > 0) & (points3d_cam2[2, :] > 0)\n            if np.sum(in_front) >= 0.5 * inliers:\n                t = t.flatten()  # Ensure t is (3,)\n                if validate_pose(R, t, f\"{img0_name}-{img1_name}\"):\n                    return R, t, inliers\n\n    print(f\"Falling back to homography for pair {img0_name} and {img1_name}\")\n    H, mask = cv2.findHomography(mkpts0, mkpts1, cv2.RANSAC, 0.5, maxIters=5000)\n    if H is None or H.shape != (3, 3):\n        print(f\"Homography estimation failed for pair {img0_name} and {img1_name}\")\n        return None, None, 0\n\n    inliers = np.sum(mask)\n    if inliers < 5:\n        print(f\"Too few inliers from homography for pair {img0_name} and {img1_name}: {inliers}\")\n        return None, None, 0\n\n    num, Rs, ts, normals = cv2.decomposeHomographyMat(H, K)\n    best_R, best_t, best_inliers = None, None, 0\n    for i in range(num):\n        R, t = Rs[i], ts[i]\n        P0 = K @ np.hstack((np.eye(3), np.zeros((3, 1))))\n        P1 = K @ np.hstack((R, t))\n        points4d = cv2.triangulatePoints(P0, P1, mkpts0.T, mkpts1.T)\n        points3d = points4d[:3] / (points4d[3] + 1e-8)\n        points3d_cam2 = R @ points3d + t.reshape(3, 1)\n        in_front = (points3d[2, :] > 0) & (points3d_cam2[2, :] > 0)\n        inliers_count = np.sum(in_front)\n        if inliers_count > best_inliers:\n            best_inliers = inliers_count\n            best_R, best_t = R, t\n\n    if best_R is None or best_t is None:\n        print(f\"No valid pose from homography for pair {img0_name} and {img1_name}\")\n        return None, None, 0\n\n    best_t = best_t.flatten()  # Ensure t is (3,)\n    if validate_pose(best_R, best_t, f\"{img0_name}-{img1_name}\"):\n        return best_R, best_t, best_inliers\n    return None, None, 0\n\ndef reprojection_error(params, points3d, observations, K, img_names, poses):\n    \"\"\"Compute reprojection error for bundle adjustment.\"\"\"\n    num_cameras = len(img_names)\n    translations = params[:num_cameras * 3].reshape(num_cameras, 3)\n    focal_length = params[num_cameras * 3]\n    cx = params[num_cameras * 3 + 1]\n    cy = params[num_cameras * 3 + 2]\n\n    K_opt = np.array([\n        [focal_length, 0, cx],\n        [0, focal_length, cy],\n        [0, 0, 1]\n    ], dtype=np.float32)\n\n    errors = []\n    for i, img_name in enumerate(img_names):\n        R, _ = poses[img_name]\n        t = translations[i]\n        for j, obs in enumerate(observations):\n            for img_obs, pt in obs:\n                if img_obs == img_name:\n                    proj = project(points3d[j:j+1], R, t, K_opt)\n                    if not np.all(np.isfinite(proj)):\n                        continue\n                    errors.append(proj[0] - pt)\n\n    if len(errors) == 0:\n        print(\"No reprojection errors computed. Returning zeros.\")\n        return np.zeros(num_cameras * 3 + 3)\n    errors = np.concatenate(errors)\n    if not np.all(np.isfinite(errors)):\n        return np.zeros_like(errors)\n    return errors\n\n# Phase 4: Feature Matching\ndef match_image_pair(args):\n    \"\"\"Match keypoints between a pair of images.\"\"\"\n    idx0, idx1, global_idx0, global_idx1, image_paths = args\n    img0 = cv2.imread(image_paths[global_idx0], cv2.IMREAD_GRAYSCALE)\n    img1 = cv2.imread(image_paths[global_idx1], cv2.IMREAD_GRAYSCALE)\n\n    if img0 is None or img1 is None:\n        print(f\"Failed to load images: {image_paths[global_idx0]} or {image_paths[global_idx1]}\")\n        return idx0, idx1, None, None, None, None\n\n    img0 = cv2.equalizeHist(img0)\n    img1 = cv2.equalizeHist(img1)\n\n    orig_size0 = img0.shape[::-1]\n    orig_size1 = img1.shape[::-1]\n\n    img0_resized = cv2.resize(img0, (320, 240))\n    img1_resized = cv2.resize(img1, (320, 240))\n\n    # Create SIFT and FLANN matcher inside the function\n    sift = cv2.SIFT_create(\n        nfeatures=5000,  # Increased to get more features\n        contrastThreshold=0.01,\n        edgeThreshold=15\n    )\n    FLANN_INDEX_KDTREE = 1\n    index_params = dict(algorithm=FLANN_INDEX_KDTREE, trees=5)\n    search_params = dict(checks=50)\n    flann = cv2.FlannBasedMatcher(index_params, search_params)\n\n    kp0, des0 = sift.detectAndCompute(img0_resized, None)\n    kp1, des1 = sift.detectAndCompute(img1_resized, None)\n    if des0 is None or des1 is None:\n        return idx0, idx1, None, None, None, None\n\n    matches = flann.knnMatch(des0, des1, k=2)\n    good_matches = []\n    for m, n in matches:\n        if m.distance < 0.6 * n.distance:\n            good_matches.append(m)\n\n    if not good_matches:\n        print(f\"No good matches between {Path(image_paths[global_idx0]).name} and {Path(image_paths[global_idx1]).name}\")\n        return idx0, idx1, None, None, None, None\n\n    mkpts0 = np.array([kp0[m.queryIdx].pt for m in good_matches])\n    mkpts1 = np.array([kp1[m.trainIdx].pt for m in good_matches])\n    kpts0 = np.array([kp.pt for kp in kp0])\n    kpts1 = np.array([kp.pt for kp in kp1])\n\n    scale0 = (orig_size0[0] / 320, orig_size0[1] / 240)\n    scale1 = (orig_size1[0] / 320, orig_size1[1] / 240)\n    kpts0[:, 0] *= scale0[0]\n    kpts0[:, 1] *= scale0[1]\n    kpts1[:, 0] *= scale1[0]\n    kpts1[:, 1] *= scale1[1]\n    mkpts0[:, 0] *= scale0[0]\n    mkpts0[:, 1] *= scale0[1]\n    mkpts1[:, 0] *= scale1[0]\n    mkpts1[:, 1] *= scale1[1]\n\n    img0_name = Path(image_paths[global_idx0]).name\n    img1_name = Path(image_paths[global_idx1]).name\n\n    return idx0, idx1, img0_name, img1_name, (kpts0, kpts1), (mkpts0, mkpts1)\n\ndef run_sfm_on_cluster(args):\n    \"\"\"Run Structure-from-Motion (SfM) on a cluster of images using incremental SfM.\"\"\"\n    cluster_indices, image_paths, scene_name, cluster_idx, clusters, overlap = args\n    print(f\"Starting cluster {cluster_idx} for scene {scene_name}\")\n\n    if not cluster_indices or not image_paths:\n        print(\"Empty cluster or image paths. Using sequential poses.\")\n        poses = {}\n        for idx in range(len(cluster_indices)):\n            if idx >= len(cluster_indices) or cluster_indices[idx] >= len(image_paths):\n                continue\n            img_name = Path(image_paths[cluster_indices[idx]]).name\n            R = np.eye(3, dtype=np.float32)\n            t = np.array([0.0, 0.0, idx * 0.1], dtype=np.float32)\n            poses[img_name] = (R, t)\n        image_width, image_height = 320, 240\n        focal_length = max(image_width, image_height) * 1.2\n        cx, cy = image_width / 2, image_height / 2\n        K = np.array([\n            [focal_length, 0, cx],\n            [0, focal_length, cy],\n            [0, 0, 1]\n        ], dtype=np.float32)\n        return poses, np.array([], dtype=np.float32).reshape(0, 3), K, []\n\n    keypoints_dict = {}\n    matches_dict = {}\n    image_sizes = {}\n    point_to_images = []\n\n    start_time = time.time()\n    window_size = 7\n    image_pairs = []\n    for i in range(len(cluster_indices)):\n        for j in range(i + 1, min(i + window_size + 1, len(cluster_indices))):\n            image_pairs.append((i, j))\n\n    edges = []\n    max_workers = min(multiprocessing.cpu_count(), 4)\n    with ProcessPoolExecutor(max_workers=max_workers) as executor:\n        futures = [\n            executor.submit(match_image_pair, (idx0, idx1, cluster_indices[idx0], cluster_indices[idx1], image_paths))\n            for idx0, idx1 in image_pairs\n        ]\n        for future in futures:\n            idx0, idx1, img0_name, img1_name, keypoints, matches = future.result()\n            if matches is None:\n                continue\n            kpts0, kpts1 = keypoints\n            mkpts0, mkpts1 = matches\n            if img0_name not in keypoints_dict:\n                keypoints_dict[img0_name] = kpts0\n                image_sizes[img0_name] = (kpts0.shape[0] * 320 / mkpts0.shape[0], kpts0.shape[1] * 240 / mkpts0.shape[1]) if mkpts0.shape[0] > 0 else (320, 240)\n            if img1_name not in keypoints_dict:\n                keypoints_dict[img1_name] = kpts1\n                image_sizes[img1_name] = (kpts1.shape[0] * 320 / mkpts1.shape[0], kpts1.shape[1] * 240 / mkpts1.shape[1]) if mkpts1.shape[0] > 0 else (320, 240)\n            matches_dict[(img0_name, img1_name)] = (mkpts0, mkpts1)\n            priority = 0\n            if cluster_idx < len(clusters) - 1:\n                overlap_indices = clusters[cluster_idx][-overlap:]\n                if cluster_indices[idx0] in overlap_indices or cluster_indices[idx1] in overlap_indices:\n                    priority = 1\n            edges.append((len(mkpts0), idx0, idx1, priority))\n            print(f\"Matches between {img0_name} and {img1_name}: {len(mkpts0)}\")\n\n    print(f\"Feature matching for cluster took {time.time() - start_time:.2f} seconds.\")\n\n    edges.sort(reverse=True)\n    poses = {}\n    added_images = set()\n\n    # Initialize the first pair\n    for num_matches, idx0, idx1, priority in edges:\n        if num_matches < 5:\n            continue\n        if idx0 >= len(cluster_indices) or idx1 >= len(cluster_indices):\n            continue\n        global_idx0 = cluster_indices[idx0]\n        global_idx1 = cluster_indices[idx1]\n        if global_idx0 >= len(image_paths) or global_idx1 >= len(image_paths):\n            continue\n        img0_name = Path(image_paths[global_idx0]).name\n        img1_name = Path(image_paths[global_idx1]).name\n        mkpts0, mkpts1 = matches_dict[(img0_name, img1_name)]\n        image_width, image_height = image_sizes[img0_name]\n        focal_length = max(image_width, image_height) * 1.2\n        cx, cy = image_width / 2, image_height / 2\n        K = np.array([\n            [focal_length, 0, cx],\n            [0, focal_length, cy],\n            [0, 0, 1]\n        ], dtype=np.float32)\n        R, t, inliers = recover_pose_with_fallback(mkpts0, mkpts1, K, img0_name, img1_name)\n        if R is None or t is None:\n            continue\n        t = normalize_translation(t, img0_name)\n        poses[img0_name] = (np.eye(3, dtype=np.float32), np.zeros(3, dtype=np.float32))\n        poses[img1_name] = (R, t)\n        added_images.update([img0_name, img1_name])\n        break\n\n    if not poses:\n        print(f\"No initial pair found for cluster {cluster_idx}. Using sequential poses.\")\n        for idx in range(len(cluster_indices)):\n            if idx >= len(cluster_indices) or cluster_indices[idx] >= len(image_paths):\n                continue\n            img_name = Path(image_paths[cluster_indices[idx]]).name\n            R = np.eye(3, dtype=np.float32)\n            t = np.array([0.0, 0.0, idx * 0.1], dtype=np.float32)\n            poses[img_name] = (R, t)\n            added_images.add(img_name)\n        image_width, image_height = 320, 240\n        focal_length = max(image_width, image_height) * 1.2\n        cx, cy = image_width / 2, image_height / 2\n        K = np.array([\n            [focal_length, 0, cx],\n            [0, focal_length, cy],\n            [0, 0, 1]\n        ], dtype=np.float32)\n        return poses, np.array([], dtype=np.float32).reshape(0, 3), K, []\n\n    # Incremental SfM: Add images one at a time\n    points3d = []\n    point_to_images = []\n    for idx in range(len(cluster_indices)):\n        if idx >= len(cluster_indices) or cluster_indices[idx] >= len(image_paths):\n            continue\n        global_idx = cluster_indices[idx]\n        img_name = Path(image_paths[global_idx]).name\n        if img_name in added_images:\n            continue\n\n        # Find the best image to match against\n        best_pair = None\n        best_num_matches = 0\n        best_idx0 = None\n        for idx0 in range(len(cluster_indices)):\n            if idx0 >= len(cluster_indices) or cluster_indices[idx0] >= len(image_paths):\n                continue\n            img0_name = Path(image_paths[cluster_indices[idx0]]).name\n            if img0_name not in added_images:\n                continue\n            # Check both (img0_name, img_name) and (img_name, img0_name) for matches\n            pair_key = (img0_name, img_name) if (img0_name, img_name) in matches_dict else (img_name, img0_name)\n            if pair_key not in matches_dict:\n                # Match on-the-fly if not already matched\n                _, _, _, _, _, matches = match_image_pair((idx0, idx, cluster_indices[idx0], global_idx, image_paths))\n                if matches is None:\n                    matches_dict[(img0_name, img_name)] = (np.array([]), np.array([]))\n                    num_matches = 0\n                else:\n                    mkpts0, mkpts1 = matches\n                    matches_dict[(img0_name, img_name)] = (mkpts0, mkpts1)\n                    num_matches = len(mkpts0)\n            else:\n                mkpts0, mkpts1 = matches_dict[pair_key]\n                num_matches = len(mkpts0)\n\n            if num_matches > best_num_matches:\n                best_num_matches = num_matches\n                best_pair = (img0_name, img_name)\n                best_idx0 = idx0\n\n        if best_pair is None or best_num_matches < 5:\n            print(f\"Could not add {img_name} to cluster {cluster_idx}. Using sequential pose.\")\n            R = np.eye(3, dtype=np.float32)\n            t = np.array([0.0, 0.0, len(added_images) * 0.1], dtype=np.float32)\n            poses[img_name] = (R, t)\n            added_images.add(img_name)\n            continue\n\n        img0_name, img1_name = best_pair\n        # Use pair_key to access matches_dict\n        pair_key = (img0_name, img1_name) if (img0_name, img1_name) in matches_dict else (img1_name, img0_name)\n        mkpts0, mkpts1 = matches_dict[pair_key]\n        if len(mkpts0) < 5:\n            print(f\"Too few matches ({len(mkpts0)}) for pair {img0_name} and {img1_name}. Using sequential pose for {img1_name}.\")\n            R = np.eye(3, dtype=np.float32)\n            t = np.array([0.0, 0.0, len(added_images) * 0.1], dtype=np.float32)\n            poses[img1_name] = (R, t)\n            added_images.add(img1_name)\n            continue\n\n        R0, t0 = poses[img0_name]\n        R, t, inliers = recover_pose_with_fallback(mkpts0, mkpts1, K, img0_name, img1_name)\n        if R is None or t is None:\n            print(f\"Failed to estimate pose for {img1_name}. Using sequential pose.\")\n            R = np.eye(3, dtype=np.float32)\n            t = np.array([0.0, 0.0, len(added_images) * 0.1], dtype=np.float32)\n            poses[img1_name] = (R, t)\n            added_images.add(img1_name)\n            continue\n\n        t = normalize_translation(t, img1_name)\n        t0 = normalize_translation(t0, img0_name)\n        R = R0 @ R\n        t = R0 @ t + t0\n        t = normalize_translation(t, img1_name)\n        if not validate_pose(R, t, img1_name):\n            print(f\"Invalid pose for {img1_name}. Using sequential pose.\")\n            R = np.eye(3, dtype=np.float32)\n            t = np.array([0.0, 0.0, len(added_images) * 0.1], dtype=np.float32)\n        poses[img1_name] = (R, t)\n        added_images.add(img1_name)\n\n        # Triangulate points with the best pair\n        if len(mkpts0) >= 5:\n            R1, t1 = poses[img1_name]\n            t0 = normalize_translation(t0, img0_name)\n            t1 = normalize_translation(t1, img1_name)\n            if validate_pose(R0, t0, img0_name) and validate_pose(R1, t1, img1_name):\n                P0 = K @ np.hstack((R0, t0.reshape(3, 1)))\n                P1 = K @ np.hstack((R1, t1.reshape(3, 1)))\n                points4d = cv2.triangulatePoints(P0, P1, mkpts0.T, mkpts1.T)\n                points4d = points4d[:3] / (points4d[3] + 1e-8)\n                points4d = points4d.T\n                if np.all(np.isfinite(points4d)):\n                    for i in range(len(points4d)):\n                        points3d.append(points4d[i])\n                        point_to_images.append([(img0_name, mkpts0[i]), (img1_name, mkpts1[i])])\n\n    points3d = np.array(points3d, dtype=np.float32) if points3d else np.array([], dtype=np.float32).reshape(0, 3)\n    print(f\"Reconstructed {len(points3d)} 3D points in cluster {cluster_idx}\")\n\n    # Local Bundle Adjustment (skip for large clusters to save time)\n    if len(poses) >= 3 and len(point_to_images) >= 10 and len(poses) < 20:\n        start_time = time.time()\n        img_names = list(poses.keys())\n        translations = np.array([normalize_translation(poses[img_name][1], img_name) for img_name in img_names], dtype=np.float32)\n        if translations.shape != (len(img_names), 3):\n            print(f\"Mismatch in translations shape: {translations.shape}. Expected ({len(img_names)}, 3). Skipping bundle adjustment.\")\n        else:\n            initial_focal_length = K[0, 0]\n            initial_cx, initial_cy = K[0, 2], K[1, 2]\n            params = np.hstack((translations.ravel(), initial_focal_length, initial_cx, initial_cy))\n            max_nfev = min(20, len(poses) * 5)\n            result = least_squares(\n                reprojection_error,\n                params,\n                args=(points3d, point_to_images, K, img_names, poses),\n                max_nfev=max_nfev,\n                ftol=1e-5,\n                xtol=1e-5\n            )\n            optimized_params = result.x\n            num_cameras = len(img_names)\n            translations = optimized_params[:num_cameras * 3].reshape(num_cameras, 3)\n            focal_length = optimized_params[num_cameras * 3]\n            cx = optimized_params[num_cameras * 3 + 1]\n            cy = optimized_params[num_cameras * 3 + 2]\n            K = np.array([\n                [focal_length, 0, cx],\n                [0, focal_length, cy],\n                [0, 0, 1]\n            ], dtype=np.float32)\n            for i, img_name in enumerate(img_names):\n                R, _ = poses[img_name]\n                t = translations[i]\n                t = normalize_translation(t, img_name)\n                if not validate_pose(R, t, img_name):\n                    R = np.eye(3, dtype=np.float32)\n                    t = np.array([0.0, 0.0, len(poses) * 0.1], dtype=np.float32)\n                poses[img_name] = (R, t)\n            print(f\"Local bundle adjustment took {time.time() - start_time:.2f} seconds.\")\n    else:\n        print(\"Skipping local bundle adjustment due to insufficient poses/observations or large cluster size.\")\n\n    print(f\"Cluster {cluster_idx} processed {len(added_images)} cameras\")\n    return poses, points3d, K, point_to_images\n\n# Phase 5: Reprojection Error\ndef compute_reprojection_error(img_name, points3d, observations, R, t, K):\n    \"\"\"Compute the average reprojection error for an image.\"\"\"\n    errors = []\n    for j, obs in enumerate(observations):\n        for img_obs, pt in obs:\n            if img_obs == img_name:\n                proj = project(points3d[j:j+1], R, t, K)\n                if not np.all(np.isfinite(proj)):\n                    continue\n                error = np.linalg.norm(proj[0] - pt)\n                if np.isfinite(error):\n                    errors.append(error)\n    return np.mean(errors) if errors else float('inf')\n\n# Phase 6: Main Loop and Submission\nall_poses = {}\nall_Ks = {}\nprocessed_images = set()\npose_status = {}\n\nfor (dataset, inferred_scene), group in grouped_images:\n    print(f\"\\nProcessing dataset: {dataset}, inferred scene: {inferred_scene}\")\n    start_time = time.time()\n\n    image_dir = f\"/kaggle/input/image-matching-challenge-2025/test/{dataset}\"\n    if not os.path.exists(image_dir):\n        print(f\"Dataset directory {image_dir} does not exist in test set. Skipping.\")\n        summary['failed_groups'].append((dataset, inferred_scene, \"Dataset not found in test directory\"))\n        summary['images_not_found'] += len(group)\n        continue\n\n    image_paths = []\n    for _, row in group.iterrows():\n        img_name = row['image']\n        img_path = f\"{image_dir}/{img_name}\"\n        if Path(img_path).exists():\n            image_paths.append(img_path)\n            processed_images.add(img_name)\n        else:\n            print(f\"Image not found: {img_path}\")\n            summary['images_not_found'] += 1\n\n    if not image_paths:\n        print(f\"No images found for dataset {dataset}, scene {inferred_scene}. Skipping.\")\n        summary['failed_groups'].append((dataset, inferred_scene, \"No images found\"))\n        continue\n\n    if len(image_paths) < 2:\n        print(f\"Dataset {dataset}, scene {inferred_scene} has fewer than 2 images. Using sequential poses.\")\n        for idx, img_path in enumerate(image_paths):\n            img_name = Path(img_path).name\n            R = np.eye(3, dtype=np.float32)\n            t = np.array([0.0, 0.0, idx * 0.1], dtype=np.float32)\n            all_poses[img_name] = (R, t)\n            image_width, image_height = 320, 240\n            focal_length = max(image_width, image_height) * 1.2\n            cx, cy = image_width / 2, image_height / 2\n            K = np.array([\n                [focal_length, 0, cx],\n                [0, focal_length, cy],\n                [0, 0, 1]\n            ], dtype=np.float32)\n            all_Ks[img_name] = K\n            pose_status[img_name] = 'sequential'\n            summary['sequential_poses'] += 1\n        summary['failed_groups'].append((dataset, inferred_scene, \"Fewer than 2 images\"))\n        continue\n\n    clusters, noise_indices = cluster_images(image_paths, inferred_scene, min_matches_threshold=10)\n    print(f\"Clusters: {clusters}\")\n    print(f\"Outlier images: {[Path(image_paths[idx]).name for idx in noise_indices if idx < len(image_paths)]}\")\n\n    overlap = 2\n\n    all_points3d = []\n    all_observations = []\n\n    print(f\"Processing clusters for dataset {dataset}, scene {inferred_scene}\")\n    for cluster_idx, cluster_indices in enumerate(clusters):\n        result = run_sfm_on_cluster((cluster_indices, image_paths, inferred_scene, cluster_idx, clusters, overlap))\n        poses, points3d, K, cluster_point_to_images = result\n        for img_name in poses:\n            R, t = poses[img_name]\n            t = normalize_translation(t, img_name)\n            if not validate_pose(R, t, img_name):\n                print(f\"Invalid pose for {img_name} after SfM. Using sequential pose.\")\n                R = np.eye(3, dtype=np.float32)\n                t = np.array([0.0, 0.0, len(all_poses) * 0.1], dtype=np.float32)\n            all_poses[img_name] = (R, t)\n            all_Ks[img_name] = K\n            pose_status[img_name] = 'computed'\n            summary['computed_poses'] += 1\n        if len(points3d) > 0:\n            all_points3d.append(points3d)\n        all_observations.extend(cluster_point_to_images)\n\n    # Handle outliers incrementally\n    for idx in noise_indices:\n        if idx >= len(image_paths):\n            continue\n        img_name = Path(image_paths[idx]).name\n        best_pair = None\n        best_num_matches = 0\n        for i in range(max(0, idx - 3), min(len(image_paths), idx + 4)):\n            if i == idx or Path(image_paths[i]).name not in all_poses:\n                continue\n            img0_name = Path(image_paths[i]).name\n            pair_key = (img0_name, img_name) if (img0_name, img_name) in matches_dict else (img_name, img0_name)\n            if pair_key not in matches_dict:\n                _, _, _, _, _, matches = match_image_pair((i, idx, i, idx, image_paths))\n                if matches is None:\n                    matches_dict[(img0_name, img_name)] = (np.array([]), np.array([]))\n                    num_matches = 0\n                else:\n                    mkpts0, mkpts1 = matches\n                    matches_dict[(img0_name, img_name)] = (mkpts0, mkpts1)\n                    num_matches = len(mkpts0)\n            else:\n                mkpts0, mkpts1 = matches_dict[pair_key]\n                num_matches = len(mkpts0)\n            if num_matches > best_num_matches:\n                best_num_matches = num_matches\n                best_pair = (img0_name, img_name)\n\n        if best_pair is None or best_num_matches < 5:\n            print(f\"Could not register outlier image {img_name}. Using sequential pose.\")\n            R = np.eye(3, dtype=np.float32)\n            t = np.array([0.0, 0.0, len(all_poses) * 0.1], dtype=np.float32)\n            all_poses[img_name] = (R, t)\n            image_width, image_height = 320, 240\n            focal_length = max(image_width, image_height) * 1.2\n            cx, cy = image_width / 2, image_height / 2\n            K = np.array([\n                [focal_length, 0, cx],\n                [0, focal_length, cy],\n                [0, 0, 1]\n            ], dtype=np.float32)\n            all_Ks[img_name] = K\n            pose_status[img_name] = 'sequential'\n            summary['sequential_poses'] += 1\n            continue\n\n        img0_name, img1_name = best_pair\n        pair_key = (img0_name, img1_name) if (img0_name, img1_name) in matches_dict else (img1_name, img0_name)\n        mkpts0, mkpts1 = matches_dict[pair_key]\n        image_width, image_height = 320, 240\n        focal_length = max(image_width, image_height) * 1.2\n        cx, cy = image_width / 2, image_height / 2\n        K = np.array([\n            [focal_length, 0, cx],\n            [0, focal_length, cy],\n            [0, 0, 1]\n        ], dtype=np.float32)\n        R, t, inliers = recover_pose_with_fallback(mkpts0, mkpts1, K, img0_name, img1_name)\n        if R is None or t is None:\n            print(f\"Failed to register outlier image {img_name}. Using sequential pose.\")\n            R = np.eye(3, dtype=np.float32)\n            t = np.array([0.0, 0.0, len(all_poses) * 0.1], dtype=np.float32)\n            all_poses[img_name] = (R, t)\n            all_Ks[img_name] = K\n            pose_status[img_name] = 'sequential'\n            summary['sequential_poses'] += 1\n            continue\n        R0, t0 = all_poses[img0_name]\n        t0 = normalize_translation(t0, img0_name)\n        t = normalize_translation(t, img1_name)\n        R = R0 @ R\n        t = R0 @ t + t0\n        t = normalize_translation(t, img1_name)\n        if not validate_pose(R, t, img1_name):\n            print(f\"Invalid pose for outlier {img1_name}. Using sequential pose.\")\n            R = np.eye(3, dtype=np.float32)\n            t = np.array([0.0, 0.0, len(all_poses) * 0.1], dtype=np.float32)\n            pose_status[img_name] = 'sequential'\n            summary['sequential_poses'] += 1\n        all_poses[img_name] = (R, t)\n        all_Ks[img_name] = K\n        pose_status[img_name] = 'computed'\n        summary['computed_poses'] += 1\n\n    if all_points3d:\n        merged_points3d = np.concatenate(all_points3d, axis=0)\n    else:\n        merged_points3d = np.array([], dtype=np.float32).reshape(0, 3)\n\n    print(\"Skipping global bundle adjustment to save time.\")\n\n    reproj_threshold = 10.0\n    for img_name in list(all_poses.keys()):\n        R, t = all_poses[img_name]\n        t = normalize_translation(t, img_name)\n        if not validate_pose(R, t, img_name):\n            print(f\"Invalid pose for {img_name} before reprojection check. Using sequential pose.\")\n            R = np.eye(3, dtype=np.float32)\n            t = np.array([0.0, 0.0, len(all_poses) * 0.1], dtype=np.float32)\n            all_poses[img_name] = (R, t)\n            pose_status[img_name] = 'sequential'\n            summary['sequential_poses'] += 1\n        if img_name not in all_Ks:\n            print(f\"Camera intrinsics (K) missing for {img_name}. Assigning default K.\")\n            image_width, image_height = 320, 240\n            focal_length = max(image_width, image_height) * 1.2\n            cx, cy = image_width / 2, image_height / 2\n            K = np.array([\n                [focal_length, 0, cx],\n                [0, focal_length, cy],\n                [0, 0, 1]\n            ], dtype=np.float32)\n            all_Ks[img_name] = K\n        K = all_Ks[img_name]\n        error = compute_reprojection_error(img_name, merged_points3d, all_observations, R, t, K)\n        if error > reproj_threshold:\n            print(f\"Image {img_name} has high reprojection error ({error:.2f}). Resetting to sequential pose.\")\n            R = np.eye(3, dtype=np.float32)\n            t = np.array([0.0, 0.0, len(all_poses) * 0.1], dtype=np.float32)\n            all_poses[img_name] = (R, t)\n            pose_status[img_name] = 'sequential'\n            summary['sequential_poses'] += 1\n            summary['high_reprojection_errors'] += 1\n            image_width, image_height = 320, 240\n            focal_length = max(image_width, image_height) * 1.2\n            cx, cy = image_width / 2, image_height / 2\n            K = np.array([\n                [focal_length, 0, cx],\n                [0, focal_length, cy],\n                [0, 0, 1]\n            ], dtype=np.float32)\n            all_Ks[img_name] = K\n\n    for img_path in image_paths:\n        img_name = Path(img_path).name\n        if img_name not in all_poses:\n            print(f\"Image {img_name} was not processed. Using sequential pose.\")\n            R = np.eye(3, dtype=np.float32)\n            t = np.array([0.0, 0.0, len(all_poses) * 0.1], dtype=np.float32)\n            all_poses[img_name] = (R, t)\n            image_width, image_height = 320, 240\n            focal_length = max(image_width, image_height) * 1.2\n            cx, cy = image_width / 2, image_height / 2\n            K = np.array([\n                [focal_length, 0, cx],\n                [0, focal_length, cy],\n                [0, 0, 1]\n            ], dtype=np.float32)\n            all_Ks[img_name] = K\n            pose_status[img_name] = 'sequential'\n            summary['sequential_poses'] += 1\n\n    print(f\"Dataset {dataset}, scene {inferred_scene} processed in {time.time() - start_time:.2f} seconds\")\n    summary['successful_groups'] += 1\n\n# Handle unprocessed images\nunprocessed_images = set(sample_submission['image']) - processed_images\nif unprocessed_images:\n    print(f\"Warning: The following images were not processed: {unprocessed_images}\")\n    for img_name in unprocessed_images:\n        R = np.eye(3, dtype=np.float32)\n        t = np.array([0.0, 0.0, len(all_poses) * 0.1], dtype=np.float32)\n        all_poses[img_name] = (R, t)\n        image_width, image_height = 320, 240\n        focal_length = max(image_width, image_height) * 1.2\n        cx, cy = image_width / 2, image_height / 2\n        K = np.array([\n            [focal_length, 0, cx],\n            [0, focal_length, cy],\n            [0, 0, 1]\n        ], dtype=np.float32)\n        all_Ks[img_name] = K\n        pose_status[img_name] = 'sequential'\n        summary['sequential_poses'] += 1\n\n# Generate submission file\nsubmission_rows = []\nfor _, row in sample_submission.iterrows():\n    img_path = row['image_id']\n    img_name = row['image']\n    if img_name in all_poses:\n        R, t = all_poses[img_name]\n        t = normalize_translation(t, img_name)\n        if not validate_pose(R, t, img_name):\n            print(f\"Invalid pose for {img_name} in submission. Using sequential pose.\")\n            R = np.eye(3, dtype=np.float32)\n            t = np.array([0.0, 0.0, len(submission_rows) * 0.1], dtype=np.float32)\n            pose_status[img_name] = 'sequential'\n            summary['sequential_poses'] += 1\n        q = rotation_matrix_to_quaternion(R)\n        t = t.reshape(3)\n        if not np.all(np.isfinite(q)) or not np.all(np.isfinite(t)):\n            print(f\"Non-finite pose for {img_name}. Using sequential pose.\")\n            q = rotation_matrix_to_quaternion(np.eye(3))\n            t = np.array([0.0, 0.0, len(submission_rows) * 0.1], dtype=np.float64)\n            pose_status[img_name] = 'sequential'\n            summary['sequential_poses'] += 1\n    else:\n        print(f\"Image {img_name} not found in poses. Using sequential pose.\")\n        q = rotation_matrix_to_quaternion(np.eye(3))\n        t = np.array([0.0, 0.0, len(submission_rows) * 0.1], dtype=np.float64)\n        pose_status[img_name] = 'sequential'\n        summary['sequential_poses'] += 1\n    submission_rows.append([img_path] + q.tolist() + t.tolist())\n\nsubmission_columns = [\n    'image_id',\n    'rotation_w', 'rotation_x', 'rotation_y', 'rotation_z',\n    'translation_x', 'translation_y', 'translation_z'\n]\nsubmission_df = pd.DataFrame(submission_rows, columns=submission_columns)\nsubmission_df.to_csv('submission.csv', index=False)\nprint(\"Submission file created: submission.csv\")\n\n# Validate submission\nprint(\"Validating submission...\")\nassert len(submission_df) == len(sample_submission), f\"Submission has {len(submission_df)} rows, expected {len(sample_submission)}\"\nassert set(submission_df.columns) == set(submission_columns), \"Submission columns do not match expected columns\"\n\nfor i, row in submission_df.iterrows():\n    q = np.array([\n        row['rotation_w'], row['rotation_x'], row['rotation_y'], row['rotation_z']\n    ], dtype=np.float64)\n    t = np.array([\n        row['translation_x'], row['translation_y'], row['translation_z']\n    ], dtype=np.float64)\n    assert not np.any(np.isnan(q)), f\"NaN values in quaternion at row {i}\"\n    assert not np.any(np.isnan(t)), f\"NaN values in translation at row {i}\"\n    assert not np.any(np.isinf(q)), f\"Infinite values in quaternion at row {i}\"\n    assert not np.any(np.isinf(t)), f\"Infinite values in translation at row {i}\"\n    norm = np.linalg.norm(q)\n    assert abs(norm - 1.0) < 1e-6, f\"Quaternion not normalized at row {i}: norm={norm}\"\n    R = quaternion_to_rotation_matrix(q)\n    det_R = np.linalg.det(R)\n    assert abs(det_R - 1.0) < 1e-6, f\"Rotation matrix determinant not 1 at row {i}: det={det_R}\"\n    orthogonality = np.linalg.norm(R.T @ R - np.eye(3))\n    assert orthogonality < 1e-6, f\"Rotation matrix not orthogonal at row {i}: orthogonality={orthogonality}\"\n\nprint(\"Submission validated successfully.\")\n\n# Print summary\nprint(\"\\n=== Processing Summary ===\")\nprint(f\"Total dataset/scene groups: {summary['total_groups']}\")\nprint(f\"Successful groups: {summary['successful_groups']}\")\nprint(f\"Failed groups: {summary['failed_groups']}\")\nprint(f\"Total images: {summary['total_images']}\")\nprint(f\"Computed poses: {summary['computed_poses']}\")\nprint(f\"Sequential poses: {summary['sequential_poses']}\")\nprint(f\"Images not found: {summary['images_not_found']}\")\nprint(f\"High reprojection errors: {summary['high_reprojection_errors']}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-06T17:08:55.411712Z","iopub.execute_input":"2025-04-06T17:08:55.412081Z","iopub.status.idle":"2025-04-06T17:10:24.483552Z","shell.execute_reply.started":"2025-04-06T17:08:55.412050Z","shell.execute_reply":"2025-04-06T17:10:24.482622Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Phase 1: Setup and Initial Functions\nimport numpy as np\nimport cv2\nimport pandas as pd\nfrom pathlib import Path\nimport time\nfrom scipy.spatial.transform import Rotation\nfrom scipy.optimize import least_squares\nfrom concurrent.futures import ProcessPoolExecutor\nimport multiprocessing\nimport os\n\n# Utility functions\ndef rotation_matrix_to_quaternion(R):\n    \"\"\"Convert a 3x3 rotation matrix to a quaternion [w, x, y, z].\"\"\"\n    if not np.all(np.isfinite(R)) or R.shape != (3, 3):\n        return np.array([1.0, 0.0, 0.0, 0.0], dtype=np.float64)\n    rot = Rotation.from_matrix(R)\n    q = rot.as_quat()  # Returns [x, y, z, w]\n    return np.array([q[3], q[0], q[1], q[2]], dtype=np.float64)  # Reorder to [w, x, y, z]\n\ndef quaternion_to_rotation_matrix(q):\n    \"\"\"Convert a quaternion [w, x, y, z] to a 3x3 rotation matrix.\"\"\"\n    if not np.all(np.isfinite(q)):\n        return np.eye(3, dtype=np.float64)\n    rot = Rotation.from_quat([q[1], q[2], q[3], q[0]])  # Reorder [x, y, z, w]\n    return rot.as_matrix()\n\ndef project(points3d, R, t, K):\n    \"\"\"Project 3D points to 2D using camera intrinsics and extrinsics.\"\"\"\n    if not (np.all(np.isfinite(points3d)) and np.all(np.isfinite(R)) and np.all(np.isfinite(t)) and np.all(np.isfinite(K))):\n        return np.zeros((points3d.shape[0], 2), dtype=np.float32)\n    points3d = points3d.T\n    P = K @ np.hstack((R, t.reshape(3, 1)))\n    points2d = P @ np.vstack((points3d, np.ones((1, points3d.shape[1]))))\n    points2d = points2d[:2] / (points2d[2] + 1e-8)\n    return points2d.T\n\ndef validate_pose(R, t, img_name=\"Unknown\"):\n    \"\"\"Validate that R and t have the correct shapes and are finite.\"\"\"\n    if R.shape != (3, 3):\n        print(f\"Invalid rotation matrix shape for {img_name}: {R.shape}\")\n        return False\n    if t.shape != (3,):\n        print(f\"Invalid translation vector shape for {img_name}: {t.shape}\")\n        return False\n    if not (np.all(np.isfinite(R)) and np.all(np.isfinite(t))):\n        print(f\"Non-finite values in pose for {img_name}: R={R}, t={t}\")\n        return False\n    return True\n\ndef normalize_translation(t, img_name=\"Unknown\"):\n    \"\"\"Ensure translation vector is a 1D array of shape (3,).\"\"\"\n    if t is None:\n        print(f\"Translation vector is None for {img_name}. Using default.\")\n        return np.array([0.0, 0.0, 0.0], dtype=np.float32)\n    t = np.array(t, dtype=np.float32).flatten()\n    if t.size != 3:\n        print(f\"Cannot normalize translation vector for {img_name}: size={t.size}\")\n        return np.array([0.0, 0.0, 0.0], dtype=np.float32)\n    return t\n\n# Load sample submission\nsample_submission = pd.read_csv('/kaggle/input/image-matching-challenge-2025/sample_submission.csv')\nprint(\"Columns in sample_submission.csv:\", sample_submission.columns.tolist())\n\nimage_path_column = 'image'\nif image_path_column not in sample_submission.columns:\n    raise ValueError(\"Expected 'image' column in sample_submission.csv\")\n\n# Function to infer the correct scene\ndef infer_scene_from_image(dataset, image_name):\n    if dataset == 'ETs':\n        if image_name.startswith('et_et') or image_name.startswith('another_et_another_et') or image_name.startswith('outliers_out_et'):\n            return 'et'\n    elif dataset == 'amy_gardens':\n        if image_name.startswith('peach_'):\n            return 'peach'\n    elif dataset == 'stairs':\n        if image_name.startswith('stairs_split_1_'):\n            return 'stairs_split_1'\n        elif image_name.startswith('stairs_split_2_'):\n            return 'stairs_split_2'\n    elif dataset == 'pt_stpeters_stpauls':\n        if image_name.startswith('st_peters_square_'):\n            return 'st_peters_square'\n        return 'pt_stpeters_stpauls'\n    elif dataset == 'imc2024_dioscuri_baalshamin':\n        return 'dioscuri_baalshamin'\n    elif dataset == 'imc2024_lizard_pond':\n        return 'lizard_pond'\n    elif dataset == 'imc2023_haiper':\n        return 'haiper'\n    elif dataset == 'imc2023_heritage':\n        return 'heritage'\n    elif dataset == 'imc2023_theather_imc2024_church':\n        return 'theather_imc2024_church'\n    elif dataset == 'fbk_vineyard':\n        return 'fbk_vineyard'\n    elif dataset == 'pt_brandenburg_british_buckingham':\n        return 'brandenburg_british_buckingham'\n    elif dataset == 'pt_piazzasanmarco_grandplace':\n        return 'piazzasanmarco_grandplace'\n    elif dataset == 'pt_sacrecoeur_trevi_tajmahal':\n        return 'sacrecoeur_trevi_tajmahal'\n    return dataset\n\n# Add inferred scene column\nsample_submission['inferred_scene'] = sample_submission.apply(\n    lambda row: infer_scene_from_image(row['dataset'], row['image']), axis=1\n)\n\n# Group images by dataset and inferred scene\ngrouped_images = sample_submission.groupby(['dataset', 'inferred_scene'])\nprint(\"Grouped dataset and inferred scene combinations:\", list(grouped_images.groups.keys()))\n\n# Initialize summary tracking\nsummary = {\n    'total_groups': len(grouped_images),\n    'successful_groups': 0,\n    'failed_groups': [],\n    'total_images': len(sample_submission),\n    'computed_poses': 0,\n    'sequential_poses': 0,\n    'sequential_poses_due_to_missing_data': 0,\n    'sequential_poses_due_to_processing': 0,\n    'images_not_found': 0,\n    'high_reprojection_errors': 0\n}\n\n# Track processed images to avoid double-counting\ncounted_images_missing = set()\ncounted_images_processing = set()\n\n# Phase 2: Feature Matching (Using SIFT)\ndef cluster_images(image_paths, scene_name, min_matches_threshold=5):\n    \"\"\"Cluster images using a sliding window approach for efficiency.\"\"\"\n    print(f\"Clustering images for scene: {scene_name}\")\n    start_time = time.time()\n    num_images = len(image_paths)\n    if num_images < 2:\n        print(\"Fewer than 2 images. Using sequential clustering.\")\n        return [list(range(num_images))], []\n\n    window_size = 10\n    edges = []\n\n    sift = cv2.SIFT_create(\n        nfeatures=5000,\n        contrastThreshold=0.01,\n        edgeThreshold=15\n    )\n    FLANN_INDEX_KDTREE = 1\n    index_params = dict(algorithm=FLANN_INDEX_KDTREE, trees=5)\n    search_params = dict(checks=50)\n    flann = cv2.FlannBasedMatcher(index_params, search_params)\n\n    for i in range(num_images):\n        for j in range(i + 1, min(i + window_size + 1, num_images)):\n            img0 = cv2.imread(image_paths[i], cv2.IMREAD_GRAYSCALE)\n            img1 = cv2.imread(image_paths[j], cv2.IMREAD_GRAYSCALE)\n            if img0 is None or img1 is None:\n                print(f\"Failed to load images: {image_paths[i]} or {image_paths[j]}\")\n                continue\n            img0 = cv2.equalizeHist(img0)\n            img1 = cv2.equalizeHist(img1)\n            img0 = cv2.resize(img0, (320, 240))\n            img1 = cv2.resize(img1, (320, 240))\n            kp0, des0 = sift.detectAndCompute(img0, None)\n            kp1, des1 = sift.detectAndCompute(img1, None)\n            if des0 is None or des1 is None:\n                continue\n            matches = flann.knnMatch(des0, des1, k=2)\n            good_matches = []\n            for m, n in matches:\n                if m.distance < 0.6 * n.distance:  # Reverted to 0.6\n                    good_matches.append(m)\n            num_matches = len(good_matches)\n            if num_matches >= min_matches_threshold:\n                edges.append((num_matches, i, j))\n\n    if not edges:\n        print(\"No matches found between any image pairs. Using sequential clustering.\")\n        return [list(range(num_images))], []\n\n    from collections import defaultdict\n    graph = defaultdict(list)\n    for _, i, j in edges:\n        graph[i].append(j)\n        graph[j].append(i)\n\n    def find_component(start, graph, visited):\n        component = []\n        stack = [start]\n        while stack:\n            node = stack.pop()\n            if node not in visited:\n                visited.add(node)\n                component.append(node)\n                stack.extend(n for n in graph[node] if n not in visited)\n        return component\n\n    visited = set()\n    clusters = []\n    for i in range(num_images):\n        if i not in visited:\n            component = find_component(i, graph, visited)\n            if len(component) >= 2:\n                clusters.append(sorted(component))\n\n    noise_indices = [i for i in range(num_images) if i not in visited]\n\n    if not clusters:\n        print(\"No clusters found. Using sequential clustering.\")\n        clusters = [list(range(num_images))]\n\n    print(f\"Clustering took {time.time() - start_time:.2f} seconds\")\n    return clusters, noise_indices\n\n# Phase 3: Pose Estimation\ndef recover_pose_with_fallback(mkpts0, mkpts1, K, img0_name, img1_name):\n    \"\"\"Recover pose between two images with robust fallbacks.\"\"\"\n    if len(mkpts0) < 5:\n        print(f\"Too few matches ({len(mkpts0)}) between {img0_name} and {img1_name}. Cannot estimate pose.\")\n        return None, None, 0\n\n    max_iters = 5000\n    E, mask = cv2.findEssentialMat(\n        mkpts0, mkpts1, K, method=cv2.RANSAC, prob=0.999, threshold=0.5, maxIters=max_iters\n    )\n    if E is None or E.shape != (3, 3):\n        print(f\"Essential matrix estimation failed for pair {img0_name} and {img1_name}. Trying homography.\")\n    else:\n        inliers, R, t, mask = cv2.recoverPose(E, mkpts0, mkpts1, K, mask=mask)\n        if inliers >= 5:\n            P0 = K @ np.hstack((np.eye(3), np.zeros((3, 1))))\n            P1 = K @ np.hstack((R, t.reshape(3, 1)))\n            points4d = cv2.triangulatePoints(P0, P1, mkpts0.T, mkpts1.T)\n            points3d = points4d[:3] / (points4d[3] + 1e-8)\n            points3d_cam2 = R @ points3d + t.reshape(3, 1)\n            in_front = (points3d[2, :] > 0) & (points3d_cam2[2, :] > 0)\n            if np.sum(in_front) >= 0.5 * inliers:\n                t = t.flatten()\n                if validate_pose(R, t, f\"{img0_name}-{img1_name}\"):\n                    return R, t, inliers\n\n    print(f\"Falling back to homography for pair {img0_name} and {img1_name}\")\n    H, mask = cv2.findHomography(mkpts0, mkpts1, cv2.RANSAC, 0.5, maxIters=5000)\n    if H is None or H.shape != (3, 3):\n        print(f\"Homography estimation failed for pair {img0_name} and {img1_name}\")\n        return None, None, 0\n\n    inliers = np.sum(mask)\n    if inliers < 5:\n        print(f\"Too few inliers from homography for pair {img0_name} and {img1_name}: {inliers}\")\n        return None, None, 0\n\n    num, Rs, ts, normals = cv2.decomposeHomographyMat(H, K)\n    best_R, best_t, best_inliers = None, None, 0\n    for i in range(num):\n        R, t = Rs[i], ts[i]\n        P0 = K @ np.hstack((np.eye(3), np.zeros((3, 1))))\n        P1 = K @ np.hstack((R, t))\n        points4d = cv2.triangulatePoints(P0, P1, mkpts0.T, mkpts1.T)\n        points3d = points4d[:3] / (points4d[3] + 1e-8)\n        points3d_cam2 = R @ points3d + t.reshape(3, 1)\n        in_front = (points3d[2, :] > 0) & (points3d_cam2[2, :] > 0)\n        inliers_count = np.sum(in_front)\n        if inliers_count > best_inliers:\n            best_inliers = inliers_count\n            best_R, best_t = R, t\n\n    if best_R is None or best_t is None:\n        print(f\"No valid pose from homography for pair {img0_name} and {img1_name}\")\n        return None, None, 0\n\n    best_t = best_t.flatten()\n    if validate_pose(best_R, best_t, f\"{img0_name}-{img1_name}\"):\n        return best_R, best_t, best_inliers\n    return None, None, 0\n\ndef reprojection_error(params, points3d, observations, K, img_names, poses):\n    \"\"\"Compute reprojection error for bundle adjustment.\"\"\"\n    num_cameras = len(img_names)\n    translations = params[:num_cameras * 3].reshape(num_cameras, 3)\n    focal_length = params[num_cameras * 3]\n    cx = params[num_cameras * 3 + 1]\n    cy = params[num_cameras * 3 + 2]\n\n    K_opt = np.array([\n        [focal_length, 0, cx],\n        [0, focal_length, cy],\n        [0, 0, 1]\n    ], dtype=np.float32)\n\n    errors = []\n    for i, img_name in enumerate(img_names):\n        R, _ = poses[img_name]\n        t = translations[i]\n        for j, obs in enumerate(observations):\n            for img_obs, pt in obs:\n                if img_obs == img_name:\n                    proj = project(points3d[j:j+1], R, t, K_opt)\n                    if not np.all(np.isfinite(proj)):\n                        continue\n                    errors.append(proj[0] - pt)\n\n    if len(errors) == 0:\n        print(\"No reprojection errors computed. Returning zeros.\")\n        return np.zeros(num_cameras * 3 + 3)\n    errors = np.concatenate(errors)\n    if not np.all(np.isfinite(errors)):\n        return np.zeros_like(errors)\n    return errors\n\n# Phase 4: Feature Matching\ndef match_image_pair(args):\n    \"\"\"Match keypoints between a pair of images.\"\"\"\n    idx0, idx1, global_idx0, global_idx1, image_paths = args\n    img0 = cv2.imread(image_paths[global_idx0], cv2.IMREAD_GRAYSCALE)\n    img1 = cv2.imread(image_paths[global_idx1], cv2.IMREAD_GRAYSCALE)\n\n    if img0 is None or img1 is None:\n        print(f\"Failed to load images: {image_paths[global_idx0]} or {image_paths[global_idx1]}\")\n        return idx0, idx1, None, None, None, None\n\n    img0 = cv2.equalizeHist(img0)\n    img1 = cv2.equalizeHist(img1)\n\n    orig_size0 = img0.shape[::-1]\n    orig_size1 = img1.shape[::-1]\n\n    img0_resized = cv2.resize(img0, (320, 240))\n    img1_resized = cv2.resize(img1, (320, 240))\n\n    sift = cv2.SIFT_create(\n        nfeatures=5000,\n        contrastThreshold=0.01,\n        edgeThreshold=15\n    )\n    FLANN_INDEX_KDTREE = 1\n    index_params = dict(algorithm=FLANN_INDEX_KDTREE, trees=5)\n    search_params = dict(checks=50)\n    flann = cv2.FlannBasedMatcher(index_params, search_params)\n\n    kp0, des0 = sift.detectAndCompute(img0_resized, None)\n    kp1, des1 = sift.detectAndCompute(img1_resized, None)\n    if des0 is None or des1 is None:\n        return idx0, idx1, None, None, None, None\n\n    matches = flann.knnMatch(des0, des1, k=2)\n    good_matches = []\n    for m, n in matches:\n        if m.distance < 0.6 * n.distance:  # Reverted to 0.6\n            good_matches.append(m)\n\n    if not good_matches:\n        print(f\"No good matches between {Path(image_paths[global_idx0]).name} and {Path(image_paths[global_idx1]).name}\")\n        return idx0, idx1, None, None, None, None\n\n    mkpts0 = np.array([kp0[m.queryIdx].pt for m in good_matches])\n    mkpts1 = np.array([kp1[m.trainIdx].pt for m in good_matches])\n    kpts0 = np.array([kp.pt for kp in kp0])\n    kpts1 = np.array([kp.pt for kp in kp1])\n\n    scale0 = (orig_size0[0] / 320, orig_size0[1] / 240)\n    scale1 = (orig_size1[0] / 320, orig_size1[1] / 240)\n    kpts0[:, 0] *= scale0[0]\n    kpts0[:, 1] *= scale0[1]\n    kpts1[:, 0] *= scale1[0]\n    kpts1[:, 1] *= scale1[1]\n    mkpts0[:, 0] *= scale0[0]\n    mkpts0[:, 1] *= scale0[1]\n    mkpts1[:, 0] *= scale1[0]\n    mkpts1[:, 1] *= scale1[1]\n\n    img0_name = Path(image_paths[global_idx0]).name\n    img1_name = Path(image_paths[global_idx1]).name\n\n    return idx0, idx1, img0_name, img1_name, (kpts0, kpts1), (mkpts0, mkpts1)\n\ndef run_sfm_on_cluster(args):\n    \"\"\"Run Structure-from-Motion (SfM) on a cluster of images using incremental SfM.\"\"\"\n    cluster_indices, image_paths, scene_name, cluster_idx, clusters, overlap = args\n    print(f\"Starting cluster {cluster_idx} for scene {scene_name}\")\n\n    if not cluster_indices or not image_paths:\n        print(\"Empty cluster or image paths. Using sequential poses.\")\n        poses = {}\n        for idx in range(len(cluster_indices)):\n            if idx >= len(cluster_indices) or cluster_indices[idx] >= len(image_paths):\n                continue\n            img_name = Path(image_paths[cluster_indices[idx]]).name\n            if img_name in counted_images_processing:\n                continue\n            R = np.eye(3, dtype=np.float32)\n            t = np.array([0.0, 0.0, idx * 0.1], dtype=np.float32)\n            poses[img_name] = (R, t)\n            counted_images_processing.add(img_name)\n            summary['sequential_poses_due_to_processing'] += 1\n        image_width, image_height = 320, 240\n        focal_length = max(image_width, image_height) * 1.2\n        cx, cy = image_width / 2, image_height / 2\n        K = np.array([\n            [focal_length, 0, cx],\n            [0, focal_length, cy],\n            [0, 0, 1]\n        ], dtype=np.float32)\n        return poses, np.array([], dtype=np.float32).reshape(0, 3), K, []\n\n    keypoints_dict = {}\n    matches_dict = {}\n    image_sizes = {}\n    point_to_images = []\n\n    start_time = time.time()\n    window_size = 7\n    image_pairs = []\n    for i in range(len(cluster_indices)):\n        for j in range(i + 1, min(i + window_size + 1, len(cluster_indices))):\n            image_pairs.append((i, j))\n\n    edges = []\n    max_workers = min(multiprocessing.cpu_count(), 4)\n    with ProcessPoolExecutor(max_workers=max_workers) as executor:\n        futures = [\n            executor.submit(match_image_pair, (idx0, idx1, cluster_indices[idx0], cluster_indices[idx1], image_paths))\n            for idx0, idx1 in image_pairs\n        ]\n        for future in futures:\n            idx0, idx1, img0_name, img1_name, keypoints, matches = future.result()\n            if matches is None:\n                continue\n            kpts0, kpts1 = keypoints\n            mkpts0, mkpts1 = matches\n            if img0_name not in keypoints_dict:\n                keypoints_dict[img0_name] = kpts0\n                image_sizes[img0_name] = (kpts0.shape[0] * 320 / mkpts0.shape[0], kpts0.shape[1] * 240 / mkpts0.shape[1]) if mkpts0.shape[0] > 0 else (320, 240)\n            if img1_name not in keypoints_dict:\n                keypoints_dict[img1_name] = kpts1\n                image_sizes[img1_name] = (kpts1.shape[0] * 320 / mkpts1.shape[0], kpts1.shape[1] * 240 / mkpts1.shape[1]) if mkpts1.shape[0] > 0 else (320, 240)\n            matches_dict[(img0_name, img1_name)] = (mkpts0, mkpts1)\n            priority = 0\n            if cluster_idx < len(clusters) - 1:\n                overlap_indices = clusters[cluster_idx][-overlap:]\n                if cluster_indices[idx0] in overlap_indices or cluster_indices[idx1] in overlap_indices:\n                    priority = 1\n            edges.append((len(mkpts0), idx0, idx1, priority))\n            print(f\"Matches between {img0_name} and {img1_name}: {len(mkpts0)}\")\n\n    print(f\"Feature matching for cluster took {time.time() - start_time:.2f} seconds.\")\n\n    edges.sort(reverse=True)\n    poses = {}\n    added_images = set()\n\n    for num_matches, idx0, idx1, priority in edges:\n        if num_matches < 5:\n            continue\n        if idx0 >= len(cluster_indices) or idx1 >= len(cluster_indices):\n            continue\n        global_idx0 = cluster_indices[idx0]\n        global_idx1 = cluster_indices[idx1]\n        if global_idx0 >= len(image_paths) or global_idx1 >= len(image_paths):\n            continue\n        img0_name = Path(image_paths[global_idx0]).name\n        img1_name = Path(image_paths[global_idx1]).name\n        mkpts0, mkpts1 = matches_dict[(img0_name, img1_name)]\n        image_width, image_height = image_sizes[img0_name]\n        focal_length = max(image_width, image_height) * 1.2\n        cx, cy = image_width / 2, image_height / 2\n        K = np.array([\n            [focal_length, 0, cx],\n            [0, focal_length, cy],\n            [0, 0, 1]\n        ], dtype=np.float32)\n        R, t, inliers = recover_pose_with_fallback(mkpts0, mkpts1, K, img0_name, img1_name)\n        if R is None or t is None:\n            continue\n        t = normalize_translation(t, img0_name)\n        poses[img0_name] = (np.eye(3, dtype=np.float32), np.zeros(3, dtype=np.float32))\n        poses[img1_name] = (R, t)\n        added_images.update([img0_name, img1_name])\n        break\n\n    if not poses:\n        print(f\"No initial pair found for cluster {cluster_idx}. Using sequential poses.\")\n        for idx in range(len(cluster_indices)):\n            if idx >= len(cluster_indices) or cluster_indices[idx] >= len(image_paths):\n                continue\n            img_name = Path(image_paths[cluster_indices[idx]]).name\n            if img_name in counted_images_processing:\n                continue\n            R = np.eye(3, dtype=np.float32)\n            t = np.array([0.0, 0.0, idx * 0.1], dtype=np.float32)\n            poses[img_name] = (R, t)\n            added_images.add(img_name)\n            counted_images_processing.add(img_name)\n            summary['sequential_poses_due_to_processing'] += 1\n        image_width, image_height = 320, 240\n        focal_length = max(image_width, image_height) * 1.2\n        cx, cy = image_width / 2, image_height / 2\n        K = np.array([\n            [focal_length, 0, cx],\n            [0, focal_length, cy],\n            [0, 0, 1]\n        ], dtype=np.float32)\n        return poses, np.array([], dtype=np.float32).reshape(0, 3), K, []\n\n    points3d = []\n    point_to_images = []\n    for idx in range(len(cluster_indices)):\n        if idx >= len(cluster_indices) or cluster_indices[idx] >= len(image_paths):\n            continue\n        global_idx = cluster_indices[idx]\n        img_name = Path(image_paths[global_idx]).name\n        if img_name in added_images:\n            continue\n\n        best_pair = None\n        best_num_matches = 0\n        best_idx0 = None\n        for idx0 in range(len(cluster_indices)):\n            if idx0 >= len(cluster_indices) or cluster_indices[idx0] >= len(image_paths):\n                continue\n            img0_name = Path(image_paths[cluster_indices[idx0]]).name\n            if img0_name not in added_images:\n                continue\n            pair_key = (img0_name, img_name) if (img0_name, img_name) in matches_dict else (img_name, img0_name)\n            if pair_key not in matches_dict:\n                _, _, _, _, _, matches = match_image_pair((idx0, idx, cluster_indices[idx0], global_idx, image_paths))\n                if matches is None:\n                    matches_dict[(img0_name, img_name)] = (np.array([]), np.array([]))\n                    num_matches = 0\n                else:\n                    mkpts0, mkpts1 = matches\n                    matches_dict[(img0_name, img_name)] = (mkpts0, mkpts1)\n                    num_matches = len(mkpts0)\n            else:\n                mkpts0, mkpts1 = matches_dict[pair_key]\n                num_matches = len(mkpts0)\n\n            if num_matches > best_num_matches:\n                best_num_matches = num_matches\n                best_pair = (img0_name, img_name)\n                best_idx0 = idx0\n\n        if best_pair is None or best_num_matches < 5:\n            print(f\"Could not add {img_name} to cluster {cluster_idx}. Using sequential pose.\")\n            if img_name in counted_images_processing:\n                continue\n            R = np.eye(3, dtype=np.float32)\n            t = np.array([0.0, 0.0, len(added_images) * 0.1], dtype=np.float32)\n            poses[img_name] = (R, t)\n            added_images.add(img_name)\n            counted_images_processing.add(img_name)\n            summary['sequential_poses_due_to_processing'] += 1\n            continue\n\n        img0_name, img1_name = best_pair\n        pair_key = (img0_name, img1_name) if (img0_name, img1_name) in matches_dict else (img1_name, img0_name)\n        mkpts0, mkpts1 = matches_dict[pair_key]\n        if len(mkpts0) < 5:\n            print(f\"Too few matches ({len(mkpts0)}) for pair {img0_name} and {img1_name}. Using sequential pose for {img1_name}.\")\n            if img1_name in counted_images_processing:\n                continue\n            R = np.eye(3, dtype=np.float32)\n            t = np.array([0.0, 0.0, len(added_images) * 0.1], dtype=np.float32)\n            poses[img1_name] = (R, t)\n            added_images.add(img1_name)\n            counted_images_processing.add(img1_name)\n            summary['sequential_poses_due_to_processing'] += 1\n            continue\n\n        R0, t0 = poses[img0_name]\n        R, t, inliers = recover_pose_with_fallback(mkpts0, mkpts1, K, img0_name, img1_name)\n        if R is None or t is None:\n            print(f\"Failed to estimate pose for {img1_name}. Using sequential pose.\")\n            if img1_name in counted_images_processing:\n                continue\n            R = np.eye(3, dtype=np.float32)\n            t = np.array([0.0, 0.0, len(added_images) * 0.1], dtype=np.float32)\n            poses[img1_name] = (R, t)\n            added_images.add(img1_name)\n            counted_images_processing.add(img1_name)\n            summary['sequential_poses_due_to_processing'] += 1\n            continue\n\n        t = normalize_translation(t, img1_name)\n        t0 = normalize_translation(t0, img0_name)\n        R = R0 @ R\n        t = R0 @ t + t0\n        t = normalize_translation(t, img1_name)\n        if not validate_pose(R, t, img1_name):\n            print(f\"Invalid pose for {img1_name}. Using sequential pose.\")\n            if img1_name in counted_images_processing:\n                continue\n            R = np.eye(3, dtype=np.float32)\n            t = np.array([0.0, 0.0, len(added_images) * 0.1], dtype=np.float32)\n            counted_images_processing.add(img1_name)\n            summary['sequential_poses_due_to_processing'] += 1\n        poses[img1_name] = (R, t)\n        added_images.add(img1_name)\n\n        if len(mkpts0) >= 5:\n            R1, t1 = poses[img1_name]\n            t0 = normalize_translation(t0, img0_name)\n            t1 = normalize_translation(t1, img1_name)\n            if validate_pose(R0, t0, img0_name) and validate_pose(R1, t1, img1_name):\n                P0 = K @ np.hstack((R0, t0.reshape(3, 1)))\n                P1 = K @ np.hstack((R1, t1.reshape(3, 1)))\n                points4d = cv2.triangulatePoints(P0, P1, mkpts0.T, mkpts1.T)\n                points4d = points4d[:3] / (points4d[3] + 1e-8)\n                points4d = points4d.T\n                if np.all(np.isfinite(points4d)):\n                    for i in range(len(points4d)):\n                        points3d.append(points4d[i])\n                        point_to_images.append([(img0_name, mkpts0[i]), (img1_name, mkpts1[i])])\n\n    points3d = np.array(points3d, dtype=np.float32) if points3d else np.array([], dtype=np.float32).reshape(0, 3)\n    print(f\"Reconstructed {len(points3d)} 3D points in cluster {cluster_idx}\")\n\n    if len(poses) >= 3 and len(point_to_images) >= 10 and len(poses) < 20:\n        start_time = time.time()\n        img_names = list(poses.keys())\n        translations = np.array([normalize_translation(poses[img_name][1], img_name) for img_name in img_names], dtype=np.float32)\n        if translations.shape != (len(img_names), 3):\n            print(f\"Mismatch in translations shape: {translations.shape}. Expected ({len(img_names)}, 3). Skipping bundle adjustment.\")\n        else:\n            initial_focal_length = K[0, 0]\n            initial_cx, initial_cy = K[0, 2], K[1, 2]\n            params = np.hstack((translations.ravel(), initial_focal_length, initial_cx, initial_cy))\n            max_nfev = min(20, len(poses) * 5)\n            result = least_squares(\n                reprojection_error,\n                params,\n                args=(points3d, point_to_images, K, img_names, poses),\n                max_nfev=max_nfev,\n                ftol=1e-5,\n                xtol=1e-5\n            )\n            optimized_params = result.x\n            num_cameras = len(img_names)\n            translations = optimized_params[:num_cameras * 3].reshape(num_cameras, 3)\n            focal_length = optimized_params[num_cameras * 3]\n            cx = optimized_params[num_cameras * 3 + 1]\n            cy = optimized_params[num_cameras * 3 + 2]\n            K = np.array([\n                [focal_length, 0, cx],\n                [0, focal_length, cy],\n                [0, 0, 1]\n            ], dtype=np.float32)\n            for i, img_name in enumerate(img_names):\n                R, _ = poses[img_name]\n                t = translations[i]\n                t = normalize_translation(t, img_name)\n                if not validate_pose(R, t, img_name):\n                    if img_name in counted_images_processing:\n                        continue\n                    R = np.eye(3, dtype=np.float32)\n                    t = np.array([0.0, 0.0, len(poses) * 0.1], dtype=np.float32)\n                    counted_images_processing.add(img_name)\n                    summary['sequential_poses_due_to_processing'] += 1\n                poses[img_name] = (R, t)\n            print(f\"Local bundle adjustment took {time.time() - start_time:.2f} seconds.\")\n    else:\n        print(\"Skipping local bundle adjustment due to insufficient poses/observations or large cluster size.\")\n\n    print(f\"Cluster {cluster_idx} processed {len(added_images)} cameras\")\n    return poses, points3d, K, point_to_images\n\n# Phase 5: Reprojection Error\ndef compute_reprojection_error(img_name, points3d, observations, R, t, K):\n    \"\"\"Compute the average reprojection error for an image.\"\"\"\n    errors = []\n    for j, obs in enumerate(observations):\n        for img_obs, pt in obs:\n            if img_obs == img_name:\n                proj = project(points3d[j:j+1], R, t, K)\n                if not np.all(np.isfinite(proj)):\n                    continue\n                error = np.linalg.norm(proj[0] - pt)\n                if np.isfinite(error):\n                    errors.append(error)\n    return np.mean(errors) if errors else float('inf')\n\n# Phase 6: Main Loop and Submission\nall_poses = {}\nall_Ks = {}\nprocessed_images = set()\npose_status = {}\n\nfor (dataset, inferred_scene), group in grouped_images:\n    print(f\"\\nProcessing dataset: {dataset}, inferred scene: {inferred_scene}\")\n    start_time = time.time()\n\n    image_dir = f\"/kaggle/input/image-matching-challenge-2025/test/{dataset}\"\n    if not os.path.exists(image_dir):\n        print(f\"Dataset directory {image_dir} does not exist in test set. Skipping.\")\n        summary['failed_groups'].append((dataset, inferred_scene, \"Dataset not found in test directory\"))\n        for _, row in group.iterrows():\n            img_name = row['image']\n            if img_name not in counted_images_missing:\n                summary['images_not_found'] += 1\n                summary['sequential_poses_due_to_missing_data'] += 1\n                counted_images_missing.add(img_name)\n        continue\n\n    image_paths = []\n    for _, row in group.iterrows():\n        img_name = row['image']\n        img_path = f\"{image_dir}/{img_name}\"\n        if Path(img_path).exists():\n            image_paths.append(img_path)\n            processed_images.add(img_name)\n        else:\n            print(f\"Image not found: {img_path}\")\n            if img_name not in counted_images_missing:\n                summary['images_not_found'] += 1\n                summary['sequential_poses_due_to_missing_data'] += 1\n                counted_images_missing.add(img_name)\n\n    if not image_paths:\n        print(f\"No images found for dataset {dataset}, scene {inferred_scene}. Skipping.\")\n        summary['failed_groups'].append((dataset, inferred_scene, \"No images found\"))\n        continue\n\n    if len(image_paths) < 2:\n        print(f\"Dataset {dataset}, scene {inferred_scene} has fewer than 2 images. Using sequential poses.\")\n        for idx, img_path in enumerate(image_paths):\n            img_name = Path(img_path).name\n            if img_name in counted_images_processing:\n                continue\n            R = np.eye(3, dtype=np.float32)\n            t = np.array([0.0, 0.0, idx * 0.1], dtype=np.float32)\n            all_poses[img_name] = (R, t)\n            image_width, image_height = 320, 240\n            focal_length = max(image_width, image_height) * 1.2\n            cx, cy = image_width / 2, image_height / 2\n            K = np.array([\n                [focal_length, 0, cx],\n                [0, focal_length, cy],\n                [0, 0, 1]\n            ], dtype=np.float32)\n            all_Ks[img_name] = K\n            pose_status[img_name] = 'sequential'\n            counted_images_processing.add(img_name)\n            summary['sequential_poses_due_to_processing'] += 1\n        summary['failed_groups'].append((dataset, inferred_scene, \"Fewer than 2 images\"))\n        continue\n\n    clusters, noise_indices = cluster_images(image_paths, inferred_scene, min_matches_threshold=5)\n    print(f\"Clusters: {clusters}\")\n    print(f\"Outlier images: {[Path(image_paths[idx]).name for idx in noise_indices if idx < len(image_paths)]}\")\n\n    overlap = 2\n\n    all_points3d = []\n    all_observations = []\n\n    print(f\"Processing clusters for dataset {dataset}, scene {inferred_scene}\")\n    for cluster_idx, cluster_indices in enumerate(clusters):\n        result = run_sfm_on_cluster((cluster_indices, image_paths, inferred_scene, cluster_idx, clusters, overlap))\n        poses, points3d, K, cluster_point_to_images = result\n        for img_name in poses:\n            R, t = poses[img_name]\n            t = normalize_translation(t, img_name)\n            if not validate_pose(R, t, img_name):\n                print(f\"Invalid pose for {img_name} after SfM. Using sequential pose.\")\n                if img_name in counted_images_processing:\n                    continue\n                R = np.eye(3, dtype=np.float32)\n                t = np.array([0.0, 0.0, len(all_poses) * 0.1], dtype=np.float32)\n                counted_images_processing.add(img_name)\n                summary['sequential_poses_due_to_processing'] += 1\n            all_poses[img_name] = (R, t)\n            all_Ks[img_name] = K\n            if img_name not in pose_status:\n                pose_status[img_name] = 'computed'\n                summary['computed_poses'] += 1\n        if len(points3d) > 0:\n            all_points3d.append(points3d)\n        all_observations.extend(cluster_point_to_images)\n\n    for idx in noise_indices:\n        if idx >= len(image_paths):\n            continue\n        img_name = Path(image_paths[idx]).name\n        best_pair = None\n        best_num_matches = 0\n        for i in range(max(0, idx - 3), min(len(image_paths), idx + 4)):  # Reduced range to ±3\n            if i == idx or Path(image_paths[i]).name not in all_poses:\n                continue\n            img0_name = Path(image_paths[i]).name\n            pair_key = (img0_name, img_name) if (img0_name, img_name) in matches_dict else (img_name, img0_name)\n            if pair_key not in matches_dict:\n                _, _, _, _, _, matches = match_image_pair((i, idx, i, idx, image_paths))\n                if matches is None:\n                    matches_dict[(img0_name, img_name)] = (np.array([]), np.array([]))\n                    num_matches = 0\n                else:\n                    mkpts0, mkpts1 = matches\n                    matches_dict[(img0_name, img_name)] = (mkpts0, mkpts1)\n                    num_matches = len(mkpts0)\n            else:\n                mkpts0, mkpts1 = matches_dict[pair_key]\n                num_matches = len(mkpts0)\n            if num_matches > best_num_matches:\n                best_num_matches = num_matches\n                best_pair = (img0_name, img_name)\n\n        if best_pair is None or best_num_matches < 5:\n            print(f\"Could not register outlier image {img_name}. Using sequential pose.\")\n            if img_name in counted_images_processing:\n                continue\n            R = np.eye(3, dtype=np.float32)\n            t = np.array([0.0, 0.0, len(all_poses) * 0.1], dtype=np.float32)\n            all_poses[img_name] = (R, t)\n            image_width, image_height = 320, 240\n            focal_length = max(image_width, image_height) * 1.2\n            cx, cy = image_width / 2, image_height / 2\n            K = np.array([\n                [focal_length, 0, cx],\n                [0, focal_length, cy],\n                [0, 0, 1]\n            ], dtype=np.float32)\n            all_Ks[img_name] = K\n            pose_status[img_name] = 'sequential'\n            counted_images_processing.add(img_name)\n            summary['sequential_poses_due_to_processing'] += 1\n            continue\n\n        img0_name, img1_name = best_pair\n        pair_key = (img0_name, img1_name) if (img0_name, img1_name) in matches_dict else (img1_name, img0_name)\n        mkpts0, mkpts1 = matches_dict[pair_key]\n        image_width, image_height = 320, 240\n        focal_length = max(image_width, image_height) * 1.2\n        cx, cy = image_width / 2, image_height / 2\n        K = np.array([\n            [focal_length, 0, cx],\n            [0, focal_length, cy],\n            [0, 0, 1]\n        ], dtype=np.float32)\n        R, t, inliers = recover_pose_with_fallback(mkpts0, mkpts1, K, img0_name, img1_name)\n        if R is None or t is None:\n            print(f\"Failed to register outlier image {img_name}. Using sequential pose.\")\n            if img_name in counted_images_processing:\n                continue\n            R = np.eye(3, dtype=np.float32)\n            t = np.array([0.0, 0.0, len(all_poses) * 0.1], dtype=np.float32)\n            all_poses[img_name] = (R, t)\n            all_Ks[img_name] = K\n            pose_status[img_name] = 'sequential'\n            counted_images_processing.add(img_name)\n            summary['sequential_poses_due_to_processing'] += 1\n            continue\n        R0, t0 = all_poses[img0_name]\n        t0 = normalize_translation(t0, img0_name)\n        t = normalize_translation(t, img1_name)\n        R = R0 @ R\n        t = R0 @ t + t0\n        t = normalize_translation(t, img1_name)\n        if not validate_pose(R, t, img1_name):\n            print(f\"Invalid pose for outlier {img1_name}. Using sequential pose.\")\n            if img1_name in counted_images_processing:\n                continue\n            R = np.eye(3, dtype=np.float32)\n            t = np.array([0.0, 0.0, len(all_poses) * 0.1], dtype=np.float32)\n            pose_status[img_name] = 'sequential'\n            counted_images_processing.add(img1_name)\n            summary['sequential_poses_due_to_processing'] += 1\n        all_poses[img_name] = (R, t)\n        all_Ks[img_name] = K\n        if img_name not in pose_status:\n            pose_status[img_name] = 'computed'\n            summary['computed_poses'] += 1\n\n    if all_points3d:\n        merged_points3d = np.concatenate(all_points3d, axis=0)\n    else:\n        merged_points3d = np.array([], dtype=np.float32).reshape(0, 3)\n\n    print(\"Skipping global bundle adjustment to save time.\")\n\n    reproj_threshold = 20.0  # Increased to 20.0\n    for img_name in list(all_poses.keys()):\n        R, t = all_poses[img_name]\n        t = normalize_translation(t, img_name)\n        if not validate_pose(R, t, img_name):\n            print(f\"Invalid pose for {img_name} before reprojection check. Using sequential pose.\")\n            if img_name in counted_images_processing:\n                continue\n            R = np.eye(3, dtype=np.float32)\n            t = np.array([0.0, 0.0, len(all_poses) * 0.1], dtype=np.float32)\n            all_poses[img_name] = (R, t)\n            if pose_status.get(img_name) == 'computed':\n                summary['computed_poses'] -= 1\n                counted_images_processing.add(img_name)\n                summary['sequential_poses_due_to_processing'] += 1\n            pose_status[img_name] = 'sequential'\n        if img_name not in all_Ks:\n            print(f\"Camera intrinsics (K) missing for {img_name}. Assigning default K.\")\n            image_width, image_height = 320, 240\n            focal_length = max(image_width, image_height) * 1.2\n            cx, cy = image_width / 2, image_height / 2\n            K = np.array([\n                [focal_length, 0, cx],\n                [0, focal_length, cy],\n                [0, 0, 1]\n            ], dtype=np.float32)\n            all_Ks[img_name] = K\n        K = all_Ks[img_name]\n        error = compute_reprojection_error(img_name, merged_points3d, all_observations, R, t, K)\n        if error > reproj_threshold and pose_status.get(img_name) == 'computed':\n            print(f\"Image {img_name} has high reprojection error ({error:.2f}). Resetting to sequential pose.\")\n            R = np.eye(3, dtype=np.float32)\n            t = np.array([0.0, 0.0, len(all_poses) * 0.1], dtype=np.float32)\n            all_poses[img_name] = (R, t)\n            pose_status[img_name] = 'sequential'\n            summary['computed_poses'] -= 1\n            if img_name not in counted_images_processing:\n                counted_images_processing.add(img_name)\n                summary['sequential_poses_due_to_processing'] += 1\n            summary['high_reprojection_errors'] += 1\n            image_width, image_height = 320, 240\n            focal_length = max(image_width, image_height) * 1.2\n            cx, cy = image_width / 2, image_height / 2\n            K = np.array([\n                [focal_length, 0, cx],\n                [0, focal_length, cy],\n                [0, 0, 1]\n            ], dtype=np.float32)\n            all_Ks[img_name] = K\n\n    for img_path in image_paths:\n        img_name = Path(img_path).name\n        if img_name not in all_poses:\n            print(f\"Image {img_name} was not processed. Using sequential pose.\")\n            if img_name in counted_images_processing:\n                continue\n            R = np.eye(3, dtype=np.float32)\n            t = np.array([0.0, 0.0, len(all_poses) * 0.1], dtype=np.float32)\n            all_poses[img_name] = (R, t)\n            image_width, image_height = 320, 240\n            focal_length = max(image_width, image_height) * 1.2\n            cx, cy = image_width / 2, image_height / 2\n            K = np.array([\n                [focal_length, 0, cx],\n                [0, focal_length, cy],\n                [0, 0, 1]\n            ], dtype=np.float32)\n            all_Ks[img_name] = K\n            pose_status[img_name] = 'sequential'\n            counted_images_processing.add(img_name)\n            summary['sequential_poses_due_to_processing'] += 1\n\n    print(f\"Dataset {dataset}, scene {inferred_scene} processed in {time.time() - start_time:.2f} seconds\")\n    summary['successful_groups'] += 1\n\n# Handle unprocessed images\nunprocessed_images = set(sample_submission['image']) - processed_images\nif unprocessed_images:\n    print(f\"Warning: The following images were not processed: {unprocessed_images}\")\n    for img_name in unprocessed_images:\n        if img_name in counted_images_missing:\n            continue\n        R = np.eye(3, dtype=np.float32)\n        t = np.array([0.0, 0.0, len(all_poses) * 0.1], dtype=np.float32)\n        all_poses[img_name] = (R, t)\n        image_width, image_height = 320, 240\n        focal_length = max(image_width, image_height) * 1.2\n        cx, cy = image_width / 2, image_height / 2\n        K = np.array([\n            [focal_length, 0, cx],\n            [0, focal_length, cy],\n            [0, 0, 1]\n        ], dtype=np.float32)\n        all_Ks[img_name] = K\n        pose_status[img_name] = 'sequential'\n        counted_images_missing.add(img_name)\n        summary['sequential_poses_due_to_missing_data'] += 1\n\n# Generate submission file\nsubmission_rows = []\nfor _, row in sample_submission.iterrows():\n    img_path = row['image_id']\n    img_name = row['image']\n    if img_name in all_poses:\n        R, t = all_poses[img_name]\n        t = normalize_translation(t, img_name)\n        if not validate_pose(R, t, img_name):\n            print(f\"Invalid pose for {img_name} in submission. Using sequential pose.\")\n            if img_name in counted_images_processing:\n                continue\n            R = np.eye(3, dtype=np.float32)\n            t = np.array([0.0, 0.0, len(submission_rows) * 0.1], dtype=np.float32)\n            if pose_status.get(img_name) == 'computed':\n                summary['computed_poses'] -= 1\n                counted_images_processing.add(img_name)\n                summary['sequential_poses_due_to_processing'] += 1\n            pose_status[img_name] = 'sequential'\n        q = rotation_matrix_to_quaternion(R)\n        t = t.reshape(3)\n        if not np.all(np.isfinite(q)) or not np.all(np.isfinite(t)):\n            print(f\"Non-finite pose for {img_name}. Using sequential pose.\")\n            q = rotation_matrix_to_quaternion(np.eye(3))\n            t = np.array([0.0, 0.0, len(submission_rows) * 0.1], dtype=np.float64)\n            if pose_status.get(img_name) == 'computed':\n                summary['computed_poses'] -= 1\n                counted_images_processing.add(img_name)\n                summary['sequential_poses_due_to_processing'] += 1\n            pose_status[img_name] = 'sequential'\n    else:\n        print(f\"Image {img_name} not found in poses. Using sequential pose.\")\n        if img_name in counted_images_missing:\n            continue\n        q = rotation_matrix_to_quaternion(np.eye(3))\n        t = np.array([0.0, 0.0, len(submission_rows) * 0.1], dtype=np.float64)\n        pose_status[img_name] = 'sequential'\n        counted_images_missing.add(img_name)\n        summary['sequential_poses_due_to_missing_data'] += 1\n    submission_rows.append([img_path] + q.tolist() + t.tolist())\n\nsubmission_columns = [\n    'image_id',\n    'rotation_w', 'rotation_x', 'rotation_y', 'rotation_z',\n    'translation_x', 'translation_y', 'translation_z'\n]\nsubmission_df = pd.DataFrame(submission_rows, columns=submission_columns)\nsubmission_df.to_csv('submission.csv', index=False)\nprint(\"Submission file created: submission.csv\")\n\n# Validate submission\nprint(\"Validating submission...\")\nassert len(submission_df) == len(sample_submission), f\"Submission has {len(submission_df)} rows, expected {len(sample_submission)}\"\nassert set(submission_df.columns) == set(submission_columns), \"Submission columns do not match expected columns\"\n\nfor i, row in submission_df.iterrows():\n    q = np.array([\n        row['rotation_w'], row['rotation_x'], row['rotation_y'], row['rotation_z']\n    ], dtype=np.float64)\n    t = np.array([\n        row['translation_x'], row['translation_y'], row['translation_z']\n    ], dtype=np.float64)\n    assert not np.any(np.isnan(q)), f\"NaN values in quaternion at row {i}\"\n    assert not np.any(np.isnan(t)), f\"NaN values in translation at row {i}\"\n    assert not np.any(np.isinf(q)), f\"Infinite values in quaternion at row {i}\"\n    assert not np.any(np.isinf(t)), f\"Infinite values in translation at row {i}\"\n    norm = np.linalg.norm(q)\n    assert abs(norm - 1.0) < 1e-6, f\"Quaternion not normalized at row {i}: norm={norm}\"\n    R = quaternion_to_rotation_matrix(q)\n    det_R = np.linalg.det(R)\n    assert abs(det_R - 1.0) < 1e-6, f\"Rotation matrix determinant not 1 at row {i}: det={det_R}\"\n    orthogonality = np.linalg.norm(R.T @ R - np.eye(3))\n    assert orthogonality < 1e-6, f\"Rotation matrix not orthogonal at row {i}: orthogonality={orthogonality}\"\n\nprint(\"Submission validated successfully.\")\n\n# Compute total sequential poses\nsummary['sequential_poses'] = summary['sequential_poses_due_to_missing_data'] + summary['sequential_poses_due_to_processing']\n\n# Print summary\nprint(\"\\n=== Processing Summary ===\")\nprint(f\"Total dataset/scene groups: {summary['total_groups']}\")\nprint(f\"Successful groups: {summary['successful_groups']}\")\nprint(f\"Failed groups: {summary['failed_groups']}\")\nprint(f\"Total images: {summary['total_images']}\")\nprint(f\"Computed poses: {summary['computed_poses']}\")\nprint(f\"Sequential poses: {summary['sequential_poses']}\")\nprint(f\"  - Due to missing data: {summary['sequential_poses_due_to_missing_data']}\")\nprint(f\"  - Due to processing: {summary['sequential_poses_due_to_processing']}\")\nprint(f\"Images not found: {summary['images_not_found']}\")\nprint(f\"High reprojection errors: {summary['high_reprojection_errors']}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-06T17:48:35.831970Z","iopub.execute_input":"2025-04-06T17:48:35.832341Z","iopub.status.idle":"2025-04-06T17:51:42.624123Z","shell.execute_reply.started":"2025-04-06T17:48:35.832314Z","shell.execute_reply":"2025-04-06T17:51:42.622792Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}