{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.12.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":91498,"databundleVersionId":11655853,"sourceType":"competition"}],"dockerImageVersionId":31234,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# **Colmap-Based Vocab Tree Performance**","metadata":{}},{"cell_type":"markdown","source":"\n## **🎯 Overview**\nA **production-ready image retrieval system** leveraging COLMAP's mature computer vision pipeline with pre-trained Flickr vocabulary tree for large-scale image matching.\n\n## **🔄 Pipeline Workflow**\n\n### **1. Environment Setup with COLMAP Integration**\n```python\n# System-level setup with proper dependencies\nsetup_environment()  # Fixes NumPy, installs COLMAP, virtual display\n\n# COLMAP is installed via apt-get: sudo apt-get install colmap\n# Vocabulary tree: 500MB pre-trained model (256,000 visual words)\n```\n\n### **2. Download Pre-trained Vocabulary Tree**\n- **112.4MB compressed file** (not 500MB as expected - likely compressed)\n- **256,000 visual words** trained on Flickr 100K dataset\n- **Used by COLMAP's vocab_tree_matcher**\n\n### **3. COLMAP Feature Extraction Pipeline**\n```\nCOLMAP feature_extractor → database.db (SQLite)\n    ↓\n    • SIFT features (max 8192 per image)\n    • Camera models (SIMPLE_RADIAL)\n    • Multi-threading (4 threads)\n    • Output: 446,291 total features (avg 5,872 per image)\n```\n\n### **4. Vocabulary Tree Matching**\n```bash\ncolmap vocab_tree_matcher \\\n  --database_path database.db \\\n  --VocabTreeMatching.vocab_tree_path vocab_tree.bin \\\n  --VocabTreeMatching.num_images 10 \\\n  --VocabTreeMatching.num_nearest_neighbors 10 \\\n  --VocabTreeMatching.num_checks 256\n```\n\n### **5. Results Analysis**\n- **182 image pairs** with successful matches\n- **34,473 total feature matches**\n- **Average 189.4 matches per pair**\n- **Maximum 1,556 matches** for best pair\n\n## **📊 Key Performance Metrics**\n\n### **COLMAP vs Custom Implementation Comparison**\n\n| Metric | COLMAP Pipeline | Custom Pipeline |\n|--------|-----------------|-----------------|\n| **Vocabulary Size** | 256,000 words | 4,096 words |\n| **Features/Image** | 5,872 (avg) | 500 (fixed) |\n| **Total Features** | 446,291 | 25,000 |\n| **Processing Time** | ~2-3 minutes | ~30 seconds |\n| **Match Quality** | **189.4 matches/pair** | Score-based |\n| **Memory Usage** | 112MB (vocab) | 2.4MB (vocab) |\n| **Dependencies** | COLMAP + SQLite | Pure Python |\n\n### **Sample Match Results from COLMAP:**\n```\nQuery Image: church_00092.png\n→ Multiple matches with 200-500 feature correspondences per pair\n\nQuery Image: church_00087.png  \n→ Strong matches within same scene\n\nQuery Image: church_00013.png\n→ Cross-scene matches with verification\n```\n\n## **🔧 Technical Components Breakdown**\n\n### **A. COLMAP's Vocabulary Tree Architecture**\n```\nPre-trained Flickr Tree (256K words)\n    ↓\nImage → SIFT → Visual Words → Inverted Index\n    ↓\nDatabase (SQLite) → Two-view Geometries\n    ↓\nRANSAC Verification → Geometric Consistency\n```\n\n### **B. Database Structure (SQLite)**\n```sql\n-- Key tables in COLMAP database:\n1. images (image_id, name, camera_id)\n2. keypoints (image_id, rows, cols, data)\n3. descriptors (image_id, rows, cols, data)  \n4. two_view_geometries (pair_id, rows, data)\n5. matches (pair_id, rows, cols, data)\n```\n\n### **C. Pair ID Encoding**\n```python\n# COLMAP's efficient pair encoding:\npair_id = image_id1 * 2147483647 + image_id2\n# Decode: image_id2 = pair_id % 2147483647\n#         image_id1 = pair_id // 2147483647\n```\n\n## **⚡ Performance Optimization Features**\n\n### **1. Multi-threading**\n- **4 threads** for SIFT extraction\n- **Parallel matching** using vocabulary tree\n\n### **2. Database Optimization**\n- **SQLite for persistent storage**\n- **Efficient binary data storage** for features\n- **Indexed queries** for fast retrieval\n\n### **3. Geometric Verification**\n- **RANSAC-based outlier rejection**\n- **Epipolar geometry constraints**\n- **Fundamental matrix estimation**\n\n## **🎯 Use Cases & Applications**\n\n### **1. Large-Scale 3D Reconstruction**\n```bash\n# Full COLMAP pipeline for photogrammetry\ncolmap feature_extractor\ncolmap vocab_tree_matcher\ncolmap mapper\ncolmap bundle_adjuster\ncolmap model_converter\n```\n\n### **2. Image-Based Localization**\n- **Visual place recognition**\n- **Augmented reality tracking**\n- **Drone navigation**\n\n### **3. Content-Based Image Retrieval**\n- **Duplicate detection** at scale\n- **Visual search engines**\n- **Image organization systems**\n\n## **✅ Advantages of COLMAP Pipeline**\n\n### **1. Production-Grade Robustness**\n- **Tested on millions of images**\n- **Robust to scale changes, rotations**\n- **Geometric verification reduces false positives**\n\n### **2. Comprehensive Feature Set**\n- **Full SIFT implementation** with adjustable parameters\n- **Multiple camera models** supported\n- **Bundle adjustment** integration\n\n### **3. Mature Ecosystem**\n- **CLI tools** for automation\n- **Python bindings** available\n- **Active community** support\n\n## **⚠️ Limitations & Considerations**\n\n### **1. Computational Requirements**\n- **Memory intensive** with large datasets\n- **Slower initialization** due to COLMAP setup\n- **Disk I/O** for database operations\n\n### **2. Dependency Management**\n- **COLMAP installation** required\n- **System-level dependencies** (OpenGL, CUDA optional)\n- **Version compatibility** issues\n\n### **3. Black Box Components**\n- **Limited customization** of core algorithms\n- **Proprietary optimizations** not exposed\n- **Debugging challenges** in compiled code\n\n## **📈 Performance Analysis**\n\n### **Your Results Show:**\n1. **High-quality matches**: 189.4 average matches per pair\n2. **Comprehensive coverage**: 182 matching pairs from 76 images\n3. **Scalable processing**: Handled 446K features efficiently\n4. **Reliable verification**: Geometric consistency checks\n\n### **Visualization Outputs:**\n1. **Similarity matrix heatmap** showing cluster patterns\n2. **Top match visualization** with feature counts\n3. **Statistical summaries** for performance evaluation\n\n## **🔮 Integration Possibilities**\n\n### **1. Cloud Deployment**\n```python\n# COLMAP in containerized environments\ndocker run -v $(pwd):/data colmap feature_extractor ...\n```\n\n### **2. Hybrid Approaches**\n```python\n# Use COLMAP for feature extraction + custom matching\ncolmap_features = extract_with_colmap()\ncustom_matching = my_vocab_tree.match(colmap_features)\n```\n\n### **3. Real-time Applications**\n```python\n# Streaming pipeline for live image matching\nwhile True:\n    new_image = capture_frame()\n    features = colmap.extract_features(new_image)\n    matches = vocab_tree.match(features)\n    display_results(matches)\n```\n\n## **🎯 Conclusion**\n\nThe **COLMAP-based vocabulary tree pipeline** demonstrates:\n\n### **Strengths:**\n- **Industrial-grade performance** with 189+ matches per pair\n- **Proven scalability** on large datasets (100K+ images)\n- **Geometric verification** for high accuracy\n- **Complete ecosystem** for 3D reconstruction\n\n### **Best For:**\n- **Research projects** needing reliable baselines\n- **Production systems** requiring robust matching\n- **Large-scale applications** (1000+ images)\n- **3D reconstruction pipelines**\n\n### **Trade-offs:**\n- **Complexity vs performance**\n- **Dependencies vs reliability**  \n- **Setup time vs matching quality**\n\n**Bottom Line**: COLMAP provides a **battle-tested, production-ready solution** for vocabulary tree-based image matching, while the custom implementation offers **simplicity and full control** for research and experimentation.","metadata":{}},{"cell_type":"code","source":"\"\"\"\nCOLMAP Vocab Tree Performance Verification Notebook\nUsing the actual vocab_tree_flickr100K_words256K.bin for image similarity search\n\"\"\"\n\nimport os\nimport sys\nimport subprocess\nimport urllib.request\nimport sqlite3\nimport numpy as np\nimport cv2\nimport matplotlib.pyplot as plt\nfrom pathlib import Path\nfrom PIL import Image\nimport shutil\nfrom tqdm import tqdm\nimport glob","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-31T13:47:23.739419Z","iopub.execute_input":"2025-12-31T13:47:23.740126Z","iopub.status.idle":"2025-12-31T13:47:24.160384Z","shell.execute_reply.started":"2025-12-31T13:47:23.740081Z","shell.execute_reply":"2025-12-31T13:47:24.159327Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def setup_environment():\n    \"\"\"\n    Setup environment with clean NumPy installation at the beginning\n    \"\"\"\n    print(\"Setting up environment for Kaggle...\")\n    WORK_DIR = '/kaggle/working/gaussian_splatting'\n\n    # ========================================================================\n    # STEP 0: Clean NumPy installation BEFORE importing anything\n    # ========================================================================\n    print(\"=\"*70)\n    print(\"STEP 0: Fixing NumPy compatibility (clean install)\")\n    print(\"=\"*70)\n    \n    try:\n        # Uninstall NumPy completely\n        print(\"Uninstalling NumPy 2.x...\")\n        subprocess.run([\n            sys.executable, '-m', 'pip', 'uninstall', '-y', 'numpy'\n        ], check=True, capture_output=True)\n        print(\"✓ NumPy uninstalled\")\n        \n        # Install NumPy 1.x\n        print(\"Installing NumPy 1.x...\")\n        subprocess.run([\n            sys.executable, '-m', 'pip', 'install', 'numpy<2'\n        ], check=True, capture_output=True)\n        print(\"✓ NumPy 1.x installed\")\n        \n        # Reinstall key packages that depend on NumPy\n        print(\"Reinstalling NumPy-dependent packages...\")\n        packages_to_reinstall = [\n            'scikit-learn',\n            'scipy',\n            'matplotlib',\n            'pandas'\n        ]\n        \n        for pkg in packages_to_reinstall:\n            try:\n                subprocess.run([\n                    sys.executable, '-m', 'pip', 'install', '--force-reinstall',\n                    '--no-deps', pkg\n                ], check=True, capture_output=True)\n                print(f\"✓ Reinstalled {pkg}\")\n            except subprocess.CalledProcessError:\n                print(f\"⚠ Failed to reinstall {pkg} (may not be critical)\")\n        \n        # Verify NumPy version\n        result = subprocess.run([\n            sys.executable, '-c', 'import numpy; print(numpy.__version__)'\n        ], capture_output=True, text=True)\n        numpy_version = result.stdout.strip()\n        print(f\"\\n✓ NumPy version now: {numpy_version}\")\n        \n        if numpy_version.startswith('1.'):\n            print(\"✓ NumPy fix successful!\")\n        else:\n            print(f\"⚠ Warning: NumPy version is {numpy_version}, expected 1.x\")\n            \n    except subprocess.CalledProcessError as e:\n        print(f\"⚠ NumPy fix encountered issues: {e}\")\n        print(\"Continuing anyway...\")\n\n    # ========================================================================\n    # STEP 1: System packages and dependencies\n    # ========================================================================\n    print(\"\\n\" + \"=\"*70)\n    print(\"STEP 1: Installing system packages\")\n    print(\"=\"*70)\n    \n    # Virtual display setup\n    try:\n        print(\"Setting up virtual display...\")\n        subprocess.run(['apt-get', 'update', '-qq'], check=True, capture_output=True)\n        subprocess.run(['apt-get', 'install', '-y', '-qq', 'xvfb'], \n                      check=True, capture_output=True)\n        \n        os.environ['QT_QPA_PLATFORM'] = 'offscreen'\n        os.environ['DISPLAY'] = ':99'\n        subprocess.Popen(['Xvfb', ':99', '-screen', '0', '1024x768x24'], \n                         stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL)\n        print(\"✓ Virtual display setup\")\n    except Exception as e:\n        print(f\"⚠ Virtual display skipped: {e}\")\n\n    # Install COLMAP\n    print(\"\\nInstalling COLMAP...\")\n    try:\n        subprocess.run(['apt-get', 'install', '-y', '-qq', 'colmap'], \n                       check=True, capture_output=True)\n        print(\"✓ COLMAP installed\")\n    except subprocess.CalledProcessError as e:\n        print(f\"⚠ COLMAP warning: {e}\")\n\n    # Install build dependencies\n    print(\"\\nInstalling build dependencies...\")\n    try:\n        subprocess.run([\n            'apt-get', 'install', '-y', '-qq',\n            'build-essential', 'cmake', 'git', 'libopenblas-dev'\n        ], check=True, capture_output=True)\n        print(\"✓ Build dependencies installed\")\n    except subprocess.CalledProcessError as e:\n        print(f\"⚠ Build dependencies warning: {e}\")\n\nsetup_environment()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-31T13:47:24.162269Z","iopub.execute_input":"2025-12-31T13:47:24.162546Z","iopub.status.idle":"2025-12-31T13:48:52.844832Z","shell.execute_reply.started":"2025-12-31T13:47:24.162520Z","shell.execute_reply":"2025-12-31T13:48:52.843475Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(\"=\"*70)\nprint(\"COLMAP Vocab Tree Performance Verification System\")\nprint(\"=\"*70)\n\n# ========================================\n# 1. Download Vocab Tree\n# ========================================\n\ndef download_vocab_tree():\n    \"\"\"Download official Vocab Tree file (500MB)\"\"\"\n    vocab_url = \"https://demuc.de/colmap/vocab_tree_flickr100K_words256K.bin\"\n    vocab_path = \"vocab_tree_flickr100K_words256K.bin\"\n    \n    if os.path.exists(vocab_path):\n        file_size = os.path.getsize(vocab_path) / (1024**2)\n        print(f\"✓ Vocab tree already exists: {vocab_path} ({file_size:.1f}MB)\")\n        return vocab_path\n    \n    print(f\"\\n📥 Downloading Vocab Tree... (~500MB)\")\n    print(f\"URL: {vocab_url}\")\n    print(\"⏱️  This may take 5-10 minutes...\")\n    \n    try:\n        def report_progress(block_num, block_size, total_size):\n            downloaded = block_num * block_size\n            percent = min(100, (downloaded / total_size) * 100)\n            mb_downloaded = downloaded / (1024**2)\n            mb_total = total_size / (1024**2)\n            print(f\"\\rProgress: {percent:.1f}% ({mb_downloaded:.1f}MB / {mb_total:.1f}MB)\", end='')\n        \n        urllib.request.urlretrieve(vocab_url, vocab_path, reporthook=report_progress)\n        print(f\"\\n✅ Download complete: {vocab_path}\")\n        return vocab_path\n    except Exception as e:\n        print(f\"\\n❌ Download failed: {e}\")\n        raise\n\n# Download vocab tree\nvocab_tree_path = download_vocab_tree()\nprint(vocab_tree_path)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-31T13:48:52.846124Z","iopub.execute_input":"2025-12-31T13:48:52.846455Z","iopub.status.idle":"2025-12-31T13:48:59.743902Z","shell.execute_reply.started":"2025-12-31T13:48:52.846427Z","shell.execute_reply":"2025-12-31T13:48:59.743097Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def use_local_images(image_folder):\n    \"\"\"Use local images instead\"\"\"\n    if not os.path.exists(image_folder):\n        raise ValueError(f\"Folder does not exist: {image_folder}\")\n    \n    image_files = []\n    for ext in ['*.jpg', '*.jpeg', '*.png', '*.JPG', '*.JPEG', '*.PNG']:\n        image_files.extend(glob.glob(os.path.join(image_folder, ext)))\n    \n    image_files.sort()\n    print(f\"✓ Found {len(image_files)} images in {image_folder}/\")\n    return image_folder, image_files\n\n# Choose image source\nUSE_SAMPLE_IMAGES = True  # Set to False to use your own images\nimage_dir, image_paths = use_local_images(\"/kaggle/input/image-matching-challenge-2025/train/imc2023_theather_imc2024_church\")\n\n# ========================================\n# 3. Display Images\n# ========================================\n\ndef display_images(image_paths, cols=5, figsize=(15, 8)):\n    \"\"\"Display images in a grid\"\"\"\n    n_images = len(image_paths)\n    rows = (n_images + cols - 1) // cols\n    \n    fig, axes = plt.subplots(rows, cols, figsize=figsize)\n    axes = axes.flatten() if isinstance(axes, np.ndarray) else [axes]\n    \n    for idx, img_path in enumerate(image_paths):\n        if idx < len(axes):\n            img = cv2.imread(img_path)\n            if img is not None:\n                img_rgb = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)\n                axes[idx].imshow(img_rgb)\n                axes[idx].set_title(os.path.basename(img_path), fontsize=9)\n            axes[idx].axis('off')\n    \n    for idx in range(n_images, len(axes)):\n        axes[idx].axis('off')\n    \n    plt.tight_layout()\n    plt.show()\n\nprint(\"\\n📷 Test Images:\")\ndisplay_images(image_paths[0:20])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-31T13:48:59.744912Z","iopub.execute_input":"2025-12-31T13:48:59.745253Z","iopub.status.idle":"2025-12-31T13:49:02.602568Z","shell.execute_reply.started":"2025-12-31T13:48:59.745212Z","shell.execute_reply":"2025-12-31T13:49:02.601375Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ========================================\n# 4. Setup COLMAP Environment\n# ========================================\n\ndef setup_colmap_environment():\n    \"\"\"Setup environment for COLMAP\"\"\"\n    env = os.environ.copy()\n    env['QT_QPA_PLATFORM'] = 'offscreen'\n    \n    # Check if COLMAP is installed\n    try:\n        result = subprocess.run(['colmap', '-h'], \n                              capture_output=True, \n                              text=True, \n                              env=env)\n        print(\"✓ COLMAP is available\")\n        return env\n    except FileNotFoundError:\n        print(\"❌ COLMAP not found!\")\n        print(\"Please install COLMAP:\")\n        print(\"  Ubuntu/Debian: sudo apt-get install colmap\")\n        print(\"  macOS: brew install colmap\")\n        print(\"  Or build from source: https://colmap.github.io/install.html\")\n        raise\n\nenv = setup_colmap_environment()\n\n# ========================================\n# 5. Run COLMAP Feature Extraction\n# ========================================\n\ndef extract_features(image_dir, database_path, env):\n    \"\"\"Extract SIFT features using COLMAP\"\"\"\n    print(\"\\n🔍 Extracting SIFT features with COLMAP...\")\n    \n    # Remove existing database\n    if os.path.exists(database_path):\n        os.remove(database_path)\n    \n    try:\n        subprocess.run([\n            'colmap', 'feature_extractor',\n            '--database_path', database_path,\n            '--image_path', image_dir,\n            '--ImageReader.single_camera', '0',\n            '--ImageReader.camera_model', 'SIMPLE_RADIAL',\n            '--SiftExtraction.max_num_features', '8192',\n            '--SiftExtraction.num_threads', '4'\n        ], check=True, env=env, capture_output=True, text=True)\n        \n        print(\"✅ Feature extraction complete\")\n        return True\n    except subprocess.CalledProcessError as e:\n        print(f\"❌ Feature extraction failed: {e.stderr}\")\n        return False\n\ncolmap_workspace = \"colmap_workspace\"\nos.makedirs(colmap_workspace, exist_ok=True)\ndatabase_path = os.path.join(colmap_workspace, \"database.db\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-31T13:49:02.603936Z","iopub.execute_input":"2025-12-31T13:49:02.604370Z","iopub.status.idle":"2025-12-31T13:49:02.668255Z","shell.execute_reply.started":"2025-12-31T13:49:02.604326Z","shell.execute_reply":"2025-12-31T13:49:02.667525Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if extract_features(image_dir, database_path, env):\n    # ========================================\n    # 6. Vocab Tree Matching (FIXED)\n    # ========================================\n    \n    def vocab_tree_matching(database_path, vocab_tree_path, env, num_images=10):\n        \"\"\"Use vocab tree to match similar images\"\"\"\n        print(f\"\\n🌳 Running Vocab Tree matching...\")\n        print(f\"   Finding top {num_images} similar images for each query\")\n        \n        try:\n            # Use vocab_tree_matcher instead of vocab_tree_retriever\n            subprocess.run([\n                'colmap', 'vocab_tree_matcher',\n                '--database_path', database_path,\n                '--VocabTreeMatching.vocab_tree_path', vocab_tree_path,\n                '--VocabTreeMatching.num_images', str(num_images),\n                '--VocabTreeMatching.num_nearest_neighbors', '10',\n                '--VocabTreeMatching.num_checks', '256',\n                '--VocabTreeMatching.num_images_after_verification', str(num_images)\n            ], check=True, env=env, capture_output=True, text=True)\n            \n            print(\"✅ Vocab tree matching complete\")\n            return True\n        except subprocess.CalledProcessError as e:\n            print(f\"❌ Vocab tree matching failed: {e.stderr}\")\n            return False\n    \n    if vocab_tree_matching(database_path, vocab_tree_path, env, num_images=10):\n        # ========================================\n        # 7. Extract and Visualize Results\n        # ========================================\n        \n        def decode_pair_id(pair_id):\n            \"\"\"Decode COLMAP pair_id to image_id1 and image_id2\"\"\"\n            # COLMAP encodes: pair_id = image_id1 * 2147483647 + image_id2\n            # where image_id1 < image_id2\n            image_id2 = pair_id % 2147483647\n            image_id1 = pair_id // 2147483647\n            return int(image_id1), int(image_id2)\n        \n        def extract_image_pairs(database_path):\n            \"\"\"Extract image pair matches from COLMAP database\"\"\"\n            conn = sqlite3.connect(database_path)\n            cursor = conn.cursor()\n            \n            # Get image names\n            cursor.execute(\"SELECT image_id, name FROM images\")\n            images = {row[0]: row[1] for row in cursor.fetchall()}\n            \n            # Get number of matches for each pair (using pair_id)\n            cursor.execute(\"\"\"\n                SELECT pair_id, rows\n                FROM two_view_geometries\n                WHERE rows > 0\n                ORDER BY rows DESC\n            \"\"\")\n            \n            # Decode pair_ids to get image_id1 and image_id2\n            match_counts = []\n            for pair_id, num_matches in cursor.fetchall():\n                img_id1, img_id2 = decode_pair_id(pair_id)\n                match_counts.append((img_id1, img_id2, num_matches))\n            \n            conn.close()\n            \n            print(f\"\\n📊 Database Statistics:\")\n            print(f\"   Total images: {len(images)}\")\n            print(f\"   Image pairs with matches: {len(match_counts)}\")\n            \n            return images, match_counts\n        \n        def visualize_similar_images(database_path, image_dir, top_n=5):\n            \"\"\"Visualize similar image pairs\"\"\"\n            images, match_counts = extract_image_pairs(database_path)\n            \n            if not match_counts:\n                print(\"⚠️  No matches found\")\n                return\n            \n            # Group by query image\n            query_matches = {}\n            for img_id1, img_id2, num_matches in match_counts:\n                if img_id1 not in query_matches:\n                    query_matches[img_id1] = []\n                query_matches[img_id1].append((img_id2, num_matches))\n            \n            # Visualize for each query image (show first 3 queries)\n            for query_id in list(query_matches.keys())[:min(3, len(query_matches))]:\n                matches = sorted(query_matches[query_id], \n                               key=lambda x: x[1], reverse=True)[:top_n]\n                \n                print(f\"\\n🔎 Query Image: {images[query_id]}\")\n                \n                fig, axes = plt.subplots(1, top_n + 1, figsize=(20, 4))\n                \n                # Display query image\n                query_path = os.path.join(image_dir, images[query_id])\n                query_img = cv2.imread(query_path)\n                if query_img is not None:\n                    query_img = cv2.cvtColor(query_img, cv2.COLOR_BGR2RGB)\n                    axes[0].imshow(query_img)\n                    axes[0].set_title(f\"Query\\n{images[query_id]}\", \n                                    fontweight='bold', fontsize=10)\n                    axes[0].axis('off')\n                \n                # Display similar images\n                for idx, (match_id, num_matches) in enumerate(matches):\n                    match_path = os.path.join(image_dir, images[match_id])\n                    match_img = cv2.imread(match_path)\n                    if match_img is not None:\n                        match_img = cv2.cvtColor(match_img, cv2.COLOR_BGR2RGB)\n                        axes[idx + 1].imshow(match_img)\n                        axes[idx + 1].set_title(\n                            f\"Match #{idx+1}\\n{images[match_id]}\\n\"\n                            f\"({num_matches} matches)\",\n                            fontsize=9\n                        )\n                        axes[idx + 1].axis('off')\n                \n                plt.tight_layout()\n                plt.show()\n        \n        # ========================================\n        # 8. Display Results\n        # ========================================\n        \n        print(\"\\n\" + \"=\"*70)\n        print(\"📈 VOCAB TREE PERFORMANCE RESULTS\")\n        print(\"=\"*70)\n        \n        visualize_similar_images(database_path, image_dir, top_n=5)\n        \n        # ========================================\n        # 9. Create Similarity Matrix\n        # ========================================\n        \n        def create_similarity_matrix(database_path):\n            \"\"\"Create similarity matrix from match counts\"\"\"\n            conn = sqlite3.connect(database_path)\n            cursor = conn.cursor()\n            \n            cursor.execute(\"SELECT image_id FROM images ORDER BY image_id\")\n            image_ids = [row[0] for row in cursor.fetchall()]\n            n = len(image_ids)\n            \n            # Initialize matrix\n            similarity_matrix = np.zeros((n, n))\n            \n            # Fill with match counts (decode pair_id)\n            cursor.execute(\"SELECT pair_id, rows FROM two_view_geometries\")\n            for pair_id, matches in cursor.fetchall():\n                img_id1, img_id2 = decode_pair_id(pair_id)\n                try:\n                    idx1 = image_ids.index(img_id1)\n                    idx2 = image_ids.index(img_id2)\n                    similarity_matrix[idx1, idx2] = matches\n                    similarity_matrix[idx2, idx1] = matches\n                except ValueError:\n                    # Skip if image_id not found in list\n                    continue\n            \n            conn.close()\n            \n            return similarity_matrix, image_ids\n        \n        print(\"\\n📊 Creating similarity matrix...\")\n        sim_matrix, img_ids = create_similarity_matrix(database_path)\n        \n        # Normalize for visualization\n        max_matches = np.max(sim_matrix)\n        if max_matches > 0:\n            sim_matrix_norm = sim_matrix / max_matches\n        else:\n            sim_matrix_norm = sim_matrix\n        \n        # Visualize\n        plt.figure(figsize=(14, 12))\n        plt.imshow(sim_matrix_norm, cmap='YlOrRd', interpolation='nearest')\n        plt.colorbar(label='Normalized Match Count')\n        plt.title('Image Similarity Matrix (Vocab Tree Results)', \n                 fontsize=14, fontweight='bold')\n        plt.xlabel('Image Index')\n        plt.ylabel('Image Index')\n        \n        # Get image names for labels\n        conn = sqlite3.connect(database_path)\n        cursor = conn.cursor()\n        cursor.execute(\"SELECT image_id, name FROM images ORDER BY image_id\")\n        image_names = [row[1][:15] for row in cursor.fetchall()]\n        conn.close()\n        \n        plt.xticks(range(len(image_names)), image_names, rotation=90, ha='right', fontsize=8)\n        plt.yticks(range(len(image_names)), image_names, fontsize=8)\n        plt.tight_layout()\n        plt.show()\n        \n        # ========================================\n        # 10. Summary Statistics\n        # ========================================\n        \n        print(\"\\n\" + \"=\"*70)\n        print(\"📊 SUMMARY STATISTICS\")\n        print(\"=\"*70)\n        \n        total_matches = int(np.sum(sim_matrix) / 2)  # Divide by 2 because matrix is symmetric\n        avg_matches = np.mean(sim_matrix[sim_matrix > 0]) if np.any(sim_matrix > 0) else 0\n        max_match = int(np.max(sim_matrix))\n        \n        print(f\"Vocab Tree File: {vocab_tree_path}\")\n        print(f\"File Size: {os.path.getsize(vocab_tree_path) / (1024**2):.1f} MB\")\n        print(f\"Number of Images: {len(image_paths)}\")\n        print(f\"Total Image Pairs Matched: {total_matches}\")\n        print(f\"Average Feature Matches per Pair: {avg_matches:.1f}\")\n        print(f\"Maximum Feature Matches: {max_match}\")\n        \n        # Get feature counts per image\n        conn = sqlite3.connect(database_path)\n        cursor = conn.cursor()\n        cursor.execute(\"SELECT rows FROM keypoints\")\n        feature_counts = [row[0] for row in cursor.fetchall()]\n        conn.close()\n        \n        print(f\"Average Features per Image: {np.mean(feature_counts):.1f}\")\n        print(f\"Total Features Extracted: {sum(feature_counts)}\")\n        \n        print(\"\\n✅ Vocab Tree performance verification complete!\")\n\nelse:\n    print(\"\\n❌ Feature extraction failed. Cannot proceed with vocab tree matching.\")\n\nprint(\"\\n\" + \"=\"*70)\nprint(\"DEMO COMPLETE\")\nprint(\"=\"*70)","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}