{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.11","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":91498,"databundleVersionId":11655853,"isSourceIdPinned":false,"sourceType":"competition"},{"sourceId":10073121,"sourceType":"datasetVersion","datasetId":6208911},{"sourceId":12029591,"sourceType":"datasetVersion","datasetId":7568826},{"sourceId":12033698,"sourceType":"datasetVersion","datasetId":7571696}],"dockerImageVersionId":31041,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Image Matching Challenge 2025 - Main Submission Pipeline 🚀\n\n**Team:** Maschinelles Lernen ([Dave](https://www.kaggle.com/daveanalyst), [Davin](https://www.kaggle.com/epimorbit), [bansalraman](https://www.kaggle.com/bansalraman)) \n**Version:** `0.2` (Date: `[30-05-2025]`) \n**Objective:** Develop a pipeline to reconstruct 3D scenes from image collections by clustering images into scenes and estimating camera poses (R, T).\n\n---\n\n## **Pipeline Overview**\n\nThis notebook orchestrates the full end-to-end pipeline for the [Image Matching Challenge 2025](https://www.kaggle.com/competitions/image-matching-challenge-2025). The major stages include:\n\n1.  **Setup & Configuration:** Installs packages, clones our codebase from GitHub, and sets up paths.\n2.  **Data Loading & Preparation:** Loads the test image list (`sample_submission.csv`) and prepares image paths for processing.\n3.  **Image Preprocessing (Shared Utility):** Standardizes images (loading, resizing) using `src/data/preprocessing.py`.\n4.  **Global Feature Extraction (DINOv2):** Extracts global image embeddings using `src/features/global_dino_extractor.py` for overall image understanding.\n5.  **Clustering (HDBSCAN):** Groups images into potential scenes or identifies outliers using `src/clustering/hdbscan_clusterer.py` based on global embeddings.\n6.  **Image Pair Selection:** Intelligently selects promising image pairs *within each cluster* for local matching using DINOv2 embedding similarity via `src/matching_strategies/pair_selector.py`.\n7.  **Local Matching & SfM:** ALIKED+LightGlue for pairwise matching (`src/features/local_matcher.py` ) and COLMAP for 3D reconstruction and pose (R,T) estimation (`src/sfm/colmap_runner.py` ). *(Initially, this stage will be a placeholder outputting NaN poses).*\n8.  **Submission File Generation:** Formats the cluster assignments and estimated poses into `submission.csv`.\n\n*Modules referenced (e.g., `src/...`) are part of our team's GitHub repository, which is cloned in the setup phase.*","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"## 1. Setup & Configuration\n\nThis section handles the initial setup required for the notebook to run:\n\n*   **Package Installation:** Installs or upgrades essential Python libraries like `timm` (for Vision Transformer models), `hdbscan` and `umap-learn` (for clustering), `Pillow` (for image manipulation), and `tqdm` (for progress bars). We rely on Kaggle's pre-installed versions for common libraries like `pandas`, `numpy`, `torch`, and `sklearn` unless specific version conflicts arise.\n*   **Code Repository Cloning:**\n    *   Retrieves a GitHub Personal Access Token (PAT) from Kaggle Secrets (if available, for cloning private repositories).\n    *   Clones our team's GitHub repository (`imc2025-team-maschinelles-lernen-1`) into the Kaggle working directory. This provides access to our custom Python modules located in the `src/` directory.\n    *   Currently, it clones the `main` branch, assuming all necessary PRs for pipeline components are merged there.\n*   **Python Path Configuration:** Adds the `src/` directory from our cloned repository to Python's `sys.path`. This allows us to import our custom modules (e.g., `from data.preprocessing import ...`).\n*   **Repository Structure Verification:** Checks for the presence of key module directories within the cloned `src/` folder and lists their contents. This helps confirm that the correct code has been cloned and is accessible.\n\nThe output of the following code cell will confirm these setup steps. Any warnings about missing directories indicate that the corresponding Pull Requests for those modules might not yet be merged into the `main` branch of our GitHub repository.","metadata":{}},{"cell_type":"code","source":"# Cell 1: Setup, Package Installation, Offline Model Weights & SRC Code from Dataset\n\n# --- Essential Imports for Setup ---\nimport os\nimport shutil\nimport sys\n# from kaggle_secrets import UserSecretsClient # No longer needed for GITHUB_PAT for cloning\n\nprint(\"===============================================================================\")\nprint(\"=== Stage 1: Installing/Updating Key Packages (Minimal) ===\")\nprint(\"===============================================================================\")\n#!pip install --upgrade timm hdbscan umap-learn Pillow tqdm --quiet \nprint(\"Key package installation/upgrade attempt complete.\")\nprint(\"-\" * 60)\n\n\nprint(\"\\n===============================================================================\")\nprint(\"=== Stage 2: Setting up Offline Model Weights ===\")\nprint(\"===============================================================================\")\n# *** ACTION: VERIFY AND REPLACE 'public-dino-v2-weights-slug' WITH THE ACTUAL SLUG ***\nDINOV2_PUBLIC_DATASET_SLUG = \"dino-small-pretrained\" \nDINOV2_PUBLIC_DATASET_PATH = f\"/kaggle/input/{DINOV2_PUBLIC_DATASET_SLUG}\"\nDINOV2_WEIGHT_FILENAME_IN_DATASET = \"dinov2_vits14_pretrain.pth\"\nEXPECTED_DINOV2_FILENAME_IN_CACHE = \"dinov2_vits14_pretrain.pth\"\n\nPYTORCH_HUB_CACHE_DIR = \"/root/.cache/torch/hub/checkpoints/\"\nos.makedirs(PYTORCH_HUB_CACHE_DIR, exist_ok=True)\n\nsource_dino_weight_path = os.path.join(DINOV2_PUBLIC_DATASET_PATH, DINOV2_WEIGHT_FILENAME_IN_DATASET)\ntarget_dino_weight_path = os.path.join(PYTORCH_HUB_CACHE_DIR, EXPECTED_DINOV2_FILENAME_IN_CACHE)\n\nprint(f\"DINOv2 Source weight path: {source_dino_weight_path}\")\nif os.path.exists(source_dino_weight_path):\n    if source_dino_weight_path != target_dino_weight_path:\n        shutil.copyfile(source_dino_weight_path, target_dino_weight_path)\n        print(\"DINOv2 weights copied to PyTorch Hub cache successfully.\")\nelse:\n    print(f\"CRITICAL WARNING: DINOv2 weight file NOT FOUND at {source_dino_weight_path}.\")\n\n# --- ALIKED & LightGlue Weights (from your team's private Kaggle Dataset) ---\n# *** ACTION: VERIFY AND REPLACE 'imc2025-team-model-weights' WITH YOUR TEAM'S ACTUAL DATASET SLUG ***\nTEAM_MODEL_WEIGHTS_DATASET_SLUG = \"imc2025-team-model-weights\"\nTEAM_MODEL_WEIGHTS_DATASET_PATH = f\"/kaggle/input/{TEAM_MODEL_WEIGHTS_DATASET_SLUG}\"\n\nALIKED_WEIGHT_FILENAME = \"aliked-n16.pth\" \nLIGHTGLUE_WEIGHT_FILENAME = \"aliked_lightglue_v0-1_arxiv.pth\" \n\nALIKED_WEIGHT_KAGGLE_PATH = os.path.join(TEAM_MODEL_WEIGHTS_DATASET_PATH, ALIKED_WEIGHT_FILENAME)\nLIGHTGLUE_WEIGHT_KAGGLE_PATH = os.path.join(TEAM_MODEL_WEIGHTS_DATASET_PATH, LIGHTGLUE_WEIGHT_FILENAME) if LIGHTGLUE_WEIGHT_FILENAME else None\n\nif not os.path.exists(ALIKED_WEIGHT_KAGGLE_PATH): print(f\"CRITICAL WARNING: ALIKED weight file NOT FOUND at {ALIKED_WEIGHT_KAGGLE_PATH}.\")\nif LIGHTGLUE_WEIGHT_KAGGLE_PATH and not os.path.exists(LIGHTGLUE_WEIGHT_KAGGLE_PATH): print(f\"CRITICAL WARNING: LightGlue weight file NOT FOUND at {LIGHTGLUE_WEIGHT_KAGGLE_PATH}.\")\nprint(\"-\" * 60)\n\n\nprint(\"\\n===============================================================================\")\nprint(\"=== Stage 3: Accessing SRC Code from Kaggle Dataset ===\")\nprint(\"===============================================================================\")\n\n# Cell 1: Setup, Package Installation, Offline Model Weights & SRC Code from Dataset\n\n# ... (Your Stage 1: Pip Installs - keep as is or minimal) ...\n# ... (Your Stage 2: Offline Model Weights - keep as is, ensuring DINOv2 weights are copied) ...\n\nprint(\"\\n===============================================================================\")\nprint(\"=== Stage 3: Accessing SRC Code from Kaggle Dataset ===\")\nprint(\"===============================================================================\")\n# *** ACTION: VERIFY AND REPLACE 'imc2025-team-src-code-final' WITH YOUR SRC CODE DATASET SLUG ***\nSRC_CODE_DATASET_SLUG = \"imc2025-team-src-code-final\" \nSRC_CODE_DATASET_PATH = f\"/kaggle/input/{SRC_CODE_DATASET_SLUG}\"\n\n# Path to the 'src' directory *directly inside* your Kaggle Dataset\n# Since os.listdir showed ['src'], it means src/ is directly under the dataset path\nPATH_TO_SRC_IN_DATASET = os.path.join(SRC_CODE_DATASET_PATH, 'src')\n\nif os.path.exists(PATH_TO_SRC_IN_DATASET) and os.path.isdir(PATH_TO_SRC_IN_DATASET):\n    print(f\"Found 'src' directory directly in Kaggle Dataset: {PATH_TO_SRC_IN_DATASET}\")\n    # No unzipping needed. The SRC_PATH for sys.path will be this direct path.\nelse:\n    print(f\"CRITICAL ERROR: 'src' directory NOT FOUND at {PATH_TO_SRC_IN_DATASET}.\")\n    print(f\"         Ensure your dataset '{SRC_CODE_DATASET_SLUG}' contains an 'src' folder at its root.\")\n    print(f\"         Contents of '{SRC_CODE_DATASET_PATH}':\")\n    if os.path.exists(SRC_CODE_DATASET_PATH):\n        print(os.listdir(SRC_CODE_DATASET_PATH))\n    else:\n        print(f\"         Dataset path '{SRC_CODE_DATASET_PATH}' itself not found.\")\nprint(\"-\" * 60)\n\n# --- Stage 4: Setting up Python Path & Verifying Modules ---\nprint(\"\\n===============================================================================\")\nprint(\"=== Stage 4: Setting up Python Path & Verifying Modules ===\")\nprint(\"===============================================================================\")\n# SRC_PATH_FROM_DATASET will now be the direct path from the Kaggle Dataset\nSRC_PATH_FROM_DATASET = PATH_TO_SRC_IN_DATASET # Use the path identified in Stage 3\n\nif os.path.exists(SRC_PATH_FROM_DATASET) and os.path.isdir(SRC_PATH_FROM_DATASET):\n    # We want to add the directory *containing* your modules (data, features, etc.)\n    # which is SRC_PATH_FROM_DATASET itself if it is .../input/dataset_slug/src/\n    sys.path.insert(0, SRC_PATH_FROM_DATASET) \n    print(f\"'{SRC_PATH_FROM_DATASET}' added to sys.path.\")\n    print(f\"Contents of {SRC_PATH_FROM_DATASET} (should be your module folders like 'data', 'features', etc.):\")\n    print(os.listdir(SRC_PATH_FROM_DATASET))\nelse:\n    print(f\"CRITICAL WARNING: Source directory '{SRC_PATH_FROM_DATASET}' does not exist or is not a directory.\")\n    print(f\"         This usually means the dataset structure is not as expected or it wasn't added correctly.\")\n    print(f\"         Imports from 'src' (actually from its submodules) will fail.\")\nprint(\"-\" * 60)\n\n# --- Verify expected module directories within src/ ---\nprint(\"Checking for expected module directories within src/:\")\n# Update this list based on PRs merged to the 'src' folder you uploaded to the dataset\nexpected_module_dirs = {\n    'data': True, \n    'features': True, \n    'clustering': True, \n    'matching_strategies': True,\n    'sfm': True # Set to True if Davin's refactored script is in the 'src/sfm/' you uploaded\n}\nall_critical_modules_found = True\nif os.path.exists(SRC_PATH_FROM_DATASET) and os.path.isdir(SRC_PATH_FROM_DATASET):\n    for subdir_name, is_critical_for_this_run in expected_module_dirs.items():\n        path_to_check = os.path.join(SRC_PATH_FROM_DATASET, subdir_name)\n        if os.path.exists(path_to_check) and os.path.isdir(path_to_check):\n            print(f\"  Found: '{subdir_name}/'\")\n        else:\n            if is_critical_for_this_run:\n                print(f\"  CRITICAL WARNING: Expected CRITICAL module directory NOT FOUND: '{path_to_check}'\")\n                all_critical_modules_found = False\n            else:\n                print(f\"  INFO: Module directory not found: '{path_to_check}' (is_critical={is_critical_for_this_run})\")\n    if all_critical_modules_found:\n        print(\"All critical expected module directories appear to be found in src/.\")\n    else:\n        print(\"CRITICAL WARNING: Not all essential module directories were found. Subsequent imports might fail.\")\nelse:\n    print(f\"Skipping module directory check as SRC_PATH_FROM_DATASET ('{SRC_PATH_FROM_DATASET}') was not found.\")\n    all_critical_modules_found = False\n\nprint(\"-\" * 60)\nprint(\"Setup cell complete. Review any CRITICAL warnings above carefully.\")\n\nPIPELINE_SETUP_OK = True\n# Check if the critical src path itself was found and if critical modules were found\nif not (os.path.exists(SRC_PATH_FROM_DATASET) and os.path.isdir(SRC_PATH_FROM_DATASET)) or \\\n   not all_critical_modules_found:\n    PIPELINE_SETUP_OK = False\n    print(\"\\n!!! PIPELINE SETUP HAS CRITICAL ISSUES - SUBSEQUENT CELLS MAY FAIL OR PRODUCE INVALID RESULTS !!!\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-02T17:11:07.180114Z","iopub.execute_input":"2025-06-02T17:11:07.180593Z","iopub.status.idle":"2025-06-02T17:11:07.377372Z","shell.execute_reply.started":"2025-06-02T17:11:07.180560Z","shell.execute_reply":"2025-06-02T17:11:07.376747Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2. Imports & Path Definitions\n\nThis cell handles the following:\n\n*   **Import Core Libraries:** Imports standard Python libraries like `pandas`, `numpy`, `os`, `sys`, `torch`, `timm`, `PIL`, `torchvision`, and `tqdm`.\n*   **Import Custom Modules:** Imports our team's custom-developed modules from the cloned GitHub repository's `src/` directory. This includes:\n    *   `data.preprocessing`: For image loading and resizing utilities.\n    *   `features.global_dino_extractor`: For extracting DINOv2 global image embeddings.\n    *   `clustering.hdbscan_clusterer`: For performing HDBSCAN clustering on the embeddings.\n    *   *(Placeholder for `matching_strategies.pair_selector`)*\n*   **Define Key Paths:** Sets up essential directory paths for the Kaggle environment:\n    *   `KAGGLE_INPUT_DIR`: Points to the read-only directory where competition data (test images, `sample_submission.csv`) will be provided during scoring.\n    *   `KAGGLE_WORKING_DIR`: Points to the writable directory where our notebook can save intermediate files (like extracted embeddings, cluster assignments) and the final `submission.csv`.\n    *   `CLONED_REPO_PATH`: Path to our cloned GitHub repository within `/kaggle/working/`.\n    *   `FEATURES_OUTPUT_DIR` & `CLUSTERING_OUTPUT_DIR`: Defines subdirectories (e.g., within the cloned repo or `/kaggle/working/`) for storing outputs from the feature extraction and clustering stages if we generate them on-the-fly. These directories are created if they don't exist.\n\nThis setup ensures all necessary tools and paths are ready for the subsequent pipeline stages.","metadata":{}},{"cell_type":"code","source":"# Cell 2: Imports & Path Definitions\n\nprint(\"=== Stage 2.1: Importing Core Libraries and Custom Modules ===\")\n\n# --- Standard Libraries ---\nimport pandas as pd\nimport numpy as np\nimport os\nimport sys\nimport torch \nimport timm  \nfrom PIL import Image \nfrom torchvision import transforms # Often used by timm's model transforms\nfrom tqdm.auto import tqdm \n\n# --- Import Our Custom Modules ---\n# These imports assume the PRs for these modules have been merged into the \n# TARGET_BRANCH (e.g., 'main') that was cloned in the previous setup cell.\n\n# REPO_NAME and TARGET_BRANCH should be defined in Cell 1 (Setup Cell)\n# If not, define them here or ensure they are passed correctly.\nif 'REPO_NAME' not in globals():\n    print(\"CRITICAL WARNING: REPO_NAME not defined from Setup Cell. Using default, but this may be incorrect.\")\n    REPO_NAME = \"imc2025-team-maschinelles-lernen-1\" # Fallback, but best to define in Cell 1\n\nif 'TARGET_BRANCH' not in globals():\n    print(\"CRITICAL WARNING: TARGET_BRANCH not defined from Setup Cell. Using default 'main'.\")\n    TARGET_BRANCH = \"main\" # Fallback\n\nMODULE_BASE_PATH = f'/kaggle/working/{REPO_NAME}/src'\n\nprint(f\"Attempting to import modules from: {MODULE_BASE_PATH} (cloned from branch: '{TARGET_BRANCH}')\")\n\ntry:\n    from data.preprocessing import load_image_pil, resize_image_maintain_aspect_ratio \n    print(\"- Successfully imported: data.preprocessing\")\nexcept ImportError as e:\n    print(f\"- WARNING: Could not import from data.preprocessing: {e}.\")\n    print(f\"  Ensure 'src/data/preprocessing.py' exists in the cloned repo ('{TARGET_BRANCH}' branch) and its PR is merged.\")\n\ntry:\n    from features.global_dino_extractor import DinoV2EmbeddingExtractor \n    print(\"- Successfully imported: features.global_dino_extractor\")\nexcept ImportError as e:\n    print(f\"- WARNING: Could not import from features.global_dino_extractor: {e}.\")\n    print(f\"  Ensure 'src/features/global_dino_extractor.py' exists and PR is merged to '{TARGET_BRANCH}'.\")\n\ntry:\n    from clustering.hdbscan_clusterer import run_hdbscan_clustering, load_embeddings \n    print(\"- Successfully imported: clustering.hdbscan_clusterer\")\nexcept ImportError as e:\n    print(f\"- WARNING: Could not import from clustering.hdbscan_clusterer: {e}.\")\n    print(f\"  Ensure 'src/clustering/hdbscan_clusterer.py' exists and PR is merged to '{TARGET_BRANCH}'.\")\n\ntry:\n    from matching_strategies.pair_selector import select_pairs_by_embedding_similarity\n    print(\"- Successfully imported: matching_strategies.pair_selector\")\nexcept ImportError as e:\n    print(f\"- WARNING: Could not import from matching_strategies.pair_selector: {e}.\")\n    print(f\"  Ensure 'src/matching_strategies/pair_selector.py' exists and PR is merged to '{TARGET_BRANCH}'.\")\n\n# Placeholder for Davin's/Raman's SfM module - uncomment when its PR is ready and merged\n# try:\n#     from sfm.scene_reconstructor import reconstruct_one_scene # Example name\n#     print(\"- Successfully imported: sfm.scene_reconstructor\")\n# except ImportError as e:\n#     print(f\"- INFO: sfm.scene_reconstructor not yet imported (PR might not be merged or module not ready): {e}\")\n\nprint(\"-\" * 60)\n\n# --- Define Key Directory Paths ---\nprint(\"=== Stage 2.2: Defining Key Directory Paths ===\")\n\nKAGGLE_INPUT_DIR = '/kaggle/input/image-matching-challenge-2025/'\nKAGGLE_WORKING_DIR = '/kaggle/working/' # Writable directory\n\n# Define paths for saving intermediate outputs generated by this notebook\n# These will be in /kaggle/working/, which is cleared after the session but available for submission output.\nFEATURES_NPZ_DIR = os.path.join(KAGGLE_WORKING_DIR, 'features_output', 'dino_embeddings')\nCLUSTERING_CSV_DIR = os.path.join(KAGGLE_WORKING_DIR, 'features_output', 'clustering_results')\n# Path for the final submission file\nSUBMISSION_DIR = KAGGLE_WORKING_DIR # submission.csv goes directly in /kaggle/working/\n\nprint(f\"Competition Input Directory: {KAGGLE_INPUT_DIR}\")\nprint(f\"Notebook Working Directory: {KAGGLE_WORKING_DIR}\")\nprint(f\"Path for DINOv2 embeddings output (NPZ): {FEATURES_NPZ_DIR}\")\nprint(f\"Path for Clustering results output (CSV): {CLUSTERING_CSV_DIR}\")\nprint(f\"Path for final submission.csv: {SUBMISSION_DIR}\")\n\n# Ensure these output directories exist\nos.makedirs(FEATURES_NPZ_DIR, exist_ok=True)\nos.makedirs(CLUSTERING_CSV_DIR, exist_ok=True)\nprint(\"Output directories for features and clustering ensured.\")\nprint(\"-\" * 60)\nprint(\"Imports and Path Definitions cell complete. Review any warnings carefully.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-02T17:11:07.378594Z","iopub.execute_input":"2025-06-02T17:11:07.378865Z","iopub.status.idle":"2025-06-02T17:11:07.388855Z","shell.execute_reply.started":"2025-06-02T17:11:07.378847Z","shell.execute_reply":"2025-06-02T17:11:07.388114Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n## 3. Data Loading & Test Image Path Preparation\n\nThis stageimage-matching-challenge-2025/`) and lists all the images in the hidden test set for which we need to generate predictions. We load this CSV into a pandas DataFrame.\n*   **Construct Full Image Paths:**\n    *   The `sample_submission.csv` contains `dataset` and `image` (filename) columns.\n    *   The competition provides test images in a **flat directory structure** located at `/kaggle/input/image-matching-challenge-2025/test/`.\n    *   Our path construction logic directly combines this base path with the image filename from the CSV. The `dataset` column is used for logical grouping and final submission output but not for path construction to the image file itself prepares the list of test images that our pipeline needs to process.\n\n*   **Load `sample_submission.csv`:** During a Kaggle submission run, this file is provided in the input directory (`/kaggle/input/image-matching-challenge-2025/`) and lists all the images in the hidden test set for which we need to generate predictions. We load this CSV into a pandas DataFrame.\n*   **Construct Full Image Paths:**\n    *   The `sample_submission.csv` contains `dataset` and `image` (filename) columns.\n    *   We define a function (`get_actual_test_image_path`) to construct the full, absolute path to each test image file.\n    *   **Test Data Structure:** It is assumed that all hidden test images are located directly within the `/kaggle/input/image-matching-challenge-2025/test/` directory. The `dataset` column from `sample_submission.csv` will be used for logical grouping and final submission output, but not for forming the image file path.\n*   **Prepare List for Processing:** A DataFrame `test_set_df` is created containing the `dataset` identifier, `image` filename, the constructed `full_path` to the image, and an `image_id_combined` (e.g., `dataset__image_filename`) for unique identification throughout the pipeline.\n\nThe output of the following code cell will show the number of images to process and example paths. Any warnings about paths not existing during a submission run would indicate an unexpected change in the hidden test data structure.\n","metadata":{}},{"cell_type":"code","source":"# Cell 3: Data Loading & Test Image Path Preparation\n\nimport pandas as pd\nimport os # Should already be imported\n\nprint(\"=== Stage 3: Loading sample_submission.csv and Preparing Test Image Paths ===\")\n\n# Initialize test_set_df to ensure it's always defined with correct columns\ntest_set_df = pd.DataFrame(columns=['dataset', 'image', 'full_path', 'image_id_combined'])\nsample_submission_df = pd.DataFrame() # Initialize as empty\nALL_IMAGES_LISTED_IN_SAMPLE_SUB_PREPARED = False # Flag\n\n# KAGGLE_INPUT_DIR should be defined in Cell 2 (Imports & Path Definitions)\nif 'KAGGLE_INPUT_DIR' not in globals():\n    print(\"CRITICAL ERROR: KAGGLE_INPUT_DIR not defined. Please run Cell 2 first.\")\nelse:\n    SAMPLE_SUBMISSION_PATH = os.path.join(KAGGLE_INPUT_DIR, 'sample_submission.csv')\n    \n    print(f\"Attempting to load sample_submission.csv from: {SAMPLE_SUBMISSION_PATH}\")\n    if not os.path.exists(SAMPLE_SUBMISSION_PATH):\n        print(f\"CRITICAL ERROR: sample_submission.csv not found at {SAMPLE_SUBMISSION_PATH}\")\n        print(\"       This is essential for knowing which test images to process. Pipeline cannot effectively proceed.\")\n    else:\n        try:\n            sample_submission_df = pd.read_csv(SAMPLE_SUBMISSION_PATH)\n            print(f\"Loaded sample_submission.csv with {len(sample_submission_df)} images to process.\")\n\n            # --- Define the function to get actual test image paths (ROBUST VERSION) ---\n            def get_actual_test_image_path_robust(row):\n                image_filename = str(row['image'])\n                dataset_name = str(row['dataset']) # Still useful for image_id_combined and final submission\n\n                # Common possible locations for test images in Kaggle\n                # KAGGLE_INPUT_DIR is /kaggle/input/image-matching-challenge-2025/\n                possible_base_dirs = [\n                    os.path.join(KAGGLE_INPUT_DIR, 'test_images'), # e.g., test_images/img.png\n                    os.path.join(KAGGLE_INPUT_DIR, 'test'),        # e.g., test/img.png\n                    os.path.join(KAGGLE_INPUT_DIR, 'images'),      # e.g., images/img.png\n                    # Fallback for sample data that might be nested like train data\n                    os.path.join(KAGGLE_INPUT_DIR, 'test', dataset_name) # e.g., test/dataset_X/img.png\n                ]\n                \n                for base_dir_option in possible_base_dirs:\n                    # For flat structures, image is directly in base_dir_option\n                    # For nested (last option), image is in base_dir_option (which is .../test/dataset_name)\n                    # This logic assumes the last option is the only one that would use the dataset_name in path\n                    if base_dir_option.endswith(dataset_name): # Check if it's the nested path attempt\n                         constructed_path = os.path.join(base_dir_option, image_filename)\n                    else: # Flat structure attempts\n                         constructed_path = os.path.join(base_dir_option, image_filename)\n                    \n                    if os.path.exists(constructed_path):\n                        # print(f\"DEBUG: Found image {image_filename} at {constructed_path}\") # Verbose, for debugging only\n                        return constructed_path\n                \n                # If no path found after trying all options\n                # print(f\"Warning: Image {image_filename} (dataset {dataset_name}) not found in common test locations.\")\n                return None\n\n            if not sample_submission_df.empty:\n                print(\"\\nConstructing full paths for test images using robust search...\")\n                sample_submission_df['full_path'] = sample_submission_df.apply(get_actual_test_image_path_robust, axis=1)\n                sample_submission_df['image_id_combined'] = sample_submission_df['dataset'].astype(str) + \"__\" + sample_submission_df['image'].astype(str)\n                \n                test_set_df = sample_submission_df[['dataset', 'image', 'full_path', 'image_id_combined']].copy()\n                print(f\"Prepared 'test_set_df'. First 5 rows:\")\n                print(test_set_df.head())\n\n                num_resolved_paths = test_set_df['full_path'].notna().sum()\n                print(f\"\\nSuccessfully resolved paths for {num_resolved_paths} / {len(test_set_df)} images.\")\n                \n                if num_resolved_paths < len(test_set_df):\n                    print(\"WARNING: Some image paths could not be resolved. These images will be skipped.\")\n                    print(\"Example rows with missing full_path:\")\n                    print(test_set_df[test_set_df['full_path'].isna()].head())\n                \n                if num_resolved_paths > 0 :\n                     ALL_IMAGES_LISTED_IN_SAMPLE_SUB_PREPARED = True\n                else: # No paths resolved at all\n                     print(\"CRITICAL WARNING: No image paths were resolved. Subsequent steps will have no data.\")\n\n            else: # sample_submission_df was empty\n                print(\"sample_submission.csv was empty after loading. Cannot prepare test image paths.\")\n        \n        except Exception as e:\n            print(f\"ERROR during data loading or path preparation in Cell 3: {e}\")\n            # Ensure test_set_df is empty if an error occurs\n            test_set_df = pd.DataFrame(columns=['dataset', 'image', 'full_path', 'image_id_combined'])\n\n\nif not ALL_IMAGES_LISTED_IN_SAMPLE_SUB_PREPARED:\n    print(\"\\nWARNING: Test set preparation was not fully successful. 'test_set_df' might be empty or incomplete.\")\n\nprint(\"-\" * 60)\nprint(\"Cell 3: Data Loading & Test Image Path Preparation complete. Review warnings carefully.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-02T17:11:07.389597Z","iopub.execute_input":"2025-06-02T17:11:07.389829Z","iopub.status.idle":"2025-06-02T17:11:09.141880Z","shell.execute_reply.started":"2025-06-02T17:11:07.389813Z","shell.execute_reply":"2025-06-02T17:11:09.141087Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 4. Global Feature Extraction (DINOv2) - Test Set\n\nThis stage processes the test images identified in the previous step to extract global DINOv2 embeddings.\n\n*   **Initialize Extractor:** An instance of our `DinoV2EmbeddingExtractor` class (from `src/features/global_dino_extractor.py`) is created. We are using the ViT-Small ('s') DINOv2 model by default. The necessary model weights should have been copied to the PyTorch Hub cache in the setup cell for offline access.\n*   **Iterate and Extract:**\n    *   The code loops through each test image path from `test_set_df`.\n    *   For each valid image path:\n        *   The `get_embedding()` method of the extractor is called. This method internally handles loading the image, applying DINOv2-specific transformations (including resizing to the model's required input size, e.g., 224x224, and normalization), and passing it through the DINOv2 model to obtain the global embedding vector (384 dimensions for ViT-Small).\n        *   *(Note: An optional initial down-sizing using our shared `preprocessing.py` could be added here for extremely large raw images before DINOv2's transforms, but for now, we rely on DINOv2's internal transforms to handle input sizing from the raw path.)*\n    *   Embeddings are stored in a dictionary keyed by `image_id_combined` (`dataset__image_filename`).\n*   **Save Embeddings:** The collected embeddings are saved to a compressed NPZ file (`test_dino_embeddings.npz`) in the `/kaggle/working/` directory. This file will be the input for the subsequent clustering stage. If no embeddings are extracted (e.g., due to all image paths being invalid during an interactive run), an empty NPZ file is saved as a placeholder to prevent downstream errors.\n\nDuring interactive runs, if test image paths are not found (due to differences between sample data and the assumed hidden test set structure), warnings will be printed, and the resulting embeddings file might be empty or contain fewer embeddings than expected. This is normal for interactive development; the paths should resolve correctly during actual submission.","metadata":{}},{"cell_type":"code","source":"# Cell 4: Global Feature Extraction (DINOv2) - Test Set\n\n# Ensure necessary variables from previous cells are available\n# KAGGLE_WORKING_DIR, DinoV2EmbeddingExtractor (class), test_set_df, FEATURES_DIR (optional, for output path)\nif 'KAGGLE_WORKING_DIR' not in globals() or \\\n   'DinoV2EmbeddingExtractor' not in globals() or \\\n   'test_set_df' not in globals():\n    \n    print(\"CRITICAL ERROR: Prerequisite variables (KAGGLE_WORKING_DIR, DinoV2EmbeddingExtractor, or test_set_df) not defined.\")\n    print(\"                 Please ensure Cell 1 (Setup), Cell 2 (Imports/Paths), and Cell 3 (Data Loading) have run successfully.\")\n    # Create dummy to prevent crash, but this stage will effectively be skipped / produce no output\n    test_embeddings = {} \n    test_embeddings_npz_path = \"\" # Will cause issues later if not properly defined\n    DINO_EXTRACTION_SUCCESSFUL = False\nelse:\n    DINO_EXTRACTION_SUCCESSFUL = True\n    print(\"=== Stage 4: Initializing DINOv2 Extractor for Test Set ===\")\n    try:\n        # Ensure you are using the desired model size. 's' for ViT-Small.\n        # Model weights should be pre-loaded into cache from Kaggle Dataset in Cell 1.\n        dino_extractor_test = DinoV2EmbeddingExtractor(model_size='s') \n    except Exception as e:\n        print(f\"ERROR: Could not initialize DinoV2EmbeddingExtractor: {e}\")\n        print(\"       Ensure model weights were copied to cache in Cell 1, or check model name.\")\n        dino_extractor_test = None \n        DINO_EXTRACTION_SUCCESSFUL = False\n\n    test_embeddings = {}\n    processed_count = 0\n    error_processing_count = 0 # Renamed from error_count for clarity\n    path_not_found_count = 0   # Renamed from not_found_count\n\n    if dino_extractor_test and not test_set_df.empty:\n        print(f\"\\nExtracting DINOv2 embeddings for {len(test_set_df)} listed test images...\")\n        \n        for idx, row in tqdm(test_set_df.iterrows(), total=len(test_set_df), desc=\"Extracting Test Embeddings\"):\n            full_path = row['full_path']\n            unique_id = row['image_id_combined'] # 'dataset__image'\n\n            if pd.notna(full_path) and os.path.exists(full_path):\n                # --- Optional initial resize for extremely large images ---\n                # try:\n                #     img_pil = load_image_pil(full_path) # Assumes load_image_pil is imported\n                #     if img_pil:\n                #         if max(img_pil.size) > 2500: # Example threshold\n                #             # Assumes resize_image_maintain_aspect_ratio is imported\n                #             img_pil = resize_image_maintain_aspect_ratio(img_pil, target_max_dimension=1280) \n                #         # How to pass PIL image to get_embedding? Modify get_embedding or save temp file.\n                #         # For now, simplified: pass full_path. DINOv2's transform will handle typical inputs.\n                #         embedding = dino_extractor_test.get_embedding(full_path) # If get_embedding takes path\n                #         # Or if get_embedding was modified to take PIL: embedding = dino_extractor_test.get_embedding(img_pil)\n                # except Exception as e_preprocess:\n                #     print(f\"Error during pre-resize for {unique_id}: {e_preprocess}\")\n                #     embedding = None # Fallback or skip\n                # --- End Optional initial resize ---\n                \n                # Current approach: pass full_path, let DINOv2's transform handle sizing\n                embedding = dino_extractor_test.get_embedding(full_path) \n                \n                if embedding is not None:\n                    test_embeddings[unique_id] = embedding\n                    processed_count += 1\n                else:\n                    # Error typically printed by get_embedding, just count it\n                    error_processing_count += 1\n            else:\n                # This print can be verbose. Keep it commented for submission unless debugging.\n                # print(f\"Warning: Test image path not found or invalid: {full_path} for ID {unique_id}\")\n                path_not_found_count += 1\n        \n        print(f\"\\n--- Test Embedding Extraction Summary ---\")\n        print(f\"Successfully extracted embeddings for: {processed_count} images.\")\n        if path_not_found_count > 0:\n            print(f\"Paths not found or invalid for:     {path_not_found_count} images (expected during interactive runs with sample_submission).\")\n        if error_processing_count > 0:\n            print(f\"Errors during embedding extraction for: {error_processing_count} images.\")\n\n    elif test_set_df.empty:\n        print(\"test_set_df is empty. No images to process for feature extraction.\")\n        DINO_EXTRACTION_SUCCESSFUL = False\n    else: # dino_extractor_test is None\n        print(\"DINOv2 Extractor not initialized. Skipping feature extraction.\")\n        DINO_EXTRACTION_SUCCESSFUL = False\n\n    # Define output path for test embeddings using FEATURES_DIR from Cell 2\n    if 'FEATURES_DIR' in globals():\n        test_embeddings_npz_path = os.path.join(FEATURES_DIR, 'test_dino_embeddings_vits.npz') # Added model size\n    else:\n        print(\"WARNING: FEATURES_DIR not defined from Cell 2. Saving NPZ to KAGGLE_WORKING_DIR.\")\n        test_embeddings_npz_path = os.path.join(KAGGLE_WORKING_DIR, 'test_dino_embeddings_vits.npz')\n\n\n    if test_embeddings: \n        print(f\"\\nSaving {len(test_embeddings)} test embeddings to {test_embeddings_npz_path}...\")\n        np.savez_compressed(test_embeddings_npz_path, **test_embeddings)\n        print(\"Test embeddings saved successfully.\")\n    else:\n        print(\"No test embeddings were extracted (or an error occurred). NPZ file not saved (or will be empty).\")\n        # Ensure an empty NPZ exists if subsequent cells expect it\n        if not os.path.exists(test_embeddings_npz_path):\n             try:\n                 np.savez_compressed(test_embeddings_npz_path) # Save an empty NPZ\n                 print(f\"Saved an empty NPZ file at {test_embeddings_npz_path} as a placeholder.\")\n             except Exception as e_save_empty:\n                 print(f\"Could not save empty NPZ: {e_save_empty}\")\n                 DINO_EXTRACTION_SUCCESSFUL = False # Mark as failed if can't even save empty\n\n# Final check\nif not DINO_EXTRACTION_SUCCESSFUL and not test_embeddings : # If it failed AND no embeddings\n    print(\"\\nWARNING: DINOv2 feature extraction was not successful or produced no embeddings.\")\n    print(\"         Subsequent clustering steps may fail or produce no results.\")\n    # Define test_embeddings_npz_path as None or empty string if it truly failed and no placeholder was made.\n    # This helps downstream cells check if they should even attempt to load it.\n    if not os.path.exists(test_embeddings_npz_path):\n        test_embeddings_npz_path = None \n\nprint(\"-\" * 60)\nprint(\"Cell 4: Global Feature Extraction (DINOv2) for Test Set complete.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-02T17:11:09.143663Z","iopub.execute_input":"2025-06-02T17:11:09.144400Z","iopub.status.idle":"2025-06-02T17:11:41.237074Z","shell.execute_reply.started":"2025-06-02T17:11:09.144380Z","shell.execute_reply":"2025-06-02T17:11:41.236513Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 5. Clustering (HDBSCAN) - Test Set\n\nWith the global DINOv2 embeddings extracted for the test images, this stage performs clustering to group these images into potential scenes or identify them as outliers.\n\n*   **Load Embeddings:** The `test_dino_embeddings.npz` file (generated in the previous step and saved in `/kaggle/working/`) is loaded. This file contains a dictionary mapping `image_id_combined` to its DINOv2 embedding vector.\n*   **Clustering Algorithm:** We utilize our custom `run_hdbscan_clustering` function from `src/clustering/hdbscan_clusterer.py`.\n    *   This function can optionally apply UMAP for dimensionality reduction (currently configured with `n_components=30` and `cosine` metric, based on earlier experiments) before running HDBSCAN.\n    *   HDBSCAN is then applied with parameters tuned during development (e.g., `min_cluster_size=5`, `metric='euclidean'` on UMAP output).\n*   **Output:**\n    *   The `run_hdbscan_clustering` function returns an array of cluster labels for each input embedding.\n    *   A pandas DataFrame (`test_clustering_results_df`) is created, mapping each `image_id_combined` to its `predicted_scene_label_raw` (the numerical output from HDBSCAN, where -1 typically indicates noise/outliers).\n    *   A summary of the found clusters and noise points is printed.\n\nIf the previous embedding extraction step resulted in no embeddings (e.g., due to test images not being found in an interactive run), this clustering step will be skipped or will process an empty set. The resulting `test_clustering_results_df` will then be used to populate the 'scene' column in our final `submission.csv`.","metadata":{}},{"cell_type":"code","source":"# Cell 5: Clustering (HDBSCAN) - Test Set\n\nprint(\"=== Stage 5: HDBSCAN Clustering on Test Set Embeddings ===\")\n\n# Ensure necessary variables and functions from previous cells/imports are available\n# KAGGLE_WORKING_DIR, load_embeddings (func), run_hdbscan_clustering (func), pd (module), os (module)\n# test_embeddings_npz_path (from previous cell if DINO extraction was successful)\n# FEATURES_DIR (from cell 2, if using consistent output paths)\n\n# Initialize to a known state\ntest_clustering_results_df = pd.DataFrame(columns=['image_id_combined', 'predicted_scene_label_raw']) \nCLUSTERING_SUCCESSFUL = False \n\n# Check prerequisites\nprereqs_met = True\nif 'KAGGLE_WORKING_DIR' not in globals():\n    print(\"ERROR: KAGGLE_WORKING_DIR not defined.\")\n    prereqs_met = False\nif 'load_embeddings' not in globals(): # Function from your src/clustering/hdbscan_clusterer.py\n    print(\"ERROR: load_embeddings function not imported/defined.\")\n    prereqs_met = False\nif 'run_hdbscan_clustering' not in globals(): # Function from your src/clustering/hdbscan_clusterer.py\n    print(\"ERROR: run_hdbscan_clustering function not imported/defined.\")\n    prereqs_met = False\nif 'pd' not in globals(): # Should be imported in Cell 2\n    print(\"ERROR: pandas (pd) not imported.\")\n    prereqs_met = False\nif 'os' not in globals(): # Should be imported in Cell 2\n    print(\"ERROR: os module not imported.\")\n    prereqs_met = False\n# Check if the embeddings NPZ path was set by the previous cell\nif 'test_embeddings_npz_path' not in globals() or not test_embeddings_npz_path:\n    print(\"ERROR: test_embeddings_npz_path not defined from previous DINOv2 extraction cell.\")\n    prereqs_met = False\n\n\nif not prereqs_met:\n    print(\"       Prerequisite variables/functions missing. Please run previous cells, especially Cell 2 (Imports/Paths) and Cell 4 (DINOv2 Extraction).\")\n    print(\"       Skipping clustering.\")\nelse:\n    # --- Define Parameters for Clustering (Based on your best experimental results) ---\n    # These should match the parameters you found effective during your local/Colab experimentation.\n    print(\"\\nDefining clustering parameters...\")\n    # UMAP Parameters (if use_umap=True in run_hdbscan_clustering function)\n    UMAP_N_NEIGHBORS = 15\n    UMAP_N_COMPONENTS = 30 \n    UMAP_MIN_DIST = 0.0\n    UMAP_METRIC = 'cosine'    \n\n    # HDBSCAN Parameters\n    HDBSCAN_MIN_CLUSTER_SIZE = 5 # This is a key parameter to tune\n    HDBSCAN_METRIC = 'euclidean' \n    HDBSCAN_MIN_SAMPLES = None   \n    # HDBSCAN_CLUSTER_SELECTION_EPSILON = 0.0 # Optional advanced param\n\n    # Path to the embeddings file generated in the previous step\n    # test_embeddings_npz_path was defined in Cell 4\n    \n    print(f\"Attempting to run HDBSCAN clustering on embeddings from: {test_embeddings_npz_path}\")\n\n    if os.path.exists(test_embeddings_npz_path):\n        test_image_ids_clust, test_embeddings_matrix_clust = load_embeddings(test_embeddings_npz_path)\n        \n        if test_image_ids_clust is not None and test_embeddings_matrix_clust is not None and test_embeddings_matrix_clust.size > 0:\n            print(f\"Successfully loaded {len(test_image_ids_clust)} embeddings for clustering (Shape: {test_embeddings_matrix_clust.shape}).\")\n            \n            test_cluster_labels = run_hdbscan_clustering(\n                test_embeddings_matrix_clust,\n                use_umap=True, # Your script default, confirm this is intended\n                umap_n_neighbors=UMAP_N_NEIGHBORS, \n                umap_n_components=UMAP_N_COMPONENTS, \n                umap_min_dist=UMAP_MIN_DIST, \n                umap_metric=UMAP_METRIC,\n                hdbscan_min_cluster_size=HDBSCAN_MIN_CLUSTER_SIZE,\n                hdbscan_metric=HDBSCAN_METRIC,\n                hdbscan_min_samples=HDBSCAN_MIN_SAMPLES\n            )\n            \n            if test_cluster_labels is not None:\n                if len(test_image_ids_clust) == len(test_cluster_labels):\n                    test_clustering_results_df = pd.DataFrame({\n                        'image_id_combined': test_image_ids_clust, \n                        'predicted_scene_label_raw': test_cluster_labels\n                    })\n                    print(\"\\nClustering complete for test set.\")\n                    print(\"First 5 rows of clustering results:\")\n                    print(test_clustering_results_df.head())\n                    print(\"\\nCluster label counts (HDBSCAN output; -1 indicates noise/outliers):\")\n                    if not test_clustering_results_df.empty:\n                        print(test_clustering_results_df['predicted_scene_label_raw'].value_counts().sort_index())\n                    else:\n                        print(\"Clustering results DataFrame is empty.\")\n                    CLUSTERING_SUCCESSFUL = True\n                else:\n                    print(f\"ERROR: Mismatch in length of image_ids ({len(test_image_ids_clust)}) and cluster_labels ({len(test_cluster_labels)}).\")\n            else:\n                print(\"HDBSCAN clustering function returned None (likely failed internally) on test set.\")\n        else:\n            print(\"Failed to load test embeddings or embeddings file was effectively empty. Skipping clustering.\")\n            if test_embeddings_matrix_clust is not None and test_embeddings_matrix_clust.size == 0:\n                 print(\"   Reason: Embeddings matrix from NPZ was empty (e.g., no test images found or processed in previous DINOv2 step).\")\n    else:\n        print(f\"Test embeddings NPZ file not found at {test_embeddings_npz_path}. Skipping clustering.\")\n\n# Ensure test_clustering_results_df is defined even if clustering fails, for downstream cells\n# This was already good, just re-confirming its placement.\nif 'test_clustering_results_df' not in locals(): # Should have been defined as empty at the top if prereqs failed\n    print(\"Defining test_clustering_results_df as empty due to critical earlier errors in this cell.\")\n    test_clustering_results_df = pd.DataFrame(columns=['image_id_combined', 'predicted_scene_label_raw'])\n    CLUSTERING_SUCCESSFUL = False # Redundant if already false, but safe.\n\n\nprint(\"-\" * 60)\nprint(\"Cell 5: Clustering (HDBSCAN) for Test Set complete.\")\nif CLUSTERING_SUCCESSFUL:\n    print(\"Clustering was successful and results are in 'test_clustering_results_df'.\")\nelse:\n    print(\"Clustering was NOT successful or was skipped. 'test_clustering_results_df' will be empty or reflect no clusters.\")\n    print(\"Subsequent pipeline steps (Pair Selection, SfM) might not have meaningful input.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-02T17:11:41.237935Z","iopub.execute_input":"2025-06-02T17:11:41.238204Z","iopub.status.idle":"2025-06-02T17:11:41.251301Z","shell.execute_reply.started":"2025-06-02T17:11:41.238178Z","shell.execute_reply":"2025-06-02T17:11:41.250552Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 6. Image Pair Selection (within Clusters using DINOv2 Embeddings)\n\nAfter assigning test images to clusters (or labeling them as noise/outliers), this stage focuses on intelligently selecting which pairs of images *within each valid cluster* should be processed by the more computationally expensive local feature matcher (ALIKED+LightGlue). Simply matching all possible pairs within large clusters can be too slow.\n\n*   **Input:**\n    *   `test_clustering_results_df`: DataFrame containing `image_id_combined` and their `predicted_scene_label_raw` from the HDBSCAN stage.\n    *   `test_dino_embeddings.npz`: The file containing DINOv2 embeddings for all test images (used to calculate similarity for pair selection).\n*   **Process:**\n    1.  **Load All Test Embeddings:** The full dictionary of test image embeddings is loaded from the NPZ file.\n    2.  **Iterate Through Clusters:** The code loops through each dataset and then each valid cluster (i.e., not noise points labeled -1) identified in the previous stage.\n    3.  **Call `select_pairs_by_embedding_similarity`:** For each cluster, our custom function from `src/matching_strategies/pair_selector.py` is used. This function:\n        *   Takes the list of `image_id_combined`s belonging to the current cluster.\n        *   Takes the dictionary of all test image embeddings.\n        *   Calculates pairwise cosine similarity between the DINOv2 embeddings of images *within the current cluster*.\n        *   Selects pairs based on a strategy (e.g., top-K most similar neighbors for each image, and/or a minimum similarity threshold).\n    4.  **Store Selected Pairs:** The selected pairs for each cluster are stored, likely in a dictionary or list structure, for the next stage.\n*   **Output:** A data structure (e.g., a dictionary mapping `cluster_id` to a list of `(image_id1, image_id2)` tuples) containing the filtered pairs that will be passed to the local feature matching module.\n\nThis step aims to significantly reduce the number of image pairs that need detailed local matching, optimizing runtime while focusing on pairs most likely to yield good matches for 3D reconstruction. If a cluster has too few images or if the pair selector finds no suitable pairs (e.g., due to low similarity), strategies like attempting all-pairs for very small clusters or skipping matching for that cluster are considered.","metadata":{}},{"cell_type":"code","source":"# Cell 6: Image Pair Selection (within Clusters using DINOv2 Embeddings) - Test Set\n\nprint(\"=== Stage 6: Image Pair Selection within Clusters (Test Set) ===\")\n\n# Ensure necessary variables and functions are available\n# KAGGLE_WORKING_DIR, test_clustering_results_df, select_pairs_by_embedding_similarity (func),\n# pd, os, np modules should be defined/imported from previous cells.\n# FEATURES_DIR (from cell 2, if using consistent output paths for NPZ)\n\n# Initialize to a known state\nall_selected_pairs_per_cluster_scene = {} \nPAIR_SELECTION_SUCCESSFUL = False\n\n# Check prerequisites\nprereqs_met_cell6 = True\nif 'KAGGLE_WORKING_DIR' not in globals(): print(\"ERROR: KAGGLE_WORKING_DIR not defined.\"); prereqs_met_cell6 = False\nif 'test_clustering_results_df' not in globals(): print(\"ERROR: test_clustering_results_df not defined.\"); prereqs_met_cell6 = False\nif 'select_pairs_by_embedding_similarity' not in globals(): print(\"ERROR: select_pairs_by_embedding_similarity function not imported.\"); prereqs_met_cell6 = False\nif 'pd' not in globals(): print(\"ERROR: pandas (pd) not imported.\"); prereqs_met_cell6 = False\nif 'os' not in globals(): print(\"ERROR: os module not imported.\"); prereqs_met_cell6 = False\nif 'np' not in globals(): print(\"ERROR: numpy (np) not imported.\"); prereqs_met_cell6 = False\nif 'FEATURES_DIR' not in globals(): # Assuming FEATURES_DIR is where test_dino_embeddings.npz is\n    print(\"WARNING: FEATURES_DIR not defined from Cell 2. Will try KAGGLE_WORKING_DIR for embeddings NPZ.\")\n    # Fallback if FEATURES_DIR isn't globally set from Cell 2, though it should be.\n    # This assumes test_embeddings_npz_path was defined in Cell 4 relative to KAGGLE_WORKING_DIR if FEATURES_DIR was missing.\n    if 'test_embeddings_npz_path' not in globals(): # If even that is missing\n        print(\"ERROR: Path to embeddings NPZ also not found.\")\n        prereqs_met_cell6 = False\n\nif not prereqs_met_cell6:\n    print(\"       Prerequisite variables/functions missing for Cell 6. Please run previous cells.\")\n    print(\"       Skipping pair selection.\")\nelse:\n    print(\"\\nStarting Image Pair Selection within clusters...\")\n    \n    # --- Load the full DINOv2 test embeddings ---\n    # These were generated and saved in Cell 4 (Global Feature Extraction)\n    # Use FEATURES_DIR if defined in Cell 2, otherwise use KAGGLE_WORKING_DIR as a fallback\n    # (This logic assumes test_embeddings_npz_path was correctly defined in Cell 4)\n    if 'test_embeddings_npz_path' not in globals(): # Should have been defined in Cell 4\n         npz_dir_base = FEATURES_DIR if 'FEATURES_DIR' in globals() else KAGGLE_WORKING_DIR\n         test_embeddings_npz_path = os.path.join(npz_dir_base, 'test_dino_embeddings_vits.npz') # Reconstruct if needed\n\n    all_test_embeddings_dict = {} \n\n    print(f\"Attempting to load DINOv2 embeddings from: {test_embeddings_npz_path}\")\n    if os.path.exists(test_embeddings_npz_path):\n        try:\n            loaded_embeddings_data = np.load(test_embeddings_npz_path, allow_pickle=False) # allow_pickle=False for security if not needed\n            if len(loaded_embeddings_data.files) == 0: # Check for empty NPZ\n                print(\"Warning: Embeddings NPZ file is empty (no arrays found).\")\n            else:\n                for key in loaded_embeddings_data.files:\n                    all_test_embeddings_dict[key] = loaded_embeddings_data[key]\n                print(f\"Successfully loaded {len(all_test_embeddings_dict)} DINOv2 embeddings for pair selection.\")\n            loaded_embeddings_data.close()\n        except Exception as e:\n            print(f\"Error loading embeddings NPZ {test_embeddings_npz_path}: {e}\")\n    else:\n        print(f\"ERROR: Test embeddings NPZ file not found at {test_embeddings_npz_path}. Cannot perform pair selection.\")\n    \n    # Fallback for small clusters if pair selector is too restrictive\n    MAX_IMAGES_FOR_ALL_PAIRS_FALLBACK = 10 \n    PAIR_SELECTOR_TOP_K = 10 \n    PAIR_SELECTOR_SIM_THRESHOLD = 0.7 \n\n    if not test_clustering_results_df.empty and all_test_embeddings_dict:\n        if 'dataset_name' not in test_clustering_results_df.columns:\n            print(\"Adding 'dataset_name' column to test_clustering_results_df for processing.\")\n            test_clustering_results_df['dataset_name'] = test_clustering_results_df['image_id_combined'].apply(lambda x: x.split('__')[0])\n\n        for dataset_name_iter in test_clustering_results_df['dataset_name'].unique():\n            print(f\"\\nSelecting pairs for dataset: {dataset_name_iter}\")\n            dataset_clusters_df_iter = test_clustering_results_df[test_clustering_results_df['dataset_name'] == dataset_name_iter]\n            \n            for cluster_label_iter in sorted(dataset_clusters_df_iter['predicted_scene_label_raw'].unique()): # Sorted for consistent processing order\n                if cluster_label_iter == -1: \n                    continue # Skip noise points for pair selection\n\n                scene_key = f\"{dataset_name_iter}__{cluster_label_iter}\" \n                print(f\"  Processing cluster {cluster_label_iter} (scene_key: {scene_key}) for pair selection...\")\n                \n                image_ids_in_cluster_list = dataset_clusters_df_iter[\n                    dataset_clusters_df_iter['predicted_scene_label_raw'] == cluster_label_iter\n                ]['image_id_combined'].tolist()\n                \n                if len(image_ids_in_cluster_list) < 2:\n                    print(f\"    Cluster {cluster_label_iter} has < 2 images. No pairs to select.\")\n                    all_selected_pairs_per_cluster_scene[scene_key] = []\n                    continue\n                \n                selected_pairs = select_pairs_by_embedding_similarity(\n                    image_ids_in_cluster_list, \n                    all_test_embeddings_dict, \n                    top_k=PAIR_SELECTOR_TOP_K, \n                    similarity_threshold=PAIR_SELECTOR_SIM_THRESHOLD\n                )\n                \n                if not selected_pairs and len(image_ids_in_cluster_list) <= MAX_IMAGES_FOR_ALL_PAIRS_FALLBACK:\n                    print(f\"    Pair selector found 0 pairs for cluster {cluster_label_iter} (size {len(image_ids_in_cluster_list)}).\")\n                    print(f\"    Attempting all-pairs fallback as cluster size <= {MAX_IMAGES_FOR_ALL_PAIRS_FALLBACK}.\")\n                    from itertools import combinations # Import here as it's only for fallback\n                    all_possible_pairs = [tuple(sorted(p)) for p in combinations(image_ids_in_cluster_list, 2)]\n                    selected_pairs = list(set(all_possible_pairs)) \n                    print(f\"    Using {len(selected_pairs)} pairs from all-pairs fallback.\")\n                elif not selected_pairs:\n                    print(f\"    Pair selector found 0 pairs for cluster {cluster_label_iter} (size {len(image_ids_in_cluster_list)}), and cluster too large for fallback. No pairs selected.\")\n\n                all_selected_pairs_per_cluster_scene[scene_key] = selected_pairs\n                print(f\"    Selected {len(selected_pairs)} pairs for cluster {cluster_label_iter}.\")\n        \n        if all_selected_pairs_per_cluster_scene: # Check if any pairs were actually selected across all clusters\n            PAIR_SELECTION_SUCCESSFUL = True\n        print(\"\\nPair selection process complete for all clusters.\")\n        \n    elif test_clustering_results_df.empty:\n        print(\"Clustering results (test_clustering_results_df) are empty. Skipping pair selection.\")\n    elif not all_test_embeddings_dict: # Check if dict is empty, implying no embeddings loaded\n        print(\"DINOv2 test embeddings dictionary is empty. Skipping pair selection.\")\n    else: # Should not be reached if above conditions are met\n        print(\"Unknown state, skipping pair selection.\")\n\n\n# Ensure all_selected_pairs_per_cluster_scene is defined for downstream cells\nif 'all_selected_pairs_per_cluster_scene' not in locals():\n    print(\"Defining all_selected_pairs_per_cluster_scene as empty due to critical earlier errors in this cell.\")\n    all_selected_pairs_per_cluster_scene = {}\n    # PAIR_SELECTION_SUCCESSFUL should already be False if this path is taken\n\n# Example: Print out some selected pairs\nif PAIR_SELECTION_SUCCESSFUL and all_selected_pairs_per_cluster_scene:\n    print(\"\\n--- Example of Selected Pairs (first few scenes with pairs) ---\")\n    inspected_count = 0\n    for scene_key_example, pairs_list_example in all_selected_pairs_per_cluster_scene.items():\n        if inspected_count < 3 and pairs_list_example: \n            print(f\"Scene Key: {scene_key_example}, Number of pairs: {len(pairs_list_example)}\")\n            print(f\"  First up to 3 pairs: {pairs_list_example[:min(3, len(pairs_list_example))]}\")\n            inspected_count += 1\n        elif inspected_count >=3:\n            break\n    if inspected_count == 0:\n        print(\"No scenes with selected pairs to show as example (all pair lists might be empty).\")\nelif not all_selected_pairs_per_cluster_scene: # If dict is empty\n     print(\"No pairs were selected across any clusters, or pair selection was skipped.\")\n\n\nprint(\"-\" * 60)\nprint(\"Cell 6: Image Pair Selection complete.\")\nif PAIR_SELECTION_SUCCESSFUL and any(all_selected_pairs_per_cluster_scene.values()): # Check if any list of pairs is non-empty\n    print(\"Pair selection ran and found pairs for at least one cluster.\")\nelse:\n    print(\"Pair selection was NOT successful or found NO pairs for any cluster. Subsequent local matching will have no input.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-02T17:11:41.252295Z","iopub.execute_input":"2025-06-02T17:11:41.252636Z","iopub.status.idle":"2025-06-02T17:11:41.273603Z","shell.execute_reply.started":"2025-06-02T17:11:41.252609Z","shell.execute_reply":"2025-06-02T17:11:41.272986Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 7. Local Feature Matching & Structure from Motion (SfM) - Placeholder\n\nThis section is where Davin's comprehensive module for local feature matching (ALIKED+LightGlue with Rotation TTA) and SfM (using COLMAP via his `database.py` and `h5_to_db.py` logic, and `pycolmap`) will be integrated.\n\n*   **Input to Davin's Module (per scene cluster):**\n    *   A list of all `image_id_combined`s belonging to the current cluster.\n    *   A list of selected `(image_id1_combined, image_id2_combined)` pairs for that cluster (from our `pair_selector` stage).\n    *   Access to the shared `preprocessing.py` utility for image resizing.\n    *   (Internally, his module will also need paths to ALIKED/LightGlue model weights).\n*   **Davin's Module Process (Conceptual):**\n    1.  Initialize ALIKED & LightGlue.\n    2.  For each unique image involved in the selected pairs for the current cluster:\n        *   Load and preprocess the image (using `preprocessing.py` to resize to target ALIKED dimension, e.g., 1024px).\n        *   Extract ALIKED keypoints and descriptors.\n        *   Cache/Store these features (e.g., in a per-cluster `keypoints.h5`).\n    3.  For each selected `(image_id1, image_id2)` pair:\n        *   Load their features.\n        *   Perform LightGlue matching, incorporating **Rotation TTA** (testing 0, 90, 180, 270 deg rotations for one image).\n        *   Perform geometric verification (RANSAC).\n        *   Cache/Store verified matches (e.g., in a per-cluster `matches.h5`).\n    4.  Use the `h5_to_db.py` logic to populate a COLMAP database (`cluster_X_colmap.db`) for the current cluster using the extracted keypoints and matches.\n    5.  Run `pycolmap.incremental_mapping` (or other COLMAP commands) using this database and the cluster's images.\n    6.  Parse the COLMAP reconstruction output.\n*   **Output from Davin's Module (per scene cluster):**\n    *   A dictionary mapping each successfully registered `image_id_combined` within that cluster to its estimated `rotation_matrix` (as a 3x3 NumPy array) and `translation_vector` (as a 1x3 NumPy array).\n    *   For images that couldn't be registered or for which an error occurred, it should indicate failure (e.g., return `None` or pre-filled NaN arrays for R, T).\n\n**Current Status in this Notebook:**\nSince Davin's refactored `.py` script is still under development, this cell will currently act as a **placeholder**. It will simulate this stage by assigning `NaN` (Not a Number) values for all rotation matrices and translation vectors. The `scene` labels from our clustering stage will still be used.\n\nOnce Davin's script (`src/reconstruction/scene_reconstructor.py` or similar) is ready and merged into our GitHub `main` branch, we will replace the placeholder logic in this cell with actual calls to his module.","metadata":{}},{"cell_type":"code","source":"# Cell 7: Local Feature Matching & Structure from Motion (SfM)\n\nprint(\"=== Stage 7: Local Feature Matching & SfM ===\")\n\n# Initialize to a known state\nall_image_final_poses = {} \nSFM_STAGE_ATTEMPTED = False # Flag to indicate if we tried to run the actual SfM\nSFM_OVERALL_SUCCESS = False # Flag if at least one cluster produced poses\n\n# Ensure necessary variables from previous cells are available\nprereqs_met_cell7 = True\nif 'all_selected_pairs_per_cluster_scene' not in globals(): \n    print(\"ERROR: 'all_selected_pairs_per_cluster_scene' not defined from Cell 6 (Pair Selection).\")\n    prereqs_met_cell7 = False\nif 'test_clustering_results_df' not in globals() or test_clustering_results_df.empty: \n    print(\"ERROR: 'test_clustering_results_df' not defined or empty from Cell 5 (Clustering).\")\n    prereqs_met_cell7 = False\nif 'test_set_df' not in globals() or test_set_df.empty: # Needed for path_lookup_for_sfm\n    print(\"ERROR: 'test_set_df' not defined or empty from Cell 3 (Data Loading).\")\n    prereqs_met_cell7 = False\n# KAGGLE_WORKING_DIR, ALIKED_WEIGHT_KAGGLE_PATH, LIGHTGLUE_WEIGHT_KAGGLE_PATH should be from Cell 1\nif 'KAGGLE_WORKING_DIR' not in globals(): print(\"ERROR: KAGGLE_WORKING_DIR missing.\"); prereqs_met_cell7 = False\nif 'ALIKED_WEIGHT_KAGGLE_PATH' not in globals(): print(\"ERROR: ALIKED_WEIGHT_KAGGLE_PATH missing.\"); prereqs_met_cell7 = False\n# LIGHTGLUE_WEIGHT_KAGGLE_PATH might be optional if LightGlue loads with ALIKED features\n\n# --- Try to import Davin's module ---\nDAVIN_MODULE_READY = False\nreconstruct_scene_cluster_func = None # Placeholder for Davin's function\nPREPROCESSING_MODULE = None\n\nif prereqs_met_cell7:\n    try:\n        from sfm.scene_reconstructor import reconstruct_scene_cluster # This is the target function\n        from data import preprocessing as team_preprocessing_module # Our shared preprocessing\n        \n        reconstruct_scene_cluster_func = reconstruct_scene_cluster\n        PREPROCESSING_MODULE = team_preprocessing_module\n        print(\"Successfully imported 'reconstruct_scene_cluster' from 'sfm.scene_reconstructor' and 'preprocessing' module.\")\n        DAVIN_MODULE_READY = True\n    except ImportError as e:\n        print(f\"WARNING: Could not import Davin's SfM module (sfm.scene_reconstructor): {e}\")\n        print(\"         Ensure the PR for this module is merged to the cloned branch and it defines 'reconstruct_scene_cluster'.\")\n        print(\"         Proceeding with placeholder (NaN) poses for SfM.\")\n    except Exception as e_import: # Catch other potential import-related errors\n        print(f\"An unexpected error occurred during SfM module import: {e_import}\")\n        print(\"         Proceeding with placeholder (NaN) poses for SfM.\")\n\n\nif not prereqs_met_cell7:\n    print(\"Prerequisites missing for Cell 7. SfM stage will be fully skipped, assigning NaN poses by default if possible.\")\n    # Attempt to fill all_image_final_poses with NaNs if test_set_df exists from sample_submission\n    if 'test_set_df' in globals() and not test_set_df.empty:\n        print(\"Populating all poses with NaNs due to prerequisite failure...\")\n        for idx, row in test_set_df.iterrows():\n            img_id_comb = row['image_id_combined']\n            # Try to get cluster label if available, else default to outliers\n            scene_lbl = \"outliers\"\n            if 'test_clustering_results_df' in globals() and not test_clustering_results_df.empty and img_id_comb in test_clustering_results_df['image_id_combined'].values:\n                raw_lbl = test_clustering_results_df.loc[test_clustering_results_df['image_id_combined'] == img_id_comb, 'predicted_scene_label_raw'].iloc[0]\n                scene_lbl = f\"cluster{int(raw_lbl)}\" if raw_lbl != -1 else \"outliers\"\n            \n            all_image_final_poses[img_id_comb] = {\n                'R_arr': np.full((3,3), np.nan), 'T_arr': np.full((3,), np.nan),\n                'scene_label_final': scene_lbl, 'registered': False}\n    SFM_STAGE_ATTEMPTED = False # Did not even attempt Davin's module\nelse:\n    print(\"\\nStarting Local Feature Matching & SfM Stage...\")\n    \n    # Create path_lookup_for_sfm from test_set_df (has 'image_id_combined' and 'full_path')\n    path_lookup_for_sfm = pd.Series(test_set_df.full_path.values, index=test_set_df.image_id_combined).to_dict()\n\n    # Define COLMAP options (Raman's input, or sensible defaults)\n    colmap_options = {\n        \"min_model_size\": 3,      # Min images to make a model\n        \"max_num_models\": 1,      # Try to get one best model\n        \"ba_global_max_num_iterations\": 50, # Default 100, reduce for speed if needed\n        # Add other pycolmap.IncrementalPipelineOptions based on Raman's research\n    }\n    \n    # --- ALIKED/LightGlue models - These should be initialized ONCE in the notebook if passed as objects ---\n    # If Davin's script initializes them internally using paths, this block is not needed here.\n    # For now, assume Davin's script takes paths to weights and initializes them.\n    # ALIKED_WEIGHT_KAGGLE_PATH and LIGHTGLUE_WEIGHT_KAGGLE_PATH are defined in Cell 1.\n\n    processed_sfm_clusters_count = 0\n    for scene_key, selected_pairs_for_scene_ids in tqdm(all_selected_pairs_per_cluster_scene.items(), desc=\"Processing Clusters for SfM\"):\n        dataset_name_sfm, cluster_label_raw_str_sfm = scene_key.split('__')\n        cluster_label_raw_sfm = int(cluster_label_raw_str_sfm)\n        \n        # This loop iterates over keys from all_selected_pairs_per_cluster_scene,\n        # which should NOT contain the noise cluster (-1) if pair_selector skips it.\n        # If pair_selector could potentially output for -1, then add: if cluster_label_raw_sfm == -1: continue\n        \n        final_scene_label = f\"cluster{cluster_label_raw_sfm}\" # e.g., \"cluster0\"\n        \n        current_cluster_image_ids_combined = test_clustering_results_df[\n            (test_clustering_results_df['dataset_name'] == dataset_name_sfm) &\n            (test_clustering_results_df['predicted_scene_label_raw'] == cluster_label_raw_sfm)\n        ]['image_id_combined'].tolist()\n\n        print(f\"  Processing {final_scene_label} (Dataset: {dataset_name_sfm}) with {len(current_cluster_image_ids_combined)} images, {len(selected_pairs_for_scene_ids)} selected pairs...\")\n\n        if not current_cluster_image_ids_combined or len(current_cluster_image_ids_combined) < 2:\n            print(f\"    Skipping SfM for {final_scene_label}: Not enough images in cluster ({len(current_cluster_image_ids_combined)}). Assigning NaN poses.\")\n            for img_id in current_cluster_image_ids_combined:\n                all_image_final_poses[img_id] = {'R_arr': np.full((3,3), np.nan), 'T_arr': np.full((3,), np.nan), \n                                                   'scene_label_final': final_scene_label, 'registered': False}\n            continue\n        \n        if not selected_pairs_for_scene_ids: # No pairs from selector (even after fallback if implemented there)\n            print(f\"    No pairs selected by pair_selector for {final_scene_label}. Assigning NaN poses.\")\n            for img_id in current_cluster_image_ids_combined:\n                all_image_final_poses[img_id] = {'R_arr': np.full((3,3), np.nan), 'T_arr': np.full((3,), np.nan),\n                                                   'scene_label_final': final_scene_label, 'registered': False}\n            continue\n\n        SFM_STAGE_ATTEMPTED = True # We are attempting SfM for at least one cluster\n\n        if DAVIN_MODULE_READY and reconstruct_scene_cluster_func is not None:\n            # Define a unique output directory for this cluster's SfM temporary files\n            cluster_sfm_temp_dir = os.path.join(KAGGLE_WORKING_DIR, \"sfm_temp\", scene_key)\n            # shutil.rmtree(cluster_sfm_temp_dir, ignore_errors=True) # Clean up previous run for this cluster\n            # os.makedirs(cluster_sfm_temp_dir, exist_ok=True)\n\n            print(f\"    Calling Davin's SfM Module for {final_scene_label}...\")\n            try:\n                # Call the imported function from Davin's refactored script\n                cluster_poses_from_sfm = reconstruct_scene_cluster_func(\n                    image_ids_in_cluster=current_cluster_image_ids_combined,\n                    all_image_paths_lookup=path_lookup_for_sfm,\n                    selected_image_pairs_ids=selected_pairs_for_scene_ids,\n                    preprocessing_module=PREPROCESSING_MODULE, # Pass the imported module\n                    aliked_weights_path=ALIKED_WEIGHT_KAGGLE_PATH, # Pass path to weights\n                    lightglue_weights_path=LIGHTGLUE_WEIGHT_KAGGLE_PATH, # Pass path to weights\n                    base_output_dir_for_sfm_run=cluster_sfm_temp_dir, # Temp dir for this cluster's SfM\n                    target_aliked_input_size=1024, # This should be an agreed-upon hyperparameter\n                    colmap_mapper_options_dict=colmap_options,\n                    rotation_tta=True # Enable Rotation TTA by default, make it a param if needed\n                )\n                # cluster_poses_from_sfm is expected to be {'image_basename.png': {'R': R_array, 'T': T_array, 'registered': True/False}}\n\n                # Merge results\n                num_registered_this_cluster = 0\n                for img_id_combined_original in current_cluster_image_ids_combined:\n                    img_basename_key = img_id_combined_original.split('__')[-1] # Key used in Davin's output\n                    pose_info = cluster_poses_from_sfm.get(img_basename_key)\n                    \n                    if pose_info and isinstance(pose_info, dict) and pose_info.get('registered', False):\n                        all_image_final_poses[img_id_combined_original] = {\n                            'R_arr': pose_info['R'], \n                            'T_arr': pose_info['T'],\n                            'scene_label_final': final_scene_label,\n                            'registered': True\n                        }\n                        num_registered_this_cluster +=1\n                        SFM_OVERALL_SUCCESS = True # At least one image got a pose\n                    else: \n                        all_image_final_poses[img_id_combined_original] = {\n                            'R_arr': np.full((3,3), np.nan), 'T_arr': np.full((3,), np.nan),\n                            'scene_label_final': final_scene_label,\n                            'registered': False\n                        }\n                print(f\"    SfM for {final_scene_label} registered {num_registered_this_cluster} images.\")\n\n            except Exception as e_sfm_call:\n                print(f\"    ERROR calling/running Davin's SfM module for {final_scene_label}: {e_sfm_call}\")\n                # Fallback: assign NaN poses for all images in this cluster\n                for img_id in current_cluster_image_ids_combined:\n                    all_image_final_poses[img_id] = {'R_arr': np.full((3,3), np.nan), 'T_arr': np.full((3,), np.nan), \n                                                       'scene_label_final': final_scene_label, 'registered': False}\n        else: # Placeholder logic if Davin's module is not ready/imported\n            # print(f\"    Using placeholder: Assigning NaN poses for {final_scene_label}.\")\n            for img_id in current_cluster_image_ids_combined:\n                all_image_final_poses[img_id] = {\n                    'R_arr': np.full((3,3), np.nan), 'T_arr': np.full((3,), np.nan),\n                    'scene_label_final': final_scene_label, 'registered': False\n                }\n        processed_sfm_clusters_count += 1\n\n    # Handle original noise points from clustering (images that were never in a scene cluster)\n    if 'test_clustering_results_df' in globals() and not test_clustering_results_df.empty:\n        original_noise_ids = test_clustering_results_df[\n            test_clustering_results_df['predicted_scene_label_raw'] == -1\n        ]['image_id_combined'].tolist()\n        for img_id_noise in original_noise_ids:\n            if img_id_noise not in all_image_final_poses: \n                 all_image_final_poses[img_id_noise] = {\n                    'R_arr': np.full((3,3), np.nan), 'T_arr': np.full((3,), np.nan), \n                    'scene_label_final': \"outliers\", 'registered': False\n                }\n        print(f\"\\nProcessed {len(original_noise_ids)} original noise points from clustering (marked as 'outliers').\")\n    \n    print(f\"\\nFinished processing {processed_sfm_clusters_count} clusters for SfM.\")\n\n\n# Ensure all_image_final_poses is defined for the next cell, even if all above failed\nif 'all_image_final_poses' not in locals():\n    print(\"CRITICAL ERROR: all_image_final_poses was not defined after SfM stage. Fallback to empty dict.\")\n    all_image_final_poses = {}\n\nprint(\"-\" * 60)\nprint(\"Cell 7: Local Feature Matching & SfM stage complete.\")\nif SFM_STAGE_ATTEMPTED and SFM_OVERALL_SUCCESS:\n    num_actually_registered_total = sum(1 for pose_info in all_image_final_poses.values() if pose_info.get('registered', False))\n    print(f\"SfM processing attempted and at least one image was registered. Total images with poses: {num_actually_registered_total} / {len(all_image_final_poses)}.\")\nelif SFM_STAGE_ATTEMPTED: # Attempted but no image got registered\n    print(f\"SfM processing attempted but NO images were successfully registered. All poses will be NaN.\")\nelif 'all_image_final_poses' in locals() and all_image_final_poses: # Placeholder logic ran\n    print(f\"SfM placeholder logic ran. Pose information (currently NaNs) stored for {len(all_image_final_poses)} images.\")\nelse:\n    print(\"SfM stage was NOT successful or was skipped. 'all_image_final_poses' may be empty.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-02T17:11:41.274411Z","iopub.execute_input":"2025-06-02T17:11:41.274790Z","iopub.status.idle":"2025-06-02T17:11:41.381904Z","shell.execute_reply.started":"2025-06-02T17:11:41.274765Z","shell.execute_reply":"2025-06-02T17:11:41.381132Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 8. Submission File Generation\n\nThis is the final stage of the pipeline, where all processed information is compiled into the `submission.csv` file in the format required by the competition.\n\n*   **Input:**\n    *   `sample_submission.csv`: Loaded from the Kaggle input directory. This provides the template, ensuring all required test images are included in the output and in the correct order. It contains `dataset` and `image` (filename) columns.\n    *   `all_image_final_poses`: A Python dictionary generated in the previous (Local Matching & SfM) stage. This dictionary maps each `image_id_combined` (e.g., `dataset__image_filename`) to its:\n        *   `scene_label_final`: The predicted scene/cluster label (e.g., \"outliers\", \"cluster0\", \"cluster1\").\n        *   `R_arr`: The 3x3 rotation matrix as a NumPy array (or an array of NaNs if pose estimation failed or was skipped).\n        *   `T_arr`: The 3x1 translation vector as a NumPy array (or an array of NaNs).\n*   **Process:**\n    1.  The `sample_submission.csv` is loaded as the base for our output file.\n    2.  An `image_id_combined` is created for each row in this template to allow merging/lookup with our processed results.\n    3.  The code iterates through each row of the submission template. For each image:\n        *   It looks up the `scene_label_final`, `R_arr`, and `T_arr` from our `all_image_final_poses` dictionary.\n        *   If information is found:\n            *   The `scene_label_final` is used directly.\n            *   The NumPy arrays for R and T are flattened and converted into semicolon-separated strings. If they are NaN arrays (because pose estimation was a placeholder or failed for that image/cluster), they are converted to the required \"nan;nan;...\" string format.\n        *   If, for any reason, an image from the `sample_submission.csv` template is *not* found in our `all_image_final_poses` dictionary (which ideally shouldn't happen if all prior stages account for every image), it defaults to \"outliers\" with NaN poses and a warning is printed.\n*   **Output:**\n    *   A `submission.csv` file is saved to `/kaggle/working/`. This file adheres to the specified columns: `dataset,scene,image,rotation_matrix,translation_vector`.\n\nThis cell ensures that our predictions for scene clustering and pose estimation are correctly formatted and submitted for all required test images.","metadata":{}},{"cell_type":"code","source":"# Cell 8: Submission File Generation\n\nprint(\"=== Stage 8: Generating Final submission.csv File ===\")\n\nSUBMISSION_CREATED_SUCCESSFULLY = False\nFINAL_SUBMISSION_DF_COLUMNS = ['dataset', 'scene', 'image', 'rotation_matrix', 'translation_vector']\n\n# KAGGLE_INPUT_DIR and KAGGLE_WORKING_DIR should be defined in Cell 2\n# all_image_final_poses should be defined and populated (even if with NaNs) from Cell 7\n\nif 'KAGGLE_INPUT_DIR' not in globals() or 'KAGGLE_WORKING_DIR' not in globals():\n    print(\"CRITICAL ERROR: KAGGLE_INPUT_DIR or KAGGLE_WORKING_DIR not defined. Cannot proceed.\")\nelse:\n    submission_template_df_path = os.path.join(KAGGLE_INPUT_DIR, 'sample_submission.csv')\n    \n    if not os.path.exists(submission_template_df_path):\n        print(f\"CRITICAL ERROR: sample_submission.csv (template) not found at {submission_template_df_path}.\")\n        print(\"         Cannot generate submission. Notebook will likely fail scoring.\")\n    else:\n        try:\n            submission_template_df = pd.read_csv(submission_template_df_path)\n            print(f\"Loaded submission template with {len(submission_template_df)} required submission rows.\")\n\n            if submission_template_df.empty:\n                print(\"ERROR: Submission template (sample_submission.csv) is empty.\")\n            elif 'all_image_final_poses' not in globals() or not isinstance(all_image_final_poses, dict):\n                print(\"ERROR: 'all_image_final_poses' dictionary not found or not a dictionary from Cell 7.\")\n                print(\"       Defaulting all images to 'outliers' with NaN poses based on template.\")\n                # Create a default submission based on the template\n                submission_template_df['scene'] = \"outliers\"\n                submission_template_df['rotation_matrix'] = \";\".join([\"nan\"] * 9)\n                submission_template_df['translation_vector'] = \";\".join([\"nan\"] * 3)\n                final_submission_df = submission_template_df[FINAL_SUBMISSION_DF_COLUMNS].copy()\n            else:\n                # Proceed with populating from all_image_final_poses\n                submission_template_df['image_id_combined'] = submission_template_df['dataset'].astype(str) + \"__\" + submission_template_df['image'].astype(str)\n\n                output_pred_scenes = []\n                output_rot_matrices_str = []\n                output_trans_vectors_str = []\n\n                nan_rotation_str = \";\".join([\"nan\"] * 9)\n                nan_translation_str = \";\".join([\"nan\"] * 3)\n                \n                processed_image_ids = set(all_image_final_poses.keys())\n                template_image_ids = set(submission_template_df['image_id_combined'].tolist())\n                \n                missing_from_poses_dict = template_image_ids - processed_image_ids\n                extra_in_poses_dict = processed_image_ids - template_image_ids\n\n                if extra_in_poses_dict:\n                    print(f\"WARNING: {len(extra_in_poses_dict)} image_ids found in 'all_image_final_poses' that are NOT in sample_submission.csv. These will be ignored.\")\n                    # print(f\"  Example extra IDs: {list(extra_in_poses_dict)[:5]}\")\n\n\n                print(f\"Populating submission data. {len(all_image_final_poses)} entries in all_image_final_poses.\")\n                \n                for idx, template_row in tqdm(submission_template_df.iterrows(), total=len(submission_template_df), desc=\"Formatting Submission\"):\n                    img_id_comb = template_row['image_id_combined']\n                    pose_data = all_image_final_poses.get(img_id_comb) \n\n                    if pose_data and isinstance(pose_data, dict):\n                        scene_label = pose_data.get('scene_label_final', 'outliers') # Default to outliers\n                        output_pred_scenes.append(scene_label)\n                        \n                        # Ensure 'registered' key exists, default to False if not (e.g. placeholder did not set it)\n                        is_registered = pose_data.get('registered', False) \n                        \n                        r_arr = pose_data.get('R_arr')\n                        has_valid_r = isinstance(r_arr, np.ndarray) and r_arr.shape == (3,3) and not np.all(np.isnan(r_arr))\n                        \n                        t_arr = pose_data.get('T_arr')\n                        has_valid_t = isinstance(t_arr, np.ndarray) and t_arr.size == 3 and not np.all(np.isnan(t_arr))\n\n                        if is_registered and has_valid_r and has_valid_t : # Only output R,T if explicitly registered and valid\n                            output_rot_matrices_str.append(\";\".join(map(str, r_arr.flatten(order='C'))))\n                            output_trans_vectors_str.append(\";\".join(map(str, t_arr.flatten())))\n                        else: # Not registered, or R/T invalid, or missing\n                            output_rot_matrices_str.append(nan_rotation_str)\n                            output_trans_vectors_str.append(nan_translation_str)\n                            if scene_label != \"outliers\" and not (is_registered and has_valid_r and has_valid_t):\n                                # If it was assigned to a cluster but has no valid pose, change to \"outliers\"\n                                # This is based on the forum discussion to avoid 0.0 scores\n                                output_pred_scenes[-1] = \"outliers\" \n                                # print(f\"  DevInfo: Image {img_id_comb} in scene {scene_label} had no valid pose, changed to 'outliers'.\")\n                    else:\n                        # This image from sample_submission.csv was not in all_image_final_poses at all.\n                        # This indicates an issue upstream (e.g. never got an embedding or cluster label).\n                        # print(f\"Warning: No pose info for {img_id_comb} in all_image_final_poses. Defaulting to 'outliers' and NaN pose.\")\n                        output_pred_scenes.append(\"outliers\") \n                        output_rot_matrices_str.append(nan_rotation_str)\n                        output_trans_vectors_str.append(nan_translation_str)\n                \n                # Create the final DataFrame with the correct columns and order\n                final_submission_df = pd.DataFrame({\n                    'dataset': submission_template_df['dataset'],\n                    'scene': output_pred_scenes,\n                    'image': submission_template_df['image'],\n                    'rotation_matrix': output_rot_matrices_str,\n                    'translation_vector': output_trans_vectors_str\n                })\n                \n                output_submission_path = os.path.join(KAGGLE_WORKING_DIR, 'submission.csv')\n                final_submission_df.to_csv(output_submission_path, index=False)\n                print(f\"\\nFinal submission file created successfully at: {output_submission_path}\")\n                print(\"First 5 rows of submission.csv:\")\n                print(final_submission_df.head())\n                if not final_submission_df.empty:\n                    print(\"\\nScene distribution in final submission:\")\n                    print(final_submission_df['scene'].value_counts(dropna=False).sort_index())\n                SUBMISSION_CREATED_SUCCESSFULLY = True\n\n        except Exception as e:\n            print(f\"ERROR during submission file generation: {e}\")\n            # Attempt to create a fallback dummy if a major error occurred\n            try:\n                if 'submission_template_df' not in locals() or submission_template_df.empty: # if template itself failed to load\n                    # Try to load it again for the dummy\n                    if os.path.exists(submission_template_df_path):\n                         submission_template_df = pd.read_csv(submission_template_df_path)\n                    else: # Cannot even load template for dummy\n                         print(\"Cannot create fallback, template submission CSV missing.\")\n                         raise RuntimeError(\"Cannot create any submission file.\")\n\n\n                print(\"Attempting to create fallback DUMMY submission due to error...\")\n                fallback_df = pd.DataFrame()\n                fallback_df['dataset'] = submission_template_df['dataset']\n                fallback_df['scene'] = 'outliers'\n                fallback_df['image'] = submission_template_df['image']\n                fallback_df['rotation_matrix'] = \";\".join([\"nan\"] * 9)\n                fallback_df['translation_vector'] = \";\".join([\"nan\"] * 3)\n                \n                output_submission_path = os.path.join(KAGGLE_WORKING_DIR, 'submission.csv')\n                fallback_df.to_csv(output_submission_path, index=False)\n                print(f\"CRITICAL WARNING: Created a fallback DUMMY submission.csv at {output_submission_path} due to an error during actual submission generation.\")\n            except Exception as e_fallback:\n                print(f\"Error creating even the fallback dummy submission: {e_fallback}\")\n\n\nprint(\"-\" * 60)\nprint(\"Cell 8: Submission File Generation complete.\")\nif SUBMISSION_CREATED_SUCCESSFULLY:\n    print(\"Submission file was generated.\")\nelse:\n    print(\"Submission file was NOT generated successfully or a fallback dummy was created. Check errors above.\")\n    print(\"Ensure a 'submission.csv' file exists in /kaggle/working/ for Kaggle to pick up.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-02T17:11:41.382698Z","iopub.execute_input":"2025-06-02T17:11:41.382976Z","iopub.status.idle":"2025-06-02T17:11:41.543072Z","shell.execute_reply.started":"2025-06-02T17:11:41.382953Z","shell.execute_reply":"2025-06-02T17:11:41.542451Z"}},"outputs":[],"execution_count":null}]}