{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.12.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceType":"competition","sourceId":49349,"databundleVersionId":5447706}],"dockerImageVersionId":31287,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"","metadata":{"trusted":true,"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# **3D Reconstruction VGGT-SLAM**\n### **VGGT-SLAM 2.0: Real time Dense Feed-forward Scene Reconstruction**\nhttps://arxiv.org/abs/2601.19887","metadata":{}},{"cell_type":"markdown","source":"### What is VGGT-SLAM?\n\nVGGT-SLAM is a dense RGB SLAM system developed at MIT's SPARK Lab and accepted at NeurIPS 2025. It builds a consistent 3D map of a scene from **ordinary monocular video**, requiring no camera calibration and no additional training.\n\nThe system is built on top of **VGGT (Visual Geometry Grounded Transformer)**, a feed-forward neural network that estimates dense point clouds, depth maps, feature tracks, and camera poses directly from raw images. VGGT is powerful but memory-limited — on a 24 GB GPU it can process only ~60 frames at a time. VGGT-SLAM overcomes this by dividing the video into **overlapping submaps**, each reconstructed independently by VGGT, and then stitching them together into a globally consistent map.\n","metadata":{}},{"cell_type":"markdown","source":"**VGGT-SLAM 2.0** is a state-of-the-art, **real-time RGB feed-forward SLAM system** designed to build globally consistent 3D maps from uncalibrated camera images, such as those from an iPhone. Developed as a significant improvement over its predecessor, it addresses critical issues like **high-dimensional drift** and **planar degeneracy**—problems that often caused earlier systems to fail or warp when facing flat surfaces like floors or walls.\n\n### Key Technical Innovations\n*   **Enhanced Submap Alignment:** Unlike the original version that used a complex 15-degree-of-freedom (DoF) alignment, VGGT-SLAM 2.0 utilizes a new factor graph design. It directly enforces that overlapping frames between submaps must share the same position, rotation, and camera calibration, which reduces **pose error by approximately 23%** compared to the previous version.\n*   **Intelligent Loop Closure Verification:** The system leverages a unique discovery within the **attention layers of the VGGT model (specifically layer 22)**. This allows the system to verify if two images actually overlap without requiring additional training, effectively rejecting false positive matches and ensuring more reliable global map optimization.\n*   **3D Open-Set Object Detection:** Beyond mapping, the system can be easily adapted for **semantic tasks**. Users can search for objects within the reconstructed 3D map using simple text queries (e.g., \"pink chair\" or \"backpack\"), and the system will identify the object and produce a 3D oriented bounding box around it.\n\n### Performance and Versatility\nThe system is built for **real-world robotics applications**, demonstrating its ability to run online onboard a ground robot using a **Jetson Thor**. It has been successfully tested in diverse environments, ranging from cluttered apartments and office spaces to large-scale 4,200 square foot barns and outdoor driving sequences spanning over 2 kilometers. \n\nBecause it does not require prior camera calibration or specialized training, VGGT-SLAM 2.0 provides a flexible and maintainable solution for dense scene reconstruction and navigation.","metadata":{}},{"cell_type":"markdown","source":"## GPU Check","metadata":{}},{"cell_type":"code","source":"import subprocess\nresult = subprocess.run(['nvidia-smi'], capture_output=True, text=True)\nprint(result.stdout if result.returncode == 0 else '⚠️ No GPU detected. Please enable a GPU accelerator.')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-07T13:11:10.815114Z","iopub.execute_input":"2026-03-07T13:11:10.815516Z","iopub.status.idle":"2026-03-07T13:11:10.879357Z","shell.execute_reply.started":"2026-03-07T13:11:10.815461Z","shell.execute_reply":"2026-03-07T13:11:10.878602Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## System Dependencies","metadata":{}},{"cell_type":"code","source":"%%bash\nset -e\napt-get update -qq\napt-get install -y -qq \\\n    git python3-pip libboost-all-dev cmake gcc g++ unzip ffmpeg \\\n    libgl1-mesa-glx libglib2.0-0 > /dev/null\necho '✅ System packages installed'","metadata":{"trusted":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2026-03-07T13:11:10.881079Z","iopub.execute_input":"2026-03-07T13:11:10.88153Z","iopub.status.idle":"2026-03-07T13:12:01.287813Z","shell.execute_reply.started":"2026-03-07T13:11:10.881496Z","shell.execute_reply":"2026-03-07T13:12:01.287198Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Clone VGGT-SLAM Repository","metadata":{}},{"cell_type":"code","source":"import os\n\nWORK_DIR = '/kaggle/working'\nREPO_DIR = f'{WORK_DIR}/VGGT-SLAM'\n\nif not os.path.exists(REPO_DIR):\n    !git clone https://github.com/MIT-SPARK/VGGT-SLAM {REPO_DIR}\nelse:\n    print('✅ Repository already cloned')\n\nos.chdir(REPO_DIR)\nprint(f'Working directory: {os.getcwd()}')","metadata":{"trusted":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2026-03-07T13:12:01.288691Z","iopub.execute_input":"2026-03-07T13:12:01.28899Z","iopub.status.idle":"2026-03-07T13:12:04.09656Z","shell.execute_reply.started":"2026-03-07T13:12:01.288968Z","shell.execute_reply":"2026-03-07T13:12:04.095624Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Install Python Dependencies","metadata":{}},{"cell_type":"code","source":"%%bash\nset -e\ncd /kaggle/working/VGGT-SLAM\npython -c 'import torch; print(f\"PyTorch {torch.__version__}, CUDA: {torch.cuda.is_available()}\")'\npip install -q -r requirements.txt\necho '✅ Python packages installed'","metadata":{"trusted":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2026-03-07T13:12:04.097953Z","iopub.execute_input":"2026-03-07T13:12:04.098491Z","iopub.status.idle":"2026-03-07T13:16:24.048163Z","shell.execute_reply.started":"2026-03-07T13:12:04.098463Z","shell.execute_reply":"2026-03-07T13:16:24.047545Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Build Third-Party Dependencies (GTSAM, VGGT)\n\n`setup.sh` assumes conda, so we patch it for the Kaggle environment.","metadata":{}},{"cell_type":"code","source":"import os, re\n\nos.chdir('/kaggle/working/VGGT-SLAM')\n\nwith open('setup.sh') as f:\n    lines = f.readlines()\n\npatched = []\nfor line in lines:\n    s = line.rstrip()\n    # Skip conda lines\n    if re.match(r'\\s*(conda create|conda activate)', s):\n        patched.append(f'# [kaggle skip] {s}\\n')\n        continue\n    # Make git clone idempotent\n    m = re.match(r'(\\s*)git clone (.*)', s)\n    if m:\n        indent, args = m.group(1), m.group(2).split()\n        clone_dir = args[-1] if len(args) >= 2 else os.path.basename(args[0]).replace('.git', '')\n        patched.append(\n            f'{indent}if [ ! -d \"{clone_dir}\" ]; then\\n'\n            f'{indent}  git clone {\" \".join(args)}\\n'\n            f'{indent}else\\n'\n            f'{indent}  echo \"[skip] {clone_dir} already cloned\"\\n'\n            f'{indent}fi\\n'\n        )\n        continue\n    patched.append(line)\n\nwith open('setup_kaggle.sh', 'w') as f:\n    f.writelines(patched)\n\nprint('✅ setup_kaggle.sh generated')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-07T13:16:24.049163Z","iopub.execute_input":"2026-03-07T13:16:24.049372Z","iopub.status.idle":"2026-03-07T13:16:24.058358Z","shell.execute_reply.started":"2026-03-07T13:16:24.049353Z","shell.execute_reply":"2026-03-07T13:16:24.057612Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"%%bash\nset -e\ncd /kaggle/working/VGGT-SLAM\necho '=== Running setup_kaggle.sh (first run: 20–40 min) ==='\nbash setup_kaggle.sh\necho '✅ Setup complete'\n###","metadata":{"trusted":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2026-03-07T13:16:24.059453Z","iopub.execute_input":"2026-03-07T13:16:24.05981Z","iopub.status.idle":"2026-03-07T13:18:35.277507Z","shell.execute_reply.started":"2026-03-07T13:16:24.059728Z","shell.execute_reply":"2026-03-07T13:18:35.276724Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# open3d conflicts with Kaggle's numpy — use matplotlib instead (already installed)\nimport matplotlib\nprint(f'✅ matplotlib {matplotlib.__version__} ready')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-07T13:18:35.279802Z","iopub.execute_input":"2026-03-07T13:18:35.28002Z","iopub.status.idle":"2026-03-07T13:18:35.285072Z","shell.execute_reply.started":"2026-03-07T13:18:35.28Z","shell.execute_reply":"2026-03-07T13:18:35.284401Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Image Data\n\nSet `MAX_FRAMES = None` to process the full video.","metadata":{}},{"cell_type":"code","source":"import subprocess, os, shutil\nimport os, shutil, math\n\nMAX_FRAMES = None  # 151, None = all frames used\nIMAGE_DIR  = \"/kaggle/input/competitions/image-matching-challenge-2023/train/haiper/bike/images_full\"\nFRAMES_DIR = '/kaggle/working/input_frames'\n\nos.makedirs(FRAMES_DIR, exist_ok=True)\n\n\nall_frames = sorted([f for f in os.listdir(IMAGE_DIR) if f.lower().endswith(('.jpg', '.jpeg', '.png'))])\nprint(f'📁 Image folder : {IMAGE_DIR}')\nprint(f'   Total images : {len(all_frames)}')\n\ntargets = all_frames[:MAX_FRAMES] if MAX_FRAMES else all_frames\nfor f in targets:\n    shutil.copy(os.path.join(IMAGE_DIR, f), os.path.join(FRAMES_DIR, f))\nprint(f'✅ Copied {len(targets)} images → {FRAMES_DIR}')\n\nn_frames = len([f for f in os.listdir(FRAMES_DIR) if f.lower().endswith(('.jpg', '.jpeg', '.png'))])\nprint(f'\\n📸 Total frames to use: {n_frames}')\n\nSUBMAP_SIZE_PREVIEW = 10\nprint(f'   Estimated submaps: {math.ceil(n_frames / SUBMAP_SIZE_PREVIEW)} ({n_frames} ÷ {SUBMAP_SIZE_PREVIEW})')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-07T13:18:35.286204Z","iopub.execute_input":"2026-03-07T13:18:35.286462Z","iopub.status.idle":"2026-03-07T13:18:36.130178Z","shell.execute_reply.started":"2026-03-07T13:18:35.286441Z","shell.execute_reply":"2026-03-07T13:18:36.129508Z"},"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport matplotlib.image as mpimg\n\nframe_files = sorted([f for f in os.listdir(FRAMES_DIR) if f.endswith(('.jpg', '.png', 'jpeg'))])\nsample_idx  = [0, len(frame_files)//4, len(frame_files)//2, min(len(frame_files)-1, 3*len(frame_files)//4)]\n\nfig, axes = plt.subplots(1, 4, figsize=(16, 4))\nfig.suptitle('Input Frame Samples', fontsize=14)\nfor ax, idx in zip(axes, sample_idx):\n    img = mpimg.imread(os.path.join(FRAMES_DIR, frame_files[idx]))\n    ax.imshow(img)\n    ax.set_title(f'Frame {idx+1}/{len(frame_files)}')\n    ax.axis('off')\nplt.tight_layout()\nplt.savefig('/kaggle/working/input_preview.png', dpi=100, bbox_inches='tight')\nplt.show()\nprint('✅ Saved: /kaggle/working/input_preview.png')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-07T13:21:40.927292Z","iopub.execute_input":"2026-03-07T13:21:40.92806Z","iopub.status.idle":"2026-03-07T13:21:42.955196Z","shell.execute_reply.started":"2026-03-07T13:21:40.92803Z","shell.execute_reply":"2026-03-07T13:21:42.954484Z"},"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Download Model Checkpoint (SALAD / DINOv2-B)","metadata":{}},{"cell_type":"code","source":"###\nimport os, urllib.request\n\nCKPT_DIR  = '/root/.cache/torch/hub/checkpoints'\nCKPT_PATH = f'{CKPT_DIR}/dino_salad.ckpt'\nCKPT_URL  = 'https://github.com/serizba/salad/releases/download/v1.0.0/dino_salad.ckpt'\n\nos.makedirs(CKPT_DIR, exist_ok=True)\n\nif os.path.exists(CKPT_PATH):\n    print(f'✅ Already downloaded: {CKPT_PATH} ({os.path.getsize(CKPT_PATH)/1e6:.0f} MB)')\nelse:\n    print(f'📥 Downloading: {CKPT_URL}')\n    def progress(block_num, block_size, total_size):\n        pct = block_num * block_size / total_size * 100 if total_size > 0 else 0\n        if block_num % 500 == 0:\n            print(f'  {pct:.1f}%  ({block_num*block_size/1e6:.0f} / {total_size/1e6:.0f} MB)', flush=True)\n    urllib.request.urlretrieve(CKPT_URL, CKPT_PATH, reporthook=progress)\n    print(f'✅ Download complete: {CKPT_PATH} ({os.path.getsize(CKPT_PATH)/1e6:.0f} MB)')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-07T13:21:49.75513Z","iopub.execute_input":"2026-03-07T13:21:49.755424Z","iopub.status.idle":"2026-03-07T13:22:00.309036Z","shell.execute_reply.started":"2026-03-07T13:21:49.755398Z","shell.execute_reply":"2026-03-07T13:22:00.308249Z"},"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Run VGGT-SLAM\n\nMemory-saving settings for P100 / T4 (16 GB VRAM):\n- **`RESIZE_WIDTH`**: resize input frames to reduce attention memory\n- **`SUBMAP_SIZE = 8`**: fewer frames per batch → lower VRAM peak\n\nIf you still get OOM, try `RESIZE_WIDTH = 518` first; reduce further if needed.","metadata":{}},{"cell_type":"code","source":"###\nimport subprocess, os, time\n\nos.chdir('/kaggle/working/VGGT-SLAM')\n\nLOG_DIR = '/kaggle/working/slam_output'\nos.makedirs(LOG_DIR, exist_ok=True)\n\n# ── VRAM-saving parameters ────────────────────────────────\nRESIZE_WIDTH = 1036   #14*x\nSUBMAP_SIZE  = 6\nMAX_LOOPS    = 1\nUSE_SIM3     = False\n# ──────────────────────────────────────────────────────────\n\n# ── Step 1: Resize frames with ffmpeg ────────────────────\nif RESIZE_WIDTH:\n    RESIZED_DIR = '/kaggle/working/input_frames_resized'\n    frames  = sorted([f for f in os.listdir(FRAMES_DIR) if f.endswith(('.jpg', '.png','jpeg'))])\n    already = len([f for f in os.listdir(RESIZED_DIR) if f.endswith(('.jpg', '.png','jpeg'))]) \\\n              if os.path.exists(RESIZED_DIR) else 0\n    if already == len(frames):\n        print(f'✅ Resize already done ({already} frames) — skipping')\n    else:\n        os.makedirs(RESIZED_DIR, exist_ok=True)\n        print(f'🔄 Resizing {len(frames)} frames to width {RESIZE_WIDTH}px ...')\n        vf = f'scale={RESIZE_WIDTH}:trunc(ow/a/8)*8'\n        for fname in frames:\n            r = subprocess.run(\n                ['ffmpeg', '-i', os.path.join(FRAMES_DIR, fname),\n                 '-vf', vf, '-q:v', '3',\n                 os.path.join(RESIZED_DIR, fname),\n                 '-y', '-loglevel', 'error'],\n                capture_output=True\n            )\n            if r.returncode != 0:\n                print(f'  ⚠️ {fname}:', r.stderr.decode())\n        done = len([f for f in os.listdir(RESIZED_DIR) if f.endswith(('.jpg', '.png','jpeg'))])\n        print(f'✅ Resize done: {done}/{len(frames)} frames → {RESIZED_DIR}')\n    INPUT_DIR = RESIZED_DIR\nelse:\n    INPUT_DIR = FRAMES_DIR\n    print('ℹ️  No resize — using original resolution')\n\n# ── Step 2: Run VGGT-SLAM ────────────────────────────────\ncmd = [\n    'python3', 'main.py',\n    '--image_folder', INPUT_DIR,\n    '--max_loops',    str(MAX_LOOPS),\n    '--submap_size',  str(SUBMAP_SIZE),\n    '--log_results',\n    '--log_path', os.path.join(LOG_DIR, 'poses.txt'),\n]\nif USE_SIM3:\n    cmd.append('--use_sim3')\n\nprint('\\nCommand:', ' '.join(cmd))\nprint('=' * 60)\nstart = time.time()\nproc  = subprocess.Popen(cmd, stdout=subprocess.PIPE, stderr=subprocess.STDOUT, text=True)\nfor line in proc.stdout:\n    print(line, end='')\nproc.wait()\nelapsed = time.time() - start\nprint('=' * 60)\nif proc.returncode == 0:\n    print(f'✅ VGGT-SLAM complete!  Elapsed: {elapsed/60:.1f} min')\n    for root, dirs, files in os.walk(LOG_DIR):\n        for f in files:\n            p = os.path.join(root, f)\n            print(f'  📄 {p}  ({os.path.getsize(p):,} bytes)')\nelse:\n    print(f'❌ Error (code={proc.returncode})')\n    print('💡 If OOM: lower RESIZE_WIDTH (448 → 384 → 320) or use Colab A100')\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-07T13:22:31.010207Z","iopub.execute_input":"2026-03-07T13:22:31.010531Z","iopub.status.idle":"2026-03-07T13:26:15.454526Z","shell.execute_reply.started":"2026-03-07T13:22:31.010505Z","shell.execute_reply":"2026-03-07T13:26:15.453851Z"},"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Output File Summary","metadata":{}},{"cell_type":"code","source":"import os\n\nsearch_dirs = [LOG_DIR, '/kaggle/working/VGGT-SLAM', '/kaggle/working']\nprint('=== Output File Search ===')\nfor d in search_dirs:\n    if not os.path.exists(d):\n        continue\n    print(f'\\n📂 {d}')\n    for root, dirs, files in os.walk(d):\n        dirs[:] = [x for x in dirs if x not in\n                   ('vggt_slam', 'evals', 'scripts', 'assets', '__pycache__', '.git')]\n        for file in files:\n            if file.endswith(('.ply', '.npy', '.npz', '.txt', '.log', '.png', 'jpeg','.json')):\n                fp   = os.path.join(root, file)\n                size = os.path.getsize(fp)\n                print(f'  {fp.replace(d, \"\")}  ({size/1e3:.1f} KB)')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-07T13:26:25.403649Z","iopub.execute_input":"2026-03-07T13:26:25.403965Z","iopub.status.idle":"2026-03-07T13:26:25.422545Z","shell.execute_reply.started":"2026-03-07T13:26:25.40394Z","shell.execute_reply":"2026-03-07T13:26:25.421798Z"},"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Static 3D Visualization: Trajectory + Point Cloud","metadata":{}},{"cell_type":"code","source":"LOG_DIR = '/kaggle/working/slam_output'\nIMG_DIR = '/kaggle/working/input_frames_resized'\n\nfrom PIL import Image\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport glob, os, random\nimport struct","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-07T13:26:25.424135Z","iopub.execute_input":"2026-03-07T13:26:25.424361Z","iopub.status.idle":"2026-03-07T13:26:25.428798Z","shell.execute_reply.started":"2026-03-07T13:26:25.42434Z","shell.execute_reply":"2026-03-07T13:26:25.427986Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Load trajectory (format: frame_id tx ty tz qx qy qz qw)\nposes = []\nwith open(f'{LOG_DIR}/poses.txt') as f:\n    for line in f:\n        vals = list(map(float, line.strip().split()))\n        poses.append((int(vals[0]), vals[1], vals[2], vals[3]))\nposes.sort(key=lambda x: x[0])\ntraj = np.array([[p[1], p[2], p[3]] for p in poses])\nprint(f'Trajectory: {len(traj)} keyframes')\n\n# Load point cloud\nnpz_files = sorted(glob.glob(f'{LOG_DIR}/poses_logs/*.npz'))\nall_xyz   = []\nfor npz_path in npz_files:\n    npz = np.load(npz_path)\n    all_xyz.append(npz['pointcloud'][npz['mask']])\npts = np.vstack(all_xyz)\nfor axis in range(3):\n    lo, hi = np.percentile(pts[:, axis], [2, 98])\n    pts = pts[(pts[:, axis] >= lo) & (pts[:, axis] <= hi)]\nif len(pts) > 800_000:\n    idx = sorted(random.sample(range(len(pts)), 800_000))\n    pts = pts[idx]\n\n# Plot\nfig = plt.figure(figsize=(18, 6))\nfor i, (elev, azim, title) in enumerate([\n    (30,  45,  'Perspective'),\n    (90, -90,  'Top-down'),\n    (0,   0,   'Front'),\n]):\n    ax = fig.add_subplot(1, 3, i+1, projection='3d')\n    ax.scatter(pts[:,0], pts[:,1], pts[:,2], c='lightgray', s=0.1, alpha=0.3, label='Point cloud')\n    ax.plot(traj[:,0], traj[:,1], traj[:,2], c='red', lw=2, label='Trajectory')\n    ax.scatter(traj[:,0], traj[:,1], traj[:,2], c='blue', s=20, zorder=5)\n    ax.scatter(*traj[0],  c='lime',   s=80, marker='^', zorder=7, label='Start')\n    ax.scatter(*traj[-1], c='orange', s=80, marker='s', zorder=7, label='End')\n    ax.view_init(elev=elev, azim=azim)\n    ax.set_title(title, fontsize=10)\n    ax.set_xlabel('X'); ax.set_ylabel('Y'); ax.set_zlabel('Z')\n\nhandles, labels = ax.get_legend_handles_labels()\nfig.legend(handles, labels, loc='lower center', ncol=4, fontsize=9)\nplt.suptitle('VGGT-SLAM — Camera Trajectory + Point Cloud Map', fontsize=13, fontweight='bold')\nplt.tight_layout()\nout = '/kaggle/working/slam_result.png'\nplt.savefig(out, dpi=150, bbox_inches='tight')\nplt.show()\nprint(f'✅ Saved: {out}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-07T13:26:25.429668Z","iopub.execute_input":"2026-03-07T13:26:25.429886Z","iopub.status.idle":"2026-03-07T13:30:04.845977Z","shell.execute_reply.started":"2026-03-07T13:26:25.429866Z","shell.execute_reply":"2026-03-07T13:30:04.845264Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def _quat_to_matrix(qx, qy, qz, qw):\n    \"\"\"\n    Converts a quaternion to a rotation matrix (NumPy only).\n    Expects quaternion components in [x, y, z, w] format.\n    \"\"\"\n    return np.array([\n        [1-2*(qy**2+qz**2),  2*(qx*qy-qw*qz),  2*(qx*qz+qw*qy)],\n        [  2*(qx*qy+qw*qz),1-2*(qx**2+qz**2),  2*(qy*qz-qw*qx)],\n        [  2*(qx*qz-qw*qy),  2*(qy*qz+qw*qx),1-2*(qx**2+qy**2)]\n    ], dtype=np.float64)\n\n\ndef _matrix_to_quat(R):\n    \"\"\"\n    Converts a rotation matrix to a quaternion [x, y, z, w].\n    Uses Shepperd's algorithm for numerical stability (NumPy only).\n    \"\"\"\n    trace = R[0,0] + R[1,1] + R[2,2]\n    \n    if trace > 0:\n        s = 0.5 / np.sqrt(trace + 1.0)\n        w = 0.25 / s\n        x = (R[2,1] - R[1,2]) * s\n        y = (R[0,2] - R[2,0]) * s\n        z = (R[1,0] - R[0,1]) * s\n    elif R[0,0] > R[1,1] and R[0,0] > R[2,2]:\n        s = 2.0 * np.sqrt(1.0 + R[0,0] - R[1,1] - R[2,2])\n        w = (R[2,1] - R[1,2]) / s\n        x = 0.25 * s\n        y = (R[0,1] + R[1,0]) / s\n        z = (R[0,2] + R[2,0]) / s\n    elif R[1,1] > R[2,2]:\n        s = 2.0 * np.sqrt(1.0 + R[1,1] - R[0,0] - R[2,2])\n        w = (R[0,2] - R[2,0]) / s\n        x = (R[0,1] + R[1,0]) / s\n        y = 0.25 * s\n        z = (R[1,2] + R[2,1]) / s\n    else:\n        s = 2.0 * np.sqrt(1.0 + R[2,2] - R[0,0] - R[1,1])\n        w = (R[1,0] - R[0,1]) / s\n        x = (R[0,2] + R[2,0]) / s\n        y = (R[1,2] + R[2,1]) / s\n        z = 0.25 * s\n        \n    return np.array([x, y, z, w], dtype=np.float64)\n\n\ndef matrix_to_quaternion_translation(matrix_3x4: np.ndarray):\n    \"\"\"3×4 [R | t] → COLMAP quaternion [w, x, y, z] + translation t. (wo scipy)\"\"\"\n    R = matrix_3x4[:3, :3]\n    t = matrix_3x4[:3, 3]\n    q = _matrix_to_quat(R)   # [x, y, z, w]\n    qvec = np.array([q[3], q[0], q[1], q[2]])  # COLMAP: [w, x, y, z]\n    return qvec, t\n\n\ndef load_poses_txt(poses_txt_path):\n    \"\"\"poses.txt → {frame_id: (qvec_colmap, tvec)} (wo scipy)\"\"\"\n    poses = {}\n    with open(poses_txt_path) as f:\n        for line in f:\n            vals = list(map(float, line.strip().split()))\n            if len(vals) < 8:\n                continue\n            frame_id = int(vals[0])\n            tx, ty, tz = vals[1], vals[2], vals[3]\n            qx, qy, qz, qw = vals[4], vals[5], vals[6], vals[7]\n\n            R_wc = _quat_to_matrix(qx, qy, qz, qw)\n            t_wc = np.array([tx, ty, tz])\n            R_cw = R_wc.T\n            t_cw = -R_cw @ t_wc\n\n            q_cw = _matrix_to_quat(R_cw)   # [x, y, z, w]\n            qvec = np.array([q_cw[3], q_cw[0], q_cw[1], q_cw[2]])  # COLMAP: [w,x,y,z]\n\n            poses[frame_id] = (qvec, t_cw)\n\n    return poses\n\n\ndef write_next_bytes(fid, data, format_str):\n    if isinstance(data, (list, tuple, np.ndarray)):\n        fid.write(struct.pack(\"<\" + format_str, *data))\n    else:\n        fid.write(struct.pack(\"<\" + format_str, data))\n\n\n\ndef write_cameras_binary(cameras, path):\n    with open(path, \"wb\") as fid:\n        write_next_bytes(fid, len(cameras), \"Q\")\n        for cam_id, cam in cameras.items():\n            write_next_bytes(fid, cam_id, \"I\")\n            write_next_bytes(fid, 1, \"I\")          # PINHOLE\n            write_next_bytes(fid, cam['width'],  \"Q\")\n            write_next_bytes(fid, cam['height'], \"Q\")\n            for p in cam['params']:\n                write_next_bytes(fid, float(p), \"d\")\n\n\ndef write_images_binary(images_data, path):\n    with open(path, \"wb\") as fid:\n        write_next_bytes(fid, len(images_data), \"Q\")\n        for img_id, img in images_data.items():\n            write_next_bytes(fid, img_id,         \"I\")\n            write_next_bytes(fid, img['qvec'],    \"dddd\")\n            write_next_bytes(fid, img['tvec'],    \"ddd\")\n            write_next_bytes(fid, img['camera_id'], \"I\")\n            for char in img['name']:\n                write_next_bytes(fid, char.encode(\"utf-8\"), \"c\")\n            write_next_bytes(fid, b\"\\x00\", \"c\")\n            write_next_bytes(fid, len(img['xys']), \"Q\")\n            for xy, pid in zip(img['xys'], img['point3D_ids']):\n                write_next_bytes(fid, xy, \"dd\")\n                write_next_bytes(fid, pid, \"Q\")\n\n\ndef write_points3d_binary(points3D, path):\n    with open(path, \"wb\") as fid:\n        write_next_bytes(fid, len(points3D), \"Q\")\n        for pid, pt in enumerate(points3D):\n            write_next_bytes(fid, pid,        \"Q\")\n            write_next_bytes(fid, pt['xyz'],  \"ddd\")\n            write_next_bytes(fid, pt['rgb'],  \"BBB\")\n            write_next_bytes(fid, pt['error'],\"d\")\n            write_next_bytes(fid, len(pt['image_ids']), \"Q\")\n            for iid, p2d_idx in zip(pt['image_ids'], pt['point2D_idxs']):\n                write_next_bytes(fid, int(iid),    \"I\")\n                write_next_bytes(fid, int(p2d_idx),\"I\")\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-07T13:30:04.847069Z","iopub.execute_input":"2026-03-07T13:30:04.848Z","iopub.status.idle":"2026-03-07T13:30:04.867536Z","shell.execute_reply.started":"2026-03-07T13:30:04.847967Z","shell.execute_reply":"2026-03-07T13:30:04.866976Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def convert_slam_to_colmap(npz_files, img_dir, poses_txt, output_dir,\n                            intrinsic_K=None, verbose=True):\n    \"\"\"\n    Converts VGGT-SLAM npz data + poses.txt into a COLMAP sparse reconstruction.\n    \n    intrinsic_K: Camera intrinsic parameters as a (3,3) numpy array.\n                  If None, a simple estimation is used (fx=fy=W).\n    \"\"\"\n    from pathlib import Path\n    import shutil\n\n    output_dir = Path(output_dir)\n    sparse_dir = output_dir / \"sparse\" / \"0\"\n    images_dir = output_dir / \"images\"\n    sparse_dir.mkdir(parents=True, exist_ok=True)\n    images_dir.mkdir(parents=True, exist_ok=True)\n\n    img_files = sorted([f for f in os.listdir(img_dir) if f.endswith(('.jpg','.png','jpeg'))])\n    id_to_img = {i+1: os.path.join(img_dir, f) for i, f in enumerate(img_files)}\n\n    poses = load_poses_txt(poses_txt)\n    if verbose:\n        print(f\"  Loaded {len(poses)} poses\")\n\n    cameras     = {}\n    images_data = {}\n    points3D    = []\n\n    cam_idx = 0\n    for npz_path in sorted(npz_files):\n        # Extract frame ID from filename\n        frame_id = int(float(os.path.basename(npz_path).replace('.npz', '')))\n\n        if frame_id not in poses:\n            if verbose:\n                print(f\"  ⚠️ frame {frame_id}: pose not found, skipping\")\n            continue\n\n        npz  = np.load(npz_path)\n        pc   = npz['pointcloud']   # (H, W, 3)\n        mask = npz['mask']         # (H, W)\n        H, W = mask.shape\n\n        # -- Intrinsics ----------------------------------------------\n        if intrinsic_K is not None:\n            K = intrinsic_K\n        else:\n            # Default fallback: focal length equals width\n            K = np.array([[W, 0, W/2],\n                          [0, W, H/2],\n                          [0,  0,  1]], dtype=np.float64)\n\n        fx, fy = float(K[0,0]), float(K[1,1])\n        cx, cy = float(K[0,2]), float(K[1,2])\n\n        cameras[cam_idx] = {\n            'model' : 'PINHOLE',\n            'width' : W,\n            'height': H,\n            'params': [fx, fy, cx, cy]\n        }\n\n        qvec, tvec = poses[frame_id]\n        img_path   = id_to_img.get(frame_id, id_to_img.get(cam_idx+1))\n        img_name   = os.path.basename(img_path) if img_path else f\"frame_{frame_id:04d}.jpg\"\n\n        images_data[cam_idx + 1] = {\n            'qvec'       : qvec,\n            'tvec'       : tvec,\n            'camera_id'  : cam_idx,\n            'name'       : img_name,\n            'xys'        : np.empty((0, 2)),\n            'point3D_ids': np.empty((0,), dtype=np.int64),\n        }\n\n        # -- 3D Point Cloud -----------------------------------------\n        pts = pc[mask]\n        if img_path and os.path.exists(img_path):\n            img  = np.array(Image.open(img_path).convert('RGB').resize((W, H), Image.LANCZOS))\n            cols = img[mask]\n        else:\n            # Fallback to gray if image is missing\n            cols = np.full((len(pts), 3), 128, dtype=np.uint8)\n\n        for pt, col in zip(pts, cols):\n            if np.all(np.isfinite(pt)):\n                points3D.append({\n                    'xyz'         : pt.astype(np.float64),\n                    'rgb'         : col.astype(np.uint8),\n                    'error'        : 0.0,\n                    'image_ids'   : np.array([], dtype=np.int32),\n                    'point2D_idxs': np.array([], dtype=np.int32),\n                })\n\n        if verbose:\n            print(f\"  frame {frame_id:3d}: {len(pts):,} points\")\n\n        if img_path and os.path.exists(img_path):\n            shutil.copy2(img_path, images_dir / img_name)\n\n        cam_idx += 1\n\n    # Export to COLMAP binary format\n    write_cameras_binary(cameras, sparse_dir / \"cameras.bin\")\n    write_images_binary(images_data, sparse_dir / \"images.bin\")\n    write_points3d_binary(points3D, sparse_dir / \"points3D.bin\")\n\n    if verbose:\n        print(f\"\\n✅ COLMAP Export Complete: {sparse_dir}\")\n        print(f\"   cameras: {len(cameras)}, images: {len(images_data)}, points: {len(points3D):,}\")\n\n    return sparse_dir\n\n\n# Usage Example\ncolmap_dir = convert_slam_to_colmap(\n    npz_files   = npz_files,\n    img_dir     = IMG_DIR,\n    poses_txt   = '/kaggle/working/slam_output/poses.txt',\n    output_dir  = '/kaggle/working/slam_colmap',\n    intrinsic_K = None,   # e.g., np.array([[fx,0,cx],[0,fy,cy],[0,0,1]]) \n    verbose     = True,\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-07T13:30:04.869439Z","iopub.execute_input":"2026-03-07T13:30:04.869798Z","iopub.status.idle":"2026-03-07T13:30:24.374279Z","shell.execute_reply.started":"2026-03-07T13:30:04.869743Z","shell.execute_reply":"2026-03-07T13:30:24.373524Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## def colmap_points3d_to_ply","metadata":{}},{"cell_type":"code","source":"from pathlib import Path","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-07T13:33:27.968237Z","iopub.execute_input":"2026-03-07T13:33:27.968796Z","iopub.status.idle":"2026-03-07T13:33:27.972448Z","shell.execute_reply.started":"2026-03-07T13:33:27.968739Z","shell.execute_reply":"2026-03-07T13:33:27.971629Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def colmap_points3d_to_ply(sparse_dir, ply_out):\n    \"\"\"points3D.bin → colored PLY (wo scipy, wo npz)\"\"\"\n    import struct\n\n    path = Path(sparse_dir) / \"points3D.bin\"\n    pts, cols = [], []\n\n    with open(path, \"rb\") as f:\n        n_pts = struct.unpack(\"<Q\", f.read(8))[0]\n        for _ in range(n_pts):\n            pid   = struct.unpack(\"<Q\", f.read(8))[0]\n            xyz   = struct.unpack(\"<ddd\", f.read(24))\n            rgb   = struct.unpack(\"<BBB\", f.read(3))\n            error = struct.unpack(\"<d\",   f.read(8))[0]\n            n_obs = struct.unpack(\"<Q\",   f.read(8))[0]\n            f.read(n_obs * 8)   # image_id(I) + point2D_idx(I) × n_obs\n            pts.append(xyz)\n            cols.append(rgb)\n\n    pts  = np.array(pts,  dtype=np.float32)\n    cols = np.array(cols, dtype=np.uint8)\n\n    valid = np.ones(len(pts), dtype=bool)\n    for axis in range(3):\n        lo, hi = np.percentile(pts[:, axis], [1, 99])\n        valid &= (pts[:, axis] >= lo) & (pts[:, axis] <= hi)\n    pts = pts[valid]; cols = cols[valid]\n\n    with open(ply_out, 'w') as f:\n        f.write(f'ply\\nformat ascii 1.0\\nelement vertex {len(pts)}\\n')\n        f.write('property float x\\nproperty float y\\nproperty float z\\n')\n        f.write('property uchar red\\nproperty uchar green\\nproperty uchar blue\\nend_header\\n')\n        for p, c in zip(pts, cols):\n            f.write(f'{p[0]:.6f} {p[1]:.6f} {p[2]:.6f} {c[0]} {c[1]} {c[2]}\\n')\n\n    print(f'✅ PLY saved: {ply_out}  ({os.path.getsize(ply_out)//1024:,} KB)')\n    return ply_out\n\nply_out = colmap_points3d_to_ply(\n    sparse_dir = '/kaggle/working/slam_colmap/sparse/0',\n    ply_out    = '/kaggle/working/slam_combined_color.ply'\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-07T13:33:27.973884Z","iopub.execute_input":"2026-03-07T13:33:27.974145Z","iopub.status.idle":"2026-03-07T13:33:34.51343Z","shell.execute_reply.started":"2026-03-07T13:33:27.974117Z","shell.execute_reply":"2026-03-07T13:33:34.512669Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## SLAM Animation: Incremental Map Building\n\nEach frame of the GIF shows the point cloud and trajectory growing as new keyframes are added — the canonical SLAM visualization.","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport matplotlib.pyplot as plt\nimport matplotlib.animation as animation\nfrom PIL import Image\nimport glob, os, random\n\n\n# Load trajectory\nposes = []\nwith open(f'{LOG_DIR}/poses.txt') as f:\n    for line in f:\n        vals = list(map(float, line.strip().split()))\n        poses.append((int(vals[0]), vals[1], vals[2], vals[3]))\nposes.sort(key=lambda x: x[0])\n\n# Map frame index → image path\nimg_files = sorted([f for f in os.listdir(IMG_DIR) if f.endswith(('.jpg', '.png', 'jpeg'))])\nid_to_img = {i+1: os.path.join(IMG_DIR, f) for i, f in enumerate(img_files)}\n\n# Load per-keyframe colored point clouds\npc_per_frame  = {}\ncol_per_frame = {}\nfor frame_id, tx, ty, tz in poses:\n    path = f'{LOG_DIR}/poses_logs/{float(frame_id)}.npz'\n    if not os.path.exists(path):\n        continue\n    npz  = np.load(path)\n    pc, mask = npz['pointcloud'], npz['mask']\n    H, W = mask.shape\n    pts  = pc[mask]\n\n    img_path = id_to_img.get(frame_id)\n    if img_path and os.path.exists(img_path):\n        img  = np.array(Image.open(img_path).resize((W, H))) / 255.0\n        cols = img[mask]\n    else:\n        cols = None\n\n    # Outlier removal\n    valid = np.ones(len(pts), dtype=bool)\n    for axis in range(3):\n        lo, hi = np.percentile(pts[:, axis], [2, 98])\n        valid &= (pts[:, axis] >= lo) & (pts[:, axis] <= hi)\n    pts  = pts[valid]\n    cols = cols[valid] if cols is not None else None\n\n    # Downsample per frame\n    if len(pts) > 4000:\n        idx  = sorted(random.sample(range(len(pts)), 4000))\n        pts  = pts[idx]\n        cols = cols[idx] if cols is not None else None\n\n    pc_per_frame[frame_id]  = pts\n    col_per_frame[frame_id] = cols\n    print(f'  frame {frame_id:3d}: {len(pts):,} pts  color={\"yes\" if cols is not None else \"no\"}')\n\n# Global axis bounds\nall_pts = np.vstack(list(pc_per_frame.values()))\ndef axis_range(data, margin=0.1):\n    lo, hi = data.min(), data.max()\n    d = (hi - lo) * margin\n    return lo - d, hi + d\nxl = axis_range(all_pts[:,0])\nyl = axis_range(all_pts[:,1])\nzl = axis_range(all_pts[:,2])\n\n# Build animation\nfig = plt.figure(figsize=(10, 8))\nax  = fig.add_subplot(111, projection='3d')\n\ndef update(frame_idx):\n    ax.cla()\n    ax.set_xlim(*xl); ax.set_ylim(*yl); ax.set_zlim(*zl)\n    ax.set_xlabel('X'); ax.set_ylabel('Y'); ax.set_zlabel('Z')\n\n    current = poses[:frame_idx+1]\n    fid     = current[-1][0]\n\n    # Cumulative colored point cloud\n    for f_id, _, _, _ in current:\n        if f_id in pc_per_frame:\n            p = pc_per_frame[f_id]\n            c = col_per_frame[f_id] if col_per_frame[f_id] is not None else 'lightgray'\n            ax.scatter(p[:,0], p[:,1], p[:,2], c=c, s=0.3, alpha=0.4)\n\n    # Trajectory\n    traj = np.array([[p[1], p[2], p[3]] for p in current])\n    ax.plot(traj[:,0], traj[:,1], traj[:,2], c='red', lw=2)\n    ax.scatter(traj[:,0], traj[:,1], traj[:,2], c='blue', s=15, zorder=5)\n\n    # Current & start markers\n    ax.scatter(*traj[-1], c='yellow', s=120, marker='*', zorder=6)\n    ax.scatter(*traj[0],  c='lime',   s=80,  marker='^', zorder=6)\n\n    ax.set_title(f'VGGT-SLAM  frame {fid}  ({frame_idx+1}/{len(poses)} keyframes)',\n                 fontsize=11, fontweight='bold')\n    ax.view_init(elev=25, azim=45 + frame_idx * 2)\n\nani = animation.FuncAnimation(fig, update, frames=len(poses), interval=800, repeat=True)\n\nout = '/kaggle/working/slam_animation_color.gif'\nani.save(out, writer='pillow', fps=1.5, dpi=100)\nplt.close()\nprint(f'✅ Animation saved: {out}  ({os.path.getsize(out)//1024:,} KB)')\nprint('→ Find it in the Files panel on the right to preview or download')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-07T13:33:34.514539Z","iopub.execute_input":"2026-03-07T13:33:34.514903Z","iopub.status.idle":"2026-03-07T13:33:55.2778Z","shell.execute_reply.started":"2026-03-07T13:33:34.51488Z","shell.execute_reply":"2026-03-07T13:33:55.276832Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Final Output Summary","metadata":{}},{"cell_type":"code","source":"import os\n\noutput_files = [\n    ('/kaggle/working/input_preview.png',           'Input frame preview'),\n    ('/kaggle/working/slam_result.png',             'Trajectory + point cloud (static)'),\n    ('/kaggle/working/slam_animation_color.gif',    'Incremental SLAM animation (colored)'),\n    ('/kaggle/working/slam_combined_color.ply',     'Colored point cloud (MeshLab / CloudCompare)'),\n    ('/kaggle/working/slam_output/poses.txt',       'Camera poses (frame_id tx ty tz qx qy qz qw)'),\n]\n\nprint('=' * 65)\nprint('📁 Output Files')\nprint('=' * 65)\nfor path, desc in output_files:\n    if os.path.exists(path):\n        size = os.path.getsize(path)\n        size_str = f'{size/1e6:.1f} MB' if size > 1e6 else f'{size/1e3:.1f} KB'\n        print(f'  ✅ {os.path.basename(path):<40s} {size_str}')\n        print(f'     {desc}')\n    else:\n        print(f'  ❌ {os.path.basename(path):<40s} (not generated)')\n\nprint()\nprint('💡 Tips:')\nprint('  - Open .ply files with MeshLab or CloudCompare for interactive 3D viewing')\nprint('  - For real-time visualization, run the viser server locally')\nprint('  - Download files via the Files panel (right sidebar) in Kaggle')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-07T13:33:55.279938Z","iopub.execute_input":"2026-03-07T13:33:55.280219Z","iopub.status.idle":"2026-03-07T13:33:55.28697Z","shell.execute_reply.started":"2026-03-07T13:33:55.280188Z","shell.execute_reply":"2026-03-07T13:33:55.286179Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Free GPU/CPU memory before rendering\nimport gc, torch\ntorch.cuda.empty_cache()\ngc.collect()\nprint(f\"🧹 Memory cleared  (GPU free: {torch.cuda.memory_reserved()/1e9:.1f} GB reserved)\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-07T13:33:55.288369Z","iopub.execute_input":"2026-03-07T13:33:55.288679Z","iopub.status.idle":"2026-03-07T13:33:57.022697Z","shell.execute_reply.started":"2026-03-07T13:33:55.288656Z","shell.execute_reply":"2026-03-07T13:33:57.022033Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 4. Display the converted GIF in the notebook\nfrom IPython.display import display, HTML\nOUTPUT_GIF=\"slam_animation_color.gif\"\nprint(f\"🖼️ Displaying the converted GIF:\")\ndisplay(HTML(f'<img src={OUTPUT_GIF}>'))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-07T13:33:57.02364Z","iopub.execute_input":"2026-03-07T13:33:57.024051Z","iopub.status.idle":"2026-03-07T13:33:57.029228Z","shell.execute_reply.started":"2026-03-07T13:33:57.024028Z","shell.execute_reply":"2026-03-07T13:33:57.028626Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## display_ply_viewer","metadata":{}},{"cell_type":"code","source":"!pip install -q open3d","metadata":{"trusted":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2026-03-07T13:33:57.030124Z","iopub.execute_input":"2026-03-07T13:33:57.030363Z","iopub.status.idle":"2026-03-07T13:34:00.750041Z","shell.execute_reply.started":"2026-03-07T13:33:57.030343Z","shell.execute_reply":"2026-03-07T13:34:00.749182Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\"\"\"\nPLY Viewer for Kaggle Notebook\nDisplay PLY files from /kaggle/working directory\n\"\"\"\n\nfrom IPython.display import HTML, display\nimport base64\n\ndef display_ply_viewer(ply_file_path):\n    \"\"\"\n    Display PLY file in Kaggle notebook\n    \n    Args:\n        ply_file_path: Path to PLY file (e.g., '/kaggle/working/model.ply')\n    \"\"\"\n    \n    # Read PLY file and encode to Base64\n    with open(ply_file_path, 'rb') as f:\n        ply_data = f.read()\n        ply_base64 = base64.b64encode(ply_data).decode('utf-8')\n    \n    # Create HTML viewer\n    html_content = f\"\"\"\n    <!DOCTYPE html>\n    <html>\n    <head>\n        <meta charset=\"UTF-8\">\n        <title>PLY Viewer</title>\n        <style>\n            body {{ margin: 0; font-family: Arial, sans-serif; }}\n            #container {{ width: 100%; height: 600px; position: relative; }}\n            #info {{\n                position: absolute;\n                top: 10px;\n                left: 10px;\n                background: rgba(0,0,0,0.7);\n                color: white;\n                padding: 10px;\n                border-radius: 5px;\n                font-size: 12px;\n                z-index: 100;\n            }}\n            .controls {{\n                position: absolute;\n                top: 10px;\n                right: 10px;\n                background: rgba(0,0,0,0.7);\n                padding: 10px;\n                border-radius: 5px;\n                z-index: 100;\n            }}\n            .button {{\n                background: #4CAF50;\n                color: white;\n                border: none;\n                padding: 8px 12px;\n                margin: 2px;\n                border-radius: 3px;\n                cursor: pointer;\n                font-size: 13px;\n            }}\n            .button:hover {{\n                background: #45a049;\n            }}\n        </style>\n    </head>\n    <body>\n        <div id=\"container\">\n            <div id=\"info\">Loading...</div>\n            <div class=\"controls\">\n                <button id=\"reset-view\" class=\"button\">Reset View</button>\n            </div>\n        </div>\n\n        <script type=\"importmap\">\n        {{\n          \"imports\": {{\n            \"three\": \"https://unpkg.com/three@0.160.0/build/three.module.js\",\n            \"three/examples/jsm/loaders/PLYLoader.js\": \"https://unpkg.com/three@0.160.0/examples/jsm/loaders/PLYLoader.js\",\n            \"three/examples/jsm/controls/OrbitControls.js\": \"https://unpkg.com/three@0.160.0/examples/jsm/controls/OrbitControls.js\"\n          }}\n        }}\n        </script>\n        <script type=\"module\">\n            import * as THREE from 'three';\n            import {{ PLYLoader }} from 'three/examples/jsm/loaders/PLYLoader.js';\n            import {{ OrbitControls }} from 'three/examples/jsm/controls/OrbitControls.js';\n\n            let scene, camera, renderer, controls;\n            let currentPointCloud = null;\n\n            function init() {{\n                const container = document.getElementById('container');\n                \n                scene = new THREE.Scene();\n                scene.background = new THREE.Color(0x1a1a1a);\n\n                camera = new THREE.PerspectiveCamera(60, container.clientWidth / container.clientHeight, 0.1, 10000);\n                camera.position.set(5, 5, 10);\n\n                renderer = new THREE.WebGLRenderer({{ antialias: true }});\n                renderer.setSize(container.clientWidth, container.clientHeight);\n                container.appendChild(renderer.domElement);\n\n                controls = new OrbitControls(camera, renderer.domElement);\n                controls.enableDamping = true;\n                controls.dampingFactor = 0.05;\n\n                const ambientLight = new THREE.AmbientLight(0xffffff, 0.6);\n                scene.add(ambientLight);\n\n                const directionalLight = new THREE.DirectionalLight(0xffffff, 0.8);\n                directionalLight.position.set(10, 20, 15);\n                scene.add(directionalLight);\n\n                const gridHelper = new THREE.GridHelper(100, 50, 0x444444, 0x222222);\n                scene.add(gridHelper);\n\n                document.getElementById('reset-view').addEventListener('click', resetView);\n\n                loadPLY();\n                animate();\n            }}\n\n            function loadPLY() {{\n                const plyBase64 = '{ply_base64}';\n                const binaryString = atob(plyBase64);\n                const bytes = new Uint8Array(binaryString.length);\n                for (let i = 0; i < binaryString.length; i++) {{\n                    bytes[i] = binaryString.charCodeAt(i);\n                }}\n\n                const loader = new PLYLoader();\n                try {{\n                    const geometry = loader.parse(bytes.buffer);\n                    displayGeometry(geometry);\n                }} catch (error) {{\n                    document.getElementById('info').innerHTML = 'Error: ' + error.message;\n                }}\n            }}\n\n            function displayGeometry(geometry) {{\n                const hasColors = geometry.attributes.color !== undefined;\n                \n                // Fixed point size: 0.002\n                const material = new THREE.PointsMaterial({{\n                    size: 0.002,\n                    vertexColors: hasColors,\n                    sizeAttenuation: true\n                }});\n\n                if (!hasColors) {{\n                    material.color = new THREE.Color(0x00aaff);\n                }}\n\n                const pointCloud = new THREE.Points(geometry, material);\n                scene.add(pointCloud);\n                currentPointCloud = pointCloud;\n\n                geometry.computeBoundingBox();\n                const bbox = geometry.boundingBox;\n                const center = new THREE.Vector3();\n                bbox.getCenter(center);\n                const size = bbox.getSize(new THREE.Vector3());\n                const maxDim = Math.max(size.x, size.y, size.z);\n\n                const fov = camera.fov * (Math.PI / 180);\n                let cameraDistance = Math.abs(maxDim / Math.tan(fov / 2)) * 1.5;\n                \n                camera.position.set(cameraDistance, cameraDistance * 0.7, cameraDistance);\n                controls.target.copy(center);\n                controls.update();\n\n                const pointCount = geometry.attributes.position.count;\n                document.getElementById('info').innerHTML = \n                    `<strong>Point Cloud</strong><br>\n                     Points: ${{pointCount.toLocaleString()}}<br>\n                     Size: (${{size.x.toFixed(2)}}, ${{size.y.toFixed(2)}}, ${{size.z.toFixed(2)}})`;\n            }}\n\n            function resetView() {{\n                if (currentPointCloud) {{\n                    currentPointCloud.geometry.computeBoundingBox();\n                    const bbox = currentPointCloud.geometry.boundingBox;\n                    const center = new THREE.Vector3();\n                    bbox.getCenter(center);\n                    const size = bbox.getSize(new THREE.Vector3());\n                    const maxDim = Math.max(size.x, size.y, size.z);\n                    \n                    const fov = camera.fov * (Math.PI / 180);\n                    let cameraDistance = Math.abs(maxDim / Math.tan(fov / 2)) * 1.5;\n                    \n                    camera.position.set(cameraDistance, cameraDistance * 0.7, cameraDistance);\n                    controls.target.copy(center);\n                    controls.update();\n                }}\n            }}\n\n            function animate() {{\n                requestAnimationFrame(animate);\n                controls.update();\n                renderer.render(scene, camera);\n            }}\n\n            init();\n        </script>\n    </body>\n    </html>\n    \"\"\"\n    \n    display(HTML(html_content))\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-07T13:34:00.751599Z","iopub.execute_input":"2026-03-07T13:34:00.752292Z","iopub.status.idle":"2026-03-07T13:34:00.760988Z","shell.execute_reply.started":"2026-03-07T13:34:00.752264Z","shell.execute_reply":"2026-03-07T13:34:00.760412Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# you can change position and zoom ratio.\n#display_ply_viewer('/kaggle/working/slam_combined_color.ply')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-07T13:34:00.761803Z","iopub.execute_input":"2026-03-07T13:34:00.762034Z","iopub.status.idle":"2026-03-07T13:34:00.776404Z","shell.execute_reply.started":"2026-03-07T13:34:00.762003Z","shell.execute_reply":"2026-03-07T13:34:00.7758Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## convert_slam_to_colmap","metadata":{}},{"cell_type":"code","source":"# ============================================================================\n# COLMAP binary writers (unchanged from original)\n# ============================================================================\n\nimport numpy as np\nimport struct\nfrom pathlib import Path","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-07T13:34:00.778397Z","iopub.execute_input":"2026-03-07T13:34:00.778665Z","iopub.status.idle":"2026-03-07T13:34:00.788274Z","shell.execute_reply.started":"2026-03-07T13:34:00.778645Z","shell.execute_reply":"2026-03-07T13:34:00.78758Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\nFRAMES_DIR = '/kaggle/working/input_frames'\nimage_dir  = FRAMES_DIR\noutput_dir = \"/kaggle/working/output\"\n\n# ── Step 2 (Process3): VGGT → COLMAP ───────────────────────────────────────\nprint(\"\\n\" + \"=\" * 70)\nprint(\"Step 2: VGGT → COLMAP Conversion\")\nprint(\"=\" * 70)\n\ncolmap_dir = os.path.join(output_dir, \"colmap\")\n\nsparse_dir = convert_slam_to_colmap(\n                npz_files  = npz_files,\n                img_dir    = IMG_DIR,\n                poses_txt  = '/kaggle/working/slam_output/poses.txt',\n                output_dir = colmap_dir,\n                verbose    = True)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-07T13:34:00.789187Z","iopub.execute_input":"2026-03-07T13:34:00.789549Z","iopub.status.idle":"2026-03-07T13:34:20.29231Z","shell.execute_reply.started":"2026-03-07T13:34:00.78952Z","shell.execute_reply":"2026-03-07T13:34:20.291679Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"!ls /kaggle/working/output/colmap/sparse/0","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-07T13:34:20.293186Z","iopub.execute_input":"2026-03-07T13:34:20.293423Z","iopub.status.idle":"2026-03-07T13:34:20.434439Z","shell.execute_reply.started":"2026-03-07T13:34:20.293402Z","shell.execute_reply":"2026-03-07T13:34:20.433572Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}