{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.11.13"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":49349,"databundleVersionId":5447706,"sourceType":"competition"}],"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true},"papermill":{"default_parameters":{},"duration":1485.608701,"end_time":"2025-09-29T08:34:31.655338","environment_variables":{},"exception":null,"input_path":"__notebook__.ipynb","output_path":"__notebook__.ipynb","parameters":{},"start_time":"2025-09-29T08:09:46.046637","version":"2.6.0"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **IMC2023 Fountain: Gaussian Splatting**\n\n\nhttps://www.kaggle.com/code/stpeteishii/imc2023-fountain-gaussian-splatting","metadata":{"papermill":{"duration":0.002721,"end_time":"2025-09-29T08:09:50.553226","exception":false,"start_time":"2025-09-29T08:09:50.550505","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"This script implements a complete 3D reconstruction pipeline using Gaussian Splatting on Kaggle. Here's how it works:","metadata":{"papermill":{"duration":0.002073,"end_time":"2025-09-29T08:09:50.557816","exception":false,"start_time":"2025-09-29T08:09:50.555743","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"## **Pipeline Overview**\n\n**Input:** Multi-view images + camera poses from COLMAP Structure-from-Motion\n**Output:** Novel view synthesis video rendered from a trained 3D Gaussian model\n\n## Step-by-Step Process\n\n### 1. Environment Setup\n- Clones the official 3D Gaussian Splatting repository\n- Installs PyTorch and dependencies\n- Compiles CUDA extensions: `diff-gaussian-rasterization` (custom rasterizer) and `simple-knn` (k-nearest neighbors)\n\n### 2. Data Preparation\n**Challenge:** The dataset uses COLMAP's SIMPLE_RADIAL camera model (focal length + 1 radial distortion parameter), but Gaussian Splatting requires PINHOLE (no distortion).\n\n**Solution:** Convert camera models by:\n- Extracting focal length (f), principal point (cx, cy) from SIMPLE_RADIAL\n- Dropping the radial distortion coefficient\n- Writing as PINHOLE format: `fx=f, fy=f, cx, cy`\n\n**Structure created:**\nActual Data Structure:\n```\n/kaggle/input/image-matching-challenge-2023/train/haiper/fountain/\n├── images/          # image subset\n├── images_full/     # all 231 images (actually used)\n└── sfm/             # COLMAP SfM results\n    ├── cameras.txt  # camera parameters\n    ├── images.txt   # camera poses\n    └── points3D.txt # 3D point cloud\n```\n\nProcessing Flow:\n1. Copy all images from images_full/\n2. Convert camera data from sfm/ to PINHOLE format\n3. Build data structure for Gaussian Splatting\n\n### 3. Training\nTrains a 3D Gaussian Splatting model for 3000 iterations (~12 minutes on Kaggle GPU):\n\n**What it learns:**\n- Position, rotation, scale of each 3D Gaussian primitive\n- Opacity and color (spherical harmonics) of each Gaussian\n- Starts from ~123k sparse COLMAP points, optimizes to represent the scene\n\n**Key parameters:**\n- Adaptive density control (splits/clones Gaussians in under-represented areas)\n- Photometric loss against training views\n- Uses LLFF dataset split (8:1 train/test ratio)\n\n### 4. Rendering\nRenders novel views from the trained model:\n- Loads trained Gaussians at iteration 3000\n- Renders 202 training views + 29 test views\n- Uses differentiable tile-based rasterization (GPU-accelerated)\n- Outputs 1440×1920 PNG images\n\n### 5. Video Generation\n- Uses ffmpeg to compile rendered frames into MP4 (30 fps, H.264)\n- Converts to GIF for notebook display (10 fps, downscaled to 720px width)\n- Displays inline using IPython.display\n\n## Technical Challenges Solved\n\n1. **Camera model mismatch** - COLMAP format vs Gaussian Splatting requirements\n2. **Data path confusion** - Found correct image folder (images_full vs images)\n3. **Render output location** - Discovered test/ours_3000/renders structure\n4. **Memory constraints** - Reduced iterations (3000 vs default 7000+) for Kaggle limits\n\n## Why This Works\n\nGaussian Splatting represents scenes as millions of 3D Gaussians (ellipsoids) rather than neural networks. This enables:\n- Fast training (minutes vs hours for NeRF)\n- Real-time rendering capability\n- Explicit 3D representation\n- High-quality novel view synthesis\n\nThe script successfully bridges the gap between standard COLMAP reconstructions and Gaussian Splatting, making it accessible on Kaggle's platform.","metadata":{"papermill":{"duration":0.001872,"end_time":"2025-09-29T08:09:50.561564","exception":false,"start_time":"2025-09-29T08:09:50.559692","status":"completed"},"tags":[]}},{"cell_type":"code","source":"\"\"\"\nGaussian Splatting Video Generator for Kaggle\nThis script processes images and camera data to create a Gaussian Splatting video\n\"\"\"\n\nimport os\nimport sys\nimport subprocess\nimport shutil\nfrom pathlib import Path\n\n# Configuration\nINPUT_PATH = '/kaggle/input/image-matching-challenge-2023/train/haiper/fountain'\nWORK_DIR = '/kaggle/working/gaussian_splatting'\nOUTPUT_DIR = '/kaggle/working/output'\n\ndef setup_environment():\n    \"\"\"Install required packages and clone Gaussian Splatting repository\"\"\"\n    print(\"Setting up environment...\")\n    \n    # Clone 3D Gaussian Splatting repository\n    if not os.path.exists(WORK_DIR):\n        print(\"Cloning Gaussian Splatting repository...\")\n        subprocess.run([\n            'git', 'clone', '--recursive',\n            'https://github.com/graphdeco-inria/gaussian-splatting.git',\n            WORK_DIR\n        ], check=True)\n    \n    os.chdir(WORK_DIR)\n    \n    # Install pip packages\n    print(\"Installing Python packages...\")\n    subprocess.run([sys.executable, '-m', 'pip', 'install', '-q', 'torch', 'torchvision', \n                    'torchaudio', 'plyfile', 'tqdm', 'opencv-python', 'pillow'], check=True)\n    \n    # Build submodules\n    print(\"Building submodules...\")\n    subprocess.run([sys.executable, '-m', 'pip', 'install', 'submodules/diff-gaussian-rasterization'], \n                   check=True, cwd=WORK_DIR)\n    subprocess.run([sys.executable, '-m', 'pip', 'install', 'submodules/simple-knn'], \n                   check=True, cwd=WORK_DIR)\n\ndef convert_cameras_to_pinhole(input_file, output_file):\n    \"\"\"Convert camera models to PINHOLE format\"\"\"\n    print(f\"Reading camera file: {input_file}\")\n    \n    with open(input_file, 'r') as f:\n        lines = f.readlines()\n    \n    converted_count = 0\n    with open(output_file, 'w') as f:\n        for line in lines:\n            if line.startswith('#') or line.strip() == '':\n                f.write(line)\n            else:\n                parts = line.strip().split()\n                if len(parts) >= 4:\n                    cam_id = parts[0]\n                    model = parts[1]\n                    width = parts[2]\n                    height = parts[3]\n                    params = parts[4:]\n                    \n                    print(f\"Camera {cam_id}: model={model}, size={width}x{height}, params={len(params)}\")\n                    \n                    # Convert to PINHOLE (fx, fy, cx, cy)\n                    if model == \"PINHOLE\":\n                        f.write(line)\n                    elif model == \"SIMPLE_PINHOLE\":\n                        # SIMPLE_PINHOLE: f, cx, cy -> fx, fy, cx, cy\n                        f_val = params[0]\n                        cx = params[1]\n                        cy = params[2]\n                        f.write(f\"{cam_id} PINHOLE {width} {height} {f_val} {f_val} {cx} {cy}\\n\")\n                        converted_count += 1\n                    elif model == \"SIMPLE_RADIAL\":\n                        # SIMPLE_RADIAL: f, cx, cy, k -> ignore distortion, use f, cx, cy\n                        f_val = params[0]\n                        cx = params[1]\n                        cy = params[2]\n                        # k = params[3]  # radial distortion, ignored for PINHOLE\n                        f.write(f\"{cam_id} PINHOLE {width} {height} {f_val} {f_val} {cx} {cy}\\n\")\n                        converted_count += 1\n                    elif model == \"RADIAL\":\n                        # RADIAL: f, cx, cy, k1, k2 -> use f, cx, cy\n                        f_val = params[0]\n                        cx = params[1]\n                        cy = params[2]\n                        f.write(f\"{cam_id} PINHOLE {width} {height} {f_val} {f_val} {cx} {cy}\\n\")\n                        converted_count += 1\n                    elif model in [\"OPENCV\", \"OPENCV_FISHEYE\", \"FULL_OPENCV\"]:\n                        # Use fx, fy, cx, cy from parameters\n                        fx = params[0]\n                        fy = params[1]\n                        cx = params[2]\n                        cy = params[3]\n                        f.write(f\"{cam_id} PINHOLE {width} {height} {fx} {fy} {cx} {cy}\\n\")\n                        converted_count += 1\n                    else:\n                        # Default: estimate from image size\n                        print(f\"Warning: Unknown model '{model}', using estimated parameters\")\n                        fx = fy = max(float(width), float(height))\n                        cx = float(width) / 2\n                        cy = float(height) / 2\n                        f.write(f\"{cam_id} PINHOLE {width} {height} {fx} {fy} {cx} {cy}\\n\")\n                        converted_count += 1\n                else:\n                    f.write(line)\n    \n    print(f\"Converted {converted_count} cameras to PINHOLE format\")\n\ndef prepare_colmap_data():\n    \"\"\"Convert SfM data to COLMAP format\"\"\"\n    print(\"Preparing COLMAP data...\")\n    \n    data_dir = f\"{WORK_DIR}/data/fountain\"\n    os.makedirs(f\"{data_dir}/sparse/0\", exist_ok=True)\n    os.makedirs(f\"{data_dir}/images\", exist_ok=True)\n    \n    # Clarify image copying logic\n    print(\"Copying images...\")\n    src_images = f\"{INPUT_PATH}/images_full\"\n    \n    if not os.path.exists(src_images):\n        print(f\"Warning: {src_images} not found, trying /images instead\")\n        src_images = f\"{INPUT_PATH}/images\"\n    \n    if not os.path.exists(src_images):\n        raise FileNotFoundError(f\"Image directory not found: {src_images}\")\n    \n    img_count = 0\n    for img_file in os.listdir(src_images):\n        if img_file.lower().endswith(('.jpeg', '.jpg', '.png')):\n            shutil.copy(\n                f\"{src_images}/{img_file}\",\n                f\"{data_dir}/images/{img_file}\"\n            )\n            img_count += 1\n    \n    print(f\"Copied {img_count} images from {src_images}\")\n\n    # Read and convert camera file to PINHOLE format\n    print(\"Converting camera data to PINHOLE format...\")\n    convert_cameras_to_pinhole(\n        f\"{INPUT_PATH}/sfm/cameras.txt\",\n        f\"{data_dir}/sparse/0/cameras.txt\"\n    )\n    \n    # Copy other SfM files\n    shutil.copy(f\"{INPUT_PATH}/sfm/images.txt\", f\"{data_dir}/sparse/0/images.txt\")\n    shutil.copy(f\"{INPUT_PATH}/sfm/points3D.txt\", f\"{data_dir}/sparse/0/points3D.txt\")\n    \n    print(f\"Data prepared at: {data_dir}\")\n    \n    return data_dir\n\ndef train_gaussian_splatting(data_dir, iterations=3000):\n    \"\"\"Train Gaussian Splatting model\"\"\"\n    print(f\"Training Gaussian Splatting model for {iterations} iterations...\")\n    \n    model_path = f\"{WORK_DIR}/output/fountain\"\n    \n    cmd = [\n        sys.executable, 'train.py',\n        '-s', data_dir,\n        '-m', model_path,\n        '--iterations', str(iterations),\n        '--eval'\n    ]\n    \n    subprocess.run(cmd, cwd=WORK_DIR, check=True)\n    \n    return model_path\n\ndef render_video(model_path, output_video_path, iteration=3000):\n    \"\"\"Render video from trained model\"\"\"\n    print(\"Rendering video...\")\n    \n    # Render images - render both train and test cameras\n    cmd = [\n        sys.executable, 'render.py',\n        '-m', model_path,\n        '--iteration', str(iteration)\n    ]\n    \n    print(f\"Running render command: {' '.join(cmd)}\")\n    subprocess.run(cmd, cwd=WORK_DIR, check=True)\n    \n    # Create video from rendered images using ffmpeg\n    print(\"Creating video from rendered frames...\")\n    \n    # Find the correct render directory - check multiple possible locations\n    possible_dirs = [\n        f\"{model_path}/test/ours_{iteration}/renders\",\n        f\"{model_path}/test/ours_{iteration}/gt\",\n        f\"{model_path}/train/ours_{iteration}/renders\",\n        f\"{model_path}/render/ours_{iteration}/renders\"\n    ]\n    \n    render_dir = None\n    for test_dir in possible_dirs:\n        if os.path.exists(test_dir):\n            render_dir = test_dir\n            print(f\"Found render directory: {render_dir}\")\n            break\n    \n    if not render_dir:\n        print(f\"Warning: Render directories not found. Searched:\")\n        for d in possible_dirs:\n            print(f\"  - {d}\")\n        \n        # Try to explore what directories exist\n        test_path = f\"{model_path}/test\"\n        if os.path.exists(test_path):\n            subdirs = os.listdir(test_path)\n            print(f\"Available directories in {test_path}: {subdirs}\")\n            for subdir in subdirs:\n                full_path = os.path.join(test_path, subdir)\n                if os.path.isdir(full_path):\n                    contents = os.listdir(full_path)\n                    print(f\"  {subdir}/: {contents}\")\n                    if 'renders' in contents:\n                        render_dir = os.path.join(full_path, 'renders')\n                        print(f\"Using: {render_dir}\")\n                        break\n    \n    if render_dir and os.path.exists(render_dir):\n        render_imgs = sorted([f for f in os.listdir(render_dir) if f.endswith('.png')])\n        \n        if render_imgs:\n            print(f\"Found {len(render_imgs)} rendered images\")\n            # Create video with ffmpeg\n            subprocess.run([\n                'ffmpeg', '-y',\n                '-framerate', '30',\n                '-pattern_type', 'glob',\n                '-i', f\"{render_dir}/*.png\",\n                '-c:v', 'libx264',\n                '-pix_fmt', 'yuv420p',\n                '-crf', '18',\n                output_video_path\n            ], check=True)\n            \n            print(f\"Video saved to: {output_video_path}\")\n            return True\n        else:\n            print(\"Warning: No rendered PNG images found!\")\n            return False\n    else:\n        print(f\"Error: Could not find render directory\")\n        return False\n\ndef create_gif_and_display(video_path, gif_path):\n    \"\"\"Convert MP4 to GIF and display in notebook\"\"\"\n    print(\"Creating animated GIF...\")\n    \n    # Convert MP4 to GIF using ffmpeg\n    subprocess.run([\n        'ffmpeg', '-y',\n        '-i', video_path,\n        '-vf', 'setpts=8*PTS,fps=10,scale=720:-1:flags=lanczos',\n        '-loop', '0',\n        gif_path\n    ], check=True)\n\n    if os.path.exists(gif_path):\n        size_mb = os.path.getsize(gif_path) / (1024 * 1024)\n        print(f\"GIF created: {gif_path} ({size_mb:.2f} MB)\")\n        \n        # Display animated GIF in notebook\n        #print(\"\\nDisplaying animation:\")\n        #from IPython.display import Image\n        #Image(open(gif_path, 'rb').read())\n        \n        return True\n\n    return False\n\ndef main():\n    \"\"\"Main execution function\"\"\"\n    print(\"=\"*60)\n    print(\"Gaussian Splatting Video Generator\")\n    print(\"=\"*60)\n    \n    try:\n        # Step 1: Setup environment\n        setup_environment()\n        \n        # Step 2: Prepare data\n        data_dir = prepare_colmap_data()\n        \n        # Step 3: Train model (reduce iterations for faster processing on Kaggle)\n        model_path = train_gaussian_splatting(data_dir, iterations=3000)\n        \n        # Step 4: Render video\n        os.makedirs(OUTPUT_DIR, exist_ok=True)\n        output_video = f\"{OUTPUT_DIR}/gaussian_splatting_fountain.mp4\"\n        success = render_video(model_path, output_video, iteration=3000)\n        \n        if success:\n            print(\"=\"*60)\n            print(f\"SUCCESS! Video generated at: {output_video}\")\n            print(\"=\"*60)\n            \n            # Display video info\n            if os.path.exists(output_video):\n                size_mb = os.path.getsize(output_video) / (1024 * 1024)\n                print(f\"Video size: {size_mb:.2f} MB\")\n            \n            # Step 5: Create and display GIF\n            output_gif = f\"{OUTPUT_DIR}/gaussian_splatting_fountain.gif\"\n            create_gif_and_display(output_video, output_gif)\n            \n        else:\n            print(\"=\"*60)\n            print(\"WARNING: Rendering completed but no video was generated.\")\n            print(\"The model was trained successfully. Check the output directories.\")\n            print(\"=\"*60)\n        \n    except Exception as e:\n        print(f\"ERROR: {str(e)}\")\n        import traceback\n        traceback.print_exc()\n\nif __name__ == \"__main__\":\n    main()","metadata":{"execution":{"iopub.execute_input":"2025-09-29T08:09:50.566802Z","iopub.status.busy":"2025-09-29T08:09:50.56652Z","iopub.status.idle":"2025-09-29T08:34:30.400415Z","shell.execute_reply":"2025-09-29T08:34:30.399559Z"},"papermill":{"duration":1479.838375,"end_time":"2025-09-29T08:34:30.401821","exception":false,"start_time":"2025-09-29T08:09:50.563446","status":"completed"},"tags":[]},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import gc\ngc.collect()","metadata":{"execution":{"iopub.execute_input":"2025-09-29T08:34:30.485358Z","iopub.status.busy":"2025-09-29T08:34:30.48495Z","iopub.status.idle":"2025-09-29T08:34:30.543466Z","shell.execute_reply":"2025-09-29T08:34:30.542669Z"},"papermill":{"duration":0.107852,"end_time":"2025-09-29T08:34:30.545179","exception":false,"start_time":"2025-09-29T08:34:30.437327","status":"completed"},"tags":[]},"outputs":[],"execution_count":null},{"cell_type":"code","source":"gif_path='/kaggle/working/output/gaussian_splatting_fountain.gif'\nfrom IPython.display import Image\nImage(open(gif_path, 'rb').read())","metadata":{"execution":{"iopub.execute_input":"2025-09-29T08:34:30.602727Z","iopub.status.busy":"2025-09-29T08:34:30.602485Z","iopub.status.idle":"2025-09-29T08:34:30.774166Z","shell.execute_reply":"2025-09-29T08:34:30.772912Z"},"papermill":{"duration":0.335421,"end_time":"2025-09-29T08:34:30.91022","exception":false,"start_time":"2025-09-29T08:34:30.574799","status":"completed"},"tags":[]},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Technical Challenges\n\n### 1. Camera Model Conversion\nThe most difficult part was bridging the gap between COLMAP's diverse camera models and Gaussian Splatting's requirements.\n\n- SIMPLE_RADIAL has 4 parameters (f, cx, cy, k)\n- PINHOLE has 4 parameters (fx, fy, cx, cy)\n\nSimply dropping the distortion coefficient k can reduce reconstruction accuracy for images with significant lens distortion. Ideally, images should be undistorted beforehand, but we simplified by skipping this step.\n\n### 2. Understanding Data Structure\nThe Kaggle dataset had `images/` (23 images) and `images_full/` (231 images), with SfM data referencing all 231. It took time to identify this mismatch.\n\n### 3. Locating Render Output\nGaussian Splatting's render.py can output to multiple directories, making it difficult to determine where the PNGs were actually generated.\n\n## What If You Only Have Photos?\n\n**If you only have photos without camera information, additional steps are required:**\n\n### Option 1: Run SfM with COLMAP\n```python\n# 1. Feature extraction\nsubprocess.run([\n    'colmap', 'feature_extractor',\n    '--database_path', 'database.db',\n    '--image_path', 'images/'\n])\n\n# 2. Feature matching\nsubprocess.run([\n    'colmap', 'exhaustive_matcher',\n    '--database_path', 'database.db'\n])\n\n# 3. Sparse reconstruction\nsubprocess.run([\n    'colmap', 'mapper',\n    '--database_path', 'database.db',\n    '--image_path', 'images/',\n    '--output_path', 'sparse/'\n])\n```\n\nCamera models are automatically estimated (typically SIMPLE_RADIAL or RADIAL).\n\n### Option 2: Commercial Software\nUse Metashape or RealityCapture for more accurate SfM, then export results to COLMAP format.\n\n### Option 3: Modern Approaches\n- **InstantNGP/Nerfstudio**: Includes camera pose estimation\n- **COLMAP + Gaussian Splatting integrated scripts**: Some implementations run SfM automatically\n\n### Challenges and Limitations\n\n**Issues when starting from photos only:**\n\n1. **Computation time**: SfM can take hours depending on image count\n2. **Failure risk**: SfM fails with low texture, reflections, or moving objects\n3. **Scale ambiguity**: Monocular cameras don't provide absolute scale (no metric units)\n4. **Memory limits**: Kaggle's free tier struggles with large datasets (500+ images)\n\n**Required photo conditions:**\n- 50%+ overlap between adjacent images\n- Varied viewing angles\n- Consistent lighting (minimal exposure changes)\n- Sharp, blur-free images\n- Static scenes (no moving objects)\n\nThis script assumes SfM is already complete. To start from photos only, you'd need to add the COLMAP preprocessing steps.","metadata":{"papermill":{"duration":0.143791,"end_time":"2025-09-29T08:34:31.20492","exception":false,"start_time":"2025-09-29T08:34:31.061129","status":"completed"},"tags":[]}}]}