{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.12.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":71885,"databundleVersionId":8143495,"sourceType":"competition"}],"dockerImageVersionId":31259,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# *Image Matching Challenge 2024 – 2D to 3D Reconstruction Pipeline*","metadata":{}},{"cell_type":"markdown","source":"\n\n## Overview\nThis notebook implements a classical Structure-from-Motion (SfM) pipeline using:\n\n- SIFT feature extraction\n- BFMatcher with Lowe Ratio Test\n- Essential Matrix estimation via RANSAC\n- Camera pose recovery\n- 3D triangulation\n\nThe goal is to reconstruct camera poses and sparse 3D structure from unordered image pairs.\n\n---\n\n## Pipeline Steps\n\n1. Load dataset images\n2. Extract local features (SIFT)\n3. Perform feature matching\n4. Estimate relative pose\n5. Triangulate 3D points\n6. Export submission file\n7. Save deployable model configuration\n\n---\n\nThis implementation is designed to be:\n\n- Internet-free (Kaggle compliant)\n- Lightweight\n- Suitable for Streamlit / HuggingFace deployment","metadata":{}},{"cell_type":"markdown","source":"## *Import Libraries*","metadata":{}},{"cell_type":"code","source":"import os\nimport cv2\nimport numpy as np\nimport pandas as pd\nfrom glob import glob\nfrom tqdm import tqdm\nimport matplotlib.pyplot as plt","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-13T16:13:24.673062Z","iopub.execute_input":"2026-02-13T16:13:24.673498Z","iopub.status.idle":"2026-02-13T16:13:24.678979Z","shell.execute_reply.started":"2026-02-13T16:13:24.673467Z","shell.execute_reply":"2026-02-13T16:13:24.677828Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## *Dataset Load*","metadata":{}},{"cell_type":"code","source":"DATA_PATH = \"/kaggle/input/image-matching-challenge-2024\"\n\nimage_paths = sorted(\n    glob(f\"{DATA_PATH}/**/images/*.png\", recursive=True)\n)\n\nprint(\"Total images:\", len(image_paths))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-13T16:13:35.708314Z","iopub.execute_input":"2026-02-13T16:13:35.708652Z","iopub.status.idle":"2026-02-13T16:13:37.137789Z","shell.execute_reply.started":"2026-02-13T16:13:35.708623Z","shell.execute_reply":"2026-02-13T16:13:37.136921Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## *Eda*","metadata":{}},{"cell_type":"code","source":"resolutions = []\n\nfor img_path in image_paths[:200]:\n    img = cv2.imread(img_path)\n    if img is not None:\n        h, w = img.shape[:2]\n        resolutions.append(h*w)\n\nplt.hist(resolutions, bins=30)\nplt.title(\"Image Resolution Distribution\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-13T16:14:07.741860Z","iopub.execute_input":"2026-02-13T16:14:07.742260Z","iopub.status.idle":"2026-02-13T16:14:15.869212Z","shell.execute_reply.started":"2026-02-13T16:14:07.742227Z","shell.execute_reply":"2026-02-13T16:14:15.868126Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## *Feature Matching*","metadata":{}},{"cell_type":"code","source":"sift = cv2.SIFT_create(nfeatures=3000)\n\ndef match_images(img_path1, img_path2):\n    img1 = cv2.imread(img_path1, cv2.IMREAD_GRAYSCALE)\n    img2 = cv2.imread(img_path2, cv2.IMREAD_GRAYSCALE)\n\n    kp1, des1 = sift.detectAndCompute(img1, None)\n    kp2, des2 = sift.detectAndCompute(img2, None)\n\n    if des1 is None or des2 is None:\n        return None, None\n\n    bf = cv2.BFMatcher(cv2.NORM_L2, crossCheck=True)\n    matches = bf.match(des1, des2)\n\n    matches = sorted(matches, key=lambda x: x.distance)\n\n    pts1 = np.float32([kp1[m.queryIdx].pt for m in matches])\n    pts2 = np.float32([kp2[m.trainIdx].pt for m in matches])\n\n    return pts1, pts2","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-13T16:14:59.852681Z","iopub.execute_input":"2026-02-13T16:14:59.853048Z","iopub.status.idle":"2026-02-13T16:14:59.866041Z","shell.execute_reply.started":"2026-02-13T16:14:59.853018Z","shell.execute_reply":"2026-02-13T16:14:59.864968Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## *Pose Estimation*","metadata":{}},{"cell_type":"code","source":"def estimate_pose(pts1, pts2):\n    if pts1 is None or len(pts1) < 8:\n        return None, None\n\n    E, mask = cv2.findEssentialMat(\n        pts1,\n        pts2,\n        method=cv2.RANSAC,\n        prob=0.999,\n        threshold=1.0\n    )\n\n    if E is None:\n        return None, None\n\n    _, R, t, _ = cv2.recoverPose(E, pts1, pts2)\n    return R, t","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-13T16:15:23.559807Z","iopub.execute_input":"2026-02-13T16:15:23.560222Z","iopub.status.idle":"2026-02-13T16:15:23.566537Z","shell.execute_reply.started":"2026-02-13T16:15:23.560189Z","shell.execute_reply":"2026-02-13T16:15:23.565711Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## *Incremental Pose Building*","metadata":{}},{"cell_type":"code","source":"poses = {}\nposes[image_paths[0]] = (np.eye(3), np.zeros((3,1)))\n\nfor i in tqdm(range(len(image_paths)-1)):\n    img1 = image_paths[i]\n    img2 = image_paths[i+1]\n\n    pts1, pts2 = match_images(img1, img2)\n    R, t = estimate_pose(pts1, pts2)\n\n    if R is not None:\n        poses[img2] = (R, t)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-13T16:15:48.793249Z","iopub.execute_input":"2026-02-13T16:15:48.794140Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## *Submission*","metadata":{}},{"cell_type":"code","source":"rows = []\n\nfor path in image_paths:\n    if path in poses:\n        R, t = poses[path]\n        R_flat = \";\".join(map(str, R.flatten()))\n        t_flat = \";\".join(map(str, t.flatten()))\n    else:\n        R_flat = \"nan;\"*8 + \"nan\"\n        t_flat = \"nan;nan;nan\"\n\n    parts = path.split(\"/\")\n    dataset = parts[-4]\n    scene = parts[-3]\n\n    rows.append([path, dataset, scene, R_flat, t_flat])\n\nsubmission = pd.DataFrame(\n    rows,\n    columns=[\n        \"image_path\",\n        \"dataset\",\n        \"scene\",\n        \"rotation_matrix\",\n        \"translation_vector\"\n    ]\n)\n\nsubmission.to_csv(\"submission.csv\", index=False)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## *Lowe Ratio Test*","metadata":{}},{"cell_type":"code","source":"def match_images(img_path1, img_path2):\n    img1 = cv2.imread(img_path1, cv2.IMREAD_GRAYSCALE)\n    img2 = cv2.imread(img_path2, cv2.IMREAD_GRAYSCALE)\n\n    sift = cv2.SIFT_create(nfeatures=3000)\n\n    kp1, des1 = sift.detectAndCompute(img1, None)\n    kp2, des2 = sift.detectAndCompute(img2, None)\n\n    bf = cv2.BFMatcher()\n\n    matches = bf.knnMatch(des1, des2, k=2)\n\n    good = []\n    for m, n in matches:\n        if m.distance < 0.75 * n.distance:\n            good.append(m)\n\n    pts1 = np.float32([kp1[m.queryIdx].pt for m in good])\n    pts2 = np.float32([kp2[m.trainIdx].pt for m in good])\n\n    return pts1, pts2","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## *Triangulation*","metadata":{}},{"cell_type":"code","source":"def triangulate_points(R, t, pts1, pts2):\n    P1 = np.hstack((np.eye(3), np.zeros((3,1))))\n    P2 = np.hstack((R, t))\n\n    points_4d = cv2.triangulatePoints(\n        P1, P2,\n        pts1.T,\n        pts2.T\n    )\n\n    points_3d = points_4d[:3] / points_4d[3]\n    return points_3d.T","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## *3D*","metadata":{}},{"cell_type":"code","source":"import plotly.express as px\n\ndef visualize_3d(points_3d):\n    fig = px.scatter_3d(\n        x=points_3d[:,0],\n        y=points_3d[:,1],\n        z=points_3d[:,2]\n    )\n    fig.show()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## *Save Model*","metadata":{}},{"cell_type":"code","source":"import json\nimport joblib\nimport cv2\n\n# Model parametreleri\nmodel_config = {\n    \"feature_detector\": \"SIFT\",\n    \"nfeatures\": 3000,\n    \"ratio_test_threshold\": 0.75,\n    \"ransac_threshold\": 1.0,\n    \"min_match_count\": 8\n}\n\n# Config kaydet\nwith open(\"model_config.json\", \"w\") as f:\n    json.dump(model_config, f, indent=4)\n\n# SIFT objesini kaydetmeye gerek yok ama sembolik olarak dump edelim\nsift = cv2.SIFT_create(nfeatures=3000)\njoblib.dump(model_config, \"sfm_model.pkl\")\n\nprint(\"Model saved successfully.\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Final Summary\n\n## What Was Implemented\n\n- Robust feature matching using Lowe Ratio Test\n- Essential matrix estimation with RANSAC\n- Camera pose recovery\n- 3D point triangulation\n- Submission file generation\n- Deployment-ready model configuration export\n\n---\n\n## Limitations\n\n- No global pose graph optimization\n- No bundle adjustment\n- Sequential matching strategy\n\n---\n\n## Deployment\n\nThe saved model configuration can be used inside a Streamlit application for:\n\n- Interactive 2D image upload\n- Real-time 3D reconstruction\n- Visualization using Plotly\n\nThis completes a deployable end-to-end classical SfM pipeline.","metadata":{}}]}