{"metadata":{"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.12.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Improved ACCIDENT @ CVPR Pipeline\n## Original 10-Video Calibration Comparison\n\n**Project goal:** compare a held-out synthetic-supervised extension with the original notebook baseline on exactly the same 10 calibration videos.\n\n**Improved method:**\n1. Original frame-difference temporal localization\n2. Gated YOLO vehicle-pair spatial estimate blended with optical flow\n3. Logistic-regression probe on CLIP ViT-L/14 embeddings\n4. Full remaining synthetic training set; the 10 evaluation videos are excluded\n\n**Kaggle setup:**\n- Attach the `accident` competition dataset\n- Enable a **GPU** (T4) and Internet for first-run package/model downloads\n- Restart and run all cells; no 2,027-video test inference or submission is required\n","metadata":{}},{"cell_type":"markdown","source":"## 0. Environment Setup\n","metadata":{}},{"cell_type":"code","source":"# [CONFIG] Course-project comparison against original notebook\nRUN_FULL_TEST = False\nCALIBRATION_N = 10             # Same first 10 synthetic videos as original code\nSEED = 42\n\nYOLO_MODEL = 'yolov8n.pt'\nYOLO_CONF = 0.20\nYOLO_CLASSES = [2, 3, 5, 7]\nYOLO_IMGSZ = 512\nYOLO_VID_STRIDE = 3\nCLIP_BACKBONE = 'ViT-L/14'     # Match the original notebook exactly\nCLIP_CONTEXT_FRAMES = 8\nTRAIN_CONTEXT_FRAMES = 3       # Full train set, 3 frames/video for practical runtime\n\n# Spatial gate / blend\nYOLO_MAX_PAIR_DIST = 0.12\nYOLO_SPATIAL_BLEND = 0.65\n\nprint('[STATUS] Config ready — original 10-video calibration comparison')\nprint(f'  CALIBRATION_N={CALIBRATION_N} | CLIP={CLIP_BACKBONE} | YOLO={YOLO_MODEL}')\nprint('  Baseline inference will NOT be rerun; original notebook metrics are reused.')\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-20T11:05:51.489756Z","iopub.execute_input":"2026-07-20T11:05:51.490319Z","iopub.status.idle":"2026-07-20T11:05:51.499732Z","shell.execute_reply.started":"2026-07-20T11:05:51.490285Z","shell.execute_reply":"2026-07-20T11:05:51.499106Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [SETUP] Core imports\nimport math, warnings, pathlib, subprocess, time\nfrom collections import defaultdict\n\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport cv2\nfrom PIL import Image as PILImage\n\nwarnings.filterwarnings('ignore')\nnp.random.seed(SEED)\nsns.set_theme(style='whitegrid', context='notebook')\nplt.rcParams.update({'figure.figsize': (10, 5), 'axes.titlesize': 13})\nprint('[STATUS] Core imports complete')\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-20T11:05:51.500995Z","iopub.execute_input":"2026-07-20T11:05:51.501314Z","iopub.status.idle":"2026-07-20T11:05:52.811203Z","shell.execute_reply.started":"2026-07-20T11:05:51.501282Z","shell.execute_reply":"2026-07-20T11:05:52.810409Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [SETUP] Install deps + GPU detection\ndef pip_install(spec: str):\n    r = subprocess.run(['pip', 'install', '-q', spec], capture_output=True, text=True)\n    print(f'[PIP] {spec} | exit={r.returncode}')\n    if r.returncode != 0:\n        print(r.stderr[-500:])\n\npip_install('ultralytics')\npip_install('git+https://github.com/openai/CLIP.git')\n\nimport torch\nfrom ultralytics import YOLO\nimport clip\n\nCUDA_OK = torch.cuda.is_available()\nDEVICE = 'cuda' if CUDA_OK else 'cpu'\nYOLO_DEVICE = 0 if CUDA_OK else 'cpu'\nUSE_HALF = False  # disabled: ultralytics warns deprecated; FP32 is fine on T4\n\nif CUDA_OK:\n    torch.backends.cudnn.benchmark = True\n    print(f'[SUCCESS] CUDA ON | {torch.cuda.get_device_name(0)} | torch={torch.__version__}')\n    print(f'  YOLO_DEVICE={YOLO_DEVICE}')\nelse:\n    print('[WARN] CUDA OFF - enable GPU T4 then Restart.')\n\nCLIP_AVAILABLE = True\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-20T11:05:52.812864Z","iopub.execute_input":"2026-07-20T11:05:52.813456Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [SETUP] Competition paths\n_candidates = [\n    pathlib.Path('/kaggle/input/competitions/accident'),\n    pathlib.Path('/kaggle/input/accident'),\n]\nBASE_DIR = next((c for c in _candidates if c.exists()), _candidates[0])\nSYNTHETIC_DIR = BASE_DIR / 'sim_dataset'\nREAL_VIDEOS_DIR = BASE_DIR / 'videos'\nOUTPUT_DIR = pathlib.Path('/kaggle/working')\nOUTPUT_DIR.mkdir(parents=True, exist_ok=True)\n\nLABELS_CSV = SYNTHETIC_DIR / 'labels.csv'\nSYNTHETIC_VIDEOS_DIR = SYNTHETIC_DIR / 'videos'\nCOLLISION_TYPE_DIRS = ['head-on', 'rear-end', 'sideswipe', 'single', 't-bone']\n\nSAMPLE_SUBMISSION = BASE_DIR / 'sample_submission.csv'\nif not SAMPLE_SUBMISSION.exists():\n    video_paths = sorted(REAL_VIDEOS_DIR.glob('*.mp4')) if REAL_VIDEOS_DIR.exists() else []\n    sample_df = pd.DataFrame({\n        'path': [f'videos/{vp.name}' for vp in video_paths],\n        'accident_time': 10.0,\n        'center_x': 0.5,\n        'center_y': 0.5,\n        'type': 'rear-end',\n    })\n    SAMPLE_SUBMISSION = OUTPUT_DIR / 'sample_submission.csv'\n    sample_df.to_csv(SAMPLE_SUBMISSION, index=False)\n    print(f'[WARN] created {SAMPLE_SUBMISSION} | rows={len(sample_df)}')\n\npath_checks = [\n    ('BASE_DIR', BASE_DIR),\n    ('SYNTHETIC_DIR', SYNTHETIC_DIR),\n    ('SYNTHETIC_VIDEOS_DIR', SYNTHETIC_VIDEOS_DIR),\n    ('REAL_VIDEOS_DIR', REAL_VIDEOS_DIR),\n    ('labels.csv', LABELS_CSV),\n    ('sample_submission.csv', SAMPLE_SUBMISSION),\n]\naudit = pd.DataFrame([\n    {'label': l, 'path': str(pp), 'exists': pp.exists()} for l, pp in path_checks\n])\ndisplay(audit)\nassert audit['exists'].all(), 'Missing competition files'\nprint('[SUCCESS] Paths OK | BASE_DIR =', BASE_DIR)\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 1. Data Loading\n","metadata":{}},{"cell_type":"code","source":"# [LOAD] Labels + videos\nlabels_df = pd.read_csv(LABELS_CSV)\nsample_sub = pd.read_csv(SAMPLE_SUBMISSION) if SAMPLE_SUBMISSION.exists() else pd.DataFrame()\n\nsynthetic_videos = []\nfor subdir in COLLISION_TYPE_DIRS:\n    sp = SYNTHETIC_VIDEOS_DIR / subdir\n    if sp.exists():\n        synthetic_videos.extend(sorted(sp.glob('*.mp4')))\nreal_videos = sorted(REAL_VIDEOS_DIR.glob('*.mp4')) if REAL_VIDEOS_DIR.exists() else []\n\nlabels_clean = labels_df.copy()\nif 'type' in labels_clean.columns:\n    labels_clean['type'] = labels_clean['type'].astype(str).str.strip().str.lower()\nfor col in ['center_x', 'center_y']:\n    if col in labels_clean.columns:\n        labels_clean[col] = labels_clean[col].clip(0, 1)\n\nprint(f'labels={len(labels_clean)} | synthetic_videos={len(synthetic_videos)} | real_test={len(real_videos)}')\ndisplay(labels_clean.head(3))\nprint('columns:', list(labels_clean.columns))\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [LOAD] Exact calibration set used by the original notebook\n\ndef video_stem(p) -> str:\n    return pathlib.Path(p).stem\n\nstem_to_video = {video_stem(v): v for v in synthetic_videos}\n\nif 'rgb_path' in labels_clean.columns:\n    labels_clean['video_stem'] = labels_clean['rgb_path'].apply(lambda p: pathlib.Path(p).stem)\nelif 'path' in labels_clean.columns:\n    labels_clean['video_stem'] = labels_clean['path'].apply(lambda p: pathlib.Path(p).stem)\nelse:\n    raise KeyError('labels.csv must contain rgb_path or path')\n\neval_pool = labels_clean[labels_clean['video_stem'].isin(stem_to_video)].copy()\nprint(f'[STATUS] Label rows with resolvable video: {len(eval_pool)} / {len(labels_clean)}')\n\n# Original code: synthetic_videos[:10]. Preserve that exact order.\neval_videos = synthetic_videos[:min(CALIBRATION_N, len(synthetic_videos))]\neval_stems = [video_stem(v) for v in eval_videos]\nlabels_by_stem = eval_pool.drop_duplicates('video_stem').set_index('video_stem')\nmissing_stems = [s for s in eval_stems if s not in labels_by_stem.index]\nassert not missing_stems, f'Missing calibration labels: {missing_stems}'\neval_labels = labels_by_stem.loc[eval_stems].reset_index()\n\nassert len(eval_labels) == CALIBRATION_N\nassert eval_labels['video_stem'].tolist() == eval_stems\nprint('[STATUS] Exact original calibration set:')\ndisplay(eval_labels[['video_stem', 'type', 'accident_time']])\nprint(f'[STATUS] CALIBRATION set size = {len(eval_videos)}')\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2. Competition Scoring Functions\n\nSame formulas as the paper / challenge:\n- Temporal Gaussian σ_t=2.0\n- Spatial Gaussian σ_s=0.1\n- Classification top-1\n- Harmonic mean H = 3 / (1/T + 1/S + 1/C)\n","metadata":{}},{"cell_type":"code","source":"# [EVAL] Official-style scoring helpers\n\ndef temporal_score(pred_time: float, gt_time: float, sigma: float = 2.0) -> float:\n    return float(np.exp(-0.5 * ((pred_time - gt_time) / sigma) ** 2))\n\n\ndef spatial_score(pred_x, pred_y, gt_x, gt_y, sigma: float = 0.1) -> float:\n    dist2 = (pred_x - gt_x) ** 2 + (pred_y - gt_y) ** 2\n    return float(np.exp(-0.5 * dist2 / (sigma ** 2)))\n\n\ndef classification_score(pred_type: str, gt_type: str) -> int:\n    return int(str(pred_type).strip().lower() == str(gt_type).strip().lower())\n\n\ndef harmonic_mean(t: float, s: float, c: float) -> float:\n    if t <= 0 or s <= 0 or c <= 0:\n        return 0.0\n    return 3.0 / (1.0 / t + 1.0 / s + 1.0 / c)\n\n\ndef score_row(pred, gt):\n    T = temporal_score(pred['accident_time'], gt['accident_time'])\n    S = spatial_score(pred['center_x'], pred['center_y'], gt['center_x'], gt['center_y'])\n    C = classification_score(pred['type'], gt['type'])\n    H = harmonic_mean(T, S, C)\n    return {'T': T, 'S': S, 'C': C, 'H': H}\n\nprint('[STATUS] Scoring functions defined')\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 3. Baseline Pipeline (Paper Reproduction)\n\nImplements Thakur & Talele (2026):\n1. Frame difference → rolling mean → z-score peak (τ=1.5)\n2. Farneback optical-flow magnitude → P90 threshold → weighted centroid\n3. CLIP **ViT-L/14**, matching the original notebook, with the original five prompts per class\n","metadata":{}},{"cell_type":"code","source":"# [BASELINE] Temporal: frame-difference anomaly\n\ndef bl_compute_frame_diff_series(video_path, resize_w=320, resize_h=180):\n    cap = cv2.VideoCapture(str(video_path))\n    diffs, prev = [], None\n    while True:\n        ret, frame = cap.read()\n        if not ret:\n            break\n        gray = cv2.cvtColor(cv2.resize(frame, (resize_w, resize_h)), cv2.COLOR_BGR2GRAY).astype(np.float32)\n        if prev is not None:\n            diffs.append(np.mean(np.abs(gray - prev)))\n        prev = gray\n    cap.release()\n    return np.asarray(diffs, dtype=np.float32)\n\n\ndef bl_score_temporal_anomaly(diff_series, smooth_window=5):\n    series = pd.Series(diff_series)\n    smoothed = series.rolling(window=smooth_window, min_periods=1, center=True).mean().values\n    mu, sigma = smoothed.mean(), smoothed.std() + 1e-8\n    return (smoothed - mu) / sigma\n\n\ndef baseline_predict_accident_time(video_path, smooth_window=5, z_threshold=1.5) -> float:\n    cap = cv2.VideoCapture(str(video_path))\n    fps = cap.get(cv2.CAP_PROP_FPS)\n    n_frames = int(cap.get(cv2.CAP_PROP_FRAME_COUNT))\n    cap.release()\n    if fps <= 0 or n_frames == 0:\n        return 0.0\n    diffs = bl_compute_frame_diff_series(video_path)\n    if len(diffs) == 0:\n        return n_frames / fps / 2.0\n    anomaly = bl_score_temporal_anomaly(diffs, smooth_window)\n    candidates = np.where(anomaly > z_threshold)[0]\n    if len(candidates) == 0:\n        peak_frame = int(np.argmax(anomaly))\n    else:\n        peak_frame = int(candidates[np.argmax(anomaly[candidates])])\n    return round(peak_frame / fps, 4)\n\nprint('[STATUS] Baseline temporal ready')\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [BASELINE] Spatial: Farneback OF weighted centroid\n\ndef bl_compute_flow_magnitude_map(video_path, resize_w=320, resize_h=180,\n                                  n_frames_context=30, center_frame=None, flow_percentile=90.0):\n    cap = cv2.VideoCapture(str(video_path))\n    total = int(cap.get(cv2.CAP_PROP_FRAME_COUNT))\n    if center_frame is not None:\n        start = max(0, center_frame - n_frames_context // 2)\n    else:\n        start = max(0, total // 3)\n    cap.set(cv2.CAP_PROP_POS_FRAMES, start)\n    mag_accum = np.zeros((resize_h, resize_w), dtype=np.float32)\n    prev, count = None, 0\n    while count < n_frames_context:\n        ret, frame = cap.read()\n        if not ret:\n            break\n        gray = cv2.cvtColor(cv2.resize(frame, (resize_w, resize_h)), cv2.COLOR_BGR2GRAY)\n        if prev is not None:\n            flow = cv2.calcOpticalFlowFarneback(\n                prev, gray, None, 0.5, 3, 15, 3, 5, 1.2, 0\n            )\n            mag, _ = cv2.cartToPolar(flow[..., 0], flow[..., 1])\n            mag_accum += mag\n        prev = gray\n        count += 1\n    cap.release()\n    if mag_accum.max() > 0:\n        thresh = np.percentile(mag_accum, flow_percentile)\n        mag_accum[mag_accum < thresh] = 0.0\n    return mag_accum\n\n\ndef baseline_predict_impact_location(video_path, accident_time=None, n_frames_context=30):\n    RW, RH = 320, 180\n    center_frame = None\n    if accident_time is not None:\n        cap = cv2.VideoCapture(str(video_path))\n        fps = cap.get(cv2.CAP_PROP_FPS)\n        cap.release()\n        if fps > 0:\n            center_frame = int(accident_time * fps)\n    mag = bl_compute_flow_magnitude_map(\n        video_path, RW, RH, n_frames_context, center_frame\n    )\n    total = mag.sum()\n    if total < 1e-6:\n        return 0.5, 0.5\n    ys, xs = np.mgrid[0:RH, 0:RW]\n    cx = float((xs * mag).sum() / total) / RW\n    cy = float((ys * mag).sum() / total) / RH\n    return round(cx, 6), round(cy, 6)\n\nprint('[STATUS] Baseline spatial ready')\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [BASELINE] CLIP prompts — paper Table 1 (5 prompts / class)\nBASELINE_PROMPTS = {\n    'rear-end': [\n        'a car colliding into the back of another car',\n        'rear-end collision between two vehicles on a road',\n        'vehicle hitting the back of a stationary car from behind',\n        'one car rear-ending another car at a traffic light',\n        'a vehicle crashing into the tail of the car ahead',\n    ],\n    't-bone': [\n        'a car hitting the side of another car at an intersection',\n        't-bone collision at a crossroads between two vehicles',\n        'side impact crash where one car strikes another perpendicularly',\n        'a vehicle running a red light and hitting the side of crossing traffic',\n        'perpendicular collision between two cars at a junction',\n    ],\n    'head-on': [\n        'two cars colliding head-on from opposite directions',\n        'frontal collision between two vehicles on a road',\n        'head-on crash between two cars driving toward each other',\n        'two vehicles smashing front-to-front on a highway',\n        'a car crossing the center line and hitting an oncoming vehicle head-on',\n    ],\n    'sideswipe': [\n        'two vehicles scraping alongside each other while driving',\n        'sideswipe collision between cars changing lanes',\n        'glancing blow between two cars moving in the same direction',\n        'a car drifting into the adjacent lane and scraping another vehicle',\n        'two vehicles brushing sides while traveling parallel on a road',\n    ],\n    'single': [\n        'a single car crashing into a wall or barrier',\n        'one vehicle running off the road and hitting an obstacle',\n        'a car losing control and crashing into a pole or guardrail',\n        'a single vehicle spinning out and hitting a roadside object',\n        'one car veering off the road and crashing without involving another vehicle',\n    ],\n}\n\nclip_model, clip_preprocess = clip.load(CLIP_BACKBONE, device=DEVICE)\nclip_model.eval()\n\n\ndef encode_prompt_bank(prompt_dict):\n    # Mean-pool L2-normalized text embeddings per class (paper Eq. 7).\n    feats = {}\n    with torch.no_grad():\n        for ctype, prompts in prompt_dict.items():\n            tokens = clip.tokenize(prompts, truncate=True).to(DEVICE)\n            f = clip_model.encode_text(tokens).float()\n            f = f / f.norm(dim=-1, keepdim=True)\n            feats[ctype] = f.mean(dim=0)\n    return feats\n\nBASELINE_TEXT_FEATURES = encode_prompt_bank(BASELINE_PROMPTS)\nprint(f'[SUCCESS] CLIP {CLIP_BACKBONE} loaded | baseline text features ready')\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [BASELINE] Classification + full inference\n\ndef extract_frames_around_peak(video_path, peak_time_s, n_context_frames=8, fps=None):\n    cap = cv2.VideoCapture(str(video_path))\n    if fps is None:\n        fps = cap.get(cv2.CAP_PROP_FPS) or 20.0\n    total = int(cap.get(cv2.CAP_PROP_FRAME_COUNT))\n    peak_frame = int(peak_time_s * fps)\n    half = n_context_frames // 2\n    idxs = range(max(0, peak_frame - half), min(total, peak_frame + half + 1))\n    frames = []\n    for idx in idxs:\n        cap.set(cv2.CAP_PROP_POS_FRAMES, idx)\n        ret, frame = cap.read()\n        if ret:\n            frames.append(PILImage.fromarray(cv2.cvtColor(frame, cv2.COLOR_BGR2RGB)))\n    cap.release()\n    return frames\n\n\ndef clip_classify(video_path, peak_time_s, text_features, n_context_frames=8) -> str:\n    pil_frames = extract_frames_around_peak(video_path, peak_time_s, n_context_frames)\n    if not pil_frames:\n        return 'rear-end'\n    with torch.no_grad():\n        imgs = torch.stack([clip_preprocess(f) for f in pil_frames]).to(DEVICE)\n        img_f = clip_model.encode_image(imgs).float()\n        img_f = img_f / img_f.norm(dim=-1, keepdim=True)\n        img_f = img_f.mean(dim=0)\n    scores = {k: float(img_f @ v) for k, v in text_features.items()}\n    return max(scores, key=scores.get)\n\n\ndef baseline_run_inference(video_path) -> dict:\n    t = baseline_predict_accident_time(video_path)\n    cx, cy = baseline_predict_impact_location(video_path, accident_time=t)\n    k = clip_classify(video_path, t, BASELINE_TEXT_FEATURES)\n    return {\n        'path': str(video_path),\n        'accident_time': t,\n        'center_x': cx,\n        'center_y': cy,\n        'type': k,\n        'method': 'baseline',\n    }\n\nprint('[STATUS] baseline_run_inference ready')\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 4. Improved Pipeline — Fair Original-Calibration Protocol\n\n1. **Temporal:** original frame-difference method\n2. **Spatial:** gated YOLO vehicle-pair midpoint blended with original optical flow\n3. **Classification:** logistic-regression probe on CLIP ViT-L/14 embeddings\n4. **Training:** every labeled synthetic video except the exact 10 held-out calibration videos\n5. **Evaluation:** the same `synthetic_videos[:10]` used by the original notebook\n\nThis is a held-out supervised improvement, not a pure zero-shot method. The notebook reports that distinction explicitly.\n","metadata":{}},{"cell_type":"code","source":"# [IMPROVED v7] YOLO window detector (no half= dep warnings)\nyolo = YOLO(YOLO_MODEL)\n_dummy = np.zeros((YOLO_IMGSZ, YOLO_IMGSZ, 3), dtype=np.uint8)\n_ = yolo.predict(source=_dummy, device=YOLO_DEVICE, verbose=False, imgsz=YOLO_IMGSZ)\ntry:\n    param_device = next(yolo.model.parameters()).device\nexcept Exception:\n    param_device = 'unknown'\nprint(f'[SUCCESS] YOLO ready | weights_device={param_device} | YOLO_DEVICE={YOLO_DEVICE}')\n\n\ndef _video_meta(video_path):\n    cap = cv2.VideoCapture(str(video_path))\n    fps = cap.get(cv2.CAP_PROP_FPS) or 20.0\n    total = int(cap.get(cv2.CAP_PROP_FRAME_COUNT))\n    W = max(1, int(cap.get(cv2.CAP_PROP_FRAME_WIDTH)))\n    H = max(1, int(cap.get(cv2.CAP_PROP_FRAME_HEIGHT)))\n    cap.release()\n    return float(fps), int(total), W, H\n\n\ndef yolo_detect_window(video_path, center_frame, half_window=12, stride=3,\n                       conf=YOLO_CONF, classes=YOLO_CLASSES, imgsz=YOLO_IMGSZ):\n    fps, total, W, H = _video_meta(video_path)\n    lo = max(0, int(center_frame) - int(half_window))\n    hi = min(total - 1, int(center_frame) + int(half_window))\n    cap = cv2.VideoCapture(str(video_path))\n    by_frame = {}\n    for f in range(lo, hi + 1, max(1, int(stride))):\n        cap.set(cv2.CAP_PROP_POS_FRAMES, f)\n        ok, frame = cap.read()\n        if not ok:\n            continue\n        res = yolo.predict(\n            source=frame, conf=conf, classes=classes,\n            device=YOLO_DEVICE, imgsz=imgsz, verbose=False,\n        )[0]\n        boxes = []\n        if res.boxes is not None and len(res.boxes) > 0:\n            for box in res.boxes.xyxy.cpu().numpy():\n                x1, y1, x2, y2 = box.tolist()\n                boxes.append({\n                    'cx': ((x1 + x2) * 0.5) / W,\n                    'cy': ((y1 + y2) * 0.5) / H,\n                })\n        by_frame[f] = boxes\n    cap.release()\n    return by_frame, fps, total\n\nprint('[STATUS] YOLO window detector ready (v7)')\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [IMPROVED v7] Temporal=baseline | Spatial=GATED YOLO (only when close pair)\n\ndef _closest_pair(boxes):\n    if not boxes:\n        return None\n    if len(boxes) == 1:\n        return boxes[0]['cx'], boxes[0]['cy'], 1.0, False\n    best_d, best = 1e9, None\n    for i in range(len(boxes)):\n        for j in range(i + 1, len(boxes)):\n            d = math.hypot(boxes[i]['cx'] - boxes[j]['cx'], boxes[i]['cy'] - boxes[j]['cy'])\n            if d < best_d:\n                best_d = d\n                best = (\n                    0.5 * (boxes[i]['cx'] + boxes[j]['cx']),\n                    0.5 * (boxes[i]['cy'] + boxes[j]['cy']),\n                    d, True,\n                )\n    return best\n\n\ndef yolo_spatial_candidate(by_frame, peak_frame, max_pair_dist=0.12):\n    # Only accept YOLO if we see a truly close vehicle pair near peak.\n    best = None  # (score, cx, cy, dist)\n    for f, boxes in by_frame.items():\n        pair = _closest_pair(boxes)\n        if pair is None:\n            continue\n        cx, cy, dist, has_pair = pair\n        if not has_pair or dist > max_pair_dist:\n            continue\n        score = -dist - 0.01 * abs(f - peak_frame)\n        if best is None or score > best[0]:\n            best = (score, cx, cy, dist)\n    if best is None:\n        return None, None, None\n    return best[1], best[2], best[3]\n\n\ndef improved_predict_time_space(video_path, max_pair_dist=0.12, blend=0.7):\n    t = baseline_predict_accident_time(video_path)\n    ofx, ofy = baseline_predict_impact_location(video_path, accident_time=t)\n\n    fps, total, _, _ = _video_meta(video_path)\n    peak_frame = int(np.clip(round(t * fps), 0, max(0, total - 1)))\n    by_frame, _, _ = yolo_detect_window(\n        video_path, peak_frame,\n        half_window=12, stride=max(2, YOLO_VID_STRIDE), conf=YOLO_CONF,\n    )\n    n_boxes = int(sum(len(v) for v in by_frame.values()))\n    yx, yy, dist = yolo_spatial_candidate(by_frame, peak_frame, max_pair_dist=max_pair_dist)\n\n    used_yolo = False\n    if yx is not None:\n        # Blend with OF — v6 evidence: pure YOLO raised mean_S but hurt H on C=1 videos\n        cx = round(min(1.0, max(0.0, blend * yx + (1.0 - blend) * ofx)), 6)\n        cy = round(min(1.0, max(0.0, blend * yy + (1.0 - blend) * ofy)), 6)\n        used_yolo = True\n    else:\n        cx, cy = ofx, ofy\n\n    return t, cx, cy, {\n        'n_tracks': n_boxes,\n        'fallback_space': (not used_yolo),\n        'used_yolo_spatial': used_yolo,\n        'pair_dist': dist,\n        'time_source': 'baseline',\n        'peak_frame': peak_frame,\n        'fps': fps,\n    }\n\nprint('[STATUS] improved temporal/spatial v7 ready (gated YOLO)')\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [IMPROVED] Held-out synthetic-supervised CLIP classifier\n# Train a linear head on CLIP ViT-L/14 embeddings near GT accident time.\n# The exact 10 original calibration videos are excluded from training.\n\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.preprocessing import LabelEncoder\n\nENSEMBLE_BASE_PROMPTS = {\n    'rear-end': [\n        'a car colliding into the back of another car',\n        'rear-end collision between two vehicles on a road',\n        'vehicle hitting the rear bumper of the car ahead',\n        'one car rear-ending another car in a traffic queue',\n        'a following vehicle crashing into a lead vehicle from behind',\n        'same-direction crash where the rear car hits the front car',\n    ],\n    't-bone': [\n        'a car hitting the side of another car at an intersection',\n        't-bone collision at a crossroads between two vehicles',\n        'perpendicular side-impact crash between two cars',\n        'a vehicle striking the passenger door of crossing traffic',\n        'right-angle collision at a junction',\n        'side impact where one car T-bones another',\n    ],\n    'head-on': [\n        'two cars colliding head-on from opposite directions',\n        'frontal collision between two vehicles on a road',\n        'head-on crash between cars driving toward each other',\n        'two vehicles smashing front-to-front',\n        'opposite-direction frontal collision on a highway',\n        'cars colliding nose-to-nose after crossing into oncoming traffic',\n    ],\n    'sideswipe': [\n        'two vehicles scraping alongside each other while driving',\n        'sideswipe collision between cars changing lanes',\n        'glancing lateral blow between parallel vehicles',\n        'a car drifting into an adjacent lane and scraping another vehicle',\n        'same-direction side scrape between two cars',\n        'lateral contact between two vehicles traveling parallel',\n    ],\n    'single': [\n        'a single car crashing into a wall or barrier',\n        'one vehicle running off the road and hitting an obstacle',\n        'a car losing control and hitting a pole or guardrail',\n        'a single-vehicle crash with no other vehicle involved',\n        'one car veering off the roadway into a roadside object',\n        'solo vehicle impact against infrastructure',\n    ],\n}\nENSEMBLE_PROMPTS = {\n    k: list(dict.fromkeys(BASELINE_PROMPTS[k] + ENSEMBLE_BASE_PROMPTS[k]))\n    for k in BASELINE_PROMPTS\n}\nENSEMBLE_TEXT_FEATURES = encode_prompt_bank(ENSEMBLE_PROMPTS)\ndisplay(pd.DataFrame([{'type': k, 'n_prompts': len(v)} for k, v in ENSEMBLE_PROMPTS.items()]))\n\n\ndef clip_image_embedding(video_path, peak_time_s, n_context_frames=CLIP_CONTEXT_FRAMES):\n    pil_frames = extract_frames_around_peak(video_path, peak_time_s, n_context_frames)\n    if not pil_frames:\n        return None\n    with torch.no_grad():\n        imgs = torch.stack([clip_preprocess(f) for f in pil_frames]).to(DEVICE, non_blocking=CUDA_OK)\n        feat = clip_model.encode_image(imgs).float()\n        feat = feat / feat.norm(dim=-1, keepdim=True)\n        feat = feat.mean(dim=0)\n    return feat.detach().cpu().numpy().astype(np.float32)\n\n\n# ---- Full synthetic train pool, excluding the exact 10 calibration videos ----\ncalibration_stems = set(eval_labels['video_stem'].tolist())\ntrain_labels = eval_pool[~eval_pool['video_stem'].isin(calibration_stems)].copy()\ntrain_labels = train_labels.sample(frac=1, random_state=SEED).reset_index(drop=True)\n\nassert calibration_stems.isdisjoint(set(train_labels['video_stem']))\nassert len(train_labels) + len(eval_labels) == len(eval_pool)\nprint(f'[STATUS] Full classifier training set: {len(train_labels)} videos')\nprint(f'[STATUS] Held-out original calibration set: {len(eval_labels)} videos')\ndisplay(train_labels['type'].value_counts().rename('n').to_frame())\n\nTRAIN_CACHE_PATH = OUTPUT_DIR / 'clip_vitl14_full_train_excluding_original10.npz'\n\n\ndef build_xy(label_rows, desc='set'):\n    X, y, kept = [], [], 0\n    for i, (_, row) in enumerate(label_rows.iterrows()):\n        stem = row['video_stem']\n        vp = stem_to_video.get(stem)\n        if vp is None:\n            continue\n        emb = clip_image_embedding(\n            vp,\n            float(row['accident_time']),\n            n_context_frames=TRAIN_CONTEXT_FRAMES,\n        )\n        if emb is None:\n            continue\n        X.append(emb)\n        y.append(str(row['type']).strip().lower())\n        kept += 1\n        if (i + 1) % 50 == 0 or (i + 1) == len(label_rows):\n            print(f'[CLIP-emb {desc}] {i+1}/{len(label_rows)}')\n    return np.stack(X), np.array(y), kept\n\nif TRAIN_CACHE_PATH.exists():\n    cached = np.load(TRAIN_CACHE_PATH, allow_pickle=False)\n    X_train = cached['X'].astype(np.float32)\n    y_train = cached['y'].astype(str)\n    n_tr = len(y_train)\n    print(f'[CACHE] Loaded full-train embeddings: {X_train.shape}')\nelse:\n    print('[STATUS] Extracting CLIP embeddings for full synthetic train set...')\n    X_train, y_train, n_tr = build_xy(train_labels, 'full-train')\n    np.savez_compressed(TRAIN_CACHE_PATH, X=X_train, y=y_train)\n    print(f'[CACHE] Saved {TRAIN_CACHE_PATH}')\n\nprint(f'[SUCCESS] Train embeddings: {X_train.shape}')\n\nlabel_encoder = LabelEncoder()\ny_train_id = label_encoder.fit_transform(y_train)\n\nclf = LogisticRegression(\n    max_iter=2000,\n    multi_class='multinomial',\n    class_weight='balanced',\n    C=1.0,\n    solver='lbfgs',\n)\nclf.fit(X_train, y_train_id)\ntrain_acc = float((clf.predict(X_train) == y_train_id).mean())\nprint(f'[SUCCESS] LogisticRegression trained | train_acc={train_acc:.3f} | classes={list(label_encoder.classes_)}')\n\n\ndef classify_v7(video_path, peak_time_s):\n    # 1) Supervised CLIP probe (primary)\n    emb = clip_image_embedding(video_path, peak_time_s)\n    if emb is not None:\n        pred_id = int(clf.predict(emb.reshape(1, -1))[0])\n        proba = clf.predict_proba(emb.reshape(1, -1))[0]\n        conf = float(proba.max())\n        pred = label_encoder.inverse_transform([pred_id])[0]\n        # 2) If probe uncertain, fall back to zero-shot baseline CLIP\n        if conf >= 0.35:\n            return pred, conf, 'probe'\n    zs = clip_classify(video_path, peak_time_s, BASELINE_TEXT_FEATURES, n_context_frames=CLIP_CONTEXT_FRAMES)\n    return zs, 0.0, 'zeroshot'\n\nprint('[STATUS] classify_v7 ready')\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [IMPROVED v7] Full inference\n\ndef improved_run_inference(video_path) -> dict:\n    t, cx, cy, dbg = improved_predict_time_space(\n        video_path,\n        max_pair_dist=YOLO_MAX_PAIR_DIST,\n        blend=YOLO_SPATIAL_BLEND,\n    )\n    k, conf, src = classify_v7(video_path, t)\n    return {\n        'path': str(video_path),\n        'accident_time': t,\n        'center_x': cx,\n        'center_y': cy,\n        'type': k,\n        'method': 'improved_v7',\n        'n_tracks': dbg.get('n_tracks'),\n        'fallback_space': dbg.get('fallback_space'),\n        'used_yolo_spatial': dbg.get('used_yolo_spatial'),\n        'clf_conf': conf,\n        'clf_src': src,\n    }\n\nprint('[STATUS] improved_run_inference v7 ready')\nprint(f'[STATUS] CUDA_OK={CUDA_OK} DEVICE={DEVICE} YOLO_DEVICE={YOLO_DEVICE}')\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 5. Original Baseline vs Improved — Same 10 Videos\n\nOnly the improved method is run here. The baseline row reuses the original notebook output: T=0.4382, S=0.1680, C=0.0000, H=0.0000.\n","metadata":{}},{"cell_type":"code","source":"# [COMPARE] Run ONLY the improved pipeline on original 10-video calibration set\n# Baseline predictions are not recomputed; cell 23 reuses original notebook metrics.\n\ndef gt_from_label_row(row):\n    return {\n        'accident_time': float(row['accident_time']),\n        'center_x': float(row['center_x']),\n        'center_y': float(row['center_y']),\n        'type': str(row['type']).strip().lower(),\n    }\n\n\ndef run_improved_eval(videos, label_rows):\n    rows = []\n    t0 = time.time()\n    for i, (vp, (_, lab)) in enumerate(zip(videos, label_rows.iterrows()), start=1):\n        pred = improved_run_inference(vp)\n        gt = gt_from_label_row(lab)\n        sc = score_row(pred, gt)\n        rows.append({\n            'video_stem': video_stem(vp),\n            'gt_type': gt['type'],\n            'pred_type': pred['type'],\n            'pred_t': pred['accident_time'],\n            'gt_t': gt['accident_time'],\n            'pred_cx': pred['center_x'], 'pred_cy': pred['center_y'],\n            'gt_cx': gt['center_x'], 'gt_cy': gt['center_y'],\n            'fallback_space': pred.get('fallback_space'),\n            'used_yolo_spatial': pred.get('used_yolo_spatial'),\n            'clf_src': pred.get('clf_src'),\n            **sc,\n        })\n        print(f'[IMPROVED] {i}/{len(videos)} | {vp.name} | type={pred[\"type\"]}')\n    df = pd.DataFrame(rows)\n    df.attrs['elapsed_sec'] = time.time() - t0\n    return df\n\n\nprint('[STATUS] Baseline inference skipped — reusing saved original metrics.')\nprint(f'[STATUS] Running improved pipeline on {len(eval_videos)} original calibration videos...')\nimproved_eval = run_improved_eval(eval_videos, eval_labels)\nprint('[SUCCESS] Improved calibration inference complete')\ndisplay(improved_eval[['video_stem', 'gt_type', 'pred_type', 'T', 'S', 'C', 'H']])\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [COMPARE] Original saved baseline vs improved on the same 10 videos\n# Source: original notebook cell 58 output (do not rerun baseline inference).\nORIGINAL_BASELINE = {\n    'method': 'Baseline (original notebook)',\n    'N': 10,\n    'mean_T': 0.4382,\n    'mean_S': 0.1680,\n    'mean_C': 0.0000,\n    'mean_H': 0.0000,\n    'C_accuracy_%': 0.00,\n    'elapsed_sec': np.nan,\n}\n\n\ndef summarize_improved(df):\n    return {\n        'method': 'Improved (YOLO + CLIP full-train probe)',\n        'N': len(df),\n        'mean_T': round(float(df['T'].mean()), 4),\n        'mean_S': round(float(df['S'].mean()), 4),\n        'mean_C': round(float(df['C'].mean()), 4),\n        'mean_H': round(float(df['H'].mean()), 4),\n        'C_accuracy_%': round(100 * float(df['C'].mean()), 2),\n        'elapsed_sec': round(float(df.attrs.get('elapsed_sec', np.nan)), 1),\n    }\n\n\nimproved_summary = summarize_improved(improved_eval)\ncomparison_summary = pd.DataFrame([ORIGINAL_BASELINE, improved_summary])\n\nmetric_cols = ['mean_T', 'mean_S', 'mean_C', 'mean_H']\ndelta = comparison_summary.iloc[1][metric_cols].astype(float) - comparison_summary.iloc[0][metric_cols].astype(float)\ndelta_row = {\n    'method': 'Delta (improved - baseline)',\n    'N': 10,\n    'mean_T': round(float(delta['mean_T']), 4),\n    'mean_S': round(float(delta['mean_S']), 4),\n    'mean_C': round(float(delta['mean_C']), 4),\n    'mean_H': round(float(delta['mean_H']), 4),\n    'C_accuracy_%': round(improved_summary['C_accuracy_%'] - ORIGINAL_BASELINE['C_accuracy_%'], 2),\n    'elapsed_sec': np.nan,\n}\ncomparison_summary = pd.concat([comparison_summary, pd.DataFrame([delta_row])], ignore_index=True)\n\nprint('========== ORIGINAL BASELINE vs IMPROVED (SAME 10 VIDEOS) ==========')\ndisplay(comparison_summary)\ncomparison_summary.to_csv(OUTPUT_DIR / 'comparison_summary_original10.csv', index=False)\nprint('[SUCCESS] Saved /kaggle/working/comparison_summary_original10.csv')\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [PRESENTATION] Aggregate score chart\nmetrics = ['mean_T', 'mean_S', 'mean_C', 'mean_H']\nx = np.arange(len(metrics))\nbaseline_values = comparison_summary.iloc[0][metrics].astype(float).values\nimproved_values = comparison_summary.iloc[1][metrics].astype(float).values\nw = 0.35\n\nfig, ax = plt.subplots(figsize=(8, 4.5))\nax.bar(x - w/2, baseline_values, w, label='Original baseline')\nax.bar(x + w/2, improved_values, w, label='Improved')\nax.set_xticks(x)\nax.set_xticklabels(['Temporal (T)', 'Spatial (S)', 'Classification (C)', 'Harmonic (H)'])\nax.set_ylim(0, 1)\nax.set_ylabel('Mean score')\nax.set_title('Same original 10-video calibration set')\nax.legend()\nax.grid(axis='y', alpha=0.25)\nplt.tight_layout()\nplt.savefig(OUTPUT_DIR / 'comparison_original10.png', dpi=160, bbox_inches='tight')\nplt.show()\nprint('[SUCCESS] Saved comparison_original10.png')\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [PRESENTATION] Improved confusion matrix and detail CSV\nfrom sklearn.metrics import confusion_matrix\n\nlabels = COLLISION_TYPE_DIRS\ncm = confusion_matrix(improved_eval['gt_type'], improved_eval['pred_type'], labels=labels)\nfig, ax = plt.subplots(figsize=(6, 5))\nsns.heatmap(cm, annot=True, fmt='d', cmap='Blues',\n            xticklabels=labels, yticklabels=labels, ax=ax)\nax.set_xlabel('Predicted')\nax.set_ylabel('Ground truth')\nax.set_title('Improved classifier — original 10 videos')\nplt.tight_layout()\nplt.savefig(OUTPUT_DIR / 'improved_confusion_original10.png', dpi=160, bbox_inches='tight')\nplt.show()\n\nimproved_eval.to_csv(OUTPUT_DIR / 'improved_eval_original10.csv', index=False)\nprint('[SUCCESS] Saved improved confusion matrix and detail CSV')\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 6. Optional Qualitative Output\n\nSkipped because the original notebook did not save per-video baseline predictions, and baseline inference is intentionally not repeated.\n","metadata":{}},{"cell_type":"code","source":"# Baseline per-video predictions were intentionally not recomputed.\n# Therefore, no side-by-side qualitative plot is generated in this fair/fast comparison run.\nprint('[STATUS] Qualitative baseline plot skipped (baseline inference not rerun).')\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 7. Course-Project Artifacts\n\nThis notebook is configured for the report comparison only. It does not run the 2,027-video test set or require a Kaggle submission.\n","metadata":{}},{"cell_type":"code","source":"# [SUBMIT] Inference on real CCTV test set (improved pipeline)\nif RUN_FULL_TEST:\n    print(f'[STATUS] Full test inference on {len(real_videos)} videos...')\n    test_results = []\n    t0 = time.time()\n    for i, vp in enumerate(real_videos):\n        result = improved_run_inference(vp)\n        result['path'] = 'videos/' + vp.name\n        test_results.append(result)\n        if (i + 1) % 20 == 0 or (i + 1) == len(real_videos):\n            elapsed = time.time() - t0\n            print(f'[STATUS] {i+1}/{len(real_videos)} | {elapsed/60:.1f} min')\n    test_preds_df = pd.DataFrame(test_results)\n\n    if not sample_sub.empty:\n        submission_df = sample_sub[['path']].merge(\n            test_preds_df[['path', 'accident_time', 'center_x', 'center_y', 'type']],\n            on='path', how='left'\n        )\n        submission_df['accident_time'] = submission_df['accident_time'].fillna(10.0)\n        submission_df['center_x'] = submission_df['center_x'].fillna(0.5)\n        submission_df['center_y'] = submission_df['center_y'].fillna(0.5)\n        submission_df['type'] = submission_df['type'].fillna('rear-end')\n    else:\n        submission_df = test_preds_df[['path', 'accident_time', 'center_x', 'center_y', 'type']].copy()\n\n    out = OUTPUT_DIR / 'submission.csv'\n    submission_df.to_csv(out, index=False)\n    print(f'[SUCCESS] Wrote {out} | rows={len(submission_df)}')\n    display(submission_df.head())\n    display(submission_df['type'].value_counts().rename('count').to_frame())\nelse:\n    print('[SUCCESS] Course-project comparison complete; full test inference skipped.')\n    print('Artifacts in /kaggle/working:')\n    print('  - comparison_summary_original10.csv')\n    print('  - improved_eval_original10.csv')\n    print('  - comparison_original10.png')\n    print('  - improved_confusion_original10.png')\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 8. Interpretation and Limitation\n\nThe comparison is directly aligned with the original notebook: same CLIP backbone and same 10 calibration videos. The improved classifier is supervised on the remaining synthetic data, so present it as a **held-out synthetic-supervised extension**, not as zero-shot CLIP.\n\nThe 10-video calibration set is small and may not represent every collision class. Report the resulting table as reproduction-aligned evidence, not a full benchmark claim.\n","metadata":{}}]}