{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.12.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# ACCIDENT @ CVPR: Zero-Shot CCTV Traffic Accident Understanding\n\n**Competition:** ACCIDENT @ CVPR 2026 -- AUTOPILOT Workshop (Kaggle)\n\n**Notebook Focus:**\n- Temporal localization of accident onset from fixed-view CCTV video clips\n- Spatial localization of impact point via open-vocabulary grounding and optical flow\n- Zero-shot collision-type classification using pre-trained vision-language models","metadata":{}},{"cell_type":"markdown","source":"## Introduction\n\nTraffic accident analysis from fixed CCTV infrastructure presents challenges distinct from general video understanding. Accident events are temporally sparse, spatially unpredictable, and visually subtle relative to normal traffic flow. The benchmark defined by ACCIDENT @ CVPR 2026 compounds these challenges by excluding labeled real training footage, requiring methods that generalize without dataset-specific fine-tuning.\n\nThis notebook constructs a zero-shot inference pipeline whose **primary signal is per-vehicle detection and tracking**: an accident is an event between two specific objects, so the colliding pair's kinematics (contact time, joint deceleration, box-intersection centroid, approach-angle geometry) answer all three benchmark questions more directly than any whole-frame signal can. A chain of zero-shot vision-language models backs every stage up when tracking is unreliable:\n\n| Stage | Question | Primary signal (tracking) | Fallback (zero-shot) |\n|-------|----------|---------------------------|----------------------|\n| 1 -- *When* | accident time | first box contact + joint deceleration peak | frame-diff anchor + bounded PE refinement |\n| 2 -- *What* | collision type | pre-impact velocity angle + contact geometry | CLIP + Qwen2.5-VL soft ensemble |\n| 3 -- *Where* | impact point | box-intersection centroid at impact | OWLv2 grounding + flow centroid |\n\nThe synthetic CARLA dataset is used for pipeline calibration and EDA. Submissions target the real CCTV test set. Every design decision below is validated (or rejected) on a *diverse* calibration set covering all five collision types -- an earlier revision tuned on a head-on-only subset and regressed badly on the other four types (see Section 7).","metadata":{}},{"cell_type":"markdown","source":"## Table of Contents\n\n1. [Data Acquisition](#1-data-acquisition)\n2. [Data Inspection](#2-data-inspection)\n3. [Data Cleaning](#3-data-cleaning)\n4. [Exploratory Data Analysis](#4-exploratory-data-analysis)\n5. [Feature Engineering](#5-feature-engineering)\n6. [Modeling](#6-modeling)\n7. [Evaluation](#7-evaluation)\n8. [Test Inference and Submission](#8-test-inference-and-submission)\n9. [Conclusion](#9-conclusion)\n10. [References](#10-references)","metadata":{}},{"cell_type":"markdown","source":"---\n## 0. Environment Setup","metadata":{}},{"cell_type":"code","source":"# [SETUP] Standard library imports -- order: stdlib, third-party, local\nimport re\nimport json\nimport gzip\nimport time\nimport warnings\nimport pathlib\nfrom collections import defaultdict\n\n# [SETUP] Numerical and data processing\nimport numpy as np\nimport pandas as pd\n\n# [SETUP] Visualization stack\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# [SETUP] Computer vision\nimport cv2\nfrom PIL import Image as PILImage\n\n# [SETUP] Suppress non-critical runtime warnings\nwarnings.filterwarnings('ignore')\n\nprint('[STATUS] Core imports complete')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:28:46.020037Z","iopub.execute_input":"2026-08-03T11:28:46.020759Z","iopub.status.idle":"2026-08-03T11:28:47.565747Z","shell.execute_reply.started":"2026-08-03T11:28:46.020718Z","shell.execute_reply":"2026-08-03T11:28:47.564919Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [SETUP] Global matplotlib and seaborn configuration\nsns.set_theme(style='whitegrid', context='notebook')\n\nplt.rcParams.update({\n    'figure.figsize'  : (10, 6),\n    'axes.titlesize'  : 13,\n    'axes.labelsize'  : 11,\n    'xtick.labelsize' : 10,\n    'ytick.labelsize' : 10,\n    'legend.fontsize' : 10,\n    'figure.dpi'      : 120,\n    'savefig.bbox'    : 'tight',\n})\n\n# [SETUP] Consistent color palette for all plots\nPALETTE = {\n    'primary'    : '#1f77b4',\n    'secondary'  : '#ff7f0e',\n    'tertiary'   : '#2ca02c',\n    'quaternary' : '#d62728',\n    'quinary'    : '#9467bd',\n    'senary'     : '#8c564b',\n}\n\nprint('[STATUS] Plot configuration applied')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:28:47.566885Z","iopub.execute_input":"2026-08-03T11:28:47.567215Z","iopub.status.idle":"2026-08-03T11:28:47.574204Z","shell.execute_reply.started":"2026-08-03T11:28:47.567194Z","shell.execute_reply":"2026-08-03T11:28:47.573631Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [SETUP] Reproducibility seed for all stochastic operations\nSEED = 42\nnp.random.seed(SEED)\n\n# [SETUP] Root paths -- auto-detected, do not hardcode.\n#\n# Kaggle mounts data at a DIFFERENT prefix depending on how it was attached:\n#   - competition data (Add Input -> Competition) -> /kaggle/input/competitions/<slug>/\n#   - a regular Dataset (Add Input -> Dataset)     -> /kaggle/input/<slug>/\n# A previous revision hardcoded the Dataset-style path\n# (/kaggle/input/accident) with a comment noting the competition-style path\n# was the correct one -- and never applied its own fix. Every downstream\n# exists() check then silently returned False and produced 0 videos. Probe\n# both known layouts, and if neither matches, search all of /kaggle/input for\n# any directory that actually contains a sim_dataset/ or videos/ folder, so\n# this survives being forked, renamed, or re-attached differently.\n_CANDIDATE_ROOTS = [\n    pathlib.Path('/kaggle/input/competitions/accident'),\n    pathlib.Path('/kaggle/input/accident'),\n]\n\ndef _looks_like_dataset_root(p: pathlib.Path) -> bool:\n    return p.is_dir() and ((p / 'sim_dataset').exists() or (p / 'videos').exists())\n\nBASE_DIR = next((r for r in _CANDIDATE_ROOTS if _looks_like_dataset_root(r)), None)\n\nif BASE_DIR is None:\n    _input_root = pathlib.Path('/kaggle/input')\n    if _input_root.exists():\n        for _entry in sorted(_input_root.iterdir()):\n            if _looks_like_dataset_root(_entry):\n                BASE_DIR = _entry\n                break\n            # competition-style datasets can nest one level deeper, e.g.\n            # /kaggle/input/competitions/<slug>/\n            if _entry.is_dir():\n                for _sub in sorted(_entry.iterdir()):\n                    if _looks_like_dataset_root(_sub):\n                        BASE_DIR = _sub\n                        break\n            if BASE_DIR is not None:\n                break\n\nif BASE_DIR is None:\n    print('[ERROR] Could not auto-detect the dataset root under /kaggle/input.')\n    print('        Attach the \"accident\" competition data via Add Input, or if it is')\n    print('        already attached, verify the actual layout below and set BASE_DIR')\n    print('        manually in this cell.')\n    _input_root = pathlib.Path('/kaggle/input')\n    if _input_root.exists():\n        print(f'[STATUS] Contents of {_input_root}:')\n        for _entry in sorted(_input_root.iterdir()):\n            print(f'    {_entry}')\n    else:\n        print('[STATUS] /kaggle/input does not exist -- not running on Kaggle, or no data attached.')\n    BASE_DIR = pathlib.Path('/kaggle/input/accident')  # placeholder so later cells don't NameError\n\nSYNTHETIC_DIR   = BASE_DIR / 'sim_dataset'    # synthetic CARLA dataset\nREAL_VIDEOS_DIR = BASE_DIR / 'videos'         # real CCTV test videos\nOUTPUT_DIR      = pathlib.Path('/kaggle/working')\nOUTPUT_DIR.mkdir(parents=True, exist_ok=True)\n\n# [SETUP] Key file paths\nLABELS_CSV            = SYNTHETIC_DIR / 'labels.csv'                # synthetic supervision index\nTEST_METADATA_CSV     = BASE_DIR      / 'test_metadata.csv'         # real test scene tags\nSAMPLE_SUBMISSION     = BASE_DIR      / 'sample_submission.csv'     # output schema reference\nANNOTATION_CLASSES    = SYNTHETIC_DIR / 'annotation_classes.yaml'   # segmentation class map\nSYNTHETIC_VIDEOS_DIR  = SYNTHETIC_DIR / 'videos'                    # {head-on,rear-end,sideswipe,single,t-bone}/\nVIDEO_ANNOTATIONS_DIR = SYNTHETIC_DIR / 'video_annotations'         # per-video annotation files\n\nCOLLISION_TYPES = ['head-on', 'rear-end', 'sideswipe', 'single', 't-bone']\n\n\ndef resolve_video_path(rgb_path: str) -> pathlib.Path:\n    \"\"\"Resolve an rgb_path value from labels.csv to an absolute video path.\n\n    rgb_path values are relative to sim_dataset/ (e.g. videos/head-on/clip.mp4),\n    but earlier dataset revisions used paths relative to the competition root.\n    Try both roots and return whichever exists, so downstream code never\n    silently operates on a non-existent path (an earlier revision of this\n    notebook lost 20/20 calibration videos to exactly that bug).\n    \"\"\"\n    p = pathlib.Path(rgb_path)\n    if p.is_absolute():\n        return p\n    for root in (SYNTHETIC_DIR, BASE_DIR):\n        candidate = root / p\n        if candidate.exists():\n            return candidate\n    return SYNTHETIC_DIR / p  # best guess; existence is asserted before inference\n\n\nprint('[STATUS] Path constants initialized')\nprint(f'  BASE_DIR        : {BASE_DIR}')\nprint(f'  SYNTHETIC_DIR   : {SYNTHETIC_DIR}')\nprint(f'  REAL_VIDEOS_DIR : {REAL_VIDEOS_DIR}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:28:47.575058Z","iopub.execute_input":"2026-08-03T11:28:47.575291Z","iopub.status.idle":"2026-08-03T11:28:47.592314Z","shell.execute_reply.started":"2026-08-03T11:28:47.575231Z","shell.execute_reply":"2026-08-03T11:28:47.591599Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [STATUS] File availability audit -- all critical paths verified before downstream steps\npath_checks = [\n    ('BASE_DIR',              BASE_DIR),\n    ('SYNTHETIC_DIR',         SYNTHETIC_DIR),\n    ('SYNTHETIC_VIDEOS_DIR',  SYNTHETIC_VIDEOS_DIR),\n    ('VIDEO_ANNOTATIONS_DIR', VIDEO_ANNOTATIONS_DIR),\n    ('REAL_VIDEOS_DIR',       REAL_VIDEOS_DIR),\n    ('labels.csv',            LABELS_CSV),\n    ('annotation_classes.yaml', ANNOTATION_CLASSES),\n    ('test_metadata.csv',     TEST_METADATA_CSV),\n    ('sample_submission.csv', SAMPLE_SUBMISSION),\n]\n\naudit_df = pd.DataFrame([\n    {\n        'label'  : label,\n        'path'   : str(p),\n        'exists' : p.exists(),\n        'kind'   : 'dir' if p.is_dir() else ('file' if p.is_file() else 'missing'),\n    }\n    for label, p in path_checks\n])\n\n# All rows should show exists=True before proceeding\ndisplay(audit_df.style.map(\n    lambda v: 'color: green; font-weight: bold' if v is True\n              else ('color: red; font-weight: bold' if v is False else ''),\n    subset=['exists']\n))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:28:47.593324Z","iopub.execute_input":"2026-08-03T11:28:47.593605Z","iopub.status.idle":"2026-08-03T11:28:47.704772Z","shell.execute_reply.started":"2026-08-03T11:28:47.593576Z","shell.execute_reply":"2026-08-03T11:28:47.703964Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [SETUP] Install model dependencies (requires internet on Kaggle)\n#   - ultralytics    : YOLOv8 vehicle detector (tracking stage)\n#   - CLIP           : zero-shot classification (Stage 2)\n#   - perception_models : Perception Encoder for temporal refinement (Stage 1)\n#   - transformers stack : Qwen2.5-VL (Stage 2) and OWLv2 (Stage 3)\nimport subprocess\nimport sys\nimport json as _json\n\n\ndef _sh(cmd):\n    r = subprocess.run(cmd, capture_output=True, text=True)\n    return r.returncode, r.stdout, r.stderr\n\n\ndef _torch_smoke_test():\n    \"\"\"Import torch and run ONE real op on cuda, in a fresh subprocess.\n\n    torch.cuda.is_available() returning True is not sufficient evidence the\n    GPU actually works: a torch build whose compiled kernels don't cover this\n    GPU's compute capability still reports CUDA as \"available\" and only fails\n    once a real kernel launch is attempted. Also captures get_device_name /\n    get_device_capability / get_arch_list, so an architecture mismatch is\n    diagnosed directly instead of guessed at from the exception text alone.\n    Runs in a subprocess -- not `import torch` in this cell -- because once\n    this kernel process has torch in sys.modules, a later pip install that\n    replaces the on-disk package would not change what an already-imported\n    `torch` in this process reports.\n    \"\"\"\n    code = (\n        \"import torch, json\\n\"\n        \"info = {'version': torch.__version__, 'cuda_available': torch.cuda.is_available()}\\n\"\n        \"try:\\n\"\n        \"    if info['cuda_available']:\\n\"\n        \"        info['device_name'] = torch.cuda.get_device_name(0)\\n\"\n        \"        info['device_capability'] = list(torch.cuda.get_device_capability(0))\\n\"\n        \"        info['torch_arch_list'] = torch.cuda.get_arch_list()\\n\"\n        \"        _ = (torch.tensor([1.0, 2.0], device='cuda') * 2).cpu()\\n\"\n        \"        _ = torch.prod(torch.tensor([2, 3, 4], dtype=torch.int64, device='cuda')).cpu()\\n\"\n        \"        info['cuda_op_ok'] = True\\n\"\n        \"    else:\\n\"\n        \"        info['cuda_op_ok'] = None\\n\"\n        \"except Exception as e:\\n\"\n        \"    info['cuda_op_ok'] = False\\n\"\n        \"    info['cuda_op_error'] = f'{type(e).__name__}: {e}'\\n\"\n        \"print(json.dumps(info))\\n\"\n    )\n    rc, out, err = _sh([sys.executable, '-c', code])\n    lines = [l for l in out.strip().splitlines() if l.strip()]\n    if lines:\n        try:\n            return _json.loads(lines[-1])\n        except _json.JSONDecodeError:\n            pass\n    return {'version': None, 'cuda_available': None, 'cuda_op_ok': False,\n            'raw_stdout': out[-500:], 'raw_stderr': err[-500:]}\n\n\ndef _diagnose_arch_mismatch(info):\n    \"\"\"Compare the GPU's actual compute capability against what this torch\n    build shipped kernels for, so the report says WHY, not just THAT it failed.\"\"\"\n    cap = info.get('device_capability')\n    arch_list = info.get('torch_arch_list') or []\n    if not cap or not arch_list:\n        return None\n    sm = f'sm_{cap[0]}{cap[1]}'\n    compiled_sms = {a.split('_')[0] + '_' + a.split('_')[1][:2] for a in arch_list if 'sm_' in a}\n    if sm not in compiled_sms and not any(a.startswith(f'compute_{cap[0]}{cap[1]}') for a in arch_list):\n        return (f\"GPU is {info.get('device_name')} (compute capability {cap[0]}.{cap[1]}, i.e. {sm}), \"\n                f\"but this torch build only has compiled kernels for: {arch_list}. \"\n                f\"{sm} is NOT in that list -- this is an architecture mismatch, not a driver problem.\")\n    return None\n\n\nprint('[STATUS] Baseline torch smoke test (before any installs)...')\n_baseline = _torch_smoke_test()\nprint(f'[STATUS] Baseline: {_baseline}')\n\n_repaired = False\nif _baseline.get('cuda_op_ok') is False:\n    _diag = _diagnose_arch_mismatch(_baseline)\n    if _diag:\n        print(f'[DIAGNOSIS] {_diag}')\n    else:\n        print('[DIAGNOSIS] cuda_available=True but a real op fails, and device_capability/'\n              'arch_list were not both captured (the failure happened before those calls) -- '\n              'consistent with either an architecture mismatch or a driver/runtime version '\n              'mismatch between this torch build and the host.')\n    print('[WARNING] A real CUDA op already fails before any installs run -- this is a '\n          'pre-existing problem with the environment\\'s preinstalled torch, not something '\n          'the installs below will cause. Attempting a repair: reinstalling torch/torchvision/'\n          'torchaudio from the stable cu121 channel, which has shipped sm_75 (T4) kernels '\n          'continuously and is a safer bet than whatever cu128 build shipped with this image.')\n    _rc, _out, _err = _sh([\n        sys.executable, '-m', 'pip', 'install', '-q', '--force-reinstall',\n        'torch', 'torchvision', 'torchaudio',\n        '--index-url', 'https://download.pytorch.org/whl/cu121',\n    ])\n    print(f'[STATUS] Repair install exit code: {_rc}')\n    if _rc != 0:\n        print(_err[-1500:])\n    _baseline = _torch_smoke_test()\n    print(f'[STATUS] Post-repair smoke test: {_baseline}')\n    if _baseline.get('cuda_op_ok') is True:\n        _repaired = True\n        print('[SUCCESS] Repair worked -- proceeding with the cu121 build for the rest of this session.')\n    else:\n        _diag2 = _diagnose_arch_mismatch(_baseline)\n        print('[ERROR] Repair did not fix it.', (_diag2 or ''))\n        print('        This looks like a Kaggle platform/accelerator issue rather than something')\n        print('        pip can fix from inside the notebook. Try: Session -> Factory reset, or')\n        print('        switch the accelerator (Settings -> Accelerator) to a different GPU option')\n        print('        and back, then Restart Kernel and re-run from the top. Do not proceed to')\n        print('        load any model on GPU until this smoke test reports cuda_op_ok=True.')\n\n# Pin torch/torchvision/torchaudio/triton to whatever is on disk RIGHT NOW (post-repair, if a\n# repair happened above), so nothing pulled in as a transitive dependency of ultralytics / CLIP /\n# perception_models / the transformers stack below is allowed to silently swap them again.\n_CONSTRAINTS_PATH = '/kaggle/working/_torch_constraints.txt'\n_pinned = []\nfor _pkg in ('torch', 'torchvision', 'torchaudio', 'triton'):\n    _rc, _out, _err = _sh([sys.executable, '-m', 'pip', 'show', _pkg])\n    if _rc == 0:\n        for _line in _out.splitlines():\n            if _line.startswith('Version:'):\n                _pinned.append(f\"{_pkg}=={_line.split(':', 1)[1].strip()}\")\nwith open(_CONSTRAINTS_PATH, 'w') as f:\n    f.write('\\n'.join(_pinned) + '\\n')\nprint(f'[STATUS] Pinning during install: {_pinned}')\n\n\ndef pip_install(args, label):\n    result = subprocess.run(\n        ['pip', 'install', '-q', '-c', _CONSTRAINTS_PATH] + args,\n        capture_output=True, text=True)\n    print(f'[STATUS] {label} install exit code: {result.returncode}')\n    if result.returncode != 0:\n        # A constraint conflict means this package genuinely needs a different\n        # torch than what's on disk. Surface it now and stop -- silently retrying\n        # without -c would defeat the whole point of pinning.\n        print(result.stderr[-2000:])\n    return result.returncode\n\n\npip_install(['ultralytics'], 'ultralytics (YOLOv8)')\npip_install(['git+https://github.com/openai/CLIP.git'], 'CLIP')\n\nresult = subprocess.run(\n    ['git', 'clone', 'https://github.com/facebookresearch/perception_models.git'],\n    capture_output=True, text=True\n)\nprint('[STATUS] perception_models clone exit code:', result.returncode)\npip_install(['-e', 'perception_models'], 'perception_models')\n\n# -U is not optional. Kaggle ships a transformers that predates Qwen3-VL, and\n# `pip install transformers` on an already-satisfied requirement is a no-op --\n# it exits 0, prints nothing, and leaves the old version in place. That is\n# exactly how a previous run reported \"install exit code: 0\" and then failed\n# with \"Transformers does not recognize qwen3_vl\", silently fell back to\n# constants for all 2027 clips, and produced a submission that scored nothing.\npip_install(['-U', 'transformers>=4.57', 'accelerate', 'qwen-vl-utils', 'bitsandbytes'],\n            'transformers stack (upgraded for Qwen3-VL)')\n\nprint('[STATUS] Post-install torch smoke test...')\n_after = _torch_smoke_test()\nprint(f'[STATUS] After installs: {_after}')\n\nif _baseline.get('version') != _after.get('version'):\n    print(f\"[ERROR] torch was changed by the installs above despite the -c constraints \"\n          f\"file: {_baseline.get('version')} -> {_after.get('version')}. One of the \"\n          f\"packages above must have forced this through some path the constraints file \"\n          f\"doesn't cover. Do not proceed to load models until this is understood -- \"\n          f\"Restart Kernel and re-run first.\")\nelif _after.get('cuda_op_ok') is False:\n    print(f\"[ERROR] torch version is unchanged ({_after.get('version')}) but a real CUDA \"\n          f\"op now fails: {_after.get('cuda_op_error')}. Since the pre-install baseline \"\n          f\"above already {'was fixed by the repair' if _repaired else 'had the same problem'}, \"\n          f\"this is consistent with an unstable environment rather than these installs -- \"\n          f\"do not proceed to Section 6.10 (Qwen3-VL) or CLIP loading below.\")\nelif _after.get('cuda_op_ok') is True:\n    print('[STATUS] torch unchanged (or successfully repaired) and a real CUDA op still '\n          'works -- safe to build on.')\n\nimport transformers\nfrom transformers.models.auto.configuration_auto import CONFIG_MAPPING_NAMES\n\nprint(f'[STATUS] transformers {transformers.__version__}')\nQWEN3_VL_OK = 'qwen3_vl' in CONFIG_MAPPING_NAMES\nprint(f'[STATUS] qwen3_vl architecture recognized: {QWEN3_VL_OK}')\nif not QWEN3_VL_OK:\n    print('[FATAL] This transformers cannot load Qwen3-VL. Section 6.10 will not run.')\n    print('        Restart the kernel after this cell (an upgrade does not take effect')\n    print('        in an already-imported module), or install from source:')\n    print('          pip install -U git+https://github.com/huggingface/transformers.git')\n\nimport torch\nDEVICE = 'cuda' if torch.cuda.is_available() else 'cpu'\nprint(f'[STATUS] torch {torch.__version__} | device: {DEVICE}')\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:28:47.706522Z","iopub.execute_input":"2026-08-03T11:28:47.707352Z","iopub.status.idle":"2026-08-03T11:33:59.432403Z","shell.execute_reply.started":"2026-08-03T11:28:47.707301Z","shell.execute_reply":"2026-08-03T11:33:59.431578Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"---\n## 1. Data Acquisition\n\nThis section loads all structured metadata from disk: the synthetic labels index, per-video annotation references, and the real test metadata. Raw binary video assets are accessed later via path references rather than bulk loading.","metadata":{}},{"cell_type":"code","source":"# [LOAD] Synthetic labels index -- primary supervision signal for pipeline calibration\nlabels_df = pd.read_csv(LABELS_CSV) if LABELS_CSV.exists() else pd.DataFrame()\n\n# [LOAD] Real test metadata -- coarse scene tags provided for analysis, not scoring\ntest_df = pd.read_csv(TEST_METADATA_CSV) if TEST_METADATA_CSV.exists() else pd.DataFrame()\n\ndim_report = pd.DataFrame({\n    'dataset' : ['synthetic_labels', 'test_metadata'],\n    'rows'    : [len(labels_df), len(test_df)],\n    'columns' : [labels_df.shape[1] if not labels_df.empty else 0,\n                 test_df.shape[1] if not test_df.empty else 0],\n})\ndisplay(dim_report)\nprint('[SUCCESS] Metadata loaded')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:33:59.433370Z","iopub.execute_input":"2026-08-03T11:33:59.433890Z","iopub.status.idle":"2026-08-03T11:33:59.490795Z","shell.execute_reply.started":"2026-08-03T11:33:59.433865Z","shell.execute_reply":"2026-08-03T11:33:59.490067Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [LOAD] Sample submission -- defines expected output schema for the final CSV\nsample_sub = pd.read_csv(SAMPLE_SUBMISSION) if SAMPLE_SUBMISSION.exists() else pd.DataFrame()\n\nprint('Submission columns:', list(sample_sub.columns))\ndisplay(sample_sub.head(3))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:33:59.491705Z","iopub.execute_input":"2026-08-03T11:33:59.492471Z","iopub.status.idle":"2026-08-03T11:34:00.417428Z","shell.execute_reply.started":"2026-08-03T11:33:59.492446Z","shell.execute_reply":"2026-08-03T11:34:00.416627Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [LOAD] Enumerate video files\n# Synthetic structure: sim_dataset/videos/{head-on,rear-end,sideswipe,single,t-bone}/*.mp4\nsynthetic_videos = []\nif SYNTHETIC_VIDEOS_DIR.exists():\n    for subdir in COLLISION_TYPES:\n        synthetic_videos.extend(sorted((SYNTHETIC_VIDEOS_DIR / subdir).glob('*.mp4')))\n\n# Real CCTV test videos: flat directory under videos/\nreal_videos = sorted(REAL_VIDEOS_DIR.glob('*.mp4')) if REAL_VIDEOS_DIR.exists() else []\n\ninventory_rows = [{'split': 'real_test', 'collision_type': 'all', 'count': len(real_videos)}]\nfor subdir in COLLISION_TYPES:\n    sp = SYNTHETIC_VIDEOS_DIR / subdir\n    inventory_rows.append({\n        'split': 'synthetic', 'collision_type': subdir,\n        'count': len(list(sp.glob('*.mp4'))) if sp.exists() else 0,\n    })\n\ndisplay(pd.DataFrame(inventory_rows))\nprint(f'[STATUS] Total synthetic videos : {len(synthetic_videos)}')\nprint(f'[STATUS] Total real test videos : {len(real_videos)}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:34:00.418515Z","iopub.execute_input":"2026-08-03T11:34:00.418823Z","iopub.status.idle":"2026-08-03T11:34:00.551466Z","shell.execute_reply.started":"2026-08-03T11:34:00.418789Z","shell.execute_reply":"2026-08-03T11:34:00.550684Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [LOAD] Annotation files inventory\n# video_annotations/ contains subdirs named like Town03_head-on_clear_00.json/\n# Each subdir holds the actual annotation file(s) -- recurse and filter to files only\nannotation_files = [\n    p for p in sorted(VIDEO_ANNOTATIONS_DIR.rglob('*')) if p.is_file()\n] if VIDEO_ANNOTATIONS_DIR.exists() else []\n\nann_subdirs = [\n    p for p in sorted(VIDEO_ANNOTATIONS_DIR.glob('*')) if p.is_dir()\n] if VIDEO_ANNOTATIONS_DIR.exists() else []\n\ndisplay(pd.DataFrame({\n    'metric' : ['annotation_dir_exists', 'annotation_subdirs', 'annotation_files', 'sample_subdir_name'],\n    'value'  : [VIDEO_ANNOTATIONS_DIR.exists(), len(ann_subdirs), len(annotation_files),\n                ann_subdirs[0].name if ann_subdirs else 'n/a'],\n}))\n\n# [LOAD] Inspect schema of the first annotation file (handles .json and .json.gz)\nif annotation_files:\n    first_ann = annotation_files[0]\n    open_fn = gzip.open if first_ann.suffix == '.gz' else open\n    with open_fn(first_ann, 'rt', encoding='utf-8') as f:\n        ann_sample = json.load(f)\n    top_keys = list(ann_sample.keys()) if isinstance(ann_sample, dict) else f'list[{len(ann_sample)}]'\n    print(f'[STATUS] {first_ann.name} top-level keys:', top_keys)\nelse:\n    print('[STATUS] No annotation files found')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:34:00.552404Z","iopub.execute_input":"2026-08-03T11:34:00.552706Z","iopub.status.idle":"2026-08-03T11:34:10.491606Z","shell.execute_reply.started":"2026-08-03T11:34:00.552667Z","shell.execute_reply":"2026-08-03T11:34:10.490715Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"---\n## 2. Data Inspection\n\nSystematic audit of all loaded data structures: schema review, null quantification, statistical profiling, and video-level properties (resolution, FPS, duration) extracted from a sample of clips for hardware-aware planning.","metadata":{}},{"cell_type":"code","source":"# [INSPECT] Schema of synthetic labels -- column names, dtypes, null counts\nif not labels_df.empty:\n    display(pd.DataFrame({\n        'column'    : labels_df.columns,\n        'dtype'     : labels_df.dtypes.values,\n        'null_count': labels_df.isnull().sum().values,\n        'null_pct'  : (labels_df.isnull().mean().values * 100).round(2),\n    }))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:34:10.492659Z","iopub.execute_input":"2026-08-03T11:34:10.492972Z","iopub.status.idle":"2026-08-03T11:34:10.507003Z","shell.execute_reply.started":"2026-08-03T11:34:10.492937Z","shell.execute_reply":"2026-08-03T11:34:10.506330Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [INSPECT] Statistical summary and sample rows\nif not labels_df.empty:\n    display(labels_df.describe().T.round(4))\n    display(labels_df.head(5))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:34:10.507930Z","iopub.execute_input":"2026-08-03T11:34:10.508337Z","iopub.status.idle":"2026-08-03T11:34:10.563113Z","shell.execute_reply.started":"2026-08-03T11:34:10.508314Z","shell.execute_reply":"2026-08-03T11:34:10.562485Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [INSPECT] Collision type distribution in synthetic labels\nif not labels_df.empty and 'type' in labels_df.columns:\n    type_counts = labels_df['type'].value_counts().reset_index()\n    type_counts.columns = ['collision_type', 'count']\n    type_counts['pct'] = (type_counts['count'] / type_counts['count'].sum() * 100).round(2)\n    display(type_counts)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:34:10.564144Z","iopub.execute_input":"2026-08-03T11:34:10.564540Z","iopub.status.idle":"2026-08-03T11:34:10.577415Z","shell.execute_reply.started":"2026-08-03T11:34:10.564507Z","shell.execute_reply":"2026-08-03T11:34:10.576602Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [INSPECT] Schema of real test metadata\nif not test_df.empty:\n    display(pd.DataFrame({\n        'column'    : test_df.columns,\n        'dtype'     : test_df.dtypes.values,\n        'null_count': test_df.isnull().sum().values,\n    }))\n    display(test_df.head(5))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:34:10.578402Z","iopub.execute_input":"2026-08-03T11:34:10.578699Z","iopub.status.idle":"2026-08-03T11:34:10.604735Z","shell.execute_reply.started":"2026-08-03T11:34:10.578668Z","shell.execute_reply":"2026-08-03T11:34:10.603903Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [INSPECT] Video-level property extraction from a sample of synthetic clips\n# Reads container metadata only -- no full video decoding\n\ndef extract_video_meta(video_path: pathlib.Path) -> dict:\n    \"\"\"Return FPS, frame count, width, height, duration for a single video file.\"\"\"\n    cap = cv2.VideoCapture(str(video_path))\n    if not cap.isOpened():\n        return {}\n    fps      = cap.get(cv2.CAP_PROP_FPS)\n    n_frames = int(cap.get(cv2.CAP_PROP_FRAME_COUNT))\n    width    = int(cap.get(cv2.CAP_PROP_FRAME_WIDTH))\n    height   = int(cap.get(cv2.CAP_PROP_FRAME_HEIGHT))\n    cap.release()\n    return {\n        'path'      : video_path.name,\n        'fps'       : round(fps, 2),\n        'n_frames'  : n_frames,\n        'width'     : width,\n        'height'    : height,\n        'duration_s': round(n_frames / fps, 2) if fps > 0 else 0.0,\n    }\n\nSAMPLE_N = min(20, len(synthetic_videos))\nvideo_meta_df = pd.DataFrame(\n    [r for r in (extract_video_meta(p) for p in synthetic_videos[:SAMPLE_N]) if r]\n)\n\nif not video_meta_df.empty:\n    display(video_meta_df.describe().T.round(2))\nelse:\n    print('[STATUS] No video metadata extracted -- check synthetic video paths')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:34:10.607713Z","iopub.execute_input":"2026-08-03T11:34:10.608067Z","iopub.status.idle":"2026-08-03T11:34:11.466586Z","shell.execute_reply.started":"2026-08-03T11:34:10.608043Z","shell.execute_reply":"2026-08-03T11:34:11.465765Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"---\n## 3. Data Cleaning\n\nTargeted cleaning of the synthetic labels index: duplicate removal, type-label normalization, out-of-bounds coordinate clamping, and target validation. Video paths are resolved with `resolve_video_path` (which checks both candidate roots) so every retained record points at a file that actually exists.","metadata":{}},{"cell_type":"code","source":"# [CLEAN] Work on a copy; drop exact duplicates; normalize type labels\nlabels_clean = labels_df.copy() if not labels_df.empty else pd.DataFrame()\npre_clean_rows = len(labels_clean)\n\nif not labels_clean.empty:\n    labels_clean = labels_clean.drop_duplicates()\n    if 'type' in labels_clean.columns:\n        labels_clean['type'] = labels_clean['type'].str.strip().str.lower()\n    print(f'[STATUS] Rows: {pre_clean_rows} -> {len(labels_clean)} after duplicate removal')\n    print('[STATUS] Type values:', sorted(labels_clean['type'].unique()))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:34:11.467512Z","iopub.execute_input":"2026-08-03T11:34:11.468215Z","iopub.status.idle":"2026-08-03T11:34:11.484321Z","shell.execute_reply.started":"2026-08-03T11:34:11.468179Z","shell.execute_reply":"2026-08-03T11:34:11.483669Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [CLEAN] Clamp spatial coordinates to [0, 1]; drop rows with missing or invalid targets\ncoord_cols       = ['center_x', 'center_y', 'x1', 'y1', 'x2', 'y2']\nrequired_targets = ['accident_time', 'center_x', 'center_y', 'type']\n\nif not labels_clean.empty:\n    for col in coord_cols:\n        if col in labels_clean.columns:\n            oob = ((labels_clean[col] < 0) | (labels_clean[col] > 1)).sum()\n            labels_clean[col] = labels_clean[col].clip(0.0, 1.0)\n            if oob > 0:\n                print(f'[STATUS] Clamped {oob} out-of-bounds values in column: {col}')\n\n    missing_mask = labels_clean[required_targets].isnull().any(axis=1)\n    labels_clean = labels_clean[~missing_mask].reset_index(drop=True)\n    print(f'[STATUS] Rows dropped due to missing targets: {missing_mask.sum()}')\n\n    if 'accident_time' in labels_clean.columns: \n        labels_clean = labels_clean[labels_clean['accident_time'] >= 0].reset_index(drop=True)\n\ndisplay(pd.DataFrame({\n    'stage'   : ['pre_clean', 'post_clean'],\n    'rows'    : [pre_clean_rows, len(labels_clean)],\n    'removed' : [0, pre_clean_rows - len(labels_clean)],\n}))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:34:11.485530Z","iopub.execute_input":"2026-08-03T11:34:11.485815Z","iopub.status.idle":"2026-08-03T11:34:11.517071Z","shell.execute_reply.started":"2026-08-03T11:34:11.485791Z","shell.execute_reply":"2026-08-03T11:34:11.516177Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [CLEAN] Cast frame index, resolve absolute video paths, encode type labels\nif not labels_clean.empty:\n    if 'accident_frame' in labels_clean.columns:\n        labels_clean['accident_frame'] = labels_clean['accident_frame'].astype(int)\n\n    if 'rgb_path' in labels_clean.columns:\n        labels_clean['abs_video_path'] = labels_clean['rgb_path'].map(\n            lambda p: str(resolve_video_path(p))\n        )\n        labels_clean['video_exists'] = labels_clean['abs_video_path'].map(\n            lambda p: pathlib.Path(p).exists()\n        )\n        n_missing = (~labels_clean['video_exists']).sum()\n        print(f'[STATUS] Records with unresolvable video paths: {n_missing}')\n        labels_clean = labels_clean[labels_clean['video_exists']].reset_index(drop=True)\n        display(labels_clean[['rgb_path', 'abs_video_path']].head(3))\n\n    if 'type' in labels_clean.columns:\n        TYPE_VOCAB  = sorted(labels_clean['type'].unique())\n        TYPE_TO_IDX = {t: i for i, t in enumerate(TYPE_VOCAB)}\n        IDX_TO_TYPE = {i: t for t, i in TYPE_TO_IDX.items()}\n        labels_clean['type_idx'] = labels_clean['type'].map(TYPE_TO_IDX)\n        display(pd.DataFrame({'type': TYPE_VOCAB, 'idx': range(len(TYPE_VOCAB))}))\n\n    print('[SUCCESS] Cleaning complete')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:34:11.518079Z","iopub.execute_input":"2026-08-03T11:34:11.518398Z","iopub.status.idle":"2026-08-03T11:34:14.203519Z","shell.execute_reply.started":"2026-08-03T11:34:11.518365Z","shell.execute_reply":"2026-08-03T11:34:14.202622Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"---\n## 4. Exploratory Data Analysis\n\nDistributions of all three prediction targets (temporal, spatial, type), scene-level tag distributions in the real test metadata, and correlation structure among numeric features.","metadata":{}},{"cell_type":"code","source":"# [EDA] Collision type frequency\nif not labels_clean.empty and 'type' in labels_clean.columns:\n    type_freq = labels_clean['type'].value_counts().reset_index()\n    type_freq.columns = ['collision_type', 'count']\n\n    fig, ax = plt.subplots(figsize=(10, 5))\n    sns.barplot(data=type_freq, x='collision_type', y='count',\n                palette=list(PALETTE.values())[:len(type_freq)], ax=ax)\n    ax.set_title('Collision Type Frequency (Synthetic Dataset)')\n    ax.set_xlabel('Collision Type')\n    ax.set_ylabel('Count')\n    ax.bar_label(ax.containers[0], fontsize=9)\n    plt.tight_layout()\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:34:14.204513Z","iopub.execute_input":"2026-08-03T11:34:14.204828Z","iopub.status.idle":"2026-08-03T11:34:14.497436Z","shell.execute_reply.started":"2026-08-03T11:34:14.204805Z","shell.execute_reply":"2026-08-03T11:34:14.496635Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [EDA] Accident time distribution -- histogram with KDE, and box plot per type\nif not labels_clean.empty and 'accident_time' in labels_clean.columns:\n    fig, axes = plt.subplots(1, 2, figsize=(14, 5))\n\n    sns.histplot(labels_clean['accident_time'], bins=30, kde=True,\n                 color=PALETTE['primary'], ax=axes[0])\n    axes[0].set_title('Accident Time Distribution (seconds)')\n    axes[0].set_xlabel('Accident Time (s)')\n    axes[0].set_ylabel('Count')\n\n    if 'type' in labels_clean.columns:\n        sns.boxplot(data=labels_clean, x='type', y='accident_time',\n                    palette='Set2', ax=axes[1])\n        axes[1].set_title('Accident Time by Collision Type')\n        axes[1].set_xlabel('Collision Type')\n        axes[1].set_ylabel('Accident Time (s)')\n        axes[1].tick_params(axis='x', rotation=30)\n\n    plt.tight_layout()\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:34:14.498436Z","iopub.execute_input":"2026-08-03T11:34:14.498742Z","iopub.status.idle":"2026-08-03T11:34:14.966081Z","shell.execute_reply.started":"2026-08-03T11:34:14.498704Z","shell.execute_reply":"2026-08-03T11:34:14.965166Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [EDA] Spatial distribution of impact points -- 2D scatter and KDE\nif not labels_clean.empty and all(c in labels_clean.columns for c in ['center_x', 'center_y']):\n    fig, axes = plt.subplots(1, 2, figsize=(14, 5))\n\n    for t_idx, (t_name, group) in enumerate(labels_clean.groupby('type')):\n        axes[0].scatter(group['center_x'], group['center_y'],\n                        label=t_name, alpha=0.5, s=20,\n                        color=list(PALETTE.values())[t_idx % len(PALETTE)])\n    axes[0].legend(fontsize=8)\n    axes[0].set_title('Impact Point Distribution (Normalized Frame Coordinates)')\n    axes[0].set_xlabel('center_x (0=left, 1=right)')\n    axes[0].set_ylabel('center_y (0=top, 1=bottom)')\n    axes[0].set_xlim(0, 1)\n    axes[0].set_ylim(0, 1)\n    axes[0].invert_yaxis()  # match image coordinate convention\n\n    sns.kdeplot(data=labels_clean, x='center_x', y='center_y',\n                fill=True, cmap='Blues', ax=axes[1])\n    axes[1].set_title('Impact Point Density (KDE)')\n    axes[1].set_xlim(0, 1)\n    axes[1].set_ylim(0, 1)\n    axes[1].invert_yaxis()\n\n    plt.tight_layout()\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:34:14.967223Z","iopub.execute_input":"2026-08-03T11:34:14.967558Z","iopub.status.idle":"2026-08-03T11:34:17.444239Z","shell.execute_reply.started":"2026-08-03T11:34:14.967520Z","shell.execute_reply":"2026-08-03T11:34:17.443407Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [EDA] Bounding box size distributions\nif not labels_clean.empty and all(c in labels_clean.columns for c in ['x1', 'y1', 'x2', 'y2']):\n    labels_clean['bbox_w']    = labels_clean['x2'] - labels_clean['x1']\n    labels_clean['bbox_h']    = labels_clean['y2'] - labels_clean['y1']\n    labels_clean['bbox_area'] = labels_clean['bbox_w'] * labels_clean['bbox_h']\n\n    fig, axes = plt.subplots(1, 3, figsize=(16, 5))\n    for ax, col, title in zip(\n        axes,\n        ['bbox_w', 'bbox_h', 'bbox_area'],\n        ['Bounding Box Width', 'Bounding Box Height', 'Bounding Box Area'],\n    ):\n        sns.histplot(labels_clean[col], bins=30, kde=True, color=PALETTE['secondary'], ax=ax)\n        ax.set_title(title)\n        ax.set_xlabel(col)\n        ax.set_ylabel('Count')\n\n    plt.tight_layout()\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:34:17.445682Z","iopub.execute_input":"2026-08-03T11:34:17.446240Z","iopub.status.idle":"2026-08-03T11:34:18.059175Z","shell.execute_reply.started":"2026-08-03T11:34:17.446211Z","shell.execute_reply":"2026-08-03T11:34:18.058464Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [EDA] Accident time as fraction of clip duration -- relative temporal position\nif not labels_clean.empty and all(c in labels_clean.columns for c in ['accident_time', 'duration']):\n    labels_clean['time_fraction'] = labels_clean['accident_time'] / labels_clean['duration']\n\n    fig, ax = plt.subplots(figsize=(10, 5))\n    sns.histplot(labels_clean['time_fraction'], bins=40, kde=True,\n                 color=PALETTE['tertiary'], ax=ax)\n    ax.set_title('Accident Time as Fraction of Clip Duration')\n    ax.set_xlabel('accident_time / duration (0=clip start, 1=clip end)')\n    ax.set_ylabel('Count')\n    ax.axvline(0.5, color=PALETTE['quaternary'], linestyle='--', label='midpoint')\n    ax.legend()\n    plt.tight_layout()\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:34:18.060088Z","iopub.execute_input":"2026-08-03T11:34:18.060416Z","iopub.status.idle":"2026-08-03T11:34:18.355029Z","shell.execute_reply.started":"2026-08-03T11:34:18.060374Z","shell.execute_reply":"2026-08-03T11:34:18.354302Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [EDA] Scene condition distributions from real test metadata\nSCENE_TAG_COLS = ['lighting', 'weather', 'layout']  # adjust to actual column names\n\nif not test_df.empty:\n    present_tags = [c for c in SCENE_TAG_COLS if c in test_df.columns]\n    if present_tags:\n        fig, axes = plt.subplots(1, len(present_tags), figsize=(6 * len(present_tags), 5))\n        if len(present_tags) == 1:\n            axes = [axes]\n        for ax, col in zip(axes, present_tags):\n            tag_freq = test_df[col].value_counts().reset_index()\n            tag_freq.columns = [col, 'count']\n            sns.barplot(data=tag_freq, x=col, y='count', palette='Set2', ax=ax)\n            ax.set_title(f'Test Set: {col.title()} Distribution')\n            ax.set_xlabel(col.title())\n            ax.set_ylabel('Count')\n            ax.tick_params(axis='x', rotation=30)\n        plt.tight_layout()\n        plt.show()\n    else:\n        print('[STATUS] Scene tag columns not found -- column names may differ')\nelse:\n    print('[STATUS] test_df empty -- skipping scene tag plots')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:34:18.356010Z","iopub.execute_input":"2026-08-03T11:34:18.356395Z","iopub.status.idle":"2026-08-03T11:34:18.513174Z","shell.execute_reply.started":"2026-08-03T11:34:18.356356Z","shell.execute_reply":"2026-08-03T11:34:18.512510Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [EDA] Correlation matrix across numeric features\nnumeric_cols = ['accident_time', 'accident_frame', 'center_x', 'center_y',\n                'x1', 'y1', 'x2', 'y2', 'duration', 'no_frames',\n                'height', 'width', 'bbox_w', 'bbox_h', 'bbox_area', 'time_fraction']\n\nif not labels_clean.empty:\n    present_numeric = [c for c in numeric_cols if c in labels_clean.columns]\n    corr_matrix = labels_clean[present_numeric].corr()\n\n    fig, ax = plt.subplots(figsize=(12, 9))\n    sns.heatmap(corr_matrix, annot=True, fmt='.2f', cmap='coolwarm',\n                center=0, linewidths=0.5, ax=ax)\n    ax.set_title('Correlation Matrix: Numeric Features (Synthetic Dataset)')\n    plt.tight_layout()\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:34:18.514097Z","iopub.execute_input":"2026-08-03T11:34:18.514574Z","iopub.status.idle":"2026-08-03T11:34:19.317799Z","shell.execute_reply.started":"2026-08-03T11:34:18.514548Z","shell.execute_reply":"2026-08-03T11:34:19.316999Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [EDA] Sample frame extraction from a synthetic video for visual inspection\ndef sample_frames(video_path: pathlib.Path, n_frames: int = 6) -> list:\n    \"\"\"Extract n_frames evenly spaced RGB frames from a video.\"\"\"\n    cap = cv2.VideoCapture(str(video_path))\n    total = int(cap.get(cv2.CAP_PROP_FRAME_COUNT))\n    frames = []\n    for idx in np.linspace(0, total - 1, n_frames, dtype=int):\n        cap.set(cv2.CAP_PROP_POS_FRAMES, int(idx))\n        ret, frame = cap.read()\n        if ret:\n            frames.append(cv2.cvtColor(frame, cv2.COLOR_BGR2RGB))\n    cap.release()\n    return frames\n\nif synthetic_videos:\n    sampled = sample_frames(synthetic_videos[0], n_frames=6)\n    if sampled:\n        fig, axes = plt.subplots(1, len(sampled), figsize=(18, 4))\n        for i, (ax, frame) in enumerate(zip(axes, sampled)):\n            ax.imshow(frame)\n            ax.set_title(f'Frame Sample {i+1}', fontsize=10)\n            ax.axis('off')\n        fig.suptitle(f'Sampled Frames: {synthetic_videos[0].name}')\n        plt.tight_layout()\n        plt.show()\nelse:\n    print('[STATUS] No synthetic videos found for frame sampling')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:34:19.318867Z","iopub.execute_input":"2026-08-03T11:34:19.319575Z","iopub.status.idle":"2026-08-03T11:34:21.944557Z","shell.execute_reply.started":"2026-08-03T11:34:19.319548Z","shell.execute_reply":"2026-08-03T11:34:21.943739Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"---\n## 5. Feature Engineering\n\nPer-video feature extractors used by the inference pipeline:\n\n1. **Frame-difference time series** -- sharp intensity changes indicating a collision (temporal anchor).\n2. **Optical-flow magnitude map** -- cumulative Farneback flow around the detected accident time (spatial signal and fallback).\n3. **Frame window extraction** -- shared utility that pulls RGB frames from an arbitrary time window for the vision-language models.","metadata":{}},{"cell_type":"code","source":"# [FEATURE] Frame-difference series -- captures sharp intensity changes indicating collision\n\ndef compute_frame_diff_series(video_path: pathlib.Path,\n                              resize_w: int = 320,\n                              resize_h: int = 180) -> np.ndarray:\n    \"\"\"Mean absolute frame difference per consecutive pair.\n\n    Returns a 1D array of length (n_frames - 1).\n    \"\"\"\n    cap = cv2.VideoCapture(str(video_path))\n    diffs, prev_gray = [], None\n\n    while True:\n        ret, frame = cap.read()\n        if not ret:\n            break\n        frame_small = cv2.resize(frame, (resize_w, resize_h))\n        gray = cv2.cvtColor(frame_small, cv2.COLOR_BGR2GRAY).astype(np.float32)\n        if prev_gray is not None:\n            diffs.append(np.mean(np.abs(gray - prev_gray)))\n        prev_gray = gray\n\n    cap.release()\n    return np.array(diffs, dtype=np.float32)\n\n\ndef score_temporal_anomaly(diff_series: np.ndarray, smooth_window: int = 5) -> np.ndarray:\n    \"\"\"Rolling-mean smoothing followed by z-score normalization.\"\"\"\n    smoothed = (pd.Series(diff_series)\n                .rolling(window=smooth_window, min_periods=1, center=True)\n                .mean().values)\n    return (smoothed - smoothed.mean()) / (smoothed.std() + 1e-8)\n\n\nprint('[STATUS] compute_frame_diff_series / score_temporal_anomaly defined')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:34:21.945658Z","iopub.execute_input":"2026-08-03T11:34:21.945979Z","iopub.status.idle":"2026-08-03T11:34:21.954652Z","shell.execute_reply.started":"2026-08-03T11:34:21.945955Z","shell.execute_reply":"2026-08-03T11:34:21.953828Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [FEATURE] Optical-flow magnitude map -- spatial localization signal\n\ndef compute_flow_magnitude_map(video_path: pathlib.Path,\n                               resize_w: int = 320,\n                               resize_h: int = 180,\n                               n_frames_context: int = 30,\n                               center_frame: int = None,\n                               flow_percentile: float = 90.0) -> np.ndarray:\n    \"\"\"Cumulative Farneback optical-flow magnitude over a window of frames\n    centered on center_frame, with percentile thresholding to suppress\n    background motion. Returns a 2D array of shape (resize_h, resize_w).\n    \"\"\"\n    cap = cv2.VideoCapture(str(video_path))\n    total = int(cap.get(cv2.CAP_PROP_FRAME_COUNT))\n\n    if center_frame is not None:\n        start_frame = max(0, center_frame - n_frames_context // 2)\n    else:\n        start_frame = max(0, total // 3)\n    cap.set(cv2.CAP_PROP_POS_FRAMES, start_frame)\n\n    mag_accum = np.zeros((resize_h, resize_w), dtype=np.float32)\n    prev_gray, count = None, 0\n\n    while count < n_frames_context:\n        ret, frame = cap.read()\n        if not ret:\n            break\n        frame_small = cv2.resize(frame, (resize_w, resize_h))\n        gray = cv2.cvtColor(frame_small, cv2.COLOR_BGR2GRAY)\n        if prev_gray is not None:\n            flow = cv2.calcOpticalFlowFarneback(\n                prev_gray, gray, None,\n                pyr_scale=0.5, levels=3, winsize=15,\n                iterations=3, poly_n=5, poly_sigma=1.2, flags=0)\n            mag, _ = cv2.cartToPolar(flow[..., 0], flow[..., 1])\n            mag_accum += mag\n        prev_gray = gray\n        count += 1\n\n    cap.release()\n\n    if mag_accum.max() > 0:\n        thresh = np.percentile(mag_accum, flow_percentile)\n        mag_accum[mag_accum < thresh] = 0.0\n\n    return mag_accum\n\n\nprint('[STATUS] compute_flow_magnitude_map defined')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:34:21.955663Z","iopub.execute_input":"2026-08-03T11:34:21.955989Z","iopub.status.idle":"2026-08-03T11:34:22.107117Z","shell.execute_reply.started":"2026-08-03T11:34:21.955950Z","shell.execute_reply":"2026-08-03T11:34:22.106343Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [FEATURE] Frame window extraction -- shared by all vision-language stages\n\ndef extract_frames_window(video_path: pathlib.Path,\n                          t_center: float,\n                          t_before: float = 1.0,\n                          t_after: float = 1.0,\n                          n_frames: int = 8,\n                          max_side: int = None) -> list:\n    \"\"\"Extract n_frames RGB PIL images evenly spaced over\n    [t_center - t_before, t_center + t_after] seconds.\n\n    An asymmetric window (t_before > t_after) lets classification models see\n    the approach trajectories that disambiguate collision types, not just the\n    post-impact wreckage. max_side optionally downscales frames to control\n    VLM vision-token count.\n    \"\"\"\n    cap = cv2.VideoCapture(str(video_path))\n    fps   = cap.get(cv2.CAP_PROP_FPS)\n    total = int(cap.get(cv2.CAP_PROP_FRAME_COUNT))\n    if fps <= 0 or total == 0:\n        cap.release()\n        return []\n\n    start_idx = max(0, int((t_center - t_before) * fps))\n    end_idx   = min(total - 1, int((t_center + t_after) * fps))\n    frame_idxs = sorted(set(np.linspace(start_idx, end_idx, n_frames, dtype=int)))\n\n    pil_frames = []\n    for idx in frame_idxs:\n        cap.set(cv2.CAP_PROP_POS_FRAMES, int(idx))\n        ret, frame = cap.read()\n        if not ret:\n            continue\n        img = PILImage.fromarray(cv2.cvtColor(frame, cv2.COLOR_BGR2RGB))\n        if max_side is not None:\n            w, h = img.size\n            scale = max_side / max(w, h)\n            if scale < 1.0:\n                img = img.resize((int(w * scale), int(h * scale)))\n        pil_frames.append(img)\n\n    cap.release()\n    return pil_frames\n\n\nprint('[STATUS] extract_frames_window defined')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:34:22.108404Z","iopub.execute_input":"2026-08-03T11:34:22.108751Z","iopub.status.idle":"2026-08-03T11:34:22.123360Z","shell.execute_reply.started":"2026-08-03T11:34:22.108718Z","shell.execute_reply":"2026-08-03T11:34:22.122510Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [FEATURE] Visualize frame difference and anomaly score for a synthetic sample\nif synthetic_videos:\n    sample_path = synthetic_videos[0]\n    diff_series = compute_frame_diff_series(sample_path)\n    anomaly     = score_temporal_anomaly(diff_series)\n\n    fig, axes = plt.subplots(2, 1, figsize=(14, 7), sharex=True)\n\n    axes[0].plot(diff_series, color=PALETTE['primary'], linewidth=0.8)\n    axes[0].fill_between(range(len(diff_series)), 0, diff_series,\n                         alpha=0.2, color=PALETTE['primary'])\n    axes[0].set_title('Raw Frame Difference Series')\n    axes[0].set_ylabel('Mean Absolute Difference')\n\n    axes[1].plot(anomaly, color=PALETTE['quaternary'], linewidth=0.8)\n    axes[1].axhline(2.0, linestyle='--', color='grey', linewidth=0.8, label='z=2 threshold')\n    axes[1].set_title('Temporal Anomaly Score (Z-Score)')\n    axes[1].set_xlabel('Frame Index')\n    axes[1].set_ylabel('Anomaly Z-Score')\n    axes[1].legend()\n\n    plt.tight_layout()\n    plt.show()\nelse:\n    print('[STATUS] No synthetic videos -- skipping frame diff visualization')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:34:22.124422Z","iopub.execute_input":"2026-08-03T11:34:22.124842Z","iopub.status.idle":"2026-08-03T11:34:23.872448Z","shell.execute_reply.started":"2026-08-03T11:34:22.124816Z","shell.execute_reply":"2026-08-03T11:34:23.871473Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [FEATURE] Visualize optical-flow magnitude map for a synthetic sample\nif synthetic_videos:\n    flow_map = compute_flow_magnitude_map(synthetic_videos[0])\n\n    fig, ax = plt.subplots(figsize=(10, 6))\n    im = ax.imshow(flow_map, cmap='inferno', aspect='auto')\n    plt.colorbar(im, ax=ax, label='Cumulative Optical Flow Magnitude')\n    ax.set_title('Optical Flow Magnitude Map (Spatial Localization Signal)')\n    ax.set_xlabel('X (pixels)')\n    ax.set_ylabel('Y (pixels)')\n    plt.tight_layout()\n    plt.show()\nelse:\n    print('[STATUS] No synthetic videos -- skipping flow map visualization')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:34:23.873612Z","iopub.execute_input":"2026-08-03T11:34:23.873936Z","iopub.status.idle":"2026-08-03T11:34:25.161132Z","shell.execute_reply.started":"2026-08-03T11:34:23.873911Z","shell.execute_reply":"2026-08-03T11:34:25.160191Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [FEATURE] CLIP text prompt templates for each collision type (zero-shot scoring)\nCOLLISION_PROMPTS = {\n    'rear-end': [\n        'a car colliding into the back of another car',\n        'rear-end collision between two vehicles on a road',\n        'vehicle hitting the back of a stationary car from behind',\n        'one car rear-ending another car at a traffic light',\n        'a vehicle crashing into the tail of the car ahead',\n    ],\n    't-bone': [\n        'a car hitting the side of another car at an intersection',\n        't-bone collision at a crossroads between two vehicles',\n        'side impact crash where one car strikes another perpendicularly',\n        'a vehicle running a red light and hitting the side of crossing traffic',\n        'perpendicular collision between two cars at a junction',\n    ],\n    'head-on': [\n        'two cars colliding head-on from opposite directions',\n        'frontal collision between two vehicles on a road',\n        'head-on crash between two cars driving toward each other',\n        'two vehicles smashing front-to-front on a highway',\n        'a car crossing the center line and hitting an oncoming vehicle head-on',\n    ],\n    'sideswipe': [\n        'two vehicles scraping alongside each other while driving',\n        'sideswipe collision between cars changing lanes',\n        'glancing blow between two cars moving in the same direction',\n        'a car drifting into the adjacent lane and scraping another vehicle',\n        'two vehicles brushing sides while traveling parallel on a road',\n    ],\n    'single': [\n        'a single car crashing into a wall or barrier',\n        'one vehicle running off the road and hitting an obstacle',\n        'a car losing control and crashing into a pole or guardrail',\n        'a single vehicle spinning out and hitting a roadside object',\n        'one car veering off the road and crashing without involving another vehicle',\n    ],\n}\n\ndisplay(pd.DataFrame([\n    {'collision_type': k, 'n_prompts': len(v)} for k, v in COLLISION_PROMPTS.items()\n]))\nprint('[STATUS] Collision prompt templates defined')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:34:25.162187Z","iopub.execute_input":"2026-08-03T11:34:25.162714Z","iopub.status.idle":"2026-08-03T11:34:25.172956Z","shell.execute_reply.started":"2026-08-03T11:34:25.162690Z","shell.execute_reply":"2026-08-03T11:34:25.172328Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"---\n## 6. Modeling\n\nThe pipeline is built around one structural insight: **an accident is an event between two specific objects**, so global-frame signals (whole-frame differencing, whole-frame optical flow, whole-frame CLIP) are fundamentally noise-limited. The primary signal is therefore **per-vehicle detection and tracking**, and all three questions are answered from the kinematics of the colliding pair:\n\n| Stage | Question | Primary signal (tracking) | Fallback (zero-shot models) |\n|-------|----------|---------------------------|------------------------------|\n| 1 -- *When* | accident time | first box contact + joint deceleration peak | frame-diff anchor + bounded PE refinement |\n| 2 -- *What* | collision type | angle between pre-impact track velocities + contact geometry | CLIP + Qwen2.5-VL soft ensemble |\n| 3 -- *Where* | impact point | centroid of the box-intersection region at impact | OWLv2 grounding + flow centroid |\n\nTwo design rules carried over from the failed experiments documented in Section 7b:\n\n1. **Geometry is only trusted when its preconditions hold** (long, confident, fast-moving tracks). An earlier angle classifier computed on raw optical flow failed precisely because the flow field mixed background traffic and compression noise into the velocity estimate; track velocities come from the two involved objects only.\n2. **No hard overrides.** The geometric type hint enters a soft ensemble with CLIP and Qwen; the tracking impact time is cross-checked against the frame-difference anchor and discarded entirely when the two disagree by more than a gate (a distant \"contact\" is usually a false pair from occlusion, not a crash).","metadata":{}},{"cell_type":"markdown","source":"### 6.1 Model Loading\n\nTwo models load by default -- YOLOv8s (~0.5 GB, tracking) and CLIP ViT-B/32 (~0.6 GB, baseline classification fallback). `LOAD_LEGACY_MODELS = False` (Section 6.1 switches) keeps the Perception Encoder, Qwen2.5-VL-3B, and OWLv2 off by default -- Section 7b measured all three losing to a constant predictor, and the code that called them has been removed from the notebook (previously ~4.9 GB loaded for signals nothing used). Their measured numbers are quoted in Section 6.1's model-switch cell and in Sections 6.4 / 6.6 below, kept as documented negative results.\n","metadata":{}},{"cell_type":"code","source":"# [MODEL] Model loading switches\n#\n# The VRAM audit measured 5.40 GB held by the five original models before\n# Qwen3-VL-8B even started loading -- which left ~3.8 GB of a 16 GB T4 for\n# activations and produced an OOM on the first clip. Qwen3-VL itself is only\n# ~6.8 GB in 4-bit; the rest was models nothing calls any more.\n#\n# Section 7b measured each of them against its constant: the Perception Encoder\n# (T=0.31 vs 0.38), OWLv2/optical-flow (S=0.15 vs 0.22), and Qwen2.5-VL-3B\n# (C=0.15, below the 0.20 chance level) all lost. The loading code and the\n# functions that called them (Stage 1b, 2b, 2c-ensemble, 3b) have been removed\n# from the notebook -- see the correction notes in Sections 6.4 and 6.6. Only\n# their measured numbers are kept, here and in those two sections.\nLOAD_YOLO = True             # still used by Section 7's comparison, ~0.5 GB\nLOAD_CLIP = True             # still used by Section 7's comparison, ~0.6 GB\n\n# [MODEL] YOLOv8 -- fast per-frame vehicle detector (primary tracking signal)\nYOLO_AVAILABLE = True\ntry:\n    from ultralytics import YOLO\n    yolo_model = YOLO('yolov8s.pt')   # COCO-pretrained; small = good speed/accuracy balance\n    print('[SUCCESS] YOLOv8s loaded')\nexcept Exception as e:\n    YOLO_AVAILABLE = False\n    print(f'[ERROR] YOLOv8 unavailable: {e}')\n\n# COCO class ids for vehicles: car, motorcycle, bus, truck\nVEHICLE_CLASS_IDS = {2, 3, 5, 7}","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:34:25.173785Z","iopub.execute_input":"2026-08-03T11:34:25.174592Z","iopub.status.idle":"2026-08-03T11:34:25.810246Z","shell.execute_reply.started":"2026-08-03T11:34:25.174535Z","shell.execute_reply":"2026-08-03T11:34:25.809343Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [MODEL] CLIP ViT-B/32 -- zero-shot classification backbone\ntry:\n    import clip\n    clip_model, clip_preprocess = clip.load('ViT-B/32', device=DEVICE)\n    clip_model.eval()\n    CLIP_AVAILABLE = True\nexcept Exception as e:\n    CLIP_AVAILABLE = False\n    print(f'[ERROR] CLIP unavailable: {e}')\n\nif CLIP_AVAILABLE:\n    # Pre-encode all collision-type text prompts once.\n    # Following the CLIP paper's prompt-ensembling recipe, the per-type mean\n    # embedding is RE-NORMALIZED after averaging. Without this, types whose\n    # prompts happen to cluster tightly get a systematically larger-norm mean\n    # vector and therefore inflated similarity scores.\n    TYPE_TEXT_FEATURES = {}\n    with torch.no_grad():\n        for ctype, prompts in COLLISION_PROMPTS.items():\n            tokens   = clip.tokenize(prompts).to(DEVICE)\n            features = clip_model.encode_text(tokens)                    # (n_prompts, 512)\n            features = features / features.norm(dim=-1, keepdim=True)\n            mean_feat = features.mean(dim=0)\n            TYPE_TEXT_FEATURES[ctype] = mean_feat / mean_feat.norm()\n    print(f'[SUCCESS] CLIP ViT-B/32 loaded | text features for {len(TYPE_TEXT_FEATURES)} types pre-computed')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:34:25.811312Z","iopub.execute_input":"2026-08-03T11:34:25.811788Z","iopub.status.idle":"2026-08-03T11:34:34.176089Z","shell.execute_reply.started":"2026-08-03T11:34:25.811759Z","shell.execute_reply":"2026-08-03T11:34:34.175239Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [MODEL] VRAM audit -- the two loaded models (YOLOv8s + CLIP) must fit a single T4 (16 GB)\nif torch.cuda.is_available():\n    print(f'[STATUS] GPU: {torch.cuda.get_device_name(0)}')\n    print(f'[STATUS] VRAM allocated: {torch.cuda.memory_allocated(0)/1e9:.2f} GB | '\n          f'reserved: {torch.cuda.memory_reserved(0)/1e9:.2f} GB | '\n          f'total: {torch.cuda.get_device_properties(0).total_memory/1e9:.2f} GB')\nelse:\n    print('[STATUS] No CUDA device -- everything will run on CPU (very slow)')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:34:34.177054Z","iopub.execute_input":"2026-08-03T11:34:34.177544Z","iopub.status.idle":"2026-08-03T11:34:34.183103Z","shell.execute_reply.started":"2026-08-03T11:34:34.177516Z","shell.execute_reply":"2026-08-03T11:34:34.182372Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 6.2 Vehicle Detection and Tracking (SORT-lite)\n\nDetections are associated frame-to-frame with a minimal SORT-style tracker. The mathematical components:\n\n**Association (Hungarian assignment).** Given predicted track boxes $\\hat{B}_i$ and detections $D_j$, solve the linear assignment problem\n\n$$\\min_{x_{ij}} \\sum_{i,j} c_{ij}\\, x_{ij}, \\qquad c_{ij} = 1 - \\mathrm{IoU}(\\hat{B}_i, D_j), \\qquad \\mathrm{IoU}(A,B) = \\frac{|A \\cap B|}{|A \\cup B|}$$\n\nvia the Kuhn-Munkres algorithm (`scipy.optimize.linear_sum_assignment`), rejecting matches below an IoU gate.\n\n**Motion prediction (constant velocity).** Each track's next box is its last box translated by the mean of its recent center displacements -- a zeroth-order Kalman surrogate that is adequate at 10 Hz sampling.\n\n**Velocity estimation (central differences on smoothed centers).** After moving-average smoothing of the center sequence $c_i$, per-sample velocity uses the nonuniform central difference\n\n$$v_i = \\frac{c_{i+1} - c_{i-1}}{t_{i+1} - t_{i-1}}$$\n\n(`np.gradient`), which is second-order accurate and does not phase-shift the estimate the way a forward difference does.","metadata":{}},{"cell_type":"code","source":"# [TRACK] IoU, box gap, and a SORT-style tracker with two-stage association\nfrom scipy.optimize import linear_sum_assignment\n\n\ndef iou_xyxy(a, b) -> float:\n    \"\"\"Intersection-over-union of two [x1, y1, x2, y2] boxes.\"\"\"\n    ix1, iy1 = max(a[0], b[0]), max(a[1], b[1])\n    ix2, iy2 = min(a[2], b[2]), min(a[3], b[3])\n    inter = max(0.0, ix2 - ix1) * max(0.0, iy2 - iy1)\n    area_a = (a[2] - a[0]) * (a[3] - a[1])\n    area_b = (b[2] - b[0]) * (b[3] - b[1])\n    return inter / (area_a + area_b - inter + 1e-9)\n\n\ndef box_gap(a, b) -> float:\n    \"\"\"Euclidean separation between two boxes (0 when they touch or overlap).\"\"\"\n    gx = max(0.0, max(a[0] - b[2], b[0] - a[2]))\n    gy = max(0.0, max(a[1] - b[3], b[1] - a[3]))\n    return float(np.hypot(gx, gy))\n\n\nclass SortLiteTracker:\n    \"\"\"Minimal SORT-style multi-object tracker with ByteTrack-style association.\n\n    Constant-velocity prediction on box centers + Hungarian assignment with an\n    IoU gate. CCTV cameras are static and vehicles move smoothly, so a full\n    Kalman filter adds little at 10 Hz sampling.\n\n    Association runs in two passes over confidence-split detections. This is the\n    one idea from ByteTrack that matters for this task: at the moment of impact\n    the vehicles occlude and deform, detector confidence collapses, and a\n    single-pass tracker with a hard confidence floor drops both tracks at\n    exactly the frame the whole pipeline is trying to measure. The second pass\n    sustains existing tracks from low-confidence boxes; those boxes are never\n    allowed to spawn new tracks, so the noise does not leak in.\n    \"\"\"\n\n    def __init__(self, iou_gate: float = 0.30, iou_gate_low: float = 0.20,\n                 max_missed: int = 10, conf_high: float = 0.50):\n        self.iou_gate     = iou_gate\n        self.iou_gate_low = iou_gate_low\n        self.max_missed   = max_missed   # 1.0 s at 10 Hz: a crash occludes for longer\n        self.conf_high    = conf_high    # than the old 0.5 s, which split tracks in two\n        self.tracks       = {}   # id -> {'boxes': [], 'frames': [], 'confs': [], 'missed': int}\n        self.finished     = {}\n        self._next_id     = 0\n\n    def _predict(self, tr):\n        \"\"\"Last box translated by the mean of the last <=3 center displacements.\"\"\"\n        boxes = tr['boxes']\n        if len(boxes) < 2:\n            return boxes[-1]\n        centers = np.array([[(b[0] + b[2]) / 2, (b[1] + b[3]) / 2] for b in boxes[-4:]])\n        d = np.diff(centers, axis=0).mean(axis=0)\n        b = boxes[-1]\n        return [b[0] + d[0], b[1] + d[1], b[2] + d[0], b[3] + d[1]]\n\n    def _match(self, tids, detections, gate):\n        \"\"\"Hungarian match of track ids to detections above an IoU gate.\n\n        Returns (pairs, unmatched_track_ids, unmatched_det_indices).\n        \"\"\"\n        if not tids or not detections:\n            return [], set(tids), set(range(len(detections)))\n\n        cost = np.ones((len(tids), len(detections)))\n        for i, tid in enumerate(tids):\n            pred = self._predict(self.tracks[tid])\n            for j, (box, _) in enumerate(detections):\n                cost[i, j] = 1.0 - iou_xyxy(pred, box)\n\n        rows, cols = linear_sum_assignment(cost)\n        pairs, matched_t, matched_d = [], set(), set()\n        for i, j in zip(rows, cols):\n            if 1.0 - cost[i, j] >= gate:\n                pairs.append((tids[i], j))\n                matched_t.add(tids[i])\n                matched_d.add(j)\n        return pairs, set(tids) - matched_t, set(range(len(detections))) - matched_d\n\n    def _append(self, tid, det, frame_idx):\n        box, conf = det\n        tr = self.tracks[tid]\n        tr['boxes'].append(list(box))\n        tr['frames'].append(frame_idx)\n        tr['confs'].append(conf)\n        tr['missed'] = 0\n\n    def update(self, detections, frame_idx):\n        \"\"\"detections: list of ([x1,y1,x2,y2], conf) for one sampled frame.\"\"\"\n        high = [d for d in detections if d[1] >= self.conf_high]\n        low  = [d for d in detections if d[1] <  self.conf_high]\n\n        # Pass 1: confident detections, strict gate.\n        pairs_h, unmatched_t, unmatched_h = self._match(\n            list(self.tracks.keys()), high, self.iou_gate)\n        for tid, j in pairs_h:\n            self._append(tid, high[j], frame_idx)\n\n        # Pass 2: whatever is left gets a shot at the low-confidence boxes.\n        pairs_l, unmatched_t, _ = self._match(\n            sorted(unmatched_t), low, self.iou_gate_low)\n        for tid, j in pairs_l:\n            self._append(tid, low[j], frame_idx)\n\n        # Age out tracks that matched nothing in either pass.\n        for tid in sorted(unmatched_t):\n            self.tracks[tid]['missed'] += 1\n            if self.tracks[tid]['missed'] > self.max_missed:\n                self.finished[tid] = self.tracks.pop(tid)\n\n        # New tracks spawn from confident detections only.\n        for j in sorted(unmatched_h):\n            box, conf = high[j]\n            self.tracks[self._next_id] = {\n                'boxes': [list(box)], 'frames': [frame_idx],\n                'confs': [conf], 'missed': 0,\n            }\n            self._next_id += 1\n\n    def all_tracks(self, min_len: int = 5) -> dict:\n        \"\"\"All tracks (finished + live) with at least min_len observations.\"\"\"\n        out = {}\n        for tid, tr in {**self.finished, **self.tracks}.items():\n            if len(tr['frames']) >= min_len:\n                out[tid] = {\n                    'boxes' : np.array(tr['boxes'], dtype=np.float32),\n                    'frames': np.array(tr['frames'], dtype=np.int64),\n                    'confs' : np.array(tr['confs'], dtype=np.float32),\n                }\n        return out\n\n\nprint('[STATUS] SortLiteTracker defined (two-stage association)')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:34:34.184220Z","iopub.execute_input":"2026-08-03T11:34:34.184617Z","iopub.status.idle":"2026-08-03T11:34:34.205563Z","shell.execute_reply.started":"2026-08-03T11:34:34.184590Z","shell.execute_reply":"2026-08-03T11:34:34.204734Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [TRACK] Run detection + tracking over a video, then compute per-track kinematics\n\nfrom scipy.signal import savgol_filter\n\n# Two vehicles at the moment of impact overlap heavily -- often past the default\n# NMS threshold of 0.7, which then deletes one of them as a duplicate box. That\n# removes the collision pair from the exact frame the pipeline exists to find.\n# 0.9 keeps both; the cost is a few extra duplicate boxes elsewhere, which the\n# tracker's IoU gate absorbs.\nYOLO_NMS_IOU  = 0.9\nYOLO_CONF_MIN = 0.10   # floor for entering the tracker at all; the tracker's\n                       # own conf_high splits high/low internally\n\n\ndef extract_vehicle_tracks(video_path: pathlib.Path,\n                           sample_fps: float = 10.0,\n                           conf_thresh: float = YOLO_CONF_MIN,\n                           min_track_len: int = 5):\n    \"\"\"Detect vehicles on frames sampled at sample_fps and link them into tracks.\n\n    Returns (tracks, fps, (W, H)) where tracks maps id -> arrays of boxes\n    (pixels), frame indices, and detection confidences.\n    \"\"\"\n    cap = cv2.VideoCapture(str(video_path))\n    fps      = cap.get(cv2.CAP_PROP_FPS)\n    n_frames = int(cap.get(cv2.CAP_PROP_FRAME_COUNT))\n    W = int(cap.get(cv2.CAP_PROP_FRAME_WIDTH))\n    H = int(cap.get(cv2.CAP_PROP_FRAME_HEIGHT))\n    if fps <= 0 or n_frames == 0:\n        cap.release()\n        return {}, fps, (W, H)\n\n    step = max(1, int(round(fps / sample_fps)))\n    tracker = SortLiteTracker()\n\n    # Sequential decode (no seeking) -- much faster than per-frame cap.set\n    frame_idx = 0\n    while True:\n        ret, frame = cap.read()\n        if not ret:\n            break\n        if frame_idx % step == 0:\n            results = yolo_model(frame, iou=YOLO_NMS_IOU, conf=conf_thresh,\n                                 verbose=False)[0]\n            detections = []\n            for box, cls, conf in zip(results.boxes.xyxy.cpu().numpy(),\n                                      results.boxes.cls.cpu().numpy(),\n                                      results.boxes.conf.cpu().numpy()):\n                if int(cls) in VEHICLE_CLASS_IDS:\n                    detections.append((box.tolist(), float(conf)))\n            tracker.update(detections, frame_idx)\n        frame_idx += 1\n\n    cap.release()\n    return tracker.all_tracks(min_len=min_track_len), fps, (W, H)\n\n\ndef track_kinematics(track: dict, fps: float, smooth_window: int = 5) -> dict:\n    \"\"\"Smoothed center trajectory and central-difference velocity for one track.\"\"\"\n    boxes = track['boxes']\n    t  = track['frames'] / fps\n    cx = (boxes[:, 0] + boxes[:, 2]) / 2\n    cy = (boxes[:, 1] + boxes[:, 3]) / 2\n\n    if len(cx) >= smooth_window:\n        # Savitzky-Golay, NOT np.convolve(mode='same') and not a rolling mean.\n        # Both distort the ends of the trajectory: convolve zero-pads, dragging\n        # the first and last samples hundreds of pixels toward the origin, and a\n        # centred rolling mean shrinks them by averaging over a truncated window.\n        # Tracks typically END at the collision (the vehicle stops, the tracker\n        # loses it) and _impact_time searches for the peak joint deceleration --\n        # so an edge artifact lands precisely on the signal being measured and\n        # fabricates an impact. savgol with mode='interp' fits a polynomial to\n        # the edge samples instead of inventing data, and reproduces a constant\n        # velocity exactly (measured: 0.0 px/s error, against 50 for a rolling\n        # mean and 680 for convolve).\n        #\n        # savgol assumes uniform spacing; a track with missed frames is not\n        # strictly uniform. The velocity below therefore still uses np.gradient\n        # against the real timestamps, and the residual error at a gap is\n        # unbiased rather than systematic.\n        cx = savgol_filter(cx, smooth_window, 2, mode='interp')\n        cy = savgol_filter(cy, smooth_window, 2, mode='interp')\n\n    vx = np.gradient(cx, t)\n    vy = np.gradient(cy, t)\n    return {\n        't': t, 'cx': cx, 'cy': cy, 'vx': vx, 'vy': vy,\n        'speed': np.hypot(vx, vy),\n        'frames': track['frames'], 'boxes': boxes,\n        'conf': float(np.median(track['confs'])),\n    }\n\n\nprint('[STATUS] extract_vehicle_tracks / track_kinematics defined')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:34:34.206653Z","iopub.execute_input":"2026-08-03T11:34:34.207030Z","iopub.status.idle":"2026-08-03T11:34:34.241176Z","shell.execute_reply.started":"2026-08-03T11:34:34.207003Z","shell.execute_reply":"2026-08-03T11:34:34.240442Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 6.3 Collision Analysis from Track Kinematics\n\n**Contact detection.** For each pair of concurrently tracked vehicles, the box gap $g(t)$ is monitored on their shared time support; contact is the first sample where $g(t)$ falls below 1% of the frame diagonal.\n\n**Impact time refinement.** Within $\\pm0.6$ s of contact, the impact instant is the peak of joint deceleration\n\n$$t^{*} = \\arg\\max_{t} \\left[ -\\frac{d}{dt}\\left( \\lVert v_A(t) \\rVert + \\lVert v_B(t) \\rVert \\right) \\right]$$\n\n-- physically, the moment momentum is exchanged.\n\n**Impact point.** The centroid of the intersection of the two boxes at $t^{*}$ (boxes dilated slightly when they only touch), normalized by frame size.\n\n**Type from geometry.** With pre-impact mean velocities $\\bar{v}_A, \\bar{v}_B$ over $[t^{*}-1.2, t^{*}-0.2]$ s, the approach angle is\n\n$$\\theta = \\arccos\\left( \\frac{\\bar{v}_A \\cdot \\bar{v}_B}{\\lVert \\bar{v}_A \\rVert\\, \\lVert \\bar{v}_B \\rVert} \\right)$$\n\nand the decision uses $\\theta$ plus the contact bearing (is the struck point ahead of, beside, or behind each vehicle relative to its own motion):\n\n| Condition | Type |\n|-----------|------|\n| $\\theta \\geq 135^{\\circ}$ | head-on |\n| $\\theta \\leq 40^{\\circ}$, contact along the leader's motion axis | rear-end |\n| $\\theta \\leq 40^{\\circ}$, contact lateral | sideswipe |\n| $60^{\\circ} \\leq \\theta \\leq 120^{\\circ}$ | t-bone |\n| otherwise (ambiguous bands) | no hint -- defer to VLM ensemble |\n\n**Single-vehicle crashes** need no partner: a track whose speed collapses by more than 60% within a short window, with no other vehicle in contact, is a `single` candidate.\n\n**Gating.** Geometry is only computed when both tracks are long ($\\geq 6$ samples), confident (median detection confidence $\\geq 0.4$), and moving ($\\lVert \\bar{v} \\rVert$ above a floor) -- exactly the preconditions whose absence sank the raw-optical-flow version of this idea.","metadata":{}},{"cell_type":"code","source":"# [TRACK] Collision-pair analysis -- When / What / Where from track kinematics\n\nMIN_SPEED_PX_S   = 15.0   # velocity direction is meaningless below this speed floor\nMIN_TRACK_CONF   = 0.40\nMIN_PAIR_SAMPLES = 6\n\n\ndef _pair_contact(kinA: dict, kinB: dict, frame_diag: float):\n    \"\"\"First contact between two tracks on their shared time support.\n\n    Returns (t_contact, idxA, idxB) or None. Contact = box gap below 1% of\n    the frame diagonal.\n    \"\"\"\n    shared, ia, ib = np.intersect1d(kinA['frames'], kinB['frames'], return_indices=True)\n    if len(shared) < MIN_PAIR_SAMPLES:\n        return None\n\n    gap_thresh = 0.01 * frame_diag\n    gaps = np.array([box_gap(kinA['boxes'][ia[k]], kinB['boxes'][ib[k]])\n                     for k in range(len(shared))])\n    for k in range(len(shared)):\n        if gaps[k] > gap_thresh:\n            continue\n        # Below the threshold but still closing means the vehicles have not\n        # touched yet (e.g. a slow lateral drift before a sideswipe) --\n        # contact is declared at the closest approach, not at first proximity.\n        # Once boxes actually touch the gap sits at 0 and cannot keep\n        # decreasing, so this always terminates at or before the true contact.\n        still_approaching = (k + 1 < len(shared)\n                             and gaps[k + 1] < gaps[k] - 0.02 * gap_thresh)\n        if not still_approaching:\n            return kinA['t'][ia[k]], ia[k], ib[k]\n    return None\n\n\ndef _impact_time(kinA: dict, kinB: dict, t_contact: float) -> float:\n    \"\"\"Peak joint deceleration within +/-0.6 s of contact: the moment of\n    momentum exchange, sharper than the contact sample itself.\"\"\"\n    grid = np.union1d(kinA['t'], kinB['t'])\n    grid = grid[(grid >= t_contact - 0.6) & (grid <= t_contact + 0.6)]\n    if len(grid) < 3:\n        return t_contact\n\n    joint_speed = (np.interp(grid, kinA['t'], kinA['speed'])\n                   + np.interp(grid, kinB['t'], kinB['speed']))\n    decel = -np.gradient(joint_speed, grid)\n    return float(grid[int(np.argmax(decel))])\n\n\ndef _impact_point(boxA, boxB, W: int, H: int) -> tuple:\n    \"\"\"Centroid of the box intersection at impact, normalized to [0, 1].\n    Boxes are dilated by 2% of the frame diagonal when they only touch.\"\"\"\n    pad = 0.02 * np.hypot(W, H)\n    for p in (0.0, pad):\n        ix1 = max(boxA[0] - p, boxB[0] - p)\n        iy1 = max(boxA[1] - p, boxB[1] - p)\n        ix2 = min(boxA[2] + p, boxB[2] + p)\n        iy2 = min(boxA[3] + p, boxB[3] + p)\n        if ix2 > ix1 and iy2 > iy1:\n            return (float(np.clip((ix1 + ix2) / 2 / W, 0, 1)),\n                    float(np.clip((iy1 + iy2) / 2 / H, 0, 1)))\n    # Disjoint even after dilation: midpoint of the two centers\n    cxa, cya = (boxA[0] + boxA[2]) / 2, (boxA[1] + boxA[3]) / 2\n    cxb, cyb = (boxB[0] + boxB[2]) / 2, (boxB[1] + boxB[3]) / 2\n    return (float(np.clip((cxa + cxb) / 2 / W, 0, 1)),\n            float(np.clip((cya + cyb) / 2 / H, 0, 1)))\n\n\ndef _pre_impact_velocity(kin: dict, t_impact: float):\n    \"\"\"Mean velocity over [t_impact - 1.2, t_impact - 0.2] s (None if unobserved).\"\"\"\n    mask = (kin['t'] >= t_impact - 1.2) & (kin['t'] <= t_impact - 0.2)\n    if mask.sum() < 2:\n        return None\n    return np.array([kin['vx'][mask].mean(), kin['vy'][mask].mean()])\n\n\ndef _type_from_geometry(vA, vB, cA, cB) -> str:\n    \"\"\"Collision type from approach angle + contact bearing; None when ambiguous.\"\"\"\n    nA, nB = np.linalg.norm(vA), np.linalg.norm(vB)\n    if nA < MIN_SPEED_PX_S or nB < MIN_SPEED_PX_S:\n        return None\n\n    cos_theta = np.clip(np.dot(vA, vB) / (nA * nB), -1.0, 1.0)\n    theta = np.degrees(np.arccos(cos_theta))\n\n    if theta >= 135.0:\n        return 'head-on'\n    if theta <= 40.0:\n        # Same direction: distinguish rear-end (contact along the motion axis)\n        # from sideswipe (lateral contact) via the bearing of B from A\n        u = (cB - cA) / (np.linalg.norm(cB - cA) + 1e-9)\n        longitudinal = abs(np.dot(vA / nA, u))\n        return 'rear-end' if longitudinal >= 0.7 else 'sideswipe'\n    if 60.0 <= theta <= 120.0:\n        return 't-bone'\n    return None  # 40-60 and 120-135 degree bands: defer to the VLM ensemble\n\n\n# Feature block for the supervised type classifier of Section 6.8.\n#\n# Every feature here is an angle, a ratio, or a speed normalized by the frame\n# diagonal. Nothing is in raw pixels and nothing touches image content. That is\n# deliberate: these are the quantities that survive the CARLA-to-CCTV domain\n# shift. A classifier trained on CARLA *pixels* latches onto render style and\n# collapses on real footage; a classifier trained on \"the two velocity vectors\n# met at 93 degrees and the slower one was doing 0.2 diagonals/second\" is\n# reading physics that renders and real cameras agree on.\nFEATURE_COLS = [\n    'theta_deg', 'cos_theta', 'speed_ratio', 'speed_a_norm', 'speed_b_norm',\n    'rel_speed_norm', 'longitudinal', 'drop_a', 'drop_b', 'area_ratio',\n    'evidence', 'is_single', 'n_tracks',\n]\n\n\ndef _speed_at(kin, t):\n    return float(np.interp(t, kin['t'], kin['speed']))\n\n\ndef _pair_features(kA, kB, t_imp, vA, vB, cA, cB, frame_diag, evidence, n_tracks):\n    \"\"\"Domain-invariant kinematics of a collision pair. NaN where undefined --\n    HistGradientBoostingClassifier consumes NaN natively, so an absent signal\n    stays absent instead of being imputed into a lie.\"\"\"\n    f = {c: np.nan for c in FEATURE_COLS}\n    f['evidence']  = float(evidence)\n    f['is_single'] = 0.0\n    f['n_tracks']  = float(n_tracks)\n\n    a_pre,  b_pre  = _speed_at(kA, t_imp - 0.6), _speed_at(kB, t_imp - 0.6)\n    a_post, b_post = _speed_at(kA, t_imp + 0.6), _speed_at(kB, t_imp + 0.6)\n    f['speed_a_norm'] = a_pre / frame_diag\n    f['speed_b_norm'] = b_pre / frame_diag\n    f['speed_ratio']  = min(a_pre, b_pre) / (max(a_pre, b_pre) + 1e-9)\n    f['drop_a'] = (a_pre - a_post) / (a_pre + 1e-9)\n    f['drop_b'] = (b_pre - b_post) / (b_pre + 1e-9)\n\n    areaA = float((kA['boxes'][:, 2] - kA['boxes'][:, 0]).mean()\n                  * (kA['boxes'][:, 3] - kA['boxes'][:, 1]).mean())\n    areaB = float((kB['boxes'][:, 2] - kB['boxes'][:, 0]).mean()\n                  * (kB['boxes'][:, 3] - kB['boxes'][:, 1]).mean())\n    f['area_ratio'] = min(areaA, areaB) / (max(areaA, areaB) + 1e-9)\n\n    if vA is not None and vB is not None:\n        nA, nB = np.linalg.norm(vA), np.linalg.norm(vB)\n        f['rel_speed_norm'] = float(np.linalg.norm(vA - vB) / frame_diag)\n        if nA >= MIN_SPEED_PX_S and nB >= MIN_SPEED_PX_S:\n            cos_t = float(np.clip(np.dot(vA, vB) / (nA * nB), -1.0, 1.0))\n            f['cos_theta'] = cos_t\n            f['theta_deg'] = float(np.degrees(np.arccos(cos_t)))\n            u = (cB - cA) / (np.linalg.norm(cB - cA) + 1e-9)\n            f['longitudinal'] = float(abs(np.dot(vA / nA, u)))\n    return f\n\n\ndef _single_features(kin, t0, frame_diag, evidence, n_tracks):\n    f = {c: np.nan for c in FEATURE_COLS}\n    f['evidence']  = float(evidence)\n    f['is_single'] = 1.0\n    f['n_tracks']  = float(n_tracks)\n    pre, post = _speed_at(kin, t0 - 0.4), _speed_at(kin, t0 + 0.4)\n    f['speed_a_norm'] = pre / frame_diag\n    f['drop_a'] = (pre - post) / (pre + 1e-9)\n    return f\n\n\ndef analyze_video_collision(video_path: pathlib.Path) -> dict:\n    \"\"\"Full tracking-based collision analysis for one video.\n\n    Returns {'found': False, 'evidence': 0.0} when no collision is detected,\n    else {'found': True, 'mode': 'pair'|'single', 't_impact': float,\n     'center': (cx, cy), 'type_hint': str|None, 'evidence': float}.\n\n    'evidence' is the normalized kinetic evidence for the contact and is the\n    gate the pipeline uses to decide whether to trust this result at all.\n    \"\"\"\n    if not YOLO_AVAILABLE:\n        return {'found': False, 'evidence': 0.0, 'features': None, 'n_tracks': 0}\n\n    tracks, fps, (W, H) = extract_vehicle_tracks(video_path)\n    if not tracks or fps <= 0:\n        return {'found': False, 'evidence': 0.0, 'features': None, 'n_tracks': 0}\n\n    frame_diag = float(np.hypot(W, H))\n    kins = {tid: track_kinematics(tr, fps) for tid, tr in tracks.items()}\n    kins = {tid: k for tid, k in kins.items() if k['conf'] >= MIN_TRACK_CONF}\n\n    # --- Pair collisions: pick the contact with the strongest kinetic evidence ---\n    best = None\n    tids = sorted(kins.keys())\n    for i in range(len(tids)):\n        for j in range(i + 1, len(tids)):\n            kA, kB = kins[tids[i]], kins[tids[j]]\n            contact = _pair_contact(kA, kB, frame_diag)\n            if contact is None:\n                continue\n            t_contact, ia, ib = contact\n            t_imp = _impact_time(kA, kB, t_contact)\n\n            # Kinetic evidence = approach speed x post-contact speed drop.\n            # Parked/slow pairs and drive-past occlusions score near zero.\n            pre  = (np.interp(t_imp - 0.6, kA['t'], kA['speed'])\n                    + np.interp(t_imp - 0.6, kB['t'], kB['speed']))\n            post = (np.interp(t_imp + 0.6, kA['t'], kA['speed'])\n                    + np.interp(t_imp + 0.6, kB['t'], kB['speed']))\n            evidence = pre * max(0.0, pre - post)\n            if pre < MIN_SPEED_PX_S:\n                continue\n            # Filter out passing vehicles that don't decelerate\n            speed_drop_ratio = (pre - post) / (pre + 1e-9)\n            if speed_drop_ratio < 0.15:\n                continue\n            if best is None or evidence > best['evidence']:\n                best = {'kA': kA, 'kB': kB, 'ia': ia, 'ib': ib,\n                        't_impact': t_imp, 'evidence': evidence}\n\n    if best is not None:\n        kA, kB = best['kA'], best['kB']\n        t_imp  = best['t_impact']\n        center = _impact_point(kA['boxes'][best['ia']], kB['boxes'][best['ib']], W, H)\n\n        vA = _pre_impact_velocity(kA, t_imp)\n        vB = _pre_impact_velocity(kB, t_imp)\n        cA = np.array([np.interp(t_imp, kA['t'], kA['cx']),\n                       np.interp(t_imp, kA['t'], kA['cy'])])\n        cB = np.array([np.interp(t_imp, kB['t'], kB['cx']),\n                       np.interp(t_imp, kB['t'], kB['cy'])])\n        type_hint = None\n        if vA is not None and vB is not None:\n            type_hint = _type_from_geometry(vA, vB, cA, cB)\n\n        # Evidence is normalized by the frame diagonal squared so the gate is\n        # resolution-independent and comparable across videos. It replaces the\n        # old \"agrees with the frame-difference anchor\" gate: that anchor scores\n        # T=0.31 against a constant's 0.52, so it was validating tracking\n        # against a signal weaker than a constant.\n        evidence = float(best['evidence'] / (frame_diag ** 2))\n        return {'found': True, 'mode': 'pair', 't_impact': round(float(t_imp), 4),\n                'center': center, 'type_hint': type_hint, 'evidence': evidence,\n                'n_tracks': len(kins),\n                'features': _pair_features(kA, kB, t_imp, vA, vB, cA, cB,\n                                           frame_diag, evidence, len(kins))}\n\n    # --- Single-vehicle crash: speed collapse with no partner in contact ---\n    best_single = None\n    for tid, kin in kins.items():\n        if len(kin['t']) < MIN_PAIR_SAMPLES:\n            continue\n        for k in range(len(kin['t'])):\n            t0 = kin['t'][k]\n            pre_mask  = (kin['t'] >= t0 - 0.8) & (kin['t'] < t0)\n            post_mask = (kin['t'] > t0) & (kin['t'] <= t0 + 0.8)\n            if pre_mask.sum() < 2 or post_mask.sum() < 2:\n                continue\n            pre, post = kin['speed'][pre_mask].mean(), kin['speed'][post_mask].mean()\n            if pre < 2 * MIN_SPEED_PX_S:\n                continue\n            drop = (pre - post) / (pre + 1e-9)\n            if drop >= 0.6 and (best_single is None or drop > best_single['drop']):\n                box = kin['boxes'][k]\n                best_single = {\n                    'drop': drop, 't_impact': round(float(t0), 4), 'kin': kin, 't0': t0,\n                    'evidence': float(pre * (pre - post) / (frame_diag ** 2)),\n                    'center': (float(np.clip((box[0] + box[2]) / 2 / W, 0, 1)),\n                               float(np.clip((box[1] + box[3]) / 2 / H, 0, 1))),\n                }\n\n    if best_single is not None:\n        return {'found': True, 'mode': 'single',\n                't_impact': best_single['t_impact'],\n                'center': best_single['center'], 'type_hint': 'single',\n                'evidence': best_single['evidence'], 'n_tracks': len(kins),\n                'features': _single_features(best_single['kin'], best_single['t0'],\n                                             frame_diag, best_single['evidence'],\n                                             len(kins))}\n\n    # n_tracks is reported even on failure: \"tracking found nothing\" and\n    # \"the detector saw no vehicles at all\" are different diagnoses, and on\n    # unlabelled real footage this is the only way to tell them apart.\n    return {'found': False, 'evidence': 0.0, 'features': None,\n            'n_tracks': len(kins)}\n\n\nprint('[STATUS] analyze_video_collision defined')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:34:34.242173Z","iopub.execute_input":"2026-08-03T11:34:34.242832Z","iopub.status.idle":"2026-08-03T11:34:34.282119Z","shell.execute_reply.started":"2026-08-03T11:34:34.242787Z","shell.execute_reply":"2026-08-03T11:34:34.281039Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 6.4 Stage 1 -- When: Temporal Localization (fallback chain)\n\n> **Correction -- the Perception-Encoder refinement described below is removed from the pipeline.** `run_inference_baseline` calls the plain frame-difference anchor (`predict_accident_time`) directly; the bounded-PE correction function (`predict_accident_time_refined`) was never wired into any pipeline that is actually called and has been deleted from the notebook. Kept here as a documented negative result: the *coarse-to-fine bounded-correction pattern itself* (anchor + capped delta) is what Section 6.10's `stage2_time_refine` reuses for the VLM, on evidence this cell originally provided.\n\nWhen tracking finds a collision, its impact time is used **only if it agrees with the frame-difference anchor within a gate** ($|t_{track} - t_{base}| \\leq 2.5$ s). Agreement of two independent signals is strong evidence; disagreement means the tracked \"contact\" is probably an occlusion crossing, so the whole tracking result is discarded for that video.\n\nThe (removed) PE refinement was a contrastive prompt ensemble -- mean similarity to accident prompts minus mean similarity to normal-traffic prompts -- bounded to $\\delta=2$ s:\n\n$$t_{final} = t_{base} + \\mathrm{clip}(t_{PE} - t_{base},\\ -\\delta,\\ +\\delta)$$\n\nOn calibration the bounded version scored T=0.61 vs 0.44 for the anchor alone and 0.16 for an unbounded PE scan -- the number that motivated capping corrections everywhere else in the notebook, even though the PE signal itself did not make it into the final pipeline.\n","metadata":{}},{"cell_type":"code","source":"# [MODEL] Stage 1a -- frame-difference anchor\n\ndef predict_accident_time(video_path: pathlib.Path,\n                          smooth_window: int = 5,\n                          z_threshold: float = 1.5) -> float:\n    \"\"\"Accident time in seconds via peak detection on frame-difference z-scores.\n\n    Among frames exceeding z_threshold, selects the strongest anomaly;\n    falls back to the global argmax if no frame exceeds the threshold.\n    \"\"\"\n    cap = cv2.VideoCapture(str(video_path))\n    fps      = cap.get(cv2.CAP_PROP_FPS)\n    n_frames = int(cap.get(cv2.CAP_PROP_FRAME_COUNT))\n    cap.release()\n\n    if fps <= 0 or n_frames == 0:\n        return 0.0\n\n    diff_series = compute_frame_diff_series(video_path)\n    if len(diff_series) == 0:\n        return n_frames / fps / 2.0  # midpoint fallback\n\n    anomaly = score_temporal_anomaly(diff_series, smooth_window)\n\n    candidates = np.where(anomaly > z_threshold)[0]\n    if len(candidates) == 0:\n        peak_frame = int(np.argmax(anomaly))\n    else:\n        peak_frame = int(candidates[np.argmax(anomaly[candidates])])\n\n    return round(peak_frame / fps, 4)\n\n\nprint('[STATUS] predict_accident_time defined')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:34:34.283403Z","iopub.execute_input":"2026-08-03T11:34:34.284141Z","iopub.status.idle":"2026-08-03T11:34:34.298397Z","shell.execute_reply.started":"2026-08-03T11:34:34.284111Z","shell.execute_reply":"2026-08-03T11:34:34.297580Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [MODEL] Stage 1c -- Object Size Dynamics (OSD): official ACCIDENT@CVPR baseline\n# signal, ensembled with the frame-diff z-score to anchor the VLM pipeline (6.10)\n\nfrom scipy.signal import medfilt\n\n\ndef compute_object_size_series(video_path: pathlib.Path,\n                               sample_fps: float = 10.0,\n                               conf_thresh: float = YOLO_CONF_MIN):\n    \"\"\"Total vehicle bounding-box area per sampled frame.\n\n    Collisions produce sudden, context-independent changes in tracked-object\n    geometry -- boxes overlap, merge, or resize sharply from impact deformation\n    and mutual occlusion -- so total box area is a direct proxy for a collision\n    event. This is the official ACCIDENT@CVPR baseline temporal signal (Object\n    Size Dynamics). It reuses the YOLO detector already loaded for tracking\n    (Section 6.2) rather than pixel intensity, so it stays robust to\n    compression artifacts and low light the way frame-differencing (Stage 1a)\n    is not -- the two signals fail in different conditions, which is exactly\n    why they are worth ensembling below rather than picking one.\n\n    Returns (frame_idxs, areas, fps). areas[i] is 0.0 for a sampled frame with\n    no detected vehicle, not a gap in the series -- the caller's smoothing\n    step is expected to be robust to that (see score_osd_anomaly).\n    \"\"\"\n    if not YOLO_AVAILABLE:\n        return np.array([]), np.array([]), 0.0\n\n    cap = cv2.VideoCapture(str(video_path))\n    fps      = cap.get(cv2.CAP_PROP_FPS)\n    n_frames = int(cap.get(cv2.CAP_PROP_FRAME_COUNT))\n    if fps <= 0 or n_frames == 0:\n        cap.release()\n        return np.array([]), np.array([]), fps\n\n    step = max(1, int(round(fps / sample_fps)))\n    frame_idxs, areas = [], []\n    frame_idx = 0\n    while True:\n        ret, frame = cap.read()\n        if not ret:\n            break\n        if frame_idx % step == 0:\n            results = yolo_model(frame, iou=YOLO_NMS_IOU, conf=conf_thresh, verbose=False)[0]\n            total_area = 0.0\n            for box, cls in zip(results.boxes.xyxy.cpu().numpy(),\n                                results.boxes.cls.cpu().numpy()):\n                if int(cls) in VEHICLE_CLASS_IDS:\n                    x1, y1, x2, y2 = box\n                    total_area += max(0.0, x2 - x1) * max(0.0, y2 - y1)\n            frame_idxs.append(frame_idx)\n            areas.append(total_area)\n        frame_idx += 1\n    cap.release()\n    return np.array(frame_idxs), np.array(areas, dtype=np.float32), fps\n\n\ndef score_osd_anomaly(area_series: np.ndarray, smooth_window: int = 5) -> np.ndarray:\n    \"\"\"Median-filter smoothing, then z-score -- deliberately NOT the rolling\n    MEAN used for frame-diff (score_temporal_anomaly). A missed detection\n    drops the raw area to 0 for a single frame; a mean smears that dip across\n    the whole window, a median filter discards it as the outlier it is.\n    \"\"\"\n    n = len(area_series)\n    if n < 3:\n        return np.zeros(n, dtype=np.float32)\n    w = smooth_window if smooth_window % 2 == 1 else smooth_window + 1\n    w = min(w, n if n % 2 == 1 else n - 1)\n    smoothed = medfilt(area_series, kernel_size=w) if w >= 3 else area_series\n    return (smoothed - smoothed.mean()) / (smoothed.std() + 1e-8)\n\n\ndef predict_accident_time_osd(video_path: pathlib.Path,\n                              sample_fps: float = 10.0,\n                              smooth_window: int = 5,\n                              z_threshold: float = 1.5):\n    \"\"\"Accident time via peak detection on the OSD z-score series. Same\n    threshold-then-argmax rule as Stage 1a (predict_accident_time), so the two\n    signals combine on equal footing in the ensemble below. Returns None\n    (not a fallback time) when YOLO is unavailable or detects nothing, so the\n    caller can tell 'no OSD signal' apart from 'OSD says t=0.0'.\n    \"\"\"\n    frame_idxs, areas, fps = compute_object_size_series(video_path, sample_fps)\n    if len(areas) == 0 or fps <= 0:\n        return None\n\n    anomaly = score_osd_anomaly(areas, smooth_window)\n    candidates = np.where(anomaly > z_threshold)[0]\n    peak = int(candidates[np.argmax(anomaly[candidates])]) if len(candidates) else int(np.argmax(anomaly))\n    return round(float(frame_idxs[peak] / fps), 4)\n\n\ndef predict_accident_time_ensemble(video_path: pathlib.Path,\n                                   smooth_window: int = 5,\n                                   z_threshold: float = 1.5) -> float:\n    \"\"\"Official-baseline ensemble: mean of the frame-diff z-score time (Stage\n    1a) and the OSD time. Both are classical, VLM-free signals, so this is a\n    cheap anchor for Section 6.10's VLM pipeline -- it is not fooled by camera\n    shake (frame-diff's failure mode) or by the detector losing every vehicle\n    to occlusion at the exact moment of impact (OSD's failure mode)\n    simultaneously, which is the point of combining rather than choosing one.\n    \"\"\"\n    t_diff = predict_accident_time(video_path, smooth_window, z_threshold)\n    t_osd  = predict_accident_time_osd(video_path, sample_fps=10.0,\n                                       smooth_window=smooth_window, z_threshold=z_threshold)\n    if t_osd is None:\n        return t_diff\n    return round((t_diff + t_osd) / 2.0, 4)\n\n\nprint('[STATUS] compute_object_size_series / predict_accident_time_osd / '\n      'predict_accident_time_ensemble defined')\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:34:34.299508Z","iopub.execute_input":"2026-08-03T11:34:34.300160Z","iopub.status.idle":"2026-08-03T11:34:34.315043Z","shell.execute_reply.started":"2026-08-03T11:34:34.300134Z","shell.execute_reply":"2026-08-03T11:34:34.314301Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 6.5 Stage 2 -- What: Collision Type Classification\n\n> **Correction -- this section is built on a false premise, and it measures worse than the baseline it replaced.** The claim below that \"the harmonic mean is 0 whenever the type is wrong\" is **wrong**: the leaderboard averages C across videos *before* the harmonic mean (see Section 7), so a wrong type costs 1/N of C, not the whole video. Classification is in fact the *lowest*-leverage component of the three (+0.022 for C: 0.30→0.45, versus +0.111 for S: 0.156→0.45). The ensemble described here scores **C = 0.15 against the CLIP-only baseline's 0.30** -- below the 0.20 chance level on a balanced set, and below the 0.359 of a constant \"rear-end\" predictor. It collapses to `rear-end` on 18/20 calibration videos. Retained as a documented negative result; see the diagnosis below and Section 7b.\n>\n> **Why it collapses.** The three ensemble terms are not on the same scale. Qwen has only 3 prompts, so `vote_fraction` is quantised to {0, ⅓, ⅔, 1} and a unanimous vote contributes the full 0.35. CLIP's probability mass is spread over 5 classes, so its argmax contributes only 0.35 × p_max ≈ 0.21. Whenever Qwen is unanimous it wins outright -- even against a geometry hint at 0.30. And the three prompts (`count-first`, `motion`, `geometry`) ask nearly the same question over the same frames, so they are not three independent votes but one vote counted three times, which is exactly what makes unanimity the default. Qwen is fed 6 stills at `max_side=448` from 1920×1080 far-field CCTV, where vehicles span a few dozen pixels: it cannot resolve the vehicles, let alone their relative motion, so it falls back on a language prior in which `rear-end` is the most common accident type. The `t-bone → rear-end` scene rule and the `return 'rear-end'` fallbacks push the same direction.\n\nClassification was assumed to dominate the leaderboard metric. The final classifier is a **soft ensemble of three independent zero-shot signals**:\n\n$$\\mathrm{score}(c) = w_{geo}\\,\\mathbb{1}[c = c_{track}] + w_{clip}\\,P_{CLIP}(c) + w_{vlm}\\,\\frac{\\#\\,\\mathrm{votes}(c)}{\\#\\,\\mathrm{votes}}$$\n\nwith $(w_{geo}, w_{clip}, w_{vlm}) = (0.30, 0.35, 0.35)$ when a geometric hint exists, and $(0, 0.5, 0.5)$ otherwise. Signals:\n\n1. **Track geometry** (Section 6.3) -- most reliable when available, but abstains on ambiguous angle bands and poor tracks.\n2. **CLIP prompt-ensemble probabilities** -- image features mean-pooled over frames around the impact, softmaxed against re-normalized per-type prompt embeddings.\n3. **Qwen2.5-VL votes** -- three structured prompts over a chronological frame sequence spanning the *approach* (asymmetric window), with class definitions spelled out and a count-first prompt that targets the `single` class.\n\nA scene-plausibility rule then remaps `t-bone` to `rear-end` when the scene is a highway or tunnel (no perpendicular cross-traffic exists there).\n\n**Compute-adaptive fast path (T4 budget).** Qwen is by far the most expensive signal (~10 s/video for three prompts on a T4). When the track geometry and CLIP *independently agree*, the label is accepted without querying Qwen: two agreeing independent signals already outvote a single potential dissent, so the extra call cannot change the argmax under the ensemble weights. This cuts total runtime substantially on the easy majority of videos while spending full compute only where the signals conflict.","metadata":{}},{"cell_type":"code","source":"# [MODEL] Stage 2a -- CLIP zero-shot type probabilities\n\ndef clip_type_probabilities(video_path: pathlib.Path, peak_time_s: float) -> dict:\n    \"\"\"Per-type probability distribution from CLIP similarity.\n\n    Image features are mean-pooled over frames in a +/-0.75 s window around\n    the impact (multi-frame pooling suppresses single-frame motion blur),\n    then compared against the pre-computed per-type text embeddings.\n    Returns {type: probability} or None when CLIP is unavailable.\n    \"\"\"\n    if not CLIP_AVAILABLE:\n        return None\n\n    pil_frames = extract_frames_window(video_path, peak_time_s,\n                                       t_before=0.75, t_after=0.75, n_frames=8)\n    if not pil_frames:\n        return None\n\n    with torch.no_grad():\n        img_tensors  = torch.stack([clip_preprocess(f) for f in pil_frames]).to(DEVICE)\n        img_features = clip_model.encode_image(img_tensors)              # (n_frames, 512)\n        img_features = img_features / img_features.norm(dim=-1, keepdim=True)\n        img_feature  = img_features.mean(dim=0)\n        img_feature  = img_feature / img_feature.norm()\n\n    types = list(TYPE_TEXT_FEATURES.keys())\n    sims  = np.array([(img_feature @ TYPE_TEXT_FEATURES[t]).item() for t in types])\n\n    # Softmax with CLIP's standard temperature (100) turns raw cosine\n    # similarities (narrow ~0.2-0.3 range) into a usable distribution.\n    logits = 100.0 * sims\n    probs  = np.exp(logits - logits.max())\n    probs /= probs.sum()\n    return dict(zip(types, probs))\n\n\ndef predict_collision_type_clip(video_path: pathlib.Path, peak_time_s: float) -> str:\n    \"\"\"Argmax of the CLIP distribution; used standalone by the baseline pipeline.\"\"\"\n    probs = clip_type_probabilities(video_path, peak_time_s)\n    if probs is None:\n        return 'rear-end'  # most common type; last-resort fallback\n    return max(probs, key=probs.get)\n\n\nprint('[STATUS] clip_type_probabilities / predict_collision_type_clip defined')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:34:34.315993Z","iopub.execute_input":"2026-08-03T11:34:34.316737Z","iopub.status.idle":"2026-08-03T11:34:34.332579Z","shell.execute_reply.started":"2026-08-03T11:34:34.316696Z","shell.execute_reply":"2026-08-03T11:34:34.331942Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 6.6 Stage 3 -- Where: Impact Point Localization (fallback chain)\n\n> **Correction -- OWLv2 grounding described below is removed from the pipeline.** `run_inference_baseline` calls `predict_impact_location_flow` (optical-flow centroid) directly; `predict_impact_location_grounded` was never wired into any pipeline that is actually called and has been deleted from the notebook. Section 7b measured OWLv2/optical-flow at S=0.156 against a constant's 0.216 -- it lost. Kept as a documented negative result.\n\nWhen tracking is trusted, the impact point is the box-intersection centroid from Section 6.3 -- it answers \"where did these two vehicles touch\" directly. Otherwise the optical-flow weighted centroid is used -- a real motion-based estimate rather than the uninformative frame center.\n","metadata":{}},{"cell_type":"code","source":"# [MODEL] Stage 3a -- optical-flow weighted centroid (last-resort spatial signal)\n\ndef predict_impact_location_flow(video_path: pathlib.Path,\n                                 accident_time: float = None,\n                                 n_frames_context: int = 30) -> tuple:\n    \"\"\"Normalized (center_x, center_y) as the weighted centroid of the\n    cumulative optical-flow magnitude map centered on the accident time.\n    \"\"\"\n    RESIZE_W, RESIZE_H = 320, 180\n\n    center_frame = None\n    if accident_time is not None:\n        cap = cv2.VideoCapture(str(video_path))\n        fps = cap.get(cv2.CAP_PROP_FPS)\n        cap.release()\n        if fps > 0:\n            center_frame = int(accident_time * fps)\n\n    mag_map = compute_flow_magnitude_map(\n        video_path, resize_w=RESIZE_W, resize_h=RESIZE_H,\n        n_frames_context=n_frames_context, center_frame=center_frame,\n    )\n\n    total_mag = mag_map.sum()\n    if total_mag < 1e-6:\n        return 0.5, 0.5  # no motion detected anywhere\n\n    ys, xs = np.mgrid[0:RESIZE_H, 0:RESIZE_W]\n    cx = float((xs * mag_map).sum() / total_mag) / RESIZE_W\n    cy = float((ys * mag_map).sum() / total_mag) / RESIZE_H\n    return round(cx, 6), round(cy, 6)\n\n\nprint('[STATUS] predict_impact_location_flow defined')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:34:34.338157Z","iopub.execute_input":"2026-08-03T11:34:34.338659Z","iopub.status.idle":"2026-08-03T11:34:34.347608Z","shell.execute_reply.started":"2026-08-03T11:34:34.338632Z","shell.execute_reply":"2026-08-03T11:34:34.346793Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 6.7 The Constant Prior (and why it is the fallback)\n\nThe scoring sigma is comparable to the spread of the labels themselves -- `center_x` has std 0.13 and `center_y` std 0.18 against a spatial sigma of 0.1 -- so the training set's own prior is a strong predictor in its own right. **Any model whose error exceeds the label spread scores worse than a constant.**\n\nThat makes the constant the correct fallback for every stage. The previous design fell back from a weak model to an even weaker one: when tracking was distrusted, spatial localization went to OWLv2 grounding and then to an optical-flow centroid, both of which measure below simply predicting the centre of the frame. Every such fallback was subtracting score.\n\nConstants are fitted on the **full synthetic training set**, which the organizers provide for exactly this purpose (\"pretraining and debugging\"), and never on the calibration split -- that would be fitting on the eval set. They are found by grid search rather than by taking the mean: the Gaussian score saturates, so its optimum is mode-seeking and outliers pull the mean away from it.","metadata":{}},{"cell_type":"code","source":"# [MODEL] The constant prior -- fitted on the synthetic training set\n\n# The real metric, from the benchmark paper via the 1st-place writeup\n# (arXiv:2605.29325, Sec. 2.3). The Kaggle Evaluation page only says\n# \"Gaussian-style similarity\"; the specification is in the papers.\n#\n# An earlier revision of this notebook guessed SIGMA_T = 2.0 and an isotropic\n# SIGMA_S = 0.1. Both are wrong, and wrong in the flattering direction:\n#\n#   T is averaged over THREE tolerances, and 2.0 s is the most generous of the\n#   three. Any signal with multi-second error scores near zero at 0.5 s.\n#   S is ANISOTROPIC, with (sigma_x, sigma_y) set to the mean annotated bbox\n#   width and height -- so vertical error is tolerated ~40% more than\n#   horizontal. Independent check: the host's Molmo-7B baseline reports\n#   S=0.488, which is unreachable under an isotropic 0.1.\nSIGMA_T_LIST = (0.5, 1.0, 2.0)\nSIGMA_X = float((labels_clean['x2'] - labels_clean['x1']).mean())\nSIGMA_Y = float((labels_clean['y2'] - labels_clean['y1']).mean())\nprint(f'[METRIC] sigma_t = {SIGMA_T_LIST}  |  sigma_x = {SIGMA_X:.4f}  sigma_y = {SIGMA_Y:.4f}')\n\n# Kept only so the constant grid-search below has a scalar to optimise against;\n# the scoring functions in Section 7 use the real definition above.\nSIGMA_T = 2.0\nSIGMA_S = 0.1\n\n\ndef fit_constant_predictor(train_df: pd.DataFrame) -> dict:\n    \"\"\"Constants maximising each Gaussian component on train_df.\"\"\"\n    t = train_df['accident_time'].to_numpy()\n    x = train_df['center_x'].to_numpy()\n    y = train_df['center_y'].to_numpy()\n\n    t_grid = np.arange(t.min(), t.max() + 1e-9, 0.05)\n    t_best = float(t_grid[np.argmax([np.exp(-0.5 * ((c - t) / SIGMA_T) ** 2).mean()\n                                     for c in t_grid])])\n\n    xy_grid = np.arange(0.30, 0.71, 0.01)\n    best_s, xy_best = -1.0, (0.5, 0.5)\n    for cx in xy_grid:\n        d2x = (cx - x) ** 2\n        for cy in xy_grid:\n            s = float(np.exp(-0.5 * (d2x + (cy - y) ** 2) / SIGMA_S ** 2).mean())\n            if s > best_s:\n                best_s, xy_best = s, (float(cx), float(cy))\n\n    return {'accident_time': t_best,\n            'center_x': xy_best[0],\n            'center_y': xy_best[1],\n            'type': train_df['type'].value_counts().idxmax()}\n\n\nCONST = fit_constant_predictor(labels_clean)\nprint(f'[STATUS] Constants fitted on {len(labels_clean)} synthetic videos: '\n      f\"t={CONST['accident_time']:.2f}s  \"\n      f\"xy=({CONST['center_x']:.2f}, {CONST['center_y']:.2f})  type={CONST['type']}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:34:34.348575Z","iopub.execute_input":"2026-08-03T11:34:34.348772Z","iopub.status.idle":"2026-08-03T11:34:34.419177Z","shell.execute_reply.started":"2026-08-03T11:34:34.348750Z","shell.execute_reply":"2026-08-03T11:34:34.418483Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 6.8 Sim-to-Real Sanity Check (no labels required)\n\nEverything downstream — the 2211-video feature extraction, the classifier, the whole object-centric design — rests on one unverified assumption: **that tracking fires on real CCTV at all**. Nothing in this notebook has ever tested that. Calibration runs on CARLA, and CARLA is not the target domain.\n\nThis is checkable without a single label. The test clips are real CCTV; run the tracker on them and count. If a collision pair is found in 45% of CARLA clips and 5% of real ones, then on the real test set the pipeline is a constant with extra steps, and no amount of classifier tuning changes that. Finding this out costs ~30 minutes here, against several hours of feature extraction and a full inference run downstream.\n\nThe same probe runs on an equal sample of synthetic clips, so the comparison is like-for-like: same code, same N, two domains.\n\n**What the numbers mean:**\n\n- **Fire rate.** Real ≈ synthetic means the design transfers. Real ≪ synthetic means it does not.\n- **`n_tracks`.** Separates two very different diagnoses: the detector sees no vehicles at all (YOLO fails on this footage — different resolution, camera angle, compression) versus it sees vehicles but never finds a collision among them (the contact logic is what fails).\n- **Predicted distributions.** Compared against the *synthetic label* prior. `CONST` is fitted on CARLA and fires on most videos, so if real predictions cluster somewhere else entirely, the fallback is aimed at the wrong place — and the fallback is most of the score.\n\nThis measures whether the machinery *runs* on real footage, not whether it is *right*: that needs labels, which is the point of the competition. A healthy fire rate is necessary, not sufficient.","metadata":{}},{"cell_type":"code","source":"# [DIAG] Does any of this work on real CCTV? -- no labels needed\nSANITY_N = 100   # per domain; ~10 s/video => ~35 min total\n\ndef _probe(paths, domain):\n    out = []\n    for k, vp in enumerate(paths):\n        c = analyze_video_collision(pathlib.Path(vp))\n        out.append({'domain': domain, 'path': str(vp),\n                    'found': bool(c['found']), 'mode': c.get('mode'),\n                    'n_tracks': int(c.get('n_tracks', 0)),\n                    'evidence': float(c.get('evidence', 0.0)),\n                    't': c['t_impact'] if c['found'] else np.nan,\n                    'x': c['center'][0] if c['found'] else np.nan,\n                    'y': c['center'][1] if c['found'] else np.nan})\n        if (k + 1) % 20 == 0:\n            print(f'  {domain}: {k+1}/{len(paths)}')\n        if torch.cuda.is_available():\n            torch.cuda.empty_cache()\n    return out\n\n_syn_sample = labels_clean.sample(min(SANITY_N, len(labels_clean)),\n                                  random_state=SEED)['abs_video_path'].tolist()\n_real_sample = real_videos[:SANITY_N]\n\nprint(f'[STATUS] Probing {len(_syn_sample)} synthetic + {len(_real_sample)} real videos')\nsanity = pd.DataFrame(_probe(_syn_sample, 'synthetic') + _probe(_real_sample, 'real'))\n\nsummary = sanity.groupby('domain').agg(\n    fire_rate=('found', 'mean'),\n    n_tracks_mean=('n_tracks', 'mean'),\n    n_tracks_zero=('n_tracks', lambda s: float((s == 0).mean())),\n    evidence_median=('evidence', lambda s: float(s[s > 0].median()) if (s > 0).any() else np.nan),\n    n=('found', 'size'),\n).round(4)\nprint('\\n[SANITY] tracking behaviour by domain')\ndisplay(summary)\n\nsyn_rate  = float(summary.loc['synthetic', 'fire_rate'])\nreal_rate = float(summary.loc['real', 'fire_rate'])\nratio = real_rate / (syn_rate + 1e-9)\nprint(f'[VERDICT] real fire rate is {ratio:.0%} of synthetic '\n      f'({real_rate:.0%} vs {syn_rate:.0%})')\nif ratio < 0.5:\n    print('[VERDICT] TRACKING DOES NOT TRANSFER. On the real test set this pipeline is')\n    print('          mostly the constant. Do not spend hours on feature extraction or a')\n    print('          classifier until the detector/tracker works on this footage --')\n    print('          check n_tracks_zero above to see whether YOLO sees vehicles at all.')\nelse:\n    print('[VERDICT] Tracking fires on real footage at a comparable rate. The')\n    print('          object-centric design is worth the downstream investment.')\n\n# --- distributions: real predictions vs the synthetic label prior ---\nfig, axes = plt.subplots(1, 3, figsize=(16, 4.5))\n\nfor dom, colr in [('synthetic', PALETTE['primary']), ('real', PALETTE['secondary'])]:\n    ev = sanity[(sanity.domain == dom) & (sanity.evidence > 0)]['evidence']\n    if len(ev):\n        sns.kdeplot(np.log10(ev), ax=axes[0], label=dom, color=colr, fill=True, alpha=.3)\naxes[0].set_title('Evidence where tracking fired (log10)')\naxes[0].set_xlabel('log10(evidence)'); axes[0].legend()\n\nfor dom, colr in [('synthetic', PALETTE['primary']), ('real', PALETTE['secondary'])]:\n    tt = sanity[(sanity.domain == dom)]['t'].dropna()\n    if len(tt):\n        sns.kdeplot(tt, ax=axes[1], label=f'{dom} (predicted)', color=colr, fill=True, alpha=.3)\nsns.kdeplot(labels_clean['accident_time'], ax=axes[1], label='synthetic GT',\n            color='black', linestyle='--')\naxes[1].axvline(CONST['accident_time'], color='crimson', linestyle=':',\n                label=f\"CONST={CONST['accident_time']:.1f}s\")\naxes[1].set_title('Predicted accident time'); axes[1].set_xlabel('seconds'); axes[1].legend(fontsize=8)\n\nfor dom, mk, colr in [('synthetic', 'o', PALETTE['primary']), ('real', '^', PALETTE['secondary'])]:\n    d = sanity[sanity.domain == dom].dropna(subset=['x', 'y'])\n    axes[2].scatter(d['x'], d['y'], s=22, alpha=.55, marker=mk, label=dom, color=colr)\naxes[2].scatter([CONST['center_x']], [CONST['center_y']], s=180, marker='*',\n                color='crimson', label='CONST', zorder=5)\naxes[2].set_xlim(0, 1); axes[2].set_ylim(0, 1); axes[2].invert_yaxis()\naxes[2].set_title('Predicted impact location'); axes[2].legend(fontsize=8)\n\nplt.suptitle('Sim-to-Real Sanity Check', fontsize=13, y=1.02)\nplt.tight_layout(); plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:34:34.420216Z","iopub.execute_input":"2026-08-03T11:34:34.420659Z","iopub.status.idle":"2026-08-03T11:46:15.765329Z","shell.execute_reply.started":"2026-08-03T11:34:34.420629Z","shell.execute_reply":"2026-08-03T11:46:15.764247Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 6.9 Supervised Collision-Type Classifier — NEGATIVE RESULT\n\n> **Measured, and it fails.** Grouped-by-map accuracy is **0.1464** — below the 0.20 chance level, and far below the 0.3766 of a constant `rear-end`. The out-of-fold confusion matrix shows the model systematically **inverting head-on and rear-end** on held-out Towns (395 rear-end videos predicted head-on; 302 head-on predicted rear-end): it is fitting per-Town spurious correlations that reverse sign on an unseen camera.\n>\n> **Why.** Tracking finds *a* contact in 76% of clips, not *the* collision. There are 9.4 tracked vehicles per synthetic clip and 17.0 per real one — 36 to 136 candidate pairs — and the evidence heuristic picks the wrong one. The proof is the `single` class: all four single-vehicle calibration clips report a *pair* contact, which is impossible by definition, and they carry the four lowest evidence scores in the set.\n>\n> **This was a known result before it was run.** The 1st-place writeup lists under negative results: *\"A CNN classifier trained on CARLA, an optical-flow + detector hybrid, and a frame-offset ensemble all improved sim validation but decreased the LB score.\"* The one thing that went right here is `GroupKFold` on `map`: a random split would have put near-duplicate renders on both sides, reported a healthy number, and shipped a model that inverts two of its five classes.\n\nThe 2211 labelled synthetic videos have so far been used only for EDA and to draw a 20-video calibration split. That is the single largest thing left on the table: the organizers provide them explicitly for *\"pretraining and debugging\"*, so training on them is squarely inside the rules. **Zero-shot in this competition means no real labels, not no training** — the withheld split is the real one, and CARLA is not it.\n\nThe zero-shot classifier being replaced scores C=0.30 against a 0.20 chance level and a 0.359 constant. A gradient-boosted tree over 2211 labelled examples has a far better claim on this problem than five words of CLIP prompt.\n\n**Why kinematic features and not CLIP embeddings.** Both are trainable on the same 2211 videos, but they fail differently under domain shift. A probe on CLIP image features can reach high accuracy on CARLA by reading render style — lighting model, texture statistics, the particular look of the simulator — and none of that exists in real CCTV. The features here are angles, ratios, and speeds normalized by the frame diagonal: \"the two velocity vectors met at 93° and the slower vehicle was doing 0.2 diagonals/second\" is physics, and physics is the same in CARLA and on a real street. This is the one design choice that has to survive the sim-to-real gap, so it gets the conservative option.\n\n**Validation must be grouped by `map`.** The dataset has ~240 camera positions across six Towns. A random split puts near-duplicate renders — same town, same camera, same junction, different weather — on both sides, and reports an accuracy that means nothing. `GroupKFold` on `map` holds out entire Towns, which is the closest proxy available for \"a camera you have never seen\". It will report a *lower* number than a random split. That lower number is the honest one.\n\n**This does not measure the render-to-real gap.** Nothing computed inside CARLA can. Grouping by Town measures generalization across scene and camera only; the true gap to CCTV is larger and unmeasurable from here.","metadata":{}},{"cell_type":"code","source":"# [FEAT] Extract track kinematics for every synthetic video -- checkpointed\n#\n# This is the expensive cell: ~10 s/video on a T4. Results are cached to disk so\n# it is paid once. Set FEATURE_SUBSET_N to prototype on fewer videos first.\n#\n# ORDER OF OPERATIONS: run the Section 7b ablation before this. If tracking does\n# not fire on real footage, these features do not exist at inference time on the\n# test set and the whole classifier is dead weight -- 4 minutes of measurement\n# de-risks several hours of extraction.\n\nFEATURE_SUBSET_N = None          # None = all 2211; set e.g. 600 to prototype\nFEATURES_PATH    = OUTPUT_DIR / 'track_features.csv'\nFEATURES_EVERY   = 100\n\nfeat_rows, feat_done = [], set()\nif FEATURES_PATH.exists():\n    _prev = pd.read_csv(FEATURES_PATH)\n    feat_rows = _prev.to_dict('records')\n    feat_done = set(_prev['rgb_path'])\n    print(f'[STATUS] Checkpoint: {len(feat_done)} videos already extracted -- resuming')\n\n_todo = labels_clean if FEATURE_SUBSET_N is None else labels_clean.head(FEATURE_SUBSET_N)\n_todo = _todo[~_todo['rgb_path'].isin(feat_done)]\nprint(f'[STATUS] Extracting features for {len(_todo)} video(s)')\n\n_t0 = time.time()\nfor _i, (_, row) in enumerate(_todo.iterrows()):\n    vp = pathlib.Path(row['abs_video_path'])\n    coll = analyze_video_collision(vp)\n    rec = {'rgb_path': row['rgb_path'], 'type': row['type'], 'map': row['map'],\n           'weather': row['weather'], 'camera_position': row['camera_position'],\n           'found': bool(coll['found']), 'mode': coll.get('mode'),\n           # Tracking's own T/S predictions are recorded too: this pass visits\n           # every video, so it measures what tracking is worth on n=2211 rather\n           # than the n=20 of Section 7b, where one video moves C by 0.05.\n           't_track': coll['t_impact'] if coll['found'] else np.nan,\n           'x_track': coll['center'][0] if coll['found'] else np.nan,\n           'y_track': coll['center'][1] if coll['found'] else np.nan}\n    rec.update(coll.get('features') or {c: np.nan for c in FEATURE_COLS})\n    feat_rows.append(rec)\n\n    if (_i + 1) % FEATURES_EVERY == 0 or (_i + 1) == len(_todo):\n        pd.DataFrame(feat_rows).to_csv(FEATURES_PATH, index=False)\n        _el = time.time() - _t0\n        _eta = _el / (_i + 1) * (len(_todo) - _i - 1) / 3600\n        print(f'[STATUS] {_i+1}/{len(_todo)} | {_el/(_i+1):.1f}s/video | ETA {_eta:.2f} h')\n    if torch.cuda.is_available():\n        torch.cuda.empty_cache()\n\nfeatures_df = pd.DataFrame(feat_rows)\nprint(f'[STATUS] {len(features_df)} rows | tracking found a collision in '\n      f\"{features_df['found'].mean():.0%}\")\nprint('[STATUS] Coverage by type (a class tracking never sees cannot be learned):')\ndisplay(features_df.groupby('type')['found'].agg(['mean', 'count']).round(3))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T11:46:15.766561Z","iopub.execute_input":"2026-08-03T11:46:15.766885Z","iopub.status.idle":"2026-08-03T13:44:00.666828Z","shell.execute_reply.started":"2026-08-03T11:46:15.766847Z","shell.execute_reply":"2026-08-03T13:44:00.665926Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [CALIB] What tracking is actually worth -- measured on every extracted video\nm = features_df[features_df['found']].merge(\n    labels_clean[['rgb_path', 'accident_time', 'center_x', 'center_y']],\n    on='rgb_path', how='inner')\nprint(f'[STATUS] Measuring on {len(m)} videos where tracking fired')\n\n# A tracked contact is the instant the BOXES touch. The dataset's accident_time\n# may be defined differently -- on a synthetic t-bone the boxes overlap ~0.5 s\n# before the centres coincide, and any such systematic offset is free score on T.\n# Measured on the training set and subtracted; this is calibration on provided\n# labels, not on the test set.\nresid = (m['t_track'] - m['accident_time']).to_numpy()\nTIME_OFFSET = float(np.median(resid))\nprint(f'[CALIB] t_track - accident_time: median={TIME_OFFSET:+.3f}s  '\n      f'IQR=[{np.quantile(resid, .25):+.3f}, {np.quantile(resid, .75):+.3f}]')\n\ndef _Tm(p, g): return float(np.exp(-0.5 * ((p - g) / SIGMA_T) ** 2).mean())\ndef _Sm(px, py, gx, gy):\n    return float(np.exp(-0.5 * ((px - gx) ** 2 + (py - gy) ** 2) / SIGMA_S ** 2).mean())\n\ngt_t, gt_x, gt_y = m['accident_time'].to_numpy(), m['center_x'].to_numpy(), m['center_y'].to_numpy()\ncov = float(features_df['found'].mean())\nn_all = len(features_df)\n\nrows = [\n    ('T  constant',            _Tm(np.full(len(m), CONST['accident_time']), gt_t), 1.0),\n    ('T  tracking (raw)',      _Tm(m['t_track'].to_numpy(), gt_t), cov),\n    ('T  tracking (offset)',   _Tm(m['t_track'].to_numpy() - TIME_OFFSET, gt_t), cov),\n    ('S  constant',            _Sm(np.full(len(m), CONST['center_x']),\n                                   np.full(len(m), CONST['center_y']), gt_x, gt_y), 1.0),\n    ('S  tracking',            _Sm(m['x_track'].to_numpy(), m['y_track'].to_numpy(), gt_x, gt_y), cov),\n]\n# Constant scores on the same subset, for the coverage-weighted blend below\nT_con_all = _Tm(np.full(len(m), CONST['accident_time']), gt_t)\nS_con_all = _Sm(np.full(len(m), CONST['center_x']), np.full(len(m), CONST['center_y']), gt_x, gt_y)\n\nout = []\nfor name, val, c in rows:\n    base = T_con_all if name.startswith('T') else S_con_all\n    blended = val * c + base * (1 - c) if c < 1.0 else val\n    out.append({'signal': name, 'coverage': round(c, 3),\n                'score_where_it_fires': round(val, 4),\n                'blended_with_constant': round(blended, 4)})\nprint()\ndisplay(pd.DataFrame(out))\nprint(f'[VERDICT] tracking beats the constant on T: '\n      f\"{out[2]['blended_with_constant'] > out[0]['score_where_it_fires']}\")\nprint(f'[VERDICT] tracking beats the constant on S: '\n      f\"{out[4]['blended_with_constant'] > out[3]['score_where_it_fires']}\")\nprint('[NOTE] Measured inside CARLA. It says nothing about whether tracking '\n      'fires at all on real CCTV -- only running it on the real videos does.')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T13:44:00.668086Z","iopub.execute_input":"2026-08-03T13:44:00.668575Z","iopub.status.idle":"2026-08-03T13:44:00.694375Z","shell.execute_reply.started":"2026-08-03T13:44:00.668546Z","shell.execute_reply":"2026-08-03T13:44:00.693562Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [MODEL] Train the collision-type classifier -- GroupKFold by map\nfrom sklearn.ensemble import HistGradientBoostingClassifier\nfrom sklearn.model_selection import GroupKFold\nfrom sklearn.metrics import accuracy_score, confusion_matrix \n\n# Only videos where tracking fired have features. Videos where it did not are\n# not trainable here and fall back to the constant at inference time.\ntrain_df = features_df[features_df['found']].copy()\nprint(f'[STATUS] Trainable: {len(train_df)}/{len(features_df)} videos '\n      f'({len(train_df)/len(features_df):.0%})')\n\nX = train_df[FEATURE_COLS].to_numpy(dtype=float)   # NaN kept: HGB handles it natively\ny = train_df['type'].to_numpy()\ngroups = train_df['map'].to_numpy()\n\nn_splits = min(5, train_df['map'].nunique())\ngkf = GroupKFold(n_splits=n_splits)\noof = np.empty(len(train_df), dtype=object)\n\nfor fold, (tr, va) in enumerate(gkf.split(X, y, groups)):\n    clf = HistGradientBoostingClassifier(\n        max_iter=300, learning_rate=0.06, max_leaf_nodes=15,\n        l2_regularization=1.0, early_stopping=True, validation_fraction=0.15,\n        random_state=SEED)\n    clf.fit(X[tr], y[tr])\n    oof[va] = clf.predict(X[va])\n    held = sorted(set(groups[va]))\n    print(f'  fold {fold}: held-out {held} | n={len(va):4d} | '\n          f'acc={accuracy_score(y[va], oof[va]):.4f}')\n\ncv_acc = accuracy_score(y, oof)\nprint(f'\\n[CV] Grouped-by-map accuracy = {cv_acc:.4f}')\nprint(f'[REF] chance = 0.20 | constant \"{CONST[\"type\"]}\" on this subset = '\n      f'{(y == CONST[\"type\"]).mean():.4f} | CLIP zero-shot measured 0.30')\nprint('[NOTE] This is accuracy on videos where tracking fired. Coverage-weighted '\n      f'C over all videos = {cv_acc * features_df[\"found\"].mean():.4f} with a '\n      'constant elsewhere -- compare THAT against the constant, not the number above.')\n\nprint('\\n[CONFUSION] out-of-fold, rows = truth')\ndisplay(pd.crosstab(pd.Series(y, name='truth'), pd.Series(oof, name='pred')))\n\n# Fit the deployed model on everything once the CV estimate is in hand.\nTYPE_CLASSIFIER = HistGradientBoostingClassifier(\n    max_iter=300, learning_rate=0.06, max_leaf_nodes=15, l2_regularization=1.0,\n    early_stopping=True, validation_fraction=0.15, random_state=SEED).fit(X, y)\nTYPE_CLASSIFIER_CV_ACC = float(cv_acc)\nprint(f'\\n[STATUS] TYPE_CLASSIFIER fitted on all {len(X)} trackable videos')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T13:44:00.695553Z","iopub.execute_input":"2026-08-03T13:44:00.696097Z","iopub.status.idle":"2026-08-03T13:44:02.932774Z","shell.execute_reply.started":"2026-08-03T13:44:00.696070Z","shell.execute_reply":"2026-08-03T13:44:02.931729Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 6.10 Qwen3-VL Three-Stage Pipeline\n\nSections 6.2–6.9 built an object-centric pipeline on the thesis that *an accident is an event between two objects, so tracking answers all three questions more directly than any whole-frame signal*. **The measurements in Sections 6.8, 6.9, and 7b refute that thesis on this data.** They are kept as a documented negative result; this section replaces them.\n\nThe 1st-place solution (arXiv:2605.29325) reaches 0.57080 with **no training at all**: a frozen Qwen3-VL checkpoint called three times per clip. Their Table 2 puts a *single* VLM call at **0.42238** — nearly double what five models and a tracker achieve here. Their reported ablations also contradict this notebook's design in three specific places:\n\n- **\"Synthetic-data transfer. A CNN classifier trained on CARLA, an optical-flow + detector hybrid, and a frame-offset ensemble all improved sim validation but decreased the LB score.\"** Section 6.9 is exactly this, and reproduced it: grouped-by-map accuracy **0.1464**, below the 0.20 chance level.\n- **\"Overlaying object-detection tracking trails on input frames consistently degraded performance.\"** Tracking is not the signal.\n- Detection survives only as a **bbox snap** post-process, worth +0.0005/+0.0013.\n\nThe three stages, and why the split matters:\n\n1. **Stage 1 — full-video scan.** The whole clip at 4 fps with a timestamp burned into every frame, returning a joint JSON of time, location, and type. The burned-in label lets the model *copy* a timestamp instead of estimating one from visual context.\n2. **Stage 2 — time refinement.** A dense window around the Stage-1 time, blended back as a *bounded* correction: `t = t_base + α·clip(t_refined − t_base, ±δ)` with α=0.35, δ=1.5 s. Direct replacement scored worse. **Section 6.4 of this notebook independently derived the same formula** for the Perception Encoder refinement, and measured the same failure for the unbounded version (T fell to 0.16). That instinct was right.\n3. **Stage 3 — spatial grounding.** A single frame at the refined time, no timestamp overlay, asking for a point on Qwen3-VL's native [0,1000] scale. This is the **largest single gain in their entire development path: +0.09356**. Under a joint query the model must split attention across time and space, and the coordinates collapse onto a coarse grid.\n\n**`scene_layout` is a provided input, not just metadata.** `test_df` carries it for every test clip (`highway`, `signalized_intersection`, `tunnel`, …) — the exact vocabulary the winner's scene-hint dictionary keys on. This notebook loaded it and labelled it *\"provided for analysis, not scoring\"*. It feeds two things: a hint in the Stage-1 prompt (+0.00216) and the t-bone rule (+0.00466).\n\n**Hardware.** Their 32B-FP8 leg needs ~13 h on an RTX PRO 6000 and does not fit a T4. `VLM_CHECKPOINT` below defaults to **Qwen3-VL-8B**, which their backbone sweep puts at 0.365 single-call at 512 px against 0.387 for the 32B — most of the quality at a quarter of the size. **Runtime on a T4 is the main open risk and is not yet measured: use the `SUBSET_N` dry run before committing to a full pass.**","metadata":{}},{"cell_type":"code","source":"# [MODEL] Qwen3-VL -- the backbone that replaces the tracking pipeline\n#\n# The reference solution serves Qwen3-VL-32B-Instruct-FP8 through vLLM. Two\n# reasons this cell does not: FP8 needs SM89+ (a T4 is SM75), and 32B does not\n# fit 16 GB. transformers + 4-bit is what this environment is known to run --\n# the existing Qwen2.5-VL cell already does it. vLLM would be substantially\n# faster and is worth switching to if 2027 clips prove too slow.\nfrom transformers import AutoModelForImageTextToText, AutoProcessor, BitsAndBytesConfig\n\nVLM_CHECKPOINT = 'Qwen/Qwen3-VL-8B-Instruct'\n\n# Reference config (32B): 4 fps, <=128 frames, 960 px. Their Table 2 measures a\n# single call at 768 px / 2 fps = 0.42238, and going 768->960 px and 2->4 fps was\n# worth +0.01969. The conservative setting is the default here because a T4 has\n# to hold the KV cache for every frame; raise them if the dry run has headroom.\n# 64 frames at 768 px OOM'd a T4 on the first clip. 512 px is not a guess: it is\n# the resolution of the reference backbone sweep, where this same 8B checkpoint\n# scored 0.365 single-call. Frame count is the other half of the activation\n# budget; raise both only if the memory report above shows headroom.\n# Frame budget. The earlier 1 fps / 24 f / 512 px was set blind, to survive an\n# OOM whose real cause was 4.9 GB of dead legacy models plus device_map pinning\n# everything to cuda:0. With those fixed the loader reports ~10 GB free on the\n# smaller of the two T4s, so the budget can go back up -- and it matters: the\n# reference measures 768 px / 2 fps at 0.42238, and their 512 px backbone sweep\n# puts the 8B at 0.365. Running below their baseline configuration gives up\n# score before the pipeline even starts.\n#\n# Raise 'whole_limit' toward 128 and 'image_max_side' toward 960 if the dry run\n# shows headroom; drop them first if anything OOMs.\n# A Qwen-VL \"visual token\" covers a 28x28 pixel block, so a 768 px frame of a\n# 16:9 clip costs 423 tokens and 48 of them cost 20,304 -- which is what the OOM\n# on every calibration clip was. The retry then halved the frame count twice,\n# so Stage 1 answered from 12 frames spread over ~20 s: roughly one frame every\n# 1.7 s, which can miss the collision entirely. That is the likely cause of both\n# the wrong times and the collapsed types.\n#\n# The two stages want different things, so they no longer share a budget:\n#   Stage 1 answers \"when\" and \"what\" -> needs TEMPORAL COVERAGE, so many frames\n#           at low resolution (32 x 448 px = 4,608 tokens).\n#   Stage 3 answers \"where\" in one frame -> needs SPATIAL DETAIL, so a single\n#           frame at high resolution (768 px = 423 tokens, cheap at n=1).\n# The reference reports the same asymmetry from the other side: going 4->10 fps\n# cost them -0.009, and rendering the grounding frame at 1024/1280 px degraded\n# spatial accuracy. More pixels is not monotonically better.\nVLM_CFG = {\n    'whole_fps': 2.0,\n    'whole_limit': 32,\n    'image_max_side': 448,\n    'grounding_max_side': 768,\n    'split_threshold_sec': 32.0,\n    'pass_duration_sec': 32.0,\n    'time_refine': True,\n    'time_refine_window': 2.0,      # dense window half-width, seconds\n    'time_refine_fps': 4.0,\n    'time_refine_limit': 10,\n    'time_refine_context_before': 8.0,\n    'time_refine_context_after': 4.0,\n    'time_refine_context_fps': 0.5,\n    'time_refine_context_limit': 4,\n    'time_refine_blend': 0.35,      # alpha -- correction is down-weighted\n    'time_refine_max_shift_sec': 1.5,   # delta_max -- and hard-capped\n    'use_grounding': True,\n}\n\nVLM_AVAILABLE = True\ntry:\n    # A previous failed/partial load can leave tensors alive in this namespace,\n    # which torch.cuda.empty_cache() alone will not reclaim -- it only returns\n    # cache the allocator itself is not still holding a live reference to. Drop\n    # any stale handle before loading again, in case this cell is re-run without\n    # a kernel restart.\n    for _name in ('vlm_model', 'vlm_processor'):\n        if _name in globals():\n            del globals()[_name]\n    import gc\n    gc.collect()\n    if torch.cuda.is_available():\n        torch.cuda.empty_cache()\n\n    _bnb = (BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_compute_dtype=torch.float16)\n            if DEVICE == 'cuda' else None)\n    vlm_processor = AutoProcessor.from_pretrained(VLM_CHECKPOINT)\n\n    # device_map='auto' shards the model across BOTH T4s by pre-load free memory,\n    # but it balances STATIC WEIGHTS only -- it has no way to know that decode-time\n    # KV-cache growth (which scales with visual-token count, i.e. whole_limit x\n    # image_max_side) will land on whichever GPU ends up owning the later decoder\n    # layers. On this run it concentrated most of the model's ~6.8 GB *and* the\n    # growing KV-cache onto GPU 1, which then had nothing left for a 938 MiB\n    # allocation. The fix is not \"prefer GPU 1\" (that would just hardcode today's\n    # imbalance and break on a single-GPU box, e.g. a laptop) -- it is to pin the\n    # WHOLE model onto whichever single GPU has the most free memory right now,\n    # since 6.8 GB comfortably fits alone on one 14-16 GB T4 with room to spare\n    # for activations. No sharding, no cross-GPU imbalance to reason about.\n    if torch.cuda.is_available():\n        _n_gpu = torch.cuda.device_count()\n        _free_by_gpu = [torch.cuda.mem_get_info(g)[0] for g in range(_n_gpu)]\n        _best_gpu = int(max(range(_n_gpu), key=lambda g: _free_by_gpu[g]))\n        _vlm_device_map = {'': f'cuda:{_best_gpu}'}\n        print(f'[STATUS] GPU free memory: '\n              f'{[f\"{f/1e9:.2f} GB\" for f in _free_by_gpu]} -> pinning VLM to cuda:{_best_gpu}')\n    else:\n        _vlm_device_map = {'': 'cpu'}\n\n    vlm_model = AutoModelForImageTextToText.from_pretrained(\n        VLM_CHECKPOINT, quantization_config=_bnb, device_map=_vlm_device_map,\n        dtype=torch.float16, trust_remote_code=True).eval()\n    print(f'[SUCCESS] {VLM_CHECKPOINT} loaded')\n\n    # Confirm bitsandbytes actually replaced the Linear layers. If the config is\n    # silently ignored the model loads in fp16, takes ~16 GB, and the OOM appears\n    # later and further away, looking like a frame-budget problem instead.\n    _n4bit = sum(1 for m in vlm_model.modules() if 'Linear4bit' in type(m).__name__)\n    print(f'[STATUS] Linear4bit layers: {_n4bit}  (0 means quantization did NOT apply)')\nexcept Exception as e:\n    VLM_AVAILABLE = False\n    print(f'[ERROR] Qwen3-VL unavailable: {type(e).__name__}: {e}')\n\nif torch.cuda.is_available():\n    for _g in range(torch.cuda.device_count()):\n        _free, _tot = torch.cuda.mem_get_info(_g)\n        print(f'[STATUS] GPU{_g} {torch.cuda.get_device_name(_g)}: '\n              f'{(_tot-_free)/1e9:.2f} GB used / {_tot/1e9:.2f} GB total, '\n              f'{_free/1e9:.2f} GB free for activations')\n\nif not VLM_AVAILABLE:\n    raise RuntimeError(\n        'Qwen3-VL is not usable, so Section 6.10 cannot run. Do NOT continue to the '\n        'submission cell: run_inference_vlm would return normalize_prediction defaults '\n        'for every clip -- a constant dressed up as a prediction. Fix the environment '\n        '(see the install cell: transformers must be >=4.57 AND the kernel restarted '\n        'afterwards) before going on.')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T13:44:02.933984Z","iopub.execute_input":"2026-08-03T13:44:02.934635Z","iopub.status.idle":"2026-08-03T13:46:37.420082Z","shell.execute_reply.started":"2026-08-03T13:44:02.934593Z","shell.execute_reply":"2026-08-03T13:46:37.419434Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [MODEL] Prompts and frame sampling for the three stages\n\n# scene_layout is a provided column of test_df. Keys match its vocabulary.\nSCENE_TYPE_HINTS = {\n    'highway': 'Common accident types here: rear-end (following too closely), single (loss of control, hitting barrier/guardrail), or sideswipe (lane change).',\n    'signalized_intersection': 'Common accident types here: t-bone (running red light, perpendicular collision), rear-end (stopping suddenly), or single (running off road).',\n    'simple_intersection': 'Common accident types here: t-bone (failure to yield at unsignalized intersection), rear-end, or single.',\n    'grade_separated_intersection': 'Common accident types here: single (loss of control on ramp/merge), rear-end (merging traffic), or t-bone.',\n    'city_street': 'Common accident types here: t-bone (intersection), rear-end, single (hitting parked car/pole), or sideswipe.',\n    'tunnel': 'Common accident types here: rear-end (reduced visibility), single (hitting wall), or sideswipe.',\n    'parking_lot': 'Common accident types here: single (hitting pole/wall), rear-end (backing up), or sideswipe.',\n    'roundabout': 'Common accident types here: t-bone (failure to yield), sideswipe (lane change in roundabout), or single.',\n}\n\n# Two vehicles cannot meet perpendicularly in these layouts.\nT_BONE_IMPOSSIBLE_LAYOUTS = {'highway', 'tunnel', 'grade_separated_intersection'}\n\n# Scene metadata keyed by submission path, for the hint and the rule.\nSCENE_BY_PATH = (test_df.set_index('path')['scene_layout'].to_dict()\n                 if not test_df.empty and 'scene_layout' in test_df.columns else {})\n\n\ndef build_stage1_prompt(duration: float, scene_hint: str) -> str:\n    return (\n        'These sequential CCTV frames show a traffic accident. '\n        'Each frame has a burned-in timestamp (t=xx.xx s). '\n        f'The full video is {duration:.1f} seconds long.\\n\\n'\n        'Your task: Find the EXACT MOMENT, LOCATION, and TYPE of the collision.\\n\\n'\n        \"1) COLLISION TIME -- Carefully examine each frame's timestamp. \"\n        'Find the frame where vehicles first make contact or a vehicle first hits an object. '\n        'Report the timestamp in seconds.\\n\\n'\n        '2) COLLISION POSITION -- Where in the frame does the impact happen? '\n        'Report as normalized coordinates: center_x (0=left, 1=right), center_y (0=top, 1=bottom).\\n\\n'\n        \"3) ACCIDENT TYPE -- Classify the collision type by watching the vehicles' \"\n        'approach angles and movements across all frames:\\n'\n        '  - rear-end = a following vehicle strikes the one ahead (both traveling same direction)\\n'\n        '  - head-on = two vehicles collide front-to-front (traveling opposite directions)\\n'\n        '  - sideswipe = vehicles traveling parallel make side-to-side contact\\n'\n        \"  - t-bone = perpendicular collision (one vehicle hits the other's side at ~90 degrees)\\n\"\n        '  - single = only one vehicle involved (hits object, barrier, pole, or loses control)\\n'\n        f'{scene_hint}\\n'\n        'Look at:\\n'\n        '- How many vehicles are involved (one = single)\\n'\n        '- The direction each vehicle is traveling before impact\\n'\n        '- The angle of collision (same direction = rear-end, opposite = head-on, '\n        'perpendicular = t-bone, parallel side = sideswipe)\\n\\n'\n        'Output ONLY this JSON:\\n'\n        '{\"accident_time\": <seconds>, \"center_x\": <0-1>, \"center_y\": <0-1>, '\n        '\"type\": \"rear-end|head-on|t-bone|sideswipe|single\"}'\n    )\n\n\nTIME_REFINE_PROMPT = (\n    'These sequential CCTV frames show a traffic accident. '\n    'Some frames provide broader context before/after the event, and other frames densely '\n    'cover a short window around the current predicted collision time. '\n    'Each frame has a burned-in timestamp (t=xx.xx s).\\n\\n'\n    'Your task: identify the EXACT timestamp when vehicles first make physical contact, '\n    'or when a vehicle first hits an object. '\n    'Do NOT choose the peak impact, aftermath, debris, spin, or when vehicles are already '\n    'clearly entangled. Use the broader context to understand motion, but report ONLY the '\n    'first contact time.\\n\\n'\n    'Output ONLY this JSON:\\n'\n    '{\"accident_time\": <seconds>}'\n)\n\nGROUNDING_PROMPT = (\n    'This CCTV frame shows a traffic accident. '\n    'Point to the EXACT location where the collision or impact occurs. '\n    'Output the collision point coordinates as JSON:\\n'\n    '{\"point\": [x, y]}\\n'\n    'where x is horizontal (0=left edge, 1000=right edge) and '\n    'y is vertical (0=top edge, 1000=bottom edge).'\n)\n\n\ndef sample_frames_stamped(video_path, target_fps, t_start=None, t_end=None,\n                          limit=64, max_side=768, burn=True):\n    \"\"\"Frames at target_fps in [t_start, t_end], each with its timestamp drawn on.\n\n    The burned-in label is the point: it lets the model read a timestamp off the\n    pixels and copy it into the answer, instead of inferring elapsed time from\n    visual context. Returns [(t_seconds, PIL.Image), ...].\n    \"\"\"\n    cap = cv2.VideoCapture(str(video_path))\n    fps = cap.get(cv2.CAP_PROP_FPS)\n    n   = int(cap.get(cv2.CAP_PROP_FRAME_COUNT))\n    if fps <= 0 or n <= 0:\n        cap.release()\n        return []\n    duration = n / fps\n    t0 = 0.0 if t_start is None else max(0.0, t_start)\n    t1 = duration if t_end is None else min(duration, t_end)\n    if t1 <= t0:\n        t1 = min(duration, t0 + 1.0)\n\n    times = np.arange(t0, t1 + 1e-6, 1.0 / max(target_fps, 0.1))\n    if len(times) > limit:\n        times = times[np.linspace(0, len(times) - 1, limit).round().astype(int)]\n\n    out = []\n    for t in times:\n        cap.set(cv2.CAP_PROP_POS_FRAMES, int(np.clip(round(t * fps), 0, n - 1)))\n        ok, frame = cap.read()\n        if not ok:\n            continue\n        h, w = frame.shape[:2]\n        if max(h, w) > max_side:\n            sc = max_side / float(max(h, w))\n            frame = cv2.resize(frame, (int(round(w * sc)), int(round(h * sc))),\n                               interpolation=cv2.INTER_AREA)\n        if burn:\n            cv2.rectangle(frame, (8, 8), (250, 46), (0, 0, 0), -1)\n            cv2.putText(frame, f't={t:.2f}s', (14, 35), cv2.FONT_HERSHEY_SIMPLEX,\n                        0.9, (255, 255, 255), 2, cv2.LINE_AA)\n        out.append((float(t), PILImage.fromarray(cv2.cvtColor(frame, cv2.COLOR_BGR2RGB))))\n    cap.release()\n    return out\n\n\ndef _extract_json(text):\n    m = re.search(r'\\{.*\\}', text, flags=re.DOTALL)\n    if not m:\n        return None\n    try:\n        return json.loads(m.group(0))\n    except json.JSONDecodeError:\n        return None\n\n\ndef vlm_generate(prompt_text, images, max_new_tokens=256):\n    \"\"\"Greedy decode over a list of PIL frames. Greedy, not sampled: the\n    reference reports self-consistency at temperature 0.1-0.6 changing fewer\n    than ~10 predictions in 2027 clips, and CoT collapsing the temporal score\n    from 0.42 to 0.06.\"\"\"\n    if not VLM_AVAILABLE:\n        return ''\n    content = [{'type': 'image', 'image': im} for im in images]\n    content.append({'type': 'text', 'text': prompt_text})\n    inputs = vlm_processor.apply_chat_template(\n        [{'role': 'user', 'content': content}], tokenize=True,\n        add_generation_prompt=True, return_dict=True, return_tensors='pt',\n    ).to(vlm_model.device)\n    with torch.no_grad():\n        out = vlm_model.generate(**inputs, max_new_tokens=max_new_tokens, do_sample=False)\n    text = vlm_processor.decode(out[0, inputs['input_ids'].shape[1]:], skip_special_tokens=True)\n    del inputs, out\n    if DEVICE == 'cuda':\n        torch.cuda.empty_cache()\n    return text\n\n\nprint('[STATUS] prompts / sample_frames_stamped / vlm_generate defined')\n\ndef vlm_generate_oom_safe(prompt_text, frames, max_new_tokens=256, min_frames=4):\n    \"\"\"vlm_generate with OOM retry: halves the frame list and retries instead of\n    crashing. frames is a list of (t, PIL.Image) tuples, as returned by\n    sample_frames_stamped -- passing the tuples (not just the images) lets this\n    halve consistently no matter which stage calls it. Returns '' if even\n    min_frames still OOMs, so a caller can fall back to a default rather than die.\n\n    Every direct vlm_generate() call in this notebook that runs inside a loop\n    over many videos (Stage 2, Stage 3, the debug dump) should go through this\n    instead: stage1_full_scan already had its own retry loop, but Stage 2/3 and\n    the diagnostic cells did not, and a single oversized clip could kill an\n    otherwise-checkpointed multi-hour run.\n    \"\"\"\n    import gc\n    cur = list(frames)\n    while True:\n        gc.collect()\n        if torch.cuda.is_available():\n            torch.cuda.empty_cache()\n        try:\n            return vlm_generate(prompt_text, [im for _, im in cur], max_new_tokens)\n        except torch.cuda.OutOfMemoryError:\n            gc.collect()\n            torch.cuda.empty_cache()\n            if len(cur) <= min_frames:\n                print(f'[WARNING] OOM persists at {len(cur)} frames -- giving up, returning empty string')\n                return ''\n            cur = cur[::2] if len(cur) > 2 * min_frames else cur[:min_frames]\n            print(f'[WARNING] OOM -- retrying with {len(cur)} frames')\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T13:46:37.422086Z","iopub.execute_input":"2026-08-03T13:46:37.422943Z","iopub.status.idle":"2026-08-03T13:46:37.453117Z","shell.execute_reply.started":"2026-08-03T13:46:37.422913Z","shell.execute_reply":"2026-08-03T13:46:37.452112Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Fix -- NVRTC builtins missing (`libnvrtc-builtins.so.13.0`)\n\nBaseline torch is `2.11.0+cu130`, and the simple smoke-test op (`tensor * 2`)\npassed -- that op uses a PRECOMPILED kernel. The canary's `generate()` call hit\na reduction op (`prod`) that has no precompiled kernel for its dtype/shape and\nfalls back to JIT-compiling one via NVRTC at runtime, which needs\n`libnvrtc-builtins.so.13.0` on disk. That file is missing or mismatched --\nlikely because this torch build was already sitting in the Kaggle image\n(not a normal `pip install torch`), so its CUDA 13.0 runtime sub-packages\n(`nvidia-cuda-nvrtc-cu13` etc.) never went through pip's own dependency\nresolution. Run the cell below, then re-run the canary (no kernel restart\nneeded -- NVRTC is `dlopen`'d lazily on first JIT call, not at import time;\nrestart only if the retry below still fails).","metadata":{}},{"cell_type":"code","source":"# [FIX] Install/repair the matching NVRTC runtime for torch 2.11.0+cu130\nimport subprocess, sys, glob\n\n_torch_cuda = None\ntry:\n    import torch\n    _torch_cuda = torch.version.cuda\nexcept Exception as e:\n    print(f'[WARNING] could not import torch to read version: {e}')\nprint(f'[STATUS] torch reports CUDA {_torch_cuda}')\n\n_rc, _out, _err = subprocess.run(\n    [sys.executable, '-m', 'pip', 'show', 'nvidia-cuda-nvrtc-cu13'],\n    capture_output=True, text=True\n).returncode, '', ''\n_show = subprocess.run([sys.executable, '-m', 'pip', 'show', 'nvidia-cuda-nvrtc-cu13'],\n                       capture_output=True, text=True)\nprint(f'[STATUS] nvidia-cuda-nvrtc-cu13 currently installed: {_show.returncode == 0}')\nif _show.returncode == 0:\n    print(_show.stdout)\n\n# Find whatever libnvrtc-builtins*.so actually exists on disk right now, for comparison\n_found = glob.glob(sys.prefix + '/lib/python*/site-packages/nvidia/cuda_nvrtc/**/libnvrtc-builtins*.so*',\n                   recursive=True)\nprint(f'[STATUS] libnvrtc-builtins*.so found on disk: {_found}')\n\nprint('[STATUS] Reinstalling nvidia-cuda-nvrtc-cu13 (force, matching pip resolver -- '\n      'not --no-deps, so pip can pick a build compatible with installed torch)...')\n_repair = subprocess.run(\n    [sys.executable, '-m', 'pip', 'install', '-q', '--force-reinstall',\n     'nvidia-cuda-nvrtc-cu13'],\n    capture_output=True, text=True\n)\nprint(f'[STATUS] Reinstall exit code: {_repair.returncode}')\nif _repair.returncode != 0:\n    print(_repair.stderr[-2000:])\n\n_found_after = glob.glob(sys.prefix + '/lib/python*/site-packages/nvidia/cuda_nvrtc/**/libnvrtc-builtins*.so*',\n                         recursive=True)\nprint(f'[STATUS] libnvrtc-builtins*.so found after reinstall: {_found_after}')\n\n# Retry the exact op class that failed -- a JIT reduction, not the plain smoke-test op\ntry:\n    import torch\n    _t = torch.tensor([2, 3, 4], dtype=torch.int64, device='cuda')\n    _p = torch.prod(_t)\n    print(f'[SUCCESS] JIT reduction op works: prod([2,3,4]) = {_p.item()}')\nexcept Exception as e:\n    print(f'[ERROR] JIT reduction still fails: {type(e).__name__}: {e}')\n    print('[NEXT] If this still fails, Restart Kernel and re-run from the top -- the')\n    print('       repair above only replaces the file on disk, and if any process')\n    print('       already opened a handle to the old/missing one this session, a')\n    print('       fresh process is the only way to pick up the reinstalled library.')\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T13:46:37.454287Z","iopub.execute_input":"2026-08-03T13:46:37.454626Z","iopub.status.idle":"2026-08-03T13:46:46.261312Z","shell.execute_reply.started":"2026-08-03T13:46:37.454566Z","shell.execute_reply":"2026-08-03T13:46:46.260637Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Fix v2 -- `nvidia-cuda-nvrtc-cu13` khong co wheel dung san, bo huong vay lai NVRTC\n\n`pip install nvidia-cuda-nvrtc-cu13` that bai ngay o buoc BUILD WHEEL (khong phai\nloi mang/OOM) -- nghia la khong co prebuilt wheel cho package nay tren PyPI voi\nplatform/Python hien tai (CUDA 13.0 con qua moi). Vay lai dung file .so thieu la\nngo cut thuc su, khong phai do lam sai buoc nao.\n\n**Doi huong**: dung dung nhanh tu-sua da co san trong cell 8 (comment da noi ro:\n\"a safer bet than whatever cu128 build shipped with this image\") nhung LAN NAY\nCHAY NO CHU DONG, khong doi baseline smoke-test tu bao loi -- vi da chung minh\nsmoke-test hien tai (`tensor * 2`) qua don gian de bat loi JIT nay, no \"pass gia\".\nCell duoi ha thang ve `cu121` (torch/torchvision/torchaudio on dinh, da chay\nsm_75/T4 lien tuc nhieu nam) va nang cap chinh smoke-test de bao gom phep JIT-\nreduction (`torch.prod`) cho lan sau khong bi lua nua.\n\n**BAT BUOC Restart Kernel sau khi cell nay chay xong** -- torch cu130 da duoc\nimport vao process nay roi, pip reinstall khong thay the duoc module dang nam\ntrong bo nho; phai la process moi (kernel moi) moi nap dung ban cu121.","metadata":{}},{"cell_type":"code","source":"# [FIX v2] Ha torch ve cu121 on dinh -- KHONG doi smoke-test bao loi truoc,\n# chu dong ha vi da xac nhan cu130 hong o JIT-reduction va khong sua duoc.\nimport subprocess, sys\n\nprint('[STATUS] Force-reinstalling torch/torchvision/torchaudio tu cu121 index...')\n_rc = subprocess.run(\n    [sys.executable, '-m', 'pip', 'install', '-q', '--force-reinstall',\n     'torch', 'torchvision', 'torchaudio',\n     '--index-url', 'https://download.pytorch.org/whl/cu121'],\n    capture_output=True, text=True\n)\nprint(f'[STATUS] cu121 install exit code: {_rc.returncode}')\nif _rc.returncode != 0:\n    print(_rc.stderr[-2000:])\nelse:\n    # Re-pin the constraints file used by later installs (ultralytics/CLIP/transformers)\n    # to the NEW cu121 versions, so nothing pulls cu130 back in as a transitive dep.\n    _pinned = []\n    for _pkg in ('torch', 'torchvision', 'torchaudio'):\n        _show = subprocess.run([sys.executable, '-m', 'pip', 'show', _pkg],\n                               capture_output=True, text=True)\n        for _line in _show.stdout.splitlines():\n            if _line.startswith('Version:'):\n                _pinned.append(f\"{_pkg}=={_line.split(':', 1)[1].strip()}\")\n    with open('/kaggle/working/_torch_constraints.txt', 'w') as f:\n        f.write('\\n'.join(_pinned) + '\\n')\n    print(f'[STATUS] Re-pinned constraints for downstream installs: {_pinned}')\n\nprint()\nprint('[ACTION REQUIRED] Restart Kernel now, then run from the top.')\nprint('  Cell 8 will re-pin these same versions (constraints file already updated),')\nprint('  so ultralytics / CLIP / perception_models / transformers installs will not')\nprint('  pull torch back to cu130.')\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T13:46:46.262368Z","iopub.execute_input":"2026-08-03T13:46:46.262731Z","iopub.status.idle":"2026-08-03T13:48:57.403305Z","shell.execute_reply.started":"2026-08-03T13:46:46.262669Z","shell.execute_reply":"2026-08-03T13:48:57.402517Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [CHECK] The canary -- does the VLM actually generate?\n#\n# Loading without raising is not the same as working. A previous run loaded\n# nothing, returned '' from every call, and filled all 2027 rows with\n# normalize_prediction's defaults (0.35*duration, 0.5/0.5, rear-end): a constant\n# wearing a submission's clothes, produced after 24 minutes of GPU time.\n#\n# An earlier attempt to prevent that put this check in the loader cell, guarded\n# by \"if 'vlm_generate' in globals()\" -- and since vlm_generate is defined below\n# the loader, the guard silently skipped the check on every run. A safety check\n# that can quietly not run is not a safety check. This cell has no guard.\n_probe = PILImage.fromarray(np.full((448, 448, 3), 127, np.uint8))\n_out = vlm_generate('Reply with exactly: OK', [_probe], max_new_tokens=8)\nprint(f'[CANARY] generation returned: {_out!r}')\nassert _out.strip(), 'The VLM loaded but generates nothing. Do not continue.'\nprint('[CANARY] PASS -- the model is alive')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T13:48:57.404222Z","iopub.execute_input":"2026-08-03T13:48:57.404615Z","iopub.status.idle":"2026-08-03T13:48:58.673599Z","shell.execute_reply.started":"2026-08-03T13:48:57.404559Z","shell.execute_reply":"2026-08-03T13:48:58.672855Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [MODEL] The three stages\n\ndef _clip01(v, lo=0.0, hi=1.0):\n    return float(max(lo, min(hi, v)))\n\n\ndef _safe_float(v, default):\n    try:\n        return float(v)\n    except (TypeError, ValueError):\n        return float(default)\n\n\ndef normalize_prediction(pred, duration):\n    \"\"\"Coerce a raw VLM JSON into a valid submission row.\n\n    The default time is 0.35*duration, not 0.5: accidents sit earlier in a clip\n    than the midpoint. This is a per-clip parse failure handler ONLY -- if it is\n    firing on every clip, the model is not running and the caller must fail loudly\n    rather than emit these defaults 2027 times. run_inference_vlm enforces that.\n    \"\"\"\n    out = dict(pred or {})\n    out['accident_time'] = _clip01(_safe_float(out.get('accident_time'), duration * 0.35),\n                                   0.0, duration)\n    out['center_x'] = _clip01(_safe_float(out.get('center_x'), 0.5))\n    out['center_y'] = _clip01(_safe_float(out.get('center_y'), 0.5))\n    t = str(out.get('type') or 'rear-end').strip().lower()\n    out['type'] = t if t in COLLISION_TYPES else 'rear-end'\n    return out\n\n\ndef stage1_full_scan(video_path, duration, scene_layout=None):\n    \"\"\"Joint (time, location, type) over the whole clip. Already a complete answer.\n\n    Retries on OOM with half the frames rather than dying on one long 4K clip:\n    the real test set ranges from 5.9 s to 30 s and up to 3840x2160, so the\n    activation cost per clip varies by more than an order of magnitude. The\n    retry prints -- a clip silently answered from 6 frames is worth knowing about.\n    \"\"\"\n    hint = SCENE_TYPE_HINTS.get(scene_layout or '', '')\n    limit = VLM_CFG['whole_limit']\n    while True:\n        frames = sample_frames_stamped(video_path, VLM_CFG['whole_fps'], limit=limit,\n                                       max_side=VLM_CFG['image_max_side'], burn=True)\n        if not frames:\n            return {**normalize_prediction({}, duration), '_parsed': False}\n        try:\n            raw = vlm_generate(build_stage1_prompt(duration, hint), [im for _, im in frames])\n            break\n        except torch.cuda.OutOfMemoryError:\n            torch.cuda.empty_cache()\n            if limit <= 6:\n                raise\n            limit //= 2\n            print(f'[WARNING] OOM on {video_path.name} -- retrying with {limit} frames')\n    parsed = _extract_json(raw)\n    out = normalize_prediction(parsed, duration)\n    out['_parsed'] = parsed is not None\n    return out\n\n\ndef stage2_time_refine(video_path, t_base, duration):\n    \"\"\"Bounded correction to t_base -- never a replacement.\n\n    A dense window around t_base plus a few sparse context anchors outside it, so\n    the model still sees the motion leading into and out of the event. The\n    correction is down-weighted by alpha and hard-capped at delta_max because\n    ACCS is a harmonic mean: one large error costs more than many small ones buy.\n    Section 6.4 derived this same bounded form for the PE refinement and measured\n    the unbounded variant regressing T to 0.16.\n    \"\"\"\n    if not VLM_CFG['time_refine']:\n        return t_base\n    w = VLM_CFG['time_refine_window']\n    dense = sample_frames_stamped(video_path, VLM_CFG['time_refine_fps'],\n                                  t_base - w, t_base + w,\n                                  limit=VLM_CFG['time_refine_limit'],\n                                  max_side=VLM_CFG['image_max_side'], burn=True)\n    ctx = sample_frames_stamped(video_path, VLM_CFG['time_refine_context_fps'],\n                                t_base - VLM_CFG['time_refine_context_before'],\n                                t_base + VLM_CFG['time_refine_context_after'],\n                                limit=VLM_CFG['time_refine_context_limit'],\n                                max_side=VLM_CFG['image_max_side'], burn=True)\n    merged = sorted({round(t, 3): im for t, im in (ctx + dense)}.items())\n    if not merged:\n        return t_base\n    j = _extract_json(vlm_generate_oom_safe(TIME_REFINE_PROMPT, merged, 64))\n    if not j or 'accident_time' not in j:\n        return t_base\n    t_ref = _safe_float(j['accident_time'], t_base)\n    delta = np.clip(t_ref - t_base, -VLM_CFG['time_refine_max_shift_sec'],\n                    VLM_CFG['time_refine_max_shift_sec'])\n    return float(np.clip(t_base + VLM_CFG['time_refine_blend'] * delta, 0.0, duration))\n\n\ndef stage3_grounding(video_path, t_final):\n    \"\"\"Point at the impact in ONE frame at t_final. No timestamp overlay here --\n    the model has nothing left to read but the collision.\n\n    Largest single gain in the reference development path (+0.09356). Under the\n    Stage-1 joint query the model splits attention between when and where, and\n    the coordinates quantise onto a coarse grid; with the time already pinned,\n    this asks only 'where'.\n    \"\"\"\n    if not VLM_CFG['use_grounding']:\n        return None\n    # One frame, so it can afford the resolution Stage 1 cannot.\n    frames = sample_frames_stamped(video_path, 1.0, t_final, t_final + 1e-3, limit=1,\n                                   max_side=VLM_CFG.get('grounding_max_side',\n                                                        VLM_CFG['image_max_side']),\n                                   burn=False)\n    if not frames:\n        return None\n    text = vlm_generate_oom_safe(GROUNDING_PROMPT, frames, 64, min_frames=1)\n\n    j = _extract_json(text)\n    pt = None\n    if j:\n        if isinstance(j.get('point'), (list, tuple)) and len(j['point']) >= 2:\n            pt = (_safe_float(j['point'][0], np.nan), _safe_float(j['point'][1], np.nan))\n        elif 'center_x' in j and 'center_y' in j:\n            pt = (_safe_float(j['center_x'], np.nan), _safe_float(j['center_y'], np.nan))\n    if pt is None:\n        m = re.search(r'\\((\\d+(?:\\.\\d+)?)\\s*,\\s*(\\d+(?:\\.\\d+)?)\\)', text)\n        if m:\n            pt = (float(m.group(1)), float(m.group(2)))\n    if pt is None or any(np.isnan(v) for v in pt):\n        return None\n    x, y = pt\n    if x > 1.0 or y > 1.0:      # Qwen3-VL's native grounding scale is [0, 1000]\n        x, y = x / 1000.0, y / 1000.0\n    return (_clip01(x), _clip01(y))\n\n\ndef apply_scene_type_postfix(accident_type, scene_layout):\n    \"\"\"t-bone is geometrically impossible where traffic cannot cross.\"\"\"\n    if accident_type == 't-bone' and scene_layout in T_BONE_IMPOSSIBLE_LAYOUTS:\n        return 'rear-end'\n    return accident_type\n\n\nprint('[STATUS] stage1_full_scan / stage2_time_refine / stage3_grounding defined')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T13:48:58.674884Z","iopub.execute_input":"2026-08-03T13:48:58.675234Z","iopub.status.idle":"2026-08-03T13:48:58.692537Z","shell.execute_reply.started":"2026-08-03T13:48:58.675206Z","shell.execute_reply":"2026-08-03T13:48:58.691757Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Fix -- bien/ham bi xoa nham trong dot don dead-code truoc\n\n`TYPE_DEFINITIONS`, `CLASSIFICATION_PROMPTS`, `_parse_type_label` von nam trong\ncell \"Stage 2b -- Qwen2.5-VL multi-prompt voting\" (da xoa vi 2 ham chinh cua no,\n`_qwen_generate` va `qwen_type_votes`, thuc su chet). Nhung 3 ten nay VAN duoc\n`classify_type_cascade` (Section 6.5, dang song) dung -- kiem tra lai lan nay\nlam dung (soat TAT CA ten dinh nghia trong 7 cell da xoa, khong chi ham, ma ca\nhang so/dict), xac nhan day la landmine DUY NHAT trong 7 cell do; 6 cell con\nlai (44,45,46,55,60,63 theo so cu) an toan, khong con gi tham chieu toi.","metadata":{}},{"cell_type":"code","source":"# [FIX] TYPE_DEFINITIONS / CLASSIFICATION_PROMPTS / _parse_type_label -- khoi\n# phuc dung nguyen ban goc, dat truoc classify_type_cascade (Section 6.5) can no.\nTYPE_DEFINITIONS = (\n    \"head-on = two vehicles traveling in opposite directions collide front-to-front; \"\n    \"rear-end = the front of one vehicle hits the back of another traveling the same direction; \"\n    \"sideswipe = two roughly parallel vehicles make glancing side contact; \"\n    \"t-bone = the front of one vehicle strikes the side of another at a right angle; \"\n    \"single = exactly one vehicle crashes (wall, pole, barrier, or off the road), no second vehicle involved.\"\n)\n\nCLASSIFICATION_PROMPTS = {\n    'count-first': (\n        \"These frames from a CCTV camera show a traffic accident in chronological order. \"\n        \"First count how many moving vehicles are directly involved in the crash. \"\n        \"If exactly one vehicle crashes on its own, the answer is 'single'. \"\n        f\"Otherwise classify the collision. Definitions: {TYPE_DEFINITIONS} \"\n        \"Answer with only one label: head-on, rear-end, sideswipe, single, or t-bone.\"\n    ),\n    'motion': (\n        \"Watch the direction each vehicle travels across these chronological frames \"\n        \"up to the moment of impact. Based on their directions of travel at contact, \"\n        f\"classify the collision. Definitions: {TYPE_DEFINITIONS} \"\n        \"Answer with only one label: head-on, rear-end, sideswipe, single, or t-bone.\"\n    ),\n    'geometry': (\n        \"Focus on the angle and first contact point between the vehicles at the moment \"\n        f\"of collision in these frames. Definitions: {TYPE_DEFINITIONS} \"\n        \"Answer with only one label: head-on, rear-end, sideswipe, single, or t-bone.\"\n    ),\n}\n\n\ndef _parse_type_label(text: str) -> str:\n    normalized_text = text.lower().replace(' ', '-').replace('_', '-')\n    for label in COLLISION_TYPES:\n        if label in normalized_text:\n            return label\n    return None\n\n\nprint('[STATUS] TYPE_DEFINITIONS / CLASSIFICATION_PROMPTS / _parse_type_label restored')\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T13:48:58.693614Z","iopub.execute_input":"2026-08-03T13:48:58.694022Z","iopub.status.idle":"2026-08-03T13:48:58.709315Z","shell.execute_reply.started":"2026-08-03T13:48:58.693982Z","shell.execute_reply":"2026-08-03T13:48:58.708561Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [MODEL] Stage 2 (cascade) -- entropy/margin-gated classification, Qwen3-VL\n\nfrom collections import Counter\n\n# stage1_full_scan's `type` comes from ONE JSON call with no cross-check --\n# the same failure diagnosed for Qwen2.5-VL in Section 7b (a model that\n# settles on one class far more often than the label distribution supports).\n# Section 6.5's fix for that model was an unconditional multi-prompt ensemble\n# on every clip; that is not affordable here (Qwen3-VL-8B is already the\n# session's VRAM ceiling -- Section 6.1's audit). Instead: escalate only when\n# cheap signals disagree, per the metadata-aware cascade baseline -- most\n# clips exit after two extra short calls confirm what Stage 1 already said.\nELIMINATION_PROMPT_TEMPLATE = (\n    \"These CCTV frames show a traffic accident. Two earlier analyses of this \"\n    \"same clip disagreed: one concluded '{a}', the other concluded '{b}'. \"\n    f\"Definitions: {TYPE_DEFINITIONS} \"\n    \"Look carefully at how many vehicles are involved, the direction each one \"\n    \"is traveling, and the angle of contact, then decide which of the two \"\n    \"labels is correct. Answer with only one label: {a} or {b}.\"\n)\n\n\ndef _vote_entropy(counts) -> float:\n    total = sum(counts.values())\n    if total == 0:\n        return 0.0\n    probs = [c / total for c in counts.values() if c > 0]\n    return float(-sum(p * np.log2(p) for p in probs))\n\n\ndef classify_type_cascade(video_path: pathlib.Path, t_final: float, type_from_stage1: str) -> str:\n    \"\"\"Entropy/margin-gated classification.\n\n    votes = [stage1's own type (free -- already computed), 'motion' prompt,\n    'geometry' prompt]. If >=2 of 3 agree and entropy stays low, that majority\n    is the answer -- two extra cheap VLM calls, no more. Only on a genuine\n    3-way split does this pay for one more, explicit elimination call between\n    the top two candidates, rather than trusting a plurality vote on 3 items\n    (which is really just \"pick whichever 2 of the 3 short prompts happened\n    to agree\").\n    \"\"\"\n    if not VLM_AVAILABLE:\n        return type_from_stage1\n\n    frames = sample_frames_stamped(video_path, 2.0, t_final - 1.5, t_final + 1.0,\n                                   limit=6, max_side=VLM_CFG['image_max_side'], burn=False)\n    if not frames:\n        return type_from_stage1\n\n    votes = [type_from_stage1]\n    for prompt in CLASSIFICATION_PROMPTS.values():\n        label = _parse_type_label(vlm_generate_oom_safe(prompt, frames, 10, min_frames=2))\n        if label is not None:\n            votes.append(label)\n\n    counts = Counter(votes)\n    top_label, top_n = counts.most_common(1)[0]\n\n    # Full or majority (>=2 of <=3) agreement with low entropy: trust it, done.\n    if top_n >= 2 and _vote_entropy(counts) <= 1.0:\n        return top_label\n\n    # Every vote disagreed (entropy at its max for 3 items) -- escalate with\n    # one elimination call between the two most-recent, most-informed votes:\n    # Stage 1 saw the whole clip, 'geometry' looks at contact angle directly.\n    candidates = list(dict.fromkeys([type_from_stage1, votes[-1]]))\n    if len(candidates) < 2:\n        return top_label\n    a, b = candidates[0], candidates[1]\n    prompt = ELIMINATION_PROMPT_TEMPLATE.format(a=a, b=b)\n    label = _parse_type_label(vlm_generate_oom_safe(prompt, frames, 10, min_frames=2))\n    return label if label in (a, b) else top_label\n\n\nprint('[STATUS] classify_type_cascade defined (entropy/margin gate)')\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T13:48:58.710136Z","iopub.execute_input":"2026-08-03T13:48:58.710453Z","iopub.status.idle":"2026-08-03T13:48:58.725694Z","shell.execute_reply.started":"2026-08-03T13:48:58.710409Z","shell.execute_reply":"2026-08-03T13:48:58.724899Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [MODEL] Assembled Qwen3-VL pipeline\n\nPARSE_FAIL_STREAK = 0\nPARSE_FAIL_ABORT = 10\n\n\nTEMPORAL_ANCHOR_DELTA_MAX = 3.0   # seconds; same bounded-correction idea as\n                                  # Section 6.4's PE refinement, roles reversed:\n                                  # here the classical (VLM-free) estimate is the\n                                  # anchor and the VLM reading is the bounded\n                                  # correction, so a misread timestamp cannot drag\n                                  # the prediction arbitrarily far from a z-score\n                                  # + OSD estimate that never saw the VLM's answer.\n\n\ndef run_inference_vlm(video_path: pathlib.Path, sub_path: str = None) -> dict:\n    \"\"\"Three staged VLM calls, then classical-signal anchoring and the entropy\n    cascade, then the scene rule. No training, no tracking.\"\"\"\n    if not VLM_AVAILABLE:\n        raise RuntimeError('run_inference_vlm called with no usable VLM: every row '\n                           'would be normalize_prediction defaults, i.e. a constant.')\n    cap = cv2.VideoCapture(str(video_path))\n    fps, n = cap.get(cv2.CAP_PROP_FPS), int(cap.get(cv2.CAP_PROP_FRAME_COUNT))\n    cap.release()\n    duration = (n / fps) if fps > 0 else 20.0\n\n    scene = SCENE_BY_PATH.get(sub_path or ('videos/' + video_path.name))\n\n    pred = stage1_full_scan(video_path, duration, scene)\n\n    # A run where Stage 1 never parses is the constant-submission failure again,\n    # just slower. Ten in a row means stop and look, not carry on for 8 hours.\n    global PARSE_FAIL_STREAK\n    if pred.get('_parsed'):\n        PARSE_FAIL_STREAK = 0\n    else:\n        PARSE_FAIL_STREAK += 1\n        if PARSE_FAIL_STREAK >= PARSE_FAIL_ABORT:\n            raise RuntimeError(\n                f'Stage 1 failed to parse JSON on {PARSE_FAIL_STREAK} consecutive clips. '\n                'The predictions being written are normalize_prediction defaults. '\n                'Inspect the raw VLM output before continuing.')\n\n    # Anchor Stage 1's time against the classical (z-score + OSD) ensemble\n    # before the VLM's own bounded refinement runs, so a bad timestamp read\n    # in Stage 1 is caught by a signal that never looked at pixels the VLM\n    # hallucinated over -- this is the \"temporal drift\" failure mode.\n    t_classical = predict_accident_time_ensemble(video_path)\n    t_vlm = pred['accident_time']\n    correction = float(np.clip(t_vlm - t_classical, -TEMPORAL_ANCHOR_DELTA_MAX,\n                               TEMPORAL_ANCHOR_DELTA_MAX))\n    pred['accident_time'] = float(np.clip(t_classical + correction, 0.0, duration))\n\n    t_final = stage2_time_refine(video_path, pred['accident_time'], duration)\n    pred['accident_time'] = t_final\n\n    pt = stage3_grounding(video_path, t_final)\n    if pt is not None:                      # grounding_weight = 1.0: it replaces\n        pred['center_x'], pred['center_y'] = pt\n\n    # Entropy/margin cascade: only escalates past two extra cheap calls when\n    # they disagree with each other and with Stage 1 -- see classify_type_cascade.\n    pred['type'] = classify_type_cascade(video_path, t_final, pred['type'])\n    pred['type'] = apply_scene_type_postfix(pred['type'], scene)\n\n    return {'path': str(video_path), 'accident_time': pred['accident_time'],\n            'center_x': pred['center_x'], 'center_y': pred['center_y'],\n            'type': pred['type'], 'scene_layout': scene}\n\n\nprint('[STATUS] run_inference_vlm defined')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T13:48:58.726601Z","iopub.execute_input":"2026-08-03T13:48:58.726938Z","iopub.status.idle":"2026-08-03T13:48:58.741332Z","shell.execute_reply.started":"2026-08-03T13:48:58.726898Z","shell.execute_reply":"2026-08-03T13:48:58.740453Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 6.11 Assembled Pipelines\n\n`run_inference_vlm` (Section 6.10) is the pipeline. `run_inference_final` below is the tracking-based design, retained so Section 7 can score all three against each other and against the constant floor; its measured result is a negative one.","metadata":{}},{"cell_type":"code","source":"# [MODEL] Baseline pipeline -- classical signals + CLIP only (kept for comparison)\n\ndef run_inference_baseline(video_path: pathlib.Path) -> dict:\n    acc_time = predict_accident_time(video_path)\n    cx, cy   = predict_impact_location_flow(video_path, accident_time=acc_time)\n    col_type = predict_collision_type_clip(video_path, acc_time)\n    return {\n        'path': str(video_path), 'accident_time': acc_time,\n        'center_x': cx, 'center_y': cy, 'type': col_type,\n    }\n\n\n# [MODEL] Final pipeline -- tracking when its own evidence supports it, constant otherwise\n#\n# What changed, and why:\n#\n#  1. The gate. Tracking used to be trusted only when its impact time agreed\n#     with the frame-difference anchor. That anchor scores T=0.31 against a\n#     constant's 0.52 -- it was validating tracking against a signal weaker\n#     than a constant. The gate is now the tracker's own kinetic evidence.\n#  2. The fallbacks are constants. Previously a distrusted video fell through to\n#     OWLv2 / optical-flow (S=0.156) and the frame-difference anchor (T=0.314),\n#     both below the constant floor. Falling back to a weaker model than the\n#     prior subtracts score on every video it touches.\n#  3. Qwen, the Perception Encoder, and OWLv2 are gone. Qwen dragged C from\n#     0.30 to 0.15 (below the 0.20 chance level) by collapsing to 'rear-end' on\n#     18/20 videos, and it dominated the runtime. Section 7b re-measures every\n#     removed signal against its constant rather than taking this on faith.\n\nTRACK_EVIDENCE_MIN = 0.0    # set from the Section 7b sweep before the real run\n\n# Set by Section 6.8. None => fall back to the hand-written geometry rule, so\n# the pipeline still runs end-to-end if feature extraction was skipped.\nTYPE_CLASSIFIER = globals().get('TYPE_CLASSIFIER')\n\n# Systematic offset between a tracked box-contact and the dataset's definition\n# of accident_time, measured in Section 6.8. 0.0 if that section was skipped.\nTIME_OFFSET = globals().get('TIME_OFFSET', 0.0)\n\n# CLIP scores C=0.30, which beats the constant on the BALANCED calibration split\n# (0.20) but loses to it under the natural synthetic prior (0.359). The real test\n# prior is unknown, so this stays an explicit switch rather than a silent choice.\nUSE_CLIP_TYPE_FALLBACK = True\n\n\ndef _type_fallback(video_path: pathlib.Path, acc_time: float) -> str:\n    if USE_CLIP_TYPE_FALLBACK and CLIP_AVAILABLE:\n        return predict_collision_type_clip(video_path, acc_time)\n    return CONST['type']\n\n\ndef _predict_type(coll: dict, video_path: pathlib.Path, acc_time: float) -> str:\n    \"\"\"Type from the supervised classifier when its features exist, else the\n    hand-written geometry rule, else the zero-shot fallback.\n\n    The trained model supersedes _type_from_geometry: the rule hard-codes angle\n    bands (>=135 head-on, <=40 rear-end/sideswipe, 60-120 t-bone) and abstains\n    in between, while the tree learns the same boundaries from 2211 labelled\n    examples and never abstains.\n    \"\"\"\n    feats = coll.get('features')\n    if TYPE_CLASSIFIER is not None and feats is not None:\n        x = np.array([[feats.get(c, np.nan) for c in FEATURE_COLS]], dtype=float)\n        return str(TYPE_CLASSIFIER.predict(x)[0])\n    return coll.get('type_hint') or _type_fallback(video_path, acc_time)\n\n\ndef run_inference_final(video_path: pathlib.Path) -> dict:\n    coll = analyze_video_collision(video_path)\n    trusted = coll['found'] and coll.get('evidence', 0.0) >= TRACK_EVIDENCE_MIN\n\n    if trusted:\n        acc_time = coll['t_impact'] - TIME_OFFSET\n        cx, cy   = coll['center']\n        col_type = _predict_type(coll, video_path, acc_time)\n    else:\n        acc_time = CONST['accident_time']\n        cx, cy   = CONST['center_x'], CONST['center_y']\n        col_type = _type_fallback(video_path, acc_time)\n\n    return {\n        'path': str(video_path), 'accident_time': acc_time,\n        'center_x': cx, 'center_y': cy, 'type': col_type,\n        'track_used': bool(trusted),\n    }\n\n\nprint('[STATUS] run_inference_baseline / run_inference_final defined')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T13:48:58.742241Z","iopub.execute_input":"2026-08-03T13:48:58.742668Z","iopub.status.idle":"2026-08-03T13:48:58.758094Z","shell.execute_reply.started":"2026-08-03T13:48:58.742610Z","shell.execute_reply":"2026-08-03T13:48:58.757221Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"---\n## 7. Evaluation\n\nScoring follows the official Evaluation page: Gaussian temporal similarity (T), Gaussian spatial similarity (S), and top-1 collision-type accuracy (C). Each component is **averaged across videos first**, and the leaderboard score is the harmonic mean of those three averages:\n\n$$\\text{ACCIDENT score} = \\frac{3}{\\frac{1}{\\bar{T}} + \\frac{1}{\\bar{S}} + \\frac{1}{\\bar{C}}}$$\n\n**This ordering matters, and an earlier revision had it backwards.** It computed a harmonic mean per video and averaged *that*. Because C is 0/1, that variant zeroes out every misclassified video and is bounded above by accuracy -- a different and much harsher metric. It reported 0.087 for the baseline where the real metric gives **0.2321**, and 0.0000 for the final pipeline where the real metric gives **0.1739**. It also produced the false conclusion that classification dominates the score (see Section 6.5). Under the real metric the harmonic mean penalizes the **weakest** component, which for this pipeline is spatial localization.\n\nMeasured leverage on the baseline (T=0.314, S=0.156, C=0.300):\n\n| improvement | score becomes | gain |\n|---|---|---|\n| S: 0.156 → 0.45 | 0.3433 | **+0.111** |\n| C: 0.300 → 0.45 | 0.2539 | +0.022 |\n| T: 0.314 → 0.45 | 0.2507 | +0.019 |\n\nSpatial has roughly five times the leverage of classification -- the reverse of what this notebook was built to assume.\n\n**The sigma values are not published.** The Evaluation page says only \"Gaussian-style similarity\". `SIGMA_T = 2.0` and `SIGMA_S = 0.1` are assumptions. They matter enormously: a constant spatial prediction scores 0.095 at σ=0.05 but 0.622 at σ=0.20, which inverts the priority order above. Every conclusion here is conditional on them.\n\n**The constant-prediction floor (Section 7a).** Because σ_S is comparable to the spread of the labels themselves (center_x std 0.13, center_y std 0.18), the dataset's own prior is a strong predictor: any model whose error exceeds the label spread scores *worse than a constant*. The floor is fitted on the synthetic training set and evaluated on the calibration split. **Any component that cannot beat its own constant is subtracting value and should be replaced by that constant** -- including as a fallback, where this pipeline currently falls back from a weak model to an even weaker one.\n\n**Calibration set caveats.** Calibration uses a diverse set of 20 videos -- 4 per collision type. This guards against the Section 7b failure (tuning on a head-on-only subset reached 0.80 there and 0.10 on the diverse set). But two limits should be stated plainly: the balanced split makes the measured C a *balanced* accuracy, which is **not** what the leaderboard measures (the real test set has its own unknown class prior); and n=20 is far too small to separate differences of ~0.15 in C, which is 3 videos.","metadata":{}},{"cell_type":"code","source":"# [EVAL] Competition scoring functions -- matches the official Evaluation page\n#\n# The leaderboard averages each component across videos FIRST, then takes the\n# harmonic mean of those three averages:\n#\n#     ACCIDENT score = 3 / (1/mean(T) + 1/mean(S) + 1/mean(C))\n#\n# An earlier revision computed a per-video harmonic mean and averaged *that*.\n# Because C is 0/1, that variant zeroes every misclassified video, so it is\n# bounded above by accuracy -- a different, far harsher metric. It reported\n# 0.087 for the baseline where the real metric gives 0.232, and it is what\n# motivated the (false) claim that classification dominates the score.\n\n# SIGMA_T / SIGMA_S are defined with the constant prior in Section 6.7.\n\ndef temporal_score(pred_time: float, gt_time: float) -> float:\n    \"\"\"Gaussian temporal similarity averaged over the three tolerances.\"\"\"\n    return float(np.mean([np.exp(-0.5 * ((pred_time - gt_time) / s) ** 2)\n                          for s in SIGMA_T_LIST]))\n\n\ndef spatial_score(pred_x: float, pred_y: float, gt_x: float, gt_y: float) -> float:\n    \"\"\"Anisotropic Gaussian spatial similarity; sigmas are the mean bbox size.\"\"\"\n    return float(np.exp(-0.5 * (((pred_x - gt_x) / SIGMA_X) ** 2\n                                + ((pred_y - gt_y) / SIGMA_Y) ** 2)))\n\n\ndef classification_score(pred_type: str, gt_type: str) -> int:\n    \"\"\"1 if predicted type matches ground truth, else 0.\"\"\"\n    return int(pred_type.strip().lower() == gt_type.strip().lower())\n\n\ndef accident_score(eval_df: pd.DataFrame) -> float:\n    \"\"\"Official leaderboard score for an already-scored prediction set.\"\"\"\n    T, S, C = eval_df['T'].mean(), eval_df['S'].mean(), eval_df['C'].mean()\n    if min(T, S, C) <= 0:\n        return 0.0\n    return float(3.0 / (1.0 / T + 1.0 / S + 1.0 / C))\n\n\ndef score_predictions(df_pred: pd.DataFrame, labels_ref: pd.DataFrame) -> pd.DataFrame:\n    \"\"\"Join predictions to ground truth on video stem; return per-video T/S/C.\n\n    Deliberately no per-video H column -- the metric is not defined per video.\n    Pass the returned frame to accident_score().\n    \"\"\"\n    df_pred = df_pred.copy()\n    df_pred['video_stem'] = df_pred['path'].map(lambda p: pathlib.Path(p).stem)\n\n    labels_ref = labels_ref.copy()\n    labels_ref['video_stem'] = labels_ref['rgb_path'].map(lambda p: pathlib.Path(p).stem)\n\n    merged = df_pred.merge(\n        labels_ref[['video_stem', 'accident_time', 'center_x', 'center_y', 'type']],\n        on='video_stem', how='inner', suffixes=('_pred', '_gt'),\n    )\n    merged['T'] = merged.apply(\n        lambda r: temporal_score(r['accident_time_pred'], r['accident_time_gt']), axis=1)\n    merged['S'] = merged.apply(\n        lambda r: spatial_score(r['center_x_pred'], r['center_y_pred'],\n                                r['center_x_gt'], r['center_y_gt']), axis=1)\n    merged['C'] = merged.apply(\n        lambda r: classification_score(r['type_pred'], r['type_gt']), axis=1)\n    return merged\n\n\nprint('[STATUS] Scoring functions defined (average components, then harmonic mean)')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T13:48:58.759031Z","iopub.execute_input":"2026-08-03T13:48:58.759368Z","iopub.status.idle":"2026-08-03T13:48:58.774041Z","shell.execute_reply.started":"2026-08-03T13:48:58.759303Z","shell.execute_reply":"2026-08-03T13:48:58.773414Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [EVAL] Build the diverse calibration set -- N_PER_TYPE videos per collision type\nN_PER_TYPE = 4\n\ndiverse_rows = []\nfor coll_type in COLLISION_TYPES:\n    subset = labels_clean[labels_clean['type'] == coll_type]\n    n_sample = min(N_PER_TYPE, len(subset))\n    if n_sample == 0:\n        print(f'[WARNING] No videos found for type={coll_type}, skipping')\n        continue\n    diverse_rows.append(subset.sample(n_sample, random_state=SEED))\n    print(f'[STATUS] {coll_type}: sampled {n_sample}/{len(subset)} videos')\n\ndiverse_labels_df = pd.concat(diverse_rows, ignore_index=True)\ndiverse_videos = [pathlib.Path(p) for p in diverse_labels_df['abs_video_path']]\n\n# Every path was existence-checked during cleaning; assert anyway before spending GPU time\nassert all(vp.exists() for vp in diverse_videos), 'Some calibration videos are missing on disk'\nprint(f'[STATUS] Diverse calibration set: {len(diverse_videos)} videos across '\n      f\"{diverse_labels_df['type'].nunique()} types\")\ndisplay(diverse_labels_df[['rgb_path', 'type']])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T13:48:58.774983Z","iopub.execute_input":"2026-08-03T13:48:58.775329Z","iopub.status.idle":"2026-08-03T13:48:58.840315Z","shell.execute_reply.started":"2026-08-03T13:48:58.775302Z","shell.execute_reply":"2026-08-03T13:48:58.839555Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 7a. The Constant-Prediction Floor\n\nThe constant prior fitted in Section 6.7 is scored here on the calibration split. It is the threshold every component must clear: **anything that loses to its own constant is subtracting score and should be replaced by that constant**, including as a fallback.","metadata":{}},{"cell_type":"code","source":"# [EVAL] Score the constant prior on the calibration split -- the floor\nconst_df = pd.DataFrame({'path': diverse_labels_df['rgb_path'].to_numpy(),\n                         'accident_time': CONST['accident_time'],\n                         'center_x': CONST['center_x'],\n                         'center_y': CONST['center_y'],\n                         'type': CONST['type']})\nconst_eval = score_predictions(const_df, diverse_labels_df)\n\nprint(f\"[FLOOR] T={const_eval['T'].mean():.4f}  S={const_eval['S'].mean():.4f}  \"\n      f\"C={const_eval['C'].mean():.4f}  ->  ACCIDENT score = {accident_score(const_eval):.4f}\")\n\n# The calibration split is balanced 4-per-type, so its C understates what the\n# same constant scores under the natural class prior. The real test prior is\n# unknown, so neither number is the leaderboard's C.\nnatural_prior = labels_clean['type'].value_counts(normalize=True).max()\nprint(f\"[NOTE] C on this balanced split = {const_eval['C'].mean():.4f}; under the natural \"\n      f'synthetic prior the same constant would score {natural_prior:.4f}. '\n      'The real test prior is unknown.')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T13:48:58.841219Z","iopub.execute_input":"2026-08-03T13:48:58.841593Z","iopub.status.idle":"2026-08-03T13:48:58.861670Z","shell.execute_reply.started":"2026-08-03T13:48:58.841539Z","shell.execute_reply":"2026-08-03T13:48:58.860801Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 7b. Per-Signal Ablation\n\nEvery signal measured **in isolation against its own constant**, on one pass over the calibration set. This is the measurement the pipeline was missing: without it there is no way to tell which of five models is carrying the score and which is destroying it, and the ensemble weights in Section 6.5 were tuned blind.\n\nThe final column is the actionable one -- what the score becomes when that signal is used where it fires and the constant is used everywhere else. A signal whose hybrid column does not beat the pure constant should be deleted, not reweighted.\n\nThe last block sweeps the tracking evidence gate. `TRACK_EVIDENCE_MIN` in Section 6.8 should be set from it before the real run.\n\n**n=20.** Every number here has an error bar of roughly one video (0.05 in C). Treat the ordering as a signal and the exact values as noise.","metadata":{}},{"cell_type":"code","source":"# [DIAG] Per-signal ablation -- one pass per video, every signal recorded\ndiag_rows = []\nfor i, vp in enumerate(diverse_videos):\n    coll   = analyze_video_collision(vp)\n    t_base = predict_accident_time(vp)\n    t_ref  = coll['t_impact'] if coll['found'] else t_base\n    fx, fy = predict_impact_location_flow(vp, accident_time=t_ref)\n    probs  = clip_type_probabilities(vp, t_ref)\n\n    diag_rows.append({\n        'stem'      : vp.stem,\n        'found'     : bool(coll['found']),\n        'evidence'  : float(coll.get('evidence', 0.0)),\n        'mode'      : coll.get('mode'),\n        't_track'   : coll['t_impact'] if coll['found'] else np.nan,\n        'x_track'   : coll['center'][0] if coll['found'] else np.nan,\n        'y_track'   : coll['center'][1] if coll['found'] else np.nan,\n        'type_track': coll.get('type_hint') if coll['found'] else None,\n        't_base'    : t_base,\n        'x_flow'    : fx, 'y_flow': fy,\n        'type_clip' : max(probs, key=probs.get) if probs else None,\n    })\n    print(f'[STATUS] {i+1}/{len(diverse_videos)}: {vp.name} | found={coll[\"found\"]} '\n          f'| evidence={coll.get(\"evidence\", 0.0):.4f}')\n    if torch.cuda.is_available():\n        torch.cuda.empty_cache()\n\ndiag = pd.DataFrame(diag_rows)\nprint(f\"\\n[STATUS] Tracking found a collision in {diag['found'].mean():.0%} of videos\")\n\n# --- align ground truth to diag row order ---\ngt = diverse_labels_df.copy()\ngt['stem'] = gt['rgb_path'].map(lambda p: pathlib.Path(p).stem)\ng = gt.set_index('stem').loc[diag['stem']]\ngt_t = g['accident_time'].to_numpy()\ngt_x = g['center_x'].to_numpy()\ngt_y = g['center_y'].to_numpy()\ngt_c = g['type'].to_numpy()\n\nfound = diag['found'].to_numpy()\nn = len(diag)\n\ndef sc_T(pred): return np.exp(-0.5 * ((pred - gt_t) / SIGMA_T) ** 2)\ndef sc_S(px, py): return np.exp(-0.5 * ((px - gt_x) ** 2 + (py - gt_y) ** 2) / SIGMA_S ** 2)\n# pd.isna, not 'is not None': a DataFrame column holding str and None comes\n# back with the None coerced to float NaN, and NaN is not None. Using an\n# identity check here silently counts every missing hint as a live prediction.\ndef sc_C(pred): return np.array([1.0 if (not pd.isna(p) and p == c) else 0.0\n                                 for p, c in zip(pred, gt_c)])\n\n# Constants, evaluated everywhere\nT_const = sc_T(np.full(n, CONST['accident_time']))\nS_const = sc_S(np.full(n, CONST['center_x']), np.full(n, CONST['center_y']))\nC_const = sc_C([CONST['type']] * n)\n\n# Track signals, NaN-safe: only defined where found\nt_tr = np.where(found, diag['t_track'].to_numpy(), CONST['accident_time'])\nx_tr = np.where(found, diag['x_track'].to_numpy(), CONST['center_x'])\ny_tr = np.where(found, diag['y_track'].to_numpy(), CONST['center_y'])\n# Normalize the NaN that pandas substituted for None back to None, so the\n# coverage mask below is not inflated and the videos without a geometric hint\n# fall back to the constant instead of scoring a hard 0.\nc_tr = [t if (f and not pd.isna(t)) else None\n        for f, t in zip(found, diag['type_track'])]\n\nrows = []\ndef add(component, signal, per_video, mask=None):\n    \"\"\"Report a signal on its coverage, and blended with the constant elsewhere.\"\"\"\n    cov = float(mask.mean()) if mask is not None else 1.0\n    on_cov = float(per_video[mask].mean()) if (mask is not None and mask.any()) \\\n             else (float(per_video.mean()) if mask is None else float('nan'))\n    const = {'T': T_const, 'S': S_const, 'C': C_const}[component]\n    hybrid = float(np.where(mask, per_video, const).mean()) if mask is not None \\\n             else float(per_video.mean())\n    rows.append({'component': component, 'signal': signal, 'coverage': cov,\n                 'score_on_coverage': on_cov, 'hybrid_with_constant': hybrid})\n\nadd('T', 'constant',            T_const)\nadd('T', 'frame-diff anchor',   sc_T(diag['t_base'].to_numpy()))\nadd('T', 'tracking',            sc_T(t_tr), mask=found)\nadd('S', 'constant',            S_const)\nadd('S', 'optical-flow centroid', sc_S(diag['x_flow'].to_numpy(), diag['y_flow'].to_numpy()))\nadd('S', 'tracking',            sc_S(x_tr, y_tr), mask=found)\nadd('C', 'constant',            C_const)\nadd('C', 'CLIP zero-shot',      sc_C(diag['type_clip']))\nadd('C', 'track geometry',      sc_C(c_tr),\n    mask=np.array([c is not None for c in c_tr]))\n\nablation = pd.DataFrame(rows).round(4)\nprint('\\n[ABLATION] every signal against its own constant')\ndisplay(ablation)\n\nprint('[VERDICT] does the signal beat its constant when blended?')\nfor comp in ['T', 'S', 'C']:\n    floor = ablation[(ablation.component == comp) &\n                     (ablation.signal == 'constant')]['hybrid_with_constant'].iloc[0]\n    for _, r in ablation[(ablation.component == comp) &\n                         (ablation.signal != 'constant')].iterrows():\n        verdict = 'KEEP  ' if r.hybrid_with_constant > floor else 'DELETE'\n        print(f'  {verdict} {comp} {r.signal:22s} {r.hybrid_with_constant:.4f} '\n              f'vs constant {floor:.4f}')\n\n# --- evidence gate sweep: where should TRACK_EVIDENCE_MIN sit? ---\nif found.any():\n    S_tr_pv, T_tr_pv = sc_S(x_tr, y_tr), sc_T(t_tr)\n    ev = diag['evidence'].to_numpy()\n    sweep = []\n    for q in [0.0, 0.25, 0.50, 0.75]:\n        thr = float(np.quantile(ev[found], q)) if q > 0 else 0.0\n        m = found & (ev >= thr)\n        sweep.append({\n            'threshold': round(thr, 5), 'kept': f'{m.sum()}/{n}',\n            'T': round(float(np.where(m, T_tr_pv, T_const).mean()), 4),\n            'S': round(float(np.where(m, S_tr_pv, S_const).mean()), 4),\n        })\n    print('\\n[SWEEP] TRACK_EVIDENCE_MIN -- tracking where evidence >= threshold, constant elsewhere')\n    display(pd.DataFrame(sweep))\n    print(f\"[BASELINE] constant everywhere: T={T_const.mean():.4f}  S={S_const.mean():.4f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T13:48:58.862678Z","iopub.execute_input":"2026-08-03T13:48:58.862933Z","iopub.status.idle":"2026-08-03T13:52:01.137116Z","shell.execute_reply.started":"2026-08-03T13:48:58.862893Z","shell.execute_reply":"2026-08-03T13:52:01.136471Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [EVAL] Run baseline and final pipelines on the diverse calibration set\ncalib_results_baseline, calib_results_final = [], []\n\nfor i, vp in enumerate(diverse_videos):\n    res_base  = run_inference_baseline(vp)\n    res_final = run_inference_final(vp)\n    calib_results_baseline.append(res_base)\n    calib_results_final.append(res_final)\n\n    gt_type = diverse_labels_df.iloc[i]['type']\n    print(f'[STATUS] {i+1}/{len(diverse_videos)}: {vp.name} | gt={gt_type} | '\n          f'baseline={res_base[\"type\"]} | final={res_final[\"type\"]} | '\n          f'track={\"yes\" if res_final.get(\"track_used\") else \"no\"}')\n    if torch.cuda.is_available():\n        torch.cuda.empty_cache()\n\ncalib_df_baseline = pd.DataFrame(calib_results_baseline)\ncalib_df_final    = pd.DataFrame(calib_results_final)\n\ntrack_rate = calib_df_final['track_used'].mean()\nprint(f'[STATUS] Tracking trusted on {track_rate:.0%} of calibration videos '\n      '(the rest used the zero-shot fallback chain)')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T13:52:01.138110Z","iopub.execute_input":"2026-08-03T13:52:01.138554Z","iopub.status.idle":"2026-08-03T13:55:15.106614Z","shell.execute_reply.started":"2026-08-03T13:52:01.138522Z","shell.execute_reply":"2026-08-03T13:55:15.105750Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [EVAL] Score both pipelines against the constant floor\neval_baseline = score_predictions(calib_df_baseline, diverse_labels_df)\neval_final    = score_predictions(calib_df_final,    diverse_labels_df)\n\nprint(f'[COMPARISON] Diverse calibration set (n={len(eval_final)})')\ncomparison = pd.DataFrame({\n    'constant floor': [const_eval['T'].mean(), const_eval['S'].mean(),\n                       const_eval['C'].mean(), accident_score(const_eval)],\n    'baseline'      : [eval_baseline['T'].mean(), eval_baseline['S'].mean(),\n                       eval_baseline['C'].mean(), accident_score(eval_baseline)],\n    'final'         : [eval_final['T'].mean(), eval_final['S'].mean(),\n                       eval_final['C'].mean(), accident_score(eval_final)],\n}, index=['T', 'S', 'C', 'ACCIDENT score']).round(4)\ndisplay(comparison)\n\n# Per-component verdict: anything that loses to its own constant is subtracting\n# value and should simply be replaced by that constant.\nprint('[VERDICT] Does each component beat its constant?')\nfor comp in ['T', 'S', 'C']:\n    floor = const_eval[comp].mean()\n    for name, ev in [('baseline', eval_baseline), ('final', eval_final)]:\n        got = ev[comp].mean()\n        flag = 'BEATS' if got > floor else 'LOSES TO'\n        print(f'  {comp} {name:8s}: {got:.4f}  {flag} constant {floor:.4f}')\n\nprint('[BREAKDOWN] Classification accuracy (C) by ground-truth type -- final pipeline')\nbreakdown = eval_final.groupby('type_gt')['C'].agg(['mean', 'count']).round(4)\nbreakdown.columns = ['accuracy', 'n_videos']\ndisplay(breakdown)\n\nprint('[CONFUSION] Predicted vs ground-truth type -- final pipeline')\ndisplay(pd.crosstab(eval_final['type_gt'], eval_final['type_pred']))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T13:55:15.107595Z","iopub.execute_input":"2026-08-03T13:55:15.107964Z","iopub.status.idle":"2026-08-03T13:55:15.154919Z","shell.execute_reply.started":"2026-08-03T13:55:15.107905Z","shell.execute_reply":"2026-08-03T13:55:15.154308Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [EVAL] Component score distributions vs. the constant floor -- final pipeline\nif not eval_final.empty:\n    fig, axes = plt.subplots(1, 3, figsize=(15, 4.5))\n\n    score_cols  = ['T', 'S', 'C']\n    score_names = ['Temporal Score', 'Spatial Score', 'Classification Score']\n    colors      = [PALETTE['primary'], PALETTE['secondary'], PALETTE['tertiary']]\n\n    for ax, col, name, color in zip(axes, score_cols, score_names, colors):\n        sns.histplot(eval_final[col], bins=20, kde=True, color=color, ax=ax)\n        ax.axvline(eval_final[col].mean(), linestyle='--', color='black',\n                   linewidth=1.2, label=f'pipeline={eval_final[col].mean():.3f}')\n        ax.axvline(const_eval[col].mean(), linestyle=':', color='crimson',\n                   linewidth=1.6, label=f'constant={const_eval[col].mean():.3f}')\n        ax.set_title(name, fontsize=12)\n        ax.set_xlabel('Score', fontsize=10)\n        ax.set_ylabel('Count', fontsize=10)\n        ax.legend(fontsize=9)\n\n    plt.suptitle(f'Component Scores vs. Constant Floor -- final pipeline '\n                 f'({accident_score(eval_final):.4f} vs floor {accident_score(const_eval):.4f})',\n                 fontsize=13, y=1.03)\n    plt.tight_layout()\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T13:55:15.155940Z","iopub.execute_input":"2026-08-03T13:55:15.156229Z","iopub.status.idle":"2026-08-03T13:55:15.773629Z","shell.execute_reply.started":"2026-08-03T13:55:15.156191Z","shell.execute_reply":"2026-08-03T13:55:15.772721Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [EVAL] Score sensitivity analysis -- Gaussian scoring curves\nfig, axes = plt.subplots(1, 2, figsize=(14, 5))\n\ntime_errors = np.linspace(-15, 15, 300)  # seconds from ground truth\nfor sigma in [1.0, 2.0, 4.0]:\n    axes[0].plot(time_errors, np.exp(-0.5 * (time_errors / sigma) ** 2),\n                 label=f'sigma={sigma}s', linewidth=1.8)\naxes[0].set_title('Temporal Score vs. Time Error')\naxes[0].set_xlabel('Prediction Error (seconds)')\naxes[0].set_ylabel('Temporal Score T')\naxes[0].axhline(0.5, linestyle=':', color='grey', linewidth=0.8)\naxes[0].legend()\n\ndist_errors = np.linspace(0, 1.0, 300)  # normalized Euclidean distance\nfor sigma in [0.05, 0.10, 0.20]:\n    axes[1].plot(dist_errors, np.exp(-0.5 * (dist_errors / sigma) ** 2),\n                 label=f'sigma={sigma}', linewidth=1.8)\naxes[1].set_title('Spatial Score vs. Euclidean Location Error')\naxes[1].set_xlabel('Normalized Euclidean Distance')\naxes[1].set_ylabel('Spatial Score S')\naxes[1].axhline(0.5, linestyle=':', color='grey', linewidth=0.8)\naxes[1].legend()\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T13:55:15.774757Z","iopub.execute_input":"2026-08-03T13:55:15.775147Z","iopub.status.idle":"2026-08-03T13:55:16.114419Z","shell.execute_reply.started":"2026-08-03T13:55:15.775120Z","shell.execute_reply":"2026-08-03T13:55:16.113582Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 7d-bis. What Is the VLM Actually Saying?\n\nThe first calibration run predicted **`t-bone` for all seven clips it reached**, four of them `head-on` and three `rear-end` — C=0/7. Times and coordinates varied per clip, so the model was responding to the video rather than emitting a constant; only the type collapsed.\n\nThat has two candidate explanations and they need different fixes:\n\n1. **The model is answering badly.** A collapsed class is the documented failure mode of small VLMs on this task — the reference reports *\"InternVL3.5-8B and Gemma 3-27B mapped most predictions to rear-end\"* and *\"Cosmos-Reason2 missed the t-bone class entirely.\"* Here it is 4-bit 8B, and after the OOM retries it saw 12 frames over ~20 s.\n2. **The model is answering fine and we are mis-reading it.** JSON in a code fence, a label like `\"T-bone\"` or `\"t_bone\"`, a refusal, or truncation at `max_new_tokens` would all be silently coerced to a default by `normalize_prediction`.\n\nGuessing between these is pointless when the raw string is one print away. This cell dumps exactly what comes back, before any parsing.","metadata":{}},{"cell_type":"code","source":"# [DIAG] Dump the raw Stage-1 output -- no parsing, no coercion\n#\n# FIX: this cell used to call vlm_generate() directly with no OOM protection.\n# stage1_full_scan() (Section 6.10) already retries on OOM by halving the frame\n# count -- this cell skipped that safety net, so a video needing more frames\n# (32 here vs. the first video's 25) could exceed whatever VRAM was still\n# fragmented/held from the previous generate() call and crash the whole cell\n# instead of degrading gracefully. Now routed through vlm_generate_oom_safe\n# (defined in Section 6.10 / cell 73), the same helper Stage 2 and Stage 3 use.\nimport gc\n\nDEBUG_N = 3\n\nfor vp, gt in zip(diverse_videos[:DEBUG_N], diverse_labels_df['type'][:DEBUG_N]):\n    gc.collect()\n    if torch.cuda.is_available():\n        torch.cuda.empty_cache()\n\n    cap = cv2.VideoCapture(str(vp))\n    _fps, _n = cap.get(cv2.CAP_PROP_FPS), int(cap.get(cv2.CAP_PROP_FRAME_COUNT))\n    cap.release()\n    duration = (_n / _fps) if _fps > 0 else 20.0\n\n    frames = sample_frames_stamped(vp, VLM_CFG['whole_fps'], limit=VLM_CFG['whole_limit'],\n                                   max_side=VLM_CFG['image_max_side'], burn=True)\n    if not frames:\n        print(f'[WARNING] No frames extracted for {vp.name} -- skipping')\n        continue\n\n    print('=' * 70)\n    print(f'{vp.name}  |  ground truth: {gt}  |  duration {duration:.1f}s')\n    print(f'{len(frames)} frames @ {frames[0][1].size} '\n          f'(~{frames[0][1].size[0] * frames[0][1].size[1] // 784} visual tokens each, '\n          f'~{len(frames) * frames[0][1].size[0] * frames[0][1].size[1] // 784} total)')\n    print(f'timestamps: {[round(t, 1) for t, _ in frames[:8]]}...')\n\n    raw = vlm_generate_oom_safe(build_stage1_prompt(duration, ''), frames)\n    if not raw:\n        print('>>> Generation failed (OOM even after retries) -- skipping this clip')\n        continue\n\n    print(f'\\n--- RAW ---\\n{raw!r}\\n')\n    parsed = _extract_json(raw)\n    print(f'parsed : {parsed}')\n    print(f'final  : {normalize_prediction(parsed, duration)}')\n    if parsed is None:\n        print('>>> JSON DID NOT PARSE -- every field below is a default, not a prediction')\n\n    gc.collect()\n    if torch.cuda.is_available():\n        torch.cuda.empty_cache()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T13:55:16.115417Z","iopub.execute_input":"2026-08-03T13:55:16.115848Z","iopub.status.idle":"2026-08-03T13:56:38.632378Z","shell.execute_reply.started":"2026-08-03T13:55:16.115820Z","shell.execute_reply":"2026-08-03T13:56:38.631430Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 7d. Qwen3-VL Pipeline on the Calibration Set\n\nTwenty videos, three VLM calls each, scored against the same constant floor as everything else. **This is the number that decides whether the full run is worth 8 hours**, and no version of this notebook has ever produced it: every score in Section 7 so far belongs to the tracking pipeline.\n\nReference points, all on the real test set: the constant floor here is ~0.24, the tracking pipeline ~0.24, a *single* VLM call is 0.42238, and the winning three-stage 32B system is 0.57080.\n\nThree caveats on reading the number:\n\n- **The calibration clips are CARLA, not CCTV.** This measures the pipeline, not the sim-to-real transfer, and the VLM was never tuned on either.\n- **No scene hint fires here.** `scene_layout` exists only for the real test clips, so both the Stage-1 hint (+0.00216) and the t-bone rule (+0.00466) are inactive on synthetic. The real-set score should be slightly *higher* than what this cell reports.\n- **n=20.** One video is 0.05 of C. Read the gap, not the digits.","metadata":{}},{"cell_type":"code","source":"# [EVAL] Qwen3-VL three-stage pipeline on the calibration set\nif not VLM_AVAILABLE:\n    raise RuntimeError('VLM unavailable -- run Section 6.10 first.')\n\nvlm_calib_rows = []\n_t0 = time.time()\nfor _i, vp in enumerate(diverse_videos):\n    vlm_calib_rows.append(run_inference_vlm(vp))\n    _el = time.time() - _t0\n    print(f'[STATUS] {_i+1}/{len(diverse_videos)}: {vp.name} | '\n          f'{vlm_calib_rows[-1][\"accident_time\"]:.2f}s '\n          f'({vlm_calib_rows[-1][\"center_x\"]:.2f}, {vlm_calib_rows[-1][\"center_y\"]:.2f}) '\n          f'{vlm_calib_rows[-1][\"type\"]} | {_el/(_i+1):.1f}s/video')\n    if torch.cuda.is_available():\n        torch.cuda.empty_cache()\n\ncalib_df_vlm = pd.DataFrame(vlm_calib_rows)\neval_vlm = score_predictions(calib_df_vlm, diverse_labels_df)\n\n# Per-video seconds here is the honest input to the full-run projection: 2027\n# clips at this rate is the real cost, and Kaggle sessions cap at 12 h.\n_sec = (time.time() - _t0) / len(diverse_videos)\nprint(f'\\n[TIMING] {_sec:.1f}s/video -> {_sec*len(real_videos)/3600:.1f} h for '\n      f'{len(real_videos)} test clips')\n\nprint(f'\\n[COMPARISON] Calibration set (n={len(eval_vlm)})')\n_cmp = pd.DataFrame({\n    'constant floor': [const_eval['T'].mean(), const_eval['S'].mean(),\n                       const_eval['C'].mean(), accident_score(const_eval)],\n    'tracking (6.2-6.9)': [eval_final['T'].mean(), eval_final['S'].mean(),\n                           eval_final['C'].mean(), accident_score(eval_final)],\n    'Qwen3-VL (6.10)': [eval_vlm['T'].mean(), eval_vlm['S'].mean(),\n                        eval_vlm['C'].mean(), accident_score(eval_vlm)],\n}, index=['T', 'S', 'C', 'ACCIDENT score']).round(4)\ndisplay(_cmp)\n\n_floor = accident_score(const_eval)\n_got = accident_score(eval_vlm)\nprint(f'[VERDICT] Qwen3-VL {_got:.4f} vs constant floor {_floor:.4f} -- '\n      f'{\"BEATS\" if _got > _floor else \"LOSES TO\"} it')\nfor comp in ['T', 'S', 'C']:\n    f, g = const_eval[comp].mean(), eval_vlm[comp].mean()\n    print(f'  {comp}: {g:.4f} vs constant {f:.4f}  '\n          f'{\"BEATS\" if g > f else \"LOSES TO\"}')\n\nprint('\\n[CONFUSION] predicted vs ground-truth type')\ndisplay(pd.crosstab(eval_vlm['type_gt'], eval_vlm['type_pred']))\n\n# A pipeline that has quietly stopped working looks exactly like a constant.\nif calib_df_vlm['type'].nunique() <= 1:\n    print('[WARNING] every clip got the same type -- check the raw VLM output')\nif float(calib_df_vlm[['center_x', 'center_y']].std().mean()) < 1e-6:\n    print('[WARNING] every clip got the same point -- Stage 3 is not returning coordinates')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T13:56:38.633443Z","iopub.execute_input":"2026-08-03T13:56:38.633808Z","iopub.status.idle":"2026-08-03T14:18:34.228101Z","shell.execute_reply.started":"2026-08-03T13:56:38.633765Z","shell.execute_reply":"2026-08-03T14:18:34.227216Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 7c. Negative Result: Optical-Flow Angle Classification (Removed)\n\nAn earlier revision classified collision type from the angle between the two dominant optical-flow direction clusters just before impact (KMeans on high-magnitude flow vectors, angle thresholds per type). It is documented here as a negative result so it is not re-tried:\n\n- On a head-on-only calibration subset it reached **0.80 accuracy** and looked promising.\n- On the diverse 5-type set, the measured angles for `single`, `sideswipe`, and `t-bone` videos clustered at **~150-178 degrees** -- the \"opposite directions\" signature -- because compression noise and background traffic dominate the flow field. The signal does not separate the classes it was meant to separate.\n- Precision of a confident head-on commit was **0.09-0.14** across thresholds 140-175 degrees, and a full grid search over all four thresholds topped out at **0.25 accuracy** -- below the CLIP-only baseline (0.35).\n- Folded into the pipeline, it dragged overall classification from C=0.35 down to C=0.10.\n\n*(The C values quoted above are from earlier runs on different calibration splits; the ordering is the point, not the absolute numbers. Note also that the constant 'rear-end' predictor scores C=0.359 under the natural synthetic prior -- above every zero-shot classifier tried here, including the CLIP baseline.)*\n\nThe general lesson transferred to the current design: **every component must be validated on all five types before being trusted**, and hard rule-based overrides lose to soft probabilistic ensembles when the underlying signal is noisy.\n\nNote that the *tracking-based* geometry of Section 6.3 is not the same idea resurrected: raw optical flow mixed background traffic and compression noise into a single velocity field, whereas track velocities come from the two involved objects only, the signal abstains outside its confident angle bands, is gated on track quality, is cross-checked against an independent temporal anchor, and enters a soft ensemble instead of overriding it.","metadata":{}},{"cell_type":"markdown","source":"---\n## 8. Test Inference and Submission\n\nThe final pipeline runs on the real CCTV test set.\n\n- **Runtime.** The previous design projected ~13.3 h for 2027 videos at ~20.5 s each, over Kaggle's 12 h session limit, because Qwen alone spent ~10 s per video on three prompts. With Qwen, the Perception Encoder, and OWLv2 removed, only YOLO tracking and (optionally) CLIP remain. Re-measure with the `SUBSET_N` dry run rather than trusting this paragraph: the projection printed by the cell below is the number that counts.\n- `SUBSET_N` is the dry-run switch: process the first N videos to measure per-video latency, project the full-run ETA, and surface OOM early — then set it to `None` for the full run.\n- Predictions are **checkpointed** to `preds_checkpoint.csv` every 25 videos, so a session that dies resumes instead of restarting. Delete the checkpoint to force a clean run.\n- **Unmatched rows fall back to `CONST`** (Section 6.7), not to a hardcoded guess.","metadata":{}},{"cell_type":"code","source":"# [SUBMIT] Full test set inference -- final pipeline, checkpointed for 12 h sessions\n# ============================================================\n# SUBSET_N: dry-run switch.\n#   SUBSET_N = 100  -> first 100 videos only: measures s/video, projects the\n#                      full-run ETA, and catches OOM early.\n#   SUBSET_N = None -> full run over all real test videos (the real submission).\n# ============================================================\nSUBSET_N = 10   # <-- 10 first: confirms the VLM runs on real clips and\n                #     measures s/video. Then None for the full run.\n\n# The Qwen3-VL pipeline of Section 6.10. run_inference_final (tracking) is kept\n# only for the Section 7 comparison -- it measures below the constant floor and\n# must not be submitted. Swap this to compare, but read Section 7b first.\nSUBMIT_PIPELINE = run_inference_vlm\n\nif SUBMIT_PIPELINE is run_inference_vlm and not VLM_AVAILABLE:\n    raise RuntimeError('Refusing to run: the VLM is unavailable, so every prediction '\n                       'would be a constant. Fix Section 6.10 first.')\n\nCHECKPOINT_PATH = OUTPUT_DIR / 'preds_checkpoint.csv'\nCHECKPOINT_EVERY = 25\n\nvideos_to_run = real_videos[:SUBSET_N] if SUBSET_N is not None else real_videos\nis_subset_run = SUBSET_N is not None and SUBSET_N < len(real_videos)\n\n# Resume from checkpoint: previously processed videos are skipped, so a run\n# interrupted by the session limit continues in the next session.\ntest_results, done_paths = [], set()\nif CHECKPOINT_PATH.exists():\n    prev = pd.read_csv(CHECKPOINT_PATH)\n    test_results = prev.to_dict('records')\n    done_paths = set(prev['path'])\n    print(f'[STATUS] Checkpoint found: {len(done_paths)} videos already processed -- resuming')\n\nprint(f'[STATUS] Final-pipeline inference on {len(videos_to_run)} video(s) '\n      f'{\"(SUBSET DRY RUN)\" if is_subset_run else \"(FULL RUN)\"} '\n      f'out of {len(real_videos)} total')\n\nper_video_times = []\nrun_start = time.time()\n\nfor i, vp in enumerate(videos_to_run):\n    sub_path = 'videos/' + vp.name  # submission path format\n    if sub_path in done_paths:\n        continue\n\n    t0 = time.time()\n    # sub_path is passed so Stage 1 can look up scene_layout for this clip\n    result = SUBMIT_PIPELINE(vp, sub_path)\n    per_video_times.append(time.time() - t0)\n\n    result['path'] = sub_path\n    test_results.append(result)\n    done_paths.add(sub_path)\n\n    if len(test_results) % CHECKPOINT_EVERY == 0:\n        pd.DataFrame(test_results).to_csv(CHECKPOINT_PATH, index=False)\n\n    if (i + 1) % 10 == 0 or (i + 1) == len(videos_to_run):\n        avg_t   = sum(per_video_times) / len(per_video_times)\n        elapsed = time.time() - run_start\n        remaining = avg_t * max(0, len(videos_to_run) - len(test_results))\n        print(f'[STATUS] {len(test_results)}/{len(videos_to_run)} | avg {avg_t:.2f}s/video | '\n              f'elapsed {elapsed/60:.1f} min | remaining ~{remaining/3600:.2f} h | '\n              f'full-set projection ~{avg_t * len(real_videos)/3600:.2f} h')\n        if avg_t * len(real_videos) > 11 * 3600:\n            print('[WARNING] Full-set projection exceeds one 12 h Kaggle session -- '\n                  'the checkpoint lets you split the run across sessions safely')\n\n    if torch.cuda.is_available():\n        torch.cuda.empty_cache()\n\npd.DataFrame(test_results).to_csv(CHECKPOINT_PATH, index=False)\n\ntest_preds_df = pd.DataFrame(test_results)\ntest_preds_df = test_preds_df[test_preds_df['path'].isin(\n    {'videos/' + vp.name for vp in videos_to_run})].reset_index(drop=True)\nassert len(test_preds_df) == len(videos_to_run), 'Some videos were dropped during inference'\nprint(f'[SUCCESS] Inference complete on {len(test_preds_df)} video(s) '\n      f'(this session: {len(per_video_times)}, {(time.time() - run_start)/60:.1f} min)')\nif 'track_used' in test_preds_df.columns:\n    print(f\"[STATUS] Tracking trusted on {test_preds_df['track_used'].mean():.0%} of test videos \"\n          '-- a sharp drop vs calibration signals a synthetic-to-real d","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T14:18:34.229198Z","iopub.execute_input":"2026-08-03T14:18:34.229525Z","iopub.status.idle":"2026-08-03T14:26:52.535988Z","shell.execute_reply.started":"2026-08-03T14:18:34.229499Z","shell.execute_reply":"2026-08-03T14:26:52.535330Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [SUBMIT] Align predictions with the sample submission schema and save\nif is_subset_run:\n    # Dry run: save under a distinct name so it is never confused with the real submission\n    submission_df = test_preds_df[['path', 'accident_time', 'center_x', 'center_y', 'type']].copy()\n    SUBMISSION_PATH = OUTPUT_DIR / f'submission_dryrun_{len(videos_to_run)}.csv'\n    print(f'[STATUS] SUBSET DRY RUN -- {SUBMISSION_PATH.name} is for validation only. '\n          'Set SUBSET_N = None and re-run for the real submission.')\nelif not sample_sub.empty:\n    submission_df = sample_sub[['path']].merge(\n        test_preds_df[['path', 'accident_time', 'center_x', 'center_y', 'type']],\n        on='path', how='left',\n    )\n    # Unmatched rows fall back to CONST -- the constants fitted on the synthetic\n    # training set in Section 6.7. The previous values here (median of our own\n    # predictions, 0.5/0.5, hardcoded 'rear-end') were guesses that happened to\n    # land near the fitted optimum; there is no reason to guess when the prior\n    # has been measured.\n    n_unmatched = submission_df['type'].isnull().sum()\n    if n_unmatched > 0:\n        print(f'[WARNING] {n_unmatched} submission rows had no matching prediction '\n              '-- filling with the fitted constant prior')\n    submission_df = submission_df.fillna({\n        'accident_time': CONST['accident_time'],\n        'center_x': CONST['center_x'],\n        'center_y': CONST['center_y'],\n        'type': CONST['type'],\n    })\n    assert len(submission_df) == len(sample_sub), 'Submission row count mismatch'\n    SUBMISSION_PATH = OUTPUT_DIR / 'submission.csv'\nelse:\n    submission_df = test_preds_df[['path', 'accident_time', 'center_x', 'center_y', 'type']].copy()\n    SUBMISSION_PATH = OUTPUT_DIR / 'submission.csv'\n\nsubmission_df.to_csv(SUBMISSION_PATH, index=False)\n# Degenerate-output check. A pipeline that has silently stopped working produces\n# a submission with almost no variety -- one type, one point, one time. That is\n# indistinguishable from a valid file until the leaderboard says nothing moved.\n_n_types = submission_df['type'].nunique()\n_xy_std = float(submission_df[['center_x', 'center_y']].std().mean())\n_t_std = float(submission_df['accident_time'].std())\nprint(f'[CHECK] distinct types={_n_types} | xy std={_xy_std:.4f} | time std={_t_std:.3f}')\nif _n_types <= 1 or _xy_std < 1e-6:\n    print('[WARNING] This submission is (near) constant -- one collision type and/or a '\n          'single impact point for every clip. That is what a dead model looks like. '\n          'Check the VLM canary in Section 6.10 before submitting.')\n\nprint(f'[SUCCESS] Saved: {SUBMISSION_PATH} ({len(submission_df)} rows)')\ndisplay(submission_df.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T14:26:52.537063Z","iopub.execute_input":"2026-08-03T14:26:52.537388Z","iopub.status.idle":"2026-08-03T14:26:52.554687Z","shell.execute_reply.started":"2026-08-03T14:26:52.537361Z","shell.execute_reply":"2026-08-03T14:26:52.553765Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [SUBMIT] Predicted collision type distribution\nif not submission_df.empty and 'type' in submission_df.columns:\n    pred_type_dist = submission_df['type'].value_counts().reset_index()\n    pred_type_dist.columns = ['collision_type', 'count']\n\n    fig, ax = plt.subplots(figsize=(10, 5))\n    sns.barplot(data=pred_type_dist, x='collision_type', y='count', palette='Set2', ax=ax)\n    ax.set_title('Predicted Collision Type Distribution (Test Set)')\n    ax.set_xlabel('Predicted Collision Type')\n    ax.set_ylabel('Count')\n    ax.bar_label(ax.containers[0], fontsize=9)\n    plt.tight_layout()\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T14:26:52.555802Z","iopub.execute_input":"2026-08-03T14:26:52.556095Z","iopub.status.idle":"2026-08-03T14:26:52.725407Z","shell.execute_reply.started":"2026-08-03T14:26:52.556070Z","shell.execute_reply":"2026-08-03T14:26:52.724461Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [SUBMIT] Predicted accident time distribution\nif not submission_df.empty and 'accident_time' in submission_df.columns:\n    fig, ax = plt.subplots(figsize=(10, 5))\n    sns.histplot(submission_df['accident_time'], bins=30, kde=True,\n                 color=PALETTE['primary'], ax=ax)\n    ax.set_title('Predicted Accident Time Distribution (Test Set)')\n    ax.set_xlabel('Predicted Accident Time (seconds)')\n    ax.set_ylabel('Count')\n    plt.tight_layout()\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T14:26:52.726573Z","iopub.execute_input":"2026-08-03T14:26:52.726949Z","iopub.status.idle":"2026-08-03T14:26:52.967705Z","shell.execute_reply.started":"2026-08-03T14:26:52.726920Z","shell.execute_reply":"2026-08-03T14:26:52.966933Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# [SUBMIT] Predicted impact location scatter\nif not submission_df.empty and all(c in submission_df.columns for c in ['center_x', 'center_y']):\n    fig, ax = plt.subplots(figsize=(8, 6))\n    scatter = ax.scatter(submission_df['center_x'], submission_df['center_y'],\n                         alpha=0.4, s=25, c=submission_df['center_y'], cmap='viridis')\n    plt.colorbar(scatter, ax=ax, label='center_y')\n    ax.set_title('Predicted Impact Location Distribution (Test Set)')\n    ax.set_xlabel('center_x (0=left, 1=right)')\n    ax.set_ylabel('center_y (0=top, 1=bottom)')\n    ax.set_xlim(0, 1)\n    ax.set_ylim(0, 1)\n    ax.invert_yaxis()\n    plt.tight_layout()\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T14:26:52.968684Z","iopub.execute_input":"2026-08-03T14:26:52.968953Z","iopub.status.idle":"2026-08-03T14:26:53.172902Z","shell.execute_reply.started":"2026-08-03T14:26:52.968927Z","shell.execute_reply":"2026-08-03T14:26:53.172128Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"---\n## 9. Conclusion\n\n### What was wrong\n\n- **The metric was implemented incorrectly, and that error drove the design.** `score_predictions` computed a harmonic mean per video and averaged it. The competition averages T, S, and C across videos *first*, then takes the harmonic mean of those three numbers. Because C is 0/1, the per-video variant zeroes every misclassified video and is bounded above by accuracy. It reported 0.087 for the baseline (real value **0.2321**) and 0.0000 for the final pipeline (real value **0.1739**).\n- **Which made classification look decisive when it is the weakest lever.** The claim \"the harmonic mean is 0 whenever the type is wrong\" is false. On the baseline, raising C from 0.30 to 0.45 is worth **+0.022**; raising S from 0.156 to 0.45 is worth **+0.111**. Nearly all the modelling effort — a three-signal ensemble, prompt engineering, a scene-plausibility rule — went into the component with one fifth the leverage.\n- **The pipeline scored below a constant on every component.** Because the scoring sigma is comparable to the spread of the labels themselves, predicting the dataset's central constant is strong: T=0.518 against the frame-difference anchor's 0.314, S≈0.295 against OWLv2/optical-flow's 0.156, C=0.359 against CLIP's 0.300 (and the final ensemble's 0.150, below the 0.20 chance level). Estimated constant-only score ≈ **0.370**, against 0.232 for the baseline. Five models and ~11.5 h of T4 produced less than three numbers.\n- **Fallbacks made it worse, not better.** Every stage fell back from a weak model to a weaker one. A fallback must be the prior.\n- **The ensemble collapsed for a structural reason.** Qwen's three prompts are quantised to {0, ⅓, ⅔, 1}, so a unanimous vote contributed the full 0.35, while CLIP's probability mass spread over five classes gave its argmax ≈0.21. Qwen won whenever it was unanimous — which was almost always, because the three prompts ask the same question of six 448px stills in which the vehicles span a few dozen pixels, so it fell back on a language prior where `rear-end` is the most common accident. Result: `rear-end` on 18/20 videos.\n- **Two tracking bugs.** NMS at the default `iou=0.7` deleted one of the two colliding vehicles as a duplicate at the exact moment of impact. `np.convolve(mode='same')` zero-padded the trajectory ends, fabricating a deceleration of up to 680 px/s precisely where `_impact_time` searches for one.\n\n### What changed\n\nMetric corrected. Constant prior fitted on the synthetic training set and installed as the fallback for every stage. NMS `iou=0.9`; two-stage confidence-split association so tracks survive the confidence collapse at impact; Savitzky-Golay smoothing. Tracking gated on its own kinetic evidence rather than on agreement with an anchor weaker than a constant. Qwen, the Perception Encoder, and OWLv2 removed. A supervised classifier over domain-invariant track kinematics, trained on the 2211 labelled synthetic videos the organizers provide for pretraining — previously used only for EDA.\n\n### What is still unmeasured\n\nThe numbers above are from the calibration run of the previous design, plus analytic estimates from the label statistics. Pending:\n\n1. **Section 6.8** — does tracking fire on real CCTV at all? Every downstream claim depends on it and none of it has been tested outside CARLA.\n2. **Section 7b** — per-signal ablation, and the `TRACK_EVIDENCE_MIN` sweep.\n3. **Section 6.9** — grouped-by-map CV accuracy, and the coverage-weighted C that follows.\n4. **The scoring sigmas are assumptions.** The Evaluation page says only \"Gaussian-style\". At σ_S=0.05 a constant scores 0.095; at σ_S=0.20, 0.622. The priority order between S and C inverts across that range.\n\n### Method notes\n\n- **Measure against the prior, not against zero.** Every component here beat \"doing nothing\" and lost to \"predicting the average\". Without a floor, a pipeline can be elaborate, expensive, and worse than a constant while every individual piece looks like it is contributing.\n- **Validate on the axis that will shift.** Section 7b's earlier failure (0.80 on a head-on-only subset, 0.10 on a diverse one) and the `map`-grouped CV here are the same lesson at different scales: a split that does not vary what the test set varies reports a number that means nothing.\n- **Prefer physics to pixels across a domain gap.** Angles and ratios transfer from CARLA to CCTV; render statistics do not.","metadata":{}},{"cell_type":"markdown","source":"---\n## 10. References\n\n1. **Competition:** [ACCIDENT @ CVPR 2026 -- Kaggle](https://kaggle.com/competitions/accident)\n2. **Workshop:** [AUTOPILOT @ CVPR 2026](https://autopilot-workshop.github.io)\n3. **SORT:** Bewley, A. et al. \"Simple Online and Realtime Tracking.\" ICIP 2016. [arXiv:1602.00763](https://arxiv.org/abs/1602.00763)\n4. **Hungarian algorithm:** Kuhn, H. W. \"The Hungarian Method for the Assignment Problem.\" Naval Research Logistics Quarterly, 1955.\n5. **YOLOv8:** Jocher, G. et al. Ultralytics YOLOv8. [https://github.com/ultralytics/ultralytics](https://github.com/ultralytics/ultralytics)\n6. **CLIP:** Radford, A. et al. \"Learning Transferable Visual Models From Natural Language Supervision.\" ICML 2021. [arXiv:2103.00020](https://arxiv.org/abs/2103.00020)\n7. **Perception Encoder:** Bolya, D. et al. \"Perception Encoder: The best visual embeddings are not at the output of the network.\" 2025. [arXiv:2504.13181](https://arxiv.org/abs/2504.13181)\n8. **Qwen2.5-VL:** Bai, S. et al. \"Qwen2.5-VL Technical Report.\" 2025. [arXiv:2502.13923](https://arxiv.org/abs/2502.13923)\n9. **OWLv2:** Minderer, M. et al. \"Scaling Open-Vocabulary Object Detection.\" NeurIPS 2023. [arXiv:2306.09683](https://arxiv.org/abs/2306.09683)\n10. **Optical Flow (Farneback):** Farneback, G. \"Two-Frame Motion Estimation Based on Polynomial Expansion.\" SCIA 2003. [Springer](https://link.springer.com/chapter/10.1007/3-540-45103-X_50)\n11. **CARLA Simulator:** Dosovitskiy, A. et al. \"CARLA: An Open Urban Driving Simulator.\" CoRL 2017. [arXiv:1711.03938](https://arxiv.org/abs/1711.03938)\n12. **OpenCV Documentation:** [https://docs.opencv.org](https://docs.opencv.org)\n13. **Traffic accident video datasets survey:** Bouzid, Y. et al. \"Anomaly Detection in Traffic Surveillance Videos: A Survey.\" Sensors 2022.","metadata":{}},{"cell_type":"markdown","source":"### 7d-ter. Ablation rieng cho 3 lop temporal cua VLM (chua tung do)\n\n`run_inference_vlm` xep chong 3 lop cho `accident_time`: (1) `stage1_full_scan`\ntho, (2) neo boi `predict_accident_time_ensemble` (co dien, +-3s), (3)\n`stage2_time_refine` (dense window, blend=0.35, cap=1.5s). Muc 7b chi ablate\ntin hieu co dien, chua bao gio tach rieng 3 lop nay -- cell duoi day do T o\ntung lop, tren dung 20 video calibration, KHONG sua ham goc nao.","metadata":{}},{"cell_type":"code","source":"# [DIAG] Do T rieng biet o tung lop cua VLM temporal cascade -- khong sua\n# stage1_full_scan / predict_accident_time_ensemble / stage2_time_refine,\n# chi goi lai chung theo dung thu tu run_inference_vlm dang dung (cell 78),\n# nhung luu lai gia tri TRUNG GIAN sau moi lop de so sanh T rieng.\nif not VLM_AVAILABLE:\n    raise RuntimeError('VLM unavailable -- run Section 6.10 first.')\n\n_TEMPORAL_ANCHOR_DELTA_MAX = 3.0  # phai khop dung TEMPORAL_ANCHOR_DELTA_MAX o cell 78\n\nrows_layered = []\nfor vp, gt_t in zip(diverse_videos, diverse_labels_df['accident_time']):\n    cap = cv2.VideoCapture(str(vp))\n    fps, n = cap.get(cv2.CAP_PROP_FPS), int(cap.get(cv2.CAP_PROP_FRAME_COUNT))\n    cap.release()\n    duration = (n / fps) if fps > 0 else 20.0\n\n    sub_path = 'videos/' + vp.name\n    scene = SCENE_BY_PATH.get(sub_path)\n\n    # Lop 1: stage1 tho, khong neo, khong refine\n    pred1 = stage1_full_scan(vp, duration, scene)\n    t_stage1 = pred1['accident_time']\n\n    # Lop 2: + neo classical (dung dung cong thuc trong run_inference_vlm)\n    t_classical = predict_accident_time_ensemble(vp)\n    correction = float(np.clip(t_stage1 - t_classical, -_TEMPORAL_ANCHOR_DELTA_MAX,\n                               _TEMPORAL_ANCHOR_DELTA_MAX))\n    t_anchored = float(np.clip(t_classical + correction, 0.0, duration))\n\n    # Lop 3: + stage2 dense refine\n    t_refined = stage2_time_refine(vp, t_anchored, duration)\n\n    rows_layered.append({\n        'video': vp.name, 'gt': gt_t,\n        't_stage1_raw': t_stage1, 't_classical': t_classical,\n        't_anchored': t_anchored, 't_refined_final': t_refined,\n    })\n    print(f\"[STATUS] {vp.name[:35]:35s} gt={gt_t:6.2f} | \"\n          f\"stage1={t_stage1:6.2f} classical={t_classical:6.2f} \"\n          f\"anchored={t_anchored:6.2f} refined={t_refined:6.2f}\")\n\nlayered_df = pd.DataFrame(rows_layered)\n\nprint(\"\\n[VERDICT] T rieng tung lop (Gaussian temporal_score trung binh, khong qua accident_score):\")\nfor col in ['t_stage1_raw', 't_classical', 't_anchored', 't_refined_final']:\n    t_vals = [temporal_score(p, g) for p, g in zip(layered_df[col], layered_df['gt'])]\n    print(f\"  {col:18s}: T = {np.mean(t_vals):.4f}\")\n\nprint(\"\\n[SO SANH] Neu 't_anchored' THUA 't_stage1_raw' -- neo classical dang keo tut,\")\nprint(\"nen bo hoac giam TEMPORAL_ANCHOR_DELTA_MAX. Neu 't_refined_final' THUA 't_anchored'\")\nprint(\"-- stage2_time_refine dang keo tut, nen kiem tra time_refine_window co du rong khong\")\nprint(\"(cua so +-2s co chua ca t_true khong) hoac giam time_refine_blend/max_shift qua chat.\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T14:26:53.173924Z","iopub.execute_input":"2026-08-03T14:26:53.174377Z","iopub.status.idle":"2026-08-03T14:42:58.822185Z","shell.execute_reply.started":"2026-08-03T14:26:53.174345Z","shell.execute_reply":"2026-08-03T14:42:58.821405Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 7d-quater. Thu nghiem NumPro-style: danh so khung hinh thay vi burn giay thap phan\n\nDua tren \"Number it: Temporal Grounding Videos like Flipping Manga\" (CVPR 2025,\narXiv:2411.10332, github.com/yongliang-wu/NumPro) -- doi burn-in tu `t=8.55s`\n(giay lien tuc) sang so thu tu khung (\"Frame 7\") va hoi model chon SO KHUNG thay\nvi hoi truc tiep so giay. Khong sua ham goc (`sample_frames_stamped`,\n`build_stage1_prompt`) -- viet ham song song de A/B tren dung 20 video\ncalibration, so voi t_stage1_raw hien tai (T=0.2386) truoc khi quyet dinh thay\nthe ban chinh.","metadata":{}},{"cell_type":"code","source":"# [DIAG] NumPro-style: burn frame-index thay vi giay, hoi model chon so khung\ndef sample_frames_numbered(video_path, target_fps, t_start=None, t_end=None,\n                           limit=32, max_side=768):\n    \"\"\"Nhu sample_frames_stamped nhung burn SO THU TU KHUNG (1, 2, 3, ...) thay\n    vi t=xx.xxs. Tra ve ([times_seconds], [PIL.Image]) -- times[i] la giay that\n    cua khung so (i+1), dung de quy doi nguoc sau khi model tra loi so khung.\"\"\"\n    cap = cv2.VideoCapture(str(video_path))\n    fps = cap.get(cv2.CAP_PROP_FPS)\n    n = int(cap.get(cv2.CAP_PROP_FRAME_COUNT))\n    if fps <= 0 or n <= 0:\n        cap.release()\n        return [], []\n    duration = n / fps\n    t0 = 0.0 if t_start is None else max(0.0, t_start)\n    t1 = duration if t_end is None else min(duration, t_end)\n    if t1 <= t0:\n        t1 = min(duration, t0 + 1.0)\n\n    times = np.arange(t0, t1 + 1e-6, 1.0 / max(target_fps, 0.1))\n    if len(times) > limit:\n        times = times[np.linspace(0, len(times) - 1, limit).round().astype(int)]\n\n    imgs, kept_times = [], []\n    for idx, t in enumerate(times, start=1):\n        cap.set(cv2.CAP_PROP_POS_FRAMES, int(np.clip(round(t * fps), 0, n - 1)))\n        ok, frame = cap.read()\n        if not ok:\n            continue\n        h, w = frame.shape[:2]\n        if max(h, w) > max_side:\n            sc = max_side / float(max(h, w))\n            frame = cv2.resize(frame, (int(round(w * sc)), int(round(h * sc))),\n                               interpolation=cv2.INTER_AREA)\n        # NumPro: so lon, ro, vi tri co dinh -- nhu danh so trang manga\n        cv2.rectangle(frame, (8, 8), (90, 56), (0, 0, 0), -1)\n        cv2.putText(frame, str(idx), (16, 46), cv2.FONT_HERSHEY_SIMPLEX,\n                    1.4, (255, 255, 255), 3, cv2.LINE_AA)\n        imgs.append(PILImage.fromarray(cv2.cvtColor(frame, cv2.COLOR_BGR2RGB)))\n        kept_times.append(float(t))\n    cap.release()\n    return kept_times, imgs\n\n\nNUMBERED_TIME_PROMPT_TEMPLATE = (\n    \"These {n} sequential CCTV frames are numbered 1 to {n}, shown in order like \"\n    \"manga panels. Find the FRAME NUMBER where vehicles first make contact, or a \"\n    \"vehicle first hits an object. Do not choose the peak impact or aftermath -- \"\n    \"the first moment of contact only.\\n\\n\"\n    'Output ONLY this JSON: {{\"collision_frame\": <integer 1-{n}>}}'\n)\n\n\ndef stage1_numbered_time(video_path, duration):\n    times, imgs = sample_frames_numbered(video_path, 2.0, 0.0, duration, limit=32)\n    if not imgs:\n        return duration * 0.35\n    prompt = NUMBERED_TIME_PROMPT_TEMPLATE.format(n=len(imgs))\n    raw = vlm_generate_oom_safe(prompt, list(zip(times, imgs)), max_new_tokens=32, min_frames=4)\n    parsed = _extract_json(raw)\n    if not parsed or 'collision_frame' not in parsed:\n        return duration * 0.35\n    idx = int(np.clip(_safe_float(parsed['collision_frame'], 1), 1, len(times)))\n    return times[idx - 1]\n\n\nrows_numpro = []\nfor vp, gt_t in zip(diverse_videos, diverse_labels_df['accident_time']):\n    cap = cv2.VideoCapture(str(vp))\n    fps, n = cap.get(cv2.CAP_PROP_FPS), int(cap.get(cv2.CAP_PROP_FRAME_COUNT))\n    cap.release()\n    duration = (n / fps) if fps > 0 else 20.0\n\n    t_numpro = stage1_numbered_time(vp, duration)\n    rows_numpro.append({'video': vp.name, 'gt': gt_t, 't_numpro': t_numpro})\n    print(f\"[STATUS] {vp.name[:35]:35s} gt={gt_t:6.2f} numpro={t_numpro:6.2f}\")\n\nnumpro_df = pd.DataFrame(rows_numpro)\nt_scores = [temporal_score(p, g) for p, g in zip(numpro_df['t_numpro'], numpro_df['gt'])]\nprint(f\"\\n[VERDICT] NumPro-style (frame-number burn-in) T = {np.mean(t_scores):.4f}\")\nprint(f\"          t_stage1_raw hien tai (giay lien tuc)     T = 0.2386\")\nprint(f\"          constant                                   T = 0.3800\")\nprint(\"\\n[QUYET DINH] Neu T o day > 0.2386 ro ret -- doi han sang burn frame-so,\")\nprint(\"giu nguyen JSON schema tra ve nhung doi 'accident_time' thanh 'collision_frame'\")\nprint(\"+ 1 buoc quy doi, roi thu lai coarse-to-fine (stage2_time_refine) TREN BIEU\")\nprint(\"DIEN MOI nay -- co the refine se het net-negative khi hoat dong tren so khung\")\nprint(\"thay vi giay thap phan.\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T14:42:58.823144Z","iopub.execute_input":"2026-08-03T14:42:58.823434Z","iopub.status.idle":"2026-08-03T14:55:19.138768Z","shell.execute_reply.started":"2026-08-03T14:42:58.823408Z","shell.execute_reply":"2026-08-03T14:55:19.137855Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 8. Chot pipeline de nop bai (uu tien thoi gian, khong nghien cuu them)\n\nAp dung 3 ket qua da co bang chung THAT (khong phai suy doan), CHI THEM cell moi\n-- khong xoa/sua cell nao khac, an toan tuyet doi voi phan con lai cua notebook:\n\n1. **NumPro frame-numbering** cho Stage 1 (da do: T 0.2386 -> 0.2929 tren chinh\n   20 video calibration, du dang bi OOM cat giam frame).\n2. **Bo neo classical + stage2_time_refine** (da do ca hai deu net-negative:\n   0.2386 (tho) -> 0.2314 (neo) -> 0.2262 (refine) -- moi lop sua deu keo tut).\n3. **Ha `whole_limit` 32 -> 16** -- log thuc te cho thay hau het video da tu OOM\n   ve ~16 frame sau 1 lan retry; dat thang 16 tranh 1 lan goi model bi lang phi\n   moi video, cong don tren 2027 video that co the tiet kiem dang ke thoi gian.","metadata":{}},{"cell_type":"code","source":"# [PATCH] Ghi de stage1_full_scan + run_inference_vlm -- KHONG xoa dinh nghia cu,\n# Python/Jupyter don gian dung ban MOI NHAT khi cell nay chay sau cell 73/77.\n# An toan tuyet doi: khong dung cell nao, khong xoa ten nao, chi ghi de 2 ham.\n\nVLM_CFG['whole_limit'] = 16   # tu 32 -- giam so lan OOM-retry lang phi tren 2027 video that\n\n\ndef sample_frames_numbered(video_path, target_fps, t_start=None, t_end=None,\n                           limit=32, max_side=768):\n    \"\"\"Nhu sample_frames_stamped nhung burn SO THU TU KHUNG (NumPro, CVPR 2025)\n    thay vi t=xx.xxs. Tra ve (times_seconds, PIL.Images) -- times[i] la giay\n    that cua khung so (i+1).\"\"\"\n    cap = cv2.VideoCapture(str(video_path))\n    fps = cap.get(cv2.CAP_PROP_FPS)\n    n = int(cap.get(cv2.CAP_PROP_FRAME_COUNT))\n    if fps <= 0 or n <= 0:\n        cap.release()\n        return [], []\n    duration = n / fps\n    t0 = 0.0 if t_start is None else max(0.0, t_start)\n    t1 = duration if t_end is None else min(duration, t_end)\n    if t1 <= t0:\n        t1 = min(duration, t0 + 1.0)\n\n    times = np.arange(t0, t1 + 1e-6, 1.0 / max(target_fps, 0.1))\n    if len(times) > limit:\n        times = times[np.linspace(0, len(times) - 1, limit).round().astype(int)]\n\n    imgs, kept_times = [], []\n    for idx, t in enumerate(times, start=1):\n        cap.set(cv2.CAP_PROP_POS_FRAMES, int(np.clip(round(t * fps), 0, n - 1)))\n        ok, frame = cap.read()\n        if not ok:\n            continue\n        h, w = frame.shape[:2]\n        if max(h, w) > max_side:\n            sc = max_side / float(max(h, w))\n            frame = cv2.resize(frame, (int(round(w * sc)), int(round(h * sc))),\n                               interpolation=cv2.INTER_AREA)\n        cv2.rectangle(frame, (8, 8), (90, 56), (0, 0, 0), -1)\n        cv2.putText(frame, str(idx), (16, 46), cv2.FONT_HERSHEY_SIMPLEX,\n                    1.4, (255, 255, 255), 3, cv2.LINE_AA)\n        imgs.append(PILImage.fromarray(cv2.cvtColor(frame, cv2.COLOR_BGR2RGB)))\n        kept_times.append(float(t))\n    cap.release()\n    return kept_times, imgs\n\n\ndef build_stage1_prompt_numbered(n_frames: int, scene_hint: str) -> str:\n    return (\n        f'These {n_frames} sequential CCTV frames show a traffic accident, '\n        f'numbered 1 to {n_frames} like manga panels shown in order.\\n\\n'\n        'Your task: Find the FRAME NUMBER, LOCATION, and TYPE of the collision.\\n\\n'\n        '1) COLLISION FRAME -- Find the frame number where vehicles first make '\n        'contact or a vehicle first hits an object. Report the frame number '\n        '(an integer), NOT a time in seconds.\\n\\n'\n        '2) COLLISION POSITION -- Where in the frame does the impact happen? '\n        'Report as normalized coordinates: center_x (0=left, 1=right), '\n        'center_y (0=top, 1=bottom).\\n\\n'\n        \"3) ACCIDENT TYPE -- Classify the collision type by watching the vehicles' \"\n        f'approach angles and movement. {scene_hint}\\n\\n'\n        'Output ONLY this JSON: {\"collision_frame\": <int>, \"center_x\": <float>, '\n        '\"center_y\": <float>, \"type\": \"<rear-end|head-on|sideswipe|t-bone|single>\"}'\n    )\n\n\ndef stage1_full_scan(video_path, duration, scene_layout=None):\n    \"\"\"NumPro version: burns frame numbers instead of timestamps, asks the model\n    for a frame index instead of a continuous second value (measured +0.054 T on\n    the calibration set vs the timestamp version, even while OOM-degraded).\"\"\"\n    hint = SCENE_TYPE_HINTS.get(scene_layout or '', '')\n    limit = VLM_CFG['whole_limit']\n    while True:\n        times, imgs = sample_frames_numbered(video_path, VLM_CFG['whole_fps'],\n                                             0.0, duration, limit=limit,\n                                             max_side=VLM_CFG['image_max_side'])\n        if not imgs:\n            return {**normalize_prediction({}, duration), '_parsed': False}\n        try:\n            raw = vlm_generate(build_stage1_prompt_numbered(len(imgs), hint), imgs)\n            break\n        except torch.cuda.OutOfMemoryError:\n            torch.cuda.empty_cache()\n            if limit <= 6:\n                raise\n            limit //= 2\n            print(f'[WARNING] OOM on {video_path.name} -- retrying with {limit} frames')\n    parsed = _extract_json(raw)\n    out = normalize_prediction(parsed or {}, duration)\n    if parsed and 'collision_frame' in parsed:\n        idx = int(np.clip(_safe_float(parsed['collision_frame'], 1), 1, len(times)))\n        out['accident_time'] = times[idx - 1]\n    out['_parsed'] = parsed is not None\n    return out\n\n\nprint('[STATUS] stage1_full_scan patched (NumPro frame-numbering), whole_limit=16 -- run_inference_vlm redefinition below in Section 10 is what actually takes effect; this cell no longer redefines it (dead code removed).')\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T14:55:19.139770Z","iopub.execute_input":"2026-08-03T14:55:19.140031Z","iopub.status.idle":"2026-08-03T14:55:19.154518Z","shell.execute_reply.started":"2026-08-03T14:55:19.139992Z","shell.execute_reply":"2026-08-03T14:55:19.153599Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 9. Thu lai coarse-to-fine (stage2_time_refine) TREN bieu dien frame-so\n\n`stage2_time_refine` ban giay lien tuc da do la net-negative (0.2314 -> 0.2262).\nGia thuyet: no hai vi hoat dong tren giay thap phan (nhieu), khong phai vi\ncoarse-to-fine sai nguyen tac (ReVisionLLM/SlowFocus, da doc o 7d-quater, xac\nnhan day la kien truc chuan). Ban duoi day lam lai dung y tuong nhung hoi FRAME\nSO trong cua so dense, khong hoi giay -- test rieng truoc khi dua vao pipeline.","metadata":{}},{"cell_type":"code","source":"# [DIAG] Coarse-to-fine tren bieu dien frame-so -- test truoc khi thay the\ndef stage2_time_refine_numbered(video_path, t_base, duration):\n    \"\"\"Nhu stage2_time_refine nhung dense window + context deu burn SO KHUNG,\n    hoi model chon SO KHUNG trong cua so hep, khong hoi giay truc tiep.\"\"\"\n    if not VLM_CFG['time_refine']:\n        return t_base\n    w = VLM_CFG['time_refine_window']\n    t0 = max(0.0, t_base - VLM_CFG['time_refine_context_before'])\n    t1 = min(duration, t_base + VLM_CFG['time_refine_context_after'])\n    # Sample toan bo cua so (context + dense) VOI MOT LAN danh so lien tuc, de\n    # frame-number la duy nhat va anh xa ro rang ve giay -- khong tach 2 lan\n    # sample rieng nhu ban goc (se bi trung so neu ghep sau).\n    times, imgs = sample_frames_numbered(video_path, VLM_CFG['time_refine_fps'],\n                                         t0, t1, limit=VLM_CFG['time_refine_limit']\n                                         + VLM_CFG['time_refine_context_limit'],\n                                         max_side=VLM_CFG['image_max_side'])\n    if not imgs:\n        return t_base\n\n    prompt = (\n        f\"These {len(imgs)} sequential CCTV frames are numbered 1 to {len(imgs)}, \"\n        \"shown in order like manga panels, zoomed into the moments around a \"\n        \"suspected collision. Find the FRAME NUMBER where vehicles first make \"\n        \"contact, or a vehicle first hits an object.\\n\\n\"\n        f'Output ONLY this JSON: {{\"collision_frame\": <integer 1-{len(imgs)}>}}'\n    )\n    j = _extract_json(vlm_generate_oom_safe(prompt, list(zip(times, imgs)), 32))\n    if not j or 'collision_frame' not in j:\n        return t_base\n    idx = int(np.clip(_safe_float(j['collision_frame'], 1), 1, len(times)))\n    t_ref = times[idx - 1]\n\n    delta = np.clip(t_ref - t_base, -VLM_CFG['time_refine_max_shift_sec'],\n                    VLM_CFG['time_refine_max_shift_sec'])\n    return float(np.clip(t_base + VLM_CFG['time_refine_blend'] * delta, 0.0, duration))\n\n\nrows_refine_numbered = []\nfor vp, gt_t in zip(diverse_videos, diverse_labels_df['accident_time']):\n    cap = cv2.VideoCapture(str(vp))\n    fps, n = cap.get(cv2.CAP_PROP_FPS), int(cap.get(cv2.CAP_PROP_FRAME_COUNT))\n    cap.release()\n    duration = (n / fps) if fps > 0 else 20.0\n\n    sub_path = 'videos/' + vp.name\n    scene = SCENE_BY_PATH.get(sub_path)\n\n    pred = stage1_full_scan(vp, duration, scene)   # da la ban NumPro\n    t_stage1 = pred['accident_time']\n    t_refined = stage2_time_refine_numbered(vp, t_stage1, duration)\n\n    rows_refine_numbered.append({'video': vp.name, 'gt': gt_t,\n                                 't_stage1': t_stage1, 't_refined': t_refined})\n    print(f\"[STATUS] {vp.name[:35]:35s} gt={gt_t:6.2f} \"\n          f\"stage1={t_stage1:6.2f} refined={t_refined:6.2f}\")\n\nrefine_df = pd.DataFrame(rows_refine_numbered)\nt_s1 = np.mean([temporal_score(p, g) for p, g in zip(refine_df['t_stage1'], refine_df['gt'])])\nt_rf = np.mean([temporal_score(p, g) for p, g in zip(refine_df['t_refined'], refine_df['gt'])])\nprint(f\"\\n[VERDICT] Stage1 (NumPro, khong refine)      T = {t_s1:.4f}\")\nprint(f\"          + coarse-to-fine tren frame-so       T = {t_rf:.4f}\")\nprint(f\"\\n[QUYET DINH] Neu t_rf > t_s1 -- refine da het net-negative, dua vao\")\nprint(f\"run_inference_vlm (goi sau stage1_full_scan, truoc stage3_grounding).\")\nprint(f\"Neu van <= t_s1 -- bo han y tuong refine o day, giu ban da chot lan truoc\")\nprint(f\"(chi Stage1 NumPro, khong refine, nhu cell muc 8 da lam).\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T14:55:19.155828Z","iopub.execute_input":"2026-08-03T14:55:19.156125Z","iopub.status.idle":"2026-08-03T15:04:32.522390Z","shell.execute_reply.started":"2026-08-03T14:55:19.156100Z","shell.execute_reply":"2026-08-03T15:04:32.521529Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 10. Ghep coarse-to-fine (frame-so) vao pipeline chinh -- ket qua tot nhat tu truoc den gio\n\nDo duoc: Stage1 NumPro don le T=0.4035, + refine frame-so T=0.4385 -- **THANG\nconstant (0.38) lan dau tien trong toan bo du an**. Ghi de `run_inference_vlm`\nlan nua (van chi THEM cell, khong xoa gi) de goi `stage2_time_refine_numbered`\nsau Stage 1, truoc Stage 3 grounding.","metadata":{}},{"cell_type":"code","source":"# [PATCH v2] Ghi de run_inference_vlm lan nua -- them stage2_time_refine_numbered\n# (da do T=0.4385, thang ca stage1 rieng le lan constant). Day la ban CHOT CUOI CUNG.\n\ndef run_inference_vlm(video_path: pathlib.Path, sub_path: str = None) -> dict:\n    \"\"\"Stage 1 (NumPro frame-numbering) -> Stage 2 refine (frame-so, coarse-to-\n    fine) -> Stage 3 grounding -> type cascade + scene rule. Da do T=0.4385 tren\n    calibration, thang constant (0.38) -- lan dau tien trong du an.\"\"\"\n    if not VLM_AVAILABLE:\n        raise RuntimeError('run_inference_vlm called with no usable VLM: every row '\n                           'would be normalize_prediction defaults, i.e. a constant.')\n    cap = cv2.VideoCapture(str(video_path))\n    fps, n = cap.get(cv2.CAP_PROP_FPS), int(cap.get(cv2.CAP_PROP_FRAME_COUNT))\n    cap.release()\n    duration = (n / fps) if fps > 0 else 20.0\n\n    scene = SCENE_BY_PATH.get(sub_path or ('videos/' + video_path.name))\n\n    pred = stage1_full_scan(video_path, duration, scene)\n\n    global PARSE_FAIL_STREAK\n    if pred.get('_parsed'):\n        PARSE_FAIL_STREAK = 0\n    else:\n        PARSE_FAIL_STREAK += 1\n        if PARSE_FAIL_STREAK >= PARSE_FAIL_ABORT:\n            raise RuntimeError(\n                f'Stage 1 failed to parse JSON on {PARSE_FAIL_STREAK} consecutive clips. '\n                'The predictions being written are normalize_prediction defaults. '\n                'Inspect the raw VLM output before continuing.')\n\n    t_final = stage2_time_refine_numbered(video_path, pred['accident_time'], duration)\n\n    pt = stage3_grounding(video_path, t_final)\n    if pt is not None:\n        pred['center_x'], pred['center_y'] = pt\n\n    pred['type'] = classify_type_cascade(video_path, t_final, pred['type'])\n    pred['type'] = apply_scene_type_postfix(pred['type'], scene)\n\n    return {'path': str(video_path), 'accident_time': t_final,\n            'center_x': pred['center_x'], 'center_y': pred['center_y'],\n            'type': pred['type'], 'scene_layout': scene}\n\n\nprint('[STATUS] run_inference_vlm (final): NumPro Stage1 + frame-number coarse-to-fine '\n      '+ Stage3 grounding + type cascade -- calibration T=0.4385, S=(chua do lai), '\n      'C=(chua do lai)')\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T15:04:32.524143Z","iopub.execute_input":"2026-08-03T15:04:32.524891Z","iopub.status.idle":"2026-08-03T15:04:32.532836Z","shell.execute_reply.started":"2026-08-03T15:04:32.524862Z","shell.execute_reply":"2026-08-03T15:04:32.532066Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 11. Submission ben vung hon cho chay da phien (~25h, 2027 video)\n\nBan goc o muc \"[SUBMIT] Full test set inference\" da co checkpoint CSV moi 25\nvideo + resume qua done_paths -- giu nguyen phan do. Cell nay CHI THEM 3 lop\nbao ve, vi 44.6s/video x 2027 = 25.1h la chac chan can >=3 phien:\n\n1. `try/except` quanh TUNG video -- 1 video loi khong lam chet ca vong lap,\n   fallback ve `normalize_prediction` mac dinh cho rieng video do, ghi log ro\n   video nao/loi gi de xem lai sau.\n2. Checkpoint moi 5 video (tu 25) -- giam toi da con ~3-4 phut mat neu session\n   bi ngat dot ngot (khong phai do het 12h ma do loi ha tang Kaggle).\n3. Tu dong DUNG SOM va luu, khong doi den khi bi kill cung, neu da chay qua\n   `SESSION_TIME_BUDGET_HOURS` (mac dinh 11h, duoi 12h Kaggle mot chut de con\n   thoi gian ghi file an toan) -- tranh truong hop dang ghi CSV thi bi ngat\n   giua chung, hong file checkpoint.","metadata":{}},{"cell_type":"code","source":"# [SUBMIT v2] Nhu ban goc nhung them try/except per-video + checkpoint day hon\n# + tu dung som truoc khi het gio session. Doi CHECKPOINT_PATH khac ten de\n# KHONG doc nham checkpoint cua ban submit cu (dinh dang co the khac neu ban\n# cu da chay mot phan).\nSUBSET_N = None    # None = full run that. Doi lai vi du 100 neu muon dry-run truoc.\nSUBMIT_PIPELINE = run_inference_vlm\n\nif SUBMIT_PIPELINE is run_inference_vlm and not VLM_AVAILABLE:\n    raise RuntimeError('Refusing to run: the VLM is unavailable, so every prediction '\n                       'would be a constant. Fix Section 6.10 first.')\n\nCHECKPOINT_PATH_V2 = OUTPUT_DIR / 'preds_checkpoint_v2.csv'\nCHECKPOINT_EVERY_V2 = 5          # tu 25 -- mat toi da ~3-4 phut neu bi ngat dot ngot\nSESSION_TIME_BUDGET_HOURS = 11.0  # duoi 12h Kaggle, chua thoi gian ghi file an toan\n\nvideos_to_run = real_videos[:SUBSET_N] if SUBSET_N is not None else real_videos\nis_subset_run = SUBSET_N is not None and SUBSET_N < len(real_videos)\n\ntest_results, done_paths, failed_videos = [], set(), []\nif CHECKPOINT_PATH_V2.exists():\n    prev = pd.read_csv(CHECKPOINT_PATH_V2)\n    test_results = prev.to_dict('records')\n    done_paths = set(prev['path'])\n    print(f'[STATUS] Checkpoint found: {len(done_paths)} videos already processed -- resuming')\n\nprint(f'[STATUS] Final-pipeline inference on {len(videos_to_run)} video(s) '\n      f'{\"(SUBSET DRY RUN)\" if is_subset_run else \"(FULL RUN)\"} '\n      f'out of {len(real_videos)} total, session budget {SESSION_TIME_BUDGET_HOURS}h')\n\nper_video_times = []\nrun_start = time.time()\nstopped_early = False\n\nfor i, vp in enumerate(videos_to_run):\n    sub_path = 'videos/' + vp.name\n    if sub_path in done_paths:\n        continue\n\n    elapsed_h = (time.time() - run_start) / 3600.0\n    if elapsed_h >= SESSION_TIME_BUDGET_HOURS:\n        print(f'[STATUS] Session time budget ({SESSION_TIME_BUDGET_HOURS}h) reached at '\n              f'{len(test_results)}/{len(videos_to_run)} videos -- stopping cleanly, '\n               'saving checkpoint, resume in the next session.')\n        stopped_early = True\n        break\n\n    t0 = time.time()\n    try:\n        result = SUBMIT_PIPELINE(vp, sub_path)\n    except Exception as e:\n        # Mot video loi KHONG duoc lam chet ca phien -- 2026 video con lai con\n        # gia tri hon 1 video bi loi. Dung default an toan, ghi lai de xem sau.\n        print(f'[WARNING] {vp.name} failed ({type(e).__name__}: {e}) -- using default prediction')\n        cap = cv2.VideoCapture(str(vp))\n        fps, n = cap.get(cv2.CAP_PROP_FPS), int(cap.get(cv2.CAP_PROP_FRAME_COUNT))\n        cap.release()\n        duration = (n / fps) if fps > 0 else 20.0\n        default_pred = normalize_prediction({}, duration)\n        result = {'path': sub_path, 'accident_time': default_pred['accident_time'],\n                 'center_x': default_pred['center_x'], 'center_y': default_pred['center_y'],\n                 'type': default_pred['type'], 'scene_layout': None}\n        failed_videos.append({'path': sub_path, 'error': f'{type(e).__name__}: {e}'})\n\n    per_video_times.append(time.time() - t0)\n    result['path'] = sub_path\n    test_results.append(result)\n    done_paths.add(sub_path)\n\n    if len(test_results) % CHECKPOINT_EVERY_V2 == 0:\n        pd.DataFrame(test_results).to_csv(CHECKPOINT_PATH_V2, index=False)\n\n    if (i + 1) % 10 == 0 or (i + 1) == len(videos_to_run):\n        avg_t = sum(per_video_times) / max(len(per_video_times), 1)\n        elapsed = time.time() - run_start\n        remaining = avg_t * max(0, len(videos_to_run) - len(test_results))\n        print(f'[STATUS] {len(test_results)}/{len(videos_to_run)} | avg {avg_t:.2f}s/video | '\n              f'elapsed {elapsed/60:.1f} min | remaining ~{remaining/3600:.2f} h | '\n              f'failed so far: {len(failed_videos)}')\n\n    if torch.cuda.is_available():\n        torch.cuda.empty_cache()\n\npd.DataFrame(test_results).to_csv(CHECKPOINT_PATH_V2, index=False)\nif failed_videos:\n    pd.DataFrame(failed_videos).to_csv(OUTPUT_DIR / 'failed_videos.csv', index=False)\n    print(f'[STATUS] {len(failed_videos)} video(s) failed and used default predictions -- '\n          f'see failed_videos.csv')\n\ntest_preds_df = pd.DataFrame(test_results)\nprint(f'\\n[SUCCESS] {len(test_preds_df)}/{len(videos_to_run)} video(s) processed this run '\n      f'({\"STOPPED EARLY, resume next session\" if stopped_early else \"COMPLETE\"})')\nif not stopped_early:\n    assert len(test_preds_df) == len(videos_to_run), 'Some videos were dropped during inference'\ndisplay(test_preds_df.head())\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-03T15:04:32.533633Z","iopub.execute_input":"2026-08-03T15:04:32.533825Z","execution_failed":"2026-08-03T23:28:15.649Z"}},"outputs":[],"execution_count":null}]}